Multi-parameter coupling optimization control method based on deep reinforcement learning
By employing a multi-parameter coupling optimization control method based on deep reinforcement learning, the valve and pump settings of the chilled water system are adjusted in real time, solving the problems of hydraulic imbalance and energy consumption in high-rise buildings and achieving a balance between hydraulic equilibrium and energy efficiency.
Patent Information
- Application Number
- CN202511935684.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2045-12-22
AI Technical Summary
Chilled water systems in high-rise buildings suffer from hydraulic imbalance, leading to uneven heating and cooling. Existing technologies struggle to effectively address the combination of hydraulic imbalance and energy consumption optimization in multi-story high-rise buildings.
A multi-parameter coupled optimization control method based on deep reinforcement learning is adopted. By collecting the opening degree and differential pressure data of the terminal valve in real time, a hydraulic state model is constructed, the differential pressure compensation signal of the partition is extended, and hydraulic constraints are introduced into the deep reinforcement learning strategy to generate control commands to adjust the settings of the chiller, pump and fan, so as to achieve hydraulic balance and energy consumption optimization.
It achieves dynamic adjustment of hydraulic balance and energy consumption optimization in high-rise buildings, avoiding water shortage in high-rise areas and oversupply in low-rise areas, and ensuring comfort and energy efficiency within the building.
Smart Images

Figure CN121364640A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of chilled water system regulation, in particular to a multi-parameter coupling optimization control method based on deep reinforcement learning. BACKGROUND
[0002] The chilled water system is a complex pipe network system, and there is coupling between each loop. The change of the flow of a certain loop often causes the flow of each loop to influence each other, so many factors need to be considered in the regulation work. At the same time, due to too many unknown factors on site, the characteristics of the water pump, valve and other elements have certain correlation with the hydraulic characteristics of the pipe network, so it is very difficult to implement. The chilled water system of the central air conditioning of the current large public building is usually large in scale and complex in pipe. Due to a series of problems generated in the design and construction stage and the later operation stage, the system generally has the phenomenon of hydraulic imbalance. The hydraulic imbalance of the chilled water system will cause the flow distribution of each loop and the terminal equipment of the system to not reach the design value, which directly leads to the phenomenon of uneven cooling and heating in each air conditioning area.
[0003] At present, in high-rise office buildings or commercial and residential complexes, the static pressure difference of the chilled water pipe network changes significantly with the floor height. Even if the reinforcement learning optimizes the overall energy consumption, it is still relatively easy for the terminal valve at the upper floor to fail to obtain sufficient water, which makes it difficult to guarantee the indoor comfort of the building. Therefore, it is necessary to provide a multi-parameter coupling optimization control method based on deep reinforcement learning for the setting of special compensation control of the three-dimensional high-rise hydraulic imbalance. SUMMARY
[0004] Therefore, it is necessary to propose a multi-parameter coupling optimization control method based on deep reinforcement learning in view of the above technical problems.
[0005] The present application adopts the following technical solutions.
[0006] The present application discloses a multi-parameter coupling optimization control method based on deep reinforcement learning, which comprises: Real-time collection of the terminal valve opening and pressure difference data of the chilled water pipe network in different height partitions of the building, combination of the pump house outlet pressure and the building height difference, and construction of a hydraulic state model of the building height partition; Based on the hydraulic state model, key compensation parameters are extracted, and a partition differential pressure compensation signal is expanded in the initial action space of the chilled water pipe network to construct a multi-parameter action space; Based on the multi-parameter action space, a deep reinforcement learning strategy is trained in a simulation environment, and a hydraulic constraint is introduced into the deep reinforcement learning strategy; The deep reinforcement learning strategy is called to generate multiple control instructions according to the real-time partition hydraulic state to control the settings of the chiller, pump and fan in the chilled water pipe network and the partition differential pressure compensation signal. monitoring the water supply temperature and flow rate at the end of the high zone after execution of the control instruction in real time to verify whether the hydraulic state of the different height partitions of the building meets the preset target threshold value; The different height partitions of the building are a high zone, a middle zone, and a low zone. The key compensation parameters are minimum pressure difference data at the end of the high zone and valve margins of the middle zone and the low zone. The initial action space is chilled water supply temperature, water pump speed, fan speed, and valve opening degree.
[0007] Further, the end valve opening degree and pressure difference data of the chilled water pipe network in the different height partitions of the building are collected in real time, and the pump house outlet pressure and the building height difference are combined to construct a hydraulic state model of the building height partitions, including: A differential pressure sensor is arranged at the end of the chilled water pipe network in the high zone, the middle zone, and the low zone of the building, and the differential pressure data of the different height partitions of the building are collected in real time through the differential pressure sensor, and the valve opening degree at the end of each partition and the pump house outlet pressure are obtained; The differential pressure data, valve opening degree, and pump house outlet pressure are filtered and normalized to construct a standardized monitoring parameter set.
[0008] Further, the end valve opening degree and pressure difference data of the chilled water pipe network in the different height partitions of the building are collected in real time, and the pump house outlet pressure and the building height difference are combined to construct a hydraulic state model of the building height partitions, including: Based on the height values of the high zone, the middle zone, and the low zone of the building, the theoretical static pressure of the different height partitions is calculated, and the effective differential pressure is determined according to the differential pressure data collected by the differential pressure sensor, the effective differential pressure being the sum of the theoretical static pressure and the differential pressure data; Based on the effective differential pressure, a valve flow characteristic curve is constructed to estimate the water supply of each partition, and a water supply vector of the whole building is obtained; According to the water supply vector and the preset target water supply of each partition, the valve flow characteristic curve and the valve opening degree are combined to generate the hydraulic state model.
[0009] Further, based on the hydraulic state model, key compensation parameters are extracted, and partition differential pressure compensation signals are expanded in the initial action space of the chilled water pipe network to construct a multi-parameter action space, including: Based on the hydraulic state model, the minimum pressure difference data at the end of the high zone and the valve margins of the middle zone and the low zone are extracted in time sequence to construct a high zone end pressure difference sequence and a middle and low zone valve opening degree sequence; In a set monitoring period, the minimum value of the high zone end pressure difference is calculated according to the high zone end pressure difference sequence, and the middle and low zone valve margin parameter set is determined based on the middle and low zone valve opening degree sequence; An initial action space of the chilled water pipe network is acquired, and the partition differential pressure compensation signal is determined according to a minimum value of the set of low-mid zone valve margin parameters and the high zone end pressure difference, the partition differential pressure compensation signal is written into the initial action space, and an extended multi-parameter action space is obtained.
[0010] Further, based on the multi-parameter action space, a deep reinforcement learning strategy is trained in a simulation environment, and a hydraulic constraint is introduced into the deep reinforcement learning strategy, including: The indoor temperature and humidity of each height partition of the building are acquired, and the indoor temperature and humidity are written into the multi-parameter action space, and the multi-parameter action space and the indoor temperature and humidity are normalized to a preset interval to obtain a normalized state space set; Based on the state space set, a reward function with the total power consumption of the chilled water pipe network system as an optimization objective is constructed, and a peak penalty, a hydraulic constraint and a comfort constraint are introduced into the reward function to obtain a composite reward function; The hydraulic constraint is that the high zone end pressure difference is not allowed to be lower than a set threshold and the low-mid zone valve opening is not allowed to exceed a set upper limit; and the comfort constraint is that the indoor temperature and humidity satisfy a set temperature and humidity range.
[0011] Further, based on the multi-parameter action space, a deep reinforcement learning strategy is trained in a simulation environment, and a hydraulic constraint is introduced into the deep reinforcement learning strategy, further including: The multi-parameter action space is used as a candidate action of the deep reinforcement learning strategy, and the state space set is used as an input of a deep learning framework, and deep reinforcement learning training is performed in a simulation environment to obtain a deep reinforcement learning strategy; The composite reward function is used as an optimization objective, and the composite reward function is introduced into the deep reinforcement learning strategy.
[0012] Further, the deep reinforcement learning strategy is called to generate a plurality of control instructions according to a real-time partition hydraulic state to control the settings of the chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal, including: During execution of the deep reinforcement learning strategy in the chilled water pipe network, indoor temperature and humidity, water supply, effective differential pressure, valve opening and pump house outlet pressure of each height partition of the building are collected in real time to construct a real-time hydraulic state set; The real-time hydraulic state set is input into the deep reinforcement learning strategy to output an optimal action instruction, and a partition differential pressure compensation signal is parsed in response to the optimal action instruction to obtain a partition differential pressure adjustment parameter; The control instruction to be executed is generated according to the partition differential pressure adjustment parameter, and a control instruction set for controlling the settings of the chiller, the pump and the fan in the chilled water pipe network and the partition differential pressure compensation signal is obtained.
[0013] Further, the water supply temperature and flow at the end of the high zone after the execution of the control instruction are monitored in real time to verify whether the hydraulic states of the different height partitions of the building meet preset target thresholds, including: In response to the control instruction, the control instruction is executed according to the partition differential pressure adjustment parameter to adjust the chilled water supply temperature, the water pump rotating speed, the fan rotating speed, the valve opening degree and the differential pressure target of each height partition, and hydraulic state feedback data is collected; The decision function based on the high zone water supply constraint, the medium and low zone valve margin constraint, the comfort constraint and the energy efficiency constraint is constructed according to the hydraulic state feedback data, the comprehensive decision result is output by calling the decision function, and whether the comprehensive decision result meets the preset condition is verified.
[0014] The second aspect of the present application discloses a multi-parameter coupled optimization control device based on deep reinforcement learning, which is used to realize the multi-parameter coupled optimization control method based on deep reinforcement learning of any one of the first aspect, and the device comprises: The building partition modeling module is used to collect the end valve opening degree and differential pressure data of the chilled water pipe network in different height partitions of the building in real time, combine the pump house outlet pressure and the building height difference, and construct a hydraulic state model of the building height partition; The action space construction module is used to extract key compensation parameters based on the hydraulic state model, and expand the partition differential pressure compensation signal in the initial action space of the chilled water pipe network to construct a multi-parameter action space; The policy reinforcement training module is used to train a deep reinforcement learning policy in a simulation environment based on the multi-parameter action space, and introduce a hydraulic constraint in the deep reinforcement learning policy; The instruction generation module is used to call the deep reinforcement learning policy to generate a plurality of control instructions according to real-time partition hydraulic states to control the settings of the chiller, the pump and the fan in the chilled water pipe network and the partition differential pressure compensation signal; The monitoring and verification module is used to monitor the water supply temperature and flow at the end of the high zone after the execution of the control instruction in real time to verify whether the hydraulic states of the different height partitions of the building meet preset target thresholds; Wherein, the different height partitions of the building are high zone, medium zone and low zone respectively; the key compensation parameters are the minimum differential pressure data at the end of the high zone and the valve margin of the medium zone and the low zone; and the initial action space is the chilled water supply temperature, the water pump rotating speed, the fan rotating speed and the valve opening degree.
[0015] The third aspect of the present application discloses a terminal comprising a processor and a storage medium; The storage medium is used for storing instructions. The processor is used for operating according to the instructions to perform the steps of the method of the first aspect.
[0016] The fourth aspect of the present application discloses a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the method of the first aspect.
[0017] The present application has the following advantages: (1) The present application arranges a differential pressure sensor in the chilled water pipe network, collects the typical end valve opening degree and pressure difference of high, medium and low zones, combines the pump house outlet pressure and floor height difference to form a hydraulic state model of floor zoning, compares the opening degree of each zone valve with the actual water supply amount to determine whether there is obvious imbalance in hydraulic distribution, and determines the influence of static pressure difference in high-rise buildings to provide a basis for subsequent compensation control. At the same time, based on the action space of initial chilled water supply temperature, pump speed, fan speed and valve opening degree, the partition differential pressure compensation signal is expanded and added as a new control variable, so that the control object can explicitly contain the hydraulic balance target, solving the problem of three-dimensional hydraulic balance and introducing an operable compensation means for water cooling optimization.
[0018] (2) The present application uses parameter action space to train deep reinforcement learning strategy in a simulation environment. The training target not only includes energy consumption minimization and comfort maintenance, but also introduces a hydraulic constraint, i.e. the end pressure difference of the high zone cannot be lower than the threshold value and the valve opening degree of the medium and low zones cannot exceed the upper limit, so that the reinforcement learning strategy can consider energy consumption optimization and hydraulic balance in the multi-floor cooling scene at the same time, and the hydraulic balance target is included in the optimization process, effectively avoiding the imbalance phenomenon of water shortage in the high zone and over-supply in the low zone of the building.
[0019] (3) In the actual system operation, the present application calls the reinforcement learning strategy to generate control instructions in combination with real-time partition hydraulic state. These instructions not only include the conventional settings of chillers, pumps and fans, but also include the regulation and control of partition differential pressure compensation signals, such as dynamically adjusting the start-stop logic of high-zone water supply pumps and setting the differential pressure maintenance point. Through online reasoning and real-time feedback, the water distribution of different floor zones is kept in dynamic balance, solving the water supply imbalance caused by floor height difference in the operation of the water cooling system, and realizing the regulation of hydraulic. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0021] Figure 1 is a flowchart of a multi-parameter coupling optimization control method based on deep reinforcement learning provided by the present application.
[0022] Figure 2 is a structural schematic diagram of a multi-parameter coupling optimization control device based on deep reinforcement learning provided by the present application. DETAILED DESCRIPTION
[0023] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the protection scope of the present application.
[0024] As shown in Figure 1 , in one embodiment, a multi-parameter coupling optimization control method based on deep reinforcement learning comprises the following steps: Step S110, collecting the end valve opening degree and pressure difference data of the chilled water pipe network in different height partitions of the building in real time, combining the pump house outlet pressure and the building height difference, and constructing a hydraulic state model of the building height partition.
[0025] Among them, the different height partitions of the building are high area, middle area and low area.
[0026] In some embodiments, the multi-parameter coupling optimization control method based on deep reinforcement learning provided by the present application specifically comprises the following steps: Step S111, setting a differential pressure sensor at the end of the chilled water pipe network in the high area, middle area and low area of the building, and collecting the pressure difference data of the different height partitions of the building in real time through the differential pressure sensor, while obtaining the valve opening degree at the end of each partition and the pump house outlet pressure.
[0027] Step S112, filtering and normalizing the pressure difference data, valve opening degree and pump house outlet pressure to construct a standardized monitoring parameter set.
[0028] In some embodiments, the application provides a multi-parameter coupled optimization control method based on deep reinforcement learning, and step S110 specifically further comprises the following steps: Step S113, based on the height values of the high, middle and low zones of the building, the theoretical static pressure of different height zones is calculated, and the effective differential pressure is determined according to the differential pressure data collected by the differential pressure sensor, which is the sum of the theoretical static pressure and the differential pressure data.
[0029] Step S114, based on the effective differential pressure, a valve flow characteristic curve is constructed to estimate the water supply of each zone, and the overall water supply vector of the building is obtained.
[0030] Step S115, according to the water supply vector and the preset target water supply of each zone, the valve flow characteristic curve and the valve opening are combined to generate a hydraulic state model.
[0031] In specific embodiments, the application provides a multi-parameter coupled optimization control method based on deep reinforcement learning, comprising steps 1-5: Step 1, whole floor hydraulic state monitoring and zone modeling.
[0032] In the chilled water pipe network, differential pressure sensors are arranged to collect typical end valve opening and differential pressure of high, middle and low zones, and combined with the pump house outlet pressure and floor height difference, a hydraulic state model of the floor zoning is formed. By comparing the opening degree of each zone valve with the actual water supply, it is determined whether there is obvious imbalance in hydraulic distribution, so as to clarify the influence of static pressure difference in high-rise buildings and provide basis for subsequent compensation control. Including the following sub-steps: Sub-step 1.1, original hydraulic parameter collection and preprocessing.
[0033] Specifically, the building is divided into high, middle and low zones according to height, and differential pressure sensors are arranged at the typical ends of the chilled water pipe network in the high, middle and low zones of the floor to obtain the differential pressure values of each zone in real time. At the same time, the end valve opening and pump house outlet pressure (200-600 kPa) of each zone are collected, and then the collected differential pressure values, valve opening and pump house outlet pressure of each zone are filtered and normalized to obtain a standardized monitoring parameter set.
[0034] Sub-step 1.2, floor height difference and static pressure compensation calculation.
[0035] Specifically, since the static pressure difference in high-rise buildings increases linearly with height, assuming that the floor height corresponding to a certain zone is known, the theoretical static pressure is the product of the floor height and the water density and gravitational acceleration. Then, the differential pressure value actually measured by the differential pressure sensor at the corresponding floor height is added to the theoretical static pressure at the floor height, which is the effective differential pressure at the floor height. According to this method, the effective differential pressures of all zones of the building are calculated, and finally an effective differential pressure set of the building is constructed.
[0036] Sub-step 1.3, fitting the valve opening degree and water supply quantity.
[0037] Specifically, the valve flow characteristic curve is constructed according to the effective differential pressure set, and the water supply quantity of each sub-area is estimated based on the curve, and the expression is:
[0038] In the formula, represents the water supply quantity of the i-th sub-area at time t; is the valve flow coefficient, ranging from 0.5 to 2.0; is the effective differential pressure of the i-th sub-area at time t; is the valve opening degree of the i-th sub-area at time t.
[0039] Finally, the water supply quantity of each sub-area of the whole floor is calculated according to the above formula, and then the corresponding water supply quantity vector of the whole floor is output, which is composed of the water supply quantity of each sub-area.
[0040] Sub-step 1.4, construction and imbalance determination of sub-area hydraulic state model.
[0041] Specifically, the target water supply quantity of each sub-area is defined, which is determined by design load calculation or operation experience. If the water supply quantity of the corresponding sub-area is less than the target water supply quantity and the valve opening degree is close to the value 1, it indicates that the sub-area is under-supplied. If the water supply quantity of the corresponding sub-area is greater than the target water supply quantity and the valve opening degree is close to 0.2, it indicates that the sub-area is over-supplied. Therefore, the target water supply quantity is combined with the aforementioned water supply vector, effective differential pressure and valve opening degree, and the required sub-area hydraulic state model is constructed.
[0042] Step S120, based on the hydraulic state model, extracting key compensation parameters and expanding sub-area differential pressure compensation signals in the initial action space of the chilled water pipe network to construct a multi-parameter action space.
[0043] Among them, the key compensation parameters are the minimum pressure difference data at the end of the high area and the valve margin of the middle and low areas; the initial action space is the chilled water supply temperature, the water pump speed, the fan speed and the valve opening degree.
[0044] In some embodiments, the multi-parameter coupling optimization control method based on deep reinforcement learning provided by the present application specifically includes the following steps: Step S121, based on the hydraulic state model, extracting the minimum pressure difference data at the end of the high area and the valve margin of the middle and low areas in time sequence to construct the high area end pressure difference sequence and the middle and low area valve opening degree sequence.
[0045] Step S122, within a set monitoring period, calculate the minimum value of the high zone end pressure difference according to the high zone end pressure difference sequence, and determine a set of low-middle zone valve margin parameters based on the low-middle zone valve opening sequence.
[0046] Step S123, obtain the initial action space of the chilled water pipe network, and determine a partition differential pressure compensation signal according to the set of low-middle zone valve margin parameters and the minimum value of the high zone end pressure difference, to write the partition differential pressure compensation signal into the initial action space to obtain an expanded multi-parameter action space.
[0047] In specific embodiments, the method for multi-parameter coupled optimization control based on deep reinforcement learning provided by the present application comprises the following steps: Sub-step 2.1, analysis of the partition hydraulic state model and extraction of indicators.
[0048] Specifically, based on the aforementioned constructed partition hydraulic state model, the minimum end pressure difference of the high zone is extracted as a key indicator for measuring the supply capacity of the high zone; and the low-middle zone valve opening is extracted as a key indicator for judging whether the valve is excessively throttled.
[0049] Sub-step 2.2, calculation of the minimum differential pressure of the high zone end.
[0050] Specifically, within a set monitoring period T, the minimum value of the high zone end differential pressure in the period is calculated, in this example, T is an optimization control period, which is 15-60 min, and if the minimum value of the high zone differential pressure is lower than a set threshold (30-50 kPa), it indicates that there is a risk of insufficient supply in the high zone.
[0051] Sub-step 2.3, calculation of the low-middle zone valve margin.
[0052] Specifically, the valve margin is defined as the difference between the valve adjustable range and the current opening, and the valve margin of the low-middle zone within the monitoring period T is determined as the low-middle zone valve margin parameter, in this example, the valve margin value ranges from 0 to 1. At the same time, the low-middle zone valve margin parameter is associated with the minimum differential pressure of the high zone to determine whether the condition for improving the supply of the high zone through partition differential pressure adjustment is met.
[0053] Sub-step 2.4, expansion of the control variable and reconstruction of the action space.
[0054] Specifically, first, an original action space is determined, including chilled water supply temperature, water pump speed, fan speed, and valve opening degree of each zone. Then, a zone differential pressure compensation signal is added to the original action space, the zone differential pressure compensation signal being determined by the minimum differential pressure of the high zone and the valve margin, and used for dynamically adjusting the water supply pump start-stop or differential pressure set point, to finally obtain an expanded action space.
[0055] Step S130, based on the multi-parameter action space, training a deep reinforcement learning strategy in a simulation environment, and introducing a hydraulic constraint in the deep reinforcement learning strategy.
[0056] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present application specifically comprises the following steps in step S130: Step S131, obtaining indoor temperature and humidity of each height zone of the building, and writing the indoor temperature and humidity into the multi-parameter action space, and simultaneously normalizing the multi-parameter action space and the indoor temperature and humidity to a preset interval to obtain a normalized state space set.
[0057] Step S132, based on the state space set, constructing a reward function with total power consumption of the chilled water pipe network system as an optimization objective, and introducing a peak penalty, a hydraulic constraint and a comfort constraint in the reward function to obtain a composite reward function.
[0058] Wherein, the hydraulic constraint is that the differential pressure at the end of the high zone is not allowed to be lower than a set threshold, and the valve opening degree of the middle and low zones is not allowed to exceed a set upper limit; the comfort constraint is that the indoor temperature and humidity meet a set temperature and humidity range.
[0059] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present application specifically further comprises the following steps in step S130: Step S133, taking the multi-parameter action space as a candidate action of the deep reinforcement learning strategy, and taking the state space set as an input of the deep learning framework, and performing deep reinforcement learning training in a simulation environment to obtain a deep reinforcement learning strategy.
[0060] Step S134, taking the composite reward function as an optimization objective, and introducing the composite reward function into the deep reinforcement learning strategy.
[0061] In specific embodiments, the present application provides a multi-parameter coupled optimal control method based on deep reinforcement learning, step 3, deep reinforcement learning strategy training and hydraulic constraint introduction. The multi-parameter action space output in step 2 is used to train the deep reinforcement learning strategy in a simulation environment. The training target includes not only energy consumption minimization and comfort maintenance, but also additional introduction of hydraulic constraints, i.e. the high zone end pressure difference cannot be lower than the threshold value and the low zone valve opening cannot exceed the upper limit. In this way, the reinforcement learning strategy can consider energy optimization and hydraulic balance in the multi-floor cooling scene at the same time, and the hydraulic balance target is included in the optimization process to avoid the phenomenon of high zone water shortage and low zone over-supply. The following sub-steps are included: Sub-step 3.1, state space construction and normalization.
[0062] Specifically, in the multi-floor cooling scene, the state space not only includes the aforementioned extended action space, but also writes the temperature and humidity in the building into the extended action space to balance the cooling energy consumption and indoor comfort, and normalizes the temperature and humidity and the hydraulic parameters in the extended action space to the interval [0, 1] to obtain the normalized state space.
[0063] Sub-step 3.2, reward function construction and energy consumption target introduction.
[0064] Specifically, the reward function takes the total power consumption of the chilled water pipe network system as the main optimization target, while considering the peak penalty (i.e. the peak penalty coefficient , which is 5-10 in this example), combining the cooling machine power , the water pump power , the fan power , and the upper limit of power demand (range 200-2000kW), the constructed reward function The expression is:
[0065] In the formula, is the total power of the system.
[0066] Sub-step 3.3, introduction of hydraulic constraints and comfort constraints.
[0067] Specifically, on the basis of the total power consumption as the optimization target in sub-step 3.2, the hydraulic constraint cost and the comfort constraint cost are introduced, and both the hydraulic constraint cost and the comfort constraint cost have been normalized according to the reference value and are dimensionless after processing. The reward function expression after introducing the hydraulic constraint and the comfort constraint is updated as:
[0068] In the formula, is the hydraulic constraint cost at time t, is a weight coefficient of the water force constraint cost; is a comfort constraint cost at time t, is a weight coefficient of the comfort constraint cost; is a composite reward function after introducing the water force constraint cost and the comfort constraint cost.
[0069] wherein the water force constraint cost is expressed as:
[0070] wherein, is a high-zone target differential pressure, is a high-zone effective differential pressure at time t; is a valve opening degree of a middle zone at time t, is a maximum value of the middle-zone valve opening degree; is a valve opening degree of a low zone at time t, is a maximum value of the low-zone valve opening degree.
[0071] the comfort constraint cost is expressed as:
[0072] wherein, is a total number of indoor spaces; is an indoor temperature of the ith zone at time t, is a set comfort temperature threshold, which is 24°C in this example; is a tolerance, which is equal to 1°C in this example; is an indoor humidity of the ith zone at time t, is an upper limit of humidity, which is 60% in this example.
[0073] Step 3.4, policy training and water compensation mechanism generation.
[0074] Specifically, the foregoing extended action space is taken as a policy selectable action, the normalized state space is taken as an input, and the composite reward function is taken as an optimization target. Deep reinforcement learning training is performed in a simulation space, and finally an optimal policy is obtained as a deep reinforcement learning policy.
[0075] wherein the extended action space contains a zone differential pressure compensation signal, and thus the finally generated deep reinforcement learning policy can generate a compensation action for high-level water supply deficiency, such as adjusting a differential pressure set point or starting and stopping a water supply pump, on the basis of conventional control of chilled water supply and return water temperature, pump, and fan speed.
[0076] Step S140, calling the deep reinforcement learning policy to generate a plurality of control instructions according to real-time zone water force states, to control the settings of chillers, pumps, and fans in the chilled water pipe network and the zone differential pressure compensation signal.
[0077] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present application specifically comprises the following steps in step S140: Step S141, during the execution of the deep reinforcement learning strategy in the chilled water pipe network, the indoor temperature and humidity, water supply, effective differential pressure, valve opening of each height partition of the building, and pump house outlet pressure are collected in real time to construct a real-time hydraulic state set.
[0078] Step S142, inputting the real-time hydraulic state set into the deep reinforcement learning strategy to output an optimal action instruction, and analyzing a partition differential pressure compensation signal in response to the optimal action instruction to obtain a partition differential pressure adjustment parameter.
[0079] Step S143, generating a control instruction to be executed according to the partition differential pressure adjustment parameter to obtain a control instruction set for controlling the settings of the chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal.
[0080] In specific embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present application comprises the following sub-steps in step 4, online reasoning and adaptive adjustment of partition differential pressure. In actual system operation, the deep reinforcement learning strategy output in step 3 is called in combination with real-time partition hydraulic states to generate control instructions. These instructions not only include the conventional settings of chillers, pumps and fans, but also include partition differential pressure compensation signals, such as dynamically adjusting the start-stop logic of high-zone water supply pumps and setting the differential pressure maintenance point. Through online reasoning and real-time feedback, the water distribution of different floor partitions is kept in dynamic balance, solving the water supply imbalance caused by the height difference of floors in operation, and realizing real-time hydraulic regulation. The following sub-steps are included: Sub-step 4.1, strategy calling and real-time state input.
[0081] Specifically, during system operation, the monitoring data of each partition on the floor, i.e., indoor temperature and humidity, water supply, effective differential pressure, valve opening and pump house outlet pressure, are collected in real time to construct a real-time state set.
[0082] Sub-step 4.2, action reasoning and compensation signal generation.
[0083] Specifically, the deep reinforcement learning strategy output as described above is called with the real-time state set as input, and then the optimal action is output, and finally the optimal action instruction set of the whole floor is obtained to control the chilled water supply temperature setting, water pump speed, fan speed, valve opening and high-zone differential pressure set point adjustment or water supply pump operation according to the partition differential pressure compensation signal.
[0084] Sub-step 4.3, compensation logic analysis and partition differential pressure adjustment.
[0085] Specifically, the partition differential pressure compensation signal is analyzed, and the high zone target differential pressure is defined as the product of the current high zone measured differential pressure and the compensation proportional coefficient (0.5-1.5) multiplied by the differential pressure compensation signal (-20-+20 kPa). If the high zone target differential pressure is lower than the set threshold (30-50 kPa), the high zone water supply pump is triggered to run, and if the average value of the low zone valve opening is less than 0.4, the low zone differential pressure set point is lowered to release the excess for the high zone. Finally, the partition adjustment parameter set is output, including the target differential pressures of the high, medium and low zones.
[0086] Sub-step 4.4, control instruction issuing and real-time feedback.
[0087] Specifically, according to the aforementioned partition adjustment parameter set, executable control instructions are generated, which are used to: adjust the water supply temperature setting on the cold side to ensure that the cold output matches the high zone demand; adjust the high zone water supply pump start-stop or frequency setting according to the high zone target pressure difference on the pump side, and adjust the medium and low zone main pump pressure setting according to the medium and low zone target pressure difference; adjust the fan speed according to the action instruction set on the fan side to stabilize the end air handling capacity in cooperation with water distribution; and issue valve opening instructions on the valve side to appropriately close the valve in the over-supply area and open the valve in the lack of supply area.
[0088] Step S150, real-time monitoring of the water supply temperature and flow of the high zone end after the execution of the control instructions to verify whether the hydraulic state of the building partitions at different heights meets the preset target threshold.
[0089] In some embodiments, the present application provides a multi-parameter coupling optimization control method based on deep reinforcement learning, and step S150 specifically includes the following steps: Step S151, in response to the control instructions, executing the control instructions according to the partition differential pressure adjustment parameters to adjust the chilled water supply temperature, water pump speed, fan speed, valve opening and differential pressure target of each height partition, while collecting hydraulic state feedback data.
[0090] Step S152, constructing a judgment function based on the high zone water supply constraint, the low zone valve margin constraint, the comfort constraint and the energy efficiency constraint according to the hydraulic state feedback data to output a comprehensive judgment result and verify whether the comprehensive judgment result meets the preset condition.
[0091] In a specific embodiment, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by this invention includes step 5: closed-loop feedback verification and end-point comfort assurance. After executing the zone differential pressure compensation control command output in step 4, the water supply temperature and flow rate of the high-zone end are monitored in real time and compared with the target threshold. When the temperature and humidity of the high-zone end remain within the comfort range, while the valve opening in the middle and low zones does not exceed the available margin, and the total system energy consumption is not higher than the benchmark value, it can be determined that the hydraulic balance and energy efficiency optimization goals are achieved simultaneously. The feedback result is input again into the deep reinforcement learning strategy for continuous strategy improvement, ultimately achieving the unification of comfort and energy efficiency across all floors, further solving the problem of hydraulic imbalance at the end of high-rise complexes. This includes the following sub-steps: Sub-step 5.1: Execute control commands and monitor data collection.
[0092] Specifically, the aforementioned zone differential pressure compensation control command is executed at the control layer, and feedback data is collected in real time after execution, including indoor temperature and humidity, water supply, effective differential pressure, valve opening degree and total system power.
[0093] Sub-step 5.2, comfort and energy efficiency assessment.
[0094] Specifically, based on the collected feedback data, a judgment function for comfort and energy efficiency is constructed to determine whether the water supply in the high zone is guaranteed, whether the valve margin in the middle and low zones meets the opening requirements, whether the indoor comfort meets the set standards, and whether the system energy efficiency exceeds the benchmark, based on the judgment results, and to impose conditional constraints on the chilled water pipeline system.
[0095] Sub-step 5.3: Feedback input and continuous strategy improvement.
[0096] Specifically, if the judgment result obtained in sub-step 5.2 meets the set conditions, then the high-zone water supply, the valve margin in the middle and low zones, indoor comfort, and system energy efficiency are all determined to meet the constraints, and the system operating status that meets the standards is output. Furthermore, if any constraint is not met, the judgment result and its monitoring data are fed back to the simulation space as incremental training samples to continuously optimize the deep reinforcement learning strategy.
[0097] The following describes the multi-parameter coupled optimization control device based on deep reinforcement learning provided by the present invention. The multi-parameter coupled optimization control device based on deep reinforcement learning described below can be referred to in correspondence with the multi-parameter coupled optimization control method based on deep reinforcement learning described above.
[0098] like Figure 2 As shown, in one embodiment, a multi-parameter coupled optimization control device based on deep reinforcement learning includes a building zoning modeling module, an action space construction module, a strategy reinforcement training module, an instruction generation module, and a monitoring and verification module.
[0099] The building partition modeling module is configured to collect data of the opening degree of the terminal valve and the differential pressure of the chilled water pipe network in different height partitions of the building in real time, and construct a hydraulic state model of the building height partitions in combination with the outlet pressure of the pump room and the height difference of the building.
[0100] The action space construction module is configured to extract key compensation parameters based on the hydraulic state model, and expand the partition differential pressure compensation signal in the initial action space of the chilled water pipe network to construct a multi-parameter action space.
[0101] The strategy reinforcement training module is configured to train a deep reinforcement learning strategy in a simulation environment based on the multi-parameter action space, and introduce a hydraulic constraint in the deep reinforcement learning strategy.
[0102] The instruction generation module is configured to call the deep reinforcement learning strategy to generate a plurality of control instructions according to the real-time partition hydraulic state, so as to control the settings of the chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal.
[0103] The monitoring and verification module is configured to monitor the water supply temperature and flow rate at the end of the high zone after execution of the control instructions in real time, so as to verify whether the hydraulic state of the different height partitions of the building meets a preset target threshold.
[0104] The different height partitions of the building are a high zone, a middle zone and a low zone; the key compensation parameters are the minimum differential pressure data of the end of the high zone and the valve margin of the middle zone and the low zone; and the initial action space is the chilled water supply temperature, the water pump speed, the fan speed and the valve opening degree.
[0105] The technical features of the above embodiments can be combined in any manner, and to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.
[0106] The above-described embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.
Claims
1. A multi-parameter coupling optimal control method based on deep reinforcement learning, characterized in that, The method comprises: Real-time acquisition of end valve opening and differential pressure data of chilled water pipe network in different height partitions of the building, combination of pump house outlet pressure and building height difference, construction of a hydraulic state model of the building height partitions; Based on the hydraulic state model, key compensation parameters are extracted, and a partition differential pressure compensation signal is expanded in the initial action space of the chilled water pipe network to construct a multi-parameter action space; Based on the multi-parameter action space, a deep reinforcement learning strategy is trained in a simulation environment, and a hydraulic constraint is introduced into the deep reinforcement learning strategy; The deep reinforcement learning strategy is called to generate a plurality of control instructions according to real-time partition hydraulic states to control the settings of chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal; Real-time monitoring of the water supply temperature and flow rate of the high-end terminal after the execution of the control instructions to verify whether the hydraulic states of the different height partitions of the building meet the preset target threshold; Wherein, the different height partitions of the building are high, medium and low zones; the key compensation parameters are the minimum differential pressure data of the high-end terminal and the valve margin of the medium and low zones; and the initial action space is the chilled water supply temperature, the water pump speed, the fan speed and the valve opening.
2. The deep-reinforcement learning based multi-parameter coupled optimal control method according to claim 1, wherein, The real-time acquisition of end valve opening and differential pressure data of chilled water pipe network in different height partitions of the building, combination of pump house outlet pressure and building height difference, construction of a hydraulic state model of the building height partitions, comprises: Differential pressure sensors are arranged at the ends of the chilled water pipe network in the high, medium and low zones of the building, and the differential pressure data of the different height partitions of the building are acquired in real time through the differential pressure sensors, while the valve opening of the end terminal of each partition and the pump house outlet pressure are obtained; The differential pressure data, valve opening and pump house outlet pressure are filtered and normalized to construct a standardized monitoring parameter set.
3. The deep-reinforcement learning based multi-parameter coupled optimal control method according to claim 2, wherein, The real-time acquisition of end valve opening and differential pressure data of chilled water pipe network in different height partitions of the building, combination of pump house outlet pressure and building height difference, construction of a hydraulic state model of the building height partitions, further comprises: Based on the height values of the high, medium and low zones of the building, the theoretical static pressure of different height partitions is calculated, and the effective differential pressure is determined according to the differential pressure data acquired by the differential pressure sensor, the effective differential pressure being the sum of the theoretical static pressure and the differential pressure data; Based on the effective differential pressure, a valve flow characteristic curve is constructed to estimate the water supply of each partition and obtain a water supply vector of the whole building; According to the water supply vector and the preset target water supply of each partition, in combination with the valve flow characteristic curve and the valve opening, the hydraulic state model is generated.
4. The deep reinforcement learning based multi-parameter coupled optimal control method according to claim 1, characterized in that, The real-time acquisition of end valve opening and differential pressure data of chilled water pipe network in different height partitions of the building, combination of pump house outlet pressure and building height difference, construction of a hydraulic state model of the building height partitions, further comprises: Based on the hydraulic state model, the minimum differential pressure data of the high-end terminal of the building and the valve margin of the medium and low zones are extracted in sequence to construct a high-end terminal differential pressure sequence and a medium-low zone valve opening sequence; In a set monitoring period, the minimum value of the high-end terminal differential pressure is calculated according to the high-end terminal differential pressure sequence, and a set of medium-low zone valve margin parameters is determined based on the medium-low zone valve opening sequence; An initial action space of the chilled water pipe network is acquired, and a partition differential pressure compensation signal is determined according to a minimum value of the set of low-mid zone valve margin parameters and the high zone end pressure difference, the partition differential pressure compensation signal is written into the initial action space to obtain an expanded multi-parameter action space.
5. The deep reinforcement learning based multi-parameter coupled optimal control method according to claim 1, wherein, The deep reinforcement learning strategy is trained in a simulation environment based on the multi-parameter action space, and a hydraulic constraint is introduced into the deep reinforcement learning strategy, including: Indoor temperature and humidity of each height partition of the building are acquired, and the indoor temperature and humidity are written into the multi-parameter action space, and the multi-parameter action space and the indoor temperature and humidity are normalized to a preset interval to obtain a normalized state space set; A reward function with total power consumption of the chilled water pipe network system as an optimization objective is constructed based on the state space set, and a peak penalty, a hydraulic constraint and a comfort constraint are introduced into the reward function to obtain a composite reward function; The hydraulic constraint is that the high zone end pressure difference is not allowed to be lower than a set threshold and the low-mid zone valve opening is not allowed to exceed a set upper limit; and the comfort constraint is that the indoor temperature and humidity satisfy a set temperature and humidity range.
6. The deep-reinforcement learning based multi-parameter coupled optimal control method according to claim 5, characterized in that, The deep reinforcement learning strategy is trained in a simulation environment based on the multi-parameter action space, and a hydraulic constraint is introduced into the deep reinforcement learning strategy, further including: The multi-parameter action space is taken as a candidate action of the deep reinforcement learning strategy, and the state space set is taken as an input of a deep learning framework, and deep reinforcement learning training is performed in a simulation environment to obtain a deep reinforcement learning strategy; The composite reward function is taken as an optimization objective, and the composite reward function is introduced into the deep reinforcement learning strategy.
7. The deep reinforcement learning based multi-parameter coupled optimal control method according to claim 1, characterized in that, The deep reinforcement learning strategy is called to generate a plurality of control instructions according to real-time partition hydraulic states to control settings of chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal, including: During execution of the deep reinforcement learning strategy in the chilled water pipe network, indoor temperature and humidity, water supply, effective differential pressure, valve opening and pump house outlet pressure of each height partition of the building are collected in real time to construct a real-time hydraulic state set; The real-time hydraulic state set is input into the deep reinforcement learning strategy to output optimal action instructions, and a partition differential pressure compensation signal is parsed in response to the optimal action instructions to obtain a partition differential pressure adjustment parameter; Control instructions to be executed are generated according to the partition differential pressure adjustment parameter to obtain a control instruction set for controlling settings of chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal.
8. The deep reinforcement learning based multi-parameter coupled optimal control method according to claim 7, characterized in that, The water supply temperature and flow at the high zone end after execution of the control instructions are monitored in real time to verify whether the hydraulic states of different height partitions of the building satisfy a preset target threshold, including: In response to the control instructions, the control instructions are executed according to the partition differential pressure adjustment parameter to adjust chilled water supply temperature, water pump speed, fan speed, valve opening and differential pressure target of each height partition, and hydraulic state feedback data are collected; A decision function is constructed based on high area water supply constraints, medium and low area valve surplus constraints, comfort constraints and energy efficiency constraints according to the hydraulic state feedback data, a comprehensive decision result is output by calling the decision function, and it is verified whether the comprehensive decision result meets a preset condition. 9.A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is configured to store instructions; The processor is configured to operate according to the instructions to perform the steps of the multi-parameter coupled optimization control method based on deep reinforcement learning according to any one of claims 1-8.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the multi-parameter coupled optimization control method based on deep reinforcement learning according to any one of claims 1-8.
Citation Information
Patent Citations
Central air conditioning system operation strategy optimization method based on big data and dynamic simulation
CN114383299A
Multi-cold-source annular cold supply system pressure difference control method based on reinforcement learning
CN116819945A
Hydraulic balance control method
CN118820674A
Intelligent control method of evaporative cooling air conditioning system based on deep reinforcement learning
CN120194407A
Building air conditioning system control optimization method based on mechanism-reinforcement learning and intelligent agent
CN120491485A