Multivariable collaborative optimization method for energy-saving operation strategy of heat pump machine room of public building
By employing a multivariate collaborative optimization method, combined with data prediction and deep learning models, the control strategy for heat pump rooms is optimized, solving the problem of balancing energy efficiency and comfort in heat pump rooms and achieving efficient energy consumption and comfort management.
Patent Information
- Application Number
- CN202511487842.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-01-06
AI Technical Summary
Existing energy-saving operation optimization strategies for heat pump rooms are difficult to adaptively adjust according to outdoor weather conditions and dynamic load demands, resulting in poor system energy efficiency and an inability to simultaneously meet indoor comfort requirements. Furthermore, the collaborative optimization and control of multiple devices is complex, making it difficult to achieve a balance between efficient energy consumption and comfort.
A multivariate collaborative optimization method is adopted. By collecting equipment operation status monitoring data, using Pearson correlation coefficient to screen influencing factors, combining variational mode decomposition and long short-term memory network for load forecasting, a Markov decision model and an Actor-Critic architecture deterministic policy gradient model are established to optimize the control strategy of the heat pump system.
It achieves multi-device collaborative optimization control of heat pumps and water pumps, improves system operating efficiency, can respond to load changes in real time, dynamically adjust parameters, meet multiple objectives of minimizing energy consumption and indoor comfort, and quantitatively evaluate the optimization effect.
Smart Images

Figure CN121276984A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of energy optimization technology, specifically involving a multivariate collaborative optimization method for energy-saving operation strategies of heat pump rooms in public buildings. Background Technology
[0002] In recent years, the total carbon emissions from the entire building process have accounted for approximately 50% of the nation's total carbon emissions, making it a major energy consumer. Public buildings, while accounting for about 19% of the total building area, account for approximately 38% of total building carbon emissions. Refrigeration, air conditioning, heating, and cooling systems in public buildings constitute the largest share of building energy consumption and are the primary targets for energy-saving operation optimization. Heat pump units, due to their cooling and heating functions and high efficiency, are increasingly being used as heat sources in various public buildings. Energy-saving operation optimization strategies for heat pump rooms in public buildings, in addition to reducing the energy consumption of the heat pump system, also need to meet indoor thermal comfort requirements, making it a key area of building energy-saving technology.
[0003] Currently, energy-saving operation optimization strategies for heat pump rooms have evolved from manual and automatic control based on experience (such as rule-based control methods and PID control) to the stage of intelligent energy operation and maintenance (such as reinforcement learning, fuzzy control, and model predictive control). Commonly used energy-saving operation strategies include automatic addition and subtraction of pumps, variable outlet water temperature of units, optimized load distribution among multiple units, energy optimization during transitional seasons, and pump frequency. However, in the actual operation of large-scale public building projects, the optimized control strategies for various equipment in heat pump rooms fail to achieve ideal results and are difficult to adapt to changes in outdoor meteorological environment, dynamic load demand, and operating conditions. The fundamental reasons for this are as follows:
[0004] (1) During the operation of a heat pump system, the coupling relationship between the heat pump unit, water pump, and terminal equipment is complex and variable, with many controllable parameters. Any change in control parameters may affect the system's energy consumption, making it difficult to find a suitable corresponding control and operating conditions, and there is always a deviation between the operating conditions and load requirements. For example, a common energy-saving operation strategy is to adjust the outlet water temperature or temperature difference on the load side of the heat pump unit based on the outdoor meteorological environment, which can significantly improve the coefficient of performance (COP) of the heat pump unit. However, changes in the outlet water temperature setting or temperature difference will change the circulating water volume, which may cause the circulating water pump to deviate from its optimal efficiency point, resulting in poor overall system energy efficiency. Similarly, when a water pump frequency conversion strategy is adopted, the energy efficiency of the heat pump unit and terminal equipment may decrease. Therefore, a multi-device multi-variable collaborative optimization strategy is beneficial to improving the overall energy efficiency of the heat pump system.
[0005] (2) The energy system of large public buildings is a complex nonlinear system with many control variables, high thermal inertia, and strong time lag. The building's cooling / heating load changes dynamically with the disturbances of uncertain factors such as outdoor meteorological parameters, pedestrian flow, heat dissipation equipment, thermal performance of the building envelope, and building construction. Compared with the characteristics of gradual changes in monthly, seasonal, and annual loads, the hourly and minute-level cooling / heating loads are highly volatile. The heat pump room (source side) cannot respond to the dynamic changes in load demand (load side) in a timely manner, resulting in the inability to simultaneously meet indoor comfort and system energy efficiency. Since the energy-saving optimization control of the system requires accurate prediction of cooling / heating loads at the hourly and minute levels, it is urgent to construct a short-term dynamic load prediction method with high accuracy, strong generalization, and low prediction error.
[0006] (3) In the past, the setting of control parameters for energy-saving optimization strategies relied heavily on manual intervention. For example, the threshold settings for unit start-up and shutdown, and outlet water temperature were generally based on fixed rules or experience, making it difficult to respond promptly to the instantaneous fluctuations in building load demand. Moreover, the energy-saving optimization strategy for heat pump rooms needs to simultaneously meet indoor comfort and minimize system energy consumption, which is a multi-objective optimization. Energy-saving optimization cannot improve the quality of the control strategy based on changes in operating conditions, and cannot adapt to complex and ever-changing operating conditions through intelligent real-time decision-making. For example, in extremely cold winter weather, maintenance personnel will increase the number of water pumps to increase flow to meet heating demand. Once the control strategy is fixed, it is difficult to adapt and flexibly adjust according to actual operating conditions, which may lead to low load rate operation of heat pump units. Therefore, an adaptive and self-decision-making supply-demand balance for efficient and collaborative optimization of heat pump room operation should be established.
[0007] (4) Optimization problems have a large search space but few constraints. During the optimization process, a large number of meaningless solutions will consume a lot of computation time, and the lack of constraints on the optimization objective may even lead to extreme solutions. Due to the contradictions between different building characteristics and performance, such as natural lighting and building shading, indoor environmental comfort and building energy efficiency, the actual energy-saving operation decision of public buildings becomes more complex. Relying on human experience to make decisions often leads to erroneous conclusions. Therefore, it is necessary to use a simulation platform to build a dynamic simulation model of the system, quantitatively evaluate the energy-saving effect of the optimized operation strategy of the heat pump room, realize closed-loop control and real-time optimization through cross-platform interaction, and verify the reliability of the energy-saving operation strategy. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this application proposes a multivariate collaborative optimization method for energy-saving operation strategies of heat pump rooms in public buildings.
[0009] In a first aspect, the present invention provides a multivariate collaborative optimization method for energy-saving operation strategies of heat pump room in public buildings, including:
[0010] Collect operational status monitoring data of various equipment in the heat pump room and calculate the real-time heat load or real-time cooling load of the building heat pump system hourly.
[0011] Obtain the factors that affect real-time heat load or real-time cooling load;
[0012] The Pearson correlation coefficient is calculated based on the average value of the real-time heat load or the average value of the real-time cooling load. The Pearson correlation coefficient characterizes the degree of correlation between each influencing factor and the real-time heat load or the real-time cooling load. The influencing factors corresponding to the values of the first M Pearson correlation coefficients are used as input factors with strong correlation.
[0013] Based on strongly correlated input factors, variational mode decomposition algorithm is used to decompose and denoise real-time heat load or real-time cold load to obtain Z modes and residual data.
[0014] The Z modalities and residual data are input into the pre-trained Long Short-Term Memory network to obtain Z prediction results. The Z prediction results are added together to obtain the final prediction result.
[0015] The final prediction result is used as the prediction result of the dynamic simulation model of the heat pump system of the public building. The process of establishing the dynamic simulation model of the heat pump system of the public building includes: based on the established mathematical model of the heat pump unit and the mathematical model of the energy consumption of the water pump equipment, the dynamic simulation model of the heat pump system of the public building is established using the TRNSYS simulation platform.
[0016] The heat pump system of public buildings is defined as a quadruple of the Markov decision model;
[0017] A deterministic policy gradient model with an Actor-Critic architecture is used to optimize the quadruplets of a Markov decision model, thereby obtaining the optimal control strategy for the dynamic simulation model of a public building heat pump system. The Actor-Critic architecture includes an Actor network and a Critic network, which are different neural network structures. The Actor network outputs the action space of the quadruplets based on their current spatial state, representing the behavioral selection of the current control strategy. The Critic network evaluates the control strategy selected by the Actor network and provides the corresponding reward for the selected control strategy. The deterministic policy gradient model of the Actor-Critic architecture uses gradient descent to minimize the mean squared loss function and gradient ascent to update the Actor network parameters, thereby maximizing the Q-value of the Critic network.
[0018] The strongly correlated input factors include: outdoor temperature, outdoor relative humidity, solar radiation, and indoor occupancy rate.
[0019] The four-tuples of the Markov decision model include:
[0020] The spatial state of the quadruple is defined as either strongly correlated input factors and real-time heat load or strongly correlated input factors and real-time cooling load.
[0021] The control variables of the controller in the heat pump system of public buildings are used as the action space of the quadruple;
[0022] The state transition probability of the current spatial state to the next spatial state is taken as the state transition probability of the quadruple.
[0023] Minimizing the energy consumption of heat pump systems in public buildings and maximizing indoor thermal comfort are the rewards of the quaternion.
[0024] The training process of the deterministic policy gradient model of the Actor-Critic architecture includes:
[0025] N training samples are extracted from the experience replay pool at the current iteration number. The training samples are quadruplets of Markov decision models.
[0026] Add noise to the training samples and then input the noisy training samples into the Actor network.
[0027] Based on the current control strategy, in the training samples with added noise, the Actor network selects the corresponding action in the spatial state of the next quadruple at the next time step.
[0028] Based on the corresponding action selected in the next time step, the Critic network calculates the corresponding reward value of the spatial state of the quadruple in the next time step. The corresponding reward value consists of the reward of the quadruple at the current iteration number, the discount factor, and the target value corresponding to the network parameters of the Critic network.
[0029] The Critic network parameters are updated by minimizing the mean squared loss function using gradient descent, and the Actor network parameters are updated using gradient ascent.
[0030] The updated Actor network is constructed using the updated Actor network parameters, and the updated Critic network is constructed using the updated Critic network parameters. The updated Actor network and the updated Critic network are used as the deterministic policy gradient model of the Actor-Critic architecture.
[0031] The corresponding return value is calculated using the following formula:
[0032] ;
[0033] in, This is the reward value corresponding to the i-th iteration. The reward of the quadruple in the i-th iteration This is the discount factor. Let be the spatial state of the quadruple in the (i+1)th iteration. For the action of the quadruple in the (i+1)th iteration, These are the target network parameters for the Critic network. For the space state of the quadruple is And the action of the quadruple is At that time, the target Q value corresponding to the target network parameters of the Critic network.
[0034] The gradient descent method is used to minimize the mean squared loss function to update the Critic network parameters, and the calculation formula is as follows:
[0035] ;
[0036] ;
[0037] ;
[0038] in, This is the reward value corresponding to the i-th iteration. Let be the spatial state of the quadruple in the i-th iteration. For the action of the quadruple in the i-th iteration, The mean squared loss function is... The learning rate of the Critic network. As the expected value, To update the gradient of the Critic network parameters, For Critic network parameters, For the updated Critic network parameters, For the space state of the quadruple is And the action of the quadruple is At that time, the Q-value corresponding to the target network parameters of the Critic network, This represents the total number of iterations.
[0039] The gradient ascent method is used to update the Actor network parameters in order to maximize the Q value of the Critic network. The calculation formula is as follows:
[0040] ;
[0041] ;
[0042] ;
[0043] in, This is the reward value corresponding to the i-th iteration. Let be the spatial state of the quadruple in the i-th iteration. For the action of the quadruple in the i-th iteration, The learning rate of the Actor network. The loss function of the Actor network. As the expected value, To update the gradient of the Actor network parameters, The gradient of the action performed by the Actor network. For Actor network parameters, For the updated Actor network parameters, For the space state of the quadruple is And the action of the quadruple is At that time, the Q-value corresponding to the target network parameters of the Critic network, This represents the total number of iterations. This represents the deterministic action generated by the Actor network when the spatial state is s. This indicates that the spatial state of the quadruple is And the action of the quadruple is At that time, the Q-value corresponding to the target network parameters of the Critic network, where, For the network parameters of the Critic network, Indicates that the spatial state is The Actor network parameters are Critic network parameters At that time, the Actor network output action is This is a process of repeated iterations.
[0044] The training process of the deterministic policy gradient model of the Actor-Critic architecture also includes:
[0045] The target network parameters of the Actor network and the Critic network are updated using a soft update method, and the calculation formula is as follows:
[0046] ;
[0047] ;
[0048] in, This is the soft update coefficient. For Critic network parameters, These are the target network parameters for the Critic network. For Actor network parameters, These are the target network parameters for the Actor network.
[0049] The deterministic policy gradient model employing the Actor-Critic architecture optimizes the four-tuple of the Markov decision model to obtain the optimal control strategy for the dynamic simulation model of the public building heat pump system, including:
[0050] The spatial state s of the quadruple at time t t Input into the deterministic policy gradient model of the Actor-Critic architecture;
[0051] Based on the spatial state of the quadruple at time t, the updated Actor network selects the deterministic action a at time t. t ;
[0052] Deterministic action a t And explore the Critic network after updating with noisy input to obtain the reward value r at time t. t ;
[0053] Deterministic action a t Input the dynamic simulation model of the heat pump system of the public building to obtain the spatial state s of the quadruple at time t+1. t+1 ;
[0054] The quadruple (s) t , a t , r t , s t+1 Store the deterministic action a at time t in the experience replay pool. t The optimal control strategy for the dynamic simulation model of a heat pump system in a public building.
[0055] Secondly, this application proposes an electronic device, including: one or more processors, and a memory, the memory being used to store instructions, which, when executed by the one or more processors, cause the one or more processors to execute the multivariate collaborative optimization method for the energy-saving operation strategy of the heat pump room in public buildings.
[0056] Thirdly, this application proposes a computer-readable storage medium storing executable instructions that, when executed, cause a processor to perform the multivariate collaborative optimization method for the energy-saving operation strategy of the heat pump room in public buildings.
[0057] Beneficial effects:
[0058] This application proposes a multivariate collaborative optimization method for energy-saving operation strategies of heat pump rooms in public buildings. This method achieves multi-device, multivariate collaborative optimization control of heat pumps and water pumps, significantly improving the overall energy efficiency of the system during operation. Based on the multivariate collaborative optimization method, when accurate control parameters cannot be obtained for complex multivariates, this application's method can not only efficiently handle multivariate optimization problems in a high-dimensional action space, but also accumulate experience through continuous interaction with the environment, continuously improving the quality of the strategy to achieve a globally optimal solution. It can adjust parameters in real time, enabling the system to dynamically respond to changes in building load and continuously make adaptive decisions based on outdoor meteorological environment, dynamic load demand, and operating conditions, while achieving multiple objectives such as minimizing building energy consumption and indoor comfort. By using a simulation platform to construct a dynamic simulation model of the system, the energy-saving effect of the optimized operation strategy for the heat pump room can be quantitatively evaluated. Attached Figure Description
[0059] Figure 1 Flowchart of a multivariate collaborative optimization method for energy-saving operation strategy of heat pump room in public buildings according to an embodiment of the present invention;
[0060] Figure 2 A schematic diagram of the building heat pump system according to an embodiment of the present invention;
[0061] Figure 3 A schematic diagram of measured hourly data of winter heat load according to an embodiment of the present invention;
[0062] Figure 4 A schematic diagram illustrating the principle of the DDPG optimization algorithm in this embodiment of the invention;
[0063] Figure 5 A schematic diagram of the joint optimization dynamic simulation model of this invention embodiment;
[0064] Figure 6 A schematic diagram of the simulation process of an embodiment of the present invention. Detailed Implementation
[0065] The specific implementation methods of this application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0066] By collecting building energy consumption monitoring data, a high-precision short-term dynamic load prediction model was constructed, and mathematical models of different equipment in the heat pump system were established. A dynamic simulation model of the heat pump room was built using the TRNSYS simulation platform and programming optimization algorithms, achieving high-precision heat pump system operation simulation and advanced control strategies. Simultaneously, data-driven artificial intelligence optimization algorithms were developed using tools such as Python and MATLAB, and a system optimization control strategy was designed. Finally, reinforcement learning algorithms were combined to optimize the operation strategy of the public building heat pump system. Closed-loop control and real-time optimization were achieved through cross-platform interaction. A complete system dynamic model was also constructed using the simulation platform. Energy consumption indicators and energy efficiency ratios under different optimization strategies were compared, quantitatively evaluating the energy-saving effect of the optimization strategies and verifying the stability of the established optimization model.
[0067] Example 1:
[0068] This embodiment provides a multivariate collaborative optimization method for energy-saving operation strategies of heat pump room in public buildings, such as... Figure 1 As shown, it includes:
[0069] Step S1: Collect the operating status monitoring data of each device in the heat pump room, and calculate the real-time heat load or real-time cooling load of the building heat pump system hourly.
[0070] This embodiment takes the optimized control of energy-saving operation of winter heating in a public building as an example. The building is a comprehensive commercial building in Shenyang City, covering an area of 15,982 square meters, with a total building area of 107,737 square meters. The standard floor area is approximately 1,400 square meters, and the building has 26 floors in total, 24 above ground and 2 underground. Floors 1-3 are low-rise commercial facilities, and floors 4-24 are high-rise office areas. The heating season is from November 1st to March 31st of the following year, with an effective heating area of approximately 80,000 square meters in winter. The winter heating system of this public building uses parallel operation of screw-type water source heat pump units. The building heat pump system is as follows... Figure 2 As shown, the equipment parameters are listed in Tables 1 and 2.
[0071] Table 1: Performance parameters of water source heat pump units;
[0072]
[0073] Table 2: Water Pump Performance Parameters;
[0074]
[0075] The components of a heat pump system include:
[0076] (1) Equipment sensing layer: It consists of system equipment (heat pump unit, water pump), monitoring equipment (temperature / pressure sensor, smart meter) and controller (PLC, frequency converter). The PLC controller is directly connected to the heat pump unit and sensors to collect key parameters such as supply and return water temperature, flow rate, and power consumption in real time. At the same time, the frequency converter is integrated to realize the speed regulation of the circulating water pump and the submersible pump.
[0077] (2) Data transmission layer: The collected data is uploaded after standardized encapsulation to form the original database. The transmission layer can transmit various types of collected data to the intermediate server. This layer includes servers, gateways, data interfaces, etc. The data is transmitted to the edge server after encryption. Transmission delay is eliminated through timestamp alignment and data verification.
[0078] (3) Data processing layer: The main function is to perform data operations on the collected data through data mining, cloud computing and other technologies, such as calculating system energy-saving indicators and data cleaning, to remove obviously erroneous or duplicate data, to ensure the basic accuracy and usability of the data, and to provide a foundation for subsequent data analysis.
[0079] (4) Application layer: It can be displayed in the form of models or charts, and can also realize decision support and automatic statistical analysis. It can automatically generate simple reports to help operation and maintenance personnel quickly understand the basic situation of the data and facilitate operation and maintenance personnel management.
[0080] The parameters of each monitoring device are shown in Table 3.
[0081] Table 3: Parameters of each monitoring device;
[0082]
[0083] Table 4: Data Types for Building Energy Consumption Monitoring;
[0084]
[0085] In this embodiment, the real-time heat load is calculated based on real-time data from temperature and flow sensors, as shown in equation (1):
[0086] (1)
[0087] in, Real-time heat load value (kW); The specific heat of the heating hot water at constant pressure is taken as 4.19 kJ / (kg·℃); m is the total flow rate of the load-side pipeline (kg / s), measured by the flow meter; ΔT is the temperature difference between the supply and return water (℃), measured by the temperature sensor of the supply and return water main pipe. Figure 3 This is the measured hourly data of the building's winter heat load for 2022-2023.
[0088] This embodiment uses building heat load prediction as an example. The method in this embodiment is also effective for cooling load prediction; the cooling load can be calculated in the same way, and will not be described in detail here.
[0089] Step S2: Obtain the factors that affect the real-time heat load or real-time cooling load;
[0090] In this embodiment, outdoor temperature, outdoor relative humidity, solar radiation, wind speed, wind direction, and number of people indoors are selected as factors that affect real-time heat load or real-time cooling load.
[0091] Step S3: Calculate the Pearson correlation coefficient based on the average value of the real-time heat load or the average value of the real-time cooling load. The Pearson correlation coefficient characterizes the degree of correlation between each influencing factor and the real-time heat load or the real-time cooling load. The influencing factors corresponding to the values of the first M Pearson correlation coefficients are used as input factors with strong correlation.
[0092] Step S4: Based on strongly correlated input factors, the variational mode decomposition algorithm is used to decompose and denoise the real-time heat load or real-time cooling load to obtain Z modes and residual data;
[0093] In this embodiment, based on outdoor temperature, outdoor relative humidity, solar radiation, wind speed, wind direction, and the number of people indoors, the Pearson correlation coefficient is used to quantitatively calculate the correlation strength of each influencing factor, automatically eliminating weakly correlated factors and selecting input factors with strong correlation to the prediction of cooling / heating loads. A VMD (Variational Mode Decomposition) algorithm is used to decompose and denoise the real-time cooling / heating load time series data of the building. The decomposition results of the real-time cooling / heating load data include Z modes and residual data. During prediction, each mode and residual are used as the output for prediction, and the final prediction result is the sum of the predicted values obtained each time.
[0094] Unobservable factors affecting building heat load, such as random equipment operation, constant personnel changes, and operational adjustment methods, lead to noise interference in time-series data. This noise, mixed in with the monitoring data, significantly impacts the load. This embodiment employs the VMD algorithm for data denoising. The advantages of this algorithm are twofold: firstly, by capturing the coordinated variation patterns among multiple signals, it accurately extracts and separates each modal component, enabling simultaneous data processing in both time and frequency dimensions. This facilitates efficient decoupling of multivariate signals and improves the stability of time-series data. Secondly, VMD possesses adaptive frequency decomposition characteristics, adaptively allocating suitable frequency bands to each mode, significantly reducing multimodal aliasing and improving the accuracy of short-term heat load prediction.
[0095] The principle is to decompose a complex signal into sub-signals with fixed bandwidth, which are called Intrinsic Mode Functions (IMFs). In addition, in order to extract the features most correlated with the heat load data and reduce feature redundancy, it is necessary to calculate the Pearson correlation coefficient between the IMF and the heat load under each influencing factor, thereby completing the extraction of supplementary input features for load forecasting.
[0096] In the load forecasting modeling process, the various influencing factors exhibit significant heterogeneity in the real-time cooling / heating load time series of buildings. Using factors with a high degree of influence as forecast input can improve forecast accuracy, but introducing weakly correlated factors may introduce more noise and redundant information, degrading model performance and reducing its generalization ability, thus reducing the model's predictive power. Therefore, to improve the model's forecast accuracy, input factors are screened based on their feature contribution before forecasting.
[0097] Pearson correlation coefficient The correlation between each influencing factor and the building's cooling / heating load is represented by the formula (2):
[0098] (2)
[0099] in, Here, n is the Pearson correlation coefficient; n is the sample size; and n is the vector. and These represent the average values of vectors X (data on each influencing factor) and Y (real-time data on building heat load), respectively. The value range is [-1, 1]. When the value is close to 1 or -1, the correlation between the two variables X and Y is higher; when the value is closer to 0, the correlation between the two variables is lower.
[0100] In this embodiment, two weakly correlated factors, wind direction and wind speed, are first filtered out to prevent the uncertainty of prediction results caused by redundant input data and to improve computational efficiency. Although there is a thermodynamic correlation between the theoretical value of the airtightness of the building envelope and wind pressure infiltration, actual monitoring data shows that in buildings with an airtightness level ≥8, when encountering specific temperature-velocity coupling conditions (such as low temperature and low wind speed / high temperature and high wind speed), the temperature field and the thermal disturbance generated by human activities will produce a nonlinear coupling effect, and its thermal inertia parameter significantly covers the transient influence of wind speed. This thermodynamic shielding effect results in a low variance contribution rate of wind environment parameters in the multiple regression model, verifying the rationality of the feature selection decision.
[0101] In this embodiment, a Pearson correlation coefficient strength table was set up according to Table 5, and weakly correlated influencing factors were automatically deleted. Finally, four factors with high correlation were selected: outdoor temperature, outdoor relative humidity, solar radiation, and indoor occupancy rate.
[0102] Table 5: Thresholds for Pearson correlation coefficients of the degree of correlation between influencing factors;
[0103]
[0104] To extract data features, reduce mode aliasing in real-time heat load time-series data, and improve the stability of real-time heat load data noise reduction processing, the VMD algorithm is used for decomposition, and the intrinsic mode function (IMF) and Pearson correlation coefficient are calculated for each influencing factor.
[0105] Step S5: Input the Z modalities and residual data into the pre-trained Long Short-Term Memory network to obtain Z prediction results. Add the Z prediction results together to obtain the final prediction result.
[0106] In this embodiment, an LSTM (Long Short-Term Memory) method is used to predict the dynamic short-term heat load of buildings. Historical data on building outdoor environment and cooling / heating load from the past two years are selected as the dataset for this prediction. The data recording interval is 1 hour, totaling 3624 data samples, which are divided into training and test sets in a 7:3 ratio. The parameters and structure settings of the pre-trained LSTM network are detailed below:
[0107] The pre-trained Long Short-Term Memory (LSTM) network consists of an input layer (receiving time-series data as input), an LSTM layer with 60 hidden units (processing the input time-series data and memorizing long and short-term dependencies), a ReLU activation layer (introducing nonlinearity to enhance the network's expressive power), and a regression layer (calculating the loss to guide network training and optimization). Training uses the Adam optimization algorithm. The parameter settings for the pre-trained LSM network are shown in Table 6.
[0108] Table 6: Parameter settings for LSTM neural networks;
[0109]
[0110] Its iterative process is as follows:
[0111] (1) Divide all samples into (number of samples / 12) rounds;
[0112] (2) Take one round of samples and input them into the pre-trained Long Short-Term Memory network;
[0113] (3) Calculate the loss function of the pre-trained Long Short-Term Memory network;
[0114] (4) Calculate the gradient through backpropagation;
[0115] (5) Update the parameters of the pre-trained Long Short-Term Memory network;
[0116] (6) Repeat steps (2) to (5) in order until all samples have been traversed (one training session is completed);
[0117] (7) Repeat the training 200 times, or meet the early stop condition.
[0118] In practical engineering projects, online real-time prediction of building energy consumption load requires not only high prediction accuracy but also a fast response speed, which is a crucial factor in practical applications. The prediction results are presented using January 2022 as an example. The horizontal axis represents hours, with 12 hours equaling one day. For example, January 1st represents 1:00 AM to 12:00 PM, corresponding to 7:00 AM to 7:00 PM, totaling 30 days and 360 hours. All models perform predictions 1, 2, and 3 hours in advance, i.e., 1 hour, 2 hours, and 3 hours in advance.
[0119] Step S6: Use the final prediction result as the prediction result of the dynamic simulation model of the heat pump system of the public building. The process of establishing the dynamic simulation model of the heat pump system of the public building includes: using the TRNSYS simulation platform to establish the dynamic simulation model of the heat pump system of the public building based on the established mathematical model of the heat pump unit and the mathematical model of the energy consumption of the water pump equipment.
[0120] In this embodiment, the mathematical model of the heat pump unit is established as follows:
[0121] The heating capacity of a heat pump unit at full load under different operating conditions is It can be expressed by equation (3):
[0122] (3)
[0123] in, The heating capacity (kJ / h) is the heating capacity when the machine is running at full load under rated operating conditions. This is a correction factor for the heating capacity under actual operating conditions.
[0124] The outlet water temperature on the condenser side can be calculated using equation (4):
[0125] (4)
[0126] in, The condenser outlet water temperature (°C); Measure the return water temperature (°C) of the condenser. Condenser side water flow rate (m³) 3 / h). Water pump power, The condenser releases heat, Specific heat of cooling water at constant pressure.
[0127] The semi-empirical mathematical model of the coefficient of performance of a heat pump unit can be expressed by equation (5):
[0128] (5)
[0129] in, The input power (kW) of the unit under rated operating conditions; This is a correction factor for the unit's input power under full load; This is a correction factor for the unit's input power under partial load. COP is the coefficient of performance (COP) for heat pump units.
[0130] , , The empirical formulas are shown in equations (6) to (8):
[0131] (6)
[0132] (7)
[0133] (8)
[0134] Among them, a1~f1, a2~f2, a3~e3, a4~e4, and a5 are the curve fitting coefficients for the unit operation; This is the ratio of the evaporator-side water flow rate to the rated water flow rate. This is the ratio of the condenser-side water flow rate to the rated water flow rate. It is the ratio of the evaporator-side outlet water temperature to the rated outlet water temperature; PLR is the ratio of the condenser-side return water temperature to the rated return water temperature. PLR is the partial load rate.
[0135] Solving equations (9), (10), and (11), the data samples for each fitting coefficient are from the energy monitoring system. Using MATLAB, multivariate nonlinear regression analysis is performed, and the equations are further solved as follows:
[0136] (9)
[0137] (10)
[0138] (11)
[0139] The coefficient of determination R² for the regression model is 0.9972. The coefficient of determination R² for the regression model is 0.9981. The coefficient of determination R² of the regression model is 0.9989, indicating a good fit.
[0140] In this embodiment, the mathematical model of the water pump unit is established as follows:
[0141] The heat pump system includes a load-side circulating water pump and a source-side submersible pump. The load-side pump is responsible for circulating the building's heating water, while the source-side pump is responsible for circulating the groundwater from the source side. This heat pump system employs a primary pump design; when the pump's operating conditions change, its speed is dynamically adjusted according to actual needs. Based on the similarity law of pumps, we know that:
[0142] (12)
[0143] Where n0 is the design speed of the water pump, and n1 is the actual operating speed of the water pump (r / min). The shaft power of the water pump is designed. The actual operating shaft power of the water pump (m) 3 / h); actual head of water pump Water pump design head, actual flow rate of water pump Water pump design flow rate.
[0144] During the operation of a water pump with variable flow rate, its energy consumption and flow rate exhibit a quadratic function relationship, which can be expressed by equation (13):
[0145] (13)
[0146] in, Let be the power of the variable frequency water pump (kW); a, b, and c are the water pump energy consumption fitting coefficients, and G is the water pump flow rate.
[0147] The fitting results of the energy consumption curves of the load-side pump and the water source-side pump were obtained by MATLAB calculation, as shown in equations (14)-(16):
[0148] (14)
[0149] (15)
[0150] (16)
[0151] in, Power (kW) of circulating water pumps 1# and 2# on the load side; Power of the No. 3 circulating water pump on the load side (kW); The flow rates (m³ / s) of circulating water pumps #1 and #2 on the load side 3 / h); The flow rate of the No. 3 circulating water pump on the load side (m³) 3 / h); The power of the submersible pump on the water source side (kW); The flow rate on the water source side (m³) 3 / h).
[0152] The coefficients of determination R of the regression model for circulating water pumps #1 and #2 on the load side 2 The coefficient of determination R0 of the regression model for circulating water pump #3 is 0.9979. 2 The coefficient of determination R0 of the submersible pump regression model on the water source side is 0.9964. 2 The value is 0.9947, which indicates a good fit.
[0153] Based on the mathematical models of the heat pump unit and the energy consumption of the water pump equipment established above, a dynamic simulation model of the heat pump system for public buildings is obtained using the TRNSYS simulation platform, including:
[0154] (1) Module configuration: Add the required modules for the building energy system, mainly including water source heat pump module, circulating water pump module, and submersible pump module, as shown in Table 7.
[0155] (2) Importing external data: Add external files to guide system control, mainly including weather and environmental data and building heat load data.
[0156] (3) Parameter setting: Set the parameters according to the actual performance parameters of each device in Table 1 and Table 2, and fill in the fitting coefficients of the mathematical model of each device.
[0157] (4) Control strategy editing: Use the calculator module or programming module to convert the actual operation scheme of the system into a mathematical model to control the equipment. It can generate heat pump start and stop signals based on load and weather data, adjust the water pump speed to achieve variable flow optimization, or adjust the heat pump unit outlet water temperature to achieve variable outlet water temperature optimization.
[0158] (5) Simulation result printing settings: Set the printing component, which can print the required simulation results.
[0159] (6) System topology connection: Connect each component according to the actual situation of the system, connect each module port according to the physical process (such as heat pump outlet → water pump inlet → terminal equipment), and set the signal transmission formula.
[0160] (7) Run the dynamic simulation model of the heat pump system of the public building and obtain the simulation results.
[0161] Establish the component connections for a dynamic simulation model of a heat pump system in a public building. In the TRNSYS system, the connections between system components follow strict logical relationships. Interactions between different modules require setting reasonable input and output parameters and matching appropriate data formats to ensure correct information flow. Furthermore, the control module plays a crucial role in system operation, adjusting the working states of each component to ensure the simulation operates according to the established physical laws and mathematical models.
[0162] The numbering and function of the main modules in the dynamic simulation model of the heat pump system of this public building are shown in Table 7.
[0163] Table 7: Main modules in the simulation model;
[0164]
[0165] In the simulation system, the heat pump unit module (Type 225) has a framework for its input and output variables during system cycling. For this component, only the heat pump unit's on / off state and the outlet temperature of the heating hot water are controllable signals; other variables are performance parameters passed from the previous component or set before operation. For the water pump module, only its on / off state and operating frequency are controllable signals. Changes in the heat pump system's flow rate can be achieved by altering the operating frequency; other variables are parameters passed from the previous component or set before operation.
[0166] Step S7: Define the heat pump system of the public building as a quadruple in the Markov decision model;
[0167] Based on Markov Decision Process (MDP) and Deep Deterministic Policy Gradient Algorithm (DDPG), an energy-saving optimization control process for heat pump systems in public buildings is constructed.
[0168] The four-tuples of the Markov decision model include:
[0169] The spatial state of the quadruple is defined as either strongly correlated input factors and real-time heat load or strongly correlated input factors and real-time cooling load.
[0170] The control variables of the controller in the heat pump system of public buildings are used as the action space of the quadruple;
[0171] The state transition probability of the current spatial state to the next spatial state is taken as the state transition probability of the quadruple.
[0172] Minimizing the energy consumption of heat pump systems in public buildings and maximizing indoor thermal comfort are the rewards of the quaternion.
[0173] In this embodiment, the operation process of the heat pump system in a public building is defined as a Markov decision process (MDP), and the MPD model of the heat pump system is defined as a quadruple {s, a, p, r}:
[0174] (1) State space s: The state space described by the outdoor environment of the building heat pump room and the relevant parameters inside the building is the environmental information at the current moment. The outdoor temperature, outdoor relative humidity, solar radiation, occupancy rate and heat load are taken as the state space.
[0175] (2) Action space a: Consists of the controllable instructions of the system controller, containing a set of all possible actions. The control variables for energy-saving operation optimization are selected based on the actual situation of the system. In this case, the water supply temperature of the heat pump unit and the operating frequency signals of the three water pumps are set as actions.
[0176] (3) State transition probability p: refers to the state transition probability from state s at the current time t to state s' at the next time. This term usually includes a penalty term, that is, the immediate reward obtained after performing control action a in state s. Its determination involves the actual environmental state after the agent performs the control action. The agent usually uses the Monte Carlo method to make an unbiased estimate of the transition probability through multiple sampling and observation.
[0177] The state transition probability can be expressed by equation (17):
[0178] (17)
[0179] in, This represents the state transition probability in action space a from state s at time t to state s' at the next time step. Let t be the action space. Let be the state space at time t. The state space at time t+1.
[0180] In reinforcement learning, when the state space is a time series, the state transition during simulation optimization can be regarded as deterministic. The next state of the system is only related to the current state. It can be regarded as a special case of MDP and can be represented by equation (18):
[0181] (18)
[0182] in, The state s at time t transitions to the state at the next time step. State transition probability, This is the state space at time 1.
[0183] (4) Reward r: refers to the feedback on whether the action performed in the current state achieves the target effect. In this embodiment, the goal is to minimize the energy consumption of the heat pump system and maximize the indoor thermal comfort of the building. The lower the energy consumption, the higher the reward. The indoor thermal comfort of the building is used to calculate the penalty term to control the reward and avoid the situation where the energy consumption decreases after the action is performed but the comfort exceeds the range T.
[0184] (19)
[0185] Where λ is the thermal comfort penalty coefficient; Let be the system energy consumption at time t. This optimization algorithm aims to maximize the total reward. However, minimizing energy cost contradicts maintaining the ideal temperature; reducing cost or energy consumption leads to a decrease in reward. Therefore, to attempt to balance these two objectives and maintain consistency with the optimization goal, all terms are set to negative values. Let T be the indoor temperature. and The upper and lower limits of temperature are set to maintain indoor comfort, with an upper limit of 25℃ and a lower limit of 21℃. The design value for the building's indoor temperature in winter is 23℃. Therefore, the closer the T value is to 23℃, the more comfortable the room temperature is. Exceeding the upper or lower limit indicates discomfort. This value is used as a penalty parameter in the calculation to ensure indoor comfort.
[0186] The MPD model of a heat pump system clarifies the agent, environment, and reward mechanism in the reinforcement learning process. As the agent interacts with the environment, it performs different actions and obtains behavioral evaluations, which serve as guidance for subsequent actions. When a behavior receives a positive reward, the agent is more likely to repeat that behavior under similar environmental conditions; conversely, when a behavior is penalized, poorly rewarded, or receives a negative reward, its probability decreases. Through continuous adjustments, the agent gradually learns to select the optimal action under various meteorological environments, dynamic loads, and operating conditions. Therefore, the MPD model of a heat pump system is essentially a mapping of a complex, nonlinear, high-dimensional discrete space.
[0187] Step S8: The deterministic policy gradient model of the Actor-Critic architecture is used to optimize the quadruplets of the Markov decision model to obtain the optimal control strategy for the dynamic simulation model of the public building heat pump system. The Actor-Critic architecture includes an Actor network and a Critic network, which are different neural network structures. The Actor network outputs the action space of the quadruplets based on the current spatial state of the quadruplets, representing the behavior selection of the current control strategy. The Critic network evaluates the control strategy selected by the Actor network and provides the Actor network with the corresponding reward for the selected control strategy. The deterministic policy gradient model of the Actor-Critic architecture uses gradient descent to minimize the mean square loss function to update the Critic network parameters and gradient ascent to update the Actor network parameters to maximize the Q value of the Critic network.
[0188] In this embodiment, after establishing a dynamic simulation model of the heat pump system of a public building, although the efficiency value of each device at the next moment can be predicted, in a scenario with complex interactions of multiple variables, it is impossible to know how each device should be controlled at each moment to achieve the optimal value. To address this situation, a deterministic strategy gradient model with an Actor-Critic architecture is proposed.
[0189] The advantage of the deterministic policy gradient model in the Actor-Critic architecture is that:
[0190] (1) Traditional policy gradient methods are computationally inefficient because they require integral operations in a high-dimensional policy distribution space. Moreover, the complexity of the discrete action space optimization process is high when dealing with complex systems. Therefore, this embodiment adopts deterministic policy gradient (DPG). By using deterministic policies to replace traditional random policies (i.e. step S8), the policy distribution integral operation is effectively avoided, and the algorithm efficiency is significantly improved.
[0191] (2) While deterministic policies effectively improve computational efficiency, they still face challenges of policy oscillation and training instability in high-dimensional state and action spaces. To address this, this embodiment introduces the DDPG (Deep Deterministic Policy Gradient) algorithm, which integrates deep neural networks and an Actor-Critic architecture to construct a dual-network collaborative mechanism: experience replay and a target network. Experience replay, by storing historical state transition samples and randomly sampling them for training, breaks the temporal correlation between data, alleviating the variance problem in policy updates. The policy network (Actor) directly outputs deterministic action instructions, while the value network (Critic) dynamically evaluates the value of actions through temporal difference error (TD). By minimizing the TD error, the Actor network ensures that the policy gradually approaches the optimum.
[0192] The training process of the deterministic policy gradient model of the Actor-Critic architecture includes:
[0193] Step S100: Extract N training samples from the experience replay pool at the current iteration number. The training samples are quadruplets of the Markov decision model.
[0194] Step S101: Add noise to the training samples and input the noise-added training samples into the Actor network;
[0195] In this embodiment, noise can be added during DDPG training to promote action exploration, break the determinism of the policy, help discover better actions, and prevent the algorithm from getting trapped in local optima too early. Noise, such as Ornstein-Uhlenbeck noise, is added to the deterministic actions generated by the policy network. See equation (20), and the noise update is shown in equation (21):
[0196] (20)
[0197] (twenty one)
[0198] in, Let be the noise at time t. The noise at time t-1, The noise attenuation coefficient determines the rate of noise change. To update the time step, set it to 0.01; This is the noise expansion factor, controlling the fluctuation amplitude of the noise. randn(size(action)) can generate a random number that follows a standard normal distribution with the same size as the action space, used to introduce randomness into the noise.
[0199] Step S102: Based on the current control strategy, in the training samples after adding noise, the Actor network selects the corresponding action in the spatial state of the next quadruple at the next time step;
[0200] Step S103: Based on the corresponding action selected in the next time step, the Critic network calculates the corresponding reward value of the spatial state of the quadruple in the next time step. The corresponding reward value consists of the reward of the quadruple in the current iteration number, the discount factor, and the target value corresponding to the network parameters of the Critic network.
[0201] Step S104: Minimize the mean squared loss function using gradient descent to update the Critic network parameters, and update the Actor network parameters using gradient ascent.
[0202] During model training, each iteration extracts a batch of sample data of size N from the experience pool. The Actor target network selects an action in the next state of the sample based on the current policy. This action is used by the Critic target network to calculate the target Q-value for the next state, expressed as:
[0203] (twenty two)
[0204] in, This is the reward value corresponding to the i-th iteration. The reward of the quadruple in the i-th iteration This is the discount factor. Let be the spatial state of the quadruple in the (i+1)th iteration. For the action of the quadruple in the (i+1)th iteration, These are the target network parameters for the Critic network. For the space state of the quadruple is And the action of the quadruple is At that time, the target Q value corresponding to the target network parameters of the Critic network.
[0205] During training, the Critic network updates its parameters by minimizing the mean squared loss function, which is:
[0206] (twenty three)
[0207] in, Expected value is a mathematical description of the average value of a random variable.
[0208] Its gradient calculation follows formula (24):
[0209] (twenty four)
[0210] The TD error can be expressed by equation (25):
[0211] (25)
[0212] in, The TD error, which combines estimates of immediate and future rewards, measures the difference between the current value estimate and the target value estimate. It can be used to update the value function to make it closer to the target value.
[0213] The network parameters are updated using the average gradient of N samples, and the Critic network parameters are updated via gradient descent. The process follows formula (26):
[0214] (26)
[0215] in, The learning rate of the Critic network. For Critic network parameters, These are the updated Critic network parameters.
[0216] The goal of an Actor network is to learn how to select actions to maximize the Q-value, i.e., optimize the policy parameters. It can be expressed by equation (27):
[0217] (27)
[0218] Based on the Critic network parameters and the empirical pool samples, its gradient calculation follows formula (28):
[0219] (28)
[0220] Update Actor network parameters using gradient ascent. The process follows formula (29):
[0221] (29)
[0222] in, The learning rate of the Actor network. For Actor network parameters, These are the updated Actor network parameters. To update the gradient of the Actor network parameters, The gradient of the action performed by the Actor network. This represents a deterministic action generated in the policy network (Actor), where s represents the state in the network. The task of the Critic network is to evaluate the actions chosen by the Actor network and provide feedback to the Actor network. This is typically represented by a Q-value function. It estimates the reward for a given state and action, where These are the online network parameters for Critic. The task of an Actor network is to output an action based on the current state, representing the behavior selection of the current policy, which can be determined by... It indicates. Among them, For the actor's online network parameters.
[0223] This is a process of repeated cycles and iterations.
[0224] Step S105: Use the updated Actor network parameters to form an updated Actor network, use the updated Critic network parameters to form an updated Critic network, and use the updated Actor network and the updated Critic network as the deterministic policy gradient model of the Actor-Critic architecture.
[0225] The training process of the deterministic policy gradient model of the Actor-Critic architecture also includes:
[0226] The parameters of the target network are gradually synchronized with the online network through soft updates. Because the parameters in the target network are updated slowly using soft updates, its output is more stable, and the calculation of the target value using the target network is naturally more stable, thus further ensuring a smoother learning process for the network. The soft update process can be represented by equations (30) and (31):
[0227] (30)
[0228] (31)
[0229] in, ≪1 (take) =0.001) is the soft update coefficient; These are the target network parameters for the Actor.
[0230] The principle of DDPG optimization algorithm is as follows: Figure 4 As shown.
[0231] The deterministic policy gradient model employing the Actor-Critic architecture optimizes the four-tuple of the Markov decision model to obtain the optimal control strategy for the dynamic simulation model of the public building heat pump system, including:
[0232] The spatial state s of the quadruple at time tt Input into the deterministic policy gradient model of the Actor-Critic architecture;
[0233] Based on the spatial state of the quadruple at time t, the updated Actor network selects the deterministic action a at time t. t ;
[0234] Deterministic action a t And explore the Critic network after updating with noisy input to obtain the reward value r at time t. t ;
[0235] Deterministic action a t Input the dynamic simulation model of the heat pump system of the public building to obtain the spatial state s of the quadruple at time t+1. t+1 ;
[0236] The quadruple (s) t , a t , r t , s t+1 Store the deterministic action a at time t in the experience replay pool. t The optimal control strategy for the dynamic simulation model of a heat pump system in a public building.
[0237] In practice, the interaction process between each time step and the environment includes:
[0238] (1) System initialization: Input system parameters, set the neural network structure of the Actor network and the Critic network, and initialize the network parameters and spatial state s0. Among them, the Critic network is used to update the network parameters using a deterministic policy gradient; (2) Actor selects action: The Actor network selects action based on the current state s0. t Generate deterministic action a t Add exploration noise The purpose of this step is to ensure that the strategy fully explores the action space in the early stages of training.
[0239] (3) Execute actions in the environment: Collect interaction data, construct state transition tuples, and assign action a t Input simulation environment (TRNSYS simulation model), observation reward r t and the next state s t+1 .
[0240] (4) Store state transition data: store the quadruple (s t , a t , r t , s t+1 The purpose of storing the data in the experience replay pool is to break the temporal correlation of the data and support offline strategy learning.
[0241] In this embodiment, the network training process is performed at regular time intervals:
[0242] (5) Sampling small batches of data: Randomly sample N small batches of data (s) from the experience pool. i , a i , r i , s i+1 By using batch data to calculate gradients, variance is reduced and training stability is improved.
[0243] (6) Calculate the target Q value y i The target Actor network (where the Actor network generates actions, and the target Actor network is a result achieved during the iteration process) generates the next state action 'a'. i +1=μ′(s i +1), the target network parameters of the Critic network are used to calculate the target Q value, i.e., Q′(si+1, ai+1), with the aim of constructing a temporal difference (TD) objective for the loss function of the Critic network. Here, μ′ represents an action selection of the actor network, and each action selection is random and variable.
[0244] (7) Update the Critic network: Minimize the loss function using gradient descent and update the Critic parameters. This enables the Critic network to accurately evaluate the value of actions.
[0245] (8) Update the Actor network: Maximize the Q-value through gradient ascent and update the Actor parameters. The goal is to optimize the strategy to generate actions with higher Q values.
[0246] (9) Soft update target network: Parameters of the soft update target network and The parameters of the online network and the target network are gradually mixed to stabilize the calculation of the target Q value and avoid training oscillations.
[0247] (10) Repeat steps (2) to (9) until the termination condition is met.
[0248] In this embodiment, steps (7) and (8) update the network by updating the network parameters. Step (9) prevents network oscillation by using soft updates. By using soft updates to update slowly, the output will be more stable, and the calculation of the target value using the target network will also be more stable, thereby further ensuring a smoother learning process for the network.
[0249] Based on the constructed energy-saving optimization control strategy, a joint optimization dynamic simulation model of the system was established using TRNSYS and its operation was simulated.
[0250] The inputs (state space: outdoor temperature, outdoor relative humidity, solar irradiance, occupancy rate, and heat load; variables used to calculate returns: system energy consumption and indoor temperature) and outputs (action space: unit water supply temperature, #1 circulating pump frequency signal, #2 circulating pump frequency signal, and #3 circulating pump frequency signal) collected in step S1 are connected to the DDPG energy-saving optimization algorithm program written in MATLAB through the Type155 module of the TRNSYS simulation platform.
[0251] The system co-optimization dynamic simulation model established using TRNSYS, such as Figure 5 As shown in the figure. In this embodiment, the entire process from input data to the final simulation result requires no manual intervention and can be fully automated. After the optimization algorithm is connected to the input and output of the simulation system, the simulation model is started, and the simulation process is as follows: Figure 6 As shown.
[0252] Based on the actual building conditions, a joint optimization dynamic simulation model of the system was established on the TRNSYS platform. The inlet and outlet water temperature difference, energy efficiency coefficient, heating capacity, and equipment energy consumption of the heat pump system on the load side and water source side were compared under different energy-saving optimization strategies (addition and subtraction control strategy, water pump frequency conversion control strategy, and unit variable outlet water temperature control strategy). The energy-saving effect of heat pump units and systems under different energy-saving optimization strategies was evaluated.
[0253] The booster / subtractor control strategy is a commonly used energy-saving control strategy, used to compare its effectiveness with the method in this embodiment. The booster / subtractor control logic adjusts the number of units in operation based on continuous load changes. When the heat load exceeds 0.2 times the rated heat capacity of the unit, the first unit is activated; when it exceeds 0.9 times the rated heat capacity of the first unit, the second unit is activated, and the load is evenly distributed. When the load exceeds 0.9 times the rated heat capacity of two units, the third unit is activated, and so on. This strategy does not consider optimized control of terminal devices. The parallel water pumps on the load side have a fixed and unadjustable inlet water temperature on the source side, and both the circulating water pump and the submersible pump on the source side operate at full frequency. The calculated load value during the design phase is large, potentially leading to oversized equipment selection.
[0254] Simulation results show that the system's hot water supply temperature is generally maintained at around 45℃, while the system return water temperature ranges from 40.02℃ to 43.5℃, with an average return water temperature of 42.57℃. The average supply and return water temperature difference is 2.43℃, significantly lower than the set temperature difference of 5℃. The deep water well temperature on the source side remains relatively stable at around 15℃, with the outlet water temperature ranging from 11.25℃ to 13.54℃. The system's average outlet water temperature is 12.56℃, and the average inlet and outlet water temperature difference is small, only 2.44℃, also significantly lower than the set temperature difference of 6℃. The average COP for the heating season is summarized in Table 8. The unit's hourly COP ranges from 3.29 to 3.54, with an average COP of 3.49.
[0255] Table 8: Average COP for each month and heating season;
[0256]
[0257] Water pump frequency conversion control strategy. Energy saving is achieved by changing the pump motor frequency, thereby altering the pump speed. All pumps operate at 50Hz, simulating constant temperature difference frequency conversion operation, with the supply and return water temperature difference set at 5℃. The frequency conversion control strategy is as follows: When the load is low, to maintain a stable 5-degree temperature difference, the pump speed ratio will be less than 0.6. In this case, the frequency converter will operate at a speed ratio of 0.6, abandoning the control of the 5-degree supply and return water temperature difference. When the load changes, the pump frequency adjustment range is between 30Hz and 50Hz, meaning the output frequency conversion signal is between 0.6 and 1, with the supply water temperature set at 45℃. When the load is very high, exceeding the pump's overload range, the pump frequency cannot exceed the limit; therefore, the frequency conversion signal is output at the maximum value of 1.
[0258] Table 9: Average COP for each month and heating season;
[0259]
[0260] The hourly COP of the units ranged from 3.44 to 3.79, with an average COP of 3.70, which was slightly higher than the addition / reduction strategy.
[0261] The unit employs a variable outlet water temperature control strategy. The outlet water temperature is adjusted based on the outdoor temperature, reducing the compressor's workload, operating frequency, and power consumption during periods of low load. This particular McQuay unit has an outlet water temperature range of 40-50℃. To avoid excessively low or high temperatures affecting the unit's lifespan, a variable temperature range of 42-48℃ is selected. In Shenyang, the calculated outdoor air conditioning temperature for winter is -20.7℃, and the starting temperature for heating is defined as 5℃. When the outdoor temperature is >5℃, the unit's outlet water temperature is adjusted to 42℃; when -20.7℃ < outdoor temperature ≤5℃, the outlet water temperature changes linearly within the range of 42℃-48℃; and when the outdoor temperature ≤-20.7℃, the unit operates at an outlet water temperature of 48℃.
[0262] Table 10: Comparison of results for different strategies;
[0263]
[0264] Comparative analysis revealed that the energy-saving rate of this embodiment is significantly higher than that of other control strategies.
[0265] Example 2:
[0266] This embodiment proposes an electronic device, including: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the multivariate collaborative optimization method for the energy-saving operation strategy of the heat pump room of the public building.
[0267] The electronic device may be a mobile phone, computer, or tablet computer, etc., and includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the multivariate collaborative optimization method for energy-saving operation strategy of public building heat pump room as described in the embodiment. It is understood that the electronic device may also include input / output (I / O) interfaces and communication components.
[0268] The processor is used to execute all or part of the steps in the multivariate collaborative optimization method for energy-saving operation strategies of public building heat pump room as described in the above embodiments. The memory is used to store various types of data, which may include, for example, instructions for any application or method in an electronic device, as well as application-related data.
[0269] The processor can be implemented as an Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic components, and is used to execute the multivariate collaborative optimization method for energy-saving operation strategy of public building heat pump room described in the above embodiments.
[0270] Example 3:
[0271] This embodiment proposes a computer-readable storage medium that stores executable instructions. When these instructions are executed, if they are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0272] The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the multivariate collaborative optimization method for energy-saving operation strategy of public building heat pump room described in various embodiments of this application.
[0273] The aforementioned storage media include: flash memory, hard disks, multimedia cards, card-type memory (e.g., SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR) memory), random access memory (RAM), static random-access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, disks, optical discs, servers, APP (Application) app stores, and other media capable of storing program verification codes. These media store computer programs, which, when executed by a processor, can implement the various steps of the multi-variable collaborative optimization method for the energy-saving operation strategy of public building heat pump room described above.
[0274] Example 4:
[0275] This embodiment proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the multivariate collaborative optimization method for the energy-saving operation strategy of the heat pump room in public buildings.
[0276] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a computer program product.
[0277] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0278] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of equivalent technology of this disclosure, then the intent of this disclosure also includes such modifications and variations.
Claims
1. A multivariate collaborative optimization method for energy-saving operation strategy of public building heat pump machine room, characterized in that, The method comprises the following steps: Collecting the operation state monitoring data of each device in the heat pump machine room, and calculating the real-time heat load or real-time cold load of the building heat pump system hour by hour; Obtaining factors that have an influence on the real-time heat load or real-time cold load; According to the average value of the real-time heat load or the average value of the real-time cold load, calculating the Pearson correlation coefficient, which represents the degree of correlation between each influencing factor and the real-time heat load or real-time cold load, and taking the influencing factors corresponding to the values of the first M Pearson correlation coefficients as the input factors of strong correlation; Based on the input factors of strong correlation, using the variational mode decomposition algorithm to decompose and denoise the real-time heat load or real-time cold load, and obtaining Z modes and residual data; Inputting the Z modes and residual data into the pre-trained long short-term memory network to obtain Z prediction results, and adding the Z prediction results to obtain the final prediction result; Taking the final prediction result as the prediction result of the dynamic simulation model of the public building heat pump system, and the establishment process of the dynamic simulation model of the public building heat pump system comprises the following steps: according to the established mathematical model of the heat pump unit and the mathematical model of the water pump device energy consumption, using the TRNSYS simulation platform to establish the dynamic simulation model of the public building heat pump system; Defining the public building heat pump system as a four-tuple of a Markov decision model; Using the deterministic policy gradient model of the Actor-Critic architecture to optimize the four-tuple of the Markov decision model to obtain the optimal control strategy of the dynamic simulation model of the public building heat pump system, wherein the Actor-Critic architecture comprises an Actor network and a Critic network, and the Actor network and the Critic network are different neural network structures, wherein the Actor network is used to output the action space of the four-tuple according to the spatial state of the current four-tuple, representing the behavior selection of the current control strategy, and the Critic network is used to evaluate the control strategy selected by the Actor network and provide the selected control strategy to the Actor network. The return; the deterministic policy gradient model of the Actor-Critic architecture uses the gradient descent method to minimize the mean square loss function to update the Critic network parameters, and uses the gradient ascent method to update the Actor network parameters to maximize the Q value of the Critic network.
2. The multi-variable collaborative optimization method for energy saving operation strategy of public building heat pump machine room according to claim 1, characterized in that, The input factors of strong correlation include outdoor temperature, outdoor relative humidity, solar radiation, and personnel in room rate.
3. The multi-variable collaborative optimization method for energy saving operation strategy of public building heat pump machine room according to claim 1, characterized in that, The four-tuple of the Markov decision model comprises: Taking the input factors of strong correlation, the real-time heat load or the input factors of strong correlation, and the real-time cold load as the spatial state of the four-tuple; Taking the control variable of the controller in the public building heat pump system as the action space of the four-tuple; Taking the state transition probability of the spatial state at the current time to the spatial state at the next time as the state transition probability of the four-tuple; Taking the minimization of the operation energy consumption of the public building heat pump system and the maximization of the indoor thermal comfort of the building as the return of the four-tuple; The deterministic policy gradient model of the Actor-Critic architecture, the training process comprises: N training samples are extracted from the experience replay pool at the current iteration, the training samples being quadruples of the Markov decision model; Noise is added to the training samples, and the training samples with added noise are input into the Actor network; Based on the current control policy, the Actor network selects the corresponding action in the spatial state of the next time quadruple in the training samples with added noise; Based on the corresponding action selected at the next time, the Critic network calculates the corresponding return value of the spatial state of the next time quadruple, the corresponding return value being composed of the return of the quadruple at the current iteration, the discount factor, and the target value corresponding to the network parameters of the Critic network; The Critic network parameters are updated by minimizing the mean square loss function using the gradient descent method, and the Actor network parameters are updated using the gradient ascent method; The updated Actor network parameters form an updated Actor network, and the updated Critic network parameters form an updated Critic network, and the updated Actor network and the updated Critic network are used as the deterministic policy gradient model of the Actor-Critic architecture.
4. The multi-variable co-simulation optimization method for energy saving operation strategy of public building heat pump machine room according to claim 3, characterized in that, The corresponding return value is calculated as follows: ; wherein, is a return value corresponding to the i-th iteration, is a return of the quadruple of the i-th iteration, is a discount factor, is a spatial state of the quadruple of the i+1-th iteration, is an action of the quadruple of the i+1-th iteration, is a target network parameter of the Critic network, is a target Q value corresponding to the target network parameter of the Critic network when the spatial state of the quadruple is and the action of the quadruple is .
5. The multi-variable co-simulation optimization method for energy saving operation strategy of public building heat pump machine room according to claim 1, characterized in that, The Critic network parameters are updated by minimizing the mean square loss function using the gradient descent method, and the Actor network parameters are updated using the gradient ascent method. ; ; ; wherein, is the return value corresponding to the i-th iteration, is the spatial state of the quadruple for the i-th iteration, is the action of the quadruple for the i-th iteration, is the mean square loss function, is the learning rate of the Critic network, is the expected value, is the gradient for updating the parameters of the Critic network, is the parameter of the Critic network, is the updated parameter of the Critic network, is the Q value corresponding to the target network parameter of the Critic network when the spatial state of the quadruple is and the action of the quadruple is is the total number of iterations. 6. The multi-variable co-simulation optimization method for energy saving operation strategy of public building heat pump machine room according to claim 1, characterized in that, The Critic network parameters are updated by minimizing the mean square loss function using the gradient descent method, and the Actor network parameters are updated using the gradient ascent method. ; ; ; wherein, is the return value corresponding to the i-th iteration, is the spatial state of the quadruple for the i-th iteration, is the action of the quadruple for the i-th iteration, is the learning rate of the Actor network, is the loss function of the Actor network, is the expected value, is the gradient for updating the Actor network parameters, is the gradient for the Actor network to perform the action, is the Actor network parameter, is the updated Actor network parameter, is the Q value corresponding to the target network parameter of the Critic network when the spatial state of the quadruple is and the action of the quadruple is is the total number of iterations, denotes the deterministic action generated by the Actor network when the spatial state is s, denotes the Q value corresponding to the target network parameter of the Critic network when the spatial state of the quadruple is and the action of the quadruple is is the network parameter of the Critic network, denotes the action output by the Actor network when the spatial state is , the Actor network parameter is , and the network parameter of the Critic network is . 7. The multi-variable co-simulation optimization method for energy saving operation strategy of public building heat pump machine room according to claim 1, characterized in that, The deterministic policy gradient model of the Actor-Critic architecture, the training process further comprises: The target network parameters of the Actor network and the target network parameters of the Critic network are updated using a soft update method, and the calculation formula is as follows: ; ; wherein, is a soft update coefficient, is a Critic network parameter, is a target network parameter of the Critic network, is an Actor network parameter, is a target network parameter of the Actor network.
8. The method of claim 1, wherein, The deterministic policy gradient model of the Actor-Critic architecture is used to optimize the quadruples of the Markov decision model, and the optimal control policy of the dynamic simulation model of the public building heat pump system is obtained, comprising: the spatial state s of the quad at time t t in the deterministic policy gradient model of the input actor-critic architecture; According to the spatial state of the quadruplet at the t-th moment, the updated Actor network selects the deterministic action a at the t-th moment t ; Deterministic action a t And the Critic network is updated by exploration noise input, and the reward value r at the t time is obtained t ; Determining a deterministic action a t Inputting a dynamic simulation model of the public building heat pump system, obtaining a spatial state s of the fourth tuple at the t+1 time t+1 ; The quadruple (s t , a t , r t , s t+1 ) is stored in the experience replay pool, and the deterministic action a t at the t-th moment is taken as the optimal control strategy of the dynamic simulation model of the public building heat pump system.
9. An electronic device, comprising: Comprising: One or more processors, and a memory for storing instructions, when the instructions are executed by the one or more processors, the one or more processors execute the multivariate collaborative optimization method of the public building heat pump room energy-saving operation strategy according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, It stores executable instructions, which when executed make the processor execute the multivariate collaborative optimization method of the public building heat pump room energy-saving operation strategy according to any one of claims 1-8.