Control strategy planning method and system for heating ventilation air conditioner

By constructing an air conditioning strategy optimization network model and combining it with air conditioning thermal and electricity cost prediction models, the problem of ignoring electricity costs in HVAC control methods is solved, enabling intelligent control during peak electricity consumption periods and reducing air conditioning power to save electricity costs.

CN120868577APending Publication Date: 2025-10-31GUANGZHOU CITY UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510931748.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing HVAC control methods only take indoor temperature as the optimization target and ignore electricity costs, which makes it impossible to adaptively adjust according to the peak and off-peak electricity price characteristics of the power grid, resulting in continuous high-power operation during peak electricity consumption periods.

Method used

An air conditioning strategy optimization network model is constructed using the near-end strategy optimization PPO algorithm. Combined with the air conditioning thermal model and the electricity cost prediction model, intelligent control of air conditioning is achieved through multi-objective collaborative optimization, dynamically adjusting the air conditioning power to balance thermal comfort and electricity economy.

Benefits of technology

While ensuring a stable indoor temperature, the power of the air conditioner is dynamically adjusted based on the electricity price signal, especially during peak electricity consumption periods, to proactively reduce the power of the air conditioner and achieve the effect of saving electricity costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120868577A_ABST
    Figure CN120868577A_ABST
Patent Text Reader

Abstract

The invention discloses a regulation and control strategy planning method and system for a heating ventilation air conditioner. The method comprises the following steps that a near-end strategy optimization PPO algorithm is adopted to construct an air conditioner strategy optimization network model; constructing an air conditioner thermal model and an air conditioner electricity charge prediction model; the obtained environment parameters, time parameters and air conditioner parameters are input into an air conditioner thermal model for indoor temperature prediction; based on the indoor temperature prediction value at the t + 1 moment and an air conditioner electricity charge prediction model, an air conditioner strategy optimization network model is trained, so that an air conditioner regulation and control strategy output by the air conditioner strategy optimization network model is converged; and collecting real-time environment parameters, and inputting the real-time environment parameters into the trained air conditioner strategy optimization network model for air conditioner regulation and control strategy planning. The method solves the problems that an existing heating ventilation air conditioner regulation and control method only takes indoor temperature as an optimization target and lacks quantitative management of power consumption cost, so that a system cannot adaptively adjust a strategy according to peak and valley electricity prices of a power grid, and conditions such as continuous high-power operation occur in the peak period of power consumption occur.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of HVAC control strategy planning technology, specifically a method and system for HVAC control strategy planning. Background Technology

[0002] With the development of society and the economy and the improvement of people's living standards, the demand for indoor environmental comfort is becoming increasingly prominent. In residential, commercial complex, and other building environments, heating, ventilation, and air conditioning (HVAC) has become a core facility for ensuring indoor environmental quality. Existing technology discloses an HVAC control method, which specifically inputs real-time acquired environmental and time parameters into a preset HVAC setpoint temperature prediction model, outputs predicted HVAC setpoint parameters using the model, and executes a control strategy based on the acquired predicted HVAC setpoint parameters using a local controller. While this method achieves dynamic control of indoor temperature and meets the thermal comfort needs of people to a certain extent, its control logic only uses indoor temperature as the optimization target, neglecting the quantitative management of electricity costs. This results in the HVAC system being unable to adaptively adjust its control strategy according to the peak and off-peak electricity pricing characteristics of the power grid, such as the price increase during morning and evening peak hours, leading to situations such as continuous high-power operation during peak electricity consumption periods. Summary of the Invention

[0003] To address the aforementioned shortcomings, this invention proposes a control strategy planning method and system for HVAC systems. The aim is to solve the problem that existing HVAC control methods only take indoor temperature as the optimization target and lack quantitative management of electricity costs, resulting in the system's inability to adaptively adjust its strategy according to the peak and off-peak electricity prices, leading to situations such as continuous high-power operation during peak electricity consumption periods.

[0004] To achieve this objective, the present invention adopts the following technical solution:

[0005] A method for planning a control strategy for heating, ventilation, and air conditioning systems includes the following steps:

[0006] Step S1: Construct an air conditioning strategy optimization network model using the near-end strategy optimization (PPO) algorithm. The air conditioning strategy optimization network model is used to output the corresponding air conditioning control strategy under the current environmental conditions and evaluate its value to measure long-term benefits.

[0007] Step S2: Construct an air conditioning thermal model and an air conditioning electricity cost prediction model. The air conditioning thermal model is used to simulate the trend of indoor temperature change over time; the air conditioning electricity cost prediction model is used to provide economic reward signals to the air conditioning strategy optimization network model.

[0008] Step S3: Obtain environmental parameters, time parameters, and air conditioning parameters, and input all environmental parameters, time parameters, and air conditioning parameters into the air conditioning thermal model to predict the indoor temperature, and output the predicted indoor temperature value at time t+1. Among them, the environmental parameters include the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, thermal effect of indoor heat source, and outdoor temperature, and the air conditioning parameters include the air conditioning power at time t.

[0009] Step S4: Based on the indoor temperature prediction value at time t+1 and the air conditioning electricity cost prediction model, train the air conditioning strategy optimization network model so that the air conditioning control strategy output by the air conditioning strategy optimization network model converges, and obtain the trained air conditioning strategy optimization network model.

[0010] Step S5: Collect real-time environmental parameters and input them into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy and output the optimal air conditioning control strategy.

[0011] Preferably, step S1 includes the following sub-step: calculating the loss function of the PPO algorithm, wherein the loss function of the PPO algorithm is used for policy optimization, and the specific calculation formula is as follows:

[0012]

[0013] Among them, L PPO (θ) represents the loss function used by the PPO algorithm to optimize the strategy; π θ (α t |s t ) indicates that the new strategy is in state s t Take action α t The probability of; This indicates that the old policy is in state s. t Take action α t The probability of ; ∈ represents the hyperparameter of the adjustment policy update range; E t (x) represents the expected value function; clip(x) represents the clipping function; min(x) represents the minimum value function; The value represents the estimated advantage function, i.e., the generalized advantage function value, and its specific mathematical expression is as follows:

[0014]

[0015] Where, δ t λ represents the time differential error; γ represents the discount factor; λ represents the hyperparameter of generalized advantage; L represents a natural number.

[0016] Preferably, in step S2, the mathematical expression of the air conditioning thermal model is as follows:

[0017]

[0018] in, θ represents the predicted indoor temperature at time t+1. t The value represents the indoor temperature at time t; Δt represents the time interval; C represents the indoor specific heat capacity; and R represents the indoor thermal resistance. η represents the outdoor temperature at time t; η represents the thermal effect of the indoor heat source; q t This represents the air conditioning power at time t.

[0019] Preferably, in step S2, the mathematical expression of the air conditioning electricity cost prediction model is as follows:

[0020] Price = Z1*q1 + Z2*q2 + Z t *q t +...(t=1,2...24);

[0021] Among them, Z t q represents the electricity cost at time t; t The value represents the air conditioner power at time t; Price represents the total electricity cost of the air conditioner.

[0022] Preferably, step S4 specifically includes the following sub-steps:

[0023] Step S41: Input the environmental parameters, time parameters, and air conditioning parameters into the air conditioning strategy optimization network model for processing, and output the corresponding control strategy under the current environmental state;

[0024] Step S42: Based on the predicted indoor temperature at time t+1 and the predicted air conditioning electricity cost model, quantitatively evaluate the electricity cost and thermal comfort of the corresponding control strategy under the current environmental conditions, so as to generate reward and penalty feedback signals.

[0025] Step S43: Based on the reward and penalty feedback signal, determine whether the corresponding control strategy under the current environmental state has converged. If yes, it means that the performance of the air conditioning strategy optimization network model has reached a stable optimal state; if no, it means that the performance of the air conditioning strategy optimization network model is not good, and continue to execute steps S41-S43 until the output control strategy converges.

[0026] Another aspect of this application provides a control strategy planning system for heating, ventilation, and air conditioning (HVAC), the system comprising:

[0027] The first construction module is used to construct an air conditioning strategy optimization network model using the near-end strategy optimization (PPO) algorithm. The air conditioning strategy optimization network model is used to output the air conditioning control strategy corresponding to the current environmental state and evaluate its value to measure long-term benefits.

[0028] The second building module is used to build an air conditioning thermal model, which is used to simulate the trend of indoor temperature change over time.

[0029] The third construction module is used to build an air conditioning electricity cost prediction model, which is used to provide economic reward signals to the air conditioning strategy optimization network model.

[0030] The acquisition module is used to acquire environmental parameters, time parameters, and air conditioning parameters. The environmental parameters include the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, thermal effect of indoor heat source, and outdoor temperature. The air conditioning parameters include the air conditioning power at time t.

[0031] The indoor temperature prediction module is used to input environmental parameters, time parameters, and air conditioning parameters into the air conditioning thermal model to predict the indoor temperature and output the predicted indoor temperature value at time t+1.

[0032] The model training module is used to train the air conditioning strategy optimization network model based on the indoor temperature prediction value and air conditioning electricity cost prediction model at time t+1, so that the air conditioning control strategy output by the air conditioning strategy optimization network model converges, and thus obtain the trained air conditioning strategy optimization network model.

[0033] The data acquisition module is used to collect real-time environmental parameters;

[0034] The air conditioning control strategy planning module is used to input real-time environmental parameters into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy and output the optimal air conditioning control strategy.

[0035] Preferably, the first building module includes:

[0036] The calculation submodule is used to calculate the loss function of the PPO algorithm, which is used for policy optimization. The specific calculation formula is as follows:

[0037]

[0038] Among them, L PPO (θ) represents the loss function used by the PPO algorithm to optimize the strategy; π θ (α t |s t ) indicates that the new strategy is in state s t Take action t The probability of; This indicates that the old policy is in state s. t Take action t The probability of ; ∈ represents the hyperparameter of the adjustment policy update range; E t (x) represents the expected value function; clip(x) represents the clipping function; min(x) represents the minimum value function; The value represents the estimated advantage function, i.e., the generalized advantage function value, and its specific mathematical expression is as follows:

[0039]

[0040] Where, δ t λ represents the time differential error; γ represents the discount factor; λ represents the hyperparameter of generalized advantage; L represents a natural number.

[0041] Preferably, in the second building module, the mathematical expression of the air conditioning thermal model is as follows:

[0042]

[0043] in, θ represents the predicted indoor temperature at time t+1. t The value represents the indoor temperature at time t; Δt represents the time interval; C represents the indoor specific heat capacity; and R represents the indoor thermal resistance. η represents the outdoor temperature at time t; η represents the thermal effect of the indoor heat source; q t This represents the air conditioning power at time t.

[0044] Preferably, in the third building module, the mathematical expression of the air conditioning electricity cost prediction model is as follows:

[0045] Price = Z1*q1 + Z2*q2 + Z t *q t +...(t=1,2...24);

[0046] Among them, Z t q represents the electricity cost at time t; t The value represents the air conditioner power at time t; Price represents the total electricity cost of the air conditioner.

[0047] Preferably, the model training module includes:

[0048] The model processing submodule is used to input environmental parameters, time parameters, and air conditioning parameters into the air conditioning strategy optimization network model for processing; and output the corresponding control strategy under the current environmental state.

[0049] The quantitative evaluation submodule is used to quantitatively evaluate the electricity cost and thermal comfort of the corresponding control strategy under the current environmental conditions based on the indoor temperature prediction value at time t+1 and the air conditioning electricity cost prediction model, so as to generate reward and penalty feedback signals.

[0050] The judgment submodule is used to determine whether the corresponding control strategy under the current environmental state has converged based on the reward and punishment feedback signal. If it has, it means that the performance of the air conditioning strategy optimization network model has reached a stable optimal state; if not, it means that the performance of the air conditioning strategy optimization network model is not good, and the model processing submodule, the quantitative evaluation submodule, and the judgment submodule continue to be executed until the output control strategy converges.

[0051] The technical solutions provided in this application embodiment may include the following beneficial effects:

[0052] This solution introduces the PPO algorithm to construct an air conditioning strategy optimization network model, and combines it with an air conditioning thermal model and an air conditioning electricity cost prediction model to achieve intelligent air conditioning control with multi-objective collaborative optimization. Compared with traditional HVAC control methods that only consider indoor temperature as the optimization objective, the air conditioning strategy optimization network model based on the PPO algorithm proposed in this solution simultaneously takes into account thermal comfort and electricity economy. Under the premise of ensuring stable indoor temperature, it dynamically adjusts the air conditioning power based on electricity price signals, especially during peak electricity consumption periods, it will actively reduce the air conditioning power to save electricity costs. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating the steps involved in planning a control strategy for HVAC systems. Detailed Implementation

[0054] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0055] A method for planning a control strategy for heating, ventilation, and air conditioning systems includes the following steps:

[0056] Step S1: Construct an air conditioning strategy optimization network model using the near-end strategy optimization (PPO) algorithm. The air conditioning strategy optimization network model is used to output the corresponding air conditioning control strategy under the current environmental conditions and evaluate its value to measure long-term benefits.

[0057] Step S2: Construct an air conditioning thermal model and an air conditioning electricity cost prediction model. The air conditioning thermal model is used to simulate the trend of indoor temperature change over time; the air conditioning electricity cost prediction model is used to provide economic reward signals to the air conditioning strategy optimization network model.

[0058] Step S3: Obtain environmental parameters, time parameters, and air conditioning parameters, and input all environmental parameters, time parameters, and air conditioning parameters into the air conditioning thermal model to predict the indoor temperature, and output the predicted indoor temperature value at time t+1. Among them, the environmental parameters include the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, thermal effect of indoor heat source, and outdoor temperature, and the air conditioning parameters include the air conditioning power at time t.

[0059] Step S4: Based on the indoor temperature prediction value at time t+1 and the air conditioning electricity cost prediction model, train the air conditioning strategy optimization network model so that the air conditioning control strategy output by the air conditioning strategy optimization network model converges, and obtain the trained air conditioning strategy optimization network model.

[0060] Step S5: Collect real-time environmental parameters and input them into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy and output the optimal air conditioning control strategy.

[0061] This solution proposes a control strategy planning method for HVAC systems, such as... Figure 1As shown, the first step is to construct an air conditioning strategy optimization network model using the Proximal Policy Optimization (PPO) algorithm. This model outputs the corresponding air conditioning control strategy under the current environmental conditions and evaluates its value to measure long-term benefits. In this embodiment, the PPO algorithm is a reinforcement learning algorithm suitable for solving optimization problems in continuous action spaces. By using the PPO algorithm to construct the air conditioning strategy optimization network model, the optimal air conditioning control strategy can be learned based on the environmental conditions, and long-term benefits can be measured through value evaluation. Further, the air conditioning strategy optimization network model includes a policy network and a value network. The second step is to construct an air conditioning thermal model and an air conditioning electricity cost prediction model. The thermal model simulates the trend of indoor temperature change over time; the electricity cost prediction model provides economic reward signals to the air conditioning strategy optimization network model. In this embodiment, the thermal model is a mathematical model describing the thermal behavior of the air conditioning system. By constructing the thermal model, the dynamic trend of indoor temperature change over time can be simulated, providing key input parameters for the air conditioning strategy optimization network model. The electricity cost prediction model is a mathematical model specifically designed for predicting air conditioning electricity costs. By constructing an air conditioning electricity cost prediction model, an economic reward signal is provided for the air conditioning strategy optimization network model, which helps to consider electricity costs during air conditioning control and achieve energy-saving goals. The third step is to obtain environmental parameters, time parameters, and air conditioning parameters, and input them into the air conditioning thermal model to predict indoor temperature, outputting the predicted indoor temperature value at time t+1. The environmental parameters include the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, thermal effect of indoor heat sources, and outdoor temperature. The air conditioning parameters include the air conditioning power at time t. In this embodiment, by predicting the indoor temperature at time t+1, a physical foundation environment is provided for the subsequent training of the air conditioning strategy optimization network model. The fourth step is to train the air conditioning strategy optimization network model based on the indoor temperature prediction value at time t+1 and the air conditioning electricity cost prediction model, so that the air conditioning control strategy output by the air conditioning strategy optimization network model converges, thus obtaining the trained air conditioning strategy optimization network model. In this embodiment, by training the air conditioning strategy optimization network model based on the indoor temperature prediction value at time t+1 and the electricity cost prediction model, the air conditioning strategy optimization network model can learn how to adjust the air conditioning control strategy according to the changes in indoor temperature and electricity costs, so as to achieve a convergence state. That is, the air conditioning control strategy can improve the user's thermal comfort and reduce electricity costs, thereby obtaining the optimal control strategy model.The fifth step is to collect real-time environmental parameters and input them into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy and output the optimal air conditioning control strategy. In this embodiment, by inputting the real-time collected environmental parameters into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy, the HVAC system can quickly and accurately generate control strategies based on the current real-time environmental conditions and flexibly respond to actual environmental changes.

[0062] This solution introduces the PPO algorithm to construct an air conditioning strategy optimization network model, and combines it with an air conditioning thermal model and an air conditioning electricity cost prediction model to achieve intelligent air conditioning control with multi-objective collaborative optimization. Compared with traditional HVAC control methods that only consider indoor temperature as the optimization objective, the air conditioning strategy optimization network model based on the PPO algorithm proposed in this solution simultaneously takes into account thermal comfort and electricity economy. Under the premise of ensuring stable indoor temperature, it dynamically adjusts the air conditioning power based on electricity price signals, especially during peak electricity consumption periods, it will actively reduce the air conditioning power to save electricity costs.

[0063] Preferably, step S1 includes the following sub-step: calculating the loss function of the PPO algorithm, wherein the loss function of the PPO algorithm is used for policy optimization, and the specific calculation formula is as follows:

[0064]

[0065] Among them, L PPO (θ) represents the loss function used by the PPO algorithm to optimize the strategy; π θ (α t |s t ) indicates that the new strategy is in state s t Take action α t The probability of; This indicates that the old policy is in state s. t Take action α t The probability of ; ∈ represents the hyperparameter of the adjustment policy update range; E t (x) represents the expected value function; clip(x) represents the clipping function; min(x) represents the minimum value function; The value represents the estimated advantage function, i.e., the generalized advantage function value, and its specific mathematical expression is as follows:

[0066]

[0067] Where, δ t λ represents the time differential error; γ represents the discount factor; λ represents the hyperparameter of generalized advantage; L represents a natural number.

[0068] In this embodiment, the probability ratio between the new and old strategies is limited by the clip function. The update range (1-∈, 1+∈) avoids the policy deviating excessively from historical experience due to a single gradient update, making policy iteration smoother. In the scenario of air conditioning control, this can prevent the air conditioning system from going out of control due to sudden policy changes, thus improving application security. Further explanation: the inherent logic of PPO policy optimization is to optimize the generalized advantage function value... Maximize, when When the expected value of the current action is higher, the policy network in the air conditioning policy optimization network model will encourage the execution of such actions through a probability boosting mechanism; conversely, if the expected value is lower, the policy network will encourage the execution of such actions. In such cases, the frequency of such actions can be suppressed through policy updates.

[0069] Preferably, in step S2, the mathematical expression of the air conditioning thermal model is as follows:

[0070]

[0071] in, θ represents the predicted indoor temperature at time t+1. t The value represents the indoor temperature at time t; Δt represents the time interval; C represents the indoor specific heat capacity; and R represents the indoor thermal resistance. η represents the outdoor temperature at time t; η represents the thermal effect of the indoor heat source; q t This represents the air conditioning power at time t.

[0072] In this embodiment, the air conditioning thermal model comprehensively considers the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, the thermal effect of the indoor heat source, outdoor temperature, and air conditioning power at time t. By quantifying the interaction mechanism of various thermodynamic elements, the air conditioning thermal model provides reliable temperature change predictions for optimizing air conditioning control strategies.

[0073] Preferably, in step S2, the mathematical expression of the air conditioning electricity cost prediction model is as follows:

[0074] Price = Z1*q1 + Z2*q2 + Z t *q t +...(t=1,2...24);

[0075] Among them, Z t q represents the electricity cost at time t; t The value represents the air conditioner power at time t; Price represents the total electricity cost of the air conditioner.

[0076] In this embodiment, by establishing a time-segmented air conditioning electricity cost prediction model, the electricity cost of the HVAC system at different times can be accurately quantified, providing data support for the control of electricity costs.

[0077] Preferably, step S4 specifically includes the following sub-steps:

[0078] Step S41: Input the environmental parameters, time parameters, and air conditioning parameters into the air conditioning strategy optimization network model for processing, and output the corresponding control strategy under the current environmental state;

[0079] Step S42: Based on the predicted indoor temperature at time t+1 and the predicted air conditioning electricity cost model, quantitatively evaluate the electricity cost and thermal comfort of the corresponding control strategy under the current environmental conditions, so as to generate reward and penalty feedback signals.

[0080] Step S43: Based on the reward and penalty feedback signal, determine whether the corresponding control strategy under the current environmental state has converged. If yes, it means that the performance of the air conditioning strategy optimization network model has reached a stable optimal state; if no, it means that the performance of the air conditioning strategy optimization network model is not good, and continue to execute steps S41-S43 until the output control strategy converges.

[0081] Specifically, in step S41, environmental parameters, time parameters, and air conditioning parameters are input into the air conditioning strategy optimization network model for processing, ensuring that the output control strategy can simultaneously consider real-time environmental characteristics and temporal patterns. In step S42, by evaluating the power cost and thermal comfort of the control strategy, energy waste caused by overemphasizing comfort can be avoided. In one embodiment, after the HVAC system executes a control strategy, based on the predicted indoor temperature at time t+1 and the air conditioning electricity cost prediction model, it is found that the indoor temperature drops to 25°C, but the power cost increases by 1 yuan. It can be determined that the thermal comfort of the control strategy is normal but the power consumption is high, and the iteration direction of the control strategy will be suppressed through a penalty signal. In step S43, the performance of the air conditioning strategy optimization network model is evaluated by judging the convergence of the control strategy, which helps to clarify whether the air conditioning strategy optimization network model is moving closer to the optimal solution during the training process.

[0082] Another aspect of this application provides a control strategy planning system for heating, ventilation, and air conditioning (HVAC), the system comprising:

[0083] The first construction module is used to construct an air conditioning strategy optimization network model using the near-end strategy optimization (PPO) algorithm. The air conditioning strategy optimization network model is used to output the air conditioning control strategy corresponding to the current environmental state and evaluate its value to measure long-term benefits.

[0084] The second building module is used to build an air conditioning thermal model, which is used to simulate the trend of indoor temperature change over time.

[0085] The third construction module is used to build an air conditioning electricity cost prediction model, which is used to provide economic reward signals to the air conditioning strategy optimization network model.

[0086] The acquisition module is used to acquire environmental parameters, time parameters, and air conditioning parameters. The environmental parameters include the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, thermal effect of indoor heat source, and outdoor temperature. The air conditioning parameters include the air conditioning power at time t.

[0087] The indoor temperature prediction module is used to input environmental parameters, time parameters, and air conditioning parameters into the air conditioning thermal model to predict the indoor temperature and output the predicted indoor temperature value at time t+1.

[0088] The model training module is used to train the air conditioning strategy optimization network model based on the indoor temperature prediction value and air conditioning electricity cost prediction model at time t+1, so that the air conditioning control strategy output by the air conditioning strategy optimization network model converges, and thus obtain the trained air conditioning strategy optimization network model.

[0089] The data acquisition module is used to collect real-time environmental parameters;

[0090] The air conditioning control strategy planning module is used to input real-time environmental parameters into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy and output the optimal air conditioning control strategy.

[0091] This solution presents a HVAC control strategy planning system. Through the coordinated efforts of a first construction module, a second construction module, a third construction module, an acquisition module, an indoor temperature prediction module, a model training module, a data acquisition module, and an HVAC control strategy planning module, it achieves intelligent HVAC control with multi-objective collaborative optimization. Compared to traditional HVAC control methods that only consider indoor temperature as the optimization objective, the HVAC strategy optimization network model based on the PPO algorithm proposed in this solution, after training, simultaneously considers thermal comfort and energy economy. While ensuring stable indoor temperature, it dynamically adjusts the air conditioning power based on electricity price signals, especially actively reducing air conditioning power during peak electricity consumption periods to save on electricity costs.

[0092] Preferably, the first building module includes:

[0093] The calculation submodule is used to calculate the loss function of the PPO algorithm, which is used for policy optimization. The specific calculation formula is as follows:

[0094]

[0095] Among them, L PPO (θ) represents the loss function used by the PPO algorithm to optimize the strategy; π θ (α t |s t ) indicates that the new strategy is in state s t Take action α t The probability of; This indicates that the old policy is in state s. t Take action α t The probability of ; ∈ represents the hyperparameter of the adjustment policy update range; E t (x) represents the expected value function; clip(x) represents the clipping function; min(x) represents the minimum value function; The value represents the estimated advantage function, i.e., the generalized advantage function value, and its specific mathematical expression is as follows:

[0096]

[0097] Where, δ t λ represents the time differential error; γ represents the discount factor; λ represents the hyperparameter of generalized advantage; L represents a natural number.

[0098] In this embodiment, the probability ratio between the old and new strategies is limited by the clip function in the calculation submodule. The update range (1-∈, 1+∈) avoids the policy from deviating too much from historical experience due to a single gradient update, making the policy iteration smoother. In the scenario of air conditioning control, it can prevent the air conditioning system from going out of control due to policy mutation and improve the application security.

[0099] Preferably, in the second building module, the mathematical expression of the air conditioning thermal model is as follows:

[0100]

[0101] in, θ represents the predicted indoor temperature at time t+1. t The value represents the indoor temperature at time t; Δt represents the time interval; C represents the indoor specific heat capacity; and R represents the indoor thermal resistance. η represents the outdoor temperature at time t; η represents the thermal effect of the indoor heat source; q t This represents the air conditioning power at time t.

[0102] In this embodiment, the air conditioning thermal model comprehensively considers the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, the thermal effect of the indoor heat source, outdoor temperature, and air conditioning power at time t. By quantifying the interaction mechanism of various thermodynamic elements, the air conditioning thermal model provides reliable temperature change predictions for optimizing air conditioning control strategies.

[0103] Preferably, in the third construction module, the mathematical expression of the air conditioning electricity cost prediction model is as follows:

[0104] Price = Z1*q1 + Z2*q2 + Z t *q t +...(t=1,2...24);

[0105] Among them, Z t q represents the electricity cost at time t; t The value represents the air conditioner power at time t; Price represents the total electricity cost of the air conditioner.

[0106] In this embodiment, by establishing a time-segmented air conditioning electricity cost prediction model, the electricity cost of the HVAC system at different times can be accurately quantified, providing data support for the control of electricity costs.

[0107] Preferably, the model training module includes:

[0108] The model processing submodule is used to input environmental parameters, time parameters, and air conditioning parameters into the air conditioning strategy optimization network model for processing; and output the corresponding control strategy under the current environmental state.

[0109] The quantitative evaluation submodule is used to quantitatively evaluate the electricity cost and thermal comfort of the corresponding control strategy under the current environmental conditions based on the indoor temperature prediction value at time t+1 and the air conditioning electricity cost prediction model, so as to generate reward and penalty feedback signals.

[0110] The judgment submodule is used to determine whether the corresponding control strategy under the current environmental state has converged based on the reward and punishment feedback signal. If it has, it means that the performance of the air conditioning strategy optimization network model has reached a stable optimal state; if not, it means that the performance of the air conditioning strategy optimization network model is not good, and the model processing submodule, the quantitative evaluation submodule, and the judgment submodule continue to be executed until the output control strategy converges.

[0111] In this embodiment, by executing the model processing submodule, it is ensured that the output control strategy can simultaneously take into account real-time environmental characteristics and temporal patterns. By executing the quantitative evaluation submodule, energy waste caused by overemphasizing comfort can be avoided. By executing the judgment submodule, it is helpful to determine whether the air conditioning strategy optimization network model is moving closer to the optimal solution during training.

[0112] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0113] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for planning control strategies for heating, ventilation, and air conditioning systems, characterized in that: Includes the following steps: Step S1: Construct an air conditioning strategy optimization network model using the near-end strategy optimization (PPO) algorithm. The air conditioning strategy optimization network model is used to output the corresponding air conditioning control strategy under the current environmental conditions and evaluate its value to measure long-term benefits. Step S2: Construct an air conditioning thermal model and an air conditioning electricity cost prediction model. The air conditioning thermal model is used to simulate the trend of indoor temperature change over time; the air conditioning electricity cost prediction model is used to provide economic reward signals to the air conditioning strategy optimization network model. Step S3: Obtain environmental parameters, time parameters, and air conditioning parameters, and input all environmental parameters, time parameters, and air conditioning parameters into the air conditioning thermal model to predict the indoor temperature, and output the predicted indoor temperature value at time t+1. Among them, the environmental parameters include the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, thermal effect of indoor heat source, and outdoor temperature, and the air conditioning parameters include the air conditioning power at time t. Step S4: Based on the indoor temperature prediction value at time t+1 and the air conditioning electricity cost prediction model, train the air conditioning strategy optimization network model so that the air conditioning control strategy output by the air conditioning strategy optimization network model converges, and obtain the trained air conditioning strategy optimization network model. Step S5: Collect real-time environmental parameters and input them into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy and output the optimal air conditioning control strategy.

2. The method for planning a control strategy for HVAC according to claim 1, characterized in that: Step S1 includes the following sub-steps: calculating the loss function of the PPO algorithm, where the loss function of the PPO algorithm is used for policy optimization, and the specific calculation formula is as follows: Among them, L PPO (θ) represents the loss function used by the PPO algorithm to optimize the strategy; π θ (ɑ t |s t ) indicates that the new strategy is in state s t Take action α t The probability of; This indicates that the old policy is in state s. t Take action t The probability of ; ∈ represents the hyperparameter of the adjustment policy update range; E t (x) represents the expected value function; clip(x) represents the clipping function; min(x) represents the minimum value function; The value represents the estimated advantage function, i.e., the generalized advantage function value, and its specific mathematical expression is as follows: Where, δ t λ represents the time differential error; γ represents the discount factor; λ represents the hyperparameter of generalized advantage; L represents a natural number.

3. The method for planning a control strategy for HVAC according to claim 1, characterized in that: In step S2, the mathematical expression of the air conditioning thermal model is as follows: in, θ represents the predicted indoor temperature at time t+1. t The value represents the indoor temperature at time t; Δt represents the time interval; C represents the indoor specific heat capacity; and R represents the indoor thermal resistance. η represents the outdoor temperature at time t; η represents the thermal effect of the indoor heat source; q t This represents the air conditioning power at time t.

4. The method for planning a control strategy for HVAC according to claim 1, characterized in that: In step S2, the mathematical expression for the air conditioning electricity cost prediction model is as follows: Price=Z1*q1+Z2*q2+Z t *q t +...(t=1,2...24); Among them, Z t q represents the electricity cost at time t; t The value represents the air conditioner power at time t; Price represents the total electricity cost of the air conditioner.

5. The method for planning a control strategy for HVAC according to claim 1, characterized in that: Step S4 specifically includes the following sub-steps: Step S41: Input the environmental parameters, time parameters, and air conditioning parameters into the air conditioning strategy optimization network model for processing, and output the corresponding control strategy under the current environmental state; Step S42: Based on the predicted indoor temperature at time t+1 and the predicted air conditioning electricity cost model, quantitatively evaluate the electricity cost and thermal comfort of the corresponding control strategy under the current environmental conditions, so as to generate reward and penalty feedback signals. Step S43: Based on the reward and penalty feedback signal, determine whether the corresponding control strategy under the current environmental state has converged. If yes, it means that the performance of the air conditioning strategy optimization network model has reached a stable optimal state; if no, it means that the performance of the air conditioning strategy optimization network model is not good, and continue to execute steps S41-S43 until the output control strategy converges.

6. A control strategy planning system for HVAC, using the control strategy planning method for HVAC as described in any one of claims 1-5, characterized in that: The system includes: The first construction module is used to construct an air conditioning strategy optimization network model using the near-end strategy optimization (PPO) algorithm. The air conditioning strategy optimization network model is used to output the air conditioning control strategy corresponding to the current environmental state and evaluate its value to measure long-term benefits. The second building module is used to build an air conditioning thermal model, which is used to simulate the trend of indoor temperature change over time. The third construction module is used to build an air conditioning electricity cost prediction model, which is used to provide economic reward signals to the air conditioning strategy optimization network model. The acquisition module is used to acquire environmental parameters, time parameters, and air conditioning parameters. The environmental parameters include the indoor temperature at time t, indoor specific heat capacity, indoor thermal resistance, thermal effect of indoor heat source, and outdoor temperature. The air conditioning parameters include the air conditioning power at time t. The indoor temperature prediction module is used to input environmental parameters, time parameters, and air conditioning parameters into the air conditioning thermal model to predict the indoor temperature and output the predicted indoor temperature value at time t+1. The model training module is used to train the air conditioning strategy optimization network model based on the indoor temperature prediction value and air conditioning electricity cost prediction model at time t+1, so that the air conditioning control strategy output by the air conditioning strategy optimization network model converges, and thus obtain the trained air conditioning strategy optimization network model. The data acquisition module is used to collect real-time environmental parameters; The air conditioning control strategy planning module is used to input real-time environmental parameters into the trained air conditioning strategy optimization network model to plan the air conditioning control strategy and output the optimal air conditioning control strategy.

7. The control strategy planning system for HVAC according to claim 6, characterized in that: The first building module includes: The calculation submodule is used to calculate the loss function of the PPO algorithm, which is used for policy optimization. The specific calculation formula is as follows: Among them, L PPO (θ) represents the loss function used by the PPO algorithm to optimize the strategy; π θ (α t |s t ) indicates that the new strategy is in state s t Take action α t The probability of; This indicates that the old policy is in state s. t Take action α t The probability of ; ∈ represents the hyperparameter of the adjustment policy update range; E t (x) represents the expected value function; clip(x) represents the clipping function; min(x) represents the minimum value function; The value represents the estimated advantage function, i.e., the generalized advantage function value, and its specific mathematical expression is as follows: Where, δ t λ represents the time differential error; γ represents the discount factor; λ represents the hyperparameter of generalized advantage; L represents a natural number.

8. The control strategy planning system for HVAC according to claim 6, characterized in that: In the second building module, the mathematical expression of the air conditioning thermal model is as follows: in, θ represents the predicted indoor temperature at time t+1. t The value represents the indoor temperature at time t; Δt represents the time interval; C represents the indoor specific heat capacity; and R represents the indoor thermal resistance. η represents the outdoor temperature at time t; η represents the thermal effect of the indoor heat source; q t This represents the air conditioning power at time t.

9. The control strategy planning system for HVAC according to claim 6, characterized in that: In the third building module, the mathematical expression for the air conditioning electricity cost prediction model is as follows: Price=Z1*q1+Z2*q2+Z t *q t +...(t=1,2...24); Among them, Z t q represents the electricity cost at time t; t The value represents the air conditioner power at time t; Price represents the total electricity cost of the air conditioner.

10. A control strategy planning system for HVAC according to claim 6, characterized in that: The model training module includes: The model processing submodule is used to input environmental parameters, time parameters, and air conditioning parameters into the air conditioning strategy optimization network model for processing; and output the corresponding control strategy under the current environmental state. The quantitative evaluation submodule is used to quantitatively evaluate the electricity cost and thermal comfort of the corresponding control strategy under the current environmental conditions based on the indoor temperature prediction value at time t+1 and the air conditioning electricity cost prediction model, so as to generate reward and penalty feedback signals. The judgment submodule is used to determine whether the corresponding control strategy under the current environmental state has converged based on the reward and punishment feedback signal. If it has, it means that the performance of the air conditioning strategy optimization network model has reached a stable optimal state; if not, it means that the performance of the air conditioning strategy optimization network model is not good, and the model processing submodule, the quantitative evaluation submodule, and the judgment submodule continue to be executed until the output control strategy converges.