Air conditioner system water pump real-time variable frequency control method and system based on adaptive reinforcement learning

CN121140160BActive Publication Date: 2026-09-18SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511514275.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-09-18
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

[0006]本发明的目的在于克服现有技术的不足,提供一种适用于空调系统水泵的变频控制方法及装置,能够克服传统控制方法适应性差的问题,并解决现有控制策略在面对空调系统非线性、动态性强以及外部扰动频繁等工况时,难以实现持续自适应优化、节能效果受限的技术瓶颈,实现水泵的智能化控制

Benefits of technology

本发明提出的自适应强化学习空调水泵变频控制方法及装置,有效解决了传统控制策略难以适应空调系统非线性与动态变化的难题。通过构建数据驱动的孪生模型并在虚拟环境中预训练强化学习策略,避免了传统模型预测控制中依赖精确物理建模、建模复杂度高、维护成本大的问题,显著降低了对专业领域知识的依赖。同时,将初始策略部署至真实系统并保留在线探索能力,使控制策略能够根据室外气候、室内负荷及运行工况的实时变化持续自适应优化,克服了基于固定规则或静态模型的控制方法无法动态调整、节能效果受限的缺陷。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121140160B_ABST
    Figure CN121140160B_ABST
Patent Text Reader

Abstract

The application discloses a kind of self-adapting reinforcement learning's air conditioning system water pump real-time variable frequency control method and system, comprising: constructing the air conditioning system data-driven twin model based on historical data;Initial control strategy is generated in pre-training in twin model;Initial control strategy is deployed to real air conditioning system and carries out online adaptive control, current state input strategy is collected, and optimal action or random action based on strategy is selected as second action according to preset probability;Safety check is carried out to second action, and is applied to water pump adjustment operating frequency after application.The application solves the problem that traditional control is difficult to adapt to system nonlinearity and dynamic change, avoids relying on accurate physical modeling, reduces professional threshold, realizes strategy continuous optimization by online exploration, overcomes the defect that static control cannot be dynamically adjusted, introduces safety check to guarantee system operation, promotes reinforcement learning engineering landing, improves energy efficiency and stability, meets the efficient energy-saving demand of intelligent building.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the optimization control of air conditioning systems, specifically to an adaptive reinforcement learning method and system for real-time variable frequency control of water pumps in air conditioning systems. Background Technology

[0002] Air conditioning systems face multiple challenges at the control level due to their nonlinear and dynamic characteristics. Outdoor disturbances such as weather changes, indoor disturbances such as variations in occupant density and work / rest schedules, and adjustments to system operating parameter setpoints all affect their performance. Traditional control methods relying on manual operation are ill-suited to adapting to dynamic environmental changes outside the system, severely hindering the improvement of overall energy efficiency. With the widespread adoption of smart building technologies, building automation systems' real-time monitoring and control capabilities for air conditioning system operation, along with the vast amounts of accumulated operational data, provide valuable resources and a solid foundation for equipment performance analysis and system optimization.

[0003] Current building energy efficiency requirements place higher demands on air conditioning system control methods. Traditional rule-based control methods based on fixed schedules and feedback have significant limitations: the lack of predictive information makes it difficult to compensate for the response lag of parameters such as water temperature during system operation; and the rigid design of control parameters or rules makes it unable to adapt to different building characteristics and diverse climatic conditions, resulting in suboptimal control performance. Model predictive control (MMCC) has shown advantages in the field of energy-saving control of air conditioning systems by establishing system models, predicting environmental disturbances, and using optimization algorithms to solve for the optimal control strategy. However, its differentiated characteristics in different scale scenarios lead to an exponential increase in the complexity of solution adaptation. Accurately modeling the air conditioning system of a specific building is time-consuming and labor-intensive, requiring extensive domain expertise. Moreover, existing MMCC cases generally suffer from short testing cycles and incomplete seasonal coverage, making it difficult to fully verify the long-term stability and maintenance requirements of the system. Any changes in the external environment or the system itself require fine-tuning or calibration of the model, resulting in high maintenance costs. Therefore, traditional static control strategies such as MMCC lack the ability to dynamically self-adjust and cannot proactively adapt to changes in equipment physical performance and external operating conditions.

[0004] Model-free reinforcement learning technology eliminates the need for precise modeling of complex controlled objects and environments. It continuously updates and iterates control strategies through exploration and experience, thereby reducing the reliance on domain-specific knowledge for control strategy formulation. This provides an innovative solution for the operation and control of diverse air conditioning systems. Leveraging its model-free and data-driven characteristics, reinforcement learning algorithms are widely applicable to various technical scenarios, including boilers, heat pumps, district heating, and renewable energy and thermal storage devices, achieving seamless integration of parameters between components. While reinforcement learning applications have experienced rapid development in central air conditioning system control in recent years, they remain largely confined to simulation experiments and the generation of static control strategies. Few engineering applications of reinforcement learning retain action exploration within real-world air conditioning systems to achieve online adaptive learning.

[0005] To address the challenges of energy-saving operation control in air conditioning systems and further unleash the application potential of reinforcement learning technology in air conditioning systems, this invention proposes an adaptive reinforcement learning-based real-time variable frequency control method for water pumps in air conditioning systems. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a variable frequency control method and device for water pumps in air conditioning systems. This invention can overcome the problem of poor adaptability of traditional control methods and solve the technical bottleneck of existing control strategies in achieving continuous adaptive optimization and limited energy-saving effects when facing nonlinear, dynamic, and frequent external disturbances in air conditioning systems. This invention enables intelligent control of water pumps.

[0007] The objective of this invention is achieved through the following technical solution: An adaptive reinforcement learning-based real-time variable frequency control method for water pumps in an air conditioning system includes the following steps: A simplified data-driven twin model of the air conditioning system is constructed. The twin model is based on the historical data of the air conditioning system of the existing building and is constructed through machine learning algorithms to reflect the core operating rules of the air conditioning system. Pre-training is performed in the twin model to generate an initial control policy. The pre-training includes: inputting target samples into the twin model and performing iterative calculations in a virtual reinforcement learning environment to obtain a sample set. Each sample in the sample set includes: a first state of the virtual reinforcement learning environment, a first action obtained according to a preset probability, a reward value obtained after applying the first action, and a second state. The initial control strategy is deployed to a real air conditioning system for online adaptive control. The online adaptive control includes: during the operation of the real air conditioning system, collecting the current state and inputting it into the initial control strategy; selecting a second action according to a preset probability, wherein the second action is the optimal action or a random action based on the initial control strategy; performing a safety check on the second action; if the safety check passes, applying the second action to the water pump in the real air conditioning system to adjust the operating frequency of the water pump.

[0008] An adaptive reinforcement learning-based real-time variable frequency control method for air conditioning system water pumps can significantly improve the energy efficiency ratio and operating efficiency of air conditioning systems. By constructing a twin model based on historical data, this method can accurately simulate the core operating rules of the air conditioning system, thereby generating an effective initial control strategy. Pre-training in a virtual reinforcement learning environment allows for iterative optimization of the control strategy, ensuring rapid adaptation and continuous improvement in real-world applications. In the online adaptive control phase, the system status is collected in real time and combined with the initial control strategy for action selection, while safety checks are performed to ensure the safety and reliability of the control. Finally, by adjusting the water pump's operating frequency in real time, this method can dynamically optimize the performance of the air conditioning system according to actual operating conditions, reducing unnecessary energy consumption.

[0009] As a preferred approach, the security check includes the following steps: The twin model predicts the end-point energy consumption under the second action. If the end-point energy consumption exceeds a preset threshold... If so, the second action is deemed unsafe; If the end-point energy consumption does not exceed the preset threshold Then, the difference between the cooling capacity output by the twin model and the input building cooling load is calculated. If the difference exceeds a preset threshold, the calculation continues. If so, the second action is deemed unsafe; If the second action passes the above two checks, it is determined to be a safe action and can be applied to the water pump in a real air conditioning system.

[0010] As a preferred method, the formula for calculating the reward value is: ; in, For air source heat pump energy consumption, For water pump energy consumption, For end-point energy consumption, For cooling capacity, For building cooling load, , , , These are customizable and adjustable hyperparameters.

[0011] As a preferred approach, it also includes: when the online control encounters a completely new state scenario, calling the twin model to initialize the Q-table value for the new state, so that Q-learning can complete the first policy decision based on simulation experience.

[0012] As a preferred approach, it also includes: in a real air conditioning system, if all actions generated by random exploration fail the safety check, the water pump frequency is forcibly set to the upper limit value, and a system warning is issued.

[0013] As a preferred approach, it also includes: continuously collecting online data in a real air conditioning system according to a preset sampling interval, and executing control actions when a preset control time point is reached.

[0014] As a preferred approach, it also includes setting a buffer time in a real air conditioning system to avoid generating invalid learning samples, during which control commands will not be issued.

[0015] As a preferred approach, the method further includes: in a real air conditioning system, updating the Q-value table in the initial control strategy based on the actual energy consumption index generated by the second action, so as to achieve incremental optimization of the control strategy.

[0016] An adaptive reinforcement learning-based real-time variable frequency control system for an air conditioning system water pump includes: The building unit is used to construct a simplified data-driven twin model of the air conditioning system. The twin model is based on the historical data of the air conditioning system of the existing building and is constructed through machine learning algorithms to reflect the core operating rules of the air conditioning system. A pre-training unit is used to pre-train the Siamese model to generate an initial control policy. The pre-training includes: inputting target samples into the Siamese model and performing iterative calculations in a virtual reinforcement learning environment to obtain a sample set. Each sample in the sample set includes: a first state of the virtual reinforcement learning environment, a first action obtained according to a preset probability, a reward value obtained after applying the first action, and a second state. An online control unit is used to deploy the initial control strategy to a real air conditioning system for online adaptive control. The online adaptive control includes: during the operation of the real air conditioning system, collecting the current state and inputting it into the initial control strategy; selecting a second action according to a preset probability, wherein the second action is the optimal action or a random action based on the initial control strategy; performing a safety check on the second action; if the safety check passes, applying the second action to the water pump in the real air conditioning system to adjust the operating frequency of the water pump.

[0017] As a preferred approach, the pre-training unit adopts the ε-greedy strategy of Q-learning, and gradually reduces reinforcement learning hyperparameters such as random exploration probability ε, fixed number of rounds, learning rate, and discount factor during the pre-training process to complete the pre-training to generate the initial control strategy. The security check includes the following steps: The twin model predicts the end-point energy consumption under the second action. If the end-point energy consumption exceeds a preset threshold... If so, the second action is deemed unsafe; If the end-point energy consumption does not exceed the preset threshold Then, the difference between the cooling capacity output by the twin model and the input building cooling load is calculated. If the difference exceeds a preset threshold, the calculation continues. If so, the second action is deemed unsafe; If the second action passes the above two checks, it is determined to be a safe action and can be applied to the water pump in a real air conditioning system.

[0018] The present invention has at least the following beneficial effects: The adaptive reinforcement learning-based variable frequency control method and device for air conditioning water pumps proposed in this invention effectively solves the problem that traditional control strategies are difficult to adapt to the nonlinearity and dynamic changes of air conditioning systems. By constructing a data-driven twin model and pre-training the reinforcement learning strategy in a virtual environment, the invention avoids the problems of relying on precise physical modeling, high modeling complexity, and high maintenance costs in traditional model predictive control, significantly reducing the dependence on domain-specific knowledge. Simultaneously, by deploying the initial strategy to the real system and retaining online exploration capabilities, the control strategy can continuously and adaptively optimize according to real-time changes in outdoor climate, indoor load, and operating conditions, overcoming the shortcomings of control methods based on fixed rules or static models, which cannot be dynamically adjusted and have limited energy-saving effects. Attached Figure Description

[0019] To reveal the technical details of the embodiments of the present invention, the accompanying drawings involved in the embodiments will be briefly described below. It should be emphasized that these drawings only present several embodiments of the present invention and should not be considered as defining the scope of the invention. For those skilled in the art, other related drawings can still be derived based on these drawings without inventive effort.

[0020] Figure 1 A schematic diagram of the framework for a simplified data-driven twin model of an air conditioning system; Figure 2 This is a schematic diagram of the control method flow for an embodiment. Detailed Implementation

[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.

[0022] In the following description, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. However, it should be understood that the present disclosure is not limited to the specific forms shown herein. Rather, it should be understood to encompass various variations, equivalents, and / or alternatives to the embodiments of the present disclosure. In illustrating the drawings, the same reference numerals will be used to denote similar components.

[0023] In the various embodiments of this disclosure, the terms "first," "second," "the first," or "the second" are intended to modify different components and not to indicate order and / or importance, nor do they constitute a limitation on the respective components. For example, a first user equipment and a second user equipment represent different user equipments, although they both fall under the category of user equipment. Similarly, a first component may be named a second component, and a second component may be named a first component, without changing their essential attributes within the scope of this disclosure.

[0024] In this disclosure, terminology is used to describe specific embodiments and does not constitute a limitation thereof. In this context, the use of the singular form also encompasses the plural form, unless otherwise expressly stated herein. In the course of description, terms such as “comprising” or “having” are intended to indicate the presence of features, quantities, steps, operations, structural components, parts, or combinations thereof, and do not preclude the possibility or addition of one or more other features, quantities, steps, operations, structural components, parts, or combinations thereof.

[0025] It should be clarified that while the following description provides detailed specific information to aid in a comprehensive understanding of the exemplary embodiments, those skilled in the art will recognize that the exemplary embodiments can be implemented even without these specific details. For example, the system may be illustrated using block diagrams to avoid excessive detail that could obscure the clarity of the example. In other cases, to maintain the clarity of the example, unnecessary details of well-known processes, structures, and techniques may be omitted.

[0026] Variable frequency drive (VFD) control of air conditioning system water pumps should consider adapting to the overall operation of the system, rather than adapting to a specific physical quantity (temperature difference or pressure difference between the supply and return water) as in traditional methods. In other words, the purpose of VFD control should be to achieve energy-efficient operation of the system while meeting the building's cooling and heating needs, rather than tracking pre-set supply and return water temperature or pressure differences. The real-time VFD control method for water pumps involved in this invention monitors three state variables in real time: the current outdoor dry-bulb temperature, the load demand at the next control moment, and the water supply temperature of the system. It adaptively learns the optimal water pump frequency setpoint under different states using reinforcement learning technology to guide the real-time operation control of the air conditioning system. In actual operation, after setting a certain water pump frequency, the algorithm monitors the overall energy consumption of the air conditioning system resulting from the action in the current state, using this as the input value for the reward function to calculate the reward value, measuring the merits of the action for the control objective; and using this experience to update the existing control strategy in real time. Meanwhile, real-time on-site control allows for the exploration of better-performing actions with a certain probability. That is, in a given state, instead of using previously accumulated "experience" (existing control strategies), the pump frequency setting is randomly tried to break through the local optima of the existing control strategy and obtain a better-performing control strategy. Considering the impact of random settings on the stability and safety of system operation, random actions must pass safety checks. Through continuous online operation of this method, the reinforcement learning algorithm will evolve the control strategy towards a direction that better suits the characteristics of the system and provides a better pump frequency value. Considering the advantages of Q-learning reinforcement learning algorithms, such as strong interpretability, stable and easy convergence, low online learning resource consumption, and strong transparency and adjustability, this invention is based on the Q-learning algorithm structure. An example is as follows: (1) Establishment of a simplified data-driven twin model of the air conditioning system: Based on the historical data of the existing building's air conditioning system, various data-driven models that can reflect the core operating laws of the air conditioning system are constructed through machine learning algorithms. From the perspective of supply and demand balance, the interaction between the operating parameters of the central air conditioning system is decoupled, and the various data-driven models are connected to establish a simplified data-driven twin model of the air conditioning system (hereinafter referred to as the "system model"), which is used for the generation of initial control strategies and safety checks of online adaptive control. The system model consists of multiple sub-models that fit different physical laws. Due to the dynamic coupling relationship between the operating parameters of the equipment (i.e., the output of the operating result of one device will directly serve as the input of the operating conditions of another device), the input and output of each sub-model will depend on and constrain each other. This correlation mechanism enables the simplified data-driven twin model of the air conditioning system to achieve dynamic simulation through variable transfer. Different air conditioning systems require different sub-models. In practical applications, it is necessary to refer to the specific form of the air conditioning system, the types and quality of available historical operating parameters, and fine-tune the model selection and its input and output parameters. However, it is essential to ensure that the output parameters of the simplified data-driven twin model of the air conditioning system are highly consistent with the objectives of the control task, including but not limited to reflecting the core elements of the air conditioning system's operation such as energy consumption, heat transfer, and fluid distribution network resistance.

[0027] (2) Initial control strategy generation: Determine the reinforcement learning formulaic expression of the target control task, that is, determine the observed state parameters, the action parameters to be executed (the number of water pumps and their respective operating frequencies), and the reward function consistent with the control objective. Among them, the reward function is designed based on reducing the total energy consumption of the air source heat pump air conditioning system, and action constraints are established in terms of the output range of the system model and the balance between cooling supply and demand. Before being deployed to the real system, the Q-learning algorithm will first perform control pre-training on the system model. By inputting real state parameters from historical operating data (consistent with the reinforcement learning formulaic expression, such as outdoor meteorological parameters, load demand, heat pump water supply temperature, etc.), it gives the suggested value of water pump frequency setting under the current state, and continuously evolves its own control strategy by accepting the feedback of the system model as a virtual environment to the actions given by Q-learning. The learning of Q-learning is essentially through a classic random exploration method that balances the relationship between exploration and experience utilization - the "ε-greedy strategy". It ensures that Q-learning has a probability of 1-ε to select the optimal behavior in the existing strategy, and at the same time has a probability of ε to randomly select behavior to explore possible better options. In the early stages of learning, Q-learning lacks sufficient "experience" to determine the optimal action for a given state. Therefore, extensive random exploration is required to build initial problem-solving experience. As the number of explorations increases, Q-learning's reward estimations for each action become more accurate. At this point, exploration becomes time-consuming and unnecessary, so the exploration rate decreases as the learning process progresses until a fixed minimum exploration rate is reached. Through continuous trial and error and feedback updates within the system model, Q-learning learns a preliminary control strategy. During this process, Q-learning learns to execute actions with higher reward values ​​while avoiding penalized actions that do not meet the constraints, based on explicit constraints set in the reward function. This initial control strategy is generated after pre-training, transferring the control "experience" accumulated in the system model to a real-world target air conditioning system for online real-time control.

[0028] (3) Online Adaptive Control: The system model is essentially an operational pattern extracted from historical operating data under limited operating conditions. It often cannot accurately and completely describe all the influencing factors of air conditioning system operation in the real world. In the process of generating the initial control strategy, Q-learning, after pre-training, effectively avoids actions that fail to meet control requirements or lose the stability of air conditioning system operation, significantly improving the control safety and stability of online reinforcement learning. In the online adaptive control stage, Q-learning changes from accepting virtual feedback from the system model to accepting real feedback from the real physical world, and further fine-tunes its own control strategy by accepting real feedback. The code implementation of this method is usually carried out in the form of a program in cooperation with the building's own operation monitoring and management platform to complete the control task. The program continuously collects online data according to the pre-set sampling interval. Considering the response lag of operating parameters such as water temperature after the air conditioning system is turned on, a buffer time is usually set to avoid generating invalid learning samples. Control commands will not be issued during the buffer time; the buffer time is generally reflected by a fixed number of sample collections, and control commands are issued after the buffer time ends. At the same time, the program will monitor the time in real time to determine whether the control time point has been reached. When the set control hour arrives, the program invokes the Q-learning algorithm, using the most recently collected data as input for the state parameters. It then assigns a pump frequency based on the existing control strategy. When assigning the pump frequency, the algorithm randomly selects an action setpoint with a given probability. This means that the online adaptive phase retains the exploration mechanism of Q-learning. By setting a small ε value, it ensures that the actions taken have a low probability of not relying on the "experience" accumulated by Q-learning in the past, but rather randomly exploring new setpoints. This allows Q-learning to potentially further evolve its control strategy towards increasing reward values ​​through random exploration.

[0029] In one embodiment, we introduce a "load prediction error feedback mechanism," which replaces the fixed random exploration probability in traditional reinforcement learning. Dynamically adjust the prediction error based on the cooling load at the next moment. Real-time adjustments enable the system to proactively enhance its exploration capabilities during sudden load changes and reduce its exploration frequency when operating smoothly, effectively balancing energy efficiency and algorithm convergence speed.

[0030] Specifically, the adjusted random exploration probability Defined by the following dimensionless formula: ; It is the dynamically adjusted random exploration probability, with a value range of [0,1]. This is the preset baseline exploration probability, with a default value of 0.1; This is the sensitivity coefficient, with a value ranging from [0,1]. It is set by the user based on the system's response characteristics, and a recommended value is 0.5. Load forecasting error is defined as the difference between the predicted cooling load and the actual cooling load, expressed in kW. This is the current building's actual cooling load, expressed in kW.

[0031] The "error feedback" dynamic exploration mechanism incorporates load prediction errors into the probability adjustment process, enabling efficient adaptive operation of reinforcement learning in building energy conservation control. When the system encounters sudden load changes (such as abrupt weather changes or dense crowds), leading to prediction errors... When significantly increased, the adjusted exploration probability Automatic improvement enhances the algorithm's exploratory capabilities, helping to quickly discover new optimal control strategies and significantly improving the system's robustness in dealing with sudden operating conditions; while when the system operation tends to stabilize and the load prediction error approaches zero, Gradually returning to the preset baseline value This method effectively reduces unnecessary random disturbances, accelerates strategy convergence, and improves control accuracy and energy efficiency. By feeding back external load disturbance information to the exploration mechanism in real time, it achieves a dynamic balance between exploration and utilization, significantly enhancing the algorithm's adaptability and long-term optimization performance in complex, unsteady building load environments.

[0032] It is important to note that all actions generated by random exploration must pass an online safety check based on the system model before they can be applied to the equipment to adjust the operating parameters of the water pump frequency. For example... Figure 2 As shown, actions that fail the safety check are removed from the candidate action set (using a random sampling mechanism without replacement). The system will continuously resample randomly until an action that passes the safety check is obtained. If all actions fail the safety check, the water pump is forced to start at its maximum frequency, and a system warning is issued. Furthermore, when the online control encounters a completely new state scenario (i.e., Q-learning has no historical action experience in this state), this method calls a simplified data-driven twin model of the air conditioning system to initialize the Q-table values ​​for the new state, enabling Q-learning to make its first policy decision based on simulation experience.

[0033] This embodiment formulates a variable frequency control strategy for the water pump based on the energy-saving and high-efficiency requirements of the overall air conditioning system operation. This method uses Q-learning to progressively update the control strategy, dynamically adapting to changes in the internal and external operating conditions of the target air conditioning system. Pre-training Q-learning with historical operating data to obtain the initial control strategy effectively avoids the insecurity and instability introduced by the early-stage action exploration in reinforcement learning for the actual operation of the air conditioning system. A safe and secure online reinforcement learning mechanism is constructed, which, while preserving the characteristics of actual operation, continuously explores better control strategies through the Q-learning algorithm, potentially accelerating the optimization of the system's energy-saving and high-efficiency operation control strategy.

[0034] In another embodiment, an adaptive reinforcement learning-based real-time variable frequency control method for an air conditioning system water pump includes the following steps: A simplified data-driven twin model of the air conditioning system is constructed. The twin model is based on the historical data of the air conditioning system of the existing building and is constructed through machine learning algorithms to reflect the core operating rules of the air conditioning system. The process of constructing a simplified data-driven twin model of the air conditioning system is based on historical data of the existing building's air conditioning system and is accomplished through machine learning algorithms. The core objective of this model is to reflect the core patterns of the air conditioning system in actual operation, thereby providing data support and theoretical basis for subsequent control strategies.

[0035] Specifically, the first step is to collect a large amount of historical operational data. This data includes key parameters such as outdoor dry-bulb temperature, building cooling load demand, and chilled water supply temperature, as well as the energy consumption data and equipment operating status of the air conditioning system under these conditions. This data reflects the operating status of the air conditioning system under different operating conditions and forms the basis for building the twin model. Then, machine learning algorithms are used to analyze and process this data. Machine learning algorithms can uncover hidden patterns and regularities from massive amounts of data, such as the energy consumption characteristics of the air conditioning system under different outdoor temperatures and load demands, and the impact of supply water temperature on energy consumption.

[0036] The construction of the simplified data-driven twin model of the air conditioning system uses massive amounts of operational data accumulated on-site as the sole input, requiring no additional manual experiments. Specifically, during its continuous operation, the building automation system automatically recorded outdoor dry-bulb temperature, building cooling load, chilled water supply temperature, and the corresponding total energy consumption of the air conditioning system at a fixed sampling period of one minute. All raw data is directly stored in the local database without manual screening or secondary processing.

[0037] Subsequently, the random forest regression algorithm was used to train the data offline. This algorithm falls under the category of "machine learning algorithms," and its essence is an ensemble model composed of a large number of decision trees. During the training phase, the algorithm autonomously finds the mapping relationship between each input variable and the system energy consumption without any human intervention, and outputs it in the form of a model file. This model file is the core component of the simplified data-driven twin model of the air conditioning system, and can be directly embedded into the subsequent reinforcement learning pre-training stage for energy consumption prediction and action safety verification in a virtual environment.

[0038] Through the above process, the forest regression algorithm reveals the following nonlinear law: when the outdoor dry-bulb temperature is between 30℃ and 34℃ and the building cooling load is fixed at approximately 1200kW, if the chilled water supply temperature increases from 7℃ to 8℃, the total energy consumption of the system shows a monotonically decreasing trend; when the supply water temperature continues to rise above 8.5℃, the energy consumption no longer decreases, but instead rebounds due to the limitation of the terminal heat exchange capacity. This law is entirely discovered autonomously by the algorithm, without the need for manually setting thresholds or piecewise functions, and can be directly used to guide the design of reinforcement learning reward functions and the setting of safety boundaries.

[0039] A twin model, built using machine learning algorithms, can transform the complex physical processes of an air conditioning system into a mathematical model. This model acts like a "digital twin" of the air conditioning system, simulating its operation in a computer and predicting its energy consumption and performance based on input parameters such as outdoor dry-bulb temperature, load demand, and water supply temperature. For example, given the current outdoor temperature and the building's cooling load demand, the model can predict how much energy the air conditioning system needs to consume to meet these demands and how the water supply temperature should be adjusted for optimal operating efficiency.

[0040] This data-driven approach avoids the reliance on precise physical models of the system found in traditional modeling methods, making model construction more flexible and efficient. The working principle of a twin model is to establish a mapping relationship between input parameters (such as outdoor dry-bulb temperature, load demand, and water supply temperature) and output results (such as energy consumption and equipment operating status) by learning from historical data. When new parameters are input, the model can quickly and accurately predict the corresponding output results based on previously learned patterns. This predictive capability allows twin models to be used in various application scenarios, such as optimizing control strategies, fault diagnosis, and energy-saving assessments. In this way, twin models provide strong support for the intelligent management and optimized control of air conditioning systems, helping to improve system operating efficiency and energy-saving effects.

[0041] Pre-training is performed in the twin model to generate an initial control policy. The pre-training includes: inputting target samples into the twin model and performing iterative calculations in a virtual reinforcement learning environment to obtain a sample set. Each sample in the sample set includes: a first state of the virtual reinforcement learning environment, a first action obtained according to a preset probability, a reward value obtained after applying the first action, and a second state. The initial control strategy is deployed to a real air conditioning system for online adaptive control. The online adaptive control includes: during the operation of the real air conditioning system, collecting the current state and inputting it into the initial control strategy; selecting a second action according to a preset probability, wherein the second action is either the optimal action based on the initial control strategy or a random action; performing a safety check on the second action; if the safety check passes, applying the second action to the water pump in the real air conditioning system to adjust the operating frequency of the water pump; and updating the initial control strategy based on the actual energy consumption feedback generated by the second action in the real air conditioning system to achieve adaptive optimization of the control strategy.

[0042] In the context of reinforcement learning, the first state refers to the environmental information received by the algorithm at the current decision-making moment, such as outdoor dry-bulb temperature, building load, and chilled water supply temperature; the first action is the pump frequency selected by the algorithm based on this state; after executing the first action, the system enters a new operating condition, and the environmental information collected at this time constitutes the second state; the second action is the pump frequency selected by the algorithm again in the second state based on the updated strategy. These four elements form a chain relationship of state, action, new state, and new action, used to calculate rewards and iteratively optimize the control strategy, thereby transforming each actual operation into a source of experience for the next decision.

[0043] In a preferred embodiment, during the pre-training process, an ε-greedy strategy of Q-learning is adopted, and reinforcement learning hyperparameters such as random exploration probability ε, fixed number of rounds, learning rate, and discount factor are gradually reduced as the pre-training process progresses, so as to complete the pre-training to generate an initial control strategy.

[0044] The pre-training phase employs the Q-learning algorithm in conjunction with an ε-greedy strategy, conducted within a virtual environment driven by a simplified data-driven twin model of the air conditioning system. Initially, lacking prior knowledge of the environment, Q-learning executes random actions with a high probability ε to fully traverse the state-action space and rapidly accumulate experience. As training rounds progress, ε is gradually reduced according to a preset decay rule, causing the algorithm to increasingly favor selecting the currently estimated optimal action using the acquired Q-value table. Simultaneously, hyperparameters such as the learning rate and discount factor are set to fixed values ​​before training begins and remain unchanged throughout the pre-training process to balance convergence speed and policy stability. After a predetermined number of training rounds, the Q-value table converges, forming an initial control strategy that can be directly deployed on a real air conditioning system. This strategy avoids high-risk initial exploration on real equipment and lays a reliable foundation for continuous optimization in the subsequent online adaptive phase.

[0045] In a preferred embodiment, the security check includes the following steps: The twin model predicts the end-point energy consumption under the second action. If the end-point energy consumption exceeds a preset threshold... If so, the second action is deemed unsafe; If the end-point energy consumption does not exceed the preset threshold Then, the difference between the cooling capacity output by the twin model and the input building cooling load is calculated. If the difference exceeds a preset threshold, the calculation continues. If so, the second action is deemed unsafe; If the second action passes the above two checks, it is determined to be a safe action and can be applied to the water pump in a real air conditioning system.

[0046] Before the second action provided by Q-learning online is sent to the actual water pump, it must first pass through a double safety gate consisting of a twin model driven by simplified data from the air conditioning system. The first gate allows the twin model to quickly assess the energy consumption of the end-user device under this action: if the predicted value is higher than a preset upper limit... If this action would cause the terminal to be in an overloaded or abnormally high energy consumption state, it would be directly judged as unsafe. If it passes the first level, it enters the second level of the gate, where the twin model further calculates the deviation between the cooling capacity of the air conditioning system and the real-time building cooling load. Once the deviation exceeds the preset tolerance, it will be considered unsafe. This indicates that the action might lead to insufficient or excessive cooling, and it will also be rejected. Only actions that pass both levels of logic verification are considered safe instructions that will neither trigger excessive energy consumption at the terminal nor disrupt the supply-demand balance, thus allowing them to be immediately sent to the real system for execution, ensuring that every step of online learning is completed within a controllable and tolerable range.

[0047] In a preferred embodiment, the formula for calculating the reward value is: ;in, For air source heat pump energy consumption, For water pump energy consumption, For end-point energy consumption, For cooling capacity, For building cooling load, , , , These are customizable and adjustable hyperparameters.

[0048] The design philosophy behind the incentive value is to minimize the total system energy consumption while ensuring that terminal energy consumption does not exceed the limit and that cooling capacity is basically matched with demand. Specifically, a twin model first provides energy consumption predictions for the air source heat pump, water pump, and terminal components. As long as the terminal energy consumption is below the upper limit... Furthermore, the deviation between the cooling capacity and the building load is less than the tolerance limit. The sum of these three energy consumption components (in kW) is taken as a negative value as a reward, with a larger reward value for lower energy consumption; once the end-point energy consumption exceeds k1, a fixed negative penalty is imposed. (Unit: kW) This prevents the algorithm from attempting the action further; if the terminal energy consumption is acceptable but the cooling capacity deviation is too large, another fixed negative value penalty is used instead. (Unit: kW) This reminds the algorithm to pay attention to supply and demand balance. Taking an office building in Chengdu as an example, the upper limit of end-point energy consumption... It can be set at around 120kW, with a tolerance for cooling capacity deviation. 50kW is acceptable, with two penalty values. , Typically, the power is set to 1000kW and 800kW respectively, which can reflect the difference in safety risks and leave enough gradient for the algorithm to learn.

[0049] In a preferred embodiment, during the online adaptive control process, the preset probability is a fixed value to maintain the continuous learning capability of the Q-learning algorithm. During the online adaptive control phase, Q-learning still needs to occasionally "test the waters" to discover better strategies; therefore, the system reserves an ever-present "exploration channel" for it. Specifically, the probability of randomly trying new actions is set to a fixed small value, such as 10%, which does not decay over time in subsequent runs. In this way, the algorithm uses verified optimal experience in most daily decisions to ensure stable operation; however, at the fixed trigger point, it will still actively break away from existing strategies and randomly change the pump frequency to continuously test whether there are more energy-efficient settings. Because this probability is constant, Q-learning can maintain its self-updating ability in the long-term process of equipment aging, climate change, or load pattern migration, without falling into local optima due to the exploration rate dropping to zero.

[0050] In a preferred embodiment, the method further includes: when the online control encounters a completely new state scenario, calling the twin model to initialize the Q-table value for the new state, so that Q-learning can complete the first policy decision based on simulation experience.

[0051] When a completely new combination of operating conditions appears on-site, the Q-learning system's Q-table contains no action value for that state, and direct decision-making would lead to blind actions. At this point, the system immediately invokes the deployed air conditioning system's simplified data to drive the twin model, sending the new state into the model for a "virtual trial run." Based on energy consumption patterns learned from historical data, the twin model quickly assesses the system energy consumption and supply-demand deviation corresponding to the frequencies of each candidate pump in that state, and converts these simulation results into initial action values, filling them into the corresponding cells of the Q-table in one go. In this way, Q-learning can make an initial reasonable decision based on the "prior experience" provided by the twin model without risky trials on real equipment. Subsequent real-world feedback from online operation gradually corrects these initial values, ensuring the strategy is both robust at the start and continuously refined in the future.

[0052] In a preferred embodiment, the system further includes: in a real air conditioning system, if all actions generated by random exploration fail the safety check, the pump frequency is forcibly set to an upper limit value, and a system warning is issued.

[0053] When the algorithm explores randomly online, if the currently sampled pump frequency is still judged to be out of limit by the safety mechanism after being verified one by one by the twin model, the system will immediately remove the frequency from the candidate pool and continue sampling. This process is repeated until the candidate pool is empty and no action is taken. Then, a fallback strategy is triggered: the pump frequency is forcibly raised to the preset upper limit and operated at maximum capacity. At the same time, a system warning is sent to the management platform to ensure that the building's cooling load can still be met under extreme conditions and to prevent cooling interruption due to "no action to take".

[0054] In a preferred embodiment, the method further includes: continuously collecting online data according to a preset sampling interval in a real air conditioning system, and performing a control action when a preset control time point is reached.

[0055] The on-site program automatically reads real-time data such as outdoor temperature, load, and water supply temperature according to a fixed sampling period (e.g., once per minute), forming a continuous state stream. Simultaneously, the system clock continuously compares the current time with the set control cycle, such as every fifteen minutes at the top of the hour. Only when the sampled data has been updated and the clock falls exactly on the control cycle does the program send the latest sample to the Q-learning algorithm, which then provides and issues a new pump frequency. At other times, the program simply caches the data without taking any action, thus ensuring the real-time nature of the decision-making process while avoiding oscillations caused by frequent adjustments.

[0056] In a preferred embodiment, the system further includes setting a buffer time in a real air conditioning system to avoid generating invalid learning samples, during which control commands are not issued.

[0057] When a water pump is first started or its operating conditions change abruptly, the water temperature and flow rate are not yet stable. Data collected at this time may mistake the "transient process" for a "steady-state effect," leading to distortion in subsequent learning samples. To avoid this problem, the system automatically initiates a fixed-duration buffer (e.g., 15 minutes) after issuing the start command. During this buffer period, the program continues sampling, but all data is only used for recording and does not trigger any control actions or write to the training set. Only after the buffer ends and the system truly enters a steady state does the program resume its normal sampling-decision process. This ensures that each piece of data used to update the strategy corresponds to a real and repeatable steady-state condition, preventing misleading reward signals caused by transient disturbances.

[0058] In a preferred embodiment, the method further includes: in a real air conditioning system, updating the Q-value table in the initial control strategy based on the actual energy consumption index generated by the second action, so as to achieve incremental optimization of the control strategy.

[0059] Once the second action is actually executed, the building automation system immediately collects the actual energy consumption of the air source heat pump, water pump, and terminal devices, and calculates the true cost of the action. This measured result is then sent back to the Q-learning core module as direct feedback on the value of the action. The module compares the feedback with previous estimates; if the actual performance is better than expected, the score of the action is increased; otherwise, it is decreased, and the corresponding cell in the Q-value table is updated simultaneously, making the strategy more inclined to high-scoring actions in the next round of decision-making. With the accumulation of each action, measurement, and update loop, the Q-value table is continuously refined, and the control strategy gradually converges to a more energy-efficient and stable operating trajectory, completing the incremental evolution from the initial strategy to the optimal strategy on site.

[0060] An adaptive reinforcement learning-based real-time variable frequency control system for an air conditioning system water pump includes: The building unit is used to construct a simplified data-driven twin model of the air conditioning system. The twin model is based on the historical data of the air conditioning system of the existing building and is constructed through machine learning algorithms to reflect the core operating rules of the air conditioning system. A pre-training unit is used to pre-train the Siamese model to generate an initial control policy. The pre-training includes: inputting target samples into the Siamese model and performing iterative calculations in a virtual reinforcement learning environment to obtain a sample set. Each sample in the sample set includes: a first state of the virtual reinforcement learning environment, a first action obtained according to a preset probability, a reward value obtained after applying the first action, and a second state. An online control unit is used to deploy the initial control strategy to a real air conditioning system for online adaptive control. The online adaptive control includes: during the operation of the real air conditioning system, collecting the current state and inputting it into the initial control strategy; selecting a second action according to a preset probability, wherein the second action is either the optimal action based on the initial control strategy or a random action; performing a safety check on the second action; if the safety check passes, applying the second action to the water pump in the real air conditioning system to adjust the operating frequency of the water pump; and updating the initial control strategy based on the actual energy consumption feedback generated by the second action in the real air conditioning system to achieve adaptive optimization of the control strategy.

[0061] The device operates using a progressive approach: offline modeling, virtual training, and online deployment. First, the construction unit feeds the building's air conditioning system's long-term accumulated operational logs into a machine learning engine, automatically generating a digital twin model that responds in real-time to changes in temperature, load, and flow rate. This model fully replicates the energy consumption patterns of the on-site equipment within the computer. Next, the pre-training unit allows the Q-learning algorithm to repeatedly test and err in this safe sandbox: the model acts as a virtual environment, and the algorithm continuously adjusts the pump frequency and receives energy consumption feedback, gradually accumulating "good actions" into an initial control table. After pre-training, the online control unit implants this experience table into the on-site controller, allowing it to continue "experimenting" with a fixed, low probability, while using the validated optimal strategy the rest of the time. Each action undergoes a double safety check by the twin model to ensure it won't overload the terminal or cause cooling imbalance before execution. After execution, on-site energy consumption data is immediately transmitted back, updating the control table and achieving closed-loop incremental optimization. Through the collaboration of these three units, the device can ensure safety from the outset with mature strategies and continuously evolve during long-term operation, continuously reducing energy consumption.

[0062] In a preferred embodiment, the pre-training unit adopts an ε-greedy strategy of Q-learning, and gradually reduces reinforcement learning hyperparameters such as random exploration probability ε, fixed number of rounds, learning rate, and discount factor during the pre-training process to complete the pre-training to generate the initial control strategy. The security check includes the following steps: The twin model predicts the end-point energy consumption under the second action. If the end-point energy consumption exceeds a preset threshold... If so, the second action is deemed unsafe; If the end-point energy consumption does not exceed the preset threshold Then, the difference between the cooling capacity output by the twin model and the input building cooling load is calculated. If the difference exceeds a preset threshold, the calculation continues. If so, the second action is deemed unsafe; If the second action passes the above two checks, it is determined to be a safe action and can be applied to the water pump in a real air conditioning system.

[0063] During the pre-training phase, the pre-training unit allows Q-learning to repeatedly practice within the Siamese model: the algorithm first randomly tries various pump frequencies with a relatively high probability, quickly accumulating experience with "good actions" and "bad actions," and then gradually narrows the random range, causing the policy to converge to the frequency combination with the lowest energy consumption. Simultaneously, key hyperparameters such as the learning rate and discount factor are set to empirical values ​​at the beginning of training and remain unchanged, accelerating convergence and preventing overfitting, ultimately outputting an initial control policy that can be directly deployed. When this policy is transferred to the real system, every new action generated by the online control unit must first pass through a double safety check performed by the Siamese model: the model first predicts the end-effector energy consumption; if it exceeds the safety limit... If the energy consumption is acceptable, the test is immediately rejected; if the energy consumption is acceptable, the difference between the cooling capacity and the actual load is checked, and if the difference is too large, it is also considered unsafe. Only when both of these checks are passed consecutively is it considered a safety instruction, which is then issued to the on-site water pumps, thus ensuring energy-saving potential while eliminating operational risks.

[0064] In one embodiment of the present invention, an adaptive reinforcement learning-based real-time variable frequency control method for an air conditioning system water pump includes the following steps: This example describes a building in Chengdu, Sichuan Province, where the controlled object is an air-source heat pump air conditioning system. A total of three variable frequency water pumps (two in operation and one on standby) can be controlled. The number of pumps in operation is considered. Np The operating frequency is constant at 2, and the operating frequency of each water pump is set to the same value. .

[0065] (1) For the controlled object in this example, the simplified data-driven twin model of the air conditioning system can be composed of four interconnected models, such as Figure 1These are three models: a water pump energy consumption-flow model, an air conditioning terminal energy consumption model, an air conditioning terminal heat exchange model, and an air source heat pump energy consumption model. These models respectively consider the energy consumption characteristics, heat transfer characteristics, and fluid distribution network resistance characteristics of the equipment. Under a specific building load demand value (taking cooling load as an example), the fluid distribution network of the central air conditioning system will exhibit different resistance characteristics due to the automatic adjustment of valves on the terminal equipment, such as the automatic regulating valves on the water side of the fan coil units. Therefore, the pump model is based on the set parameters of two devices (including the number of pumps activated). N p and water pump frequency setting value f p And a state parameter (building cooling load) CL To estimate the total energy consumption of the water pump. P Pump and the total flow rate of the central air conditioning water system F For total flow F The estimated values ​​are used as output parameters for the pump model, and then shared as input parameters among the subsequent three models. The next end-point energy consumption model uses the supply water temperature... T s and building cooling load CL and total flow F The input parameter is used to estimate the total energy consumption of all air conditioning terminal equipment inside the building. P t The terminal heat transfer model is based on P t , F and T s Estimate return water temperature T r ,Then T r Used to calculate the actual cooling supply of a building. Q , representing the actual cooling capacity generated by the air source heat pump, of which , The specific heat capacity of water. Finally, the energy consumption model of the air source heat pump is based on... Q , T r , F and outdoor dry bulb temperature T d Output air source heat pump energy consumption P ASHP Outdoor dry-bulb temperature affects the air-side heat exchange efficiency of an air-source heat pump, while T s , F and Q This affects energy consumption by influencing the refrigerant's vapor compression refrigeration cycle. P tThis is a key factor in the heat transfer model of air conditioning terminals because almost all the energy consumption of air conditioning terminal equipment is used to transfer the cooling capacity from the water system to the air. For example, in a fan coil unit, the equipment's energy consumption mainly comes from the operation of the fan, which is the core of mechanically driven heat transfer. The working principle of a fan coil unit is to circulate indoor air through a coil containing chilled or hot water, carrying away the cooling and heating capacity transported by the fluid distribution network, thereby regulating the air temperature of the surrounding space. The temperature control of the fan coil unit is achieved through user-adjustable indoor set temperature and fan speed. Fan speed adjustment is achieved through the airflow in the coil, while an electric two-way valve automatically adjusts the water flow inside the pipe based on feedback from the temperature sensor. This dual adjustment mechanism can dynamically respond to actual room temperature changes: when the room temperature deviates from the set value, the coordinated adjustment of fan speed and water flow can quickly adjust the delivery of cooling and heating capacity, thereby accurately matching the room's load demand; when the room temperature reaches the set value, the system automatically reduces the fan speed or water flow to maintain temperature stability and avoid energy waste, thus ensuring that the space temperature always remains within the user-set target range. In conclusion, N p , f p and T s As an action of an intelligent agent, and CL and T d The environmental states observed by the agent collectively constitute the input to the entire pre-training framework, initiating the computation process. After computation, three energy consumption results are generated. P Pump , P t and P ASHP Output from a virtual environment based on a pre-trained framework and provided to the agent to calculate reward values ​​under different states and actions.

[0066] (2) After establishing the system model, determine the reinforcement learning formula for the control task: the state parameter is set as "outdoor dry-bulb temperature". T d Building cooling load CL chilled water supply temperature T s The action parameter is set to "pump frequency". f pump "Given the characteristic of Q-learning in processing discrete state variables to discrete action variables, the state parameters and action parameters must be discretized: "

[0067]

[0068]

[0069]

[0070] The reward function is designed to reduce the total energy consumption of the air conditioning system and to constrain unsafe and unstable actions. The design is as follows: ;in, , , These are customizable hyperparameters. This reinforcement learning formula is applied throughout the entire reinforcement learning process, including both pre-training and online training phases.

[0071] (3) This formula expresses the pre-training of the simplified data-driven twin model of the air conditioning system under virtual conditions, with the input state conditions using real hourly state samples from historical data. Q-learning is employed. -greedy( The strategy is a (greedy) approach, and the probability of random exploration is gradually reduced during the pre-training process. The reinforcement learning hyperparameters, such as fixed number of rounds, learning rate, and discount factor, are used to complete pre-training and generate the initial control policy.

[0072] (4) Deploy the initial control strategy to the real system for online control. The online control method is as follows: Figure 2 As shown in the diagram, the program in this case collects data every minute. The program's control time points are hh:00, hh:15, hh:30, and hh:45 every hour. The program will detect whether the water pump is on. If the water pump is not on, no control will be performed. The first control after the water pump is on requires the water pump to run continuously for 15 minutes and reach the control point before control is performed. The control time interval and buffer time can be customized according to the specific project.

[0073] (5) When the pump control point is reached and control is required, the program will call the Q-learning algorithm. First, it collects the action taken at the previous control moment and the actual energy consumption feedback it brings, in order to calculate the reward value brought by the action. Then, it collects the current state and calculates the action to be taken. If the control action is random exploration, the value will go through a safety check process. First, the virtual model will calculate the end energy consumption under this value. If the end energy consumption exceeds the value set in the reward function, the program will proceed accordingly. If the value is not met, it is considered unsafe; if the end-point energy consumption does not exceed the value set in the reward function... Then continue calculating the difference between the cooling capacity output by the system model and the input building cooling load. If the difference exceeds... If the value is not deemed safe, it is considered unsafe. Actions not rejected by the safety check framework are considered safe and executable, and the value can be used in the equipment. During the execution of the pump frequency command, the system continuously monitors until the next baseline control moment, collects the actual energy consumption index generated by the action to update the reward function, and feeds back the complete state-action-reward sequence to the Q-learning algorithm framework to complete the incremental update of the Q-value table and achieve progressive optimization of the control strategy.

[0074] The specific process is as follows: 1. Deploy the initial control strategy, pre-trained in a simplified data-driven twin model of the air conditioning system, into the real system and configure a fine-tuned fixed exploration probability. (10% in this case) to maintain the continuous learning ability of the Q-learning algorithm.

[0075] 2. The program will run continuously, collecting data at 1-minute intervals. Assuming the program detects that the water pump has been turned on at 10:02, it will trigger the control timing calculation logic. According to preset rules, the water pump must first complete a buffer period (15 minutes in this case). Since the next control point is 10:15, the water pump has only been running for 13 minutes and has not yet completed this buffer period. The system will automatically skip this control window until the buffer requirement is met at the 10:30 baseline time, at which point the first control action will be executed.

[0076] 3. At the baseline control time of 10:30, the system collects the real-time state vector S={outdoor dry-bulb temperature: 30℃, building cooling load: 1602kW, water supply temperature: 7.5℃} as the state input for the Q-learning agent. Subsequently, the Q-learning agent will output the optimal control action a=41Hz based on the control strategy it has learned. This action needs to pass a two-level safety verification framework: first, the data-driven twin model predicts the end-point energy consumption; if the first-level verification criterion is met (predicted end-point energy consumption < 41Hz), then... Then proceed to the second level of verification—calculate the absolute value of the deviation between the predicted cooling capacity and the building's cooling load. When |ΔQ| < The action is deemed safe upon detection. The control action, triggered by the safety detection framework, is then sent to the actual water pump equipment for execution.

[0077] 4. When performing step 3, the program has a fixed probability ( =10%) to conduct exploration, that is, in each control cycle there is Instead of basing the probability on the optimal policy output of the Q-table, the pump frequency value (e.g., 35Hz) is randomly sampled from the action space. This value is then used to control the pump after passing a safety check. This mechanism ensures continuous exploration of the policy space within the safety constraints, preventing the pump from getting trapped in locally optimal control policies.

[0078] 5. If the random action a=35Hz given in step 4 fails the safety check, Q-learning will resample a new action a'=36Hz from the discrete action pool that does not include the previous action. This process will continue to iterate until a safe action is obtained. If all 21 preset frequency levels (discrete distribution of 30-50Hz) fail the verification, the pump frequency will be forcibly set to the upper limit of 50Hz, and an abnormal alarm signal will be triggered.

[0079] 6. At the next control point, 10:45, the system traces back to the actual energy consumption consequences of the pump frequency action a=41Hz executed in step 3 (10:30). It collects energy consumption data from each subsystem of the air conditioning system via the sensor network, uses the collected energy consumption data to calculate the reward R, and feeds this reward R, along with the state S and action a at control point 10:30, as a tuple (S, a, R) to the Q-learning algorithm. This completes the incremental update of the Q-value table, achieving progressive optimization of the policy space. The decision-making process of step 3 is then repeated, forming a closed-loop control cycle.

[0080] This solution constructs an online adaptive reinforcement learning-based variable frequency control architecture for air conditioning system pumps. Through an iterative mechanism of real-time state perception, action decision-making, and reward feedback, it can dynamically match the optimal pump frequency configuration under the current operating conditions. This strategy is equipped with a self-learning system that can dynamically update the control strategy based on energy consumption feedback values, avoiding control strategy failure due to changes in equipment physical performance. A dual safety verification framework ensures that control actions are reasonable while meeting the building's cooling load requirements, ultimately achieving the energy-saving control goal under dynamic environments.

[0081] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention. The above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A real-time variable frequency control method for water pumps in an air conditioning system using adaptive reinforcement learning, characterized in that, Includes the following steps: A simplified data-driven twin model of the air conditioning system is constructed. The twin model is based on the historical data of the air conditioning system of the existing building and is constructed through machine learning algorithms to reflect the core operating rules of the air conditioning system. Pre-training is performed in the twin model to generate an initial control policy. The pre-training includes: inputting target samples into the twin model and performing iterative calculations in a virtual reinforcement learning environment to obtain a sample set. Each sample in the sample set includes: a first state of the virtual reinforcement learning environment, a first action obtained according to a preset probability, a reward value obtained after applying the first action, and a second state. The initial control strategy is deployed to a real air conditioning system for online adaptive control. The online adaptive control includes: during the operation of the real air conditioning system, collecting the current state and inputting it into the initial control strategy; selecting a second action according to a preset probability, wherein the second action is the optimal action or a random action based on the initial control strategy; performing a safety check on the second action; if the safety check passes, applying the second action to the water pump in the real air conditioning system to adjust the operating frequency of the water pump. The security check includes the following steps: The terminal energy consumption under the second action is predicted by the twin model. If the terminal energy consumption exceeds the preset threshold k1, the second action is determined to be unsafe. If the terminal energy consumption does not exceed the preset threshold k1, the difference between the cooling capacity output by the twin model and the input building cooling load will continue to be calculated. If the difference exceeds the preset threshold k2, the second action will be deemed unsafe. If the second action passes the above two checks, it is determined to be a safe action and can be applied to the water pump in a real air conditioning system. The formula for calculating the reward value is: ; in, For air source heat pump energy consumption, For water pump energy consumption, For end-point energy consumption, For cooling capacity, For building cooling load, , , , These are customizable and adjustable hyperparameters.

2. The adaptive reinforcement learning-based real-time variable frequency control method for air conditioning system water pumps according to claim 1, characterized in that, Also includes: When the online control encounters a completely new state scenario, the twin model is invoked to initialize the Q-table values ​​for the new state, enabling Q-learning to make the first policy decision based on simulation experience.

3. The adaptive reinforcement learning-based real-time variable frequency control method for air conditioning system water pumps according to claim 1, characterized in that, Also includes: In a real air conditioning system, if all actions generated by random exploration fail the safety check, the water pump frequency is forcibly set to the upper limit, and a system warning is issued.

4. The adaptive reinforcement learning-based real-time variable frequency control method for air conditioning system water pumps according to claim 1, characterized in that, Also includes: In a real air conditioning system, online data is continuously collected according to a preset sampling interval, and control actions are executed when the preset control time point is reached.

5. A real-time variable frequency control method for water pumps in an air conditioning system based on adaptive reinforcement learning, as described in claim 1 or 4, characterized in that... Also includes: In real air conditioning systems, a buffer time is set to avoid generating invalid learning samples, and control commands will not be issued during the buffer time.

6. The adaptive reinforcement learning-based real-time variable frequency control method for air conditioning system water pumps according to claim 1, characterized in that, Also includes: In a real air conditioning system, the Q-value table in the initial control strategy is updated based on the actual energy consumption index generated by the second action, so as to achieve incremental optimization of the control strategy.

7. A real-time variable frequency control system for an air conditioning system water pump based on adaptive reinforcement learning, characterized in that, include: The building unit is used to construct a simplified data-driven twin model of the air conditioning system. The twin model is based on the historical data of the air conditioning system of the existing building and is constructed through machine learning algorithms to reflect the core operating rules of the air conditioning system. A pre-training unit is used to pre-train the Siamese model to generate an initial control policy. The pre-training includes: inputting target samples into the Siamese model and performing iterative calculations in a virtual reinforcement learning environment to obtain a sample set. Each sample in the sample set includes: a first state of the virtual reinforcement learning environment, a first action obtained according to a preset probability, a reward value obtained after applying the first action, and a second state. An online control unit is used to deploy the initial control strategy to a real air conditioning system for online adaptive control. The online adaptive control includes: during the operation of the real air conditioning system, collecting the current state and inputting it into the initial control strategy; selecting a second action according to a preset probability, wherein the second action is the optimal action or a random action based on the initial control strategy; performing a safety check on the second action; if the safety check passes, applying the second action to the water pump in the real air conditioning system to adjust the operating frequency of the water pump.

8. The adaptive reinforcement learning-based real-time frequency conversion control system for air conditioning system water pumps according to claim 7, characterized in that, The pre-training unit adopts the ε-greedy strategy of Q-learning, and gradually reduces reinforcement learning hyperparameters such as random exploration probability ε, fixed number of rounds, learning rate, and discount factor during the pre-training process to complete the pre-training and generate the initial control strategy. The security check includes the following steps: The terminal energy consumption under the second action is predicted by the twin model. If the terminal energy consumption exceeds the preset threshold k1, the second action is determined to be unsafe. If the terminal energy consumption does not exceed the preset threshold k1, the difference between the cooling capacity output by the twin model and the input building cooling load will continue to be calculated. If the difference exceeds the preset threshold k2, the second action will be deemed unsafe. If the second action passes the above two checks, it is determined to be a safe action and can be applied to the water pump in a real air conditioning system.

Citation Information

Patent Citations

  • Central air conditioner chilled water control system based on frequency conversion control adjustment

    CN118654364A

  • Internet of things control system for heating and energy and water saving of air conditioning

    WO2014161505A1