Office building heating ventilation and air conditioning system adaptive energy-saving operation method and device

CN122813340APending Publication Date: 2026-09-25CHENGWU POWER SUPPLY CO STATE GRID SHANDONG ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610968985.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]针对现有技术的不足,本发明提供一种办公楼宇暖通空调系统自适应节能运行方法及装置,具备多源异构数据融合感知、基于数字孪生的热环境实时预测、深度强化学习驱动的柔性需求响应以及模型在线自进化迭代等优点,解决了传统控制系统因感知维度单一导致的“盲控”与能源浪费、缺乏建筑热惯性利用而造成的舒适度牺牲、以及离线模型无法适应环境变化导致的控制失效与频繁人工标定的问题

Benefits of technology

一、本申请通过融合多模态视觉感知数据(人员位置、数量、活动姿态及衣物厚度系数 Kdo),突破了传统单一温湿度传感器无法捕捉局部冷热不均和人员动态分布的技术局限。与传统仅依赖固定点位传感器的盲控模式不同,本申请将视觉元数据直接映射至RC数字孪生模型的内部热源项 Qinternal,i,使模型能够根据区域内人员的实时活动状态(静坐60W/m²至剧烈运动150 W/m²)和衣物厚度动态修正散热功率。系统据此实时识别高活跃区、静坐区及无人区,并动态调整变风量末端的风向、风速及阀门开度,彻底解决了人走风不停或人热风未至的供需错配问题,实现了真正的按需供能,大幅提升了空间热环境的均匀性与舒适度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813340A_ABST
    Figure CN122813340A_ABST
Patent Text Reader

Abstract

The present application relates to the HVAC (Heating, Ventilation and Air Conditioning) technical field, in particular to an office building HVAC system adaptive energy-saving operation method and device, which solves the problems of blind control and energy waste caused by single sensing dimension, control failure caused by offline model unable to adapt to environmental changes, and comfort sacrifice caused by lack of building thermal inertia utilization in the prior art. The method comprises: collecting multi-source heterogeneous data in real time, constructing a lightweight digital twin thermal environment model and dynamically updating the model parameters using online incremental learning; setting the HVAC system as an agent, outputting the optimal control action sequence based on the deep reinforcement learning algorithm and the multi-objective reward function containing energy consumption, comfort, carbon emissions and demand response income; dynamically dividing virtual microclimate zones based on visual perception data and implementing differentiated control; executing the control action and monitoring the deviation to trigger the model correction mechanism. The present application realizes on-demand energy supply by integrating visual perception and digital twin technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heating, ventilation and air conditioning system technology, and specifically to an adaptive energy-saving operation method and device for an office building heating, ventilation and air conditioning system. Background Technology

[0002] Energy-saving renovations of public buildings, especially office buildings, have become a key focus of the industry. As a major source of building energy consumption, the level of intelligence in the control strategy of heating, ventilation, and air conditioning (HVAC) systems directly determines energy-saving effects and user comfort. However, existing HVAC control systems still face significant technical bottlenecks in practical applications, making it difficult to meet current demands for refined, adaptive, and flexible control.

[0003] (i) Blind control and energy waste caused by a single perception dimension Traditional control systems primarily rely on a small number of temperature and humidity sensors deployed in fixed locations to acquire environmental data. They cannot capture real-time variations in temperature in localized indoor areas, and lack the ability to sense dynamic load factors such as the specific distribution of people, their activity levels, and the thickness of their clothing. This blind control mode means the system can only adjust based on preset schedules or fixed thresholds, failing to respond in real-time to changes in the flow of people. For example, the system may continue to operate at high load after people leave the area, or significant heat buildup may occur in densely populated areas, resulting in continuous airflow even when people are gone, or insufficient warm air reaching those who are present, leading to significant energy waste and comfort blind spots.

[0004] (ii) Control failure caused by the offline model's inability to adapt to environmental changes Existing AI control models mostly employ offline training methods, resulting in a static operating state once deployed. They cannot self-update based on environmental parameter drift caused by aging building envelopes, renovations, or changes in occupant habits. When building physical characteristics change, the deviation between model predictions and actual thermal environments increases rapidly, leading to control failure and often requiring frequent manual recalibration.

[0005] (iii) Sacrifice of comfort due to lack of utilization of building thermal inertia When responding to grid demand response (DR) commands, the lack of accurate thermal inertia prediction based on digital twin models often leads to abrupt shutdowns or large temperature adjustments, severely sacrificing user comfort. The system also lacks the ability to use building thermal inertia for flexible adjustments, making it difficult to achieve an optimal balance between energy costs and comfort.

[0006] Therefore, there is an urgent need for an adaptive energy-saving operation method that can integrate real-time visual perception and dynamic digital twin models, has millisecond-level response speed, on-demand power supply capability, and online self-evolution mechanism, in order to solve the above-mentioned problems of missing perception and rigid control. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides an adaptive energy-saving operation method and device for HVAC systems in office buildings. It has advantages such as multi-source heterogeneous data fusion sensing, real-time thermal environment prediction based on digital twins, flexible demand response driven by deep reinforcement learning, and online self-evolution and iteration of models. It solves the problems of "blind control" and energy waste caused by the single sensing dimension of traditional control systems, the sacrifice of comfort caused by the lack of utilization of building thermal inertia, and the control failure and frequent manual calibration caused by the inability of offline models to adapt to environmental changes.

[0008] This invention is achieved through the following technical solution: An adaptive energy-saving operation method for HVAC systems in office buildings is provided, which is implemented using a layered computing architecture. The layered computing architecture includes a perception layer, an edge computing layer, and a cloud training layer. The method includes the following steps: S1, the perception layer collects multi-source heterogeneous data in real time. This data includes indoor and outdoor environmental physical parameters, HVAC equipment operating status parameters, electricity market signals, and multimodal visual perception data acquired through non-contact sensors. The multimodal visual perception data is preprocessed and locally anonymized at edge computing nodes, with only structured metadata extracted and uploaded to the cloud or central controller. This structured metadata includes personnel coordinates, number of people, activity postures, and clothing thickness coefficient K. do The edge computing layer deploys a lightweight digital twin thermal environment model inference engine and a deep reinforcement learning policy execution module to issue control commands in milliseconds; the cloud training layer is responsible for the offline pre-training of the model, incremental learning updates, and global optimization of the multi-objective reward function, and the layers interact with each other through a secure communication protocol. S2, a lightweight digital twin thermal environment model is constructed based on multi-source heterogeneous data. The building space is discretized into several thermal nodes, and a set of heat balance equations based on thermal resistance-heat capacity RC networks is established. The set of heat balance equations includes the internal heat source power Q at each thermal node. internal,i The heat gain from solar radiation, Q solar,i and the heating or cooling power Q of the HVAC system HVAC,i Included as an independent item; mapping the activity pose and clothing thickness coefficient in the structured metadata to the internal heat source item Q of the RC network. internal,i This allows the internal heat source term to be dynamically adjusted between a baseline value of 60 W / m² and a value of 150 W / m² for strenuous activity, depending on the activity level of people in the area, and is based on the clothing thickness coefficient K. do The heat dissipation power of the human body is corrected; an online incremental learning algorithm is used to dynamically update the thermal resistance matrix R and thermal capacity vector C of the RC network with the residual between the model prediction value and the actual measurement value at historical time as the optimization target. S3 sets the HVAC system as a deep reinforcement learning agent, constructs a four-objective comprehensive reward function including energy consumption cost, comfort deviation penalty, carbon emission cost, and grid demand response benefit. Combined with the prediction results of the digital twin thermal environment model, it outputs the optimal control action sequence including pre-cooling / pre-heating, supply air temperature setpoint, fresh air valve opening, and variable air volume terminal air speed setting. When it detects that the electricity price is at its peak or the grid carbon factor is higher than the preset benchmark, it calls the digital twin thermal environment model to predict the building's thermal inertia in the future period and fine-tunes the indoor set temperature towards the upper or lower limit of the comfort range, instead of executing shutdown or large temperature adjustment. S4, based on the real-time location, number and activity status of people in multimodal visual perception data, dynamically divides virtual microclimate zones in the digital space, and determines the wind direction, wind speed and valve opening adjustment parameters of the variable air volume terminal according to the load characteristics of each zone. S5 executes the optimal control action sequence, monitors the deviation between the actual thermal environment and the predicted value of the digital twin thermal environment model in real time. If the deviation exceeds the preset threshold, the model correction mechanism is triggered to synchronously update the thermal resistance and thermal capacity parameters of the RC network in S2 and the weights of the deep reinforcement learning strategy network in S3. Implicit user feedback data is also collected, including the adjustment frequency of the temperature control panel, the window opening status, and the personal fan usage record. The user satisfaction index is calculated based on the implicit user feedback data, and the weight coefficient of the comfort deviation penalty term in S3 is dynamically adjusted based on the satisfaction index.

[0009] Furthermore, in step S2, the online incremental learning algorithm is either Recursive Least Squares (RLS) or Gradient Descent, and the historical prediction error Δe=∣T is collected once at preset intervals. pred -T real | When the error exceeds the preset threshold of 0.5℃ multiple times, the parameter identification module is triggered to fine-tune the thermal resistance matrix R and the thermal capacity vector C to minimize the sum of squared residuals. When the deviation exceeds the threshold and triggers the model correction mechanism, the RC network parameters and the deep reinforcement learning strategy network weights are updated simultaneously.

[0010] Furthermore, the specific method for dynamically dividing the virtual microclimate zones in step S4 is as follows: based on the real-time personnel density distribution map, the office area is divided into three categories: high-activity zone, quiet office zone, and unmanned zone using the K-means clustering algorithm; for the high-activity zone, the number of fresh air exchanges is increased and the supply air temperature is reduced; for the quiet office zone, the basic ventilation volume and constant temperature control are maintained; for the unmanned zone, when the continuous detection time exceeds 30 minutes, the variable air volume terminal valve of the area is closed or the system is switched to the minimum maintenance operation mode.

[0011] Furthermore, the four-objective comprehensive reward function in step S3 is expressed as: ; in, to These are the weighting coefficients for energy consumption costs, comfort deviation penalties, carbon emission costs, and grid demand response benefits, respectively.

[0012] Furthermore, in step S3, when it is detected that the outdoor weather conditions are suitable and the outdoor enthalpy is lower than the indoor set enthalpy, the usage ratio of natural ventilation or free cooling mode is increased, and the output of mechanical refrigeration equipment is limited.

[0013] Furthermore, it also includes steps for optimizing fresh air during the transition season: determining the outdoor air enthalpy H. out Compared with the indoor set enthalpy value H in The size relationship, if H out <H in If the outdoor air quality index (AQI) meets the preset threshold, the free fresh air cooling mode is activated and the fresh air valve opening is dynamically maximized; if the outdoor AQI exceeds the standard, the mode is switched to internal circulation mode, and the minimum fresh air volume required to maintain indoor air quality is calculated based on the mapping relationship between CO2 concentration and personnel density, and the fresh air valve opening is dynamically adjusted.

[0014] Furthermore, it also includes coordinated heat storage and release control in winter heating scenarios: during off-peak electricity periods and when outdoor temperatures are low, the indoor set temperature is dynamically increased to the upper limit of the comfort zone, utilizing the heat capacity of the building envelope and furniture for sensible heat storage; after entering the daytime peak electricity consumption period, the heating output is gradually reduced based on the indoor temperature decay rate predicted by the digital twin model, and the building releases stored heat to maintain room temperature; for high-density areas such as large conference rooms, the output power of floor radiant heating or hot air curtains is dynamically increased based on the visually perceived density of people, combined with local air supply to avoid local overheating; in sparsely populated areas, the heating output is reduced while maintaining the basic temperature set value.

[0015] The present invention also provides an apparatus using an adaptive energy-saving operation method for office building HVAC systems, comprising: The perception layer is used to collect multi-source heterogeneous data in real time, including indoor and outdoor environmental physical parameters, equipment operating status parameters, power market signals and multimodal visual perception data obtained by non-contact sensors. The raw visual data is locally desensitized at the edge computing node and only structured metadata is uploaded. The digital twin layer, deployed on edge computing nodes, is used to build a lightweight RC hot node network model and map the activity pose and clothing thickness coefficient in the structured metadata to the internal heat source term Q of the RC network. internal,i This allows the internal heat source term to be dynamically adjusted according to the activity status of people in the area, and an online incremental learning algorithm is used to dynamically update the thermal resistance matrix and heat capacity vector; the RC heat balance equations will Q internal,i Q solar,iand Q HVAC,i Included as a separate item; The decision layer, deployed on edge computing nodes, is used to build the HVAC system into a deep reinforcement learning agent. The agent's state space consists of indoor temperature, outdoor temperature, real-time electricity price, number of people, CO2 concentration, grid carbon factor, and time period. The optimization objective is a comprehensive reward function with four objectives, including energy consumption cost, comfort deviation penalty, carbon emission cost, and grid demand response benefit. Combined with the prediction results of the digital twin layer, it outputs the optimal control action sequence, including pre-cooling / pre-heating, supply air temperature setpoint, fresh air valve opening, and variable air volume terminal wind speed setting. In the grid demand response scenario, the digital twin layer is called to predict the building's thermal inertia to perform pre-cooling / pre-heating and comfort zone fine-tuning. The cloud-based training layer is used for offline pre-training, incremental learning updates, and global optimization of multi-objective reward functions for deep reinforcement learning agents. The execution and feedback layer, deployed on edge computing nodes, is used to execute the optimal control action sequence, monitor the deviation between the actual thermal environment and the predicted value, and synchronously update the RC network parameters and the deep reinforcement learning strategy network weights when the deviation exceeds the threshold. It also calculates the user satisfaction index based on implicit user feedback data and dynamically adjusts the weight coefficient of the comfort deviation penalty term based on the satisfaction index.

[0016] The beneficial effects of this invention are: I. This application integrates multimodal visual perception data (person location, number, activity posture, and clothing thickness coefficient K) do This breakthrough overcomes the technical limitations of traditional single temperature and humidity sensors, which cannot capture localized uneven heating and cooling and the dynamic distribution of people. Unlike the traditional blind control mode that relies solely on fixed-point sensors, this application directly maps visual metadata to the internal heat source item Q of the RC digital twin model. internal,i This allows the model to dynamically adjust the heat dissipation power based on the real-time activity status of people in the area (60W / m² when sitting to 150W / m² during vigorous exercise) and the thickness of their clothing. The system thus identifies high-activity areas, sitting areas, and unoccupied areas in real time, and dynamically adjusts the airflow direction, speed, and valve opening of the variable air volume (VAV) terminals. This completely solves the supply-demand mismatch problem of continuous airflow when people leave or insufficient warm airflow when people are present, achieving true on-demand energy supply and significantly improving the uniformity and comfort of the spatial thermal environment.

[0017] II. This application employs a lightweight digital twin thermal environment model (RC thermal resistance-thermal capacity network) combined with an online incremental learning algorithm. The model discretizes the building space into thermal nodes and incorporates internal heat sources, solar radiation heat gain, and HVAC power as independent terms into the heat balance equations. The system collects prediction errors at preset intervals. When the error exceeds 0.5℃ multiple times consecutively, it triggers recursive least squares (RLS) or gradient descent to fine-tune the thermal resistance matrix and thermal capacity vector. More importantly, when the deviation between the actual thermal environment and the model prediction exceeds a preset threshold, the system simultaneously triggers a dual-track correction mechanism involving the RC network parameters and the deep reinforcement learning strategy network weights. This dual online self-evolutionary mechanism enables the system to automatically adapt to environmental characteristic drift caused by aging building envelopes, renovations, and seasonal changes, maintaining long-term high-precision control performance without manual intervention. This overcomes the shortcomings of traditional offline models, which fail once deployed and require frequent manual recalibration.

[0018] Third, this application sets the HVAC system as a deep reinforcement learning agent, constructing a four-objective comprehensive reward function that includes energy consumption costs, comfort deviation penalties, carbon emission costs, and grid demand response benefits. When responding to grid demand response (DR) commands, the system no longer adopts traditional strategies that sacrifice comfort, such as abrupt shutdowns or drastic temperature adjustments. Instead, it uses a digital twin model to accurately predict building thermal inertia and proactively performs pre-cooling / pre-heating and fine-tuning of the comfort range (e.g., adjusting from 24℃ to 25.5℃ in summer). This flexible adjustment mechanism effectively reduces energy costs and carbon emission indicators while ensuring the indoor environment remains within the comfort range, achieving a win-win situation for economic benefits, social responsibility, and user experience.

[0019] IV. Based on multimodal visual perception data, this application uses the K-means clustering algorithm to dynamically divide the office area into three virtual microclimate zones: a high-activity zone, a quiet office zone, and an unoccupied zone, and implements differentiated control. For the high-activity zone, the number of fresh air exchanges is increased and the supply air temperature is reduced to quickly remove accumulated heat; for the quiet office zone, basic ventilation and constant temperature control are maintained; for the unoccupied zone, the terminal valve is closed or the system switches to the minimum maintenance operation mode after continuous monitoring for more than 30 minutes. This differentiated control based on the actual distribution of personnel avoids energy waste in the traditional uniform energy supply mode and achieves refined on-demand energy supply in the spatial dimension.

[0020] Fifth, this application collects implicit user feedback data (temperature control panel adjustment frequency, window opening status, and personal fan usage records), calculates user satisfaction indicators based on this data, and dynamically adjusts the weight coefficients of the comfort deviation penalty term in the deep reinforcement learning reward function. This mechanism enables the control strategy to continuously adapt to changes in user behavior habits, achieving self-iteration and optimization of the control strategy, and avoiding a decline in user satisfaction due to fixed weight settings.

[0021] In summary, this application integrates visual perception and digital twin technologies to construct a closed-loop adaptive energy-saving operation system encompassing perception, modeling, decision-making, execution, and feedback. This system simultaneously addresses three major technical issues of traditional control systems: "blind control" and energy waste due to a single perception dimension; control failure caused by offline models' inability to adapt to environmental changes; and the sacrifice of comfort due to a lack of utilization of building thermal inertia. This achieves synergistic optimization of HVAC systems across four dimensions: energy saving, comfort, low carbon emissions, and demand response. Attached Figure Description

[0022] Figure 1 This is a flowchart of the adaptive energy-saving operation method for HVAC systems in office buildings according to the present invention.

[0023] Figure 2 This is a system architecture diagram provided for an embodiment of the present invention. Detailed Implementation

[0024] To clearly illustrate the technical features of this solution, the following detailed implementation method will be used to explain the solution. Example

[0025] like Figure 2 As shown in the figure, this embodiment provides an adaptive energy-saving operation method for HVAC systems in office buildings. This method is implemented based on a hierarchical computing architecture, and the overall system mainly consists of a perception layer, an edge computing layer, and a cloud training layer.

[0026] The perception layer acquires multi-source heterogeneous data through various sensors deployed in the office area, including indoor and outdoor environmental physical parameters (temperature, humidity, CO2 concentration), HVAC equipment operating status parameters (fan frequency, water valve opening, chilled water supply and return water temperature), real-time electricity market signals (electricity price, grid carbon factor), and multimodal visual perception data acquired through non-contact sensors (such as high-definition cameras or millimeter-wave radar).

[0027] The edge computing layer is equipped with an inference engine for a lightweight digital twin thermal environment model and a deep reinforcement learning policy execution module, which is responsible for local data preprocessing, privacy desensitization, real-time thermal environment prediction, and millisecond-level control command issuance.

[0028] The cloud-based training layer is responsible for offline pre-training of the model, incremental learning updates, and global optimization of the multi-objective reward function. Data exchange between layers occurs via a secure communication protocol, forming a closed-loop control system.

[0029] like Figure 1 As shown, the specific operation process includes the following steps: Step S1: Real-time acquisition and fusion of multi-source heterogeneous data The system synchronously collects various data at a preset frequency (e.g., 1Hz). Physical data includes indoor and outdoor dry / wet bulb temperatures, CO2 concentration, real-time electricity price signals, and grid carbon factor. Equipment data includes fan frequency, water valve opening, chilled water supply and return temperatures, etc., and is read via BACnet or Modbus protocols.

[0030] The visual data processing flow is as follows: Edge computing nodes receive raw image streams or point cloud data from indoor cameras, perform preprocessing and privacy desensitization operations locally, and do not upload any raw images containing faces or sensitive information; only structured metadata is extracted and uploaded. This metadata includes the three-dimensional coordinates (x, y, z) of people, real-time people count N, activity posture classification (sitting, walking, or vigorous exercise), and clothing thickness coefficient K estimated based on infrared thermal imaging. do All data is then formatted and input into the subsequent control model.

[0031] Step S2: Construction and online updating of a lightweight digital twin thermal environment model Based on the aforementioned multi-source data, a lightweight digital twin thermal environment model is constructed and run. The core of this model lies in discretizing the building space into a network of thermal nodes and establishing a linearized set of heat balance equations based on a thermal resistance-heat capacity (RC) network to describe the heat transfer process. The specific heat balance equations are expressed as follows: ; in, Representing the The current temperature of each node. The heat capacity of this node. The thermal resistance between nodes, For internal heat source power, Heat is gained from solar radiation. The heating or cooling capacity of the HVAC system.

[0032] To achieve online self-evolution of the model, the system maps the visual data obtained in step S1 to the model parameters in real time: when an increase in population density and "vigorous movement" is detected in a certain area, the internal heat source term Q of that area is automatically mapped. internal,i The value was dynamically adjusted from a baseline of 60 W / m² to 150 W / m²; simultaneously, based on the clothing thickness coefficient K... do Correcting the body's heat dissipation efficiency. Heat gain from solar radiation, Q. solar,i Data can be collected in real time through outdoor weather stations or irradiance sensors, or predicted by a digital twin model based on building orientation, time, and weather conditions. Furthermore, the system employs an online incremental learning algorithm, collecting historical prediction errors at regular intervals (e.g., every 10 minutes). If the error exceeds a preset threshold (e.g., 0.5℃) multiple times consecutively, the parameter identification module is triggered. Using recursive least squares (RLS) or gradient descent, the thermal resistance matrix R and thermal capacity vector C in the model are fine-tuned to minimize the sum of squared residuals. This mechanism enables the model to automatically adapt to thermal performance drift caused by aging building envelopes, renovations, or seasonal changes, maintaining high-precision prediction capabilities over the long term without human intervention.

[0033] Step S3: Generation of optimal control action sequence based on deep reinforcement learning After obtaining accurate thermal environment predictions, the system sets the HVAC system as an agent and uses a deep reinforcement learning algorithm (selected from one or more combinations of Q-learning, Deep Q-Network (DQN), Proximal Policy Optimization (PPO), or Soft Actor Critic (SAC) algorithm) to generate the optimal control action sequence. The agent's state space is defined by the input vector. The system comprises multi-dimensional information including current indoor temperature, outdoor weather conditions, electricity price, population distribution, carbon emission factors, and time period. The action space outputs optimal control commands for a future preset time period, including pre-cooling / pre-heating and supply air temperature setpoint T. set Fresh air valve opening V new And the airflow setting of the variable air volume terminal (VAV).

[0034] To balance energy consumption, comfort, and social responsibility, the system constructs a multi-objective comprehensive reward function as follows: ; in, to These are the weighting coefficients for energy consumption costs, comfort deviation penalties, carbon emission costs, and grid demand response benefits, respectively. The initial values ​​can be set according to the actual scenario (e.g., 0.4, 0.3, 0.2, 0.1).

[0035] In actual decision-making, when it is detected that the electricity price is at its peak or the grid carbon factor is higher than the preset benchmark, the intelligent agent calls the digital twin model to predict the building's thermal inertia at future moments, and actively adjusts the indoor set temperature towards the upper or lower limit of the comfort range (for example, from 24℃ to 25.5℃ in summer), using the building's thermal inertia to reduce the output of mechanical cooling equipment, rather than taking a brutal shutdown strategy; conversely, when it is detected that the outdoor weather conditions are suitable and the outdoor enthalpy is lower than the indoor set enthalpy, the system increases the proportion of natural ventilation or free cooling mode.

[0036] Step S4: Dynamic microclimate zoning and differentiated control To further address the energy waste caused by the "blind control" of traditional control systems, the system implements dynamic microclimate zoning control based on multimodal visual perception data. Specifically, based on real-time personnel density distribution maps, the system uses the K-means clustering algorithm to dynamically divide the office area into three virtual microclimate zones: high-activity zones, sedentary work zones, and unoccupied zones. For high-activity zones (personnel density greater than 0.8 people / m² and with vigorous activity), the system determines adjustment parameters to increase the fresh air exchange rate and decrease the supply air temperature to quickly remove accumulated heat. For sedentary work zones, basic ventilation and constant temperature control parameters are maintained. For unoccupied zones, if continuous monitoring exceeds a preset duration (e.g., 30 minutes), the system determines to close the terminal valves in that area or switch to the minimum maintenance operation mode. This differentiated control based on the actual distribution of personnel achieves true on-demand energy supply.

[0037] Step S5: Execution, Monitoring, and Feedback Correction The system executes the optimal control sequence described above and enters the real-time monitoring and feedback correction phase. During execution, the system monitors the deviation between the actual thermal environment and the predicted values ​​of the digital twin model in real time. If the deviation exceeds a preset threshold, the model correction mechanism is immediately triggered, synchronously updating the thermal resistance and thermal capacity parameters of the RC network in step S2 and the weights of the deep reinforcement learning strategy network in step S3 to ensure the accuracy of the control strategy. In addition, the system also collects implicit user feedback data, including the adjustment frequency of the temperature control panel, the window opening status, and personal fan usage records. Based on this data, a user satisfaction index is calculated, and the weight coefficient w2 of the comfort deviation penalty term in step S3 is dynamically corrected accordingly. The corrected weights are then applied to the training of the strategy network in subsequent time steps, thereby achieving self-iteration and optimization of the control strategy. Example

[0038] This embodiment further illustrates the adaptive control strategy of the present invention under different seasonal operating conditions, specifically covering winter heating scenarios and transitional season fresh air optimization scenarios, to verify the robustness and energy-saving effect of the system under different environmental constraints.

[0039] Winter heating scenarios: In winter heating scenarios, the system uses the solar radiation heat gain curve predicted by the digital twin model and the heat dissipation power distribution of people, combined with the strategy of generating a deep reinforcement learning agent, to execute heat storage-heat release coordinated control based on building thermal inertia.

[0040] Specifically, when the system detects that electricity prices are at their lowest and outdoor temperatures are low, it proactively adjusts the indoor set temperature T. setThe system dynamically adjusts to the upper limit of the comfort zone (e.g., 26°C), utilizing the heat capacity of the building envelope (walls, floors) and furniture for sensible heat storage. As the daytime peak electricity consumption period approaches or outdoor temperatures rise, the system gradually reduces the output of heating equipment based on the model-predicted indoor temperature decay rate, instead relying on the latent heat released by the building to maintain room temperature within the comfort zone. Furthermore, for high-density areas such as large conference rooms, the system dynamically increases the output power of floor radiant heating or air curtains based on real-time visual perception data identifying the density of people, and employs localized air supply strategies to prevent localized overheating. In sparsely populated areas, the system automatically reduces heating output while maintaining the base temperature setpoint, achieving precise control of energy supply on demand.

[0041] Transitional Season Fresh Air Optimization Scenarios In the scenario of fresh air optimization during the transition season, the system prioritizes determining the outdoor air enthalpy value H. out Compared with the indoor set enthalpy value H in The difference. If we judge H... out <H in If the outdoor Air Quality Index (AQI) meets the preset threshold, the system activates a free fresh air cooling mode. At this time, the visual sensing module monitors the population density distribution in each area in real time, dynamically maximizing the opening of the fresh air valve to utilize natural cooling sources. If the outdoor AQI exceeds the safety threshold, the system immediately switches to internal circulation mode and calculates the minimum fresh air volume required to maintain indoor air quality based on the mapping relationship between CO2 concentration gradient and population density, dynamically adjusting the fresh air valve opening. This process minimizes heat load loss due to excessive fresh air while ensuring indoor air quality (IAQ).

[0042] Through the implementation of the two typical operating conditions described above, this invention not only solves the control blind spot problem caused by a single sensing dimension, but also proves that the system can automatically adjust the control strategy according to the dynamic changes in seasonal characteristics, meteorological conditions, and personnel activity patterns, effectively balancing energy consumption costs, comfort, and social carbon emission responsibility. Experimental data show that, compared with the traditional setpoint control strategy, this embodiment can further reduce the operating energy consumption of the HVAC system in winter heat storage and transitional season fresh air utilization scenarios, while maintaining stable user satisfaction indicators.

[0043] Of course, the above description is not limited to the examples above. Technical features not described in this invention can be implemented by or using existing technology, and will not be repeated here. The above embodiments and drawings are only used to illustrate the technical solutions of this invention and are not intended to limit this invention. This invention has been described in detail with reference to preferred embodiments. Those skilled in the art should understand that any changes, modifications, additions or substitutions made by those skilled in the art within the scope of this invention do not depart from the spirit of this invention and should also fall within the scope of protection of the claims of this invention.

Claims

1. An adaptive energy-saving operation method for an office building HVAC system, characterized in that: The method employs a layered computing architecture, comprising a perception layer, an edge computing layer, and a cloud training layer, and includes the following steps: S1, the perception layer collects multi-source heterogeneous data in real time. This data includes indoor and outdoor environmental physical parameters, HVAC equipment operating status parameters, electricity market signals, and multimodal visual perception data acquired through non-contact sensors. The multimodal visual perception data is preprocessed and locally anonymized at edge computing nodes, with only structured metadata extracted and uploaded to the cloud or central controller. This structured metadata includes personnel coordinates, number of people, activity postures, and clothing thickness coefficient K. do The edge computing layer deploys a lightweight digital twin thermal environment model inference engine and a deep reinforcement learning policy execution module to issue control commands in milliseconds; the cloud training layer is responsible for the offline pre-training of the model, incremental learning updates, and global optimization of the multi-objective reward function, and the layers interact with each other through a secure communication protocol. S2, a lightweight digital twin thermal environment model is constructed based on multi-source heterogeneous data. The building space is discretized into several thermal nodes, and a set of heat balance equations based on thermal resistance-heat capacity RC networks is established. The set of heat balance equations includes the internal heat source power Q at each thermal node. internal,i The heat gain from solar radiation, Q solar,i and the heating or cooling power Q of the HVAC system HVAC,i Included as an independent item; mapping the activity pose and clothing thickness coefficient in the structured metadata to the internal heat source item Q of the RC network. internal,i This allows the internal heat source term to be dynamically adjusted between a baseline value of 60 W / m² and a value of 150 W / m² for strenuous activity, depending on the activity level of people in the area, and is based on the clothing thickness coefficient K. do The heat dissipation power of the human body is corrected; an online incremental learning algorithm is used to dynamically update the thermal resistance matrix R and thermal capacity vector C of the RC network with the residual between the model prediction value and the actual measurement value at historical time as the optimization target. S3 sets the HVAC system as a deep reinforcement learning agent, constructs a four-objective comprehensive reward function including energy consumption cost, comfort deviation penalty, carbon emission cost, and grid demand response benefit. Combined with the prediction results of the digital twin thermal environment model, it outputs the optimal control action sequence including pre-cooling / pre-heating, supply air temperature setpoint, fresh air valve opening, and variable air volume terminal air speed setting. When it detects that the electricity price is at its peak or the grid carbon factor is higher than the preset benchmark, it calls the digital twin thermal environment model to predict the building's thermal inertia in the future period and fine-tunes the indoor set temperature towards the upper or lower limit of the comfort range, instead of executing shutdown or large temperature adjustment. S4, based on the real-time location, number and activity status of people in multimodal visual perception data, dynamically divides virtual microclimate zones in the digital space, and determines the wind direction, wind speed and valve opening adjustment parameters of the variable air volume terminal according to the load characteristics of each zone. S5 executes the optimal control action sequence, monitors the deviation between the actual thermal environment and the predicted value of the digital twin thermal environment model in real time. If the deviation exceeds the preset threshold, the model correction mechanism is triggered to synchronously update the thermal resistance and thermal capacity parameters of the RC network in S2 and the weights of the deep reinforcement learning strategy network in S3. Implicit user feedback data is also collected, including the adjustment frequency of the temperature control panel, the window opening status, and the personal fan usage record. The user satisfaction index is calculated based on the implicit user feedback data, and the weight coefficient of the comfort deviation penalty term in S3 is dynamically adjusted based on the satisfaction index.

2. The adaptive energy-saving operation method for office building HVAC systems according to claim 1, characterized in that: In step S2, the online incremental learning algorithm is either Recursive Least Squares (RLS) or Gradient Descent, and the historical prediction error Δe = |T| is collected at preset intervals. pred -T real | When the error exceeds the preset threshold of 0.5℃ multiple times, the parameter identification module is triggered to fine-tune the thermal resistance matrix R and the thermal capacity vector C to minimize the sum of squared residuals. When the deviation exceeds the threshold and triggers the model correction mechanism, the RC network parameters and the deep reinforcement learning strategy network weights are updated simultaneously.

3. The adaptive energy-saving operation method for office building HVAC systems according to claim 1, characterized in that: The specific method for dynamically dividing the virtual microclimate zones in step S4 is as follows: Based on the real-time personnel density distribution map, the office area is divided into three categories: high-activity zone, quiet office zone, and unmanned zone using the K-means clustering algorithm; for the high-activity zone, the number of fresh air exchanges is increased and the supply air temperature is reduced; for the quiet office zone, the basic ventilation volume and constant temperature control are maintained; for the unmanned zone, when the continuous detection time exceeds 30 minutes, the variable air volume terminal valve of the area is closed or the system is switched to the minimum maintenance operation mode.

4. The adaptive energy-saving operation method for office building HVAC systems according to claim 1, characterized in that: The four-objective comprehensive reward function in step S3 is expressed as follows: ; in, to These are the weighting coefficients for energy consumption costs, comfort deviation penalties, carbon emission costs, and grid demand response benefits, respectively.

5. The adaptive energy-saving operation method for office building HVAC systems according to claim 1, characterized in that: In step S3, when it is detected that the outdoor weather conditions are suitable and the outdoor enthalpy is lower than the indoor set enthalpy, the usage ratio of natural ventilation or free cooling mode is increased, and the output of mechanical refrigeration equipment is limited.

6. The adaptive energy-saving operation method for office building HVAC systems according to claim 1, characterized in that: It also includes the steps for optimizing fresh air during the transition season: determining the outdoor air enthalpy H. out Compared with the indoor set enthalpy value H in The size relationship, if H out <H in Furthermore, if the outdoor air quality index (AQI) meets the preset threshold, the free fresh air cooling mode will be activated and the fresh air valve opening will be dynamically maximized. If the outdoor AQI exceeds the standard, switch to internal circulation mode and calculate the minimum fresh air volume required to maintain indoor air quality based on the mapping relationship between CO2 concentration and personnel density, and dynamically adjust the opening of the fresh air valve.

7. The adaptive energy-saving operation method for office building HVAC systems according to claim 1, characterized in that: It also includes coordinated heat storage and release control in winter heating scenarios: during off-peak electricity periods and when outdoor temperatures are low, the indoor set temperature is dynamically increased to the upper limit of the comfort zone, utilizing the heat capacity of the building envelope and furniture for sensible heat storage; after entering the daytime peak electricity consumption period, the heating output is gradually reduced based on the indoor temperature decay rate predicted by the digital twin model, and the building releases stored heat to maintain room temperature; for high-density areas such as large conference rooms, the output power of floor radiant heating or hot air curtains is dynamically increased based on the visual perception of the number of people gathered, in conjunction with local air supply to avoid local overheating; in sparsely populated areas, the heating output is reduced while maintaining the basic temperature set value.

8. An apparatus for using the adaptive energy-saving operation method for an office building HVAC system according to any one of claims 1 to 7, characterized in that, include: The perception layer is used to collect multi-source heterogeneous data in real time, including indoor and outdoor environmental physical parameters, equipment operating status parameters, power market signals and multimodal visual perception data obtained by non-contact sensors. The raw visual data is locally desensitized at the edge computing node and only structured metadata is uploaded. The digital twin layer, deployed on edge computing nodes, is used to build a lightweight RC hot node network model and map the activity pose and clothing thickness coefficient in the structured metadata to the internal heat source term Q of the RC network. internal,i This allows the internal heat source term to be dynamically adjusted according to the activity status of people in the area, and an online incremental learning algorithm is used to dynamically update the thermal resistance matrix and heat capacity vector; the RC heat balance equations will Q internal,i Q solar,i and Q HVAC,i Included as a separate item; The decision layer, deployed on edge computing nodes, is used to build the HVAC system into a deep reinforcement learning agent. The agent's state space consists of indoor temperature, outdoor temperature, real-time electricity price, number of people, CO2 concentration, grid carbon factor, and time period. The optimization objective is a comprehensive reward function with four objectives, including energy consumption cost, comfort deviation penalty, carbon emission cost, and grid demand response benefit. Combined with the prediction results of the digital twin layer, it outputs the optimal control action sequence, including pre-cooling / pre-heating, supply air temperature setpoint, fresh air valve opening, and variable air volume terminal wind speed setting. In the grid demand response scenario, the digital twin layer is called to predict the building's thermal inertia to perform pre-cooling / pre-heating and comfort zone fine-tuning. The cloud-based training layer is used for offline pre-training, incremental learning updates, and global optimization of multi-objective reward functions for deep reinforcement learning agents. The execution and feedback layer, deployed on edge computing nodes, is used to execute the optimal control action sequence, monitor the deviation between the actual thermal environment and the predicted value, and synchronously update the RC network parameters and the deep reinforcement learning strategy network weights when the deviation exceeds the threshold. It also calculates the user satisfaction index based on implicit user feedback data and dynamically adjusts the weight coefficient of the comfort deviation penalty term based on the satisfaction index.