A subway station air conditioning system automatic control method based on reinforcement learning

By constructing a virtual sandbox using digital twin technology and reinforcement learning algorithms, the automated control of the subway station air conditioning system was achieved, solving the problems of high energy consumption and inaccurate control of the air conditioning system, reducing energy consumption and improving adaptability.

CN116857778BActive Publication Date: 2025-12-16CHINA RAILWAY FIRST SURVEY & DESIGN INST GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310793296.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-12-16
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

The air conditioning system in subway stations cannot achieve automatic adaptive control, resulting in problems such as substandard indoor temperature and humidity, uneven air supply, poor air quality, and high system energy consumption. Existing control methods cannot effectively reduce energy consumption.

Method used

A virtual sandbox is constructed using digital twin technology, and combined with reinforcement learning algorithms, enabling the computer to generate the optimal control scheme through self-learning, thereby achieving automated control of the air conditioning system.

Benefits of technology

It improves the adaptive control accuracy of air conditioning systems, reduces energy consumption, reduces the need for manual adjustments, and has generalization capabilities, making it suitable for different systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116857778B_ABST
    Figure CN116857778B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of metro station electromechanical system control, and particularly relates to a metro station air conditioning system automatic control method based on reinforcement learning. The present application takes metro digital twin as a virtual sand table, uses a reinforcement learning algorithm to enable a computer to deduce and generate an executable and operable optimal control scheme at low cost and high efficiency, effectively solves the high energy consumption problem of building air conditioning systems, and has important technical value and practical significance. The goal of automatically deducing and generating an executable and operable optimal control scheme is achieved, thereby reducing the number of repeated adjustments, effectively reducing the work difficulty and personnel cost of personnel, and solving the pain point problem of high energy consumption of metro station air conditioning systems. Meanwhile, the present application has a certain generalization ability and can be applied to other air conditioning system forms or other electromechanical systems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of metro station electromechanical system control, relates to digital twin building technology and optimization decision-making technology based on reinforcement learning, and in particular to an automatic control method for a metro station air conditioning system based on reinforcement learning. BACKGROUND

[0002] The air conditioning system of an underground station cannot achieve the requirements of automatic adaptability control. The operation and maintenance personnel adjust and control according to the control mode table, resulting in local problems such as substandard indoor temperature and humidity, uneven air supply, and poor indoor air quality, and systematic problems such as the self-control level not reaching the expected degree and wind and water imbalance. Not only does it not meet the air conditioning comfort requirements, but it also causes high system energy consumption. The fundamental reason is that the air conditioning system contains equipment and fluid pipe networks, has characteristics such as large lag, large inertia, and nonlinearity, and the air conditioning systems of different stations differ greatly in system form, topological structure, equipment type, and operation strategy. Most station air conditioning systems are difficult to control to the best state, resulting in a system in a sick state. According to statistics, if the air conditioning system is controlled in detail, energy saving can reach 30%-50%. Therefore, adaptive control is the key to whether the metro air conditioning system can achieve the design intent and meet the project requirements of the owner.

[0003] Traditional metro station control is mainly based on a control time mode table or uses a corresponding control system, but manual or previous control systems cannot automatically adjust according to load changes, resulting in energy waste. If computer automation adaptive control can be achieved, it will greatly reduce air conditioning system energy consumption. Compared with other digital twins and reinforcement learning air conditioning systems, this patent is based on a simplified physical model optimization. An accurate virtual sand table reflecting the global operating state of the air conditioning system is established, and then an artificial intelligence algorithm is used to let the computer continuously try and error and self-learn on the sand table, ultimately automatically generating a reasonable control scheme. Reinforcement learning algorithm is used for control. After the reinforcement learning model is trained, it is closer to reality, and the optimization result has higher precision. It does not depend on the theoretical knowledge and practical experience of air conditioning control operators themselves and can achieve automatic control for different systems.

[0004] Metro digital twin is a new technology that applies digital twin technology to metro air conditioning. Simply put, it is based on a physical model and uses various sensors to obtain data in all directions, and completes mapping in a virtual space to reflect the entire life cycle process of the corresponding physical air conditioning system. With the continuous improvement of infrastructure such as the Internet of Things and cloud computing, digital twin technology has achieved great success in various fields, greatly promoting the transformation of product design, production, operation, and maintenance, etc.

[0005] Reinforcement learning is one of the paradigms and methodologies of the new generation of artificial intelligence, which is used to describe and solve the problem that an agent learns a strategy through interaction with the environment to maximize the reward or achieve a specific goal. Some complex reinforcement learning algorithms have universal intelligence to a certain extent to solve complex problems, and can reach the level of human beings in chess and electronic games. Therefore, reinforcement learning can enable computers to constantly try and error on the sand table, and ultimately realize the self-learning control ability of the system, which is one of the feasible solutions to solve the energy consumption problem of complex air conditioning system control. SUMMARY

[0006] The present application aims to propose an automatic control method for subway station air conditioning system based on reinforcement learning, taking subway digital twin as a virtual sand table, using reinforcement learning algorithm to enable computers to deduce and generate executable and operable optimal control scheme at low cost and high efficiency, effectively solving the high energy consumption problem of building air conditioning system, and having important technical value and practical significance.

[0007] To achieve the above purpose, the present application proposes an automatic control method for subway station air conditioning system based on reinforcement learning, comprising the following steps:

[0008] An automatic control method for subway station air conditioning system based on reinforcement learning, comprising the following steps:

[0009] Step 1: Based on digital twin technology, construct multi-physical field simulation module, digital modeling module, data analysis module and machine learning module to realize digital twin model accurately reflecting the global running state of air conditioning system;

[0010] Step 2: Set up a control agent, build an interactive environment of the agent and the air conditioning system twin model, and design the corresponding environment state, the agent's selectable action and the reward function of the agent by the environment;

[0011] Step 3: According to DoubleQ-learning reinforcement learning algorithm, establish the initial decision model of the control agent, including action selection network, action evaluation network and corresponding value function;

[0012] Step 4: The control agent and the environment interact to realize self-learning: data is obtained from the experience pool, and the time difference algorithm is used to update the action selection network and the action evaluation network;

[0013] Step 5: For various problems existing in the air conditioning system, repeat steps 3 and 4, so that the control agent constantly tries different control strategies, accumulates data and continuously updates the action selection network and the action evaluation network, and finally completes the training of the agent's decision model;

[0014] Step 6: Based on the trained agent decision-making model, given the control task and the current state of the air conditioning system, the control agent automatically generates a control scheme to guide the actual staff to deploy the control strategy into the actual system.

[0015] Furthermore, the digital twin model constructed in step 1 includes: a multiphysics simulation module, a digital modeling module, a data analysis module, and a machine learning module.

[0016] Furthermore, the environmental state s in step 2 t Including but not limited to: the indoor temperature T in each area i of the subway station at time t. i,t CO2 concentration in each region i i,t Outdoor temperature OT t Passenger flow intensity R t Wind speed v t Building orientation L.

[0017] Furthermore, in step 2, the agent's action a at time t... t Including but not limited to: adjusting the air supply volume G in various areas i of the subway station i,t Adjust the fresh air volume M in each zone i i,t Adjusting the opening degree V of the air valve that affects the air supply volume in each area i,t Optimize the supply air temperature setpoint T set,t Optimize the chilled water supply temperature setpoint CHW set,t Optimize the cooling tower outlet water temperature setpoint CW set,t ;

[0018] The selection range for the air supply volume of the air conditioning system in zone i at time t is: and These are the minimum and maximum air supply volumes for region i, respectively.

[0019] The selection range for the fresh air volume of the air conditioning system in zone i at time t is: and M i max These are the minimum and maximum fresh air volumes for region i, respectively.

[0020] The selection range for the opening degree of the air valve in region i at time t is: 0 ≤ V i,t The range for selecting the supply air temperature setpoint at time t is ≤1. and These are the lower and upper limits of the supply air temperature setpoint for zone i, respectively. The selection range of the chilled water supply temperature setpoint at time t is: and These are the lower and upper limits of the chilled water supply temperature setpoint, respectively. The selection range of the cooling tower outlet water temperature setpoint at time t is: and are the lower and upper limits of the cooling tower outlet water temperature set value, respectively;

[0021] The station energy load data is obtained through a station intelligent control system.

[0022] Further, the component terms of the reward function in step 2 include but are not limited to an energy consumption term cost() and a comfort penalty term penalty(). t The sum of the device energy consumptions of the post-air-conditioning system in the new state s t+1 , including the chiller, chilled water pump, cooling water pump, cooling tower, air handling unit, the comfort penalty term includes but is not limited to a temperature penalty term, a humidity penalty term, and an air quality penalty term, and the calculation formula of the reward function is:

[0023] r t =-c0costa t ,s t+1 -c1penalty 1 s t+1 -…-c n penalty n s t+1

[0024] wherein c0, c1, …, c n are weight coefficients of the respective terms.

[0025] Further, the action selection network value function in step 3 is Qs t ,a t ; θ, for selecting an action to be taken in the current state; and the action evaluation network value function is Q's t ,a t ; θ', for evaluating the cumulative reward brought by the action.

[0026] wherein θ and θ' are parameters of the two networks, respectively, and the two networks have the same structure but different parameters.

[0027] Further, the interaction between the control agent and the environment in step 4 includes the following steps:

[0028] 4.1: The agent first obtains the environment state s t at time t, then outputs an action a t using the action selection network and controls the air conditioning system, and then the environment feeds back the reward r t corresponding to the action and the state s t+1 at the next time, and the agent stores the four-tuple s t ,a t ,r ts t+1 store into experience pool;

[0029] 4.2: get quadruple from experience pool, update action selection network value function Q using time difference algorithm, specific update formula is:

[0030] Qs t ,a t ; theta <- Qs t ,a t ; theta + alpha r t + gamma Q's t+1 ,a * ; theta - Qs t ,a t ; theta

[0031]

[0032] Similarly, the update formula of the action evaluation network is:

[0033] Q's t ,a t ; theta' <- Q's t ,a t ; theta' + alpha r t + gamma Qs t+1 ,a * ; theta - Q's t ,a t ; theta'

[0034]

[0035] Further, the various problems existing in the air conditioning system in step 4 include but are not limited to: indoor temperature not up to standard, hydraulic imbalance, wind imbalance, uneven cooling and heating, uneven air supply and indoor air quality not up to standard.

[0036] Further, the standard for judging whether the agent decision model is trained in step 4 is that the agent successfully completes various control tasks or reaches a predetermined number of training times.

[0037] Compared with the prior art, the subway station air conditioning system automatic control method based on reinforcement learning has the advantages that: compared with the traditional control method, the method is on the virtual control sand table constructed by the subway digital twin, and uses the reinforcement learning algorithm to make the computer interact with the environment, continuously learns and updates itself, and finally forms an intelligent agent that can solve various control tasks. After the reinforcement learning model is trained, it is closer to the actual situation, the optimization result has higher precision, and it does not depend on the theoretical knowledge and practical experience of the air conditioning control personnel itself, and can realize automatic control for different systems. And automatically deduce the target of the executable and operable optimal control scheme, thereby reducing the number of repeated adjustments, effectively reducing the work difficulty and personnel cost of personnel, solving the pain point problem of high energy consumption of the subway station air conditioning system. At the same time, the present application has a certain generalization ability and can be applied to other air conditioning system forms or other mechanical and electrical systems. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The flowchart of the automatic control method provided by the present application. DETAILED DESCRIPTION

[0039] The present application will be further described below through specific examples and drawings. The embodiments of the present application are to enable those skilled in the art to better understand the present application and do not limit the present application in any way.

[0040] As shown in the figure, the present application provides a subway station air conditioning system automatic control method based on reinforcement learning, which includes control agent training (steps 1-5) and control agent application (step 6) two stages, and specifically includes the following six steps: Figure 1 Step 1: based on the digital twin technology, construct a digital twin model that accurately reflects the global running state of the air conditioning system as a virtual sand table for trial and error and learning;

[0041] Step 2: set up a control agent, build an interactive environment of the agent and the air conditioning system twin model, and design the corresponding environment state, the agent's selectable action and the agent's reward function;

[0042] Step 3: according to the DoubleQ-learning reinforcement learning algorithm, establish an initial decision model of the control agent, including an action selection network, an action evaluation network and a corresponding value function;

[0043]

[0044] ​Step 4: Control the agent and the environment to interact, constantly update the action selection network and the action evaluation network; control the agent to obtain the current environment state, output the action to be taken by using the action selection network, control the air conditioning system according to the action, obtain the new environment state and the reward feedback of the action, send the current environment state, the action, the new environment state and the reward feedback to the experience pool, and update the action selection network and the action evaluation network using the time difference algorithm;

[0045] Step 5: For various problems existing in the air conditioning system, repeat step 3, so that the control agent constantly tries different control strategies, accumulates data and continuously updates the action selection network and the action evaluation network, and finally completes the training of the agent decision model; based on the trained agent decision model, given the control task, input the current state of the air conditioning system, and the control agent automatically generates a control scheme.

[0046] Taking the full-air air conditioning system of a subway station as an example, the environment state in step 2 above is: t =(T i,t ,CO i,t ,OT t ,R t ,t) , where T i,t is the indoor temperature of region i at time t, CO i,t is the CO2 concentration of region i at time t, OT t is the outdoor temperature at time t, R t is the passenger flow intensity at time t, and t is the time.

[0047] The agent action in step 2 above is: t =(G i,t ,M i,t ,V i,t ,T set,t ) , where G i,t is the supply air volume of region i at time t, M i,t is the fresh air volume of region i at time t, V i,t is the corresponding air valve opening degree affecting the supply air volume of region i at time t, and T set,t is the supply air temperature set value at time t. The selection range of the supply air volume of air conditioning system region i is: and are the minimum supply air volume and the maximum supply air volume of region i, respectively. The selection range of the fresh air volume of air conditioning system region i is: and are the minimum fresh air volume and the maximum fresh air volume of region i, respectively. The selection range of the air valve opening degree of region i is: 0≤V i,t ≤1. The selection range of the supply air temperature set value is: and These are the lower and upper limits of the air supply temperature setting for zone i, respectively.

[0048] To reduce the energy consumption of the air conditioning system while maintaining indoor temperature and carbon dioxide concentration within a comfortable range, the reward function in step 2 above consists of three parts. The first part controls the agent to perform action a. t The rear air conditioning system is in a new state. t+1 The first term is the sum of equipment energy consumption (including chillers, chilled water pumps, cooling water pumps, cooling towers, air handling units, etc.). The second term is the penalty term when the indoor temperature in each area exceeds the comfort range. The third term is the penalty term when the CO2 concentration in each area exceeds the limit value. The formula for calculating the reward function is:

[0049]

[0050]

[0051]

[0052] In the formula, c1, c2, and c2 are the weight coefficients of the three terms, respectively. This represents the upper limit of the indoor thermal comfort temperature for region i. T i This represents the lower limit of the indoor thermal comfort temperature for region i. This is the limit value for CO2 concentration in region i.

[0053] The action selection network value function Q(s) in step 3 above t ,a t ;θ) is used to select the action to be taken in the current state, and the action evaluation network value function Q'(s) t ,a t θ') is used to evaluate the cumulative reward brought by the action, where θ and θ' are the parameters of the two networks, which have the same structure but different parameters.

[0054] The interaction process in step 4 above is as follows: the control agent acquires the current environmental state, uses the action selection network to output the action to be taken, controls the air conditioning system according to the action, obtains the new environmental state and the reward feedback for the action, and simultaneously sends the current environmental state, action, new environmental state, and reward feedback to the experience pool, and uses the temporal difference algorithm to update the action selection network and action evaluation network.

[0055] In step 4 above, when controlling the interaction between the agent and the environment, the agent first obtains the environmental state s at time t. t Then, the action selection network is used to output action a. t It also controls the air conditioning system, and then the environment provides a reward r corresponding to this action. t and the state s at the next moment t+1At the same time, the agent stores the quadruple <s t ,a t ,r t ,s t+1 > into the experience pool.

[0056] The quadruple is obtained from the experience pool, and the action selection network value function Q is updated using the time difference algorithm, and the specific update formula is:

[0057] Q(s t ,a t ; θ) ← Q(s t ,a t ; θ) + α(r t + γQ(s t+1 ,a * ; θ) - Q(s t ,a t ; θ)

[0058]

[0059] Similarly, the update formula of the action evaluation network is:

[0060] Q(s t ,a t ; θ) ← Q(s t ,a t ; θ) + α(r t + γQ(s t+1 ,a * ; θ) - Q(s t ,a t ; θ)

[0061]

[0062] The various problems existing in the air conditioning system in the above step 4 include but are not limited to: indoor temperature not up to standard, water imbalance, wind imbalance, uneven cooling and heating, uneven air supply, and indoor air quality not up to standard.

[0063] The standard for determining whether the agent decision model is trained in the above step 4 is that the agent successfully completes various control tasks such as indoor temperature not up to standard, water imbalance, wind imbalance, or reaches a predetermined number of training times.

[0064] The above-described embodiments are only a preferred scheme of the present application, and are not intended to limit the present application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present application.

Claims

1. A subway station air conditioning system automatic control method based on reinforcement learning, characterized in that, Comprising the following steps: Step 1: Based on digital twin technology, build multi-physics simulation module, digital modeling module, data analysis module, machine learning module, realize the digital twin model that accurately reflects the global operation state of the air conditioning system; Step 2: Set up control agent, build the interaction environment of the agent and the air conditioning system twin model, and design the corresponding environment state, agent's available actions and the reward function of the agent by the environment; Step 3: According to the Double Q-learning reinforcement learning algorithm, establish the initial decision model of the control agent, including action selection network, action evaluation network and corresponding value function; Step 4: The control agent and the environment interact to realize self-learning: obtain data from the experience pool, and update the action selection network and the action evaluation network using the time difference algorithm; Step 5: For various problems existing in the air conditioning system, repeat steps 3 and 4 to make the control agent constantly try different control strategies, accumulate data and continuously update the action selection network and the action evaluation network, and finally complete the training of the agent decision model; Step 6: Based on the trained agent decision model, give the control task, input the current air conditioning system state, and the control agent automatically generates the control scheme to guide the actual staff to deploy the control strategy to the actual system.

2. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 1, wherein: The digital twin model built in step 1 includes: multi-physics simulation module, digital modeling module, data analysis module, machine learning module.

3. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 2, wherein: The environmental state s in step 2 t including but not limited to: the indoor temperature T of each area i of the subway station at time t i,t , the CO2 concentration CO of each area i i,t , the outdoor temperature OT t , the passenger flow intensity R t , the wind speed v t , the building orientation L.

4. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 2 or 3, wherein: The step 2 intelligent agent action a at time t t Including but not limited to: adjusting the air supply amount G of each area i of the subway station i,t Adjusting the fresh air amount M of each area i i,t Adjusting the air valve opening V affecting the air supply amount of each area i i,t Optimizing the air supply temperature set value T set,t Optimizing the chilled water supply temperature set value CHW set,t Optimizing the cooling tower outlet water temperature set value CW set,t ; The selection range of the air supply amount of the air conditioning system at time t in region i is: and are the minimum air supply amount and the maximum air supply amount of region i, respectively. The selection range of the fresh air volume of the air conditioning system at time t in region i is: and are the minimum fresh air volume and the maximum fresh air volume of region i, respectively. The selection range of the air valve opening of region i at time t is: 0≤V i,t and are the lower limit and the upper limit of the supply air temperature set value of region i respectively, the selection range of the chilled water supply temperature set value at time t is: and are the lower limit and the upper limit of the chilled water supply temperature set value respectively, the selection range of the cooling tower outlet water temperature set value at time t is: and are the lower limit and the upper limit of the cooling tower outlet water temperature set value respectively;​ The station energy load data is obtained through the station intelligent control system.

5. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 4, wherein: The component items of the reward function in step 2 include but are not limited to an energy consumption item cost(), a comfort penalty item penalty(); the energy consumption item is for the intelligent agent to execute an action a t The sum of the equipment energy consumptions of the post air conditioning system in the new state s includes a chiller, a chilled water pump, a cooling water pump, a cooling tower and an air handling unit t+1 The comfort penalty item includes but is not limited to a temperature penalty item, a humidity penalty item and an air quality penalty item, and the calculation formula of the reward function is: r t = -c0cost(a t ,s t+1 )-c1penalty 1 (s t+1 )-…-c n penalty n (s t+1 ) where c0, c1,..., c n are the weight coefficients of each term, respectively.

6. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 5, wherein: The action selection network value function in step 3 is Q(s t ,a t ; θ) for selecting an action to take in the current state; and the action evaluation network value function is Q'(s t ,a t ; θ') for evaluating the cumulative reward resulting from the action. Where θ and θ' are the parameters of the two networks respectively, the structures of the two networks are the same, and the parameters are different.

7. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 6, wherein: The interaction between the control agent and the environment in step 4 includes the following steps: 4.1: The agent first obtains the environment state s at time t t Then it outputs an action a using the action selection network t And controls the air conditioning system. Then the environment will feedback the reward r corresponding to the action t And the state s at the next time t+1 At the same time, the agent stores the quadruple <s t ,a t ,r t ,s t+1 > into the experience pool; 4.2: Obtain the four-tuple from the experience pool, and update the action selection network value function Q using the time difference algorithm, the specific update formula is: Q(s t ,a t ; θ) ← Q(s t ,a t ; θ) + α(r t + γQ'(s t+1 ,a * ; θ') - Q(s t ,a t ; θ)) Similarly, the update formula of the action evaluation network is: Q'(s t ,a t ; θ) ← Q'(s t ,a t ; θ) + α(r t + γ Q(s t+1 ,a * ; θ) - Q'(s t ,a t ; θ)) 8. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 7, wherein: The various problems existing in the air conditioning system in step 4 include but are not limited to: indoor temperature not meeting the standard, hydraulic imbalance, wind imbalance, uneven cooling and heating, uneven air supply and unqualified indoor air quality.

9. The automatic control method of the subway station air conditioning system based on reinforcement learning according to claim 8, wherein: The criterion for judging whether the intelligent agent decision model is trained in step 4 is that the intelligent agent successfully completes various control tasks or reaches a predetermined number of training times.

Citation Information

Patent Citations

  • Incomplete information game strategy optimization method based on double deep Q network learning

    CN114089627A

  • Direct expansion air conditioning system suitable for subway station and control method of direct expansion air conditioning system

    CN114992729A