A Feng Shui linkage control method and system based on reinforcement learning
Through the Feng Shui linkage control method based on reinforcement learning, the adaptability and energy efficiency issues of the subway station air-conditioning system in the face of sudden changes were solved, the system's flexible, energy-saving and comfortable operation was achieved, and the overall control efficiency and passenger satisfaction were improved.
Patent Information
- Application Number
- CN202510969383.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-15
AI Technical Summary
The air-conditioning systems in subway stations lack adaptability when faced with sudden changes in passenger flow or extreme weather, have low energy efficiency, insufficient comfort, limited data processing capabilities, poor overall coordination, and backward algorithms, leading to energy waste and passenger dissatisfaction.
A feng shui linkage control method based on reinforcement learning is adopted. By constructing a feng shui linkage control model, using the MPO-MADDPG algorithm and continuous learning algorithm, combining intelligent agents, value networks, strategy networks and hierarchical experience replay pools, the coordinated control and real-time optimization of the feng shui system are achieved.
It improves the flexibility and adaptability of the air-conditioning system, reduces high-energy consumption operation, maintains temperature stability, enhances passenger comfort, improves the intelligence level and system efficiency, and avoids frequent starts and stops.
Smart Images

Figure CN120466801B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of subway station air-conditioning system control, and specifically relates to a feng shui linkage control method and system based on reinforcement learning. Background Art
[0002] Subway station air conditioning systems typically consist of air and water systems. The air system controls the air quality and temperature within the subway station by regulating air volume and supply air temperature through modular air conditioning units. The water system produces chilled water through chillers and other equipment, and influences the air supply temperature by regulating water flow and temperature. With the rapid development of urban rail transit, energy consumption in subway station air conditioning systems has become increasingly prominent. Therefore, efficient and accurate control of subway station ventilation and air conditioning systems to achieve energy-saving and environmentally friendly use is a key development direction in this field.
[0003] In the field of subway station air conditioning system control technology, although some technologies and methods have been applied, they still have many defects, as follows:
[0004] 1) Lack of adaptability: Traditional subway air conditioning control systems typically use fixed schedules or preset temperature ranges for adjustment. This static control method is not flexible and adaptable enough to deal with sudden changes in passenger flow or extreme weather conditions.
[0005] 2) Inefficient energy consumption: Traditional control methods, which fail to take into account real-time environmental and load conditions, can cause air conditioning systems to operate at high energy consumption levels for extended periods, resulting in wasted energy and unnecessary costs.
[0006] 3) Insufficient comfort: Passenger experience in subway stations may be affected by temperature fluctuations, especially between peak and off-peak hours, which may lead to increased passenger dissatisfaction and complaints;
[0007] 4) Limited data processing capabilities: Faced with massive amounts of sensor data and complex interactions, traditional analysis methods struggle to respond quickly and accurately, limiting the system's intelligence and decision-making speed.
[0008] 5) Poor overall coordination: Existing technologies do not consider the interplay and collaboration between the various components of a subway air conditioning control system, resulting in a lack of overall coordination.
[0009] 6) Outdated algorithms: The algorithms used in existing technologies require a large amount of historical data and cannot be used in scenarios where historical data is difficult to obtain, such as new projects. Due to the time lag in the control of subway station air-conditioning systems, the existing algorithms converge too slowly and are prone to control instability during the convergence and control process. Summary of the Invention
[0010] In order to solve the problems of lack of adaptability, low energy efficiency, insufficient comfort, limited data processing capability, poor overall coordination and backward algorithm in the existing technology, the purpose of the present invention is to provide a Feng Shui linkage control method and system based on reinforcement learning.
[0011] The technical solution adopted in the present invention is:
[0012] A feng shui linkage control method based on reinforcement learning comprises the following steps:
[0013] Based on several historical subway station air conditioning system monitoring data, a reinforcement learning algorithm was used to build a Feng Shui linkage control model for subway station air conditioning systems.
[0014] Based on the collected real-time subway station air conditioning system monitoring data, the wind and water linkage control model is used to perform wind and water linkage control and obtain the real-time wind and water linkage control strategy of the subway station air conditioning system;
[0015] A continuous learning algorithm is used to update the Feng Shui linkage control model, and Feng Shui linkage control is performed on the subway station air-conditioning system based on the real-time Feng Shui linkage control strategy.
[0016] Furthermore, based on several historical subway station air conditioning system monitoring data, a reinforcement learning algorithm was used to construct a feng shui linkage control model for subway station air conditioning systems, including the following steps:
[0017] Collecting and preprocessing a number of historical subway station air conditioning system monitoring data to obtain a number of preprocessed historical subway station air conditioning system monitoring data;
[0018] Use reinforcement learning algorithms to build an initial Feng Shui linkage control model for subway station air conditioning systems;
[0019] Based on several pre-processed historical subway station air-conditioning system monitoring data, the initial Feng Shui linkage control model is trained to obtain the final Feng Shui linkage control model.
[0020] Furthermore, the Feng Shui linkage control model is constructed based on the MPO-MADDPG algorithm, and the Feng Shui linkage control model includes a meta-strategy optimization module constructed based on the MPO algorithm and a reinforcement learning module constructed based on the MADDPG algorithm, which are connected in sequence. The reinforcement learning module is provided with an intelligent agent, a value network, a policy network and a hierarchical experience replay pool. The intelligent agent is respectively connected to the value network, the policy network, the hierarchical experience replay pool and the meta-strategy optimization module.
[0021] Furthermore, the intelligent agents include a water system intelligent agent, an A-side large system intelligent agent, and a B-side large system intelligent agent, and the water system intelligent agent, the A-side large system intelligent agent, and the B-side large system intelligent agent are respectively connected to the value network, the policy network, the hierarchical experience replay pool, and the meta-policy optimization module;
[0022] The strategy network includes a first current strategy network, a first target strategy network, a second current strategy network, a second target strategy network, a third current strategy network, and a third target strategy network. The first current strategy network and the first target strategy network are both connected to the water system agent, the second current strategy network and the second target strategy network are both connected to the A-end large system agent, and the third current strategy network and the third target strategy network are both connected to the B-end large system agent.
[0023] The value network includes the first current value network, the first target value network, the second current value network, the second target value network, the third current value network and the third target value network. The first current value network and the first target value network are both connected to the water system intelligent entity, the second current value network and the second target value network are both connected to the A-end large system intelligent entity, and the third current value network and the third target value network are both connected to the B-end large system intelligent entity.
[0024] Furthermore, a reinforcement learning algorithm was used to construct an initial Feng Shui linkage control model for the subway station air conditioning system, which included the following steps:
[0025] Use the MPO algorithm to build the initial meta-strategy optimization module, and use the optimization target generated by the Feng Shui linkage control strategy as the meta-strategy optimization scenario;
[0026] Using the MADDPG algorithm, we built an initial reinforcement learning module, set up a layered experience replay pool, and used the Feng Shui linkage control strategy generation problem as the simulation environment.
[0027] Set the objective function set, reward function, action space, and state space for the agent in the initial reinforcement learning module, set the action-value function for the value network, and set the policy function for the policy network;
[0028] According to the elastic weight connection mechanism of the continuous learning algorithm, the loss function of the reinforcement learning module is adjusted to obtain the adjusted loss function;
[0029] By integrating the initial meta-strategy optimization module and the initial reinforcement learning module, the initial Feng Shui linkage control model of the subway station air-conditioning system is obtained.
[0030] Furthermore, the historical monitoring data of the subway station air-conditioning system includes the historical outdoor temperature of the subway station air-conditioning system, the historical outdoor humidity, the historical subway station passenger flow factor, the historical number of cold source open sets, the historical chiller outlet water temperature, the historical refrigeration pump frequency, the historical cooling pump frequency, the historical cooling tower fan frequency, the historical A-end platform average temperature, the historical A-end group air-water valve opening, the historical A-end group air supply fan frequency, the historical B-end platform average temperature, the historical B-end group air-water valve opening, the historical B-end group air supply fan frequency, and the historical system total energy efficiency ratio;
[0031] The real-time subway station air-conditioning system monitoring data includes the real-time outdoor temperature of the subway station air-conditioning system, the real-time outdoor humidity, the real-time subway station passenger flow factor, the real-time number of cold source open sets, the real-time chiller outlet water temperature, the real-time refrigeration pump frequency, the real-time cooling pump frequency, the real-time cooling tower fan frequency, the real-time A-end platform average temperature, the real-time A-end group air-water valve opening, the real-time A-end group air supply fan frequency, the real-time B-end platform average temperature, the real-time B-end group air-water valve opening, the real-time B-end group air supply fan frequency and the real-time system total energy efficiency ratio.
[0032] Furthermore, the initial Feng Shui linkage control model is trained based on several pre-processed historical subway station air conditioning system monitoring data to obtain the final Feng Shui linkage control model, including the following steps:
[0033] Based on several pre-processed historical subway station air conditioning system monitoring data, the initial meta-strategy optimization module of the initial Feng Shui linkage control model is trained in different scenarios to obtain the final meta-strategy optimization module;
[0034] Use the final meta-strategy optimization module to initialize the policy network and value network of the initial reinforcement learning module in different scenarios to obtain the initial policy network and initial value network;
[0035] Traverse all the objective functions in the objective function set and train the initial agent of the initial reinforcement learning module in different scenarios based on some pre-processed historical subway station air conditioning system monitoring data;
[0036] During the training process, the back propagation method is used to optimize the value network parameters of the initial value network, and the gradient ascent method is used to optimize the policy network parameters of the initial policy network;
[0037] Using the adjusted loss function, obtain the first loss value in the training process until the first loss value is less than the loss value threshold or the first iteration number reaches the number threshold, and obtain the final reinforcement learning module;
[0038] Collect several historical Feng Shui linkage control experiences generated during the training process of the reinforcement learning module, and store the several historical Feng Shui linkage control experiences in a layered experience revisit pool.
[0039] Furthermore, based on the collected real-time subway station air conditioning system monitoring data, a Feng Shui linkage control model is used to perform Feng Shui linkage control, and a real-time Feng Shui linkage control strategy for the subway station air conditioning system is obtained, which includes the following steps:
[0040] Collecting real-time subway station air conditioning system monitoring data, preprocessing it to obtain preprocessed real-time subway station air conditioning system monitoring data, extracting real-time data features of the preprocessed real-time subway station air conditioning system monitoring data, and inputting the real-time data features into the Feng Shui linkage control model;
[0041] According to the real-time data characteristics, the meta-strategy optimization module of the Feng Shui linkage control model is used to adjust the policy network and value network of the reinforcement learning module of the Feng Shui linkage control model to obtain the adjusted policy network and adjusted value network;
[0042] Randomly extract a number of historical Feng Shui linkage control experiences from the stratified experience revisit pool, and update the action space of the adjusted strategy network based on the historical Feng Shui linkage control experiences to obtain the updated action space;
[0043] According to the real-time data characteristics, the state space of the adjusted policy network is updated to obtain the updated state space, and the updated state space is input into the adjusted value network;
[0044] Use the agent to control the adjusted value network, obtain the reward value of each possible action in the updated action space according to the reward function, and update the Q value of the possible action according to the reward value to obtain the updated Q value of each possible action;
[0045] The intelligent agent is used to control the adjusted policy network. According to the updated Q value of each possible action in the updated state space, the probability distribution of the executed action is output. Based on the probability distribution of the executed action, the real-time Feng Shui linkage control strategy of the subway station air-conditioning system is obtained.
[0046] Furthermore, a continuous learning algorithm is used to update the Feng Shui linkage control model, and Feng Shui linkage control is performed on the subway station air conditioning system based on the real-time Feng Shui linkage control strategy, including the following steps:
[0047] Collect the real-time feng shui linkage control experience corresponding to the real-time feng shui linkage control strategy generated by the feng shui linkage control model;
[0048] Mixing the real-time Feng Shui linkage control experience with a number of historical Feng Shui linkage control experiences randomly selected from the stratified experience revisit pool to obtain a number of mixed Feng Shui linkage control experiences;
[0049] Based on several hybrid Feng Shui linkage control experiences, the Feng Shui linkage control model is continuously trained, and the adjusted loss function is used to obtain the second loss value in the continuous training process until the second loss value is less than the loss value threshold or the second iteration number reaches the number threshold, thereby obtaining an updated Feng Shui linkage control model;
[0050] According to the real-time Feng Shui linkage control strategy, the real-time Feng Shui linkage control instructions of the subway station air-conditioning system are generated, and the real-time Feng Shui linkage control instructions are executed to perform Feng Shui linkage control on the subway station air-conditioning system.
[0051] A feng shui linkage control system based on reinforcement learning is used to implement a feng shui linkage control method, comprising a model building unit, a feng shui linkage control unit and a continuous learning unit connected in sequence.
[0052] The beneficial effects of the present invention are:
[0053] The present invention provides a Feng Shui linkage control method and system based on reinforcement learning. By introducing reinforcement learning and continuous learning algorithms, the control strategy can be dynamically adjusted according to real-time environmental changes and passenger flow, thereby improving flexibility and adaptability. By utilizing real-time data and historical experience, the operating status of the air-conditioning system can be more accurately predicted and adjusted, avoiding long-term high-energy consumption operation and achieving energy conservation and emission reduction. Through refined control and real-time adjustment, the temperature stability in the subway station can be maintained, and the comfort and satisfaction of passengers can be improved. With the help of advanced intelligent algorithm models, massive sensor data can be efficiently processed, and responses can be made quickly, thereby improving the level of intelligence. By comprehensively considering the interaction between the A-end large system and the B-end large system of the water system and the wind system, overall coordinated control can be achieved, and the efficiency and stability of the entire air-conditioning system can be improved. The online reinforcement learning method used pre-trains the strategy function of each intelligent agent through a small amount of expert experience imitation learning, so that the reinforcement learning algorithm also has a stable and reliable control result output in the early stage of interaction with the actual system. The improved experience replay mechanism and reward function make the control output as stable and reliable as possible during the convergence process, avoiding frequent starts and stops.
[0054] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flow chart of the Feng Shui linkage control method based on reinforcement learning in the present invention.
[0056] Figure 2 It is a structural block diagram of the Feng Shui linkage control system based on reinforcement learning in the present invention. DETAILED DESCRIPTION
[0057] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0058] Example 1:
[0059] like Figure 1 As shown, this embodiment provides a Feng Shui linkage control method based on reinforcement learning, comprising the following steps:
[0060] S1: Based on several historical subway station air conditioning system monitoring data, a reinforcement learning algorithm is used to construct a Feng Shui linkage control model for subway station air conditioning systems, including the following steps:
[0061] S1-1: Collect and pre-process a number of historical subway station air conditioning system monitoring data to obtain a number of pre-processed historical subway station air conditioning system monitoring data;
[0062] Historical subway station air conditioning system monitoring data includes the historical outdoor temperature of the subway station air conditioning system T out , historical outdoor humidity RH out , Historical subway station passenger flow factors PPF , Historical number of cold source openings Sets , Historical chiller outlet water temperature T ch , Historical refrigeration pump frequency F chp , Historical cooling pump frequency F cwp , Historical cooling tower fan frequency F ct , Historical average temperature of the A-end platform T a , Historical A-end group air and water valve opening V a , Historical A-end air supply fan frequency F a , Historical average temperature of the B-end platform T b , Historical B-end group air and water valve opening V b , Historical B-end air supply fan frequency F b and historical system total energy efficiency ratio EER ;
[0063] Preprocessing includes data cleaning, Gaussian denoising, upper and lower limit outlier screening, and normalization to improve data quality and provide data support for subsequent model training;
[0064] S1-2: Use reinforcement learning algorithms to build an initial Feng Shui linkage control model for the subway station air conditioning system;
[0065] The Feng Shui linkage control model is based on the Meta-Policy Optimization (MPO)-Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. It includes a sequentially connected meta-policy optimization module based on the MPO algorithm and a reinforcement learning module based on the MADDPG algorithm. The reinforcement learning module is equipped with an agent, a value network, a policy network, and a hierarchical experience replay pool. The agent is connected to the value network, policy network, hierarchical experience replay pool, and meta-policy optimization module respectively.
[0066] The meta-strategy optimization module is used to strengthen the network parameters of the value network and policy network in the learning module so that these parameters can quickly adapt to new and unseen data analysis results, thereby improving the generalization ability of the model. Even under unseen data analysis results, the value network and policy network can be updated based on previous learning experience, thereby improving the adaptability of the generation of Feng Shui linkage control strategies; the objective function set of the reinforcement learning module can handle multiple conflicting goals, such as minimizing Feng Shui linkage control costs, maximizing Feng Shui linkage control accuracy, maximizing Feng Shui linkage control efficiency, etc., and generate Feng Shui linkage control strategies that balance these goals. The intelligent agent learns historical Feng Shui linkage control strategies through a layered experience replay pool. , continuously optimize its own strategy generation ability. The agent controls the policy network based on the learned experience to generate more effective value strategies. The design of the layered experience replay pool and the agent enables the model to continuously learn and optimize, improving the quality of strategy generation. The reinforcement learning module adopts a group exploration method, which can avoid falling into the local optimal solution to a certain extent. The policy network outputs the distribution probability of actions under a given state, and the value network estimates the action value in the current state, that is, the Q value; the reinforcement learning module realizes the collaboration and learning between multiple agents, optimizes their respective strategies and value estimates by interacting with the environment, and comprehensively considers the cooperation between the water system and the wind system, as well as the game between the wind system at end A and the wind system at end B;
[0067] The intelligent agents include the water system agent, the A-side large system agent, and the B-side large system agent, and the water system agent, the A-side large system agent, and the B-side large system agent are respectively connected to the value network, the strategy network, the hierarchical experience replay pool, and the meta-strategy optimization module;
[0068] The strategy network includes a first current strategy network, a first target strategy network, a second current strategy network, a second target strategy network, a third current strategy network, and a third target strategy network. The first current strategy network and the first target strategy network are both connected to the water system agent, the second current strategy network and the second target strategy network are both connected to the A-end large system agent, and the third current strategy network and the third target strategy network are both connected to the B-end large system agent.
[0069] First, the current policy network: Function: Provides the water system agent with the action probability distribution or action suggestions under the current state; Role: Based on the current state information, outputs the actions that the water system agent should take to optimize the performance of the water system;
[0070] First target policy network: Function: Provides the water system agent with action probability distribution or action suggestions under the target state; Role: Used to stabilize the learning process, by providing a stable target to update the first current policy network and reduce fluctuations in the learning process;
[0071] Second Current Policy Network: Function: Provides the A-side Large System Agent with the action probability distribution or action suggestions in the current state; Role: Based on the current state information, outputs the action that the A-side Large System Agent should take to optimize the performance of the A-side Large System; Second Target Policy Network: Function: Provides the A-side Large System Agent with the action probability distribution or action suggestions in the target state; Role: Used to stabilize the learning process, by providing a stable target to update the second current policy network, reducing fluctuations in the learning process;
[0072] Third, the current policy network: Function: Provides the B-side large system agent with the action probability distribution or action suggestions under the current state; Role: Based on the current state information, outputs the action that the B-side large system agent should take to optimize the performance of the B-side large system;
[0073] The third target policy network: Function: Provides the B-side large system intelligent agent with the action probability distribution or action suggestions under the target state; Role: Used to stabilize the learning process, by providing a stable target to update the third current policy network and reduce fluctuations in the learning process;
[0074] In each agent, the current policy network and the target policy network typically work together in the following manner: the current policy network is used to generate a real-time control policy, that is, to determine the action to be taken based on the current state; the target policy network is used to generate a target policy, that is, to provide a long-term stable policy goal, which is used to update the current policy network to achieve better convergence and stability; the parameters of the target policy network are usually a delayed version of the parameters of the current policy network, and the parameters are periodically copied from the current policy network to maintain a certain degree of stability;
[0075] The value network includes a first current value network, a first target value network, a second current value network, a second target value network, a third current value network, and a third target value network. The first current value network and the first target value network are both connected to the water system intelligent agent, the second current value network and the second target value network are both connected to the A-end large system intelligent agent, and the third current value network and the third target value network are both connected to the B-end large system intelligent agent.
[0076] First, the current value network: Function: Provides the water system agent with a value estimate (Q-value) of the current state or state-action pair; Role: Based on the current state information, it outputs the expected reward of the water system agent under each possible action to guide the water system agent's decision-making;
[0077] First Target Value Network: Function: Provides a stable value estimation target for the water system agent, used to update the first current value network; Effect: Reduces fluctuations in value estimation and improves the stability and convergence of the water system agent's learning process;
[0078] Second Current Value Network: Function: Provides the value estimate of the current state or state-action pair for the large system agent on the A side. Role: Based on the current state information, it outputs the expected reward of the large system agent on the A side under each possible action, which is used to guide the decision-making of the large system agent on the A side.
[0079] Second Target Value Network: Function: Provides a stable value estimation target for the large system agent on the A side, used to update the second current value network; Effect: Reduces fluctuations in value estimation and improves the stability and convergence of the learning process of the large system agent on the A side;
[0080] The third current value network: Function: Provides the B-side large system agent with a value estimate of the current state or state-action pair; Role: Based on the current state information, it outputs the expected return of the B-side large system agent under each possible action to guide the decision-making of the B-side large system agent;
[0081] The third target value network: Function: Provides a stable value estimation target for the B-side large system agent, which is used to update the third current value network; Effect: Reduces the fluctuation of value estimation and improves the stability and convergence of the learning process of the B-side large system agent;
[0082] In each agent, the current value network and the target value network work together in the following way: the current value network outputs value estimates based on the current state or state-action pair, and the agent chooses actions based on these estimates; the target value network provides a stable value estimation target for calculating the loss function and updating the parameters of the current value network; the parameters of the current value network are copied to the target value network periodically or according to certain conditions to maintain the stability of the target value network while gradually reflecting the latest learning results;
[0083] Using a reinforcement learning algorithm, we constructed an initial Feng Shui linkage control model for the subway station air conditioning system, including the following steps:
[0084] S1-2-1: Use the MPO algorithm to build the initial meta-strategy optimization module, and use the optimization target generated by the Feng Shui linkage control strategy as the meta-strategy optimization scenario;
[0085] S1-2-2: Use the MADDPG algorithm to build the initial reinforcement learning module, set up a layered experience replay pool, and use the Feng Shui linkage control strategy generation problem as the simulation environment;
[0086] S1-2-3: Set the objective function set, reward function, action space, and state space for the agent in the initial reinforcement learning module, set the action-value function for the value network, and set the policy function for the policy network; F a
[0087] The first state space of the water system agent , the first action space ,in, is the status value, is the action value; the second state space of the large system agent at end A is , the second action space =[ , ]; B-side large system intelligent agent state variables ;Action variables =[ , ]; global state space ;
[0088] First Current Policy Network , First Target Strategy Network , Second Current Policy Network , Second target strategy network , the third current strategy network and the third target strategy network ,in, are the current policy network parameters, These are target strategy network parameters;
[0089] First Current Value Network , First Target Value Network , Second Current Value Network , Second target value network , the third current value network and the third target value network ,in, are the current policy network parameters, These are target strategy network parameters;
[0090] The formula of the reward function is:
[0091]
[0092]
[0093]
[0094]
[0095] in, They are the A-end platform temperature control reward, the B-end platform temperature control reward, and the total system energy efficiency reward; These are all parameters ranging from 0 to 1, set according to actual needs; Set the temperature for the A and B end stations respectively; Current t Moment and t -1 Number of cold source sets at a moment;
[0096] Reward function for the water system agent:
[0097] r 1 =
[0098] in, r 1 is the reward value of the water system agent;
[0099] The reward function of the large system agent on the A side:
[0100] r 2 =
[0101] in, r 2 is the reward value of the large system agent on the A side;
[0102] The reward function of the B-side large system agent:
[0103] r 3 =
[0104] in, r 3 is the reward value of the B-side large system agent;
[0105] Set up a tiered experience replay pool:
[0106] Create a shared experience replay pool D, storing tuples ( s,a 1 ,a 2 ,a 3 ,r 1 ,r 2 ,r 3 ,s′ ), where s′ is the next state, and the experience replay pool D is divided into m storage spaces, that is, according to (Right now s 1) Divide the experience replay pool D into m storage blocks with equal spacing in three dimensions (the storage blocks of the remaining agents are set according to the corresponding state space dimensions), thus obtaining a layered experience replay pool;
[0107] Exploration and collection of tiered experience replay pools:
[0108] Each agent i Generate actions based on the current policy and noise mechanism, i Indicator for the agent:
[0109] in, For intelligent agents i The action space, is Gaussian noise, used for exploration; the state transition data generated by the interaction is converted into Store it in the storage block of the experience pool and save the current state Updated to ;
[0110] Differentiated Sampling Experience Replay Mechanism: Computation Storage blocks in the experience replay pool j Euclidean distance between the middle positions (m in total) :
[0111]
[0112] in,( x,y,z )for location, ( x j ,y j ,z j ) is a storage blockj central location; j Indicates the amount of storage blocks;
[0113] According to Euclidean distance Sort from small to large (storage blocks with the same distance are randomly sorted), starting from the smallest distance k ( k ≥1) storage blocks, each sample is extracted from this k The probabilities of extracting from the storage blocks are:
[0114]
[0115] S1-2-4: Adjust the loss function of the reinforcement learning module according to the elastic weight connection mechanism of the continuous learning algorithm to obtain the adjusted loss function;
[0116] S1-2-4: Integrate the initial meta-strategy optimization module and the initial reinforcement learning module to obtain the initial Feng Shui linkage control model of the subway station air conditioning system;
[0117] S1-3: Based on several pre-processed historical subway station air conditioning system monitoring data, the initial Feng Shui linkage control model is trained to obtain the final Feng Shui linkage control model, including the following steps:
[0118] S1-3-1: Based on several pre-processed historical subway station air conditioning system monitoring data, the initial meta-strategy optimization module of the initial Feng Shui linkage control model is trained under different scenarios to obtain the final meta-strategy optimization module;
[0119] S1-3-2: Use the final meta-strategy optimization module to initialize the policy network and value network of the initial reinforcement learning module in different scenarios to obtain the initial policy network and initial value network;
[0120] S1-3-3: Traverse all the objective functions in the objective function set and train the initial agent of the initial reinforcement learning module in different scenarios based on some pre-processed historical subway station air conditioning system monitoring data;
[0121] S1-3-4: During the training process, use the backpropagation method to optimize the value network parameters of the initial value network, and use the gradient ascent method to optimize the policy network parameters of the initial policy network;
[0122] Use the back propagation method to optimize the value network parameters of the initial value network:
[0123] Calculate each agent i The formula for the Q target value is: in, For intelligent agents i Q target value; is the reward for the current state; is the discount factor, , , All are the states of the next moment; is the action value function of the target value network;
[0124] The formula for minimizing the timing difference error is:
[0125]
[0126] in, Is the loss function used to update the agent i Parameters; is the batch size; For the b The target value of each sample; b is the sample indicator amount; is the action-value function of the target value network;
[0127] Update via backpropagation ;
[0128] Use the gradient ascent method to optimize the policy network parameters of the initial policy network:
[0129] For each agent i , calculate the policy gradient:
[0130]
[0131] in, Relative to the parameter Policy gradient; For action relative to Action-value function gradient; is the action-value function of the current value network; Relative to the parameter Strategy gradient;
[0132] Parameter update: Update using gradient ascent method :
[0133] in, is the learning rate;
[0134] Formula for target network soft update:
[0135]
[0136]
[0137] in, τ is the update coefficient (take 0.01-0.05);
[0138] S1-3-5: Use the adjusted loss function to obtain the first loss value in the training process until the first loss value is less than the loss value threshold or the first iteration number reaches the number threshold, and obtain the final reinforcement learning module;
[0139] S1-3-6: Collect several historical Feng Shui linkage control experiences generated during the training of the reinforcement learning module, and store the several historical Feng Shui linkage control experiences in the layered experience revisit pool;
[0140] S2: Based on the collected real-time subway station air conditioning system monitoring data, the Feng Shui linkage control model is used to perform Feng Shui linkage control to obtain the real-time Feng Shui linkage control strategy of the subway station air conditioning system, including the following steps:
[0141] S2-1: Collecting and preprocessing real-time subway station air conditioning system monitoring data to obtain preprocessed real-time subway station air conditioning system monitoring data, extracting real-time data features from the preprocessed real-time subway station air conditioning system monitoring data, and inputting the real-time data features into the Feng Shui linkage control model;
[0142] Real-time subway station air-conditioning system monitoring data includes the real-time outdoor temperature, real-time outdoor humidity, real-time subway station passenger flow factor, real-time number of cooling source open sets, real-time chiller outlet water temperature, real-time refrigeration pump frequency, real-time cooling pump frequency, real-time cooling tower fan frequency, real-time A-end platform average temperature, real-time A-end group air-water valve opening, real-time A-end group air supply fan frequency, real-time B-end platform average temperature, real-time B-end group air-water valve opening, real-time B-end group air supply fan frequency, and real-time system total energy efficiency ratio;
[0143] S2-2: Based on the real-time data characteristics, the meta-strategy optimization module of the Feng Shui linkage control model is used to adjust the policy network and value network of the reinforcement learning module of the Feng Shui linkage control model to obtain the adjusted policy network and adjusted value network;
[0144] S2-3: Randomly extract a number of historical Feng Shui linkage control experiences from the stratified experience revisit pool, and update the action space of the adjusted strategy network based on the historical Feng Shui linkage control experiences to obtain an updated action space;
[0145] S2-4: Update the state space of the adjusted policy network according to the real-time data characteristics to obtain an updated state space, and input the updated state space into the adjusted value network;
[0146] S2-5: Use the agent to control the adjusted value network, obtain the reward value of each possible action in the updated action space according to the reward function, and update the Q value of the possible action according to the reward value to obtain the updated Q value of each possible action;
[0147] S2-6: Use the agent to control the adjusted policy network, output the probability distribution of executing the action based on the updated Q value of each possible action in the updated state space, and obtain the real-time Feng Shui linkage control strategy of the subway station air conditioning system based on the probability distribution of executing the action;
[0148] S3: Use a continuous learning algorithm to update the Feng Shui linkage control model and perform Feng Shui linkage control on the subway station air conditioning system based on the real-time Feng Shui linkage control strategy. This includes the following steps:
[0149] S3-1: Collect the real-time Feng Shui linkage control experience corresponding to the Feng Shui linkage control strategy generated by the Feng Shui linkage control model;
[0150] S3-2: Mix the real-time Feng Shui linkage control experience with a number of historical Feng Shui linkage control experiences randomly selected from the stratified experience revisit pool to obtain a number of mixed Feng Shui linkage control experiences;
[0151] S3-3: Based on several hybrid Feng Shui linkage control experiences, continuously train the Feng Shui linkage control model, use the adjusted loss function, and obtain a second loss value during the continuous training process until the second loss value is less than a loss value threshold or the second iteration number reaches a number threshold, thereby obtaining an updated Feng Shui linkage control model;
[0152] S3-4: Generate a real-time Feng Shui linkage control instruction for the subway station air-conditioning system according to the real-time Feng Shui linkage control strategy, and execute the real-time Feng Shui linkage control instruction to perform Feng Shui linkage control on the subway station air-conditioning system.
[0153] Example 2:
[0154] like Figure 2 As shown, this embodiment provides a feng shui linkage control system based on reinforcement learning, which is used to implement a feng shui linkage control method, including a model building unit, a feng shui linkage control unit and a continuous learning unit connected in sequence.
[0155] A model building unit is used to build a Feng Shui linkage control model for the subway station air conditioning system based on a number of historical subway station air conditioning system monitoring data using a reinforcement learning algorithm;
[0156] The Feng Shui linkage control unit is used to perform Feng Shui linkage control based on the collected real-time subway station air conditioning system monitoring data and use the Feng Shui linkage control model to obtain the real-time Feng Shui linkage control strategy of the subway station air conditioning system;
[0157] The continuous learning unit is used to use the continuous learning algorithm to update the Feng Shui linkage control model, and to perform Feng Shui linkage control on the subway station air-conditioning system according to the real-time Feng Shui linkage control strategy.
[0158] The present invention provides a Feng Shui linkage control method and system based on reinforcement learning. By introducing reinforcement learning and continuous learning algorithms, the control strategy can be dynamically adjusted according to real-time environmental changes and passenger flow, thereby improving flexibility and adaptability. By utilizing real-time data and historical experience, the operating status of the air-conditioning system can be more accurately predicted and adjusted, avoiding long-term high-energy consumption operation and achieving energy conservation and emission reduction. Through refined control and real-time adjustment, the temperature stability in the subway station can be maintained, and the comfort and satisfaction of passengers can be improved. With the help of advanced intelligent algorithm models, massive sensor data can be efficiently processed, and responses can be made quickly, thereby improving the level of intelligence. By comprehensively considering the interaction between the A-end large system and the B-end large system of the water system and the wind system, overall coordinated control can be achieved, and the efficiency and stability of the entire air-conditioning system can be improved. The online reinforcement learning method used pre-trains the strategy function of each intelligent agent through a small amount of expert experience imitation learning, so that the reinforcement learning algorithm also has a stable and reliable control result output in the early stage of interaction with the actual system. The improved experience replay mechanism and reward function make the control output as stable and reliable as possible during the convergence process, avoiding frequent starts and stops.
[0159] The present invention is not limited to the above optional embodiments. Anyone can derive various other forms of products based on the teachings of the present invention. The above specific embodiments should not be construed as limiting the scope of protection of the present invention. The scope of protection of the present invention shall be based on the scope defined in the claims, and the description can be used to interpret the claims.
Claims
1. A Feng Shui linkage control method based on reinforcement learning, characterized by: The steps include: Based on several historical subway station air conditioning system monitoring data, a reinforcement learning algorithm was used to build a Feng Shui linkage control model for subway station air conditioning systems. The Feng Shui linkage control model is constructed based on the MPO-MADDPG algorithm, and includes a meta-strategy optimization module constructed based on the MPO algorithm and a reinforcement learning module constructed based on the MADDPG algorithm, which are connected in sequence. The reinforcement learning module is provided with an intelligent agent, a value network, a policy network, and a hierarchical experience replay pool. The intelligent agent is respectively connected to the value network, the policy network, the hierarchical experience replay pool, and the meta-strategy optimization module; The intelligent agents include a water system agent, an A-side large system agent, and a B-side large system agent, and the water system agent, the A-side large system agent, and the B-side large system agent are respectively connected to the value network, the strategy network, the hierarchical experience replay pool, and the meta-strategy optimization module; The strategy network includes a first current strategy network, a first target strategy network, a second current strategy network, a second target strategy network, a third current strategy network, and a third target strategy network. The first current strategy network and the first target strategy network are both connected to the water system intelligent agent, the second current strategy network and the second target strategy network are both connected to the A-end large system intelligent agent, and the third current strategy network and the third target strategy network are both connected to the B-end large system intelligent agent; The value network includes a first current value network, a first target value network, a second current value network, a second target value network, a third current value network and a third target value network. The first current value network and the first target value network are both connected to the water system intelligent agent, the second current value network and the second target value network are both connected to the A-end large system intelligent agent, and the third current value network and the third target value network are both connected to the B-end large system intelligent agent; Based on the collected real-time subway station air conditioning system monitoring data, the wind and water linkage control model is used to perform wind and water linkage control and obtain the real-time wind and water linkage control strategy of the subway station air conditioning system; Using a continuous learning algorithm, the Feng Shui linkage control model is updated. Based on the real-time Feng Shui linkage control strategy, Feng Shui linkage control is implemented for the subway station air conditioning system. The steps include: Collect the real-time feng shui linkage control experience corresponding to the real-time feng shui linkage control strategy generated by the feng shui linkage control model; Mixing the real-time Feng Shui linkage control experience with a number of historical Feng Shui linkage control experiences randomly selected from the stratified experience revisit pool to obtain a number of mixed Feng Shui linkage control experiences; Based on several hybrid Feng Shui linkage control experiences, the Feng Shui linkage control model is continuously trained, and the adjusted loss function is used to obtain the second loss value in the continuous training process until the second loss value is less than the loss value threshold or the second iteration number reaches the number threshold, thereby obtaining an updated Feng Shui linkage control model; According to the real-time Feng Shui linkage control strategy, the real-time Feng Shui linkage control instructions of the subway station air-conditioning system are generated, and the real-time Feng Shui linkage control instructions are executed to perform Feng Shui linkage control on the subway station air-conditioning system.
2. The feng shui linkage control method based on reinforcement learning according to claim 1, characterized in that: Based on several historical subway station air conditioning system monitoring data, a reinforcement learning algorithm was used to construct a Feng Shui linkage control model for subway station air conditioning systems. The model includes the following steps: Collecting and preprocessing a number of historical subway station air conditioning system monitoring data to obtain a number of preprocessed historical subway station air conditioning system monitoring data; Use reinforcement learning algorithms to build an initial Feng Shui linkage control model for subway station air conditioning systems; Based on several pre-processed historical subway station air-conditioning system monitoring data, the initial Feng Shui linkage control model is trained to obtain the final Feng Shui linkage control model.
3. The feng shui linkage control method based on reinforcement learning according to claim 2, characterized in that: Using a reinforcement learning algorithm, we constructed an initial Feng Shui linkage control model for the subway station air conditioning system, including the following steps: Use the MPO algorithm to build the initial meta-strategy optimization module, and use the optimization target generated by the Feng Shui linkage control strategy as the meta-strategy optimization scenario; Using the MADDPG algorithm, we built an initial reinforcement learning module, set up a layered experience replay pool, and used the Feng Shui linkage control strategy generation problem as the simulation environment. Set the objective function set, reward function, action space, and state space for the agent in the initial reinforcement learning module, set the action-value function for the value network, and set the policy function for the policy network; According to the elastic weight connection mechanism of the continuous learning algorithm, the loss function of the reinforcement learning module is adjusted to obtain the adjusted loss function; By integrating the initial meta-strategy optimization module and the initial reinforcement learning module, the initial Feng Shui linkage control model of the subway station air-conditioning system is obtained.
4. The feng shui linkage control method based on reinforcement learning according to claim 3 is characterized in that: The historical subway station air conditioning system monitoring data includes the historical outdoor temperature of the subway station air conditioning system, the historical outdoor humidity, the historical subway station passenger flow factor, the historical number of cold source opening sets, the historical chiller water outlet temperature, the historical refrigeration pump frequency, the historical cooling pump frequency, the historical cooling tower fan frequency, the historical A-end platform average temperature, the historical A-end group air-water valve opening, the historical A-end group air supply fan frequency, the historical B-end platform average temperature, the historical B-end group air-water valve opening, the historical B-end group air supply fan frequency and the historical system total energy efficiency ratio; The real-time subway station air-conditioning system monitoring data includes the real-time outdoor temperature of the subway station air-conditioning system, the real-time outdoor humidity, the real-time subway station passenger flow factor, the real-time number of cold source opening sets, the real-time chiller outlet water temperature, the real-time refrigeration pump frequency, the real-time cooling pump frequency, the real-time cooling tower fan frequency, the real-time A-end platform average temperature, the real-time A-end group air-water valve opening, the real-time A-end group air supply fan frequency, the real-time B-end platform average temperature, the real-time B-end group air-water valve opening, the real-time B-end group air supply fan frequency and the real-time system total energy efficiency ratio.
5. The feng shui linkage control method based on reinforcement learning according to claim 4 is characterized in that: Based on several pre-processed historical subway station air conditioning system monitoring data, the initial Feng Shui linkage control model is trained to obtain the final Feng Shui linkage control model, which includes the following steps: Based on several pre-processed historical subway station air conditioning system monitoring data, the initial meta-strategy optimization module of the initial Feng Shui linkage control model is trained in different scenarios to obtain the final meta-strategy optimization module; Use the final meta-strategy optimization module to initialize the policy network and value network of the initial reinforcement learning module in different scenarios to obtain the initial policy network and initial value network; Traverse all the objective functions in the objective function set and train the initial agent of the initial reinforcement learning module in different scenarios based on some pre-processed historical subway station air conditioning system monitoring data; During the training process, the back propagation method is used to optimize the value network parameters of the initial value network, and the gradient ascent method is used to optimize the policy network parameters of the initial policy network; Using the adjusted loss function, obtain the first loss value in the training process until the first loss value is less than the loss value threshold or the first iteration number reaches the number threshold, and obtain the final reinforcement learning module; Collect several historical Feng Shui linkage control experiences generated during the training process of the reinforcement learning module, and store the several historical Feng Shui linkage control experiences in a layered experience revisit pool.
6. The feng shui linkage control method based on reinforcement learning according to claim 5, characterized in that: Based on the collected real-time monitoring data of the subway station air conditioning system, the Feng Shui linkage control model is used to perform Feng Shui linkage control, and the real-time Feng Shui linkage control strategy of the subway station air conditioning system is obtained, which includes the following steps: Collecting real-time subway station air conditioning system monitoring data, preprocessing it to obtain preprocessed real-time subway station air conditioning system monitoring data, extracting real-time data features of the preprocessed real-time subway station air conditioning system monitoring data, and inputting the real-time data features into the Feng Shui linkage control model; According to the real-time data characteristics, the meta-strategy optimization module of the Feng Shui linkage control model is used to adjust the policy network and value network of the reinforcement learning module of the Feng Shui linkage control model to obtain the adjusted policy network and adjusted value network; Randomly extract a number of historical Feng Shui linkage control experiences from the stratified experience revisit pool, and update the action space of the adjusted strategy network based on the historical Feng Shui linkage control experiences to obtain the updated action space; According to the real-time data characteristics, the state space of the adjusted policy network is updated to obtain the updated state space, and the updated state space is input into the adjusted value network; Use the agent to control the adjusted value network, obtain the reward value of each possible action in the updated action space according to the reward function, and update the Q value of the possible action according to the reward value to obtain the updated Q value of each possible action; The intelligent agent is used to control the adjusted policy network. According to the updated Q value of each possible action in the updated state space, the probability distribution of the executed action is output. Based on the probability distribution of the executed action, the real-time Feng Shui linkage control strategy of the subway station air-conditioning system is obtained.
7. A feng shui linkage control system based on reinforcement learning, used to implement the feng shui linkage control method according to any one of claims 1 to 6, characterized in that: It includes a model building unit, a Feng Shui linkage control unit and a continuous learning unit that are connected in sequence.
Citation Information
Patent Citations
Subway station air conditioning system energy-saving control method based on deep reinforcement learning
CN113283156A
Multi-mode intelligent control system for subway station air conditioner
CN116772386A