A large public building demand response intelligent regulation method based on a multi-modal reinforcement learning agent system
By using a multimodal reinforcement learning intelligent agent system and combining multi-source heterogeneous data, a multi-agent collaborative mechanism is constructed, which solves the problems of high energy consumption and insufficient intelligent management and control in the energy system of large public buildings, and realizes the high efficiency, safety and green energy saving of the energy system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2025-01-10
- Publication Date
- 2026-05-05
AI Technical Summary
The energy systems of large public buildings are characterized by high energy consumption, low automation, and insufficient intelligent management, making it difficult to cope with complex and ever-changing operating environments and market fluctuations. Furthermore, traditional methods are insufficient to maximize energy efficiency, minimize operating costs, and improve environmental benefits.
A multimodal reinforcement learning intelligent agent system is adopted. Through a multi-agent collaborative mechanism and combined with multi-source heterogeneous data, an intelligent agent system is constructed to realize the dynamic regulation of the electricity-carbon-green certificate market, including electricity market demand response, carbon market demand response and green certificate market demand response, and regulation of equipment such as cold and heat sources, air conditioning, and lighting.
It has improved the operation and maintenance efficiency and safety and reliability of the energy system, achieved green energy conservation, ensured the normal operation and comfort of large buildings, and improved energy utilization efficiency and economic benefits.
Smart Images

Figure CN120031292B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of energy systems for large public buildings and artificial intelligence, specifically involving an intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system. Background Technology
[0002] Large-scale building energy systems face challenges such as high energy consumption, low automation levels, limited intelligent management, and complex subsystem structures, all of which urgently require solutions. Demand response technology is gradually becoming a key means to achieve green, low-carbon, and intelligent development of energy in large public buildings. Building demand response technology can be used to regulate energy consumption in large public buildings during peak electricity consumption periods, improve energy efficiency through intelligent control, reduce operating costs, and achieve energy savings. However, traditional building energy management methods mainly rely on rule-based static control strategies, which are ill-suited to cope with real-time dynamic energy market fluctuations and the complex and ever-changing operating environment within buildings. Furthermore, with the rapid development of the electricity market, carbon trading market, and green certificate market, large public buildings need to participate in demand response optimization across multiple markets under constraints such as economic efficiency, environmental benefits, and user comfort.
[0003] In recent years, the rise of artificial intelligence technology, especially reinforcement learning, has provided new pathways for the intelligent optimization of complex systems. Reinforcement learning, through interactive learning, can adaptively optimize control strategies, while multimodal reinforcement learning further leverages the characteristics of multi-source heterogeneous data in buildings (such as environmental conditions, equipment operating status, and market price dynamics), providing efficient solutions for demand response in large public buildings. Furthermore, through multi-agent collaborative reinforcement learning technology, unified optimization of global and local objectives can be achieved across the building's multi-layered structure (such as the market layer and the equipment layer). This approach supports buildings' participation in electricity market trading, carbon market, and green certificate market while dynamically regulating internal building equipment (such as air conditioning, heating, and lighting systems) to maximize energy efficiency, minimize operating costs, and enhance environmental benefits.
[0004] To address the aforementioned challenges, it is urgent to find a solution that leverages artificial intelligence to collaboratively optimize the operation and maintenance management strategies of large-scale building integrated energy systems, thereby improving their operational efficiency and reliability, achieving green energy conservation, and ensuring the safe and comfortable operation of large-scale buildings across various scenarios. Summary of the Invention
[0005] This invention proposes an intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system. It aims to overcome the limitations of traditional methods by introducing multimodal data processing and multi-agent collaborative mechanisms to construct an intelligent agent system that can adapt to the dynamic market of electricity, carbon and green certificates and complex operating environments. This provides an efficient and intelligent solution for large public buildings to participate in demand response in the electricity, carbon and green certificate market.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A method for intelligent demand response control of large public buildings based on a multimodal reinforcement learning agent system includes the following steps:
[0008] S1. Based on the operation and maintenance characteristics of large public buildings, determine the scope of large public buildings participating in demand response control and identify key controllable equipment; based on the scope of participation in demand response control and key controllable equipment, determine the types of intelligent agents involved and their respective functions to obtain a multimodal reinforcement learning intelligent agent system.
[0009] S2, using intelligent sensors to realize multimodal data perception and acquisition of large public buildings, constructing a multimodal data set D of large public buildings, and extracting features from set D to obtain a feature set F of large public buildings multimodal data; combining existing large language models and visual language models to initially construct each agent of the multimodal reinforcement learning agent system in step S1, and using deep reinforcement learning algorithms to initially train each agent;
[0010] S3, based on the multimodal data feature set 𝐹 and the agents constructed in step S2, a multimodal reinforcement learning agent system is trained in a centralized manner using a deep learning reinforcement algorithm, and the collaborative control of the agents constructed in step S2 is realized based on a distributed execution mechanism;
[0011] S4, Deploy the multimodal reinforcement learning agent system into the large public building;
[0012] S5, the multimodal reinforcement learning intelligent agent system participates in the intelligent regulation of demand response of large public buildings. The specific method is as follows: establish a two-layer regulation model of demand response of large public buildings in the electricity-carbon-green certificate market, collaboratively solve the two-layer regulation model of demand response of large public buildings in the electricity-carbon-green certificate market based on the multimodal reinforcement learning intelligent agent system, and regulate the key controllable equipment in step S1 based on the model solution results; and maintain the multimodal reinforcement learning intelligent agent system.
[0013] Furthermore, in step S1:
[0014] The scope of demand response regulation for large public buildings includes electricity market demand response regulation, carbon market demand response, and green certificate market demand response regulation. Key controllable equipment includes cooling and heating power equipment, central air conditioning equipment, energy storage equipment, lighting equipment, and renewable energy equipment.
[0015] The multimodal reinforcement learning agent system includes an upper-layer agent system, a lower-layer agent system, and a communication and coordination agent system; the upper-layer agent system and the lower-layer agent system operate in global coordination through the communication and coordination agent system.
[0016] The upper-level intelligent agent system includes an intelligent agent for demand response in the electricity market, an intelligent agent for demand response in the carbon market, and an intelligent agent for demand response in the green certificate market;
[0017] The lower-level intelligent agent system includes environmental perception intelligent agents, energy management intelligent agents, user behavior intelligent agents, fault detection and maintenance intelligent agents, policy generation intelligent agents, and multimodal data fusion intelligent agents.
[0018] Each intelligent agent possesses different functional characteristics and can be extended to build corresponding intelligent agents according to different scenario requirements.
[0019] The functions of the upper-layer intelligent agent system are as follows:
[0020] 1) The electricity market demand response intelligent agent minimizes electricity costs by optimizing electricity procurement and electricity consumption control strategies, thereby supporting grid stability and improving building economic efficiency;
[0021] 2) Carbon market demand response agents are responsible for monitoring and optimizing building carbon emissions, reducing carbon emission costs and promoting the achievement of green and low-carbon goals by participating in carbon quota trading and emission reduction strategy formulation;
[0022] 3) The green certificate market demand response intelligent agent improves the utilization rate of renewable energy by rationally planning the purchase and sale of green certificates, while ensuring policy compliance and optimizing transaction costs.
[0023] The functions of the lower-layer intelligent agent system are:
[0024] 4) The environmental sensing agent is responsible for collecting and analyzing environmental data in real time, including information such as temperature, humidity, and light intensity, to provide accurate basic data support for demand response regulation;
[0025] 5) The user behavior intelligent agent analyzes user behavior patterns (such as working hours, activity areas, etc.) to intelligently regulate the indoor environment and dynamically adjust energy distribution and equipment operation to meet user needs;
[0026] 6) The multimodal data fusion intelligent agent is responsible for processing multi-source heterogeneous data (images, audio, video, etc.) from large public buildings, including environmental data, equipment operation status data, user behavior data, and external market dynamic information. Through efficient data fusion and feature extraction methods, it provides comprehensive and accurate decision support for intelligent demand response regulation.
[0027] 7) The energy management agent is responsible for monitoring and optimizing the building's energy consumption, such as electricity and heat, and developing efficient energy use strategies to reduce energy consumption.
[0028] 8) The strategy-generating agent uses deep reinforcement learning algorithms to optimize intelligent decision-making based on multimodal data and preset targets, thereby realizing the dynamic adjustment and execution of demand response strategies.
[0029] 9) The fault detection and maintenance intelligent agent is responsible for monitoring the operating status of building equipment, promptly detecting and responding to potential faults, reducing energy waste, improving equipment operating efficiency, and extending service life.
[0030] The functions of the communication coordination intelligent agent system are:
[0031] 10) The communication and coordination intelligent agent system is responsible for information interaction and collaboration in the multi-agent system by coordinating the tasks and data flow of upper and lower-level intelligent agents, ensuring the consistency and synergy of the behavior of each intelligent agent.
[0032] Through the global collaboration of these intelligent agents, the entire system can achieve precise, efficient, and intelligent management of building demand response under complex and ever-changing environmental and market conditions.
[0033] It should be noted that the above outlines the general composition of a multimodal enhanced intelligent agent system for large public buildings, and corresponding intelligent agents can be constructed according to different scenario requirements.
[0034] Furthermore, step S2 specifically includes the following steps:
[0035] Step S21 involves using intelligent sensors to achieve multimodal data perception and acquisition for large public buildings. These intelligent sensors include temperature and humidity sensors, gas sensors, infrared sensors, intelligent cameras, and radio frequency identification (RFID) sensors. Multimodal data for large public buildings refers to diverse types of data reflecting the operational status of various aspects of the building. This includes environmental data, energy consumption data, indoor occupant behavior data, electricity market data, and carbon market data. The diverse data types include text, audio, video, and images, and are uploaded to the existing integrated energy intelligent management and control platform for large public buildings via wireless communication technology.
[0036] Among them, environmental data refers to indoor temperature and humidity, indoor CO2 concentration, and outdoor meteorological data (temperature, humidity, wind speed, etc.); energy consumption data refers to the real-time energy consumption of electrical equipment in various areas of large public buildings and the real-time energy consumption statistics of each cold and heat source equipment; indoor personnel behavior data refers to the real-time distribution characteristics of indoor personnel in large public buildings, including the number of personnel in each functional area, personnel flow, and real-time personnel load; electricity market data includes real-time information such as grid load, electricity price, load fluctuation, and grid stability; carbon market data includes real-time carbon trading volume, real-time carbon tax, carbon emissions, and real-time carbon emission factors.
[0037] Step S22: Construct a multimodal data set D for large public buildings. The multimodal data set D consists of historical data collected by the existing integrated energy intelligent management and control platform for large public buildings and existing knowledge sets for large public buildings. The existing knowledge sets for large public buildings include existing knowledge on electricity-carbon coupling demand response, green building energy efficiency optimization, energy system integration and management, renewable energy utilization, energy conversion and storage technologies, and intelligent energy monitoring and dispatching systems. The carriers of this existing knowledge can be e-books, existing standards, audio, video, or case collections formed during the operation and maintenance of large public buildings, or even voice conversations between operation and maintenance personnel.
[0038] The multimodal dataset D of the large public building is represented in tuple form as follows:
[0039]
[0040] Among them, D i Let represent the dataset for the i-th mode, and n be the total number of modes. The dataset for each mode... This can be further expressed as:
[0041]
[0042] in, It is the value of the k-th data point of the i-th mode at time t, m i It is the total number of data points for the i-th mode.
[0043] Step S23: Use intelligent algorithms to perform feature processing and fusion on the multimodal dataset D to obtain the feature set F of multimodal data of large public buildings. The feature processing and fusion specifically includes denoising, time alignment, standardization and normalization, feature extraction and feature fusion.
[0044] Among these, denoising refers to using algorithms such as mean filtering, median filtering, and Gaussian filtering to clean up noise and interference in various smart sensor data, thereby improving the quality of multimodal data; time alignment refers to using techniques such as timestamp synchronization, interpolation methods, or dynamic time warping to ensure that data from different modalities are consistent in time, that is, to ensure that the readings of different sensors are at the same point in time, such as ensuring the temporal correspondence between visual and sound information in video and audio data; standardization and normalization refers to standardizing the multimodal data to make the data dimensions consistent; feature extraction refers to using wavelet transform and Fourier transform to extract time-frequency domain feature information from the original data, such as electricity consumption patterns, temperature changes, and human activity patterns; feature fusion is the process of integrating the features obtained from the above feature extraction using techniques such as weighted averaging, splicing, and deep learning fusion, thereby forming a more comprehensive and robust feature representation.
[0045] The feature set F of the multimodal data of the large public building is represented in the form of a tuple as follows:
[0046]
[0047] Among them, F i Let represent the set of data features for the i-th mode, and n be the total number of modes. The set of data features for each mode can be further represented as:
[0048]
[0049] Among them, f ik M is the value of the k-th feature of the i-th mode at time t. i This represents the total number of data features for the i-th modality. It should be noted that the above steps outline the general process for multimodal data feature processing and fusion in large public buildings. Appropriate algorithms and technologies need to be selected and adjusted according to the characteristics of different modal data and application scenarios.
[0050] Step S24: Based on existing large language models, visual language models, and deep reinforcement learning algorithms, initially construct individual intelligent agents for large public buildings. Specific steps include:
[0051] Step S241 involves providing the multimodal data feature set F of large public buildings as input features to the existing large language model and visual language model, thus initially constructing the various agents in the multimodal reinforcement learning agent system (specifically including an electricity market demand response agent, a carbon market demand response agent, a green certificate market demand response agent, an environmental perception agent, an energy management agent, a multimodal data fusion agent, a user behavior agent, a fault detection and maintenance agent, and a strategy generation agent). Existing large language models include GPT-4 and Gemini, which possess human-like decision-making and reasoning abilities; visual language models such as CLIP, Qwen2-VL, and PaliGemma combine visual perception of images and videos with text-based reasoning to process and mine complex multimodal data.
[0052] Step S242: Use a deep reinforcement learning algorithm to perform preliminary training on each agent;
[0053] Parameter definition, including defining the global state. The local state space of the i-th single agent The action space of the i-th single agent and the reward function of the i-th single agent ;
[0054] Global state It is used to describe the operational status of the entire large public building, specifically including real-time status information of ambient temperature, ambient humidity, equipment operating status, personnel activity level, electricity market, carbon market, and green certificate market.
[0055] The local state of the i-th single agent is Its observed local state space It is a subset of the global state, that is:
[0056]
[0057] Where T represents ambient temperature, H represents ambient humidity, E represents equipment operating status, and P represents personnel activity level. For electricity market transaction volume, For carbon market trading volume, For green certificate market trading volume;
[0058] All possible actions performed by the i-th single agent The set is represented as the action space. ,Right now:
[0059]
[0060] In the formula, For the joint action of the i-th single agent, Represents the amount of environmental temperature regulation in large public buildings. Represents the lighting adjustment amount of large public buildings. Represents the heating regulation capacity of large public buildings. For the trading volume of large public buildings in the electricity market, For the trading volume of large public buildings in the carbon market, This refers to the transaction volume of large public buildings in the green certificate market.
[0061] reward function It is defined based on the current state and actions performed by the i-th single agent, and is used to guide the learning process of the agent.
[0062]
[0063] In the formula, It is a reward based on energy consumption. It is a reward based on indoor comfort. It is a reward for participating in the electricity market demand response. It is a reward for participating in carbon market trading. It is a reward for participating in the green certificate market.
[0064] The essence of training the i-th agent using deep reinforcement learning algorithms is learning. The strategy corresponding to maximizing expected cumulative reward Specifically, it is represented by the following mathematical model:
[0065]
[0066] In the formula, It is the learning strategy of the i-th agent. Performance metrics; The expected value is represented by the initial state s0 and the action. The probability distribution is calculated from ρ(⋅∣). ) is in the strategy The state distribution under; It is the discount factor for the i-th agent at time t, used to balance the importance of immediate rewards and future rewards. .
[0067] The agent operates in a simulated environment and adjusts its behavioral strategies based on feedback. At each step, the agent selects an action and receives a reward based on the action's effects (such as energy consumption, comfort, etc.). The simulated environment refers to a virtual or computational model used to test and train the agent's behavior. This environment maps certain physical characteristics of the real world as closely as possible, allowing the agent to interact within it and learn how to achieve specific goals.
[0068] Furthermore, in step S3:
[0069] Given the complexity of large public buildings, a global reward function is constructed using a distributed execution mechanism and a multimodal data feature set 𝐹. This enables the collaborative control of the multimodal reinforcement learning-based intelligent agent system. Specifically, it includes the following steps:
[0070] Step S31: Establish the global joint action of the multimodal reinforcement learning agent system. ;
[0071] Global Joint Actions of Multimodal Reinforcement Learning Agent Systems Represented as:
[0072]
[0073] in, This represents the joint action of the i-th agent;
[0074] Constructing the global action space of a multimodal reinforcement learning agent system Global action space of multimodal reinforcement learning intelligent agent system It is the Cartesian product of the action spaces of all intelligent agents, specifically expressed as:
[0075]
[0076] in, Let i represent the action space of the i-th agent.
[0077] Step S32, define union function , Indicates the global state and global joint actions The reward for the i-th agent;
[0078] =
[0079] in, This represents the Q-network parameters (such as weights and biases) of the i-th agent.
[0080] Policy network of the i-th agent By optimizing the gradient of the policy network Make the objective function Maximize, policy gradient Represented as:
[0081]
[0082] in, Represents the gradient. It is the gradient of the joint Q-function with respect to the action of the i-th agent; These are the parameters of the policy network; It is a policy network Regarding parameters The gradient.
[0083] Joint Q function Through the mean square error function renew:
[0084]
[0085] Where the target value for:
[0086]
[0087] This indicates the global joint action of the multimodal reinforcement learning agent system in the next state; The next state represents the global state of the multimodal reinforcement learning agent system. Let be the Q-network parameters of the i-th agent in the next state.
[0088] Cooperative control of multimodal reinforcement learning agent systems through global objectives Represented as:
[0089]
[0090] in, Represents the global expected value. It is the global reward function, defined as the weighted sum of the rewards of all agents:
[0091]
[0092] It is the weight of the i-th agent.
[0093] During the training phase, the agent's policy and Q-function are trained and optimized intensively. The distributed update rule for the policy of each agent i is as follows:
[0094]
[0095] The parameter update rule for the Q function is as follows:
[0096]
[0097] in, This is the learning rate.
[0098] Step S33: In distributed execution, the agent bases its actions on local observations. and strategy Decision-making actions are jointly implemented to achieve the coordinated operation of various agents in a multimodal reinforcement learning agent system.
[0099] Based on multimodal reinforcement learning intelligent agent systems, multi-scenario demand response control strategies can be designed, including dynamic equipment scheduling, energy consumption optimization, and user comfort assurance.
[0100] Furthermore, in step S4:
[0101] The trained multimodal reinforcement learning intelligent agent system is deployed into a practical integrated energy intelligent management and control platform and integrated with the building's control systems (such as air conditioning, lighting, and heating systems). This includes deploying the upper-level intelligent agent system on the integrated energy intelligent management and control platform to handle electricity-carbon-green certificate market transactions; and embedding the lower-level intelligent agent into the equipment controller to dynamically regulate various devices. Specifically, this includes the following steps:
[0102] Step S41: The upper-layer intelligent agent system integrates with the integrated energy intelligent management and control platform. The data interface can connect with the building management system (BMS) and energy market platform through protocols (such as OPC UA, Modbus, or RESTful API) to obtain market prices (electricity, carbon, green certificates), building energy consumption forecast data, and historical transaction record information. On the other hand, it submits transaction requests to the energy market through the integrated energy intelligent management and control platform, such as green certificate purchase, carbon emission quota trading, and electricity procurement. Then, it dynamically adjusts the operation of the lower-layer intelligent agent system based on the transaction results.
[0103] Step S42, Integration of the lower-level intelligent agent system with the device controller, Embedded deployment: The trained lower-level intelligent agent is embedded in the device controller, including PLC (Programmable Logic Controller) and IoT devices.
[0104] Furthermore, step S5 specifically includes the following steps:
[0105] Step S51: Establish a two-tiered regulatory model for the demand response of large public buildings in the electricity-carbon-green certificate market, which includes an upper-level economic model and a lower-level day-ahead scheduling model.
[0106] The upper-level economic model maximizes profits by involving large public buildings in the electricity market, carbon market, and green certificate market demand response. The objective is to establish a higher-level economic model. Specifically, the upper-level economic model is as follows:
[0107]
[0108] In the formula, The electricity purchase price for large public building b at time t. It is the marginal cost of large public building b at time t. This refers to the electricity consumption of large public building b at time t. The price of a green certificate for a large public building (b) at time t. This refers to the volume of green certificate transactions purchased by large public building B at time t. The carbon price of large public building b at time t. It is the carbon emissions of a large public building b at time t.
[0109] The lower-level day-ahead dispatch model minimizes the cost of various energy devices in the demand response of large public buildings participating in the electricity market, carbon market, and green certificate market. The objective involves determining the market clearing price and the trading volume of each market participant to achieve supply and demand balance in the carbon electricity market and the green certificate market. The lower-level day-ahead scheduling model is as follows:
[0110]
[0111] in, This represents the day-ahead scheduling cost of large public building b within time t. These are the day-ahead scheduling decision variables for each key controllable device in a large public building b within time t; This represents the total cost of electricity in the electricity market for a large public building b within time t, including electricity purchase costs and demand response costs. This represents the total carbon market cost of large public building b within time t, including the cost of purchasing carbon allowances. This represents the total cost or benefit of the green certificate market for a large public building b within time t, depending on whether it buys or sells green certificates.
[0112] This can be expressed as:
[0113]
[0114] in, This refers to the day-ahead scheduling decision of the cold and heat source power equipment in large public building b; This refers to the day-ahead scheduling decision of air conditioning equipment in large public building b; This refers to the day-ahead scheduling decision of energy storage equipment in large public building b; This refers to the day-ahead scheduling decision for lighting equipment in a large public building (b). This refers to the day-ahead scheduling decisions for renewable energy equipment in large public buildings (b).
[0115] The constraints of the lower-level day-ahead dispatch model include power supply and demand balance constraints, carbon emission limit constraints, green certificate trading limit constraints, and energy equipment output constraints.
[0116] Electricity supply and demand balance constraints:
[0117]
[0118] in, This represents the amount of electricity purchased by large public building b at time t. It represents the electricity sales volume of large public building b at time t. This refers to the electricity purchased or sold by large public building B. This means that it holds true for all times t.
[0119] Carbon emission limits and constraints:
[0120]
[0121] in, This refers to the carbon emissions of a large public building (b) at time t. It is the carbon allowance for large public buildings (b).
[0122] Restrictions on Green Certificate Trading:
[0123]
[0124] in, It represents the volume of green certificate transactions for large public building B at time t. It is the renewable energy generation of a large public building b at time t.
[0125] Energy equipment output constraints:
[0126]
[0127] and These refer to the lower and upper limits of the output of the i-th energy device in a large public building, respectively.
[0128] The aforementioned constraints ensure that when large public buildings participate in the electricity-carbon market, their electricity purchases meet their electricity demand, their carbon emissions do not exceed the prescribed quotas, and their green certificate trading volume matches their building renewable energy trading volume. These constraints are further integrated into the objective function of the lower-level model to ensure that large public buildings participate reasonably in the electricity-carbon-green certificate market demand response.
[0129] Step S52: Solve the upper-level economic model and the lower-level day-ahead scheduling model using the upper-level intelligent agent system and the lower-level intelligent agent system respectively. Then, achieve information interaction and collaborative solving between the upper and lower-level intelligent agent systems through communication coordination. Solve the upper-level economic model to obtain... , , Solve the lower-level day-ahead scheduling model to obtain .
[0130] By establishing the upper-level mean square error function of the upper-level intelligent agent system Objective function of the upper-level intelligent agent system The policy network gradient of the upper-level intelligent agent system Then, by solving the upper-level economic model, we can obtain... , , Specifically:
[0131] Establish the upper-level mean square error function of the upper-level intelligent agent system. ,
[0132]
[0133]
[0134] In the formula, The mean square error function of the upper layer of the upper-layer intelligent agent system; This represents the expected value of the upper-level intelligent agent system. For the joint Q-function of the upper-level agent system in a multimodal reinforcement learning agent system; This represents the global state of the upper-level intelligent agent system. For the joint actions of the upper-level intelligent agent system, it is represented as ; For adjustment The learning parameters of the function; The target value for the upper-level intelligent agent system; The reward function established by the upper-level intelligent agent system based on the upper-level economic model, i.e. ; The discount factor of the upper-level intelligent agent system is used to balance the rewards of the upper-level intelligent agent system; This is the joint Q-function of the upper-level agent system in the multimodal reinforcement learning agent system in the next state.
[0135] Objective function of upper-level intelligent agent system for:
[0136]
[0137] Furthermore, the policy update process of the upper-layer intelligent agent system is based on the gradient of the policy network of the upper-layer intelligent agent system. Implementation, specifically:
[0138]
[0139] In the formula, The objective function of the upper-level intelligent agent system Regarding its strategy network parameters The policy network gradient is used to guide policy updates to maximize... ; It refers to the network parameters of the upper-level intelligent agent system regarding its policy. Expectations; These are the policy network parameters for the upper-level intelligent agent system. For the upper-layer intelligent agent system policy network; For the policy network of the upper-level intelligent agent system The gradient; This is the gradient of the joint Q-function of the upper-level intelligent agent system.
[0140] By establishing the lower-level mean square error function of the lower-level intelligent agent system Objective function of the lower-level intelligent agent system The policy network gradient of the lower-level intelligent agent system Solving the lower-level day-ahead scheduling model yields... Specifically:
[0141] Lower-level intelligent agent system upper-level mean square error function for,
[0142]
[0143]
[0144] In the formula, The mean square error function of the lower-level intelligent agent system; This represents the expected value of the lower-level intelligent agent system. For the joint Q-function of the lower-level agent system in a multimodal reinforcement learning agent system; This represents the global state of the lower-level intelligent agent system. For the joint actions of the lower-level intelligent agent system, it is represented as ; For adjustment The learning parameters of the function; The target value for the lower-level intelligent agent system; The reward function established by the lower-level intelligent agent system based on the lower-level day-ahead scheduling model, i.e. ; The discount factor of the lower-level intelligent agent system is used to balance the rewards of the lower-level intelligent agent system; This is the joint Q-function of the lower-level agent system in the multimodal reinforcement learning agent system in the next state.
[0145] Objective function of the lower-level intelligent agent system for:
[0146]
[0147] Furthermore, the policy update process of the lower-level intelligent agent system is based on the policy network gradient of the lower-level intelligent agent system. Implementation, specifically:
[0148]
[0149] In the formula, The objective function of the upper-level intelligent agent system Regarding its strategy network parameters The policy network gradient is used to guide policy updates to maximize... ; It refers to the network parameters of the lower-level intelligent agent system regarding its policy. Expectations; These are the policy network parameters for the lower-level intelligent agent system. For the policy network of the lower-level intelligent agent system; Policy network for lower-level intelligent agent systems The gradient; This represents the gradient of the joint Q-function of the lower-level intelligent agent system.
[0150] Based on the model solution results , , , Adjust the aforementioned key adjustable equipment;
[0151] Upper-layer agents output market trading strategies (such as electricity procurement and carbon emission allowances); lower-layer agents provide feedback on equipment operation results and environmental conditions (such as actual energy consumption, carbon emissions, and indoor comfort). Based on real-time market fluctuations and equipment operating status, both upper and lower-layer agents adjust their respective strategies.
[0152] In practical applications, the upper-level intelligent agent system responds to and regulates demand in real time based on real-time monitoring of environmental conditions (including grid load, electricity price, and building energy consumption). For example, during peak grid periods, the lower-level intelligent agent system can reduce building energy consumption by reducing air conditioning load and delaying the operation of non-critical equipment.
[0153] The system dynamically adjusts and optimizes the behavior of the lower-level intelligent agent system based on real-time feedback: if a certain control strategy leads to reduced user comfort or lack of energy efficiency, the multimodal reinforcement learning intelligent agent system autonomously adjusts to achieve the best results. , , , .
[0154] The system collects environmental data (temperature, humidity, light intensity, CO2 concentration, etc.) in real time through intelligent sensors; it also uses a multimodal reinforcement learning-based intelligent agent system to collaboratively solve a dual-layer control model for the market demand response of large public buildings, electricity, carbon, and green certificates, and sends control signals to the equipment based on the solution results.
[0155] The key adjustable equipment is then controlled based on the model solution results.
[0156] Step S53 involves maintaining the multimodal reinforcement learning agent system. Specifically, this involves monitoring energy usage within the building in real-time using smart sensors, and further optimizing and maintaining the behavior of the multimodal reinforcement learning agent system based on feedback information. This includes, but is not limited to: regularly inspecting and maintaining the system to ensure optimal performance of the agent during long-term use; regularly evaluating the system's performance, checking indicators such as building energy efficiency, user comfort, and grid load balance; and regularly analyzing system operation logs to assess the agent's decision-making effectiveness and response speed.
[0157] The beneficial effects of this invention are:
[0158] Optimize building energy management, especially demand response systems for buildings, by leveraging multimodal data and reinforcement learning.
[0159] (1) This invention improves the demand response accuracy of large public buildings in the electricity-carbon-green certificate market by combining multimodal data fusion with reinforcement learning. The method of this invention can make full use of multimodal data inside the building (such as indoor temperature and humidity, CO2 concentration, light intensity, equipment operating status, etc.) and external market dynamics (such as electricity prices, carbon trading prices and green certificate prices), and fuse multimodal data through deep learning technology to provide high-quality input features. The intelligent agent system based on multimodal reinforcement learning can adapt to complex and ever-changing environments, accurately optimize demand response strategies, and significantly improve the accuracy and flexibility of energy regulation.
[0160] (2) This invention utilizes a multimodal reinforcement learning intelligent agent system collaborative mechanism. The upper-level intelligent agent system optimizes the building's trading strategies in the electricity market, carbon trading market, and green certificate market in real time, while the lower-level intelligent agent system dynamically regulates the building's internal equipment (such as air conditioning, lighting, and heating). The collaborative work of the upper-level and lower-level intelligent agent systems enables large public buildings to minimize operating costs and maximize economic benefits while meeting indoor comfort requirements.
[0161] (3) This invention deeply integrates energy regulation of large public buildings with electricity, carbon emissions, and green certificate markets, significantly reducing carbon emissions while optimizing the energy use efficiency of large public buildings, and providing intelligent solutions for green and low-carbon buildings. The method can effectively support the intelligent upgrading of large public buildings to participate in the demand response of the electricity-carbon-green certificate market. Attached Figure Description
[0162] Figure 1 This is a flowchart of the intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system, as described in this invention.
[0163] Figure 2 This is a schematic diagram of the multi-agent composition in the intelligent management and optimization method for large-scale building integrated energy systems based on large language models, as described in this invention.
[0164] Figure 3 This is a schematic diagram of the single-agent construction process of the intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system, as described in this invention.
[0165] Figure 4 This is a schematic diagram of a large-scale public building demand response intelligent control method based on a multimodal reinforcement learning intelligent agent system, as described in this invention. Detailed Implementation
[0166] like Figure 1 This is a flowchart of the intelligent control method for demand response of large airport buildings based on a multimodal reinforcement learning intelligent agent system, according to the present invention.
[0167] As airports are typical public buildings, this invention takes a large-scale integrated energy system for an airport located in a frigid region as an example to describe the method of this invention in detail, including the following steps:
[0168] S1. Based on the operation and maintenance characteristics of large airport public buildings, determine the scope of participation in demand response control of large airport public buildings and identify key controllable equipment; based on the scope of participation in demand response control and key controllable equipment, determine the types of intelligent agents involved and their respective functions to obtain a multimodal reinforcement learning intelligent agent system.
[0169] S2, using intelligent sensors to realize multimodal data perception and collection of large airport public buildings, constructing a multimodal data set D of large airport public buildings, and extracting features from set D to obtain a feature set F of large airport public buildings multimodal data, combining existing large language models and visual language models to initially construct each agent of the multimodal reinforcement learning agent system in step S1, and using deep reinforcement learning algorithms to initially train each initially constructed agent;
[0170] S3, based on the multimodal data feature set 𝐹 and the agents constructed in step S2, a multimodal reinforcement learning agent system is trained in a centralized manner using a deep learning reinforcement algorithm, and the collaborative control of the agents constructed in step S2 is realized based on a distributed execution mechanism;
[0171] S4, Deploy the multimodal reinforcement learning agent system into the large airport public building;
[0172] S5. The multimodal reinforcement learning intelligent agent system participates in the intelligent regulation of demand response of large airport public buildings. The specific method is as follows: establish a two-layer regulation model of demand response of large airport public buildings in the electricity-carbon-green certificate market, collaboratively solve the two-layer regulation model of demand response of large airport public buildings in the electricity-carbon-green certificate market based on the multimodal reinforcement learning intelligent agent system, and regulate the key controllable equipment in step S1 based on the model solution results; and maintain the multimodal reinforcement learning intelligent agent system.
[0173] In step S1, the scope of demand response regulation for large airport public buildings includes electricity market demand response regulation, carbon market demand response, and green certificate market demand response regulation. Key controllable equipment includes cold and heat source power equipment, central air conditioning equipment, energy storage equipment, lighting equipment, and renewable energy equipment.
[0174] like Figure 2 As shown, the multimodal reinforcement learning agent system includes an upper-layer agent system, a lower-layer agent system, and a communication and coordination agent system; the upper-layer agent system and the lower-layer agent system operate in global coordination through the communication and coordination agent system.
[0175] The upper-level intelligent agent system includes an intelligent agent for demand response in the electricity market, an intelligent agent for demand response in the carbon market, and an intelligent agent for demand response in the green certificate market;
[0176] The lower-level intelligent agent system includes environmental perception intelligent agents, energy management intelligent agents, user behavior intelligent agents, fault detection and maintenance intelligent agents, policy generation intelligent agents, and multimodal data fusion intelligent agents.
[0177] Each intelligent agent possesses different functional characteristics and can be extended to build corresponding intelligent agents according to different scenario requirements.
[0178] The functions of the upper-layer intelligent agent system are as follows:
[0179] 1) The electricity market demand response intelligent agent minimizes electricity costs by optimizing electricity procurement and electricity consumption control strategies, thereby supporting grid stability and improving building economic efficiency;
[0180] 2) Carbon market demand response agents are responsible for monitoring and optimizing building carbon emissions, reducing carbon emission costs and promoting the achievement of green and low-carbon goals by participating in carbon quota trading and emission reduction strategy formulation;
[0181] 3) The green certificate market demand response intelligent agent improves the utilization rate of renewable energy by rationally planning the purchase and sale of green certificates, while ensuring policy compliance and optimizing transaction costs.
[0182] The function of the lower-level intelligent agent system is as follows:
[0183] 4) The environmental sensing agent is responsible for collecting and analyzing environmental data in real time, including information such as temperature, humidity, and light intensity, to provide accurate basic data support for demand response regulation;
[0184] 5) The user behavior intelligent agent analyzes user behavior patterns (such as working hours, activity areas, etc.) to intelligently regulate the indoor environment and dynamically adjust energy distribution and equipment operation to meet user needs;
[0185] 6) The multimodal data fusion intelligent agent is responsible for processing multi-source heterogeneous data (images, audio, video, etc.) from large airport public buildings, including environmental data, equipment operation status data, user behavior data, and external market dynamic information. Through efficient data fusion and feature extraction methods, it provides comprehensive and accurate decision support for intelligent demand response regulation.
[0186] 7) The energy management agent is responsible for monitoring and optimizing the building's energy consumption, such as electricity and heat, and developing efficient energy use strategies to reduce energy consumption.
[0187] 8) The strategy-generating agent uses deep reinforcement learning algorithms to optimize intelligent decision-making based on multimodal data and preset targets, thereby realizing the dynamic adjustment and execution of demand response strategies.
[0188] 9) The fault detection and maintenance intelligent agent is responsible for monitoring the operating status of building equipment, promptly detecting and responding to potential faults, reducing energy waste, improving equipment operating efficiency, and extending service life.
[0189] The functions of the communication coordination intelligent agent system are:
[0190] 10) The communication and coordination intelligent agent system is responsible for information interaction and collaboration in the multi-agent system by coordinating the tasks and data flow of upper and lower-level intelligent agents, ensuring the consistency and synergy of the behavior of each intelligent agent.
[0191] Through the global collaboration of these intelligent agents, the entire system can achieve precise, efficient, and intelligent management of building demand response under complex and ever-changing environmental and market conditions.
[0192] It should be noted that the above outlines the general composition of a multimodal enhanced intelligent agent system for large airport public buildings, and corresponding intelligent agents can be constructed according to different scenario requirements.
[0193] like Figure 3 As shown, step S2 specifically includes the following steps:
[0194] Step S21 involves using intelligent sensors to achieve multimodal data perception and collection for large airport public buildings. These intelligent sensors include temperature and humidity sensors, gas sensors, infrared sensors, intelligent cameras, and radio frequency identification (RFID) sensors. Multimodal data for large airport public buildings refers to diverse types of data reflecting the operational status of various aspects of these buildings. This operational status includes environmental data, energy consumption data, indoor personnel behavior data, electricity market data, and carbon market data. The diverse data types include text, audio, video, and images, and are uploaded to the existing integrated energy intelligent management and control platform for large airport public buildings via wireless communication technology.
[0195] Among them, environmental data refers to indoor temperature and humidity, indoor CO2 concentration, and outdoor meteorological data (temperature, humidity, wind speed, etc.); energy consumption data refers to the real-time energy consumption of electrical equipment in various areas of large airport public buildings and the real-time energy consumption statistics of various cold and heat source equipment; indoor personnel behavior data refers to the real-time distribution characteristics of indoor personnel in large airport public buildings, including the number of personnel in each functional area, personnel flow, and real-time personnel load; electricity market data includes real-time information such as grid load, electricity price, load fluctuation, and grid stability; carbon market data includes real-time carbon trading volume, real-time carbon tax, carbon emissions, and real-time carbon emission factors.
[0196] Step S22: Construct a multimodal data set D for large airport public buildings. The multimodal data set D consists of historical data collected by the existing integrated energy intelligent management and control platform for large airport public buildings and existing knowledge sets for large airport public buildings. The existing knowledge sets for large airport public buildings include existing knowledge on electricity-carbon coupling demand response, green building energy efficiency optimization, energy system integration and management, renewable energy utilization, energy conversion and storage technologies, and intelligent energy monitoring and dispatching systems. The carriers of this existing knowledge can be e-books, existing standards, audio, video, or case collections formed during the operation and maintenance of large airport public buildings, or even voice conversations between operation and maintenance personnel.
[0197] The multimodal dataset D of the large airport public buildings is represented in tuple form as follows:
[0198]
[0199] Among them, D i Let represent the dataset for the i-th mode, and n be the total number of modes. The dataset for each mode... This can be further expressed as:
[0200]
[0201] in, It is the value of the k-th data point of the i-th mode at time t, m i It is the total number of data points for the i-th mode.
[0202] Step S23: Use intelligent algorithms to perform feature processing and fusion on the multimodal dataset D to obtain the feature set F of multimodal data of large airport public buildings. The feature processing and fusion specifically includes denoising, time alignment, standardization and normalization, feature extraction and feature fusion.
[0203] Among these, denoising refers to using algorithms such as mean filtering, median filtering, and Gaussian filtering to clean up noise and interference in various smart sensor data, thereby improving the quality of multimodal data; time alignment refers to using techniques such as timestamp synchronization, interpolation methods, or dynamic time warping to ensure that data from different modalities are consistent in time, that is, to ensure that the readings of different sensors are at the same point in time, such as ensuring the temporal correspondence between visual and sound information in video and audio data; standardization and normalization refers to standardizing the multimodal data to make the data dimensions consistent; feature extraction refers to using wavelet transform and Fourier transform to extract time-frequency domain feature information from the original data, such as electricity consumption patterns, temperature changes, and human activity patterns; feature fusion is the process of integrating the features obtained from the above feature extraction using techniques such as weighted averaging, splicing, and deep learning fusion, thereby forming a more comprehensive and robust feature representation.
[0204] The multimodal data feature set F of the large airport public buildings is represented in tuple form as follows:
[0205]
[0206] Among them, F i Let represent the set of data features for the i-th mode, and n be the total number of modes. The set of data features for each mode can be further represented as:
[0207]
[0208] Among them, f ik M is the value of the k-th feature of the i-th mode at time t. i This represents the total number of data features for the i-th modality. It should be noted that the above steps outline the general process for multimodal data feature processing and fusion in large airport public buildings. Appropriate algorithms and technologies need to be selected and adjusted according to the characteristics of different modal data and application scenarios.
[0209] Step S24, the steps for initially constructing individual intelligent agents for large airport public buildings based on existing large language models, visual language models, and deep reinforcement learning algorithms, specifically include:
[0210] Step S241 refers to providing the multimodal data feature set F of large airport public buildings as input features to the existing large language model and visual language model, and initially constructing the various agents in the multimodal reinforcement learning intelligent agent system (specifically including electricity market demand response agent, carbon market demand response agent, green certificate market demand response agent, environmental perception agent, energy management agent, multimodal data fusion agent, user behavior agent, fault detection and maintenance agent, and policy generation agent). Existing large language models include GPT-4, Gemini, etc., which have decision-making and reasoning abilities similar to humans; visual language models such as CLIP, Qwen2-VL, PaliGemma, etc., combine visual perception of images and videos with text-based reasoning to process and mine complex multimodal data.
[0211] Step S242: Use a deep reinforcement learning algorithm to perform preliminary training on each agent;
[0212] Parameter definition, including defining the global state. The local state space of the i-th single agent The action space of the i-th single agent and the reward function of the i-th single agent .
[0213] Global state It is used to describe the operational status of the entire large airport public building, specifically including real-time status information of ambient temperature, ambient humidity, equipment operating status, personnel activity level, electricity market, carbon market, and green certificate market.
[0214] The local state of the i-th single agent is Its observed local state space It is a subset of the global state, that is:
[0215]
[0216] Where T represents ambient temperature, H represents ambient humidity, E represents equipment operating status, and P represents personnel activity level. For electricity market transaction volume, For carbon market trading volume, For green certificate market trading volume;
[0217] All possible actions performed by the i-th single agent The set is represented as the action space. ,Right now:
[0218]
[0219] In the formula, For the joint action of the i-th agent, Represents the temperature regulation capacity of public buildings in large airports. Represents the lighting adjustment amount of public buildings in large airports. Represents the amount of heating regulation. For the trading volume of large airport public buildings in the electricity market, For the trading volume of large airport public buildings in the carbon market, The transaction volume of large airport public buildings in the green certificate market.
[0220] reward function It is defined based on the current state and actions performed by the i-th single agent, and is used to guide the learning process of the agent.
[0221]
[0222] In the formula, It is a reward based on energy consumption. It is a reward based on indoor comfort. It is a reward for participating in the electricity market demand response. It is a reward for participating in carbon market trading. It is a reward for participating in the green certificate market.
[0223] The essence of training the i-th agent using deep reinforcement learning algorithms is learning. The strategy corresponding to maximizing expected cumulative reward Specifically, it is represented by the following mathematical model:
[0224]
[0225] In the formula, It is the learning strategy of the i-th agent. Performance metrics; The expected value is represented by the initial state s0 and the action. The probability distributions (these distributions are determined by the policy) (Determined) Calculated to obtain; ρ(⋅∣ ) is in the strategy The state distribution under; It is the discount factor for the i-th agent at time t, used to balance the importance of immediate rewards and future rewards. .
[0226] The agent operates in a simulated environment and adjusts its behavioral strategies based on feedback. At each step, the agent selects an action and receives a reward based on the action's effects (such as energy consumption, comfort, etc.). The simulated environment refers to a virtual or computational model used to test and train the agent's behavior. This environment maps certain physical characteristics of the real world as closely as possible, allowing the agent to interact within it and learn how to achieve specific goals.
[0227] In step S3, based on the complexity of large public buildings, a global reward function is constructed using a distributed execution mechanism and the multimodal data feature set 𝐹. This enables the collaborative control of the multimodal reinforcement learning-based intelligent agent system. For example... Figure 4 As shown, step S3 specifically includes the following steps:
[0228] Step S31: Establish the global joint action of the multimodal reinforcement learning agent system. ;
[0229] Global Joint Actions of Multimodal Reinforcement Learning Agent Systems Represented as:
[0230]
[0231] in, This represents the joint action of the i-th agent;
[0232] Constructing the global action space of a multimodal reinforcement learning agent system Global action space of multimodal reinforcement learning intelligent agent system It is the Cartesian product of the action spaces of all intelligent agents, specifically expressed as:
[0233]
[0234] in, Let i represent the action space of the i-th single agent.
[0235] Step S32, define union function , Indicates the global state and global joint actions The reward for the i-th agent:
[0236] =
[0237] in, These are the Q-network parameters (such as weights and biases) of the i-th agent.
[0238] Policy network of the i-th agent By optimizing the gradient of the policy network Make the objective function Maximize, policy gradient Represented as:
[0239]
[0240] in, Represents the gradient. It is the gradient of the joint Q-function with respect to the action of the i-th agent; These are the parameters of the policy network; It is a policy network Regarding parameters The gradient.
[0241] Joint Q function Through the mean square error function renew:
[0242]
[0243] Among them, the target value for:
[0244]
[0245] This indicates the global joint action of the multimodal reinforcement learning agent system in the next state; The next state represents the global state of the multimodal reinforcement learning agent system. Let be the Q-network parameters of the i-th agent in the next state.
[0246] Cooperative control of multimodal reinforcement learning agent systems through global objectives Represented as:
[0247]
[0248] in, Represents the global expected value. It is the global reward function, defined as the weighted sum of the rewards of all agents:
[0249]
[0250] It is the weight of the i-th agent.
[0251] During the training phase, the agent's policy and Q-function are trained centrally. The distributed update rule for the policy of each agent i is as follows:
[0252]
[0253] The parameter update rule for the Q function is as follows:
[0254]
[0255] in, This is the learning rate.
[0256] Step S33: In distributed execution, the agent bases its actions on local observations. and strategy Decision-making actions are jointly implemented to achieve the coordinated operation of various agents in a multimodal reinforcement learning agent system.
[0257] Based on multimodal reinforcement learning intelligent agent systems, multi-scenario demand response control strategies can be designed, including dynamic equipment scheduling, energy consumption optimization, and user comfort assurance.
[0258] In step S4, the trained multimodal reinforcement learning intelligent agent system is deployed into the actual integrated energy intelligent management and control platform, and integrated with the building's control systems (such as air conditioning, lighting, heating systems, etc.). This includes the upper-level intelligent agent system being deployed in the integrated energy intelligent management and control platform to handle electricity-carbon-green certificate market transactions; and the lower-level intelligent agent being embedded in the equipment controller to dynamically regulate each device. Specifically, this includes the following steps:
[0259] Step S41: The upper-layer intelligent agent system integrates with the integrated energy intelligent management and control platform. The data interface can connect with the building management system (BMS) and energy market platform through protocols (such as OPC UA, Modbus, or RESTful API) to obtain market prices (electricity, carbon, green certificates), building energy consumption forecast data, and historical transaction record information. On the other hand, it submits transaction requests to the energy market through the integrated energy intelligent management and control platform, such as green certificate purchase, carbon emission quota trading, and electricity procurement. Then, it dynamically adjusts the operation of the lower-layer intelligent agent system based on the transaction results.
[0260] Step S42, Integration of the lower-level intelligent agent system with the device controller, Embedded deployment: The trained lower-level intelligent agent is embedded in the device controller, including PLC (Programmable Logic Controller) and IoT devices.
[0261] Step S5 specifically includes the following steps:
[0262] Step S51: Establish a two-tiered regulatory model for the demand response of large airport public buildings to the electricity-carbon-green certificate market, specifically including an upper-level economic model and a lower-level day-ahead scheduling model.
[0263] The upper-level economic model maximizes profits by having large airport public buildings participate in the demand response of the electricity market, carbon market, and green certificate market. The objective is to establish a higher-level economic model. Specifically, the upper-level economic model is as follows:
[0264]
[0265] In the formula, The electricity purchase price for large airport public building b at time t. The marginal cost of large airport public building b at time t. This refers to the electricity consumption of large airport public building b at time t. The price of green certificates for large airport public buildings (b) at time t. This refers to the volume of green certificate transactions purchased by large airport public building b at time t. The carbon price of a large airport public building (b) at time t. It is the carbon emissions of a large airport public building b at time t;
[0266] The lower-level day-ahead dispatch model minimizes the cost of various energy devices in the demand response of large airport public buildings participating in the electricity market, carbon market, and green certificate market. The objective involves determining the market clearing price and the trading volume of each market participant to achieve supply and demand balance in the carbon electricity market and the green certificate market. The lower-level day-ahead scheduling model is as follows:
[0267]
[0268] in, This represents the day-ahead scheduling cost of large airport public building b at time t. These are the day-ahead scheduling decision variables for each key controllable device in the public building b of a large airport at time t. This represents the total electricity market cost of a large airport public building at time t, including electricity purchase cost and demand response cost. This represents the total carbon market cost of a large airport public building b at time t, which may include the cost of purchasing carbon allowances. This represents the total cost or benefit of the green certificate market for a large airport public building b at time t, depending on whether it purchases or sells green certificates.
[0269] This can be expressed as:
[0270]
[0271] This refers to the day-ahead scheduling decisions for the cold and heat source power equipment in public buildings (b) of large airports. This refers to the day-ahead scheduling decision of the air conditioning equipment in public building b of a large airport; This refers to the day-ahead scheduling decision-making of energy storage equipment in public buildings b of large airports; This refers to the day-ahead scheduling decision for lighting equipment in public buildings (b) of large airports; This refers to the day-ahead scheduling decisions for renewable energy equipment in public buildings (b) of large airports.
[0272] The constraints of the lower-level day-ahead dispatch model include power supply and demand balance constraints, carbon emission limit constraints, green certificate trading limit constraints, and energy equipment output constraints.
[0273] Electricity supply and demand balance constraints:
[0274]
[0275] in, This refers to the electricity purchases of large airport public building b at time t. It represents the electricity sales volume of large airport public building b at time t. This refers to the electricity purchased or sold by large airport public buildings (b). This means that it holds true for all times t.
[0276] Carbon emission limits and constraints:
[0277]
[0278] in, This refers to the carbon emissions of a large airport public building (b) at time t. It is the carbon allowance for public buildings b at large airports.
[0279] Restrictions on Green Certificate Trading:
[0280]
[0281] in, It represents the green certificate transaction volume of large airport public building b at time t. It is the renewable energy generation of the large airport public building b at time t.
[0282] Energy equipment output constraints:
[0283]
[0284] and These refer to the lower and upper limits of the output of the i-th energy device in a large airport public building, respectively.
[0285] The aforementioned constraints ensure that when large airport public buildings participate in the electricity-carbon market, their electricity purchases meet their electricity demand, carbon emissions do not exceed the prescribed quotas, and green certificate trading volume matches building renewable energy trading. These constraints are further integrated into the objective function of the lower-level model to ensure that large airport public buildings participate reasonably in the electricity-carbon-green certificate market demand response.
[0286] Step S52: Solve the upper-level economic model and the lower-level day-ahead scheduling model using the upper-level intelligent agent system and the lower-level intelligent agent system respectively. Then, achieve information interaction and collaborative solving between the upper and lower-level intelligent agent systems through communication coordination. Solve the upper-level economic model to obtain... , , Solve the lower-level day-ahead scheduling model to obtain .
[0287] By establishing the upper-level mean square error function of the upper-level intelligent agent system Objective function of the upper-level intelligent agent system The policy network gradient of the upper-level intelligent agent system Then, by solving the upper-level economic model, we can obtain... , , Specifically:
[0288] Upper-layer intelligent agent system upper-layer mean square error function for,
[0289]
[0290]
[0291] In the formula, The mean square error function of the upper layer of the upper-layer intelligent agent system; This represents the expected value of the upper-level intelligent agent system. For the joint Q-function of the upper-level agent system in a multimodal reinforcement learning agent system; This represents the global state of the upper-level intelligent agent system. For the joint actions of the upper-level intelligent agent system, it is represented as ; For adjustment The learning parameters of the function; The target value for the upper-level intelligent agent system; The reward function established by the upper-level intelligent agent system based on the upper-level economic model, i.e. ; The discount factor of the upper-level intelligent agent system is used to balance the rewards of the upper-level intelligent agent system; This is the joint Q-function of the lower-level agent system in the multimodal reinforcement learning agent system in the next state.
[0292] Objective function of upper-level intelligent agent system for:
[0293]
[0294] Furthermore, the policy update process of the upper-layer intelligent agent system is based on the gradient of the policy network of the upper-layer intelligent agent system. Implementation, specifically:
[0295]
[0296] In the formula, The objective function of the upper-level intelligent agent system Regarding its strategy network parameters The policy network gradient is used to guide policy updates to maximize... ; It refers to the network parameters of the upper-level intelligent agent system regarding its policy. Expectations; These are the policy network parameters for the upper-level intelligent agent system. For the upper-layer intelligent agent system policy network; For the policy network of the upper-level intelligent agent system The gradient; This is the gradient of the joint Q-function of the upper-level intelligent agent system.
[0297] By establishing the lower-level mean square error function of the lower-level intelligent agent system Objective function of the lower-level intelligent agent system The policy network gradient of the lower-level intelligent agent system Solving the lower-level day-ahead scheduling model yields... Specifically:
[0298] Lower-level intelligent agent system upper-level mean square error function for,
[0299]
[0300]
[0301] In the formula, The mean square error function of the lower-level intelligent agent system; This represents the expected value of the lower-level intelligent agent system. For the joint Q-function of the lower-level agent system in a multimodal reinforcement learning agent system; This represents the global state of the lower-level intelligent agent system. For the joint actions of the lower-level intelligent agent system, it is represented as ; For adjustment The learning parameters of the function; The target value for the lower-level intelligent agent system; The reward function established by the lower-level intelligent agent system based on the lower-level day-ahead scheduling model, i.e. ; The discount factor of the lower-level intelligent agent system is used to balance the rewards of the lower-level intelligent agent system; This is the joint Q-function of the lower-level agent system in the multimodal reinforcement learning agent system in the next state.
[0302] Objective function of the lower-level intelligent agent system for:
[0303]
[0304] Furthermore, the policy update process of the upper-layer intelligent agent system is based on the policy network gradient of the lower-layer intelligent agent system. Implementation, specifically:
[0305]
[0306] In the formula, The objective function of the upper-level intelligent agent system Regarding its strategy network parameters The policy network gradient is used to guide policy updates to maximize... ; It refers to the network parameters of the lower-level intelligent agent system regarding its policy. Expectations; These are the policy network parameters for the lower-level intelligent agent system. For the policy network of the lower-level intelligent agent system; Policy network for lower-level intelligent agent systems The gradient; This represents the gradient of the joint Q-function of the lower-level intelligent agent system.
[0307] Based on the model solution results , , , Adjust the aforementioned key adjustable equipment;
[0308] Upper-layer agents output market trading strategies (such as electricity procurement and carbon emission allowances); lower-layer agents provide feedback on equipment operation results and environmental conditions (such as actual energy consumption, carbon emissions, and indoor comfort). Based on real-time market fluctuations and equipment operating status, both upper and lower-layer agents adjust their respective strategies.
[0309] In practical applications, the upper-level intelligent agent system responds to and regulates demand in real time based on real-time monitoring of environmental conditions (including grid load, electricity price, and building energy consumption). For example, during peak grid periods, the lower-level intelligent agent system can reduce building energy consumption by reducing air conditioning load and delaying the operation of non-critical equipment.
[0310] The system dynamically adjusts and optimizes the behavior of the lower-level intelligent agent system based on real-time feedback: if a certain control strategy leads to reduced user comfort or lack of energy efficiency, the multimodal reinforcement learning intelligent agent system autonomously adjusts to achieve the best results. , , , .
[0311] The system collects environmental data (temperature, humidity, light intensity, CO2 concentration, etc.) in real time through intelligent sensors; it also uses a multimodal reinforcement learning intelligent agent system to collaboratively solve a dual-layer control model for the market demand response of large airport public buildings, electricity, carbon, and green certificates, and sends control signals to the equipment based on the solution results.
[0312] The key adjustable equipment is then controlled based on the model solution results.
[0313] Step S53 involves maintaining the multimodal reinforcement learning agent system. Specifically, this involves monitoring energy usage within the building in real-time using smart sensors, and further optimizing and maintaining the behavior of the multimodal reinforcement learning agent system based on feedback information. This includes, but is not limited to: regularly inspecting and maintaining the system to ensure optimal performance of the agent during long-term use; regularly evaluating the system's performance, checking indicators such as building energy efficiency, user comfort, and grid load balance; and regularly analyzing system operation logs to assess the agent's decision-making effectiveness and response speed.
Claims
1. A method for intelligent demand response control of large public buildings based on a multimodal reinforcement learning intelligent agent system, characterized in that, Includes the following steps: S1. Based on the operation and maintenance characteristics of large public buildings, determine the scope of large public buildings participating in demand response control and identify key controllable equipment; based on the scope of participation in demand response control and key controllable equipment, determine the types of intelligent agents involved and their respective functions to obtain a multimodal reinforcement learning intelligent agent system. S2, using intelligent sensors to realize the perception and acquisition of multimodal data of large public buildings, constructing a multimodal data set D of large public buildings, and extracting features from set D to obtain a feature set F of multimodal data of large public buildings; By combining existing large language models and visual language models, the agents of the multimodal reinforcement learning agent system in step S1 are initially constructed, and the agents are initially trained using deep reinforcement learning algorithms. S3, based on the multimodal data feature set 𝐹 and the agents constructed in step S2, a multimodal reinforcement learning agent system is trained in a centralized manner using a deep learning reinforcement algorithm, and the collaborative control of the agents constructed in step S2 is realized based on a distributed execution mechanism; S4, Deploy the multimodal reinforcement learning agent system into the large public building; S5. The multimodal reinforcement learning intelligent agent system participates in the intelligent regulation of demand response of large public buildings. The specific method is as follows: establish a two-layer regulation model of demand response of large public buildings in the electricity-carbon-green certificate market, collaboratively solve the two-layer regulation model of demand response of large public buildings in the electricity-carbon-green certificate market based on the multimodal reinforcement learning intelligent agent system, and regulate the key controllable equipment in step S1 based on the model solution results; and maintain the multimodal reinforcement learning intelligent agent system. Step S2 includes the following steps: Step S21: Employ intelligent sensors to achieve multimodal data perception and acquisition for large public buildings. These intelligent sensors include temperature sensors, gas sensors, infrared sensors, intelligent cameras, and radio frequency identification (RFID) sensors. Multimodal data for large public buildings refers to diverse types of data reflecting the operational status of various aspects of the building. This operational status includes environmental data, energy data, indoor occupant behavior data, electricity market data, and carbon market data. The diverse data types include text, audio, video, and image multimodal data, which are uploaded to the existing integrated energy intelligent management and control platform for large public buildings via wireless communication technology. Step S22: Construct a multimodal data set D for large public buildings; the multimodal data set D consists of historical data collected by the existing integrated energy intelligent management and control platform for large public buildings and a set of existing knowledge for large public buildings; the set of existing knowledge for large public buildings includes existing knowledge on electricity-carbon coupling demand response, green building energy efficiency optimization, energy system integration and management, covering renewable energy utilization, energy conversion and storage technologies, and intelligent energy monitoring and dispatching systems. The carriers of this existing knowledge are electronic books, existing standards, audio, video, or case collections formed during the operation and maintenance of large public buildings, or voice conversations between operation and maintenance personnel. The multimodal dataset D of the large public building is represented in the form of tuples: Among them, D i Let n represent the dataset of the i-th mode, where n is the total number of modes; and let n represent the dataset of each mode. Further expressed as: in, It is the value of the k-th data point of the i-th mode at time t, m i It is the total number of data points in the i-th mode; Step S23: Utilize intelligent algorithms to perform feature processing and fusion on the multimodal dataset D, specifically including denoising, time alignment, standardization and normalization, feature extraction, and feature fusion, to obtain the feature set F of the multimodal data of large public buildings. F is represented in tuple form. Among them, F i Let F represent the set of data features for the i-th mode, where n is the total number of modes; F represents the set of data features for the i-th mode. i Represented as: in, M is the value of the k-th feature of the i-th mode at time t. i It is the total number of data features of the i-th modality; Step S24 involves initially constructing individual agents for large public buildings based on existing large language models, visual language models, and deep reinforcement learning algorithms. Specifically: First, the multimodal data feature set F of large public buildings is used as input features and provided to the existing large language model and visual language model to initially construct each agent in the multimodal reinforcement learning agent system; Then, deep reinforcement learning algorithms are used to perform preliminary training on each agent. The specific method is as follows: Define parameters: global state The local state space observed by the i-th single agent The action space of the i-th single agent and the reward function of the i-th single agent ; Global state It is used to describe the operation and maintenance status of the entire large public building, specifically including real-time status information of ambient temperature, ambient humidity, equipment operating status, personnel activity level, electricity market, carbon market, and green certificate market; The local state of the i-th single agent is Its observed local state space It is a subset of the global state, that is: Where T represents ambient temperature, H represents ambient humidity, E represents equipment operating status, and P represents personnel activity level. For electricity market information, For carbon market information, Information on the green certificate market; All possible actions performed by the i-th single agent The set is represented as the action space. ,Right now: In the formula, For the joint action of the i-th agent, Represents the amount of environmental temperature regulation in large public buildings. Represents the lighting adjustment amount of large public buildings. Represents the heating regulation capacity of large public buildings. For the trading volume of large public buildings in the electricity market, For the trading volume of large public buildings in the carbon market, The transaction volume of large public buildings in the green certificate market; reward function It is defined based on the current state and actions performed by the i-th single agent, and is used to guide the learning process of the agent; In the formula, It is a reward based on energy consumption. It is a reward based on indoor comfort. It is a reward for participating in the electricity market demand response. It is a reward for participating in carbon market trading. It is a reward for participating in the green certificate market trading; The essence of training the i-th agent using a deep reinforcement learning algorithm is learning. The strategy corresponding to maximizing expected cumulative reward Specifically, it is represented by the following mathematical model: In the formula, It is the learning strategy of the i-th agent. Performance metrics; The expected value is represented by the initial state s0 and the action. The probability distribution is calculated from ρ(⋅∣). ) is in the strategy The state distribution under; It is the discount factor for the i-th agent at time t, used to balance the importance of immediate rewards and future rewards.
2. The intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system according to claim 1, characterized in that, In step S1: The scope of demand response regulation for the large public buildings mentioned specifically includes demand response regulation in the electricity market, carbon market, and green certificate market. The key controllable equipment includes cold and heat source power equipment, central air conditioning equipment, energy storage equipment, lighting equipment, and renewable energy equipment. The multimodal reinforcement learning agent system includes an upper-layer agent system, a lower-layer agent system, and a communication and coordination agent system; the upper-layer agent system and the lower-layer agent system operate in global coordination through the communication and coordination agent system. The upper-layer intelligent agent system includes an intelligent agent for electricity market demand response, an intelligent agent for carbon market demand response, and an intelligent agent for green certificate market demand response; the lower-layer intelligent agent system includes an intelligent agent for environmental perception, an intelligent agent for energy management, an intelligent agent for user behavior, an intelligent agent for fault detection and maintenance, an intelligent agent for strategy generation, and an intelligent agent for multimodal data fusion.
3. The intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system according to claim 1, characterized in that, Step S3 specifically includes the following steps: Step S31: Establish the global joint action of the multimodal reinforcement learning agent system. , is represented as: in, This represents the joint action of the i-th agent; Constructing the global action space of a multimodal reinforcement learning agent system , The Cartesian product of the action space of all agents, specifically represented as in, This represents the action space of the i-th agent; Step S32, define union function , Indicates the global state and global joint actions The reward for the i-th agent: = in, This represents the Q-network parameters of the i-th agent; Policy network of the i-th agent Through policy gradient Optimize, policy gradient Represented as: in, Represents the gradient. It is the gradient of the joint Q-function with respect to the action of the i-th agent; These are the parameters of the policy network; It is a policy network Regarding parameters The gradient; Joint Q function Through mean square error renew: Where the target value for: in, This indicates the global joint action of the multimodal reinforcement learning agent system in the next state; The next state represents the global state of the multimodal reinforcement learning agent system. The Q-network parameters for the i-th agent in the next state; This represents the discount factor for the i-th agent; In a multimodal reinforcement learning agent system, the collaborative control of each agent is achieved through a global objective. Represented as: in, Represents the global expected value. It is the global reward function, defined as the weighted sum of the rewards of all agents: It is the weight of the i-th agent; During the training phase, the agent's policy and Q-function are trained and optimized in a concentrated manner; the distributed update rule for the policy of each agent i is as follows: The parameter update rule for the Q function is as follows: in, The learning rate; Step S33: In distributed execution, the agent bases its actions on local observations. and strategy Decision-making actions are jointly implemented to achieve the coordinated operation of various agents in a multimodal reinforcement learning agent system.
4. The intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system according to claim 3, characterized in that, Step S4 includes the following steps: Step S41: The upper-level intelligent agent system integrates with the integrated energy intelligent management and control platform. The data interface connects with the building management system and energy market platform through protocols to obtain market prices, building energy consumption forecast data, and historical transaction record information. On the other hand, it submits transaction requests to the energy market through the integrated energy intelligent management and control platform. Then, it dynamically adjusts the operation of the lower-level intelligent agent system based on the transaction results. Step S42, Integration of the lower-level intelligent agent system with the device controller, embedded deployment: The trained lower-level intelligent agent is embedded in the device controller.
5. The intelligent control method for demand response of large public buildings based on a multimodal reinforcement learning intelligent agent system according to claim 4, characterized in that, Step S5 includes the following steps: Step S51: Establish a two-tiered regulatory model for the demand response of large public buildings in the electricity-carbon-green certificate market, including an upper-level economic model and a lower-level day-ahead scheduling model. The upper-level economic model maximizes profits by involving large public buildings in the electricity market, carbon market, and green certificate market demand response. The goal is specifically: In the formula, The electricity purchase price for large public building b at time t. It is the marginal cost of large public building b at time t. This refers to the electricity consumption of large public building b at time t. The price of a green certificate for a large public building (b) at time t. This refers to the volume of green certificate transactions purchased by large public building B at time t. The carbon price of large public building b at time t. It is the carbon emissions of large public building b at time t; The lower-level day-ahead dispatch model minimizes the cost of various energy devices in the demand response of large public buildings participating in the electricity market, carbon market, and green certificate market. The goal is specifically: in, This represents the day-ahead scheduling cost of large public building b within time t. These are the day-ahead scheduling decision variables for each key controllable device in a large public building b within time t; This represents the total cost of electricity in the electricity market for a large public building b within time t, including electricity purchase costs and demand response costs. This represents the total carbon market cost of large public building b within time t, including the cost of purchasing carbon allowances. This represents the total cost or benefit of the green certificate market for a large public building b during time t, depending on whether it buys or sells green certificates. Represented as: in, This refers to the day-ahead scheduling decision of the cold and heat source power equipment in large public building b; This refers to the day-ahead scheduling decision of air conditioning equipment in large public building b; This refers to the day-ahead scheduling decision of energy storage equipment in large public building b; This refers to the day-ahead scheduling decision for lighting equipment in a large public building (b). This refers to the day-ahead scheduling decisions for renewable energy equipment in large public buildings (b). The constraints of the lower-level day-ahead dispatch model include power supply and demand balance constraints, carbon emission limit constraints, green certificate trading limit constraints, and energy equipment output constraints. The constraints on the balance between electricity supply and demand are: in, This represents the amount of electricity purchased by large public building b at time t. It represents the electricity sales volume of large public building b at time t. This refers to the electricity purchased or sold by large public building B. This means that it holds true for all times t; Carbon emission limits are: in, This refers to the carbon emissions of a large public building (b) at time t. It is the carbon allowance for large public building b; The restrictions on green certificate trading are as follows: in, It represents the volume of green certificate transactions for large public building B at time t. It is the renewable energy generation of a large public building b at time t; Energy equipment output constraints: and These refer to the lower and upper limits of the output of the i-th energy device in a large public building, respectively. Step S52: Solve the upper-level economic model and the lower-level day-ahead scheduling model using the upper-level intelligent agent system and the lower-level intelligent agent system respectively. Then, achieve information interaction and collaborative solving between the upper and lower-level intelligent agent systems through communication and coordination of the intelligent agent system. Solve the upper-level economic model to obtain... , , Solve the lower-level day-ahead scheduling model to obtain ; By establishing the mean square error function of the upper-level intelligent agent system Objective function of the upper-level intelligent agent system The policy network gradient of the upper-level intelligent agent system Then, by solving the upper-level economic model, we can obtain... , , Specifically: Mean square error function of upper-level intelligent agent system for, In the formula, The mean square error function value of the upper layer of the upper-layer intelligent agent system; This represents the expected value of the upper-level intelligent agent system. For the joint Q-function of the upper-level agent system in a multimodal reinforcement learning agent system; This represents the global state of the upper-level intelligent agent system. For the joint actions of the upper-level intelligent agent system, it is represented as ; For adjustment The learning parameters of the function; The target value for the upper-level intelligent agent system; The reward function established by the upper-level intelligent agent system based on the upper-level economic model, i.e. ; The discount factor of the upper-level intelligent agent system is used to balance the rewards of the upper-level intelligent agent system; For the joint Q-function of the upper-level agent system in the multimodal reinforcement learning agent system in the next state; Objective function of upper-level intelligent agent system for: The policy update process of the upper-layer intelligent agent system is based on the gradient of the policy network of the upper-layer intelligent agent system. Implementation, specifically: In the formula, The objective function of the upper-level intelligent agent system Regarding its strategy network parameters The policy network gradient is used to guide policy updates to maximize... ; It refers to the network parameters of the upper-level intelligent agent system regarding its policy. Expectations; These are the policy network parameters for the upper-level intelligent agent system. For the upper-layer intelligent agent system policy network; For the policy network of the upper-level intelligent agent system The gradient; The gradient of the joint Q-function of the upper-level intelligent agent system; By establishing the lower-level mean square error function of the lower-level intelligent agent system Objective function of the lower-level intelligent agent system The policy network gradient of the lower-level intelligent agent system Solving the lower-level day-ahead scheduling model yields... Specifically: Lower-level intelligent agent system upper-level mean square error function for, In the formula, The mean square error function of the lower-level intelligent agent system; This represents the expected value of the lower-level intelligent agent system. For the joint Q-function of the lower-level agent system in a multimodal reinforcement learning agent system; This represents the global state of the lower-level intelligent agent system. For the joint actions of the lower-level intelligent agent system, it is represented as ; For adjustment The learning parameters of the function; The target value for the lower-level intelligent agent system; The reward function established by the lower-level intelligent agent system based on the lower-level day-ahead scheduling model, i.e. ; The discount factor of the lower-level intelligent agent system is used to balance the rewards of the lower-level intelligent agent system; For the joint Q-function of the lower-level agent system in the multimodal reinforcement learning agent system in the next state; Objective function of the lower-level intelligent agent system for: The policy update process of the lower-level intelligent agent system is based on the gradient of the policy network of the lower-level intelligent agent system. Implementation, specifically: In the formula, The objective function of the upper-level intelligent agent system Regarding its strategy network parameters The policy network gradient is used to guide policy updates to maximize... ; It refers to the network parameters of the lower-level intelligent agent system regarding its policy. Expectations; These are the policy network parameters for the lower-level intelligent agent system. For the policy network of the lower-level intelligent agent system; Policy network for lower-level intelligent agent systems The gradient; The gradient of the joint Q-function of the lower-level intelligent agent system; Based on the model solution results , , , Adjust the aforementioned key adjustable equipment; Step S53: Maintain the multimodal reinforcement learning agent system. Specifically, monitor the energy usage in the building in real time using smart sensors, and further optimize and maintain the behavior of the multimodal reinforcement learning agent system based on feedback information.
Citation Information
Patent Citations
Micro-grid energy optimization method and system, electronic equipment and medium
CN116885799A
Household energy demand response optimization method and system based on deep reinforcement learning
CN117057553A
Energy consumption monitoring and optimizing method and system based on large model and multiple agents
CN118916778A