Micro energy network collaborative optimization method based on multi-agent deep reinforcement learning

By establishing a collaborative framework of source, grid, load, and storage based on multi-agent deep reinforcement learning, the problems of weak model adaptability and communication resource overload in micro energy grids are solved, and efficient multi-objective optimization and stable scheduling under carbon constraints are achieved.

CN121749209APending Publication Date: 2026-03-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-27

Smart Images

  • Figure CN121749209A_ABST
    Figure CN121749209A_ABST
Patent Text Reader

Abstract

The invention relates to a comprehensive energy system optimization control technology, in particular to a micro-energy grid collaborative optimization method based on multi-agent deep reinforcement learning, which comprises the following steps: calculating operation cost and carbon emission according to a source domain model, an energy storage model and a load model; a source side-renewable agent, a source side-adjustable device agent, a network side agent and an energy storage agent are arranged in the system, updating is carried out by taking global optimum as a target during training, and a local agent selects an optimal action by combining a local state of the local agent with local safety punishment. And storing the global state, the local state of each agent, the joint action of all agents, the global reward and the next global state in an experience pool as experience. According to the method, cost reduction, carbon reduction and efficiency improvement are taken into consideration under the scene of renewable output fluctuation and price / carbon parameter time varying, and the method has the advantages of being high in self-adaption, stable in convergence, high in intelligent cooperation capacity and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to integrated energy system optimization and control technology, specifically to the field of distributed collaborative scheduling and low-carbon optimization of micro energy networks, and particularly to a collaborative optimization method for micro energy networks based on multi-agent deep reinforcement learning. Background Technology

[0002] Microgrids, with their multi-energy complementarity, have become an important carrier for the low-carbon transformation of the power system. However, current control technologies face three major challenges: First, traditional centralized optimization methods based on deterministic models rely on accurate prediction information, making it difficult to adapt to the strong output fluctuations caused by the high proportion of renewable energy integration, resulting in significant scheduling deviations. Second, single-agent reinforcement learning frameworks cannot effectively characterize the dynamic coupling relationships of heterogeneous devices such as gas turbines, photovoltaics, and energy storage, leading to low coordination efficiency. Third, the full-information interaction mechanism used in conventional multi-agent reinforcement learning involves a large amount of redundant data exchange, significantly increasing communication burden and reducing real-time response speed. These methods show significant limitations when coordinating multi-objective optimization, especially in scenarios balancing economy and energy efficiency under carbon constraints. While existing distributed decision-making schemes can alleviate communication pressure, they have not yet solved the problem of information overload between devices. Therefore, it is urgent to develop a cooperative scheduling framework adapted to the heterogeneous architecture of microgrids, which can improve the accuracy of multi-objective coordination while reducing communication overhead. Summary of the Invention

[0003] To overcome the shortcomings of existing technologies, such as weak model adaptability, heterogeneous collaboration failure, and communication resource overload, this invention proposes a micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning, which specifically includes the following steps:

[0004] S1. Perform micro-energy grid environment modeling, namely, establish a source domain model for quantifying energy production, an energy storage model for quantifying energy storage, and a load model for quantifying energy use, and calculate operating costs and carbon emissions based on the source domain model, energy storage model, and load model;

[0005] S2. In the source domain model, set up a source-side renewable intelligent agent and a source-side adjustable device intelligent agent, wherein the source-side renewable intelligent agent is used for setting the renewable energy generation limit ratio, and the source-side adjustable device intelligent agent is used for setting the power, heat or cooling capacity of adjustable devices.

[0006] S3. Set up a grid-side intelligent agent for purchasing or selling electric power; set up an energy storage intelligent agent for setting the charging and discharging power of electric energy storage, heat storage, and cold storage.

[0007] S4. Construct a global state vector. During training, update the vector with the goal of achieving the global optimum. Local agents select the optimal action based on their local state and local safety penalty. Then, store the global state, the local states of each agent, the joint action of all agents, the global reward, and the next global state as experience in the experience pool.

[0008] Compared with the prior art, the present invention has the following beneficial effects:

[0009] 1. This invention establishes a multi-agent collaborative framework for sources, grids, loads, and storage, which operates under a centralized training and distributed execution (CTDE) mechanism. In the centralized evaluator, an attention mechanism is used to selectively aggregate key information across agents. This improves the strategy convergence stability and online response speed, reduces communication burden and single-point failure risk, and ensures the steady-state operation and dynamic adjustment capability of the system in scenarios such as renewable power output fluctuations, load uncertainty, and electricity price disturbances.

[0010] 2. This invention incorporates time-period carbon emission intensity / carbon cost parameters along with electricity prices, supply and demand balance, equipment / network boundaries, and energy storage SOC into targets and constraints, sets a total carbon emission control line for the cycle, and continuously verifies it during rolling execution. Based on this, it uses parameterized action boundaries and standardized data / constraint interfaces to issue instructions, thereby prioritizing the use of low-carbon periods and local clean power output without violating carbon constraints, reducing overall energy purchase costs, and facilitating on-site setup and capacity expansion. It can also be quickly migrated and deployed according to changes in equipment scale, electricity pricing mechanisms, or policy interpretations. Attached Figure Description

[0011] Figure 1 This is a flowchart of a micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to the present invention;

[0012] Figure 2 This is a diagram of the multi-agent training framework based on the MAATD3 algorithm of this invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] This invention proposes a collaborative optimization method for micro-energy networks based on multi-agent deep reinforcement learning, specifically including the following steps:

[0015] S1. Perform micro-energy grid environment modeling, namely, establish a source domain model for quantifying energy production, an energy storage model for quantifying energy storage, and a load model for quantifying energy use, and calculate operating costs and carbon emissions based on the source domain model, energy storage model, and load model;

[0016] S2. In the source domain model, set up a source-side renewable intelligent agent and a source-side adjustable device intelligent agent, wherein the source-side renewable intelligent agent is used for setting the renewable energy generation limit ratio, and the source-side adjustable device intelligent agent is used for setting the power, heat or cooling capacity of adjustable devices.

[0017] S3. Set up a grid-side intelligent agent for purchasing or selling electric power; set up an energy storage intelligent agent for setting the charging and discharging power of electric energy storage, heat storage, and cold storage.

[0018] S4. Construct a global state vector. During training, update the vector with the goal of achieving the global optimum. Local agents select the optimal action based on their local state and local safety penalty. Then, store the global state, the local states of each agent, the joint action of all agents, the global reward, and the next global state as experience in the experience pool.

[0019] Please see Figures 1-2 This embodiment provides a collaborative optimization method for micro-energy networks based on multi-agent deep reinforcement learning, and the specific steps are as follows:

[0020] Step S1: Construct a micro-energy grid environment model covering the four domains of source, grid, load, and storage, and establish a corresponding data acquisition and rolling update mechanism.

[0021] The environmental model is used to unify the parameters of supporting equipment, operating boundaries, energy flow coupling relationships, and settlement and carbon emission elements, providing an executable and verifiable physical basis for subsequent multi-agent decision-making and training. In this embodiment, data collection is completed based on the existing monitoring / control system: the power supply side records the unit output and start-up / shutdown information; the energy storage side obtains charging / discharging and supply / storage of heat / cold energy, as well as charging / storage status, through the BMS / EMS and thermal / cold energy storage control units; the energy consumption side collects real-time electricity / heat / cold energy demand, and simultaneously obtains the time window and planned quantity of transferable tasks when demand response is involved; the external side accesses time-of-use electricity pricing and grid-related carbon intensity or equivalent carbon cost parameters, and can combine basic forecast information such as wind speed and irradiance. Equipment and system modeling includes:

[0022] (1) Establish the source domain model:

[0023] In this embodiment, the source domain model can be divided into two types according to whether it is a renewable energy source. In this embodiment, renewable energy mainly includes wind power and photovoltaic power, while non-renewable energy sources are referred to as adjustable devices.

[0024] In this embodiment, the source domain models for wind power and photovoltaic power are represented as follows:

[0025]

[0026]

[0027] in, The power generation capacity of the photovoltaic equipment at time t; This parameter represents the power generation limit for photovoltaic (PV) equipment at time t. If the park does not restrict the power generation limit for PV equipment, this parameter can be set to 0. The available output of the photovoltaic equipment at time t; The available output of the wind power equipment at time t; This parameter represents the power generation limit for wind power equipment at time t. If the park does not restrict the power generation limit for photovoltaic equipment, this parameter can be set to 0. The available wind power output at time t.

[0028] In this embodiment, an adjustable device for non-renewable energy sources is used. (The adjustable device can be a CHP, boiler, heat pump, electric refrigeration / heating, etc.). This embodiment takes the generation of electrical energy, heat energy, and cold energy by an adjustable device that generates non-renewable energy as an example, namely:

[0029]

[0030]

[0031]

[0032] in, , , They are respectively at time t with equivalent energy After inputting an adjustable device that generates electrical energy, heat energy, or cold energy, the corresponding power of the electrical energy, heat energy, or cold energy output by the adjustable device is displayed. , , These are the conversion efficiencies of adjustable devices that generate electrical energy, thermal energy, and cold energy, respectively.

[0033] (2) Establish an energy storage model:

[0034] Taking electrical energy storage as an example (the same applies to thermal / cold energy storage):

[0035]

[0036] in, The stored electrical energy at time t+1; Let be the electrical energy stored at time t; The charging efficiency of energy storage devices; The charging power of the energy storage device at time t; The interval between two moments; The discharge efficiency of the energy storage device; Let t be the discharge power of the energy storage device at time t.

[0037] Similarly, given the remaining energy of the device at the previous moment, the charging and discharging efficiency from the previous moment to the current moment, and the charging and discharging power, the energy storage at the current moment can be calculated.

[0038] (3) Establish a load model:

[0039] The load model consists of the demand sequence of various energy sources in the system. This embodiment takes electricity, heat, and cooling as examples to construct the predicted load sequences for electricity, heat, and cooling:

[0040]

[0041] in, For time t, the load sequence on the three energy sources: electricity, heat, and cold; Let t be the electrical load; The thermal load at time t; Let t be the load of cooling energy.

[0042] The system constraints for establishing the load model include:

[0043] ① Three-energy power balance constraint: The energy balance constraint can be expressed as:

[0044]

[0045] The thermal energy equilibrium constraint can be expressed as:

[0046]

[0047] The balance constraint of cold electrical energy can be expressed as:

[0048]

[0049] in, , , These are the loss or efficiency conversion items for the power grid, heating network, and cooling network at time t, respectively.

[0050] ② Interface or line constraints:

[0051]

[0052] Among them, P grid (t) represents the power of buying and selling with the main network (positive buying and negative selling); This refers to the capacity of the interface (or line).

[0053] ③ Fulfillable domain and dynamic boundary of the device:

[0054] Feasible domain for electrical equipment (same for cold / hot):

[0055]

[0056] in, , These are the upper and lower limits of the equipment power. This refers to the ramp-up limit for energy storage devices. , These are the minimum and maximum energy storage boundaries for the energy storage device.

[0057] The calculation of operating costs and carbon emissions based on source area models, energy storage models, and load models includes:

[0058] (1) Operating costs:

[0059]

[0060] Where C(t) is the operating cost; In order to coordinate with the power purchase and sale on the main grid, [] + This indicates a positive value operation, meaning the positive part within the parentheses is used for electricity purchase; if the number within the parentheses is not positive, then 0 is output. [] - This indicates a negative value operation, meaning the negative part within the parentheses is used for electricity sales; if the number within the parentheses is not negative, then 0 is output. For electricity purchase price, The electricity price; Let f be the fuel cost function for device k.

[0061] (2) Carbon emissions:

[0062]

[0063] Among them, Q carbon (t) represents carbon emissions; Let t be the carbon emission factor of the power grid. This is the equivalent emission factor for local generating units.

[0064] Step S2: Based on the model and constraints in S1, this step defines the state and actions of the agent; and realizes distributed decision control of multiple agents.

[0065] To cover the four links of "source, grid, load, and storage", the following intelligent agents are set up (load is regarded as an exogenous predicted quantity and only enters the global state S). t (not used as a decision variable)

[0066] (1) Source-side renewable (RES) smart agent: used for setting the curtailment ratio of photovoltaic and wind power. When curtailment is not implemented on site, the action of this smart agent is fixed at 0.

[0067] (2) Source-side adjustable device (CONV) agent: used for power / heating / cooling settings of adjustable devices (such as CHP / tri-generation, boiler, heat pump, electric refrigeration / electric heating).

[0068] (3) Grid-side (GRID) agent: used for purchasing / selling power with the main grid.

[0069] (4) Energy Storage (ESS) Intelligent Agent: Used for setting the charging and discharging power of electrical energy storage, heat storage and cold storage.

[0070] The states in this invention are divided into two categories: global states (used for centralized evaluation during the training phase) and local observations of each agent (used for individual decision-making during the execution phase). The global state is defined as follows:

[0071]

[0072] in, Represents the global state vector at time t; Let t be the available photovoltaic output. Let t be the available wind power output. Forecast the electrical load at time t; For heat load prediction at time t; Forecast of cooling load at time t; The time-of-use electricity price at time t; Let be the carbon emission factor of the power grid at time t; For the set of constraints and boundary parameters; The energy storage for all energy sources at time t.

[0073] The local state of each agent is defined as follows:

[0074] (1) The local state of the RES agent is:

[0075]

[0076] (2) CONV agent:

[0077]

[0078] in, , These are the minimum and maximum electrical power output constraints for device k, respectively. , These are the minimum and maximum thermal power output constraints for device k, respectively. , These are the minimum and maximum cooling power output constraints for device k, respectively. , , The upper limit of the ramp rate for electrical, thermal, and cooling power; For multi-energy coupling conversion parameters.

[0079] (3) GRID agent:

[0080]

[0081] in, This refers to the minimum capacity of the interface or line. This refers to the maximum capacity of the interface or line.

[0082] (4) ESS agent:

[0083]

[0084] in, Limits on the maximum charging and discharging power of energy storage devices; Limits are imposed on the maximum heat charging rate and maximum heat dissipation rate of thermal energy storage equipment; Limitations on the maximum charging and discharging rates of cold energy storage equipment; , , These refer to the charge / discharge energy conversion efficiency of electric, thermal, and cold energy storage devices, respectively.

[0085] The actions of each agent are defined as follows:

[0086] (1) RES agent:

[0087]

[0088]

[0089] in, Let be the active power curtailment ratio of photovoltaic power at time t; The active curtailment ratio of wind power at time t; when curtailment is not implemented on-site, the following is set: .

[0090] (2) CONV agent:

[0091]

[0092] in, Let t be the setpoint value of the electrical power output of the adjustable device k. The setpoint for the heat output of the adjustable device k at time t; Let be the cooling power output setting value of adjustable device k at time t. All settings must meet constraints such as the device's upper and lower capacity limits, ramp rate, coupling, and power balance.

[0093] (3) GRID agent:

[0094]

[0095] in, This represents the power exchanged with the main grid; a positive value indicates electricity purchase, and a negative value indicates electricity sale.

[0096] (4) ESS agent:

[0097]

[0098] in, The power command for the integrated charging and discharging of the energy storage device at time t, the value range of which is... ,when When the device performs a charging action, The device performs a discharge action at that time; , These are integrated instructions for the storage and release of thermal and cold energy storage devices, respectively, with control logic consistent with that of electric energy storage.

[0099] Step S3: This step constructs a global reward function aimed at optimizing the overall team performance, which may be supplemented by key local penalties. The reward is guided by "cost reduction, carbon reduction, and utilization improvement," and is geared towards a multi-agent centralized training and decentralized execution framework.

[0100] Specifically, the global reward function in this invention is:

[0101]

[0102] Where C(t) is the system operating cost at time t; Q carbon (t) represents the carbon emissions of the system at time t; U(t) represents the renewable energy absorption rate. The weighting coefficient must include a normalization factor to eliminate dimensional differences between amount, weight, and ratio.

[0103] The renewable energy absorption rate U(t) is expressed as:

[0104]

[0105] in, , Power output for grid-connected photovoltaic and wind power, , These are the predicted available power outputs for photovoltaic and wind power, respectively. To prevent extremely small positive numbers with a denominator of zero.

[0106] Further, localized security penalties:

[0107]

[0108] Where SOC(t) is the normalized energy storage rate of the electric / thermal / cold energy storage device, SOC min SOC max These represent the lower and upper safety limits for various types of energy storage; p soc A value greater than 0 represents a penalty coefficient, which is significantly higher than the standard cost time scale, in order to effectively prevent out-of-bounds behavior.

[0109] This local penalty item By employing reward-based shaping technology, the physical safety constraints of energy storage devices are transformed into value signals for reinforcement learning.

[0110] During centralized training in S5, when an agent explores an out-of-bounds state, the negative penalty significantly reduces the cumulative expected reward (Q-value) at the current time step. Through policy gradient backpropagation in the multi-agent TD3 algorithm, the critic network guides the actor network to correct its policy parameters, reducing the probability density of out-of-bounds actions. This allows the agent to autonomously converge to a safe and feasible region through trial and error without the need for manually writing complex rules.

[0111] Step S4: This step describes the process by which agents explore, experiment, and collect learning samples in a simulation environment. All agents i, based on their local observations... and current strategy (Actor Network) outputs its respective actions. The joint action is applied to the microgrid environment model established in step one. The environment model evolves to the next state according to physical rules and calculates and returns the global reward defined in step three. The complete experience tuple (including the global state, local observations of each agent, joint action, global reward, next global state, etc.) is stored in the experience replay buffer.

[0112] Furthermore, action generation and environment advancement: To enhance the agent's exploration capabilities during the training phase, each agent outputs deterministic actions in the policy network (Actor). Based on this, Gaussian exploration noise needs to be superimposed. ,Right now The action a t On the one hand, it is used to store the experience replay pool, and on the other hand, it is denormalized and mapped to physical values ​​before being applied to the environment.

[0113] At each time step t, each agent, based on its observation information... Generate Actions :

[0114] Energy Storage Intelligent Agent (ESS): Generates charging and discharging power.

[0115]

[0116] Grid agent: generates electricity purchase / sale capacity.

[0117]

[0118] Source-side renewable (RES) smart agents: limiting the generation ratio of photovoltaic and wind power.

[0119]

[0120] in,

[0121] Source-side adjustable device (CONV) agent: generates power output or heat / cooling output for each controlled device.

[0122] (Output of each device k)

[0123] The generated joint action is represented as:

[0124]

[0125] The joint action is applied to the environment model, and the environment progresses to the next state S according to the physical rules in S1. t+1 .

[0126] Furthermore, the empirical data is stored: at each time t, the environment generates a state S based on the agent's actions. t+1 The system calculates the reward R(t). Complete experience data is stored in the experience replay pool D, and each experience sample M in the replay pool... t The definition is as follows:

[0127]

[0128] Among them, S t S t+1 Let be the global states at time t and time t+1, respectively; } represents the set of local observations for all agents i; } represents the set of normalized joint actions performed by all agents with noise; R(t) is the global reward at the current time step; D tThis is a round end flag, set to 1 when the round ends or a termination condition is triggered, and 0 otherwise; this empirical data is then used by the off-policy method during the training phase.

[0129] Step S5: This step is based on a centralized training, decentralized execution (CTDE) framework, using a multi-agent deep reinforcement learning algorithm (MAATD3) to optimize the collaborative learning process among agents. The training framework is as follows: Figure 2 As shown, the computational efficiency of the critic network is improved through a self-attention mechanism, and the interaction information between agents is effectively aggregated during training. This method improves the accuracy and efficiency of Q-value calculation, ensuring that the system can stably learn the optimal cooperative strategy during training.

[0130] Furthermore, during training, a batch of experience data is randomly sampled from the experience replay pool; this data will be used to train the centralized commentator network and the decentralized enforcer network.

[0131] Furthermore, the centralized critic network is updated. Its primary task is to guide agent learning by evaluating the long-term rewards of joint actions. A multi-head self-attention mechanism is introduced into the critic network. This mechanism dynamically calculates the mutual influence weights between agents based on the current global state and joint actions, selectively aggregating key information. The specific calculation process is as follows:

[0132] Furthermore, firstly, the local observations of the i-th agent are... With action Concatenate the data and generate the query vector Q by sharing the fully connected layer mapping. i ; all other intelligent agents As the subject of attention, its local observation With action Mapped to key vector K j AND value vector V j .

[0133] Furthermore, the correlation between agents is calculated using a dot product scaling model, and the attention weights are obtained by normalization, as shown in the following formula:

[0134]

[0135] Among them, Q i The query vector represents the state characteristics and decision intention of the current agent i. It is obtained by concatenating the local observations and actions of agent i and mapping them through a linear layer with learnable parameters. The key vector and value vector are obtained similarly by concatenating the local observations and actions and mapping them through a linear layer with learnable parameters. K is the key vector matrix (containing the key vectors of all neighboring agents).j , N is the number of agents in the system, representing the feature indices of other agents; V is the value vector matrix (containing the value vectors of all neighboring agents). j ), representing the actual state information content of other intelligent agents; d k is the feature dimension of the key vector.

[0136] Furthermore, the computational output of the above attention mechanism is defined as the aggregated collaborative feature X of agent i. i ,Right now:

[0137]

[0138] This process enables the commentator network to dynamically adjust its focus on information from other agents based on the agents' local observations, thereby improving the accuracy of Q-value estimation.

[0139] Furthermore, the Q value is calculated as follows:

[0140]

[0141] Among them, e i This invention utilizes X to represent the observation-action characteristics of the current agent; the commentator network can more effectively aggregate information from various agents, thereby improving the collaborative efficiency of multi-agent systems. i and e i The two data points are combined to fully cover the system's global state and joint actions; therefore, global data is used for updating here.

[0142] Furthermore, the decentralized actor updates its policy network parameters using a policy gradient algorithm to ensure that the agent's output actions maximize the value assessment of the critic network. The decentralized actor updates as follows:

[0143]

[0144] in, Let be the gradient of the policy objective function, and let represent the policy network parameters of the i-th agent. The direction of updates; Let $\mathematical expectation$ be based on the experience replay pool, and $\mathematical expectation$ be calculated for the average gradient within the brackets after randomly sampling a batch of historical data (including observations $o$ and actions $a$) from the experience replay pool. Let be the gradient of the action value function with respect to the action, representing the action a of agent i in order to obtain a higher value score Q, given the current global state S and joint action A. i How should it be adjusted? Let be the gradient of the policy network with respect to the network parameters, and let represent the weight parameters within the policy network. How do minute changes affect the output action a? i of.

[0145] Furthermore, during training, the following key metrics are monitored to determine whether convergence has occurred:

[0146] Average round reward: Check whether the average reward of all agents is stable to ensure that the system is gradually optimized during training.

[0147] Average cost: Calculate the average cost over a period of time to ensure that costs gradually decrease and tend to stabilize.

[0148] Average carbon emissions: Calculate carbon emissions to ensure that the system not only reduces costs during optimization but also meets carbon emission control targets.

[0149] When these metrics stabilize or reach the preset number of training cycles over multiple consecutive training cycles, the training process is considered to have converged, and training can be stopped to enter the deployment phase.

[0150] Step S6: This step deploys the trained agent actuators to the integrated energy system control center, ensuring coordinated operation across all domains (source, grid, load, and storage). During operation, the agent makes real-time decisions based on global and local observations, generates control commands through feasible domain verification, and distributes them to the devices. The system achieves cross-domain collaboration and global optimization, avoiding high-frequency central communication. The system synchronously outputs economic, carbon, and energy efficiency assessment results, generating charts or geographic visualizations. When performance deviates from predetermined targets or environmental changes occur, the system will automatically trigger parameter adaptation or retraining.

[0151] Furthermore, after training, the agent's scheduling strategy is deployed to the control center of the integrated energy system. Each actuator will be uniformly coordinated at the control center according to the system's needs and operating status. The actuators in the source domain include controllers for renewable energy (photovoltaics, wind power, etc.) and adjustable devices (such as boilers, combined heat and power, etc.), grid actuators are responsible for regulating the purchase and sale of electricity, and energy storage actuators regulate charging and discharging according to the energy storage status.

[0152] Furthermore, the agent generates control commands based on global and local observation information and ensures that the commands comply with physical and operational constraints through feasible domain verification. Each agent makes decisions based on its local operating environment and load requirements, while ensuring that actions do not exceed the device's operational limits or cause system instability. When generated actions exceed the constraints, the system automatically adjusts or prunes them to ensure that the commands are legal and effective. Simultaneously, the agent updates the control strategy in real time when making decisions to allocate resources optimally.

[0153] Furthermore, the system will generate specific control instruction sets based on real-time scheduling requirements, including:

[0154] Unit output setting: Adjusts the output of renewable energy equipment such as photovoltaic and wind power to match grid demand.

[0155] Energy storage charging and discharging power: Adjust the charging and discharging power of the energy storage device according to the status of the energy storage device and the load demand.

[0156] Adjustable load adjustment: Dispatch loads that can be moved (such as air conditioners, electric equipment, etc.) to achieve flexible load adjustment.

[0157] In addition, the system will output economic, carbon emission, and energy efficiency assessment results based on the execution results, and generate charts or geographic visualizations to help operators monitor the system's operating status in real time.

[0158] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A collaborative optimization method for micro-energy networks based on multi-agent deep reinforcement learning, characterized in that, Specifically, the following steps are included: S1. Perform micro-energy grid environment modeling, namely, establish a source domain model for quantifying energy production, an energy storage model for quantifying energy storage, and a load model for quantifying energy use, and calculate operating costs and carbon emissions based on the source domain model, energy storage model, and load model; S2. In the source domain model, set up a source-side renewable intelligent agent and a source-side adjustable device intelligent agent, wherein the source-side renewable intelligent agent is used for setting the renewable energy generation limit ratio, and the source-side adjustable device intelligent agent is used for setting the power, heat or cooling capacity of adjustable devices. S3. Set up a grid-side intelligent agent for purchasing or selling power; set up an energy storage intelligent agent for setting the charging and discharging power of electrical energy storage, heat storage, and cold storage. S4. Construct a global state vector. During training, update the vector with the goal of achieving the global optimum. Local agents select the optimal action based on their local state and local safety penalty. Then, store the global state, the local states of each agent, the joint action of all agents, the global reward, and the next global state as experience in the experience pool.

2. The micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, The global state vector is defined as: ; in, Represents the global state vector at time t; Let t be the available photovoltaic output. Let t be the available wind power output. Forecast the electrical load at time t; For heat load prediction at time t; Forecast of cooling load at time t; The time-of-use electricity price at time t; Let be the carbon emission factor of the power grid at time t; For the set of constraints and boundary parameters; The energy storage for all energy sources at time t.

3. A micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 1 or 2, characterized in that, The local state of the source-side regenerative agent is: ; The local state of the source-side adjustable device agent is: ; The local state of the network-side agent is: ; The local state of the energy storage agent is: ; in, For time t, the local state of the source-side adjustable device intelligent agent; Let t be the available photovoltaic output. Let t be the available wind power output. For time t, the local state of the source-side adjustable device intelligent agent; The minimum electrical power output constraint for device k; The maximum electrical power output constraint for device k; Let be the ramp-up limit of the electrical power of device k; The minimum thermal power output constraint for device k; The maximum thermal power output constraint for device k; The ramp-up limit for the thermal power of device k; The minimum cooling power output constraint for device k; The maximum cooling power output constraint for device k; This represents the maximum ramp-up limit for the cooling power of device k; For the necessary electro-thermal-cold coupling parameters; For time t, the local state of the network-side agent; For the purchase and sale of electricity with the main grid; Let be the carbon emission factor of the power grid at time t; This refers to the minimum capacity of the interface or line. This refers to the maximum capacity of the interface or line. Let t be the local state of the energy storage agent; Energy storage for all energy sources at time t; The purchase price of electricity at time t; The electricity price at time t; The maximum charging power of the energy storage device; This refers to the maximum discharge power of the energy storage device. This refers to the maximum heat charging rate of the thermal energy storage device. This refers to the maximum heat release rate of the thermal energy storage device. This represents the maximum charging and cooling rate of the cold energy storage equipment. This represents the maximum cooling rate of the cold energy storage device. Energy conversion efficiency when charging energy storage devices; The energy conversion efficiency of an energy storage device during discharge; Energy conversion efficiency during the charging of thermal energy storage equipment; The energy conversion efficiency of a thermal energy storage device when it releases heat. Energy conversion efficiency during the charging of cold storage equipment; This refers to the energy conversion efficiency of a cold energy storage device when it releases cold air.

4. The micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 3, characterized in that, The action definition of a source-side adjustable device agent is: ; The action definition of a source-side adjustable device agent is: ; The action definition of a network-side agent is: ; The action of an energy storage agent is defined as follows: ; in, For the action of the source-side renewable intelligent agent at time t; Let be the active power curtailment ratio of photovoltaic power at time t; Let be the active curtailment ratio of wind power at time t; For the action of the intelligent agent on the source side at time t - adjustable device; For the actions of the network-side intelligent agent at any given time; For the action of the energy storage intelligent agent at time t; The electrical power output of the adjustable device k at time t; The thermal power output of the adjustable device k at time t; The cooling power output of the time-adjustable device k is t. The power command for the integrated charging and discharging of the energy storage device at time t, the value range of which is... ,when When the device performs a charging action, The device performs a discharge action at that time; This is an integrated heat storage and release command for thermal energy storage equipment. The value range of this command is: ,when When the equipment performs the heat storage action, The equipment then performs a heat release action; This is an integrated heat / cold storage command for cold energy storage devices. The value range of this command is: ,when When the equipment performs a cold storage operation, The equipment then performs a cooling action.

5. A micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 3 or 4, characterized in that, The global reward function is: ; in, Let be the global reward function at time t; Let be the operating cost at time t; Let be the carbon emissions at time t; The renewable energy consumption rate is the indicator for time t. , , These are the weighting coefficients for operating costs, carbon emissions, and renewable energy integration rates, respectively.

6. The micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 5, characterized in that, During training in a local agent, when the locally observed energy storage rate exceeds a set safety range, a local penalty term is used to reduce the cumulative expected reward. This local penalty term is: ; in, p is the local penalty term at time t; soc >0 represents the penalty coefficient; This represents the lower limit of safe energy storage for current energy storage devices; The normalized energy storage rate of the current energy storage device at time t; This represents the current safe energy storage limit for energy storage devices; This indicates a positive value operation; if the value in the parentheses is positive, the value will be output directly, otherwise 0 will be output.

7. The micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, The training process for multi-agent systems includes: The decentralized commentator network takes the local observations of agents as input and weights the local observation states of each agent through an attention mechanism; This allows decentralized commentator networks to dynamically adjust their focus on information from other agents based on the agent's local observations, thereby improving the accuracy of Q-value estimation. The Q-value of a centralized commentator network is expressed as: ; The process by which decentralized executors update network parameters includes: ; in, This represents the state vector at time t of a centralized critic network with network parameters θ. Joint actions The Q value below; ; The observed states and actions of all agents in the system, excluding the current agent i, are aggregated based on an attention mechanism. For agent i, the observed state and actions; The gradient of the policy objective function; The mathematical expectation is based on the experience replay pool; Let be the gradient of the action value function with respect to the action; This represents the gradient of the policy network with respect to the network parameters.

8. The micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 7, characterized in that, During training, the commentator network for each agent uses an attention mechanism to process the local state observed by that agent. Specifically, it generates query vectors, key vectors, and value vectors based on each agent's local observations and uses the calculated attention weights to weight the agent's local observations.

9. The micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, In the source domain model, the energy generated by renewable energy devices is represented as: ; The energy produced by non-renewable equipment is: ; in, The power generation capacity of renewable energy equipment at time t; The power generation limit for renewable energy equipment at time t. If the park does not restrict the power generation limit for photovoltaic equipment, this parameter can be set to 0. The available output of renewable energy equipment at time t; The energy generated by non-renewable energy devices at time t; Energy conversion efficiency of non-renewable energy equipment; The equivalent energy input of the non-renewable energy equipment at time t; Energy storage models include: ; in, The stored energy of the energy storage device at time t+1; Let be the energy stored in the energy storage device at time t; The charging efficiency of energy storage devices; The charging power of the energy storage device at time t; The interval between two moments; The discharge efficiency of energy storage devices; Let t be the discharge power of the energy storage device; The load model consists of the sequence of energy load demands in the system. The constraints of the load model include at least the balance constraints between energy sources in the system, the minimum and maximum transmission constraints of each interface and line, and the feasible domain and dynamic boundary of the equipment.

10. The micro-energy network collaborative optimization method based on multi-agent deep reinforcement learning according to claim 9, characterized in that, The calculation of operating costs and carbon emissions based on source area models, energy storage models, and load models includes: ; ; in, Time-of-use electricity pricing; The electricity price; Let f be the fuel cost function of device k; Carbon emission factor of power grid; This is the equivalent emission factor for local generating units.