Micro-grid energy dispatching carbon emission reduction system based on reinforcement learning dual-objective optimization
By applying a dual-objective optimization method based on reinforcement learning in the microgrid, the problem that traditional scheduling methods cannot simultaneously reduce carbon emissions and optimize energy use is solved, and the intelligent scheduling and carbon emission reduction effects of the microgrid are achieved.
Patent Information
- Application Number
- CN202510015783.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional microgrid scheduling methods cannot simultaneously achieve effective reduction of carbon emissions and optimized use of green energy in complex environments, and it is difficult to meet the intelligent and adaptive needs of modern microgrids.
The dual-objective optimization method based on reinforcement learning is adopted, and through data acquisition, processing and intelligent scheduling modules, the space-time graph convolutional neural network and deep reinforcement learning algorithm are used to dynamically optimize the energy scheduling of the microgrid to achieve carbon emission reduction and efficient energy utilization.
It has achieved intelligence and adaptability of microgrid energy scheduling, optimized carbon emissions and energy use, reduced operating costs, and supported the carbon neutrality goal of industrial parks.
Smart Images

Figure CN119990592A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microgrid energy dispatching, and in particular to a microgrid energy dispatching carbon emission reduction system based on reinforcement learning dual-objective optimization. Background Art
[0002] Microgrid is a small distribution system that integrates distributed power sources, energy storage systems, power conversion equipment and power loads. In recent years, it has attracted widespread attention due to its flexibility and sustainability. With the gradual strengthening of global carbon emission control and the rapid development of microgrid technology, the integrated management of microgrid energy and carbon emissions has become a hot topic of research. However, traditional scheduling methods often show limitations in complex microgrid environments and cannot simultaneously achieve effective reduction of carbon emissions and optimal use of green energy.
[0003] Therefore, an intelligent and adaptable dispatching method is urgently needed to meet the needs of modern microgrids. As an important energy consumer and carbon emitter, the widespread application of microgrids in industrial parks provides a new path for the park to achieve carbon neutrality. In order to further improve the energy dispatching efficiency of industrial park microgrids, reinforcement learning has gradually attracted attention as an effective tool to solve complex nonlinear problems. Reinforcement learning models the sequential decision-making problem as an interactive process between the intelligent agent and the external environment, so that the intelligent agent can continuously optimize the decision-making strategy through trial and error learning, thereby obtaining the maximum cumulative benefit in sequential decision-making. In recent years, the combination of reinforcement learning and deep learning has greatly improved the performance of the algorithm, and has achieved remarkable results in the fields of games, robot control, natural language processing and computer vision. Based on this background, the application of deep reinforcement learning technology to the energy and carbon emission management of microgrids is not only of theoretical significance, but also can provide technical support for the green and sustainable development of industrial parks. Summary of the invention
[0004] In order to solve the problems raised in the above background technology, the purpose of the present invention is to provide a microgrid energy scheduling carbon emission reduction system based on reinforcement learning dual-objective optimization, including a data acquisition module, a data processing module, an intelligent scheduling module, an execution feedback module and a human-computer interaction module, characterized in that: the data acquisition module: collects equipment energy consumption data, environmental parameters, carbon emission coefficient and other information in the microgrid in real time through sensor equipment; the data processing module: cleans, extracts features and standardizes the collected data; the intelligent scheduling module: uses a spatiotemporal graph convolutional neural network (STGCN) to extract information features, and further adopts a deep reinforcement learning algorithm based on actors and critics, and optimizes the model data dynamically updated by the time window to obtain an optimization strategy and form a scheduling instruction set; the execution feedback module: dynamically schedules the operating status of the equipment in the microgrid according to the scheduling instruction set output by the intelligent scheduling module; the human-computer interaction module: provides a user interface for displaying the usage of carbon quotas, the system operating status, and allows users to input adjustment requirements;
[0005] The specific steps for building a system model in the intelligent scheduling module are as follows: S1. Set the campus microgrid model, which is equipped with a distributed energy system, an energy storage system, and load equipment. S2. Determine the required energy and load range of each device in the microgrid. S3. Establish an objective function and its dynamic weight adjustment model. S4. Set the state-action-reward model: the state space of the microgrid includes the power supply and demand relationship between the microgrid nodes, the action space of the microgrid includes the scheduling of the microgrid power, and the reward model of the microgrid is designed based on the aforementioned state space and action space. S5. Establish a strategy model based on reinforcement learning of actors and critics. S6. According to the model optimization results, allocate power and form a scheduling instruction set.
[0006] The step S2 is specifically the following steps:
[0007] S21. Obtain the capacity range of the distributed energy system:
[0008] Assuming that the total number of distributed energy (DG) systems is N, the capacity of the kth distributed energy system is The minimum capacity is The maximum capacity is The range is
[0009]
[0010] S22. Calculate the power status of the energy storage device:
[0011] SOC s (t) represents the power state of the energy storage device, and the change of the power state of the energy storage device is:
[0012]
[0013] Among them, η s is the charging and discharging efficiency of the energy storage device, It indicates the charge and discharge time. is the charging power of the energy storage device, is the maximum charging power of the energy storage device, and its power range is Represents the discharge power of the energy storage device, is the maximum discharge power of the energy storage device, and its power range is SOC s,min It is the minimum value of the energy storage device, SOC s,max It is the maximum value of the energy storage device. The capacity limit of the energy storage device is: SOC s,min ≤SOC s (t)≤SOC s,max .
[0014] S23. Calculate the total load power of the load equipment:
[0015]
[0016] Among them PD i Represents the peak load of the i-node unit, PF i represents the power factor at time t, f i represents the load fluctuation probability distribution of node i, L m (t) represents the total load power of the microgrid.
[0017] S24. Calculate the demand response power limit range:
[0018] The power demand of the zth demand response block in the microgrid is denoted as q z (t),u z (t) is a binary variable indicating whether the zth block response module responds. The demand response needs to meet the following constraints:
[0019]
[0020] u z-1 (t)≥u z (t),for z=2,…,Z
[0021] S25. Calculate the total load of the microgrid:
[0022] The microgrid needs to meet the following supply and demand power balance conditions:
[0023]
[0024] Where L m(t) represents the total load of the microgrid, It indicates the amount of electricity purchased by microgrid m from the main grid (the park distribution network from the city grid).
[0025] The step S3 is specifically the following steps:
[0026] S31. Establish a model based on the economic target, and the calculation method is as follows:
[0027]
[0028] Is related to power generation The related nonlinear equations. The function can be a quadratic polynomial, which represents the power generation cost of distributed generation.
[0029]
[0030] Among them, the coefficient N t Represents a dispatch cycle, N k Represents the total number of distributed energy devices, η m represents the network loss during power transmission, λ(t) represents the time-varying retail electricity price at the common coupling point, Represents the cost of power exchange between the microgrid and the main grid, represents the internal dispatching cost of the microgrid, It represents the unit price of dispatch. is a binary variable with a value of 0 or 1, indicating the zth demand response module Whether it is scheduled, ρ s |SOC s (t)-SOC s (t-1)| represents the loss cost of energy storage equipment, ρ s A coefficient that indicates the performance degradation and shortened life of energy storage equipment caused by changes in electrical quantity.
[0031] S32. Establish a model based on environmental goals. The calculation method is as follows: The total carbon emissions of the park in a scheduling cycle E m and carbon emission cost C carbon It can be expressed as:
[0032]
[0033] Where E m (t) represents the total carbon emissions of the microgrid at time step t, ETS(t) represents the carbon emission market trading system pricing at time step t, process (t) represents the total amount of greenhouse gas fugitive emissions from the production process at time step t, which can be obtained by the following formula:
[0034] EF j Represents the process carbon emission factor of the jth equipment.
[0035] S33. The comprehensive objective function combines economic and carbon reduction goals and is a weighted reward function that minimizes total greenhouse gas emissions and current net costs, expressed as:
[0036] R t =ω1·(-C cost,t )+ω2·(-C carbon,t )
[0037] C cost,t : Economic cost of energy-consuming equipment within the time step, C carbon,t : Carbon emission cost within the time step, ω1, ω2 are weight parameters.
[0038] Dynamically adjust weights ω1 and ω2 according to real-time changes in carbon quota utilization or economic costs:
[0039]
[0040]
[0041] The step S5 is specifically the following steps:
[0042] S51, using the learning strategy function π(A|S), the probability of selecting action a in state s to maximize the expected long-term accumulated reward G t :
[0043]
[0044] Where γ∈[0,1] is the discount factor, E is the expected value operator, t is the current time point, k is a counter, and R is the reward obtained. The optimization algorithm based on the dual-objective optimization of operating cost and total carbon emissions is as follows:
[0045] in It is the decision-making behavior space of the microgrid.
[0046] S52. Build an LSTM deep learning network to predict environmental meteorological data.
[0047] S53. Build an STGCN model for extracting microgrid network and node features. The STGCN model consists of an input layer, an ST-Conv Block layer, a GCN-Block layer and an output layer.
[0048] S54. Build a deep reinforcement learning network: Use the intelligent agent to use the actor-critic algorithm, dynamic parameter update, combine the policy gradient and temporal difference methods, and use the Adam optimization algorithm to update the weight matrix.
[0049] The step S53 is specifically as follows:
[0050] S501, the enterprise graph structure is represented as an adjacency matrix and a feature matrix, and the input layer inputs training data;
[0051] S502, using the He initialization method, using Gaussian distribution to initialize weight parameters, the variance is Where is the input dimension of the weight matrix, and the bias parameter can be initialized to 0 or any small value. The formula is as follows, Among them, W represents the weight parameter and n represents the input dimension of the weight matrix.
[0052] S503, construct the ST-Block layer, which consists of multiple one-dimensional convolutional layers and residual blocks to process time series data. The formula is as follows: X (k) =ReLU(LX (k-1) W (k-1) ). Among them, X (k-1) is the feature matrix of the k-1th layer, W (k-1) is the weight matrix of the k-1th layer, L is the graph Laplacian matrix, and ReLU is the activation function.
[0053] S504, construct the GCN-Block layer, which is composed of multiple graph convolutional layers and residual blocks, to process spatial data. The formula is as follows: Among them, X t represents the feature matrix at time t, A represents the adjacency matrix of the graph structure, I is the identity matrix, is the degree matrix, W (1) and W (2) are the weight matrices of the two convolutional layers respectively, and σ represents the activation function.
[0054] S505, the output layer includes a time domain convolution layer and a fully connected layer. The convolution kernel size of the time domain convolution layer is The number is C o , mapping the output to The fully connected layer is in Output
[0055] Compared with the existing technology, the present invention treats the energy system of the park as a unified network, coordinates them together, and performs topology, taking into account the mutual influence, mutual coupling, and mutual coordination of multiple systems. The deep reinforcement learning algorithm can quickly adapt to environmental changes, achieve precise scheduling, be environmentally friendly and efficient, optimize the use of carbon quotas, address overall carbon emissions, and reduce operating costs through intelligent scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is the topology diagram of the park microgrid in this embodiment;
[0057] Figure 2 Intelligent optimization module reinforcement learning algorithm architecture;
[0058] Figure 3 System overall architecture diagram. DETAILED DESCRIPTION
[0059] The present invention will be further described below with reference to the accompanying drawings.
[0060] like Figures 1 to 3 , a microgrid energy dispatching carbon emission reduction system based on reinforcement learning dual-objective optimization, including a data acquisition module, a data processing module, an intelligent dispatching module, an execution feedback module and a human-computer interaction module, characterized in that: the data acquisition module: collects equipment energy consumption data, environmental parameters, carbon emission coefficient and other information in the microgrid in real time through sensor equipment; the data processing module: cleans, extracts features and standardizes the collected data; the intelligent dispatching module: uses a spatiotemporal graph convolutional neural network (STGCN) to extract features with rich information between nodes, and then enters the deep reinforcement learning algorithm based on actors and critics, and uses the model data dynamically updated in the time window to perform intraday rolling optimization, obtain an optimization strategy that maximizes the dual benefits of economy and carbon emission reduction, and form a dispatch instruction; the execution feedback module: dynamically dispatches the operating status of the equipment in the microgrid according to the dispatch instruction set output by the intelligent dispatching module; the human-computer interaction module: provides a user interface for displaying the use of carbon quotas, the system operating status, and allows users to input adjustment requirements
[0061] The specific steps of building the system model in the intelligent scheduling module are as follows: S1. Set the campus microgrid model, which is equipped with a distributed energy system, an energy storage system, and load equipment. S2. Determine the required energy and load range of each device in the microgrid. S3. Establish a reward function and its dynamic weight adjustment model. S4. Set the state-action-reward model: the state space of the microgrid includes the power supply and demand relationship between the microgrid nodes, such as the amount of power a node needs or can provide, the action space of the microgrid includes the scheduling of microgrid power, such as the power scheduling range allowed at time t, and the reward model of the microgrid is designed based on the aforementioned state space and action space. For example, when the power supply and demand between the nodes in the microgrid are balanced, a certain reward is given. S5. Establish a strategy model based on reinforcement learning of actors and critics. S6. According to the optimization of economic and environmental goals, a scheduling instruction set is formed to reasonably allocate power to increase user satisfaction with electricity consumption and reduce the peak-to-average ratio.
[0062] The step S2 is specifically the following steps:
[0063] S21. Obtain the capacity range of the distributed energy system:
[0064] Assuming that the total number of distributed energy (DG) systems is N, the capacity of the kth distributed energy system is The minimum capacity is The maximum capacity is
[0065]
[0066] S22. Calculate the power status of the energy storage device:
[0067] SOC s (t) represents the power state of the energy storage device, and the change of the power state of the energy storage device is:
[0068]
[0069] Among them, η s is the charging and discharging efficiency of the energy storage device, It indicates the charge and discharge time.
[0070] is the charging power of the energy storage device, is the maximum charging power of the energy storage device, and its power range is Represents the discharge power of the energy storage device, is the maximum discharge power of the energy storage device, and its power range is SOC s,min It is the minimum value of the energy storage device, SOC s,maxSOC is the maximum value of the energy storage device. In order to avoid large fluctuations in the total power of the park distribution network, the power of the energy storage system needs to be controlled to make the charging and discharging power of the distribution network relatively stable. The capacity limit of the energy storage device is: SOC s,min ≤SOC s (t)≤SOC s,max .
[0071] S23. Calculate the total load power of the load equipment. For large fixed equipment in the park, including but not limited to air conditioners, elevators, water pumps, lighting, charging piles, etc.:
[0072]
[0073] Among them PD i Represents the peak load of the i-node unit, PF i represents the power factor at time t, f i represents the load fluctuation probability distribution of node i, L m (t) represents the total load power of the microgrid.
[0074] S24. Calculate the demand response power limit range:
[0075] The power demand of the zth demand response block in the microgrid is denoted as q z (t),u z (t) is a binary variable indicating whether the zth block response module responds. The demand response needs to meet the following constraints:
[0076]
[0077] u z-1 (t)≥u z (t),for z=2,…,Z
[0078] S25. Calculate the total load of the microgrid:
[0079] The microgrid needs to meet the following supply and demand power balance conditions:
[0080]
[0081] Where L m (t) represents the total load of the microgrid, It indicates the amount of electricity purchased by microgrid m from the main grid (the park distribution network from the city grid).
[0082] Step S3 is specifically the following steps:
[0083] S31. Establish a model based on economic objectives to improve the economic benefits of energy-consuming equipment and minimize energy consumption costs. The calculation method is as follows:
[0084]
[0085] Is related to power generation The related nonlinear equations. The function can be a quadratic polynomial, which represents the power generation cost of distributed generation.
[0086]
[0087] Among them, the coefficient N t Represents a dispatch cycle, N k Represents the total number of distributed energy devices, η m (t) represents the network loss during power transmission, λ(t) represents the time-varying retail electricity price at the common coupling point, Represents the cost of power exchange between the microgrid and the main grid, represents the internal dispatching cost of the microgrid, It represents the unit price of dispatch. is a binary variable with a value of 0 or 1, indicating the zth demand response module Whether it is scheduled, ρ s |SOC s (t)-SOC s (t-1)| represents the loss cost of energy storage equipment, ρ s A coefficient that indicates the performance degradation and shortened life of energy storage equipment caused by changes in electrical quantity.
[0088] S32. Establish a model based on environmental goals to reduce carbon emissions. The calculation method is as follows: The total carbon emissions of the park in a scheduling cycle E m and carbon emission cost C carbon Can
[0089] It is expressed as:
[0090]
[0091] Where E m (t) represents the total carbon emissions of the microgrid at time step t, unit tCO2e, ETS(t) represents the carbon emission market trading system pricing at time step t, unit ¥ / tCO2e, E process (t) represents the total amount of greenhouse gas fugitive emissions from the production process at time step t, which can be obtained by the following formula:
[0092] EF j Represents the process carbon emission factor of the jth equipment.
[0093] S33. The comprehensive objective function combines economic and carbon reduction goals and is a weighted reward function that minimizes total carbon dioxide emissions and current net costs, expressed as:
[0094] R t =ω1·(-C costt )+ω2·(-C carbon,t )
[0095] C cost,t : The economic cost of energy-consuming equipment within the time step, including electricity procurement cost and equipment operation and maintenance cost, C carbon,t : Carbon emission cost within the time step, ω1, ω2 are weight parameters.
[0096] In order to better adapt to different scenarios, the weights ω1 and ω2 are adjusted dynamically, and change in real time according to the carbon quota utilization rate or economic cost:
[0097]
[0098]
[0099] Among them C cost Refers to economic cost, C carbon Refers to the cost of carbon emissions.
[0100] The step S5 is specifically the following steps:
[0101] S51, using the learning strategy function π(A|S), the probability of selecting action a in state s to maximize the expected long-term accumulated reward G t :
[0102]
[0103] where γ∈[0, 1] is a discount factor used to balance short-term and long-term goals. E is the expected value operator, t is the current time point, k is a counter, and R is the reward obtained.
[0104] With each state transition of the park microgrid, there will be changes in carbon emissions. Through reinforcement learning training, adjusting the configurable items of the park equipment nodes, a feasible path with less carbon emissions can be deduced. The optimization algorithm based on the dual-objective optimization of operating cost and total carbon emissions is as follows:
[0105] in It is the decision-making behavior space of the microgrid.
[0106] S52. Build an LSTM deep learning network to predict environmental meteorological data.
[0107] S53. Build the STGCN model for carbon emission prediction, which consists of an input layer, an ST-ConvBlock layer, a GCN-Block layer, and an output layer.
[0108] S54. Build a deep reinforcement learning network: Use an intelligent agent to use the actor-critic algorithm and dynamically update parameters, where the actor refers to the strategy function π θ (a|s), the commentator is the scoring function V π (s). The actor performs a series of actions (a) according to the strategy, and the critic assigns a score to each action taken by the actor to evaluate the quality of the strategy. The critic's scoring function guides the direction of the actor's policy parameter update. That is, in each step of the loop, according to the direction guided by the critic network, in order to minimize the value of the scoring function, the actor network will generate an action a in state s to interact with the environment. The policy gradient and temporal difference methods are combined to update the weight matrices of the actor and critic respectively. The policy gradient and temporal difference methods are combined, and the Adam optimization algorithm is used to update the weight matrix.
[0109] The step S53 is specifically as follows:
[0110] S501, the enterprise graph structure is represented as an adjacency matrix and a feature matrix, and the input layer inputs training data;
[0111] S502, using the He initialization method, can effectively accelerate the convergence of the model and improve the accuracy of the model. Gaussian distribution is used to initialize the weight parameters with a variance of Where is the input dimension of the weight matrix, and the bias parameter can be initialized to 0 or any small value. The formula is as follows, Among them, W represents the weight parameter and n represents the input dimension of the weight matrix.
[0112] S503, construct the ST-Block layer, which consists of multiple one-dimensional convolutional layers and residual blocks to process time series data. The formula is as follows: X (k) =ReLU(LX (k-1) W (k-1) ). Among them, X (k-1) is the feature matrix of the k-1th layer, W (k-1) is the weight matrix of the k-1th layer, L is the graph Laplacian matrix, and ReLU is the activation function.
[0113] The combination of multiple one-dimensional convolutional layers and residual blocks can improve the feature extraction ability of the model, thereby better predicting time series data. The number and size of one-dimensional convolutional layers and residual blocks can be adjusted according to the characteristics of the data.
[0114] S504, construct the GCN-Block layer, which is composed of multiple graph convolutional layers and residual blocks, to process spatial data. The formula is as follows: Among them, X t represents the feature matrix at time t, A represents the adjacency matrix of the graph structure, is the identity matrix, is the degree matrix, W (1) and W (2) are the weight matrices of the two convolutional layers respectively, and σ represents the activation function.
[0115] The combination of multiple graph convolutional layers and residual blocks can improve the feature extraction ability of the model, thereby better predicting spatial data. The number and size of graph convolutional layers and residual blocks can be adjusted according to the characteristics of the data.
[0116] S505, the output layer includes a time domain convolution layer and a fully connected layer. The convolution kernel size of the time domain convolution layer is The number is C o , mapping the output to The fully connected layer is in Output
[0117] Introducing the concept of carbon neutrality in the entire cycle of park planning, construction, management and operation, combined with the microgrid carbon emission monitoring system, can achieve accurate statistics and integration of infrastructure energy consumption and material transformation. By establishing a local energy network (microgrid), various energy systems (mains power, photovoltaics, energy storage, etc.) can be connected, energy scheduling can be unified, and its control strategy can be optimized to achieve the overall optimal operating efficiency. Furthermore, this system supports the optimization of the dual goals of economy and carbon emission reduction, and integrates energy conservation, emission reduction, carbon fixation and carbon sink measures through digital means, so as to achieve low-carbon industry, green energy, shared facilities and resource recycling, and ultimately achieve carbon emission reduction in the microgrid system, and even achieve carbon neutrality.
Claims
1. A microgrid energy dispatching carbon emission reduction system based on reinforcement learning dual-objective optimization, including a data acquisition module, a data processing module, an intelligent dispatching module, an execution feedback module and a human-computer interaction module, characterized in that: The data acquisition module collects the energy consumption data, environmental parameters, carbon emission coefficient and other information of the equipment in the microgrid in real time through the sensor equipment; the data processing module cleans, extracts features and performs standardization processing on the collected data; Intelligent scheduling module: Use the spatiotemporal graph convolutional neural network (STGCN) to extract information features, further adopt the deep reinforcement learning algorithm based on actors and critics, use the model data dynamically updated in the time window for optimization, obtain the optimization strategy, and form a scheduling instruction set; Execution feedback module: According to the scheduling instruction set output by the intelligent scheduling module, dynamically schedule the operation status of the equipment in the microgrid; Human-computer interaction module: Provide a user interface to display the use of carbon quotas, the system operation status, and allow users to input adjustment requirements; The specific steps of building the system model in the intelligent scheduling module are as follows: S1. Set the park microgrid model, which is equipped with a distributed energy system, an energy storage system, and load equipment. S2. Determine the required energy and load range of each device in the microgrid. S3. Establish the objective function and its dynamic weight adjustment model. S4. Set the state-action-reward model: The state space of the microgrid includes the power supply and demand relationship between the microgrid nodes, the action space of the microgrid includes the scheduling of the microgrid power, and the reward model of the microgrid is designed according to the aforementioned state space and action space. S5. Establish a strategy model based on reinforcement learning of actors and critics. S6. Allocate power based on the model optimization results and form a scheduling instruction set.
2. According to claim 1, a microgrid energy scheduling carbon emission reduction system based on reinforcement learning dual-objective optimization is characterized by: The step S2 is specifically the following steps: S21. Obtain the capacity range of the distributed energy system: Assuming that the total number of distributed energy (DG) systems is N, the capacity of the kth distributed energy system is The minimum capacity is The maximum capacity is The range is S22. Calculate the power status of the energy storage device: SOC s (t) represents the power state of the energy storage device, and the change of the power state of the energy storage device is: Among them, η s is the charging and discharging efficiency of the energy storage device, It indicates the charge and discharge time. is the charging power of the energy storage device, is the maximum charging power of the energy storage device, and its power range is Represents the discharge power of the energy storage device, is the maximum discharge power of the energy storage device, and its power range is SOC s,min It is the minimum value of the energy storage device, SOC s,max It is the maximum value of the energy storage device. The capacity limit of the energy storage device is: SOC s,min ≤SOC s (t)≤SOC s,max . S23. Calculate the total load power of the load equipment: Among them PD i Represents the peak load of the i-node unit, PF i represents the power factor at time t, f i represents the load fluctuation probability distribution of node i, L m (t) represents the total load power of the microgrid. S24. Calculate the demand response power limit range: The power demand of the zth demand response block in the microgrid is denoted as q z (t),u z (t) is a binary variable indicating whether the zth block response module responds. The demand response needs to meet the following constraints: u z-1 (t)≥u z (t),for z=2,…,Z S25. Calculate the total load of the microgrid: The microgrid needs to meet the following supply and demand power balance conditions: Where L m (t) represents the total load of the microgrid, It indicates the amount of electricity purchased by microgrid m from the main grid (the park distribution network from the city grid).
3. According to claim 1, a microgrid energy scheduling carbon emission reduction system based on reinforcement learning dual-objective optimization is characterized by: The step S3 is specifically the following steps: S31. Establish a model based on the economic target, and the calculation method is as follows: Is related to power generation The related nonlinear equations. The function can be a quadratic polynomial, which represents the power generation cost of distributed generation. Among them, the coefficient N t Represents a dispatch cycle, N k Represents the total number of distributed energy devices, η m represents the network loss during power transmission, λ(t) represents the time-varying retail electricity price at the common coupling point, Represents the cost of power exchange between the microgrid and the main grid, represents the internal dispatching cost of the microgrid, It represents the unit price of dispatch. is a binary variable with a value of 0 or 1, indicating the zth demand response module Whether it is scheduled, ρ s |SOC s (t)-SOC s (t-1)| represents the loss cost of energy storage equipment, ρ s A coefficient that indicates the performance degradation and shortened life of energy storage equipment caused by changes in electrical quantity. S32. Establish a model based on environmental goals. The calculation method is as follows: The total carbon emissions of the park in a scheduling cycle E m and carbon emission cost C carbon It can be expressed as: Where E m (t) represents the total carbon emissions of the microgrid at time step t, ETS(t) represents the carbon emission market trading system pricing at time step t, and E process (t) represents the total amount of greenhouse gas fugitive emissions from the production process at time step t, which can be obtained by the following formula: EF j Represents the process carbon emission factor of the jth equipment. S33. The comprehensive objective function combines economic and carbon reduction goals and is a weighted reward function that minimizes total carbon dioxide emissions and current net costs, expressed as: R t =ω1·(-C cost,t )+ω2·(-C carbon,t ) C cost,t : Economic cost of energy-consuming equipment within the time step, C carbon,t : Carbon emission cost within the time step, ω1, ω2 are weight parameters. Dynamically adjust weights ω1 and ω2 according to real-time changes in carbon quota utilization or economic costs:
4. According to claim 1, a microgrid energy scheduling carbon emission reduction system based on reinforcement learning dual-objective optimization is characterized by: The step S5 is specifically the following steps: S51, using the learning strategy function π(A|S), the probability of selecting action A in state S to maximize the expected long-term accumulated reward G t : where γ∈[0, 1] is the discount factor, E is the expected value operator, t is the current time point, k is a counter, and R is the reward obtained. The optimization algorithm based on the dual-objective optimization of operating cost and total carbon emissions is as follows: in It is the decision-making behavior space of the microgrid. S52. Build an LSTM deep learning network to predict environmental meteorological data. S53. Build an STGCN model for extracting microgrid network and node features. The STGCN model consists of an input layer, an ST-Conv Block layer, a GCN-Block layer and an output layer. S54. Build a deep reinforcement learning network: Use the intelligent agent to use the actor-critic algorithm, dynamic parameter update, combine the policy gradient and temporal difference methods, and use the Adam optimization algorithm to update the weight matrix.
5. According to claim 1, a microgrid energy scheduling carbon emission reduction system based on reinforcement learning dual-objective optimization is characterized by: The step S53 is specifically as follows: S501, the enterprise graph structure is represented as an adjacency matrix and a feature matrix, and the input layer inputs training data; S502, using the He initialization method, using Gaussian distribution to initialize weight parameters, the variance is Where is the input dimension of the weight matrix, and the bias parameter can be initialized to 0 or any small value. The formula is as follows, Among them, W represents the weight parameter and n represents the input dimension of the weight matrix. S503, construct the ST-Block layer, which consists of multiple one-dimensional convolutional layers and residual blocks to process time series data. The formula is as follows: X (k) =ReLU(LX (k-1) W (k-1) ). Among them, X (k-1) is the feature matrix of the k-1th layer, W (k-1) is the weight matrix of the k-1th layer, L is the graph Laplacian matrix, and ReLU is the activation function. S504, construct the GCN-Block layer, which is composed of multiple graph convolutional layers and residual blocks, to process spatial data. The formula is as follows: Among them, X t represents the feature matrix at time t, A represents the adjacency matrix of the graph structure, I is the identity matrix, is the degree matrix, W (1) and W (2) are the weight matrices of the two convolutional layers respectively, and σ represents the activation function. S505, the output layer includes a time domain convolution layer and a fully connected layer. The convolution kernel size of the time domain convolution layer is The number is C o , mapping the output to The fully connected layer is in Output
Citation Information
Cited By
Optimal configuration method for hybrid energy storage capacity of highway micro-grid
CN120810719A
Pumped storage power station carbon optimization method based on deep reinforcement learning
CN120875257A
Park comprehensive energy scheduling method, system, equipment and medium
CN121052606A
Bridge full life cycle carbon emission reduction cost efficiency optimization method and device based on reinforcement learning, computer equipment and readable storage medium
CN121525913A
Reinforcement learning-based bridge life cycle carbon emission reduction cost efficiency optimization method and device, computer equipment and readable storage medium
CN121525913B