Demand response scheduling decision method and apparatus based on electricity-carbon coupling, device, and system

By employing an electricity-carbon joint demand response scheduling decision-making method, and utilizing Markov decision-making and deep reinforcement learning to optimize the user-side demand response model, the problem of insufficient incentive mechanisms in existing technologies is solved. This enables coordinated adjustment of user-side load and distributed energy resources, thereby improving system flexibility and resource utilization efficiency.

WO2026123453A1PCT designated stage Publication Date: 2026-06-18GUANGZHOU INST OF ENERGY CONVERSION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GUANGZHOU INST OF ENERGY CONVERSION CHINESE ACAD OF SCI
Filing Date
2025-02-06
Publication Date
2026-06-18

Smart Images

  • Figure CN2025075928_18062026_PF_FP_ABST
    Figure CN2025075928_18062026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of distributed energy response scheduling, and discloses a demand response scheduling decision method and apparatus based on electricity-carbon coupling, a device, and a system. The method comprises: acquiring information data of a user end, wherein the information data comprises an electricity-carbon coupling locational marginal price signal corresponding to a distribution system operator, the information data is determined on the basis of a locational marginal price model of electricity-carbon coupling at a distribution system end, and the locational marginal price model is determined by the distribution system operator by using a distribution system branch flow algorithm; determining a user-side demand response model, wherein the user-side demand response model is constructed on the basis of the information data and with the objective of minimizing a power factor; and using Markov decision to perform mathematical transformation on the user-side demand response model, and using a deep reinforcement learning method to perform iterative optimization calculation to obtain a scheduling decision scheme, so as to coordinate a user-side load of the user end with distributed generation. The present application aims to implement coordinated adjustment and scheduling of distributed energy and user loads.
Need to check novelty before this filing date? Find Prior Art

Description

Demand response scheduling decision-making method, apparatus, equipment and system based on electricity-carbon co-operation

[0001] This application claims priority to Chinese Patent Application No. 202411832960.4, filed on December 13, 2024, entitled "Method, Apparatus and Device for Demand Response Scheduling Based on Electricity-Carbon Integration", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of distributed energy response dispatch, and in particular to a demand response dispatch decision-making method, apparatus, equipment and system based on electricity-carbon joint. Background Technology

[0003] With the widespread integration of distributed energy resources into power systems, the system's flexibility and resilience have been significantly enhanced, bringing considerable environmental benefits. However, existing research still has shortcomings in the coordinated optimization of demand response and distributed energy resources, especially in the design of mechanisms to effectively incentivize users to actively participate in demand response. Currently, most methods, when scheduling and optimizing power systems, focus on traditional power flow models and economic dispatch strategies, failing to fully consider user behavioral incentives in a diversified electricity market environment.

[0004] Existing methods attempt to achieve demand response through simple electricity price signals, but in the complex and volatile electricity market, user response is often unsatisfactory. Overly simplistic incentive mechanisms and dispatch strategies fail to fully mobilize user enthusiasm and may lead to unbalanced resource allocation in practical applications. Due to significant differences in user geographical location, electricity demand, and the distributed energy resources connected, existing mechanisms still have limitations in terms of fairness and response effectiveness. Furthermore, while some studies attempt to optimize dispatch strategies on the user side, they are mostly based on idealized assumptions and lack in-depth modeling of user demand response behavior and consideration of practical constraints. Effective strategies and mechanisms to incentivize user participation in demand response are scarce, failing to fully realize the potential of distributed energy resources and user load adjustment. Summary of the Invention

[0005] The purpose of this application is to provide a demand response scheduling decision-making method, device, equipment and system based on electricity-carbon joint, which can realize the coordinated adjustment and scheduling of distributed energy and user load.

[0006] To achieve the above objectives, this application provides the following solution:

[0007] Firstly, this application provides a demand response scheduling decision-making method based on electricity-carbon joint principles, which is applied to a distributed power grid framework system. The distributed power grid framework system includes: a communication layer and a power layer connected together; the communication layer includes a demand response aggregator for connecting user terminals and the distribution network; the power layer is used for allocating and managing the power flow of the distribution network; the demand response scheduling decision-making method based on electricity-carbon joint principles includes: acquiring user terminal information data; the information data includes: marginal electricity price signals of the electricity-carbon joint nodes corresponding to the operators at the distribution network end; the information data is... The nodal marginal electricity price model is determined based on the distribution network-side electricity-carbon linkage; wherein, the nodal marginal electricity price model is determined by the distribution network operator using the distribution network branch flow algorithm; the nodal marginal electricity price model includes: objective function and constraints; a user-side demand response model is determined; the user-side demand response model is constructed based on the information data with the goal of minimizing the power coefficient; Markov decision is used to mathematically transform the user-side demand response model, and deep reinforcement learning is used for iterative optimization to obtain a scheduling decision scheme; the scheduling decision scheme is used to coordinate user-side load and distributed energy generation at the user end.

[0008] In one exemplary embodiment, the expression for the objective function is:

[0009] Where F is the objective function; B dg B is a distributed generation collection. re B is a collection of new energy power generation facilities. ess For energy storage collection, B gd It is a collection of the upper-level power grid. Let t be the carbon price in the carbon market. Cost of distributed generation; Let t be the carbon emission intensity of distributed generation; For distributed generation power; Cost of new energy power generation; For new energy power generation; Cost of energy storage degradation; Power generation for energy storage systems; The electricity price is determined by the higher-level power grid. Carbon emission intensity of the upper-level power grid; denoted as active power of the upper-level power grid; g represents the serial number of the distributed generation equipment; r represents the serial number of the new energy generation equipment; e represents the serial number of the energy storage equipment; and b represents the serial number of the node participating in the response.

[0010] In an exemplary embodiment, the constraints include: power balance constraints, reactive power balance constraints, voltage difference constraints, power flow limits of transmission lines, voltage magnitude constraints, line current limits, limits on distributed generation power, capacity limits of distributed generators, range limits on the reactive power of distributed generators, range limits on the active power of energy storage systems, power range limits of renewable energy sources, and power factor limits of renewable energy sources.

[0011] In an exemplary embodiment, the objective function corresponding to the user-side demand response model is:

[0012] Where D is the objective function corresponding to the user-side demand response model; T is the total number of time steps; B is the total number of nodes participating in the response; t is the time step number; and b is the node number participating in the response. This is the marginal electricity price signal for the combined electricity and carbon node; The active power load of the node after demand response; P represents the maximum load of node b participating in the response. b,t Active load before demand response; ΔP b,max This represents the maximum response size.

[0013] In an exemplary embodiment, a Markov decision is used to mathematically transform the user-side demand response model, and a deep reinforcement learning method is used for iterative optimization to obtain a scheduling decision scheme, specifically including:

[0014] Based on the aforementioned user-side demand response model, Markov decision-making is employed at discrete time steps to transform the response into a decision, yielding the transformation result. This transformation result includes: an agent, a global state, local observations, a set of actions, a reward function, and a state transition function. The agent represents the user-side demand response participants at the user end. The global state is the user-side demand response state at the user end. The local observations include the active power load at nodes after demand response and the marginal electricity price signal at the combined electricity and carbon nodes. The set of actions is the set of actions by the power layer for allocating and managing the power flow of the distribution network.

[0015] A deep reinforcement learning method is adopted, and an entropy regularization strategy based on the entropy term is used to maximize the cumulative expected reward of the agent. The transformation result is iteratively optimized to obtain a scheduling decision scheme.

[0016] In an exemplary embodiment, the process of iteratively optimizing and solving using a deep reinforcement learning method specifically includes:

[0017] Acquire initial parameter data; the initial parameter data includes: weight data, local observations, action set, and reward function; the local observations, action set, and reward function are stored in the replay buffer;

[0018] At each time step, each agent selects action data from the action set based on the current local observations and determines the current cumulative expected reward based on the reward function;

[0019] Based on the current cumulative expected reward, the action data and the corresponding local observations are used as local experience and stored in the replay buffer.

[0020] Batch data is sampled from the playback buffer, based on the formulas corresponding to the policy network loss function and the value network loss function:

[0021] Update the weight data to obtain the updated weights;

[0022] Wherein, J(π) i () represents the policy network loss function; To find the mathematical symbol for the expected value; π i (a i,t |o i,t ) represents the probability distribution of the policy network for agent i; o i,t Let be the observation value of the i-th agent at time step t; For playback buffer; Q i (o i,t a i,t ) represents the soft Q-value of the value network estimate, indicating the Q-value given a state and action; a i,t Let S be the action value of the i-th agent at time step t; α is the entropy coefficient; S t Let a be the state value at time step t; 1:I,t The action before the t-th time step; r i,t Let s be the reward value of the i-th agent at time step t; t+1 Q represents the state value at time step t+1. i (s t ,a 1:I,t ) represents the Q-value of the target Q-value network, and soft updates (soft target networks) are typically used to avoid overestimation; y i,t The target value; L(Q) i ) represents the value network loss function; π i It is a probability distribution;

[0023] Based on updated weights and local experience, the global state and local observations are updated according to all selected action data to obtain a scheduling decision scheme.

[0024] Secondly, this application provides a demand response scheduling decision-making device based on electricity-carbon integration, including: a data acquisition module, a model determination module, and a scheme determination module.

[0025] The data acquisition module is used to acquire user-end information data; the information data includes: the marginal electricity price signal of the electricity-carbon joint node corresponding to the operator at the distribution network end; the information data is determined based on the node marginal electricity price model of the electricity-carbon joint node at the distribution network end; wherein, the node marginal electricity price model is determined by the operator at the distribution network end using the distribution network branch flow algorithm; the node marginal electricity price model includes: objective function and constraints.

[0026] The model determination module is used to determine the user-side demand response model; the user-side demand response model is constructed based on the information data with the goal of minimizing the power coefficient.

[0027] The scheme determination module is used to mathematically transform the user-side demand response model using Markov decision and to iteratively optimize and solve it using deep reinforcement learning methods to obtain a scheduling decision scheme; the scheduling decision scheme is used to coordinate the user-side load and distributed energy generation at the user end.

[0028] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the demand response scheduling decision method based on the electricity-carbon joint described above.

[0029] Fourthly, a distributed power distribution network framework system includes: a communication layer and a power layer connected together; the communication layer includes a demand response aggregator for connecting user terminals and the power distribution network; the power layer is used for distributing and managing the power flow of the power distribution network.

[0030] According to the specific embodiments provided in this application, this application has the following technical effects:

[0031] This application provides a method, apparatus, equipment, and system for demand response scheduling decision-making based on electricity-carbon joint principles. It determines the user-side demand response model by acquiring the electricity-carbon joint node marginal price signal corresponding to the operator at the distribution network end. Markov decision theory is used to mathematically transform the user-side demand response model, and deep reinforcement learning is used for iterative optimization to obtain a scheduling decision scheme. Since the electricity-carbon joint node marginal price signal is obtained based on the node marginal price model determined by the operator at the distribution network end using the distribution network branch flow algorithm, it fully considers the behavioral incentives of users in a diversified electricity market environment. Furthermore, the user-side demand response model is constructed with the goal of minimizing the power coefficient, which can mobilize the user's response enthusiasm. In addition, by using deep reinforcement learning for iterative optimization, electricity demand, i.e., user-side load, and connected distributed energy resources can be coordinated, avoiding significant discrepancies and fully leveraging the potential for adjustment between distributed energy resources and user load. Therefore, coordinated adjustment and scheduling of distributed energy resources and user load can be achieved. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 is a flowchart of the demand response scheduling decision-making method based on electricity-carbon joint provided in this application;

[0034] Figure 2 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] This application aims to fill the gap in existing solutions by designing an innovative demand response mechanism that provides precise user-side incentive strategies. This encourages users to proactively adjust their electricity consumption behavior based on electricity price signals, achieving more efficient resource use and greater response flexibility. This mechanism aims to strengthen user-grid interaction, improve the economic efficiency and sustainability of system operation, and promote the wider and more effective application of distributed energy resources.

[0037] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] In an exemplary embodiment, as shown in Figure 1, a demand response scheduling decision-making method based on electricity-carbon co-operation is provided. This method is executed by a computer device, specifically a terminal or server, or both. In this embodiment, the method is described using a server as an example.

[0039] This application provides a demand response scheduling decision-making method based on electricity-carbon co-operation, which is applied to a distributed power grid framework system. The distributed power grid framework system includes: a communication layer and a power layer; the communication layer includes a demand response aggregator for connecting user terminals and the distribution network; the power layer is used for allocating and managing the power flow of the distribution network.

[0040] As shown in Figure 1, the demand response scheduling decision-making method based on electricity-carbon co-operation provided in this application includes the following steps:

[0041] Step 100: Obtain user-side information data. The information data includes: the marginal electricity price signal of the electricity-carbon joint node corresponding to the operator on the distribution network side; the information data is determined based on the node marginal electricity price model of the electricity-carbon joint node on the distribution network side; wherein, the node marginal electricity price model is determined by the operator on the distribution network side using the distribution network branch flow algorithm; the node marginal electricity price model includes: objective function and constraints.

[0042] Step 200: Determine the user-side demand response model. The user-side demand response model is constructed based on information data, with the goal of minimizing the power factor.

[0043] Step 300: The user-side demand response model is mathematically transformed using Markov decision theory, and then iteratively optimized using deep reinforcement learning to obtain a scheduling decision scheme. This scheduling decision scheme is used to coordinate user-side loads and distributed energy generation.

[0044] In one embodiment, the objective function is expressed as:

[0045] In another embodiment, the expression for the objective function can be:

[0046] Where F is the objective function; B dg B is a distributed generation collection. re B is a collection of new energy power generation facilities. ess For energy storage collection, B gd It is a collection of the upper-level power grid. Let t be the carbon price in the carbon market. Cost of distributed generation; Let t be the carbon emission intensity of distributed generation; For distributed generation power; Cost of new energy power generation; For new energy power generation; Cost of energy storage degradation; Power generation for energy storage systems; The electricity price is determined by the higher-level power grid. Carbon emission intensity of the upper-level power grid; denoted as active power of the upper-level power grid; g represents the serial number of the distributed generation equipment; r represents the serial number of the new energy generation equipment; e represents the serial number of the energy storage equipment; and b represents the serial number of the node participating in the response.

[0047] The constraints include: power balance constraints, reactive power balance constraints, voltage difference constraints, power flow limits for transmission lines, voltage magnitude constraints, line current limits, power limits for distributed generation, capacity limits for distributed generators, range limits for reactive power of distributed generators, range limits for active power of energy storage systems, power range limits for renewable energy, and power factor limits for renewable energy.

[0048] The objective function corresponding to the user-side demand response model is:

[0049] Where D is the objective function corresponding to the user-side demand response model; T is the total number of time steps; B is the total number of nodes participating in the response; t is the time step number; and b is the node number participating in the response. This is the marginal electricity price signal for the combined electricity and carbon node; The active power load of the node after demand response; P represents the maximum load of node b participating in the response. b,t Active load before demand response; ΔP b,max This represents the maximum response size.

[0050] In one embodiment, a Markov decision is used to mathematically transform the user-side demand response model, and a deep reinforcement learning method is used for iterative optimization to obtain a scheduling decision scheme, specifically including:

[0051] Based on the user-side demand response model, Markov decision-making is used to transform the response into a decision at discrete time steps, resulting in the transformation result. The transformation result includes: agent, global state, local observations, action set, reward function, and state transition function. The agent represents the user-side demand response participants at the user end; the global state is the user-side demand response state at the user end; local observations include the active power load of nodes after demand response and the marginal electricity price signal of the combined electricity and carbon nodes; the action set is the set of power layer allocation and management of the power flow of the distribution network.

[0052] A deep reinforcement learning approach is adopted, with an entropy regularization strategy based on the entropy term aiming to maximize the agent's cumulative expected reward. The transformation results are iteratively optimized to obtain a scheduling decision scheme.

[0053] The process of iteratively optimizing and solving using deep reinforcement learning methods specifically includes:

[0054] Acquire initial parameter data; the initial parameter data includes: weight data, local observations, action set, and reward function; the local observations, action set, and reward function are stored in the replay buffer.

[0055] At each time step, each agent selects action data from the action set based on the current local observations and determines the current cumulative expected reward based on the reward function.

[0056] Based on the current cumulative expected reward, the action data and the corresponding local observations are used as local experience and stored in the replay buffer.

[0057] Sample batch data from the playback buffer, based on the formulas corresponding to the policy network loss function and the value network loss function:

[0058] Update the weight data to obtain the updated weights;

[0059] Wherein, J(π) i () represents the policy network loss function; To find the mathematical symbol for the expected value; π i (a i,t |o i,t ) represents the probability distribution of the policy network for agent i; o i,t Let be the observation value of the i-th agent at time step t; For playback buffer; Q i (o i,t ,a i,t ) represents the soft Q-value of the value network estimate, indicating the Q-value given a state and action; a i,tLet S be the action value of the i-th agent at time step t; α is the entropy coefficient; S t Let a be the state value at time step t; 1:I,t The action before the t-th time step; r i,t Let s be the reward value of the i-th agent at time step t; t+1 Q represents the state value at time step t+1. i (s t ,a 1:I,t ) represents the Q-value of the target Q-value network, and soft updates (soft target networks) are typically used to avoid overestimation; y i,t The target value; L(Q) i ) represents the value network loss function; π i It represents a probability distribution.

[0060] Based on updated weights and local experience, the global state and local observations are updated according to all selected action data to obtain a scheduling decision scheme.

[0061] This application aims to achieve effective coordination of user-side demand response in distribution networks, based on a distributed framework system for distribution networks. The system architecture consists of two main parts: a communication layer and a power layer. The communication layer is operated by a demand response aggregator, acting as a bridge between the user end and the distribution network. This aggregator can collect, transmit, and exchange limited information to help users make informed decisions in complex electricity market environments. The power layer involves the actual allocation and management of power flow, rationally configuring distributed energy resources (DERs), including energy demand (ED) response units, distributed generation (DG), and renewable energy-based generating units, i.e., renewable energy systems (RES), to ensure the diversity and sustainability of the energy structure.

[0062] Users receive the Joint Nodal Marginal Price (C-LMP) signal, calculated in real time by the Distribution System Operator (DSO), where LMP stands for Locational Marginal Price. This allows users to optimize their electricity consumption patterns, enabling flexible responses and maximizing energy efficiency under low-carbon conditions. The DSO plays a crucial role in this mechanism, managing controllable distributed energy resources and performing optimal power flow (OPF) calculations to ensure the stability and security of the distribution network.

[0063] Unlike traditional centralized control methods, this application assumes that all users at the user end have autonomous decision-making capabilities and independent electricity needs and preferences. Users operate within a distributed framework, and personal information is not stored in the central control system, thus ensuring privacy and data security. The introduction of the aggregator aims to bridge the information asymmetry problem among users, providing them with necessary information and reasonable incentive mechanisms to promote collaboration and resource sharing among users, thereby optimizing overall demand response behavior and improving system flexibility and user participation.

[0064] In practical applications, the operation steps of the method mentioned in this application can be as follows.

[0065] S1: Establish a user-side demand response model.

[0066] The goal of the user-side demand response model is to minimize electricity costs while optimizing response behavior. In other words, it is built with the goal of minimizing the electricity coefficient.

[0067] The objective function corresponding to the user-side demand response model can be expressed as:

[0068] Where D is the objective function corresponding to the user-side demand response model; T is the total number of time steps; B is the total number of nodes participating in the response; t is the time step number; and b is the node number participating in the response. This is the marginal electricity price signal for the combined electricity and carbon node; This refers to the active power load of the node after demand response.

[0069] Post-response load constraints:

[0070] in, The maximum load of node b participating in the response.

[0071] Among them, P b,t Active load before demand response; ΔP b,max This represents the maximum response size.

[0072] S2: The user-side demand response model is expressed as a distributed partially observable Markov decision process.

[0073] The significance of transforming the user-side demand response problem into a distributed, partially observable Markov decision process lies in providing a structured and formal description of the complex electricity demand response coordination problem, making it suitable for solving using reinforcement learning algorithms. This transformation has three main benefits: First, it provides a clear mathematical model of the problem, clarifying elements such as states, actions, and rewards, facilitating algorithmic processing. Second, by introducing partially observable characteristics, it effectively addresses uncertainties in the power system and the randomness of user behavior, thereby enhancing the model's adaptability to real-world environments. Finally, the distributed design allows users to make independent decisions with only local information, achieving decentralized control, ensuring data privacy, and improving system flexibility. In summary, this modeling step not only provides a foundation for applying multi-agent reinforcement learning but also enables users to optimize their responses in real time after receiving marginal electricity price signals from the joint electricity and carbon node, promoting improved energy efficiency and system stability.

[0074] Specifically, in discrete time steps, a distributed partially observable Markov decision process is defined as follows:<I,S,O,A,R,T1,γ> This includes I agents (representing participants in user-side demand response), a set of global states s∈S, a set of local observations O, a set of actions A, a set of reward functions R, and a state transition function T1. γ is a discount factor. S is the global state.

[0075] Let the time interval between two consecutive time steps be Δt = 1 hour.

[0076] At each time step t, each agent i, based on its local observations, that is, the observation o of the i-th agent at time step t... i,t And select an action a based on the control strategy μ(o). i,t This refers to the action value of the i-th agent at time step t. The environment transitions to the next state according to the state transition function T1, obtaining the state value s at time step (t+1). t+1 Each agent i receives a reward r. i,t And acquire new local observations. i,t+1 This process continues, forming the trajectory τ of each agent i's observations, actions, and rewards. i τ i =(o i,1 ,a i,1 ,r i,1 ,o i,2 ,…,r i,T ), its mapping is O i ×A i ×O i →R.

[0077] Among them, o i,1For the local observation of the i-th agent at the first time step; a i,1 r represents the action data of the i-th agent at the first time step. i,1 The reward for the i-th agent at the first time step; i,2 For the local observation of the i-th agent at the second time step; r i,T Let this be the reward for the i-th agent at the T-th time step.

[0078] Each agent i aims to maximize its cumulative discounted reward, which is also its cumulative expected reward. Where γ∈[0,1) is the discount factor, and T=24 hours is the daily time range. t This is the attenuation coefficient.

[0079] Observations: At time step t, the local observations of each agent i are as follows: The observation consists of two parts: exogenous states unaffected by actions, including the node locations of users with demand response capabilities. Active power demand of nodes and reactive power demand and the active power of renewable energy at the nodes The endogenous state, which serves as the action feedback signal, is modeled as the node load after the demand response.

[0080] Action: At time step t, each agent i controls its action:

[0081] in, A represents the change in active power demand response at a node; i Let i be the set of actions of agent i.

[0082] State transition: The state transition is determined by the following formula: s t+1 =T1(s t ,0 1:I,t ,a 1:I,t ,ω t ).

[0083] State transitions are influenced by the current environmental state st and the local observations O of all agents. 1:I,t and the action a before the t-th time step 1:I,t and environmental random parameters The combination of factors has an impact.

[0084] in, The active power price of the main power grid; The reactive power price of the main power grid; The active power load of the node after demand response; This is for reactive power demand.

[0085] Determined by market conditions, Whether to participate in the demand response decision is decided by the users on that day. Influenced by energy consumption behavior, and renewable energy generation This is influenced by solar radiation or wind speed. Explicitly modeling the distribution of these factors is extremely difficult. In the field of machine learning, reinforcement learning solves this problem through a data-driven approach that does not rely on an accurate model of underlying uncertainties but instead learns dynamic characteristics directly from the data source.

[0086] Reward: At the end of time step t, each agent i receives its reward r. i,t First, users connected to the grid who can provide demand response aim to maximize their revenue by providing active and reactive power services. This can be calculated based on the magnitude of the demand response (i.e., actions) and the marginal electricity price at each node connected to the grid. Second, all users need to ensure their total load meets the daily production needs of their businesses. Since the distributed partially observable Markov decision process is a dynamic decision process, requiring sufficient physical constraints on total load as part of the environment, these constraints are not directly accessible to the demand response regulating agent. Therefore, a reward shaping mechanism is needed to penalize violations of this constraint, and it is assumed that this mechanism is effective in addressing this challenge.

[0087] When users choose to participate in grid demand response. The nodal marginal electricity price of the combined electricity and carbon Calculated charging (discharging) cost (benefit), P b,t The first term represents the active power load before demand response. The second term represents the penalty incurred at the end of each day for failing to meet the sufficient load requirement, where κ is the penalty factor corresponding to the demand response term. P b,t This refers to the active load prior to demand response.

[0088] S3: C-LMP model of nodal marginal electricity price for distribution network side electricity carbon integration.

[0089] The nodal marginal price (C-LMP) model, which integrates electricity and carbon emission costs on the distribution network side, plays a crucial bridging role. It aims to combine electricity costs and carbon emission costs into an actionable, real-time price signal, guiding users and distributed energy participants to optimize their behavior. As input to user demand response models, it helps users adjust their electricity consumption strategies based on this signal. Simultaneously, in Markov Decision Processes (MDP) modeling, environmental observations are incorporated to support user decision optimization at each time step.

[0090] Under current regulations, the carbon trading market primarily focuses on direct carbon emissions in Scope 1. In this context, an advanced Continuous Emission Monitoring System (CEMS) is used, assuming grid operators can monitor data from all direct carbon emission sources within their jurisdiction in real time. Since power generation companies have already participated in the national carbon market, their generation costs already incorporate carbon emission costs; therefore, the electricity price paid by users can be considered to cover their carbon emission responsibilities. Within the same power supply area, users are considered to share a unified carbon emission responsibility. The marginal nodal carbon price (C-LMP) is calculated using real-time generation-side carbon prices to more fairly distribute environmental responsibility within the region, avoiding an excessive carbon burden borne solely by power generation companies and high-energy-consuming enterprises. This method introduces carbon price signals, which helps incentivize power generation companies to undertake low-carbon transformation and avoids the complexities of green electricity market auctions, thus streamlining the coupling between the carbon market and the electricity market.

[0091] More specifically, a distribution network branch flow algorithm operated by DSO is introduced. For each time step t, DSO will solve the following optimization problem based on the existing situation, calculate the marginal electricity price of the combined electricity and carbon generation, and then adjust the internal power generation equipment to maximize its own objective function. The objective function is a cost minimization problem, including the power cost of distributed generation (DG), photovoltaic (PV), energy storage system (ESS), and the power purchased from the upstream grid.

[0092] The constraints are as follows:

[0093] Formula (5) represents the power balance constraint for node b, ensuring that total power generation equals total load demand plus transmission power loss. The active power of the upstream power grid. For distributed generation active power, For active power of renewable energy, P represents the active power load of the node after demand response. bp,t This refers to the active power loss of the line.

[0094] B dg B is a distributed generation collection. res B is a collection of new energy power generation facilities. ess For energy storage collection, B gd For the upper-level power grid. B edd is the set of demand response nodes; bp is the demand response index; L is the node number of any connected node; B is the set of network branches; and B is the set of network nodes.

[0095] Formula (6) above represents the reactive power balance constraint at node b, ensuring that total reactive power generation equals total reactive power demand plus transmission reactive power loss, where v b,t and v p,t Let r be the square of the voltage magnitudes of any two adjacent nodes. bp and x bp These represent the resistance and reactance of the corresponding connection lines, l bp,t Let be the square of the current magnitude at any adjacent node. This refers to the reactive power of the upstream power grid. For distributed generation reactive power, Reactive power of renewable energy Q represents the reactive load of nodes after demand response. bp,t This refers to the active power loss of the line.

[0096] Formula (7) above is the node voltage difference constraint connecting bp, which represents the relationship between voltage difference and transmission power.

[0097] Formula (8) above is the power flow limit for the transmission line, ensuring that the power flow does not exceed the line capacity.

[0098] Formula (9) above is a constraint on the magnitude of the node voltage, ensuring that the node voltage is within a safe range, where v and These are the squares of the maximum and minimum node voltages, respectively.

[0099] Formula (10) above is the line current limit to ensure that the line current does not exceed the maximum capacity, where It is the maximum value of the square of the current magnitude of adjacent nodes.

[0100] Formula (11) above represents the limit on distributed generation power, ensuring that the power is within the allowable range. and These represent the minimum and maximum values ​​of distributed generation power, respectively.

[0101] Formula (12) above represents the capacity limit for distributed generators, ensuring that the apparent power does not exceed the maximum capacity. Formula (13) above defines the range limit for the reactive power of distributed generators, ensuring that the reactive power remains within the allowable range. This refers to the power angle of the distributed generator.

[0102] Formula (14) above represents the range limit of the active power of the energy storage system, ensuring that the energy storage system operates within the allowable range. and These represent the minimum and maximum values ​​for energy storage charging and discharging, respectively. e is the energy storage index number; E is the energy storage set.

[0103] Formula (15) above represents the power range of renewable energy sources, ensuring that their output power is within the allowable range, where This represents the maximum power output of renewable energy sources.

[0104] Formula (16) above represents the power factor limitation for renewable energy, which restricts reactive power output. This refers to the power angle of a renewable energy generator.

[0105] S4: Use the MASAC (Multi-Agent Soft Actor-Critic) algorithm to solve the above distributed part of the observable Markov decision process.

[0106] Multi-agent Soft Actor-Critic (MASAC) is a deep reinforcement learning method that extends the single-agent soft actor-critic (SAC) algorithm for multi-agent environments. In MASAC, each agent is trained through independent policy networks (Actors) and value networks (Critics) to optimize its policy and maximize rewards in dynamic environments. The algorithm aims to achieve more stable and robust policy learning in multi-agent systems by maximizing the cumulative expected reward for each agent while introducing an entropy term to encourage policy exploration.

[0107] Policy update: Policy π for each agent i i (a i |o i Based on its local observations i Select action a i The goal of the strategy update is to maximize the following objective function.

[0108] Where α is the entropy coefficient, used to balance exploration and utilization. It is the playback buffer.

[0109] Value function update: the commentator network Q for each agent i i (s,a 1:I The update is performed using the Temporal-Difference (TD) method. The update objective is to minimize the following mean squared error.

[0110] Among them, y i,t That is the target value.

[0111] Entropy regularization: entropy term αlogπ i (a i,t |o i,t The introduction of entropy coefficient α ensures that the policy does not converge to a deterministic policy prematurely, thereby increasing the likelihood of exploration. The entropy coefficient α can be fixed or dynamically updated through automatic policy adjustment to control the level of exploration.

[0112] Using the MASAC algorithm to solve the above distributed part, the Markov decision process can be observed as follows.

[0113] The MASAC algorithm steps for I agents are as follows:

[0114] Parameter initialization: Initialize parameters θ and φ for each agent's policy network and critic network, and copy them to their respective target network weights θ′ and φ′.

[0115] Experience Acquisition: Initialize a shared replay buffer for all agents. The agent interacts with the environment, collects observations, actions, and rewards, and stores them in the playback buffer.

[0116] Main loop:

[0117] At each time step t, each agent i selects an action based on the current observation and interacts with the environment: for each round (i.e., each day), the loop starts from 1 to M:

[0118] Random process initialization: Initialize a random process for action exploration. in Let Variance be the variance.

[0119] State and observation initialization: Record the experience of each agent and update the global s0 and local state o. i,0 .

[0120] Time step loop: For each time step (i.e., every hour) t=1 to T, the loop begins.

[0121] Action selection: For each agent i, based on the current observation o i,t Select action a i,t =μ(o i,t ).

[0122] Execute actions: Perform all actions of the agents a t =[a 1,t ,…,a I,t It is applied to power distribution networks.

[0123] Power Flow Calculation and Nodal Marginal Price Acquisition: Distribution System Operator (DSO) solves branch power flow algorithm (4) and obtains... λ b,t For each node's LMP.

[0124] Rewards and observation records: For each agent, observe the current reward r. i,t And the next observation o i,t+1 Then local experience (o i,t ,a i,t ,r i,t ,o i,t+1 It is passed to the demand response aggregator.

[0125] Experience storage: The demand response aggregator connects C-LMP and local experience and stores it in the replay buffer. middle.

[0126] State and observation updates: Update global state s t →s t+1 and local observation o i,t →o i,t+1 .

[0127] Mini-batch sampling: from the playback buffer Sampling a small batch of data Where j is the sequence number and J is the total number of sampled data. j and λ j All data are from small batches of sampled data.

[0128] Critics Network Update: Update the online critics network.

[0129] Network weight update: Updates the online network weight.

[0130] End the cycle.

[0131] By employing reinforcement learning, this application achieves intelligent regulation and profit maximization of user-side response behavior. Specifically, the reinforcement learning algorithm enables users (or agents) to learn and optimize their electricity consumption strategies through dynamic interaction with the environment. After receiving real-time market signals such as the Coordinated Electricity-Carbon Marginal Price (C-LMP), users respond intelligently through the trained model, flexibly adjusting their electricity consumption patterns to adapt to changes in prices and carbon emissions. The reward function guides users to maximize profits when optimizing their behavior, such as increasing electricity consumption during low-price or low-carbon periods and reducing consumption during high-price periods, thereby effectively reducing costs. The adaptive nature of reinforcement learning allows users to adjust their decisions in real time even in the face of market fluctuations, maintaining long-term profit maximization. Simultaneously, Multiagent Reinforcement Learning (MARL) further supports collaborative responses from users within a distributed framework, improving the overall efficiency and stability of the system. Therefore, this solution method not only achieves intelligent user response but also promotes efficient and low-carbon operation of the distribution network.

[0132] In summary, this application aims to optimize the marginal electricity price strategy of the joint electricity-carbon node in an active distribution network. By constructing a distributed partially observable Markov decision process and applying the Multi-Agent Soft Actor-Critic (MASAC) algorithm, intelligent control and profit maximization of user-side response behavior are achieved. The implementation of the computational process enables the demand response aggregator to make precise real-time adjustments based on the joint electricity-carbon regional marginal price mechanism, promoting autonomous collaboration and flexible response on the user side. This mechanism effectively coordinates user-side load and distributed energy generation, ensuring the stability and low-carbon development of the distribution network, while providing economic incentives for users to actively participate. This invention demonstrates significant advantages in achieving fair sharing of environmental responsibility, reducing overall carbon emissions, and improving energy utilization efficiency, promoting the important role of the distribution network in the green energy transition. The mechanism and algorithm proposed in this application are expected to be applied to a wider range of power system scenarios, further promoting the maturity and popularization of demand response technology, and providing new solutions for realizing an intelligent and low-carbon energy system.

[0133] Based on the same inventive concept, this application also provides an electric-carbon combined demand response scheduling decision-making device for implementing the above-mentioned electric-carbon combined demand response scheduling decision-making method. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more electric-carbon combined demand response scheduling decision-making device embodiments provided below can be found in the limitations of the electric-carbon combined demand response scheduling decision-making method described above, and will not be repeated here.

[0134] In one exemplary embodiment, a demand response scheduling decision-making device based on electricity-carbon integration is provided, comprising: a data acquisition module, a model determination module, and a scheme determination module.

[0135] The data acquisition module is used to acquire user-side information data. This information data includes: the marginal electricity price signal of the electricity-carbon joint node corresponding to the operator on the distribution network side; the information data is determined based on the node marginal electricity price model of the electricity-carbon joint node on the distribution network side; wherein, the node marginal electricity price model is determined by the operator on the distribution network side using the distribution network branch flow algorithm; the node marginal electricity price model includes: objective function and constraints.

[0136] The model determination module is used to determine the user-side demand response model. The user-side demand response model is constructed based on the aforementioned information data, with the goal of minimizing the power factor.

[0137] The scheme determination module uses Markov decision theory to mathematically transform the user-side demand response model and employs deep reinforcement learning for iterative optimization to obtain a scheduling decision scheme. This scheduling decision scheme is used to coordinate user-side loads and distributed energy generation.

[0138] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram is shown in Figure 2. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores user-side demand response scheduling decision data. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a demand response scheduling decision method based on an electric-carbon joint approach.

[0139] Those skilled in the art will understand that the structure shown in Figure 2 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0140] In this application, all actions involving the acquisition of signals, information, or data are carried out in compliance with the relevant data protection laws and regulations of the country where the application is located, and with the authorization granted by the owner of the relevant device. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant regulations.

[0141] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0142] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0143] Based on the same inventive concept, this application also provides a distributed power grid framework system, including: a communication layer and a power layer connected together; the communication layer includes a demand response aggregator for connecting user terminals and the power grid; the power layer is used for distributing and managing the power flow of the power grid. This distributed power grid framework system can apply the aforementioned demand response scheduling decision-making method based on electricity-carbon joint principles.

[0144] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A demand response scheduling decision-making method based on electricity-carbon co-operation, characterized in that, The demand response scheduling decision-making method based on electricity-carbon joint is applied to the distributed framework system of the distribution network. The distributed power grid framework system includes: a communication layer and a power layer; the communication layer includes a demand response aggregator for connecting user terminals and the power grid; the power layer is used for allocating and managing the power flow of the power grid; the demand response scheduling decision-making method based on electricity-carbon joint includes: Acquire user-end information data; the information data includes: the marginal electricity price signal of the electricity carbon joint node corresponding to the operator at the distribution network end; the information data is determined based on the node marginal electricity price model of the electricity carbon joint on the distribution network side; wherein, the node marginal electricity price model is determined by the operator at the distribution network end using the distribution network branch flow algorithm; the node marginal electricity price model includes: objective function and constraints; Determine the user-side demand response model; the user-side demand response model is constructed based on the information data with the goal of minimizing the power coefficient; The user-side demand response model is mathematically transformed using Markov decision-making, and iteratively optimized using deep reinforcement learning to obtain a scheduling decision scheme. This scheduling decision scheme is used to coordinate user-side loads and distributed energy generation at the user end.

2. The demand response scheduling decision-making method based on electricity-carbon co-location as described in claim 1, characterized in that, The expression for the objective function is: Where F is the objective function; B dg B is a distributed generation collection. re B is a collection of new energy power generation facilities. ess For energy storage collection, B gd For the upper-level power grid; Cost of distributed generation; Let t be the carbon price in the carbon market. Let t be the carbon emission intensity of distributed generation; For distributed generation power; Cost of new energy power generation; For new energy power generation; Cost of energy storage degradation; Power generation for energy storage systems; The electricity price is determined by the higher-level power grid. Carbon emission intensity of the upper-level power grid; denoted as active power of the upper-level power grid; g represents the serial number of the distributed generation equipment; r represents the serial number of the new energy generation equipment; e represents the serial number of the energy storage equipment; and b represents the serial number of the node participating in the response.

3. The demand response scheduling decision-making method based on electricity-carbon co-operation according to claim 1, characterized in that, The constraints include: power balance constraints, reactive power balance constraints, voltage difference constraints, power flow limits for transmission lines, voltage magnitude constraints, line current limits, limits on distributed generation power, capacity limits for distributed generators, range limits on reactive power of distributed generators, range limits on active power of energy storage systems, power range limits for renewable energy, and power factor limits for renewable energy.

4. The demand response scheduling decision-making method based on electricity-carbon co-operation according to claim 1, characterized in that, The objective function corresponding to the user-side demand response model is: Where D is the objective function corresponding to the user-side demand response model; T is the total number of time steps; B is the total number of nodes participating in the response; t is the time step number; and b is the node number participating in the response. This is the marginal electricity price signal for the combined electricity and carbon node; P represents the active power load of the node after demand response. b ed.max P represents the maximum load of node b participating in the response. b,t Active load before demand response; ΔP b,max This represents the maximum response size.

5. The demand response scheduling decision-making method based on electricity-carbon co-location as described in claim 1, characterized in that, The user-side demand response model is mathematically transformed using Markov decision theory, and iteratively optimized using deep reinforcement learning to obtain a scheduling decision scheme, including: Based on the aforementioned user-side demand response model, Markov decision-making is employed at discrete time steps to transform the response into a decision, yielding the transformation result. This transformation result includes: an agent, a global state, local observations, a set of actions, a reward function, and a state transition function. The agent represents the user-side demand response participants at the user end. The global state is the user-side demand response state at the user end. The local observations include the active power load at nodes after demand response and the marginal electricity price signal at the combined electricity and carbon nodes. The set of actions is the set of actions performed by the power layer to allocate and manage the power flow of the distribution network. A deep reinforcement learning method is adopted, based on the entropy regularization strategy of the entropy term, with the goal of maximizing the cumulative expected reward of the agent, and the transformation result is iteratively optimized to obtain the scheduling decision scheme.

6. The demand response scheduling decision-making method based on electricity-carbon co-operation according to claim 5, characterized in that, The process of iteratively optimizing and solving using deep reinforcement learning methods includes: Acquire initial parameter data; the initial parameter data includes: weight data, local observations, action set, and reward function; the local observations, action set, and reward function are stored in the replay buffer; At each time step, each agent selects action data from the action set based on the current local observations and determines the current cumulative expected reward based on the reward function; Based on the current cumulative expected reward, the action data and the corresponding local observations are used as local experience and stored in the replay buffer. Batch data is sampled from the playback buffer, based on the formulas corresponding to the policy network loss function and the value network loss function: Update the weight data to obtain the updated weights; Wherein, J(π) i () represents the policy network loss function; To find the mathematical symbol for the expected value; π i (a i,t |o i,t ) represents the probability distribution of the policy network for agent i; o i,t Let be the observation value of the i-th agent at time step t; For playback buffer; Q i (o i,t ,a i,t ) represents the soft Q-value of the value network estimate, indicating the Q-value given a state and action; a i,t Let S be the action value of the i-th agent at time step t; α is the entropy coefficient; S t Let a be the state value at time step t; 1:I,t The action before the t-th time step; r i,t Let s be the reward value of the i-th agent at time step t; t+1 Q represents the state value at time step t+1. i (s t ,a 1:I,t ) represents the Q-value of the target Q-value network; y i,t The target value; L(Q) i ) represents the value network loss function; π i It is a probability distribution; Based on updated weights and local experience, the global state and local observations are updated according to all selected action data to obtain a scheduling decision scheme.

7. A demand response scheduling decision-making device based on electricity-carbon co-operation, characterized in that, The demand response scheduling decision-making device based on electricity-carbon co-operation includes: The data acquisition module is used to acquire user-end information data; the information data includes: the marginal electricity price signal of the electricity carbon joint node corresponding to the operator at the distribution network end; the information data is determined based on the node marginal electricity price model of the electricity carbon joint on the distribution network side; wherein, the node marginal electricity price model is determined by the operator at the distribution network end using the distribution network branch flow algorithm; the node marginal electricity price model includes: objective function and constraints; The model determination module is used to determine the user-side demand response model; the user-side demand response model is constructed based on the information data with the goal of minimizing the power coefficient. The scheme determination module is used to mathematically transform the user-side demand response model using Markov decision and to iteratively optimize and solve it using deep reinforcement learning methods to obtain a scheduling decision scheme; the scheduling decision scheme is used to coordinate the user-side load and distributed energy generation at the user end.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the demand response scheduling decision method based on the electric-carbon joint method as described in any one of claims 1-6.

9. A distributed grid system, characterized in that, The distributed power grid framework system includes: a communication layer and a power layer; the communication layer includes a demand response aggregator for connecting user terminals and the power grid; the power layer is used to allocate and manage the power flow of the power grid.