Active power distribution network interactive carbon reduction decision-making agent construction method based on carbon flow distribution
By constructing an interactive carbon reduction decision-making intelligent agent for active distribution networks, and utilizing carbon flow distribution matrices and reinforcement learning algorithms, precise carbon flow tracking and interactive carbon reduction of active distribution networks were achieved, improving the adaptability and robustness of low-carbon dispatching, and optimizing the consumption of new energy sources and network economics.
Patent Information
- Application Number
- CN202510852312.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies cannot accurately and in real-time dynamically track and interactively make carbon reduction decisions for carbon flows in active distribution networks coupled with massive operating modes, limited measurement, and green electricity trading. They cannot construct accurate power flow and lack a multi-stakeholder interaction mechanism based on dynamic carbon flows, resulting in poor low-carbon dispatching performance.
A carbon flow distribution-based interactive carbon reduction decision-making agent for active distribution networks is constructed. Through closed-loop interaction between the agent and the distribution network, a Markov decision process framework is designed. The agent is trained using a reinforcement learning algorithm driven by mechanism-data fusion. Combining the carbon flow distribution matrix and green electricity trading data, the renewable energy absorption rate and network line loss are optimized to achieve multi-timescale decision-making.
It significantly improves the adaptability and robustness of carbon reduction decisions, reduces the carbon emission intensity of the system, supports low-carbon scheduling at multiple time scales including day-ahead, intraday, and real-time, and optimizes the consumption of new energy sources and network economics.
Smart Images

Figure CN120974688A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-carbon dispatching technology for active distribution networks, and in particular to a method for constructing an interactive carbon reduction decision-making intelligent agent for active distribution networks based on carbon flow distribution. Background Technology
[0002] Building a new power system with new energy sources as the mainstay is the only way to support energy transformation and achieve the "dual carbon" strategic goal. Real-time, accurate, and comprehensive calculation of carbon emissions at each stage of the "source-grid-load-storage" system is the foundation and prerequisite for tapping the potential for carbon reduction in the power sector, guiding interactive carbon reduction among power users, and promoting the low-carbon transformation of the power economy. The development characteristics of the new power system are accelerated clean energy production, highly electrified energy consumption, increasingly platform-based energy allocation, and increasingly efficient energy utilization. The profound adjustment of the energy landscape will bring profound changes to the power system. The power supply structure will shift from being dominated by controllable, continuous coal-fired power generation to being dominated by new energy power generation with strong uncertainty and weak controllable output. Load characteristics will shift from traditional rigid, purely consumption-oriented to flexible, combining production and consumption. The grid structure will shift from a traditional grid dominated by unidirectional, tiered transmission to an energy internet that includes large AC / DC hybrid grids, microgrids, local DC grids, and adjustable loads.
[0003] Accurate and comprehensive accounting of carbon emissions related to the power system is a prerequisite for guiding power emission reduction. With the exponential growth in the scale of power system analysis and computation, represented by carbon analysis and optimization decision-making, existing calculation methods face challenges such as difficulty in integrating massive amounts of data and slow speed of extrapolation, analysis, and optimization decision-making. There is an urgent need for innovative calculation models and methods. Leveraging the powerful computational extrapolation and optimization reasoning capabilities driven by artificial intelligence, and integrating model mechanisms with data-driven approaches, research is being conducted on methods such as intelligent tracking of carbon flows under limited measurement, generation of green electricity trading auxiliary strategies based on carbon flow correction, and interactive carbon reduction decision-making by multiple stakeholders. This research aims to overcome the spatial and temporal fine-grainedness, accuracy, and performance bottlenecks in the coupled calculation of carbon emissions in active distribution networks, facilitating accurate carbon assessment across the entire network and at the user side, guiding users to actively participate in green electricity trading for carbon reduction, optimizing source-grid-load-storage resources for interactive carbon reduction, and enhancing the observability, descriptibility, and controllability of carbon flows.
[0004] Current technologies cannot accurately and in real-time dynamically track and interactively reduce carbon flows in active distribution networks under massive operating modes, limited measurements, and green electricity trading coupling, posing significant application challenges. Current mainstream grid carbon flow calculations still rely on precise grid topology information. When there are limited measurements or insufficient accuracy in the distribution network, it is impossible to construct accurate power flow for carbon flow calculation, thus failing to further guide carbon reduction on the distribution network-user side. Furthermore, it does not consider the spatiotemporal interaction and transfer of carbon flows in green electricity trading. Low-carbon dispatch focuses on directly absorbing renewable energy, without forming a multi-agent interaction mechanism based on dynamic carbon flows. In addition, pure model-driven algorithms rely on precise information and grid boundary conditions, while pure data-driven methods risk poor generalization ability. A breakthrough in mechanism-data fusion-driven methods is needed for carbon flow tracking and interactive carbon reduction in active distribution networks. Therefore, it is necessary to research mechanism-data fusion-driven methods for carbon flow tracking and interactive carbon reduction in active distribution networks based on electricity-carbon coupling calculation technology, and to construct an intelligent agent method for interactive carbon reduction decision-making in active distribution networks based on carbon flow distribution, thereby improving the low-carbon operation level of active distribution networks. Summary of the Invention
[0005] The technical problem to be solved by this invention is to construct a low-carbon scheduling method based on carbon flow distribution, which is an active distribution network interactive carbon reduction decision-making intelligent agent. Through the closed-loop interaction between the intelligent agent and the distribution network, it supports multi-time scale decision-making, significantly improves the adaptability and robustness of carbon reduction decision-making, effectively reduces the carbon emission intensity of the system, and provides technical support for building a new type of low-carbon and highly resilient power system.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0007] A method for constructing an interactive carbon reduction decision-making agent for active power distribution networks based on carbon flow distribution, proposed according to the present invention, includes:
[0008] A dynamic tracking model for carbon flow distribution in an active distribution network is constructed. Based on carbon flow theory, a carbon flow distribution matrix is established for multiple links including source, grid, load, and storage. The real-time carbon flow intensity of each node in the distribution network is calculated through a node-level carbon emission intensity quantification model.
[0009] Design a Markov decision process framework for an interactive carbon reduction decision-making agent, defining a state space, action space, and reward function; the state space includes real-time carbon flow data, predicted values of renewable energy output, load demand fluctuations, and dynamic changes in network topology; the action space includes controllable unit output adjustment, energy storage charging and discharging strategies, demand-side response commands, and green electricity trading auxiliary strategies.
[0010] The agent is trained using a reinforcement learning algorithm driven by mechanism data fusion. It integrates physical model constraints and historical operation data, and optimizes the agent's strategy through offline pre-training and online dynamic fine-tuning to ensure the feasibility of the action space.
[0011] Based on dynamic carbon flow distribution, multi-objective collaborative optimization decision-making is generated, with the core objective of minimizing carbon emissions. Simultaneously, the renewable energy absorption rate, network line loss and operational economy are optimized, and real-time scheduling instructions are output.
[0012] Through closed-loop interaction between the intelligent agent and the power distribution network control system, iterative updates of carbon flow tracking, decision optimization, and dynamic execution are achieved, improving the adaptability and robustness of carbon reduction decisions and supporting multi-timescale decision-making at the day-ahead, intraday, and real-time levels.
[0013] As a method for constructing an interactive carbon reduction decision-making agent for active power distribution networks based on carbon flow distribution according to the present invention, the method for constructing the carbon flow distribution matrix is as follows:
[0014] Based on power system flow calculation and carbon emission responsibility allocation model, carbon emissions from the generation side are allocated to the load side according to the nodal power contribution. The calculation formula is as follows:
[0015]
[0016] Among them, C j The carbon flux intensity at node j, in kgCO2, P ij Power components from node i to node j, in kW and C. i Carbon emissions at node i, in kgCO2, L kj This represents the power loss in transmission line k caused by electricity consumption at node j, in kW and L. k This represents the total power loss of transmission line k, in kW. This represents the carbon emissions from line k, expressed in kgCO2. N and M represent the total number of nodes and lines, respectively.
[0017] Dynamically adjust the carbon emission factor C of distributed generation (DG). DG,t The calculation formula is updated based on green electricity trading data, as follows:
[0018] C DG,t =η green P DG,t I base (2)
[0019] Among them, I base Regional baseline carbon emission factor, unit: kgCO2 / kWh, η green The green electricity trading correction factor is calculated using the following formula:
[0020]
[0021] Where T is the scheduling period, P DG,t For the real-time output of the DG unit, P green,t Power for green electricity trading.
[0022] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the Markov decision process of the interactive carbon reduction decision-making agent includes:
[0023] (1) Decision state space: refers to the physical quantities that the agent can observe in the environment, mathematically represented as:
[0024] S={P L ,V,I,S oc ,t} (4)
[0025] Among them, P L V, I, S oc t represents the load power, node voltage, branch current, and energy storage charge / discharge state S, respectively. OC Time, where S OC It is a relative measure of the energy stored in the battery, used to indicate the expected available power before the system is recharged, thereby ensuring that the battery is within a safe operating range;
[0026] (2) Decision action space: refers to the variables controlled by the agent in the system, mathematically represented as:
[0027] A = {P} G ,P es} (5)
[0028] Among them, P G P es These represent the output of a conventional thermal power unit and the power of its energy storage charging and discharging, respectively.
[0029] (3) Decision reward function: The low-carbon economic dispatch objective of an active power distribution network system integrating wind, solar, and energy storage is to minimize the total carbon emissions of the system. This includes the proposed network losses and wind and solar curtailment penalties. The problem of minimizing the total carbon emissions of the system is transformed into the form of maximizing the cumulative reward through deep reinforcement learning. At the same time, the power balance constraints of the system are executed. Then, the reward function r in the process of agent-environment interaction is defined. t for:
[0030] r t (s t ,a t )=-σ1F-σ2ΔP t (6)
[0031] Where F represents the state s of the agent at time t. t Next, execute action a t The reward obtained later is expressed as equation (7); ΔP tLet P be the power imbalance at time t, which aims to constrain the safe and stable operation of the entire system, and takes into account the power P generated in the electricity carbon market trading system. market,t The expression is given by equation (8); σ1 and σ2 are the reward control coefficients;
[0032]
[0033] Where F is composed of total network loss Loss and line load factor L i The penalty function F for wind and solar power curtailment punish Total carbon intensity C Σ Composition, where a, b, c, and d are harmonic weighting coefficients, P ess,i,t P load,i,t N represents the energy storage output and load size, respectively. ess This indicates the number of energy storage devices introduced into the network.
[0034] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the agent is trained using a reinforcement learning algorithm driven by mechanism-data fusion, including the use of an improved Deep Deterministic Policy Gradient (DDPG) framework, with the objective function being:
[0035]
[0036] Where s and a represent the current state and action, respectively; Indicates the experience replay cache ρ π The expectation of the randomly sampled transition sample s is used to minimize the global error between the Critic prediction and the target value; Q(s,a|θ) Q θ represents the predicted value from the Critic network, y represents the target value, and θ represents the target value. Q Let F(a) be the target network parameters under the network prediction strategy Q, and F(a) be the constraint function of the physical model. phys Let λ represent the boundary of possible actions, and λ be the penalty coefficient.
[0037] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the reward function r of the Markov decision process, which generates multi-objective collaborative optimization decisions based on dynamic carbon flow distribution, is... t The constraints include power balance, upper and lower limits of unit output, and energy storage charge and discharge rates.
[0038] The construction of expression (7) in the reward function specifically includes:
[0039] (1) Total network loss: can be broken down into the sum of total transformer loss and line loss over a period of time:
[0040]
[0041] Among them, P Tk P represents the losses of transformer k (out of K). Ll and Q Ll These represent the active power and reactive power at the end of branch l, respectively; R l and X l These are the resistance and reactance of branch l, respectively; U Ll T is the voltage at the end of branch l; l,i The corresponding elements of the power transmission allocation coefficient matrix T represent the power flow changes on branch l caused by the injection of unit power into node i:
[0042]
[0043] Where B' is the nodal admittance matrix of the system after removing the balancing node; X B The elements satisfy:
[0044]
[0045] Among them, B l For the admittance of branch l, c l,i This indicates the direction of the power flow on the l-th branch connected to node i. If the power flows out of node i, it is 1; if it flows in, it is -1; otherwise, it is 0.
[0046] (2) Line load factor L i The percentage of the average active load on the line over a certain period of time relative to the highest active load flowing through the line. A higher value indicates more efficient power utilization and avoidance of overload.
[0047]
[0048] (3) Curtailment penalty function F punish :
[0049]
[0050] Among them, P W,i,t P represents the output of a clean energy unit (wind or solar) at node i per unit time t. S,i,t Represents the load demand at node i per unit time t, expressed by the wind and solar curtailment penalty function F. punish As the total output of thermal power units in the distribution network system, the portion of the actual output of wind power or photovoltaic units that does not meet the load requirements needs to be compensated by the output of thermal power units.
[0051] (4) Total carbon intensity C Σ :
[0052]
[0053] Among them, C j The carbon flux intensity at node j is expressed in kgCO2 and C. DG,t This represents the carbon emission factor of distributed generation (DG).
[0054] As a method for constructing an interactive carbon reduction decision-making agent for active power distribution networks based on carbon flow distribution as described in this invention, the constraints specifically include the output constraints of each generating unit and energy storage:
[0055] (1) Constraints on wind power generation:
[0056] The simulation is performed using the Weibull distribution function, and the function model is expressed as follows:
[0057]
[0058] Where v is the actual measured wind speed; c is the scale parameter; k is the shape parameter; c and k are two important variables of the Weibull distribution, and their functional relationship with the wind speed mean μ and variance σ is expressed as follows:
[0059]
[0060] The functional relationship between the gamma function Γ, the mean μ, and the variance σ function is expressed as follows:
[0061]
[0062] The wind speed conditions for wind power system operation are described as follows:
[0063]
[0064] Wherein, the value of the wind turbine coefficient 'a' is P. r / (v r -v ci The value of the fan coefficient b is -av. ci ;v ci v c0 v r These represent the unit's cut-in wind speed, cut-out wind speed, and rated wind speed, respectively.
[0065] (2) Constraints on photovoltaic power generation:
[0066] Using a beta distribution:
[0067]
[0068] Where Γ is the gamma function; r is the actual irradiance; r max The maximum irradiance; the equation relating the functions α and β is expressed as:
[0069]
[0070] Where μ is the average irradiance during the test period; σ is the variance of irradiance during the test period. Considering the photovoltaic power supply installation area, the power output model of the photovoltaic power supply is expressed as:
[0071]
[0072] Among them, P M P is the output power of the photovoltaic power source, A is the installation area of the photovoltaic module, and η is the photoelectric conversion efficiency; where P Mmax This refers to the maximum output power of the photovoltaic cell; the output power model of the photovoltaic power source is expressed as:
[0073]
[0074] (3) Conventional unit constraints:
[0075] Conventional generating units are subject to power output constraints during operation, which can be described as follows:
[0076] μ MT P MT,min ≤μ MT P MT,t ≤μ MT P MT,max (twenty four)
[0077] Where, μ MT For the running status of MT, μ MT =1 indicates startup, μ MT =0 indicates shutdown; P MT,max and P MT,min These represent the upper and lower limits of the gas turbine output;
[0078] Meanwhile, conventional units are subject to ramp-up constraints, which can be described as follows:
[0079] P MT,down ≤P MT,t+1 -P MT,t ≤P MT,up (25)
[0080] Among them, P MT,down P is the downhill climbing speed of the gas turbine. MT,up This refers to the climbing speed of the gas turbine.
[0081] Furthermore, when a conventional unit needs to adjust its output due to fluctuations in wind and solar power output during a certain period, the difference between its final output during that period and the original output setting value satisfies the following:
[0082] -k MT P MT,max ≤ΔP MT,t ≤k MT PMT,max (26)
[0083] Wherein, ΔP MT,t This refers to the adjustment amount when regulating the output of a conventional unit; k MT To adjust the limits of power;
[0084] (4) Energy storage constraints:
[0085] The electrical energy stored in an energy storage device is expressed as:
[0086]
[0087] Where S is the electrical energy in the energy storage device during the current time period; S0 is the electrical energy in the energy storage device during the previous time period; P SC P SD These represent the charging and discharging power of the energy storage device; η SC η SD These represent the charge and discharge efficiencies of the energy storage device; τ SC τ SD These are the charging and discharging times of the energy storage device;
[0088] To ensure the safe use of energy storage devices, energy storage device S OC The following constraints apply:
[0089] S min ≤S oc ≤S max (28)
[0090] Among them, S max and S min These represent the upper and lower limits of electrical energy in the energy storage device;
[0091] Meanwhile, during the use of energy storage devices, the following constraints apply to their charging and discharging rates:
[0092]
[0093] Among them, P SC,max P SD,max These are the maximum charging and discharging power of the energy storage device;
[0094] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the constraints specifically include power flow equation constraints:
[0095] The standard power flow calculation model based on Distflow is adopted, which means that for any node in the entire system, the following conditions are met under the same time metric:
[0096]
[0097] Among them, P ij Q ij r is the power flowing through the branch between the two nodes. ij x ij Let be the impedance of the branch between the two nodes. This indicates the generator output at node j. This indicates the load consumption at node j. and This represents the total power flowing from node j to downstream node k, which is the sum of all downstream branches. and This represents the total power injected into node j from upstream node i, which is the sum of all upstream branches;
[0098] Introducing variable α i β ij Let the squares of the voltage at node i and the squares of the current in branch ij be represented respectively.
[0099]
[0100] Then some of the constraints can be transformed into:
[0101]
[0102] Because of the presence of nonlinear terms, the last equality constraint needs to be modified, and the second-order cone equality constraint is relaxed into an inequality constraint:
[0103]
[0104] in, These are used to limit the voltage at node i and the current in branch ij, respectively.
[0105] As described in this invention, the method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution achieves iterative updates of carbon flow tracking, decision optimization, and dynamic execution through closed-loop interaction between the agent and the distribution network control system. Specifically, this includes using a CT network to learn from the real physical environment. The CT network model specifically includes:
[0106] (1) The state transition model function is used to predict the probability distribution of the next state after taking a certain action in the current state:
[0107] s′ t+1 =P(s) t ,a t (34)
[0108] (2) The reward module function is used to predict the immediate reward obtained after taking an action in the current state:
[0109] rt ′=R(s t ,a t (35)
[0110] Among them, s t a t These represent the state and action of the agent at time t, respectively.
[0111] The closed-loop interaction between the intelligent agent and the power distribution network control system should include:
[0112] (1) Constructing a power system perception-type CT network model: Using deep learning technology, a CT network model that can accurately represent the dynamic characteristics of a power system including wind power generation and photovoltaic power generation is constructed.
[0113] (2) Efficient Environmental Simulation Based on CT Model: The constructed CT network model is used as the interactive environment for intelligent agents. By interacting with this model, intelligent agents can efficiently and cost-effectively simulate and access massive sample data of the power system under various operating states and wind and solar power output scenarios;
[0114] (3) Dynamic programming exploration and strategy evaluation of the agent: In the environment based on the CT model, the agent explores dynamic programming strategies. The agent tries different scheduling decisions. For each strategy tried, the agent predicts and evaluates the expected cumulative reward that can be obtained under the guidance of the strategy in a series of future states based on the interaction results with the CT model.
[0115] (4) Optimal strategy search objective setting: The optimization objective of the agent is set as follows: Under the premise of meeting the power system safety operation constraints, the optimal scheduling strategy is found through exploration and evaluation, so as to maximize the expected cumulative reward in the entire scheduling cycle or long-term operation.
[0116] (5) Reinforcement learning training agent: The reinforcement learning algorithm is used to train the agent using the massive sample data generated efficiently by the CT model. During the training process, the agent updates its policy parameters according to the reward and state transition information obtained by the current policy, and finally converges to the optimal scheduling policy.
[0117] (6) Apply the optimal scheduling strategy: Apply the obtained optimal scheduling strategy to the actual scheduling decision of the power system with wind and solar power to achieve the optimal system economy, security and new energy absorption capacity.
[0118] The CT network specifically enables sequence data prediction in the following ways:
[0119] (1) The CT network will input the current state s t and action a tIt is transformed into a processable feature vector, and positional encoding is added to the feature vector to ensure that the CT network can capture the sequential relationship of the feature vector;
[0120] (2) The self-attention mechanism in the network effectively captures long-distance dependencies in sequence data and extracts different types of information through multi-head attention. After passing through the self-attention network, an attention score matrix is obtained, which can represent the relationship between the inputs at each time step. Furthermore, temporal masking ensures that the inputs at each time step are only correlated with the inputs before that time step, thus reflecting causality.
[0121] (3) The feedforward neural network composed of fully connected layers and ReLU layers performs nonlinear mapping on the output of the multi-head attention network to enhance the model's representation ability;
[0122] (4) After layer normalization, linear transformation and multi-layer stacking are performed, and the final output is the predicted state change s′ of the environment. t+1 And reward r t ′.
[0123] The specific mathematical model of the CT module includes:
[0124] (1) Input part: state s t and action a t The input feature sequence data will be mapped to a dimension d. pos The feature vectors are used, and the position of each feature vector is determined by position encoding:
[0125]
[0126] Where p is the position of the feature vector, i∈{1,…,d} pos / 2};
[0127] (2) Self-attention mechanism: Extract relevant feature information from the feature vector, and multiply the input feature vector X by the weight matrix W respectively. q W k W v We obtain the query matrix q, the key matrix k, and the value matrix v:
[0128]
[0129] The relevance of data is characterized by calculating the score of each query in the query matrix q and all keys in the key matrix k through dot product attention, and obtaining the corresponding value weight coefficient for each key, thereby determining the weighted combination of values under a given query.
[0130] The self-attention score sequence e is obtained by weighted summation of the value matrix using the softmax function:
[0131]
[0132] Where att represents the self-attention mechanism, d is the dimension of the sequence; the softmax function transforms each element into a probability value in the interval (0,1) through a normalization exponent, and the sum of the probabilities of all elements is 1, which can increase the difference between values and make them easier to distinguish:
[0133]
[0134] To extract richer features, multi-head self-attention is used on top of self-attention to improve computational efficiency:
[0135]
[0136] in, This means concatenating vectors along the same dimension; the input vector first passes through a multi-head attention network to obtain the feature vector E;
[0137] (3) Feedforward network: The ReLU layer is placed as an activation function after the hidden layer of the neural network to introduce nonlinear characteristics:
[0138] f(x) = max(0,x) (41)
[0139] Fully connected layers connect all neurons in the input or hidden layers to all neurons in the next layer, performing matrix multiplication and bias addition to map input features to the output for high-level feature extraction, classification, and regression. The feature vector E is then fed into a feedforward neural network layer composed of ReLU layers and fully connected layers to obtain:
[0140] FFN(E)=max(0,EW1+b1)W2+b2 (42)
[0141] Where FFN represents the feedforward neural network layer; W1, W2, b1, and b2 are the weights and biases of the network, respectively; layer normalization is added to the multi-head attention network and the feedforward neural network to ensure the stability of the data distribution. After layer normalization, it is as follows:
[0142] layout=layernorm(x+layer(x)) (43)
[0143] Where x is the input of the layer normalization, layer is the layer normalization neural network; layernorm represents the layer normalization, layout is the layer normalization output, the output value is linearly transformed and then passed through the softmax layer to obtain the probability distribution, and the value with the highest probability is selected as the prediction output.
[0144] (4) Output section: The CT module network with multi-head attention predicts the sequence data as follows:
[0145] Out=(s′ t+1 ,r t ′)=CT(s t ,a t (44)
[0146] Among them, s t a t ,s′ t+1 r t ′ represents the state sequence, action sequence, and reward sequence starting from time t.
[0147] Meanwhile, this invention also proposes a device for constructing an active power distribution network interactive carbon reduction decision-making intelligent agent based on carbon flow distribution, specifically including:
[0148] The carbon flow distribution dynamic tracking unit is used to construct a dynamic tracking model of carbon flow distribution in an active distribution network. Based on carbon flow theory, it establishes a carbon flow distribution matrix for multiple links from source to grid to load to storage, and calculates the real-time carbon flow intensity of each node in the distribution network through a node-level carbon emission intensity quantification model.
[0149] The intelligent agent decision-making construction unit is used to design the Markov decision process framework of the interactive carbon reduction decision-making intelligent agent, and defines the state space, action space and reward function. The state space includes real-time carbon flow data, new energy output forecast values, load demand fluctuations and network topology dynamic change information. The action space includes controllable unit output adjustment, energy storage charging and discharging strategies, demand-side response commands and green electricity trading auxiliary strategies.
[0150] The agent training unit is used to train agents using reinforcement learning algorithms driven by mechanism-data fusion. It integrates physical model constraints and historical operation data, and optimizes agent strategies through offline pre-training and online dynamic fine-tuning to ensure action space feasibility.
[0151] The decision output unit is used to generate multi-objective collaborative optimization decisions based on dynamic carbon flow distribution. With the minimization of carbon emissions as the core objective, it simultaneously optimizes the renewable energy absorption rate, network line loss, and operational economy, and outputs real-time scheduling instructions through the intelligent agent.
[0152] Furthermore, the present invention proposes an electronic system comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, characterized in that the instructions are executed by the at least one processor to enable the at least one processor to perform the method of signing the present invention.
[0153] Finally, the present invention also provides a computer-readable storage medium storing computer instructions for causing the computer to perform the methods described in the present invention.
[0154] The present invention has the following technical effects:
[0155] This invention constructs a dynamic carbon flow distribution tracking model to quantify node-level carbon emission intensity. It also designs an agent framework based on Markov decision processes, defining a state space including real-time carbon flow, renewable energy output, and load fluctuations, as well as an action space encompassing unit regulation and energy storage strategies. A reward function is used to optimize the carbon emission minimization objective. The agent is then trained using a reinforcement learning algorithm that integrates physical constraints and historical data, combined with offline pre-training and online fine-tuning strategies. Finally, a multi-objective collaborative optimization decision is generated, simultaneously optimizing renewable energy consumption, line losses, and economic efficiency. Through closed-loop interaction between the agent and the distribution network, iterative updates of carbon flow tracking, decision optimization, and dynamic execution are achieved, supporting multi-timescale decision-making. This invention utilizes causal Transformer networks to improve interaction efficiency and combines carbon flow responsibility sharing and green electricity trading correction mechanisms to significantly enhance the adaptability and robustness of carbon reduction decisions, effectively reducing system carbon emission intensity and providing technical support for building a low-carbon, highly resilient new power system. Attached Figure Description
[0156] Figure 1 This is a flowchart of the method of the present invention;
[0157] Figure 2 This is a schematic diagram of the carbon emission responsibility sharing model of the present invention;
[0158] Figure 3 This is a schematic diagram of the basic structure of the CT module of the present invention;
[0159] Figure 4 This is a schematic diagram of the self-attention network mechanism structure of the present invention;
[0160] Figure 5 This is a schematic diagram of the multi-head attention network mechanism structure of the present invention. Detailed Implementation
[0161] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0162] Example 1: This invention proposes a method for constructing an interactive carbon reduction decision-making agent for active power distribution networks based on carbon flow distribution, such as... Figure 1 As shown, it includes the following steps:
[0163] 1) Construct a dynamic tracking model for carbon flow distribution in active distribution networks, establish a carbon flow distribution matrix for multiple links of source-grid-load-storage based on carbon flow theory, and calculate the real-time carbon flow intensity of each node in the distribution network through a node-level carbon emission intensity quantification model.
[0164] 2) Design a Markov decision process framework for an interactive carbon reduction decision-making agent, defining the state space, action space, and reward function; the state space includes real-time carbon flow data, predicted values of renewable energy output, load demand fluctuations, and dynamic changes in network topology; the action space includes controllable unit output adjustment, energy storage charging and discharging strategies, demand-side response commands, and green electricity trading auxiliary strategies.
[0165] 3) The agent is trained using a reinforcement learning algorithm driven by mechanism and data fusion. It integrates physical model constraints and historical operation data, and optimizes the agent's strategy through offline pre-training and online dynamic fine-tuning to ensure the feasibility of the action space.
[0166] 4) Generate multi-objective collaborative optimization decisions based on dynamic carbon flow distribution, with the core objective of minimizing carbon emissions, while simultaneously optimizing the renewable energy absorption rate, network line loss and operational economy, and outputting real-time scheduling instructions;
[0167] 5) Through the closed-loop interaction between the intelligent agent and the power distribution network control system, the iterative update of carbon flow tracking, decision optimization and dynamic execution is realized, which improves the adaptability and robustness of carbon reduction decisions and supports multi-timescale decision-making at the day-ahead, intraday and real-time scales.
[0168] As a method for constructing an interactive carbon reduction decision-making agent for active power distribution networks based on carbon flow distribution according to the present invention, the method for constructing the carbon flow distribution matrix is as follows:
[0169] Based on power system flow calculation and carbon emission responsibility sharing model, such as Figure 2 As shown, carbon emissions from the power generation side are allocated to the load side based on the node power contribution, and the calculation formula is as follows:
[0170]
[0171] Among them, C j The carbon flux intensity at node j, in kgCO2 / kWh, P ij Power components from node i to node j, in kW and C. i Carbon emissions at node i, in kgCO2, L kj This represents the power loss in transmission line k caused by electricity consumption at node j, in kW and L. k This represents the total power loss of transmission line k, in kW. This represents the carbon emissions from line k, expressed in kgCO2. N and M represent the total number of nodes and lines, respectively.
[0172] Dynamically adjust the carbon emission factor C of distributed generation (DG). DG,t The calculation formula is updated based on green electricity trading data, as follows:
[0173] CDG,t =η green P DG,t I base (46)
[0174] Among them, I base Regional baseline carbon emission factor, unit: kgCO2 / kWh, η green The green electricity trading correction factor is calculated using the following formula:
[0175]
[0176] Where T is the scheduling period, P DG,t For the real-time output of the DG unit, P green,t Power for green electricity trading.
[0177] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the Markov decision process of the interactive carbon reduction decision-making agent includes:
[0178] (1) Decision state space: refers to the physical quantities that the agent can observe in the environment, mathematically represented as:
[0179] S={P L ,V,I,S oc ,t} (48)
[0180] Among them, P L V, I, S oc t represents the load power, node voltage, branch current, and energy storage charge / discharge state S, respectively. OC , time t, where S OC It is a relative measure of the energy stored in the battery, used to indicate the expected available power before the system is recharged, thereby ensuring that the battery is within a safe operating range;
[0181] (2) Decision action space: refers to the variables controlled by the agent in the system, mathematically represented as:
[0182] A = {P} G ,P es} (49)
[0183] Among them, P G P es These represent the output of a conventional thermal power unit and the power of its energy storage charging and discharging, respectively.
[0184] (3) Decision reward function: The low-carbon economic dispatch objective of an active power distribution network system integrating wind, solar, and energy storage is to minimize the total carbon emissions of the system. This includes the proposed network losses and wind and solar curtailment penalties. The problem of minimizing the total carbon emissions of the system is transformed into the form of maximizing the cumulative reward through deep reinforcement learning. At the same time, the power balance constraints of the system are executed. Then, the reward function r in the process of agent-environment interaction is defined. t for:
[0185] r t (s t ,a t )=-σ1F-σ2ΔP t (50)
[0186] Where F represents the state s of the agent at time t. t Next, execute action a t The reward obtained later is expressed as (7); ΔP t Let P be the power imbalance at time t, which aims to constrain the safe and stable operation of the entire system, and takes into account the power P generated in the electricity carbon market trading system. market,t The expression is (8); σ1 and σ2 are the reward control coefficients;
[0187]
[0188] Where F is composed of total network loss Loss and line load factor L i The penalty function F for wind and solar power curtailment punish Total carbon intensity C Σ The composition consists of harmonic weighting coefficients a, b, c, and d.
[0189] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the agent is trained using a reinforcement learning algorithm driven by mechanism-data fusion, including the use of an improved Deep Deterministic Policy Gradient (DDPG) framework, with the objective function being:
[0190]
[0191] Where F(a) is the constraint function of the physical model, a phys Let λ represent the boundary of possible actions, and λ be the penalty coefficient.
[0192] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the reward function r of the Markov decision process, which generates multi-objective collaborative optimization decisions based on dynamic carbon flow distribution, is... t The constraints include power balance, upper and lower limits of unit output, and energy storage charge and discharge rates.
[0193] The construction of expression (7) in the reward function specifically includes:
[0194] (1) Total network loss: can be broken down into the sum of total transformer loss and line loss over a period of time:
[0195]
[0196] Among them, P Ll and Q Ll These represent the active power and reactive power at the end of branch l, respectively; R l and X l These are the resistance and reactance of branch l, respectively; U Ll T is the voltage at the end of branch l; l,i The corresponding elements of the power transmission allocation coefficient matrix T represent the power flow changes on branch l caused by the injection of unit power into node i:
[0197]
[0198] Where B' is the nodal admittance matrix of the system after removing the balancing node; X B The elements satisfy:
[0199]
[0200] Among them, B l For the admittance of branch l, c l,i This indicates the direction of the power flow on the l-th branch connected to node i. If the power flows out of node i, it is 1; if it flows in, it is -1; otherwise, it is 0.
[0201] (2) Line load factor L i The percentage of the average active load on the line over a certain period of time relative to the highest active load flowing through the line. A higher value indicates more efficient power utilization and avoidance of overload.
[0202]
[0203] (3) Curtailment penalty function F punish :
[0204]
[0205] Among them, P W,i,t P represents the output of a clean energy unit (wind or solar) at node i per unit time t. S,i,t Represents the load demand at node i per unit time t, expressed by the wind and solar curtailment penalty function F. punish As the total output of thermal power units in the distribution network system, the portion of the actual output of wind power or photovoltaic units that does not meet the load requirements needs to be compensated by the output of thermal power units.
[0206] (4) Total carbon intensity C Σ :
[0207]
[0208] As a method for constructing an interactive carbon reduction decision-making agent for active power distribution networks based on carbon flow distribution as described in this invention, the constraints specifically include the output constraints of each generating unit and energy storage:
[0209] (1) Constraints on wind power generation:
[0210] The simulation is performed using the Weibull distribution function, and the function model is expressed as follows:
[0211]
[0212] Where v is the actual measured wind speed; c is the scale parameter; k is the shape parameter; c and k are two important variables of the Weibull distribution, and their functional relationship with the wind speed mean μ and variance σ is expressed as follows:
[0213]
[0214] The functional relationship between the gamma function Γ, the mean μ, and the variance σ function is expressed as follows:
[0215]
[0216] The wind speed conditions for wind power system operation are described as follows:
[0217]
[0218] Wherein, the value of the wind turbine coefficient 'a' is P. r / (v r -v ci The value of the fan coefficient b is -av. ci ;
[0219] (2) Constraints on photovoltaic power generation:
[0220] Using a beta distribution:
[0221]
[0222] Where Γ is the gamma function; r is the actual irradiance; r max The maximum irradiance; the equation relating the functions α and β is expressed as:
[0223]
[0224] Where μ is the average irradiance during the test period; σ is the variance of irradiance during the test period. Considering the photovoltaic power supply installation area, the power output model of the photovoltaic power supply is expressed as:
[0225]
[0226] Among them, P M P is the output power of the photovoltaic power source, A is the installation area of the photovoltaic module, and η is the photoelectric conversion efficiency; where P Mmax This refers to the maximum output power of the photovoltaic cell; the output power model of the photovoltaic power source is expressed as:
[0227]
[0228] (3) Conventional unit constraints:
[0229] Conventional generating units are subject to power output constraints during operation, which can be described as follows:
[0230] μ MT P MT,min ≤μ MT P MT,t ≤μ MT P MT,max (68)
[0231] Where, μ MT For the running status of MT, μ MT =1 indicates startup, μ MT =0 indicates shutdown; P MT,max and P MT,min These represent the upper and lower limits of the gas turbine output;
[0232] Meanwhile, conventional units are subject to ramp-up constraints, which can be described as follows:
[0233] P MT,down ≤P MT,t+1 -P MT,t ≤P MT,up (69)
[0234] Among them, P MT,down P is the downhill climbing speed of the gas turbine. MT,up This refers to the climbing speed of the gas turbine.
[0235] Furthermore, when a conventional unit needs to adjust its output due to fluctuations in wind and solar power output during a certain period, the difference between its final output during that period and the original output setting value satisfies the following:
[0236] -k MT P MT,max ≤ΔP MT,t ≤k MT P MT,max (70)
[0237] Wherein, ΔP MT,t This refers to the adjustment amount when regulating the output of a conventional unit; k MT To adjust the limits of power;
[0238] (4) Energy storage constraints:
[0239] The electrical energy stored in an energy storage device is expressed as:
[0240]
[0241] Where S is the electrical energy in the energy storage device during the current time period; S0 is the electrical energy in the energy storage device during the previous time period; P SC P SD These represent the charging and discharging power of the energy storage device; η SC η SD These represent the charge and discharge efficiencies of the energy storage device; τ SC τ SD These are the charging and discharging times of the energy storage device;
[0242] To ensure the safe use of energy storage devices, energy storage device S OC The following constraints apply:
[0243] S min ≤S oc ≤S max (72)
[0244] Among them, S max and S min These represent the upper and lower limits of electrical energy in the energy storage device;
[0245] Meanwhile, during the use of energy storage devices, the following constraints apply to their charging and discharging rates:
[0246]
[0247] Among them, P SC,max P SD,max These are the maximum charging and discharging power of the energy storage device;
[0248] As a method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution as described in this invention, the constraints specifically include power flow equation constraints:
[0249] The standard power flow calculation model based on Distflow is adopted, which means that for any node in the entire system, the following conditions are met under the same time metric:
[0250]
[0251] Among them, P ij Q ijr is the power flowing through the branch between the two nodes. ij x ij Let be the impedance of the branch between the two nodes. This indicates the generator output at node j. This indicates the load consumption at node j. and This represents the total power flowing from node j to downstream node k, which is the sum of all downstream branches. and This represents the total power injected into node j from upstream node i, which is the sum of all upstream branches;
[0252] Introducing variable α i β ij Let the squares of the voltage at node i and the squares of the current in branch ij be represented respectively.
[0253]
[0254] Then some of the constraints can be transformed into:
[0255]
[0256] Because of the presence of nonlinear terms, the last equality constraint needs to be modified, and the second-order cone equality constraint is relaxed into an inequality constraint:
[0257]
[0258] As described in this invention, the method for constructing an interactive carbon reduction decision-making agent for active distribution networks based on carbon flow distribution achieves iterative updates of carbon flow tracking, decision optimization, and dynamic execution through closed-loop interaction between the agent and the distribution network control system. Specifically, this includes using a CT network to learn from the real physical environment. The CT network model specifically includes:
[0259] (1) The state transition model function is used to predict the probability distribution of the next state after taking a certain action in the current state:
[0260] s′ t+1 =P(s) t ,a t (78)
[0261] (2) The reward module function is used to predict the immediate reward obtained after taking an action in the current state:
[0262] r t ′=R(s t ,a t (79)
[0263] The closed-loop interaction between the intelligent agent and the power distribution network control system should include:
[0264] (1) Constructing a power system perception-type CT network model: Using deep learning technology, a CT network model that can accurately represent the dynamic characteristics of a power system including wind power generation and photovoltaic power generation is constructed.
[0265] (2) Efficient Environmental Simulation Based on CT Model: The constructed CT network model is used as the interactive environment for intelligent agents. By interacting with this model, intelligent agents can efficiently and cost-effectively simulate and access massive sample data of the power system under various operating states and wind and solar power output scenarios;
[0266] (3) Dynamic programming exploration and strategy evaluation of the agent: In the environment based on the CT model, the agent explores dynamic programming strategies. The agent tries different scheduling decisions. For each strategy tried, the agent predicts and evaluates the expected cumulative reward that can be obtained under the guidance of the strategy in a series of future states based on the interaction results with the CT model.
[0267] (4) Optimal strategy search objective setting: The optimization objective of the agent is set as follows: Under the premise of meeting the power system safety operation constraints, the optimal scheduling strategy is found through exploration and evaluation, so as to maximize the expected cumulative reward in the entire scheduling cycle or long-term operation.
[0268] (5) Reinforcement learning training agent: The reinforcement learning algorithm is used to train the agent using the massive sample data generated efficiently by the CT model. During the training process, the agent updates its policy parameters according to the reward and state transition information obtained by the current policy, and finally converges to the optimal scheduling policy.
[0269] (6) Apply the optimal scheduling strategy: Apply the obtained optimal scheduling strategy to the actual scheduling decision of the power system with wind and solar power to achieve the optimal system economy, security and new energy absorption capacity.
[0270] like Figure 3 As shown, the CT network for sequence data prediction specifically includes:
[0271] (1) The CT network will input the current state s t and action a t It is transformed into a processable feature vector, and positional encoding is added to the feature vector to ensure that the CT network can capture the sequential relationship of the feature vector;
[0272] (2) The self-attention mechanism in the network effectively captures long-distance dependencies in sequence data and extracts different types of information through multi-head attention. After passing through the self-attention network, an attention score matrix is obtained, which can represent the relationship between the inputs at each time step. Furthermore, temporal masking ensures that the inputs at each time step are only correlated with the inputs before that time step, thus reflecting causality.
[0273] (3) The feedforward neural network composed of fully connected layers and ReLU layers performs nonlinear mapping on the output of the multi-head attention network to enhance the model's representation ability;
[0274] (4) After layer normalization, linear transformation and multi-layer stacking are performed, and the final output is the predicted state change s′ of the environment. t+1 And reward r t ′.
[0275] The specific mathematical model of the CT module includes:
[0276] (1) Input part: state s t and action a t The input feature sequence data will be mapped to a dimension d. pos The feature vectors are used, and the position of each feature vector is determined by position encoding:
[0277]
[0278] Where p is the position of the feature vector, i∈{1,…,d} pos / 2};
[0279] (2) Self-attention mechanism: Extract relevant feature information from the feature vector, and multiply the input feature vector X by the weight matrix W respectively. q W k W v We obtain the query matrix q, the key matrix k, and the value matrix v:
[0280]
[0281] like Figure 4 As shown, the relevance of data is characterized by calculating the score of each query in the query matrix q and all keys in the key matrix k through dot product attention, and the corresponding value weight coefficient of each key is obtained, thereby determining the weighted combination of values under a given query.
[0282] The self-attention score sequence e is obtained by weighted summation of the value matrix using the softmax function:
[0283]
[0284] Where att represents the self-attention mechanism, d is the dimension of the sequence; the softmax function transforms each element into a probability value in the interval (0,1) through a normalization exponent, and the sum of the probabilities of all elements is 1, which can increase the difference between values and make them easier to distinguish:
[0285]
[0286] like Figure 5 As shown, to extract richer features, multi-head self-attention is used on top of self-attention to improve computational efficiency:
[0287]
[0288] in, This means concatenating vectors along the same dimension; the input vector first passes through a multi-head attention network to obtain the feature vector E;
[0289] (3) Feedforward network: The ReLU layer is placed as an activation function after the hidden layer of the neural network to introduce nonlinear characteristics:
[0290] f(x) = max(0,x) (85)
[0291] Fully connected layers connect all neurons in the input or hidden layers to all neurons in the next layer, performing matrix multiplication and bias addition to map input features to the output for high-level feature extraction, classification, and regression. The feature vector E is then fed into a feedforward neural network layer composed of ReLU layers and fully connected layers to obtain:
[0292] FFN(E)=max(0,EW1+b1)W2+b2 (86)
[0293] Where FFN represents the feedforward neural network layer; W1, W2, b1, and b2 are the weights and biases of the network, respectively; layer normalization is added to the multi-head attention network and the feedforward neural network to ensure the stability of the data distribution. After layer normalization, it is as follows:
[0294] layout=layernorm(x+layer(x)) (87)
[0295] Where x is the input of the layer normalization, layer is the layer normalization neural network; layernorm represents the layer normalization, layout is the layer normalization output, the output value is linearly transformed and then passed through the softmax layer to obtain the probability distribution, and the value with the highest probability is selected as the prediction output.
[0296] (4) Output section: The CT module network with multi-head attention predicts the sequence data as follows:
[0297] Out=(s′t+1 ,r t ′)=CT(s t ,a t (88)
[0298] Among them, s t a t ,s′ t+1 r t ′ represents the state sequence, action sequence, and reward sequence starting from time t.
[0299] Example 2: This example proposes a device for constructing an interactive carbon reduction decision-making intelligent agent for active power distribution networks based on carbon flow distribution, specifically including:
[0300] The carbon flow distribution dynamic tracking unit is used to construct a dynamic tracking model of carbon flow distribution in an active distribution network. Based on carbon flow theory, it establishes a carbon flow distribution matrix for multiple links from source to grid to load to storage, and calculates the real-time carbon flow intensity of each node in the distribution network through a node-level carbon emission intensity quantification model.
[0301] The intelligent agent decision-making construction unit is used to design the Markov decision process framework of the interactive carbon reduction decision-making intelligent agent, and defines the state space, action space and reward function. The state space includes real-time carbon flow data, new energy output forecast values, load demand fluctuations and network topology dynamic change information. The action space includes controllable unit output adjustment, energy storage charging and discharging strategies, demand-side response commands and green electricity trading auxiliary strategies.
[0302] The agent training unit is used to train agents using reinforcement learning algorithms driven by mechanism-data fusion. It integrates physical model constraints and historical operation data, and optimizes agent strategies through offline pre-training and online dynamic fine-tuning to ensure action space feasibility.
[0303] The decision output unit is used to generate multi-objective collaborative optimization decisions based on dynamic carbon flow distribution. With the minimization of carbon emissions as the core objective, it simultaneously optimizes the renewable energy absorption rate, network line loss, and operational economy, and outputs real-time scheduling instructions through the intelligent agent.
[0304] Example 3: This example proposes a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the method described in this invention, which will not be repeated here.
[0305] Example 4: This example also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor.
[0306] It should be noted that the processing flows of Embodiments 2 to 4 correspond to the specific steps of the method provided in Embodiment 1 of the present invention, and have the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.
[0307] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0308] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0309] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0310] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for constructing an interactive carbon reduction decision-making agent for active power distribution networks based on carbon flow distribution, characterized in that, include: A dynamic tracking model for carbon flow distribution in an active power distribution network is constructed. A carbon flow distribution matrix is established based on carbon flow theory. The real-time carbon flow intensity of each node in the power distribution network is calculated through a node-level carbon emission intensity quantification model. Design a Markov decision process framework for an interactive carbon reduction decision-making agent, defining a state space, action space, and reward function; the state space includes real-time carbon flow data, predicted values of renewable energy output, load demand fluctuations, and dynamic changes in network topology; the action space includes controllable unit output adjustment, energy storage charging and discharging strategies, demand-side response commands, and green electricity trading auxiliary strategies. The agent is trained using a reinforcement learning algorithm driven by mechanism data fusion. It integrates physical model constraints and historical operation data, and optimizes the agent's strategy through offline pre-training and online dynamic fine-tuning to ensure the feasibility of the action space. Based on dynamic carbon flow distribution, multi-objective collaborative optimization decisions are generated, with the minimization of carbon emissions as the core objective. Simultaneously, the renewable energy absorption rate, network line loss, and operational economy are optimized, and real-time scheduling instructions are output through intelligent agents.
2. The method according to claim 1, characterized in that, The method for constructing the carbon flow distribution matrix is as follows: Based on power system flow calculation and carbon emission responsibility allocation model, carbon emissions from the generation side are allocated to the load side according to the nodal power contribution. The calculation formula is as follows: Among them, C j The carbon flux intensity at node j, in kgCO2, P ij Power components from node i to node j, in kW and C. i Carbon emissions at node i, in kgCO2, L kj This represents the power loss in transmission line k caused by electricity consumption at node j, in kW and L. k This represents the total power loss of transmission line k, in kW. This represents the carbon emissions from line k, expressed in kgCO2. N and M represent the total number of nodes and lines, respectively. Dynamically adjust the carbon emission factor C of distributed generation (DG). DG,t The calculation formula is updated based on green electricity trading data, as follows: C DG,t =the green P DG,t I base (2) Among them, I base Regional baseline carbon emission factor, unit: kgCO2 / kWh, η green The green electricity trading correction factor is calculated using the following formula: Where T is the scheduling period, P DG,t For the real-time output of the DG unit, P green,t Power for green electricity trading.
3. The method according to claim 1, characterized in that, The Markov decision process of the interactive carbon reduction decision-making agent includes: (1) Decision state space: refers to the physical quantities that the agent can observe in the environment, mathematically represented as: S={P L ,V,I,S oc ,t} (4) Among them, P L V, I, S oc t represents the load power, node voltage, branch current, and energy storage charge / discharge state S, respectively. OC Time, where S OC It is a relative measure of the energy stored in the battery, used to indicate the expected available power before the system is recharged, thereby ensuring that the battery is within a safe operating range; (2) Decision action space: refers to the variables controlled by the agent in the system, mathematically represented as: A={P G ,P es } (5) Among them, P G P es These represent the output of a conventional thermal power unit and the power of its energy storage charging and discharging, respectively. (3) Decision reward function: The low-carbon economic dispatch objective of an active power distribution network system integrating wind, solar, and energy storage is to minimize the total carbon emissions of the system. This includes the proposed network losses and wind and solar curtailment penalties. The problem of minimizing the total carbon emissions of the system is transformed into the form of maximizing the cumulative reward through deep reinforcement learning. At the same time, the power balance constraints of the system are executed. Then, the reward function r in the process of agent-environment interaction is defined. t for: r t (s t ,a t )=-σ1F-σ2ΔP t (6) Where F represents the state s of the agent at time t. t Next, execute action a t The reward obtained later is expressed as equation (7); ΔP t Let P be the power imbalance at time t, which aims to constrain the safe and stable operation of the entire system, and takes into account the power P generated in the electricity carbon market trading system. market,t The expression is given by equation (8); σ1 and σ2 are the reward control coefficients; Where F is composed of total network loss Loss and line load factor L i The penalty function F for wind and solar power curtailment punish Total carbon intensity C Σ Composition, where a, b, c, and d are harmonic weighting coefficients, P ess,i,t P load,i,t N represents the energy storage output and load size, respectively. ess This indicates the number of energy storage devices introduced into the network.
4. The method according to claim 1, characterized in that, The reinforcement learning algorithm for training the agent, driven by mechanism-data fusion, includes an improved Deep Deterministic Policy Gradient (DDPG) framework, with the objective function being: Where s and a represent the current state and action, respectively; Indicates the experience replay cache ρ π The expectation of the randomly sampled transition sample s is used to minimize the global error between the Critic prediction and the target value; Q(s,a|θ) Q θ represents the predicted value from the Critic network, y represents the target value, and θ represents the target value. Q Let Q be the target network parameters under the network prediction strategy Q; F(a) is the constraint function of the physical model, where a phys Let λ represent the boundary of possible actions, and λ be the penalty coefficient.
5. The method according to claim 1, characterized in that, The reward function r of the multi-objective collaborative optimization decision-making based on dynamic carbon flow distribution, i.e., the Markov decision process, is... t The constraints include power balance, upper and lower limits of unit output, and energy storage charge and discharge rates.
6. The method according to claim 3, characterized in that, The construction of expression (7) in the reward function specifically includes: (1) Total network loss: can be broken down into the sum of total transformer loss and line loss over a period of time: Where P Tk P represents the losses of transformer k (out of K). Ll and Q Ll These represent the active power and reactive power at the end of branch l, respectively; R Ll and X Ll These are the resistance and reactance of branch l, respectively; U Ll T is the voltage at the end of branch l; l,i The corresponding elements of the power transmission allocation coefficient matrix T represent the power flow changes on branch l caused by the injection of unit power into node i: Where B' is the nodal admittance matrix of the system after removing the balancing node; X B The elements satisfy: Among them, B l For the admittance of branch l, c l,i This indicates the direction of the power flow on the l-th branch connected to node i. If the power flows out of node i, it is 1; if it flows in, it is -1; otherwise, it is 0. (2) Line load factor L i The percentage of the average active load on the line over a certain period of time relative to the highest active load flowing through the line. A higher value indicates more efficient power utilization and avoidance of overload. (3) Curtailment penalty function F punish : Among them, P W,i,t P represents the output of the clean energy unit at node i per unit time t. S,i,t Represents the load demand at node i per unit time t, expressed by the wind and solar curtailment penalty function F. punish As the total output of thermal power units in the distribution network system, the portion of the actual output of wind power or photovoltaic units that does not meet the load requirements needs to be compensated by the output of thermal power units. (4) Total carbon intensity C Σ : Among them, C j The carbon flux intensity at node j is expressed in kgCO2 and C. DG,t This represents the carbon emission factor of distributed generation (DG).
7. The method according to claim 5, characterized in that, The constraints specifically include the output constraints of each generating unit and energy storage: (1) Constraints on wind power generation: The simulation is performed using the Weibull distribution function, and the function model is expressed as follows: Where v is the actual measured wind speed; c is the scale parameter; k is the shape parameter; c and k are two important variables of the Weibull distribution, and their functional relationship with the wind speed mean μ and variance σ is expressed as follows: The functional relationship between the gamma function Γ, the mean μ, and the variance σ function is expressed as follows: The wind speed conditions for wind power system operation are described as follows: Wherein, the value of the wind turbine coefficient 'a' is P. r / (v r -v ci The value of the fan coefficient b is -av. ci ;v ci v c0 v r These represent the unit's cut-in wind speed, cut-out wind speed, and rated wind speed, respectively. (2) Constraints on photovoltaic power generation: Using a beta distribution: Where Γ is the gamma function; r is the actual irradiance; r max The maximum irradiance; the equation relating the functions α and β is expressed as: Where μ is the average irradiance during the test period; σ is the variance of irradiance during the test period. Considering the photovoltaic power supply installation area, the power output model of the photovoltaic power supply is expressed as: Among them, P M P is the output power of the photovoltaic power source, A is the installation area of the photovoltaic module, and η is the photoelectric conversion efficiency; where P Mmax This refers to the maximum output power of the photovoltaic cell; the output power model of the photovoltaic power source is expressed as: (3) Conventional unit constraints: Conventional generating units are subject to power output constraints during operation, which can be described as follows: μ MT P MT,min ≤μ MT P MT,t ≤μ MT P MT,max (24) Where, μ MT For the running status of MT, μ MT =1 indicates startup, μ MT =0 indicates shutdown; P MT,max and P MT,min These represent the upper and lower limits of the gas turbine output; Meanwhile, conventional units are subject to ramp-up constraints, which can be described as follows: P MT,down ≤P MT,t+1 -P MT,t ≤P MT,up (25) Among them, P MT,down P is the downhill climbing speed of the gas turbine. MT,up This refers to the climbing speed of the gas turbine. Furthermore, when a conventional unit needs to adjust its output due to fluctuations in wind and solar power output during a certain period, the difference between its final output during that period and the original output setting value satisfies the following: -k MT P MT,max ≤ΔP MT,t ≤k MT P MT,max (26) Where, ΔP MT,t This refers to the adjustment amount when regulating the output of a conventional unit; k MT To adjust the limits of power; (4) Energy storage constraints: The electrical energy stored in an energy storage device is expressed as: Where S is the electrical energy in the energy storage device during the current time period; S0 is the electrical energy in the energy storage device during the previous time period; P SC P SD These represent the charging and discharging power of the energy storage device; η SC η SD These represent the charge and discharge efficiencies of the energy storage device; τ SC τ SD These are the charging and discharging times of the energy storage device; To ensure the safe use of energy storage devices, energy storage device S OC The following constraints apply: S min ≤S oc ≤S max (28) Among them, S max and S min These represent the upper and lower limits of electrical energy in the energy storage device; Meanwhile, during the use of energy storage devices, the following constraints apply to their charging and discharging rates: Among them, P SC,max P SD,max These represent the maximum charging and discharging power of the energy storage device.
8. The method according to claim 5, characterized in that, The constraints specifically include power flow equation constraints: The standard power flow calculation model based on Distflow is adopted, which means that for any node in the entire system, the following conditions are met under the same time metric: Among them, P ij Q ij r is the power flowing through the branch between the two nodes. ij x ij The impedance of the branch between the two nodes is... This indicates the generator output at node j. This indicates the load consumption at node j. and This represents the total power flowing from node j to downstream node k, which is the sum of all downstream branches. and This represents the total power injected into node j from upstream node i, which is the sum of all upstream branches; Introducing variable α i β ij Let the squares of the voltage at node i and the squares of the current in branch ij be represented respectively. Then some of the constraints are transformed into: Because of the presence of nonlinear terms, the last equality constraint needs to be modified, and the second-order cone equality constraint is relaxed into an inequality constraint: in, These are used to limit the voltage at node i and the current in branch ij, respectively.
9. The method according to claim 1, characterized in that, This includes iterative updates of carbon flow tracking, decision optimization, and dynamic execution through closed-loop interaction between intelligent agents and the power distribution network control system. Specifically, it involves using CT network learning to replace the real physical environment. The CT network model specifically includes: (1) The state transition model function is used to predict the probability distribution of the next state after taking a certain action in the current state: s′ t+1 =P(s t ,a t ) (34) (2) The reward module function is used to predict the immediate reward obtained after taking an action in the current state: r′ t =R(s t ,a t ) (35) Among them, s t a t These represent the state and action of the agent at time t, respectively.
10. The method according to claim 9, characterized in that, The closed-loop interaction between the intelligent agent and the power distribution network control system should include: (1) Constructing a power system perception-type CT network model: Using deep learning technology, a CT network model that can accurately represent the dynamic characteristics of a power system including wind power generation and photovoltaic power generation is constructed. (2) Efficient environmental simulation based on CT model: The constructed CT network model is used as the interactive environment of the intelligent agent; by interacting with the model, the intelligent agent can efficiently and cost-effectively simulate and access massive sample data of the power system under various operating states and wind and solar power output scenarios. (3) Dynamic programming exploration and strategy evaluation of the agent: In the environment based on the CT model, the agent explores dynamic programming strategies. The agent tries different scheduling decisions. For each strategy tried, the agent predicts and evaluates the expected cumulative reward that can be obtained under the guidance of the strategy in a series of future states based on the interaction results with the CT model. (4) Optimal strategy search objective setting: The optimization objective of the agent is set as follows: Under the premise of meeting the power system safety operation constraints, the optimal scheduling strategy is found through exploration and evaluation, so as to maximize the expected cumulative reward in the entire scheduling cycle or long-term operation. (5) Reinforcement learning training agent: The reinforcement learning algorithm is used to train the agent using the massive sample data generated efficiently by the CT model. During the training process, the agent updates its policy parameters according to the reward and state transition information obtained by the current policy, and finally converges to the optimal scheduling policy. (6) Apply the optimal scheduling strategy: Apply the obtained optimal scheduling strategy to the actual scheduling decision of the power system with wind and solar power to achieve the optimal system economy, security and new energy absorption capacity.
11. The method according to claim 9, characterized in that, The CT network specifically enables sequence data prediction in the following ways: (1) The CT network will input the current state s t and action a t It is transformed into a processable feature vector, and positional encoding is added to the feature vector to ensure that the CT network can capture the sequential relationship of the feature vector; (2) The self-attention mechanism in the network effectively captures long-distance dependencies in sequence data and extracts different types of information through multi-head attention. After passing through the self-attention network, an attention score matrix is obtained, which can represent the relationship between the inputs at each time step. Furthermore, temporal masking ensures that the inputs at each time step are only correlated with the inputs before that time step, thus reflecting causality. (3) The feedforward neural network composed of fully connected layers and ReLU layers performs nonlinear mapping on the output of the multi-head attention network to enhance the model's representation ability; (4) After layer normalization, linear transformation and multi-layer stacking are performed, and the final output is the predicted state change s′ of the environment. t+1 And the reward r′ t .
12. The method according to claim 9, characterized in that, The specific mathematical model of the CT module includes: (1) Input part: state s t and action a t The input feature sequence data will be mapped to a dimension d. pos The feature vectors are used, and the position of each feature vector is determined by position encoding: Where p is the position of the feature vector, i∈{1,…,d} pos / 2}; (2) Self-attention mechanism: Extract relevant feature information from the feature vector, and multiply the input feature vector X by the weight matrix W respectively. q W k W v We obtain the query matrix q, the key matrix k, and the value matrix v: The relevance of data is characterized by calculating the score of each query in the query matrix q and all keys in the key matrix k through dot product attention, and obtaining the corresponding value weight coefficient for each key, thereby determining the weighted combination of values under a given query. The self-attention score sequence e is obtained by weighted summation of the value matrix using the softmax function: Where att represents the self-attention mechanism, d is the dimension of the sequence; the softmax function transforms each element into a probability value in the interval (0,1) through a normalization exponent, and the sum of the probabilities of all elements is 1, which can increase the difference between values and make them easier to distinguish: To extract richer features, multi-head self-attention is used on top of self-attention to improve computational efficiency: in, This means concatenating vectors along the same dimension; the input vector first passes through a multi-head attention network to obtain the feature vector E; (3) Feedforward network: The ReLU layer is placed as an activation function after the hidden layer of the neural network to introduce nonlinear characteristics: f(x) = max(0,x) (41) Fully connected layers connect all neurons in the input or hidden layers to all neurons in the next layer, performing matrix multiplication and bias addition to map input features to the output for high-level feature extraction, classification, and regression. The feature vector E is then fed into a feedforward neural network layer composed of ReLU layers and fully connected layers to obtain: FFN(E)=max(0,EW1+b1)W2+b2 (42) Where FFN represents the feedforward neural network layer; W1, W2, b1, and b2 are the weights and biases of the network, respectively; layer normalization is added to the multi-head attention network and the feedforward neural network to ensure the stability of the data distribution. After layer normalization, it is as follows: layout=layernorm(x+layer(x)) (43) Where x is the input of the layer normalization, layer is the layer normalization neural network; layernorm represents the layer normalization, layout is the layer normalization output, the output value is linearly transformed and then passed through the softmax layer to obtain the probability distribution, and the value with the highest probability is selected as the prediction output. (4) Output section: The CT module network with multi-head attention predicts the sequence data as follows: Out=(s′ t+1 ,r′ t )=CT(s t ,a t ) (44) Among them, s t a t ,s′ t+1 、r′ t This represents the state sequence, action sequence, and reward sequence starting from time t.
13. A device for constructing an active power distribution network interactive carbon reduction decision-making intelligent agent based on carbon flow distribution, characterized in that, include: The carbon flow distribution dynamic tracking unit is used to construct a dynamic tracking model of carbon flow distribution in an active distribution network. Based on carbon flow theory, it establishes a carbon flow distribution matrix for multiple links from source to grid to load to storage, and calculates the real-time carbon flow intensity of each node in the distribution network through a node-level carbon emission intensity quantification model. The intelligent agent decision-making construction unit is used to design the Markov decision process framework of the interactive carbon reduction decision-making intelligent agent, and defines the state space, action space and reward function. The state space includes real-time carbon flow data, new energy output forecast values, load demand fluctuations and network topology dynamic change information. The action space includes controllable unit output adjustment, energy storage charging and discharging strategies, demand-side response commands and green electricity trading auxiliary strategies. The agent training unit is used to train agents using reinforcement learning algorithms driven by mechanism-data fusion. It integrates physical model constraints and historical operation data, and optimizes agent strategies through offline pre-training and online dynamic fine-tuning to ensure action space feasibility. The decision output unit is used to generate multi-objective collaborative optimization decisions based on dynamic carbon flow distribution. With the minimization of carbon emissions as the core objective, it simultaneously optimizes the renewable energy absorption rate, network line loss, and operational economy, and outputs real-time scheduling instructions through the intelligent agent.
14. An electronic system comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, characterized in that the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-13.
15. A computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.
Citation Information
Cited By
Carbon emission real-time regulation and control method and system based on multi-source data fusion and AI decision
CN121257983A
Power system carbon potential tracking and predicting method based on physical and data dual drive
CN121880901A
Carbon potential tracking and prediction method for power system based on physical and data double driving
CN121880901B
Power distribution network-oriented power carbon flexibility evaluation method, device and equipment thereof
CN122088135A
Power carbon flexibility evaluation method and device for power distribution network and equipment thereof
CN122088135B