Regional power grid day-ahead scheduling method based on reinforcement learning and carbon-green evidence coupling

By adopting a method based on reinforcement learning and carbon-green certificate coupling in regional power grid scheduling, combining deep reinforcement learning and SA-TD3 algorithm, the problem of inefficient calculation efficiency of traditional algorithms and the problem of parallel carbon trading and green certificate trading is solved, and efficient scheduling and low-carbon economic operation are achieved.

CN119944653AActive Publication Date: 2025-05-06HEFEI UNIV OF TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510109950.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Traditional optimization algorithms have low computational efficiency in regional power grid scheduling, making it difficult to effectively deal with complex and random environments, and carbon trading and green certificate trading are carried out in parallel, lacking benign coupling.

Method used

A regional power grid recent scheduling method based on reinforcement learning and carbon-green certificate coupling is adopted, and a scheduling framework is constructed through deep reinforcement learning, combined with the carbon-green certificate coupling model, and the SA-TD3 algorithm is introduced to optimize the scheduling plan.

Benefits of technology

It has improved the ability to deal with large-scale random problems, achieved benign coupling between the carbon trading market and the green certificate trading market, promoted the low-carbon economic operation of regional power grids, and improved the power grid's regulation capabilities and new energy consumption level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944653A_ABST
    Figure CN119944653A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power systems, and particularly relates to a regional power grid day-ahead scheduling method based on reinforcement learning and carbon-green certificate coupling. Firstly, a stepped carbon transaction model is constructed based on carbon quota distribution and carbon emission calculation, a green certificate transaction model is constructed based on a green certificate transaction mechanism, a carbon transaction market and a green certificate transaction market are considered for benign coupling, and a carbon-green certificate coupling model is constructed. And then, constructing a regional power grid day-ahead scheduling optimization model based on a deep reinforcement learning algorithm. And finally, introducing the core idea of the SA algorithm into the TD3 algorithm, proposing an SA-TD3 algorithm, and completing the solving of the regional power grid day-ahead scheduling optimization model by using the SA-TD3 algorithm to obtain a day-ahead scheduling plan. According to the method, the economical efficiency can be effectively improved under the condition of ensuring that the scheduling plan is obtained more quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power systems, and more specifically, relates to a regional power grid day-ahead dispatching method based on reinforcement learning and carbon-green certificate coupling. Background Art

[0003] With the exponential growth of dispatch objects in regional power grid dispatching and the huge amount of monitoring information, traditional optimization algorithms may be difficult to solve due to the problem of low computational efficiency. Reinforcement learning can better solve the optimization problems corresponding to the interaction between intelligent agents and the environment. With the relevant research combining the perception ability of deep learning with reinforcement learning, the learning ability and application scope of reinforcement learning have been further improved. In recent years, deep reinforcement learning methods have received increasing attention in the field of power system dispatching.

[0004] In summary, the present invention proposes a method for day-ahead dispatch of regional power grids based on reinforcement learning and carbon-green certificate coupling. This method first constructs a regional power grid dispatch framework considering the carbon-green certificate coupling trading mechanism based on deep reinforcement learning theory, and then constructs the state and action space of the intelligent agent in combination with deep learning theory, and gives the day-ahead dispatch optimization objectives and constraints of the regional power grid; finally, the core idea of ​​the SA algorithm is introduced into the existing TD3 algorithm, and a method for day-ahead dispatch optimization of regional power grids based on the carbon-green certificate coupling mechanism is proposed, and the established day-ahead dispatch optimization model of regional power grids considering the carbon-green certificate coupling mechanism is solved. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention proposes a regional power grid day-ahead scheduling method based on reinforcement learning and carbon-green certificate coupling. The method utilizes the characteristics of deep reinforcement learning and processes states and actions in random and complex environments through neural networks, which can effectively improve the processing capabilities of large-scale random problems. It also considers the benign coupling of the carbon trading market and the green certificate trading market to achieve flexible conversion between green certificates and carbon rights, and promote the low-carbon economic operation of the regional power grid.

[0006] To achieve the above object, the present invention adopts the following technical solution:

[0007] The regional power grid day-ahead dispatching method based on reinforcement learning and carbon-green certificate coupling includes the following steps:

[0008] Step 1: Construct models related to carbon trading and green certificate trading. Construct a step-by-step carbon trading model based on carbon quota allocation and carbon emission calculation, construct a green certificate trading model based on the green certificate trading mechanism, and construct a carbon-green certificate coupling model based on the benign coupling between carbon trading and green certificate trading.

[0009] Step 2: Construct a regional power grid day-ahead dispatch optimization model based on deep reinforcement learning algorithm. Introduce the deep reinforcement learning framework, and correspond the four stages of power grid operation status collection, dispatch instruction formulation, dispatch instruction execution, and dispatch instruction feedback involved in the actual dispatch process of the regional power grid with the four stages of observation status, action selection, action execution, and reward feedback involved in deep reinforcement learning, and construct a regional power grid dispatch framework based on deep reinforcement learning.

[0010] Step 3: Introduce the core idea of ​​the SA algorithm (Simulated Annealing algorithm, SA) into the Twin Delayed Deep Deterministic Policy Gradient algorithm TD3 (TD3 for short), propose the SA-TD3 algorithm, and use the SA-TD3 algorithm to solve the regional power grid day-ahead dispatch optimization model and obtain the day-ahead dispatch plan.

[0011] This technical solution is further optimized, and step 1 specifically includes the construction of a carbon trading model, a green certificate trading model, and a carbon-green certificate coupling model:

[0012] Step 1.1: Build a carbon trading model.

[0013] The formula for calculating the allocation of carbon quotas using the baseline method is as follows:

[0014]

[0015] In the formula, The initial carbon emission quota obtained for thermal power unit i; τ * is the carbon emission baseline per unit of electricity of power generation enterprises formulated in the industry; T is the number of time periods in the entire dispatch cycle, ΔT is the dispatch time interval, and there are T = 24h / ΔT decision periods in the whole day, P i,t is the output of thermal power unit i at time t;

[0016] Taking into account the relationship between the carbon emissions of thermal power generating units and the types of fuel and fuel consumption, the carbon emissions of thermal power generating units i under conventional output are:

[0017]

[0018] In the formula, is the unit fuel consumption carbon emission of thermal power unit i, a i 、b i 、c i is the unit power generation burn-up coefficient of thermal power unit i in normal operation, P i max , P i minare respectively the upper and lower limits of the conventional output of thermal power unit i. In addition, the present invention takes into account that some thermal power units can perform deep peak regulation. In this state, the combustion efficiency is reduced and the additional fuel loss is increased, which will cause additional carbon emissions. The carbon emissions are expressed as:

[0019]

[0020] In the formula, is the additional carbon emissions associated with the deep peak load regulation of the units, is the coefficient of additional fuel consumption generated when thermal power unit i performs deep peak regulation, P i a is the lower limit of the output of thermal power unit i under deep peak regulation;

[0021] Carbon trading for power generation companies meets the following requirements:

[0022]

[0023] In the formula, is the carbon trading cost of thermal power unit i; CET is the price per unit carbon certificate transaction; is the actual carbon emissions of thermal power unit i. Traditional carbon trading adopts a unified pricing principle, and the carbon trading price is uniform and fixed. In this way, power generation companies are not very motivated to reduce emissions. When thermal power units produce more carbon emissions, they should be given greater penalties to promote the emission reduction of thermal power units. The higher the carbon emission quota that thermal power units need to purchase, the higher the transaction price of the corresponding interval carbon certificate purchased from the carbon certificate market, that is,

[0024]

[0025] In the formula, is the basic price of unit carbon certificate transaction; θ CET Represents the sensitivity coefficient of carbon price growth rate; d CET is the interval length of a single carbon emission level. Therefore, the tiered carbon trading cost of thermal power units can be expressed as:

[0026]

[0027] Step 1.2: Construct a green certificate trading model.

[0028] The new energy generating units in the regional power grid need to produce a certain proportion of green electricity, that is, to formulate the initial allocation benchmark value of the power generation of the new energy generating units and issue a certain green certificate quota index accordingly. The green certificate quota index and the number of green certificates obtained by the new energy generating units in the regional power grid meet the following requirements:

[0029]

[0030] In the formula, is the green certificate quota index and green certificate quantity obtained by the new energy generator j in the regional power grid, α l is the green certificate quota coefficient, is the initial allocation benchmark value and actual power generation of the new energy generator j in period t, α f The coefficient for obtaining the number of green certificates for each unit of electricity generated by new energy. The green certificate transaction introduces a tiered pricing model. The more carbon certificates a new energy unit needs to purchase, the higher the transaction price of green certificates purchased from the green certificate market for the corresponding range. The surplus green certificates are uniformly recovered by the government or the carbon certificate market at the basic price. The model satisfies:

[0031]

[0032] In the formula, is the green certificate transaction cost of new energy unit j, is the basic price of green certificate transaction, d GCT is the interval length of a single green certificate trading level, θ GCT It is the growth coefficient of green certificate price.

[0033] Step 1.3: Construct a carbon-green certificate coupling trading model.

[0034] The emission reduction of new energy is the difference between the carbon emissions generated by new energy power generation and the carbon emission standards set by the industry. Wind power and photovoltaic power generation can be regarded as zero carbon emissions, so the emission reduction of new energy power generation is:

[0035]

[0036] In the formula, The amount of carbon emissions reduced by the new energy generator j; Standard carbon emissions set for the industry; The carbon emissions generated by the new energy power generation enterprise j are taken as zero in the present invention. Therefore, the green certificate obtained by the new energy represents the carbon emission rights:

[0037]

[0038] In the formula, β is the emission reduction represented by a unit green certificate. Therefore, according to the above relationship, the conversion between green certificates and carbon emission rights is as follows:

[0039]

[0040] In the formula, Carbon emission rights converted from green certificates purchased for thermal power unit i; is the number of green certificates purchased by thermal power unit i from the green certificate market; therefore, the carbon trading and green certificate trading of thermal power unit i under the carbon-green certificate coupling mechanism can be expressed as:

[0041]

[0042] In the formula, is the cost of purchasing green certificates from the green certificate market for thermal power unit i.

[0043] The technical solution is further optimized, and the step 2 specifically includes:

[0044] By using the one-to-one correspondence between the four stages involved in the actual dispatching process of the regional power grid and the four stages involved in deep reinforcement learning, the regional power grid dispatching problem is modeled as a Markov decision process: the dispatching time of a whole day is divided into t periods, and the environmental state s t Corresponding to the operating status of the regional power grid at time t, action a t Corresponding to the dispatch plan of the regional power grid at time t+1, in the environmental state s t Next, perform action a t , the regional power grid shifts to the next environmental state s t+1 , which also corresponds to the action a of the regional power grid dispatch plan at time t+2 t+1 Finally, the agent will give feedback on the accumulated rewards, corresponding to the regional power grid's economic performance, carbon emissions and other indicators related to the dispatch plan C t Therefore, the Markov decision process modeling of the regional power grid can be expressed as:

[0045] M e ={s t ,a t ,C t ,s t+1}

[0046] The source load resources considered in this invention mainly include coal-fired power generation, gas-fired power generation, wind power generation, and photovoltaic power generation. The load-side resources mainly consider the movable load, the reducible load, and the rigid load. The set of coal-fired thermal power generation units is defined as Φ pg ={1,2,...,N g}, the set of gas-fired power generating units is Φ gas ={1,2,...,N gas}, the wind turbine generator set is Φ wind ={1,2,...,N w}, the photovoltaic power station set is Φ PV ={1,2,...,N v}. Let the output value of coal-fired power generation unit i at time t be P i,t , the output value of gas-fired power unit i at time t is The output of wind turbine i at time t is The output value of the photovoltaic power station at time t is The actual value of the load at time t is P t load , where the load that can be reduced is P t cut , the translatable load is P t sh .

[0047] Step 2.1: Construct the state value space.

[0048] In order to establish the state space, first divide the scheduling time T in a whole day into t scheduling periods, that is, T = {1, 2, ..., t}; the system state at time t is s t , which includes time t, the unit output P of the coal-fired unit at time t i,t ,i∈{1,...,N g}、The startup time of coal-fired power unit i at time t The shutdown time of coal-fired power unit i at time k The output of gas-fired power unit i at time t The startup time of gas-fired power unit i at time t The downtime of gas-fired power unit i at time t The output of the wind turbine at time k Output of photovoltaic power station at time t The load power at time t is P t load , specifically described as follows:

[0049]

[0050] To establish the action space, record the agent’s current state s at time t t The response is a t , that is, the dispatch plan executed by the regional power grid at time t, and the output adjustment of the coal-fired unit in the current state is ΔP i,t ,i∈{1,...,N g}, the output adjustment of the gas unit in the current state is The action space is composed of the output adjustment of all coal-fired units and gas-fired units in the current state, which can be expressed as:

[0051]

[0052] Step 2.2: Construct the objective function and constraints.

[0053] The day-ahead dispatching cost of the regional power grid is composed of the operating cost of thermal power units, the carbon trading cost of thermal power units, the green certificate trading cost of thermal power units, the green certificate trading cost of new energy units, the dispatching cost of reducible loads, the dispatching cost of shiftable loads, the penalty cost of abandoning new energy, and the penalty cost of load shedding. The specific expression of the cost model of the optimization target is as follows:

[0054] Operating costs of thermal power units:

[0055]

[0056] In the formula, C g is the operating cost of thermal power unit i, including the operating cost of coal-fired thermal power units and gas-fired units.

[0057] Carbon trading costs of thermal power units:

[0058]

[0059] In the formula, C CET It is the carbon trading cost of thermal power units, including the carbon trading cost of coal-fired thermal power units and the carbon trading cost of gas-fired thermal power units.

[0060] Green certificate transaction cost of thermal power units:

[0061]

[0062] In the formula, C buy-GCT It is the green certificate transaction cost of thermal power units, including the green certificate transaction cost of coal-fired thermal power units and the green certificate transaction cost of gas-fired thermal power units.

[0063] Green certificate transaction cost for new energy units:

[0064]

[0065] In the formula, C GCT It is the green certificate transaction cost of new energy units, including the green certificate transaction cost of wind turbines and the green certificate transaction cost of photovoltaic power stations.

[0066] The dispatching cost of load can be reduced:

[0067] Constructing a stepped load reduction compensation price model:

[0068]

[0069] In the formula, is the compensation price for the unit reduction of load at time t, P t cut is the amount of power that can be reduced, P t cut,max , Pt cut,min The maximum and minimum reduction amounts that can be accepted for load reduction are: is the growth coefficient function related to the reduction amount, is the initial compensation price for the curtailable load; the curtailment compensation cost for the curtailable load response reduction satisfies:

[0070]

[0071] In the formula, C pay,cut The cost of reducing the load that can be reduced is compensated.

[0072] In addition, the load that can be reduced can also provide spare capacity on the load side, making the load side resources more flexible. The spare capacity that can be reduced is determined when making the day-ahead plan, and the upper limit constraint of the spare capacity must be met:

[0073]

[0074] In the formula, The load-side reserve capacity and capacity upper limit provided for the load that can be reduced. Within the reserve capacity constraint range, the reserve cost brought by providing reserve capacity to respond to the grid dispatching requirements of the load that can be reduced meets the following requirements:

[0075]

[0076] In the formula, C re,cut The reserve cost caused by providing load-side reserve capacity for curtailable loads, The price for providing a unit of reserve capacity for curtailable load.

[0077] Therefore, the dispatch cost of load reduction includes the reduction compensation cost and the reserve cost, which satisfies:

[0078] C cut =C pay,cut +C re,cut

[0079] In the formula, C cut The dispatching cost of the load can be reduced.

[0080] Dispatch cost of shiftable loads:

[0081] The dispatch model of the shiftable load satisfies:

[0082]

[0083] Where, t sh is the start time of the movable load after scheduling, t sh- ,t sh+P is the earliest and latest operation start time that the shiftable load is willing to respond to the grid demand, t sh is the power consumption of the movable load in period t after scheduling, T sh is the continuous operation time of the translatable load, After the shiftable load is dispatched, sh +τ period of power consumption, before dispatching at t sh* +Power consumption during the τ period, t sh* ,t sh It is the start time of the movable load before and after scheduling;

[0084] Acceptable scheduling time for shiftable loads [t sh- ,t sh+ ]satisfy:

[0085]

[0086] In the formula, ρ sh is the unit compensation price of the translatable load, β sh (ρ sh ) is the acceptable translation time and compensation price of the movable load ρ sh Related related functions, The maximum number of periods during which a load can be translated.

[0087] The function of the shiftable load compensation price satisfies:

[0088]

[0089] In the formula, ρ sh- and ρ sh+ are the minimum compensation price and maximum compensation price of the unit load that can be translated when the dispatch agreement is signed, Δρ sh is the sensitivity coefficient related to the compensation price.

[0090] Therefore, the dispatch cost of the shiftable load satisfies:

[0091]

[0092] In the formula, C sh is the dispatching cost of the movable load; τ sh is a Boolean variable, τ sh =1 means the translatable load is translated, τ sh =0 means that the translatable load does not translate.

[0093] New energy disposal penalty costs:

[0094]

[0095] In the formula, C cnew is the penalty cost for abandoning new energy, λ cw , cv is the penalty coefficient for unit abandoned electricity of wind power and photovoltaic power, is the abandoned power of wind turbine i in period t, is the abandoned electricity of PV plant j at time t.

[0096] Load shedding penalty cost:

[0097]

[0098] In the formula, C cl is the penalty cost of load shedding in the regional power grid, P t cl is the load shedding amount of the regional power grid during period t.

[0099] Therefore, the optimization goal of the regional power grid day-ahead dispatch is to minimize the day-ahead dispatch cost:

[0100] minC day =min(C g +C CET +C buy-GCT +C GCT +C cut +C sh +C cnew +C cl )

[0101] In order to maintain safe and stable operation, the regional power grid needs to meet the line transmission capacity flow constraints in addition to the constraints related to the operation of the thermal power units themselves, the constraints related to the load that can be reduced, and the constraints related to the load that can be shifted.

[0102] The transmission capacity constraints are satisfied:

[0103]

[0104] Where, T l,i , T l,j , T l,k , T l,m and T l,n is the power transfer allocation coefficient; N l is the number of node loads in the regional power grid; is the load power of node m at time t; F l max is the upper limit of the power flow of line l.

[0105] The technical solution is further optimized, and the step 3 specifically includes:

[0106] Step 3.1: Introduce a regularization strategy for smoothing the target strategy.

[0107] TD3 smoothes the Q function to make the probability distribution of action selection more balanced, thereby reducing the deviation when estimating the action value and improving the performance of the algorithm. After smoothing, action a is expressed as:

[0108] a←μ'(s|θ μ' )+η

[0109] η←clip(N(0,σ),-c,c)

[0110] In the formula, μ'(s|θ μ' ) is the policy function, s is the given state, θ μ' is the parameter of the target network, η is the noise added during action selection, clip(·) is the cutoff function that limits the maximum and minimum values ​​of the noise, and N(0,σ) is a normal distribution function with a mean of 0 and a variance of σ.

[0111] Step 3.2: Introduce the core idea of ​​the SA algorithm into the TD3 algorithm.

[0112] The TD3 algorithm introduces two groups of Critic network structures and selects the smaller Q value in the two groups of networks to update the Actor network, thereby reducing the impact of Q value deviation. The network update function of the TD3 algorithm is specifically expressed as a loss function as follows:

[0113] y=-C+γmin(Q'1(s,a),Q'2(s,a)),γ∈[0,1]

[0114]

[0115] In the formula, γ is the discount factor, N is the total number of samples used, s, a and C are the sampled batch states, actions and running states respectively. is the loss function of Critic network i. The TD3 algorithm reduces the frequency of Actor network updates, thereby reducing the amount of learned erroneous information and making the learned policy network more stable. The update gradient of the Actor network is satisfy:

[0116]

[0117] The specific steps of the algorithm flow of the SA algorithm are as follows:

[0118] Step 1: Initialize k = 0, based on the initial annealing temperature T k Find the initial random solution x0, at this time x k = x0;

[0119] Step 2: For a feasible solution xk Add a random noise Δx k Get x' k =x k +Δx k , respectively calculate the evaluation function value f(x k ) and f(x' k ). If f(x k )<f(x' k ), then x' k =x k Otherwise, when Time x' k =x k +Δx k Repeat step 2 k Second-rate;

[0120] Step 3: Calculate T k+1 =Q SA T k , where Q SA is the annealing factor in the simulated annealing algorithm. Let k = k + 1, and determine whether the algorithm termination condition is met. If it is met, stop; if not, return to step 2;

[0121] The core idea of ​​the SA algorithm is introduced into the TD3 algorithm, and the first Critic network is used as the exploration network of the initial solution in the SA algorithm, and the second Critic network is used as the exploration network of the disturbance term in the SA algorithm. Therefore, the network update formula of the SA-TD3 algorithm is specifically expressed as follows:

[0122]

[0123] The SA-TD3 algorithm is used to solve the regional power grid day-ahead dispatch optimization model. The specific steps for obtaining the day-ahead dispatch plan are as follows:

[0124] Step 3.3: Initialize the policy network π φ , Valuation Network Target Network Experience pool β, set learning rate α, discount factor γ, delayed update step number d, soft update coefficient τ and maximum learning step number maxepisode, learning step number count = 0, initial annealing temperature T k , and the state space s constructed above t As algorithm input, it is input into the Actor network;

[0125] Step 3.4: According to the output of the Actor network, introduce Gaussian noise and generate the current action a under the condition of satisfying relevant constraints t =μ(s t |θ μ)+η, and schedule a t Send to the power grid simulation environment to calculate the next state s t+1 And the current running cost C t That is, considering the carbon and green certificate transaction costs of power generation units under the carbon-green certificate coupling transaction mechanism and other related day-ahead scheduling costs, the sample data (s t ,a t ,C t ,s t+1 ) is stored in the experience pool, and t=t+1 is used to update the Critic network Let count = count + 1. If the number of learning steps count is divisible by the number of delayed update steps d, then based on the SA algorithm strategy, update the Actor network θ μ , soft update target network Minimize the day-ahead scheduling cost and repeat this step, otherwise directly repeat this step until the maximum number of learning steps is reached, and output the optimized scheduling plan, that is, the output of each unit.

[0126] Different from the prior art, the beneficial effects of the present invention are mainly manifested in:

[0127] (1) In the traditional mode, carbon trading and green certificate trading in the regional power grid are carried out in parallel. The present invention allows the carbon trading market and the green certificate trading market to be benignly coupled to achieve flexible conversion between green certificates and carbon rights, and promote the low-carbon economic operation of the regional power grid.

[0128] (2) The proposed SA-TD3 algorithm improves the ability to escape from local optimality during training by introducing the core idea of ​​simulated annealing, thereby improving the training speed of the TD3 algorithm and maintaining the inherent stability of the TD3 algorithm.

[0129] (3) The integration of gas-fired thermal power units into the power grid further improves the flexibility of source-side resources, enhances the regulation capacity of the power grid, effectively improves the level of new energy consumption, reduces the proportion of maximum load shedding, and thus improves the economy of the entire regional power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0130] Figure 1 This is a diagram of the carbon-green certificate trading coupling mechanism;

[0131] Figure 2 This is a schematic diagram of the carbon-green certificate coupling model;

[0132] Figure 3 This is a schematic diagram of regional power grid dispatch optimization based on the SA-TD3 algorithm;

[0133] Figure 4 This is a flow chart of the regional power grid dispatch optimization method based on the SA-TD3 algorithm. DETAILED DESCRIPTION

[0134] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.

[0135] The present invention proposes a method for day-ahead dispatching of regional power grids based on reinforcement learning and carbon-green certificate coupling. The method considers the benign coupling between the carbon trading market and the green certificate trading market, and thereby constructs a carbon-green certificate coupling model, so that the day-ahead dispatching plan is more in line with actual needs, and realizes the flexible conversion between green certificates and carbon rights, and promotes the low-carbon economic operation of regional power grids. The proposed SA-TD3 algorithm improves the training speed of the TD3 algorithm and maintains the inherent stability of the TD3 algorithm.

[0136] See also Figure 1 , Carbon-green certificate trading coupling mechanism diagram. The government is responsible for the issuance, approval and recovery of carbon certificates and green certificates; the green certificate market conducts centralized trading of green certificates of new energy units; in addition to purchasing carbon emission quotas from the carbon certificate market, thermal power units can also purchase green certificates from the green certificate market to offset carbon emissions.

[0137] See Figure 2 As shown, the schematic diagram of the carbon-green certificate coupling model includes the following steps:

[0138] Step 1: Construct a model related to carbon trading and green certificate trading.

[0139] A stepped carbon trading model is constructed based on carbon quota allocation and carbon emission calculation, and a green certificate trading model is constructed based on the green certificate trading mechanism. This allows the carbon trading market and the green certificate trading market to be benignly coupled, realizes flexible conversion between green certificates and carbon rights, and constructs a carbon-green certificate coupling model.

[0140] Step 1.1: Build a carbon trading model.

[0141] The formula for calculating the allocation of carbon quotas using the baseline method is as follows:

[0142]

[0143] In the formula, The initial carbon emission quota obtained for thermal power unit i; τ * is the carbon emission baseline per unit of electricity of power generation enterprises formulated in the industry; T is the number of time periods in the entire dispatch cycle, ΔT is the dispatch time interval, and there are T = 24h / ΔT decision periods in the whole day; P i,t is the output of thermal power unit i at time t.

[0144] Taking into account the relationship between the carbon emissions of thermal power generating units and the type of fuel and fuel consumption, the carbon emissions of thermal power generating units i under conventional output are

[0145]

[0146] In the formula, is the unit fuel consumption carbon emission of thermal power unit i, a i 、b i 、c i is the unit power generation burn-up coefficient of thermal power unit i in normal operation, P i max , P i min are respectively the upper and lower limits of the conventional output of thermal power unit i. In addition, the present invention takes into account that some thermal power units can perform deep peak regulation. In this state, the combustion efficiency is reduced and the additional fuel loss is increased, which will cause additional carbon emissions. The carbon emissions are expressed as:

[0147]

[0148] In the formula, is the additional carbon emissions associated with the deep peak load regulation of the units, is the coefficient of additional fuel consumption generated when thermal power unit i performs deep peak regulation, P i a It is the lower limit of the output of thermal power unit i under deep peak regulation state.

[0149] Carbon trading for power generation companies meets the following requirements:

[0150]

[0151] In the formula, is the carbon trading cost of thermal power unit i; CET is the price per unit carbon certificate transaction; is the actual carbon emissions of thermal power unit i. Traditional carbon trading adopts a unified pricing principle, and the carbon trading price is uniform and fixed. In this way, power generation companies are not very motivated to reduce emissions. When thermal power units produce more carbon emissions, they should be given greater penalties to promote the emission reduction of thermal power units. The higher the carbon emission quota that thermal power units need to purchase, the higher the transaction price of the corresponding interval carbon certificate purchased from the carbon certificate market, that is,

[0152]

[0153] In the formula, is the basic price of unit carbon certificate transaction; θ CET Represents the sensitivity coefficient of carbon price growth rate; d CET is the interval length of a single carbon emission level. Therefore, the tiered carbon trading cost of thermal power units can be expressed as:

[0154]

[0155] Step 1.2: Construct a green certificate trading model.

[0156] New energy generating units within the regional power grid need to produce a certain proportion of green electricity, that is, to establish an initial allocation benchmark value for the power generation of new energy units and to issue certain green certificate quota indicators accordingly.

[0157]

[0158] In the formula, is the green certificate quota index and green certificate quantity obtained by the new energy generator j in the regional power grid, α l is the green certificate quota coefficient, is the initial allocation benchmark value and actual power generation of the new energy generator j in period t, α f The coefficient for obtaining the number of green certificates for each unit of electricity generated by new energy. The green certificate transaction introduces a tiered pricing model. The more carbon certificates a new energy unit needs to purchase, the higher the transaction price of green certificates purchased from the green certificate market for the corresponding range. The surplus green certificates are uniformly recovered by the government or the carbon certificate market at the basic price. The model satisfies:

[0159]

[0160] In the formula, is the green certificate transaction cost of new energy unit j, is the basic price of green certificate transaction, d GCT is the interval length of a single green certificate trading level, θ GCT It is the growth coefficient of green certificate price.

[0161] Step 1.3: Construct a carbon-green certificate coupling trading model.

[0162] The emission reduction of new energy is the difference between the carbon emissions generated by new energy power generation and the carbon emission standards set by the industry. Wind power and photovoltaic power generation can be regarded as zero carbon emissions, so the emission reduction of new energy power generation is:

[0163]

[0164] In the formula, The amount of carbon emissions reduced by the new energy generator j; Standard carbon emissions set for the industry; The carbon emissions generated by the new energy power generation enterprise j are taken as zero in the present invention. Therefore, the green certificate obtained by the new energy represents the carbon emission rights:

[0165]

[0166] In the formula, β is the emission reduction represented by a unit green certificate. Therefore, according to the above relationship, the conversion between green certificates and carbon emission rights is as follows:

[0167]

[0168] In the formula, Carbon emission rights converted from green certificates purchased for thermal power unit i; is the number of green certificates purchased by thermal power unit i from the green certificate market. Therefore, the carbon trading and green certificate trading of thermal power unit i under the carbon-green certificate coupling mechanism can be expressed as:

[0169]

[0170]

[0171] In the formula, is the cost of purchasing green certificates from the green certificate market for thermal power unit i.

[0172] See Figure 3 As shown in the figure, the schematic diagram of regional power grid dispatch optimization based on SA-TD3 algorithm includes the following steps:

[0173] Step 2: Construct a regional power grid day-ahead dispatch optimization model based on deep reinforcement learning algorithm. The actual dispatch process of the regional power grid involves four stages: grid operation status collection, dispatch instruction formulation, dispatch instruction execution, and dispatch instruction feedback. Deep reinforcement learning involves four stages: observation status, action selection, action execution, and reward feedback. Introduce the deep reinforcement learning framework, and use the one-to-one correspondence between the regional power grid optimization problem and the deep reinforcement learning framework to construct a regional power grid dispatch framework based on deep reinforcement learning.

[0174] By using the one-to-one correspondence between the four stages involved in the actual dispatching process of the regional power grid and the four stages involved in deep reinforcement learning, the regional power grid dispatching problem is modeled as a Markov decision process: the dispatching time of a whole day is divided into t periods, and the environmental state s t Corresponding to the operating status of the regional power grid at time t, action a t Corresponding to the dispatch plan of the regional power grid at time t+1, in the environmental state s t Next, perform action a t , the regional power grid shifts to the next environmental state s t+1 , which also corresponds to the action a of the regional power grid dispatch plan at time t+2 t+1 Finally, the agent will give feedback on the accumulated rewards, corresponding to the regional power grid's economic performance, carbon emissions and other indicators related to the dispatch plan C t Therefore, the Markov decision process modeling of the regional power grid can be expressed as:

[0175] Me ={s t ,a t ,C t ,s t+1}

[0176] The source load resources considered in this invention mainly include coal-fired power generation, gas-fired power generation, wind power generation, and photovoltaic power generation. The load-side resources mainly consider the movable load, the reducible load, and the rigid load. The set of coal-fired thermal power generation units is defined as Φ pg ={1,2,...,N g}, the set of gas-fired power generating units is Φ gas ={1,2,...,N gas}, the wind turbine generator set is Φ wind ={1,2,...,N w}, the photovoltaic power station set is Φ PV ={1,2,...,N v}. Let the output value of coal-fired power generation unit i at time t be P i,t , the output value of gas-fired power unit i at time t is The output value of wind turbine i at time t is The output value of the photovoltaic power station at time t is The actual value of the load at time t is P t load , where the load that can be reduced is P t cut , the translatable load is P t sh .

[0177] Step 2.1: Construct the state value space.

[0178] In order to establish the state space, first divide the scheduling time T in a whole day into t scheduling periods, that is, T = {1, 2, ..., t}; the system state at time t is s t , which includes time t, the unit output P of the coal-fired unit at time t i,t ,i∈{1,...,N g}、The startup time of coal-fired power unit i at time t The shutdown time of coal-fired power unit i at time k The output of gas-fired power unit i at time t The startup time of gas-fired power unit i at time t The downtime of gas-fired power unit i at time t The output of the wind turbine at time k Output of photovoltaic power station at time t The load power at time t is Pt load , specifically described as follows:

[0179]

[0180] To establish the action space, record the agent’s current state s at time t t The response is a t , that is, the dispatch plan executed by the regional power grid at time t, and the output adjustment of the coal-fired unit in the current state is ΔP i,t ,i∈{1,...,N g}, the output adjustment of the gas unit in the current state is The action space is composed of the output adjustment of all coal-fired units and gas-fired units in the current state, which can be expressed as:

[0181]

[0182] Step 2.2: Construct the objective function and constraints.

[0183] The day-ahead dispatching cost of the regional power grid is composed of the operating cost of thermal power units, the carbon trading cost of thermal power units, the green certificate trading cost of thermal power units, the green certificate trading cost of new energy units, the dispatching cost of reducible loads, the dispatching cost of shiftable loads, the penalty cost of abandoning new energy, and the penalty cost of load shedding. The specific expression of the cost model of the optimization target is as follows:

[0184] Operating costs of thermal power units:

[0185]

[0186] In the formula, C g is the operating cost of thermal power unit i, including the operating cost of coal-fired thermal power units and gas-fired units.

[0187] Carbon trading costs of thermal power units:

[0188]

[0189] In the formula, C CET It is the carbon trading cost of thermal power units, including the carbon trading cost of coal-fired thermal power units and the carbon trading cost of gas-fired thermal power units.

[0190] Green certificate transaction cost of thermal power units:

[0191]

[0192] In the formula, C buy-GCT It is the green certificate transaction cost of thermal power units, including the green certificate transaction cost of coal-fired thermal power units and the green certificate transaction cost of gas-fired thermal power units.

[0193] Green certificate transaction cost for new energy units:

[0194]

[0195] In the formula, C GCT It is the green certificate transaction cost of new energy units, including the green certificate transaction cost of wind turbines and the green certificate transaction cost of photovoltaic power stations.

[0196] The dispatching cost of load can be reduced:

[0197] Constructing a stepped load reduction compensation price model:

[0198]

[0199] In the formula, is the compensation price for the unit reduction of load at time t, P t cut is the amount of power that can be reduced, P t cut,max , P t cut,min The maximum and minimum reduction amounts that can be accepted for load reduction are: is the growth coefficient function related to the reduction amount, is the initial compensation price for the curtailable load; the curtailment compensation cost for the curtailable load response reduction satisfies:

[0200]

[0201] In the formula, C pay,cut The cost of reducing the load that can be reduced is compensated.

[0202] In addition, the load that can be reduced can also provide spare capacity on the load side, making the load side resources more flexible. The spare capacity that can be reduced is determined when making the day-ahead plan, and the upper limit constraint of the spare capacity must be met:

[0203]

[0204] In the formula, The load-side spare capacity and capacity upper limit provided for the load that can be cut. Within the spare capacity constraint range, the spare cost brought by providing spare capacity to respond to the grid dispatching requirements of the load that can be cut meets:

[0205]

[0206] In the formula, C re,cut The reserve cost caused by providing load-side reserve capacity for curtailable loads, The price for providing a unit of reserve capacity for curtailable load.

[0207] Therefore, the dispatch cost of load reduction includes the reduction compensation cost and the reserve cost, which satisfies:

[0208] C cut =C pay,cut +C re,cut

[0209] In the formula, C cut The dispatching cost of the load can be reduced.

[0210] Dispatch cost of shiftable loads:

[0211] The dispatch model of the shiftable load satisfies:

[0212]

[0213] Where, t sh is the start time of the movable load after scheduling, t sh- ,t sh+ P is the earliest and latest operation start time that the shiftable load is willing to respond to the grid demand, t sh is the power consumption of the movable load in period t after scheduling, T sh is the continuous operation time of the translatable load, After the shiftable load is dispatched, sh +τ period of power consumption, before dispatching at t sh* +Power consumption during the τ period, t sh* ,t sh It is the start time of the movable load before and after scheduling;

[0214] Acceptable scheduling time for shiftable loads [t sh- ,t sh+ ]satisfy:

[0215]

[0216] In the formula, ρ sh is the unit compensation price of the translatable load, β sh (ρ sh ) is the acceptable translation time and compensation price of the movable load ρ sh Related related functions, is the maximum number of periods during which the load can be translated;

[0217] The function of the shiftable load compensation price satisfies:

[0218]

[0219] In the formula, ρ sh- and ρsh+ are the minimum compensation price and maximum compensation price of the unit load that can be translated when the dispatch agreement is signed, Δρ sh is the sensitivity coefficient related to the compensation price;

[0220] Therefore, the dispatch cost of the shiftable load satisfies:

[0221]

[0222] In the formula, C sh is the dispatching cost of the movable load; τ sh is a Boolean variable, τ sh =1 means the translatable load is translated, τ sh =0 means that the translatable load does not translate.

[0223] New energy disposal penalty costs:

[0224]

[0225] In the formula, C cnew is the penalty cost for abandoning new energy, λ cw , cv is the penalty coefficient for unit abandoned electricity of wind power and photovoltaic power, is the abandoned power of wind turbine i in period t, is the abandoned electricity of PV plant j at time t.

[0226] Load shedding penalty cost:

[0227]

[0228] In the formula, C cl is the penalty cost of load shedding in the regional power grid, P t cl is the load shedding amount of the regional power grid during period t.

[0229] Therefore, the optimization goal of the regional power grid day-ahead dispatch is to minimize the day-ahead dispatch cost:

[0230] minC day =min(C g +C CET +C buy-GCT +C GCT +C cut +C sh +C cnew +C cl )

[0231] In order to maintain safe and stable operation, the regional power grid needs to meet the line transmission capacity flow constraints in addition to the constraints related to the operation of the thermal power units themselves, the constraints related to the load that can be reduced, and the constraints related to the load that can be shifted.

[0232] The transmission capacity constraints are satisfied:

[0233]

[0234] Where, T l,i , T l,j , T l,k , T l,m and T l,n is the power transfer allocation coefficient; N l is the number of node loads in the regional power grid; is the load power of node m at time t; F l max is the upper limit of the power flow of line l;

[0235] See Figure 4 As shown in the flowchart of the regional power grid dispatch optimization method based on the SA-TD3 algorithm, the steps are as follows:

[0236] Step 3: Introduce the core idea of ​​the SA algorithm into the TD3 algorithm, propose the SA-TD3 algorithm, and use the SA-TD3 algorithm to solve the regional power grid day-ahead dispatch optimization model and obtain the day-ahead dispatch plan.

[0237] Step 3.1: Introduce a regularization strategy for smoothing the target strategy.

[0238] TD3 smoothes the Q function to make the probability distribution of action selection more balanced, thereby reducing the deviation value when estimating the action value and improving the performance of the algorithm. It is expressed as:

[0239] a←μ'(s|θ μ' )+η

[0240] η←clip(N(0,σ),-c,c)

[0241] In the formula, μ'(s|θ μ' ) is the policy function, s is the given state, θ μ' is the parameter of the target network, η is the noise added during action selection, clip(·) is the cutoff function that limits the maximum and minimum values ​​of the noise, and N(0,σ) is a normal distribution function with a mean of 0 and a variance of σ.

[0242] Step 3.2: Introduce the core idea of ​​the SA algorithm into the TD3 algorithm.

[0243] In the DDPG algorithm, the maximum Q value is calculated through a single Critic network, which causes the Q value to be too high to accumulate. The TD3 algorithm introduces two groups of Critic network structures and updates the Actor network by selecting the smaller Q value in the two groups of networks, thereby reducing the impact of Q value deviation. The network update function of the TD3 algorithm is specifically expressed as a loss function as follows:

[0244] y=-C+γmin(Q'1(s,a),Q'2(s,a)),γ∈[0,1]

[0245]

[0246] In the formula, γ is the discount factor, N is the total number of samples used, s, a and C are the sampled batch states, actions and running states respectively. is the loss function of Critic network i. If the Actor network is updated immediately every time the target value is updated, the large error in the updated target value will continue to accumulate, thus affecting the convergence of the final optimization result. Therefore, the TD3 algorithm reduces the frequency of Actor network updates to reduce the amount of learned error information, making the learned policy network more stable. The update gradient of the Actor network is satisfy:

[0247]

[0248] However, the TD3 algorithm will undoubtedly reduce the exploration efficiency of the algorithm by choosing a lower Q value update each time, and may fall into a local optimum. The SA algorithm is a general probabilistic algorithm, which is often used to seek the optimal solution to random optimization problems with a large solution space. Its name comes from the term annealing in the metallurgical field. The core idea of ​​the SA algorithm draws on the annealing principle of metals: the theory of thermodynamics is applied to the field of statistics, and each solution within a search range is compared to a molecule in the air. Each solution within the search range has "energy" like a molecule in the air. The specific steps of the SA algorithm's algorithm flow are as follows:

[0249] Step 1: Initialize k = 0, based on the initial annealing temperature T k Find the initial random solution x0, at this time x k = x0;

[0250] Step 2: For a feasible solution x k Add a random noise Δx k Get x' k =x k +Δx k , respectively calculate the evaluation function value f(x k ) and f(x' k ). If f(xk )<f(x' k ), then x' k =x k Otherwise, when Time x' k =x k +Δx k Repeat step 2 k Second-rate;

[0251] Step 3: Calculate T k+1 =Q SA T k , where Q SA is the annealing factor in the simulated annealing algorithm. Let k = k + 1, and determine whether the algorithm termination condition is met. If it is met, stop; if not, return to step 2;

[0252] At this point, the core idea of ​​the SA algorithm is introduced into the TD3 algorithm, and the first Critic network is used as the exploration network of the initial solution in the SA algorithm, and the second Critic network is used as the exploration network of the perturbation term in the SA algorithm. Therefore, the network update formula of the SA-TD3 algorithm is specifically expressed as follows:

[0253]

[0254] The SA-TD3 algorithm is used to solve the regional power grid day-ahead dispatch optimization model. The specific steps for obtaining the day-ahead dispatch plan are as follows:

[0255] Step 3.3: Initialize the policy network π φ , Valuation Network Target Network Experience pool β, set learning rate α, discount factor γ, delayed update step number d, soft update coefficient τ and maximum learning step number maxepisode, learning step number count = 0, initial annealing temperature T k , and the state space s constructed above t As algorithm input, it is input into the Actor network;

[0256] Step 3.4: According to the output of the Actor network, introduce Gaussian noise and generate the current action a under the condition of satisfying relevant constraints t =μ(s t |θ μ )+η, and schedule a t Send to the power grid simulation environment to calculate the next state s t+1 And the current running cost C t That is, considering the carbon and green certificate transaction costs of power generation units under the carbon-green certificate coupling transaction mechanism and other related day-ahead scheduling costs, the sample data (s t,a t ,C t ,s t+1 ) is stored in the experience pool, and t=t+1 is used to update the Critic network Let count = count + 1. If the number of learning steps count is divisible by the number of delayed update steps d, then based on the SA algorithm strategy, update the Actor network θ μ , soft update target network Minimize the day-ahead scheduling cost and repeat this step, otherwise repeat this step directly until the maximum number of learning steps is reached, and output the optimized scheduling plan, that is, the output of each unit.

Claims

1. A regional power grid day-ahead dispatching method based on reinforcement learning and carbon-green certificate coupling, characterized in that: The following steps are involved: Step 1: Construct a model related to carbon trading and green certificate trading. Construct a stepped carbon trading model based on carbon quota allocation and carbon emission calculation. Construct a green certificate trading model based on the green certificate trading mechanism. Consider the benign coupling between the carbon trading market and the green certificate trading market, and construct a carbon-green certificate coupling model. Step 2: Construct a regional power grid day-ahead dispatch optimization model based on deep reinforcement learning algorithm. The actual dispatch process of the regional power grid involves four stages: grid operation status collection, dispatch instruction formulation, dispatch instruction execution, and dispatch instruction feedback. Deep reinforcement learning involves four stages: observation status, action selection, action execution, and reward feedback. Introduce the deep reinforcement learning framework, correspond the four stages of the regional power grid optimization problem to the four stages of the deep reinforcement learning framework, and construct a regional power grid dispatch framework based on deep reinforcement learning. Step 3: Introduce the core idea of ​​the SA algorithm into the TD3 algorithm, propose the SA-TD3 algorithm, and use the SA-TD3 algorithm to solve the regional power grid day-ahead dispatch optimization model and obtain the day-ahead dispatch plan.

2. The method for regional power grid day-ahead dispatching based on reinforcement learning and carbon-green certificate coupling as claimed in claim 1, characterized in that: The step 1 is as follows: Step 1.1: Construct a carbon trading model and use the baseline method to calculate the allocation formula of carbon quotas as follows: In the formula, The initial carbon emission quota obtained for thermal power unit i, τ * is the carbon emission baseline per unit of electricity of power generation enterprises formulated in the industry, T is the number of time periods in the entire dispatch cycle, ΔT is the dispatch time interval, and there are T = 24h / ΔT decision periods in the whole day, P i,t is the output of thermal power unit i at time t; Taking into account the relationship between the carbon emissions of thermal power generating units and the types of fuel and fuel consumption, the carbon emissions of thermal power generating units i under conventional output are: In the formula, is the unit fuel consumption carbon emission of thermal power unit i, a i , b i 、c i is the unit power generation burn-up coefficient of thermal power unit i in normal operation, P i max , P i min are the upper and lower limits of the conventional output of thermal power unit i, respectively; in addition, considering that some thermal power units can perform deep peak regulation, the combustion efficiency is reduced and the additional fuel loss is increased in this state, which will cause additional carbon emissions. The carbon emissions are expressed as: In the formula, is the additional carbon emissions associated with the deep peak load regulation of the units, is the coefficient of additional fuel consumption generated when thermal power unit i performs deep peak regulation, P i a is the lower limit of the output of thermal power unit i under deep peak regulation; Carbon trading for power generation companies meets the following requirements: In the formula, is the carbon trading cost of thermal power unit i; CET is the price per unit carbon certificate transaction; is the actual carbon emission of thermal power unit i. The higher the carbon emission quota that the thermal power unit needs to purchase, the higher the transaction price of the corresponding interval carbon certificate purchased from the carbon certificate market, that is, In the formula, is the basic price of carbon certificate trading, θ CET represents the sensitivity coefficient of carbon price growth rate, d CET is the interval length of a single carbon emission level; therefore, the tiered carbon trading cost of thermal power units can be expressed as: Step 1.2: Construct a green certificate trading model, and formulate the initial allocation benchmark value and green certificate quota index of the power generation of new energy units according to the green power production tasks that need to be completed by the new energy generating units in the regional power grid; In the formula, is the green certificate quota index and green certificate quantity obtained by the new energy generator j in the regional power grid, α l is the green certificate quota coefficient, is the initial allocation benchmark value and actual power generation of the new energy generator j in period t, α f The coefficient for obtaining the number of green certificates per unit of electricity generated by new energy is used. A step-by-step pricing model is introduced into the green certificate transaction. The more carbon certificates a new energy unit needs to purchase, the higher the transaction price of green certificates in the corresponding range purchased from the green certificate market. The surplus green certificates are uniformly recovered by the government or the carbon certificate market at a basic price. The model satisfies: In the formula, is the green certificate transaction cost of new energy unit j, is the basic price of green certificate transaction, d GCT is the interval length of a single green certificate trading level, θ GCT is the growth coefficient of green certificate price; Step 1.3: Construct a carbon-green certificate coupling trading model. The emission reduction of new energy is the difference between the carbon emissions generated by new energy power generation and the carbon emission standards set by the industry. Wind power and photovoltaic power generation can be regarded as zero carbon emissions. The emission reduction of new energy power generation is: In the formula, The carbon emissions reduced by new energy generators j, The standard carbon emissions set for the industry, is the carbon emission generated by the new energy power generation enterprise j, which is taken as zero; therefore, the green certificate obtained by the new energy represents the carbon emission rights: In the formula, β is the emission reduction represented by a unit green certificate. Therefore, according to the above relationship, the conversion between green certificates and carbon emission rights is as follows: In the formula, Carbon emission rights converted from green certificates purchased for thermal power unit i; is the number of green certificates purchased by thermal power unit i from the green certificate market; therefore, the carbon trading and green certificate trading of thermal power unit i under the carbon-green certificate coupling mechanism can be expressed as: In the formula, is the cost of purchasing green certificates from the green certificate market for thermal power unit i.

3. The method for regional power grid day-ahead dispatching based on reinforcement learning and carbon-green certificate coupling as claimed in claim 1, characterized in that: The step 2 is as follows: By using the one-to-one correspondence between the four stages involved in the actual dispatching process of the regional power grid and the four stages involved in deep reinforcement learning, the regional power grid dispatching problem is modeled as a Markov decision process: the dispatching time of a whole day is divided into t periods, and the environmental state s t Corresponding to the operating status of the regional power grid at time t, action a t Corresponding to the dispatch plan of the regional power grid at time t+1, in the environmental state s t Next, perform action a t , the regional power grid shifts to the next environmental state s t+1 , which also corresponds to the action a of the regional power grid dispatch plan at time t+2 t+1 Finally, the agent will give feedback on the accumulated rewards, corresponding to the regional power grid's economic performance, carbon emissions and other indicators related to the dispatch plan C t Therefore, the Markov decision process modeling of the regional power grid can be expressed as: M e ={s t ,a t ,C t ,s t+1 } The source-side resources mainly include coal-fired power generation, gas-fired power generation, wind power generation, and photovoltaic power generation. The load-side resources mainly include shiftable loads, curtailable loads, and rigid loads. The set of coal-fired thermal power generating units is defined as Φ pg ={1,2,...,N g }, the set of gas-fired power generating units is Φ gas ={1,2,...,N gas }, the wind turbine generator set is Φ wind ={1,2,...,N w }, the photovoltaic power station set is Φ PV ={1,2,...,N v }, let the output value of coal-fired power generation unit i at time t be P i,t , the output value of gas-fired power unit i at time t is The output value of wind turbine i at time t is The output value of the photovoltaic power station at time t is The actual value of the load at time t is P t load , where the load that can be reduced is P t cut , the translatable load is P t sh。 4. The method for day-ahead dispatching of a regional power grid based on reinforcement learning and carbon-green certificate coupling as claimed in claim 3, characterized in that: The specific process of step 2 is as follows: Step 2.1: Construct the state value space. To construct the state space, first divide the scheduling time T in a whole day into t scheduling periods, that is, T = {1, 2, ..., t}; the system state at time t is s t , which includes time t, the unit output P of the coal-fired unit at time t i,t ,i∈{1,...,N g }、The startup time of coal-fired power unit i at time t The shutdown time of coal-fired power unit i at time k The output of gas-fired power unit i at time t The startup time of gas-fired power unit i at time t The downtime of gas-fired power unit i at time t The output of the wind turbine at time k Output of photovoltaic power station at time t The load power at time t is P t load , specifically described as follows: To establish the action space, record the agent’s current state s at time t t The response is a t , that is, the dispatch plan executed by the regional power grid at time t, and the output adjustment of the coal-fired unit in the current state is ΔP i,t ,i∈{1,...,N g }, the output adjustment of the gas unit in the current state is The action space is composed of the output adjustment of all coal-fired units and gas-fired units in the current state, which can be expressed as: Step 2.2: Construct the objective function and constraints. The day-ahead dispatch cost of the regional power grid is composed of the operation cost of thermal power units, the carbon trading cost of thermal power units, the green certificate trading cost of thermal power units, the green certificate trading cost of new energy units, the dispatch cost of reducible loads, the dispatch cost of shiftable loads, the penalty cost of abandoning new energy, and the penalty cost of load shedding. The specific expression of the cost model of the optimization objective is as follows: Operating costs of thermal power units: In the formula, C g is the operating cost of thermal power unit i, including the operating cost of coal-fired thermal power units and gas-fired units; Carbon trading costs of thermal power units: In the formula, C CET The carbon trading cost of thermal power units, including the carbon trading cost of coal-fired thermal power units and the carbon trading cost of gas-fired thermal power units; Green certificate transaction cost of thermal power units: In the formula, C buy-GCT The green certificate transaction cost of thermal power units, including the green certificate transaction cost of coal-fired thermal power units and the green certificate transaction cost of gas-fired thermal power units; Green certificate transaction cost for new energy units: In the formula, C GCT The green certificate transaction cost of new energy units, including the green certificate transaction cost of wind turbines and the green certificate transaction cost of photovoltaic power stations; The dispatching cost of load can be reduced: Constructing a stepped load reduction compensation price model: In the formula, is the compensation price for the unit reduction of load at time t, P t cut is the amount of power that can be reduced, P t cut,max , P t cut,min The maximum and minimum reduction amounts that can be accepted for load reduction are: is the growth coefficient function related to the reduction amount, is the initial compensation price for the curtailable load; the curtailment compensation cost for the curtailable load response reduction satisfies: In the formula, C pay,cut Compensation costs for reductions in load that can be curtailed; In addition, the load that can be reduced can also provide load-side spare capacity, which must meet the upper limit constraint of spare capacity: In the formula, The load-side spare capacity and capacity upper limit provided for the load that can be curtailed; within the spare capacity constraint range, the spare cost brought by providing spare capacity to respond to the grid dispatching requirements of the curtailed load meets the following requirements: In the formula, C re,cut The reserve cost caused by providing load-side reserve capacity for curtailable loads, the price of providing a unit of reserve capacity for curtailable load; Therefore, the dispatch cost of load reduction includes the reduction compensation cost and the reserve cost, which satisfies: C cut =C pay,cut +C re,cut In the formula, C cut The dispatch cost for the load that can be reduced; Dispatch cost of shiftable loads: The dispatch model of the shiftable load satisfies: Where, t sh is the start time of the movable load after scheduling, t sh- ,t sh+ P is the earliest and latest operation start time that the shiftable load is willing to respond to the grid demand, t sh is the power consumption of the movable load in period t after scheduling, T sh is the continuous operation time of the translatable load, After the shiftable load is dispatched, sh +τ period of power consumption, before dispatching at t sh* +Power consumption during the τ period, t sh* ,t sh It is the start time of the movable load before and after scheduling; Acceptable scheduling time for shiftable loads [t sh- ,t sh+ ]satisfy: In the formula, ρ sh is the unit compensation price of the translatable load, β sh (ρ sh ) is the acceptable translation time and compensation price of the movable load ρ sh Related related functions, is the maximum number of periods during which the load can be translated; The function of the shiftable load compensation price satisfies: In the formula, ρ sh- and ρ sh+ are the minimum compensation price and maximum compensation price of the unit load that can be translated when the dispatch agreement is signed, Δρ sh is the sensitivity coefficient related to the compensation price; Therefore, the dispatch cost of the shiftable load satisfies: In the formula, C sh is the dispatching cost of the movable load; τ sh is a Boolean variable, τ sh =1 means the translatable load is translated, τ sh =0 means that the translatable load does not translate; New energy disposal penalty costs: In the formula, C cnew is the penalty cost for abandoning new energy, λ cw , cv is the penalty coefficient for unit abandoned electricity of wind power and photovoltaic power, is the abandoned power of wind turbine i in period t, is the abandoned electricity of photovoltaic power station j at time t; Load shedding penalty cost: In the formula, C cl is the penalty cost of load shedding in the regional power grid, P t cl is the load shedding power of the regional power grid during period t; Therefore, the optimization goal of the regional power grid day-ahead dispatch is to minimize the day-ahead dispatch cost: my C day =min(C g +C CET +C buy-GCT +C GCT +C cut +C sh +C cnew +C cl ) In order to maintain safe and stable operation, the regional power grid needs to meet the constraints of line transmission capacity flow in addition to the constraints related to the operation of the thermal power units themselves, the constraints related to load reduction, and the constraints related to load shifting. The transmission capacity constraints are satisfied: Where, T l,i , T l,j , T l,k , T l,m and T l,n is the power transfer allocation coefficient; N l is the number of node loads in the regional power grid; is the load power of node m at time t; F l max is the upper limit of the power flow of line l.

5. The method for regional power grid day-ahead dispatching based on reinforcement learning and carbon-green certificate coupling as claimed in claim 1, characterized in that: The step 3 specifically includes: Step 3.1: Introduce a regularization strategy for smoothing the target strategy. After smoothing, action a is expressed as: a←μ'(s|θ μ' )+η η←clip(N(0,σ),-c,c) In the formula, μ'(s|θ μ' ) is the policy function, s is the given state, θ μ' The parameters of the target network, η is the noise added during action selection, clip(·) is the cutoff function that limits the maximum and minimum values ​​of the noise, and N(0,σ) is a normal distribution function with a mean of 0 and a variance of σ. Step 3.2: Introduce the core idea of ​​the SA algorithm into the TD3 algorithm. The network update function of the TD3 algorithm is specifically expressed as the loss function as follows: y=-C+γmin(Q'1(s,a),Q'2(s,a)),γ∈[0,1] In the formula, γ is the discount factor, N is the total number of samples used, s, a and C are the sampled batch states, actions and running states respectively. is the loss function of the Critic network i. The TD3 algorithm reduces the frequency of Actor network updates to reduce the amount of erroneous information learned. The update gradient of the Actor network is satisfy: The core idea of ​​the SA algorithm is introduced into the TD3 algorithm. The first Critic network is used as the exploration network of the initial solution in the SA algorithm, and the second Critic network is used as the exploration network of the perturbation term in the SA algorithm. The network update formula of the SA-TD3 algorithm is specifically expressed as follows: The SA-TD3 algorithm is used to solve the regional power grid day-ahead dispatch optimization model and obtain the day-ahead dispatch plan. The specific steps are as follows: Step 3.3: Initialize the policy network π φ , Valuation Network Target Network Experience pool β, set learning rate α, discount factor γ, delayed update step number d, soft update coefficient τ and maximum learning step number max episode, learning step number count = 0, initial annealing temperature T k , and the state space s constructed above t As algorithm input, it is input into the Actor network; Step 3.4: According to the output of the Actor network, introduce Gaussian noise and generate the current action a under the condition of satisfying relevant constraints t =μ(s t |θ μ )+η, and schedule a t Send to the power grid simulation environment to calculate the next state s t+1 And the current running cost C t That is, considering the carbon and green certificate transaction costs of power generation units under the carbon-green certificate coupling transaction mechanism and other related day-ahead scheduling costs, the sample data (s t ,a t ,C t ,s t+1 ) is stored in the experience pool, and t=t+1 is used to update the Critic network Let count = count + 1. If the number of learning steps count is divisible by the number of delayed update steps d, then based on the SA algorithm strategy, update the Actor network θ μ , soft update target network Minimize the day-ahead scheduling cost and repeat this step, otherwise repeat this step directly until the maximum number of learning steps is reached, and output the optimized scheduling plan, that is, the output of each unit.

6. The method for regional power grid day-ahead dispatching based on reinforcement learning and carbon-green certificate coupling as claimed in claim 5, characterized in that: The specific steps of the algorithm flow of the SA algorithm are as follows: Step 1: Initialize k = 0, based on the initial annealing temperature T k Find the initial random solution x0, at this time x k = x0; Step 2: For a feasible solution x k Add a random noise Δx k Get x' k =x k +Δx k , respectively calculate the evaluation function value f(x k ) and f(x' k ), if f(x k )<f(x' k ), then x' k =x k ; Otherwise, when e-(f(x' k )-f(x k ) / T k )>rand() k =x k +Δx k Repeat step 2 k Second-rate; Step 3: Calculate T k+1 =Q SA T k , where Q SA is the annealing factor in the simulated annealing algorithm, let k=k+1, and determine whether the algorithm termination condition is met. If so, stop; if not, return to step 2.

Citation Information

Patent Citations

  • Deep reinforcement learning-based day-ahead-intra-day combined dispatching method for regional power grid

    CN115441437A

  • Hydrogen-containing comprehensive energy scheduling method and system fusing green certificate and carbon transaction

    CN116542485A

  • Micro-grid energy optimization method and system, electronic equipment and medium

    CN116885799A

  • Electricity-carbon-green evidence multi-market equilibrium analysis method based on multi-agent reinforcement learning

    CN117314040A

  • Reinforcement learning green certificate transaction method based on privacy protection in energy internet

    CN118780914A