Day-ahead dispatch method for regional power grid based on reinforcement learning and carbon-green certificate coupling
By constructing a carbon-green certificate coupling model and introducing the SA-TD3 algorithm, the problem of low computational efficiency in traditional regional power grid dispatching was solved, the benign coupling of carbon trading and green certificate trading was achieved, and the low-carbon economic operation and new energy absorption capacity of the regional power grid were improved.
Patent Information
- Application Number
- CN202510109950.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Traditional optimization algorithms have low computational efficiency in regional power grid scheduling and are unable to effectively handle complex and highly random scheduling problems. In addition, the parallel trading of carbon trading and green certificate trading lacks flexibility, which affects the low-carbon economic operation of the regional power grid.
Based on the regional power grid day-ahead dispatch method of deep reinforcement learning and carbon-green certificate coupling, a carbon trading and green certificate trading model is constructed. Combined with deep learning theory, the SA-TD3 algorithm is introduced to optimize scheduling, realize the benign coupling and flexible conversion of the carbon-green certificate market, and improve the flexibility and efficiency of scheduling.
It has improved the ability to handle large-scale random problems, promoted the low-carbon economic operation of regional power grids, improved the level of new energy consumption, reduced the maximum load shedding, and improved the economy and stability of the power grid.
Smart Images

Figure CN119944653B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power systems, and more specifically, relates to a regional power grid day-ahead scheduling method based on reinforcement learning and carbon-green certificate coupling. Background Art
[0002] With the exponential growth of dispatchable objects in regional power grid dispatching and the huge amount of monitoring information, traditional optimization algorithms may be difficult to solve due to the problem of low computational efficiency. However, reinforcement learning can better solve the optimization problems corresponding to the interaction between intelligent agents and the environment. With the relevant research combining the perception capabilities of deep learning with reinforcement learning, the learning ability and application scope of reinforcement learning have been further improved. In recent years, deep reinforcement learning methods have attracted increasing attention in the field of power system dispatching.
[0003] In summary, the present invention proposes a method for day-ahead dispatch of regional power grids based on reinforcement learning and carbon-green certificate coupling. This method first constructs a regional power grid dispatch framework based on deep reinforcement learning theory, taking into account the carbon-green certificate coupling trading mechanism. Then, the state and action space of the intelligent agent are constructed in combination with deep learning theory, and the day-ahead dispatch optimization objectives and constraints of the regional power grid are given. Finally, the core idea of the SA algorithm is introduced into the existing TD3 algorithm, and a method for day-ahead dispatch optimization of regional power grids based on the carbon-green certificate coupling mechanism is proposed. The established regional power grid day-ahead dispatch optimization model taking into account the carbon-green certificate coupling mechanism is solved. Summary of the Invention
[0004] In response to the problems existing in the prior art, the present invention proposes a regional power grid day-ahead scheduling method based on reinforcement learning and carbon-green certificate coupling. This method utilizes the characteristics of deep reinforcement learning and processes states and actions in random and complex environments through neural networks, which can effectively improve the processing capabilities of large-scale random problems. It also considers the benign coupling of the carbon trading market and the green certificate trading market to achieve flexible conversion between green certificates and carbon rights, and promote the low-carbon economic operation of the regional power grid.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The regional power grid day-ahead dispatch method based on reinforcement learning and carbon-green certificate coupling includes the following steps:
[0007] Step 1: Construct models related to carbon trading and green certificate trading. Construct a tiered carbon trading model based on carbon quota allocation and carbon emission calculations, a green certificate trading model based on the green certificate trading mechanism, and a carbon-green certificate coupling model based on the benign coupling between carbon trading and green certificate trading.
[0008] Step 2: Construct a regional power grid day-ahead dispatch optimization model based on a deep reinforcement learning algorithm. This model introduces a deep reinforcement learning framework and maps the four phases of the regional power grid's actual dispatch process—collecting grid operating status, formulating dispatch instructions, executing dispatch instructions, and providing dispatch instruction feedback—to the four phases of deep reinforcement learning—observing status, selecting actions, executing actions, and providing reward feedback. This framework for regional power grid dispatch based on deep reinforcement learning is then constructed.
[0009] Step 3: Introduce the core idea of the SA algorithm (Simulated Annealing algorithm, SA) into the Twin Delayed Deep Deterministic Policy Gradient algorithm TD3 (TD3), propose the SA-TD3 algorithm, and use the SA-TD3 algorithm to solve the regional power grid day-ahead dispatch optimization model and obtain the day-ahead dispatch plan.
[0010] This technical solution is further optimized, and step 1 specifically includes the construction of a carbon trading model, a green certificate trading model, and a carbon-green certificate coupling model:
[0011] Step 1.1: Build a carbon trading model.
[0012] The formula for calculating carbon quota allocation using the baseline method is as follows:
[0013]
[0014] Where, The initial carbon emission quota obtained for thermal power unit i; τ * is the carbon emission baseline per unit of electricity for power generation enterprises established within the industry; T is the number of time periods in the entire dispatch cycle, ΔT is the dispatch time interval, and there are a total of T = 24h / ΔT decision periods in the whole day, P i,t is the output of thermal power unit i at time t;
[0015] Taking into account the relationship between the carbon emissions of thermal power generating units and the type and fuel consumption, the carbon emissions of thermal power unit i under conventional output are:
[0016]
[0017] Where, is the unit fuel consumption carbon emission of thermal power unit i, a i 、b i 、c i is the unit power generation burn-up coefficient of thermal power unit i in normal operation state, P i max 、P i minare the upper and lower limits of the conventional output of thermal power unit i, respectively. In addition, the present invention takes into account that some thermal power units can perform deep peak regulation. In this state, the combustion efficiency is reduced and additional fuel loss is increased, which will cause additional carbon emissions. The carbon emissions are expressed as:
[0018]
[0019] Where, is the additional carbon emissions associated with the deep peak regulation of the unit, is the coefficient of additional fuel consumption generated when thermal power unit i performs deep peak regulation, P i a is the lower limit of the output of thermal power unit i in the deep peak regulation state;
[0020] Carbon trading for power generation companies meets the following requirements:
[0021]
[0022] Where, is the carbon trading cost of thermal power unit i; CET is the price per carbon certificate transaction; is the actual carbon emissions of thermal power unit i. Traditional carbon trading adopts a unified pricing principle, and the carbon trading price is uniform and fixed. Under this method, power generation companies are not very motivated to reduce emissions. When thermal power units produce more carbon emissions, they should be punished more severely to promote their emission reduction efforts. The higher the carbon emission quota that a thermal power unit needs to purchase, the higher the transaction price of the corresponding range of carbon certificates purchased from the carbon certificate market, that is,
[0023]
[0024] Where, is the basic price of carbon certificate transaction; θ CET represents the sensitivity coefficient of carbon price growth rate; d CET is the interval length of a single carbon emission level. Therefore, the tiered carbon trading cost of thermal power units can be expressed as:
[0025]
[0026] Step 1.2: Build a green certificate trading model.
[0027] The new energy generating units in the regional power grid need to complete a certain proportion of green electricity production, that is, to formulate the initial allocation benchmark value of the power generation of the new energy generating units and issue a certain green certificate quota index accordingly. The green certificate quota index and the number of green certificates obtained by the new energy generating units in the regional power grid must meet the following requirements:
[0028]
[0029] Where, is the green certificate quota index and green certificate quantity obtained by the new energy generator j in the regional power grid, α l is the green certificate quota coefficient, is the initial allocation benchmark value and actual power generation of the new energy generator set j in period t, α f The coefficient for obtaining the number of green certificates per unit of electricity generated by new energy. The green certificate transaction introduces a tiered pricing model. The more carbon certificates a new energy unit needs to purchase, the higher the transaction price of green certificates in the corresponding range purchased from the green certificate market. The excess green certificates are uniformly recovered by the government or the carbon certificate market at the base price. The model satisfies:
[0030]
[0031] Where, is the green certificate transaction cost of new energy unit j, is the basic price of green certificate transaction, d GCT is the interval length of a single green certificate trading level, θ GCT is the growth coefficient of green certificate price.
[0032] Step 1.3: Construct a carbon-green certificate coupling trading model.
[0033] The emission reduction of new energy is the difference between the carbon emissions generated by new energy power generation and the carbon emission standards set by the industry. Wind power and photovoltaic power generation can be regarded as zero carbon emissions. The emission reduction of new energy power generation is:
[0034]
[0035] Where, The amount of carbon emissions reduced by the new energy generator set j; Standard carbon emissions set for the industry; The carbon emissions generated by the new energy power generation enterprise j are taken as zero in this invention. Therefore, the carbon emission rights represented by the green certificate obtained by the new energy are:
[0036]
[0037] In the formula, β is the emission reduction represented by a unit green certificate. Therefore, based on the above relationship, the conversion between green certificates and carbon emission rights is as follows:
[0038]
[0039] Where, Carbon emission rights converted from green certificates purchased for thermal power unit i; is the number of green certificates purchased by thermal power unit i from the green certificate market; therefore, the carbon trading and green certificate trading of thermal power unit i under the carbon-green certificate coupling mechanism can be expressed as:
[0040]
[0041] Where, is the cost of purchasing green certificates from the green certificate market for thermal power unit i.
[0042] This technical solution is further optimized, and the step 2 specifically includes:
[0043] By using the one-to-one correspondence between the four stages involved in the actual dispatching process of the regional power grid and the four stages involved in deep reinforcement learning, the regional power grid dispatching problem is modeled as a Markov decision process: the dispatching time of a whole day is divided into t periods, and the environment state s is set to t Corresponding to the operating status of the regional power grid at time t, action a t Corresponding to the dispatch plan of the regional power grid at time t+1, in the environmental state s t Next, perform action a t , the regional power grid shifts to the next environmental state s t+1 , which also corresponds to the action a of the regional power grid dispatch plan at time t+2 t+1 Finally, the intelligent agent will give feedback on the accumulated rewards, which corresponds to the regional power grid's response to the economic efficiency, carbon emissions and other indicators of the dispatch plan. t Therefore, the Markov decision process modeling of the regional power grid can be expressed as:
[0044] M e ={s t ,a t ,C t ,s t+1}
[0045] The source load resources considered in this invention mainly include coal-fired power generation, gas-fired power generation, wind power generation, and photovoltaic power generation. The load-side resources mainly consider the movable load, the curtailable load, and the rigid load. The set of coal-fired thermal power generation units is defined as Φ pg ={1,2,...,N g}, the set of gas-fired power generation units is Φ gas ={1,2,...,N gas}, the wind turbine generator set is Φ wind ={1,2,...,N w}, the photovoltaic power station set is Φ PV ={1,2,...,N v The output value of coal-fired power generation unit i at time t is P i,t , the output value of gas-fired power unit i at time t is The output value of wind turbine i at time t is The output value of the photovoltaic power station at time t is The actual value of the load at time t is P t load , where the load that can be reduced is P t cut , the translatable load is P t sh .
[0046] Step 2.1: Construct the state value space.
[0047] In order to establish the state space, first divide the scheduling time T of the whole day into t scheduling periods, that is, T = {1, 2, ..., t}; the system state at time t is s t , which includes time t, the unit output P of the coal-fired unit at time t i,t ,i∈{1,...,N g}、The operating time of coal-fired power unit i at time t The shutdown time of coal-fired power unit i at time k The output of gas-fired power unit i at time t The operating time of gas-fired power unit i at time t The downtime of gas-fired power unit i at time t The output of the wind turbine at time k The output of the photovoltaic power station at time t The power P of the load at time t t load , specifically expressed as follows:
[0048]
[0049] To establish the action space, record the agent’s current state s at time t t The response is a t , that is, the dispatch plan executed by the regional power grid at time t, and the output adjustment of the coal-fired unit in the current state is ΔP i,t ,i∈{1,...,N g}, the output adjustment of the gas unit in the current state is The action space is composed of the output adjustment of all coal-fired units and gas-fired units in the current state, which can be expressed as:
[0050]
[0051] Step 2.2: Construct the objective function and constraints.
[0052] The day-ahead dispatch cost of the regional power grid is composed of the operating cost of thermal power units, the carbon trading cost of thermal power units, the green certificate trading cost of thermal power units, the green certificate trading cost of new energy units, the dispatch cost of curtailable loads, the dispatch cost of shiftable loads, the penalty cost for abandoning new energy resources, and the penalty cost for load shedding. The cost model for the optimization objective is specifically expressed as follows:
[0053] Operating costs of thermal power units:
[0054]
[0055] Where C g is the operating cost of thermal power unit i, including the operating cost of coal-fired thermal power units and gas-fired units.
[0056] Carbon trading costs of thermal power units:
[0057]
[0058] Where C CET It is the carbon trading cost of thermal power units, including the carbon trading cost of coal-fired thermal power units and the carbon trading cost of gas-fired thermal power units.
[0059] Green certificate transaction costs for thermal power units:
[0060]
[0061] Where C buy-GCT It is the green certificate transaction cost of thermal power units, including the green certificate transaction cost of coal-fired thermal power units and the green certificate transaction cost of gas-fired thermal power units.
[0062] Green certificate transaction costs for new energy units:
[0063]
[0064] Where C GCT It is the green certificate transaction cost of new energy units, including the green certificate transaction cost of wind turbine units and the green certificate transaction cost of photovoltaic power stations.
[0065] Dispatch costs that can reduce loads:
[0066] Constructing a tiered load reduction compensation price model:
[0067]
[0068] Where, P is the compensation price for the unit amount of electricity that can be reduced at time t, t cut is the amount of power that can be reduced, P t cut,max 、Pt cut,min The maximum and minimum reduction amounts that can be accepted for load reduction are: is the growth coefficient function related to the reduction amount, is the initial compensation price for curtailable load; the curtailment compensation cost of curtailable load response meets the following requirements:
[0069]
[0070] Where C pay,cut The cost of compensating for the reduction of load that can be reduced.
[0071] In addition, load reduction can also provide load-side spare capacity, making load-side resources more flexible. The spare capacity that can be reduced is determined when formulating the day-ahead plan, and the upper limit constraint of the spare capacity must be met:
[0072]
[0073] Where, The load-side spare capacity and capacity limit provided for the load that can be curtailed. Within the spare capacity constraint, the spare cost brought by providing spare capacity to respond to grid dispatch requirements for the load that can be curtailed meets the following requirements:
[0074]
[0075] Where C re,cut The backup cost caused by providing load-side backup capacity for curtailable loads, The price of providing a unit of reserve capacity for curtailable load.
[0076] Therefore, the dispatch cost of load reduction includes the reduction compensation cost and the reserve cost, which satisfies:
[0077] C cut =C pay,cut +C re,cut
[0078] Where C cut The dispatch cost of the load can be reduced.
[0079] Dispatch cost of shiftable loads:
[0080] The scheduling model of the shiftable load satisfies:
[0081]
[0082] Where, t sh is the start time of the movable load after scheduling, t sh- , t sh+P is the earliest and latest operation start time that the shiftable load is willing to respond to the grid demand, t sh T is the power consumption of the movable load in period t after scheduling, sh is the continuous operation time of the translatable load, After the load is dispatched, the sh +τ period of power consumption, before dispatch at t sh* +Power consumption during the τ period, t sh* , t sh The start time of the movable load before and after scheduling;
[0083] Acceptable scheduling time for shiftable loads [t sh- ,t sh+ ]satisfy:
[0084]
[0085] Where, ρ sh is the unit compensation price of the translatable load, β sh (ρ sh ) is the acceptable translation time and compensation price of the translatable load ρ sh Related related functions, The maximum number of translation periods for the translation load.
[0086] The function of the shiftable load compensation price satisfies:
[0087]
[0088] Where, ρ sh- and ρ sh+ are the minimum compensation price and maximum compensation price of the unit load that can be translated when the dispatch agreement is signed, Δρ sh is the sensitivity coefficient related to the compensation price.
[0089] Therefore, the dispatch cost of the shiftable load satisfies:
[0090]
[0091] Where C sh is the dispatching cost of the movable load; τ sh is a Boolean variable, τ sh =1 means that the translatable load moves, τ sh =0 means that the translatable load does not translate.
[0092] New energy disposal penalty costs:
[0093]
[0094] Where C cnew is the penalty cost for abandoning new energy, λ cw ,λ cv is the penalty coefficient for unit abandoned electricity of wind power and photovoltaic power, is the abandoned power of wind turbine i in period t, is the abandoned electricity of photovoltaic power station j at time t.
[0095] Load shedding penalty cost:
[0096]
[0097] Where C cl is the load shedding penalty cost of the regional power grid, P t cl is the load shedding amount of the regional power grid during period t.
[0098] Therefore, the optimization goal of the regional power grid's day-ahead dispatch is to minimize the day-ahead dispatch cost:
[0099] minC day =min(C g +C CET +C buy-GCT +C GCT +C cut +C sh +C cnew +C cl )
[0100] In order to maintain safe and stable operation, the regional power grid needs to meet the line transmission capacity flow constraints in addition to the relevant constraints on the operation of the thermal power units themselves, the constraints on the load that can be reduced, and the constraints on the load that can be shifted.
[0101] The transmission capacity constraints are satisfied:
[0102]
[0103] Where, T l,i 、T l,j 、T l,k 、T l,m and T l,n is the power transmission distribution coefficient; N l is the number of node loads in the regional power grid; is the load power of node m at time t; F l max is the upper limit of the power flow of line l.
[0104] This technical solution is further optimized, and step 3 specifically includes:
[0105] Step 3.1: Introduce a regularization strategy for smoothing the target strategy.
[0106] TD3 smoothes the Q function to make the probability distribution of action selection more balanced, thereby reducing the deviation when estimating action values and improving the performance of the algorithm. After smoothing, action a is expressed as:
[0107] a←μ'(s|θ μ' )+η
[0108] η←clip(N(0,σ),-c,c)
[0109] Where, μ'(s|θ μ' ) is the policy function, s is the given state, θ μ' The parameters of the target network, η is the noise added during action selection, clip(·) is the truncation function that limits the maximum and minimum values of the noise, and N(0,σ) is a normal distribution function with a mean of 0 and a variance of σ.
[0110] Step 3.2: Introduce the core idea of the SA algorithm into the TD3 algorithm.
[0111] The TD3 algorithm introduces two sets of Critic networks and selects the smaller Q value of the two networks to update the Actor network, thereby reducing the impact of Q value deviation. The network update function of the TD3 algorithm is specifically expressed as a loss function as follows:
[0112] y=-C+γmin(Q'1(s,a),Q'2(s,a)),γ∈[0,1]
[0113]
[0114] In the formula, γ is the discount factor, N is the total number of samples used, s, a and C are the sampled states, actions and running states respectively. is the loss function of Critic network i. The TD3 algorithm reduces the frequency of Actor network updates, thereby reducing the amount of learned error information and making the learned policy network more stable. The update gradient of the Actor network is satisfy:
[0115]
[0116] The specific steps of the SA algorithm are as follows:
[0117] Step 1: Initialize k = 0, based on the initial annealing temperature T k Find the initial random solution x0, at this time x k =x0;
[0118] Step 2: For a feasible solution xk Add a random noise Δx k Get x' k =x k +Δx k , calculate the evaluation function value f(x k ) and f(x' k ). If f(x k )<f(x' k ), then x' k =x k Otherwise, when Time x' k =x k +Δx k Repeat step 2 k Second-rate;
[0119] Step 3: Calculate T k+1 =Q SA T k , where Q SA is the annealing factor in the simulated annealing algorithm. Let k = k + 1, and determine whether the algorithm termination condition is met. If so, stop; if not, return to step 2;
[0120] The core idea of the SA algorithm is introduced into the TD3 algorithm. The first critic network is used as the exploration network for the initial solution in the SA algorithm, and the second critic network is used as the exploration network for the perturbation term in the SA algorithm. Therefore, the network update formula of the SA-TD3 algorithm is specifically expressed as follows:
[0121]
[0122] The SA-TD3 algorithm is used to solve the regional power grid's day-ahead dispatch optimization model. The specific steps for obtaining the day-ahead dispatch plan are as follows:
[0123] Step 3.3: Initialize the policy network π φ , valuation network Target Network Experience pool β, set learning rate α, discount factor γ, delayed update step number d, soft update coefficient τ and maximum learning step number maxepisode, learning step number count = 0, initial annealing temperature T k , and the state space s constructed above t As algorithm input, input to the Actor network;
[0124] Step 3.4: According to the output of the Actor network, introduce Gaussian noise and generate the current action a under the condition of satisfying relevant constraints t =μ(s t |θ μ)+η, and schedule the plan a t Send to the power grid simulation environment to calculate the next state s t+1 And the current running cost C t That is, considering the carbon and green certificate transaction costs of power generation units under the carbon-green certificate coupling transaction mechanism and other related day-ahead scheduling costs, the sample data (s t ,a t ,C t ,s t+1 ) is stored in the experience pool, and t=t+1 is used to update the Critic network Let count = count + 1. If the number of learning steps count is divisible by the number of delayed update steps d, then based on the SA algorithm strategy, update the Actor network θ μ , soft update target network Minimize the day-ahead scheduling cost and repeat this step. Otherwise, repeat this step directly until the maximum number of learning steps is reached, and output the optimized scheduling plan, that is, the output of each unit.
[0125] Different from the prior art, the beneficial effects of the present invention are mainly manifested in:
[0126] (1) In the traditional model, carbon trading and green certificate trading within the regional power grid are carried out in parallel. The present invention allows the carbon trading market and the green certificate trading market to be benignly coupled to achieve flexible conversion between green certificates and carbon rights, and promote the low-carbon economic operation of the regional power grid.
[0127] (2) The proposed SA-TD3 algorithm improves the ability to escape from local optimality during training by introducing the core idea of simulated annealing, thereby improving the training speed of the TD3 algorithm and maintaining the inherent stability of the TD3 algorithm;
[0128] (3) The integration of gas-fired power units into the power grid further enhances the flexibility of source-side resources, improves the regulation capacity of the power grid, effectively improves the level of new energy consumption, reduces the proportion of maximum load shedding, and thus improves the economic efficiency of the entire regional power grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0129] Figure 1 This is a diagram of the carbon-green certificate trading coupling mechanism;
[0130] Figure 2 Schematic diagram of the carbon-green certificate coupling model;
[0131] Figure 3 This is a schematic diagram of regional power grid dispatch optimization based on the SA-TD3 algorithm;
[0132] Figure 4 Flowchart of the regional power grid dispatch optimization method based on SA-TD3 algorithm. DETAILED DESCRIPTION
[0133] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.
[0134] This paper proposes a method for day-ahead scheduling of regional power grids based on reinforcement learning and carbon-green certificate coupling. This method considers the benign coupling between the carbon trading market and the green certificate trading market, and constructs a carbon-green certificate coupling model based on this. This method makes day-ahead scheduling more compliant with actual needs, enables flexible conversion between green certificates and carbon rights, and promotes low-carbon economic operation of the regional power grid. The proposed SA-TD3 algorithm improves the training speed of the TD3 algorithm while maintaining its inherent stability.
[0135] See also Figure 1 Diagram of the carbon-green certificate trading coupling mechanism. The government is responsible for the issuance, approval, and recycling of carbon and green certificates. The green certificate market centrally trades green certificates for new energy generators. In addition to purchasing carbon emission allowances from the carbon certificate market, thermal power generators can also purchase green certificates from the green certificate market to offset their carbon emissions.
[0136] See for example Figure 2 As shown in the figure, the carbon-green certificate coupling model schematic includes the following steps:
[0137] Step 1: Construct a model related to carbon trading and green certificate trading.
[0138] A stepped carbon trading model is constructed based on carbon quota allocation and carbon emission calculation, and a green certificate trading model is constructed based on the green certificate trading mechanism. This allows the carbon trading market and the green certificate trading market to be benignly coupled, realizes flexible conversion between green certificates and carbon rights, and constructs a carbon-green certificate coupling model.
[0139] Step 1.1: Build a carbon trading model.
[0140] The formula for calculating carbon allowance allocation using the baseline method is as follows:
[0141]
[0142] Where, The initial carbon emission quota obtained for thermal power unit i; τ * is the carbon emission baseline per unit of electricity for power generation enterprises established within the industry; T is the number of time periods in the entire dispatch cycle, ΔT is the dispatch time interval, and there are a total of T = 24h / ΔT decision periods throughout the day; P i,t is the output of thermal power unit i at time t.
[0143] Taking into account the relationship between the carbon emissions of thermal power generation units and the type and fuel consumption, the carbon emissions of thermal power generation unit i under conventional output are
[0144]
[0145] Where, is the unit fuel consumption carbon emission of thermal power unit i, a i 、b i 、c i is the unit power generation burn-up coefficient of thermal power unit i in normal operation state, P i max 、P i min are the upper and lower limits of the conventional output of thermal power unit i, respectively. In addition, the present invention takes into account that some thermal power units can perform deep peak regulation. In this state, the combustion efficiency is reduced and additional fuel loss is increased, which will cause additional carbon emissions. The carbon emissions are expressed as:
[0146]
[0147] Where, is the additional carbon emissions associated with the deep peak regulation of the unit, is the coefficient of additional fuel consumption generated when thermal power unit i performs deep peak regulation, P i a is the lower limit of the output of thermal power unit i in the deep peak regulation state.
[0148] Carbon trading for power generation companies meets the following requirements:
[0149]
[0150] Where, is the carbon trading cost of thermal power unit i; CET is the price per carbon certificate transaction; is the actual carbon emissions of thermal power unit i. Traditional carbon trading adopts a unified pricing principle, and the carbon trading price is uniform and fixed. Under this method, power generation companies are not very motivated to reduce emissions. When thermal power units produce more carbon emissions, they should be punished more severely to promote their emission reduction efforts. The higher the carbon emission quota that a thermal power unit needs to purchase, the higher the transaction price of the corresponding range of carbon certificates purchased from the carbon certificate market, that is,
[0151]
[0152] Where, is the basic price of carbon certificate transaction; θ CET represents the sensitivity coefficient of carbon price growth rate; d CET is the interval length of a single carbon emission level. Therefore, the tiered carbon trading cost of thermal power units can be expressed as:
[0153]
[0154] Step 1.2: Build a green certificate trading model.
[0155] New energy generating units in the regional power grid need to complete a certain proportion of green electricity production, that is, to formulate an initial allocation benchmark value for the power generation of new energy units and issue a certain green certificate quota indicator based on this.
[0156]
[0157] Where, is the green certificate quota index and green certificate quantity obtained by the new energy generator j in the regional power grid, α l is the green certificate quota coefficient, is the initial allocation benchmark value and actual power generation of the new energy generator set j in period t, α f The coefficient for obtaining the number of green certificates per unit of electricity generated by new energy. The green certificate transaction introduces a tiered pricing model. The more carbon certificates a new energy unit needs to purchase, the higher the transaction price of green certificates in the corresponding range purchased from the green certificate market. The excess green certificates are uniformly recovered by the government or the carbon certificate market at the base price. The model satisfies:
[0158]
[0159] Where, is the green certificate transaction cost of new energy unit j, is the basic price of green certificate transaction, d GCT is the interval length of a single green certificate trading level, θ GCT is the growth coefficient of green certificate price.
[0160] Step 1.3: Construct a carbon-green certificate coupling trading model.
[0161] The emission reduction of new energy is the difference between the carbon emissions generated by new energy power generation and the carbon emission standards set by the industry. Wind power and photovoltaic power generation can be regarded as zero carbon emissions. The emission reduction of new energy power generation is:
[0162]
[0163] Where, The amount of carbon emissions reduced by the new energy generator set j; Standard carbon emissions set for the industry; The carbon emissions generated by the new energy power generation enterprise j are taken as zero in this invention. Therefore, the carbon emission rights represented by the green certificate obtained by the new energy are:
[0164]
[0165] In the formula, β is the emission reduction represented by a unit green certificate. Therefore, based on the above relationship, the conversion between green certificates and carbon emission rights is as follows:
[0166]
[0167] Where, Carbon emission rights converted from green certificates purchased for thermal power unit i; is the number of green certificates purchased by thermal power unit i from the green certificate market. Therefore, under the carbon-green certificate coupling mechanism, the carbon trading and green certificate trading of thermal power unit i can be expressed as:
[0168]
[0169]
[0170] Where, is the cost of purchasing green certificates from the green certificate market for thermal power unit i.
[0171] See for example Figure 3 As shown in the figure, the regional power grid dispatch optimization diagram based on the SA-TD3 algorithm includes the following steps:
[0172] Step 2: Build a regional power grid day-ahead dispatch optimization model based on a deep reinforcement learning algorithm. The actual regional power grid dispatch process involves four stages: grid operating status collection, dispatch instruction formulation, dispatch instruction execution, and dispatch instruction feedback. Deep reinforcement learning, on the other hand, involves four stages: state observation, action selection, action execution, and reward feedback. By introducing a deep reinforcement learning framework and leveraging the one-to-one correspondence between regional power grid optimization problems and the deep reinforcement learning framework, we construct a regional power grid dispatch framework based on deep reinforcement learning.
[0173] By using the one-to-one correspondence between the four stages involved in the actual dispatching process of the regional power grid and the four stages involved in deep reinforcement learning, the regional power grid dispatching problem is modeled as a Markov decision process: the dispatching time of a whole day is divided into t periods, and the environment state s is set to t Corresponding to the operating status of the regional power grid at time t, action a t Corresponding to the dispatch plan of the regional power grid at time t+1, in the environmental state s t Next, perform action a t , the regional power grid shifts to the next environmental state s t+1 , which also corresponds to the action a of the regional power grid dispatch plan at time t+2 t+1 Finally, the intelligent agent will give feedback on the accumulated rewards, which corresponds to the regional power grid's response to the economic efficiency, carbon emissions and other indicators of the dispatch plan. t Therefore, the Markov decision process modeling of the regional power grid can be expressed as:
[0174] Me ={s t ,a t ,C t ,s t+1}
[0175] The source load resources considered in this invention mainly include coal-fired power generation, gas-fired power generation, wind power generation, and photovoltaic power generation. The load-side resources mainly consider the movable load, the curtailable load, and the rigid load. The set of coal-fired thermal power generation units is defined as Φ pg ={1,2,...,N g}, the set of gas-fired power generation units is Φ gas ={1,2,...,N gas}, the wind turbine generator set is Φ wind ={1,2,...,N w}, the photovoltaic power station set is Φ PV ={1,2,...,N v The output value of coal-fired power generation unit i at time t is P i,t , the output value of gas-fired power unit i at time t is The output value of wind turbine i at time t is The output value of the photovoltaic power station at time t is The actual value of the load at time t is P t load , where the load that can be reduced is P t cut , the translatable load is P t sh .
[0176] Step 2.1: Construct the state value space.
[0177] In order to establish the state space, first divide the scheduling time T of the whole day into t scheduling periods, that is, T = {1, 2, ..., t}; the system state at time t is s t , which includes time t, the unit output P of the coal-fired unit at time t i,t ,i∈{1,...,N g}、The operating time of coal-fired power unit i at time t The shutdown time of coal-fired power unit i at time k The output of gas-fired power unit i at time t The operating time of gas-fired power unit i at time t The downtime of gas-fired power unit i at time t The output of the wind turbine at time k The output of the photovoltaic power station at time t The power P of the load at time tt load , specifically expressed as follows:
[0178]
[0179] To establish the action space, record the agent’s current state s at time t t The response is a t , that is, the dispatch plan executed by the regional power grid at time t, and the output adjustment of the coal-fired unit in the current state is ΔP i,t ,i∈{1,...,N g}, the output adjustment of the gas unit in the current state is The action space is composed of the output adjustment of all coal-fired units and gas-fired units in the current state, which can be expressed as:
[0180]
[0181] Step 2.2: Construct the objective function and constraints.
[0182] The day-ahead dispatch cost of the regional power grid is composed of the operating cost of thermal power units, the carbon trading cost of thermal power units, the green certificate trading cost of thermal power units, the green certificate trading cost of new energy units, the dispatch cost of curtailable loads, the dispatch cost of shiftable loads, the penalty cost for abandoning new energy resources, and the penalty cost for load shedding. The cost model for the optimization objective is specifically expressed as follows:
[0183] Operating costs of thermal power units:
[0184]
[0185] Where C g is the operating cost of thermal power unit i, including the operating cost of coal-fired thermal power units and gas-fired units.
[0186] Carbon trading costs of thermal power units:
[0187]
[0188] Where C CET It is the carbon trading cost of thermal power units, including the carbon trading cost of coal-fired thermal power units and the carbon trading cost of gas-fired thermal power units.
[0189] Green certificate transaction costs for thermal power units:
[0190]
[0191] Where C buy-GCT It is the green certificate transaction cost of thermal power units, including the green certificate transaction cost of coal-fired thermal power units and the green certificate transaction cost of gas-fired thermal power units.
[0192] Green certificate transaction costs for new energy units:
[0193]
[0194] Where C GCT It is the green certificate transaction cost of new energy units, including the green certificate transaction cost of wind turbine units and the green certificate transaction cost of photovoltaic power stations.
[0195] Dispatch costs that can reduce loads:
[0196] Constructing a tiered load reduction compensation price model:
[0197]
[0198] Where, P is the compensation price for the unit amount of electricity that can be reduced at time t, t cut is the amount of power that can be reduced, P t cut,max 、P t cut,min The maximum and minimum reduction amounts that can be accepted for load reduction are: is the growth coefficient function related to the reduction amount, is the initial compensation price for curtailable load; the curtailment compensation cost of curtailable load response meets the following requirements:
[0199]
[0200] Where C pay,cut The cost of compensating for the reduction of load that can be reduced.
[0201] In addition, load reduction can also provide load-side spare capacity, making load-side resources more flexible. The spare capacity that can be reduced is determined when formulating the day-ahead plan, and the upper limit constraint of the spare capacity must be met:
[0202]
[0203] Where, The load-side spare capacity and capacity limit provided for the load that can be curtailed. Within the spare capacity constraint, the spare cost brought by providing spare capacity to respond to grid dispatch requirements for the load that can be curtailed meets the following requirements:
[0204]
[0205] Where C re,cut The backup cost caused by providing load-side backup capacity for curtailable loads, The price of providing a unit of reserve capacity for curtailable load.
[0206] Therefore, the dispatch cost of load reduction includes the reduction compensation cost and the reserve cost, which satisfies:
[0207] C cut =C pay,cut +C re,cut
[0208] Where C cut The dispatch cost of the load can be reduced.
[0209] Dispatch cost of shiftable loads:
[0210] The scheduling model of the shiftable load satisfies:
[0211]
[0212] Where, t sh is the start time of the movable load after scheduling, t sh- , t sh+ P is the earliest and latest operation start time that the shiftable load is willing to respond to the grid demand, t sh T is the power consumption of the movable load in period t after scheduling, sh is the continuous operation time of the translatable load, After the load is dispatched, the sh +τ period of power consumption, before dispatch at t sh* +Power consumption during the τ period, t sh* , t sh The start time of the movable load before and after scheduling;
[0213] Acceptable scheduling time for shiftable loads [t sh- ,t sh+ ]satisfy:
[0214]
[0215] Where, ρ sh is the unit compensation price of the translatable load, β sh (ρ sh ) is the acceptable translation time and compensation price of the translatable load ρ sh Related related functions, is the maximum number of translation periods for the translation load;
[0216] The function of the shiftable load compensation price satisfies:
[0217]
[0218] Where, ρ sh- and ρsh+ are the minimum compensation price and maximum compensation price of the unit load that can be translated when the dispatch agreement is signed, Δρ sh is the sensitivity coefficient related to the compensation price;
[0219] Therefore, the dispatch cost of the shiftable load satisfies:
[0220]
[0221] Where C sh is the dispatching cost of the movable load; τ sh is a Boolean variable, τ sh =1 means that the translatable load moves, τ sh =0 means that the translatable load does not translate.
[0222] New energy disposal penalty costs:
[0223]
[0224] Where C cnew is the penalty cost for abandoning new energy, λ cw ,λ cv is the penalty coefficient for unit abandoned electricity of wind power and photovoltaic power, is the abandoned power of wind turbine i in period t, is the abandoned electricity of photovoltaic power station j at time t.
[0225] Load shedding penalty cost:
[0226]
[0227] Where C cl is the load shedding penalty cost of the regional power grid, P t cl is the load shedding amount of the regional power grid during period t.
[0228] Therefore, the optimization goal of the regional power grid's day-ahead dispatch is to minimize the day-ahead dispatch cost:
[0229] minC day =min(C g +C CET +C buy-GCT +C GCT +C cut +C sh +C cnew +C cl )
[0230] In order to maintain safe and stable operation, the regional power grid needs to meet the line transmission capacity flow constraints in addition to the relevant constraints on the operation of the thermal power units themselves, the constraints on the load that can be reduced, and the constraints on the load that can be shifted.
[0231] The transmission capacity constraints are satisfied:
[0232]
[0233] Where, T l,i 、T l,j 、T l,k 、T l,m and T l,n is the power transmission distribution coefficient; N l is the number of node loads in the regional power grid; is the load power of node m at time t; F l max is the upper limit of the power flow of line l;
[0234] See for example Figure 4 As shown in the flowchart of the regional power grid dispatch optimization method based on the SA-TD3 algorithm, it includes the following steps:
[0235] Step 3: Introduce the core idea of the SA algorithm into the TD3 algorithm, propose the SA-TD3 algorithm, and use the SA-TD3 algorithm to solve the regional power grid day-ahead dispatch optimization model and obtain the day-ahead dispatch plan.
[0236] Step 3.1: Introduce a regularization strategy for smoothing the target strategy.
[0237] TD3 smoothes the Q function to make the probability distribution of action selection more balanced, thereby reducing the deviation value when estimating the action value and improving the performance of the algorithm. Expressed as:
[0238] a←μ'(s|θ μ' )+η
[0239] η←clip(N(0,σ),-c,c)
[0240] Where, μ'(s|θ μ' ) is the policy function, s is the given state, θ μ' The parameters of the target network, η is the noise added during action selection, clip(·) is the truncation function that limits the maximum and minimum values of the noise, and N(0,σ) is a normal distribution function with a mean of 0 and a variance of σ.
[0241] Step 3.2: Introduce the core idea of the SA algorithm into the TD3 algorithm.
[0242] In the DDPG algorithm, the maximum Q value is calculated using a single critic network, which can lead to the accumulation of excessively high Q values. The TD3 algorithm, on the other hand, introduces two critic networks and updates the actor network by selecting the smaller Q value from the two networks, thereby reducing the impact of Q value deviations. The TD3 algorithm's network update function, expressed as a loss function, is as follows:
[0243] y=-C+γmin(Q'1(s,a),Q'2(s,a)),γ∈[0,1]
[0244]
[0245] In the formula, γ is the discount factor, N is the total number of samples used, s, a and C are the sampled states, actions and running states respectively. is the loss function of Critic network i. If the Actor network is updated immediately every time the target value is updated, the large error in the updated target value will continue to accumulate, thus affecting the convergence of the final optimization result. Therefore, the TD3 algorithm reduces the frequency of Actor network updates to reduce the amount of learned error information, making the learned policy network more stable. The update gradient of the Actor network is satisfy:
[0246]
[0247] However, the TD3 algorithm's choice of a lower Q value update each time will undoubtedly reduce the algorithm's exploration efficiency and may lead to falling into a local optimum. The SA algorithm is a general probabilistic algorithm that is often used to seek the optimal solution to random optimization problems with a large solution space. Its name comes from the term annealing in the metallurgical field. The core idea of the SA algorithm draws on the annealing principle of metals: thermodynamics theory is applied to the field of statistics, and each solution within a search range is compared to a molecule in the air. Each solution within the search range has "energy" like a molecule in the air. The specific steps of the SA algorithm's algorithm flow are as follows:
[0248] Step 1: Initialize k = 0, based on the initial annealing temperature T k Find the initial random solution x0, at this time x k =x0;
[0249] Step 2: For a feasible solution x k Add a random noise Δx k Get x' k =x k +Δx k , respectively calculate the evaluation function value f(x k ) and f(x' k ). If f(xk )<f(x' k ), then x' k =x k Otherwise, when Time x' k =x k +Δx k Repeat step 2 k Second-rate;
[0250] Step 3: Calculate T k+1 =Q SA T k , where Q SA is the annealing factor in the simulated annealing algorithm. Let k = k + 1, and determine whether the algorithm termination condition is met. If so, stop; if not, return to step 2;
[0251] At this point, the core idea of the SA algorithm is introduced into the TD3 algorithm, using the first critic network as the exploration network for the initial solution in the SA algorithm, and the second critic network as the exploration network for the perturbation term in the SA algorithm. Therefore, the network update formula of the SA-TD3 algorithm is specifically expressed as follows:
[0252]
[0253] The SA-TD3 algorithm is used to solve the regional power grid's day-ahead dispatch optimization model. The specific steps for obtaining the day-ahead dispatch plan are as follows:
[0254] Step 3.3: Initialize the policy network π φ , valuation network Target Network Experience pool β, set learning rate α, discount factor γ, delayed update step number d, soft update coefficient τ and maximum learning step number maxepisode, learning step number count = 0, initial annealing temperature T k , and the state space s constructed above t As algorithm input, input to the Actor network;
[0255] Step 3.4: According to the output of the Actor network, introduce Gaussian noise and generate the current action a under the condition of satisfying relevant constraints t =μ(s t |θ μ )+η, and schedule the plan a t Send to the power grid simulation environment to calculate the next state s t+1 And the current running cost C t That is, considering the carbon and green certificate transaction costs of power generation units under the carbon-green certificate coupling transaction mechanism and other related day-ahead scheduling costs, the sample data (s t,a t ,C t ,s t+1 ) is stored in the experience pool, and t=t+1 is used to update the Critic network Let count = count + 1. If the number of learning steps count is divisible by the number of delayed update steps d, then based on the SA algorithm strategy, update the Actor network θ μ , soft update target network Minimize the day-ahead scheduling cost and repeat this step. Otherwise, repeat this step directly until the maximum number of learning steps is reached, and output the optimized scheduling plan, that is, the output of each unit.
Claims
1. A regional power grid day-ahead dispatching method based on reinforcement learning and carbon-green certificate coupling, characterized by: The following steps are involved: Step 1: Construct a model related to carbon trading and green certificate trading. Build a tiered carbon trading model based on carbon quota allocation and carbon emission calculations. Build a green certificate trading model based on the green certificate trading mechanism. Considering the benign coupling between the carbon trading market and the green certificate trading market, build a carbon-green certificate coupling model. Step 2: Build a regional power grid day-ahead dispatch optimization model based on a deep reinforcement learning algorithm. The actual dispatch process of a regional power grid involves four stages: grid operating status collection, dispatch instruction formulation, dispatch instruction execution, and dispatch instruction feedback. Deep reinforcement learning involves four stages: observing status, selecting actions, executing actions, and reward feedback. By introducing a deep reinforcement learning framework, we mapped the four stages of the regional power grid optimization problem to the four stages of the deep reinforcement learning framework, and constructed a regional power grid dispatch framework based on deep reinforcement learning. Step 3: Introduce the core idea of the SA algorithm into the TD3 algorithm, propose the SA-TD3 algorithm, and use the SA-TD3 algorithm to solve the regional power grid day-ahead dispatch optimization model and obtain the day-ahead dispatch plan.
2. The method for day-ahead dispatch of a regional power grid based on reinforcement learning and carbon-green certificate coupling according to claim 1, characterized in that: The step 1 is as follows: Step 1.1: Construct a carbon trading model and use the baseline method to calculate the allocation formula for carbon quotas as follows: Where, The initial carbon emission quota obtained for thermal power unit i, τ * is the carbon emission baseline per unit electricity volume of power generation enterprises formulated within the industry, T is the number of time periods in the entire dispatch cycle, ΔT is the dispatch time interval, and there are T = 24h / ΔT decision periods in the whole day. i,t is the output of thermal power unit i at time t; Taking into account the relationship between the carbon emissions of thermal power generating units and the type and fuel consumption, the carbon emissions of thermal power unit i under conventional output are: Where, is the unit fuel consumption carbon emission of thermal power unit i, a i 、b i 、c i is the unit power generation burn-up coefficient of thermal power unit i in normal operation state, P i max 、P i min are the upper and lower limits of the conventional output of thermal power unit i, respectively. In addition, considering that some thermal power units can perform deep peak regulation, the combustion efficiency is reduced and the additional fuel loss is increased in this state, which will cause additional carbon emissions. The carbon emissions are expressed as: Where, is the additional carbon emissions associated with the deep peak regulation of the unit, is the coefficient of additional fuel consumption generated when thermal power unit i performs deep peak regulation, P i a is the lower limit of the output of thermal power unit i in the deep peak regulation state; Carbon trading for power generation companies meets the following requirements: Where, is the carbon trading cost of thermal power unit i; CET is the price per carbon certificate transaction; is the actual carbon emission of thermal power unit i. The higher the carbon emission quota that the thermal power unit needs to purchase, the higher the transaction price of the corresponding interval carbon certificate purchased from the carbon certificate market, that is, Where, is the basic price of carbon certificate transaction, θ CET represents the sensitivity coefficient of carbon price growth rate, d CET is the interval length of a single carbon emission level; therefore, the tiered carbon trading cost of thermal power units can be expressed as: Step 1.2: Build a green certificate trading model. Based on the green power production tasks that need to be completed by new energy generators in the regional power grid, formulate the initial allocation benchmark value of new energy generator power generation and green certificate quota indicators; Where, is the green certificate quota index and green certificate quantity obtained by the new energy generator j in the regional power grid, α l is the green certificate quota coefficient, is the initial allocation benchmark value and actual power generation of the new energy generator set j in period t, α f The coefficient for obtaining the number of green certificates per unit of electricity generation from renewable energy sources is used. A tiered pricing model is introduced into the green certificate trading. The more carbon certificates a new energy unit needs to purchase, the higher the transaction price of green certificates in the corresponding range purchased from the green certificate market. The surplus green certificates are uniformly recovered by the government or the carbon certificate market at a base price. The model satisfies the following requirements: Where, is the green certificate transaction cost of new energy unit j, is the basic price of green certificate transaction, d GCT is the interval length of a single green certificate trading level, θ GCT is the growth coefficient of green certificate price; Step 1.3: Construct a carbon-green certificate coupling trading model. The emission reduction of new energy is the difference between the carbon emissions generated by new energy power generation and the carbon emission standards set by the industry. Wind power and photovoltaic power generation can be regarded as zero carbon emissions. The emission reduction of new energy power generation is: Where, The carbon emissions reduced by new energy generators j, The standard carbon emissions set for the industry, is the carbon emission generated by new energy power generation enterprise j, which is zero; therefore, the green certificate obtained by new energy represents the carbon emission rights: In the formula, β is the emission reduction represented by a unit green certificate. Therefore, according to the above relationship, the conversion between green certificates and carbon emission rights is as follows: Where, Carbon emission rights converted from green certificates purchased for thermal power unit i; is the number of green certificates purchased by thermal power unit i from the green certificate market; therefore, the carbon trading and green certificate trading of thermal power unit i under the carbon-green certificate coupling mechanism can be expressed as: Where, is the cost of purchasing green certificates from the green certificate market for thermal power unit i.
3. The method for day-ahead dispatch of a regional power grid based on reinforcement learning and carbon-green certificate coupling according to claim 1, characterized in that: The step 2 is as follows: By using the one-to-one correspondence between the four stages involved in the actual dispatching process of the regional power grid and the four stages involved in deep reinforcement learning, the regional power grid dispatching problem is modeled as a Markov decision process: the dispatching time of a whole day is divided into t periods, and the environment state s is set to t Corresponding to the operating status of the regional power grid at time t, action a t Corresponding to the dispatch plan of the regional power grid at time t+1, in the environmental state s t Next, perform action a t , the regional power grid shifts to the next environmental state s t+1 , which also corresponds to the action a of the regional power grid dispatch plan at time t+2 t+1 Finally, the intelligent agent will give feedback on the accumulated rewards, which corresponds to the regional power grid's response to the economic efficiency, carbon emissions and other indicators of the dispatch plan. t Therefore, the Markov decision process modeling of the regional power grid can be expressed as: M e ={s t ,a t ,C t ,s t+1 } Source-side resources mainly include coal-fired power generation, gas-fired power generation, wind power generation, and photovoltaic power generation. Load-side resources mainly include shiftable loads, curtailable loads, and rigid loads. The set of coal-fired power generation units is defined as Φ pg ={1,2,...,N g }, the set of gas-fired power generation units is Φ gas ={1,2,...,N gas }, the wind turbine generator set is Φ wind ={1,2,...,N w }, the photovoltaic power station set is Φ PV ={1,2,...,N v }, the output value of coal-fired power generation unit i at time t is P i,t , the output value of gas-fired power unit i at time t is The output value of wind turbine i at time t is The output value of the photovoltaic power station at time t is The actual value of the load at time t is P t load , where the load that can be reduced is P t cut , the translatable load is P t sh ; 4. The method for day-ahead dispatch of a regional power grid based on reinforcement learning and carbon-green certificate coupling according to claim 3, characterized in that: The specific process of step 2 is as follows: Step 2.1: Construct the state value space. To construct the state space, first divide the scheduling time T in a whole day into t scheduling periods, that is, T = {1, 2, ..., t}; the system state at time t is s t , which includes time t, the unit output P of the coal-fired unit at time t i,t ,i∈{1,...,N g }、The operating time of coal-fired power unit i at time t The shutdown time of coal-fired power unit i at time k The output of gas-fired power unit i at time t The operating time of gas-fired power unit i at time t The downtime of gas-fired power unit i at time t The output of the wind turbine at time k The output of the photovoltaic power station at time t The power P of the load at time t t load , specifically expressed as follows: To establish the action space, record the agent’s current state s at time t t The response is a t , that is, the dispatch plan executed by the regional power grid at time t, and the output adjustment of the coal-fired unit in the current state is ΔP i,t ,i∈{1,...,N g }, the output adjustment of the gas unit in the current state is The action space is composed of the output adjustment of all coal-fired units and gas-fired units in the current state, which can be expressed as: Step 2.2: Construct the objective function and constraints. The day-ahead dispatch cost of the regional power grid is composed of the operating cost of thermal power units, the carbon trading cost of thermal power units, the green certificate trading cost of thermal power units, the green certificate trading cost of new energy units, the dispatch cost of curtailable loads, the dispatch cost of shiftable loads, the penalty cost for abandoning new energy resources, and the penalty cost for load shedding. The cost model for the optimization objective is specifically expressed as follows: Operating costs of thermal power units: Where C g is the operating cost of thermal power unit i, including the operating cost of coal-fired thermal power units and gas-fired units; Carbon trading costs of thermal power units: Where C CET The carbon trading cost of thermal power units, including the carbon trading cost of coal-fired thermal power units and the carbon trading cost of gas-fired thermal power units; Green certificate transaction costs for thermal power units: Where C buy-GCT The green certificate transaction cost of thermal power units, including the green certificate transaction cost of coal-fired thermal power units and the green certificate transaction cost of gas-fired thermal power units; Green certificate transaction costs for new energy units: Where C GCT The green certificate transaction cost of new energy units, including the green certificate transaction cost of wind turbines and photovoltaic power stations; Dispatch costs that can reduce loads: Constructing a tiered load reduction compensation price model: Where, P is the compensation price for the unit amount of electricity that can be reduced at time t, t cut is the amount of power that can be reduced, P t cut,max 、P t cut,min The maximum and minimum reduction amounts that can be accepted for load reduction are: is the growth coefficient function related to the reduction amount, is the initial compensation price for curtailable load; the curtailment compensation cost of curtailable load response meets the following requirements: Where C pay,cut Compensation costs for curtailment of load that can be curtailed; In addition, load reduction can also provide load-side spare capacity, which must meet the upper limit constraint of spare capacity: Where, The load-side spare capacity and capacity upper limit provided for the curtailable load; within the reserve capacity constraint, the reserve cost brought about by providing reserve capacity to respond to grid dispatch requirements for curtailable loads satisfies: Where C re,cut The backup cost caused by providing load-side backup capacity for curtailable loads, the price of providing a unit of reserve capacity for curtailable load; Therefore, the dispatch cost of load reduction includes the reduction compensation cost and the reserve cost, which satisfies: C cut =C pay,cut +C re,cut Where C cut The dispatch cost for load reduction; Dispatch cost of shiftable loads: The scheduling model of the shiftable load satisfies: Where, t sh is the start time of the movable load after scheduling, t sh- , t sh+ P is the earliest and latest operation start time that the shiftable load is willing to respond to the grid demand, t sh T is the power consumption of the movable load in period t after scheduling, sh is the continuous operation time of the translatable load, After the load is dispatched, the sh +τ period of power consumption, before dispatch at t sh* +Power consumption during the τ period, t sh* , t sh The start time of the movable load before and after scheduling; Acceptable scheduling time for shiftable loads [t sh- ,t sh+ ]satisfy: Where, ρ sh is the unit compensation price of the translatable load, β sh (ρ sh ) is the acceptable translation time and compensation price of the translatable load ρ sh Related related functions, is the maximum number of translation periods for the translation load; The function of the shiftable load compensation price satisfies: Where, ρ sh- and ρ sh+ are the minimum compensation price and maximum compensation price of the unit load that can be translated when the dispatch agreement is signed, Δρ sh is the sensitivity coefficient related to the compensation price; Therefore, the dispatch cost of the shiftable load satisfies: Where C sh is the dispatching cost of the movable load; τ sh is a Boolean variable, τ sh =1 means that the translatable load moves, τ sh =0 means that the translatable load does not translate; New energy disposal penalty costs: Where C cnew is the penalty cost for abandoning new energy, λ cw ,λ cv is the penalty coefficient for unit abandoned electricity of wind power and photovoltaic power, is the abandoned power of wind turbine i in period t, is the abandoned electricity of photovoltaic power station j at time t; Load shedding penalty cost: Where C cl is the load shedding penalty cost of the regional power grid, P t cl is the load shedding power of the regional power grid during period t; Therefore, the optimization goal of the regional power grid's day-ahead dispatch is to minimize the day-ahead dispatch cost: my C day =min(C g +C CET +C buy-GCT +C GCT +C cut +C sh +C cnew +C cl ) To maintain safe and stable operation, the regional power grid must not only meet the constraints related to the operation of the thermal power units themselves, the constraints related to load reduction, and the constraints related to load shifting, but also meet the line transmission capacity and flow constraints. The transmission capacity constraints are satisfied: Where, T l,i 、T l,j 、T l,k 、T l,m and T l,n is the power transmission distribution coefficient; N l is the number of node loads in the regional power grid; is the load power of node m at time t; F l max is the upper limit of the power flow of line l.
5. The method for regional power grid day-ahead scheduling based on reinforcement learning and carbon-green certificate coupling according to claim 1, characterized in that: The step 3 specifically includes: Step 3.1: Introduce a regularization strategy to smooth the target strategy. After smoothing, action a is expressed as: a←μ'(s|θ μ' )+η η←clip(N(0,σ),-c,c) Where, μ'(s|θ μ' ) is the policy function, s is the given state, θ μ' The parameters of the target network, η is the noise added during action selection, clip(·) is the truncation function that limits the maximum and minimum values of the noise, and N(0,σ) is a normal distribution function with a mean of 0 and a variance of σ. Step 3.2: Introduce the core idea of the SA algorithm into the TD3 algorithm. The network update function of the TD3 algorithm is specifically expressed as a loss function as follows: y=-C+γmin(Q'1(s,a),Q'2(s,a)),γ∈[0,1] In the formula, γ is the discount factor, N is the total number of samples used, s, a and C are the sampled states, actions and running states respectively. is the loss function of the Critic network i. The TD3 algorithm reduces the frequency of the Actor network update to reduce the amount of learned error information. The update gradient of the Actor network is satisfy: The core idea of the SA algorithm is introduced into the TD3 algorithm. The first critic network is used as the exploration network for the initial solution in the SA algorithm, and the second critic network is used as the exploration network for the perturbation term in the SA algorithm. The network update formula of the SA-TD3 algorithm is specifically expressed as follows: The SA-TD3 algorithm is used to solve the regional power grid's day-ahead dispatch optimization model and obtain the day-ahead dispatch plan. The specific steps are as follows: Step 3.3: Initialize the policy network π φ , valuation network Target Network Experience pool β, set learning rate α, discount factor γ, delayed update step number d, soft update coefficient τ and maximum learning step number max episode, learning step number count = 0, initial annealing temperature T k , and the state space s constructed above t As algorithm input, input to the Actor network; Step 3.4: According to the output of the Actor network, introduce Gaussian noise and generate the current action a under the condition of satisfying relevant constraints t =μ(s t |θ μ )+η, and schedule the plan a t Send to the power grid simulation environment to calculate the next state s t+1 And the current running cost C t That is, considering the carbon and green certificate transaction costs of power generation units under the carbon-green certificate coupling transaction mechanism and other related day-ahead scheduling costs, the sample data (s t ,a t ,C t ,s t+1 ) is stored in the experience pool, and t=t+1 is used to update the Critic network Let count = count + 1. If the number of learning steps count is divisible by the number of delayed update steps d, then based on the SA algorithm strategy, update the Actor network θ μ , soft update target network Minimize the day-ahead scheduling cost and repeat this step. Otherwise, repeat this step directly until the maximum number of learning steps is reached, and output the optimized scheduling plan, that is, the output of each unit.
6. The method for day-ahead dispatch of a regional power grid based on reinforcement learning and carbon-green certificate coupling according to claim 5, characterized in that: The specific steps of the algorithm flow of the SA algorithm are as follows: Step 1: Initialize k = 0, based on the initial annealing temperature T k Find the initial random solution x0, at this time x k =x0; Step 2: For a feasible solution x k Add a random noise Δx k Get x' k =x k +Δx k , calculate the evaluation function value f(x k ) and f(x' k ), if f(x k )<f(x' k ), then x' k =x k ; Otherwise, when e-(f(x' k )-f(x k ) / T k )>rand() k =x k +Δx k , repeat step 2 k Second-rate; Step 3: Calculate T k+1 =Q SA T k , where Q SA is the annealing factor in the simulated annealing algorithm, let k = k + 1, and determine whether the algorithm termination condition is met. If so, stop; if not, return to step 2.
Citation Information
Patent Citations
Deep reinforcement learning-based day-ahead-intra-day combined dispatching method for regional power grid
CN115441437A
Electricity-carbon-green evidence multi-market equilibrium analysis method based on multi-agent reinforcement learning
CN117314040A