A power adjustment strategy for power grid section based on deep reinforcement learning

By employing a grid cross-section power adjustment strategy based on deep reinforcement learning, which comprehensively considers unit sensitivity, economy, and carbon emissions, the speed and accuracy issues of grid dispatching in complex networks are resolved, thereby achieving grid safety, stability, and economic optimization.

CN115864409BActive Publication Date: 2026-04-24JIANGSU HOPERUN SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU HOPERUN SOFTWARE CO LTD
Filing Date
2023-01-09
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies can meet the requirements of power grid dispatch in small network scenarios, but they cannot meet the speed and accuracy requirements in complex networks, and they ignore the economic efficiency and carbon emissions of power grid operation, which increases the difficulty of cross-sectional power adjustment.

Method used

A grid cross-section power adjustment strategy based on deep reinforcement learning is adopted. By combining generator selection module and deep reinforcement learning module, the sensitivity, economy and carbon emissions of generator units are comprehensively considered. Deep neural network is used to map state and action to optimize cross-section power adjustment.

Benefits of technology

It enables rapid and precise adjustment of cross-sectional power in complex networks, ensuring the safe and stable operation of the power grid while optimizing economic efficiency and carbon emissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115864409B_ABST
    Figure CN115864409B_ABST
Patent Text Reader

Abstract

The application provides a power grid section power adjustment strategy based on deep reinforcement learning, characterized by comprising a generator set selection and a deep reinforcement learning module, wherein the generator set selection comprises: section power adjustment, comprehensively considering the sensitivity, economy and carbon emission of the unit, and when maintaining the convergence of section power flow and quickly reaching the target value, the economy and carbon emission of the unit operation are optimal; the deep reinforcement learning module comprises training and verification of the module, in the training stage, the set state, action, reward value and reward discount coefficient are used to repeatedly input the section state into the model, and the corresponding action is generated to finally obtain the learning model fitting the state and action relationship; for the verification stage, whether the section power can quickly reach the target value and ensure the economy and carbon emission of the section as a whole is verified; the application makes the section power quickly reach the target value under the premise of ensuring the safe and stable operation of the power grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid dispatching, specifically a power grid section adjustment strategy based on deep reinforcement learning. Background Technology

[0002] With the escalating global energy crisis in recent years and the emergence of the "dual-carbon" concept, accelerating the construction of new-type power grids and improving their intelligence level have become key research focuses. Power grid dispatching, as a crucial component of smart grid construction, plays a vital role. A rational dispatching method can not only maintain the safe and stable operation of the power grid but also reduce unnecessary economic expenditures and energy waste, thereby improving the overall safety and economy of the power grid system. This paper mainly studies the power adjustment of power grid sections. As a component of power grid dispatching, the purpose of section power adjustment is primarily to ensure load balance within the system and maintain the safe, stable, and economical operation of the power grid by regulating the output of generator units within a section. On the one hand, there are various types of generator units within the smart grid, and the specific generator unit selected for actual section power adjustment requires comprehensive consideration from multiple perspectives. On the other hand, the load within the power grid is significantly affected by meteorological and human factors, resulting in overall randomness and uncertainty in the data. These factors contribute to the diversity of section power adjustment methods, exacerbating the difficulty of section power adjustment. Summary of the Invention

[0003] The purpose of this invention is to propose a power grid section adjustment strategy based on deep reinforcement learning to solve the problem that existing technologies can meet scheduling requirements in small network scenarios, but cannot meet the requirements in terms of speed and accuracy when facing complex networks; and there is also the problem that they only consider the timeliness of adjustment and the stability of grid operation, while ignoring the economic efficiency and carbon emissions of grid operation.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A power grid cross-section power adjustment strategy based on deep reinforcement learning includes a generator selection module and a deep reinforcement learning module. The generator selection module mainly includes generator screening and power compensation, dividing the two ends of the cross-section into sending and receiving ends according to the power flow direction. The power of the transmission cross-section is the sum of the power on the intermediate tie line between the sending and receiving ends. The deep reinforcement learning module mainly refers to using a deep neural network as the policy network for reinforcement learning, thereby completing the mapping relationship between the cross-section state and the required actions. Specifically, it is implemented through the following five steps, including:

[0006] S1. Calculate the sensitivity of the generator sets: While ensuring the convergence of the power flow calculation, minimize the adjustment data of the generator sets as much as possible. It is necessary to calculate the sensitivity of each generator set here.

[0007] S2. Determine the economic efficiency and carbon emissions of the generator: For this patent, when selecting a generator set for power regulation, in addition to considering the sensitivity of the generator set, its economic efficiency and carbon emissions also need to be considered. Here, the concept of carbon emission rights is introduced. Carbon emission rights refer to the amount of carbon emissions allowed by the generator set under the premise of the annual power generation hours.

[0008] S3. Determine the power regulating units: After determining the sensitivity of each generator unit, the generator units can be divided into a set of forward regulating units ψ. pos ψ, a set of reverse regulating units neg Then, based on the actions learned from reinforcement learning, a set of units is selected to participate in the cross-sectional power adjustment;

[0009] S4. Determine the power compensation unit: The method for determining the power compensation unit is generally similar to the power adjustment method. When the unit adjusts the power in the forward direction, the reverse power compensation is performed; when the unit adjusts the power in the reverse direction, the forward power compensation is performed.

[0010] S5. Establish a deep reinforcement learning module: When training with deep reinforcement learning, the agent learns by continuously interacting with the environment, thereby obtaining the optimal cross-sectional power compensation strategy.

[0011] The generator selection module specifically includes: For generator selection, not only the sensitivity of the generator set but also its operational economy and carbon emissions are considered. The generator sets are ranked from largest to smallest based on these three aspects, and the higher-ranked generator sets are prioritized for power adjustment during cross-sectional power adjustment. Secondly, to avoid power imbalance after grid cross-sectional power adjustment, a sending-end adjustment and receiving-end compensation method is adopted. The sending end selects the corresponding generator set for cross-sectional power adjustment based on the power difference, while the receiving end maintains the power balance of the smart grid system through reverse power output.

[0012] The deep reinforcement learning module specifically includes: based on the requirements of the Markov decision process quintuple (S, A, P, R, γ), the power grid power adjustment model mainly focuses on the setting of the cross-section state S, action A, reward R, and discount factor γ. During the training phase, with the reward function R and discount factor γ set, the cross-section state S is input into the reinforcement learning module. Based on the action A output by the policy network, the cross-section power is adjusted, and a new cross-section state S is obtained through power flow calculation. This process is repeated, ensuring the convergence of the power flow calculation, and relevant data is collected to complete the optimization process of the policy network. In the prediction phase, since the policy network has already been fitted, the power adjustment action is obtained by continuously inputting the state into the reinforcement learning model, ensuring power flow convergence until the cross-section power reaches the target value.

[0013] Step S1 specifically includes the following:

[0014] The specific calculation formula for the sensitivity of each unit is as follows:

[0015]

[0016]

[0017] Where: Ω represents the set of all units except the balancing unit; P init This represents the current power output of the power grid section. and This represents the change in grid cross-sectional power of unit k at its maximum and minimum power. and This refers to the sensitivity index of generator k for forward and reverse regulation. It should be noted that during forward regulation, the reverse generator unit is used for compensation, and during reverse regulation, the forward generator unit is used for compensation. For regulation, the generator unit with higher sensitivity is preferred, while for power compensation, the generator unit with lower sensitivity is preferred.

[0018] Step S2 specifically includes the following:

[0019] The calculation of carbon emission rights can be expressed as follows:

[0020] E r,i =e i T i W i (3)

[0021] Where: E r,i Carbon emission rights for unit i, kg; e i The carbon emission intensity of the corresponding unit is expressed in kg / kW·h; T i For the corresponding generating unit's power generation time, h; W i The installed capacity is expressed in kW.

[0022] To achieve energy conservation and emission reduction, before considering the participation of the unit in power adjustment, it is necessary to calculate its carbon emission rights. By evaluating the unit's output time, the future carbon emissions of the unit are calculated and compared with the carbon emission rights. If the carbon emissions are less than the carbon emission rights, the unit can participate in power adjustment and is recorded as 1; otherwise, it will not participate and is recorded as 0.

[0023] E o2,i =E o1,i +e i ·P max,i ·h (4)

[0024]

[0025] Among them, E o2,iCarbon emissions after power adjustment for the group; E o1,i This represents the unit's current carbon emissions; P max,i ρ is the maximum operating power of the unit; h is the operating time of the unit; ρ is the carbon emission weight factor, which ranges from 0 to 1.

[0026] Smart grids typically contain various types of generating units, such as thermal power units, natural gas units, and renewable energy units. Generally, renewable energy units do not participate in cross-sectional power regulation. Thermal power units have higher carbon emissions but lower costs; natural gas units have higher costs but lower carbon emissions. Therefore, when selecting units for cross-sectional power regulation, in addition to considering sensitivity, economic efficiency and carbon emissions must be comprehensively considered. Regarding economic efficiency, this patent currently only considers the cost of raw materials, specifically as follows:

[0027]

[0028] Among them, c c,i Q represents the power generation cost of the corresponding generating unit; i The raw materials required for a unit of electricity generation, where raw materials include coal, natural gas, etc.; n i For the corresponding generating unit's power generation efficiency; l i This refers to the price of raw materials.

[0029] Specifically, carbon emissions are expressed as follows:

[0030]

[0031] Among them, e c,i Q represents the carbon emissions of the corresponding unit. i n i The meaning is the same as in equation (6); g i Carbon emissions per unit of fuel.

[0032] Step S3 specifically includes the following:

[0033] For this method, economic efficiency and carbon emissions also need to be considered. Therefore, it is necessary to appropriately expand the set of participants in power adjustment, as shown in the following equation:

[0034]

[0035]

[0036] Where: a is the action value output by the reinforcement learning module; when the target value is greater than the current power of the cross section, a>0, and positive adjustment is performed using equation (8); when the target value is less than the current power of the cross section, a<0, and reverse adjustment is performed using equation (9); N′ is the number of units participating in positive adjustment, and M′ is the number of units participating in reverse adjustment; thus, a new set of units ψ′ can be obtained. pos ψ n ′ eg The relationship between this set and the original set is: By combining the sensitivity, economy, and carbon emissions of the generating units, the most suitable unit is selected to participate in power adjustment. The specific integration method can adopt a weighted average approach as follows:

[0037]

[0038] Where w1 is the weighted average coefficient of the unit's adjustment, w2 is the weighted average coefficient of carbon emissions, and w3 is the weighted average coefficient of economic efficiency, and the sum of w1, w2, and w3 is 1.

[0039] Finally, regarding z i The units are sorted from largest to smallest to select the top N (M) units for cross-sectional power adjustment, as shown below:

[0040]

[0041]

[0042] Step S4 specifically includes the following:

[0043] The specific selection method for power compensation units is as follows:

[0044]

[0045] This results in a new set of compensation units ψ′ pos (ψ n ′ eg The relationship here is the same as in step 3. Finally, the sensitivity, economy, and carbon emissions are integrated as follows:

[0046]

[0047] When performing cross-sectional power compensation, a sorting method from smallest to largest is adopted, and the first N (M) units are selected to participate in power compensation, as shown in the following figure:

[0048]

[0049] Step S5 specifically includes the following:

[0050] The process of adjusting the cross-sectional power can be regarded as a Korkov decision process, which can be represented as M={S,A,P,R,γ};

[0051] The state space S represents the state of the current power grid section after being affected by the corresponding action. For this invention, it includes the power of the current section, the target power value of the section, and the output of each generator in the region; it is represented as follows:

[0052] S = [P] init ,P tar p1, p2, ... p m (16)

[0053] Action Space A: The action space mainly represents the set of scalars controlled by the intelligent agent in the cross-sectional power adjustment. Here, the action represents the value of the unit's required output.

[0054] A = {△P} (17)

[0055] State transition probability P: The state transition probability refers to the probability that the power grid system will reach a certain state after being affected by an action issued by an agent, and is expressed as: S t+1 ←S t ×A, P∈[0,1];

[0056] Reward R: The reward represents the optimization objective of the deep reinforcement learning network. By continuously optimizing the network structure parameters, the cumulative reward value is maximized. In this patent, the main goal is to adjust the generator sets to ensure the convergence of power flow calculations and that the cross-sectional power reaches a certain target value. Therefore, R here can be expressed as the absolute value of the difference between the actual cross-sectional power and the target value, as follows:

[0057] R = -|P t -P tar | (18)

[0058] Reward discount factor γ: The discount factor represents the influence of time distance on the model, which can be expressed as: γ∈[0,1]. The larger γ is, the greater the time distance, the greater the influence of the model parameters.

[0059] In deep reinforcement learning models, deep neural networks serve as the policy function π(a|s) to fit the relationship between the cross-sectional state S and the action a. Therefore, the training objective of reinforcement learning networks is to continuously update the network parameters of the policy function to maximize the accumulated reward, max R(τ), as follows:

[0060]

[0061] Where J(π) θ ) represents the maximum expected return for reinforcement learning, γ is the discount factor, and r is the cumulative reward value;

[0062] The training objective is to find a set of parameters for the network θ such that J(π) θ To maximize network parameters, it is only necessary to solve for the partial derivatives of each parameter within the network. To improve data utilization and agent stability, resampling is used during model training. Here, the PPO algorithm, a branch of deep reinforcement learning, is used to train the agent, and its optimization objective can be expressed as:

[0063]

[0064] Where θ * To optimize the objective, D KL For the KL divergence constraint, in the formula For π θ (·|s t )and The distance between them;

[0065] It should be noted that π θ (a t |s t ), These represent the target policy and the sampling policy, respectively. and The value represents the empirical estimate. In addition, to constrain the distance between the target policy and the sampling policy, KL divergence is used to calculate the expected distance between them and serve as a constraint term in the optimization problem. The hyperparameter λ is used to balance the KL divergence penalty term.

[0066] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0067] This invention employs a method combining deep reinforcement learning to formulate power grid cross-section power adjustment. Given the large number of nodes within the power grid, the selection of generator sets for cross-section power adjustment considers not only the generator set's sensitivity but also its economic efficiency and carbon emissions. The optimal selection is achieved by comprehensively considering these three criteria. Secondly, the deep reinforcement learning algorithm is used to fully fit the relationship between the cross-section state and the adjustment action, ensuring that the generated action values, after conversion, can effectively guide the generator set's output. This allows the cross-section power to quickly reach the target value while ensuring the safe and stable operation of the power grid. Attached Figure Description

[0068] Figure 1 This invention provides a reinforcement learning framework;

[0069] Figure 2 This is a flowchart of the cross-section control strategy of the present invention. Detailed Implementation

[0070] To clarify the technical problems, technical solutions, implementation processes, and performance demonstrations, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative. The present invention is not intended to limit the scope of the invention. Various exemplary embodiments, features, and aspects of this disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements with the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0071] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0072] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0073] Example 1

[0074] like Figure 1 and Figure 2 As shown, a power grid cross-section adjustment strategy based on deep reinforcement learning includes a generator selection module and a deep reinforcement learning module. The generator selection module mainly includes generator selection and power compensation, dividing the two ends of the cross-section into sending and receiving ends according to the power flow direction. The power of the transmission cross-section is the sum of the power on the intermediate tie line between the sending and receiving ends. The deep reinforcement learning module mainly refers to using a deep neural network as the policy network for reinforcement learning, thereby completing the mapping relationship between the cross-section state and the required actions. Specifically, it is implemented through the following five steps, including:

[0075] S1. Calculate the sensitivity of the generator sets: While ensuring the convergence of the power flow calculation, minimize the adjustment data of the generator sets as much as possible. It is necessary to calculate the sensitivity of each generator set here.

[0076] S2. Determine the economic efficiency and carbon emissions of the generator: For this patent, when selecting a generator set for power regulation, in addition to considering the sensitivity of the generator set, its economic efficiency and carbon emissions also need to be considered. Here, the concept of carbon emission rights is introduced. Carbon emission rights refer to the amount of carbon emissions allowed by the generator set under the premise of the annual power generation hours.

[0077] S3. Determine the power regulating units: After determining the sensitivity of each generator unit, the generator units can be divided into a set of forward regulating units ψ. pos ψ, a set of reverse regulating units negThen, based on the actions learned from reinforcement learning, a set of units is selected to participate in the cross-sectional power adjustment;

[0078] S4. Determine the power compensation unit: The method for determining the power compensation unit is generally similar to the power adjustment method. When the unit adjusts the power in the forward direction, the reverse power compensation is performed; when the unit adjusts the power in the reverse direction, the forward power compensation is performed.

[0079] S5. Establish a deep reinforcement learning module: When training with deep reinforcement learning, the agent learns by continuously interacting with the environment, thereby obtaining the optimal cross-sectional power compensation strategy.

[0080] The generator selection module specifically includes: For generator selection, not only the sensitivity of the generator set but also its operational economy and carbon emissions are considered. The generator sets are ranked from largest to smallest based on these three aspects, and the higher-ranked generator sets are prioritized for power adjustment during cross-sectional power adjustment. Secondly, to avoid power imbalance after grid cross-sectional power adjustment, a sending-end adjustment and receiving-end compensation method is adopted. The sending end selects the corresponding generator set for cross-sectional power adjustment based on the power difference, while the receiving end maintains the power balance of the smart grid system through reverse power output.

[0081] The deep reinforcement learning module specifically includes: based on the requirements of the Markov decision process quintuple (S, A, P, R, γ), the power grid power adjustment model mainly focuses on the setting of the cross-sectional state S, action A, reward R, and discount factor γ, as shown in the specific structure. Figure 1 As shown. During the training phase, with the reward function R and discount factor γ set, the cross-sectional state S is input into the reinforcement learning module. Based on the action A output by the policy network, the cross-sectional power is adjusted, and the power flow calculation is completed to obtain a new cross-sectional state S. This process is repeated, ensuring the convergence of the power flow calculation, and relevant data is collected to complete the optimization process of the policy network. In the prediction phase, since the policy network has already been fitted, the state is continuously input into the reinforcement learning model to obtain power adjustment actions. This continues until the cross-sectional power reaches the target value, ensuring the convergence of the power flow.

[0082] Step S1 specifically includes the following:

[0083] The specific calculation formula for the sensitivity of each unit is as follows:

[0084]

[0085]

[0086] Where: Ω represents the set of all units except the balancing unit; P init This represents the current power output of the power grid section. and This represents the change in grid cross-sectional power of unit k at its maximum and minimum power. and This refers to the sensitivity index of generator k for forward and reverse regulation. It should be noted that during forward regulation, the reverse generator unit is used for compensation, and during reverse regulation, the forward generator unit is used for compensation. For regulation, the generator unit with higher sensitivity is preferred, while for power compensation, the generator unit with lower sensitivity is preferred.

[0087] Step S2 specifically includes the following:

[0088] The calculation of carbon emission rights can be expressed as follows:

[0089] E r,i =e i T i W i (3)

[0090] Where: E r,i Carbon emission rights for unit i, kg; e i The carbon emission intensity of the corresponding unit is expressed in kg / kW·h; T i For the corresponding generating unit's power generation time, h; W i The installed capacity is expressed in kW.

[0091] To achieve energy conservation and emission reduction, before considering the participation of the unit in power adjustment, it is necessary to calculate its carbon emission rights. By evaluating the unit's output time, the future carbon emissions of the unit are calculated and compared with the carbon emission rights. If the carbon emissions are less than the carbon emission rights, the unit can participate in power adjustment and is recorded as 1; otherwise, it will not participate and is recorded as 0.

[0092] E o2,i =E o1,i +e i ·P max,i ·h (4)

[0093]

[0094] Among them, E o2,i Carbon emissions after power adjustment for the group; E o1,i This represents the unit's current carbon emissions; P max,i ρ is the maximum operating power of the unit; h is the operating time of the unit; ρ is the carbon emission weight factor, which ranges from 0 to 1.

[0095] Smart grids typically contain various types of generating units, such as thermal power units, natural gas units, and renewable energy units. Generally, renewable energy units do not participate in cross-sectional power regulation. Thermal power units have higher carbon emissions but lower costs; natural gas units have higher costs but lower carbon emissions. Therefore, when selecting units for cross-sectional power regulation, in addition to considering sensitivity, economic efficiency and carbon emissions must be comprehensively considered. Regarding economic efficiency, this patent currently only considers the cost of raw materials, specifically as follows:

[0096]

[0097] Among them, c c,i Q represents the power generation cost of the corresponding generating unit; i The raw materials required for a unit of electricity generation, where raw materials include coal, natural gas, etc.; n i For the corresponding generating unit's power generation efficiency; l i This refers to the price of raw materials.

[0098] Specifically, carbon emissions are expressed as follows:

[0099]

[0100] Among them, e c,i Q represents the carbon emissions of the corresponding unit. i n i The meaning is the same as in equation (6); g i Carbon emissions per unit of fuel.

[0101] Step S3 specifically includes the following:

[0102] For this method, economic efficiency and carbon emissions also need to be considered. Therefore, it is necessary to appropriately expand the set of participants in power adjustment, as shown in the following equation:

[0103]

[0104]

[0105] Where: a is the action value output by the reinforcement learning module; when the target value is greater than the current power of the cross section, a>0, and positive adjustment is performed using equation (8); when the target value is less than the current power of the cross section, a<0, and reverse adjustment is performed using equation (9); N′ is the number of units participating in positive adjustment, and M′ is the number of units participating in reverse adjustment; thus, a new set of units ψ′ can be obtained. pos ψ n ′ eg The relationship between this set and the original set is: By combining the sensitivity, economy, and carbon emissions of the generating units, the most suitable unit is selected to participate in power adjustment. The specific integration method can adopt a weighted average approach as follows:

[0106]

[0107] Where w1 is the weighted average coefficient of the unit's adjustment, w2 is the weighted average coefficient of carbon emissions, and w3 is the weighted average coefficient of economic efficiency, and the sum of w1, w2, and w3 is 1.

[0108] Finally, regarding z i The units are sorted from largest to smallest to select the top N (M) units for cross-sectional power adjustment, as shown below:

[0109]

[0110]

[0111] Step S4 specifically includes the following:

[0112] The specific selection method for power compensation units is as follows:

[0113]

[0114] This results in a new set of compensation units ψ′ pos (ψ n ′ eg The relationship here is the same as in step 3. Finally, the sensitivity, economy, and carbon emissions are integrated as follows:

[0115]

[0116] When performing cross-sectional power compensation, a sorting method from smallest to largest is adopted, and the first N (M) units are selected to participate in power compensation, as shown in the following figure:

[0117]

[0118] Step S5 specifically includes the following:

[0119] The process of adjusting the cross-sectional power can be regarded as a Korkov decision process, which can be represented as M={S,A,P,R,γ};

[0120] The state space S represents the state of the current power grid section after being affected by the corresponding action. For this invention, it includes the power of the current section, the target power value of the section, and the output of each generator in the region; it is represented as follows:

[0121] S = [P] init ,P tar p1, p2, ... pm (16)

[0122] Action Space A: The action space mainly represents the set of scalars controlled by the intelligent agent in the cross-sectional power adjustment. Here, the action represents the value of the unit's required output.

[0123] A = {△P} (17)

[0124] State transition probability P: The state transition probability refers to the probability that the power grid system will reach a certain state after being affected by an action issued by an agent, and is expressed as: S t+1 ←S t ×A, P∈[0,1];

[0125] Reward R: The reward represents the optimization objective of the deep reinforcement learning network. By continuously optimizing the network structure parameters, the cumulative reward value is maximized. In this patent, the main goal is to adjust the generator sets to ensure the convergence of power flow calculations and that the cross-sectional power reaches a certain target value. Therefore, R here can be expressed as the absolute value of the difference between the actual cross-sectional power and the target value, as follows:

[0126] R = -P t -P tar (18)

[0127] Reward discount factor γ: The discount factor represents the influence of time distance on the model, which can be expressed as: γ∈[0,1]. The larger γ is, the greater the time distance, the greater the influence of the model parameters.

[0128] In deep reinforcement learning models, deep neural networks serve as the policy function π(a|s) to fit the relationship between the cross-sectional state S and the action a. Therefore, the training objective of reinforcement learning networks is to continuously update the network parameters of the policy function to maximize the accumulated reward, maxR(τ), as follows:

[0129]

[0130] Where J(π) θ ) represents the maximum expected return for reinforcement learning, γ is the discount factor, and r is the cumulative reward value;

[0131] The training objective is to find a set of parameters for the network θ such that J(π) θ To maximize network parameters, it is only necessary to solve for the partial derivatives of each parameter within the network. To improve data utilization and agent stability, resampling is used during model training. Here, the PPO algorithm, a branch of deep reinforcement learning, is used to train the agent, and its optimization objective can be expressed as:

[0132]

[0133] Where θ * To optimize the objective, D KL For the KL divergence constraint, in the formula For π θ (·|s t )and The distance between them;

[0134] It should be noted that π θ (a t |s t ), These represent the target policy and the sampling policy, respectively. and The value represents the empirical estimate. In addition, to constrain the distance between the target policy and the sampling policy, KL divergence is used to calculate the expected distance between them and serve as a constraint term in the optimization problem. The hyperparameter λ is used to balance the KL divergence penalty term.

[0135] In practice, this is accomplished through the following steps:

[0136] Step 1: Generator Set Selection: For cross-sectional power adjustment, this patent comprehensively considers factors such as generator sensitivity, economy, and carbon emissions to ensure optimal generator operation in terms of both economy and carbon emissions while maintaining cross-sectional power flow convergence and quickly reaching the target value. After receiving the deep reinforcement learning output action, the specific process is as follows:

[0137] (1) Calculate the carbon emission rights of the relevant units, as shown in equation (4-5). If the carbon emissions of the relevant units exceed the carbon emission rights, they will not participate in the power adjustment.

[0138] (2) Calculate the sensitivity, economy and carbon emissions of the relevant units as shown in equations (1-2) and (6-7). First, based on action a, determine the set of units to be adjusted by prioritizing the selection of units with high sensitivity, as shown in equation (8-9). Then, calculate the weighted average of sensitivity, economy and carbon emissions, as shown in equation (10). Finally, select units with larger values ​​to participate in the cross-sectional power adjustment.

[0139] (3) The selection of power compensation units is similar to that of power adjustment units. Based on the adjusted power value, the power compensation unit set is determined by prioritizing low-sensitivity units, as shown in Equation (13). Then, the weighted average of the sensitivity, economy and carbon emissions of the unit is calculated, as shown in Equation (14). Finally, the unit with the smaller value is selected to participate in the cross-sectional power compensation.

[0140] Step 2: Deep Reinforcement Learning Module: This mainly consists of two parts: training and validation. During training, the model is repeatedly fed with the cross-section state, actions, reward values, and reward discount coefficients. The resulting actions are used to participate in cross-section power adjustment and power flow calculation. Data with power flow convergence is selected, and the strategy network is optimized in the direction of maximum accumulated reward value. Finally, a deep reinforcement learning model fitting the relationship between state and action is obtained. During validation, since the network parameters have been optimized, cross-section state values ​​are added to the model. Based on the actions output by the model, generating units are selected to participate in cross-section power adjustment. This verifies whether the model can quickly reach the target power value while ensuring power flow convergence and maintaining the overall economic efficiency and carbon emissions of the cross-section.

[0141] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A power grid section adjustment strategy based on deep reinforcement learning, characterized in that, It includes a generator set selection module and a deep reinforcement learning module. The generator set selection module mainly includes generator set screening and power compensation. According to the power flow direction, the two ends of the cross section are divided into the sending end and the receiving end. The power of the transmission cross section is the superposition of the power on the intermediate tie line between the sending end and the receiving end. The deep reinforcement learning module mainly refers to using a deep neural network as the policy network for reinforcement learning, thereby completing the mapping relationship between the cross section state and the required action. Specifically, this is achieved through the following five steps, including: S1. Calculate the sensitivity of the generator sets: While ensuring the convergence of the power flow calculation, reduce the adjustment data of the generator sets. Here, we calculate the sensitivity of each generator set. S2. Determine the economic efficiency and carbon emissions of the generator: When selecting a generator set for power regulation, in addition to considering the sensitivity of the generator set, its economic efficiency and carbon emissions also need to be considered. Here, the concept of carbon emission rights is introduced. Carbon emission rights refer to the amount of carbon emissions allowed by the generator set under the premise of the annual power generation hours. S3. Determine the power regulating units: After determining the sensitivity of each generator unit, the generator units can be divided into a set of forward regulating units. Reverse regulation unit collection Then, based on the actions learned from reinforcement learning, a set of units will be selected to participate in the cross-sectional power adjustment. S4. Determine the power compensation unit: The method for determining the power compensation unit is generally similar to the power adjustment method. When the unit adjusts the power in the forward direction, the reverse power compensation is performed; when the unit adjusts the power in the reverse direction, the forward power compensation is performed. S5. Establish a deep reinforcement learning module: When training with deep reinforcement learning, the agent learns by continuously interacting with the environment, thereby obtaining the best cross-sectional power compensation strategy. The generator selection module specifically includes: The selection of generators considers not only the sensitivity of the generator set but also its operational economy and carbon emissions. Generator sets are ranked from largest to smallest based on these three factors, and those ranked higher are prioritized for power adjustment during cross-sectional power adjustments. Secondly, to avoid power imbalance after grid cross-sectional power adjustments, a sending-end adjustment and receiving-end compensation approach is adopted. The sending end selects the corresponding generator set for cross-sectional power adjustment based on the power difference, while the receiving end maintains the power balance of the smart grid system through reverse power output. The deep reinforcement learning module specifically includes: based on the Markov decision process quintuple The composition requirements dictate that power grid power regulation modeling primarily focuses on cross-sectional conditions. ,action ,award and discount factor The settings; during the training phase, the reward function is set. With discount factor Under the premise of inputting the cross-sectional state into the reinforcement learning module Actions output by the policy network The cross-sectional power is regulated, and the power flow calculation is performed to obtain the new cross-sectional state. While ensuring the convergence of power flow calculation, the reinforcement learning module is input, the above actions are repeated, and relevant data is collected to complete the optimization process of the policy network. In the prediction stage, since the fitting of the policy network has been completed, the power adjustment action is obtained by continuously inputting the state into the reinforcement learning model, until the cross-sectional power reaches the target value while ensuring the convergence of power flow.

2. The power grid section adjustment strategy based on deep reinforcement learning according to claim 1, characterized in that, Step S1 specifically includes the following: The specific calculation formula for the sensitivity of each unit is as follows: (1) (2) in: This refers to the collection of all units except the balancing unit; This represents the current power output of the power grid section. and For the unit The change in grid cross-sectional power at maximum and minimum power; and For generator The sensitivity indicators correspond to forward and reverse regulation. It should be noted that when regulating forward, the reverse unit is used for compensation, and when regulating reverse, the forward unit is used for compensation. For regulation, the unit with higher sensitivity is preferred, while for power compensation, the unit with lower sensitivity is preferred.

3. The power grid section adjustment strategy based on deep reinforcement learning according to claim 1, characterized in that, Step S2 specifically includes the following: The calculation of carbon emission rights can be expressed as follows: (3) in: For the unit Carbon emission rights, in units of ; The carbon emission intensity of the corresponding unit is expressed in units of... ; The unit is the power generation time of the corresponding unit. ; Installed capacity, unit: ; To achieve energy conservation and emission reduction, before considering the participation of the unit in power adjustment, it is necessary to calculate its carbon emission rights. By evaluating the unit's output time, the future carbon emissions of the unit are calculated and compared with the carbon emission rights. If it is less than the carbon emission rights, the unit participates in power adjustment and is recorded as 1; otherwise, it does not participate and is recorded as 0. (4) (5) in, The carbon emissions after the group participates in power adjustment; This represents the unit's current carbon emissions; This is the maximum operating power of the unit; This refers to the unit's operating time; This is the carbon emission rights factor, with a value ranging from 0 to 1; For a smart grid, it contains various types of generating units, including thermal power units, natural gas units, and renewable energy units. Renewable energy units do not participate in cross-sectional power regulation. Thermal power units have high carbon emissions but low costs. Natural gas units have high power generation costs but low carbon emissions. Therefore, when selecting units for cross-sectional power regulation, in addition to considering sensitivity, economic efficiency and carbon emissions must be comprehensively considered. Regarding economic efficiency, currently only the cost of raw materials is considered, as detailed below: (6) in, The power generation cost of the corresponding generating unit; The raw materials required for a unit of electricity generation, here referring to coal and natural gas; This refers to the power generation efficiency of the corresponding generating unit; For raw material prices; Specifically, carbon emissions are expressed as follows: (7) in, This refers to the carbon emissions of the corresponding unit; , The meaning is the same as that of equation (6); Carbon emissions per unit of fuel.

4. The power grid section adjustment strategy based on deep reinforcement learning according to claim 1, characterized in that, Step S3 specifically includes the following: Economic efficiency and carbon emissions also need to be considered, therefore the set of participants in power adjustment needs to be expanded, as shown in the following equation: (8) (9) in: The action values ​​output by the reinforcement learning module; When the target value is greater than the current power of the cross section Positive adjustment is performed using equation (8); when the target value is less than the current cross-sectional power... Reverse adjustment is performed using equation (9); The number of units participating in positive regulation, The number of units participating in reverse regulation; from this, a new set of units is derived. , The relationship between this set and the original set is: , By combining the sensitivity, economy, and carbon emissions of the generating units, the most suitable unit is selected to participate in power adjustment. The specific integration method can adopt a weighted average approach as follows: (10) Where w1 is the weighted average coefficient of the unit's adjustment, w2 is the weighted average coefficient of carbon emissions, and w3 is the weighted average coefficient of economic efficiency, and the sum of w1, w2, and w3 is 1. Finally, Sort by size from largest to smallest, and then select the top... Each generating unit participates in the cross-sectional power adjustment, as detailed below: (11) (12)。 5. The power grid section adjustment strategy based on deep reinforcement learning according to claim 1, characterized in that, Step S4 specifically includes the following: The specific selection method for power compensation units is as follows: (13) This will result in a new set of compensation units. ( The relationship here is the same as in step 3. Finally, the sensitivity, economy, and carbon emissions are integrated as follows: (14) When compensating for cross-sectional power, a sorting method from smallest to largest is adopted, and the first [unit] is selected. Each generating unit participates in power compensation, as detailed below: (15)。 6. The power grid section adjustment strategy based on deep reinforcement learning according to claim 1, characterized in that, Step S5 specifically includes the following: Viewing the cross-sectional power adjustment process as a Korkov decision process, it can be represented as follows: ; The state space This indicates the state of a current power grid section after being affected by a corresponding action, including the current power of the section, the target power value of the section, and the output of each generator in the area; it is represented as follows: (16) Action space Action space mainly represents the set of scalars controlled by the intelligent agent in the cross-sectional power adjustment. Here, action represents the value of power output required by the unit. (17) State transition probability State transition probability refers to the probability that a power grid system will reach a certain state after being affected by an action issued by an agent, and is expressed as: , ; award The reward represents the optimization goal of the deep reinforcement learning network. By continuously optimizing the network structure parameters, the cumulative reward value is maximized. This is mainly achieved by adjusting the generator sets to ensure the convergence of power flow calculations and that the cross-sectional power reaches a certain target value. Therefore, the reward here... It can be expressed as the absolute value of the difference between the actual cross-sectional power and the target value, as follows: (18) Reward Discount Coefficient The discount factor represents the impact of time proximity on the model, and can be expressed as: when The larger the value, the greater the influence of points at a greater time interval on the model parameters; In deep reinforcement learning models, deep neural networks serve as the policy function. Used to fit the cross-sectional state With action The relationship between these factors leads to the conclusion that the network training objective of reinforcement learning is to maximize the accumulated reward by continuously updating the network parameters of the policy function. , means as follows: (19) in To enhance the maximum expected return on learning, For discount factor, The cumulative value of the rewards; The training objective is to find a set of parameter networks. , making To maximize network parameters, it is only necessary to solve for the partial derivatives of each parameter within the network. To improve data utilization and agent stability, resampling is used during model training. Here, the PPO algorithm, a branch of deep reinforcement learning, is used to train the agent, and its optimization objective can be expressed as: (20) in To optimize the objective, For the KL divergence constraint, in the formula for and The distance between them; It should be noted that, , These represent the target policy and the sampling policy, respectively. and The values ​​represent empirical estimates. Furthermore, to constrain the distance between the target policy and the sampling policy, KL divergence is used to calculate the expected distance between them, which serves as a constraint term in the optimization problem; hyperparameters. Used to balance the KL divergence penalty term.

Citation Information

Patent Citations

  • Power grid transmission section power flow adjustment unit selection method based on deep learning

    CN115021270A