A method for automatic adjustment of power grid operation mode based on reinforcement learning

Through the combination of expert systems and reinforcement learning models, the active power adjustment and trend adjustment of thermal power units are optimized, which solves the high cost of exploration and new energy consumption problems in automatic adjustment of power grid operation mode, and achieves the safe and stable operation of the power grid and the maximum consumption of new energy.

CN114865714BActive Publication Date: 2025-08-08XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210456909.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-08-08
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

The existing reinforcement learning model has the problem of exponential growth in the state space and action space in the automatic adjustment of the power grid operation mode, which has led to a sharp increase in exploration costs and cannot meet the convergence requirements of trend calculations, making it difficult to effectively deal with the balance and consumption problems caused by the uncertainty of source and load on both sides of the high proportion of new energy power systems.

Method used

By designing an expert system and combining with the reinforcement learning model, the total amount of active power adjustment of thermal power units is determined, and the power switch is performed when the action space exceeds the limit, the current adjustment and the voltage adjustment of the machine end are performed, the key units are identified using the active power-line load rate sensitivity matrix, the unit regulation sequence is optimized, and the comprehensive evaluation indicators are used as rewards to reduce the exploration cost.

Benefits of technology

It realizes automatic adjustment of the power grid operation mode, ensures the safe and stable operation of the power grid, improves the consumption capacity of new energy, and reduces the operating costs of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114865714B_ABST
    Figure CN114865714B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for automatically adjusting the operation mode of a power grid based on reinforcement learning, comprising determining the total active power adjustment amount of the thermal power units at the next moment; if the action space of each thermal power unit is within the output adjustment range, the total active power adjustment amount is allocated to each thermal power unit according to the optimal unit control sequence; if the action space of each thermal power unit is lower than the lower limit of the thermal power unit action space or higher than the upper limit of the thermal power unit action space, after the power on / off operation, the total active power adjustment amount is allocated to each thermal power unit according to the optimal unit control sequence; after the allocation is completed, the power flow adjustment amount is redistributed according to line overload or critical overload, and the terminal voltage is adjusted; the optimal unit control sequence is obtained through a reinforcement learning model. This method can realize automatic adjustment of the power grid operation mode, ensure the safe and stable operation of the power grid, and achieve maximum absorption of new energy. Through the reinforcement learning model, the efficiency of obtaining the optimal unit control sequence is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of power system dispatching, and in particular to a method for automatically adjusting power grid operation modes based on reinforcement learning. Background Art

[0002] The operational challenges of new power systems dominated by renewable energy are becoming increasingly prominent. In the context of a massive influx of renewable energy, the power grid has evolved from a single or limited number of optimization objectives to a complex, multi-layered, multi-regional optimization problem. Adjusting the grid's operating mode is the most demanding and repetitive component of the calculation. Traditional manual adjustments are not only time-consuming and labor-intensive, but also impose relatively fixed renewable energy output and load settings, making it difficult to address the balance and absorption challenges inherent in the actual operational scenarios of power systems with a high proportion of renewable energy.

[0003] In recent years, with the development of artificial intelligence (AI) technology, reinforcement learning has gradually been applied to the automatic adjustment of power grid operation modes. Reinforcement learning explores the state and action spaces, using the information gained during exploration to update the action utility function, thereby generating experience to guide the automatic adjustment of power grid operation modes. However, the size of the state and action spaces of reinforcement learning models increases exponentially with the number of system nodes. This exponential growth in state and action spaces dramatically increases the exploration cost. Furthermore, power systems, especially complex ones, have high requirements for their operation modes. However, during the training of reinforcement learning models, the randomly generated new power grid operation modes often fail to meet the convergence requirements of power flow calculations, generating invalid operation modes and extremely low exploration efficiency. Therefore, directly using traditional reinforcement learning models for automatic adjustment of power grid operation modes still presents significant challenges. Summary of the Invention

[0004] In view of this, the main purpose of this application is at least to provide a specially designed expert system to solve the problems existing in the existing reinforcement learning model in the automatic adjustment of the power grid operation mode, to deal with the balance and absorption problems brought about by the uncertainty on both the source and load sides in the high proportion of new energy power systems, and to provide a new technical solution to realize the automatic adjustment of the power grid operation mode.

[0005] Based on the above objectives, the present invention proposes a method for automatically adjusting the operation mode of a power grid based on reinforcement learning, which includes the following steps:

[0006] Determine the total active power adjustment of the thermal power unit at the next moment;

[0007] If the action space of each thermal power unit is within the output adjustment range, the total active power adjustment amount will be allocated to each thermal power unit according to the optimal unit control order;

[0008] If the action space of each thermal power unit is lower than the lower limit of the action space of the thermal power unit or higher than the upper limit of the action space of the thermal power unit, then after the power on / off operation, the total active power adjustment amount will be allocated to each thermal power unit according to the optimal unit control sequence;

[0009] After the sharing is completed, the power flow adjustment amount is redistributed according to the line overload or critical overload, and the terminal voltage is adjusted;

[0010] The optimal unit control sequence is obtained through a reinforcement learning model.

[0011] In the above technical solution, the method enables automatic adjustment of the grid's operating mode, effectively resolving the balance and absorption issues caused by uncertainty in both the source and load sides of power systems with a high proportion of renewable energy. This ensures the safe and stable operation of the grid and maximizes the absorption of renewable energy. Using a reinforcement learning model, the efficiency of exploring the optimal unit control sequence can be improved.

[0012] As a further improvement to the above technical solution, in the method, the apportioned system is inspected to see whether each apportioned system verification line is overloaded or critically overloaded, and the flow adjustment amount is redistributed for the main units involved in the overloaded or critically overloaded lines to improve the safety of power grid operation. The redistribution of the flow adjustment amount includes the following steps:

[0013] Identify critical units for line load factor;

[0014] If the key unit is a new energy unit, when the load rate is greater than a first set threshold, the output of the new energy unit is reduced to a first set value; when the load rate is greater than 1 and less than or equal to the first set threshold, if the overload persists after the number of consecutive reductions reaches a set number, the output of the new energy unit is reduced to a second set value;

[0015] If the key unit is a thermal power unit, the output of the thermal power unit shall be reduced to the lower output limit of the unit.

[0016] As a further improvement to the above technical solution, in the method, the key units are determined by an active power-line load factor sensitivity matrix, thereby quickly and accurately determining the control sequence of overloaded lines or basic units, including:

[0017] Extract the row vector of the active power-line load factor sensitivity matrix;

[0018] Filter the components corresponding to the node where the unit is located;

[0019] The node-mounted unit corresponding to the component with the largest absolute value is determined as the key unit;

[0020] The active power line load rate sensitivity matrix is an m×n order matrix, where m is the number of power system branches and n is the number of power system nodes.

[0021] As a further improvement to the above technical solution, in the method, the optimal unit control sequence is obtained by inputting a basic unit control sequence into a reinforcement learning model; the basic unit control sequence is obtained by summing the column vectors of an active power-line load factor sensitivity matrix and then sorting it; the active power line load factor sensitivity matrix is an m×n matrix, where m is the number of power system branches and n is the number of power system nodes. The unit control sequence with the highest probability of obtaining the maximum reward, discovered during the training process of the reinforcement learning model, is used.

[0022] As a further improvement of the above technical solution, in the method, the active power-line load rate sensitivity matrix is extracted based on historical operating data when all units are fully powered on and there are no disconnections in the grid, so that the identification of key units and the judgment of overload are closer to the real power grid, and the automatic adjustment of the power grid is safe, effective and stable.

[0023] As a further improvement to the above technical solution, in the described method, the reinforcement learning model uses the unit control sequence as the state of the intelligent agent, the two positions in the sequence as the intelligent agent's actions, and a comprehensive evaluation index as the reward. The influencing factors of the comprehensive evaluation index include the relative absorption of renewable energy, line over-limit conditions, unit output constraints, node voltage constraints, and operating economic costs. This ensures that the optimal unit control sequence can maximize the absorption of renewable energy while ensuring the safe operation of the power grid, improve renewable energy utilization, and thus reduce power grid operating costs. Furthermore, the model only needs to learn a two-dimensional discrete action vector composed of two scalar coordinates, making convergence relatively simple.

[0024] As a further improvement to the above technical solution, in the method, the effectiveness of each exploration output grid operation mode is evaluated through reward feedback, thereby improving exploration efficiency and converting the exponential growth of exploration cost into linear growth; the reward is calculated by the following formula:

[0025]

[0026] Where: R is the reward; r i Get the value for the reward points;

[0027] When i=1,

[0028] Where: renewable t+1,j is the output of the jth group of new energy generators at time t+1; is the output upper limit of the jth group of new energy generating units at time t+1; Re is the number of new energy generating units;

[0029] When i≠1,

[0030] Where A represents a constraint; when i is 2, the constraint is the line current; when i is 3, the constraint is the unit output; when i is 4, the constraint is the node voltage; when i is 5, the constraint is the operating economic cost; the subscripts max and min represent the upper and lower limits of the corresponding constraints, respectively.

[0031] As a further improvement of the above technical solution, in the method, the total adjustment amount of the active power of the thermal power unit at the next moment is determined by the following formula:

[0032] Δthermal=thermal t+1 -thermal t

[0033] Where: thermal t is the thermal power output at the current moment t, thermal t+1 To provide thermal power for the next moment;

[0034] thermal t+1 Calculated by the following formula:

[0035]

[0036] Where:

[0037] L is the total number of loads, l is the variable of the number of loads, Re is the number of new energy units, and j is the variable of the number of new energy units;

[0038] is the total load at time t+1;

[0039] renewable t+1,j is the output of the jth group of new energy generators at time t+1;

[0040] balance t+1 The output of the balancing machine at time t+1;

[0041] loss t+1 is the network loss power at the next moment, calculated by the following formula:

[0042] loss t+1 =loss t Lfactor

[0043] Where Lfactor is the network loss estimation coefficient, which is calculated by the following formula:

[0044]

[0045] As a further improvement to the above technical solution, in the method, when the action space of the i-th thermal power unit exceeds the lower limit or upper limit of the thermal power action space, the power on / off operation is performed based on the total active power adjustment of the unit, taking into account the unit control sequence, unit capacity, and network parameters, to ensure that network losses are maintained at a low level. The power on / off operation includes:

[0046] When load fluctuations cause the thermal power adjustment to exceed the upper limit of the thermal power unit's ramp constraint, the thermal power units are started up according to the sensitivity of the line load factor from small to large. The power provided by the number of started thermal power units can compensate for the part of the thermal power adjustment that exceeds the upper limit of the ramp constraint.

[0047] When load fluctuations cause the thermal power adjustment to fall below the lower limit of the thermal power unit's ramp constraint, the thermal power units are shut down according to the sensitivity of the line load rate from large to small. The power reduction caused by shutting down the number of thermal power units can offset the fact that the thermal power adjustment falls below the lower limit of the thermal power unit's ramp constraint.

[0048] When the ratio of the actual processing to the maximum processing of all running generators exceeds a second set threshold, the generators are started up in order of sensitivity of the line load rate from small to large, so that the ratio is less than the second set threshold;

[0049] When the ratio of the actual processing to the maximum processing of all running generators is lower than the third set threshold, they are shut down in descending order of sensitivity of line load rate so that the ratio is greater than the third set threshold.

[0050] As a further improvement to the above technical solution, in the method, after the thermal power unit output is adjusted, the voltage of the generator set is adjusted to control the reactive power within the range of [-180, 100] to ensure the normal operation of the power grid and minimize network losses. The generator end voltage adjustment includes:

[0051] The voltage of the generator set is recorded as U k , and the reactive power is recorded as Q k , where k represents the generator group identifier;

[0052] If Q k ≥100, then U k Use U k The value after -0.01 is updated;

[0053] If 60≤Q k <100, then U k Use U k The value after -0.004 is updated;

[0054] If -90<Q k <60, then U k constant;

[0055] If -180<Q k ≤-90, then U k Use U k The value after +0.0015 is updated;

[0056] If Q k ≤-180, then U k Use U k The value after +0.01 is updated. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0058] Figure 1 、 one A schematic diagram of an embodiment of the combined application of expert system and reinforcement learning;

[0059] Figure 2 、 one A schematic diagram comparing the performance of a reinforcement learning model using only a reinforcement learning model and a reinforcement learning model using the disclosed method in an embodiment;

[0060] Figure 3 、 one Schematic diagram of the normal scene adjustment effect in an embodiment;

[0061] Figure 4 、 one Schematic diagram of the adjustment effect of extreme scenarios in an embodiment. DETAILED DESCRIPTION

[0062] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0063] The terms "first," "second," and "third" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or to implicitly specify the quantity of the technical features indicated. Therefore, a feature specified as "first," "second," or "third" may explicitly or implicitly include one or more of such features.

[0064] In Example 1, an expert system and a reinforcement learning model are implemented based on the method disclosed herein, and the two are combined to automatically adjust the grid operation mode. The expert system ensures the effectiveness of each exploration output of the grid operation mode, greatly improving exploration efficiency and converting the exponential growth of the exploration cost of the reinforcement learning model into a linear growth. The unit control sequence with the highest probability of obtaining the maximum reward discovered during the reinforcement learning training process guides the expert system to automatically adjust the grid, thereby achieving maximum absorption of new energy while ensuring the safe and stable operation of the grid.

[0065] In the expert system, the following method steps are implemented:

[0066] (1.1) Identify the total load at the next moment Where L is the total number of loads, l is the load number variable;

[0067] (1.2) Identify the sum of the upper limits of the output of new energy units at the next moment Where Re is the number of new energy units, j is the variable of the number of new energy units;

[0068] (1.3) If the output of each new energy unit at the next moment is set to its maximum value, then

[0069] (1.4) Calculate the network loss estimation coefficient Lfactor:

[0070]

[0071] Where: L is the total number of loads, l is the load number variable;

[0072] (1.5) Based on the loss of the previous moment t And the network loss estimation coefficient Lfactor, calculate the network loss power loss at the next moment t+1 :

[0073] loss t+1 =loss t Lfactor

[0074] (1.6) Balance the output of the balancing machine at the next moment t+1 Set it to the arithmetic mean of its upper and lower limits, leaving sufficient margin.

[0075] (1.7) Calculate the expected total thermal power output at the next moment t+1 :

[0076]

[0077] (1.7) The total active power adjustment of the thermal power unit at the next moment is determined by the following formula:

[0078] Δthermal=thermal t+1 -thermal t

[0079] Where: thermal t is the thermal power output at the current moment t, thermal t+1 To provide thermal power for the next moment.

[0080] For the number of thermal power units T, the kth thermal power unit G k , its action space ΔG k , there is a lower limit low k , low k <0, upper limit high k ,Right now:

[0081] low k <ΔG k <high k

[0082] For all thermal power units, obtain the action space of each unit. If each unit is in a reasonable output adjustment range, the total active power adjustment amount will be allocated to each thermal power unit according to the unit control order; otherwise, if the action space of each thermal power unit is lower than the lower limit of the action space of the thermal power unit or higher than the upper limit of the action space of the thermal power unit, then after the power on and off operation, the total active power adjustment amount will be allocated to each thermal power unit according to the unit control order.

[0083] When the total active power adjustment is allocated to each thermal power unit, when Δthermal>0, all thermal power units G k Set to lower limit low k ,Right now:

[0084]

[0085] The obtained Δthermal * Distribute in sequence according to the optimal unit control order. When Δthermal<0, all thermal power units G k Set to the lower limit high k ,Right now:

[0086]

[0087] The obtained Δthermal * Distribute in reverse order according to the optimal unit control sequence.

[0088] After the completion of the apportionment, the flow adjustment amount is redistributed according to the line overload or critical overload, that is, after the thermal power unit output is adjusted, the voltage u of the generator unit is adjusted. k Adjust the method to control the reactive power Q k The range is [-180, 100] to ensure the normal operation of the power grid and minimize network losses. The voltage of the generator set is recorded as u k , and the reactive power is recorded as Q k , where k represents the generator set identifier; the generator terminal voltage adjustment includes:

[0089] If Q k ≥100, then U k Use U k The value after -0.01 is updated;

[0090] If 60≤Q k <100, then U k Use U k The value after -0.004 is updated;

[0091] If -90<Q k <60, then U k constant;

[0092] If -180<Q k ≤-90, then U k Use U k The value after +0.0015 is updated;

[0093] If Q k ≤-180, then U k Use U k The value after +0.01 is updated.

[0094] In Example 1, by setting the line load rate alarm threshold, when the line current load rate exceeds the alarm threshold, it is identified as an overloaded line. When an overloaded line appears in the system, it is necessary to identify the overloaded line and find the key generator group G that affects the line overload based on the overloaded line. key .

[0095] The sum of the generator power and load at each node is defined as the net injected power of the node. Since the load rate ρ and the net injected active power P and net injected reactive power Q of the node are approximately linearly related, the following relationship exists:

[0096] Δρ=Hp ΔP+H Q ·ΔQ (1)

[0097] Where: H p Active power injected into the node-line load sensitivity matrix, H Q is the node injected reactive power-line load factor sensitivity matrix, Δρ is the line load factor variation matrix, ΔP is the node injected active power adjustment matrix, and ΔQ is the node injected reactive power adjustment matrix.

[0098] Since the influence of ΔQ on the load rate is small, it is ignored, that is, formula (1) becomes:

[0099] Δρ≈H p ·ΔP (2)

[0100] From numerical simulations or actual operation and maintenance, we obtain massive historical operation data, extract typical operation scenarios where all units are fully powered on and there are no disconnections in the grid, and generate sampling data: the node injected active power adjustment matrix ΔP and the line load factor change matrix Δρ.

[0101] Δρ=[Δρ1, Δρ2,..., Δρ x ], ΔP = [ΔP1, ΔP2,..., ΔP x ], x is the number of sampling times.

[0102] The active power-line load sensitivity matrix H in formula (2) is p , solve using the least squares method:

[0103] H p =Δρ(ΔP T ΔP) -1 ΔP T

[0104] Where: H p is an m×n matrix, where m is the number of system branches and n is the number of system nodes. p The row vector where the overloaded line is located is used to select the component corresponding to the node where the unit is located. The unit mounted on the node corresponding to the component with the largest absolute value is the key unit affecting the overloaded line.

[0105] If the key unit is a thermal power unit, the output of the thermal power unit is reduced to the lower output limit of the unit. If the key unit is a new energy unit, when the load rate is greater than the first set threshold, the output of the new energy unit is reduced to the first set value; when the load rate is greater than 1 and less than or equal to the first set threshold, if the overload persists after the number of consecutive reductions reaches the set number, the output of the new energy unit is reduced to the second set value. The first set threshold can be 1.1, 1.2, 1.3, etc., the first set value can be 9%, 10%, 11%, 12%, etc., the second set value can be 25%, 30%, 35%, etc., and the number of iterations can be 2, 3, 4, 5, etc., to ensure the safe and stable operation of the power grid and achieve maximum absorption of new energy.

[0106] The use of power-on and power-off operations can ensure that network losses are maintained at a low level. Based on the network parameter information of network topology, line capacity and line admittance, the startup sequence is specified, and the thermal power units closer to the load in the network are started first. The reverse sequence will give priority to shutting down the thermal power units farther away from the load in the network.

[0107] The system will be powered on or off when one of the following two situations occurs:

[0108] The first case: When the load fluctuates greatly and the new energy has reached its maximum absorption, the amount of thermal power required to be adjusted exceeds the ramp constraint range of the thermal power unit. Consider the thermal power on / off operation to ensure power balance.

[0109] In the second scenario, when the ratio of the sum of the actual output of all operating thermal power units to the sum of the output limits exceeds the second threshold or falls below the third threshold, a power on / off operation is considered. In the sum of the actual outputs of all operating thermal power units, both the actual output and the output limit contributed by the shutdown units are zero. The second or third threshold can be adjusted based on the actual operating conditions of the power system.

[0110] For the first case:

[0111] (I) When the load fluctuation causes the thermal power adjustment amount to exceed the upper limit of the thermal power unit's ramp constraint, that is, The thermal power units should be started up. The order of starting up is from small to large according to the sensitivity of the line load rate. The smaller the impact on the line load rate, the higher the starting priority. The number of thermal power units started up provides power Δthermal open , when the required adjustment amount of thermal power compensation exceeds the upper limit of the ramp, the startup can be terminated, that is:

[0112]

[0113] (II) When the load fluctuation causes the thermal power adjustment amount to be lower than the lower limit of the thermal power unit's ramp constraint, that is, Shut down the thermal power units. The shutdown order is the reverse order of the startup order, and the sensitivity of each thermal power unit to the line load rate is ranked from large to small. The greater the impact on the line load rate, the higher the shutdown priority. As with startup, the number of units shut down depends on the power reduction Δthermal close , which can offset the thermal power adjustment amount being lower than the ramp constraint lower limit of the thermal power unit, that is, to ensure:

[0114]

[0115] For the second case:

[0116] (III) When the ratio of the actual processing of all running generators to the maximum processing exceeds the second set threshold, the running generators are under heavy load and need to be started to share part of the load. The startup sequence is also followed. After startup, the ratio of the actual processing of all running generators to the maximum processing is less than the second set threshold, and the operation can be terminated.

[0117] (IV) When the ratio of the actual processing of all running generators to the maximum processing is lower than the third set threshold, the load is not large at this time. When the running generators are in a low-load situation, they need to be shut down. The shutdown sequence is performed in the reverse order of the startup sequence. After the shutdown, the ratio of the actual processing of all running generators to the maximum processing is greater than the third set threshold, and the process can be terminated.

[0118] In Example 1, the optimal unit control sequence is obtained using a reinforcement learning model. In this model, the unit control sequence is used as the agent's state S, and two position coordinates in the sequence are used as the agent's action A. At each time step, the agent transitions from the old state to the new state by swapping the unit positions at these two coordinates.

[0119] The factors influencing the use of comprehensive evaluation indicators include the relative absorption capacity of renewable energy, line over-limit conditions, unit output constraints, node voltage constraints, and operating economic costs. This allows the optimal unit control sequence to meet the requirements of maximizing the absorption of renewable energy while ensuring the safe operation of the power grid, improving the utilization rate of renewable energy, and thus reducing the operating costs of the power grid. Therefore, a feasible implementation of the reward is:

[0120]

[0121] Where: R is the reward; r i Get the value for the reward points;

[0122] When i=1,

[0123] Where: renewable t+1,j is the output of the jth group of new energy generators at time t+1; is the output upper limit of the jth group of new energy generating units at time t+1; Re is the number of new energy generating units;

[0124] When i≠1,

[0125] Where A represents a constraint; when i is 2, the constraint is the line current; when i is 3, the constraint is the unit output; when i is 4, the constraint is the node voltage; when i is 5, the constraint is the operating economic cost; the subscripts max and min represent the upper and lower limits of the corresponding constraints, respectively.

[0126] During model training, the agent swaps the positions of units at two random indices in the unit control sequence and outputs a new control sequence. The basic unit control sequence is input into the agent of the reinforcement learning model, which outputs the optimal unit control sequence. The method of Example 1 is then adjusted during grid operation based on the optimal unit control sequence. The agent's reward is calculated based on the adjusted system power flow.

[0127] Specifically, the reinforcement learning model's learning output is the action utility function Q: (S, A) → R. If the current (S, A) combination has not been explored, that is, there is no relevant information in Q, then two positions are randomly generated to form a random action A for exploration; if the current (S, A) combination has been explored, Q is updated using the following formula:

[0128] Q(S, A)←(1-α)Q(S, A)+α[R(S, a)+γmax a Q(S′, a)]

[0129] Among them, α is the learning rate and γ is the discount factor.

[0130] After the training is completed, the action utility function Q: (S, A) → R is rolled up into the state evaluation function V: S → R, and the unit control order corresponding to the state with the highest score is selected. This order is the final optimized unit control order.

[0131] The basic unit control sequence used in the reinforcement learning model is obtained through the following steps:

[0132] Active power-line load sensitivity matrix H p The column vectors are summed up and sorted from large to small. The relative order of each generator unit in this sorting is the basic unit control order.

[0133] In Example 2, setting the alarm threshold to less than 1 and simultaneously identifying overloaded and critically overloaded lines facilitates early protection actions, thereby improving the robustness of the control strategy. This sequence is written into the expert system to complete the closed loop.

[0134] In Example 3, after the disclosed method is implemented using Python language, the following example scenario is set: the IEEE standard example case 118 system grid is used. The system contains 118 nodes, 54 generator sets, 186 transmission lines and 91 loads, 18 of which are set as new energy units. According to the new energy output and load fluctuation characteristics, 8760 hours of new energy output and load data are generated by random simulation. The length of each time step is 5 minutes. A section is randomly selected as the starting section in each round, and the total reward accumulated within 288 consecutive time steps is used to evaluate the automatic flow adjustment scheme. If the flow cannot converge, the round is ended in advance. The reinforcement learning model DDPG (Deep Deterministic Policy Gradient) is used as the reinforcement learning model this time.

[0135] (1) Comparison of reinforcement learning models with and without expert systems

[0136] Figure 2 This is a performance comparison chart of the reinforcement learning model when testing whether the expert system is introduced into the example.

[0137] When the expert system is not introduced, the reinforcement learning model needs to directly learn the active output adjustment and terminal voltage adjustment of 54 generator sets, that is, a 108-dimensional continuous action vector, which is extremely difficult to converge. Figure 2 In the results, the performance of the model did not improve significantly after more than 600 rounds of training. In addition, when the reinforcement learning model directly randomly explores the operation mode of the power grid, the probability of exploring an effective operation mode is low. Figure 2 This is reflected in the fact that in more than 600 rounds of training, the score of the reinforcement learning model directly did not exceed 100 points, hovering at an extremely low level.

[0138] When an expert system was introduced and a reinforcement learning model was integrated with it, model performance improved significantly. This improvement stemmed from two factors: first, the reinforcement learning model indirectly influenced the grid's operating mode by guiding the expert system. The specific operating mode was generated by the expert system, ensuring high quality, with scores exceeding 400 at the outset of training. Second, the reinforcement learning model only needed to learn a two-dimensional discrete action vector consisting of two scalar coordinates, making convergence relatively simple. The model converged after just over 300 training rounds.

[0139] (2) Operation effect under normal scenarios

[0140] Figure 3 This diagram shows the effects of this automatic grid operation adjustment method under normal conditions. Under normal conditions, load fluctuations and renewable energy output fluctuations are relatively smooth. This adjustment method can fully accommodate renewable energy output while ensuring safe and stable grid operation.

[0141] (3) Operational effects in extreme scenarios

[0142] Figure 4 This diagram illustrates the effectiveness of this automatic grid operation adjustment method in an extreme scenario. In this extreme scenario, the load decreases rapidly while the output of renewable energy generators increases dramatically. In this scenario, to ensure grid stability, the output of renewable energy generators cannot be fully absorbed. This adjustment method initially curtails some wind and solar power, then promptly adjusts towards full absorption, achieving maximum absorption of renewable energy output while ensuring safe and stable grid operation.

[0143] Through the above description of the embodiments, those skilled in the art will clearly understand that the disclosed method can be implemented using software plus necessary general-purpose hardware. Of course, it can also be implemented using dedicated hardware, including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, any function performed by a computer program can be easily implemented using corresponding hardware. Moreover, the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present disclosure, software program implementation is often the preferred embodiment.

[0144] Although the embodiments of the present invention have been described above with reference to the accompanying drawings, the present invention is not limited to the above-mentioned specific embodiments and application fields. The above-mentioned specific embodiments are merely illustrative and instructive, and are not restrictive. A person skilled in the art, guided by this specification and without departing from the scope of protection of the claims of the present invention, may also devise various forms, all of which fall within the scope of protection of the present invention.

Claims

1. A method for automatically adjusting the operation mode of a power grid based on reinforcement learning, characterized in that: The method comprises the following steps: Determine the total active power adjustment of the thermal power unit at the next moment; If the action space of each thermal power unit is within the output adjustment range, the total active power adjustment amount will be allocated to each thermal power unit according to the optimal unit control order; If the action space of each thermal power unit is lower than the lower limit of the thermal power unit action space or higher than the upper limit of the thermal power unit action space, then after the power on / off operation, the total active power adjustment amount will be allocated to each thermal power unit according to the optimal unit control sequence; After the sharing is completed, the power flow adjustment amount is redistributed according to the line overload or critical overload, and the terminal voltage is adjusted; The optimal unit control sequence is obtained through a reinforcement learning model.

2. The method according to claim 1, characterized in that The power flow adjustment amount redistribution includes the following steps: Identify critical units for line load factor; If the key unit is a new energy unit, when the load rate is greater than a first set threshold, the output of the new energy unit is reduced to a first set value; when the load rate is greater than 1 and less than or equal to the first set threshold, if the overload persists after the number of consecutive reductions reaches a set number, the output of the new energy unit is reduced to a second set value; If the key unit is a thermal power unit, the output of the thermal power unit shall be reduced to the lower output limit of the unit.

3. The method according to claim 2, characterized in that The key units are determined by the active power-line load rate sensitivity matrix, including: Extract the row vector of the active power-line load factor sensitivity matrix; Filter the components corresponding to the node where the unit is located; The node-mounted unit corresponding to the component with the largest absolute value is determined as the key unit; The active power-line load rate sensitivity matrix is an m×n order matrix, where m is the number of power system branches and n is the number of power system nodes.

4. The method according to claim 1, wherein: The optimal unit control sequence is obtained by inputting the basic unit control sequence into the reinforcement learning model; The basic unit control sequence is obtained by summing the column vectors of the active power-line load factor sensitivity matrix and then sorting them; The active power-line load rate sensitivity matrix is an m×n order matrix, where m is the number of power branches and n is the number of power nodes.

5. The method according to claim 3 or 4, characterized in that: The active power-line load factor sensitivity matrix is extracted based on historical operating data when all units are fully powered on and there are no disconnections in the grid.

6. The method according to claim 1, wherein: The reinforcement learning model uses the unit control sequence as the state of the intelligent agent, the two positions in the sequence as the actions of the intelligent agent, and the comprehensive evaluation index as the reward; the influencing factors of the comprehensive evaluation index include the relative absorption capacity of new energy, line over-limit conditions, unit output constraints, node voltage constraints, and operating economic costs.

7. The method according to claim 6, characterized in that The reward is calculated as follows: Where: For rewards; Get the value for the reward points; when hour, ; Where: For the The output of new energy units is The effort of every moment; For the The output of new energy units is Output limit at any time; is the number of new energy units; when hour, Where, Indicates a constraint; when When it is 2, the constraint is the line current; when When it is 3, the constraint is the unit output; when When it is 4, the constraint is the node voltage; when When it is 5, the constraint is the operating economic cost; the subscripts max and min represent the upper and lower limits of the corresponding constraints, respectively.

8. The method according to claim 1, characterized in that The total adjustment amount of the active power of the thermal power unit at the next moment is determined by the following formula: Where: For the current moment Thermal power output, To provide thermal power for the next moment; Calculated by the following formula: Where: is the total number of loads, is the load number variable, is the number of new energy units, is the variable of the number of new energy units; for Total load at any moment; For the The output of new energy units is The effort of every moment; for Always balance the machine output; is the network loss power at the next moment, calculated by the following formula: in is the network loss estimation coefficient, which is calculated by the following formula: 。 9. The method according to claim 1, characterized in that The power on / off operation includes: When load fluctuations cause the thermal power adjustment to exceed the upper limit of the thermal power unit's ramp constraint, the thermal power units are started up according to the sensitivity of the line load factor from small to large. The power provided by the number of started thermal power units can compensate for the part of the thermal power adjustment that exceeds the upper limit of the ramp constraint. When load fluctuations cause the thermal power adjustment to fall below the lower limit of the thermal power unit's ramp constraint, the thermal power units are shut down according to the sensitivity of the line load rate from large to small. The power reduction caused by shutting down the number of thermal power units can offset the fact that the thermal power adjustment falls below the lower limit of the thermal power unit's ramp constraint. When the ratio of the actual processing to the maximum processing of all running generators exceeds a second set threshold, the generators are started up in order of sensitivity of the line load rate from small to large, so that the ratio is less than the second set threshold; When the ratio of the actual processing to the maximum processing of all running generators is lower than the third set threshold, they are shut down in descending order of sensitivity of line load rate so that the ratio is greater than the third set threshold.

10. The method according to claim 1, characterized in that The terminal voltage adjustment includes: The voltage of the generator set is recorded as , the reactive power is recorded as , where k represents the generator group identifier; like , then use The value after is updated; like , then use The value after is updated; like ,but constant; like ,but use The value after is updated; like ,but use The value after is updated.

Citation Information

Patent Citations

  • Static security aid decision making method for provincial power grid

    CN104868471A

  • Wind power plant active power control method, device and system

    CN107482692A