Power grid regulation method and device based on reinforcement learning, equipment and medium
By acquiring power grid operation data and calculating the adaptive reward feedback value R, and using power grid regulation reinforcement learning agent to generate action strategies that match the power grid operation state, the problem of mismatch between agent actions and state is solved, thereby improving the power grid's safety, economy, and low carbon emissions.
Patent Information
- Application Number
- CN202410436489.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-11
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2044-04-11
AI Technical Summary
In existing power grid control systems, the actions of intelligent agents often do not match the current state, which increases the difficulty of power grid control.
By acquiring power grid operation data and calculating the adaptive reward feedback value R, a power grid regulation reinforcement learning agent is used to generate action strategies that match the power grid operation state. The power grid action strategies are then optimized by weighting the feedback values of safety, economy, and low carbon emissions.
This improves the performance of reinforcement learning agent algorithms, making their output action strategies more compatible with the power grid's operating state, thereby enhancing the power grid's safety, economy, and low-carbon characteristics.
Smart Images

Figure CN118336695B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of power grid regulation, and particularly relates to a power grid regulation method and device based on reinforcement learning, equipment and medium. BACKGROUND
[0002] With the construction of new power systems, the proportion of new energy access to power grid operation is gradually increasing, the uncertainty faced by power grid operation has increased significantly, and the difficulty of power grid regulation is also increasing. The regulation and operation of the power grid need to constantly adjust the strategy according to the operation state of the power grid to improve the safety, economy and low carbon of the power grid operation. However, the uncertainty scheduling problem brought by multiple resource access is difficult to model and solve in real time based on traditional physical mechanism, and new generation intelligent algorithms such as reinforcement learning provide a way to solve this problem.
[0003] Reinforcement learning is a learning process through which an agent constantly interacts with the environment to obtain the maximum reward value accumulation. The setting of the reward value is particularly important in the training and solving process of reinforcement learning. Setting a reasonable reward value can make the reinforcement learning agent more adaptable to the current operation state of the power grid and better achieve the preset goal of the regulation personnel. However, most of the agent reward functions currently applied in power grid regulation are several types of fixed values preset in advance, which leads to the occurrence of mismatch between the action of the agent and the current state. SUMMARY
[0004] The purpose of the present application is to provide a power grid regulation method and device based on reinforcement learning to solve the technical problem that the action of the agent and the current state often do not match in the prior art.
[0005] In order to achieve the above purpose, the present application adopts the following technical solutions:
[0006] In a first aspect, the present application provides a power grid regulation method based on reinforcement learning, comprising:
[0007] obtaining power grid operation data from the power grid operation environment;
[0008] obtaining an adaptive reward feedback value R according to the power grid operation data;
[0009] obtaining a power grid action strategy based on reinforcement learning through a power grid regulation reinforcement learning agent according to the power grid operation data and the adaptive reward feedback value R;
[0010] verifying the power grid action strategy according to the action space, and updating the power grid operation data by executing the power grid action strategy that passes the verification.
[0011] The further improvement of the present application is that in the step of obtaining power grid operation data from the power grid operation environment, the power grid operation data includes:
[0012] Upper limit of active power output of unit i Lower limit of active power output P i Active power output value P of unit i at time t it Adjustment upper limit value Adjustment lower limit value P abi The units include generator units, new energy units, energy storage units and load-side flexible resource equivalent units;
[0013] Number N of overload elements eq_over Number N of buses with voltage out of limit bus_over Number N of transmission sections out of limit sec_over Number N of light-load devices fac_under Total power generation P G Total load power P L New energy grid-connected power generation P at time t tnew New energy available power generation P at time t tN Total power generation P of conventional power sources and new energy at time t Δ .
[0014] The further improvement of the present application is that in the step of verifying the power grid action strategy according to the action space and updating the power grid operation data by executing the verified power grid action strategy, the action space includes:
[0015] Generator unit power up A1, generator unit power down A2, new energy unit power up A3, new energy unit power down A4, energy storage unit power generation up A5, energy storage unit power generation down A6, energy storage unit charging power up A7, energy storage unit charging power down A8, load-side flexible resource equivalent unit power consumption up A9, and load-side flexible resource equivalent unit power consumption down A 10 ;
[0016] The up and down adjustment amounts of each unit are determined by the adjustment step of each unit and do not exceed the capacity upper and lower limits.
[0017] The further improvement of the present application is that the step of obtaining an adaptive reward feedback value R according to the power grid operation data specifically includes:
[0018] Obtaining power grid operation state characteristic quantities according to the power grid operation data;
[0019] Calculating safety feedback value, economy feedback value and low-carbon feedback value according to the power grid operation state characteristic quantities;
[0020] An adaptive reward feedback value R is calculated according to the safety feedback value, the economy feedback value and the low-carbon feedback value.
[0021] The further improvement of the application is that in the step of obtaining the power grid operation state characteristic quantity according to the power grid operation data, the power grid operation state characteristic quantity comprises:
[0022] The upper rotation capacity reserve rate S of the unit i upsi The lower rotation capacity reserve rate S of the unit i downsi The upper regulation reserve capacity S of the unit i upabi The lower regulation reserve capacity S of the unit i downabi ;
[0023]
[0024] S downsi =(P it -P i ) / P i
[0025]
[0026] S downabi =(P it -P abi ) / P abi
[0027] The heavy overload rate λ eq The bus voltage out-of-limit rate λ line The cross-section power flow out-of-limit rate λ sec :
[0028]
[0029]
[0030]
[0031] Wherein, N eq is the total number of elements, N bus is the total number of buses, and N sec is the total number of power transmission cross sections.
[0032] The device light load rate δ fac :
[0033]
[0034] Wherein, N fac is the total number of devices.
[0035] The power grid loss rate δ loss :
[0036]
[0037] Renewable energy power generation grid-connected rate ζ new , renewable energy power generation proportion ζ rate :
[0038]
[0039]
[0040] The further improvement of the present application is that the step of calculating the safety feedback value, the economy feedback value and the low-carbon feedback value according to the grid operation state characteristic quantity specifically comprises:
[0041] The full marks of the safety feedback value, the economy feedback value and the low-carbon feedback value are all set as G;
[0042] The safety feedback value of the current operation state of the power grid is calculated as follows:
[0043]
[0044] Wherein, m is the number of devices with over-limit standby rate, S up , S down , S upab , S down are respectively the preset upper rotation reasonable standby rate, the lower rotation reasonable standby rate, the upper regulation reasonable standby ability and the lower regulation reasonable standby ability;
[0045] The economy feedback value of the current operation state of the power grid is calculated as follows:
[0046] R2 = [1-(δ fac + δ loss )] × G
[0047] The low-carbon feedback value of the current operation state of the power grid is calculated as follows:
[0048]
[0049] The further improvement of the present application is that the step of calculating the safety feedback value, the economy feedback value and the low-carbon feedback value according to the grid operation state characteristic quantity specifically comprises:
[0050] The weight of the safety feedback value, the economy feedback value and the low-carbon feedback value is calculated respectively as follows:
[0051]
[0052] Wherein, is a preset constant;
[0053] Then the security feedback value, the economy feedback value and the low-carbon feedback value are weighted and accumulated to obtain an adaptive reward feedback value R:
[0054] R = ε1R1 + ε2R2 + ε3R3.
[0055] The further improvement of the present application is that in the step of obtaining the power grid action strategy based on reinforcement learning by the power grid regulation reinforcement learning agent according to the power grid operation data and the adaptive reward feedback value R:
[0056] The power grid regulation reinforcement learning agent generates the power grid action strategy based on reinforcement learning according to the power grid operation data, aiming to maximize the cumulative return of the adaptive reward feedback value R.
[0057] In the second aspect, the present application provides a power grid regulation device based on reinforcement learning, comprising:
[0058] The acquisition module is configured to acquire power grid operation data from a power grid operation environment;
[0059] The adaptive reward feedback module is configured to acquire an adaptive reward feedback value R according to the power grid operation data;
[0060] The power grid regulation reinforcement learning agent is configured to obtain a power grid action strategy based on reinforcement learning according to the power grid operation data and the adaptive reward feedback value R;
[0061] The verification and execution module is configured to verify the power grid action strategy according to an action space, and update the power grid operation data by executing the power grid action strategy that passes the verification.
[0062] In the third aspect, the present application provides an electronic device comprising a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to realize the power grid regulation method based on reinforcement learning.
[0063] In the fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor to realize the power grid regulation method based on reinforcement learning.
[0064] Compared with the prior art, the present application has the following unexpected beneficial effects:
[0065] The application provides a power grid regulation method based on reinforcement learning, comprising: obtaining power grid operation data from a power grid operation environment; obtaining an adaptive reward feedback value R according to the power grid operation data; obtaining a power grid action strategy based on reinforcement learning through a power grid regulation reinforcement learning agent according to the power grid operation data and the adaptive reward feedback value R; and verifying the power grid action strategy according to an action space, and updating the power grid operation data by executing the verified power grid action strategy. The application comprehensively considers the safety, economy and low-carbon requirements of the power grid, then calculates the weights of the reward feedback values of the three aspects by using the evaluation results, and then performs weighted calculation to obtain an adaptive reward feedback value R matched with the power grid operation state. The adaptive reward feedback value R is adaptive to the power grid operation state, so that the training of the reinforcement learning is more likely to meet the regulation requirements. The application provides basic support for the performance improvement of the reinforcement learning agent algorithm, so that the action strategy output by the agent is matched with the current power grid operation state, and the agent can learn an action strategy more conducive to the improvement of the power grid operation state. BRIEF DESCRIPTION OF DRAWINGS
[0066] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of the application illustrated in the drawings are presented by way of example or for explanation only. In the drawings:
[0067] Figure 1 It is an interaction process diagram of the power grid regulation reinforcement learning agent and the power grid environment;
[0068] Figure 2 It is a flowchart of the power grid regulation method based on reinforcement learning of the application;
[0069] Figure 3 It is a structural diagram of the power grid regulation device based on reinforcement learning of the application;
[0070] Figure 4 It is a structural diagram of the electronic device of the application. DETAILED DESCRIPTION
[0071] The application will be described in detail below with reference to the drawings and in combination with the embodiments. It should be noted that the embodiments in the application and the features in the embodiments can be combined with each other without conflict.
[0072] The following detailed description is exemplary and is intended to provide further explanation of the application. Unless otherwise specified, all technical terms used in the application have the same meanings as those generally understood by those skilled in the art. The terms used in the application are only used to describe the specific embodiments, and are not intended to limit the exemplary embodiments according to the application.
[0073] The power system is a real-time balance system of power supply and demand, and dispatchers need to adjust the power grid according to the operation data of the power grid, so that the power grid is operated in a controllable range, and power grid accidents caused by imbalance between supply and demand are avoided.
[0074] The power grid regulation based on reinforcement learning needs to construct a corresponding state space and action space, and set an adaptive reward function, so that the reinforcement learning agent can give the most suitable adjustment strategy according to the current state of the power grid.
[0075] Among them, the action space of the agent is mainly set according to the power grid regulation means, mainly including: generator power up A1, generator power down A2, new energy unit power up A3, new energy unit power down A4, energy storage unit power up A5, energy storage unit power down A6, energy storage unit charging power up A7, energy storage unit charging power down A8, load side flexible resource equivalent unit power up A9, load side flexible resource equivalent unit power down A 10 The up and down adjustment amount of each unit is determined according to the adjustment step of each unit, and does not exceed the upper and lower limits of its capacity.
[0076] The state space is the measurement value of the power grid obtained by the agent and the related calculation index, and the acquisition of the state value also provides the basis for the calculation of the adaptive reward value, mainly including: the active power upper limit of unit i The active power lower limit P i ; the active power value P it of unit i at time t, the adjustment upper limit value P The adjustment lower limit value P abi ; the unit includes generator, new energy unit, energy storage unit and load side flexible resource equivalent unit; the number of overload elements N eq_over , the number of buses with overvoltage N bus_over , the number of overvoltage transmission sections N sec_over . The number of light load devices N fac_under , the total power P G and the total load power P L , the new energy grid-connected power P tnew at time t, the new energy available power P tN at time t, the total power of conventional power and new energy P Δ .
[0077] The interaction process of the power grid regulation reinforcement learning agent and the power grid environment is shown in Figure 1 :
[0078] The power grid regulation and control reinforcement learning agent obtains power grid operation data at the current time of the power grid, and gives an action strategy. Under the action strategy, the power grid operation environment changes, the measurement value changes, the safety feedback value, the economy feedback value and the low-carbon feedback value are affected, and the adaptive reward feedback value R is generated according to the safety feedback value, the economy feedback value and the low-carbon feedback value. In the training process of the agent, the agent and the power grid operation environment interact continuously, and the cumulative return of the adaptive reward feedback value R is maximized as the goal, and the action strategy of the agent is continuously corrected through the reinforcement learning algorithm, so that the power grid is operated in the optimal interval range.
[0079] Embodiment 1
[0080] Referring to Figure 2 The present application provides a power grid regulation and control method based on reinforcement learning, comprising:
[0081] S1, obtaining power grid operation data from the power grid operation environment;
[0082] In an embodiment, the power grid operation data comprises:
[0083] Upper limit of active power output of unit i Lower limit of active power output P i Active power output value P it of unit i at time t, adjustment upper limit value Adjustment lower limit value P abi ; the units include generator units, new energy units, energy storage units and load side flexible resource equivalent units;
[0084] Number of overload elements N eq_over Number of buses with voltage out-of-limit N bus_over Number of out-of-limit power transmission sections N sec_over Number of light load devices N fac_under Total power generation P G Total load power P L New energy grid-connected power P tnew at time t, new energy available power P tN at time t, total power generation of conventional power source and new energy P Δ at time t.
[0085] S2, obtaining an adaptive reward feedback value R according to the power grid operation data;
[0086] In an embodiment, the adaptive reward feedback value R is obtained according to the power grid operation data, specifically comprising:
[0087] Obtaining power grid operation state characteristic quantities according to the power grid operation data;
[0088] According to the grid operation state characteristic quantity, a safety feedback value, an economy feedback value and a low-carbon feedback value are calculated.
[0089] According to the safety feedback value, the economy feedback value and the low-carbon feedback value, an adaptive reward feedback value R is obtained.
[0090] The grid operation state characteristic quantity comprises: an upper rotating capacity reserve rate S upsi of the unit i, a lower rotating capacity reserve rate S downsi of the unit i, an upper regulation reserve capacity S upabi of the unit i, and a lower regulation reserve capacity S downabi of the unit i.
[0091]
[0092] S downsi =(P it -P i ) / P i
[0093]
[0094] S downabi =(P it -P abi ) / P abi
[0095] A heavy overload rate λ eq , a bus voltage out-of-limit rate λ line , and a cross-section power flow out-of-limit rate λ sec .
[0096]
[0097]
[0098]
[0099] Wherein, N eq is the total number of elements, N bus is the total number of buses, and N sec is the total number of transmission cross-sections.
[0100] A device light load rate δ fac .
[0101]
[0102] Wherein, N fac is the total number of devices.
[0103] A grid loss rate δ loss .
[0104]
[0105] Renewable energy power generation grid-connected rate ζ new , renewable energy power generation proportion ζ rate :
[0106]
[0107]
[0108] Among them, the safety feedback value, the economy feedback value and the low-carbon feedback value are calculated according to the grid operation state characteristic quantity, and specifically include:
[0109] The full marks of the safety feedback value, the economy feedback value and the low-carbon feedback value are all set as G;
[0110] The safety feedback value of the current operation state of the power grid is calculated:
[0111]
[0112] Among them, m is the number of devices with reserve rate out of limit, S up , S down , S upab , S down are respectively the preset upper rotation reasonable reserve rate, the lower rotation reasonable reserve rate, the upper regulation reasonable reserve capacity and the lower regulation reasonable reserve capacity;
[0113] The economy feedback value of the current operation state of the power grid is calculated:
[0114] R2 = [1-(δ fac + δ loss )] × G
[0115] The low-carbon feedback value of the current operation state of the power grid is calculated:
[0116]
[0117] Among them, the adaptive reward feedback value R is calculated according to the safety feedback value, the economy feedback value and the low-carbon feedback value, and specifically includes:
[0118] The weights of the safety feedback value, the economy feedback value and the low-carbon feedback value are calculated respectively:
[0119]
[0120] Among them, is a preset constant;
[0121] Then the safety feedback value, the economy feedback value and the low-carbon feedback value are weighted and accumulated to obtain the adaptive reward feedback value R:
[0122] R = ε1R1 + ε2R2 + ε3R3.
[0123] S3, obtaining a power grid action strategy based on reinforcement learning by the power grid regulation reinforcement learning agent according to the power grid operation data and the adaptive reward feedback value R;
[0124] In a specific embodiment, the power grid regulation reinforcement learning agent aims to maximize the cumulative return of the adaptive reward feedback value R, and generates a power grid action strategy based on reinforcement learning according to the power grid operation data.
[0125] S4, verifying the power grid action strategy according to the action space, and updating the power grid operation data by executing the power grid action strategy that passes the verification.
[0126] The present application provides a power grid regulation method based on reinforcement learning, which comprehensively considers the safety, economy and low-carbon requirements of the power grid, then calculates the weights of the reward feedback values of the three aspects according to the evaluation results, and then performs weighted calculation to obtain an adaptive reward feedback value R that matches the power grid operation state. The adaptive reward feedback value R is adaptive to the power grid operation state, making the training of reinforcement learning more easily meet the regulation requirements. The present application provides basic support for the performance improvement of reinforcement learning agent algorithm, so that the action strategy output by the agent matches the current power grid operation state, and the agent can learn an action strategy that is more beneficial to the improvement of the power grid operation state.
[0127] Example 2
[0128] Please refer to Figure 3 The present application provides a power grid regulation device based on reinforcement learning, which comprises:
[0129] The acquisition module is configured to acquire power grid operation data from a power grid operation environment;
[0130] The adaptive reward feedback module is configured to acquire an adaptive reward feedback value R according to the power grid operation data;
[0131] The power grid regulation reinforcement learning agent is configured to obtain a power grid action strategy based on reinforcement learning according to the power grid operation data and the adaptive reward feedback value R;
[0132] The verification and execution module is configured to verify the power grid action strategy according to the action space, and update the power grid operation data by executing the power grid action strategy that passes the verification.
[0133] In a specific embodiment, the step of acquiring power grid operation data from a power grid operation environment comprises:
[0134] Upper limit of active power output of unit i Lower limit of active power output P i; active power output value P of unit i at time t it , upper limit value of adjustment lower limit value of adjustment P abi ; the units include generator units, new energy units, energy storage units, and load-side flexible resource equivalent units;
[0135] number N of overload elements eq_over , number N of buses with voltage out-of-limit bus_over , number N of transmission sections with out-of-limit sec_over , number N of light-load devices fac_under , total power generation P G , total load power P L , new energy grid-connected power P at time t tnew , new energy available power P at time t tN , total power generation P of conventional power sources and new energy at time t Δ .
[0136] In a specific embodiment, in the step of verifying the grid action strategy according to the action space and updating the grid operation data by executing the grid action strategy that passes the verification, the action space includes:
[0137] generator unit power up A1, generator unit power down A2, new energy unit power up A3, new energy unit power down A4, energy storage unit power generation up A5, energy storage unit power generation down A6, energy storage unit charging power up A7, energy storage unit charging power down A8, load-side flexible resource equivalent unit power up A9, and load-side flexible resource equivalent unit power down A 10 ;
[0138] The up and down adjustment amounts of each unit are determined by the adjustment step of each unit and do not exceed the capacity upper and lower limits.
[0139] In a specific embodiment, the step of obtaining an adaptive reward feedback value R according to grid operation data specifically includes:
[0140] obtaining a grid operation state characteristic quantity according to the grid operation data;
[0141] calculating a safety feedback value, an economic feedback value, and a low-carbon feedback value according to the grid operation state characteristic quantity;
[0142] calculating the adaptive reward feedback value R according to the safety feedback value, the economic feedback value, and the low-carbon feedback value.
[0143] In a specific embodiment, in the step of obtaining a grid operation state characteristic quantity according to grid operation data, the grid operation state characteristic quantity includes:
[0144] Upper rotation capacity reserve rate S of unit i upsi Lower rotation capacity reserve rate S of unit i downsi Upper regulation reserve capacity S of unit i upabi Lower regulation reserve capacity S of unit i downabi ;
[0145]
[0146] S downsi = (P it -P i ) / P i
[0147]
[0148] S downabi = (P it -P abi ) / P abi
[0149] Heavy overload rate λ eq Bus voltage out-of-limit rate λ line Sectional power flow out-of-limit rate λ sec :
[0150]
[0151]
[0152]
[0153] wherein N eq is the total number of elements, N bus is the total number of buses, and N sec is the total number of transmission sections;
[0154] Device light load rate δ fac :
[0155]
[0156] wherein N fac is the total number of devices;
[0157] Network loss rate δ loss :
[0158]
[0159] Renewable energy power generation grid connection rate ζ new Renewable energy power generation proportion ζ rate :
[0160]
[0161]
[0162] In a specific embodiment, the step of calculating the safety feedback value, the economy feedback value and the low-carbon feedback value according to the grid operation state characteristic quantity specifically comprises:
[0163] The full scores of the safety feedback value, the economy feedback value and the low-carbon feedback value are all set as G;
[0164] The safety feedback value of the current operation state of the power grid is calculated as:
[0165]
[0166] Wherein, m is the number of devices with over-limit reserve rate, S up , S down , S upab , S down are respectively the preset upper rotation reasonable reserve rate, the lower rotation reasonable reserve rate, the upper regulation reasonable reserve capacity and the lower regulation reasonable reserve capacity;
[0167] The economy feedback value of the current operation state of the power grid is calculated as:
[0168] R2 = [1-(δ fac + δ loss )] × G
[0169] The low-carbon feedback value of the current operation state of the power grid is calculated as:
[0170]
[0171] In a specific embodiment, the step of calculating the adaptive reward feedback value R according to the safety feedback value, the economy feedback value and the low-carbon feedback value specifically comprises:
[0172] The weights of the safety feedback value, the economy feedback value and the low-carbon feedback value are respectively calculated as:
[0173]
[0174] Wherein, is a preset constant;
[0175] Then, the safety feedback value, the economy feedback value and the low-carbon feedback value are weighted and accumulated to obtain the adaptive reward feedback value R:
[0176] R = ε1R1+ ε2R2+ ε3R3.
[0177] In a specific embodiment, the power grid regulation reinforcement learning agent aims to maximize the cumulative return of adaptive reward feedback value R, and generates a reinforcement learning-based power grid action strategy according to power grid operation data.
[0178] Embodiment 3
[0179] Referring to Figure 4 The electronic device 100 comprises a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0180] The memory 101 can be used to store the computer program 103, and the processor 102 can realize the steps of the reinforcement learning-based power grid regulation method of embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; and the data storage area can store data (such as audio data) created according to the use of the electronic device 100, etc. In addition, the memory 101 can include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device.
[0181] The at least one processor 102 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 102 can be a microprocessor or any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects all parts of the electronic device 100 through various interfaces and lines.
[0182] The memory 101 in the electronic device 100 stores a plurality of instructions to implement a power grid regulation method based on reinforcement learning, and the processor 102 can execute the plurality of instructions to implement:
[0183] obtain power grid operation data from a power grid operation environment;
[0184] obtain an adaptive reward feedback value R according to the power grid operation data;
[0185] obtain a power grid action strategy based on reinforcement learning by a power grid regulation reinforcement learning agent according to the power grid operation data and the adaptive reward feedback value R;
[0186] verify the power grid action strategy according to an action space, and update the power grid operation data by executing the power grid action strategy that passes the verification.
[0187] In a specific embodiment, the power grid regulation reinforcement learning agent generates a power grid action strategy based on reinforcement learning according to the power grid operation data, with the goal of maximizing the cumulative return of the adaptive reward feedback value R.
[0188] Embodiment 4
[0189] The modules / units integrated in the electronic device 100, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods of the present application can also be instructed by a computer program to related hardware to complete, and the computer program can be stored in a computer-readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, and read-only memory (ROM).
[0190] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0191] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.
[0192] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.
[0193] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.
[0194] Finally, it should be noted that the above-mentioned embodiments are merely used to illustrate the technical solutions of the present application, but are not intended to limit the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced, and any modification or replacement without departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.
Claims
1. A power grid regulation method based on reinforcement learning, characterized in that, The method comprises the following steps: obtaining power grid operation data from a power grid operation environment; obtaining an adaptive reward feedback value R according to the power grid operation data; obtaining a power grid action strategy based on reinforcement learning by a power grid regulation and control reinforcement learning agent according to the power grid operation data and the adaptive reward feedback value R; verifying the power grid action strategy according to an action space, and updating the power grid operation data by executing the power grid action strategy that passes the verification; in the step of obtaining the power grid operation data from the power grid operation environment, the power grid operation data comprises: Generating unit Upper limit of active power output Lower limit of active power output Generating unit Active power output value at Active power output value at Upper limit value of adjustment Lower limit value of adjustment The generating unit includes a generator unit, a new energy unit, an energy storage unit, and a load-side flexible resource equivalent unit. Number of overload elements Number of busbars with voltage out of limit Number of transmission sections with power out of limit Number of light-load devices Total power generated Total power consumed , Power generated by new energy at the moment , Power that can be generated by new energy at the moment , Total power generated by conventional power source and new energy at the moment ; in the step of verifying the power grid action strategy according to the action space, and updating the power grid operation data by executing the power grid action strategy that passes the verification, the action space comprises: A1, A2, A3, A4, A5, A6, A7, A8, A9, and A10, respectively. 10 ; the up and down adjustment amount of each unit is determined by the adjustment step length of each unit, and does not exceed the upper and lower limits of the capacity; the step of obtaining the adaptive reward feedback value R according to the power grid operation data specifically comprises: obtaining power grid operation state characteristic quantities according to the power grid operation data; calculating a safety feedback value, an economy feedback value and a low-carbon feedback value according to the power grid operation state characteristic quantities; obtaining the adaptive reward feedback value R by calculation according to the safety feedback value, the economy feedback value and the low-carbon feedback value; the step of obtaining the adaptive reward feedback value R by calculation according to the safety feedback value, the economy feedback value and the low-carbon feedback value specifically comprises: respectively calculating the weights of the safety feedback value, the economy feedback value and the low-carbon feedback value: wherein is a predetermined constant; then, the safety feedback value, the economy feedback value and the low-carbon feedback value are weighted and accumulated to obtain the adaptive reward feedback value R: 。 2. The reinforcement learning based power grid regulation method of claim 1, wherein, in the step of obtaining the power grid operation state characteristic quantities according to the power grid operation data, the power grid operation state characteristic quantities comprise: unit upper rotational capacity reserve ratio , unit lower rotational capacity reserve ratio , unit upper regulation reserve capacity , unit lower regulation reserve capacity ; Heavy overload rate Bus voltage out-of-limit rate Sectional current flow out-of-limit rate : wherein, is the total number of elements, is the total number of busbars, is the total number of transmission sections; Device light load rate : wherein, is the total number of devices; grid loss rate : Renewable energy generation grid connection rate Renewable energy generation share : 。 3. The reinforcement learning based power grid regulation method of claim 2, wherein, in the step of calculating the safety feedback value, the economy feedback value and the low-carbon feedback value according to the power grid operation state characteristic quantities, specifically comprising: the full marks of the safety feedback value, the economy feedback value and the low-carbon feedback value are all set as G; calculating the safety feedback value of the current operation state of the power grid: Wherein, m is the number of devices whose standby rate is out of limit, , , , are respectively a preset upper rotation reasonable standby rate, a lower rotation reasonable standby rate, an upper adjustment reasonable standby capacity and a lower adjustment reasonable standby capacity. calculating the economy feedback value of the current operation state of the power grid: calculating the low-carbon feedback value of the current operation state of the power grid: 。 4. A power grid regulating device based on reinforcement learning, characterized in that The method comprises the following steps: an obtaining module, configured to obtain power grid operation data from a power grid operation environment; an adaptive reward feedback module, configured to obtain an adaptive reward feedback value R according to the power grid operation data; a power grid regulation and control reinforcement learning agent, configured to obtain a power grid action strategy based on reinforcement learning according to the power grid operation data and the adaptive reward feedback value R; a verification and execution module, configured to verify the power grid action strategy according to an action space, and update the power grid operation data by executing the power grid action strategy that passes the verification; in the step of obtaining the power grid operation data from the power grid operation environment, the power grid operation data comprises: unit upper limit of active power lower limit of active power unit active power value at time upper limit value of adjustment lower limit value of adjustment ; the unit includes a generator unit, a new energy unit, an energy storage unit, and a load-side flexible resource equivalent unit. Number of overload components Number of busbars with voltage exceeding limits Number of transmission sections exceeding the limit Number of light-load equipment Total power generation Total load power , At any time, the grid-connected power generation capacity of new energy , Momentary renewable energy power generation , Total power generation of conventional power sources and new energy sources at all times ; in the step of verifying the power grid action strategy according to the action space, and updating the power grid operation data by executing the power grid action strategy that passes the verification, the action space comprises: A1, A2, A3, A4, A5, A6, A7, A8, A9, and A10, respectively. 10 ; the up and down adjustment amount of each unit is determined by the adjustment step length of each unit, and does not exceed the upper and lower limits of the capacity; the step of obtaining the adaptive reward feedback value R according to the power grid operation data specifically comprises: obtaining power grid operation state characteristic quantities according to the power grid operation data; According to the grid operation state characteristic quantity, a safety feedback value, an economy feedback value and a low-carbon feedback value are calculated; An adaptive reward feedback value R is calculated according to the safety feedback value, the economy feedback value and the low-carbon feedback value; The step of calculating the adaptive reward feedback value R according to the safety feedback value, the economy feedback value and the low-carbon feedback value specifically comprises: The weights of the safety feedback value, the economy feedback value and the low-carbon feedback value are respectively calculated: wherein is a predetermined constant; Then, the safety feedback value, the economy feedback value and the low-carbon feedback value are weighted and accumulated to obtain the adaptive reward feedback value R: 。 5. An electronic device, comprising: The computer readable storage medium stores at least one instruction, and the at least one instruction is executed by the processor to implement the grid regulation method based on the reinforcement learning.
6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one instruction, and the at least one instruction is executed by the processor to implement the grid regulation method based on the reinforcement learning.
Citation Information
Patent Citations
Intelligent optimization method for power grid safe operation strategy based on deep reinforcement learning
CN114048903A
Power grid real-time scheduling optimization method and system, computer equipment and storage medium
CN115241885A
Reinforcement learning method and device based on environment reward fuzzy self-adaption and medium
CN117808117A