Power grid regulation optimization method and device of pumped storage power station based on power system
By combining adaptive reinforcement learning algorithms with deep reinforcement learning algorithms, online joint optimization of virtual pumped storage power stations and power systems is achieved, solving the problem of insufficient accuracy and applicability of virtual pumped storage power station models, and improving the flexibility and efficiency of power grid regulation.
Patent Information
- Application Number
- CN202510463369.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The existing virtual pumped storage power station model has low accuracy and poor applicability, and cannot be optimized independently, resulting in suboptimal control and cannot effectively alleviate the power balance pressure of new energy access to the power grid.
Adaptive reinforcement learning algorithm is combined with deep reinforcement learning algorithms, and online joint optimization is achieved through interaction with the power system, through independent learning of optimal control strategies and combined with the physical model of virtual pumped storage power stations.
It improves the accuracy and applicability of the physical model of the virtual pumped storage power station, improves the flexibility and efficiency of power grid regulation, and effectively deals with the power fluctuations of new energy.
Smart Images

Figure CN120341985A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and device for optimizing grid regulation of a pumped-storage power station based on a power system. Background Art
[0002] With the large-scale access of new energy, the power system is facing huge power balance pressure. The virtual pumped-storage power station can provide auxiliary services such as frequency modulation and peak regulation for the power grid similar to the actual pumped-storage power station, and plays an important role in alleviating the impact of new energy access and improving the flexibility of the power grid.
[0003] In related technologies, the models of virtual pumped-storage power stations mainly have the following defects: Most models adopt simplified linear models, which are difficult to accurately describe the non-linear dynamic characteristics of the power stations, and the accuracy is low; they cannot autonomously optimize internal parameters according to real-time operation data, resulting in poor model applicability and lack of main self-learning ability; they often adopt static optimization methods, which are prone to sub-optimal control and the optimization methods are rough.
[0004] In view of this, there is an urgent need for a brand-new method for optimizing grid regulation of a pumped-storage power station to improve the model performance and better play its role in the power system. Summary of the Invention
[0005] In view of this, embodiments of the present disclosure provide a method and device for optimizing grid regulation of a pumped-storage power station based on a power system. An adaptive reinforcement learning algorithm is proposed based on the physical model of the virtual pumped-storage power station. By continuously interacting with the power system environment, the optimal control strategy is autonomously learned; then, the optimal control strategy is combined with the physical model of the virtual pumped-storage power station to realize the online joint optimization of the virtual pumped-storage power station and the power system, and the virtual pumped-storage power station and the power system jointly participate in grid regulation. This way of combining the adaptive reinforcement learning algorithm with the online joint optimization greatly improves the accuracy and applicability of the physical model of the virtual pumped-storage power station.
[0006] In a first aspect, embodiments of the present disclosure provide a method for optimizing grid regulation of a pumped-storage power station based on a power system, and the following technical solutions are adopted: Establish a physical model of a virtual pumped-storage power station; Based on the physical model of the virtual pumped-storage power station, define the state space, action space, and reward function of the adaptive reinforcement learning algorithm; Use a deep reinforcement learning algorithm to continuously interact with the actual power system, train the neural network, and autonomously learn the optimal control strategy through the trained neural network; According to the obtained real-time training data, further optimize the optimal control strategy by using an adaptive adjustment strategy; According to the operation status information of the current power grid obtained in real time, the virtual pumped storage power station issues control instructions to each energy storage unit of the power grid for execution by using the further optimized optimal control strategy; Receive the new operation status information fed back by the power grid according to the control instructions; According to the new operation status information, the virtual pumped storage power station continuously performs online learning and optimization on the optimal control strategy.
[0007] In some embodiments, a physical model of a virtual pumped storage power station is established, including: Establish a topological structure model of the virtual pumped storage power station, and define the parameters of the collection line and energy storage unit of the topological structure model; Based on the parameters of the collection line and energy storage unit of the topological structure model, establish a physical model of the energy storage unit of the virtual pumped storage power station; wherein, the physical model of the energy storage unit includes at least one of energy storage capacity, charge and discharge power, and conversion efficiency; Based on the physical model of the energy storage unit, establish a physical model of the whole power station, where the physical model of the whole power station represents the coupling relationship between the energy storage units of the physical model of the energy storage unit and the coupling relationship between each energy storage unit and the power grid.
[0008] In some embodiments, based on the physical model of the virtual pumped storage power station, the state space, action space, and reward function of the adaptive reinforcement learning algorithm are defined, including: Define the state space of the adaptive reinforcement learning algorithm, where the state space includes the state of charge of each energy storage unit and the power grid frequency; Define the action space of the adaptive reinforcement learning algorithm, where the action space includes the charge and discharge power instructions of each energy storage unit; Define the reward function based on the power station benefit and power grid performance of the virtual pumped storage power station.
[0009] In some embodiments, the deep reinforcement learning algorithm is used to continuously interact with the actual power system, train the neural network, and autonomously learn the optimal control strategy through the trained neural network, including: According to the defined value network and policy network, estimate the action state value function and the deterministic policy function respectively; Set the minimization of the temporal difference error as the update target of the value network; Set the maximization of the expected return as the update target of the policy network; Use the gradient information of the value network to optimize the policy network; By continuously interacting the deep reinforcement learning algorithm with the actual power system, trajectory data representing states, actions, rewards, and next states is generated; Using the trajectory data as training samples to train the neural network, and obtaining an optimal control strategy through the trained neural network.
[0010] In some embodiments, according to the obtained real-time training data, the optimal control strategy is further optimized using an adaptive adjustment strategy, including: Defining an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the neural network within a preset time period; Defining an adaptive exploration noise adjustment strategy to dynamically adjust the exploration noise according to the control requirements at different stages of the neural network; Defining an adaptive experience replay adjustment strategy to dynamically adjust the replay batch size and frequency according to the quality of the experience data of the neural network.
[0011] In some embodiments, according to the operation state information of the current power grid obtained in real time, the virtual pumped storage power station uses the further optimized optimal control strategy to issue control instructions to each energy storage unit of the power grid for execution, including: Through the virtual pumped storage power station, according to the operation state information of the current power grid, using the optimal control strategy to send control instructions to each energy storage unit of the actual power system, so that each energy storage unit of the actual power system sends the adjusted new operation state information to the virtual pumped storage power station according to the received control instructions; Through the virtual pumped storage power station receiving the adjusted new operation state information, and continuously optimizing the optimal control strategy according to the adjusted new operation state information.
[0012] In some embodiments, the method further includes: Through the virtual pumped storage power station, the operation state information of the current power grid is obtained in real time through a communication link to form a state observation value; Inputting the state observation value into the trained optimal policy network to obtain the charge and discharge power instructions of each energy storage unit; Through the energy management system of the virtual pumped storage power station, the charge and discharge power instructions are issued to each energy storage unit for execution to achieve coordinated control within the virtual pumped storage power station; The power grid dispatching center re-performs power flow calculation and frequency analysis according to the charge and discharge power instructions sent by the virtual pumped storage power station, combined with other power source and load information, to obtain the new operation state information of the actual power system after being adjusted by the virtual pumped storage power station; The actual power system feeds back the new operating state information to the virtual pumped - storage power station and uses the new operating state information as the input for the next decision - making moment of the virtual pumped - storage power station; The grid settlement center calculates the regulation compensation cost according to the actual regulation power and effect of the virtual pumped - storage power station; The new state observation value, the actions and feedback of the previous moment are stored in the experience replay pool through the virtual pumped - storage power station to form a complete trajectory data; A batch of trajectory data is randomly selected from the experience replay pool, and the value network and the policy network are updated by using the batch gradient descent method; Through the virtual pumped - storage power station, according to the updated policy network and combined with the new state observation value, a new charge - discharge power instruction is generated to start a new round of grid interaction and policy optimization process.
[0013] In a second aspect, an embodiment of the present disclosure also provides a model construction device for a virtual pumped - storage power station, adopting the following technical solution: A building unit, configured to build a physical model of the virtual pumped - storage power station; A design unit, configured to define a state space, an action space and a reward function of an adaptive reinforcement learning algorithm based on the physical model of the virtual pumped - storage power station; A policy generation unit, configured to continuously interact with the actual power system by using a deep reinforcement learning algorithm, train a neural network, and autonomously learn an optimal control policy through the trained neural network; A policy optimization unit, configured to further optimize the optimal control policy by using an adaptive adjustment policy according to the obtained real - time training data; An instruction issuing unit, configured to issue a control instruction to each energy storage unit of the grid for execution through the virtual pumped - storage power station by using the further optimized optimal control policy according to the real - time obtained operating state information of the current grid; A receiving unit, configured to receive the new operating state information fed back by the grid according to the control instruction; A continuous optimization unit, configured to continuously perform online learning and optimization of the optimal control policy through the virtual pumped - storage power station according to the new operating state information.
[0014] In a third aspect, an embodiment of the present disclosure also provides a computer device, adopting the following technical solution: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for constructing a model of the virtual pumped-storage power station described in any one of the above.
[0015] Fourthly, an embodiment of the present disclosure further provides a computer-readable storage medium storing computer instructions for causing a computer to execute the method for constructing a model of the virtual pumped-storage power station described in any one of the above.
[0016] Fifthly, an embodiment of the present disclosure further provides a computer program product including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0017] The grid regulation optimization method and device for a pumped-storage power station based on a power system provided by the embodiments of the present disclosure propose an adaptive reinforcement learning algorithm based on the physical model of the virtual pumped-storage power station. By continuously interacting with the power system environment, the optimal control strategy is autonomously learned; then, the optimal control strategy is combined with the physical model of the virtual pumped-storage power station to realize the online joint optimization of the virtual pumped-storage power station and the power system, and the virtual pumped-storage power station and the power system jointly participate in grid regulation. This way of combining the adaptive reinforcement learning algorithm with the online joint optimization greatly improves the accuracy and applicability of the physical model of the virtual pumped-storage power station.
[0018] The above description is only an overview of the technical solutions of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flow chart of a grid regulation optimization method for a pumped-storage power station based on a power system provided by an embodiment of the present disclosure; Figure 2 It is a schematic structural diagram of a grid regulation optimization device for a pumped-storage power station based on a power system provided by an embodiment of the present disclosure; Figure 3Schematic diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0021] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0022] It should be clear that the following describes the embodiments of the present disclosure through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0023] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0024] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure schematically. Only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape and size of the components in actual implementation. The type, quantity and proportion of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0025] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0026] As Figure 1 shown, the embodiments of the present disclosure provide a grid regulation optimization method for a pumped storage power station based on a power system, including the following steps: S101. Establish a physical model of a virtual pumped storage power station.
[0027] S102. Define the state space, action space, and reward function of the adaptive reinforcement learning algorithm based on the physical model of the virtual pumped - storage power station.
[0028] S103. Continuously interact with the actual power system using the deep reinforcement learning algorithm, train the neural network, and autonomously learn the optimal control strategy through the trained neural network.
[0029] S104. Further optimize the optimal control strategy using the adaptive adjustment strategy according to the obtained real - time training data.
[0030] S105. According to the real - time obtained operation state information of the current power grid, use the further optimized optimal control strategy through the virtual pumped - storage power station to issue control instructions to each energy storage unit of the power grid for execution.
[0031] S106. Receive the new operation state information feedback by the power grid according to the control instruction.
[0032] S107. According to the new operation state information, continuously perform online learning and optimization of the optimal control strategy through the virtual pumped - storage power station.
[0033] A grid regulation optimization method for a pumped - storage power station based on a power system provided by an embodiment of the present disclosure can propose an adaptive reinforcement learning algorithm based on the physical model of the virtual pumped - storage power station. By continuously interacting with the power system environment, the optimal control strategy is autonomously learned. Then, by combining the optimal control strategy with the physical model of the virtual pumped - storage power station, online joint optimization of the virtual pumped - storage power station and the power system is realized, and the virtual pumped - storage power station and the power system jointly participate in grid regulation. This way of combining the adaptive reinforcement learning algorithm with online joint optimization greatly improves the accuracy and applicability of the physical model of the virtual pumped - storage power station.
[0034] In some embodiments, establishing the physical model of the virtual pumped - storage power station includes: Establish the topological structure model of the virtual pumped - storage power station and define the collection line and energy storage unit parameters of the topological structure model; Based on the collection line and energy storage unit parameters of the topological structure model, establish the physical model of the energy storage unit of the virtual pumped - storage power station; wherein, the physical model of the energy storage unit includes at least one of energy storage capacity, charge - discharge power, and conversion efficiency. Based on the physical model of the energy storage unit, a physical model of the entire power station is established, where the physical model of the entire power station represents the coupling relationship between the energy storage units of the physical model of the energy storage unit and the coupling relationship between each energy storage unit and the power grid.
[0035] Optionally, establish a topological structure model of the virtual pumped-storage power station and define the collection line and energy storage unit parameters of the topological structure model, including the following steps: Define the network topological structure of the virtual pumped-storage power station, usually using a hierarchical tree structure, with the top layer being the total collection point of the virtual power station, connected to multiple collection lines below, and each collection line connecting multiple energy storage units; Define the key parameters of the collection line, including line number, resistance, reactance, load capacity, etc., for calculating line power flow and losses; Define the key parameters of each energy storage unit, including energy storage type (such as battery, flywheel), rated power, rated capacity, self-discharge rate, etc., as the basis for subsequent modeling.
[0036] Optionally, based on the collection line and energy storage unit parameters of the topological structure model, establish a physical model of the energy storage unit of the virtual pumped-storage power station (including energy storage capacity, charge and discharge power, conversion efficiency, etc.), including the following steps: For battery energy storage, use an equivalent circuit model, such as the Rint model, to describe the characteristics of battery SOC, OCV, internal resistance, etc., and consider self-discharge and capacity decay; For flywheel energy storage, use mechanical motion equations to describe the characteristics of flywheel speed, torque, losses, etc., and consider the efficiency during the energy conversion process; Add actual operation constraints to the energy storage model, such as charge and discharge rate limits, upper and lower bounds of state of charge, etc., to ensure the feasibility and safety of the energy storage model.
[0037] Optionally, establish a physical model of the entire power station to describe the coupling relationship between the energy storage units and the coupling relationship between each energy storage unit and the power grid, including the following steps: Based on the topological structure and power flow equation, establish the power balance constraint between the energy storage units to ensure the internal power balance of the virtual power station; Based on the power exchange at the power grid interface point, establish the coupling constraint between the virtual power station and the external power grid, reflecting that the virtual power station participates in power grid dispatching as a whole; Consider the losses within the power station, such as transformer losses, line losses, etc., to improve the accuracy of the model and avoid being too idealized.
[0038] For example, assume that a virtual pumped-storage power station consists of 1,000 distributed battery energy storage units, with a total of 10 collection lines, a total capacity of 50 MWh, and a total power of 10 MW. Define a three-layer topology where the top-level collection point is connected to 10 lines numbered L1 - L10. Each line has a rated voltage of 10 kV, a resistance of 0.1 Ω / km, an inductive reactance of 0.2 Ω / km, a carrying current of 100 A, and each line is evenly connected to 100 energy storage units. Each battery adopts an Rint equivalent circuit model, with a rated voltage of 400 V, a rated capacity of 50 kWh, a nominal self-discharge rate of 0.1% per day, a charging efficiency of 95%, a discharging efficiency of 96%, and an SOC range of [0.1, 0.9].
[0039] Optionally, establish a coupling model for the entire virtual pumped-storage power station to satisfy the power balance between the energy storage units and the collection lines: ; where is the output power of the i-th collection line; is the output power of the j-th battery on the i-th collection line; and respectively represent the resistance and current of the collection line; The total active power of the virtual pumped-storage power station is equal to the sum of the output powers of each collection line: .
[0040] The physical model of the virtual pumped-storage power station constructed in the embodiments of the present disclosure not only describes the network structure and dynamic characteristics inside the power station, but also reflects the interaction between the virtual pumped-storage power station and the external power grid, and can be used to guide the actual control and optimal scheduling of the virtual pumped-storage power station.
[0041] In some embodiments, based on the physical model of the virtual pumped-storage power station, define the state space, action space, and reward function of the adaptive reinforcement learning algorithm, including: Define the state space of the adaptive reinforcement learning algorithm, where the state space includes the state of charge of each energy storage unit and the grid frequency; Define the action space of the adaptive reinforcement learning algorithm, where the action space includes the charge and discharge power commands of each energy storage unit; Define the reward function based on the power station benefits and grid performance of the virtual pumped-storage power station.
[0042] In some embodiments, use the deep reinforcement learning algorithm to continuously interact with the actual power system, train the neural network, and autonomously learn the optimal control strategy through the trained neural network, including: Estimate the action-state value function and the deterministic policy function according to the defined value network and policy network respectively; Set minimizing the temporal difference error as the update objective of the value network; Set maximizing the expected return as the update objective of the policy network; Optimize the policy network using the gradient information of the value network; Generate trajectory data representing states, actions, rewards, and next states by continuously interacting the deep reinforcement learning algorithm with the actual power system; Use the trajectory data as training samples to train the neural network, and obtain the optimal control strategy through the trained neural network.
[0043] Optionally, the state space includes the state of charge (SOC) of the energy storage unit, grid frequency, etc. For the j-th energy storage unit on the i-th collection line, its SOC state can be expressed as s ij SOC ∈[0.1, 0.9]. The grid frequency state can be expressed as s f ∈[49.5, 50.5], with the unit of Hz. The entire state space S is the Cartesian product of the SOCs of each energy storage unit and the grid frequency.
[0044] Optionally, the action space includes the charge and discharge power commands of each energy storage unit. For the j-th energy storage unit on the i-th collection line, its charge and discharge power command can be expressed as a ij P ∈ [−PB ij max , PB ij max , with the unit of kW, and PB ij max is the rated power of this energy storage unit. The action space A is the Cartesian product of the charge and discharge power commands of each energy storage unit.
[0045] Optionally, the reward function considers the power plant benefits and grid performance, and guides the continuous optimization of the optimal control strategy. The power plant benefits can be represented by the revenue function Rrev, which is related to the electricity price, charge and discharge power, etc., and can be defined as R rev = ∑ t λ t (∑ i,j a ij , t P Δt), where is the electricity price at the t-th time period, and Δt is the duration of each time period.
[0046] Optionally, the grid performance can be represented by the frequency qualification rate R fre and is defined as: R fre=(1 / T)∫0 T 1 [49.8,50.2] (ft)dt, where 1 [49.8,50.2] is an indicator function, which takes the value of 1 when the frequency ft is within the qualified range of 49.8 to 50.2 Hz, and 0 otherwise.
[0047] Optionally, a deep reinforcement learning algorithm (such as DDPG, PPO, etc.) is used to learn the optimal control strategy. Taking the DDPG algorithm as an example, the value network Q(s,a|θ Q ) and the policy network μ(s∣θ μ ) are defined to estimate the action-state value function and the deterministic policy function respectively. The update objective of the value network is to minimize the temporal difference error: L(θ Q ) = E s,a,r,s′ [(Q(s,a∣θ Q ) - y) 2 , where y = r + γ Q (s’,μ(s’∣θ μ )∣θ Q′ ). The update objective of the policy network is to maximize the expected return: J(θμ) = Es[Q(s,μ(s∣θμ)∣θQ)]. The policy network is optimized using the gradient information of the value network. By continuously interacting with the virtual power plant environment, trajectory data of states, actions, rewards, and next states are generated to train the neural network, and finally the optimal control strategy μ∗(s) is obtained.
[0048] In some embodiments, according to the acquired real-time training data, the optimal control strategy is further optimized using an adaptive adjustment strategy, including: Defining an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the neural network within a preset time period; Defining an adaptive exploration noise adjustment strategy to dynamically adjust the exploration noise according to the control requirements at different stages of the neural network; Defining an adaptive experience replay adjustment strategy to dynamically adjust the replay batch size and frequency according to the quality of the experience data of the neural network.
[0049] Embodiments of the present disclosure incorporate an adaptive mechanism to online adjust hyperparameters such as the learning rate based on real-time data, define an adaptive learning rate adjustment strategy, and dynamically adjust the learning rate according to the training results (such as the return growth rate) over a period of time: if the training effect improves significantly, the learning rate is moderately increased; if the training results oscillate or do not converge, the learning rate is moderately decreased. Define an adaptive exploration noise adjustment strategy to dynamically adjust the exploration noise according to the control requirements at different stages: at the initial stage, the noise can be moderately increased to encourage exploration; at the later stage, the noise can be moderately decreased to seek stability. Define an adaptive experience replay adjustment strategy to dynamically adjust the replay batch size and frequency according to the quality of the experience data: if the data quality is high, the batch size and frequency are increased; otherwise, the batch size and frequency are decreased.
[0050] Optionally, assume that the state of the virtual power plant environment is the SOC of each energy storage unit and the grid frequency, the action space is the charge and discharge power of each energy storage unit, each time period is 15 minutes, and the reward function is set as the sum of the benefits when the frequency is qualified. Define the state space st = [s 11,t SOC , s 12,t SOC ,..., s 10,100,t SOC , s t f , the action space at = [a 11,t P , a 12,t P ,..., a 10,100,t P , the reward function r t = R rev,t ⋅ 1 [49.8,50.2] (f t ).
[0051] By constructing the value network and policy network of the DDPG algorithm, with the state space st as the input and the action space at as the output, training data is generated by interacting with the virtual power plant environment, and the neural network is iteratively optimized using the error function L(θ Q ) and the policy objective function J(θ μ ), and finally the optimal control strategy μ∗(s) is obtained.
[0052] Set the adaptive adjustment strategy: for example, when the average return growth rate of 10 consecutive episodes is lower than 0.1%, halve the learning rate; when the average return volatility of 10 consecutive episodes is higher than 5%, halve the exploration noise; when the time span of the data in the experience replay pool exceeds 1 month, increase the batch size by 20%.
[0053] Through the above adaptive reinforcement learning method, the agent can continuously self-optimize according to the actual operation data, improve the accuracy and adaptability of the optimal control strategy, and better guide the dispatching operation of the virtual pumped storage power station.
[0054] In some embodiments, according to the operation status information of the current power grid obtained in real time, the virtual pumped storage power station uses the further optimized optimal control strategy to issue control instructions to each energy storage unit of the power grid for execution, including: The virtual pumped storage power station sends control instructions to each energy storage unit of the actual power system according to the operation status information of the current power grid by using the optimal control strategy, so that each energy storage unit of the actual power system sends the adjusted new operation status information to the virtual pumped storage power station according to the received control instructions; The virtual pumped storage power station receives the adjusted new operation status information and continues to optimize the optimal control strategy according to the adjusted new operation status information.
[0055] In some embodiments, the method further includes: The virtual pumped storage power station obtains the operation status information of the current power grid in real time through the communication link to form a state observation value; The state observation value is input into the trained optimal policy network to obtain the charge and discharge power instructions of each energy storage unit; The charge and discharge power instructions are sent to each energy storage unit for execution through the energy management system of the virtual pumped storage power station to realize the coordinated control inside the virtual pumped storage power station; The power grid dispatching center re-performs power flow calculation and frequency analysis according to the charge and discharge power instructions sent by the virtual pumped storage power station, combined with other power source and load information, to obtain the new operation status information of the actual power system after the adjustment of the virtual pumped storage power station; The actual power system feeds back the new operation status information to the virtual pumped storage power station and uses the new operation status information as the input for the next decision-making moment of the virtual pumped storage power station; The power grid settlement center calculates the adjustment compensation cost according to the actual adjustment power and effect of the virtual pumped storage power station; The virtual pumped storage power station stores the new state observation value, the actions and feedback of the previous moment in the experience replay pool to form a complete trajectory data; Randomly extract a batch of trajectory data from the experience replay pool and update the value network and policy network by using the batch gradient descent method; The virtual pumped - storage power station generates new charge - discharge power commands according to the updated policy network, combined with the new state observation value, and starts a new round of grid interaction and policy optimization process.
[0056] Optionally, the embodiments of the present disclosure carry out online joint optimization of the virtual pumped - storage power station and the actual power system. The virtual pumped - storage power station sends control commands according to the optimal policy learned by itself based on the current grid state. The virtual pumped - storage power station obtains the current grid state information in real - time through a communication link. The current grid state information includes system frequency, electricity price, load level, etc., and forms a state observation value \(s_t\). The state observation value \(s_t\) is input into the trained optimal policy network \(\mu^*(s)\) to obtain the charge - discharge power commands \(a\) t =\(\mu\) ∗ (s t ) = [a 11,t P , a 12,t P ,..., a 10,100,t _P], and the charge - discharge commands \(a_t\) are sent to each energy storage unit for execution through the energy management system of the virtual pumped - storage power station, realizing the coordinated control inside the pumped - storage power station.
[0057] The actual power system receives the regulation commands of the virtual pumped - storage power station and feeds back the new operating state to the virtual pumped - storage power station. The grid dispatching center recomputes the power flow and frequency analysis according to the commands sent by the virtual pumped - storage power station, combined with other power source and load information, and obtains the new operating state \(s'\) of the actual power system after the regulation of the virtual pumped - storage power station t+1 . The actual power system feeds back the new operating state \(s'\) t+1 to the virtual pumped - storage power station, including information such as the regulated system frequency and line power flow, as the input for the next decision - making moment of the virtual pumped - storage power station. The grid settlement center calculates the regulation compensation cost according to the actual regulated electricity quantity and effect of the virtual pumped - storage power station, and the grid economically incentivizes the service of the virtual pumped - storage power station through the regulation compensation cost. The virtual pumped - storage power station continues to optimize the operation strategy in combination with the new grid state, and repeats the cycle to achieve continuous online optimization.
[0058] The virtual pumped - storage power station stores the new state observation value \(s\) t+1 and the action \(a\) t at the previous moment, and the feedback \(r\) t into the experience replay pool, forming a complete trajectory data \((s\) t , a t , r t , s t+1). The agent randomly extracts a batch of trajectory data from the experience replay pool, and uses batch gradient descent and other methods to update the value network and policy network to achieve online learning and optimization of the strategy. The virtual pumped storage power station uses the updated policy network and the new state observation s t+1 , generate new charge and discharge instructions a t+1 , starting a new round of grid interaction and strategy optimization process.
[0059] Assume that at time t, the virtual pumped storage power station obtains the current grid frequency of 49.8 Hz, the electricity price of 0.6 yuan / kWh, and the load level of 80 MW through the communication link, and generates the state observation value s accordingly. t . Virtual pumped storage power station will s t Input into the strategy network μ∗(s) to obtain the charging power instruction a of each energy storage unit t , for example t =[50kW,30kW,...,−20kW], where positive values represent charging and negative values represent discharging.
[0060] Virtual pumped storage power station uses energy management system to t The power grid dispatch center calculates the power flow based on this. When the frequency returns to 50.0Hz, the line power flow returns to normal, and a new state value s is generated accordingly. t+1 , and the new state value s t+1 Feedback to the virtual pumped storage power station. The grid settlement center calculates that the virtual pumped storage power station provides 10MWh of charging power and 2MWh of discharging power during this regulation, and accordingly gives the virtual pumped storage power station 800 yuan of regulation compensation, forming a feedback r t .
[0061] Virtual power plant will t , a t , r t ,s t+1 ) is stored in the experience replay pool, and this data is used to evaluate the value network Q(s,a|θ) Q ) and the policy network μ(s|θ μ ) is updated. Then based on the updated policy network and the new state observation s t+1 , generate new charge and discharge instructions a t+1 , starting a new round of grid interaction.
[0062] The virtual pumped-storage power station can carry out real-time information interaction and collaborative optimization with the actual power system. On the one hand, it can flexibly adjust its own operation strategy according to the grid demand. On the other hand, it can use its own regulation ability to provide auxiliary services for the grid, promoting the consumption of new energy and system balance. At the same time, the intelligent control level of the power station is continuously improved through online learning to maximize the operation benefit. This online joint optimization mode based on reinforcement learning provides a new idea for giving play to the flexible regulation potential of the virtual pumped-storage power station and constructing the ubiquitous power Internet of Things.
[0063] Optionally, for a typical power grid, such as the IEEE-39 bus system, a simulation model is constructed for verification. The IEEE-39 bus system is selected as the test power grid. This system contains 39 buses, 10 generators, and 46 lines, and it is a commonly used medium-scale power grid benchmark model. Based on the IEEE-39 bus system, the access location and capacity of the virtual pumped-storage power station are reasonably configured. For example, a virtual pumped-storage power station with a capacity of 200 MW and an energy storage of 10 MWh is connected to bus No. 8. A power flow and frequency simulation model of the IEEE-39 bus system is built and integrated with the physical model of the virtual pumped-storage power station to form a virtual power station-power grid joint simulation system. It can be implemented using professional software such as MATLAB / Simulink and PSS / E.
[0064] Optionally, different scenarios such as different new energy penetration rates and load curves are set to analyze the algorithm performance: Different new energy penetration rate scenarios are designed. For example, the cases where the wind power and photovoltaic penetration rates are 10%, 20%, and 30% are considered respectively, and they are simulated by modifying the generator output curve in the IEEE-39 bus system. Different load curve scenarios are designed. For example, typical load types such as industrial load, commercial load, and residential load are considered respectively, and they are simulated by modifying the load power curve in the IEEE-39 bus system. Under each scenario, the virtual power station-power grid joint simulation system is run, and key performance indicators such as system frequency, power balance, and economic operation are recorded, and a comparative analysis is carried out with the benchmark case without the virtual power station. Under the same new energy penetration rate and load curve scenarios, the adaptive reinforcement learning strategy and the traditional strategy of this application are run respectively, and quantitative indicators such as the system frequency qualification rate and the virtual power station revenue are recorded. The quantitative indicators under different strategies are compared and analyzed, and the improvement amplitude of the strategy of this application relative to the traditional strategy is calculated. For example, the frequency qualification rate is increased by 5 percentage points, and the power station revenue is increased by 20%.
[0065] The virtual pumped-storage power station modeled by the embodiments of the present disclosure, in the scenario of 30% new energy penetration rate, has a 5 percentage point increase in the frequency qualification rate and a 20% increase in the power station revenue compared with the traditional virtual power station model. In the working condition where the load curve presents a double-peak characteristic, the adaptive reinforcement learning algorithm can be continuously optimized according to the actual situation, avoiding the problem of performance degradation of the traditional fixed-parameter model due to load changes, and the frequency qualification rate is stable above 95%. Through online joint optimization, the embodiments of the present disclosure can enable the virtual pumped-storage power station to cooperate efficiently with the power grid. When dealing with large-scale new energy power fluctuations, the peak shaving and valley filling capacity is increased by 30% compared with the traditional scheme, greatly alleviating the power grid peak regulation pressure.
[0066] In summary, the embodiments of the present disclosure endow the power station with self-learning ability through intelligent algorithms, and realize power grid collaborative control through online optimization, effectively overcoming the deficiencies of traditional modeling methods, and providing possibilities for promoting the development of the ubiquitous power Internet of Things and building a new power system.
[0067] As Figure 2 shown, the embodiments of the present disclosure also provide a grid regulation and optimization device for a pumped-storage power station based on a power system, including: A building unit 21 configured to build a physical model of a virtual pumped-storage power station; A design unit 22 configured to define a state space, an action space, and a reward function of an adaptive reinforcement learning algorithm based on the physical model of the virtual pumped-storage power station; A policy generation unit 23 configured to continuously interact with the actual power system by using a deep reinforcement learning algorithm, train a neural network, and autonomously learn an optimal control policy through the trained neural network; A policy optimization unit 24 configured to further optimize the optimal control policy by using an adaptive adjustment policy according to the obtained real-time training data; An instruction issuing unit 25 configured to issue control instructions to each energy storage unit of the power grid for execution by the virtual pumped-storage power station by using the further optimized optimal control policy according to the operation state information of the current power grid obtained in real time; A receiving unit 26 configured to receive new operation state information fed back by the power grid according to the control instructions; A continuous optimization unit 27 configured to continuously perform online learning and optimization on the optimal control policy through the virtual pumped-storage power station according to the new operation state information.
[0068] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0069] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device executes all or part of the steps of the model construction method of the virtual pumped storage power station according to the foregoing embodiments of the present disclosure.
[0070] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience effect, this embodiment may also include well-known structures such as communication buses, interfaces, etc., and these well-known structures should also be included in the protection scope of the present disclosure.
[0071] As Figure 3 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of a computer device suitable for implementing the computer device in the embodiments of the present disclosure. Figure 3 The illustrated computer device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0072] As Figure 3 As shown, the computer device may include a processor (such as a central processor, a graphics processor, etc.), which may perform various appropriate actions and processes according to a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0073] Generally, the following devices may be connected to the I / O interface: input devices including, for example, sensors or visual information acquisition devices; output devices including, for example, display screens; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication device may allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or wiredly to exchange data. Although Figure 3A computer device with various devices is shown, but it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0074] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the method for constructing a model of a virtual pumped-storage power station according to the embodiments of the present disclosure are executed.
[0075] For a detailed description of this embodiment, reference can be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0076] A computer-readable storage medium according to an embodiment of the present disclosure stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the method for constructing a model of a virtual pumped-storage power station according to the foregoing embodiments of the present disclosure are executed.
[0077] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).
[0078] For a detailed description of this embodiment, reference can be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0079] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative and easy-to-understand purposes and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0080] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the phrase "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.
[0081] In addition, as used herein, "or" in a list of items starting with "at least one" indicates a disjunctive list, so that for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Further, the term "exemplary" does not mean that the examples described are preferred or better than other examples.
[0082] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.
[0083] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0084] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0085] The foregoing description has been presented for purposes of illustration and description. In addition, the description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those of skill in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A grid regulation optimization method for a pumped storage power station based on a power system, characterized in that Including: Establishing a physical model of a virtual pumped-storage power station; Based on the physical model of the virtual pumped-storage power station, defining the state space, action space, and reward function of an adaptive reinforcement learning algorithm; Using a deep reinforcement learning algorithm to continuously interact with the actual power system, training a neural network, and autonomously learning an optimal control strategy through the trained neural network; According to the obtained real-time training data, further optimizing the optimal control strategy using an adaptive adjustment strategy; According to the operation state information of the current power grid obtained in real time, issuing control instructions to each energy storage unit of the power grid for execution through the virtual pumped-storage power station using the further optimized optimal control strategy; Receiving new operation state information fed back by the power grid according to the control instructions; According to the new operation state information, continuously performing online learning and optimization of the optimal control strategy through the virtual pumped-storage power station.
2. The grid regulation optimization method for a pumped storage power station based on a power system according to claim 1, characterized in that Establishing a physical model of a virtual pumped-storage power station, including: Establishing a topological structure model of the virtual pumped-storage power station and defining the collector line and energy storage unit parameters of the topological structure model; Based on the collector line and energy storage unit parameters of the topological structure model, establishing a physical model of the energy storage unit of the virtual pumped-storage power station; wherein, the physical model of the energy storage unit includes at least one of energy storage capacity, charge and discharge power, and conversion efficiency; Based on the physical model of the energy storage unit, establishing a physical model of the entire power station, where the physical model of the entire power station represents the coupling relationship between the energy storage units of the physical model of the energy storage unit and the coupling relationship between each energy storage unit and the power grid.
3. The grid regulation optimization method for a pumped storage power station based on a power system according to claim 2, wherein, Based on the physical model of the virtual pumped-storage power station, defining the state space, action space, and reward function of an adaptive reinforcement learning algorithm, including: Defining the state space of the adaptive reinforcement learning algorithm, where the state space includes the state of charge of each energy storage unit and the power grid frequency; Defining the action space of the adaptive reinforcement learning algorithm, where the action space includes the charge and discharge power instructions of each energy storage unit; Defining the reward function based on the power station benefit and power grid performance of the virtual pumped-storage power station.
4. The grid regulation optimization method for a pumped-storage power station based on a power system according to claim 3, wherein Using a deep reinforcement learning algorithm to continuously interact with the actual power system, training a neural network, and autonomously learning an optimal control strategy through the trained neural network, including: According to the defined value network and policy network, respectively estimating the action-state value function and the deterministic policy function; Setting the minimization of the temporal difference error as the update target of the value network; Setting the maximization of the expected return as the update target of the policy network; Optimizing the policy network using the gradient information of the value network; By continuously interacting the deep reinforcement learning algorithm with the actual power system, generating trajectory data representing state, action, reward, and next state; Using the trajectory data as training samples to train the neural network and obtaining an optimal control strategy through the trained neural network.
5. The grid regulation optimization method for a pumped storage power station based on a power system according to claim 4, characterized in that According to the obtained real-time training data, further optimizing the optimal control strategy using an adaptive adjustment strategy, including: Define an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the neural network within a preset time period; Define an adaptive exploration noise adjustment strategy to dynamically adjust the exploration noise according to the control requirements at different stages of the neural network; Define an adaptive experience replay adjustment strategy to dynamically adjust the replay batch size and frequency according to the quality of the experience data of the neural network.
6. The grid regulation optimization method for a pumped storage power station based on a power system according to claim 4, wherein According to the operation state information of the current power grid obtained in real time, the virtual pumped-storage power station uses the further optimized optimal control strategy to issue control instructions to each energy storage unit of the power grid for execution, including: The virtual pumped-storage power station uses the optimal control strategy to send control instructions to each energy storage unit of the actual power system according to the operation state information of the current power grid, so that each energy storage unit of the actual power system sends the adjusted new operation state information to the virtual pumped-storage power station according to the received control instructions; The virtual pumped-storage power station receives the adjusted new operation state information and continues to optimize the optimal control strategy according to the adjusted new operation state information.
7. The grid regulation optimization method for a pumped storage power station based on a power system according to claim 6, characterized in that, The method further includes: The virtual pumped-storage power station obtains the operation state information of the current power grid in real time through the communication link to form a state observation value; Input the state observation value into the trained optimal policy network to obtain the charge and discharge power instructions of each energy storage unit; The energy management system of the virtual pumped-storage power station issues the charge and discharge power instructions to each energy storage unit for execution to achieve coordinated control within the virtual pumped-storage power station; The power grid dispatching center re-performs power flow calculation and frequency analysis according to the charge and discharge power instructions sent by the virtual pumped-storage power station, combined with other power source and load information, to obtain the new operation state information of the actual power system after the adjustment of the virtual pumped-storage power station; The actual power system feeds back the new operation state information to the virtual pumped-storage power station and uses the new operation state information as the input at the next decision-making moment of the virtual pumped-storage power station; The power grid settlement center calculates the adjustment compensation cost according to the actual adjustment power and effect of the virtual pumped-storage power station; The virtual pumped-storage power station stores the new state observation value, the actions and feedbacks of the previous moment into the experience replay pool to form a complete trajectory data; Randomly extract a batch of trajectory data from the experience replay pool and update the value network and policy network using the batch gradient descent method; The virtual pumped-storage power station generates new charge and discharge power instructions according to the updated policy network, combined with the new state observation value, and starts a new round of power grid interaction and policy optimization process.
8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the grid regulation optimization method of the pumped-storage power station based on the power system according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the grid regulation optimization method of the pumped-storage power station based on the power system according to any one of claims 1 to 7.
10. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.