Virtual power plant resource coordination optimization control method, system, equipment and medium
By predicting the power generation and load status of virtual power plants, and using multi-agent reinforcement learning algorithm combined with environmental information generation control strategies, the problem of derailment of the control scheme of virtual power plants in the existing technology is solved, and the effect of grid frequency stabilization and reverse power flow suppression is achieved.
Patent Information
- Application Number
- CN202411882544.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-16
AI Technical Summary
The existing virtual power plant control methods cannot effectively combine environmental information, resulting in the generated control scheme derailment from the actual virtual power grid control and cannot effectively deal with the natural variability of renewable energy.
Long-term memory network is used to predict the power generation and load status of virtual power plants in the future, and multi-agent reinforcement learning algorithm is used to generate internal resource coordination optimization control strategies for virtual power plants, and optimize and adjust them in combination with environmental information.
It realizes the generation of optimal coordinated optimization control strategies when facing complex virtual power plant environments, stabilizes the grid frequency, suppresses voltage rise caused by reverse power flow, and improves the practical feasibility of the control scheme.
Smart Images

Figure CN120016435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual power plants, and in particular to a method, system, equipment and medium for coordinating and optimizing control of resources of a virtual power plant. Background Art
[0002] In the actual operation of the power grid, renewable energy is being connected to the power system in large quantities. However, due to the natural discontinuity of renewable energy, power generation will fluctuate. In this case, other energy sources such as battery energy storage systems are needed to compensate for the natural variability of renewable energy, ensure the stability of the grid frequency and suppress the voltage rise caused by reverse power flow. However, in the process of using a variety of other energy sources, there are situations where the capacity is inconvenient to control.
[0003] Chinese patent application CN118508526A discloses a virtual power plant control method, system and equipment based on neural network. The energy storage plan and energy release plan are obtained by constructing an optimal incentive model based on neural network. However, this method only relies on relevant data in the virtual power plant to plan the control plan, and cannot interact with the environment. It ignores the impact of the environment on the virtual power grid, which easily causes the generated control plan to be derailed from the actual virtual power grid control.
[0004] Therefore, providing a method for coordinated control of internal resources of a virtual power plant that can incorporate environmental information is a problem that needs to be solved. Summary of the invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a virtual power plant resource coordination and optimization control method, system, equipment and medium.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] According to a first aspect of the present invention, a virtual power plant resource coordination optimization control method is provided, the method comprising:
[0008] Obtaining original data and constraints, wherein the original data includes the power generation of various distributed energy sources, the energy storage status of energy storage facilities, equipment status, load power consumption and weather conditions; the distributed energy power generation includes the power generation of renewable energy and other energy sources; the constraints include virtual power plant capacity constraints, overvoltage protection constraints and cost constraints;
[0009] Based on the raw data, the long short-term memory network is used to predict the power generation of various distributed energy sources, the energy storage status of energy storage facilities and the load power consumption in the future virtual power plant;
[0010] Based on the predicted power generation of various distributed energy sources, the energy storage status of energy storage facilities and the load power consumption, a multi-agent reinforcement learning algorithm is used to generate a virtual power plant internal resource coordination optimization control strategy; renewable energy, other energy and energy storage facilities are used as agents, and the action set of the agent is determined to be whether to generate electricity and the amount of electricity generated, and the state set of the agent is the power generation status and power generation of the agent at the current moment, and the reward and punishment of the agent are constructed according to the constraints;
[0011] The internal resources of the virtual power plant are optimized and adjusted according to the virtual power plant internal resource coordination and optimization control strategy.
[0012] As a preferred technical solution, the method for predicting the power generation of various distributed energy sources, the energy storage status of energy storage facilities and load power consumption in the future virtual power plant is:
[0013] Preprocessing the raw data, wherein the preprocessing includes data cleaning, verification, normalization and smoothing;
[0014] The preprocessed raw data is input into the trained long short-term memory network to output the future power generation of various distributed energy sources, the energy storage status of energy storage facilities and load power consumption.
[0015] As a preferred technical solution, the training method of the long short-term memory network is offline training and online learning.
[0016] As a preferred technical solution, the long short-term memory network includes an input layer, an LSTM layer and an output layer, and the LSTM layer includes a forget gate, an input gate, a candidate memory unit, an output gate, a memory cell and a hidden state.
[0017] As a preferred technical solution, the method for generating a virtual power plant internal resource coordination optimization control strategy is:
[0018] Initialize the parameters of the multi-agent reinforcement learning algorithm, and obtain the initial state and initial action of each agent based on the predicted power generation of various distributed energy sources and the energy storage status of energy storage facilities;
[0019] Calculate the local value function of the corresponding agent based on the initial state and initial action of each agent;
[0020] Integrate the local value function of each agent to obtain the global value function, and determine the action of each agent at the next moment based on the global value function;
[0021] The agent executes the action at the next moment and calculates the reward or punishment of the action, generating the state of the agent at the corresponding moment;
[0022] The agent updates the algorithm parameters according to the rewards and punishments and the agent state at the corresponding moment, and repeats the above operations until the algorithm converges. The best action of each agent is output according to the rewards and punishments, and the best actions of each agent are combined to generate the internal resource coordination optimization control strategy of the virtual power plant.
[0023] As a preferred technical solution, the rewards and penalties include cost rewards, capacity rewards and overload penalties;
[0024] The expression of the cost reward is:
[0025]
[0026] Where n represents the number of agents; α represents the decision parameter of whether to generate electricity in the action of the agent, which is 01 type data, and when α=1, the action of the agent is to generate electricity, and when α=1, the action of the agent is not to generate electricity; Q i represents the amount of electricity generated in the action of the ith agent; C i is the power generation cost of agent i; C max Allow a maximum value for the cost;
[0027] The expression of the capacity reward is:
[0028]
[0029] Among them, k1 represents the reward coefficient, Q max Indicates the maximum capacity limit;
[0030] The expression of the overload penalty is:
[0031] r3=k2(O 实 -O max ),
[0032] Among them, k2 represents the penalty coefficient, O 实 Indicates actual overload, O max Indicates the overload threshold;
[0033] The total reward and punishment function is: r=r1+r2-r3.
[0034] As a preferred technical solution, the training process of the multi-agent reinforcement learning algorithm is:
[0035] The agent selects a state from the state set, calculates the action under the state, and updates the reward and punishment and the state at the next moment according to the action, and stores the state, action, reward and punishment and the state at the next moment as experience in the experience replay pool, and repeats until the capacity of the experience replay pool is zero;
[0036] Randomly select a batch of experiences from the experience replay pool and divide them into training set, test set and validation set according to the preset ratio;
[0037] The multi-agent reinforcement learning algorithm is trained using the training set, and the algorithm parameters are updated using the gradient loss function. The training effect is detected using the test set and validation set, and the algorithm parameters corresponding to the best training effect are used as the final parameters.
[0038] According to a second aspect of the present invention, a virtual power plant resource coordination and optimization control system is provided, wherein the system is used in the above method and comprises:
[0039] Data collection and analysis module: used to collect the power generation of various distributed energy sources, the energy storage status of energy storage facilities, the equipment status and load power consumption in the virtual power grid, and the weather conditions in the corresponding area, and generate constraints based on the power generation of various distributed energy sources, the energy storage status of energy storage facilities, the equipment status and load power consumption;
[0040] Intelligent prediction module: used to predict the power generation of various distributed energy sources, the energy storage status of energy storage facilities and the load power consumption in the future virtual power plant based on the power generation of various distributed energy sources, the energy storage status of energy storage facilities, equipment status, load power consumption and weather conditions collected by the data collection and analysis module, and output the prediction results;
[0041] Resource planning and strategy generation module: used to output the internal resource coordination optimization control strategy of the virtual power plant by using the multi-agent reinforcement learning algorithm according to the prediction results output by the intelligent prediction module and the constraints mentioned above;
[0042] Controller: Optimize and adjust the internal resources of the virtual power plant according to the virtual power plant internal resource coordination and optimization control strategy.
[0043] According to a third aspect of the present invention, there is provided an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the method described above is implemented when the processor executes the program.
[0044] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, wherein the program implements the method described when executed by a processor.
[0045] Compared with the prior art, the present invention has the following advantages:
[0046] 1) The present invention first predicts the future power generation and load in the virtual power plant, and uses a multi-agent reinforcement learning algorithm to take the power generation facilities in the virtual power plant as algorithm agents according to the prediction results, and takes whether the power generation facilities generate electricity and how much electricity they generate as agent actions, and designs a reward and punishment function according to the agent actions, so that the algorithm can still generate the best coordinated optimization control strategy in the face of complex virtual power plant scenarios containing renewable energy, so as to cope with the natural variability of renewable energy, ensure the stability of the grid frequency and suppress the voltage rise caused by reverse power flow;
[0047] 2) The present invention takes environmental information into consideration through a multi-agent reinforcement learning algorithm, realizes the interaction between the agent and the environment, and makes the virtual power plant coordination control scheme generated by it more in line with actual operation;
[0048] 3) A training method combining offline training and online learning is adopted for the long short-term memory network, which not only shortens the training time, but also eliminates the need to build a huge network training database, saving computing costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flow chart of the virtual power plant resource coordination optimization control method of the present invention;
[0050] Figure 2 This is a framework diagram of the virtual power plant resource coordination and optimization control system of the present invention. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0052] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent.
[0053] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0054] Example 1
[0055] This embodiment provides a virtual power plant resource coordination and optimization control method to solve the problem that the natural discontinuity of renewable energy will cause fluctuations in power generation, and other energy sources are urgently needed to compensate to smooth the natural variability of renewable energy, ensure the stability of grid frequency and suppress the voltage rise caused by reverse power flow. However, in the process of using multiple other energy sources, there is a problem that the capacity is inconvenient to control.
[0056] The specific process of this method is as follows Figure 1 As shown, the following steps are included:
[0057] S1. Get the original data and constraints:
[0058] The original data includes the power generation of various distributed energy sources, the energy storage status of energy storage facilities, equipment status, load power consumption and weather conditions, and the distributed energy power generation includes the power generation of renewable energy and other energy sources; among them, the constraints include virtual power plant capacity constraints, overvoltage protection constraints and cost constraints.
[0059] S2. Based on the original data, the long short-term memory network is used to predict the power generation of various distributed energy sources, the energy storage status of energy storage facilities and load power consumption in the future virtual power plant.
[0060] S21, preprocessing the original data:
[0061] S211. Perform data cleaning on the original data to remove outliers and missing values.
[0062] S212. Verify the cleaned original data to improve the prediction accuracy of the long short-term memory network.
[0063] S213, normalizing and smoothing the verified raw data to unify the form of input data and reduce the impact of noise.
[0064] S22. Input the preprocessed raw data into the trained long short-term memory network to output the future power generation of various distributed energy sources, the energy storage status of energy storage facilities and the load power consumption.
[0065] In this embodiment, a training method combining offline training and online learning is adopted for the long short-term memory network, and the detailed steps include:
[0066] A1. Obtain historical data of the virtual power plant, including power generation of various distributed energy sources, energy storage status of energy storage facilities, equipment status, load power consumption and weather conditions, and pre-process the first historical data.
[0067] A2. Set the long short-term memory network to include an input layer, an LSTM layer, and an output layer, and the LSTM layer includes a forget gate, an input gate, a candidate memory unit, an output gate, a memory cell, and a hidden state, and initialize the parameters of each structure in the LSTM layer.
[0068] A3. Divide the preprocessed historical data into a training set and a validation set in a ratio of 7:3. Input the training set into the input layer, calculate the predicted value through forward propagation, and compare it with the real data at the corresponding time in the historical data. Calculate the loss function, calculate the gradient through the back propagation algorithm, and update the parameter values of the forget gate, input gate, candidate memory unit, output gate, memory cell, and hidden state by combining the loss function and gradient.
[0069] A4. Repeat steps A1-A3 until the performance of the LSTM network in the validation set is stable, and select the LSTM network with the most stable performance as the prediction LSTM network.
[0070] A5. Collect the original data in the virtual power plant in real time, input it into the prediction long short-term memory network, output the predicted data, and compare the predicted data with the true value at the corresponding prediction time. Feedback the comparison structure to the network LSTM layer, and the LSTM layer performs online learning based on the comparison results.
[0071] S3. Generate the internal resource coordination optimization control strategy of the virtual power plant:
[0072] Renewable energy, other energy sources and energy storage facilities are taken as intelligent agents. The action set of the intelligent agent is determined as whether to generate electricity and the amount of electricity generated. The state set of the intelligent agent is the power generation state and power generation of the intelligent agent at the current moment. The rewards and punishments of the intelligent agent are constructed according to the constraints.
[0073] S31. Initialize the parameters of the multi-agent reinforcement learning algorithm, and obtain the initial state S of each agent based on the predicted power generation of various distributed energy sources and the energy storage state of energy storage facilities. t and the initial action a t .
[0074] S32, based on the initial state S of each agent t and the initial action a t Compute the local value function of the corresponding agent.
[0075] S33, integrate the local value function of each agent to obtain the global value function, and determine the action a of each agent at the next moment according to the global value function t+1 .
[0076] S34, the agent performs action a t+1 And calculate the reward and punishment of the action, and generate the state of the agent at the corresponding moment.
[0077] Specifically, rewards and penalties include cost rewards, capacity rewards, and overload penalties; the expression of cost rewards is: n represents the number of agents; α represents the decision parameter of whether to generate electricity in the action of the agent, which is a 01 type data, and when α=1, the action of the agent is to generate electricity, and when α=1, the action of the agent is not to generate electricity; Q i represents the amount of electricity generated in the action of the ith agent; C i is the power generation cost of agent i; C max Allow the maximum value for cost; the expression for capacity reward is: k1 represents the reward coefficient, Q max represents the maximum capacity limit; the expression of overload penalty is: r3=k2(O 实 -O max ), k2 represents the penalty coefficient, O 实 Indicates actual overload, O max Represents the overload threshold; according to the above description, the total reward and punishment function of the agent is set to: r = r1 + r2 - r3.
[0078] S35. The intelligent agent updates the algorithm parameters according to the rewards and punishments and the intelligent agent status at the corresponding moment, and repeats steps S31-S35 until the algorithm converges. The best action of each intelligent agent is output according to the rewards and punishments, and the best actions of each intelligent agent are combined to generate the internal resource coordination optimization control strategy of the virtual power plant.
[0079] The training steps of the multi-agent reinforcement learning algorithm used in this embodiment are:
[0080] B1. The agent selects a state from the state set, calculates the action under the state, and updates the reward and punishment and the state at the next moment based on the action. The state, action, reward and punishment and the state at the next moment are stored as experience in the experience replay pool, and it is repeated until the capacity of the experience replay pool is zero.
[0081] B2. Randomly select a batch of experiences from the experience replay pool and divide them into training set, test set and validation set according to the ratio of 6:3:1.
[0082] B3. Use the training set to train the multi-agent reinforcement learning algorithm, and use the gradient loss function to update the algorithm parameters. Use the test set and validation set to detect the training effect, and use the algorithm parameters corresponding to the best training effect as the final parameters.
[0083] S4. Optimizing and adjusting the internal resources of the virtual power plant is achieved according to the virtual power plant internal resource coordination and optimization control strategy.
[0084] Example 2
[0085] The above is an introduction to the method embodiment. The following is a further explanation of the solution of the present invention through a system embodiment.
[0086] This embodiment provides a virtual power plant resource coordination and optimization control system, which is used to implement the method provided in the above embodiment. The structure of the system is as follows: Figure 2 As shown, including:
[0087] Data collection and analysis module: used to collect the power generation of various distributed energy sources in the virtual power grid, the energy storage status of energy storage facilities, equipment status and load power consumption, as well as the weather conditions in the corresponding area, and generate constraints based on the power generation of various distributed energy sources, the energy storage status of energy storage facilities, equipment status and load power consumption.
[0088] Intelligent prediction module: used to predict the various distributed energy power generation, energy storage status of energy storage facilities and load power consumption in the future virtual power plant based on the various distributed energy power generation, energy storage status of energy storage facilities, equipment status, load power consumption and weather conditions collected by the data collection and analysis module, and output the prediction results.
[0089] Resource planning and strategy generation module: It is used to combine the prediction results output by the intelligent prediction module with constraints and use the multi-agent reinforcement learning algorithm to output the internal resource coordination optimization control strategy of the virtual power plant.
[0090] Controller: Optimize and adjust the internal resources of the virtual power plant according to the internal resource coordination and optimization control strategy of the virtual power plant.
[0091] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0092] The present embodiment also provides an electronic device including a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0093] Multiple components in the device are connected to the I / O interface, including: input units, such as keyboards, mice, etc.; output units, such as various types of displays, speakers, etc.; storage units, such as disks, optical disks, etc.; and communication units, such as network cards, modems, wireless communication transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunication networks.
[0094] The processing unit performs the various methods and processes described above, such as methods S1 to S4, methods A1 to A5, and methods B1 to B3. For example, in some embodiments, methods S1 to S4, methods A1 to A5, and methods B1 to B3 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S4, methods A1 to A5, and methods B1 to B3 described above may be executed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S4, methods A1 to A5, and methods B1 to B3 by any other appropriate means (e.g., by means of firmware).
[0095] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0096] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0097] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0098] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A virtual power plant resource coordination optimization control method, characterized in that: The method includes: Obtaining original data and constraints, wherein the original data includes the power generation of various distributed energy sources, the energy storage status of energy storage facilities, equipment status, load power consumption and weather conditions; the distributed energy power generation includes the power generation of renewable energy and other energy sources; the constraints include virtual power plant capacity constraints, overvoltage protection constraints and cost constraints; Based on the raw data, the long short-term memory network is used to predict the power generation of various distributed energy sources, the energy storage status of energy storage facilities and the load power consumption in the future virtual power plant; Based on the predicted power generation of various distributed energy sources, the energy storage status of energy storage facilities and the load power consumption, a multi-agent reinforcement learning algorithm is used to generate a virtual power plant internal resource coordination optimization control strategy; renewable energy, other energy and energy storage facilities are used as agents, and the action set of the agent is determined to be whether to generate electricity and the amount of electricity generated, and the state set of the agent is the power generation status and power generation of the agent at the current moment, and the reward and punishment of the agent are constructed according to the constraints; The internal resources of the virtual power plant are optimized and adjusted according to the virtual power plant internal resource coordination and optimization control strategy.
2. A virtual power plant resource coordination optimization control method according to claim 1, characterized in that: The method for predicting the power generation of various distributed energy sources, the energy storage status of energy storage facilities and load power consumption in the future virtual power plant is as follows: Preprocessing the raw data, wherein the preprocessing includes data cleaning, verification, normalization and smoothing; The preprocessed raw data is input into the trained long short-term memory network to output the future power generation of various distributed energy sources, the energy storage status of energy storage facilities and load power consumption.
3. A virtual power plant resource coordination optimization control method according to claim 2, characterized in that: The training method of the long short-term memory network is offline training and online learning.
4. A virtual power plant resource coordination optimization control method according to claim 1, characterized in that: The long short-term memory network comprises an input layer, an LSTM layer and an output layer, and the LSTM layer comprises a forget gate, an input gate, a candidate memory unit, an output gate, a memory cell and a hidden state.
5. The method for coordinated optimization and control of virtual power plant resources according to claim 1, characterized in that: The method for generating the internal resource coordination optimization control strategy of the virtual power plant is: Initialize the parameters of the multi-agent reinforcement learning algorithm, and obtain the initial state and initial action of each agent based on the predicted power generation of various distributed energy sources and the energy storage status of energy storage facilities; Calculate the local value function of the corresponding agent based on the initial state and initial action of each agent; Integrate the local value function of each agent to obtain the global value function, and determine the action of each agent at the next moment based on the global value function; The agent executes the action at the next moment and calculates the reward or punishment of the action, generating the state of the agent at the corresponding moment; The agent updates the algorithm parameters according to the rewards and punishments and the agent state at the corresponding moment, and repeats the above operations until the algorithm converges. The best action of each agent is output according to the rewards and punishments, and the best actions of each agent are combined to generate the internal resource coordination optimization control strategy of the virtual power plant.
6. A virtual power plant resource coordination optimization control method according to claim 1, characterized in that: The rewards and penalties mentioned include cost rewards, capacity rewards and overload penalties; The expression of the cost reward is: Where n represents the number of agents; α represents the decision parameter of whether to generate electricity in the action of the agent, which is 01 type data, and when α=1, the action of the agent is to generate electricity, and when α=1, the action of the agent is not to generate electricity; Q i represents the amount of electricity generated in the action of the ith agent; C i is the power generation cost of agent i; C max Allow a maximum value for the cost; The expression of the capacity reward is: Among them, k1 represents the reward coefficient, Q max Indicates the maximum capacity limit; The expression of the overload penalty is: r3=k2(O 实 -O max ), Among them, k2 represents the penalty coefficient, O 实 Indicates actual overload, O max Indicates the overload threshold; The total reward and punishment function is: r=r1+r2-r3.
7. A virtual power plant resource coordination optimization control method according to claim 1, characterized in that: The training process of the multi-agent reinforcement learning algorithm is as follows: The agent selects a state from the state set, calculates the action under the state, and updates the reward and punishment and the state at the next moment according to the action, and stores the state, action, reward and punishment and the state at the next moment as experience in the experience replay pool, and repeats until the capacity of the experience replay pool is zero; Randomly select a batch of experiences from the experience replay pool and divide them into training set, test set and validation set according to the preset ratio; The multi-agent reinforcement learning algorithm is trained using the training set, and the algorithm parameters are updated using the gradient loss function. The training effect is detected using the test set and validation set, and the algorithm parameters corresponding to the best training effect are used as the final parameters.
8. A virtual power plant resource coordination and optimization control system, characterized in that: The system is used to implement any one of claims 1 to 7, comprising: Data collection and analysis module: used to collect the power generation of various distributed energy sources, the energy storage status of energy storage facilities, the equipment status and load power consumption in the virtual power grid, and the weather conditions in the corresponding area, and generate constraints based on the power generation of various distributed energy sources, the energy storage status of energy storage facilities, the equipment status and load power consumption; Intelligent prediction module: used to predict the power generation of various distributed energy sources, the energy storage status of energy storage facilities and the load power consumption in the future virtual power plant based on the power generation of various distributed energy sources, the energy storage status of energy storage facilities, equipment status, load power consumption and weather conditions collected by the data collection and analysis module, and output the prediction results; Resource planning and strategy generation module: used to output the internal resource coordination optimization control strategy of the virtual power plant by using the multi-agent reinforcement learning algorithm according to the prediction results output by the intelligent prediction module and the constraints mentioned above; Controller: Optimize and adjust the internal resources of the virtual power plant according to the virtual power plant internal resource coordination and optimization control strategy.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Virtual power plant control method, system and equipment based on neural network
CN118508526A