Virtual power plant operation strategy optimization method and equipment considering supply and demand ends

By constructing a two-layer hybrid game model on both the supply and demand sides and a multi-agent Markov game, combined with reinforcement learning methods, the problems of supply and demand separation and uncertainty in virtual power plant research were solved, and efficient optimization decision-making of virtual power plants in the electricity market was realized.

CN121458084APending Publication Date: 2026-02-03CSG POWER GENERATION (GUANGDONG) ENERGY STORAGE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511490425.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing research on virtual power plants suffers from fragmented modeling of supply and demand sides, simplification of game behavior, and insufficient handling of uncertainties, resulting in low strategy adaptability and low solution efficiency.

Method used

A two-layer hybrid game model integrating supply and demand is constructed, combining multi-agent Markov game theory and reinforcement learning methods to perform supply and demand forecasting and decision optimization by acquiring electricity market data.

Benefits of technology

It improves the synergy, dynamic adaptability, and decision-making efficiency of virtual power plant operation strategies, and enhances market adaptability and profitability in uncertain market environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121458084A_ABST
    Figure CN121458084A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual power plant operation strategy optimization method and equipment considering supply and demand ends, and the method comprises the steps: obtaining the supply side, demand side and market environment data of a power market, and building a double-layer mixed game model integrating the supply and demand ends based on the data; the model comprises a lower-layer market clearing model used for market clearing and an upper-layer virtual power plant decision model used for making a power supply strategy by a virtual power plant. Meanwhile, a supply and demand prediction model is established according to market data to predict power supply output and power demands of users. And in combination with the predicted output and demand data and a double-layer game model, modeling a dynamic interaction relationship between the virtual power plant and other market subjects by using a multi-agent Markov game, and solving the Markov game through a multi-agent reinforcement learning method so as to obtain an optimized operation decision strategy of the virtual power plant. The method can improve the collaboration and dynamic adaptability of the operation strategy of the virtual power plant, and can be widely applied to the technical field of electric power.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric power, and particularly relates to a virtual power plant operation strategy optimization method and device considering both supply and demand ends. BACKGROUND

[0002] With the rapid development of the energy internet, as an important carrier of integrating distributed power generation resources and demand-side flexible resources, the virtual power plant (VPP) plays an increasingly important role in the electricity market. It aggregates dispersed distributed energy, energy storage systems and adjustable loads to realize source-load collaborative optimization of the power system, which helps to improve the flexibility and economy of power grid operation. Under the background of deepening of electricity market reform, the virtual power plant not only participates in energy trading, but also provides ancillary services, and the effectiveness of its operation strategy directly affects market revenue and system balance. Therefore, the research on the optimization strategy of the virtual power plant under the electricity market environment has gradually become the focus.

[0003] There are still some key problems in the modeling and solving of the virtual power plant operation strategy in the existing research. On the one hand, most of the researches model the supply side and the demand side separately, and fail to fully reflect the coupling characteristics of the virtual power plant in the supply-demand coordination. On the other hand, the game behavior between market participants is mostly assumed to be static or simplified, which is difficult to reflect the dynamic competition process in the actual market. In addition, when dealing with the uncertainty of renewable energy output and load demand, the existing methods are insufficient in the integration of real-time information and intelligent prediction technology, and the traditional optimization algorithm also faces the challenges of low solving efficiency and poor convergence in the multi-agent and high-dimensional strategy space.

[0004] Therefore, the problems in the related art need to be solved urgently. SUMMARY

[0005] The present application aims to at least partly solve one of the problems in the related art.

[0006] To this end, an object of embodiments of the present application is to provide a virtual power plant operation strategy optimization method and device considering both supply and demand ends.

[0007] To achieve the above technical purpose, the technical solutions adopted by the embodiments of the present application include: On the one hand, the embodiments of the present application provide a virtual power plant operation strategy optimization method considering both supply and demand ends, which comprises: obtaining market data of an electricity market; wherein the market data comprises supply side data, demand side data and market environment data; According to the market data, a double-layer hybrid game model of both supply and demand ends is constructed, wherein the double-layer hybrid game model comprises a lower-layer market clearing model and an upper-layer virtual power plant decision model, the lower-layer market clearing model is used for market clearing to realize supply and demand balance, and the upper-layer virtual power plant decision model is used for the virtual power plant to formulate an optimized power supply strategy; According to the market data, a supply and demand prediction model is established, and output data of power supply and power demand data of users are predicted according to the supply and demand prediction model; According to the output data, the power demand data and the double-layer hybrid game model, a relationship between the virtual power plant and other market subjects in the power market is modeled through a multi-agent Markov game; The Markov game is solved through a multi-agent reinforcement learning method, and an operation decision strategy of the virtual power plant is obtained.

[0008] In addition, according to the virtual power plant operation strategy optimization method considering both supply and demand ends, the method can further have the following additional technical features: Further, in an embodiment of the present application, the double-layer hybrid game model of both supply and demand ends is constructed according to the market data, comprising: The lower-layer market clearing model is constructed with market social welfare maximization as an optimization target, and a constraint condition corresponding to the lower-layer market clearing model is set; The upper-layer virtual power plant decision model is constructed with virtual power plant operation profit maximization as an optimization target.

[0009] Further, in an embodiment of the present application, the constraint condition corresponding to the lower-layer market clearing model comprises a supply and demand power balance constraint, a virtual power plant output constraint and a demand response load constraint.

[0010] Further, in an embodiment of the present application, the supply and demand prediction model is established according to the market data, and the output data of power supply and the power demand data of users are predicted according to the supply and demand prediction model, comprising: According to the market data, a long short-term memory network model and a generalized autoregressive conditional heteroscedasticity model are established; Time sequence dependent features in the output data and the power demand data are determined through the long short-term memory network model, and mean values of the output data and the power demand data in a future period are predicted; According to a residual error of the mean value prediction, a fluctuation rate of the output data and the power demand data in the future period is predicted through the generalized autoregressive conditional heteroscedasticity model.

[0011] Further, in an embodiment of the present application, the relationship between the virtual power plant and other market participants in the electricity market is modeled by Markov game of multi-agent according to the output data, the electricity demand data and the bi-level mixed game model, including: constructing a multi-agent system according to the virtual power plant and other market participants in the electricity market; determining the state space, action space, state transition probability and reward function of the multi-agent system under Markov game according to the output data, the electricity demand data and the bi-level mixed game model.

[0012] Further, in an embodiment of the present application, the Markov game is solved by reinforcement learning method of multi-agent to obtain the operation decision strategy of the virtual power plant, including: constructing an agent unit for each agent of the virtual power plant and the other market participants; each agent unit includes an actor network and a critic network, the input of the actor network is the local observation state of the agent, and the output is a continuous decision action, the input of the critic network is the joint state and joint action of the agent, and the output is the value evaluation of the current decision strategy; updating the critic network of each agent by centralized training; applying the actor network of each agent by distributed execution to output the operation decision strategy of the virtual power plant.

[0013] On the other hand, an embodiment of the present application provides a virtual power plant operation strategy optimization device considering both supply and demand sides, the device comprising: an acquisition unit configured to acquire market data of an electricity market; wherein the market data includes supply side data, demand side data and market environment data; a construction unit configured to construct a bi-level mixed game model considering both supply and demand sides according to the market data; wherein the bi-level mixed game model includes a lower market clearing model and an upper virtual power plant decision model, the lower market clearing model is used for market clearing to achieve supply and demand balance, and the upper virtual power plant decision model is used for virtual power plant to formulate an optimized power supply strategy; a prediction unit configured to establish a supply and demand prediction model according to the market data, and predict output data of power supply and electricity demand data of users according to the supply and demand prediction model; a modeling unit configured to model the relationship between the virtual power plant and other market participants in the electricity market by Markov game of multi-agent according to the output data, the electricity demand data and the bi-level mixed game model; A solution unit is configured to solve the Markov game by using a multi-agent reinforcement learning method, so as to obtain an operation decision strategy of the virtual power plant.

[0014] Further, in an embodiment of the present application, the construction unit is specifically configured to: The lower market clearing model is constructed with the optimization target of maximizing market social welfare, and the constraint condition corresponding to the lower market clearing model is set. The upper virtual power plant decision model is constructed with the optimization target of maximizing operation profit of the virtual power plant.

[0015] In another aspect, an electronic device is provided in an embodiment of the present application, which comprises: at least one processor; at least one memory configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor is caused to implement the above-mentioned method for optimizing a virtual power plant operation strategy considering both supply and demand sides.

[0016] In another aspect, an electronic device is provided in an embodiment of the present application, which comprises:

[0017] The advantages and beneficial effects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be learned by the practice of the present application: The method disclosed in the embodiments of the present application firstly acquires supply side, demand side and market environment data of a power market, and constructs a double-layer hybrid game model considering both supply and demand sides based on the data. The model comprises a lower market clearing model for market clearing and an upper virtual power plant decision model for virtual power plant to make power supply strategy. Meanwhile, a supply and demand prediction model is established according to market data to predict power supply output and user electricity demand. In combination with the predicted output and demand data and the double-layer game model, a multi-agent Markov game is used to model dynamic interaction relationship between the virtual power plant and other market subjects, and a multi-agent reinforcement learning method is used to solve the Markov game, so as to obtain an optimized operation decision strategy of the virtual power plant. The present application effectively improves the coordination, dynamic adaptability, decision efficiency and convergence of the virtual power plant operation strategy in uncertain market environment by constructing a double-layer hybrid game model considering both supply and demand, combining supply and demand prediction with multi-agent Markov game, and using reinforcement learning to solve the game. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the present application or the prior art. It should be understood that the drawings in the following introduction are only for the convenience of clearly describing part of the embodiments of the technical solutions of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the premise that there is no conflict.

[0019] Figure 1 An implementation environment schematic diagram of a virtual power plant operation strategy optimization method considering both supply and demand ends provided in the embodiments of the present application; Figure 2 A flowchart schematic diagram of a virtual power plant operation strategy optimization method considering both supply and demand ends provided in the embodiments of the present application; Figure 3 A structure schematic diagram of an electronic device provided in the embodiments of the present application. DETAILED DESCRIPTION

[0020] The present application will be further described below in conjunction with the drawings of the specification and specific embodiments. The described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0021] In the following description, “some embodiments” are related to a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0023] With the rapid development of the energy internet, the virtual power plant (VPP) as an important carrier of integrating distributed power generation resources and demand-side flexible resources plays an increasingly important role in the electricity market. It aggregates dispersed distributed energy, energy storage systems and adjustable loads to realize source-load collaborative optimization of the power system, which helps to improve the flexibility and economy of power grid operation. Under the background of deepening of electricity market reform, virtual power plants not only participate in energy trading, but also provide ancillary services, and the effectiveness of their operation strategy directly affects market revenue and system balance. Therefore, the research on the optimization strategy of virtual power plants in the electricity market environment has gradually become the focus of attention.

[0024] Currently, although certain progress has been made in the research on virtual power plant strategies in the power market environment, there are still the following key problems: Split modeling of supply and demand ends: Existing researches mostly focus on the optimization of power generation resources on the supply side or separately consider demand side response, lack systematic modeling of the coupling characteristics of the supply and demand ends, and are difficult to reflect the core advantage of "source-load coordination" of virtual power plants.

[0025] Simplified game behavior: Some researches assume that the behavior of market participants is fixed or only use static game models, ignoring the dynamic game process between virtual power plants, traditional power suppliers, aggregators, users and other multi-agent, resulting in insufficient adaptability of the strategy to the market competition environment.

[0026] Insufficient response to uncertainty: The description of uncertain factors such as renewable energy output and user electricity demand is relatively single, mostly using historical data statistics or simple probability distribution, without fully combining real-time data and machine learning methods to improve prediction accuracy.

[0027] Defects in solving efficiency and convergence: Traditional game theory methods are prone to "dimension disaster" when dealing with high-dimensional strategy space, and existing reinforcement learning algorithms have problems such as strategy update lag and slow convergence speed in multi-agent game scenarios.

[0028] Therefore, in the embodiments of the present application, a virtual power plant operation strategy optimization method and device considering both supply and demand ends are provided. The method first acquires supply side, demand side and market environment data of the power market, and constructs a comprehensive double-layer hybrid game model of both supply and demand based on these data. The model includes a lower market clearing model for market clearing and an upper virtual power plant decision model for virtual power plant to formulate power supply strategy. At the same time, a supply and demand prediction model is established according to the market data to predict power supply output and user electricity demand. Combined with the predicted output and demand data and the double-layer game model, a multi-agent Markov game is used to model the dynamic interaction between virtual power plants and other market participants, and a multi-agent reinforcement learning method is used to solve the Markov game, so as to obtain the optimized operation decision strategy of the virtual power plant. Through the construction of a comprehensive double-layer hybrid game model of supply and demand, combined with supply and demand prediction and multi-agent Markov game, and using reinforcement learning for solving, the coordination, dynamic adaptability, decision efficiency and convergence of the virtual power plant operation strategy in the uncertain market environment are effectively improved.

[0029] Please refer to Figure 1 , Figure 1 An implementation environment schematic diagram of a virtual power plant operation strategy optimization method considering both supply and demand ends provided in the embodiments of the present application is shown. In the implementation environment, the main software and hardware subjects involved include a terminal device 110 and a background server 120. The terminal device 110 and the background server 120 are in communication connection.

[0030] Specifically, the virtual power plant operation strategy optimization method considering both supply and demand sides provided in the embodiments of the present application can be executed on the terminal device 110 side alone or based on data interaction between the terminal device 110 and the background server 120. The terminal device 110 can be a computer device, a mobile phone, a smart portable device, etc. The background server 120 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.

[0031] The terminal device 110 and the background server 120 can establish a communication connection through a wireless network or a wired network. The wireless network or the wired network uses standard communication technology and / or protocols, and the network can be set as the Internet or any other network, for example, any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network, but is not limited to these.

[0032] Of course, it can be understood that the implementation environment in Figure 1 is only some optional application scenarios of the virtual power plant operation strategy optimization method considering both supply and demand sides provided in the embodiments of the present application, and the actual application is not fixed to the software and hardware environment shown in Figure 1 .

[0033] Next, in combination with the foregoing introduction of the implementation environment, a virtual power plant operation strategy optimization method considering both supply and demand sides provided in the embodiments of the present application is introduced and described.

[0034] Please refer to Figure 2 , Figure 2 is a schematic diagram of a virtual power plant operation strategy optimization method considering both supply and demand sides provided in the embodiments of the present application. The virtual power plant operation strategy optimization method considering both supply and demand sides includes but is not limited to: Step 210, obtaining market data of a power market; wherein the market data includes supply side data, demand side data and market environment data; In step 220, a double-layer hybrid game model of both supply and demand ends is constructed according to the market data; wherein the double-layer hybrid game model comprises a lower-layer market clearing model and an upper-layer virtual power plant decision model, the lower-layer market clearing model is used for market clearing to realize supply and demand balance, and the upper-layer virtual power plant decision model is used for the virtual power plant to formulate an optimized power supply strategy; In step 230, a supply and demand prediction model is established according to the market data, and output data of power supply and power demand data of users are predicted according to the supply and demand prediction model; In step 240, a relationship between the virtual power plant and other market subjects in the power market is modeled by a multi-agent Markov game according to the output data, the power demand data and the double-layer hybrid game model; In step 250, the Markov game is solved by a multi-agent reinforcement learning method, and an operation decision strategy of the virtual power plant is obtained.

[0035] In the embodiment of the application, a virtual power plant operation strategy optimization method considering both supply and demand ends is provided. The method lays a foundation for modeling by widely collecting market data, and then constructs a double-layer game model that can reflect the actual structure of the market, and simultaneously establishes a supply and demand prediction model to cope with uncertainty. On this basis, the multi-agent Markov game theory is introduced to accurately describe the dynamic interaction relationship between the virtual power plant and other market subjects, and an advanced multi-agent reinforcement learning algorithm is used to efficiently solve the complex game process, so as to output the optimal operation strategy of the virtual power plant.

[0036] Specifically, in step 210, market data of the power market is obtained, and these data are the basis for subsequent modeling and analysis. The market data covers multi-dimensional data such as supply side (such as output cost of generator set, available capacity), demand side (such as user load curve, price elasticity) and market environment (such as historical price, policy rules), and the specific data types contained are not limited in the application.

[0037] In step 220, a double-layer hybrid game model of both supply and demand ends is constructed based on the market data. The upper layer of the model is a virtual power plant decision model, which focuses on formulating its own optimized power supply and bidding strategies in a given market environment, denoted as an upper-layer virtual power plant decision model. The lower layer is a market clearing model, which is used to simulate the process of unified clearing by the market operation agency according to the bids of all participants to realize supply and demand balance, denoted as a lower-layer market clearing model. The clearing result (such as market price) will feedback to affect the decision of the upper layer, so as to accurately describe the two-way coupling relationship between the virtual power plant strategy and the market result.

[0038] In step 230, in response to the uncertainty of renewable energy output and user demand, a supply-demand prediction model is established according to historical and real-time market data, which can predict the power supply output data (such as wind power and photovoltaic power generation) and user electricity demand data in future time periods, providing forward-looking information for decision-making.

[0039] In step 240, the predicted output data and electricity demand data are combined with the established double-layer hybrid game model, and a multi-agent Markov game theory is used for in-depth modeling. In this step, the virtual power plant and its competitors (such as traditional power generators and other aggregators) are respectively regarded as independent agents, which interact with each other in a common market environment. The decision of each agent not only depends on the current market state, but also needs to consider the possible behavior of other agents and the resulting future state transition, thereby achieving an accurate description of the continuous dynamic game process of multiple parties.

[0040] In step 250, for the above-mentioned multi-agent Markov game model, a multi-agent reinforcement learning method is used for solving. Through the continuous trial and error and learning of multiple agents in the simulated environment, the strategies of each agent are gradually optimized, and the optimal operation decision strategy of the virtual power plant in the face of uncertainty and dynamic strategies of competitors can be efficiently solved.

[0041] It can be understood that the virtual power plant operation strategy optimization method considering both supply and demand provided in the embodiments of the present application acquires market data of a power market; wherein the market data includes supply side data, demand side data and market environment data; a double-layer hybrid game model considering comprehensive supply and demand is constructed according to the market data; wherein the double-layer hybrid game model includes a lower market clearing model and an upper virtual power plant decision model, the lower market clearing model is used for market clearing to achieve supply and demand balance, and the upper virtual power plant decision model is used for virtual power plant to formulate an optimized power supply strategy; a supply-demand prediction model is established according to the market data, and output data of power supply and electricity demand data of users are predicted according to the supply-demand prediction model; the relationship between the virtual power plant and other market participants in the power market is modeled through multi-agent Markov game according to the output data, the electricity demand data and the double-layer hybrid game model; the Markov game is solved through multi-agent reinforcement learning method, and the operation decision strategy of the virtual power plant is obtained. This method effectively improves the coordination, dynamic adaptability, decision-making efficiency and convergence of the virtual power plant operation strategy in an uncertain market environment by constructing a double-layer hybrid game model considering comprehensive supply and demand, combining supply-demand prediction and multi-agent Markov game, and using reinforcement learning for solving.

[0042] The application aims to solve the key problems in the existing virtual power plant strategy research, taking "supply-demand coordination" as the core, providing input through high-precision uncertainty prediction, modeling interaction through mixed game, solving output strategy through efficient reinforcement learning, and guaranteeing landing through dynamic adaptation mechanism, forming a whole-process optimization from "data perception" to "strategy execution", and finally realizing the dual goals of virtual power plant revenue improvement and market adaptability enhancement.

[0043] In the embodiments of the application, by constructing a double-layer mixed game model of both supply and demand, combining uncertainty environment modeling based on machine learning, mixed strategy Nash equilibrium and game solving framework of reinforcement learning fusion, the decision-making ability of virtual power plant in complex market environment is improved. The logic of each step is: first, provide accurate input through uncertainty modeling, then construct a double-layer game model to depict supply-demand coordination and multi-agent interaction, and finally solve it efficiently through the fusion of mixed game and reinforcement learning, and dynamically adapt to market changes.

[0044] Some technical implementation details of the application will be described and explained in detail below.

[0045] Specifically, in some embodiments, the double-layer mixed game model of both supply and demand is constructed according to the market data, comprising: a lower market clearing model is constructed with the optimization goal of maximizing market social welfare, and the constraint conditions corresponding to the lower market clearing model are set; an upper virtual power plant decision-making model is constructed with the optimization goal of maximizing the operating profit of the virtual power plant.

[0046] The traditional method only focuses on supply-side generation optimization or demand-side response, and has the limitation of "supply-demand fragmentation", which cannot reflect the core value of virtual power plant "integration of source and load". In the embodiments of the application, through the double-layer model of "lower supply-demand balance + upper strategy optimization", the supply-demand resources are taken as a whole variable for decision-making, providing a basic framework for multi-agent game. The specific implementation logic is as follows: Step 1: A lower market clearing model containing virtual power plant supply and demand response is constructed with the goal of maximizing market social welfare, and the objective function is: (1) wherein, is a set of time periods, is a set of virtual power plants, is a set of demand response users, is the market clearing price in period t, is the power supply of virtual power plant v in period t, is the power supply cost function of virtual power plant, is the demand response load of user d in period t, The utility function of the user for power consumption.

[0047] Step 2: Set constraints.

[0048] (1) Power supply and demand balance constraint: (2) Where G is the set of traditional generator units, is the power supply of traditional units; H is the set of non-adjustable loads, is the non-adjustable load.

[0049] (2) Virtual power plant output constraint. This constraint limits the output range of each virtual power plant, ensuring that it operates within the feasible generating capacity range.

[0050] (3) Where, , is the minimum / maximum output of the virtual power plant.

[0051] (3) Demand response load constraint: (4) Where, , is the minimum / maximum response load of user d at time period t.

[0052] Step 3: Build the upper-level virtual power plant decision model (game strategy layer).

[0053] The virtual power plant aims to maximize profits, synchronously optimizes power supply pricing and demand response incentives, and realizes "source-load coordination" decision-making. The optimization of power supply pricing strategy and demand response incentive strategy , the objective function is: (5) Where, is the set of users aggregated by the virtual power plant, is the demand response incentive coefficient of user d at time period t.

[0054] Constraints: s.t. (6) Where, is the power supply pricing coefficient of virtual power plant v at time period t; , is the upper and lower limit of the pricing coefficient; , is the upper and lower limit of the incentive coefficient.

[0055] Specifically, in some embodiments, the power supply and demand prediction model is established according to the market data, and the output data of the power supply and the power demand data of the user are predicted according to the power supply and demand prediction model, comprising: According to the market data, a long short-term memory network model and a generalized autoregressive conditional heteroscedasticity model are established; The time sequence dependent characteristics in the output data and the power demand data are determined by the long short-term memory network model, and the mean value of the output data and the power demand data in the future period is predicted; According to the residual error of the mean value prediction, the volatility rate of the output data and the power demand data in the future period is predicted by the generalized autoregressive conditional heteroscedasticity model.

[0056] In the embodiments of the present application, when modeling the uncertainty of the supply and demand, the limitation of traditional "single probability distribution" is broken through, and the prediction accuracy is improved by the LSTM-GARCH fusion model to provide reliable input for the game strategy. Because the renewable energy output (such as wind power and photovoltaic) and the user power demand have strong uncertainty, accurate prediction is needed to provide input for the aforementioned game model. The strategy in the embodiments of the present application can improve the problem that the traditional method (such as ARIMA) cannot describe the time sequence dependence and volatility aggregation, as follows: Step 1: (1) Supply and demand state feature extraction. To fully capture the dynamics of supply and demand, a multi-dimensional feature space is constructed and the data is preprocessed. A multi-dimensional state space including the supply side and the demand side is constructed, and the feature vector is: (7) Among them, the supply side features include: is the average renewable energy output in period t, reflecting the average power supply capacity of distributed power supply; is the renewable energy output volatility rate, which is calculated by rolling window standard deviation to describe the uncertainty; the demand side features include: is the average user base load in period t, which is based on historical power consumption data smoothing processing; is the load volatility rate, reflecting the uncertainty of user power consumption behavior; the market environment features include: is the real-time market price, which drives the supply and demand strategy adjustment of the virtual power plant; is the period type label (such as peak / flat / valley period).

[0057] (2) Data preprocessing. The continuous features are normalized: (8) Among them, is the mean value, is the standard deviation. The discrete features (such as period type) are encoded to enhance the model's ability to identify category information.

[0058] Step 2: Uncertainty prediction based on LSTM-GARCH. In the embodiments of the present application, a long short-term memory network-generalized autoregressive conditional heteroskedasticity (LSTM-GARCH) model is used to jointly predict the mean and volatility of renewable energy output and electricity demand. LSTM captures temporal dependence, and GARCH describes volatility clustering. The joint output of the mean and volatility of supply and demand variables is as follows: (1) LSTM mean prediction layer: The input sequence is the supply and demand feature vector of the previous N periods The LSTM network is used to extract temporal dependence: (9) wherein is the LSTM hidden layer output (stores historical information); is the mean prediction value of the supply and demand features at period t; Dense is a fully connected layer (output prediction result).

[0059] Network structure: 2-layer LSTM hidden layer (128 units each) + fully connected output layer, with a linear function as the activation function (suitable for regression tasks). Loss function: mean squared error (MSE), which is expressed as: (10) wherein is the predicted mean, is the actual value, and T is the total number of periods.

[0060] (2) GARCH volatility prediction layer: 1) Predict the residual based on LSTM , and construct a GARCH(1,1) model to describe the volatility clustering effect: (11) wherein is a standard normal random variable; is the square of the volatility at period t; is the long-term variance level (non-negative), is the ARCH term coefficient (describing the influence of recent shocks on variance, ), is the GARCH term coefficient (describing the persistent influence of historical variance on current variance, , and ensures stationarity).

[0061] 2) Loss function: (12) (3) Joint forecast output. A scenario set is generated by the joint distribution of mean and volatility, which is used for robust optimization of subsequent game strategies. Supply and demand variables follow a normal distribution: (13) Among them, the supply and demand characteristics of time period t Follow the mean variance is The normal distribution of (I is the identity matrix).

[0062] Specifically, in some embodiments, the step of modeling the relationship between the virtual power plant and other market participants in the electricity market through a multi-agent Markov game based on the output data, the electricity demand data, and the two-layer hybrid game model includes: Construct a multi-agent system based on the virtual power plant and other market participants in the electricity market; Based on the output data, the electricity demand data, and the two-layer hybrid game model, the state space, action space, state transition probability, and reward function of the multi-agent system under the Markov game are determined.

[0063] In this embodiment, the limitations of traditional "static game" are overcome by using Markov games to characterize the dynamic interaction of multiple agents, and combining mixed-strategy Nash equilibrium to improve strategy adaptability. Virtual power plants need to engage in dynamic games with traditional power generators and other aggregators, and traditional static game models cannot reflect the evolution of strategies over time. In this embodiment, Markov game modeling is used to solve for the mixed-strategy Nash equilibrium, as detailed below: Step 1: Markov Game (MG) Modeling. The interaction between the virtual power plant (VPP) and other market participants (traditional power generators (G) and aggregators (A)) is abstracted as a finite-state Markov game, formally defined as: (14) in, It is a collection of intelligent agents (including virtual power plants, traditional power generators, and aggregators). For state space (integrating supply and demand states with historical strategies); For the action space (strategy variables of the virtual power plant); The state transition probability (determined by the supply and demand uncertainty model and market clearing rules); For the reward function (the instantaneous profit of the virtual power plant); It is a discount factor (balancing immediate and future rewards).

[0064] 1) Set of intelligent agents: (15) 2) State space: (16) including supply and demand status and , opponent historical strategy , etc.

[0065] 3) Action space: for virtual power plant , action is a continuous vector, where: (17) where, is the supply offer coefficient, is the demand response incentive coefficient.

[0066] 4) State transition probability : (18) which is determined by the supply and demand uncertainty model and market clearing rules.

[0067] 5) Reward function : the immediate reward of virtual power plant is: (19) where, is the power supply, is the response load of aggregated users, is the power supply cost.

[0068] Step 2: Mixed strategy Nash equilibrium (MSNE) solution. In the embodiments of the present application, the regularization policy gradient (RPG) method is introduced to find the mixed strategy Nash equilibrium in the strategy space, avoiding the local optimum of the traditional method and improving the efficiency of equilibrium solution. The steps are as follows: (1) Strategy parameterization. The strategy of virtual power plant is represented as a random strategy , using Gaussian strategy: (20) where, is the strategy parameter; is the mean network (parameterized by a fully connected neural network), is the diagonal covariance matrix.

[0069] (2) Regularization objective function. To avoid premature convergence of the strategy to a local optimum, an entropy regularization term is introduced to enhance exploration: (21) where, is the strategy entropy, is a regularization coefficient.

[0070] (3) Policy gradient update. The gradient is estimated using Monte Carlo sampling: (22) where K is the number of samples; is the advantage function (estimated by Critic network, which measures the pros and cons of the relative average return of actions), and the gradient update direction optimizes the expected reward and policy entropy at the same time.

[0071] (4) Nash equilibrium verification. Whether the equilibrium is reached is judged by calculating the best response error: (23) where is the optimal policy parameter of agent i, is the policy parameter of other agents. When (threshold, such as 10 -3 ), it is considered to reach the approximate Nash equilibrium.

[0072] Specifically, in some embodiments, the multi-agent reinforcement learning method is used to solve the Markov game to obtain the operation decision strategy of the virtual power plant, including: An agent unit is constructed for each agent in the virtual power plant and the other market subjects; each agent unit includes an actor network and a critic network, the input of the actor network is the local observation state of the agent, and the output is a continuous decision action, the input of the critic network is the joint state and joint action of the agent, and the output is the value evaluation of the current decision strategy; The critic network of each agent is updated by centralized training; The actor network of each agent is applied by distributed execution to output the operation decision strategy of the virtual power plant.

[0073] In the embodiments of the present application, the limitations of traditional "inefficient solution" are broken through, the convergence speed is improved through the MADDPG algorithm, and the real-time decision demand of the power market is adapted. The aforementioned game model has high dimension and complex multi-agent interaction, and requires an efficient solution method. In the embodiments of the present application, the multi-agent deep deterministic policy gradient (MADDPG) algorithm is used, which can improve the convergence speed compared with the traditional method. The specific process is as follows: Step 1: Distributed policy network architecture. A multi-agent deep deterministic policy gradient (MADDPG) algorithm is used to construct an independent Actor-Critic network (separate policy generation and value evaluation) for each virtual power plant.

[0074] (1) Actor network (policy network) Input: Local state (The virtual power plant's own supply and demand data, historical actions, etc.); Output: Continuous action The tanh activation function constrains actions to a specified range. , ]and[ , ]; Network structure: 2 fully connected layers (256 units each, ReLU activation) + output layer (linear activation).

[0075] (2) Critic Network (Value Network) Input: Global state With joint actions ; Output: Value of Joint Actions (Evaluate the strengths and weaknesses of the strategy); Network structure: 2 fully connected layers (512 units each, ReLU activation) + output layer (linear activation).

[0076] (3) Target network (stable training) Set up an independent target Actor / Critic network Track current network parameters through a soft update mechanism: (twenty four) in, This is a soft update coefficient (usually set to 0.001) to ensure training stability.

[0077] Step 2: Experience playback and batch training.

[0078] A central experience pool is built to store multi-agent interaction data, and batch training is used to improve efficiency.

[0079] (1) An empirical tuple is defined as: (25) when When the threshold (e.g., 10⁻³) is reached, it is considered to have reached an approximate Nash equilibrium.

[0080] (2) The training process is as follows: 1. Action Exploration: Add Gaussian noise to the policy output. To achieve the balance of exploration and utilization: (26) Batch sampling: Randomly drawing batches of data from the experience pool. (27) Critic update: Compute target Q-values: (28) Actor update: Boost expected Q-values through policy gradient: (29) Step 3: Collaborative reward mechanism design. Promote "individual-global" collaboration through hierarchical rewards to avoid suboptimal results caused by single profit targets. Design a reward function that considers both individual benefits and global collaboration: (30) Individual reward: (maximize own profit); Collaborative reward: Measure policy collaboration through cosine similarity of incentive coefficients (value range [-1, 1]); : Collaborative weight coefficient (e.g. = 0.1, balance individual and global optimization goals).

[0081] Step 4: Dynamic strategy adjustment and market adaptation. Realize real-time strategy update through online fine-tuning to adapt to market changes. Based on the latest prediction results of LSTM-GARCH, dynamically adjust strategy parameters: (31) where is the online learning rate (usually 10-4), is the cumulative reward gradient of the last periods, ensuring the strategy's rapid response to real-time market changes.

[0082] The technical solution of the present application has at least the following advantages over traditional strategies: 1. Supply and demand collaboration modeling advantage: Through a double-layer mixed game model integrating supply-side power generation optimization and demand-side load regulation, the traditional single-dimensional modeling limitations are broken, and the response capability of the virtual power plant to power market supply and demand fluctuations can be significantly improved.

[0083] 2. Robustness of uncertainty processing: Based on the LSTM-GARCH spatiotemporal uncertainty prediction model, the prediction error is reduced compared to traditional models, which can accurately depict the mean-variance dynamic characteristics of renewable energy and electricity demand, providing reliable input for strategy formulation.

[0084] 3. Efficient game solving: The solution framework combines mixed strategy Nash equilibrium and MADDPG, which has significantly improved convergence speed in multi-virtual power plant game scenarios compared to traditional methods, with strategy update period shortened to minutes, suitable for real-time bidding requirements of the electricity market.

[0085] 4. Strategy flexibility and scalability: support virtual power plant dynamic adjustment of bidding and demand response strategy combination, scalability adapts to different scale market subjects, and provides a unified strategy framework for virtual power plant participation in day-ahead market, real-time market and auxiliary service market.

[0086] The embodiment of the application also provides a virtual power plant operation strategy optimization device considering both supply and demand sides, comprising: An acquisition unit is configured to acquire market data of a power market; wherein the market data comprises supply side data, demand side data and market environment data; A construction unit is configured to construct a double-layer mixed game model of comprehensive supply and demand sides according to the market data; wherein the double-layer mixed game model comprises a lower market clearing model and an upper virtual power plant decision model, the lower market clearing model is configured to perform market clearing to achieve supply and demand balance, and the upper virtual power plant decision model is configured to enable the virtual power plant to formulate an optimized power supply strategy; A prediction unit is configured to establish a supply and demand prediction model according to the market data, and predict output data of power supply and power consumption demand data of users according to the supply and demand prediction model; A modeling unit is configured to model relationships between the virtual power plant and other market subjects in the power market through a multi-agent Markov game according to the output data, the power consumption demand data and the double-layer mixed game model; A solution unit is configured to solve the Markov game through a multi-agent reinforcement learning method to obtain an operation decision strategy of the virtual power plant.

[0087] It can be understood that the contents in the above method embodiments are applicable to the device embodiments, the device embodiments specifically realize the functions of the above method embodiments, and achieve the same beneficial effects as the above method embodiments.

[0088] With reference to Figure 3 The embodiment of the application provides an electronic device, comprising: at least one processor 310; at least one memory 320 configured to store at least one program; When the at least one program is executed by the at least one processor 310, the at least one processor 310 implements the above-mentioned virtual power plant operation strategy optimization method considering both supply and demand sides.

[0089] Similarly, the contents in the above method embodiments are applicable to the electronic device embodiments, the electronic device embodiments specifically realize the functions of the above method embodiments, and achieve the same beneficial effects as the above method embodiments.

[0090] The embodiment of the application further provides a computer readable storage medium, wherein a program executable by the processor 310 is stored, and the program executable by the processor 310 is used for executing the virtual power plant operation strategy optimization method considering both supply and demand ends when executed by the processor 310.

[0091] Similarly, the contents in the method embodiment are applicable to the computer readable storage medium embodiment, the computer readable storage medium embodiment specifically implements the functions same as the method embodiment, and the beneficial effects same as the method embodiment are achieved.

[0092] In some alternative embodiments, the functions / operations mentioned in the block diagram can not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two blocks shown in succession can actually be executed substantially simultaneously or the blocks can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example, and the purpose is to provide a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of larger operations are independently executed.

[0093] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features can be integrated in a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is unnecessary for an understanding of the present application. Rather, given the properties, functions and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be within the routine skill of the engineer, given the disclosure herein. Thus, the present application as set forth in the claims is enabled by those skilled in the art using ordinary skill, without undue experimentation. It can also be understood that the disclosed specific concepts are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.

[0094] If the functions are implemented in software, the functions can be stored in or implemented as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage medium can be any available medium that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, or twisted pair, then the coaxial cable, fiber optic cable, or twisted pair are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-Ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0095] In other words, like a human driver of a vehicle, an autonomous vehicle can be programmed to follow traffic laws and to make decisions based on its environment. For example, an autonomous vehicle can be programmed to follow a speed limit, to stop at a stop sign, to yield to a pedestrian, to merge onto a highway, to change lanes, to park, and so on. In some embodiments, an autonomous vehicle can be programmed to follow traffic laws and to make decisions based on its environment using a machine learning algorithm. For example, an autonomous vehicle can be programmed to follow a speed limit, to stop at a stop sign, to yield to a pedestrian, to merge onto a highway, to change lanes, to park, and so on using a machine learning algorithm.

[0096] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber (optical), and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.

[0097] It should be understood that portions of the application can be realized with hardware, software, firmware or a combination thereof. In the foregoing description, multiple steps or methods can be realized as software or firmware to be executed by a suitable instruction executing system. For example, if realized with hardware, and as in another embodiment, any one or a combination of the following technologies known in the art can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0098] In the above description of the present specification, the description of the terms "one embodiment", "another embodiment", or "certain embodiments" or the like means that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0099] Although the embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made thereto without departing from the principles and spirit of the present application, the scope of which is defined by the claims and their equivalents.

[0100] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the embodiments, and those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present application, and these equivalent modifications or substitutions are included in the scope defined by the claims of the present application.

Claims

1. A method for optimizing the operation strategy of a virtual power plant considering both supply and demand, characterized in that, The method includes: Acquire market data for the electricity market; wherein the market data includes supply-side data, demand-side data, and market environment data; Based on the market data, a two-layer hybrid game model integrating supply and demand is constructed. The two-layer hybrid game model includes a lower-layer market clearing model and an upper-layer virtual power plant decision model. The lower-layer market clearing model is used to clear the market to achieve supply and demand balance, and the upper-layer virtual power plant decision model is used for virtual power plants to formulate strategies for optimizing power supply. A supply and demand forecasting model is established based on the market data, and the power output data and user electricity demand data are predicted based on the supply and demand forecasting model. Based on the power output data, the electricity demand data, and the two-layer hybrid game model, the relationship between the virtual power plant and other market participants in the electricity market is modeled through a multi-agent Markov game. The Markov game is solved using a multi-agent reinforcement learning method to obtain the operation decision strategy of the virtual power plant.

2. The method for optimizing the operation strategy of a virtual power plant considering both supply and demand sides as described in claim 1, characterized in that, The step of constructing a two-layer hybrid game model integrating supply and demand based on the market data includes: With the goal of maximizing social welfare in the market, a lower-level market clearing model is constructed, and constraints corresponding to the lower-level market clearing model are set. The upper-level virtual power plant decision model is constructed with the goal of maximizing the operating profit of the virtual power plant.

3. The method for optimizing the operation strategy of a virtual power plant considering both supply and demand sides as described in claim 2, characterized in that, The constraints corresponding to the lower-level market clearing model include supply and demand power balance constraints, virtual power plant output constraints, and demand response load constraints.

4. The method for optimizing the operation strategy of a virtual power plant considering both supply and demand sides as described in claim 1, characterized in that, The step of establishing a supply and demand forecasting model based on the market data, and forecasting power output data and user electricity demand data based on the supply and demand forecasting model, includes: Based on the market data, a long short-term memory network model and a generalized autoregressive conditional heteroscedasticity model are established. The time-series dependency features in the output data and the electricity demand data are determined by the Long Short-Term Memory Network model, and the mean values ​​of the output data and the electricity demand data in future time periods are predicted. Based on the residuals predicted by the mean, the volatility of the power output data and the electricity demand data in future periods is predicted using the generalized autoregressive conditional heteroscedasticity model.

5. The method for optimizing the operation strategy of a virtual power plant considering both supply and demand sides as described in claim 1, characterized in that, The step of modeling the relationship between the virtual power plant and other market participants in the electricity market using a multi-agent Markov game based on the power output data, the electricity demand data, and the two-layer hybrid game model includes: Construct a multi-agent system based on the virtual power plant and other market participants in the electricity market; Based on the output data, the electricity demand data, and the two-layer hybrid game model, the state space, action space, state transition probability, and reward function of the multi-agent system under the Markov game are determined.

6. The method for optimizing the operation strategy of a virtual power plant considering both supply and demand sides as described in claim 1, characterized in that, The step of solving the Markov game using a multi-agent reinforcement learning method to obtain the operating decision strategy for the virtual power plant includes: For each agent in the virtual power plant and other market entities, an agent unit is constructed; each agent unit includes an actor network and a critic network. The input of the actor network is the local observation state of the agent, and the output is a continuous decision action. The input of the critic network is the joint state and joint action of the agent, and the output is a value assessment of the current decision strategy. The critic network for each agent is updated through centralized training. The actor network of each of the intelligent agents is applied in a distributed manner to output the operation decision strategy of the virtual power plant.

7. A device for optimizing the operation strategy of a virtual power plant considering both supply and demand, characterized in that, The device includes: An acquisition unit is used to acquire market data of the electricity market; wherein, the market data includes supply-side data, demand-side data, and market environment data; The construction unit is used to construct a two-layer hybrid game model that integrates supply and demand based on the market data. The two-layer hybrid game model includes a lower-layer market clearing model and an upper-layer virtual power plant decision model. The lower-layer market clearing model is used to clear the market to achieve supply and demand balance, and the upper-layer virtual power plant decision model is used for virtual power plants to formulate strategies for optimizing power supply. The forecasting unit is used to establish a supply and demand forecasting model based on the market data, and to forecast power output data and user electricity demand data based on the supply and demand forecasting model. The modeling unit is used to model the relationship between the virtual power plant and other market participants in the electricity market through a multi-agent Markov game based on the output data, the electricity demand data, and the two-layer hybrid game model. The solution unit is used to solve the Markov game using a multi-agent reinforcement learning method to obtain the operation decision strategy of the virtual power plant.

8. The virtual power plant operation strategy optimization device considering both supply and demand sides according to claim 7, characterized in that, The building unit is specifically used for: With the goal of maximizing social welfare in the market, a lower-level market clearing model is constructed, and constraints corresponding to the lower-level market clearing model are set. The upper-level virtual power plant decision model is constructed with the goal of maximizing the operating profit of the virtual power plant.

9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a method for optimizing the operation strategy of a virtual power plant considering both supply and demand sides as described in any one of claims 1-6.

10. A computer-readable storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to implement a virtual power plant operation strategy optimization method considering both supply and demand sides as described in any one of claims 1-6.