Scheduling decision-making method, device and equipment for virtual power plant and storage medium
By using long sequence prediction models and deep deterministic policy gradient models in virtual power plants, the problem of insufficient scheduling accuracy in large-scale virtual power plant environments is solved, and more efficient resource scheduling and power system optimization are achieved.
Patent Information
- Application Number
- CN202410274954.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies lack scheduling accuracy when dealing with large-scale, highly dynamic virtual power plant environments, making it difficult to effectively optimize resource allocation and the flexibility and reliability of power systems.
Using a long sequence prediction model and a deep deterministic policy gradient model, the operating status data of the virtual power plant is obtained to perform state prediction and scheduling strategy optimization. The state prediction data is used to update the rewards and losses of the scheduling network model, and the scheduling network model is trained to generate a more accurate scheduling strategy.
It improves the accuracy and efficiency of virtual power plant scheduling, can better respond to dynamic changes in the power market, and achieve efficient allocation and optimized scheduling of resources.
Smart Images

Figure CN120634068A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of virtual power plants, and in particular to scheduling decision methods, devices, equipment and storage media for virtual power plants. Background Art
[0002] In the field of power engineering technology, there has been extensive research and development, particularly in the concept and application of virtual power plants (VPPs). A VPP uses advanced information and communication technologies to centrally manage distributed power resources (such as wind, solar, small hydropower, and energy storage facilities) to provide power services similar to a single power plant. The primary goal of this concept is to optimize the allocation of energy resources, improve the flexibility and reliability of power systems, and reduce the impact of environmental impacts on power systems.
[0003] In recent years, with the widespread use of renewable energy and the diversification of electricity market demand, the management and optimization of virtual power plants have become increasingly complex. Current implementation solutions involve the use of various algorithms and techniques to predict power load and optimize resource scheduling. Among them, a common method is to use time series analysis technology to predict power load. For example, the traditional autoregressive moving average (ARMA) model and its variants are used in some cases to predict short-term or medium-term power demand. On the other hand, for resource scheduling, existing solutions include linear programming, integer programming, mixed integer linear programming and other methods. These methods are effective in dealing with specific types of scheduling problems, but may be limited when dealing with large-scale, highly dynamic virtual power plant environments, which is not conducive to improving the accuracy of scheduling. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a scheduling decision method, device, equipment and storage medium for a virtual power plant to solve the problem in the prior art that when dealing with large-scale, highly dynamic virtual power plant environments, the scheduling of virtual power plants is limited, which is not conducive to improving the accuracy of scheduling.
[0005] A first aspect of an embodiment of the present application provides a scheduling decision method for a virtual power plant, the method comprising:
[0006] Obtaining operational status data of the virtual power plant;
[0007] Inputting the operating status data into a pre-trained long sequence prediction model to obtain status prediction data of the virtual power plant;
[0008] updating a reward of a scheduling network model for outputting a scheduling policy according to the state prediction data, determining a loss of the scheduling network model according to the reward, and completing training of the scheduling network model according to the loss;
[0009] The state prediction data is input into the trained scheduling network model to obtain the scheduling strategy of the virtual power plant.
[0010] In conjunction with the first aspect, in a first possible implementation of the first aspect, updating a reward of a scheduling network model for outputting a scheduling policy according to the state prediction data includes:
[0011] Obtaining the dispatching actions and market data taken to obtain the operational status data of the virtual power plant;
[0012] The operating status data, the status prediction data, the scheduling behavior and the market data are input into a pre-trained reward model to update the reward of the scheduling network model.
[0013] In combination with the first possible implementation manner of the first aspect, in a second possible implementation manner of the first aspect, the scheduling network model is a deep deterministic policy gradient model;
[0014] The loss of the scheduling network model is determined based on the reward, including:
[0015] Inputting state prediction data and scheduling behavior into the deep deterministic policy gradient model;
[0016] According to the reward, combined with a predetermined reviewer network utilization value function and an actor network utilization strategy function, a loss corresponding to the scheduling behavior is determined.
[0017] In conjunction with the second possible implementation of the first aspect, in a third possible implementation of the first aspect, determining the loss corresponding to the scheduling behavior based on the reward, combined with a predetermined reviewer network utilization value function and an actor network utilization strategy function, includes:
[0018] Determine the commentator network usage value function for the next state based on the actor network usage strategy function for the next state;
[0019] The loss corresponding to the scheduling behavior is determined based on the reviewer network usage value function of the next state, a preset discount factor, the reviewer network usage value function of the current state, and the reward.
[0020] In combination with the first aspect, in a fourth possible implementation manner of the first aspect, before updating the reward of the scheduling network model for outputting the scheduling policy according to the state prediction data, the method further includes:
[0021] determining a mean and a variance of the state prediction data;
[0022] The state prediction data is normalized according to the mean and variance of the state prediction data.
[0023] In combination with the first aspect, in a fifth possible implementation manner of the first aspect, before updating the reward of the scheduling network model for outputting the scheduling policy according to the state prediction data, the method further includes:
[0024] Obtaining the time corresponding to the state prediction data;
[0025] By embedding a time coding function, the time corresponding to the state prediction data is encoded into a continuous value.
[0026] In combination with any one of the first aspect to the fifth possible implementation manner of the first aspect, in a sixth possible implementation manner of the first aspect, the long sequence prediction model includes an Informer deep learning model.
[0027] A second aspect of an embodiment of the present application provides a scheduling decision device for a virtual power plant, the device comprising:
[0028] an operating status data acquisition unit, configured to acquire operating status data of the virtual power plant;
[0029] A prediction unit, configured to input the operating status data into a pre-trained long sequence prediction model to obtain status prediction data of the virtual power plant;
[0030] a training unit, configured to update a reward of a scheduling network model for outputting a scheduling policy according to the state prediction data, determine a loss of the scheduling network model according to the reward, and complete training of the scheduling network model according to the loss;
[0031] The scheduling strategy generation unit is used to input the state prediction data into the trained scheduling network model to obtain the scheduling strategy of the virtual power plant.
[0032] In conjunction with the second aspect, in a first possible implementation manner of the second aspect, the training unit includes:
[0033] The data acquisition subunit acquires the dispatching behavior and market data taken by the operation status data of the virtual power plant;
[0034] The updating subunit is used to input the operating status data, the status prediction data, the scheduling behavior and the market data into the pre-trained reward model to update the reward of the scheduling network model.
[0035] In combination with the first possible implementation manner of the second aspect, in a second possible implementation manner of the second aspect, the scheduling network model is a deep deterministic policy gradient model;
[0036] The training unit comprises:
[0037] An input subunit, configured to input state prediction data and scheduling behavior into the deep deterministic policy gradient model;
[0038] The loss determination subunit is used to determine the loss corresponding to the scheduling behavior based on the reward, combined with a predetermined commentator network utilization value function and an actor network utilization strategy function.
[0039] In conjunction with the second possible implementation manner of the second aspect, in a third possible implementation manner of the second aspect, the loss determination subunit includes:
[0040] A first function determination module is used to determine the commentator network usage value function of the next state according to the actor network usage strategy function of the next state;
[0041] The loss determination module is used to determine the loss corresponding to the scheduling behavior based on the reviewer network usage value function of the next state, a preset discount factor, the reviewer network usage value function of the current state, and the reward.
[0042] In combination with the second aspect, in a fourth possible implementation of the second aspect, the apparatus further includes:
[0043] a data determination unit, configured to determine a mean and a variance of the state prediction data;
[0044] A normalization unit is used to normalize the state prediction data according to the mean and variance of the state prediction data.
[0045] In conjunction with the second aspect, in a fifth possible implementation of the second aspect, the apparatus further includes:
[0046] A time acquisition unit, configured to acquire the time corresponding to the state prediction data;
[0047] The time coding unit is used to encode the time corresponding to the state prediction data into a continuous value by embedding a time coding function.
[0048] In combination with any one of the second aspect to the fifth possible implementation of the second aspect, in a sixth possible implementation of the second aspect, the long sequence prediction model includes an Informer deep learning model.
[0049] A third aspect of an embodiment of the present application provides a virtual power plant, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method described in any one of the first aspects are implemented.
[0050] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.
[0051] Compared with the prior art, the embodiments of the present application have the following advantages: by acquiring the operating status data of the virtual power plant and inputting the operating status data into a long-sequence prediction model, the embodiments of the present application can adapt to the large-scale, highly dynamic virtual power plant environment, effectively obtain state prediction data, update the rewards of the scheduling network model based on the state prediction data, determine the losses of the scheduling network model based on the rewards, train the scheduling network model based on the losses, and input the prediction data into the scheduling network model to obtain the scheduling strategy of the virtual power plant. Because the state prediction data is used to determine the rewards and losses, the scheduling network model can more accurately adapt to the scheduling calculation of the state prediction data, thereby obtaining a more accurate scheduling strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0053] Figure 1 This is a schematic diagram of an implementation flow of a scheduling decision-making method for a virtual power plant provided in an embodiment of the present application;
[0054] Figure 2 This is a schematic diagram of an implementation flow of an update model reward provided in an embodiment of the present application;
[0055] Figure 3 This is a schematic diagram of an implementation flow for determining the loss corresponding to a scheduling behavior provided by an embodiment of the present application;
[0056] Figure 4 This is a schematic diagram of a scheduling decision system for a virtual power plant provided in an embodiment of the present application;
[0057] Figure 5 This is a schematic diagram of a scheduling decision-making device for a virtual power plant provided in an embodiment of the present application;
[0058] Figure 6 This is a schematic diagram of a virtual power plant provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0060] In order to illustrate the technical solution described in this application, specific embodiments are provided below.
[0061] In the field of power engineering technology, there has been extensive research and development, particularly in the concept and application of virtual power plants (VPPs). A VPP uses advanced information and communication technologies to centrally manage distributed power resources (such as wind, solar, small hydropower, and energy storage facilities) to provide power services similar to a single power plant. The primary goal of this concept is to optimize the allocation of energy resources, improve the flexibility and reliability of power systems, and reduce the impact of environmental impacts on power systems.
[0062] In recent years, with the widespread use of renewable energy and the diversification of electricity market demand, the management and optimization of virtual power plants (VPPs) have become increasingly complex. A key challenge is accurately forecasting power load and properly dispatching resources. Traditional approaches often rely on empirical rules or simplified mathematical models, but these methods have limitations in handling the uncertainties of renewable energy and responding to market fluctuations.
[0063] Current implementation schemes involve the use of various algorithms and techniques to predict power load and optimize resource scheduling. Among them, a common method is to use time series analysis technology to predict power load. For example, the traditional autoregressive moving average (ARMA) model and its variants are used in some cases to predict short-term or medium-term power demand. On the other hand, for resource scheduling, existing schemes include linear programming, integer programming, mixed integer linear programming and other methods. These methods are effective in dealing with specific types of scheduling problems, but may be limited when dealing with large-scale, highly dynamic virtual power plant environments. They lack the necessary flexibility and adaptability, which is not conducive to improving the accuracy of scheduling.
[0064] To solve the above problems, the present application proposes a scheduling decision method for a virtual power plant, such as Figure 1 As shown, the method includes:
[0065] In S101, the operating status data of the virtual power plant is obtained.
[0066] The virtual power plant (VPP) in the embodiments of this application is an energy management system that aggregates and coordinates distributed energy resources (such as distributed generation systems, energy storage systems, and controllable loads) through modern information technology and advanced software systems to form a virtual integrated power plant. Like traditional power plants, virtual power plants can participate in bidding transactions in the electricity market, provide grid ancillary services, and regulate the supply and demand balance of the power grid.
[0067] The operating status data of the virtual power plant may include one or more of load data, weather data related to renewable energy, renewable energy production and energy storage status.
[0068] The operating status data of the virtual inductor obtained in the embodiment of the present application can be used as the first sample data of the long sequence prediction model, that is, the operating status data can be used to train the long sequence prediction model, thereby optimizing the parameters of the long sequence prediction model, so that the long sequence prediction model can obtain more accurate status prediction data.
[0069] Alternatively, the operation status data can also be used as the second sample data of the scheduling network model for training the scheduling network model.
[0070] Alternatively, after the training of the above model is completed, it can also be used as the basis for scheduling decisions and used to calculate the scheduling strategy.
[0071] In S102, the operating status data is input into a pre-trained long sequence prediction model to obtain status prediction data of the virtual power plant.
[0072] In the embodiment of the present application, before inputting the operating status data into the pre-trained long sequence prediction model, the long sequence prediction model can also be trained using sample data so that the long sequence prediction model can calculate accurate status prediction data.
[0073] Long sequence prediction models may include Informer models, Transformer models, etc.
[0074] When training a long sequence prediction model, the first sample data can be determined first. The first sample data used to train the long sequence prediction model includes operating status data and state prediction data corresponding to the operating status data. The state prediction data can be state prediction data within a predetermined time period after the operating status data. The first sample data can be obtained by sampling. For example, the first operating status data can be sampled at time T1, and the second operating status data can be sampled within a predetermined time period after time T1 to serve as the predicted state data for the first operating status data.
[0075] The first operating status data in the first sample data can be input into the long sequence prediction model to be trained, and the third operating status data can be obtained through prediction. The deviation between the third operating status data and the second operating status data is determined, and the parameters of the long sequence prediction model are updated according to the deviation until the deviation meets the predetermined requirements. The training of the long sequence prediction model can be completed.
[0076] After the long sequence prediction model training is completed, the currently collected operation status data can be input into the long sequence prediction model to calculate the status prediction data.
[0077] The state prediction data in the embodiment of the present application may correspond to the operating state data, and may include one or more of the prediction data such as power load prediction data, distributed energy output prediction data, energy storage system state prediction, electricity price prediction, charging demand prediction, grid voltage and frequency prediction, etc.
[0078] After obtaining the predicted operation data, in order to adapt to the input requirements of the scheduling network model, the predicted operation data can be preprocessed, including standardization processing and / or feature encoding processing.
[0079] When the predicted operation data is standardized, the mean and variance of the obtained predicted operation data may be calculated first, and the pre-processed data may be determined based on the calculated mean and variance.
[0080] For example, the formula can be used to determine the pre-processed forecasted operational data:
[0081]
[0082] Where X is the original forecasted operation data, X′ is the processed forecasted operation data, μ is the mean, and σ is the variance.
[0083] When encoding the predicted operation data or the predicted operation data after standardization, the time corresponding to the predicted operation data can be obtained to obtain time series data.
[0084] For example, you can use embedding technology to:
[0085] E t =Embedding(t)
[0086] Encode time series data into continuous numerical values, where E tis the encoded temporal feature. Embedding represents the embedding encoding function, which can include linear encoding, periodic encoding, and other encoding functions. t is the original timestamp. This encoding enables the model to better understand and process temporal information, resulting in more accurate predictions of future load. Feature encoding converts preprocessed data into a format that can be effectively processed by scheduling network models, including deep deterministic policy gradient models.
[0087] In S103, the reward of the scheduling network model for outputting the scheduling policy is updated according to the state prediction data, the loss of the scheduling network model is determined according to the reward, and the training of the scheduling network model is completed according to the loss.
[0088] The scheduling network model in the embodiments of the present application may include a variety of different deep reinforcement learning network models, including the deep deterministic policy gradient model (full name in English: Deep Deterministic Policy Gradient, abbreviated as DDPG in English), TRPO (full name in English: Trust Region Policy Optimization, full name in Chinese: Trust Region Policy Optimization Model), etc.
[0089] DDPG consists of an actor network and a critic network. The actor network is responsible for generating actions in a given state, while the critic network evaluates the expected reward of the action.
[0090] Here, the actor network uses a policy function π to determine the action a to perform in a given state s.
[0091] It can be expressed by the formula:
[0092] a=π(s|θ π ), where θ π Parameters representing the actor network.
[0093] The critic network uses a value function Q to evaluate the expected return for a given state s and action a. The expected return can be expressed as: Q(s,a|θ Q ), where θ Q represents the parameters of the critic network.
[0094] The scheduling network model in the embodiment of the present application may also include a training process before use.
[0095] The second sample data used for training the scheduling network model can include operational status data and the value gained from performing different actions within the operational status. A loss function can be determined based on the determined value to determine the magnitude of the loss for different actions. The parameters of the scheduling network model are updated based on the loss until the scheduling network model outputs the optimal scheduling behavior.
[0096] Among them, the state of the scheduling network model is determined according to the state prediction data, which can be as follows: Figure 2 Shown, including:
[0097] In S201, the dispatching behavior and market data taken based on the operating status data of the virtual power plant are obtained.
[0098] Among them, the scheduling behavior may include one or more of power output intensity regulation, electric energy storage / discharge regulation, power trading regulation, and load regulation.
[0099] Market data may include, for example, electricity price data.
[0100] In S202, the operating status data, the status prediction data, the scheduling behavior and the market data are input into a pre-trained reward model to update the reward of the scheduling network model.
[0101] In the embodiment of the present application, the reward corresponding to the executed scheduling behavior can be determined by a reward model. The reward model can be:
[0102] R(s,a)=Model dynamic (s,a,MarketData,PredictedData)
[0103] Where R represents the reward function, which is used to evaluate the return or reward under a given state and action. s is the operating state data of the virtual power plant at a specific moment. a represents the scheduling behavior or action, that is, the scheduling behavior taken under a specific state. Model dynamic It is a dynamic reward model for artificial intelligence, used to adjust the reward function based on the input state prediction data. Data For market data, it reflects the current market conditions. It can automatically obtain the corresponding market data from the Internet. Data State prediction data can be generated using long-sequence prediction models based on historical and real-time data. By controlling the reward parameters in reinforcement learning using a large reward model, the speed and results of reinforcement learning training for scheduling network models can be optimized. This large reward model enables scheduling network models, such as the DDPG model, to flexibly respond to dynamic changes in the power market during resource scheduling, achieving more efficient power resource allocation and optimized scheduling.
[0104] After determining the reward of the scheduling network model, the loss of the scheduling network model can be determined based on the reward. For example, when the scheduling network model is a deep deterministic policy gradient model, the process of determining the loss can be as follows: Figure 3 Shown, including:
[0105] In S301 , state prediction data and scheduling behavior are input into the deep deterministic policy gradient model.
[0106] When state prediction data is input into a deep deterministic policy gradient model, it can be used to determine the next state after the scheduling behavior is executed, as well as the scheduling behavior predicted by the next state.
[0107] In S302 , the loss corresponding to the scheduling behavior is determined based on the reward, in combination with a predetermined reviewer network utilization value function and an actor network utilization strategy function.
[0108] Based on the current state and the scheduling behavior of the current state, the value of executing the scheduling behavior in the current state can be determined by using the value function through the evaluator network.
[0109] Based on the predicted next state and in combination with the predicted scheduling behavior of the next state, the value of the scheduling behavior of the next state is determined through the actor network using a value function.
[0110] Based on the value of the current state and the value of the next state, combined with the reward and discount factor, the loss function of the deep deterministic policy gradient model can be determined, which can be expressed as:
[0111] L=(r+γQ′(s′,π′(s′|θ π′ )|θ Q′ )-Q(s,a|θ Q )) 2 , where r is the reward, γ is the discount factor, s′ is the next state, Q′ and π′ are the commentator network usage value function and the actor network usage value function respectively. θ Q is the parameter of the reviewer network value function, θ π are the parameters of the actor network value function.
[0112] In S104, the state prediction data is input into the trained dispatching network model to obtain the dispatching strategy of the virtual power plant.
[0113] By feeding state prediction data into a trained dispatch network model, the model can more accurately understand and predict future changes, leading to more effective dispatch decisions. Conversely, without state prediction data, the model will be unable to fully account for future load fluctuations, resulting in reduced accuracy and efficiency in dispatch decisions. Furthermore, by using state prediction data to update the dispatch network model's losses in real time, the model can generate dispatch strategies more reliably and effectively.
[0114] Therefore, inputting state prediction data can effectively improve the accuracy of DDPG model training and updating, which can significantly improve the accuracy of decision-making and the overall efficiency of the system.
[0115] In addition, the scheduling decision method of the virtual power plant shown in this application can be integrated into the virtual power plant management system, and an intuitive and easy-to-use wording interface can be developed to enable operators to effectively monitor and manage the operation of the virtual power plant. Based on test results and wording feedback, the system can be continuously optimized and upgraded to make the system more adaptable.
[0116] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0117] In addition, if Figure 4 As shown, the present application also provides a schematic diagram of a scheduling decision system for a virtual power plant. The system includes a data center, a decision center, and a prediction center. The data center is used to collect virtual power plant data, including real-time or historical photovoltaic data, wind power data, weather data, nuclear power data, etc. Through the collected data, the prediction center can use a long sequence prediction model to perform state prediction and obtain state prediction data. Based on the predicted state prediction data, combined with the operating state data, it can be used to output a scheduling strategy for the decision center, so that the virtual power plant can be effectively scheduled and multi-objective optimization can be effectively achieved.
[0118] Among them, the data of the data center can be used to train and update the models of the decision center and the prediction center, so that the system can be dynamically optimized and improve the system's adaptability to the environment.
[0119] Figure 5 A schematic diagram of a scheduling decision-making device for a virtual power plant provided in an embodiment of the present application includes:
[0120] An operating status data acquisition unit 501 is configured to acquire operating status data of the virtual power plant;
[0121] The prediction unit 502 is configured to input the operation status data into a pre-trained long sequence prediction model to obtain status prediction data of the virtual power plant;
[0122] A training unit 503 is configured to update a reward of a scheduling network model for outputting a scheduling policy according to the state prediction data, determine a loss of the scheduling network model according to the reward, and complete training of the scheduling network model according to the loss;
[0123] The dispatching strategy generating unit 504 is configured to input the state prediction data into the trained dispatching network model to obtain the dispatching strategy of the virtual power plant.
[0124] Figure 5 The scheduling decision-making device of the virtual power plant shown in Figure 1 The scheduling decision method of the virtual power plant shown corresponds to this.
[0125] Figure 6 Schematic diagram of the scheduling decision-making device of the virtual power plant provided in the embodiment of the present application. Figure 6 As shown, the virtual power plant scheduling decision device 6 of this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60, such as a virtual power plant scheduling decision program. When the processor 60 executes the computer program 62, the steps of the aforementioned virtual power plant scheduling decision method embodiments are implemented. Alternatively, when the processor 60 executes the computer program 62, the functions of the modules / units in the aforementioned device embodiments are implemented.
[0126] For example, the computer program 62 may be divided into one or more modules / units, which are stored in the memory 61 and executed by the processor 60 to implement the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 62 in the scheduling decision-making device 6 of the virtual power plant.
[0127] The scheduling decision device 6 of the virtual power plant can be a computing device such as a desktop computer, a notebook, a palmtop computer, a cloud server, etc. The scheduling decision device of the virtual power plant can include, but is not limited to, a processor 60 and a memory 61. It can be understood by those skilled in the art that Figure 6 It is only an example of the scheduling decision device 6 of the virtual power plant and does not constitute a limitation of the scheduling decision device 6 of the virtual power plant. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the scheduling decision device of the virtual power plant may also include input and output devices, network access devices, buses, etc.
[0128] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0129] The memory 61 may be an internal storage unit of the scheduling decision device 6 of the virtual power plant, such as a hard disk or memory of the scheduling decision device 6 of the virtual power plant. The memory 61 may also be an external storage device of the scheduling decision device 6 of the virtual power plant, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the scheduling decision device 6 of the virtual power plant. Furthermore, the memory 61 may also include both an internal storage unit and an external storage device of the scheduling decision device 6 of the virtual power plant. The memory 61 is used to store the computer program and other programs and data required by the scheduling decision device of the virtual power plant. The memory 61 may also be used to temporarily store data that has been output or is to be output.
[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0131] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0132] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0133] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0134] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0135] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0136] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0137] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A scheduling decision method for a virtual power plant, characterized in that: The method comprises: Obtaining operational status data of the virtual power plant; Inputting the operating status data into a pre-trained long sequence prediction model to obtain status prediction data of the virtual power plant; updating a reward of a scheduling network model for outputting a scheduling policy according to the state prediction data, determining a loss of the scheduling network model according to the reward, and completing training of the scheduling network model according to the loss; The state prediction data is input into the trained scheduling network model to obtain the scheduling strategy of the virtual power plant.
2. The method according to claim 1, characterized in that Updating the reward of the scheduling network model for outputting the scheduling policy according to the state prediction data includes: Obtaining the dispatching actions and market data taken to obtain the operational status data of the virtual power plant; The operating status data, the status prediction data, the scheduling behavior and the market data are input into a pre-trained reward model to update the reward of the scheduling network model.
3. The method according to claim 2, characterized in that The scheduling network model is a deep deterministic policy gradient model; The loss of the scheduling network model is determined based on the reward, including: Inputting state prediction data and scheduling behavior into the deep deterministic policy gradient model; According to the reward, combined with a predetermined reviewer network utilization value function and an actor network utilization strategy function, a loss corresponding to the scheduling behavior is determined.
4. The method according to claim 3, characterized in that According to the reward, combined with a predetermined reviewer network utilization value function and an actor network utilization strategy function, the loss corresponding to the scheduling behavior is determined, including: Determine the commentator network usage value function for the next state based on the actor network usage strategy function for the next state; The loss corresponding to the scheduling behavior is determined based on the reviewer network usage value function of the next state, a preset discount factor, the reviewer network usage value function of the current state, and the reward.
5. The method according to claim 1, characterized in that Before updating the reward of the scheduling network model for outputting the scheduling policy according to the state prediction data, the method further includes: determining a mean and a variance of the state prediction data; The state prediction data is normalized according to the mean and variance of the state prediction data.
6. The method according to claim 1, characterized in that Before updating the reward of the scheduling network model for outputting the scheduling policy according to the state prediction data, the method further includes: Obtaining the time corresponding to the state prediction data; By embedding a time coding function, the time corresponding to the state prediction data is encoded into a continuous value.
7. The method according to any one of claims 1 to 6, characterized in that The long sequence prediction model includes an Informer deep learning model.
8. A scheduling decision-making device for a virtual power plant, characterized in that: The device comprises: an operating status data acquisition unit, configured to acquire operating status data of the virtual power plant; A prediction unit, configured to input the operating status data into a pre-trained long sequence prediction model to obtain status prediction data of the virtual power plant; a training unit, configured to update a reward of a scheduling network model for outputting a scheduling policy according to the state prediction data, determine a loss of the scheduling network model according to the reward, and complete training of the scheduling network model according to the loss; The scheduling strategy generation unit is used to input the state prediction data into the trained scheduling network model to obtain the scheduling strategy of the virtual power plant.
9. A scheduling decision device for a virtual power plant, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Updating method and device of power station dispatching information, equipment and storage medium
CN122092390A