A virtual power plant source storage load intelligent scheduling method and system based on deep reinforcement learning

CN122553253APending Publication Date: 2026-08-11XIAN THERMAL POWER RES INST CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0008]本发明的目的在于克服现有技术在处理虚拟电厂多维资源协同、应对现货市场剧烈波动方面存在计算复杂度高、对预测依赖性强,以及自适应能力差的问题,提供一种基于深度强化学习的虚拟电厂源储荷智能调度方法及系统

Benefits of technology

本发明提供一种基于深度强化学习的虚拟电厂源储荷智能调度方法及系统,通过深度强化学习模型自动生成并优化市场策略,能够从高维、复杂且动态变化的市场数据中学习最优决策模式,显著超越了传统基于规则或简单模型的策略,将调频市场与现货市场的决策耦合在一个统一的模型框架内,实现了真正的多市场协同优化,避免了分市场独立决策可能导致的收益冲突或机会损失,从而最大化综合收益。在模型输出动作后,立即基于储能系统实时物理状态进行边界校验与安全重塑,确保所有指令均在系统安全运行范围内执行。这实现了经济性追求与物理安全硬约束的无缝融合,从根本上防止了因过度追求收益而引发的设备损坏或运行风险。采用复合奖励信号(融合市场收益与边界校验结果)对模型进行在线或周期性迭代更新,使模型能够持续跟踪电力市场结构、价格规律、系统老化状态的变化,实现策略的动态自适应与自我完善。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122553253A_ABST
    Figure CN122553253A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for intelligent scheduling of virtual power plant energy sources, storage, and load based on deep reinforcement learning, belonging to the field of power grid dispatching technology. By automatically generating and optimizing market strategies through a deep reinforcement learning model, it can learn optimal decision-making patterns from high-dimensional, complex, and dynamically changing market data, significantly surpassing traditional rule-based or simple model-based strategies. It couples the decisions of the frequency regulation market and the spot market within a unified model framework, achieving true multi-market collaborative optimization and avoiding potential revenue conflicts or opportunity losses that may result from independent decisions in different markets, thereby maximizing overall returns. After the model outputs actions, boundary checks and safety reshaping are immediately performed based on the real-time physical state of the energy storage system, ensuring that all instructions are executed within the system's safe operating range. This achieves a seamless integration of economic pursuit and hard physical safety constraints, fundamentally preventing equipment damage or operational risks caused by excessive pursuit of returns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power grid dispatching technology, specifically relating to a method and system for intelligent dispatching of virtual power plant sources, storage and load based on deep reinforcement learning. Background Technology

[0002] With the deepening of the global energy structure transformation, the penetration rate of distributed energy resources (DERs), represented by wind and solar power, continues to increase. However, these distributed energy sources have significant intermittency, volatility, and randomness, and their large-scale grid connection poses a severe challenge to the safe and stable operation of the power system and real-time power balance. At the same time, flexible resources such as demand-side flexible loads and energy storage systems are becoming increasingly abundant, but they are widely distributed, have small individual capacities, and vary in characteristics, making them difficult to utilize directly and effectively by traditional grid dispatching models. Against this backdrop, the Virtual Power Plant (VPP) has emerged as an innovative technological architecture and business model. Through advanced information and communication technologies and intelligent control platforms, it aggregates, coordinates, and optimizes the control of various resources such as geographically dispersed distributed power sources, energy storage devices, and flexible adjustable loads, forming a single controllable entity equivalent to a traditional power plant that can participate in grid operation and electricity market transactions.

[0003] The core value of virtual power plants lies in achieving three major goals through intelligent coordination of "source-storage-load": first, optimizing energy system operation and improving overall energy efficiency and economics; second, enhancing trading and profitability in the electricity market; and third, promoting the local consumption and friendly integration of a high proportion of distributed renewable energy. Among these, participating in bidding and quantity decisions in the electricity spot market (including day-ahead and real-time markets) is a key way for virtual power plants to realize their core commercial value and obtain operating revenue. In this process, virtual power plants need to simultaneously complete two closely coupled tasks: externally, as market participants, submitting competitive electricity and price curves to the trading center; internally, conducting precise coordinated scheduling planning for various aggregated heterogeneous resources to ensure that actual output matches market contracts.

[0004] However, virtual power plants face severe challenges in their spot market bidding decisions due to multiple and significant uncertainties. External uncertainties primarily stem from the dramatic fluctuations in electricity spot market prices, which are influenced by multiple factors such as system load, network congestion, renewable energy output, and market dynamics, exhibiting strong nonlinearity, randomness, and time-series correlation. Internal uncertainties arise from forecasting errors in wind and solar power output, as well as the uncertainty of flexible load response behavior. These intertwined internal and external uncertainties make the decision-making environment for virtual power plants highly complex.

[0005] Currently, the scheduling and pricing strategies for virtual power plants primarily rely on traditional operations research optimization methods, such as mixed-integer linear programming, stochastic programming, and robust optimization. These methods construct precise mathematical models, incorporating physical constraints and market rules into a unified framework for solution. While effective under certain conditions, their limitations become increasingly apparent when facing the complex real-world environment of the electricity spot market: First, these models heavily depend on high-precision electricity price and renewable energy output forecasts, and forecast errors are unavoidable in the volatile spot market, leading to poor adaptability of decision-making results to the actual environment and a high risk of response delays or defaults. Second, when the aggregated "source-storage-load" resources are numerous and diverse, the number of model variables and constraints increases dramatically, placing a heavy computational burden on traditional methods and making it difficult to meet the rapid response requirements of spot markets (especially real-time markets) for minute-level or even second-level online rolling decisions. Furthermore, traditional mathematical models have limited ability to characterize the complex nonlinear dynamics of market electricity prices and the multi-agent game environment, making it difficult to achieve dynamic adaptive optimization.

[0006] To overcome the shortcomings of traditional methods, some studies have attempted to introduce heuristic algorithms (such as particle swarm optimization) or basic reinforcement learning (such as Q-learning). However, in the continuous high-dimensional action space scenario of virtual power plant "source-storage-load" coordination, these methods often suffer from slow convergence speed, low sample utilization efficiency, and susceptibility to getting trapped in local optima. More importantly, existing algorithms often lack the ability to extract and model deep features of complex coupling constraints within the system, leading to situations where the generated strategies may exhibit low coordination efficiency, high default risk, and suboptimal economic returns in actual implementation.

[0007] In summary, developing a novel decision-making technology that can adapt to highly uncertain environments, efficiently process high-dimensional continuous decision spaces, and deeply understand the complex constraints within the system has become an urgent need to enhance the competitiveness and profitability of virtual power plants in the electricity spot market, and is also key to promoting the mature development of commercial operation of virtual power plants. Summary of the Invention

[0008] The purpose of this invention is to overcome the problems of high computational complexity, strong dependence on prediction, and poor adaptability in the existing technology for handling multi-dimensional resource coordination in virtual power plants and coping with drastic fluctuations in the spot market. The invention provides a method and system for intelligent scheduling of virtual power plant sources, storage, and load based on deep reinforcement learning.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for intelligent scheduling of virtual power plant source-storage-load based on deep reinforcement learning, comprising the following steps: Collect and fuse external environmental data and internal physical state data of the energy storage system to form a time-aligned multi-dimensional state vector; The multidimensional state vector is input into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and the deep reinforcement learning model outputs a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions. Based on the real-time physical state data of the energy storage system, and combined with the theoretical charging and discharging power mapped by the frequency regulation market declaration parameters and the spot market charging and discharging instructions, boundary verification and safety reshaping are performed to obtain the boundary verification results and safety execution instructions. Based on the security execution command and the frequency modulation market declaration parameters, the coordinated settlement of the frequency modulation market and the spot market is executed in parallel to obtain the market revenue result; Based on the market return results and the boundary verification results, a composite reward signal is generated. The deep reinforcement learning model is iteratively updated using the composite reward signal, and the optimal bid-ask strategy is obtained through the iteratively updated model.

[0010] In the step of collecting and fusing external environmental data and internal physical state data of the energy storage system to form a time-aligned multidimensional state vector, the external environmental data includes the predicted power of distributed photovoltaic power, the predicted power of heterogeneous loads, the spot market clearing price, AGC frequency regulation commands, the set value of frequency regulation performance indicators, and the time-series periodic characteristics of the corresponding time step of the system; the internal physical state data includes the collected state of charge of the energy storage system, and the evolution space of the state of charge and the safety boundary of subsequent physical actions are defined by using the rated power and rated capacity of the energy storage system as the underlying physical constraint benchmark parameters.

[0011] In the step of inputting the multidimensional state vector into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and outputting a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions through the deep reinforcement learning model, the deep reinforcement model adopts a policy network-value network architecture. The policy network outputs the probability distribution of the multidimensional continuous action space based on the multidimensional state vector, including the frequency modulation declaration capacity ratio, the frequency modulation declaration price coefficient, and the battery charging and discharging ratio under the spot dual-track system. The value network is used to evaluate the expected state value corresponding to the multidimensional state vector to guide the gradient update direction of the policy network.

[0012] The training method for the deep reinforcement learning model is as follows: Construct and randomly initialize the parameters of the policy network and the value network; The latest parameters of the policy network are used to interact with the virtual power plant operating environment to collect experience data at multiple time steps. The experience data includes state vectors, action vectors, immediate rewards, and the state vector of the next time step; wherein, the immediate reward is a composite reward signal. Based on the collected empirical data, the state value is calculated through a value network, and the advantage function value of the action at each time step is calculated using a generalized advantage estimation algorithm. Based on the same batch of empirical data, the parameters of the policy network and the value network are updated in parallel: based on the advantage function value, the parameters of the policy network are updated by maximizing the pruning substitution objective function of the near-end policy optimization algorithm, and the parameters of the value network are updated by minimizing the error between the state value estimate output by the value network and the target value; the pruning substitution objective function is used to limit the range of change in the probability ratio between the new and old policies. Based on the updated policy network and value network parameters, the experience collection and parameter update process is repeated. When the relative change in the average cumulative reward obtained by the deep reinforcement learning model in multiple consecutive iterations is less than the preset convergence threshold, the policy performance is determined to be converged, and the pre-trained deep reinforcement learning model is obtained. The latest parameters include the initial policy network parameters and the updated policy network parameters.

[0013] The method for performing boundary verification and safety reshaping on the theoretical charge / discharge power based on the real-time physical state data of the energy storage system, combined with the frequency regulation market declaration parameters and the spot market charge / discharge commands, to obtain the boundary verification results and safety execution commands is as follows: The maximum safety boundary of the energy storage system is calculated based on the real-time state of charge, rated power and rated capacity of the energy storage system. The maximum safety boundary includes the physically permissible maximum charging power boundary and the maximum discharging power boundary. The theoretical charging and discharging power output by the deep reinforcement learning model is compared with the maximum safety boundary. If the theoretical charging and discharging power exceeds the maximum safety boundary, a nonlinear truncation function is invoked to impose a forced boundary constraint on the theoretical charging and discharging power that exceeds the maximum safety boundary, thereby constraining the theoretical charging and discharging power to within the safe operating range and generating a safe execution command.

[0014] Based on the security execution command and the frequency modulation market declaration parameters, the method for parallel settlement of the frequency modulation market and the spot market to obtain market revenue results is as follows: On the frequency regulation market side, the bidding ranking price is determined based on the frequency regulation bid price and the set value of the frequency regulation performance index, and the frequency regulation service revenue is obtained based on the frequency regulation market clearing price, effective frequency regulation mileage, and the set value of the frequency regulation performance index. On the spot market side, the net load is calculated based on the predicted power of heterogeneous loads, the predicted power of photovoltaics, and the aforementioned safety execution instructions. The net load is settled using a dual-track settlement rule: if the net load is positive, the purchased electricity is settled at the real-time spot market clearing price; if the net load is negative, the priority dual-track settlement logic is triggered, and the back-feeding electricity is preferentially matched to the guaranteed purchase quota based on the actual output of photovoltaics and settled at the mechanism price. The overflowing electricity exceeding the guaranteed purchase quota is settled at the real-time fluctuating spot market clearing price.

[0015] In the step of generating a composite reward signal based on the market return result and the boundary check result, iteratively updating the deep reinforcement learning model using the composite reward signal, and obtaining the optimal bid-ask-volume strategy through the iteratively updated model, the composite reward signal is derived from the market return result. and boundary check results Weighted superposition:

[0016] Where η is the weighting coefficient of the boundary verification result.

[0017] Secondly, the present invention provides a virtual power plant source-storage-load intelligent scheduling system based on deep reinforcement learning, comprising: The data acquisition module is used to collect and fuse external environmental data and internal physical state data of the energy storage system to form a time-aligned multi-dimensional state vector. The scheduling decision module is used to input the multi-dimensional state vector into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and output a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions through the deep reinforcement learning model. The safety constraint module is used to perform boundary verification and safety reshaping on the theoretical charge and discharge power based on the real-time physical state data of the energy storage system, which is a combination of the frequency regulation market declaration parameters and the spot market charge and discharge instructions, to obtain the boundary verification results and safety execution instructions. The market settlement module is used to perform parallel settlement of the frequency modulation market and the spot market according to the security execution command and the frequency modulation market declaration parameters, so as to obtain the market revenue result; The model update module is used to generate a composite reward signal based on the market return results and the boundary verification results, and to iteratively update the deep reinforcement learning model using the composite reward signal, thereby obtaining the optimal bid-ask strategy through the iteratively updated model.

[0018] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a virtual power plant source-storage-load intelligent scheduling method based on deep reinforcement learning.

[0019] Fourthly, the present invention provides a storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of a virtual power plant source-storage-load intelligent scheduling method based on deep reinforcement learning.

[0020] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a method and system for intelligent scheduling of virtual power plant energy sources, storage, and load based on deep reinforcement learning. By automatically generating and optimizing market strategies through a deep reinforcement learning model, it can learn optimal decision-making patterns from high-dimensional, complex, and dynamically changing market data, significantly surpassing traditional rule-based or simple model-based strategies. It couples the decisions of the frequency regulation market and the spot market within a unified model framework, achieving true multi-market collaborative optimization and avoiding potential revenue conflicts or opportunity losses that may result from independent decisions in different markets, thereby maximizing overall returns. After the model outputs an action, boundary checks and safety reshaping are immediately performed based on the real-time physical state of the energy storage system, ensuring that all instructions are executed within the system's safe operating range. This achieves a seamless integration of economic pursuit and hard physical safety constraints, fundamentally preventing equipment damage or operational risks caused by excessive pursuit of returns. A composite reward signal (integrating market returns and boundary check results) is used to iteratively update the model online or periodically, enabling the model to continuously track changes in the electricity market structure, price patterns, and system aging state, achieving dynamic adaptation and self-improvement of the strategy.

[0021] Furthermore, the time-aligned multidimensional state vectors provide the model with accurate decision-making basis, improving the timeliness and accuracy of decision-making. This solution creatively combines the intelligent decision-making advantages of deep reinforcement learning with the physical security rigid constraints of energy storage systems. It can not only maximize synergistic benefits in the current complex power market environment, but also provide a solid technical foundation for coping with future uncertainties through continuous learning mechanisms. It has significant economic value, safety value and long-term strategic value. Attached Figure Description

[0022] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a timing diagram of the multi-dimensional input state of the virtual power plant in Embodiment 2 of the present invention; Figure 3 This is a graph showing the convergence of PPO agent training and the number of BMS security checks in Embodiment 2 of the present invention. Figure 4This is a diagram showing the intraday source-storage-load coordinated scheduling and SOC status trajectory in Embodiment 2 of the present invention; Figure 5 This is a breakdown diagram of spot dual-track settlement electricity volume and arbitrage profits in Embodiment 2 of the present invention; Figure 6 This is a comparison chart of cumulative revenue in Embodiment 2 of the present invention; Figure 7 This is a comparison chart of bidding capacity strategies at different time periods in Embodiment 2 of the present invention; Figure 8 This is a system diagram of Embodiment 3 of the present invention. Detailed Implementation

[0023] To further understand the content of this invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.

[0024] Example 1 like Figure 1 As shown, a method for intelligent scheduling of virtual power plant sources, storage, and load based on deep reinforcement learning includes the following steps: S1: Collect and fuse external environmental data and internal physical state data of the energy storage system to form a time-aligned multi-dimensional state vector; S2: Input the multidimensional state vector into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and output a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions through the deep reinforcement learning model; S3: Based on the real-time physical state data of the energy storage system, perform boundary verification and safety reshaping on the theoretical charging and discharging power that is collaboratively mapped by the frequency regulation market declaration parameters and the spot market charging and discharging instructions, and obtain the boundary verification results and safety execution instructions; S4: Based on the security execution command and the frequency modulation market declaration parameters, perform parallel settlement of the frequency modulation market and the spot market to obtain market revenue results; S5: Based on the market return results and the boundary verification results, generate a composite reward signal, and use the composite reward signal to iteratively update the deep reinforcement learning model. The optimal bid-ask strategy is obtained through the iteratively updated model.

[0025] Specifically, in S1, external environmental data and internal physical state data of the energy storage system are collected and fused to form a time-aligned multidimensional state vector.

[0026] Collect historical and real-time electricity spot market prices, renewable energy output, heterogeneous load demand, and system time-series cyclical characteristics data. A unified data interface and normalization mechanism need to be established to ensure accurate alignment of external environmental data and internal physical state data in the time dimension, thereby providing reliable high-dimensional state input for subsequent Markov decision-making processes.

[0027] The multi-dimensional data perception stage forms the basis for the decision-making of virtual power plants participating in the complex spot and frequency regulation markets. This stage primarily collects two types of core data: external environmental data and internal physical state data. External environmental data includes predicted power from distributed photovoltaic systems, predicted power from heterogeneous loads (e.g., building foundation loads, electric vehicle charging pile loads), spot market clearing prices, AGC frequency regulation commands, frequency regulation performance index setpoints, and normalized time-series characteristics. Internal physical state data includes the real-time state of charge (SOC) of the energy storage system. The rated power and rated capacity of the energy storage system are preset as static parameters in the underlying physical constraint model to define the safety boundaries of decision-making actions. By performing time-series alignment and fusion of multi-source heterogeneous data, the system establishes a state space vector including multi-dimensional features. This high-quality state data plays three key roles in subsequent processes. First, it provides decision-making basis for the Actor network of the intelligent agent, enabling it to capture the fluctuation patterns of market electricity prices and physical operating boundaries. Second, it provides a safety constraint benchmark for the BMS physical shield, because the real-time SOC of the energy storage directly determines the upper and lower limits of the adjustable output. Third, it provides a settlement prerequisite for dual-market clearing. The combination of photovoltaic output, load demand and spot electricity price constitutes the calculation factors for internal net load and arbitrage profit assessment, thereby supporting the accurate operation of the entire collaborative dispatch model.

[0028] Specifically, in S2, the multidimensional state vector is input into a deep reinforcement learning model pre-trained based on the Proximal Policy Optimization (PPO) algorithm, and the deep reinforcement learning model outputs a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions.

[0029] Furthermore, the deep reinforcement learning model adopts an Actor-Critic architecture. The Actor policy network outputs the probability distribution of a multi-dimensional continuous action space based on the multi-dimensional state vector. This action space is precisely divided into three continuous dimensions: the frequency regulation declaration capacity ratio, the frequency regulation declaration price coefficient, and the battery charge-discharge ratio under the spot dual-track system. At the same time, the running Critic value network evaluates the long-term expected value of the current state vector.

[0030] The method for pre-training a deep reinforcement learning model using a proximal policy optimization algorithm is as follows: Construct and randomly initialize the parameters of the policy network and the value network; The latest parameters of the policy network are used to interact with the virtual power plant operating environment to collect experience data at multiple time steps. The experience data includes state vectors, action vectors, immediate rewards, and the state vector of the next time step. The immediate rewards are composite reward signals. The latest parameters include the initialized policy network parameters and the updated policy network parameters.

[0031] Based on the collected empirical data, the state value is calculated through the value network, and the generalized advantage estimation algorithm is used to calculate the advantage function value of the action at each time step.

[0032]

[0033] in, It is time difference error; It is a discount factor; It is the GAE discount factor; It is the value network's estimate of the value of state S.

[0034] Based on the same batch of empirical data, the parameters of the policy network and the value network are updated in parallel: based on the advantage function value, the parameters of the policy network are updated by maximizing the pruning substitution objective function of the near-end policy optimization algorithm, wherein the pruning substitution objective function is used to limit the range of change of the probability ratio between the new and old policies; a whole batch of data from multiple time steps is updated together based on the same batch of empirical data.

[0035]

[0036]

[0037] in, These are the parameters of the policy network; It is the probability ratio of the old and new strategies; It is the dominant function value; It is the clipping hyperparameter.

[0038] The parameters of the value network are updated by minimizing the error between the state value estimate output by the value network and the target value.

[0039] in, These are value network parameters. It is a value network for state The predicted value; It is the target value, expressed through the formula. calculate.

[0040] Based on the updated policy network and value network parameters, the experience collection and parameter update process is repeated. When the relative change in the average cumulative reward obtained by the deep reinforcement learning model in multiple consecutive iterations is less than a preset convergence threshold, the policy performance is determined to have converged, and the pre-trained deep reinforcement learning model is obtained.

[0041] This process effectively solves the problem that traditional mathematical programming and heuristic algorithms are prone to the curse of computational dimensionality when faced with multi-source heterogeneous equipment and high-dimensional continuous control variables, giving virtual power plants the ability to make rapid adaptive decisions in response to nonlinear spot price characteristics.

[0042] Specifically, in S3, based on the real-time physical state data of the energy storage system, the theoretical charging and discharging power, which is mapped by combining the frequency regulation market declaration parameters and the spot market charging and discharging instructions, is subjected to boundary verification and safety reshaping to obtain the boundary verification results and safety execution instructions.

[0043] Based on the real-time state of charge (SOC), rated power, and rated capacity of the energy storage system, calculate the physically permissible maximum safety boundary, namely the physically permissible maximum charging power boundary and maximum discharging power boundary. In a preferred embodiment, the maximum discharge power boundary ensures that the SOC after discharge is not lower than a set lower limit of 5%. Dischargeable capacity = (Current SOC - Lower SOC limit) × Rated capacity Maximum discharge power = min(rated power, discharge capacity / time resolution) The maximum charging power boundary ensures that the SOC after charging does not exceed 95% of the set upper limit. Rechargeable capacity = (Maximum SOC - Current SOC) × Rated capacity Maximum charging power = min(rated power, rechargeable capacity / time resolution) The theoretical charge / discharge power output by the deep reinforcement learning model is compared with the safety boundary; the theoretical charge / discharge power is the battery charge / discharge ratio under the spot dual-track system output by the policy network. If the theoretical charging and discharging power exceeds the maximum safety boundary, a nonlinear truncation function is invoked to impose a forced boundary constraint on the theoretical charging and discharging power that exceeds the maximum safety boundary, thereby constraining the theoretical charging and discharging power to within the safe operating range and generating a safe execution command.

[0044] Preferably, the nonlinear truncation function is a piecewise saturation function, as follows:

[0045] in, Theoretical charge / discharge power, and These are the maximum charging power threshold and the maximum discharging power threshold, respectively.

[0046] This core fault-tolerance mechanism ensures that the execution commands ultimately sent to the underlying physical inverters and controllers are firmly locked within a safe operating range, completely eliminating the fatal risk of hardware overcharging and discharging or even equipment damage caused by the inherent randomness of the reinforcement learning exploration mechanism itself.

[0047] Specifically, in S4, based on the security execution command and the frequency modulation market declaration parameters, the coordinated settlement of the frequency modulation market and the spot market is performed in parallel to obtain the market revenue result.

[0048] 1) On the frequency regulation market side, the bidding ranking price is determined based on the frequency regulation bid price and the set value of the frequency regulation performance index, and the frequency regulation service revenue is obtained based on the frequency regulation market clearing price, the set value of the frequency regulation performance index, and the effective frequency regulation mileage. Wherein, the ranking price = the declared price ÷ the performance index The calculated bidding ranking price is compared with the real-time clearing price of the FM market. If the ranking price is not higher than (or meets other market rules) the clearing price, the FM capacity for that period is determined to be awarded.

[0049] If the bid is successful, the system enters the revenue calculation phase. The revenue from frequency regulation mileage is calculated based on the actual effective regulation mileage (i.e., power change), the set value of frequency regulation performance indicators, and the unit price per mileage when responding to grid AGC commands if the actual regulation exceeds the specified control dead zone.

[0050] 2) On the spot market side, the net load is calculated based on the predicted power of heterogeneous load, the predicted power of photovoltaics, and the aforementioned safety execution instructions, and the net load is settled using the dual-track settlement rules: if the net load is positive, the purchased electricity is settled at the real-time spot market clearing price; if the net load is negative, the priority dual-track settlement logic is triggered, and the back-feeding electricity is preferentially matched to the guaranteed purchase quota based on the actual output of photovoltaics and settled at the mechanism price. The overflowing electricity exceeding the guaranteed purchase quota is settled at the real-time fluctuating spot market clearing price.

[0051] Net load = Heterogeneous load forecast power - Photovoltaic forecast power - Energy storage system safe operating power; if the result is positive, it means that the virtual power plant is in a state of purchasing electricity and needs to buy electricity from the spot market. If the result is negative, it means that the plant is in a state of surplus electricity fed into the grid and can sell electricity to the market.

[0052] When the net load is positive, all purchased electricity is settled in full based on the real-time spot market clearing price at the node, plus additional charges such as transmission and distribution prices and government funds. When the net load is negative (i.e., there is surplus electricity fed into the grid), the fed-in electricity is settled in two parts according to a pre-set or market-determined ratio: A portion of the electricity (such as a certain percentage or a fixed amount) is settled according to the fixed or stable electricity price in a pre-signed medium- or long-term contract to guarantee basic revenue. The remaining electricity participates fully in the spot market and is settled according to the spot market clearing price that fluctuates in real time.

[0053] Specifically, in S5, a composite reward signal is generated based on the market return result and the boundary verification result. The deep reinforcement learning model is iteratively updated using the composite reward signal, and the optimal bid-ask strategy is obtained through the iteratively updated model.

[0054] The compound reward signal is derived from market return results. and boundary check results Weighted superposition:

[0055] Where η is the weighting coefficient of the boundary verification result; Boundary check results include: The hard constraint penalty term is triggered when the safety constraint unit truncates the theoretical charge / discharge command. The soft constraint penalty term is triggered when the real-time state of charge (SOC) of the energy storage system approaches the preset warning boundary, and its penalty value increases non-linearly with the degree of deviation.

[0056] Example 2 Following the process in Example 1, a scenario is established where a virtual power plant participates in the integrated scheduling and bid / quota decision verification of both the spot and frequency regulation markets. During the data acquisition and environment setup phase, a time-series data set is generated based on historical meteorological and market operation data. This set includes photovoltaic forecasts, building load, electric vehicle charging pile load, spot market clearing prices, and frequency regulation ancillary service instructions (AGC instructions and Kp performance indicators), with a sampling period of 1 hour. The physical parameters of the aggregated resources in the virtual power plant are set as follows: a 5MW rated power, 20MWh capacity energy storage system with a safe operating range of SOC∈[0.05, 0.95] and a distributed photovoltaic system with a maximum theoretical output of 10MW. Figure 2 It presents the intraday fluctuation of some core multi-source heterogeneous data in the input state space over a period of 24 hours.

[0057] Figure 2The blue curve represents the spot market electricity price, exhibiting a typical duck curve pattern (the price drops to a low of 0.10 yuan / kWh between 12 PM and 3 PM due to massive grid-wide photovoltaic power generation, and then climbs to a peak of over 1.05 yuan / kWh between 6 PM and 10 PM when there is no sunlight). The orange curve represents the photovoltaic output within the microgrid, and the green histogram overlays the typical characteristics of industrial and commercial office parks (i.e., a high load base from 8 AM to 6 PM, combined with peak charging times in the morning and afternoon). This chart demonstrates the complex situation faced by virtual power plants—a volatile external market and internal source-load mismatch—and highlights the necessity of using deep reinforcement learning for high-dimensional state adaptive decision-making.

[0058] During the training of the agent (deep reinforcement learning model), by using a multi-objective composite reward function containing economic incentives and physical boundary constraints, the agent performs iterative update operations of deep reinforcement learning on the policy network and value network in the Actor-Critic dual network architecture. This successfully obtains the timing of battery charging and discharging and the long-term time dependency of spot electricity price fluctuations, demonstrating the dynamic game characteristics of the agent in the process of pursuing maximum market returns and adhering to the bottom line of BMS safety shield.

[0059] Figure 3 This diagram illustrates the convergence curve and the number of times the limit-crossing penalty was triggered for the PPO decision agent during the training process. The horizontal axis represents the number of training steps, the left vertical axis represents the agent's expected total reward, and the right vertical axis represents the number of times the "BMS physical shield interception" and "soft barrier penalty" were triggered. It can be observed that in the initial training phase, from 0 to 200,000 steps, the agent is in a period of random exploration, frequently issuing charge / discharge commands exceeding the battery's physical limits. These commands are forcibly intercepted by the BMS shield and subject to a single violation penalty of -5.0, causing the total reward to oscillate at a low level. As the gradient of the time difference error is backpropagated, the agent gradually learns the physical red line boundary, namely the quadratic penalty warning zone of SOC 0.15 and 0.85, and the red limit-crossing histogram rapidly drops to zero. After 600,000 steps, the total reward curve steadily rises and eventually converges, demonstrating that the unique "BMS shield + barrier penalty" mechanism of this invention can effectively guide artificial intelligence to find the optimal solution within an absolutely safe physical feasible domain.

[0060] Figure 4This reflects the optimal strategy and corresponding physical response trajectory of the "source-storage-load" multi-dimensional coordinated scheduling executed by the agent during a typical operating day after training and convergence using the PPO algorithm. This curve clearly shows that the agent dynamically issues charging and discharging power commands based on spot market price signals and the microgrid's net load status, mapping this control sequence to a smooth evolution of the battery's state of charge (SOC) within the safety constraint boundary. The figure illustrates the agent's actual physical actions driven by spot market price signals. During the midday period when photovoltaic power generation is high and spot electricity prices are extremely low (11:00-15:00), the agent accurately issues the maximum power charging command. The red negative bars in the figure show that the battery fully absorbs the cheap electricity overflowing from the internal photovoltaic system, effectively avoiding the risks associated with selling at low prices in the spot market. Simultaneously, the SOC curve, represented by the black broken line, rises steadily. However, when the SOC curve approaches 0.85, the agent actively slows down the charging rate, successfully preventing overcharging penalties. During the evening peak electricity price period, from 18:00 to 21:00, the intelligent agent controls the battery to discharge at full power, as can be seen from the blue positive bar chart. This supports the system's net load turning into a negative value, which means feeding back into the grid.

[0061] Figure 5 This paper provides a detailed analysis of the profit composition and distribution characteristics of the "spot dual-track arbitrage" scheme during the external grid market settlement stage. The figure clearly illustrates how the intelligent agent precisely divides the back-feeding electricity from virtual power plants into two aspects: "medium-to-long-term guaranteed quantity and price (guaranteed purchase)" and "spot market high-price settlement," thereby achieving in-depth exploitation of policy dividends and market fluctuations. The stacked bar chart in the figure presents the financial recognition of the back-feeding electricity from virtual power plants. According to the spot dual-track settlement rules in this invention, during periods when photovoltaic power is output, the grid-connected electricity is strictly limited to within 80% of the photovoltaic output, i.e., the portion represented by the blue bars, and enjoys a fixed mechanism price guarantee of 0.4161 yuan / kWh. During the evening peak hours, i.e., when there is no photovoltaic output, the back-feeding electricity generated by the discharge of energy storage batteries completely exceeds the medium-to-long-term guarantee framework and is 100% recognized as pure spot trading electricity, i.e., the portion shown by the red bars, and will be fully settled at the then-current spot peak price exceeding 1.0 yuan / kWh. This chart clearly and intuitively demonstrates that the PPO strategy of this invention successfully utilizes the time-shifting characteristic to transfer photovoltaic power, which was originally of lower value, to the spot high-price track, thus achieving excess arbitrage under the dual-track system.

[0062] Figure 6 This paper demonstrates a quantitative comparison of the economic value of the PPO collaborative and physical shield optimization strategy of this invention with a conventional benchmark strategy of uncoupled dual-track deep optimization during long-term operation. Figure 6The steep slope of the cumulative return clearly shows that the profitability of the optimized strategy is significantly higher than that of the benchmark strategy. To further explore the underlying drivers of these excess returns and extreme safety, Figure 7 A comparison was made of the bidding capacity arrangements of the two strategies at different time periods. Unlike the benchmark strategy, which uses a rigid, fixed capacity bidding approach throughout the day, the strategy of this invention has the ability to sense dynamic fluctuations. By utilizing a dynamic capacity allocation mechanism that actively bids during periods of high volatility in spot prices and net load, and proactively contracts during periods of low volatility, the system effectively avoids response defaults caused by insufficient equipment capacity while maximizing the capture of market opportunities.

[0063] Through the Figure 6 Observing the cumulative return growth slope reveals that the optimized strategy demonstrates stronger profitability than the benchmark strategy. Specific return and safety metrics are presented in Table 1.

[0064] Table 1 Data Comparison Table

[0065] The final experimental results show that, compared with the baseline strategy, the optimized strategy increases the cumulative net profit over the entire cycle by approximately 52.25%, while successfully reducing the number of BMS physical overruns and market response defaults to zero, achieving a leap forward in both economic efficiency and security. The six visualization charts above (multi-dimensional state input, training safety convergence, source-storage-load coordinated scheduling, dual-track settlement mechanism, cumulative revenue comparison, and time-of-use bidding strategy) map from low-level features to high-level decision execution, fully demonstrating the accuracy of the modeling and the robustness of the strategy. This perfectly verifies that the technical process can accurately capture the spatiotemporal coupling characteristics of spot electricity prices and source-load fluctuations through deep reinforcement learning, and, with the joint protection of the BMS physical shield and soft boundary mechanism, provides a reliable scientific basis for virtual power plants to formulate risk-sensitive and profit-multiplying coordinated scheduling schemes and time-of-use bidding strategies in complex dual-track settlement environments.

[0066] Example 3 like Figure 8 As shown, the present invention also provides an electronic device 100 for a virtual power plant source-storage-load intelligent scheduling method based on deep reinforcement learning; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0067] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the virtual power plant source-storage-load intelligent scheduling method based on deep reinforcement learning described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0068] The at least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor. The processor 102 is the control center of the electronic device 100, connecting various parts of the electronic device 100 via various interfaces and lines.

[0069] The memory 101 in the electronic device 100 stores multiple instructions to implement a virtual power plant source-storage-load intelligent scheduling method based on deep reinforcement learning. The processor 102 can execute the multiple instructions to achieve the following: Collect and fuse external environmental data and internal physical state data of the energy storage system to form a time-aligned multi-dimensional state vector; The multidimensional state vector is input into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and the deep reinforcement learning model outputs a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions. Based on the real-time physical state data of the energy storage system, the theoretical charging and discharging power based on the frequency regulation market declaration parameters and the spot market charging and discharging instructions is subjected to boundary verification and safety reshaping to obtain the boundary verification results and safety execution instructions. Based on the security execution command and the frequency modulation market declaration parameters, the coordinated settlement of the frequency modulation market and the spot market is executed in parallel to obtain the market revenue result; Based on the market return results and the boundary verification results, a composite reward signal is generated. The deep reinforcement learning model is iteratively updated using the composite reward signal, and the optimal bid-ask strategy is obtained through the iteratively updated model.

[0070] Example 4 If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).

[0071] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0072] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0073] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0074] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A virtual power plant source storage load intelligent scheduling method based on deep reinforcement learning, characterized in that, Includes the following steps: Collect and fuse external environmental data and internal physical state data of the energy storage system to form a time-aligned multi-dimensional state vector; The multidimensional state vector is input into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and the deep reinforcement learning model outputs a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions. Based on the real-time physical state data of the energy storage system, and combined with the theoretical charging and discharging power mapped by the frequency regulation market declaration parameters and the spot market charging and discharging instructions, boundary verification and safety reshaping are performed to obtain the boundary verification results and safety execution instructions. Based on the security execution command and the frequency modulation market declaration parameters, the coordinated settlement of the frequency modulation market and the spot market is executed in parallel to obtain the market revenue result; Based on the market return results and the boundary verification results, a composite reward signal is generated. The deep reinforcement learning model is iteratively updated using the composite reward signal, and the optimal bid-ask strategy is obtained through the iteratively updated model.

2. The method for intelligent scheduling of virtual power plant source-storage-load based on deep reinforcement learning according to claim 1, characterized in that, In the step of collecting and fusing external environmental data and internal physical state data of the energy storage system to form a time-aligned multidimensional state vector, the external environmental data includes distributed photovoltaic predicted power, heterogeneous load predicted power, spot market clearing price, AGC frequency regulation command, frequency regulation performance index set value, and the time-series periodic characteristics of the corresponding time step of the system. The internal physical state data includes the state of charge of the energy storage system. The rated power and rated capacity of the energy storage system are used as the underlying physical constraint benchmark parameters to define the evolution space of the state of charge and the safety boundary of subsequent physical actions.

3. The virtual power plant source storage load intelligent scheduling method based on deep reinforcement learning according to claim 2, characterized in that, In the step of inputting the multidimensional state vector into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and outputting a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions through the deep reinforcement learning model, the deep reinforcement model adopts a policy network-value network architecture. The policy network outputs the probability distribution of the multidimensional continuous action space based on the multidimensional state vector, including the frequency modulation declaration capacity ratio, the frequency modulation declaration price coefficient, and the battery charging and discharging ratio under the spot dual-track system. The value network is used to evaluate the expected state value corresponding to the multidimensional state vector to guide the gradient update direction of the policy network.

4. The virtual power plant source storage load intelligent scheduling method based on deep reinforcement learning according to claim 3, characterized in that, The training method for the deep reinforcement learning model is as follows: Construct and randomly initialize the parameters of the policy network and the value network; The latest parameters of the policy network are used to interact with the virtual power plant operating environment to collect experience data at multiple time steps. The experience data includes state vectors, action vectors, immediate rewards, and the state vector of the next time step; wherein, the immediate reward is a composite reward signal. Based on the collected empirical data, the state value is calculated through a value network, and the advantage function value of the action at each time step is calculated using a generalized advantage estimation algorithm. Based on the same batch of empirical data, the parameters of the policy network and the value network are updated in parallel: based on the advantage function value, the parameters of the policy network are updated by maximizing the pruning substitution objective function of the near-end policy optimization algorithm, and the parameters of the value network are updated by minimizing the error between the state value estimate output by the value network and the target value; the pruning substitution objective function is used to limit the range of change in the probability ratio between the new and old policies. Based on the updated policy network and value network parameters, the experience collection and parameter update process is repeated. When the relative change in the average cumulative reward obtained by the deep reinforcement learning model in multiple consecutive iterations is less than the preset convergence threshold, the policy performance is determined to be converged, and the pre-trained deep reinforcement learning model is obtained. The latest parameters include the initial policy network parameters and the updated policy network parameters.

5. The virtual power plant source storage load intelligent scheduling method based on deep reinforcement learning according to claim 4, characterized in that, The method for performing boundary verification and safety reshaping based on the real-time physical state data of the energy storage system, combined with the theoretical charge and discharge power mapped by the frequency regulation market declaration parameters and the spot market charge and discharge commands, to obtain the boundary verification results and safety execution commands is as follows: The maximum safety boundary of the energy storage system is calculated based on the real-time state of charge, rated power and rated capacity of the energy storage system. The maximum safety boundary includes the physically permissible maximum charging power boundary and the maximum discharging power boundary. The theoretical charging and discharging power output by the deep reinforcement learning model is compared with the maximum safety boundary. If the theoretical charging and discharging power exceeds the maximum safety boundary, a nonlinear truncation function is invoked to impose a forced boundary constraint on the theoretical charging and discharging power that exceeds the maximum safety boundary, thereby constraining the theoretical charging and discharging power to within the safe operating range and generating a safe execution command.

6. The virtual power plant source storage load intelligent scheduling method based on deep reinforcement learning according to claim 5, characterized in that, Based on the security execution command and the frequency modulation market declaration parameters, the method for parallel settlement of the frequency modulation market and the spot market to obtain market revenue results is as follows: On the frequency regulation market side, the bidding ranking price is determined based on the frequency regulation bid price and the set value of the frequency regulation performance index, and the frequency regulation service revenue is obtained based on the frequency regulation market clearing price, effective regulation mileage, and the set value of the frequency regulation performance index. On the spot market side, the net load is calculated based on the predicted power of heterogeneous loads, the predicted power of photovoltaics, and the aforementioned safety execution instructions. The net load is settled using a dual-track settlement rule: if the net load is positive, the purchased electricity is settled at the real-time spot market clearing price; if the net load is negative, the priority dual-track settlement logic is triggered, and the back-feeding electricity is preferentially matched to the guaranteed purchase quota based on the actual output of photovoltaics and settled at the mechanism price. The overflowing electricity exceeding the guaranteed purchase quota is settled at the real-time fluctuating spot market clearing price.

7. The virtual power plant source storage load intelligent scheduling method based on deep reinforcement learning according to claim 6, characterized in that, In the step of generating a composite reward signal based on the market return result and the boundary check result, iteratively updating the deep reinforcement learning model using the composite reward signal, and obtaining the optimal bid-ask-volume strategy through the iteratively updated model, the composite reward signal is derived from the market return result. and boundary check results Weighted superposition: Where η is the weighting coefficient of the boundary verification result.

8. A virtual power plant source storage load intelligent scheduling system based on deep reinforcement learning, characterized in that, include: The data acquisition module is used to collect and fuse external environmental data and internal physical state data of the energy storage system to form a time-aligned multi-dimensional state vector. The scheduling decision module is used to input the multi-dimensional state vector into a deep reinforcement learning model pre-trained based on a near-end policy optimization algorithm, and output a continuous action vector containing frequency modulation market declaration parameters and spot market charging and discharging instructions through the deep reinforcement learning model. The safety constraint module is used to perform boundary verification and safety reshaping based on the real-time physical state data of the energy storage system, combined with the theoretical charge and discharge power mapped by the frequency regulation market declaration parameters and the spot market charge and discharge instructions, to obtain the boundary verification results and safety execution instructions; The market settlement module is used to perform parallel settlement of the frequency modulation market and the spot market according to the security execution command and the frequency modulation market declaration parameters, so as to obtain the market revenue result; The model update module is used to generate a composite reward signal based on the market return results and the boundary verification results, and to iteratively update the deep reinforcement learning model using the composite reward signal, thereby obtaining the optimal bid-ask strategy through the iteratively updated model. 9.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, When the processor executes the computer program, it implements the steps of the intelligent scheduling method for virtual power plant source-storage-load based on deep reinforcement learning as described in any one of claims 1 to 7.

10. A storage medium having stored thereon a computer program, characterized in that When the computer program is executed by the processor, it implements the steps of the intelligent scheduling method for virtual power plant source-storage-load based on deep reinforcement learning as described in any one of claims 1 to 7.