Large-scale wind-solar-thermal storage base scheduling decision-making and optimal configuration method and system based on hierarchical reinforcement learning

By employing a hierarchical reinforcement learning approach, the high-dimensional state space and multi-timescale problems of large-scale wind, solar, thermal, and energy storage bases were addressed, enabling coordinated optimization of investment planning and operation scheduling, and improving the frequency stability and economy of the system.

CN121920591APending Publication Date: 2026-04-24THREE GORGES BAZHOU RUOQIANG ENERGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THREE GORGES BAZHOU RUOQIANG ENERGY CO LTD
Filing Date
2025-12-19
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively handle the high-dimensional continuous state space, multiple time scales, strong constraints, and high uncertainty of large-scale wind, solar, thermal, and energy storage bases, leading to problems such as state space dimension explosion, disconnect between investment planning and operation scheduling, insufficient system inertia, and frequency stability.

Method used

We employ a hierarchical reinforcement learning approach, using a Transformer architecture to predict wind and solar power probabilistically. We construct a multi-timescale state space and design a three-layer intelligent agent architecture, including high-level capacity configuration, mid-level day-ahead scheduling, and low-level real-time control. By combining an Actor-Critic network and multi-Agent learning, and embedding key constraint processing, we achieve online adaptive optimization.

Benefits of technology

It enables efficient scheduling decisions for large-scale wind, solar, thermal, and energy storage systems, reduces the state space dimension, improves the synergistic optimization of investment planning and operation scheduling, meets the requirements of frequency stability and voltage support, and enhances the reliability and economy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920591A_ABST
    Figure CN121920591A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale wind-solar-thermal storage base scheduling decision-making and optimal configuration method and system based on hierarchical reinforcement learning. The method comprises the steps that historical data are collected and preprocessed; a probability prediction model based on Transform is constructed, and wind and light power quantile prediction is output; designing a multi-time scale state space including long-term planning, short-term scheduling and real-time control; a hierarchical deep reinforcement learning framework is established, a high layer adopts a TD3 algorithm to carry out capacity configuration, a middle layer adopts an MADDPG algorithm to carry out day-ahead scheduling, and a low layer adopts a PPO algorithm to carry out real-time control; key constraints such as system inertia and short-circuit ratio are embedded through a projection method and a penalty term; training the network by adopting a priority experience playback and course learning strategy; deploying an online rolling optimization and incremental learning mechanism; according to the method, the high-dimensional state space of a large-scale base can be effectively dealt with, collaborative optimization of investment and operation is realized, and the system supporting capacity under the high proportion of new energy is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fishway technology, and in particular to a method and system for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal and energy storage bases based on hierarchical reinforcement learning. Background Technology

[0002] With the advancement of the "dual carbon" target, my country is planning and constructing large-scale wind and solar power bases in desert, Gobi, and arid regions, with individual bases reaching millions to tens of millions of kilowatts. These large bases exhibit the following characteristics: ultra-large scale, with installed capacity reaching the GW level, involving dozens to hundreds of wind and solar power stations; strong volatility, with strong spatiotemporal correlation and uncertainty in wind and solar power output; weak grid support, being far from load centers, with low grid short-circuit ratios and weak system strength; and multi-entity collaboration, involving multiple types of resources such as wind power, photovoltaics, thermal power, and energy storage.

[0003] The existing technology has the following problems: 1. Dimensional explosion of state space. Traditional methods use discrete state representation. When the base contains dozens of units, the state space grows exponentially, leading to the "curse of dimensionality". The state representation of wind, solar, thermal and storage is too coarse and cannot characterize the continuously changing state of charge of energy storage, wind and solar power output, etc., and it is difficult to extend to large-scale systems.

[0004] 2. Single Time Scale Optimization. Existing methods mostly focus on a single time scale, such as the day-ahead or real-time, lacking coordinated optimization between long-term investment planning and short-term operation scheduling. This leads to a disconnect between capacity allocation and actual operational needs, an inability to cope with the transmission of uncertainties across multiple time scales, and a limitation to the level of operation optimization, without addressing long-term decisions regarding capacity allocation.

[0005] 3. Neglecting system support capabilities. Many studies only consider power balance, ignoring key constraints under the high proportion of renewable energy. Insufficient system inertia leads to frequency stability problems, requiring the system to have sufficient inertia support; weakened short-circuit capacity affects voltage support capabilities; insufficient flexibility of thermal power and energy storage makes it impossible to track wind and solar power fluctuations, resulting in high curtailment rates or system instability.

[0006] 4. Insufficient handling of uncertainties. Traditional deterministic optimization or simple scenario methods cannot fully characterize the spatiotemporal correlation of wind and solar power, as well as the volatility of electricity prices and loads. Point forecasting methods ignore the uncertainty range of forecasts, causing the optimized scheme to deviate from the ideal state in actual operation.

[0007] 5. Limitations of other related work: Mixed Integer Programming (MIP) based methods have high computational complexity. A typical 24-hour unit combinatorial problem contains thousands of decision variables and constraints, and the solution time increases exponentially with the problem size, making it difficult to handle large-scale problems. Heuristic algorithms such as Particle Swarm Optimization (PSO) and Genetic Algorithm (GA) are prone to getting trapped in local optima and lack theoretical convergence guarantees, resulting in weak generalization ability. Traditional tabular methods of reinforcement learning, such as Q-learning and SARSA, face the curse of dimensionality, with the number of states increasing exponentially with the system size. Existing deep reinforcement learning applications are mostly aimed at single optimization objectives or simplified scenarios, lacking systematic research on multi-timescale collaboration, system constraint embedding, and online adaptation.

[0008] Therefore, there is a need for an optimal allocation method for wind, solar, thermal, and energy storage that can handle large-scale, multi-timescale, strongly constrained, and highly uncertain conditions. Summary of the Invention

[0009] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a method and system for scheduling decision-making and optimization of large-scale wind, solar, thermal and energy storage bases based on hierarchical reinforcement learning, so as to cope with the high-dimensional continuous state space of large-scale bases, realize multi-scale coordination of investment planning and operation scheduling, and effectively handle multi-source uncertainty.

[0010] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method and system for scheduling decision-making and optimization allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning, comprising the following steps: Step S1: Data acquisition and preprocessing. Obtain historical meteorological data, power grid parameters, equipment technical and economic parameters, load data and policy constraints of the target base. Perform outlier detection, missing value filling and normalization on the data, and extract time series features. Step S2: Constructing the probabilistic prediction model. A deep learning model based on the Transformer architecture is used to predict the probabilistic power of wind and solar power, outputting the P value for future periods. 10 P 50 P 90 Quantile prediction values ​​are generated by hypercubic sampling based on the predicted quantiles, and the spatiotemporal correlation of multiple stations is modeled. Scene reduction is achieved through K-means clustering. Step S3: Multi-timescale state space design, defining the long-term planning state including wind, solar, thermal and energy storage installed capacity, grid strength index, carbon quota margin and investment budget; the short-term dispatch state including wind and solar load forecast sequence, energy storage state of charge, unit operation state, system inertia and short-circuit ratio; and the real-time control state including actual output, frequency deviation and voltage deviation. Step S4: Construct a hierarchical deep reinforcement learning framework and establish a three-layer agent architecture; wherein, the high-level agent adopts Twi The Delayed DDPG algorithm is used, with the Multi-Agent DDPG algorithm used for the middle-level agents and the Proximal Policy Optimization algorithm used for the lower-level agents. Step S5: Deep neural network architecture design, designing an Actor-Critic network for each layer of the agent; Step S6: Key Constraint Embedding Method; Step S7: Implement the training algorithm; Step S8: Online adaptive and scrolling optimization; Step S9: Multi-site collaborative optimization; Step S10: Performance evaluation and index calculation.

[0011] Preferably, the probabilistic prediction model structure in step S2 is as follows: the input layer receives historical power sequences, meteorological numerical forecasts, satellite cloud images, and time codes, which are then converted into a high-dimensional representation through a temporal embedding layer and input into a Multi-head A. t The attention layer performs self-attention calculations to capture long-term dependencies, performs nonlinear transformations through a feedforward network, and finally outputs P from three independent quantile regression layers. 10 P 50 P 90 Quantile prediction values; the quantile loss function is used during training, as shown in equation (1): L quantile (pred,target,τ)=mean(max(τ×(target-pred),(τ-1)×(target-pred)))(1); Among them, L quantile τ is the quantile loss function; pred is the predicted value; target is the true value; τ is the quantile level; mean is the mean function; max is the maximum value function.

[0012] Preferably, step S4 specifically includes the following process: High-level intelligent agents adopt Tw iThe n-level Delayed DDPG algorithm has a decision-making cycle of monthly or quarterly, an action space of capacity increment, and a reward function that includes investment cost, expected operating cost, carbon cost, reliability reward, and flexibility reward. The middle-level agent uses the Multi-Agent DDPG algorithm, with a decision-making cycle of 24 hours before the day of the day. Each power generation unit is treated as an agent, and the action space includes unit combination, thermal power output setting, energy storage charging and discharging power, and reserve allocation. The reward function includes fuel cost, start-up and shutdown cost, maintenance cost, curtailment cost, reserve cost, and revenue from the energy market and ancillary services market. The lower-level agent uses the Proximal Policy Optimization algorithm, with a decision-making cycle of minutes. The action space includes thermal power and energy storage output adjustments and wind and solar curtailment. The reward function includes the square of the power deviation term, curtailment penalty, frequency deviation penalty, and ramp-up cost.

[0013] Preferably, the capacity configuration action of the high-level intelligent agent includes increasing wind power capacity ΔE. wind,cap Photovoltaic capacity increase ΔE pv,cap Energy storage power capacity increment ΔE storage,cap The reward function for the thermal power plant flexibility transformation level is shown in Equation (2); the middle-level agent adopts the MADDPG framework with centralized training and decentralized execution. The Critic of each power generation unit agent inputs the global state and the actions of all agents, while the Actor only inputs the local observable state; the PPO algorithm of the low-level agent adopts the pruning objective function, as shown in Equation (3). R high =-(C invest +C operate +C carbon +C curtail )+R reliability +R flexibility (2); Among them, R high Rewards for high-level intelligent agents; C invest For investment costs; C operate For expected operating costs; C carbon For carbon cost; C curtail As a penalty for abandoning electricity; R reliability For reliability rewards; R flexibility Rewards for flexibility; L clip =min(r(θ)×Â,clip(r(θ),1-ε,1+ε)×Â)(3); Among them, L clip Let r(θ) be the objective function for PPO pruning; r(θ) be the probability ratio of the new and old strategies;  be the advantage function estimate; ε be the pruning parameter; clip be the pruning function; and min be the minimum function.

[0014] Preferably, step S5 specifically includes the following process: The Actor network consists of an input layer, an encoding layer, a temporal processing layer, an attention layer, a feature fusion layer, and an output layer. The temporal processing layer uses LSTM or Transformer Block to handle time series dependencies, and the attention layer uses Multi-head Self-A. t Attention is used to handle multi-site or multi-unit collaboration; the Critic network contains independent state encoders and action encoders, which are spliced ​​and fused to output the Q value. A dual-Critic network design is adopted to alleviate the problem of Q value overestimation.

[0015] Preferably, the structure of the Actor network is as follows: the state vector s is input into a linear layer (Linear(dims, 256), normalized and activated by a LayerNorm layer, and then input into an LSTM(256, 128) or Transformer Block for time-series processing, and finally processed by a Multi-head Self-Amplifier. t The attention (numheads=8) layer handles the spatial dependencies of multiple stations or multiple units. After passing through the feature fusion layer Linear(128,128)+ReLU+Dropout(0.2), the output layer Linear(128,dima)+Tanh outputs the action normalized to [-1,1]. The Critic network contains a state encoder Linear(dims,256)+ReLU and an action encoder Linear(dima,128)+ReLU. After concatenation, the Q value is output through the deep network Linear(384,256)+ReLU+Dropout(0.1)→Linear(256,128)+ReLU→Linear(128,1). The dual Critic networks Critic1 and Critic2 are used. The target value is calculated as shown in Equation (4). Q target =min(Q 1,target (s',a'),Q 2,target (s',a'))(4); Among them, Q target The target Q value; Q 1,target For the first target Critic network output; Q 2,target The output of the Critic network is for the second objective; s' is the next state; a' is the next action.

[0016] Preferably, step S6 specifically includes the following process: Hard constraints, including power balance constraints, unit output upper and lower limits constraints, energy storage SOC constraints, and ramp rate constraints, are handled through a projection method, projecting actions onto the constraint feasible region in real time. Soft constraints, including system inertia constraints, short-circuit ratio constraints, frequency regulation capability constraints, and energy storage cycle life constraints, are handled by adding a penalty term to the reward function. The system inertia constraint is J. system ≥J min J system The weighted summation is calculated from the inertia contributions of the thermal power unit and energy storage; the short-circuit ratio constraint is SCR=S sc / P renewable ≥SCR min .

[0017] Preferably, the specific implementation process of the projection method is as follows: For power balance constraints, when the total output P total Overload P load When the controllable output is reduced proportionally, as shown in equation (5); for unit output constraints, the thermal power output is limited to the constraint range, as shown in equation (6), where U is the unit switching state; for energy storage SOC constraints, the state of charge at the next moment is calculated, as shown in equation (7), if it exceeds [SOC min SOC max The range is then adjusted by P. storage For the ramp constraint, the rate of change of thermal power output is limited, as shown in Equation (8); the soft constraint penalty term is shown in Equation (9). P total =P load (5); Among them, P total For total output; P load For load power; P thermal =clip(P thermal ,U×P min ,U×P max (6); Among them, P thermal For thermal power output; U represents the unit's on / off status; P min For minimum technical effort; P max To maximize technical output; clip is the clipping function; SOC next =SOC+(P storage ×Δt) / E capacity (7); Among them, SOC next The next state of charge; SOC is the current state of charge; P storage E represents the energy storage power; Δt represents the time step; E capacity Energy storage capacity; P thermal=clip(P thermal ,P t-1 -Rdown×Δt,P t-1 +Rup×Δt)(8); Among them, P t-1 Rdown represents the thermal power output at the previous moment; Rup represents the downward ramp rate; Rup represents the upward ramp rate. Penalty=λ bal |P load -P total | 2 +λ H max(0,H min -H) 2 +λ SCR max(0,SCR min -SCR) 2 +λ R max(0,R req -R avail ) 2 +λ ramp Σ|ΔP| 2 +λ cyc ×C deg (9); Where Penalty is the soft constraint penalty term; λ bal λ is the power balance penalty factor; H H is the inertia penalty factor; min The minimum inertia requirement; H is the system inertia; λ SCR SCR is the short-circuit ratio penalty factor. min Minimum short-circuit ratio requirement; SCR is the short-circuit ratio; λ R R is the spare penalty coefficient; req For backup needs; R avail Available for standby; λ ramp λ is the climbing penalty coefficient; ΔP is the power change; cyc C is the cyclic penalty coefficient; deg Costs associated with battery degradation.

[0018] Preferably, the formula for calculating the system inertia is shown in equation (10), where J thermal,i Let U be the unit inertia time constant of the i-th thermal power unit. i For unit on / off status, P i To provide power to the generator unit, S base As the baseline capacity, J storage Inertia provided for energy storage; minimum inertia requirement J min RoCoF is limited according to the rate of frequency change. limitThe load level is determined as shown in equation (11); the short-circuit ratio is calculated as shown in equation (12), where S sc The system short-circuit capacity is related to the number of operating thermal power plants, P. renewable The sum of contributions made to the scenic area; H=Σ(H i ×U i ×P i / S base )+H storage ×(P storage / S base (10); Where H is the system inertia time constant; H i U is the unit inertia time constant of the i-th thermal power unit; i The unit is in on / off state; P i To provide power to the generator unit; S base H is the baseline capacity. storage Inertia provided for energy storage; H min =ΔP max / (2×S base ×RoCoF max (11); Among them, H min Minimum inertia requirement; ΔP max For maximum power disturbance; S base Reference capacity; RoCoF max The maximum permissible rate of change of frequency; SCR=S sc / P renewable (12); Where SCR is the short-circuit ratio; S sc The system short-circuit capacity is related to the number of operating thermal power plants; P renewable The sum of contributions made to the scenery.

[0019] Preferably, step S7 specifically includes the following process: A priority experience replay buffer is adopted, and sampling priority and importance sampling weight are set based on TD error; a course learning strategy is adopted, which gradually increases the task difficulty, tightens the constraints and reduces the exploration noise in four stages; a delayed update strategy is used, and the Actor network is updated once every d steps; a soft target network is used for update, with an update coefficient τ=0.005; gradient pruning is performed to prevent gradient explosion.

[0020] Preferably, the sampling strategy of the priority experience playback buffer is as follows: the priority setting of each experience transformation is as shown in equation (13), where TD errorFor TD error, ε is a small constant to prevent zero priority, and α is the priority index with a value of 0.6; the sampling probability is shown in equation (14); the importance sampling weight is shown in equation (15), where β is the importance sampling index with an initial value of 0.4 that linearly increases to 1.0 during training, and N is the buffer size; the course learning is divided into four stages, gradually increasing the task difficulty, tightening the constraints, and reducing the exploration noise; priority=(|δ TD |+ε) α (13); Where priority is the experience priority; δ TD 1 is the TD error; ε is a small constant to prevent zero priority; α is the priority exponent, with a value of 0.6; P(i) = priority i / Σjpriority j (14); Where P(i) is the sampling probability of the i-th experience; priority i The priority of the i-th experience; w i =(N×P(i)) (-β) / max(wj)(15); Among them, w i The importance sampling weight is denoted by N; N is the buffer size; P(i) is the sampling probability; β is the importance sampling index, which is initially 0.4 and increases linearly to 1.0 during training.

[0021] Preferably, the Q-value update formula of the TD3 algorithm in step S7 is shown in equation (16), where s' is the next state, a' is the action selected by the target policy under s' plus the clipped Gaussian noise, γ is the discount factor with a value of 0.99, and α critic τ is the learning rate of the Critic; the Actor loss is shown in Equation (17), where μ is the Actor policy; the delayed update policy is to update the Actor and the target network once every d steps (d=2); the soft update formula is shown in Equation (18), where τ=0.005; gradient clipping uses torch.nn.utils.clipgradnorm, and the maximum norm is set to 1.0; Q t+1 (s,a)=Q t (s,a)+α critic ×[r+γ×min(Q 1,target (s',a'),Q 2,target (s',a'))-Q t (s,a)](16); Among them, Q t+1(s,a) represents the updated Q value; Q t (s,a) represents the current Q value; α critic γ is the Critic learning rate; r is the immediate reward; γ is the discount factor, with a value of 0.99; Q 1,target Q 2,target The target network output is s'; the next state is s'; the target policy action is a'. L actor =-E[Q1(s,μ θ (s))](17); Among them, L actor Here, E represents the Actor loss function; E represents the expectation operation; Q1 represents the first Critic network; μ θ (s) represents the action of the Actor policy in state s; θ target ←τ×θ+(1-τ)×θ target (18); Where, θ target θ represents the target network parameters; θ represents the primary network parameters; and τ represents the soft update coefficient, with a value of 0.005.

[0022] Preferably, step S8 specifically includes the following process: The trained model is deployed to the actual system. At each decision moment, it receives real-time data, updates probability predictions, and executes mid-level and low-level optimization decisions. It achieves online incremental learning, collects online experience to update the Critic network parameters regularly, keeps the Actor stable, and evaluates performance on the validation set regularly. If the performance degrades, it rolls back to the historical checkpoint to prevent catastrophic forgetting. The Kolmogorov-Smirnov test is used to detect data distribution shifts. When a significant shift is detected, the model is retrained or multiple models are integrated.

[0023] Preferably, the execution flow of the rolling optimization framework is as follows: at the current time t, receive real-time data including the actual wind power output P. wind (t), Actual photovoltaic output P pv (t), actual load P load (t) and energy storage SOC(t); call the probabilistic prediction model to update the prediction P for the future time period T. forecast [t+1:t+T]; Input the updated state smid into the middle-level agent and output the unit combination decision and energy storage plan amid; Execute low-level control once every Δt minutes, input the real-time state slow and output the power output adjustment alow; Calculate the prediction error as shown in Equation (19), and re-trigger the middle-level optimization when |error| exceeds the threshold; error=P actual -P forecast (19); Where error is the prediction error; P actual P represents the actual power. forecast To predict power.

[0024] Preferably, step S9 specifically includes the following process: For base clusters containing multiple sub-bases, a Multi-Agent reinforcement learning framework is adopted, in which each base acts as an independent agent to make distributed decisions. During training, a centralized Critic is used to input the global state and all actions, while during execution, a decentralized Actor is used to input only the local state. A coordinator is designed to decompose the global objective and coordinate emergency events.

[0025] Preferably, step S10 specifically includes the following process: Economic indicators include total investment cost, annual operating cost, levelized cost of electricity (LCOE), internal rate of return (IRR), and payback period; reliability indicators include power supply reliability, curtailment rate, energy storage cycle life, and reserve capacity adequacy; system support capability indicators include average system inertia, minimum short-circuit ratio, maximum rate of change of frequency (RoCoF), minimum frequency point, and frequency regulation response time; environmental benefit indicators include annual emission reduction, renewable energy share, and carbon emission reduction benefits; single-parameter sensitivity analysis, Monte Carlo simulation, and extreme scenario testing are conducted to verify the robustness of the scheme.

[0026] Preferably, the calculation formula for the levelized cost of electricity (LCOE) is as shown in equation (20), where C invest The total investment cost is given by CRF, which is the capital recovery factor as shown in equation (21), where r is the discount rate, n is the project life, and C... operate,annual E is the annual operating cost. generated,annual The annual power generation is given; the curtailment rate is shown in equation (22); LCOE=(C invest ×CRF+C operate ) / E generated (20); Where LCOE is the cost per kilowatt-hour; C invest Total investment cost; CRF is the capital recovery factor; C operate Annual operating cost; E generated Annual power generation; CRF = r(1+r) n / [(1+r) n -1](21); Where CRF is the capital recovery factor; r is the discount rate; and n is the project life in years. R curtail =E curtail / (E wind +E pv)×100%(22); Among them, R curtail E represents the curtailment rate. curtail E represents the amount of electricity wasted. wind E represents wind power generation. pv This refers to photovoltaic power generation.

[0027] This invention also discloses a large-scale wind, solar, thermal, and energy storage base scheduling decision-making and optimization allocation system based on hierarchical reinforcement learning, comprising: The data acquisition module is used to acquire historical meteorological data, power grid parameters, equipment technical and economic parameters, load data, and policy constraints of the target base, and to perform preprocessing. The probability prediction module is used to perform probability prediction of wind and solar power based on the Transformer architecture, output quantile prediction values, and generate representative operating scenarios. The state space design module is used to define long-term planning states, short-term scheduling states, and real-time control states. The hierarchical reinforcement learning module includes a high-level agent for capacity allocation decisions, a mid-level agent for day-ahead scheduling decisions, and a low-level agent for real-time control decisions. The neural network module includes an Actor network and a Critic network, and has an encoding layer, a temporal processing layer, and an attention layer; The constraint processing module is used to embed key constraints such as system inertia and short-circuit ratio through projection and penalty terms. The training module is used to train the neural network using priority experience replay, delayed policy update, and soft target network update. The online optimization module is used to perform rolling optimization decisions and online incremental learning; The collaborative optimization module is used to achieve collaborative optimization across multiple sites. The evaluation module is used to calculate indicators of economic efficiency, reliability, system support capability, and environmental benefits, and to perform sensitivity analysis.

[0028] Furthermore, in the hierarchical reinforcement learning module, the high-level agent adopts the TD3 algorithm, which includes an Actor network, two Critic networks and their corresponding target networks; the mid-level agent adopts the MADDPG algorithm, which includes multiple Actor networks and a centralized Critic network; and the low-level agent adopts the PPO algorithm, which includes an Actor network, a Critic network and an old policy network.

[0029] Furthermore, the constraint processing module includes a projection submodule and a penalty submodule. The projection submodule realizes real-time projection of the action onto the constraint feasible region, and the penalty submodule adds a constraint violation penalty term in the reward function.

[0030] Furthermore, the online optimization module includes a rolling optimization submodule, an incremental learning submodule, and a distribution shift detection submodule. The rolling optimization submodule realizes real-time data reception and dynamic decision-making, the incremental learning submodule realizes online experience collection and model parameter update, and the distribution shift detection submodule uses statistical testing methods to judge changes in data distribution.

[0031] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal and energy storage bases based on hierarchical reinforcement learning.

[0032] The present invention also discloses a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning.

[0033] Beneficial effects of this invention: 1. This invention effectively solves the curse of dimensionality problem in large-scale systems. By employing deep neural networks to approximate the value function and policy function, the traditional discrete state representation is replaced with a continuous state representation, reducing the dimensionality of the state space from exponential growth to linear growth. In this embodiment, which includes two thermal power units, several wind and solar power plants, and an energy storage system base, the state space can be fully characterized by only a 50-dimensional vector, while the number of states required by traditional tabular methods would reach astronomical levels. This allows the method of this invention to be extended to wind, solar, thermal, and energy storage systems of any scale.

[0034] 2. This invention achieves multi-timescale collaborative optimization of investment planning and operation scheduling. The hierarchical intelligent agent architecture organically integrates long-term capacity configuration (quarterly level), medium-term unit combination (day-ahead level), and short-term output adjustment (minute level), establishing a complete decision-making closed loop through inter-layer information transmission. High-level investment decisions consider feedback from the operational performance of mid- and low-level layers, while mid- and low-level scheduling decisions seek optimization within the capacity boundaries determined by high-level layers, avoiding the problem of disconnect between investment planning and operation scheduling in traditional methods.

[0035] 3. This invention innovatively embeds key constraints such as system inertia and short-circuit ratio, which are crucial for high renewable energy usage, into a reinforcement learning framework. Through a hybrid constraint handling strategy combining projection and penalty terms, it ensures that the optimization results meet the safety requirements of frequency stability and voltage support while pursuing economic efficiency.

[0036] 4. This invention possesses excellent engineering practicality and scalability. The method framework has good scalability, allowing for adjustments to the network structure, algorithm parameters, and constraints according to actual project needs. It is suitable for various application scenarios, including single-site, multi-site, and integrated source-grid-load-storage systems.

[0037] 5. The technical solution of this invention can be applied to wind-solar-thermal-storage systems of various scales, including single-site, multi-site, and integrated source-grid-load-storage systems. The network structure, algorithm parameters, and constraints can be adjusted according to actual conditions. Attached Figure Description

[0038] Figure 1 This is a schematic diagram of the overall process of the method of the present invention; Figure 2 This is a schematic diagram of a hierarchical deep reinforcement learning architecture; Figure 3 This is a schematic diagram of a deep neural network structure. Detailed Implementation

[0039] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0040] Example 1: As Figure 1 The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning is shown. Taking the large-scale base in the desert and Gobi Desert as an example, it mainly includes the following basic processes: I. Project Background and Basic Parameter Acquisition: This embodiment takes a 5GW wind and solar power base in the Gobi Desert region of Northwest my country as the research object. This base is one of the key large-scale wind and solar power bases under national construction. The project is located in the desert and Gobi region and has advantages such as abundant wind and solar resources, low land costs, and small ecological impact. However, it also faces technical challenges such as being far from load centers, weak grid support capacity, and large fluctuations in renewable energy output.

[0041] The base is planned to have a total installed capacity of 5000MW, including 3000MW of wind power and 2000MW of photovoltaic power. It will also include two 660MW coal-fired power units as regulating power sources and a certain scale of electrochemical energy storage system. Electricity will be transmitted to the eastern load center via an 800kV ultra-high voltage direct current line, with approximately 800MW of load absorbed locally. Due to the high proportion of renewable energy (over 80%), the system exhibits typical weak grid characteristics, with a short-circuit ratio of only 1.5, placing stringent requirements on system inertia and voltage support capabilities. Basic project parameters are shown in Table 1.

[0042] Table 1 Basic Parameters of the Project

[0043] To establish an accurate optimization model, the project team collected three years of historical operational data for the region, including hourly output records of wind and solar power, load curves, electricity market price series, and meteorological numerical forecast data. During data preprocessing, the 3σ criterion was used for outlier detection, time series interpolation was used to fill missing values, and Min-Max normalization was used to map wind and solar power output to the [0,1] interval, providing a high-quality data foundation for subsequent machine learning model training.

[0044] II. Construction of Probabilistic Prediction Model: Accurate wind and solar power forecasting is a prerequisite for optimized scheduling. Considering that traditional point forecasting methods cannot characterize the uncertainty of forecasting, this embodiment constructs a probabilistic forecasting model based on the Transformer architecture, which can simultaneously output multiple quantile forecast values, providing probability distribution information for subsequent stochastic optimization.

[0045] The prediction model employs an encoder-decoder structure. Input features include historical 24-hour power sequences, 72-hour meteorological numerical forecasts (wind speed, irradiance, temperature), and time-coded information, totaling a 168-dimensional feature vector. The core model utilizes a multi-head self-attention mechanism, employing eight attention heads in parallel computation to effectively capture the long-term temporal dependencies of wind and solar power output and the spatial correlations among multiple solar stations. The output layer is designed with three independent quantile regression branches, predicting the P10 (underestimation boundary), P50 (median prediction), and P90 (overestimation boundary) quantile values ​​for the next 24 hours, respectively. Model structure parameters are shown in Table 2.

[0046] Table 2 Probabilistic Prediction Model Structure

[0047] The model was trained using the quantile loss function and the Adam optimizer, with an initial learning rate of 1e-3, a batch size of 64, and a total training duration of 100 epochs. Evaluation results on the test set showed that the root mean square error (RMSE) of the P50 prediction was 0.08 (normalized value), and the actual coverage of the P10 and P90 predictions reached 89.5%, close to the theoretical value of 80%, indicating that the prediction interval has good calibration performance and can reliably characterize the uncertainty range of wind and solar power output.

[0048] III. Layered Reinforcement Learning Framework: like Figure 2As shown, to address the multi-timescale characteristics of optimization problems in large-scale wind, solar, thermal, and energy storage systems, this embodiment designs a three-layer intelligent agent architecture to achieve collaborative optimization from long-term capacity planning to real-time power control. This layered design draws inspiration from the hierarchical structure of human decision-making: the strategic level focuses on long-term investment returns, the tactical level optimizes day-ahead operating plans, and the operational level handles real-time disturbance responses.

[0049] The high-level agent is responsible for capacity allocation decisions, employing the Twin Delayed DDPG (TD3) algorithm with a quarterly decision-making cycle. The agent's state space includes eight dimensions: current installed capacity, grid strength indicators, carbon allowance margin, and investment budget. Its action space comprises capacity increases for wind power, solar power, and energy storage, constituting a continuous action space. The reward function comprehensively considers investment costs, expected operating costs, carbon emission costs, curtailment penalties, and reliability and flexibility rewards, guiding the agent to find a balance between economic efficiency and technical feasibility.

[0050] The middle-level agent is responsible for day-ahead scheduling decisions. The Multi-Agent DDPG (MADDPG) algorithm is used to model the two thermal power units and the energy storage system as three cooperative agents. This multi-agent design can handle the heterogeneity constraints of different types of power sources, while achieving implicit coordination among agents through a centralized training-decentralized execution framework. Each agent's Actor network only receives locally observable states, while the Critic network during training receives the global state and the actions of all agents, thus achieving globally optimal performance while ensuring execution efficiency.

[0051] The lower-level agent is responsible for real-time control decisions, employing the Proximal Policy Optimization (PPO) algorithm with a decision cycle of 5 minutes. The main task of this layer is to track the power output plan formulated by the middle layer, while simultaneously addressing real-time disturbances such as wind and solar forecast errors and load fluctuations, ensuring system power balance and frequency stability. The pruning objective function of the PPO algorithm effectively limits the policy update magnitude, avoiding performance crashes caused by aggressive updates, making it particularly suitable for real-time control scenarios in power systems with high safety requirements. The configuration parameters for each layer of agents are shown in Table 3.

[0052] Table 3 Hierarchical Agent Configuration

[0053] IV. Deep Neural Network Architecture: like Figure 3As shown, deep neural networks are the core component of hierarchical reinforcement learning frameworks, and their architecture design directly affects the decision-making quality and learning efficiency of the agent. This embodiment designs a dedicated Actor-Critic network structure for each layer of the agent, fully considering the temporal dependencies and multi-agent collaborative characteristics of the power system optimization problem.

[0054] Taking the mid-level agent as an example, the Actor network adopts a five-stage structure of encoding-temporal processing-attention-fusion-output. First, the state vector is mapped to a 256-dimensional high-dimensional space through a linear encoding layer, and then normalized by LayerNorm and activated by ReLU. Subsequently, the LSTM layer processes the temporal dependencies of the 24-hour prediction sequence with 128-dimensional hidden states. Next, the Multi-head Self-Attention layer (8 attention heads, each with a dimension of 16) captures the spatial dependencies between thermal power units and energy storage. Finally, the output layer, after feature fusion and Tanh activation, generates an action vector normalized to the [-1,1] interval.

[0055] The Critic network employs a dual-network design to mitigate the Q-value overestimation problem. Each Critic network contains an independent state encoder and an action encoder, which map the state and action to a high-dimensional feature space respectively and then concatenate and fuse them. The output state-action value estimate is then passed through a three-layer fully connected network (384→256→128→1). The smaller value of the two Critic network outputs is used when calculating the target Q-value, effectively suppressing the overly optimistic estimation of the value function. The neural network structure parameters are shown in Table 4.

[0056] Table 4 Neural Network Structure Parameters

[0057] V. Key Constraint Embedding Method: Power system optimization must strictly meet physical constraints and operational safety requirements. This embodiment employs a hybrid strategy of hard constraint projection and soft constraint penalty, effectively embedding various constraints into the reinforcement learning framework to ensure that the agent's output decisions always meet feasibility requirements.

[0058] For hard constraints such as power balance, unit output range, energy storage state of charge (SOC) boundary, and ramp rate, a projection method is used. Specifically, when the original action output by the agent violates the constraints, the algorithm projects it to the nearest feasible domain boundary. For example, when the calculated SOC of the energy storage at the next moment exceeds the safe range of [10%, 90%], the charging and discharging power is automatically adjusted so that the SOC just reaches the boundary value; when the output change of the thermal power unit exceeds the ramp rate limit (±6MW / min), the change is trimmed to the allowable range.

[0059] For soft constraints such as system inertia, short-circuit ratio, and frequency regulation reserve, a penalty term is added to the reward function. The penalty term is designed as a quadratic function of the constraint violation amount; the penalty is zero when the constraint is satisfied, and increases quadratically with the degree of violation when the constraint is violated. The penalty coefficient is adjusted experimentally to ensure that the penalty for violating the constraint is large enough to guide the agent to learn a strategy that satisfies the constraint. Specifically, the system inertia constraint requires an equivalent inertia time constant of not less than 3.0 seconds to ensure that the frequency change rate does not exceed the safety limit of 0.5 Hz / s; the short-circuit ratio constraint requires an SCR of not less than 1.0 to maintain the system's voltage support capability. The constraint parameters are shown in Table 5.

[0060] Table 5 Constraint Parameters

[0061] VI. Training Algorithm Implementation: Efficient and stable training is key to the successful application of deep reinforcement learning. This embodiment comprehensively utilizes advanced techniques such as priority experience replay, delayed policy updates, soft-target network updates, and curriculum learning, significantly improving training efficiency and final performance.

[0062] The capacity of the priority experience replay buffer is set to 1 million experience transformations, and sampling priority is assigned to each experience based on the temporal difference (TD) error. Experiences with larger TD errors contain more learning signals and are therefore assigned a higher sampling probability. At the same time, importance sampling weights are introduced to correct the distribution offset caused by priority sampling. The weight exponent β increases linearly from an initial value of 0.4 to 1.0 during training, restoring unbiased estimation in the later stages of training.

[0063] The training employs a course-based learning strategy, dividing the learning process into four stages with progressively increasing difficulty: The first stage uses simple scenarios and relaxed constraints, exploring a noise standard deviation of 0.3 to allow the agent to quickly establish basic decision-making capabilities; the second stage increases scenario diversity, restoring normal constraints, and reducing noise to 0.2; the third stage introduces challenging scenarios (such as extreme weather and equipment failure), implementing strict constraints, and reducing noise to 0.1; the fourth stage uses all scenarios and complete constraints, gradually decreasing noise from 0.05 to 0.01, fine-tuning the policy performance. The training hyperparameter configurations are shown in Table 6.

[0064] Table 6 Training Hyperparameter Configuration

[0065] VII. Performance Evaluation Results: To verify the effectiveness of the method of this invention, a comprehensive evaluation was conducted on the trained agent in a test scenario, and a comparative analysis was performed with traditional deterministic optimization methods. The evaluation indicators covered three dimensions: economy, reliability, and system support capability.

[0066] In terms of economic indicators, the levelized cost of electricity (LCOE) of the method of this invention is 0.297 yuan / kWh, which is 11.9% lower than the 0.337 yuan / kWh of the traditional method. This significant improvement mainly comes from two aspects: First, the hierarchical optimization framework realizes the coordination between investment decision-making and operation scheduling, avoiding the disconnect between capacity allocation and actual operating demand; second, multi-agent collaborative optimization fully taps the flexibility and adjustment potential of thermal power plants. The annual operating cost is reduced to 1.25 billion yuan, and the economic feasibility of the project is significantly improved.

[0067] In terms of reliability indicators, the curtailment rate decreased significantly from 7.5% using traditional methods to 3.2%, a reduction of 57%. This means that approximately 500 million kWh of clean energy can be absorbed annually, which, based on a grid connection price of 0.3 yuan / kWh, is equivalent to an increase of approximately 150 million yuan in electricity sales revenue per year. Meanwhile, the system met the inertia and short-circuit ratio constraints in all 8760 hours of simulation testing throughout the year, verifying the effectiveness of the constraint embedding method. A comparison of economic indicators is shown in Table 7.

[0068] Table 7 Comparison of Economic Indicators

[0069] To evaluate the robustness of the optimized scheme, a Monte Carlo simulation method was used to conduct a sensitivity analysis on key uncertain parameters. Probability distributions were set for four parameters: wind power capacity factor, photovoltaic capacity factor, energy storage cost, and electricity price. 1000 random samples were taken, and the distribution characteristics of LCOE were statistically analyzed. The results show that the mean LCOE is 0.301 yuan / kWh, the standard deviation is only 0.018 yuan / kWh, and the 95% confidence interval is [0.268, 0.335] yuan / kWh. This indicates that the optimized scheme has good robustness under parameter fluctuations and can adapt to future changes in the market environment and technology costs.

[0070] The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The scope of protection of the present invention should be limited to the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning, characterized in that, Includes the following steps: Step S1: Data acquisition and preprocessing. Obtain historical meteorological data, power grid parameters, equipment technical and economic parameters, load data and policy constraints of the target base. Perform outlier detection, missing value filling and normalization on the data, and extract time series features. Step S2: Constructing the probabilistic prediction model. A deep learning model based on the Transformer architecture is used to predict the probabilistic power of wind and solar power, outputting the P value for future periods. 10 P 50 P 90 Quantile prediction values ​​are generated by hypercubic sampling based on the predicted quantiles, and the spatiotemporal correlation of multiple stations is modeled. Scene reduction is achieved through K-means clustering. Step S3: Multi-timescale state space design, defining the long-term planning state including wind, solar, thermal and storage installed capacity, grid strength index, carbon quota margin and investment budget; the short-term dispatch state including wind and solar load forecast sequence, energy storage state of charge, unit operation state, system inertia and short-circuit ratio; and the real-time control state including actual output, frequency deviation and voltage deviation. Step S4: Construct a hierarchical deep reinforcement learning framework and establish a three-layer agent architecture; wherein, the high-level agent adopts Tw i The nDelayed DDPG algorithm is used, the Multi-Agent DDPG algorithm is used for the middle-level agents, and the ProximalPolicy Optimization algorithm is used for the lower-level agents. Step S5: Deep neural network architecture design, designing an Actor-Critic network for each layer of the agent; Step S6: Key Constraint Embedding Method; Step S7: Implement the training algorithm; Step S8: Online adaptive and scrolling optimization; Step S9: Multi-site collaborative optimization; Step S10: Performance evaluation and index calculation.

2. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 1, characterized in that, The probabilistic prediction model structure in step S2 is as follows: the input layer receives historical power sequences, numerical weather forecasts, satellite cloud images, and time codes, which are then converted into a high-dimensional representation through a temporal embedding layer and input into a Multi-head A. t The attention layer performs self-attention calculations to capture long-term dependencies, performs nonlinear transformations through a feedforward network, and finally outputs P from three independent quantile regression layers. 10 P 50 P 90 Quantile prediction values; the quantile loss function is used during training, as shown in equation (1): L quantile (pred,target,τ)=mean(max(τ×(target-pred),(τ-1)×(target-pred)))(1); Among them, L quantile τ is the quantile loss function; pred is the predicted value; target is the true value; τ is the quantile level; mean is the mean function; max is the maximum value function.

3. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 1, characterized in that, Step S4 specifically includes the following process: High-level intelligent agents adopt Tw i The n-level Delayed DDPG algorithm has a decision-making cycle of monthly or quarterly, an action space of capacity increment, and a reward function that includes investment cost, expected operating cost, carbon cost, reliability reward, and flexibility reward. The middle-level agent uses the Multi-Agent DDPG algorithm, with a decision-making cycle of 24 hours before the day of the day. Each power generation unit is treated as an agent, and the action space includes unit combination, thermal power output setting, energy storage charging and discharging power, and reserve allocation. The reward function includes fuel cost, start-up and shutdown cost, maintenance cost, curtailment cost, reserve cost, and revenue from the energy market and ancillary services market. The lower-level agent uses the Proximal Policy Optimization algorithm, with a decision-making cycle of minutes. The action space includes thermal power and energy storage output adjustments and wind and solar curtailment. The reward function includes the square of the power deviation term, curtailment penalty, frequency deviation penalty, and ramp-up cost.

4. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 3, characterized in that, The capacity configuration actions of the high-level intelligent agent include wind power capacity increase ΔE. wind,cap Photovoltaic capacity increase ΔE pv,cap Energy storage power capacity increment ΔE storage,cap The reward function for the thermal power plant flexibility transformation level is shown in Equation (2); the middle-level agent adopts the MADDPG framework with centralized training and decentralized execution. The Critic of each power generation unit agent inputs the global state and the actions of all agents, while the Actor only inputs the local observable state; the PPO algorithm of the low-level agent adopts the pruning objective function, as shown in Equation (3). R high =-(C invest +C operate +C carbon +C curtail )+R reliability +R flexibility (2); Among them, R high Rewards for high-level intelligent agents; C invest For investment costs; C operate For expected operating costs; C carbon For carbon cost; C curtail As a penalty for abandoning electricity; R reliability For reliability rewards; R flexibility Rewards for flexibility; L clip =min(r(θ)×Â,clip(r(θ),1-ε,1+ε)×Â)(3); Among them, L clip Let r(θ) be the objective function for PPO pruning; r(θ) be the probability ratio of the new and old strategies;  be the advantage function estimate; ε be the pruning parameter; clip be the pruning function; and min be the minimum function.

5. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 1, characterized in that, Step S5 specifically includes the following process: The Actor network consists of an input layer, an encoding layer, a temporal processing layer, an attention layer, a feature fusion layer, and an output layer. The temporal processing layer uses LSTM or Transformer Block to handle time series dependencies, and the attention layer uses Multi-head Self-A. t Attention is used to handle multi-site or multi-unit collaboration; the Critic network contains independent state encoders and action encoders, which are spliced ​​and fused to output the Q value. A dual-Critic network design is adopted to alleviate the problem of Q value overestimation.

6. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning, as described in claim 5, is characterized in that... The structure of the Actor network is as follows: the state vector s is input into the linear layer Linear(dims, 256), normalized and activated by the LayerNorm layer, and then input into the LSTM(256, 128) or Transformer Block for time-series processing, and then processed through Multi-head Self-A. t The attention (numheads=8) layer handles the spatial dependencies of multiple stations or multiple units. After passing through the feature fusion layer Linear(128,128)+ReLU+Dropout(0.2), the output layer Linear(128,dima)+Tanh outputs the action normalized to [-1,1]. The Critic network contains a state encoder Linear(dims,256)+ReLU and an action encoder Linear(dima,128)+ReLU. After concatenation, the Q value is output through the deep network Linear(384,256)+ReLU+Dropout(0.1)→Linear(256,128)+ReLU→Linear(128,1). The dual Critic networks Critic1 and Critic2 are used. The target value is calculated as shown in Equation (4). Q target =min(Q 1,target (s',a'),Q 2,target (s',a'))(4); Among them, Q target The target Q value; Q 1,target For the first target Critic network output; Q 2,target The output of the Critic network is for the second objective; s' is the next state; a' is the next action.

7. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 1, characterized in that, Step S6 specifically includes the following process: Hard constraints, including power balance constraints, unit output upper and lower limits constraints, energy storage SOC constraints, and ramp rate constraints, are handled through a projection method, projecting actions onto the constraint feasible region in real time. Soft constraints, including system inertia constraints, short-circuit ratio constraints, frequency regulation capability constraints, and energy storage cycle life constraints, are handled by adding a penalty term to the reward function. The system inertia constraint is J. system ≥J min J system The weighted summation is calculated from the inertia contributions of the thermal power unit and energy storage; the short-circuit ratio constraint is SCR=S sc / P renewable ≥SCR min .

8. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 7, characterized in that, The specific implementation process of the projection method is as follows: For power balance constraints, when the total output P... total Overload P load When the controllable output is reduced proportionally, as shown in equation (5); for unit output constraints, the thermal power output is limited to the constraint range, as shown in equation (6), where U is the unit switching state; for energy storage SOC constraints, the state of charge at the next moment is calculated, as shown in equation (7), if it exceeds [SOC min SOC max The range is then adjusted by P. storage For the ramp constraint, the rate of change of thermal power output is limited, as shown in Equation (8); the soft constraint penalty term is shown in Equation (9). P total =P load (5); Among them, P total For total output; P load For load power; P thermal =clip(P thermal ,U×P min ,U×P max )(6); Among them, P thermal For thermal power output; U represents the unit's on / off status; P min For minimum technical effort; P max To maximize technical output; clip is the clipping function; SOC next =SOC+(P storage ×Δt) / E capacity (7); Among them, SOC next The next state of charge; SOC is the current state of charge; P storage E represents the energy storage power; Δt represents the time step; E capacity Energy storage capacity; P thermal =clip(P thermal ,P t-1 -Rdown×Δt,P t-1 +Rup×Δt)(8); Among them, P t-1 Rdown represents the thermal power output at the previous moment; Rup represents the downward ramp rate; Rup represents the upward ramp rate. Penalty=l bal |P load -P total | 2 +λ H max(0,H min -H) 2 +λ SCR max(0,SCR min -SCR) 2 +λ R max(0,R req -R avail ) 2 +λ ramp S|ΔP| 2 +λ cyc ×C deg (9); Where Penalty is the soft constraint penalty term; λ bal λ is the power balance penalty factor; H H is the inertia penalty factor; min The minimum inertia requirement; H is the system inertia; λ SCR SCR is the short-circuit ratio penalty factor. min Minimum short-circuit ratio requirement; SCR is the short-circuit ratio; λ R R is the spare penalty coefficient; req For backup needs; R avail Available for standby; λ ramp λ is the climbing penalty coefficient; ΔP is the power change; cyc C is the cyclic penalty coefficient; deg Costs associated with battery degradation.

9. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 7, characterized in that, The formula for calculating the system inertia is shown in equation (10), where J thermal,i Let U be the unit inertia time constant of the i-th thermal power unit. i For unit on / off status, P i To provide power to the generator unit, S base As the baseline capacity, J storage Inertia provided for energy storage; minimum inertia requirement J min RoCoF is limited according to the rate of frequency change. limit The load level is determined as shown in equation (11); the short-circuit ratio is calculated as shown in equation (12), where S sc The system short-circuit capacity is related to the number of operating thermal power plants, P. renewable The sum of contributions made to the scenic area; H=Σ(H i ×U i ×P i / S base )+H storage ×(P storage / S base )(10); Where H is the system inertia time constant; H i U is the unit inertia time constant of the i-th thermal power unit; i The unit is in on / off state; P i To provide power to the generator unit; S base H is the baseline capacity. storage Inertia provided for energy storage; H min =ΔP max / (2×S base ×RoCoF max )(11); Among them, H min Minimum inertia requirement; ΔP max For maximum power disturbance; S base Reference capacity; RoCoF max The maximum permissible rate of change of frequency; SCR=S sc / P renewable (12); Where SCR is the short-circuit ratio; S sc The system short-circuit capacity is related to the number of operating thermal power plants; P renewable The sum of contributions made to the scenery.

10. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 1, characterized in that, The specific steps S7 are as follows Includes the following processes: A priority experience replay buffer is adopted, and sampling priority and importance sampling weight are set based on TD error; a course learning strategy is adopted, which gradually increases the task difficulty, tightens the constraints and reduces the exploration noise in four stages; a delayed update strategy is used, and the Actor network is updated once every d steps; a soft target network is used for update, with an update coefficient τ=0.005; gradient pruning is performed to prevent gradient explosion.

11. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 10, characterized in that, The sampling strategy for the priority experience playback buffer is as follows: the priority setting for each experience transformation is shown in equation (13), where TD error For TD error, ε is a small constant to prevent zero priority, and α is the priority exponent with a value of 0.6; the sampling probability is shown in equation (14); The importance sampling weights are shown in Equation (15), where β is the importance sampling index, which initially increases from 0.4 to 1.0 during training, and N is the buffer size. The course learning is divided into four stages, gradually increasing the task difficulty, tightening the constraints, and reducing the exploration noise. priority=(|d TD |+e) α (13); Where priority is the experience priority; δ TD 1 is the TD error; ε is a small constant to prevent zero priority; α is the priority exponent, with a value of 0.6; P(i)=priority i / Σjpriority j (14); Where P(i) is the sampling probability of the i-th experience; priority i The priority of the i-th experience; w i =(N×P(i)) (-β) / max(wj)(15); Among them, w i The importance sampling weight is denoted by N; N is the buffer size; P(i) is the sampling probability; β is the importance sampling index, which is initially 0.4 and increases linearly to 1.0 during training.

12. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 10, characterized in that, The Q-value update formula for the TD3 algorithm in step S7 is shown in equation (16), where s' is the next state, a' is the action selected by the target policy under s' plus the clipped Gaussian noise, γ is the discount factor with a value of 0.99, and α critic τ is the learning rate of the Critic; the Actor loss is shown in Equation (17), where μ is the Actor policy; the delayed update policy is to update the Actor and the target network once every d steps (d=2); the soft update formula is shown in Equation (18), where τ=0.005; gradient clipping uses torch.nn.utils.clipgradnorm, and the maximum norm is set to 1.0; Q t+1 (s,a)=Q t (s,a)+α critic ×[r+γ×min(Q 1,target (s',a'),Q 2,target (s',a'))-Q t (s,a)](16); Among them, Q t+1 (s,a) represents the updated Q value; Q t (s,a) represents the current Q value; α critic γ is the Critic learning rate; r is the immediate reward; γ is the discount factor, with a value of 0.99; Q 1,target Q 2,target The target network output is s'; the next state is s'; the target policy action is a'. L actor =-E[Q1(s,μ θ (s))](17); Among them, L actor Here, E represents the Actor loss function; E represents the expectation operation; Q1 represents the first Critic network; μ θ (s) represents the action of the Actor policy in state s; i target ←τ×θ+(1-τ)×θ target (18); Where, θ target θ represents the target network parameters; θ represents the primary network parameters; and τ represents the soft update coefficient, with a value of 0.

005.

13. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 1, characterized in that, Step S8 specifically includes the following process: The trained model is deployed to the actual system. At each decision moment, it receives real-time data, updates probability predictions, and executes mid-level and low-level optimization decisions. It achieves online incremental learning, collects online experience to update the Critic network parameters regularly, keeps the Actor stable, and evaluates performance on the validation set regularly. If the performance degrades, it rolls back to the historical checkpoint to prevent catastrophic forgetting. The Kolmogorov-Smirnov test is used to detect data distribution shifts. When a significant shift is detected, the model is retrained or multiple models are integrated.

14. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning, as described in claim 13, is characterized in that... The execution flow of the rolling optimization framework is as follows: At the current time t, receive real-time data including the actual wind power output P. wind (t), Actual photovoltaic output P pv (t), actual load P load (t) and energy storage SOC(t); Call the probabilistic prediction model to update the prediction P for the future time period T. forecast [t+1:t+T]; Input the updated state smid into the middle-level agent and output the unit combination decision and energy storage plan amid; Execute low-level control once every Δt minutes, input the real-time state slow and output the power output adjustment alow; Calculate the prediction error as shown in Equation (19), and re-trigger the middle-level optimization when |error| exceeds the threshold; error=P actual -P forecast (19); Where error is the prediction error; P actual P represents the actual power. forecast To predict power.

15. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 1, characterized in that, Step S9 specifically includes the following process: For base clusters containing multiple sub-bases, a Multi-Agent reinforcement learning framework is adopted, in which each base acts as an independent agent to make distributed decisions. During training, a centralized Critic is used to input the global state and all actions, while during execution, a decentralized Actor is used to input only the local state. A coordinator is designed to decompose the global objective and coordinate emergency events.

16. The method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning according to claim 1, characterized in that, The specific steps S10 are as follows Includes the following processes: Economic indicators include total investment cost, annual operating cost, levelized cost of electricity (LCOE), internal rate of return (IRR), and payback period; reliability indicators include power supply reliability, curtailment rate, energy storage cycle life, and reserve capacity adequacy. The computing system support capability indicators include average system inertia, minimum short-circuit ratio, maximum frequency change rate RoCoF, minimum frequency point, and frequency modulation response time; The environmental benefit indicators include annual emission reduction, the proportion of renewable energy, and carbon emission reduction benefits; single-parameter sensitivity analysis, Monte Carlo simulation, and extreme scenario testing are conducted to verify the robustness of the scheme.

17. A method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal, and energy storage bases based on hierarchical reinforcement learning as described in claim 16, characterized in that, The formula for calculating the levelized cost of electricity (LCOE) is shown in equation (20), where C invest The total investment cost is given by CRF, which is the capital recovery factor as shown in equation (21), where r is the discount rate, n is the project life, and C... operate,annual E is the annual operating cost. generated,annual The annual power generation is given; the curtailment rate is shown in equation (22); LCOE=(C invest ×CRF+C operate ) / E generated (20); Where LCOE is the cost per kilowatt-hour; C invest Total investment cost; CRF is the capital recovery factor; C operate Annual operating cost; E generated Annual power generation; CRF=r(1+r) n / [(1+r) n -1](21); Where CRF is the capital recovery factor; r is the discount rate; and n is the project life in years. R curtail =E curtail / (AND wind +E pv )×100%(22); Among them, R curtail E represents the curtailment rate. curtail E represents the amount of electricity wasted. wind E represents wind power generation. pv This refers to photovoltaic power generation.

18. A large-scale wind, solar, thermal, and energy storage base scheduling decision-making and optimization allocation system based on hierarchical reinforcement learning, characterized in that, include: The data acquisition module is used to acquire historical meteorological data, power grid parameters, equipment technical and economic parameters, load data, and policy constraints of the target base, and to perform preprocessing. The probability prediction module is used to perform probability prediction of wind and solar power based on the Transformer architecture, output quantile prediction values, and generate representative operating scenarios. The state space design module is used to define long-term planning states, short-term scheduling states, and real-time control states. The hierarchical reinforcement learning module includes a high-level agent for capacity allocation decisions, a mid-level agent for day-ahead scheduling decisions, and a low-level agent for real-time control decisions. The neural network module includes an Actor network and a Critic network, and has an encoding layer, a temporal processing layer, and an attention layer; The constraint processing module is used to embed key constraints such as system inertia and short-circuit ratio through projection and penalty terms. The training module is used to train the neural network using priority experience replay, delayed policy update, and soft target network update. The online optimization module is used to perform rolling optimization decisions and online incremental learning; The collaborative optimization module is used to achieve collaborative optimization across multiple sites. The evaluation module is used to calculate indicators of economic efficiency, reliability, system support capability, and environmental benefits, and to perform sensitivity analysis.

19. A large-scale wind, solar, thermal, and energy storage base scheduling decision-making and optimization allocation system based on hierarchical reinforcement learning as described in claim 11, characterized in that, In the hierarchical reinforcement learning module, the high-level agent uses the TD3 algorithm, which includes an Actor network, two Critic networks and their corresponding target networks; the mid-level agent uses the MADDPG algorithm, which includes multiple Actor networks and a centralized Critic network; and the low-level agent uses the PPO algorithm, which includes an Actor network, a Critic network and an old policy network.

20. A large-scale wind, solar, thermal, and energy storage base scheduling decision-making and optimization allocation system based on hierarchical reinforcement learning as described in claim 11, characterized in that, The constraint processing module includes a projection submodule and a penalty submodule. The projection submodule realizes real-time projection of the action onto the constraint feasible region, and the penalty submodule adds a constraint violation penalty term to the reward function.

21. The system according to claim 11, characterized in that, The online optimization module includes a rolling optimization submodule, an incremental learning submodule, and a distribution shift detection submodule. The rolling optimization submodule realizes real-time data reception and dynamic decision-making, the incremental learning submodule realizes online experience collection and model parameter update, and the distribution shift detection submodule uses statistical testing methods to judge changes in data distribution.

22. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal and energy storage bases based on hierarchical reinforcement learning as described in any one of claims 1-17.

23. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a method for scheduling decision-making and optimal allocation of large-scale wind, solar, thermal and energy storage bases based on hierarchical reinforcement learning as described in any one of claims 1-17.