Reservoir scheduling optimization method and system based on aerospace big data
By integrating aerospace big data with ground-based hydraulic monitoring data and employing a reinforcement learning-based scheduling decision-making mechanism, the problems of insufficient utilization of aerospace information and limited strategy stability in existing reservoir scheduling technologies have been solved. This has enabled efficient perception of complex hydrological scenarios and engineering-level scheduling execution, thereby improving the stability and reliability of reservoir scheduling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-05
- Publication Date
- 2026-03-31
AI Technical Summary
Existing intelligent reservoir scheduling technologies have problems such as insufficient utilization of aerospace information, limited stability of scheduling strategies, and difficulty in directly supporting engineering-level scheduling execution. In particular, under multi-objective constraints, they are prone to unstable value functions, slow strategy convergence, or fluctuations in scheduling results, making it difficult to meet the real-time, stable, and executable scheduling requirements under complex hydrological scenarios.
A reservoir scheduling optimization method and system based on aerospace big data is proposed. By deeply integrating aerospace remote sensing observation data with ground hydraulic engineering monitoring data, a multi-dimensional fusion state expression is constructed. Combined with reinforcement learning scheduling decision-making mechanism and multi-model integration, a comprehensive representation of watershed inflow conditions, underlying surface hydrological conditions and hydraulic structure operating conditions is achieved. Satisfactory value constraints and dynamic satisfaction baselines are introduced to improve the stability of scheduling strategies and engineering execution capabilities.
It significantly improves the overall perception capability of reservoir scheduling in response to complex water inflow scenarios, enhances the stability and reliability of scheduling strategies under multi-objective constraints, realizes the direct transformation of intelligent scheduling decision results to the engineering execution level, and improves the engineering feasibility and application value of reservoir scheduling optimization technology.
Smart Images

Figure CN121766733A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent scheduling and control technology, and in particular to a reservoir scheduling optimization method and system based on aerospace big data. Background Technology
[0002] Reservoir scheduling is a crucial technological link in watershed flood control, water resource allocation, ecological protection, and hydropower utilization. With the increasing impact of climate change and human activities, reservoir scheduling faces growing uncertainty in inflow and multi-objective constraints. To improve scheduling efficiency, existing technologies have introduced intelligent scheduling methods such as model predictive control, optimization algorithms, and artificial intelligence, and have to some extent achieved automatic learning from historical hydrological data and operational experience. However, existing intelligent scheduling technologies still mainly rely on surface hydrological monitoring data and historical statistical characteristics, lacking sufficient characterization of key elements such as spatial distribution of rainfall, underlying hydrological conditions, and watershed-scale runoff characteristics. This makes it difficult to fully leverage the advantages of aerospace remote sensing data in terms of spatial coverage and temporal continuity. Furthermore, some reinforcement learning-based scheduling methods are prone to problems such as unstable value functions, slow strategy convergence, or fluctuating scheduling results under multi-objective constraints, affecting the reliability of engineering applications. In addition, existing schemes often focus on offline analysis or scheduling suggestion generation, with insufficient coupling with engineering execution links such as gate control, power plant output regulation, and ecological discharge, making it difficult to meet the engineering application requirements for real-time, stable, and executable scheduling decisions under complex hydrological scenarios. Summary of the Invention
[0003] To address the shortcomings of existing intelligent reservoir scheduling technologies, such as insufficient utilization of aerospace information, limited stability of scheduling strategies, and difficulty in directly supporting engineering-level scheduling execution, this invention proposes a reservoir scheduling optimization method and system based on aerospace big data. This method is based on the deep fusion of aerospace remote sensing observation data and ground hydraulic engineering monitoring data, constructing a multi-dimensional fused state expression that comprehensively represents the watershed inflow conditions, underlying surface hydrological status, and hydraulic structure operating conditions. This allows scheduling decisions to move beyond single-point hydrological information and perceive key hydrological characteristics such as rainfall spatial distribution, confluence condition evolution, and risk area changes, thereby improving the overall perception capability of reservoir scheduling in complex inflow scenarios. Furthermore, a reinforcement learning scheduling decision-making mechanism oriented towards multi-objective constraints is introduced, through price... The introduction of satisfactory value constraints and dynamic satisfactory baselines during the value function update process ensures that the scheduling strategy remains within a stable and controllable value range during both the training and application phases. This effectively suppresses the strategy oscillation problem that easily occurs in multi-constraint scheduling scenarios. Combined with multi-model integration and value distillation mechanisms, the decision-making ability for the coordinated optimization of multiple objectives such as water supply, flood control, ecology, and power generation is gradually improved while ensuring strategy stability. Furthermore, the above-mentioned scheduling decision-making mechanism is unified and linked with engineering execution links such as reservoir gate control, power station output regulation, and ecological discharge. This enables the direct conversion of scheduling optimization decisions driven by aerospace big data into engineering control commands, thereby achieving safe, reliable, and engineering-feasible reservoir scheduling optimization operation under complex hydrological scenarios and multiple operational constraints.
[0004] This invention provides a reservoir scheduling optimization system based on aerospace big data, applied to the scheduling and operation scenarios of river basin reservoirs and reservoir groups. The river basin includes reservoir dams with gates, hydraulic structures, and hydropower station units. Rain gauges, flow meters, water level gauges, and video monitoring equipment are deployed in the upstream and downstream channels of the reservoirs. The river basin water conservancy information center deploys a scheduling optimization server cluster and data storage equipment. The system includes:
[0005] The air-space-ground monitoring and acquisition module is deployed in the basin water conservancy information center to collect integrated air-space-ground monitoring data;
[0006] The multi-source data fusion module is executed by the processor in the scheduling and optimization server cluster to construct a multi-dimensional fused state vector to describe the water inflow conditions of the basin and the operating conditions of the reservoir area.
[0007] The scheduling decision model training module is used to construct the state space, action space and multi-objective reward function of the reinforcement learning environment based on multi-dimensional fused state vectors, and to train the scheduling reinforcement learning model by introducing a Q-learning training mechanism based on satisfactory value constraints and multi-model integration.
[0008] The real-time scheduling decision module is used to receive the fusion processing results of integrated air-space-ground monitoring data in real time during actual operation, input the current fusion state vector into the trained scheduling reinforcement learning model, and output the scheduling instructions of each reservoir in the current and foreseeable period.
[0009] The scheduling execution control module is connected to the reservoir monitoring terminals, programmable logic controllers, gate hoists, and turbine generator control systems of each reservoir dam through the scheduling control network. It is used to convert scheduling instructions into gate opening adjustment, power station output setting, and ecological discharge flow control commands and execute them.
[0010] This invention also provides a reservoir scheduling optimization method based on aerospace big data. This method is executed by a scheduling optimization server cluster and specifically includes the following steps:
[0011] Step S1: Construct integrated air-space-ground monitoring data;
[0012] Step S2: Perform unified spatial projection transformation, coordinate system one and time reference alignment processing on the integrated air-space-ground monitoring data to construct a multi-dimensional fusion state vector to characterize the watershed inflow conditions, reservoir operation status and underlying hydrological characteristics.
[0013] Step S3: Based on the multi-dimensional fused state vector, construct the state space of the reinforcement learning scheduling environment. The state space includes the reservoir water level state, reservoir capacity state, inflow state, downstream key control section flow prediction value, and rainfall intensity index, runoff area humidity index, and risk area identification information extracted from the integrated air-space-ground monitoring data for the current scheduling time and several future forecast periods. Define the outflow, gate opening adjustment, power generation output adjustment, and target reservoir water level of each reservoir in each scheduling period as the action space of the reinforcement learning scheduling environment. Define the maximization of water supply guarantee rate, minimization of flood control risk, maximization of ecological flow compliance rate, and maximization of power generation benefits as joint optimization objectives, and construct a multi-objective reward function for the reinforcement learning scheduling environment accordingly. Set penalty terms for scheduling behaviors that violate reservoir capacity limits, flood control safety constraints, and ecological flow constraints, so that the multi-objective reward function simultaneously constrains reservoir capacity limits, flood control safety constraints, and ecological flow constraints, thereby forming a scheduling optimization reinforcement learning environment for subsequent scheduling strategy training.
[0014] Step S4: Establish a scheduling reinforcement learning model. In the scheduling optimization reinforcement learning environment, acquire historical integrated air-space-ground monitoring data and historical scheduling records. According to the definitions of state space and action space, map the historical scheduling process into time-series samples of states, actions, and corresponding reward results to construct a global experience replay library for reinforcement learning training. Based on the global experience replay library, sample the state-action-result sequence according to the scheduling time order. Introduce a Q-learning training mechanism based on satisfactory value constraints and multi-model integration. Apply dynamic satisfactory baseline constraints to the update process of the scheduling value function used to characterize the long-term scheduling benefits of state-action in the scheduling reinforcement learning model. Train the scheduling reinforcement learning model using parallel training of scheduling weak learners and value integration to obtain the trained scheduling reinforcement learning model. The scheduling reinforcement learning model includes a scheduling weak learner stage and a scheduling reinforcement learning main model stage during the training process. The scheduling weak learner stage is used to perform stable scheduling value learning during the training stage, and the scheduling reinforcement learning main model stage is used to integrate, distill, and form the final scheduling policy output.
[0015] Step S5: Online Scheduling Decision and Command Issuance: Real-time reception of integrated air-space-ground monitoring data; construction of a multi-dimensional fused state vector for the current scheduling moment according to the definition of state space; input of the multi-dimensional fused state vector for the current scheduling moment into the trained scheduling reinforcement learning model to generate scheduling decision results for each reservoir within the current scheduling period; generation of corresponding scheduling control commands based on the scheduling decision results; scheduling control commands include outflow setpoints, gate opening adjustment amounts, power generation output setpoints, and ecological discharge flow setpoints; transmission of scheduling control commands to monitoring terminals and programmable logic controllers (PLCs) deployed at each reservoir dam via the scheduling control network; and control operations performed by the PLCs on the floodgate opening and closing mechanisms, turbine generator units, and ecological discharge facilities, thereby achieving optimized reservoir scheduling operation driven by air-space big data.
[0016] Furthermore, a Q-learning training mechanism based on satisfactory value constraints and multi-model ensemble is introduced to train the scheduling reinforcement learning model. The process of obtaining the trained scheduling reinforcement learning model includes the following steps:
[0017] Step S41: To address the issues of value amplification and unstable propagation in the Bellman update process of the scheduling value function, a satisfactory value constraint mechanism is introduced to impose structural restrictions on the scheduling value function update process, thereby suppressing the unbounded diffusion of scheduling value during the value propagation stage.
[0018] Step S42: Based on the historical scheduling records formed by the scheduling reinforcement learning model during the training phase of the scheduling optimization reinforcement learning environment, and combined with the accumulated reward results obtained by the multi-objective reward function in multiple scheduling periods, on the basis of the satisfaction-based value constraint mechanism, a scheduling satisfaction baseline function is constructed to characterize the achievable scheduling benefit level under the current training phase, and the scheduling satisfaction baseline function is used as the value constraint reference when performing Bellman update of the scheduling value function.
[0019] Step S43: Based on steps S41 and S42, a value regularization term with a satisfaction interval constraint is introduced to restrict the output of the scheduling value function to the satisfaction interval determined by the scheduling satisfaction baseline function and its corresponding preset margin parameter; by applying a penalty constraint to the value estimate that exceeds the satisfaction interval, the drastic fluctuation of the scheduling value function between different scheduling states is reduced, the gradient noise and numerical instability in the scheduling reinforcement learning training process are reduced, and a scheduling value loss function containing a satisfaction interval constraint term is constructed.
[0020] The scheduling value loss function is composed of a satisfactory temporal difference error term and a satisfactory interval constraint regularization term. The temporal difference error term is used to make the scheduling value function fit the satisfactory target value, and the satisfactory interval constraint regularization term limits the fluctuation range of the scheduling value estimation through an interval penalty mechanism based on the scheduling satisfactory baseline function and margin parameter, thereby improving the numerical stability and convergence robustness of the training process.
[0021] Step S44: Under the constraint of the scheduling value loss function, construct multiple lightweight scheduling value networks as scheduling weak learners; based on the global experience replay library, configure an independent scheduling experience replay library for each scheduling weak learner; based on the independent scheduling experience replay library, sample scheduling experience sample sets from their corresponding scheduling experience replay libraries for each scheduling weak learner to learn the scheduling value function on the state-action-result sequence corresponding to the scheduling experience sample set; perform Q-learning training based on satisfactory value constraints in parallel for each scheduling weak learner to form the scheduling value estimation results of multiple scheduling weak learners, and perform integrated processing to obtain an integrated scheduling value for representing the scheduling state-action pair, so as to reduce the sensitivity of the single scheduling model to noisy scheduling samples or extreme scheduling paths and improve the consistency and robustness of scheduling value estimation; Q-learning training based on satisfactory value constraints is achieved by minimizing the scheduling value loss function.
[0022] Step S45: In the scheduling reinforcement learning master model stage, a scheduling reinforcement learning master model with a capacity higher than any single weak scheduling learner is constructed. The integrated scheduling value is used as a supervision signal to perform distillation training on the scheduling reinforcement learning master model, so that it inherits the stable scheduling value structure formed by multiple weak scheduling learners. After completing the distillation training, the scheduling reinforcement learning master model is further refined by scheduling policy training, so that the scheduling policy gradually releases optimization capabilities and improves the final scheduling optimization performance under multi-objective constraints while maintaining training stability. Through the above distillation training and policy refinement process, the trained scheduling reinforcement learning model is obtained.
[0023] Furthermore, the multidimensional fusion state vector includes the following state features: 1. Meteorological state features used to characterize the spatial distribution characteristics of rainfall, including gridded rainfall intensity features, radar reflectivity features, and rainband movement speed features, used to characterize the spatial distribution pattern and temporal evolution characteristics of rainfall processes within the watershed; 2. Hydrological underlying surface state features used to characterize the underlying surface conditions of the watershed, including vegetation index features, soil moisture index features, and snow cover area ratio features, used to reflect the impact of rainfall runoff conditions and snowmelt processes on water inflow formation; 3. Hydrodynamic state features used to characterize the operating conditions of reservoirs and rivers, including current reservoir water level, current reservoir capacity, inflow, outflow, and the flow and water level of downstream key control sections, used to characterize the reservoir's regulation and storage status and downstream flooding conditions; 4. Engineering state features used to characterize the operating status of hydraulic structures, including gate opening status, unit output status, sediment discharge facility operation status, and ecological discharge facility operation status, used to reflect the reservoir's scheduling execution capacity and engineering constraints.
[0024] By introducing the aforementioned multi-category state features, the multi-dimensional fusion state vector can fully express the advantages of aerospace big data in terms of spatial coverage, temporal resolution, and information dimension in a reinforcement learning scheduling environment, thereby enhancing the reservoir scheduling optimization strategy's ability to perceive complex hydrological scenarios and engineering constraints.
[0025] By adopting the above solution, the beneficial effects achieved by the present invention are as follows:
[0026] First, this invention effectively realizes the engineering application of aerospace big data in reservoir scheduling decision-making, significantly improving the overall perception capability of reservoir scheduling in complex water inflow scenarios. By unifying aerospace remote sensing observation data with ground-based hydraulic monitoring data in a spatiotemporal manner, this invention constructs a multi-dimensional fusion state expression that simultaneously reflects the spatial distribution of rainfall, the hydrological state of the underlying surface, and the operational status of the reservoir. This achieves a comprehensive characterization of the formation process and risk evolution characteristics of watershed inflow. Consequently, it enhances the reservoir scheduling's ability to perceive the spatial heterogeneity of heavy rainfall, changes in confluence conditions, and potential flood risk areas in advance. It solves the problem that existing intelligent scheduling technologies mainly rely on point-based hydrological data and lack sufficient understanding of watershed-scale hydrological processes. This strengthens the scientific rigor and foresight of reservoir scheduling decisions under extreme rainfall and complex hydrological scenarios, providing a more reliable state basis for the safe scheduling of reservoirs and reservoir groups.
[0027] Secondly, this invention effectively improves the stability and reliability of intelligent scheduling strategies under multi-objective constraints, ensuring the safe and controllable operation of reservoir scheduling optimization. By introducing a satisfactory value constraint and a dynamic satisfactory baseline mechanism into the reinforcement learning scheduling model, this invention ensures that the scheduling value function remains within a reasonable satisfactory range throughout the training and decision-making process, effectively suppressing unbounded amplification of scheduling value and severe strategy oscillations. This enhances the training stability and convergence robustness of the scheduling strategy under multi-objective collaborative constraints such as flood control, water supply, ecology, and power generation. It solves the problem of large decision fluctuations and unstable results that are common in existing reinforcement learning scheduling methods in complex reservoir scheduling scenarios, thereby strengthening the credibility and engineering application safety of intelligent scheduling strategies under long-term operation and extreme conditions.
[0028] Third, this invention enables the direct transformation of intelligent scheduling decision-making results into engineering execution, significantly enhancing the engineering feasibility and application value of reservoir scheduling optimization technology. This invention integrates scheduling decision-making results generated based on aerospace big data and stable reinforcement learning models with engineering execution processes such as reservoir gate control, power station output regulation, and ecological discharge control, achieving a unified modeling and coordinated control. This realizes the direct output of scheduling strategies from "intelligent decision-making" to "executable scheduling instructions." Consequently, it improves the response efficiency and execution consistency of the reservoir scheduling system in actual operation, solving the problem that existing intelligent scheduling technologies often remain at the level of offline analysis or decision suggestions, making it difficult to directly support on-site scheduling control. This enhances the practicality, scalability, and comprehensive engineering benefits of this invention in the scheduling and operation of river basin reservoirs and reservoir groups. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating the reservoir scheduling optimization method based on aerospace big data proposed in this invention.
[0030] Figure 2This is the convergence curve of the scheduling value loss function proposed in Example 2.
[0031] Figure 2 In the figure, the horizontal axis represents the number of training steps, and the vertical axis represents the scheduling value loss function value. As can be seen from the figure, the loss function gradually decreases in the early stage of training and maintains a stable small fluctuation in the later stage, indicating that the satisfaction value constraint and satisfaction interval regularization mechanism can improve the numerical stability and convergence robustness of the scheduling reinforcement learning training process. Detailed Implementation
[0032] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0033] Example 1, according to Figure 1 This invention provides a reservoir scheduling optimization system based on aerospace big data, applied to the scheduling and operation scenarios of river basin reservoirs and reservoir groups. The river basin includes reservoir dams with gates, hydraulic structures, and hydropower station units. Rain gauges, flow meters, water level gauges, and video monitoring equipment are deployed in the upstream and downstream channels of the reservoirs. The river basin water conservancy information center deploys a scheduling optimization server cluster and data storage equipment. The system includes:
[0034] The air-space-ground monitoring and acquisition module is deployed in the basin water conservancy information center to collect integrated air-space-ground monitoring data;
[0035] The multi-source data fusion module is executed by the processor in the scheduling and optimization server cluster to construct a multi-dimensional fused state vector to describe the water inflow conditions of the basin and the operating conditions of the reservoir area.
[0036] The scheduling decision model training module is used to construct the state space, action space and multi-objective reward function of the reinforcement learning environment based on multi-dimensional fused state vectors, and to train the scheduling reinforcement learning model by introducing a Q-learning training mechanism based on satisfactory value constraints and multi-model integration.
[0037] The real-time scheduling decision module is used to receive the fusion processing results of integrated air-space-ground monitoring data in real time during actual operation, input the current fusion state vector into the trained scheduling reinforcement learning model, and output the scheduling instructions of each reservoir in the current and foreseeable period.
[0038] The scheduling execution control module is connected to the reservoir monitoring terminals, programmable logic controllers, gate hoists, and turbine generator control systems of each reservoir dam through the scheduling control network. It is used to convert scheduling instructions into gate opening adjustment, power station output setting, and ecological discharge flow control commands and execute them.
[0039] This invention also provides a reservoir scheduling optimization method based on aerospace big data. This method is executed by a scheduling optimization server cluster and specifically includes the following steps:
[0040] Step S1: Acquisition of Space-Air Big Data: By accessing remote sensing data products from optical remote sensing satellites, synthetic aperture radar satellites, and meteorological satellites, as well as airborne / UAV aerial photography observation data, information on the spatial distribution of rainfall, snow cover, surface soil moisture distribution, vegetation cover distribution, reservoir water surface boundary changes, and the topography and confluence characteristics of the upstream runoff area are acquired to form space-air observation data reflecting the water inflow conditions and underlying surface conditions of the basin. Simultaneously, rainfall data from rain gauges deployed within the basin, river flow data from flow stations, reservoir water level data from water level gauges, water quality parameter data from water quality monitoring stations, and reservoir capacity data, inflow data, outflow data, gate opening data, and turbine generator output data from status sensors installed on hydraulic structures in the reservoir area are collected to form ground and hydraulic monitoring data reflecting the reservoir's operating conditions. The space-air observation data and the ground and hydraulic monitoring data are then integrated to construct an integrated space-air-ground monitoring data system.
[0041] Step S2: Space-Air-Ground Hydraulic Data Fusion: Perform unified spatial projection transformation, coordinate system one, and time reference alignment on the integrated space-air-ground monitoring data. Spatial resampling and temporal interpolation are performed on the three types of spatial raster data (rainfall spatial distribution field, evapotranspiration intensity distribution field, and snow cover distribution field) obtained from space-air observations, along with reservoir water level data, reservoir capacity data, inflow data, outflow data, and river flow data obtained from ground and hydraulic monitoring. This generates data under a unified spatial grid and time step to characterize the watershed inflow conditions, reservoir operating conditions, and underlying surface. A spatiotemporally consistent hydrological state dataset for surface hydrological conditions; based on the spatiotemporally consistent hydrological state dataset, taking reservoirs and their upstream catchment areas as the analysis objects, according to the preset spatial analysis units and scheduling time steps, the hydrological state features corresponding to the spatial distribution field of rainfall, evapotranspiration intensity, snow cover distribution field, reservoir water level data, reservoir capacity data, inflow data, outflow data, and river flow data are subjected to feature splicing, normalization, and standardization processing to construct a multi-dimensional fusion state vector for characterizing the watershed inflow conditions, reservoir operation status, and underlying surface hydrological characteristics;
[0042] Step S3: Construction of the Reinforcement Learning Environment for Scheduling: Based on the multi-dimensional fused state vector, a state space for the reinforcement learning scheduling environment is constructed. The state space includes the reservoir water level, reservoir capacity, inflow, and downstream key control section flow prediction values for the current scheduling time and several future forecast periods, as well as rainfall intensity, runoff area humidity, and risk area identification information extracted from integrated air-space-ground monitoring data. The outflow, gate opening adjustment, power generation output adjustment, and target reservoir water level of each reservoir in each scheduling period are defined as the action space of the reinforcement learning scheduling environment. Maximizing water supply security rate, minimizing flood control risk, maximizing ecological flow compliance rate, and maximizing power generation benefits are defined as joint optimization objectives. Based on this, a multi-objective reward function for the reinforcement learning scheduling environment is constructed, and penalty terms are set for scheduling behaviors that violate reservoir capacity limits, flood control safety constraints, and ecological flow constraints. This allows the multi-objective reward function to simultaneously constrain reservoir capacity limits, flood control safety constraints, and ecological flow constraints, thereby forming a scheduling optimization reinforcement learning environment for subsequent scheduling strategy training.
[0043] Step S4: Training the Scheduling Reinforcement Learning Model: A scheduling reinforcement learning model is established. In the scheduling optimization reinforcement learning environment, historical integrated air-space-ground monitoring data and historical scheduling records are acquired. According to the definitions of state space and action space, the historical scheduling process is mapped to time-series samples of states, actions, and corresponding reward results to construct a global experience replay library for reinforcement learning training. Based on the global experience replay library, state-action-result sequences are sampled according to scheduling time order. A Q-learning training mechanism based on satisfactory value constraints and multi-model integration is introduced. Dynamic satisfactory baseline constraints are applied to the update process of the scheduling value function, which characterizes the long-term scheduling benefits of state-action, in the scheduling reinforcement learning model. Parallel training of the scheduling weak learner and value integration are used to train the scheduling reinforcement learning model, resulting in the trained scheduling reinforcement learning model. The training process of the scheduling reinforcement learning model includes a scheduling weak learner stage and a scheduling reinforcement learning main model stage. The scheduling weak learner stage is used to perform stable scheduling value learning during the training phase, while the scheduling reinforcement learning main model stage is used to integrate, distill, and form the final scheduling policy output.
[0044] In this invention, the scheduling value function is used to characterize the long-term effects of scheduling actions during reservoir scheduling. It reflects the comprehensive impact of different scheduling decisions on water supply security, flood control safety, ecological flow target achievement, and hydropower generation efficiency operation targets in multiple scheduling periods in the future by discounting and accumulating the immediate scheduling rewards.
[0045] Step S5: Online Scheduling Decision and Command Issuance: Real-time reception of integrated air-space-ground monitoring data; construction of a multi-dimensional fused state vector for the current scheduling moment according to the definition of state space; input of the multi-dimensional fused state vector for the current scheduling moment into the trained scheduling reinforcement learning model to generate scheduling decision results for each reservoir within the current scheduling period; generation of corresponding scheduling control commands based on the scheduling decision results; scheduling control commands include outflow setpoints, gate opening adjustment amounts, power generation output setpoints, and ecological discharge flow setpoints; transmission of scheduling control commands to monitoring terminals and programmable logic controllers (PLCs) deployed at each reservoir dam via the scheduling control network; and control operations performed by the PLCs on the floodgate opening and closing mechanisms, turbine generator units, and ecological discharge facilities, thereby achieving optimized reservoir scheduling operation driven by air-space big data.
[0046] Example 2, according to Figure 2 This embodiment is based on Embodiment 1. In this embodiment, a Q-learning training mechanism based on satisfactory value constraints and multi-model ensemble is introduced to train the scheduling reinforcement learning model. The process of obtaining the trained scheduling reinforcement learning model specifically includes the following steps:
[0047] Step S41: To address the issues of value amplification and unstable propagation in the Bellman update process of the scheduling value function, a satisfactory value constraint mechanism is introduced to structurally restrict the scheduling value function update process, thereby suppressing the unbounded diffusion of scheduling value during the value propagation stage. Specifically, by setting a dynamic upper bound constraint based on the scheduling satisfaction baseline for the state-action value estimation at the next scheduling moment when updating the scheduling value function, the growth rate of the scheduling value function in the current training stage is limited, thereby suppressing the unbounded diffusion of the value function in the early stage of scheduling reinforcement learning training and improving the numerical stability and convergence robustness of the scheduling reinforcement learning model training process.
[0048] The corresponding formula for updating the satisfactory scheduling value in the satisfactory value constraint mechanism is:
[0049] ;
[0050] in, The output value of the scheduling value function (Q function) is used to characterize the scheduling state. The following dispatching action was taken. Then, the estimated value of the cumulative scheduling revenue that can be obtained in the long term with discounts; Indicates the update assignment operator; This indicates an immediate reward (one-step reward). Indicates the discount factor. Indicates the state at the next scheduling moment. Indicates the next state The following candidate actions, Indicates the next scheduling state Below, for all optional scheduling actions Calculate its corresponding scheduling value function And select the maximum value from them as the optimal scheduling value estimate for the next state; Represents the scheduling satisfaction baseline function, in the scheduling state The baseline value below; Indicates the margin parameter; This represents a satisfactory clipping operator;
[0051] Step S42: Based on the historical scheduling records formed during the training phase of the scheduling reinforcement learning model in the scheduling optimization reinforcement learning environment, and combined with the accumulated reward results obtained by the multi-objective reward function over multiple scheduling periods, a scheduling satisfaction baseline function is constructed to characterize the achievable scheduling benefit level under the current training phase, based on the satisfaction-based value constraint mechanism. The scheduling satisfaction baseline function is then used as a value constraint reference when performing Bellman updates on the scheduling value function, enabling the scheduling value constraint to be dynamically adjusted as the scheduling policy capability improves. Specifically, the scheduling satisfaction baseline function is dynamically updated during the scheduling reinforcement learning training process, allowing the scheduling value constraint threshold to adaptively adjust as the scheduling policy capability improves, thereby ensuring training stability while avoiding overly conservative fixed restrictions on the scheduling value function.
[0052] The formula for the scheduling satisfaction baseline function is as follows:
[0053] ;
[0054] in, Represents the scheduling satisfaction baseline function, in the scheduling state The baseline value below, This indicates that in a scheduling optimization reinforcement learning environment, from the state... The cumulative scheduling reward obtained after departure and execution according to the current scheduling strategy; Expressing conditional expectation;
[0055] Step S43: Based on steps S41 and S42, a value regularization term with a satisfactory interval constraint is introduced to restrict the output of the scheduling value function to the satisfactory interval (also serving as a stable interval to ensure training stability) determined by the scheduling satisfactory baseline function and its corresponding preset margin parameter. By imposing a penalty constraint on value estimates that exceed the satisfactory interval, the drastic fluctuations of the scheduling value function between different scheduling states are reduced, thereby reducing gradient noise and numerical instability in the scheduling reinforcement learning training process. A scheduling value loss function containing a satisfactory interval constraint term is constructed, so that the learning process of the scheduling value function is simultaneously affected by the temporal difference error and the stability constraint of the satisfactory interval. By imposing a regularization constraint on value estimates that deviate from the satisfactory interval, the drastic fluctuations of the scheduling value function between different scheduling states are reduced, further improving the numerical stability of the training process.
[0056] The formula for the scheduling value loss function is:
[0057] ;
[0058] in, This represents the value of the scheduling value loss function. This represents the satisfactory temporal difference objective value, used to replace the unconstrained objective value in traditional Q-learning. This represents the squared loss of the timing difference error; This represents the weight coefficient of the regularization term. Represents a linear rectified function; Represents the value regularization term of the satisfaction interval constraint;
[0059] The scheduling value loss function is composed of a satisfactory temporal difference error term and a satisfactory interval constraint regularization term, wherein the temporal difference error term is used to make the scheduling value function fit the satisfactory target value. The satisfaction interval constraint regularization term is obtained by using the scheduling satisfaction baseline function. and margin parameters The interval penalty mechanism limits the fluctuation range of scheduling value estimation, thereby improving the numerical stability and convergence robustness of the training process.
[0060] Step S44: Under the constraint of the scheduling value loss function, construct multiple lightweight scheduling value networks as scheduling weak learners; based on the global experience replay library, configure an independent scheduling experience replay library for each scheduling weak learner; based on the independent scheduling experience replay library, sample scheduling experience sample sets from their corresponding scheduling experience replay libraries for each scheduling weak learner to learn the scheduling value function on the state-action-result sequence corresponding to the scheduling experience sample set; perform Q-learning training based on satisfactory value constraints in parallel for each scheduling weak learner to form the scheduling value estimation results of multiple scheduling weak learners, and perform integrated processing to obtain an integrated scheduling value for representing the scheduling state-action pair, so as to reduce the sensitivity of the single scheduling model to noisy scheduling samples or extreme scheduling paths and improve the consistency and robustness of scheduling value estimation; Q-learning training based on satisfactory value constraints is achieved by minimizing the scheduling value loss function.
[0061] The integrated scheduling value is calculated using an integrated averaging method in the integrated processing. The calculation formula is as follows:
[0062] ;
[0063] in, This indicates the value of integrated scheduling, which is the value in the scheduling state. The following dispatching action was taken. At that time, the stable scheduling value estimate is obtained by combining the outputs of multiple weak scheduling learners; This indicates the number of weak learners to schedule. Indicates the index number of the weak learner. Indicates the first The scheduling value estimate of the output of a weak scheduling learner.
[0064] Step S45: In the scheduling reinforcement learning master model stage, a scheduling reinforcement learning master model with a capacity higher than any single weak scheduling learner is constructed. The integrated scheduling value is used as a supervision signal to perform distillation training on the scheduling reinforcement learning master model, so that it inherits the stable scheduling value structure formed by multiple weak scheduling learners. After completing the distillation training, the scheduling reinforcement learning master model is further refined by scheduling policy training, so that the scheduling policy gradually releases optimization capabilities and improves the final scheduling optimization performance under multi-objective constraints while maintaining training stability. Through the above distillation training and policy refinement process, a trained scheduling reinforcement learning model is obtained, thereby reducing the number of models while maintaining the stable perception capability of the scheduling policy for complex water inflow scenarios and multi-objective constraints.
[0065] In conventional techniques, the process of introducing a Q-learning training mechanism to train a scheduling reinforcement learning model and obtaining the trained scheduling reinforcement learning model includes the following steps:
[0066] Step C1: Construction of the scheduling reinforcement learning environment: Based on the integrated air-space-ground monitoring data, a scheduling reinforcement learning environment is constructed, defining the scheduling state space, scheduling action space, and instant reward function;
[0067] Step C2: Initialize the scheduling value function: Initialize the scheduling value function to represent the estimated scheduling value corresponding to taking different scheduling actions under different scheduling states;
[0068] Step C3: Scheduling value update based on Bellman equation: During the scheduling reinforcement learning training process, the scheduling value function is iteratively updated based on the current scheduling state, scheduling action and the obtained immediate reward, according to the Bellman update rule learned by standard Q.
[0069] Step C4: Model parameter update based on temporal difference error: Using the temporal difference error of the scheduling value function as the training objective, the parameters of the scheduling reinforcement learning model are updated by minimizing the error between the scheduling value estimate and the Bellman objective value, so as to gradually improve the accuracy of scheduling value estimation.
[0070] Step C5: Iterative training to obtain the scheduling reinforcement learning model: Repeat steps C3 and C4 until the scheduling value function converges to obtain the trained scheduling reinforcement learning model.
[0071] Example 3, based on Example 2, includes the following state features in the multidimensional fusion state vector: 1. Meteorological state features describing the spatial distribution characteristics of rainfall, including gridded rainfall intensity features, radar reflectivity features, and rainband movement speed features, used to characterize the spatial distribution pattern and temporal evolution of rainfall processes within the watershed; 2. Hydrological underlying surface state features describing the underlying surface conditions of the watershed, including vegetation index features, soil moisture index features, and snow cover area ratio features, used to reflect the impact of rainfall runoff conditions and snowmelt processes on water inflow formation; 3. Hydrodynamic state features describing the operating conditions of reservoirs and rivers, including current reservoir water level, current reservoir capacity, inflow, outflow, and the flow and water level of downstream key control sections, used to characterize the reservoir's regulation and storage status and downstream flooding conditions; 4. Engineering state features describing the operating status of hydraulic structures, including gate opening status, unit output status, sediment discharge facility operation status, and ecological discharge facility operation status, used to reflect the reservoir's scheduling execution capacity and engineering constraints.
[0072] By introducing the aforementioned multi-category state features, the multi-dimensional fusion state vector can fully express the advantages of aerospace big data in terms of spatial coverage, temporal resolution, and information dimension in a reinforcement learning scheduling environment, thereby enhancing the reservoir scheduling optimization strategy's ability to perceive complex hydrological scenarios and engineering constraints.
[0073] Example 4, based on Example 2, in which the multi-objective reward function constructed in step S3 is:
[0074] ;
[0075] in, Indicates the time of scheduling The corresponding comprehensive reward value is used to measure the overall merits of the current scheduling decision under multiple objective constraints such as water supply, flood control, ecology and power generation, and serves as the evaluation basis for the reinforcement learning model to optimize the strategy. This represents the water supply security evaluation function, used to reflect the current scheduling time. And within the preset future forecast time range, the degree to which the reservoir system meets the water supply needs of various water users, the value of which increases with the increase of the water supply security rate; This represents the flood risk assessment function, which reflects the flood risk level of the reservoir and downstream areas. It comprehensively considers at least the risk that the reservoir water level is close to or exceeds the flood control limit water level, as well as the risk that the water level at the downstream key control section exceeds the warning water level. The value of this function increases with the increase of flood risk. This represents the ecological flow compliance evaluation function, which reflects the degree to which the actual downstream discharge flow at the downstream ecological control section meets the ecological flow requirements. Its value increases as the degree of ecological flow compliance increases. This represents the power generation benefit evaluation function, which reflects the power generation revenue level of the reservoir hydropower station under the current dispatch conditions. Its value increases with the increase of hydropower generation or power generation economic benefits. The constraint violation penalty function is used to reflect the degree of violation of reservoir operation constraints by scheduling decisions. It includes at least the degree of violation of upper and lower limits of reservoir capacity, flood control safety constraints, and ecological flow constraints. When scheduling behavior violates the above constraints, the value of the function increases to impose penalties on unsafe or non-compliant scheduling decisions. , , , , The weight coefficients corresponding to each evaluation function are used to characterize the relative importance of water supply targets, flood control targets, ecological targets, power generation targets, and constraint violation penalties in the comprehensive reward function.
[0076] Example 4, based on Example 2, in this example, step S5: online scheduling decision and instruction issuance: real-time reception of integrated air-space-ground monitoring data, construction of a multi-dimensional fusion state vector at the current scheduling moment according to the definition of state space, and input of the multi-dimensional fusion state vector at the current scheduling moment into the trained scheduling reinforcement learning model to generate scheduling decision results for each reservoir during the current scheduling period; generating corresponding scheduling control instructions based on the scheduling decision results; scheduling control instructions include outflow setpoints, gate opening adjustment amounts, power generation output setpoints, and ecological discharge flow setpoints; sending the scheduling control instructions to the monitoring terminals and programmable logic controllers deployed at each reservoir dam through the scheduling control network, and having the programmable logic controllers perform control operations on the floodgate opening and closing mechanisms, turbine generator units, and ecological discharge facilities, thereby realizing optimized reservoir scheduling operation driven by air-space big data.
[0077] This embodiment takes the joint dispatching system consisting of three cascade reservoirs in the X watershed as the application object, including:
[0078] Upstream control reservoir A (primarily for flood control and power generation);
[0079] Midstream regulating reservoir B (primarily for water supply and regulation);
[0080] Downstream ecological control reservoir C (primarily for ecological discharge and flood control);
[0081] The inference calculation is completed by the training and scheduling reinforcement learning model, and the scheduling decision results of each reservoir during the current scheduling period (08:00–09:00) are output as follows:
[0082] Reservoir A scheduling decision results:
[0083] Target outbound flow rate: 1480 m³ / s;
[0084] Target opening degree of floodgates: 32%;
[0085] Target output of generator unit: 118 MW;
[0086] Ecological outflow setpoint: 90 m³ / s;
[0087] Reservoir B scheduling decision results:
[0088] Target outbound flow rate: 980 m³ / s;
[0089] Target opening degree of floodgates: 27%;
[0090] Target output of generator set: 86 MW;
[0091] Ecological outflow setpoint: 75 m³ / s;
[0092] Reservoir C dispatching decision results:
[0093] Target outbound flow rate: 620 m³ / s;
[0094] Target opening degree of floodgates: 18%;
[0095] Target output of generator set: 42 MW;
[0096] Ecological outflow setpoint: 60 m³ / s.
[0097] On-site execution and equipment control (taking Reservoir A as an example):
[0098] (1) Control and execution of flood discharge gates:
[0099] Initial gate opening: 26%;
[0100] Target opening: 32%;
[0101] The PLC controls the gate to adjust smoothly at a rate of 0.07% / s, taking approximately 90 seconds.
[0102] The measured outflow rate increased from 1260 m³ / s to 1480 m³ / s;
[0103] (2) Control and execution of the hydro-generator unit:
[0104] The power generation capacity increased from 104 MW to 118 MW;
[0105] The unit's vibration, voltage, and current are all within the safe operating range.
[0106] (3) Implementation of ecological discharge facility control:
[0107] The ecological outflow rate has been increased from 65 m³ / s to 90 m³ / s;
[0108] The measured flow rate at the downstream ecological section remained stable within ±3% of the set value.
[0109] Scheduling operation results and feedback:
[0110] The water level in Reservoir A dropped from 245.3 m to 245.1 m;
[0111] The predicted water level at key downstream control sections has decreased from 92% to 88% of the warning level.
[0112] Total power generation output increased by approximately 14% compared to before the dispatch.
[0113] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.
Claims
1. A reservoir scheduling optimization method based on aerospace big data, characterized in that, The method includes the following steps: Step S1: Construct integrated air-space-ground monitoring data; Step S2: Perform unified spatial projection transformation, coordinate system 1 and time reference alignment processing on the integrated air-space-ground monitoring data to construct a multi-dimensional fused state vector; Step S3: Based on the multi-dimensional fused state vector, construct the state space of the reinforcement learning scheduling environment; define the action space of the reinforcement learning scheduling environment; construct a multi-objective reward function for the reinforcement learning scheduling environment, and set penalty terms for scheduling behaviors that violate reservoir capacity limits, flood control safety constraints, and ecological flow constraints, thus forming a scheduling optimization reinforcement learning environment; Step S4: Establish a scheduling reinforcement learning model. In the scheduling optimization reinforcement learning environment, acquire historical integrated air-space-ground monitoring data and historical scheduling records. According to the definition of state space and action space, construct a global experience replay library. Based on the global experience replay library, introduce a Q-learning training mechanism based on satisfactory value constraints and multi-model integration. By applying dynamic satisfactory baseline constraints to the update process of the scheduling value function in the scheduling reinforcement learning model, and by using parallel training of scheduling weak learners and value integration, train the scheduling reinforcement learning model to obtain the trained scheduling reinforcement learning model. Step S5: Generate scheduling decision results for each reservoir during the current scheduling period by training the scheduling reinforcement learning model; generate corresponding scheduling control instructions based on the scheduling decision results.
2. The reservoir scheduling optimization method based on aerospace big data according to claim 1, characterized in that: The state space includes the reservoir water level status, reservoir capacity status, inflow status, downstream key control section flow forecast value, and rainfall intensity index, runoff area humidity index, and risk area identification information extracted from the integrated air-space-ground monitoring data at the current scheduling time and the future forecast period.
3. The reservoir scheduling optimization method based on aerospace big data according to claim 1, characterized in that: The action space of the reinforcement learning scheduling environment is defined as follows: the outflow, gate opening adjustment, power generation output adjustment, and target reservoir water level of each reservoir during each scheduling period are defined as the action space of the reinforcement learning scheduling environment.
4. The reservoir scheduling optimization method based on aerospace big data according to claim 1, characterized in that: The construction of a multi-objective reward function for a reinforcement learning scheduling environment is as follows: the goals of maximizing water supply security rate, minimizing flood risk, maximizing ecological flow compliance rate, and maximizing power generation benefits are defined as joint optimization objectives, and a multi-objective reward function for a reinforcement learning scheduling environment is constructed accordingly.
5. The reservoir scheduling optimization method based on aerospace big data according to claim 1, characterized in that: The scheduling reinforcement learning model includes a scheduling weak learner stage and a scheduling reinforcement learning main model stage during the training process. The scheduling weak learner stage is used to perform stable scheduling value learning during the training stage, and the scheduling reinforcement learning main model stage is used to integrate, distill and form the final scheduling policy output.
6. The reservoir scheduling optimization method based on aerospace big data according to claim 5, characterized in that: The process of training a scheduling reinforcement learning model by introducing a Q-learning training mechanism based on satisfactory value constraints and multi-model ensemble, and obtaining the trained scheduling reinforcement learning model, specifically includes the following steps: Step S41: To address the issues of value amplification and unstable propagation in the Bellman update process of the scheduling value function, a satisfactory value constraint mechanism is introduced to impose structural restrictions on the scheduling value function update process, thereby suppressing the unbounded diffusion of scheduling value during the value propagation stage. Step S42: Based on historical scheduling records and combined with the accumulated reward results obtained by the multi-objective reward function in multiple scheduling periods, a scheduling satisfaction baseline function is constructed on the basis of the satisfaction-based value constraint mechanism, and the scheduling satisfaction baseline function is used as the value constraint reference when performing Bellman update of the scheduling value function in step S41. Step S43: Based on steps S41 and S42, a value regularization term with a satisfaction interval constraint is introduced to restrict the output of the scheduling value function to the satisfaction interval determined by the scheduling satisfaction baseline function and its corresponding preset margin parameter; by imposing a penalty constraint on the value estimate that exceeds the satisfaction interval, the drastic fluctuation of the scheduling value function between different scheduling states is reduced, and a scheduling value loss function is constructed. Step S44: Under the constraint of the scheduling value loss function, construct multiple lightweight scheduling value networks as scheduling weak learners; based on the global experience replay library, configure an independent scheduling experience replay library for each scheduling weak learner; based on the independent scheduling experience replay library, sample scheduling experience sample sets from their corresponding scheduling experience replay libraries for each scheduling weak learner, learn the scheduling value function on the state-action-result sequence corresponding to the scheduling experience sample set, perform Q-learning training based on the satisfaction value constraint in parallel for each scheduling weak learner, form the scheduling value estimation results of multiple scheduling weak learners, and perform integrated processing to obtain the integrated scheduling value; Step S45: In the scheduling reinforcement learning master model stage, a scheduling reinforcement learning master model with a capacity higher than any single weak scheduling learner is constructed, and the integrated scheduling value is used as a supervision signal to perform distillation training on the scheduling reinforcement learning master model, so that it inherits the stable scheduling value structure formed by multiple weak scheduling learners; after completing the distillation training, further scheduling policy refinement training is performed on the scheduling reinforcement learning master model; through the above distillation training and policy refinement process, the trained scheduling reinforcement learning model is obtained.
7. The reservoir scheduling optimization method based on aerospace big data according to claim 6, characterized in that: The scheduling value loss function is composed of a satisfactory temporal difference error term and a satisfactory interval constraint regularization term. The temporal difference error term is used to fit the satisfactory target value to the scheduling value function, while the satisfactory interval constraint regularization term limits the fluctuation range of the scheduling value estimate through an interval penalty mechanism based on the scheduling satisfactory baseline function and margin parameters.
8. The reservoir scheduling optimization method based on aerospace big data according to claim 6, characterized in that: The Q-learning training based on the satisfaction-based value constraint is achieved by minimizing the scheduling value loss function.
9. A reservoir scheduling optimization system based on aerospace big data, used to implement the reservoir scheduling optimization method according to any one of claims 1-8, characterized in that: The system includes: The air-space-ground monitoring and acquisition module collects integrated air-space-ground monitoring data. The multi-source data fusion module constructs a multi-dimensional fused state vector based on integrated air-space-ground monitoring data; The scheduling decision model training module constructs a reinforcement learning environment based on multi-dimensional fused state vectors to train the scheduling reinforcement learning model. The real-time scheduling decision module outputs scheduling instructions by training a post-scheduling reinforcement learning model. The scheduling execution control module executes scheduling instructions.
Citation Information
Patent Citations
Cascade reservoir random optimization scheduling method based on deep Q learning
CN110930016A
Reservoir group joint optimization scheduling method based on MADDPG reinforcement learning
CN115952958A
Deep reinforcement learning short-term random optimization scheduling method for wind-light-cascade reservoir
CN116720674A
Reservoir group multi-target intelligent optimization scheduling method based on multi-constraint coupling
CN118798588A
Reservoir gate dispatching method and system
CN120069374A