Small reservoir intelligent flood discharge scheduling method based on reinforcement learning
By applying intelligent flood discharge scheduling methods based on reinforcement learning in small reservoirs, using quantum genetic algorithms and improved reinforcement learning models, the real-time response and multi-objective optimization problems of flood discharge scheduling in the existing technology are solved, and more efficient and reliable flood discharge decisions are achieved.
Patent Information
- Application Number
- CN202510281719.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art has problems such as real-time response capabilities, multi-objective optimization capabilities and complex dynamic environments in flood discharge scheduling of small reservoirs, which is difficult to meet the actual needs of intelligent flood discharge scheduling.
The intelligent flood discharge scheduling method of small reservoirs based on reinforcement learning is adopted. By obtaining real-time reservoir data, environmental state vectors are constructed, and the initial flood discharge scheduling strategy is generated using quantum genetic algorithms, and multiple rounds of iterative training are carried out through improved reinforcement learning model to optimize the flood discharge decision plan.
It significantly improves the efficiency and reliability of flood discharge scheduling in small reservoirs, and can achieve dynamic adaptive scheduling optimization under complex hydrological meteorological conditions, improving the accuracy, comprehensiveness and real-time nature of flood discharge decisions.
Smart Images

Figure CN120197889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reservoirs, and in particular, to an intelligent flood discharge scheduling method for small reservoirs based on reinforcement learning. Background Art
[0002] With the intensification of climate change and the frequent occurrence of extreme weather events, the role of small reservoirs in flood control, irrigation, and ecological maintenance has become increasingly prominent. However, due to their small storage capacity, simple equipment, and insufficient monitoring means, small reservoirs face significant technical challenges in flood discharge scheduling.
[0003] Currently, the flood discharge scheduling of small reservoirs usually adopts a scheduling method based on empirical rules or a simple rule-based control model. Traditional methods rely on the historical experience of reservoir managers or single-objective rules. Although they can meet the basic flood control requirements under normal circumstances, they have obvious deficiencies when facing complex hydrometeorological conditions and multi-objective scheduling requirements.
[0004] In recent years, automated scheduling technologies based on optimization algorithms have gradually been applied to reservoir management. For example, linear programming, genetic algorithms, or scheduling methods based on rule optimization are used. Existing methods introduce the mathematical modeling of scheduling constraints and the optimization solution of objective functions, but there are still many problems in practical applications. On the one hand, these methods rely on pre-set model assumptions for changes in environmental states and lack the ability to learn in a highly uncertain environment. On the other hand, these methods are usually solved offline and it is difficult to achieve real-time scheduling optimization.
[0005] In summary, the existing technologies have obvious deficiencies in real-time response ability, multi-objective optimization ability, and adaptability to complex dynamic environments, and it is difficult to meet the actual needs of intelligent flood discharge scheduling for small reservoirs. Therefore, there is an urgent need for a new method to achieve dynamic adaptive scheduling optimization, improve the accuracy, comprehensiveness, and real-time nature of flood discharge decisions, so as to better cope with complex hydrometeorological conditions and diverse operation objectives. Summary of the Invention
[0006] An object of the present invention is to propose an intelligent flood discharge scheduling method for small reservoirs based on reinforcement learning, which significantly improves the efficiency and reliability of flood discharge scheduling for small reservoirs.
[0007] An intelligent flood discharge scheduling method for small reservoirs based on reinforcement learning according to an embodiment of the present invention includes the following steps:
[0008] S1. Obtain real-time reservoir data to form a complete reservoir data set;
[0009] S2. Construct an environmental state vector required for reinforcement learning based on the reservoir data set;
[0010] S3. Calculate and optimize the environmental state vector using a quantum genetic algorithm to generate an initial flood discharge scheduling strategy;
[0011] S4. Establishing an improved reinforcement learning model based on the environmental state vector and the initial flood discharge scheduling strategy;
[0012] S5. Using the initial flood discharge scheduling strategy as the initial parameters of the improved reinforcement learning model, through multiple rounds of iterative training in the simulation environment, the strategy network in the improved reinforcement learning model is updated to obtain the optimal flood discharge decision-making plan;
[0013] S6. Operate the gates according to the optimal flood discharge decision plan, monitor the water level changes of small reservoirs, the safety status of the downstream basin and the satisfaction of ecological needs after the implementation, and record the monitoring results and actual operation data as feedback information;
[0014] S7. Input the feedback information into the improved reinforcement learning model again, and combine the reservoir data set and the environmental state vector to iteratively update and correct the strategy of the improved reinforcement learning model.
[0015] Optionally, the S1 includes the following steps:
[0016] S11. Obtain real-time hydrological data of small reservoirs and their surrounding areas through sensor networks and remote monitoring equipment. The data includes the reservoir water level H at a certain time t. t , the inflow flow Q at a certain time t t , the rainfall R at a certain time t t and the evaporation E at a certain time t t ;
[0017] S12. Use sensors to monitor the downstream basin status in real time and collect the water level H of the downstream basin at a certain time t. d,t and the flow rate Q of the downstream basin at a certain time t d,t ;
[0018] S13. Obtain the rainfall forecast value R within the future time interval Δt from the meteorological forecast system f,t+Δt and the real-time rainfall R at the current moment t Conduct joint modeling;
[0019] S14. Extract the water level data H collected for the i-th time in the historical dispatch record from the historical data. h,i , the inflow flow data Q collected for the i-th time in the historical dispatch record h,i , the flood discharge operation parameter A of the i-th time in the historical dispatch record h,i and the downstream water level status data H of the i-th time in the historical dispatch record d,h,i , construct a complete historical scheduling record set H;
[0020] S15. Integrate the real-time reservoir water level data, incoming flow data, rainfall data, evaporation data, downstream basin status data, weather forecast data, and the historical scheduling record set H to construct a complete reservoir data set D:
[0021] D = {H t , Q t , R t , E t , H d,t , Q d,t , R f,t+Δt , H}.
[0022] Optionally, S2 includes the following steps:
[0023] S21. Extract the key environmental status information of the small reservoir at the current moment t based on the reservoir data set D;
[0024] S22. Construct a storage capacity constraint model and a gate operation characteristic model according to the physical characteristics of the small reservoir. The storage capacity constraint model is based on the relationship between the current reservoir water level and the maximum allowable water level and is defined as:
[0025] C = H t - H max , C ≤ 0;
[0026] where C represents the reservoir water level overrun status;
[0027] The gate operation characteristic model is based on the gate opening and the maximum opening and is defined as:
[0028] G t ∈ [0, G max ;
[0029] where G t represents the actual opening of the current gate;
[0030] S23. According to the flood control safety requirements of the downstream area, using the current downstream water level H d,t and the maximum safety water level H d,max as constraints, define the downstream flood control safety model as:
[0031] S = H d,t - H d,max , S ≤ 0;
[0032] where S represents the water level overrun status of the downstream basin;
[0033] S24. Based on the extracted environmental status information, storage capacity constraint model, gate operation characteristic model, and downstream flood control safety model, define the environmental state vector E t :
[0034] E t = {H t , Q t , R t , E t , H d,t , Q d,t , R f,t+Δt , C, S, G t}}。
[0035] Optionally, S3 includes the following steps:
[0036] S31. Generate population individuals for the intelligent flood discharge scheduling of small reservoirs based on the environmental state vector E t and represent the population individuals in a multi-layer coding form:
[0037] S i = {G, Q, T};
[0038] where G represents the gate opening during the scheduling period, controlling the flood discharge capacity of each period t i to make the flood discharge volume meet the reservoir capacity constraint and flood control target, Q is the flood discharge volume vector indicating the water volume discharged from the reservoir per period, affecting the water level of the reservoir and the flood control safety of the downstream area, and T is the flood discharge period allocation, determining the duration of the flood discharge action, and optimizing it in combination with the dynamic changes of the incoming water flow Q t and the rainfall R t ;
[0039] Set the population size N to the number of candidates for the reservoir scheduling scheme;
[0040] S32. Based on the population individuals S of the multi-level scheduling scheme i , perform adaptive coding using qubits, and define the state of each qubit as:
[0041] ψ i,j = α i,j |0> + β i,j |1>;
[0042] |α i,j | 2 + |β i,j | 2 = 1;
[0043] where ψ i,j represents the quantum state of the population individual S i in the j-th dimension, specifically corresponding to the gate opening, flood discharge volume, or time interval, α i,j and β i,j are the quantum probability amplitudes of the j-th dimension respectively, used to represent the possibilities of different scheduling decisions in this dimension, |α i,j |2 and |β i,j | 2 respectively represent the probabilities of the selection states |0< and |1<, and are used for random sampling within the ranges of the gate opening and flood discharge parameters;
[0044] For the dynamic characteristics of small reservoir operation, adaptively adjust the encoding dimension d of the quantum bits:
[0045]
[0046] where T is the total operation duration, representing the overall operation time of the small reservoir operation plan, and R t is the current rainfall, representing the impact of the current external precipitation on the reservoir flood discharge demand, and Q t is the current inflow rate, representing the real-time hydrological pressure faced by the reservoir, and K is the operation complexity coefficient;
[0047] S33. Construct a dynamic optimization function f(S i , t) in combination with the actual requirements of the small reservoir flood discharge operation to dynamically balance the benefits of flood control, irrigation, and ecological objectives. The dynamic optimization function is defined as:
[0048] f(S i , t) = w1f flood (S i , t) + w2f irrigation (S i , t) + w3f eco (S i , t);
[0049] where w1, w2, and w3 are the dynamic weights of flood control, irrigation, and ecological objectives respectively, adjusted according to the current operation state of the reservoir, f flood (S i , t) is the flood control constraint benefit function, f irrigation (S i , t) is the irrigation demand benefit function, and f eco (S i , t) is the ecological water use guarantee function;
[0050] S34. Adjust the individual state through the quantum rotation gate to make the population converge to the optimal operation strategy:
[0051]
[0052] where Δθ ij is calculated based on the fitness of the current population, and is collaboratively optimized in combination with the storage capacity constraint, flood discharge volume, and gate opening strategy;
[0053] S35. When the fitness of the population meets the preset convergence condition, output the final initial flood discharge operation strategy:
[0054] S opt = {G opt , Q opt , T opt};
[0055] If the convergence condition is not met, return to step S32 for continued iteration until the termination condition is satisfied.
[0056] Optionally, the flood control constraint benefit function f flood (S i , t) is defined as the safety guarantee benefit of the flood discharge scheduling for the current reservoir water level and the downstream water level:
[0057]
[0058] Where H max is the maximum allowable safety water level of the reservoir, H safe is the safety warning water level of the reservoir, H d,max and H d,safe are the maximum safety water level and the safety warning water level of the downstream basin respectively, Q out,t is the flood discharge amount, which directly affects the change of the downstream water level, Q in,t represents the inflow discharge at the current moment;
[0059] The irrigation demand benefit function f irrigation (S i , t) measures the degree to which the scheduling scheme meets the agricultural irrigation demand:
[0060]
[0061] Where η is the water conveyance efficiency, Q irrigation represents the agricultural irrigation demand volume in the current period, Q out,t is the downstream ecological water demand, which is used to measure the minimum flow rate maintained by the river channel ecosystem;
[0062] The ecological water supply guarantee function f eco (S i , t) evaluates the water volume guarantee of the scheduling scheme for the downstream ecosystem:
[0063]
[0064] Where Q eco is the minimum water volume required for the downstream ecosystem to maintain, |Q eco - Q out,t | measures the deviation degree of the ecological water use.
[0065] Optionally, the S4 includes the following steps:
[0066] S41. Initialize the improved reinforcement learning model. Define and initialize its parameters based on the environmental state vector and the initial flood discharge scheduling strategy as follows:
[0067] θ0 = Init(w, b, E t , S opt );
[0068] where θ0 represents the set of initial parameters of the improved reinforcement learning model, including the weight w and the bias b, which are used for the initialization of the reinforcement learning policy network. Init represents the initialization function, and its input parameters are the environmental state vector E t and the initial flood discharge scheduling strategy S opt , and combine the current hydrological conditions and the preliminary optimization results of the small reservoir to generate the basic architecture of the scheduling policy network;
[0069] S42. Define the action space A as the set of gate opening G t , flood discharge period ΔT, and flood discharge volume Q out,t . The dynamic range is calculated as follows:
[0070] G t ∈[G safe , G max ;
[0071]
[0072]
[0073] where G safe represents the minimum opening that ensures no risk to the structure during gate operation. ΔT is the flood discharge period, which is calculated based on the excess part of the current reservoir water level H t and the difference between the inflow discharge Q in,t and the flood discharge volume Q out,t to complete the scheduling action within a safe time range;
[0074] S43. Based on the multi-objective requirements of the small reservoir, construct a reward function R(E t , a t ) that covers flood control benefits, satisfaction of agricultural irrigation needs, ecological water use guarantee, and downstream safety indicators:
[0075] R(E t , a t ) = α1·f flood (E t , a t ) + α2·f irrigation (E t , a t ) + α3·f eco (E t , a t) + α4·
[0076] f safety (E t , a t );
[0077] Wherein, f flood (E t , a t ) measures the contribution of the flood discharge action to the safety of the current reservoir water level, and f irrigation (E t , a t ) represents the satisfaction degree of the current flood discharge volume for the irrigation demand, and f eco (E t , a t ) is the current ecological water use guarantee, and f safety (E t , a t ) measures the impact of the flood discharge action on the downstream water level safety. α1, α2, α3, and α4 represent the weight coefficients of each objective;
[0078] S44. Combine the action space A in step S42, the reward function in step S43, and the environmental state vector, and define the improved reinforcement learning model as a quadruple:
[0079] <ε, A, P, R>;
[0080] Wherein, ε is the environmental state space, corresponding to the current and predicted hydrological information of the small reservoir, A is the action space, corresponding to the gate opening, flood discharge period, and flood discharge volume, P is the environmental state transition probability distribution, which is calculated by combining the water level evolution and downstream response in the reservoir operation process, and R is the reward function, which measures the satisfaction degree of each scheduling action for the multi-objective demand.
[0081] S45. Integrate steps S41 to S44, and based on the environmental state vector and the initial flood discharge scheduling strategy, complete the parameter setting and structure definition of the improved reinforcement learning model, and form a scheduling model that can output flood discharge decisions in real time under multi-objective constraints.
[0082] Optionally, the S5 includes the following steps:
[0083] S51. Based on the quadruple <ε, A, P, R> of the improved reinforcement learning model, model the hydrological change process of the small reservoir in the simulation environment. The simulation environment deduces the water level change, incoming flow change, and downstream basin response at different times {t, t + 1,...} according to the storage capacity constraint conditions, gate operation characteristics, and historical data;
[0084] S52. Use the initial parameter set θ0 and the initial scheduling strategy S opt Load the policy network of the improved reinforcement learning model in the simulation environment
[0085] S53. Run several rounds in the simulation environment. Each round includes the following training process:
[0086] Observe the environmental state vector E at time t t ;
[0087] According to the current policy network Select an action a t ∈ A, where θ k is the network parameter updated after the k - th round of training;
[0088] Execute the action a t After that, the simulation environment updates the reservoir and downstream basin states to E according to the state transition probability P t+1 , and calculates the reward R(E t , a t );
[0089] According to the obtained reward R(E t , a t ) and the subsequent state E t+1 , update the policy network, making the network parameter adjust from θ k to θ k+1 , and the update method adopts gradient - based optimal value function approximation or temporal - difference - based adaptive update;
[0090] S54. After completing the training of multiple rounds, judge the convergence degree of the policy network. If the preset convergence condition is met, output the new flood - release scheduling policy π θ* (a t ∣ E t ) and its corresponding network parameter θ * ; if the convergence condition is not reached, return to step S53 to continue training until a better flood - release scheduling policy under complex hydrometeorological conditions is obtained, and generate the optimal flood - release decision - making scheme:
[0091] S′ opt ={G′ opt , Q′ opt , T′ opt}.
[0092] The beneficial effects of the present invention are:
[0093] (1) The present invention introduces a quantum genetic algorithm in the initialization stage of the reinforcement learning model, and efficiently optimizes the scheduling strategy through qubit encoding combined with rotation gate operations. The traditional random initialization strategy often has a slow convergence speed and unstable initial decisions due to the overly large exploration space. However, the quantum genetic algorithm can quickly locate the global optimal or approximately optimal scheduling scheme under multi-objective constraints. The qubit encoding adaptively adjusts the dimension, enabling the algorithm to dynamically adapt to the complex environment of small reservoir flood discharge. At the same time, the efficient evolution of the population is achieved through the rotation gate update of the probability amplitude, significantly improving the quality of the initial scheduling strategy.
[0094] (2) The present invention constructs a multi-objective reward function covering flood control, irrigation, ecological maintenance, and downstream safety and introduces a dynamic weight adjustment mechanism, enabling the scheduling strategy to optimize the balance between different objectives in real time according to the current hydrological conditions of the reservoir. The reward function can give priority to ensuring the safety of the reservoir capacity during the flood control peak period, dynamically allocate water resources during the peak irrigation demand period, and ensure the minimum requirements for downstream ecological water use. Through the real-time evaluation of the downstream water level change and ecological flow deviation, the algorithm can accurately regulate the flood discharge volume, avoiding the phenomenon of resource waste or increased risk caused by a single target bias in traditional methods.
[0095] (3) The present invention realizes multi-round policy iteration optimization in the simulation environment through the closed-loop training mechanism of the reinforcement learning model, enabling the model to autonomously learn the optimal scheduling scheme under complex and variable hydrometeorological conditions. The reinforcement learning model can not only dynamically adjust the scheduling strategy through historical data and real-time monitoring data, but also improve the decision-making ability in unknown scenarios through the adaptive update of the policy network. It can quickly respond under sudden heavy rainfall or extreme meteorological conditions and output a safe and efficient flood discharge strategy in real time, effectively reducing the risk of reservoir operation. Description of the Drawings
[0096] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation to the present invention. In the drawings:
[0097] Figure 1 is a flowchart of an intelligent flood discharge scheduling method for small reservoirs based on reinforcement learning proposed by the present invention. Detailed Embodiments
[0098] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0099] Refer to Figure 1 , an intelligent flood discharge scheduling method for small reservoirs based on reinforcement learning, includes the following steps:
[0100] S1. Obtain real-time reservoir data to form a complete reservoir dataset;
[0101] S2. Construct the environmental state vector required for reinforcement learning based on the reservoir dataset;
[0102] S3. Use the quantum genetic algorithm to calculate and optimize the environmental state vector to generate an initial flood discharge scheduling strategy;
[0103] S4. Establish an improved reinforcement learning model based on the environmental state vector and the initial flood discharge scheduling strategy;
[0104] S5. Use the initial flood discharge scheduling strategy as the initial parameter of the improved reinforcement learning model, and through multiple rounds of iterative training in the simulation environment, update the policy network in the improved reinforcement learning model to obtain the optimal flood discharge decision-making plan;
[0105] S6. Perform gate operations according to the optimal flood discharge decision-making plan, monitor the water level change of the small reservoir, the safety status of the downstream basin and the satisfaction degree of ecological needs after execution, and record the monitoring results and actual operation data as feedback information;
[0106] S7. Re-input the feedback information into the improved reinforcement learning model, and combine the reservoir dataset and the environmental state vector to perform iterative update and policy correction on the improved reinforcement learning model.
[0107] In this embodiment, S1 includes the following steps:
[0108] S11. Obtain real-time hydrological data of the small reservoir and its surrounding area through the sensor network and remote monitoring equipment. The data includes the reservoir water level H at a certain moment t t , the inflow Q at a certain moment t t , the rainfall R at a certain moment t t and the evaporation E at a certain moment t t ;
[0109] S12. Use sensors to monitor the state of the downstream basin in real time, and collect the water level H of the downstream basin at a certain moment t d,t and the flow Q of the downstream basin at a certain moment t d,t ;
[0110] S13. Obtain the rainfall prediction value R within the future time interval Δt from the meteorological forecasting system f,t+Δt , and perform joint modeling with the real-time rainfall R at the current moment t ;
[0111] S14. Extract the water level data H collected for the i-th time in the historical scheduling records from the historical data h,i , the inflow data Q collected for the i-th time in the historical scheduling recordsh,i , the flood discharge operation parameter A in the i-th historical scheduling record h,i and the downstream water level status data H in the i-th historical scheduling record d,h,i , to construct a complete set of historical scheduling records H;
[0112] S15. Integrate the real-time reservoir water level data, inflow data, rainfall data, evaporation data, downstream basin status data, meteorological forecast data, and the set of historical scheduling records H to construct a complete reservoir dataset D:
[0113] D = {H t , Q t , R t , E t , H d,t , Q d,t , R f,t+Δt , H}.
[0114] In this embodiment, S2 includes the following steps:
[0115] S21. Extract the key environmental status information of the small reservoir at the current moment t based on the reservoir dataset D;
[0116] S22. Construct a storage capacity constraint model and a gate operation characteristic model according to the physical characteristics of the small reservoir. The storage capacity constraint model is based on the relationship between the current reservoir water level and the maximum allowable water level and is defined as:
[0117] C = H t - H max , C ≤ 0;
[0118] where C represents the reservoir water level overlimit status;
[0119] The gate operation characteristic model is based on the gate opening and the maximum opening and is defined as:
[0120] G t ∈ [0, G max ;
[0121] where G t represents the actual opening of the current gate;
[0122] S23. According to the flood control safety requirements of the downstream area, using the current downstream water level H d,t and the maximum safety water level H d,max as constraints, define the downstream flood control safety model as:
[0123] S = H d,t - H d,max , S ≤ 0;
[0124] Among them, S represents the water level over-limit state of the downstream basin;
[0125] S24. Define the environmental state vector E of the reinforcement learning model based on the extracted environmental state information, storage capacity constraint model, gate operation characteristic model, and downstream flood control safety model t :
[0126] E t ={H t , Q t , R t , E t , H d,t , Q d,t , R f,t+Δt , C, S, G t}.
[0127] In this embodiment, S3 includes the following steps:
[0128] S31. Generate population individuals for the intelligent flood discharge scheduling of small reservoirs based on the environmental state vector E t The population individuals are represented in a multi-layer coding form:
[0129] S i ={G, Q, T};
[0130] Among them, G represents the gate opening during the scheduling period, controlling the flood discharge capacity of each period t i to make the flood discharge volume meet the storage capacity constraint and flood control objectives. Q is the flood discharge volume vector, indicating the water volume discharged from the reservoir in each period, affecting the water level of the reservoir and the flood control safety of the downstream area. T is the flood discharge period allocation, determining the duration of the flood discharge action, and optimizing it in combination with the dynamic changes of the incoming water flow Q t and rainfall R t ;
[0131] The population size N is set to the number of candidates for the reservoir scheduling scheme;
[0132] S32. Based on the population individuals S of the multi-level scheduling scheme i , perform adaptive coding using qubits. The state of each qubit is defined as:
[0133] ψ i,j =α i,j |0> + β i,j |1>;
[0134] |α i,j | 2 + |β i,j | 2 = 1;
[0135] Among them, ψ i,j represents the population individual Si The quantum state in the j-th dimension, specifically corresponding to the gate opening, flood discharge volume, or time interval, α i,j and β i,j are respectively the quantum probability amplitudes in the j-th dimension, used to represent the possibilities of different scheduling decisions in this dimension. |α i,j | 2 and |β i,j | 2 respectively represent the probabilities of selecting the states |0> and |1>, and are used for random sampling within the ranges of the gate opening and flood discharge volume parameters;
[0136] For the dynamic characteristics of small reservoir scheduling, adaptively adjust the encoding dimension d of the quantum bits:
[0137]
[0138] where T is the total scheduling duration, representing the overall operation time of the small reservoir scheduling scheme, R t is the current rainfall, representing the impact of the current external precipitation on the reservoir flood discharge demand, Q t is the current inflow, representing the real-time hydrological pressure faced by the reservoir, and K is the scheduling complexity coefficient;
[0139] S33. Construct a dynamic optimization function f(S i , t) in combination with the actual requirements of small reservoir flood discharge scheduling to dynamically balance the benefits of flood control, irrigation, and ecological objectives. The dynamic optimization function is defined as:
[0140] f(S i , t) = w1f flood (S i , t) + w2f irrigation (S i , t) + w3f eco (S i , t);
[0141] where w1, w2, and w3 are respectively the dynamic weights of flood control, irrigation, and ecological objectives, adjusted according to the current operating state of the reservoir. f flood (S i , t) is the flood control constraint benefit function, f irrigation (S i , t) is the irrigation demand benefit function, and f eco (S i , t) is the ecological water use guarantee function;
[0142] S34. Adjust the individual state through the quantum rotation gate to make the population converge to the optimal scheduling strategy:
[0143]
[0144] Among them, Δθ ij Based on the calculation of the fitness of the current population, collaborative optimization is carried out in combination with the reservoir capacity constraint, the flood discharge volume, and the gate opening strategy;
[0145] S35. When the population fitness meets the preset convergence condition, output the final initial flood discharge scheduling strategy:
[0146] S opt ={G opt ,Q opt ,T opt};
[0147] If the convergence condition is not met, return to step S32 to continue the iteration until the termination condition is met.
[0148] In this embodiment, the flood control constraint benefit function f flood (S i ,t) is defined as the safety guarantee benefit of the flood discharge scheduling for the current reservoir water level and the downstream water level:
[0149]
[0150] Among them, H max is the maximum allowable safety water level of the reservoir, H safe is the safety warning water level of the reservoir, H d,max and H d,safe are the maximum safety water level and the safety warning water level of the downstream basin respectively, Q out,t is the flood discharge volume, which directly affects the change of the downstream water level, Q in,t represents the inflow rate at the current moment;
[0151] The irrigation demand benefit function f irrigation (S i ,t) measures the degree to which the scheduling scheme meets the agricultural irrigation demand:
[0152]
[0153] Among them, η is the water conveyance efficiency, Q irrigation represents the agricultural irrigation demand volume in the current period, Q out,t is the downstream ecological water demand volume, which is used to measure the minimum flow rate maintained by the river channel ecosystem;
[0154] The ecological water guarantee function f eco (S i ,t) evaluates the water volume guarantee of the scheduling scheme for the downstream ecosystem:
[0155]
[0156] Among them, Q ecoThe minimum water volume required to maintain the downstream ecosystem, |Q eco -Q out,t | Measures the deviation degree of ecological water use.
[0157] In this embodiment, S4 includes the following steps:
[0158] S41. Initialize the improved reinforcement learning model, and define its parameters based on the environmental state vector and the initial flood discharge scheduling strategy as:
[0159] θ0 = Init(w, b, E t , S opt );
[0160] Among them, θ0 represents the set of initial parameters of the improved reinforcement learning model, including the weight w and the bias b, which are used for the initialization of the reinforcement learning policy network. Init represents the initialization function, and its parameter input is the environmental state vector E t and the initial flood discharge scheduling strategy S opt , and combines the current hydrological conditions and the preliminary optimization results of the small reservoir to generate the basic architecture of the scheduling policy network;
[0161] S42. Define the action space A as the set of gate opening G t , flood discharge period ΔT, and flood discharge volume Q out,t , and the dynamic range is calculated as follows:
[0162] G t ∈[G safe , G max ;
[0163]
[0164] Among them, G safe represents the minimum opening degree to ensure that the gate operation does not pose a risk to the structure. ΔT is the flood discharge period, which is calculated based on the excess part of the current reservoir water level H t and the difference between the inflow rate Q in,t and the flood discharge volume Q out,t to complete the scheduling action within a safe time range;
[0165] S43. Based on the multi-objective requirements of the small reservoir, construct a reward function R(E t , a t ) that covers flood control benefits, satisfaction of agricultural irrigation needs, ecological water use guarantee, and downstream safety indicators:
[0166] R(E t , a t ) = α1·f flood (E t , a t ) + α2·firrigation (E t , a t ) + α3·f eco (E t , a t ) + α4·
[0167] f safety (E t , a t );
[0168] Wherein, f flood (E t , a t ) measures the contribution of the flood discharge action to the safety of the current reservoir water level, f irrigation (E t , a t ) represents the satisfaction degree of the current flood discharge volume to the irrigation demand, f eco (E t , a t ) is the current ecological water use guarantee, f safety (E t , a t ) measures the impact of the flood discharge action on the safety of the downstream water level, and α1, α2, α3, α4 represent the weight coefficients of each objective;
[0169] S44. Combine the action space A in step S42, the reward function in step S43, and the environmental state vector, and define the improved reinforcement learning model as a quadruple:
[0170] <ε, A, P, R>;
[0171] Wherein, ε is the environmental state space, corresponding to the current and predicted hydrological information of the small reservoir, A is the action space, corresponding to the gate opening, flood discharge period, and flood discharge volume, P is the environmental state transition probability distribution, calculated by combining the water level evolution and downstream response in the reservoir operation process, and R is the reward function, measuring the satisfaction degree of each scheduling action to the multi-objective requirements.
[0172] S45. Integrate steps S41 to S44, and based on the environmental state vector and the initial flood discharge scheduling strategy, complete the parameter setting and structure definition of the improved reinforcement learning model, and form a scheduling model that can output flood discharge decisions in real time under multi-objective constraints.
[0173] In this embodiment, S5 includes the following steps:
[0174] S51. Model the hydrological change process of a small reservoir in a simulation environment based on the quadruple <ε, A, P, R> of the improved reinforcement learning model. The simulation environment deduces the water level change, inflow change, and downstream basin response at different times {t, t+1, …} according to the storage capacity constraint conditions, gate operation characteristics, and historical data;
[0175] S52. Use the initial parameter set θ0 and the initial scheduling strategy S opt Load the policy network of the improved reinforcement learning model into the simulation environment
[0176] S53. Run several rounds in the simulation environment. Each round includes the following training process:
[0177] Observe the environmental state vector E at time t t ;
[0178] According to the current policy network Select an action a t ∈A, where θ k is the network parameter updated after the k-th round of training;
[0179] Execute the action a t After that, the simulation environment updates the states of the reservoir and the downstream basin to E t+1 , and calculates the reward R(E t , a t );
[0180] According to the obtained reward R(E t , a t ) and the subsequent state E t+1 , update the policy network, and let the network parameter be adjusted from θ k to θ k+1 , and the update method adopts gradient-based optimal value function approximation or time difference-based adaptive update;
[0181] S54. After completing the training of multiple rounds, judge the convergence degree of the policy network. If the preset convergence condition is met, output the new flood discharge scheduling strategy π θ* (a t ∣E t ) and its corresponding network parameter θ * ; if the convergence condition is not reached, return to step S53 to continue training until a better flood discharge scheduling strategy under complex hydrometeorological conditions is obtained, and generate the optimal flood discharge decision-making scheme:
[0182] S′ opt ={G′ opt , Q′ opt , T′ opt}.
[0183] Example 1:
[0184] In mid-July 2024, a small reservoir "XX Reservoir" in central China encountered rare heavy rainfall. At 8 am on July 15, the local meteorological department issued a red rainstorm warning, predicting that the cumulative rainfall in the next 48 hours would reach 400 mm, and the peak rainfall intensity would be 60 mm per hour. As one of the main flood control facilities, the storage capacity of XX Reservoir is 1 million cubic meters, the initial water level is 16 m, and there is still a 6-m margin from the maximum safety water level of 22 m. The downstream area includes an ecological reserve and 1,500 mu of farmland. The maximum allowable water level of the ecological reserve is 22 m, and the urgent water demand for the farmland is 300,000 cubic meters.
[0185] At 9 am on July 15, through the sensor network deployed in XX Reservoir, the system collected the following key data: the current reservoir water level was 16.5 m, the inflow rate was 20 cubic meters per second, and it was initially predicted that the inflow rate would gradually increase in the next 6 hours and reach a peak of 150 cubic meters per second. The water level in the downstream area was 20 m, and there was still a 2-m margin from the maximum safety water level of 22 m. Based on these real-time monitoring data, the system recorded these key information in the environmental state vector and quickly constructed a complete reservoir state in combination with the rainfall data from the weather forecast.
[0186] Through the quantum genetic algorithm, an initial scheduling strategy was quickly generated: the initial gate opening was set at 30%, the flood discharge period was set at 2 hours, and the initial flood discharge volume was 50 cubic meters per second. Subsequently, the reinforcement learning model took over the scheduling task and entered the real-time decision-making mode. At 9:30 am, the system calculated the optimal scheduling strategy based on the current state and reminded the management personnel with an alarm signal: due to the continuous increase in the inflow rate, the gate opening needed to be adjusted to 50%, and the flood discharge volume was increased to 75 cubic meters per second.
[0187] At 10 am on July 15, the rainfall intensity increased further, and the inflow rate quickly rose to 100 cubic meters per second. The reservoir water level had reached 18 m. At this time, the system detected that the inflow rate would continue to increase in the next 2 hours. Combining the water level status in the downstream area and the farmland irrigation demand, the reinforcement learning model gave the scheduling strategy: increase the gate opening to 60%, increase the flood discharge volume to 100 cubic meters per second, and at the same time remind the management personnel to control the flood discharge duration not to exceed 4 hours to ensure that the water level in the downstream ecological reserve does not exceed 22 m.
[0188] At 14:00 in the afternoon, the heavy rainfall was still continuing. The water level of the reservoir had reached 19.5 meters, and the inflow rate had reached the peak of 150 cubic meters per second. The system monitored in real time that the water level in the downstream protection area had approached 21.5 meters. The model adjusted the strategy based on multi-objective trade-off: reduced the gate opening to 40%, controlled the flood discharge volume at 60 cubic meters per second, and extended the flood discharge period to 6 hours to reduce the rising speed of the downstream water level. At the same time, the model set the flood discharge plan for the next stage according to the farmland irrigation demand and ecological water use guarantee to ensure that the irrigation demand was fully met after the rain stopped.
[0189] After the scheduling process ended, the following data in Table 1 recorded the specific performance of the method of the present invention:
[0190] Table 1 Scheduling process data
[0191]
[0192] From Table 1 above, during the dynamic change process of the water level and the inflow rate, the present invention adjusted the gate opening and the flood discharge volume through the real-time optimization of the reinforcement learning model, so that the reservoir was always in a safe operation state. During the peak stage when the inflow rate was as high as 150 cubic meters per second, the water level of the reservoir was controlled at 19.5 meters (far lower than the maximum safety water level of 22 meters). When the water level of the reservoir approached the critical value (19.5 meters), the flood discharge volume was quickly increased to 100 cubic meters per second to quickly reduce the pressure on the reservoir. When the downstream water level reached the high level of 21.5 meters, the flood discharge volume was adjusted to 60 cubic meters per second to balance the downstream ecological safety and irrigation demand. The farmland irrigation satisfaction rate and the ecological water use guarantee rate gradually increased during the scheduling process and finally reached 100% and 95%.
[0193] Comparing with the fixed-rule scheduling of the traditional method, the following Table 2 shows the performance comparison data of the two methods:
[0194] Table 2 Performance comparison data
[0195] Index Traditional method Method of the present invention <![CDATA[Peak flood discharge (m 3 / s)]]> 50 100 Maximum downstream water level (m) 21.8 21.5 Farmland irrigation satisfaction rate (%) 70 100 Decision response time (s) 300 120
[0196] In critical moments, the present invention increases the peak flood discharge rate to 100 cubic meters per second, which is twice as high as the 50 cubic meters per second of traditional methods. This not only effectively alleviates the water level pressure of the reservoir but also ensures the safety of the downstream area. At the same time, it provides more water resource support for irrigation and ecological needs. The present invention controls the maximum downstream water level at 21.5 meters, which is 0.3 meters lower than that of traditional methods. This difference fully demonstrates that the present invention can accurately predict the downstream water level changes and dynamically adjust the flood discharge strategy, significantly improving the safety of the downstream area. The farmland irrigation satisfaction rate is increased from 70% of traditional methods to 100%, and the ecological water use guarantee rate is increased from 70% to 95%. The decision response time of the present invention is shortened from 300 seconds of traditional methods to 120 seconds, a reduction of 60%. Thanks to the real-time update ability of the reinforcement learning model and the high-quality initial strategy generated by the quantum genetic algorithm, the scheduling efficiency is significantly improved.
[0197] From the implementation of the above scheduling process, it can be seen that the method of the present invention can accurately and real-time adjust the flood discharge strategy under dynamic and complex hydrometeorological conditions, effectively solving the deficiencies of traditional scheduling methods in flood control efficiency, resource utilization, and ecological protection, and greatly improving the comprehensive optimization ability of intelligent flood discharge of small reservoirs.
[0198] In the initialization stage of the reinforcement learning model, the present invention introduces the quantum genetic algorithm, and efficiently optimizes the scheduling strategy through qubit encoding combined with rotation gate operations. Traditional random initialization strategies often have slow convergence speed and unstable initial decisions due to the too large exploration space. The quantum genetic algorithm can quickly locate the global optimal or approximate optimal scheduling scheme under multi-objective constraints. The qubit encoding adaptively adjusts the dimension, enabling the algorithm to dynamically adapt to the complex environment of small reservoir flood discharge. At the same time, the efficient evolution of the population is realized through the rotation gate update of the probability amplitude, greatly improving the quality of the initial scheduling strategy.
[0199] The present invention constructs a multi-objective reward function covering flood control, irrigation, ecological maintenance, and downstream safety and introduces a dynamic weight adjustment mechanism, enabling the scheduling strategy to optimize the balance between different objectives in real-time according to the current hydrological conditions of the reservoir. The reward function can give priority to ensuring the safety of the reservoir capacity during the flood control peak period, dynamically allocate water resources during the peak irrigation demand period, and ensure the minimum requirements for downstream ecological water use. Through the real-time evaluation of the downstream water level changes and ecological flow deviation, the algorithm can accurately regulate the flood discharge volume, avoiding the phenomenon of resource waste or increased risk caused by single-sided objectives in traditional methods.
[0200] The present invention realizes multi-round policy iteration optimization in a simulation environment through the closed-loop training mechanism of the reinforcement learning model, enabling the model to autonomously learn the optimal scheduling plan under complex and variable hydrometeorological conditions. The reinforcement learning model can not only dynamically adjust the scheduling strategy through historical data and real-time monitoring data, but also improve the decision-making ability in unknown scenarios through the adaptive update of the policy network. It can quickly respond under sudden heavy rainfall or extreme meteorological conditions and output safe and efficient flood discharge strategies in real time, effectively reducing the risk of reservoir operation.
[0201] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent replacements or changes should be covered within the protection scope of the present invention.
Claims
1. A small reservoir intelligent flood discharge scheduling method based on reinforcement learning, characterized in that: The steps include: S1. Obtain real-time reservoir data to form a complete reservoir data set; S2. Construct the environment state vector required for reinforcement learning based on the reservoir dataset; S3. Calculate and optimize the environmental state vector using a quantum genetic algorithm to generate an initial flood discharge scheduling strategy; S4. Establishing an improved reinforcement learning model based on the environmental state vector and the initial flood discharge scheduling strategy; S5. Using the initial flood discharge scheduling strategy as the initial parameters of the improved reinforcement learning model, through multiple rounds of iterative training in the simulation environment, the strategy network in the improved reinforcement learning model is updated to obtain the optimal flood discharge decision-making plan; S6. Operate the gates according to the optimal flood discharge decision plan, monitor the water level changes of small reservoirs, the safety status of the downstream basin and the satisfaction of ecological needs after the implementation, and record the monitoring results and actual operation data as feedback information; S7. Input the feedback information into the improved reinforcement learning model again, and combine the reservoir data set and the environmental state vector to iteratively update and correct the strategy of the improved reinforcement learning model.
2. According to claim 1, a small reservoir intelligent flood discharge scheduling method based on reinforcement learning is characterized in that: The S1 comprises the following steps: S11. Obtain real-time hydrological data of small reservoirs and their surrounding areas through sensor networks and remote monitoring equipment. The data includes the reservoir water level H at a certain time t. t , the inflow flow Q at a certain time t t , the rainfall R at a certain time t t and the evaporation E at a certain time t t ; S12. Use sensors to monitor the downstream basin status in real time and collect the water level H of the downstream basin at a certain time t. d,t and the flow rate Q of the downstream basin at a certain time t d,t ; S13. Obtain the rainfall forecast value R within the future time interval Δt from the meteorological forecast system f,t+Δt and the real-time rainfall R at the current moment t Conduct joint modeling; S14. Extract the water level data H collected for the i-th time in the historical dispatch record from the historical data. h,i , the inflow flow data Q collected for the i-th time in the historical dispatch record h,i , the flood discharge operation parameter A of the i-th time in the historical dispatch record h,i and the downstream water level status data H of the i-th time in the historical dispatch record d,h,i , construct a complete historical scheduling record set H; S15. Integrate the real-time reservoir water level data, inflow data, rainfall data, evaporation data, downstream basin status data, weather forecast data, and historical dispatch record set H to construct a complete reservoir data set D: D={H t ,Q t ,R t ,E t ,H d,t ,Q d,t ,R f,t+Δt ,H}。 3. According to the reinforcement learning-based intelligent flood discharge scheduling method for small reservoirs in claim 1, it is characterized in that: The S2 comprises the following steps: S21. Extracting key environmental status information of a small reservoir at the current time t based on the reservoir data set D; S22. A reservoir capacity constraint model and a gate operation characteristic model are constructed based on the physical characteristics of the small reservoir. The reservoir capacity constraint model is based on the relationship between the current reservoir water level and the maximum allowable water level and is defined as: C=H t -H max ,C≤0; Among them, C represents the reservoir water level over-limit state; The gate operation characteristic model is based on the gate opening and maximum opening, and is defined as: G t ∈[0,G max ]; Among them, G t Indicates the actual opening of the current gate; S23. According to the flood control safety requirements of the downstream area, the current downstream water level H d,t and the maximum safe water level H d,max As a constraint, the downstream flood control safety model is defined as: S=H d,t -H d,max ,S≤0; Among them, S represents the water level exceeding limit state in the downstream basin; S24. Based on the extracted environmental state information, reservoir capacity constraint model, gate operation characteristic model and downstream flood control safety model, define the environmental state vector E of the reinforcement learning model t : E t ={H t ,Q t ,R t ,E t ,H d,t ,Q d,t ,R f,t+Δt ,C,S,G t }。 4. According to the reinforcement learning-based intelligent flood discharge scheduling method for small reservoirs in claim 1, it is characterized in that: The S3 comprises the following steps: S31. Based on the environment state vector E t Generate population individuals for intelligent flood discharge scheduling of small reservoirs. The population individuals are represented by multi-layer coding: S i ={G,Q,T}; Among them, G represents the gate opening within the dispatching period, and controls each period t i The flood discharge capacity of the reservoir is determined by the reservoir capacity constraint and flood control objectives. Q is the flood discharge vector, indicating the amount of water discharged from the reservoir in each period, which affects the reservoir water level and the flood control safety of the downstream area. T is the flood discharge period allocation, which determines the duration of the flood discharge action. Combined with the inflow flow Q t and rainfall R t Optimize the dynamic changes of The population size N is set to the number of candidates for reservoir operation schemes; S32. Population individual S based on multi-level scheduling scheme i , quantum bits are used for adaptive encoding, and the state of each quantum bit is defined as: ψ i,j =a i,j |0>+β i,j |1>; |α i,j | 2 +|β i,j | 2 =1; Among them, ψ i,j Represents the population individual S i The quantum state in the jth dimension specifically corresponds to the gate opening, flood discharge or time interval, α i,j and β i,j are the quantum probability amplitudes of the jth dimension, which are used to represent the possibility of different scheduling decisions in this dimension. i,j | 2 and |β i,j | 2 They represent the probability of selecting states |0> and |1>, respectively, and are used for random sampling within the range of gate opening and flood discharge parameters; In view of the dynamic characteristics of small reservoir operation, the encoding dimension d of quantum bits is adaptively adjusted: Among them, T is the total dispatching time, which represents the overall operation time of the small reservoir dispatching plan, and R t is the current rainfall, indicating the impact of current external precipitation on the reservoir flood discharge demand, Q t is the current inflow, indicating the real-time hydrological pressure faced by the reservoir, and K is the scheduling complexity coefficient; S33. Combined with the actual needs of small reservoir flood discharge scheduling, a dynamic optimization function f(S i ,t), in order to dynamically balance the benefits of flood control, irrigation and ecological objectives, the dynamic optimization function is defined as: f(S i ,t)=w1f flood (S i ,t)+w2f irrigation (S i ,t)+w3f eco (S i ,t); Among them, w1, w2, and w3 are the dynamic weights of flood control, irrigation, and ecological goals, which are adjusted according to the current operating status of the reservoir. flood (S i ,t) is the flood control constraint benefit function, f irrigation (S i ,t) is the irrigation demand benefit function, f eco (S i ,t) is the ecological water security function; S34. Adjust the individual state through the quantum rotating gate to make the population converge to the optimal scheduling strategy: Among them, Δθ ij Based on the current population fitness calculation, the reservoir capacity constraint, flood discharge volume and gate opening strategy are combined for collaborative optimization; S35. When the population fitness meets the preset convergence conditions, the final initial flood discharge scheduling strategy is output: S opt ={G opt ,Q opt ,T opt }; If the convergence condition is not met, the process returns to step S32 and continues iterating until the termination condition is met.
5. According to claim 4, a small reservoir intelligent flood discharge scheduling method based on reinforcement learning is characterized in that: The flood control constraint benefit function f flood (S i ,t) is defined as the safety benefit of flood discharge regulation on the current reservoir water level and downstream water level: Among them, H max is the maximum safe water level allowed by the reservoir, H safe is the safety warning water level of the reservoir, H d,max and H d,safe are the maximum safe water level and safe warning water level in the downstream basin, respectively, out,t is the flood discharge, which directly affects the downstream water level change, Q in,t Indicates the inflow flow at the current moment; Irrigation demand benefit function f irrigation (S i ,t)Measure the degree to which the scheduling scheme meets the agricultural irrigation needs: Among them, η is the water delivery efficiency, Q irrigation represents the agricultural irrigation demand in the current period, Q out,t It is the downstream ecological water demand, which is used to measure the minimum flow maintained by the river ecosystem; Ecological water security function f eco (S i ,t) Evaluate the water supply guarantee of the scheduling scheme for the downstream ecosystem: Among them, Q eco To maintain the minimum amount of water required for downstream ecosystems, |Q eco -Q out,t |Measure the degree of deviation in ecological water use.
6. According to the reinforcement learning-based intelligent flood discharge scheduling method for small reservoirs in claim 1, it is characterized in that: The S4 comprises the following steps: S41. The initial improved reinforcement learning model is defined based on the environmental state vector and the initial flood discharge scheduling strategy, and its parameters are initialized as: θ0=Init(w,b,E t ,S opt ); Among them, θ0 represents the initial parameter set of the improved reinforcement learning model, including weight w and bias b, which is used to initialize the reinforcement learning strategy network. Init represents the initialization function, whose parameter input is the environment state vector E t and the initial flood discharge scheduling strategy S opt , combining the current hydrological conditions of small reservoirs and preliminary optimization results, generate the basic framework of the scheduling strategy network; S42. The action space A is defined as the gate opening G t , flood discharge period ΔT, flood discharge volume Q out,t The dynamic range is calculated as follows: G t ∈[G safe ,G max ]; Among them, G safe It represents the minimum opening to ensure that gate operation does not pose a risk to the structure, ΔT is the flood discharge period, and the calculation is based on the current reservoir water level H t The excess amount and the inflow flow Q in,t and flood discharge Q out,t The difference between , enables the scheduling action to be completed within the safe time range; S43. Based on the multi-objective needs of small reservoirs, a reward function R(E) covering flood control benefits, agricultural irrigation demand satisfaction, ecological water supply guarantee and downstream safety indicators is constructed. t ,a t ): R(E t ,a t )=α1·f flood (E t ,a t )+α2·f irrigation (E t ,a t )+α3·f eco (E t ,a t )+α4· f safety (HAVE BEEN t , a t ); Among them, f flood (E t ,a t ) measures the contribution of flood discharge to the current reservoir water level safety, f irrigation (E t ,a t ) represents the satisfaction degree of the current flood discharge for irrigation demand, f eco (E t ,a t ) is the current ecological water guarantee, f safety (E t ,a t ) measures the impact of flood discharge on downstream water level safety, α1, α2, α3, α4 represent the weight coefficients of various objectives; S44. Combine the action space A in step S42, the reward function in step S43, and the environment state vector to define the improved reinforcement learning model as a four-tuple: <ε,A,P,R>; Among them, ε is the environmental state space, corresponding to the current and predicted hydrological information of the small reservoir, A is the action space, corresponding to the gate opening, flood discharge period and flood discharge amount, P is the probability distribution of environmental state transition, which is calculated in combination with the water level evolution and downstream response during the reservoir operation process, and R is the reward function, which measures the degree to which each scheduling action satisfies multiple objective requirements. S45. Integrate steps S41 to S44, based on the environmental state vector and the initial flood discharge scheduling strategy, complete the parameter setting and structure definition of the improved reinforcement learning model, and form a scheduling model that can output flood discharge decisions in real time under multi-objective constraints.
7. According to claim 1, a small reservoir intelligent flood discharge scheduling method based on reinforcement learning is characterized in that: The S5 comprises the following steps: S51. Based on the improved reinforcement learning model quadruple <ε, A, P, R>, the hydrological change process of a small reservoir is modeled in a simulation environment. The simulation environment deduces the water level change, inflow change and downstream basin response at different times {t, t+1, …} according to the reservoir capacity constraints, gate operation characteristics and historical data; S52. Using the initial parameter set θ0 and the initial scheduling strategy S opt Loading the policy network of the improved reinforcement learning model into the simulation environment S53. Run several rounds in the simulation environment, each round includes the following training process: Observe the environment state vector E at time t t ; According to the current strategy network Select action a t ∈A, where θ k is the updated network parameters after the kth round of training; Execute action a t After that, the simulation environment updates the state of the reservoir and the downstream basin to E according to the state transition probability P t+1 , and calculate the reward R(E t ,a t ); According to the reward R(E t ,a t ) and subsequent state E t+1 , update the policy network and change the network parameters from θ k Adjust to θ k+1 ,The update method adopts the optimal value function approximation based on gradient or adaptive update based on temporal difference; S54. After completing multiple rounds of training, determine the convergence degree of the strategy network, and output a new flood discharge scheduling strategy if the preset convergence conditions are met. and its corresponding network parameter θ * If the convergence condition is not reached, return to step S53 to continue training until a better flood discharge scheduling strategy is obtained under complex hydrological and meteorological conditions, and an optimal flood discharge decision plan is generated: S′ opt ={G′ opt ,Q′ opt ,T′ opt }。
Citation Information
Patent Citations
Reinforcement learning model FQI-based reservoir flood control optimal scheduling method
CN112966445A
Reservoir flood control optimal scheduling method considering forecast uncertainty based on digital twinborn
CN116050628A
Reservoir group flood control dispatching intelligent method and system and medium
CN116307533A
Temporal convolutional network-based flood control scheduling solution optimum selection method
WO2022193681A1
Cited By
Intelligent channel gate dam self-adaptive adjusting method based on Internet of Things
CN120370712A
Reservoir multi-objective optimization intelligent scheduling method based on deep reinforcement learning and deterministic strategy gradient algorithm
CN120430478A
Multi-objective optimization intelligent scheduling method for reservoirs based on deep reinforcement learning and deterministic policy gradient algorithm
CN120430478B
Block chain-based drainage basin environment water volume allocation system
CN120471404A
Ecological pattern-based ecological management evaluation method and ecological integrated monitoring system
CN120509795A