Multi-target cascade hydropower station optimal scheduling method and system based on multi-step return A3C reinforcement learning

Through the multi-step return A3C reinforcement learning method, a multi-objective reward function with adjustable dynamic weights and embedded physical constraints are constructed, which solves the multi-objective conflict and strategic short-sightedness problems in cascade hydropower station optimization scheduling, and realizes an efficient and adaptive scheduling solution.

CN120579845APending Publication Date: 2025-09-02FUJIAN HUADIAN FURUI ENERGY DEVELOPMENT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510676746.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Traditional cascade hydropower station optimization scheduling methods are difficult to achieve multi-objective collaborative optimization in high-dimensional, dynamic, and multi-constraint scenarios, and traditional algorithms have high computational complexity and are difficult to cope with dynamic environments and single-step return estimation deviations, resulting in short-sighted strategies.

Method used

A multi-step return A3C reinforcement learning method is adopted to construct a multi-objective reward function with adjustable dynamic weights, combining future multi-step cumulative rewards and state value estimation, embed physical constraints, dynamically adjust the reward function weights, and generate an adaptive multi-objective scheduling strategy.

Benefits of technology

Multi-objective collaborative optimization under high-dimensional dynamic systems is realized, the stability and engineering feasibility of scheduling strategies are improved, and rapid response and multi-objective dynamic adaptation are supported in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579845A_ABST
    Figure CN120579845A_ABST
Patent Text Reader

Abstract

The invention provides a multi-target cascade hydropower station optimal scheduling method and system based on multi-step return A3C reinforcement learning, and the method comprises the steps: taking the minimization of water level deviation, the maximization of power generation benefits, flood control safety and ecological protection as a collaborative optimization target, and constructing a multi-target reward function with an adjustable dynamic weight; future multi-step accumulated rewards are adopted as the return of improved A3C reinforcement learning, and a long-term optimized scheduling strategy is generated through joint estimation of future multi-step rewards and state values; embedding water balance, water turbine output limitation and ecological flow requirements into the reinforcement learning strategy network to ensure that a scheduling scheme meets physical feasibility and safety constraints; and dynamically adjusting the weight coefficient of the reward function according to the real-time power demand and the ecological priority, and outputting a self-adaptive multi-target balanced scheduling instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hydropower and reinforcement learning, and specifically relates to a multi-objective cascade hydropower station optimization scheduling method and system based on multi-step reward A3C reinforcement learning. Background Art

[0002] Hydropower station scheduling plays a crucial role in the energy sector. It not only regulates the rational use of water resources, ensures power supply, and safeguards flood control, but is also crucial for maximizing the utilization of hydropower resources and ensuring stable power system operation. Research on optimal scheduling of cascade hydropower stations aims to efficiently utilize water resources and power generation capacity across multiple hydropower stations, maximizing power generation revenue and ensuring stable power supply. Given the massive scale and complex hierarchies of hydropower systems, as well as the close connection between water and electricity, scheduling requires comprehensive consideration of multiple constraints, resulting in a highly complex optimization decision-making problem.

[0003] Currently, research on cascade hydropower optimization scheduling faces several major challenges: First, designing reasonable and effective optimization scheduling objectives remains challenging given the complex hydropower system, diverse electricity market mechanisms, and uncertain water resources. Second, how to incorporate non-economic factors such as environmental protection and ecological conservation into optimization scheduling decisions to achieve the unity of economic and environmental benefits is also an issue that needs urgent consideration.

[0004] Traditional cascade hydropower station optimization scheduling methods, such as dynamic programming, linear programming, and integer programming, solve optimization problems by building mathematical models. However, due to the uncertainty of hydrological data and computational complexity, they struggle to achieve global optimal solutions or effectively address uncertainty in large-scale problems. In contrast, methods based on intelligent algorithms, such as genetic algorithms and particle swarm optimization, can effectively handle more complex nonlinear, large-scale optimization problems. However, they face difficulties in parameter selection and convergence in high-dimensional systems. Summary of the Invention

[0005] In view of the defects and shortcomings of the existing technology in the optimization and scheduling of cascade hydropower stations, such as the difficulty in multi-objective coordination (such as the difficulty in balancing power generation benefits and ecological protection), the high computational complexity of traditional algorithms and their difficulty in coping with dynamic environments, and the short-sightedness of strategies caused by single-step reward estimation bias in standard reinforcement learning, the present invention provides a multi-objective cascade hydropower station optimization and scheduling method and system based on multi-step reward A3C reinforcement learning.

[0006] With the collaborative optimization goals of maximizing power generation benefits, minimizing water level deviations, ensuring flood control safety and meeting ecological flow standards, a multi-objective reward function with dynamically adjustable weights is constructed, and multi-dimensional goals such as water level deviation from the target, risk of exceeding the safe water level, and substandard ecological flow are converted into a unified reward signal; future multi-step cumulative rewards are used instead of single-step returns, combined with subsequent state value estimation to more accurately evaluate the long-term impact of actions and solve the problem of strategic short-sightedness; physical constraints such as water balance, turbine output limit and ecological flow lower limit are embedded in the reinforcement learning strategy network to ensure that the scheduling plan meets actual operation requirements; at the same time, the weight coefficient of the reward function is dynamically adjusted according to the real-time peaks and valleys of electricity demand, flood warning levels and ecological protection priorities to achieve adaptive balance of multiple objectives.

[0007] Furthermore, the algorithm's stability is enhanced through an improved loss function (integrating a policy gradient term, a value estimation term, and an entropy regularization term). A reinforcement learning environment is constructed using real-time operating parameters such as reservoir water level and power generation as state variables, and water release and output scheduling as action variables. Ultimately, an optimized scheduling strategy is generated that balances long-term benefits, ecological safety, and physical feasibility. This invention effectively addresses the optimization challenges of traditional methods in high-dimensional, dynamic, and multi-constrained scenarios, providing a highly efficient solution for the intelligent scheduling of cascade hydropower stations.

[0008] Its innovative design mainly includes:

[0009] Multi-objective dynamic optimization framework: By constructing a multi-objective reward function with adjustable dynamic weights, minimizing water level deviation, maximizing power generation benefits, flood control safety, and ecological protection are integrated into a unified optimization objective. The weight coefficients (α, β, γ, δ) are dynamically adjusted based on real-time power demand, flood warnings, and ecological priorities, achieving a multi-objective adaptive balance.

[0010] Multi-step reward reinforcement learning mechanism: This mechanism uses future multi-step cumulative rewards instead of the traditional A3C single-step reward. This mechanism combines state value estimation to optimize long-term scheduling strategies, significantly reducing reward estimation bias and improving decision stability.

[0011] Physically constrained embedded policy network: Water balance equations, turbine output limits, and ecological flow requirements are directly embedded in the reinforcement learning training process. Through piecewise linear interpolation modeling, the relationships between reservoir capacity and water level, and between flow and tailwater level are accurately correlated, ensuring that the scheduling plan meets actual physical feasibility and safety constraints.

[0012] Modular system architecture: Integrates multi-objective dynamic decision-making, constraint-embedded execution, and adaptive adjustment modules to achieve closed-loop control from data acquisition, strategy optimization to real-time scheduling, supporting efficient resource scheduling and multi-objective collaborative optimization in dynamic environments.

[0013] The present invention solves the technical bottlenecks of traditional methods such as multi-objective conflicts, insufficient long-term benefit prediction and rigid constraint processing in high-dimensional dynamic systems, and provides an intelligent and adaptive integrated optimization scheduling solution for cascade hydropower stations.

[0014] The technical solution specifically adopted by the present invention to solve the technical problem is:

[0015] A multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning:

[0016] With the coordinated optimization goals of minimizing water level deviation, maximizing power generation benefits, flood control safety, and ecological protection, a multi-objective reward function with adjustable dynamic weights is constructed.

[0017] The future multi-step cumulative reward is used as the reward for improving A3C reinforcement learning. By jointly estimating the future multi-step reward and state value, a long-term optimized scheduling strategy is generated.

[0018] Embed water balance, turbine output limits, and ecological flow requirements into a reinforcement learning strategy network to ensure that the scheduling plan meets physical feasibility and safety constraints;

[0019] The weight coefficient of the reward function is dynamically adjusted according to real-time power demand and ecological priority, and scheduling instructions with adaptive multi-objective balance are output.

[0020] Furthermore, the future multi-step cumulative reward is calculated by accumulating the immediate rewards of the next n steps starting from the current time step and superimposing the state value estimate of the nth step.

[0021] Furthermore, the multi-objective reward function includes the following collaborative objectives:

[0022] Water level deviation penalty: based on the absolute deviation between the actual water level and the target water level;

[0023] Power generation benefit reward item: positively correlated with power generation;

[0024] Flood control safety penalty: constrained by the square deviation of the water level from the maximum safety value;

[0025] Ecological flow tracking item: based on the absolute deviation between the actual ecological flow and the target value.

[0026] Furthermore, the multi-objective reward function balances multiple objectives in the following way:

[0027] Negative rewards are imposed on deviations from the day-ahead target water level to constrain the deviation between real-time scheduling and the plan;

[0028] Imposing positive incentives on power generation to encourage maximization of power generation benefits;

[0029] Imposing negative rewards for the risk of exceeding safe water levels to reduce flood safety risks;

[0030] Imposing negative rewards for failing to meet ecological flow standards;

[0031] The priority of each target item is adjusted through the weight coefficient.

[0032] Furthermore, embedding water balance, turbine output limit, and ecological flow requirements into the reinforcement learning strategy network to ensure that the scheduling plan meets physical feasibility and safety constraints includes:

[0033] The water balance constraint relates reservoir capacity and water level through a piecewise linear interpolation model: ensuring that changes in reservoir water storage conform to the dynamic balance between inflow and discharge;

[0034] Turbine output constraint: limit the power generation to within the equipment safety range;

[0035] The ecological flow constraint relates the downstream flow and tailwater level through a piecewise linear interpolation model: the target ecological flow is dynamically tracked and the water release amount is adjusted.

[0036] Furthermore, the weight coefficient of the reward function is dynamically adjusted according to the real-time power demand and ecological priority, and the scheduling instruction for adaptive multi-objective balance is outputted in the following manner:

[0037] Dynamically improve the power generation efficiency weight coefficient β according to the peak and valley changes in electricity demand;

[0038] Dynamically increase the flood control safety weight coefficient γ according to the flood warning level;

[0039] According to the priority of ecological protection, the ecological flow weight coefficient δ is dynamically adjusted.

[0040] Furthermore, the loss function used to optimize the scheduling strategy includes:

[0041] Policy gradient term: maximizes the multi-step reward advantage function;

[0042] Value estimation term: minimize the state value prediction error;

[0043] Policy diversity term: Prevents the policy from converging prematurely through entropy regularization.

[0044] Furthermore, parameters reflecting the real-time operating status of the hydropower station, including reservoir water level, power generation, inflow, power demand, ecological flow, and the day-ahead target water level, are used as state variables, and control operations including water release and output scheduling are used as action variables to generate a reinforcement learning interactive environment.

[0045] And, a multi-objective cascade hydropower station optimization scheduling system based on multi-step reward A3C reinforcement learning, including:

[0046] Multi-objective dynamic decision-making module: Integrates a multi-step reward mechanism with a dynamic weighted reward function to generate long-term optimal scheduling strategies;

[0047] Constraint embedding execution module: Feedback water balance, equipment output and ecological flow constraints to the strategy network in real time;

[0048] Adaptive adjustment module: Dynamically adjusts target weights based on external inputs including power demand and ecological indicators, and outputs adaptive scheduling instructions.

[0049] And, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0050] A non-transitory computer-readable storage medium stores a computer program, which implements the steps of the method described above when executed by a processor.

[0051] Compared with the prior art, the present invention and its preferred embodiments have at least the following beneficial effects:

[0052] Significantly improved multi-objective collaborative optimization capabilities: Through the design of a multi-objective reward function with adjustable dynamic weights, it effectively balances the conflicting demands of water level control, power generation efficiency, flood control safety, and ecological protection, resolving the lack of adaptability caused by the static weighting of multiple objectives in traditional methods.

[0053] Enhanced stability of long-term scheduling decisions: A reinforcement learning mechanism based on multi-step cumulative rewards, combined with joint estimation of future multi-step rewards and state values, significantly reduces the short-sightedness bias of single-step rewards, improving the robustness of long-term revenue forecasting and strategy optimization for cascade hydropower stations.

[0054] Precisely embedding physical constraints and ensuring project feasibility: By directly embedding water balance, equipment output limits, and ecological flow requirements into the policy network training process and using a piecewise linear interpolation model to accurately describe the relationships between reservoir capacity and water level, and between flow and tailwater level, this ensures that the scheduling plan strictly meets actual project constraints, avoiding safety risks caused by traditional reinforcement learning that ignores physical rules.

[0055] Optimization of dynamic environmental adaptability: Supports dynamic adjustment of weight coefficients based on real-time power demand, flood warning level, and ecological protection priority. Combined with the closed-loop control mechanism of the modular system architecture, it achieves rapid response and multi-objective dynamic adaptation in complex hydrological and power market environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0057] Figure 1This is a flow chart of a multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the features and advantages of the present invention more clearly understood, the following embodiments are given for detailed description:

[0059] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs.

[0060] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0061] To address the complex challenges of existing technologies, this paper proposes a multi-objective cascade hydropower station optimization scheduling method and system based on multi-step reward A3C (Advantage Actor-Critic) reinforcement learning. This method efficiently handles complex optimization problems with high-dimensional state and action spaces, accurately achieving comprehensive optimization objectives, improving the overall operational efficiency of cascade hydropower systems, and promoting the intelligent and efficient development of energy systems.

[0062] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions and their design processes:

[0063] In the first aspect, the present invention provides a multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning. First, through environmental modeling, the reservoir water level, power generation, inflow, power demand and ecological flow are used as state variables, and the hydropower station water discharge, generator unit scheduling, etc. are regarded as action variables; secondly, by designing a comprehensive reward function, the water level deviation, power generation efficiency, and ecological protection are converted into reward functions; in the training process, a multi-step reward strategy is adopted, that is, the multi-step cumulative reward starting from the current time step is calculated, and the long-term value of the state is evaluated to further reduce the estimation deviation; at the same time, the loss function is improved, the training of the policy network and the value network is optimized, and the stability of the algorithm is enhanced. The method of the present invention can improve the accuracy of scheduling and has significant advantages in the utilization of hydropower resources and the protection of the ecological environment, specifically including:

[0064] (1) Environmental modeling

[0065] Modeling the dispatching environment of cascade hydropower stations transforms the actual hydropower dispatching problem into a form that can be processed by reinforcement learning algorithms. This mainly includes state and action modeling:

[0066] Status S t is the state of the hydropower system at time t, expressed as a vector:

[0067]

[0068] in: is the reservoir water level of the i-th hydropower station; F t i is the power generation of the i-th hydropower station; Q t is the inflow at the current moment; D t is the electricity demand at the current moment; E t for ecological flow; The target water level for day-ahead dispatch.

[0069] Action A t Represents the scheduling decision taken by the agent at time t:

[0070]

[0071] in: is the water discharge of the i-th hydropower station; is the output dispatch of the i-th hydropower station.

[0072] (2) Reward Function

[0073] Designing dispatching targets that are consistent with the realities of cascade hydropower is crucial. First, excessively high reservoir water levels can easily lead to flooding or safety issues, while excessively low water levels can reduce power generation efficiency. Second, the significant deviation between real-time dispatching and planned dispatching makes it difficult to ensure stable system operation. Furthermore, while ensuring power generation efficiency, in order to meet ecological flow requirements, it is necessary to effectively incorporate ecological protection into power generation dispatching, specifically as follows:

[0074]

[0075] Therefore, the reward function R t Taking the above objectives into consideration, we transform multiple objectives into a single objective and use it as the reward function for the action:

[0076]

[0077] Among them: α, β, γ, δ are the weight coefficients of water level deviation, power generation efficiency, flood control safety and ecological protection respectively; is the maximum safe water level of the i-th hydropower station; E targetThe target ecological flow.

[0078] To more accurately evaluate state values, we use multi-step returns instead of single-step returns in standard A3C. Multi-step returns are a method for updating policy and value functions by taking into account returns from multiple time steps into the future. This provides more stable and accurate gradient estimates and helps reduce estimation bias. At each time step t, we calculate the multi-step return and multi-step advantage function:

[0079]

[0080] in: represents the n-step return starting from time t; r t+k represents the immediate reward of step k; V(S t+n θ v ) represents the state value at time t+n.

[0081] (3) Training and parameter updating of policy network and value network

[0082] The A3C algorithm consists of a policy network (Actor) and a value network (Critic). The policy network outputs the probability distribution of each action taken in the current state, and the value network estimates the value of the current state. The specific update method is as follows:

[0083] According to the policy gradient algorithm, the parameters of the policy network are updated to maximize the long-term return. The actor is updated as follows:

[0084]

[0085] Combining strategy loss, value loss, and entropy regularization to improve the Critic update loss function:

[0086]

[0087] Among them: β and η are hyperparameters used to adjust the weights of value loss and entropy regularization terms; π θ (A t |S t ) is the given state S under the policy network parameters θ t Take action A t The probability of logπ θ (A t |S t ) is the strategy π θ Next, take action A t The logarithmic probability of J(θ) is the objective function of the policy network, that is, the expected long-term return; L(θ,θ v ) is the loss function of the value network.

[0088] The parameters θ of the policy network and the parameters θ of the value network v Update separately:

[0089]

[0090] (4) Constraints

[0091] Water balance equation:

[0092]

[0093] Electric power balance equation:

[0094]

[0095] Turbine output characteristics:

[0096]

[0097] Reservoir capacity and water level constraints:

[0098]

[0099] Relationship between tailwater level and discharge flow:

[0100]

[0101]

[0102] Water level constraint:

[0103]

[0104] Downflow flow restriction:

[0105]

[0106] Turbine output constraints:

[0107]

[0108] Where: S i,t is the water storage capacity of power station i at time t; and are inflow, outflow and power generation flow respectively; E i,t is the ecological flow; P i,t is the turbine output power; is the turbine power generation efficiency; H i,t is the water head of power station i at time t; is the mth storage capacity interval indicator variable of power station i at time t, a 0-1 variable used to judge the storage capacity S i,t The interval; S i,mis the lower boundary of the mth storage capacity interval of power station i; is the lower boundary of the mth upstream water level interval of power station i; is the water level upstream of power station i at time t; k i,m is the proportional coefficient between water level and reservoir capacity in the mth interval of power station i; M is the number of reservoir capacity-water level intervals; The mth discharge flow interval indicator variable of power station i at time t, a 0-1 variable, is used to determine the discharge flow The interval; Q i,n is the lower boundary of the mth flow interval of power station i; is the lower boundary of the nth downstream water level interval of power station i; is the water level downstream of power station i at time t; v i,n is the proportional coefficient between tailwater level and discharge flow in the nth interval of power station i; N is the number of tailwater level-discharge flow intervals; are the maximum and minimum limits of upstream and downstream water levels respectively; ΔZ is the limit of water level change in adjacent time periods; Q i,min , Q i,max are the lower and upper limits of the discharge flow of power station i respectively.

[0109] Therefore, based on the above design, the present invention proposes a multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning. The specific technical process is as follows: Figure 1 As shown, the following steps are included:

[0110] S1. Environmental modeling: Reservoir water level, power generation, inflow, power demand and ecological flow are used as state variables, and the water release of the hydropower station and the scheduling of the generator units are used as action variables for modeling.

[0111] S2. Reward function design: Design a reward function that comprehensively considers water level deviation, power generation efficiency, and ecological protection to guide the algorithm to optimize the scheduling strategy.

[0112] S3. Multi-step reward calculation: Calculate the multi-step cumulative reward from the current time step t, evaluate the long-term value of the state, and reduce the return estimation bias.

[0113] S4. Improve the loss function: Combine the policy loss, value loss, and entropy regularization term to improve the loss function to enhance the stability and performance of the algorithm.

[0114] S5. Parameter update: Update the parameters of the policy network and value network through gradient descent to optimize scheduling decisions.

[0115] S6. Scheduling optimization: Apply the trained strategy network for real-time scheduling to ensure that multiple constraints such as water level deviation, power generation efficiency, and ecological flow are met.

[0116] Secondly, this embodiment provides a multi-objective cascade hydropower station optimization scheduling system using multi-step reward A3C reinforcement learning, including:

[0117] Data Acquisition and Preprocessing Module: This module collects historical operating data from the hydropower station, including reservoir capacity, generator performance, rainfall, and water flow. It then cleans, standardizes, and preprocesses the data to meet the input requirements of the reinforcement learning algorithm.

[0118] Model Building and Environmental Simulation Module: This module builds a comprehensive model of the hydropower station, describing the reservoir capacity, generator performance, and their dynamic characteristics. This module simulates the hydropower station's operating environment, including factors such as power demand, generation costs, and environmental impact, providing a training environment for the reinforcement learning algorithm.

[0119] Multi-step reward A3C reinforcement learning module: This module implements the multi-step reward-based A3C reinforcement learning algorithm, consisting of a policy network and a value network. The policy network generates scheduling policies, while the value network estimates the value of states. It uses a multi-step reward approach to improve the stability of policy updates and improves the training performance of the policy and value functions through an improved loss function.

[0120] Training and Optimization Module: Conducts asynchronous training, accelerates convergence through multiple parallel workers, and improves training efficiency using an improved experience replay mechanism. It adjusts and optimizes hyperparameters during training to ensure the effectiveness and stability of the strategy.

[0121] Dispatch Strategy Generation and Application Module: Generates multi-objective optimization dispatch strategies based on trained strategies. These strategies are applied in real time to cascade hydropower station dispatch, adjusting and optimizing dispatch strategies to accommodate varying operating conditions and target changes.

[0122] Performance Evaluation and Verification Module: This module evaluates the performance of the optimized scheduling method, including comparisons with traditional scheduling methods, and assesses its performance in terms of power generation efficiency, cost control, and environmental protection. Verification is conducted in simulations and actual operating environments to ensure the practical application and stability of the optimized scheduling method.

[0123] User Interface and Monitoring Module: Provides an intuitive user interface that allows users to configure system parameters and view scheduling results and performance indicators. It monitors system operation status and scheduling effects in real time, providing feedback and adjustment suggestions.

[0124] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0125] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which executes the above method when executed by a processor. The storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.

[0126] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0127] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.

[0128] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive various other forms of a multi-objective cascade hydropower station optimization scheduling method and system based on multi-step reward A3C reinforcement learning under the inspiration of the present invention. All equal changes and modifications made within the scope of the patent application of the present invention should fall within the scope of the present invention.

Claims

1. A multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning, characterized by: With the coordinated optimization goals of minimizing water level deviation, maximizing power generation benefits, flood control safety, and ecological protection, a multi-objective reward function with adjustable dynamic weights is constructed. The future multi-step cumulative reward is used as the reward for improving A3C reinforcement learning. By jointly estimating the future multi-step reward and state value, a long-term optimized scheduling strategy is generated. Embed water balance, turbine output limits, and ecological flow requirements into a reinforcement learning strategy network to ensure that the scheduling plan meets physical feasibility and safety constraints; The weight coefficient of the reward function is dynamically adjusted according to real-time power demand and ecological priority, and scheduling instructions with adaptive multi-objective balance are output.

2. The multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to claim 1 is characterized by: The future multi-step cumulative reward is calculated by accumulating the immediate rewards for the next n steps starting from the current time step and superimposing the state value estimate of the nth step.

3. The multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to claim 1 is characterized by: The multi-objective reward function includes the following collaborative objectives: Water level deviation penalty: based on the absolute deviation between the actual water level and the target water level; Power generation benefit reward item: positively correlated with power generation; Flood control safety penalty: constrained by the square deviation of the water level from the maximum safety value; Ecological flow tracking item: based on the absolute deviation between the actual ecological flow and the target value.

4. The multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to claim 3 is characterized by: The multi-objective reward function balances multiple objectives in the following way: Negative rewards are imposed on deviations from the day-ahead target water level to constrain the deviation between real-time scheduling and the plan; Imposing positive incentives on power generation to encourage maximization of power generation benefits; Imposing negative rewards for the risk of exceeding safe water levels to reduce flood safety risks; Imposing negative rewards for failing to meet ecological flow standards; The priority of each target item is adjusted through the weight coefficient.

5. The multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to claim 1 is characterized by: The aforementioned embedding of water balance, turbine output limit, and ecological flow requirements into the reinforcement learning strategy network ensures that the scheduling scheme meets physical feasibility and safety constraints, including: The water balance constraint relates reservoir capacity and water level through a piecewise linear interpolation model: ensuring that changes in reservoir water storage conform to the dynamic balance between inflow and discharge; Turbine output constraint: limit the power generation to within the equipment safety range; The ecological flow constraint relates the downstream flow and tailwater level through a piecewise linear interpolation model: the target ecological flow is dynamically tracked and the water release amount is adjusted.

6. The multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to claim 1 is characterized by: The dynamic adjustment of the weight coefficient of the reward function according to the real-time power demand and ecological priority and the output of the adaptive multi-objective balanced dispatch instruction are achieved by the following method: Dynamically improve the power generation efficiency weight coefficient β according to the peak and valley changes in electricity demand; Dynamically increase the flood control safety weight coefficient γ according to the flood warning level; According to the priority of ecological protection, the ecological flow weight coefficient δ is dynamically adjusted.

7. The multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to claim 1 is characterized by: The loss functions used to optimize the scheduling strategy include: Policy gradient term: maximizes the multi-step reward advantage function; Value estimation term: minimize the state value prediction error; Policy diversity term: Prevents the policy from converging prematurely through entropy regularization.

8. The multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to claim 1 is characterized by: Parameters reflecting the real-time operating status of the hydropower station, including reservoir water level, power generation, inflow, power demand, ecological flow, and the day-ahead target water level, are used as state variables, and control operations including water release and output scheduling are used as action variables to generate a reinforcement learning interactive environment.

9. A multi-objective cascade hydropower station optimization scheduling system based on multi-step reward A3C reinforcement learning, characterized by: include: Multi-objective dynamic decision-making module: Integrates a multi-step reward mechanism with a dynamic weighted reward function to generate long-term optimal scheduling strategies; Constraint embedding execution module: Feedback water balance, equipment output and ecological flow constraints to the strategy network in real time; Adaptive adjustment module: Dynamically adjusts target weights based on external inputs including power demand and ecological indicators, and outputs adaptive scheduling instructions.

10. An electronic device, characterized in that: It comprises a processor and a memory; the memory stores a computer program, and when the computer program is executed by the processor, it implements the multi-objective cascade hydropower station optimization scheduling method based on multi-step reward A3C reinforcement learning according to any one of claims 1 to 8.