Real-time load optimal distribution method between cascade hydropower stations, equipment and medium

By constructing a generalized model based on hydraulic relationships and inflow, and designing a reinforcement learning framework, combined with water level control objectives and hydropower utilization efficiency, the load allocation of cascade hydropower stations is optimized. This solves the limitations of traditional algorithms and the insufficient generalization ability of deep learning, and achieves both accuracy and efficiency in load allocation.

CN121525916APending Publication Date: 2026-02-13GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511345458.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

In existing technologies, traditional optimization algorithms are susceptible to the "curse of dimensionality," heuristic algorithms have low computational accuracy, and deep learning and reinforcement learning algorithms have insufficient generalization ability, resulting in difficulties in real-time load optimization and allocation for cascade hydropower stations and low training efficiency.

Method used

By constructing a generalized model based on hydraulic relationships and inflow, a reinforcement learning framework is designed. Combining water level control objectives and hydropower utilization efficiency, agent training and a dual-network structure are adopted to optimize the load allocation scheme. Utilizing the reservoir water balance principle and power balance equation, and considering the river flow lag time, a reward function and action selection strategy are constructed to achieve accurate simulation and optimization of load allocation.

Benefits of technology

It improved the feasibility and computational efficiency of load allocation, ensured the accuracy of power plant unit simulation and river flow prediction, enhanced the coordinated operation capability of the cascade system and the utilization efficiency of hydropower resources, and realized the stability and efficiency of the load allocation scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121525916A_ABST
    Figure CN121525916A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hydropower dispatching operation, in particular to a real-time load optimal distribution method and device between cascade hydropower stations and a medium, a hydropower station group system is generalized according to a hydraulic relationship and reservoir inflow, and a cascade reservoir group is simulated; constructing a load distribution model according to the constraint relationship, building a solving framework, designing an environment state, a reward function and an intelligent agent of the solving framework, and training the intelligent agent; and solving the load distribution model based on the solving framework, screening selectable load distribution schemes according to constraint conditions and distribution requirements, and outputting a final load distribution scheme in combination with a water level control target and water energy utilization efficiency, thereby realizing dynamic simulation of a real-time operation process of the cascade power station and optimization of an inter-station load distribution scheme. In a novel electric power system, regulation and control of multi-target tasks such as flood control and power generation can be taken into consideration, the decision-making efficiency in the model solving process is effectively improved, and quick response to load distribution under the high-intensity peak regulation and frequency modulation background is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydropower dispatching and operation technology, and in particular to a method, equipment and medium for real-time load optimization allocation among cascade hydropower stations. Background Technology

[0002] New energy sources, primarily wind and solar, are transitioning from supplementing incremental electricity consumption to becoming the main source of electricity consumption. However, due to the strong randomness and volatility of wind and solar power output, the explosive growth in new energy installed capacity, while driving the green energy transformation, also poses severe challenges to the safe and stable operation of the power system. Therefore, designing and constructing a real-time load optimization and allocation model adapted to large-scale cascade hydropower stations, and achieving synergistic optimization of decision-making accuracy and computational efficiency while meeting the requirements of grid load commands and water dispatching, has become a key challenge in the real-time dispatching of cascade hydropower under the background of new power system construction.

[0003] With the development of computers, a large number of algorithms have been applied to practical engineering. Among them, deep reinforcement learning algorithms have stood out due to their advantages such as strong environmental adaptability and fast decision-making speed. They have a solid theoretical foundation and strong feasibility for solving the real-time load optimization allocation of cascade hydropower.

[0004] However, current research by most scholars mainly focuses on energy storage optimization and water consumption rate optimization, with insufficient attention paid to real-time dynamic water level control. Some scholars consider minimizing the system's water consumption rate but do not focus on real-time water level changes. Others start from minimizing total energy consumption to improve the economic operation of hydropower stations but still do not consider the adaptive load allocation of hydropower stations under different conditions. Secondly, in terms of algorithms, traditional optimization algorithms are easily limited by the "curse of dimensionality"; while heuristic algorithms may have low computational accuracy or even produce invalid solutions; some deep learning and reinforcement learning algorithms have insufficient generalization ability and low training efficiency. Summary of the Invention

[0005] In view of the aforementioned existing problems, the present invention is proposed.

[0006] Therefore, this invention provides a method, equipment, and medium for real-time load optimization allocation between cascade hydropower stations to address the limitations of traditional optimization algorithms, which are easily affected by the "curse of dimensionality." Heuristic algorithms may suffer from low computational accuracy or even invalid solutions. Some deep learning and reinforcement learning algorithms have insufficient generalization ability and low training efficiency.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0008] In a first aspect, the present invention provides a method for real-time load optimization allocation among cascade hydropower stations, comprising:

[0009] Based on hydraulic relationships and inflow, the hydropower station group system is generalized, and the cascade reservoir group is simulated;

[0010] A load allocation model is constructed based on the constraint relationship, and a solution framework is built. The environment state, reward function, and agent of the solution framework are designed, and the agent is trained.

[0011] The load allocation model is solved based on the solution framework. Selectable load allocation schemes are selected according to constraints and allocation requirements. Finally, the load allocation scheme is output by combining water level control objectives and hydropower utilization efficiency.

[0012] As a preferred embodiment of the real-time load optimization allocation method among cascade hydropower stations described in this invention, the method includes: simulating the cascade reservoir group, including:

[0013] Simulate the power plant unit;

[0014] Simulation of the river channel unit within the interval.

[0015] As a preferred embodiment of the real-time load optimization allocation method among cascade hydropower stations described in this invention, the training of the agent includes:

[0016] The agent reads the state information and uses it as input, selects and executes actions, obtains timely feedback and the state of the next time period, and generates samples to be stored in the experience replay pool.

[0017] Input a portion of the data from the sample into the reinforcement learning policy network and the reinforcement learning value network respectively, and calculate the difference between the target value and the current value;

[0018] The difference between the calculated target value and the current value is input into the reinforcement learning policy network, and the parameters of the reinforcement learning policy network are updated using an optimization algorithm.

[0019] As a preferred embodiment of the real-time load optimization allocation method among cascade hydropower stations described in this invention, the final load allocation scheme is output by combining water level control targets and hydropower utilization efficiency, including:

[0020] The processing interval is iteratively calculated to back-calculate the outflow rate of the current period and the reservoir water level at the beginning of the next period, and to calculate the reward value of the current state.

[0021] The original state is input into the artificial neural network to obtain the value of the state, and the state is updated. The updated state is input into the artificial neural network again to obtain the value of the corresponding state, and the parameters of the artificial neural network are updated.

[0022] If the sum of the output rewards tends to converge after the round termination condition and the learning termination condition are met, then the final load allocation scheme is output.

[0023] As a preferred embodiment of the real-time load optimization allocation method among cascade hydropower stations described in this invention, the simulation of the power station unit includes:

[0024] The power station units of the cascade hydropower system are simulated. During operation, the power generation flow and flood discharge flow of the power station are superimposed to form the total discharge flow of the power station, and the dynamic changes of the discharge flow must strictly follow the principle of reservoir water balance.

[0025]

[0026] In the formula, V t+1 and V t These represent the initial reservoir storage volumes at time t and time t+1, respectively, in m³; t This represents the inflow rate of the reservoir during time period t, expressed in m³ / s. This represents the power generation flow rate during time period t, expressed in m³ / s. Δt represents the flood discharge rate during time period t, in m³ / s; Δt represents the duration of the calculation period, in seconds.

[0027] The specific scheduling parameters of each power plant unit are mainly determined through trial calculations using the "electricity-based water supply" model. The specific simulation method is as follows:

[0028] The output of each power station in the cascade is calculated based on the grid load command and the power balance equation:

[0029]

[0030] In the formula, P t P represents the total load command of the cascade hydropower system during time period t, i.e., the load value to be completed, in MW; i,t This represents the load value that power station i needs to bear during time period t, in MW; i represents the number of the cascade hydropower station; n represents the total number of power stations included in the cascade hydropower system;

[0031] Following the "electricity-determined water supply" model, the power generation flow of the hydropower station is calculated from its output equation:

[0032]

[0033] In the formula, i represents the designation of the cascade hydropower station, i = 1, 2, ..., n; N i,t K represents the power output of power station i in time period t, in MW; i This represents the output coefficient of power station i; H represents the power generation flow rate in time period t, in m³ / s. i,t Let be the average hydropower head of power station i during time period t, in meters.

[0034] The beneficial effects of this preferred technical solution are as follows: By strictly adhering to the principle of reservoir water balance and adopting an "electricity-determined water supply" operation mode for simulation calculations, the accuracy and practicality of power station unit simulation are effectively ensured. This method, based on the total load command of the power grid, rationally decomposes the output tasks of each power station through the power balance equation, and then uses the power station output equation to inversely deduce the power generation flow, achieving precise coupling between electricity demand and water resources. This simulation process ensures the physical reality of the dynamic changes in the power station's discharge flow, providing a reliable hydraulic boundary and operational basis for subsequent optimized allocation, significantly improving the feasibility of the load allocation scheme, computational efficiency, and the overall coordinated operation capability of the cascade system.

[0035] As a preferred embodiment of the real-time load optimization allocation method between cascade hydropower stations described in this invention, the method includes: simulating the river channel units within the intervals, including:

[0036] The river channel unit was simulated, and the river was divided into several segments based on the geographical distribution characteristics of the power stations. The hydrological processes of each segment showed significant spatial correlations:

[0037] I down,t =Q up,t +Q section,t

[0038] In the formula, I down,t Q represents the inflow rate to the downstream reservoir at time t, expressed in m³ / s. up,t Q represents the downstream discharge flow rate of the upstream reservoir at time t, expressed in m³ / s; section,t This represents the interval runoff between the upstream and downstream reservoirs of the river during time period t, in m3 / s.

[0039] If the time interval between two adjacent load commands within the scheduling cycle is less than the water flow propagation time between upstream and downstream power stations, then the simulation of the river channel unit in the interval needs to consider the flow arrival lag time of the river channel, and the calculation formula for the inflow of the downstream reservoir is updated as follows:

[0040] I down,t =Q up,t-τ +Q section,t

[0041]

[0042] t1≤t-τ≤t2

[0043] In the formula, I down,t Q represents the inflow rate to the downstream reservoir at time t, in m³ / s; τ represents the time it takes for the outflow from the upstream reservoir to reach the downstream power station, in seconds; up,t-τ This represents the outflow from the upstream reservoir during the time period t-τ, expressed in m³ / s. and These represent the outflow from the upstream reservoir at times t1 and t2, respectively, in m³ / s.

[0044] The beneficial effects of this preferred technical solution are as follows: Based on the geographical distribution of river sections and the adoption of a dynamic inflow calculation formula that considers flow lag time, the accuracy of downstream power station inflow prediction is significantly improved. It can effectively handle the time lag effect of water flow propagation when short-term load command intervals are small, overcoming simulation biases that may arise from traditional simplified methods. This provides a reliable hydrological and hydraulic boundary for real-time optimized allocation of cascade loads, ensuring the rationality and engineering practicality of the optimization results.

[0045] As a preferred embodiment of the real-time load optimization allocation method among cascade hydropower stations described in this invention, the intelligent agent reads state information as input, selects and executes actions, obtains timely feedback and the state for the next time period, and generates samples to be stored in the experience playback pool, including:

[0046] During time period t, the agent reads the status information s of each cascade reservoir in the environment, such as the upstream water level and inflow. t As input to the model, with the goal of maximizing the value function, action a is selected and executed. t Timely feedback r is obtained based on the reward function. t and the next time period state s t+1 The above information is combined into a single sample {s}. t ,a t ,r t ,s t+1 Store it in the experience replay pool.

[0047] As a preferred embodiment of the real-time load optimization allocation method among cascade hydropower stations described in this invention, the method includes: inputting a portion of the sample data into a reinforcement learning policy network and a reinforcement learning value network respectively, and calculating the difference between the target value and the current value, including:

[0048] The state data of the sample s t and s t+1 Input the ON and TN networks respectively, and generate Q(s) based on the network parameters. t ,a t ) and Q(s t+1 ,a t+1 ), and calculate the target Q value Q. π (s t ,a t ), and calculate the difference between it and the current Q value.

[0049] Q π (s t ,a t) = r t (s t ,a t )+γmaxQ(s t+1 ,a t+1 )

[0050] In the formula, Q π (s t ,a t ) indicates in s t In the given state, select and execute action a according to action selection strategy π. t The cumulative reward value generated afterward; γ represents the discount factor; maxQ(s) t+1 ,a t+1 ) indicates in s t+1 In the given state, the agent takes action a, calculated by the action value function Q(s,a). t+1 The maximum value that can be obtained later.

[0051] The beneficial effects of this preferred technical solution are as follows: By inputting sample data from the experience replay pool into the online policy network and the target value network respectively, and calculating the difference between their output values, a clear direction and quantitative basis are provided for the agent's policy optimization. This process utilizes the target network to generate a stable target value estimate, effectively overcoming the instability problem that may occur when evaluating a single network. Furthermore, the difference calculation intuitively reflects the gap between the current policy and the ideal state, significantly improving the convergence speed and stability of the reinforcement learning algorithm. This enables the agent to learn more efficiently from historical experience, ultimately achieving continuous optimization and performance improvement of the tiered load allocation strategy.

[0052] In a second aspect, the present invention provides an electronic device, comprising:

[0053] Memory and processor;

[0054] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of a method for real-time load optimization allocation between cascade hydropower stations.

[0055] Thirdly, the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method for real-time load optimization allocation between cascade hydropower stations.

[0056] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention utilizes a policy network and a value network to evaluate the utility of states and actions respectively, and continuously adjusts network parameters by calculating the difference between the target value and the current value, thereby effectively improving the stability of the training process and the convergence of the policy. The experience replay mechanism breaks the temporal correlation between samples, improves data utilization efficiency, and enables the agent to gradually learn the optimal load allocation strategy that balances water level control objectives and hydropower utilization efficiency under complex hydraulic constraints and power load conditions, ultimately improving the overall operational efficiency of the cascade hydropower system. Attached Figure Description

[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a schematic diagram of the overall process of a real-time load optimization allocation method among cascade hydropower stations according to an embodiment of the present invention.

[0059] Figure 2 This is a simplified model diagram of a cascade hydropower system for a real-time load optimization allocation method among cascade hydropower stations according to an embodiment of the present invention. Detailed Implementation

[0060] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0061] Example 1, referring to Figure 1 As an embodiment of the present invention, a method for real-time load optimization allocation among cascade hydropower stations is provided, comprising:

[0062] S1: Based on hydraulic relationships and inflow, the hydropower station group system is generalized, and the cascade reservoir group is simulated;

[0063] S2: Construct a load distribution model based on the constraint relationship, build a solution framework, design the environment state, reward function, and agent of the solution framework, and train the agent;

[0064] S3: Solve the load allocation model based on the solution framework, select the available load allocation schemes according to the constraints and allocation requirements, and output the final load allocation scheme by combining the water level control target and water energy utilization efficiency.

[0065] It should be noted that by accurately generalizing the reservoir system through hydraulic relationships and inflow, a realistic physical basis is laid for the optimization problem. Furthermore, by rationally designing the core elements of the reinforcement learning framework, the agent can fully learn the system's operating rules. Finally, in the solution process, multiple constraints and objectives are comprehensively considered, and the optimal allocation scheme is automatically selected and output to maximize the efficiency of hydropower utilization and the accuracy of water level control while meeting the power load demand. This significantly enhances the economy and reliability of the cascade hydropower system.

[0066] Example 2, refer to Figures 1-2 As an embodiment of the present invention, based on the above embodiment, a method for real-time load optimization allocation among cascade hydropower stations is provided.

[0067] In this embodiment of the application, step S1 generalizes the hydropower station group system based on hydraulic relationships and inflow, and simulates the cascade reservoir group, including the following steps A1-A4:

[0068] A1: The hydropower station group system is generalized based on hydraulic relationships and inflow.

[0069] A2: Simulate the power plant unit;

[0070] A3: Simulate the river unit within the interval.

[0071] Specifically, step A1, based on hydraulic relationships and inflow, generalizes the hydropower station group system as follows:

[0072] There is a close hydraulic-electric connection between the upstream and downstream hydropower stations in a river basin. The water flow forms a continuous transmission process from upstream to downstream, from "power station-river channel-reservoir". The inflow into the downstream reservoir consists of two parts: first, the discharge from the upstream power station through generator operation or gate regulation, which is controllable; second, the runoff from the natural river section between the two reservoirs, exhibiting significant spatial variability and randomness. The cascade hydropower system can be generalized as follows: Figure 2 The diagram shows a simplified model consisting of power stations at various levels and waterways.

[0073] Specifically, the simulation of the power plant unit in step A2 is manifested as follows:

[0074] The power station units of the cascade hydropower system are simulated. During operation, the power generation flow and flood discharge flow of the power station are superimposed to form the total discharge flow of the power station, and the dynamic changes of the discharge flow must strictly follow the principle of reservoir water balance.

[0075]

[0076] In the formula, V t+1 and V t These represent the initial reservoir storage volumes at time t and time t+1, respectively, in m³; t This represents the inflow rate of the reservoir during time period t, expressed in m³ / s. This represents the power generation flow rate during time period t, expressed in m³ / s. Δt represents the flood discharge rate during time period t, in m³ / s; Δt represents the duration of the calculation period, in seconds.

[0077] The specific scheduling parameters of each power plant unit are mainly determined through trial calculations using the "electricity-based water supply" model. The specific simulation method is as follows:

[0078] The output of each power station in the cascade is calculated based on the grid load command and the power balance equation:

[0079]

[0080] In the formula, P t P represents the total load command of the cascade hydropower system during time period t, i.e., the load value to be completed, in MW; i,t This represents the load value that power station i needs to bear during time period t, in MW; i represents the number of the cascade hydropower station; n represents the total number of power stations included in the cascade hydropower system;

[0081] Following the "electricity-determined water supply" model, the power generation flow of the hydropower station is calculated from its output equation:

[0082]

[0083] In the formula, i represents the designation of the cascade hydropower station, i = 1, 2, ..., n; N i,t K represents the power output of power station i in time period t, in MW; i This represents the output coefficient of power station i; H represents the power generation flow rate in time period t, in m³ / s. i,t Let be the average hydropower head of power station i during time period t, in meters.

[0084] Specifically, the simulation of the inter-section river unit in step A3 is manifested as follows:

[0085] The river channel unit was simulated, and the river was divided into several segments based on the geographical distribution characteristics of the power stations. The hydrological processes of each segment showed significant spatial correlations:

[0086] I down,t =Q up,t +Qsection,t

[0087] In the formula, I down,t Q represents the inflow rate to the downstream reservoir at time t, expressed in m³ / s. up,t Q represents the downstream discharge flow rate of the upstream reservoir at time t, expressed in m³ / s; section,t This represents the interval runoff between the upstream and downstream reservoirs of the river during time period t, in m3 / s.

[0088] If the time interval between two adjacent load commands within the scheduling cycle is less than the water flow propagation time between upstream and downstream power stations, then the simulation of the river channel unit in the interval needs to consider the flow arrival lag time of the river channel, and the calculation formula for the inflow of the downstream reservoir is updated as follows:

[0089] I down,t =Q up,t-τ +Q section,t

[0090]

[0091] t1≤t-τ≤t2

[0092] In the formula, I down,t Q represents the inflow rate to the downstream reservoir at time t, in m³ / s; τ represents the time it takes for the outflow from the upstream reservoir to reach the downstream power station, in seconds; up,t- τ represents the outflow from the upstream reservoir during the period t-τ, in m³ / s; and These represent the outflow from the upstream reservoir at times t1 and t2, respectively, in m³ / s.

[0093] In this embodiment of the application, step S2 involves constructing a load allocation model based on constraint relationships, building a solution framework, designing the environment state, reward function, and agent of the solution framework, and training the agent, including the following steps B1-B5:

[0094] B1: Construct a load allocation model based on the constraint relationships;

[0095] B2: Build the solution framework and design the environment state, reward function, and agent of the solution framework;

[0096] B3: The agent reads the state information and uses it as input, selects and executes actions, obtains timely feedback and the state of the next time period, and generates samples to be stored in the experience replay pool.

[0097] B4: Input a portion of the data from the sample into the reinforcement learning policy network and the reinforcement learning value network respectively, and calculate the difference between the target value and the current value;

[0098] B5: Input the difference between the calculated target value and the current value into the reinforcement learning policy network, and use an optimization algorithm to update the parameters of the reinforcement learning policy network.

[0099] Specifically, in step B1, constructing the load allocation model based on the constraint relationships is manifested as follows:

[0100] With the primary objective of minimizing water consumption, a load distribution model among cascade hydropower stations that takes into account constraints such as water balance constraints, flow balance constraints, power generation flow constraints, downstream flow constraints, water level constraints, water level fluctuation constraints, power output balance constraints, power station output constraints, power station output amplitude constraints, and inter-station load transfer is constructed.

[0101] With the primary objective of minimizing water consumption, a load allocation model for cascade hydropower stations that balances power dispatch and water dispatch was constructed. The specific objective function is as follows:

[0102]

[0103] In the formula, f1 represents the objective function of the load allocation model among cascade hydropower stations; T represents the total number of scheduling periods; Δt represents the average power generation flow of power plant i during the time period t, in m³ / s; Δt is the length of the calculation period, in seconds.

[0104] The specific constraints of the model are as follows:

[0105] Water balance constraints

[0106]

[0107] In the formula, V i,t+1 and V i,t These represent the initial reservoir capacity of reservoir i in time period t+1 and time period t, respectively, in m3; This represents the power generation flow rate of reservoir i during time period t, in m³ / s; Represents the flood discharge rate of reservoir i during time period t, in m³ / s; i,t Δt represents the inflow rate of reservoir i during time period t, in m³ / s; Δt represents the duration of the calculation period, in seconds.

[0108] Output balance constraints:

[0109] P t =N t

[0110]

[0111] In the formula, P t N represents the total load command of the cascade hydropower system during time period t, in MW;t N represents the total output of the cascade hydropower system in time period t, in MW; i,t The output of power station i in time period t is expressed in MW; i represents the number of the cascade hydropower station; n represents the total number of power stations included in the cascade hydropower system.

[0112] Flow balancing constraints:

[0113]

[0114] In the formula, Q i,t This represents the discharge flow rate of power station i during the time period t, expressed in m³ / s.

[0115] Power generation flow constraints:

[0116]

[0117] In the formula, This represents the maximum permissible power generation flow of power station i, expressed in m³ / s.

[0118] Downflow constraint:

[0119] Q i,min ≤Q i,t ≤Q i,max

[0120] In the formula, Q i,min Q represents the minimum discharge flow of reservoir i, which is usually selected as the discharge flow rate that meets basic ecological requirements, and is expressed in m³ / s; i,max This represents the maximum permissible discharge flow of reservoir i. It is usually selected as the maximum permissible discharge flow to ensure the flood control safety of the reservoir itself and the downstream flood control control point. The unit is m3 / s.

[0121] Water level constraints:

[0122] Z i,min ≤Z i,t ≤Z i,max

[0123] In the formula, Z i,min Z represents the minimum water level upstream of reservoir i, usually the dead water level, in meters (m); i,t Z represents the upstream water level of reservoir i at the beginning of time period t, in meters. i,max This represents the maximum allowable water level of reservoir i. During the flood control restriction period, it is the flood control restriction water level, and at other times it is the normal storage water level. The unit is meters (m).

[0124] Water level fluctuation constraints:

[0125] |Z i,t -Z i,t-1 |≤ΔZi,max

[0126] In the formula, Z i,t and Z i,t-1 ΔZ represents the upstream water level at the beginning of time period t and time period (t-1) of reservoir i, respectively, in meters; i,max This represents the maximum permissible water level difference between adjacent time periods for reservoir i, expressed in meters (m). Given the large regulating capacity of the annual regulating power station and its strong adaptability to complex environments, the core of this constraint is to control the water level fluctuations of the daily regulating power station with relatively poor regulating capacity within the cascade hydropower system, thereby achieving early warning of water level anomalies and ensuring the overall economic efficiency and safety of the cascade hydropower operation.

[0127] Power plant output constraints:

[0128] N i,min ≤N i,t ≤N i,max

[0129] In the formula, N i,min N represents the minimum output limit of power plant i, in MW; i,max This indicates the maximum output limit of power plant i, which is usually the installed capacity of the power plant, in MW.

[0130] Power plant output amplitude constraints:

[0131] |N i,t -N i,t-1 |≤ΔN i

[0132] In the formula, N i,t and N i,t-1 ΔN represents the power output of power station i in time period t and time period t-1, respectively, in MW; i This represents the maximum allowable output variation of power plant i, in MW. This constraint is used to ensure the feasibility of the load distribution scheme output by the model and to smooth the power plant output process, preventing sudden large changes from impacting the power grid.

[0133] Inter-station load transfer constraints:

[0134]

[0135] In the formula, ΔP represents the load transfer of each power station in a cascade hydropower station, in MW. t This represents the maximum allowable output variation of the cascade power stations in time period t, calculated from the total load command of the cascade power stations, and is expressed in MW. This constraint is used to avoid frequent unit start-ups and shutdowns that may result from large-scale load transfers between power stations.

[0136] Specifically, step B2 involves building the solution framework, designing the environment state, reward function, and agent of the solution framework, which are manifested in the following ways:

[0137] We build an efficient solution framework for DQN, setting state variables that are dynamically updatable and cover the environmental interaction information required for agent decision-making; and setting a reward function that satisfies the objective, with a larger reward function value resulting in better optimization; then we construct an action space suitable for the current environment and select the action with the highest value in the current state within the action space using an appropriate greed rate.

[0138] Environment state design for the efficient DQN solution framework:

[0139] As a set of key parameters describing the specific environment in which the agent is located, state variables need to comprehensively cover the environmental interaction information required for the agent's decision-making, and dynamic relationships need to be established between state variables to achieve state updates.

[0140] Taking into account the design requirements of state variables and the characteristics of load distribution among cascade hydropower stations, the output N of each power station at each stage is determined based on the current time period. i,t Inflow rate of each reservoir I i,t Initial water level Z of the time period i,t Based on this information and the hydropower calculation formula for hydropower stations, the power generation flow rate of each power station in the current period is calculated. and the initial water level Z in the next period i,t+1 Trial calculations were performed. To further improve the model's calculation speed, considering that the change in hydropower head between adjacent time periods is relatively small at the real-time scale, the hydropower head H of the previous time period can be used. i,t-1 As of the current period, the hydropower head H i,t Preliminary calculations were performed using substituting the values ​​to reduce the number of calculations required for "electricity-based water supply". This resulted in a one-dimensional tensor containing 10 elements. s as the state variable during time period t t When a cascade hydropower system contains n power stations, the states of each power station are integrated by stacking and splicing them together, finally forming a multidimensional tensor containing n×10 elements to represent the state of the entire cascade hydropower system.

[0141] Reward function design for an efficient DQN solution framework:

[0142] The reward function is used to calculate the immediate benefit that the agent can obtain from the environment after selecting and executing a specific action in the current state. This invention sets minimizing water consumption as the basic reward term. Simultaneously, operational constraints such as the water level control requirements of cascade power stations are embedded, and a dynamic penalty mechanism applies negative feedback to behaviors that exceed these limits. The basic reward term drives the agent to prioritize output combinations with lower water consumption, while the penalty term suppresses the generation of infeasible strategies. A larger reward value results in better optimization. To meet this criterion, this invention sets the basic reward term as the difference between the maximum power generation flow of each cascade power station and the power generation flow of the current time period. The specific definition of the reward function is as follows:

[0143]

[0144] P t 1 =P1+P2

[0145]

[0146] In the formula, r t 1 (s t ,a t ) represents the reward function of the load allocation model among cascade hydropower stations; P represents the base reward for time period t; t 1 ΔZ represents the penalty value for time period t. i,t~t+1 denoted as , where represents the water level difference between adjacent time periods at power station i; C represents the penalty parameter.

[0147] Agent design for an efficient DQN solution framework:

[0148] An intelligent agent is an entity with learning capabilities within a deep reinforcement learning framework, used to select and execute actions, interact with the environment, and receive feedback from it. Its optimization objective is to progressively learn behavioral strategies that maximize long-term cumulative rewards based on sample information generated during training. It mainly comprises two parts: an action space and an action selection strategy.

[0149] Actions are specific decision instructions generated by an intelligent agent based on the current environmental state and action selection strategy. For cascade hydropower systems, to enhance the fitting effect of neural networks and improve optimization efficiency, an action space can be constructed based on the output change interval ΔN between adjacent decision stages of upstream and downstream power stations in the cascade system, thereby reducing the number of actions and narrowing the action space.

[0150] Action selection strategy refers to the rules or methods by which an agent selects actions in a specific state. In the DQN model, the agent must possess two key capabilities: First, the agent needs to be able to memorize and execute the optimal action, that is, select the action that brings the maximum expected reward in a given state; second, the agent needs to explore in the face of unknown or uncertain situations in order to discover potentially better strategies. Therefore, the ε-greedy method is adopted as the agent's action selection strategy. The agent randomly samples actions in the action space with a probability of ε, and selects the action with the highest value estimated by the value function in the current state with a probability of 1-ε. The specific formula is as follows:

[0151]

[0152] In the formula, π(a|s) represents the action selection strategy; |A| represents the number of actions available in the action space; and ε represents the greed rate.

[0153] To better align with the training requirements of intelligent agents in practical production applications and improve the approximation speed of the value function, this invention sets that after a certain number of training rounds, ε decreases with the increase of learning rounds, thereby reducing the sampling repetition rate and achieving a balance between exploration and utilization.

[0154] Specifically, in step B3, the agent reads the state information and uses it as input, selects and executes actions, obtains timely feedback and the state for the next time period, and generates samples to be stored in the experience replay pool. This is specifically manifested as follows:

[0155] During time period t, the agent reads the status information s of each cascade reservoir in the environment, such as the upstream water level and inflow. t As input to the model, with the goal of maximizing the value function, action a is selected and executed. t Timely feedback r is obtained based on the reward function. t and the next time period state s t+1 The above information is combined into a single sample {s}. t ,a t ,r t ,s t+1 Store it in the experience replay pool.

[0156] Specifically, in step B4, a portion of the sample data is input into the reinforcement learning policy network and the reinforcement learning value network, respectively, to calculate the difference between the target value and the current value. The reinforcement learning policy network and the reinforcement learning value network use ON networks and TN networks, respectively.

[0157] The state data of the sample s t and s t+1 Input the ON and TN networks respectively, and generate Q(s) based on the network parameters. t ,at ) and Q(s t+1 ,a t+1 ), and calculate the target Q value Q. π (s t ,a t ), and calculate the difference between it and the current Q value.

[0158] Q π (s t ,a t ) = r t (s t ,a t )+γmax Q(s t+1 ,a t+1 )

[0159] In the formula, Q π (s t ,a t ) indicates in s t In the given state, select and execute action a according to action selection strategy π. t The cumulative reward value generated afterward; γ represents the discount factor; maxQ(s) t+1 ,a t+1 ) indicates in s t+1 In the given state, the agent takes action a, calculated by the action value function Q(s,a). t+1 The maximum value that can be obtained later.

[0160] In an alternative implementation, when the inflow into the reservoir increases rapidly during the high-water season and the water level-output relationship is highly nonlinear, the reinforcement learning strategy network in step B4 can also adopt a deep feedforward neural network, taking "the upstream water level of each reservoir, the current inflow into the reservoir, and the power grid load demand" as inputs, and outputting the Q value of different scheduling actions, thereby accurately characterizing the nonlinear water level-power mapping relationship.

[0161] In another optional implementation, when the water inflow is consistently low during the dry season and it is necessary to consider both current and future scheduling, the reinforcement learning strategy network in step B4 can also use a long short-term memory network to take "the water level sequence, the trend of inflow runoff, and the power load demand sequence of the past several periods" as input and output the Q value of each action, so that the agent can take into account the water supply guarantee for future periods when formulating the current scheduling plan.

[0162] In an alternative implementation, in the scenario of a sudden increase in water inflow and drastic changes in state during the high-water season, the reinforcement learning value network in step B4 can also adopt soft update instead of the original periodic hard update. The soft update uses a weighted average method to gradually integrate the parameters of the online network into the target network.

[0163] In another optional implementation, when the water inflow is consistently low during the dry season and the reward signal is sparse, the reinforcement learning value network in step B4 can also adopt multi-step prediction, using the cumulative reward of multiple future steps as the target, to improve the sensitivity to long-term returns and ensure the long-term scheduling effect of water resources.

[0164] Specifically, in step B5, the difference between the calculated target value and the current value is input into the reinforcement learning policy network, and an optimization algorithm is used to update the parameters of the reinforcement learning policy network. The optimization algorithm specifically uses stochastic gradient descent.

[0165] The calculation results are passed to the ON network to calculate the loss function, and the parameters of the ON neural network are updated using stochastic gradient descent.

[0166] L(θ)=E[(r t (s t ,a t )+γmaxQ(s t+1 ,a t+1 ;θ')-Q(s t ,a t ;θ)) 2 ]

[0167] In the formula, L(θ) represents the loss function; θ represents the parameters of the ON neural network.

[0168] In an optional implementation, when the water inflow fluctuates drastically during the high-water season and the gradient changes significantly, the optimization algorithm used in step B5 to update the network parameters of the reinforcement learning strategy can also be adapted to use adaptive moment estimation, which introduces first-order and second-order moment estimation to adaptively adjust the learning rate for different parameters.

[0169] In another optional implementation, when the grid load is a continuous value and the output allocation accuracy requirement is high, the optimization algorithm used in step B5 to update the network parameters of the reinforcement learning strategy can also be AdaGrad, which adaptively scales the learning rate according to the historical gradient of the parameters. For parameters that are updated frequently, the learning rate is gradually reduced, while for parameters that are updated less frequently, a larger step size is maintained.

[0170] It should be noted that by constructing a load allocation model with the goal of minimizing water consumption and embedding multiple practical operational constraints, and designing a solution framework, an efficient and stable solution to the load allocation problem of cascade hydropower stations was achieved. This framework innovatively designs a tensor representation that integrates multi-dimensional state information of the system, introduces a reward function combining penalty mechanisms and basic rewards, and adopts an action selection strategy that balances exploration and utilization. This effectively guides the agent to learn feasible and efficient scheduling strategies. Through a dual-network structure and an experience replay mechanism, the agent can continuously learn from historical interaction samples, constantly narrowing the difference between the target value and the current value. It also updates network parameters using a stochastic gradient descent algorithm, ultimately achieving rapid model convergence and strategy optimization. This significantly improves the automation and computational efficiency of the load allocation process, achieving efficient utilization of hydropower resources while ensuring grid safety and ecological requirements.

[0171] In this embodiment of the application, step S3 solves the load allocation model based on the solution framework, selects the available load allocation schemes according to the constraints and allocation requirements, and outputs the final load allocation scheme by combining the water level control target and hydropower utilization efficiency, including the following steps C1-C4:

[0172] C1: Solve the load allocation model based on the solution framework, and select the available load allocation schemes according to the constraints and allocation requirements;

[0173] C2: Iteratively calculate the processing interval, back-calculate the outflow of the current period and the reservoir water level at the beginning of the next period, and calculate the reward value of the current state;

[0174] C3: Input the original state into the artificial neural network to obtain the value of the state, update the state, input the updated state into the artificial neural network again to obtain the value of the corresponding state, and update the parameters of the artificial neural network.

[0175] C4: If the sum of the output rewards tends to converge after the round termination condition and the learning termination condition are met, then the final load allocation scheme is output.

[0176] Specifically, in step C1, the load allocation model is solved based on the solution framework, and the selectable load allocation schemes are screened according to the constraints and allocation requirements. This is specifically reflected in:

[0177] Based on the output constraints of each power station and the total output plan of the cascade, all selectable load allocation schemes are screened and output.

[0178] Read the basic scheduling information of each reservoir, and input the water level information of each power station in the cascade hydropower system for the initial time period [Z] 1,str Z 2,str Z 3,str ,...,Zn,str ]、Inflow runoff[I 1,str ,I 2,str ,I 3,str ,...,I n,str ], Power grid load command sequence [N1, N2, N3, ..., N T [Establish basic information such as], determine the calculation period T, the iteration round L, and set t=0, l=1;

[0179] Based on the current status information and the output constraints of each power station, the selectable action range is filtered and the action is selected to determine the output range of each level of downstream power station, t = t + 1.

[0180] Specifically, in step C2, the processing interval is iteratively calculated to back-calculate the outflow rate for the current period and the reservoir water level at the beginning of the next period, and the reward value for the current state is calculated as follows:

[0181] The output range is discretized with an accuracy of 0.5MW and iteratively calculated. In the initial calculation, the upper limit of the corresponding output range is taken for each downstream regulating hydropower station in the cascade hydropower series. Based on the current output plan, the output of the corresponding leading hydropower station is calculated, and the current output calculation results of each power station are recorded [N]. 1,t N 2,t N 3,t ,...,N n,t ],

[0182] Based on the initial upstream water levels of each reservoir [Z] at the beginning of the current period 1,t Z 2,t Z 3,t ,...,Z n,t Inbound flow [I] 1,t ,I 2,t ,I 3,t ,...,I n,t Load distribution [N] 1,t N 2,t N 3,t ,...,N n,t Information such as [Q] is used to inversely deduce the outflow from the reservoir for each power station during the current period using the hydropower station's hydropower calculation formula. 1,t Q 2,t Q 3,t ,...,Q n,t [Z] and the reservoir water level at the beginning of the next period 1,t+1 Z 2,t+1 Z 3,t+1 ,...,Z n,t+1 ],

[0183] Calculate the reward value r corresponding to the current state based on the current load distribution. t Determine r tIs this the optimal value with 0.5MW accuracy within the current output range? If so, record the current value for r. t If the value and the operating status of each corresponding power station are not specified, adjust the output combination of the downstream daily regulating power station and repeat step C2 until the output range is traversed.

[0184] Specifically, in step C3, the original state is input into the artificial neural network to obtain the value of the state, and the state is updated. The updated state is then input into the artificial neural network again to obtain the corresponding value of the state, and the parameters of the artificial neural network are updated. The artificial neural network used is a deep neural network.

[0185] The original state is input into a deep neural network to obtain the value Q of that state. t-1 Based on the calculation results in Step 5, update the state and calculate the cumulative reward value r for this round. sum =r sum +r t The updated state is then input into the neural network to obtain the value Q of the state after the action is performed. t Combined with Q t-1 Q t and r t The loss function is calculated, and the gradient descent method is used to update the parameters of the main neural network. The parameters are then assigned to the target neural network at regular time steps to achieve the agent's "autonomous learning".

[0186] Determine whether the cycle termination condition t=T has been met. If it is met, proceed to determine whether the learning termination condition has been met. Otherwise, return to step C1 to filter the selectable action range and select actions based on the current status information and the output constraints of each power station, and determine the output range of each level of downstream power station, t=t+1.

[0187] Determine if the learning termination condition m = M has been met. If so, output the total reward r for each round of learning. sum Otherwise, return to step C1, iteration number l = l + 1.

[0188] Output the total reward r for each learning round. sum If r sum If the model tends to converge, it completes the learning process and outputs the current scheduling scheme; otherwise, it increases the number of iterations and continues learning.

[0189] In an alternative implementation, when the scheduling process has obvious time series characteristics, the artificial neural network in step C3 can also adopt an LSTM network, using the historical state sequence as the input sequence. The LSTM unit remembers the long-term dependency and predicts the value of the current state.

[0190] In another alternative implementation, when there is a topological relationship between multiple power stations, the artificial neural network in step C3 can also adopt a graph neural network, which models the power stations as nodes in the graph, with water flow coupling as edges, and propagates the information of adjacent power stations through graph convolution to update the state value of each node.

[0191] In summary, this invention generalizes the cascade system into power station and river units based on hydraulic connections and inflow rates. It employs a rigorous water balance principle and a "water-to-electricity" model for precise simulation, particularly considering interval runoff and flow lag effects, thus laying a reliable physical foundation for the optimization model. Furthermore, it constructs a load allocation model with the objective of minimizing water consumption and embedding various electrical and water regulation constraints. An innovative DQN solution framework, incorporating state variables, reward functions, and action strategies, is designed. Stable and efficient agent training is achieved through a dual-network structure and experience replay mechanism. Finally, through iterative calculation and constraint selection, an optimized allocation scheme is output that significantly improves hydropower utilization efficiency, stabilizes water level control, and ensures coordinated operation of the cascade systems while meeting grid load demands. This scheme possesses strong engineering practicality and promotional value.

[0192] Example 3: The above is an illustrative scheme of a real-time load optimization allocation method among cascade hydropower stations.

[0193] The storage medium proposed in this embodiment and the method for real-time load optimization allocation between cascade hydropower stations proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0194] Based on the above description of the implementation methods, those skilled in the art will clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0195] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for real-time load optimization allocation among cascade hydropower stations, characterized in that, include: Based on hydraulic relationships and inflow, the hydropower station group system is generalized, and the cascade reservoir group is simulated; A load allocation model is constructed based on the constraint relationship, and a solution framework is built. The environment state, reward function, and agent of the solution framework are designed, and the agent is trained. The load allocation model is solved based on the solution framework. Selectable load allocation schemes are selected according to constraints and allocation requirements. Finally, the load allocation scheme is output by combining water level control objectives and hydropower utilization efficiency.

2. The method for real-time load optimization allocation among cascade hydropower stations as described in claim 1, characterized in that, Simulation of the cascade reservoir group includes: Simulate the power plant unit; Simulation of the river channel unit within the interval.

3. The method for real-time load optimization allocation among cascade hydropower stations as described in claim 2, characterized in that, The training of the agent includes: The agent reads the state information and uses it as input, selects and executes actions, obtains timely feedback and the state of the next time period, and generates samples to be stored in the experience replay pool. Input a portion of the data from the sample into the reinforcement learning policy network and the reinforcement learning value network respectively, and calculate the difference between the target value and the current value; The difference between the calculated target value and the current value is input into the reinforcement learning policy network, and the parameters of the reinforcement learning policy network are updated using an optimization algorithm.

4. The method for real-time load optimization allocation among cascade hydropower stations as described in claim 3, characterized in that, The final load allocation scheme, which combines water level control targets and hydropower utilization efficiency, includes: The processing interval is iteratively calculated to back-calculate the outflow rate of the current period and the reservoir water level at the beginning of the next period, and to calculate the reward value of the current state. The original state is input into the artificial neural network to obtain the value of the state, and the state is updated. The updated state is input into the artificial neural network again to obtain the value of the corresponding state, and the parameters of the artificial neural network are updated. If the sum of the output rewards tends to converge after the round termination condition and the learning termination condition are met, then the final load allocation scheme is output.

5. The method for real-time load optimization allocation among cascade hydropower stations as described in claim 4, characterized in that, The simulation of the power plant unit includes: The power station units of the cascade hydropower system are simulated. During operation, the power generation flow and flood discharge flow of the power station are superimposed to form the total discharge flow of the power station, and the dynamic changes of the discharge flow must strictly follow the principle of reservoir water balance. In the formula, V t+1 and V t These represent the initial reservoir storage volumes at time t and time t+1, respectively, in m³; t This represents the inflow rate of the reservoir during time period t, expressed in m³ / s. This represents the power generation flow rate during time period t, expressed in m³ / s. Δt represents the flood discharge rate during time period t, in m³ / s; Δt represents the duration of the calculation period, in seconds. The specific scheduling parameters of each power plant unit are mainly determined through trial calculations using the "electricity-based water supply" model. The specific simulation method is as follows: The output of each power station in the cascade is calculated based on the grid load command and the power balance equation: In the formula, P t P represents the total load command of the cascade hydropower system during time period t, i.e., the load value to be completed, in MW; i,t This represents the load value that power station i needs to bear during time period t, in MW; i represents the number of the cascade hydropower station; n represents the total number of power stations included in the cascade hydropower system; Based on the "electricity-determined water supply" model, the power generation flow of the hydropower station is calculated from its output equation: In the formula, i represents the designation of the cascade hydropower station, i = 1, 2, ..., n; N i,t K represents the power output of power station i in time period t, in MW; i This represents the output coefficient of power station i; H represents the power generation flow rate in time period t, in m³ / s. i,t Let be the average hydropower head of power station i during time period t, in meters.

6. The method for real-time load optimization allocation among cascade hydropower stations as described in claim 5, characterized in that, The simulation of the interval river unit includes: The river channel unit was simulated, and the river was divided into several segments based on the geographical distribution characteristics of the power stations. The hydrological processes of each segment showed significant spatial correlations: I down,t =Q up,t +Q section,t In the formula, I down,t Q represents the inflow rate to the downstream reservoir at time t, expressed in m³ / s. up,t Q represents the downstream discharge flow rate of the upstream reservoir at time t, expressed in m³ / s; section,t This represents the interval runoff between the upstream and downstream reservoirs of the river during time period t, in m3 / s. If the time interval between two adjacent load commands within the scheduling cycle is less than the water flow propagation time between upstream and downstream power stations, then the simulation of the river channel unit in the interval needs to consider the flow arrival lag time of the river channel, and the calculation formula for the inflow of the downstream reservoir is updated as follows: I down,t =Q up,t-τ +Q section,t t1≤t-τ≤t2 In the formula, I down,t Q represents the inflow rate to the downstream reservoir at time t, in m³ / s; τ represents the time it takes for the outflow from the upstream reservoir to reach the downstream power station, in seconds; up,t-τ This represents the outflow from the upstream reservoir during the time period t-τ, expressed in m³ / s. and These represent the outflow from the upstream reservoir at times t1 and t2, respectively, in m³ / s.

7. The method for real-time load optimization allocation among cascade hydropower stations as described in claim 6, characterized in that, The intelligent agent reads state information as input, selects and executes actions, obtains timely feedback and the state for the next time period, generates samples and stores them in the experience replay pool, including: During time period t, the agent reads the status information s of each cascade reservoir in the environment, such as the upstream water level and inflow. t As input to the model, with the goal of maximizing the value function, action a is selected and executed. t Timely feedback r is obtained based on the reward function. t and the next time period state s t+1 The above information is combined into a single sample {s}. t ,a t ,r t ,s t+1 Store it in the experience replay pool.

8. The method for real-time load optimization allocation among cascade hydropower stations as described in claim 7, characterized in that, The step of inputting a portion of the sample data into the reinforcement learning policy network and the reinforcement learning value network respectively, and calculating the difference between the target value and the current value, includes: The state data of the sample s t and s t+1 Input the ON and TN networks respectively, and generate Q(s) based on the network parameters. t ,a t ) and Q(s t+1 ,a t+1 ), and calculate the target Q value Q. π (s t ,a t ), and calculate the difference between it and the current Q value. Q π (s t ,a t )=r t (s t ,a t )+γmaxQ(s t+1 ,a t+1 ) In the formula, Q π (s t ,a t ) indicates in s t In the given state, select and execute action a according to action selection strategy π. t The cumulative reward value generated afterward; γ represents the discount factor; maxQ(s) t+1 ,a t+1 ) indicates in s t+1 In the given state, the agent takes action a, calculated by the action value function Q(s,a). t+1 The maximum value that can be obtained later.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.