Method and device for hydroelectric dispatching based on real-time hydrology and knowledge reasoning

By using quantitative calculations based on real-time hydrological data and prior knowledge in the hydrological field, and designing a hybrid reward function, the problems of robustness and low collaborative efficiency in hydropower dispatching are solved, enabling efficient collaborative dispatching and rapid response of hydropower station groups.

CN122456641APending Publication Date: 2026-07-24GUODIAN DADU RIVER POWER ENG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUODIAN DADU RIVER POWER ENG
Filing Date
2026-03-25
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing hydropower dispatching methods lack robustness in complex and volatile non-stationary environments. Each power station agent tends to seek short-term individual power generation rewards, resulting in low collaborative efficiency and difficulty in forming a long-term stable and clearly defined global collaborative dispatching strategy, thus leading to low hydropower dispatching efficiency.

Method used

Based on real-time hydrological data and prior knowledge in the hydrological field, the cumulative deviation of multiple indicators of hydropower station agents during disturbances is quantitatively calculated, a global cooperative resilience score is established, collective cooperative reward items are inferred through preference relation inversion, and a decentralized scheduling strategy is trained using a hybrid reward function and a multi-agent reinforcement learning framework to achieve cooperative scheduling of hydropower station agents.

Benefits of technology

It improves the operational efficiency of hydropower station groups and the response speed to extreme working conditions. Through the design of automated and intelligent reward functions, it guides the division of labor and cooperation among power stations, solving the problems of low robustness and low cooperation efficiency in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122456641A_ABST
    Figure CN122456641A_ABST
Patent Text Reader

Abstract

The application provides a kind of water and electricity scheduling method and device based on real-time hydrology and knowledge reasoning, the method comprises: the cumulative deviation of a plurality of indexes of hydropower station agent during disturbance is quantified, global cooperation restoring force score is obtained to establish the preference relationship between each scheduling trajectory, and the preset parameterized reward model is deduced using the preference relationship, to obtain collective cooperation reward item, then combined with the power generation income signal of hydropower station individual to determine the mixed reward function, and then realize the strategy training of hydropower station agent, to obtain decentralized scheduling strategy, to control the execution mechanism corresponding to hydropower station agent to realize water and electricity scheduling.The method and device of the application can extract and quantify the cooperation incentive structure from complex scheduling behavior, realize the dynamic adjustment of reward function, and guide the division of labor and cooperation of each power station through the mixed reward mechanism, improve the operation efficiency of hydropower station group and the response speed of extreme working condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a hydropower scheduling method and apparatus based on real-time hydrology and knowledge reasoning. Background Technology

[0002] In modern energy systems, hydropower dispatching needs to meet the load demands of the power system while also taking into account multiple constraints such as flood control, water supply, and ecological protection. It also needs to coordinate the collaborative operation of multiple cascade power stations in complex river basin systems.

[0003] Hydropower dispatching is essentially a multi-objective, dynamic, and highly uncertain multi-agent system (MAS) optimization problem. In actual operating environments, various hydropower stations form complex coupling relationships through water flow time delays, water level constraints, and power generation capacity limitations. The dispatching decision of any one station can have a cascading impact on upstream and downstream stations. Therefore, how to achieve efficient, stable, and adaptive dispatching strategies under complex hydrological environments and multi-agent collaborative conditions has always been an important research problem in the field of hydropower dispatching.

[0004] In related technologies, hydropower scheduling typically relies on human experience to pre-set fixed static reward functions for agents (e.g., setting only the maximization of single-station power generation revenue or simple water level exceedance penalties) to guide model training. This method is difficult to accurately quantify and evaluate the system's resistance and recovery capabilities when facing disturbances such as extreme hydrological fluctuations, resulting in a lack of robustness of the model in complex and volatile non-stationary environments. In addition, during training, each power station agent is prone to over-consuming reservoir capacity to obtain short-term individual power generation rewards, thereby triggering subsequent global water shortages and resource depletion. It is difficult to spontaneously form a long-term stable and clearly defined global collaborative scheduling strategy, which in turn leads to low efficiency in hydropower scheduling. Summary of the Invention

[0005] This invention provides a hydropower scheduling method and apparatus based on real-time hydrology and knowledge reasoning, which solves the problem that existing technologies rely on human experience to pre-set a fixed static reward function for training decision models for intelligent agents. This method has low robustness in complex and ever-changing non-stationary environments, and each power station agent tends to obtain short-term individual power generation rewards during training, resulting in low cooperation efficiency and thus low hydropower scheduling efficiency.

[0006] This invention provides a hydropower scheduling method based on real-time hydrology and knowledge reasoning, comprising: Based on real-time hydrological data and prior knowledge in the hydrological field, the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period is quantitatively calculated to obtain a global cooperative resilience score of multiple hydropower dispatch trajectories; the multiple indicators include at least two of the following: cumulative output sustainability, resource availability, dispatch fairness index and shortage risk index. Based on the global collaborative resilience score, a preference relationship is established between each scheduling trajectory, and the preference relationship is used to invert and infer the preset parameterized reward model to obtain the collective collaborative reward item; A hybrid reward function is determined based on the collective collaboration reward item and the power generation revenue signal of the individual hydropower station. The hybrid reward function and the multi-agent reinforcement learning framework are used to train the hydropower station agents to obtain a decentralized scheduling strategy. Hydropower scheduling is achieved by using the decentralized scheduling strategy to control the execution mechanism corresponding to the hydropower station's intelligent agent.

[0007] According to the present invention, a hydropower dispatching method based on real-time hydrology and knowledge reasoning is provided. The method quantifies the cumulative deviation of multiple indicators of a hydropower station agent during disturbances based on real-time hydrological data and prior knowledge in the hydrological domain, obtaining a global cooperative resilience score for multiple hydropower dispatching trajectories, including: Based on the real-time hydrological data and prior knowledge in the hydrological field, the index thresholds are compared and rules are matched to determine the time of disturbance occurrence, the time of most severe degradation, and the recovery endpoint in the hydropower dispatch trajectory. Calculate the cumulative deviation of the multiple indicators from the normal baseline indicators during the period from the time of the disturbance to the time of the most severe degradation to obtain the corresponding failure profile; calculate the cumulative deviation of the multiple indicators from the normal baseline indicators during the period from the time of the most severe degradation to the recovery endpoint to obtain the corresponding recovery profile; The global collaborative resilience score is obtained by calculating the failure profile and the recovery profile corresponding to each indicator.

[0008] According to the present invention, a hydropower scheduling method based on real-time hydrology and knowledge reasoning is provided, wherein the global cooperative resilience score is obtained by calculating based on the failure profile and the recovery profile corresponding to each index, including: Obtain the first time interval from the time of the disturbance to the time of the most severe degradation, and the second time interval from the time of the most severe degradation to the end point of the recovery; Based on the first time interval, the second time interval, and the failure profile and recovery profile corresponding to each indicator, calculate the individual resilience score for each indicator. The harmonic mean is used to aggregate the individual resilience scores of each indicator to obtain the global collaborative resilience score.

[0009] According to the present invention, a hydropower dispatching method based on real-time hydrology and knowledge reasoning, wherein the method utilizes the preference relationship to inversely infer a preset parameterized reward model to obtain a collective cooperation reward item, includes: The loss function of the preset parameterized reward model is determined according to the preference learning algorithm and the preference relationship; wherein, the loss function is constructed based on the reward difference or preference probability between high-scoring scheduling trajectories and low-scoring scheduling trajectories; The parameters of the preset parameterized reward model are updated by minimizing the loss function, and the output of the updated preset parameterized reward model is used as the collective collaboration reward.

[0010] According to the present invention, a hydropower scheduling method based on real-time hydrology and knowledge reasoning is provided, wherein the preset parameterized reward model includes at least one of the following: A linear model of handcrafted knowledge based on predefined expert experience features; State linear model based on the characteristics of the original state variables; A nonlinear deep neural network model based on a multilayer perceptron.

[0011] According to the present invention, a hydropower dispatching method based on real-time hydrology and knowledge reasoning is provided, wherein determining a hybrid reward function based on the collective cooperation reward item and the power generation revenue signal of individual hydropower stations includes: Configure corresponding individual adjustment weights and collaborative adjustment weights for the power generation revenue signal and the collective cooperation reward item, respectively; The hybrid reward function is obtained by linearly weighting and summing the power generation revenue signal and the collective cooperation reward item based on the individual adjustment weight and the cooperative adjustment weight.

[0012] The present invention also provides a hydropower dispatching device based on real-time hydrology and knowledge reasoning, comprising: The knowledge reasoning and evaluation module is used to quantify the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period based on real-time hydrological data and prior knowledge in the hydrological field, and obtain a global cooperative resilience score of multiple hydropower dispatch trajectories; the multiple indicators include at least two of the following: cumulative output sustainability, resource availability, dispatch fairness index and shortage risk index. The reward configuration module is used to establish the preference relationship between each scheduling trajectory based on the global collaborative resilience score, and to use the preference relationship to perform inversion inference on the preset parameterized reward model to obtain the collective collaborative reward item; The scheduling strategy training module is used to determine a hybrid reward function based on the collective cooperation reward item and the power generation revenue signal of the individual hydropower station, and to use the hybrid reward function and the multi-agent reinforcement learning framework to train the hydropower station agent to obtain a decentralized scheduling strategy. The hydropower dispatching module is used to control the execution mechanism corresponding to the hydropower station intelligent agent to realize hydropower dispatching using the decentralized dispatching strategy.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the hydropower scheduling method based on real-time hydrology and knowledge reasoning as described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the hydropower scheduling method based on real-time hydrology and knowledge reasoning as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the hydropower scheduling method based on real-time hydrology and knowledge reasoning as described above.

[0016] The hydropower dispatching method and apparatus based on real-time hydrology and knowledge reasoning provided by this invention transforms hydrological experience into quantitative indicators and evaluates the performance of dispatching strategies under extreme environments, providing objective data for the subsequent automated inference of reward functions. Furthermore, it automatically extracts and quantifies the collaborative incentive structure that is difficult for experts to describe from complex dispatching behaviors based on global collaborative resilience scores, realizing the automation and intelligence of reward function design and solving the drawbacks of manually setting weights. Through a hybrid reward mechanism, it guides the division of labor and cooperation among power stations, improving the operating efficiency of the hydropower station group and the response speed to extreme conditions. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts of the hydropower scheduling method based on real-time hydrology and knowledge reasoning provided by the present invention.

[0019] Figure 2 This is the second flowchart of the hydropower scheduling method based on real-time hydrology and knowledge reasoning provided by the present invention.

[0020] Figure 3 This is a schematic diagram of the hydropower dispatching device based on real-time hydrology and knowledge reasoning provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0023] The following is combined with Figures 1-3 This invention describes a hydropower scheduling method based on real-time hydrology and knowledge reasoning.

[0024] Figure 1 This is one of the flowcharts illustrating the hydropower scheduling method based on real-time hydrology and knowledge reasoning provided by this invention, such as... Figure 1 As shown, the method includes the following steps: Step 110: Based on real-time hydrological data and prior knowledge in the hydrological field, quantify the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period to obtain the global cooperative resilience score of multiple hydropower dispatch trajectories; the multiple indicators include at least two of the following: cumulative output sustainability, resource availability, dispatch fairness index and shortage risk index.

[0025] In this step, real-time hydrological data can be dynamic information such as water levels of reservoirs at all levels, real-time inflow, rainfall forecasts, and unit operating status obtained through watershed telemetry stations.

[0026] In this step, prior knowledge in the hydrological field may include industry standards such as reservoir operation procedures, flood control limit water levels, ecological flow baselines, and grid load stability requirements. This knowledge is transformed into quantitative reasoning rules to identify key time points when hydropower dispatching systems (such as hydropower station agents) are disturbed, such as the moment of disturbance (e.g., a sudden increase in flow), the moment of most severe degradation (e.g., water level reaching the warning line), and the recovery endpoint (e.g., output stabilization).

[0027] In this embodiment, the cumulative deviation can be characterized by the failure profile and recovery profile of the hydropower station agent. That is, within the above time interval, the cumulative deviation of the integral area of ​​the actual operating indicators (such as output and reservoir capacity) deviating from the normal benchmark value can be the integral area of ​​the actual operating indicators (such as output and reservoir capacity) deviating from the normal benchmark value within the above time interval.

[0028] In this embodiment, multiple hydropower dispatch trajectories can be multiple sets of "state-action" sequences generated by the hydropower station intelligent agent in historical operation or simulation.

[0029] In this embodiment, the Markov game modeling process of the hydropower dispatching agent is as follows: (1) First, the joint scheduling problem of the cascade hydropower station group is abstracted into a fully observed multi-agent Markov game, whose formal definition is given by the six-tuple. The composition, the elements in a six-tuple, are represented as follows: ①Agent set : Indicates the cascade basin A single or collaborative hydropower station node.

[0030] ②State Space Global environment status In hydropower scenarios, the state vector contains the real-time water levels of each reservoir. Inbound flow Outbound flow Unit output and grid demand load .

[0031] ③ Joint Action Space : Each power station agent Implement a decentralized strategy ,action This includes regulating flow through gates and allocating power generation load.

[0032] ④ State transition function : It describes the evolution of reservoir storage and release based on hydrological principles.

[0033] ⑤ Reward Structure This is the core research object of the present invention, namely the reward function dynamically generated by the knowledge reasoning module.

[0034] ⑥ Discount Factor : This is used to balance short-term power generation benefits with long-term system resilience.

[0035] In this embodiment, based on prior knowledge in areas such as hydropower dispatching procedures, hydrological evolution patterns, and cascade coordination constraints, four categories of multiple indicators with inference quantification rules can be defined to form a real-time indicator set. Each indicator value is automatically calculated by the knowledge reasoning engine in combination with real-time hydrological data, replacing the traditional manual assignment.

[0036] Among them, cumulative output sustainability is a power fluctuation tolerance threshold inferred from knowledge of stable grid load supply. It is used to compare the actual power output of the power plant with the benchmark power output in real time, and to infer the output maintenance rate and disturbance recovery speed. The calculation formula is as follows: ; in, The baseline output is derived from historical steady hydrological knowledge. To provide real-time output, The closer the value is to 1, the more stable the output force.

[0037] In this embodiment, resource availability is a reservoir safety capacity threshold inferred from knowledge of reservoir flood control limits, ecological water storage, and dry season water supply. It is used to infer the matching degree between total reservoir capacity and safety capacity in real time, and to provide early warning of excessive water storage / discharge risks. The calculation formula is as follows: ; in, The lower limit of safe reservoir capacity derived from hydrological knowledge For real-time total storage capacity, .

[0038] Specifically, the lower limit of safe storage capacity The reasoning process is as follows: First, based on the reservoir operation regulations, ecological flow guarantee standards, water supply guarantee requirements, and turbine unit operation constraints, the flood control baseline reservoir capacity is extracted. Minimum ecological reservoir capacity required Water supply guarantee reservoir capacity and the minimum head constraint reservoir capacity of the generating unit The maximum value among the four basic constraint capacities is taken as the basic safety capacity. Secondly, by combining real-time hydrological data and the forecast results of inflow into the reservoir in the foreseeable future, a hydrological trend correction coefficient is calculated. The basic safety storage capacity is dynamically adjusted to obtain the adjusted storage capacity. ,in The system adaptively adjusts based on the inflow trends; finally, for the cascade reservoir group scheduling scenario, it incorporates upstream and downstream flow time delays, water level connections, and collaborative operation constraints to obtain the minimum reservoir capacity under cascade collaborative constraints. The maximum value of the modified storage capacity and the minimum storage capacity under the tiered collaborative constraint is taken as the lower limit of the safe storage capacity obtained by the final reasoning. .

[0039] In this embodiment, the dispatch fairness index is a load allocation fairness threshold inferred based on the knowledge of coordinated dispatch of cascade power stations. It is used to infer the degree of balance in power output and reservoir capacity utilization among power stations based on the Gini coefficient model. The calculation formula is as follows: ; in, The Gini coefficient warning value is determined by knowledge-based reasoning. .

[0040] The shortage risk index is a load gap tolerance threshold derived from knowledge of power grid supply and demand balance and water supply security. It is used to infer the frequency, duration, and risk level of output gaps in real time. The calculation formula is as follows: ; in, This represents the grid demand load (derived from knowledge of electricity consumption). For the overall contribution, The smaller the value, the lower the risk.

[0041] In this embodiment, the system calculates the cumulative output sustainability index (the degree of dispersion between actual output and baseline output) and the resource availability index (the degree of matching between remaining reservoir capacity and safe reservoir capacity) under the hydropower dispatch trajectory. For example, if a dispatch trajectory exceeds the reservoir capacity limit in order to generate electricity during a flood, its "failure profile" area will increase sharply. Then, by harmonic averaging and aggregating the individual resilience scores of these indicators, the system finally obtains the global cooperative resilience score of the group of dispatch trajectories.

[0042] Step 120: Establish the preference relationship between each scheduling trajectory based on the global collaborative resilience score, and use the preference relationship to invert and infer the preset parameterized reward model to obtain the collective collaborative reward item.

[0043] In this step, the global collaborative resilience score is used to rank the different scheduling trajectories to establish a preference relationship. For example, if the score of trajectory A is higher than that of trajectory B, then A is determined to be better than B.

[0044] In this step, the preset parameterized reward model can be a neural network composed of a multilayer perceptron (MLP), which is used to learn the nonlinear mapping between state features (such as water level variability and load gap) and resilience output.

[0045] In this embodiment, inverse reinforcement learning algorithms (such as marginal preference learning MPL or probabilistic preference learning PPL) can be used to automatically solve for the reward function parameters that can represent expert preferences and collaboration logic, i.e., the collective collaboration reward item, by minimizing the reward prediction loss between high-scoring trajectories and low-scoring trajectories.

[0046] For example, in a flood control scenario, the hydropower station agent collects multiple sets of control trajectories. Trajectory A adopts a "cascade-style peak shaving" strategy, resulting in a very high resilience score. Trajectory B adopts a "self-protection and rapid release" strategy, leading to downstream water levels exceeding the standard and a lower score. The hydropower station agent uses a preference learning algorithm to identify the implicit collaborative features in trajectory A and adjusts the weight parameters of the neural network model accordingly. For example, if the model finds that the upstream reservoir capacity and the downstream flood discharge capacity reach a certain balance during the inversion process, the system will give a higher reward. Finally, the model outputs a collective collaborative reward item that can calculate the collaborative value corresponding to the current state in real time.

[0047] Step 130: Determine the hybrid reward function based on the collective collaboration reward item and the power generation revenue signal of the individual hydropower station, and use the hybrid reward function and the multi-agent reinforcement learning framework to train the hydropower station agents to obtain the decentralized scheduling strategy.

[0048] In this step, the individual power generation revenue signal can represent the instinct of a single power plant to pursue economic benefits (such as real-time electricity sales profits, avoiding fines, etc.).

[0049] In this step, the hybrid reward function integrates the collective cooperation reward and individual income by adjusting parameters, so that each power station agent considers both its own profit and system security when making decisions.

[0050] In this step, multi-agent reinforcement learning frameworks include, but are not limited to, algorithms such as PPO (Proximal Policy Optimization) or QMIX (Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning). This framework provides an evolutionary environment that allows multiple agents (hydropower station agents) to autonomously learn the optimal action plan through multiple iterative interactions driven by mixed rewards.

[0051] In the aforementioned flood control scenario, by substituting the hybrid reward function into the PPO algorithm, the total reward received by the upstream power station agent A includes not only its current power generation revenue but also a collaborative term inferred from global resilience. If power station A disregards downstream safety in order to generate more power for itself, the collective collaborative term it receives will decrease significantly, resulting in a loss of total reward. After a long period of strategy training, the upstream power station agent gradually learns to proactively empty reservoir capacity before the flood arrives, while the downstream power station agent learns to coordinate and adjust its output. Ultimately, the system generates a decentralized control strategy that allows each power station to make independent decisions based on local hydrological information.

[0052] Step 140: Use a decentralized scheduling strategy to control the execution mechanism corresponding to the hydropower station's intelligent agent to achieve hydropower scheduling.

[0053] In this step, the actuator corresponding to the hydropower station intelligent agent is the actual physical control system of the hydropower station. For example, the actuator includes the load distribution mechanism of the generator set and the opening adjustment device of the flood discharge gate.

[0054] In this embodiment, the hydropower station intelligent agent uses a decentralized scheduling strategy to analyze the local information such as water level, flow rate, and load demand observed by the current power station in real time, and outputs the optimal action command to guide the hydropower station's actuators to take corresponding actions.

[0055] The hydropower dispatching method based on real-time hydrology and knowledge reasoning provided in this invention transforms hydrological experience into quantitative indicators and evaluates the performance of dispatching strategies under extreme environments, providing objective data for the subsequent automated inference of reward functions. Furthermore, it automatically extracts and quantifies the collaborative incentive structure, which is difficult for experts to describe, from complex dispatching behaviors based on global collaborative resilience scores, thus achieving automation and intelligence in reward function design and overcoming the drawbacks of manually setting weights. By guiding the division of labor and cooperation among power stations through a hybrid reward mechanism, it improves the operational efficiency of the hydropower station group and the response speed to extreme conditions.

[0056] In some embodiments, based on real-time hydrological data and prior knowledge in the hydrological domain, the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period is quantitatively calculated to obtain a global cooperative resilience score for multiple hydropower dispatch trajectories. This includes: comparing indicator thresholds and matching rules based on real-time hydrological data and prior knowledge in the hydrological domain to determine the disturbance occurrence time, the time of most severe degradation, and the recovery endpoint in the hydropower dispatch trajectory; calculating the cumulative amount of deviation of multiple indicators from the normal benchmark indicator during the period from the disturbance occurrence time to the time of most severe degradation to obtain the corresponding failure profile; calculating the degree of cumulative deviation of multiple indicators from the normal benchmark indicator during the period from the time of most severe degradation to the recovery endpoint to obtain the corresponding recovery profile; and calculating the global cooperative resilience score based on the failure profile and recovery profile corresponding to each indicator.

[0057] In this embodiment, the index threshold comparison and rule matching refers to the process of comparing the real-time collected dynamic hydrological data with the pre-set scheduling procedures in the system (such as prior knowledge such as flood control limit water level and output fluctuation tolerance). Through comparison, the system can accurately divide the time axis into three time nodes: the moment of disturbance occurrence (representing the starting point when the system begins to be subjected to external shocks (such as the arrival of flood peaks) and the first time the operating index exceeds the limit), the moment of most severe degradation (representing the extreme point where the system is in the worst state and deviates the farthest from the normal level during the process of resisting disturbances), and the recovery endpoint (representing the moment when the system returns to the stable baseline state through adjustment).

[0058] In this embodiment, the failure profile and the recovery profile are essentially integral areas over time. The former quantifies the cumulative amount of loss or risk accumulated by the system from the moment of disturbance to the moment of most severe degradation, while the latter quantifies the degree of recovery sluggishness from the moment of most severe degradation to the end of recovery. Based on the area size of these two profiles, this embodiment performs a comprehensive calculation to obtain a global collaborative resilience score that objectively reflects the system's shock resistance and self-healing ability.

[0059] In this embodiment, in order to accurately evaluate the system resilience of the scheduling strategy under complex scenarios such as extreme hydrological fluctuations and equipment disturbances, this embodiment integrates hydrological knowledge and real-time operational data to construct a knowledge-reasoning-driven collaborative resilience index system. Through a three-layer quantitative reasoning mechanism of knowledge rule reasoning, index threshold reasoning, and trajectory state reasoning, the system achieves automated and accurate evaluation of the quality of the scheduling trajectory.

[0060] For example, in the aforementioned flood control scenario, the knowledge reasoning engine compares the inflow to historical benchmarks in real time. It detects that the inflow suddenly exceeds the warning threshold. The hydropower station's intelligent agent immediately marks this point as the disturbance occurrence. As the flood peak continues to flow in, upstream reservoirs are forced to reduce power generation to intercept the flood, causing the reservoir water level to continuously approach and even briefly touch the maximum flood control limit. Simultaneously, the power grid load gap reaches its maximum. The hydropower station's intelligent agent automatically captures this extreme point and marks it as the most severe degradation moment. Subsequently, each power station coordinates flood discharge and gradually restores unit output. When the water level falls back to the safe operating range and the flood discharge... When the force meets the baseline, it is marked as the recovery endpoint; the hydropower station agent then performs quantitative calculations: the time integral of "the difference between the actual available reservoir capacity and the safe reservoir capacity" in the first stage (from the occurrence of the disturbance to the most severe degradation) is obtained to obtain the failure profile representing the accumulation of risk; the index deviation in the second stage (from the most severe degradation to the recovery endpoint) is integral to obtain the recovery profile representing the recovery cost; if a certain scheduling trajectory can quickly control the water level during the flood season and restore power generation very quickly after the disaster, the areas of its two profiles will be small, thus obtaining a very high global cooperative resilience score after mathematical aggregation.

[0061] The hydropower scheduling method based on real-time hydrology and knowledge reasoning provided in this invention accurately divides the continuous and fuzzy hydrological disturbance process into two quantifiable physical stages: degradation and recovery. It also introduces integral calculation of failure profiles and recovery profiles, thereby achieving a fine representation of the dynamic evolution characteristics of hydropower agents in complex environments. At the same time, it provides reliable data support for subsequent inverse reinforcement learning to extract collaborative incentive logic.

[0062] In some embodiments, a global collaborative resilience score is obtained by calculating based on the failure profile and recovery profile corresponding to each indicator, including: obtaining a first time interval from the time of disturbance occurrence to the time of most severe degradation and a second time interval from the time of most severe degradation to the end of recovery; calculating the individual resilience score of each indicator based on the first time interval, the second time interval, and the failure profile and recovery profile corresponding to each indicator; and aggregating the individual resilience scores of each indicator using the harmonic mean to obtain the global collaborative resilience score.

[0063] In this embodiment, by combining the first time interval, the second time interval, and the previously obtained failure profile and recovery profile, the hydropower scheduling agent transforms the absolute integral area into a normalized single-item resilience score (usually between 0 and 1, with a higher score indicating stronger resilience of the single-item index) through a specific mathematical mapping formula.

[0064] In this embodiment, the resilience scoring function The calculation is performed through the following steps: (1) For a segment of length... scheduling trajectory Identify the time of disturbance through knowledge reasoning The most severe moment of degradation and recovery endpoint .

[0065] (2) The failure profile (essentially a dimensionless quantity) and recovery profile (essentially a dimensionless quantity) are expressed by the following formulas: ; ; in, It serves as a benchmark indicator based on historical stable hydrological data, while This is a real-time indicator under the current disturbed state, from which the resilience score of a single indicator can be derived. , is represented as: ; (3) Using the harmonic mean to pair k The indicators are aggregated to obtain a global collaborative resilience score. : ; The above aggregation method can effectively punish strategies that perform poorly in any key dimension (such as ecological flow), ensuring the bottom-line safety of hydropower agents.

[0066] For example, in a scenario of a sudden and severe rainstorm in a river basin, the hydropower agent records the first time interval of degradation caused by flood confluence as 10 hours, followed by a second time interval of 40 hours of coordinated flood discharge until the water level and output return to normal. The hydropower agent calculates the individual resilience scores for the two indicators of "cumulative output sustainability" and "resource availability (reservoir capacity safety)". Assuming that a certain scheduling trajectory maintains extremely high power generation without significantly reducing output during this period, its individual score for "cumulative output sustainability" is an extremely high 0.9; however, this approach... As a result, the reservoir failed to release flood control capacity in time, and the water level dangerously approached the dam crest, causing its "resource availability" score to be only 0.2. If a normal arithmetic mean is used during aggregation, the overall score may be as high as 0.55, and the hydropower agent may mistakenly consider this a "passable" scheduling strategy. However, this embodiment introduces a harmonic mean for aggregation, and the calculated global cooperative resilience score will be significantly reduced to about 0.32, which sends a clear penalty signal to the reinforcement learning model: the strategy has touched the bottom line of the hydropower agent's safety.

[0067] The hydropower scheduling method based on real-time hydrology and knowledge reasoning provided in this invention quantifies the rate of degradation and recovery of hydropower agents by introducing a first time interval, a second time interval, and failure and recovery profiles of each indicator; and uses the harmonic mean to achieve aggregate calculation to ensure the safety of hydropower scheduling.

[0068] In some embodiments, the collective collaboration reward item is obtained by inverting the preset parameterized reward model using preference relations, including: determining the loss function of the preset parameterized reward model according to the preference learning algorithm and preference relations; wherein the loss function is constructed based on the reward difference or preference probability between high-scoring scheduling trajectories and low-scoring scheduling trajectories; updating the parameters of the preset parameterized reward model by minimizing the loss function, and using the output of the updated preset parameterized reward model as the collective collaboration reward item.

[0069] In this embodiment, the preference learning algorithm includes a marginal preference learning (MPL) algorithm or a probabilistic preference learning algorithm, and the preference relationship is the ranking of the trajectories based on the global cooperative resilience score obtained in the preceding steps (e.g., trajectory A is better than trajectory B).

[0070] In this embodiment, the preset parameterized reward model is a neural network (such as a multilayer perceptron) or a linear feature map model containing adjustable parameters (such as weights and biases). Its input is real-time hydrological and power station operating status, and its output is a specific reward value. In this embodiment, the loss function is constructed based on the reward difference or preference probability between high-scoring and low-scoring scheduling trajectories: if it is based on the reward difference, the cumulative reward of the high-scoring trajectory predicted by the model must exceed a certain margin of the low-scoring trajectory; if it is based on the preference probability, the probability of the high-scoring trajectory being selected by the model is maximized through maximum likelihood estimation.

[0071] In this embodiment, the loss function can be minimized through optimization algorithms such as gradient descent, and the parameters of the preset parameterized reward model can be continuously updated. Finally, the output value calculated by the model in real time under any input state is the collective cooperation reward item that guides the intelligent agents to cooperate.

[0072] In this embodiment, the marginal preference learning algorithm ensures that, in the prediction results, the cumulative reward of the high-resilience trajectory exceeds that of the low-resilience trajectory by a specific marginal value.

[0073] This embodiment collects M initial trajectory sets D and establishes a preference relation based on ρ(τ): If Then it is written as This embodiment can perform inversion inference on a preset parameterized reward model through the following two switchable learning paths: (1) Marginal preference learning is adopted; this embodiment designs a dynamic marginal preference learning mechanism. This makes it proportional to the difference in resilience scores between the two trajectories. This can be expressed by the following formula: .

[0074] The optimization objective function is defined as minimizing the hinge loss: .

[0075] (2) The Probabilistic Preference Learning (PPL) algorithm is adopted; the PPL algorithm uses the idea of ​​maximum likelihood estimation to model the preferences between trajectories as a Bradley-Terry probability model; trajectories Superior The predicted probability is expressed as: ; The parameters are then updated by minimizing the negative log-likelihood function, which is expressed as: ; The method described above provides a smooth gradient descent process, which can better handle noise and outliers in hydrological data.

[0076] The hydropower scheduling method based on real-time hydrology and knowledge reasoning provided in this invention determines the loss function of a preset parameterized reward model through a preference learning algorithm and preference relations. The parameters of the preset parameterized reward model are updated by minimizing the loss function, and the output of the updated preset parameterized reward model is used as a collective collaborative reward item. This method automatically infers the deep-level collaborative incentive structure from trajectory scores containing prior hydrological knowledge through inverse reinforcement learning, helping the model to determine reward signals that meet the safety bottom line of human experts and the requirements of system resilience in complex nonlinear feature spaces.

[0077] In some embodiments, the preset parameterized reward model includes at least one of the following: a manual knowledge linear model based on predefined expert experience features; a state linear model based on features of the original state variables; or a nonlinear deep neural network model based on a multilayer perceptron.

[0078] In this embodiment, a dynamic reward function is assumed. Feature representation depending on state This embodiment provides a three-layer progressive parameterized model: (1) Handcrafted Linear Model, specifically: Model the reward as Among them, features Predefined by expert experience, such as water level variability and load gap. This method has extremely high interpretability, but depends on the quality of expert knowledge.

[0079] (2) State-based Linear Model, specifically: Use the original state variables directly As a feature, namely This reduces human intervention, but it has limitations in handling nonlinear hydrological logic.

[0080] (3) Nonlinear deep neural network model, specifically: Parameterization It utilizes multilayer perceptrons (MLPs) to automatically learn the complex dependencies between the state space and the restorative output, exhibiting the strongest fitting ability.

[0081] The hydropower scheduling method based on real-time hydrology and knowledge reasoning provided in this invention offers a multi-dimensional parameterized architecture option, ranging from strong domain interpretability (manual feature model) and low manual cost (state linear model) to extremely strong generalization fitting ability (deep neural network model). This model mechanism enables the hydropower scheduling system to fully integrate the hydrological scheduling experience of human experts in scenarios with clear features, and to leverage the advantages of deep learning in automatically extracting nonlinear features in extremely complex and high-dimensional non-stationary hydrological environments. This improves the engineering implementation capability and system robustness of the reward inference mechanism in multi-scale real hydropower projects.

[0082] In some embodiments, determining a hybrid reward function based on the collective cooperation reward item and the power generation revenue signal of an individual hydropower station includes: configuring corresponding individual adjustment weights and cooperation adjustment weights for the power generation revenue signal and the collective cooperation reward item, respectively; and performing a linear weighted summation of the power generation revenue signal and the collective cooperation reward item based on the individual adjustment weights and the cooperation adjustment weights to obtain the hybrid reward function.

[0083] In this embodiment, the individual adjustment weight and the collaborative adjustment weight are proportional coefficients that are set manually or dynamically allocated by the system. Their function is to provide a flexible numerical balance valve between maximizing individual profits and maximizing the resilience of the global system.

[0084] In this embodiment, by linearly weighting and summing the individual adjustment weights and power generation revenue signals, the cooperative adjustment weights and cooperative adjustment weights, the originally conflicting multi-objective evaluation indicators can be reduced in dimension and fused, and finally output a scalar value, namely the hybrid reward function, which is applied to the training loop of multi-agent reinforcement learning (such as PPO or QMIX) as the only optimization feedback for each agent to update the neural network strategy.

[0085] Specifically, in order to balance economic benefits and system security, a hybrid strategy was adopted: for each power station agent The final real-time total reward received The definition is as follows: ; in, Individual reward signals (such as real-time electricity sales profits, savings from penalties, etc.). These are the collective collaboration items inferred through the IRL process described above. and It is an adjustment parameter used to balance profit maximization and system resilience maximization.

[0086] In this embodiment, dynamic reward signals (i.e. power generation revenue signals) are injected into a reinforcement learning framework (such as the PPO algorithm) for agent policy training. By continuously alternating between the closed-loop process of "action execution - resilience assessment - reward correction - policy update", the hydropower station agent can spontaneously evolve a highly intelligent cooperative mode, such as peak-shifting scheduling and joint flood control response among power stations.

[0087] The hydropower dispatching method based on real-time hydrology and knowledge reasoning provided in this invention configures corresponding individual adjustment weights and collaborative adjustment weights for the power generation revenue signal and the collective collaborative reward item, respectively. Then, it performs a linear weighted summation of the power generation revenue signal and the collective collaborative reward item according to the individual adjustment weights and the collaborative adjustment weights to obtain a hybrid reward function. Through the flexible configuration of weights, the hydrological collaborative constraints are transformed into decision conditions for the intelligent agent, avoiding the tendency of each power station agent to over-consume reservoir capacity to obtain short-term individual power generation rewards during training, and ensuring that each power station agent operates stably in the long term and has a clear division of labor.

[0088] This embodiment introduces an innovative scheduling paradigm that deeply integrates collaborative resilience metrics with preference-based inverse reinforcement learning. It enhances the collaborative resilience of multi-power station systems in complex hydrological environments by inferring the collective reward function through reinforcement learning. Specifically, this embodiment first establishes a multi-dimensional hydrological assessment system based on knowledge reasoning. By quantifying key indicators such as cumulative output stability, resource availability, reservoir capacity fairness index, and shortage risk, it calculates the global resilience score of the scheduling trajectory using failure and recovery profiles. Then, to extract expert preferences from complex scheduling behaviors and improve training efficiency, this embodiment introduces marginal preference learning and probabilistic preference learning algorithms, and employs a handcrafted linear feature model or a nonlinear neural network to refine the reward function. Parametric representation and inversion inference are used to accurately capture the collaborative incentive structure of hydrological evolution in the feature space. Finally, this embodiment uses a hybrid reward strategy to integrate the inferred collective resilience term into the individual scheduling objectives. Through a multi-agent reinforcement learning framework, agents of each power station are guided to achieve division of labor and cooperation in spatial behavior while ensuring productivity, effectively alleviating the social dilemma failure problem of over-exploitation of resources. In addition, extensive experiments in various scale environments show that the hydropower scheduling method based on real-time hydrology and knowledge reasoning provided in this embodiment is consistently superior to existing advanced baseline models such as PPO and QMIX in terms of improving system lifespan, cumulative power generation, and robustness to extreme hydrological disturbances.

[0089] Figure 2 This is the second flowchart of the hydropower scheduling method based on real-time hydrology and knowledge reasoning provided by the present invention. Figure 2 In the illustrated embodiment, the method comprises three stages: (1) obtaining the collaborative resilience strength based on real-time hydrological knowledge inference; (2) learning a reward function to achieve collaborative resilience; and (3) integrating the inference reward into the learning process. Specifically, in stage (1), key hydrological evaluation indicators (including cumulative output sustainability, resource availability, scheduling fairness index, and shortage risk index) are calculated using scheduling data, and a resilience scoring function is calculated based on these evaluation indicators (the figure below is accompanied by a medal icon, indicating that the data is scored or sorted). In stage (2), the parameterized representation of the reward function is first expressed as: R(s) = f( ; ) The reward inference algorithm for preference learning and the calculated resilience score function are processed through two preference learning algorithms (marginal preference learning and probabilistic preference learning) to obtain a hybrid reward function. In stage (3), the final reward function and the learning process are applied together to the robot icon at the center (representing the learning subject) to obtain the corresponding decentralized scheduling strategy, so as to control the execution mechanism corresponding to the hydropower station intelligent agent to realize hydropower scheduling. The final reward function is expressed as: R(s) = f( ; ) .

[0090] The hydropower scheduling device based on real-time hydrology and knowledge reasoning provided by the present invention is described below. The hydropower scheduling device based on real-time hydrology and knowledge reasoning described below and the hydropower scheduling method based on real-time hydrology and knowledge reasoning described above can be referred to in correspondence.

[0091] Figure 3 This is a schematic diagram of the hydropower dispatching device based on real-time hydrology and knowledge reasoning provided by the present invention, as shown below. Figure 3 As shown, the device includes: a knowledge reasoning evaluation module 310, a reward configuration module 320, a scheduling strategy training module 330, and a water and electricity scheduling module 340.

[0092] The knowledge reasoning and evaluation module 310 is used to quantify the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period based on real-time hydrological data and prior knowledge in the hydrological field, and obtain the global cooperative resilience score of multiple hydropower dispatch trajectories; the multiple indicators include at least two of the following: cumulative output sustainability, resource availability, dispatch fairness index and shortage risk index. The reward configuration module 320 is used to establish the preference relationship between each scheduling trajectory based on the global collaborative resilience score, and to use the preference relationship to perform inversion inference on the preset parameterized reward model to obtain the collective collaborative reward item; The scheduling strategy training module 330 is used to determine the hybrid reward function based on the collective cooperation reward item and the power generation revenue signal of the individual hydropower station, and to train the hydropower station agents with the hybrid reward function and the multi-agent reinforcement learning framework to obtain a decentralized scheduling strategy. The hydropower dispatching module 340 is used to control the execution mechanism corresponding to the hydropower station intelligent agent to realize hydropower dispatching using a decentralized dispatching strategy.

[0093] The hydropower dispatching device based on real-time hydrology and knowledge reasoning provided in this invention transforms hydrological experience into quantitative indicators and evaluates the performance of dispatching strategies under extreme environments, providing objective data for the subsequent automated inference of reward functions. Furthermore, it automatically extracts and quantifies the collaborative incentive structure, which is difficult for experts to describe, from complex dispatching behaviors based on global collaborative resilience scores, thus achieving automation and intelligence in reward function design and overcoming the drawbacks of manually setting weights. By guiding the division of labor and cooperation among power stations through a hybrid reward mechanism, it improves the operational efficiency of the hydropower station group and the response speed to extreme conditions.

[0094] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can call logic instructions in the memory 430 to execute a hydropower dispatching method based on real-time hydrology and knowledge reasoning. This method includes: quantifying the cumulative deviation of multiple indicators of the hydropower station agent during disturbances based on real-time hydrological data and prior knowledge in the hydrological domain, obtaining a global cooperative resilience score for multiple hydropower dispatching trajectories; the multiple indicators include at least two of cumulative output sustainability, resource availability, dispatching fairness index, and shortage risk index; establishing preference relationships between dispatching trajectories based on the global cooperative resilience score, and using these preference relationships to inversely infer a preset parameterized reward model to obtain a collective cooperative reward item; determining a hybrid reward function based on the collective cooperative reward item and the power generation revenue signal of individual hydropower station agents, and using the hybrid reward function and a multi-agent reinforcement learning framework to train the hydropower station agent's strategy, obtaining a decentralized dispatching strategy; and using the decentralized dispatching strategy to control the execution mechanism corresponding to the hydropower station agent to achieve hydropower dispatching.

[0095] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0096] On the other hand, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the hydropower dispatching method based on real-time hydrology and knowledge reasoning provided by the above methods. The method includes: quantifying the cumulative deviation of multiple indicators of a hydropower station agent during a disturbance based on real-time hydrological data and prior knowledge in the hydrological domain to obtain a global cooperative resilience score for multiple hydropower dispatching trajectories; the multiple indicators include at least two of cumulative output sustainability, resource availability, dispatching fairness index, and shortage risk index; establishing a preference relationship between each dispatching trajectory based on the global cooperative resilience score, and using the preference relationship to inversely infer a preset parameterized reward model to obtain a collective cooperative reward item; determining a hybrid reward function based on the collective cooperative reward item and the power generation revenue signal of the individual hydropower station, and using the hybrid reward function and a multi-agent reinforcement learning framework to train the hydropower station agent's strategy to obtain a decentralized dispatching strategy; and using the decentralized dispatching strategy to control the execution mechanism corresponding to the hydropower station agent to realize hydropower dispatching.

[0097] In another aspect, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the hydropower dispatching method based on real-time hydrology and knowledge reasoning provided by the above methods. The method includes: quantifying the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period based on real-time hydrological data and prior knowledge in the hydrological domain to obtain a global cooperative resilience score for multiple hydropower dispatching trajectories; the multiple indicators include at least two of cumulative output sustainability, resource availability, dispatching fairness index, and shortage risk index; establishing a preference relationship between each dispatching trajectory based on the global cooperative resilience score, and using the preference relationship to perform inversion inference on a preset parameterized reward model to obtain a collective cooperative reward item; determining a hybrid reward function based on the collective cooperative reward item and the power generation revenue signal of the individual hydropower station, and using the hybrid reward function and a multi-agent reinforcement learning framework to train the hydropower station agent's strategy to obtain a decentralized dispatching strategy; and using the decentralized dispatching strategy to control the execution mechanism corresponding to the hydropower station agent to realize hydropower dispatching.

[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0099] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A hydropower dispatching method based on real-time hydrology and knowledge reasoning, characterized in that, include: Based on real-time hydrological data and prior knowledge in the hydrological field, the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period is quantitatively calculated to obtain a global cooperative resilience score of multiple hydropower dispatch trajectories; the multiple indicators include at least two of the following: cumulative output sustainability, resource availability, dispatch fairness index and shortage risk index. Based on the global collaborative resilience score, a preference relationship is established between each scheduling trajectory, and the preference relationship is used to invert and infer the preset parameterized reward model to obtain the collective collaborative reward item; A hybrid reward function is determined based on the collective collaboration reward item and the power generation revenue signal of the individual hydropower station. The hybrid reward function and the multi-agent reinforcement learning framework are used to train the hydropower station agents to obtain a decentralized scheduling strategy. Hydropower scheduling is achieved by using the decentralized scheduling strategy to control the execution mechanism corresponding to the hydropower station's intelligent agent.

2. The hydropower dispatching method based on real-time hydrology and knowledge reasoning according to claim 1, characterized in that, Based on real-time hydrological data and prior knowledge in the hydrological field, the cumulative deviation of multiple indicators of the hydropower station's intelligent agent during disturbances is quantitatively calculated to obtain a global cooperative resilience score for multiple hydropower dispatch trajectories, including: Based on the real-time hydrological data and prior knowledge in the hydrological field, the index thresholds are compared and rules are matched to determine the time of disturbance occurrence, the time of most severe degradation, and the recovery endpoint in the hydropower dispatch trajectory. Calculate the cumulative deviation of the multiple indicators from the normal baseline indicators during the period from the time of the disturbance to the time of the most severe degradation to obtain the corresponding failure profile; calculate the cumulative deviation of the multiple indicators from the normal baseline indicators during the period from the time of the most severe degradation to the recovery endpoint to obtain the corresponding recovery profile; The global collaborative resilience score is obtained by calculating the failure profile and the recovery profile corresponding to each indicator.

3. The hydropower dispatching method based on real-time hydrology and knowledge reasoning according to claim 2, characterized in that, The global collaborative resilience score is obtained by calculating the failure profile and the recovery profile corresponding to each indicator, including: Obtain the first time interval from the time of the disturbance to the time of the most severe degradation, and the second time interval from the time of the most severe degradation to the end point of the recovery; Based on the first time interval, the second time interval, and the failure profile and recovery profile corresponding to each indicator, calculate the individual resilience score for each indicator. The harmonic mean is used to aggregate the individual resilience scores of each indicator to obtain the global collaborative resilience score.

4. The hydropower dispatching method based on real-time hydrology and knowledge reasoning according to claim 1, characterized in that, The step of using the preference relationship to invert and infer the preset parameterized reward model to obtain the collective collaboration reward item includes: The loss function of the preset parameterized reward model is determined according to the preference learning algorithm and the preference relationship; wherein, the loss function is constructed based on the reward difference or preference probability between high-scoring scheduling trajectories and low-scoring scheduling trajectories; The parameters of the preset parameterized reward model are updated by minimizing the loss function, and the output of the updated preset parameterized reward model is used as the collective collaboration reward.

5. The hydropower dispatching method based on real-time hydrology and knowledge reasoning according to claim 1 or 4, characterized in that, The preset parameterized reward model includes at least one of the following: A linear model of handcrafted knowledge based on predefined expert experience features; State linear model based on the characteristics of original state variables; A nonlinear deep neural network model based on a multilayer perceptron.

6. The hydropower dispatching method based on real-time hydrology and knowledge reasoning according to claim 1, characterized in that, The determination of the hybrid reward function based on the collective cooperation reward item and the power generation revenue signal of the individual hydropower station includes: Configure corresponding individual adjustment weights and collaborative adjustment weights for the power generation revenue signal and the collective cooperation reward item, respectively; The hybrid reward function is obtained by linearly weighting and summing the power generation revenue signal and the collective cooperation reward item based on the individual adjustment weight and the cooperative adjustment weight.

7. A hydropower dispatching device based on real-time hydrology and knowledge reasoning, characterized in that, include: The knowledge reasoning and evaluation module is used to quantify the cumulative deviation of multiple indicators of the hydropower station agent during the disturbance period based on real-time hydrological data and prior knowledge in the hydrological field, and obtain a global cooperative resilience score of multiple hydropower dispatch trajectories; the multiple indicators include at least two of the following: cumulative output sustainability, resource availability, dispatch fairness index and shortage risk index. The reward configuration module is used to establish the preference relationship between each scheduling trajectory based on the global collaborative resilience score, and to use the preference relationship to perform inversion inference on the preset parameterized reward model to obtain the collective collaborative reward item; The scheduling strategy training module is used to determine a hybrid reward function based on the collective cooperation reward item and the power generation revenue signal of the individual hydropower station, and to use the hybrid reward function and the multi-agent reinforcement learning framework to train the hydropower station agent to obtain a decentralized scheduling strategy. The hydropower dispatching module is used to control the execution mechanism corresponding to the hydropower station intelligent agent to realize hydropower dispatching using the decentralized dispatching strategy.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the hydropower scheduling method based on real-time hydrology and knowledge reasoning as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the hydropower scheduling method based on real-time hydrology and knowledge reasoning as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the hydropower scheduling method based on real-time hydrology and knowledge reasoning as described in any one of claims 1 to 6.