Reservoir power station group dispatching rule extraction method and system based on multi-model soft scene fusion

CN122333403APending Publication Date: 2026-07-03NANJING HYDRAULIC RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING HYDRAULIC RES INST
Filing Date
2026-06-05
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies have problems such as weak scenario adaptability and high risk of sudden decision changes in the scheduling of cascade reservoir groups. In particular, under complex hydrological conditions, it is difficult to accurately characterize the hydraulic constraints and game states between cascades, and it is impossible to calculate the uncertainty risk boundary behind the output decision.

Method used

A multi-model soft scene fusion approach is adopted. By constructing multi-dimensional dimensionless feature vectors, the membership degree of soft scenes is determined. Local prediction mean and bias are generated using pre-trained prediction base models. Combined with weighted synthesis technology, the predicted outflow and prediction variance of the target reservoir are generated.

Benefits of technology

It suppresses sudden decision-making changes during operating condition switching, improves the robustness of scheduling and the accuracy of prediction, reduces uncertainty risks, and enhances scheduling decision-making capabilities under complex hydrological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333403A_ABST
    Figure CN122333403A_ABST
Patent Text Reader

Abstract

This invention constructs a method and system for extracting scheduling rules for reservoir power station groups based on multi-model soft scenario fusion. The method includes: acquiring the historical operation sequence and current operational observation status of the reservoir power station group; constructing an input feature set based on this, and extracting physical state features reflecting the hydrological situation of a single reservoir and the spatial coupling relationship of the cascade, obtaining a multi-dimensional dimensionless feature vector; determining the soft scenario membership degree based on the spatial distance mapping between this feature vector and the cluster centers of each scheduling scenario; inputting the input feature set into a pre-trained prediction base model specific to each scheduling scenario, and combining it with the estimated posterior parameters of the scenario to generate the local prediction mean and local prediction deviation for each scenario; and weighting and synthesizing the local prediction results based on the soft scenario membership degree to obtain the predicted outflow and prediction variance. This invention suppresses decision-making abrupt changes during operating condition switching, calculates prediction uncertainty, and improves the robustness of scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of water conservancy and hydropower engineering and artificial intelligence, and in particular to a method and system for extracting scheduling rules for reservoir power station groups based on multi-model soft scene fusion. Background Technology

[0002] Optimized scheduling of cascade hydropower stations is a technical means to achieve efficient utilization of hydropower resources and ensure flood control and ecological security. With the continuous expansion of hydropower development in river basins, cascade water network systems exhibit strong nonlinearity and strong coupling characteristics across multiple temporal and spatial dimensions. Accurately extracting and constructing high-fidelity intelligent scheduling rules from massive historical operational data is of engineering and technical value for guiding reservoir groups in making real-time release decisions under complex and variable hydrological conditions and balancing multiple objectives such as cascade power generation and regional water supply.

[0003] Currently, implicit stochastic optimization methods based on data mining and artificial intelligence are widely used in reservoir scheduling rule extraction. Existing conventional approaches typically use raw data such as absolute water levels and absolute inflow / outflow rates of each reservoir as input features, and employ a single machine learning algorithm to globally fit and train the entire time-series sample set. In the implementation of such methods, the input of absolute physical quantities is easily affected by the significant differences in installed capacity and storage capacity between upstream and downstream reservoirs within the cascade, making it difficult for the algorithm to accurately characterize the hydraulic constraints and game states between cascades. Furthermore, when using a single global model to fit massive samples, the model parameters often compromise to the features of the normal water period, which constitute the majority of the data. This leads to a decrease in fitting accuracy when facing sudden changes in conditions such as rapid shifts between drought and flood or extreme low water levels.

[0004] In summary, existing rule extraction methods suffer from weak scenario adaptability and high risk of sudden decision changes when dealing with the scheduling of cascade reservoir groups under variable hydrological conditions. Particularly at complex scheduling boundaries such as flood-dry season transitions and multi-reservoir coordination, the generalization and anti-interference capabilities of existing global prediction models face severe challenges, and they are unable to calculate the uncertainty risk boundary behind the output decisions. Therefore, there is an urgent need to research a method that can adapt to complex hydrological conditions and improve the robustness and risk perception level of rule extraction across multiple scenarios. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for extracting scheduling rules for reservoir power station groups based on multi-model soft scene fusion, in order to solve the aforementioned technical problems existing in the prior art.

[0006] Technical solution: A method for extracting scheduling rules for reservoir power station groups based on multi-model soft scenario fusion, including:

[0007] Obtain the historical operation sequence and current operational observation status of the reservoir power station group;

[0008] Based on historical operation sequences and current operational observation status, an input feature set for the prediction base model is constructed, and physical state features reflecting the coupling relationship between the hydrological situation of a single reservoir and the spatial relationship of the cascade are extracted to obtain a multidimensional dimensionless feature vector.

[0009] Based on the spatial distance mapping between multidimensional dimensionless feature vectors and pre-constructed cluster centers of various scheduling scenarios, the soft scenario membership degree of the current running observation state relative to each scheduling scenario is determined.

[0010] The input feature set is input into the pre-trained prediction base model specific to each scheduling scenario, and combined with the pre-estimated scenario posterior parameters, to generate the local prediction mean and local prediction bias corresponding to each scheduling scenario.

[0011] Based on the soft scene membership degree, the local prediction mean and local prediction deviation corresponding to each scheduling scene are weighted and synthesized to obtain the predicted outflow and prediction variance of the target reservoir.

[0012] Optionally, physical state features reflecting the coupling relationship between the hydrological situation of a single reservoir and the spatial relationship of the cascade system are extracted to obtain a multidimensional dimensionless feature vector, including:

[0013] Based on historical operation sequences and current operational observation status, individual reservoir hydrological situation characteristic indicators are extracted for each regulating reservoir.

[0014] Based on the hydrological situation characteristics of a single reservoir, we extract the cascade coupling characteristics between adjacent regulating reservoirs.

[0015] By combining the hydrological situation characteristic indicators of a single reservoir with the characteristic indicators of cascade coupling, a multidimensional dimensionless feature vector is obtained.

[0016] Optionally, the hydrological characteristic indicators for a single reservoir include inflow abundance / difference indicators, water storage utilization rate indicators, and inflow change trend indicators; among them, the hydrological characteristic indicators for each regulating reservoir are extracted, including:

[0017] For each regulating reservoir, the current inflow rate is extracted from the current operating observation status, and the ratio of the current inflow rate to the pre-configured average inflow rate for the same period over many years is calculated to obtain the water abundance / scarcity index.

[0018] Extract the current water storage from the current operating observation status, and calculate the ratio of the available water after deducting the pre-configured dead storage capacity from the current water storage to the pre-configured total regulating storage capacity to obtain the water storage utilization rate index;

[0019] Extract the inflow rate of the previous period from the historical operation sequence, calculate the difference between the current inflow rate and the inflow rate of the previous period, and compare this difference with the average inflow rate over the same period to obtain the trend index of water inflow change.

[0020] Optionally, based on the spatial distance mapping between multidimensional dimensionless feature vectors and pre-constructed cluster centers of each scheduling scenario, the soft scenario membership degree of the current operational observation state relative to each scheduling scenario is determined, including:

[0021] Calculate the spatial feature distance between the multidimensional dimensionless feature vector and each pre-built cluster center of the scheduling scenario;

[0022] By using a distance decay function configured based on preset bandwidth parameters, the distance of each spatial feature is mapped to the corresponding assignment weight coefficient of the scheduling scenario.

[0023] The proportion of the attribution weight coefficient of each scheduling scenario in the total attribution weight coefficient of all scheduling scenarios is determined as the soft scenario membership degree of the current running observation state relative to the corresponding scheduling scenario.

[0024] Optionally, the pre-estimated scenario posterior parameters include the scenario prediction base model weights, the linear correction intercept of the prediction base model, the linear correction slope of the prediction base model, and the local baseline prediction variance; combined with the pre-estimated scenario posterior parameters, the local prediction mean and local prediction bias corresponding to each scheduling scenario are generated, including:

[0025] For each scheduling scenario, obtain the initial outbound flow prediction value output by each prediction base model for the input feature set;

[0026] By using the linear correction intercept and linear correction slope of the prediction base model, the corresponding initial outflow flow prediction values ​​are linearly corrected to obtain the corrected prediction values ​​of each prediction base model.

[0027] The corrected prediction values ​​of each prediction base model are weighted and summed using the weights of the scenario prediction base models to obtain the local prediction mean of the current scheduling scenario.

[0028] By combining the local baseline prediction variance and the dispersion of the corrected prediction values ​​of each prediction base model relative to the local prediction mean, the local prediction bias of the current scheduling scenario is calculated.

[0029] Optionally, based on the soft scenario membership degree, the local prediction mean and local prediction deviation corresponding to each scheduling scenario are weighted and synthesized to obtain the predicted outflow and prediction variance of the target reservoir, including:

[0030] Using the soft scenario membership degree of each scheduling scenario as a weighting coefficient, the local prediction mean of all scheduling scenarios is weighted and summed to obtain the predicted outbound flow.

[0031] Based on local prediction bias and membership of each soft scene, calculate the intra-scene variance component that reflects the prediction discrepancy of the model within a single scene.

[0032] Based on the deviation of each local prediction mean from the final predicted outbound flow, and combined with the membership degree of each soft scenario, the inter-scenario variance component reflecting the uncertainty of cross-scenario affiliation is calculated.

[0033] The variance components within a scenario are added together with the variance components between scenarios to obtain the predicted variance that characterizes the final scheduling risk boundary.

[0034] Optionally, each pre-trained prediction base model is trained using a pre-constructed optimal scheduling sample set; the optimal scheduling sample set is constructed through a multi-timescale nested optimization mechanism, including:

[0035] Multi-objective optimization scheduling models at monthly, weekly, and daily scales were constructed respectively.

[0036] The monthly-scale multi-objective optimization scheduling model is solved, and the monthly target water level of each reservoir is extracted.

[0037] The target water level at the end of the month is used as an upper boundary constraint and injected into the weekly multi-objective optimization scheduling model for solution. The weekend target water level of each reservoir is then extracted from the output.

[0038] The target water level at the weekend is used as an upper-level boundary constraint and injected into the daily-scale multi-objective optimization scheduling model for solution. The daily-scale optimal scheduling trajectory sequence that satisfies the constraint is selected and compiled into an optimal scheduling sample set.

[0039] Optionally, the optimization objectives of each multi-objective optimization scheduling model specifically include:

[0040] The goals are to maximize the cascade power generation, maximize the water supply guarantee rate of the reservoir group, and maximize the ecological flow guarantee rate of the river section.

[0041] Specifically, the solution for each multi-objective optimization scheduling model is as follows:

[0042] A non-dominated sorting genetic algorithm with a preset population size is used to iteratively calculate the model considering the optimization objective, generating a non-dominated optimal solution set;

[0043] A multi-attribute decision-making model based on the approximation ideal solution ranking method is used to optimize the non-dominated optimal solution set and obtain the optimal scheduling trajectory of each reservoir at the corresponding time scale.

[0044] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the steps of the above-described method for extracting scheduling rules for reservoir power station groups based on multi-model soft scene fusion.

[0045] According to another aspect of the embodiments of this application, a reservoir power station group scheduling rule extraction system based on multi-model soft scene fusion is also provided, comprising:

[0046] Memory, used to store computer programs;

[0047] The processor is used to implement the steps of the above-mentioned method for extracting scheduling rules for reservoir power station groups based on multi-model soft scene fusion when executing computer programs.

[0048] Beneficial effects: This invention suppresses sudden changes in decision-making during operating condition switching, calculates and predicts uncertainties, and improves the robustness of scheduling. Attached Figure Description

[0049] Figure 1 This is a flowchart of the reservoir power station group scheduling rule extraction method based on multi-model soft scene fusion according to the present invention.

[0050] Figure 2 The flowchart illustrates how the optimal scheduling sample set of this invention is constructed using a multi-timescale nested optimization mechanism.

[0051] Figure 3 This is a flowchart illustrating how the multidimensional dimensionless eigenvectors are obtained in this invention.

[0052] Figure 4 This is a flowchart illustrating the extraction of hydrological characteristic indicators for each regulating reservoir in this invention.

[0053] Figure 5 This is a flowchart illustrating the process of extracting cascade coupling characteristic indicators between adjacent regulating reservoirs according to the present invention.

[0054] Figure 6 The flowchart shows how the preset bandwidth parameter configuration of this invention is adaptively determined through cross-validation. Detailed Implementation

[0055] This embodiment provides a method for extracting scheduling rules for reservoir power station groups based on multi-model soft scenario fusion, such as... Figure 1 As shown, the method includes the following steps:

[0056] Step 101: Obtain the historical operation sequence and current operation observation status of the reservoir power station group.

[0057] As an optional implementation method, the reservoir power station group takes the cascade power station reservoirs in the middle and lower reaches of the Yalong River as the research object. Its hydropower pattern includes reservoirs with multiple regulation functions such as multi-year regulation, annual regulation, seasonal regulation, and daily regulation. Historical operation sequences and current operation observation status are collected in real time through the basin's automated scheduling system, specifically including time-series data such as inflow, outflow, reservoir water level, and power output information of each level of power station.

[0058] The digital representation of historical operating sequences involves calculating power plant output. In this embodiment, the comprehensive output coefficient ηi of each power plant has been pre-determined, and its value incorporates the unit weight of water and the overall unit efficiency. The physical unit of the output coefficient ηi is kW / (m*m). 3 / s). This comprehensive output coefficient enables the establishment of a linear conversion relationship between power generation head, power generation flow rate, and output value, simplifying the complexity of the physical model.

[0059] Accordingly, the CPG generated in the historical operating sequence can be calculated based on the following formula:

[0060] CPG=∑(ηi*Hi,t*Qi,t*Δt);

[0061] Wherein, CPG is the total power generation of the cascade, ηi is the comprehensive output coefficient of the i-th power station, Hi,t is the power generation head of the i-th power station in time period t, Qi,t is the power generation flow of the i-th power station in time period t, and Δt is the length of the time period.

[0062] Step 102: Based on the historical operation sequence and the current operation observation status, construct the input feature set for the prediction base model, and extract the physical state features that reflect the coupling relationship between the hydrological situation of a single reservoir and the spatial relationship of the cascade, to obtain a multidimensional dimensionless feature vector.

[0063] In this embodiment, the input feature set is used to drive the subsequent artificial intelligence model. Considering the continuous nature of hydrological processes and the hydraulic connections between cascade reservoirs, the input feature set is constructed as an 8-dimensional real vector. Specifically, the input feature set X includes the inflow It of the current period, the inflow It1 of the previous period, the inflow It2 of the two previous periods, the initial water level Zt of the current period, the initial water level Zt1 of the previous period, the initial water level Zt2 of the two previous periods, the outflow Ot1 of the previous period, and the outflow Ot2 of the two previous periods. By introducing the previous outflow decision as an input, the model can capture the inertial characteristics of the scheduling instructions and the connection between upstream and downstream water volumes.

[0064] Furthermore, to eliminate the differences in flow rate and absolute water level between reservoirs of different sizes, this embodiment further extracts physical state features reflecting physical meaning and transforms them into a dimensionless form. The multidimensional dimensionless feature vector includes the inflow abundance and water storage level of each reservoir, as well as coupling indicators characterizing the upstream and downstream coordination relationship. Through dimensionless processing, the distance calculation in the feature space for different scheduling scenarios has physical consistency, improving the accuracy of scene recognition.

[0065] Step 103: Based on the spatial distance mapping between the multidimensional dimensionless feature vector and the pre-constructed cluster centers of each scheduling scenario, determine the soft scenario membership degree of the current running observation state relative to each scheduling scenario.

[0066] Specifically, soft scenario membership is used to characterize the degree to which the current operating condition belongs to different scheduling modes. The system pre-identifies several typical scheduling scenarios through historical samples and determines the cluster center coordinates of each scenario in the dimensionless feature space. By calculating the Euclidean distance between the multidimensional dimensionless feature vector and each cluster center, and using a continuous function for mapping, a set of continuous soft scenario memberships is obtained. This method of dividing soft scenarios avoids abrupt changes in scheduling rules caused by hard boundaries, enabling the decision-making process to smoothly adapt to the transition from the dry season to the flood season.

[0067] Step 104: Input the input feature set into the pre-trained prediction base model specific to each scheduling scenario, and combine it with the pre-estimated scenario posterior parameters to generate the local prediction mean and local prediction bias corresponding to each scheduling scenario.

[0068] In this embodiment, each scheduling scenario is configured with a set of pre-trained prediction base models, which are optimized for parameters of samples in the predetermined scenario during the offline phase. By inputting the 8-dimensional input feature set into these models in parallel, the model clusters within each scenario will output corresponding initial prediction results. Furthermore, the initial results are linearly biased and weighted ensembled using the scenario's posterior parameters to obtain the prediction mean and prediction bias that reflect the characteristics of the local scenario.

[0069] Step 105: Based on the soft scene membership degree, the local prediction mean and local prediction deviation corresponding to each scheduling scene are weighted and synthesized to obtain the predicted outflow and prediction variance of the target reservoir.

[0070] Furthermore, the system uses soft scenario membership as weights to globally synthesize the local results of all scenarios. The predicted outbound flow is output as the final scheduling suggestion. Simultaneously, the prediction variance is obtained using a variance synthesis formula, which quantitatively characterizes the uncertainty or risk range of the current scheduling decision. When the prediction variance is large, it can alert scheduling personnel that the current operating condition is in a complex scenario boundary zone, requiring enhanced manual review.

[0071] In some optional implementations, the aforementioned historical operating sequences can be subjected to outlier removal and trend smoothing after acquisition to improve the signal-to-noise ratio of the input data. Furthermore, the predicted outflow can be limited by incorporating the actual current-carrying capacity constraints of each power station to ensure that the output scheduling rules conform to the engineering physical boundaries.

[0072] In the embodiments of the present invention, the process of constructing the optimal scheduling sample set by the multi-timescale nested optimization mechanism is further described in detail.

[0073] In one possible implementation, each pre-trained prediction base model is trained using a pre-constructed optimal scheduling sample set; the optimal scheduling sample set is constructed through a multi-timescale nested optimization mechanism, such as... Figure 2 As shown, it includes the following steps:

[0074] Step 201: Based on the hydrological boundary data in the historical operation sequence and the pre-configured engineering constraint parameters, construct multi-objective optimization scheduling models including monthly, weekly and daily scales respectively.

[0075] In this embodiment, the multi-objective optimization scheduling model is established to meet the operational needs of the Yalong River cascade hydropower reservoir group. Models at three time scales—monthly, weekly, and daily—are established to address different scheduling cycle requirements. The monthly model primarily handles long-term water balance and power generation planning, the weekly model is used for medium-term load allocation, and the daily model corresponds to the actual daily operating trajectory of the power stations. All models must conform to constraints related to water balance, reservoir water level, downstream flow, power station output, and the cascade hydraulic connection equations.

[0076] Step 202: Solve the multi-objective optimization scheduling model at the monthly scale and extract the monthly target water level of each reservoir.

[0077] Specifically, the monthly-scale model aims to obtain the optimal water level control path throughout the entire scheduling period. By solving the model, the reservoir water level that should be maintained at each time point can be obtained. The water level at the end of each monthly scheduling period is extracted as the target water level for the month. This water level represents the reservoir's water storage state under the premise of maximizing long-term benefits and is a decision variable for the long-term hydropower resource allocation of cascade power stations.

[0078] Step 203: The target water level at the end of the month is injected as an upper boundary constraint into the weekly-scale multi-objective optimization scheduling model for solution, and the weekend target water level of each reservoir is extracted.

[0079] In this embodiment, the solution of the weekly-scale model is not performed independently, but is rigidly constrained by the monthly-scale scheduling results. Specifically, when solving the weekly-scale multi-objective optimization model, the final water level of the last week of the month is set as the monthly target water level output by the corresponding monthly-scale model. This nested mechanism ensures that the medium-term scheduling plan does not deviate from the long-term macro plan, taking into account both short-term benefits and long-term water storage and power generation potential. By solving the weekly model, the weekend target water level at the end of each weekly period is further extracted.

[0080] Step 204: The target water level at the weekend is used as an upper boundary constraint and injected into the daily-scale multi-objective optimization scheduling model for solution. The daily-scale optimal scheduling trajectory sequence that satisfies the constraint is selected and compiled into an optimal scheduling sample set.

[0081] Accordingly, the daily-scale model uses the weekend target water level as the control boundary at the end of the scheduling period. When solving the daily-scale model, the system searches for a series of intraday water level changes and flow trajectories that satisfy physical constraints and achieve optimal multi-objective comprehensive benefits. The daily-scale optimal scheduling trajectory sequence specifically includes the inflow, initial water level, and corresponding optimal outflow for each calculation period. This type of sequence reflects the optimal scheduling decision logic under complex constraints. After being digitized and compiled, it forms an optimal scheduling sample set for training artificial intelligence models.

[0082] In this embodiment, the input data for the optimal scheduling sample set is constructed based on measured hydrological sequences from the watershed. The data spans multiple complete hydrological years and covers typical inflow year patterns such as abundant, normal, and dry years. Those skilled in the art can construct an optimal scheduling sample set that matches their research object based on the length and quality of hydrological observation data from the actual watershed, using the aforementioned multi-timescale nested optimization mechanism.

[0083] In one possible implementation, the optimization objectives of each multi-objective optimization scheduling model specifically include: maximizing the cascade power generation, maximizing the water supply guarantee rate of the reservoir group, and maximizing the ecological flow guarantee rate of the river section.

[0084] In this embodiment, the model employs three parallel optimization objectives. The power generation objective is obtained by integrating the output of all power stations within the cascade during different time periods. The water supply objective is represented by the ARWA (Aqueous Area Water Supply Guarantee Rate), which reflects the percentage of time during which the water demand of a water supply node is met. The ecological objective is represented by the AREF (Aqueous Area Ecological Flow Guarantee Rate), used to evaluate the degree to which the flow at the control section meets the ecological flow requirements. The specific calculation formulas are as follows:

[0085] Water supply guarantee rate ARWA = (1 / (J*T))*∑∑pj,t*100%;

[0086] Where ARWA is the water supply guarantee rate of the reservoir group, J is the number of water supply nodes, T is the total number of time periods, and pj,t is the water supply status bit of the j-th node in time period t. When the actual water supply of the j-th node in time period t is not less than its water demand, pj,t is 1; otherwise, it is 0.

[0087] River section ecological flow guarantee rate AREF = (1 / (L*T))*∑∑ml,t*100%;

[0088] Where AREF is the ecological flow guarantee rate of the river section, L is the number of ecological flow control sections, T is the total number of time periods, and ml,t is the ecological flow status of the l-th control section in time period t. When the actual flow of the l-th control section in time period t is not less than its preset ecological flow requirement, ml,t is set to 1; otherwise, it is set to 0.

[0089] In one possible implementation, the multi-objective optimization scheduling model is solved by: using a non-dominated sorting genetic algorithm with a preset population size to iteratively calculate the model considering the optimization objective and generate a non-dominated optimal solution set; and using a multi-attribute decision model based on the approximation of ideal solution sorting method to optimize the non-dominated optimal solution set and obtain the optimal scheduling trajectory of each reservoir at the corresponding time scale.

[0090] To address the three competitive optimization objectives mentioned above, this embodiment employs the non-dominated sorting genetic algorithm NSGA-III. The population size and maximum number of iterations for the NSGA-III can be determined by those skilled in the art through conventional trial calculations based on computational resources and convergence requirements. The algorithm evolves through selection, crossover, and mutation operators, and utilizes a reference point mechanism to maintain population diversity, searching for a series of non-dominated optimal solutions in the target space, forming a Pareto front. Furthermore, a multi-attribute decision-making model based on the TOPSIS (Topological Approximation of Ideal Solutions) method is used to evaluate the non-dominated solution set. Specifically, by calculating the relative proximity of each solution to the positive and negative ideal solutions, the individual with the highest comprehensive score is selected from the solution set and used as the final optimal scheduling trajectory for each reservoir at the corresponding time scale.

[0091] Furthermore, to improve the stability of the solution, the optimization process can be run independently multiple times and the statistically optimal result can be taken. For the flow rate and output constraints involved, the relationship curves between water level and reservoir capacity, and between water level and output can be transformed into corresponding water level constraint boundaries during the calculation process, thereby improving the convergence speed of the algorithm and the search efficiency of feasible solutions.

[0092] In the embodiments of the present invention, the specific steps for extracting physical state features that reflect the coupling relationship between the hydrological situation of a single reservoir and the spatial relationship of the cascade are further described in detail.

[0093] In one possible implementation, physical state features reflecting the coupling relationship between the hydrological situation of a single reservoir and the spatial relationship of the cascade system are extracted to obtain a multidimensional dimensionless feature vector, such as... Figure 3 As shown, it includes the following steps:

[0094] Step 301: Based on the historical operation sequence and the current operation observation status, extract the single reservoir hydrological situation characteristic indicators of each regulating reservoir.

[0095] Specifically, in a cascade reservoir system, the backbone hydropower reservoirs with regulation capabilities primarily contribute degrees of freedom to the overall dispatch strategy, while the outflow from daily regulating or run-of-river power stations is usually constrained by upstream inflows and lacks dispatch flexibility. Therefore, this embodiment extracts characteristics for backbone reservoirs with seasonal, annual, or multi-year regulation capabilities. Single-reservoir hydrological situation characteristic indicators are used to reflect the inflow conditions and water storage adequacy of a single reservoir from a time perspective.

[0096] In one possible implementation, the hydrological characteristic indicators for a single reservoir include inflow abundance / difference indicators, water storage utilization rate indicators, and inflow trend indicators; wherein, the hydrological characteristic indicators for each regulating reservoir are extracted, such as... Figure 4 As shown, it includes:

[0097] Step 301a: For each regulating reservoir, extract the current inflow from the current operating observation status, and calculate the ratio of the current inflow to the pre-configured average inflow for the same period over many years to obtain the water abundance / shortage index.

[0098] In this embodiment, the inflow abundance / scarcity index is used to characterize the degree of deviation of the current inflow from the historical average for the same period. The specific calculation formula is as follows:

[0099] α t_m =I t_m / I tau_m ;

[0100] Where, α t_m I represents the water inflow index of reservoir m during time period t. t_m Let I be the current inflow rate of reservoir m during time period t. tau_m The pre-configured average inflow of reservoir m during the same calendar period as time t.

[0101] The above division operation can eliminate the seasonal fluctuations in inbound flow. When α t_m When α > 1, it indicates that the current period is in a high-water state; when α t_m A value less than 1 indicates that the current period is in a low-water state. Using relative ratios instead of absolute flow values ​​makes the inflow characteristics of reservoirs with different drainage areas comparable.

[0102] Step 301b: Extract the current water storage from the current operating observation status, and calculate the ratio of the available water after deducting the pre-configured dead storage capacity from the current water storage to the pre-configured total regulating storage capacity to obtain the water storage utilization rate index.

[0103] Specifically, the water storage utilization rate index is used to uniformly characterize the water storage status of each reservoir in a dimensionless space. The calculation formula is as follows:

[0104] β t_m =(Vt_m -V dead_m ) / V reg_m ;

[0105] Where, β t_m V represents the water storage utilization rate of reservoir m during time period t. t_m Let V be the current water storage of reservoir m at time t. dead_m V is the dead storage capacity of the pre-configured reservoir m. reg_m The total regulating capacity of the pre-configured reservoir m.

[0106] This calculation process maps the water level status of all regulating reservoirs to the interval between 0 and 1. When β t_m When β approaches 1, it indicates that the reservoir is approaching its normal water level; when β... t_m When the value approaches 0, it indicates that the reservoir is nearing the dead water level.

[0107] Step 301c: Extract the inflow rate of the previous period from the historical operation sequence, calculate the difference between the current inflow rate and the inflow rate of the previous period, and compare the difference with the average inflow rate of the same period over many years to obtain the inflow trend index.

[0108] To capture the dynamic characteristics of hydrological processes, this embodiment introduces an indicator of inflow trend, the specific calculation formula of which is as follows:

[0109] γ t_m =(I t_m -I t_1_m ) / I tau_m ;

[0110] Where, γ t_m I is an indicator of the inflow trend of reservoir m during time period t. t_m For the current inbound flow, I t_1_m This represents the inbound flow in the previous period, I tau_m This represents the average inbound flow over multiple years during the same period. In this embodiment, I... tau_m The average inbound flow rate over multiple years is used instead of the inbound flow rate of the previous period to prevent abnormal divergence of the indicator value when the flow rate of the previous period is close to 0. When γ t_m A value greater than 0 indicates that the flow rate is on an upward trend; conversely, a value less than 0 indicates that the flow rate is on a downward trend.

[0111] Step 302: Based on the hydrological situation characteristic indicators of a single reservoir, extract the cascade coupling characteristic indicators between adjacent regulating reservoirs.

[0112] Correspondingly, in addition to single reservoir features, the scheduling strategy of cascade reservoirs depends on the water storage and release game and coordination between upstream and downstream reservoirs. The cascade coupling feature index aims to extract spatial correlation information that cannot be covered by single reservoir features.

[0113] In one possible implementation, the cascade coupling characteristic indicators include cascade water storage gradient indicators and inflow adjustable reservoir capacity matching indicators; wherein, cascade coupling characteristic indicators between adjacent regulating reservoirs are extracted, such as... Figure 5 As shown, it specifically includes:

[0114] Step 302a: For adjacent upstream and downstream regulating reservoirs, calculate the difference between the water storage utilization rate index of the upstream reservoir and the water storage utilization rate index of the adjacent downstream reservoir to obtain the cascade water storage gradient index.

[0115] In this embodiment, the cascade water storage gradient index is used to calculate the degree of imbalance in water storage status between adjacent reservoirs. The specific calculation formula is as follows:

[0116] δ t =β t_up -β t_down ;

[0117] Where, δ t β represents the cascade water storage gradient index of adjacent upstream and downstream reservoirs at time t. t_up β is an indicator of the water storage utilization rate of the upstream reservoir. t_down This is an indicator of the water storage utilization rate of adjacent downstream reservoirs.

[0118] When δ t When δ > 0, it indicates that the upstream reservoir is relatively full while the downstream reservoir is relatively empty. In this case, the scheduling strategy tends to increase upstream outflow to replenish downstream water. t When the value is less than 0, it indicates that there is flood control pressure downstream, and the dispatching strategy tends to intercept and store floodwater upstream. This indicator constitutes the basis for determining the coordinated dispatching of upstream and downstream areas.

[0119] Step 302b: Multiply the inflow abundance / dampness index of each regulating reservoir with the water storage utilization rate index to obtain the inflow adjustable reservoir capacity matching index, which reflects the interaction between inflow intensity and available water storage space.

[0120] Accordingly, to calculate the nonlinear coupling risk between inflow intensity and reservoir capacity, this embodiment constructs an adjustable inflow reservoir capacity matching index, the specific calculation formula of which is as follows:

[0121] ε t_m =α t_m *β t_m ;

[0122] Where, ε t_m Let α be the adjustable reservoir capacity matching index for reservoir m during time period t. t_m β is an indicator of water abundance or scarcity. t_m α is an indicator of water storage utilization rate. Physically, when the inflow of water is abundant and the reservoir is near full capacity, α... t_m With β t_mBoth are at high levels, and their product ε t_m The increase indicates a heightened risk of water spillage, enabling the model to capture the scheduling boundary conditions that require flood discharge.

[0123] Step 303: The single-reservoir hydrological situation characteristic indicators and the cascade coupling characteristic indicators are spliced ​​together to obtain a multidimensional dimensionless feature vector.

[0124] Specifically, taking a cascade system containing three regulating reservoirs as an example, the inflow abundance / dampness index, water storage utilization rate index, and inflow change trend index are calculated for each reservoir, yielding 9-dimensional single-reservoir features. For adjacent reservoirs, the cascade water storage gradient index yields 2-dimensional features, and the inflow adjustable storage capacity matching index yields 3-dimensional features. These features are concatenated to form a 14-dimensional feature vector. Since all components of this feature vector have been converted into dimensionless ratios or differences, eliminating inconsistencies in physical dimensions, subsequent distance mapping clustering calculations do not require mean-variance normalization, thus avoiding distribution distortion caused by feature scaling.

[0125] This embodiment constructs multi-dimensional dimensionless features including cascade water storage gradient and inflow-adjustable reservoir capacity matching degree, eliminating the interference of differences in reservoir capacity and flow rate, and characterizing the dynamic synergy between flood control pressure and reservoir water storage potential, so as to improve the algorithm's ability to identify hydrological scenarios.

[0126] In the embodiments of the present invention, the soft scene membership mapping mechanism and the entire process of soft scene membership Bayesian model fusion prediction are further described in detail.

[0127] In one possible implementation, the preset bandwidth parameter configuration is adaptively determined through cross-validation, such as... Figure 6 As shown, it includes the following steps:

[0128] Step 401: Calculate the shortest distance from each pre-divided training sample to the cluster center of the scheduling scene it does not belong to, and extract the median value of all shortest distances.

[0129] Specifically, to avoid errors introduced by manually specifying distance decay parameters, the system performs statistical analysis on distance features offline. For all pre-divided training samples, the Euclidean distance from each training sample to the cluster centers of all other scheduling scenarios that do not contain that sample is calculated, and the minimum distance value is selected as the shortest non-member distance for that sample. After traversing all training samples, all shortest non-member distances are sorted by numerical value, and the value in the middle of the sequence is extracted as the median value. This median value reflects the general density of overlap between scheduling scenarios in the feature space.

[0130] Step 402: Construct a set of candidate bandwidth parameters based on a preset ratio multiple of the median value.

[0131] In this embodiment, the system constructs a search grid using the extracted median value. The preset ratio can be selected as 0.1, 0.2, 0.5, 1.0, 2.0, and 5.0. Multiplying the median value by each of these preset ratios yields a discrete set of values, which is then determined as the candidate bandwidth parameter set.

[0132] Step 403: Substitute each candidate bandwidth parameter into the distance decay function and use the pre-divided cross-validation set to calculate the root mean square error of the predicted outflow.

[0133] Furthermore, the system sequentially iterates through each value in the candidate bandwidth parameter set, substituting it into the model calculation process as the bandwidth scale of the Gaussian kernel function. A pre-defined cross-validation set is used as input to the model for simulation prediction, and the root mean square error (RMSE) between the predicted outbound flow and the actual target value is calculated. The RMSE serves as a target indicator for quantitatively evaluating the prediction accuracy under the current bandwidth parameters.

[0134] Step 404: Select the candidate bandwidth parameter that minimizes the root mean square error, and use it as the preset bandwidth parameter configuration for executing the distance decay function.

[0135] Accordingly, by comparing the root mean square error under different candidate values, the system extracts the set of parameters with the smallest corresponding error value and solidifies it as the preset bandwidth parameter configuration in the online prediction stage. This parameter plays an implicit regularization role in the algorithm. When the training data is sufficient and the scene boundaries are clear, the optimized bandwidth parameter value is relatively small, and the model tends to block information sharing between different scenes; when extreme hydrological years lead to sparse local samples, the optimized bandwidth parameter value is relatively large, and the model automatically enhances the smooth sharing of weights between scenes to prevent overfitting.

[0136] In one possible implementation, the soft scene membership degree of the current operational observation state relative to each scheduling scenario is determined based on the spatial distance mapping between multidimensional dimensionless feature vectors and pre-constructed cluster centers of each scheduling scenario, including:

[0137] Optionally, the spatial feature distance between the multidimensional dimensionless feature vector and each pre-built scheduling scenario cluster center is calculated.

[0138] As an optional implementation, during the online prediction phase, the system acquires the current operational observation state input and extracts a 14-dimensional multidimensional dimensionless feature vector according to the aforementioned embodiments. Further, the system calculates the squared Euclidean distances between this multidimensional dimensionless feature vector and the known cluster centers of all scheduling scenarios, using these squared distances as the spatial feature distances.

[0139] Optionally, a distance decay function configured based on preset bandwidth parameters can be used to map each spatial feature distance to the corresponding scheduling scenario's attribution weight coefficient.

[0140] In this embodiment, the distance decay function adopts the form of a Gaussian kernel function; the larger the spatial feature distance, the smaller the coefficient obtained by mapping. The specific calculation formula is as follows:

[0141] W c =exp(-0.5*D c / h 2 );

[0142] Among them, W c D represents the weighting coefficient of scheduling scenario c corresponding to the current operational observation state. c is the spatial feature distance between the multidimensional dimensionless feature vector and the cluster center of the scheduling scenario c, and h is the preset bandwidth parameter configuration.

[0143] Optionally, the proportion of the attribution weight coefficient of each scheduling scenario in the total attribution weight coefficient of all scheduling scenarios is determined as the soft scenario membership degree of the current running observation state relative to the corresponding scheduling scenario.

[0144] Furthermore, the system sums the membership weight coefficients of all scenarios corresponding to the current state, and then divides the coefficient of a single scenario by this summation to complete the normalization process. This normalization ratio is the soft scenario membership degree. To reflect the resilience of this calculation process, the following example of normalized dimensionless numerical computation is provided:

[0145] As an optional implementation, assume there are scenarios A and B in the two-dimensional scheduling space. The system calculates that the normalized squared distance of the current feature vector from the center of scenario A is 0.2, and the normalized squared distance from the center of scenario B is 0.8, while the preset bandwidth parameter is set to 1.0. Substituting into the above exponential decay formula, the non-normalized weight of scenario A is approximately exp(-0.1) = 0.905, and the non-normalized weight of scenario B is approximately exp(-0.4) = 0.670. After normalization, the soft scene membership degree of the current sample to scenario A is approximately 0.57, and the soft scene membership degree to scenario B is approximately 0.43. This example shows that when the scheduling condition is on the edge of the transition between high water and low water, the algorithm does not assign it to a single scenario, but rather allocates a set of continuously changing weight proportions, ensuring the continuity of scheduling decisions.

[0146] In one possible implementation, the pre-estimated scenario posterior parameters include scenario prediction base model weights, prediction base model linear correction intercepts, prediction base model linear correction slopes, and local baseline prediction variances. Combining the pre-estimated scenario posterior parameters, the local prediction mean and local prediction bias corresponding to each scheduling scenario are generated, including:

[0147] Optionally, for each scheduling scenario, the initial outbound flow prediction value output by each prediction base model for the input feature set can be obtained.

[0148] Accordingly, the system inputs a set of input features, including water level and flow time series, into multiple prediction base models in parallel, and extracts the initial outflow forecast value calculated by the forward propagation of the models.

[0149] Optionally, the initial outflow forecast values ​​are corrected for linear deviation by using the linear correction intercept and linear correction slope of the forecast base model, respectively, to obtain the corrected forecast values ​​of each forecast base model.

[0150] Specifically, due to the inherent systematic bias in the nonlinear fitting of each prediction base model, the system extracts the offline-estimated posterior parameters and uses a linear affine transformation to correct the initial values. The specific calculation formula is as follows:

[0151] F i_c =a i_c +b i_c *P i_c ;

[0152] Among them, F i_c To predict the corrected prediction value of base model i in scheduling scenario c, a i_c For the linear correction intercept of the prediction basis model, b i_c To predict the linear correction slope of the base model, P i_c This corresponds to the initial outbound flow forecast value.

[0153] Optionally, the corrected prediction values ​​of each prediction base model are weighted and summed using the weights of the scenario prediction base models to obtain the local prediction mean of the current scheduling scenario.

[0154] Furthermore, the system calls the confidence weights of each prediction base model extracted from the posterior, performs a weighted calculation, and the specific calculation formula is as follows:

[0155] E c =Σ(W i_c *F i_c );

[0156] Among them, E c W is the local prediction mean for scheduling scenario c. i_c For the scene prediction base model weights of prediction base model i, Fi_c Σ represents the corresponding corrected prediction value, where Σ indicates the summation operation performed on all prediction base models in the current scenario.

[0157] Optionally, the local prediction deviation of the current scheduling scenario is calculated by combining the local baseline prediction variance and the degree of dispersion of the corrected prediction values ​​of each prediction base model relative to the local prediction mean.

[0158] In this embodiment, the prediction bias within a single scene consists of two physical mechanisms: one is the inherent fitting variance of each prediction base model, and the other is the degree of disagreement among the models regarding their predictions of the same set of data. The specific calculation formula is as follows:

[0159] V c =Σ(W i_c *var i_c )+Σ(W i_c *(F i_c -E c ) 2 );

[0160] Among them, V c For the local prediction bias of scheduling scenario c, W i_c For the weights of the base model for scene prediction, var i_c F represents the local baseline prediction variance of the prediction base model i. i_c To correct the predicted value, E c This is the local predicted mean.

[0161] In one possible implementation, based on the soft scene membership degree, the local prediction mean and local prediction deviation corresponding to each scheduling scene are weighted and synthesized to obtain the predicted outflow and prediction variance of the target reservoir, including:

[0162] Optionally, the predicted outbound flow can be obtained by weighting and summing the local prediction mean of all scheduling scenarios using the soft scenario membership degree of each scheduling scenario as a weighting coefficient.

[0163] After obtaining all local means, the system performs a vector dot product summation operation with the soft scene membership degree, and finally outputs the predicted outflow as the actual reservoir control instruction issued.

[0164] Optionally, based on local prediction bias and the membership degree of each soft scene, the intra-scene variance component reflecting the prediction divergence of the model within a single scene is calculated.

[0165] To measure scheduling risk, the system initiates full variance decomposition. The variance components within the scenario are calculated as follows:

[0166] Var intra =Σ(U c *Vc );

[0167] Among them, Var intra U represents the variance component within the scene. c V represents the soft scene membership degree of the scheduling scene c. c For local prediction bias, Σ represents the summation over all partitioned scheduling scenarios.

[0168] Optionally, based on the degree of deviation of each local prediction mean relative to the final predicted outbound flow, and combined with the membership degree of each soft scenario, the inter-scenario variance component reflecting the uncertainty of cross-scenario affiliation is calculated.

[0169] The system continues to calculate the variance component between scenes. This component mainly measures the discrepancy in scene transitions caused by ambiguity in hydrological conditions. Its calculation formula is as follows:

[0170] Var inter =Σ(U c *(E c -E total ) 2 );

[0171] Among them, Var inter U represents the variance components between scenes. c For soft scene membership, E c E is the local predicted mean. total To predict outbound flow.

[0172] Optionally, the variance components within a scenario are added to the variance components between scenarios to obtain the predicted variance that characterizes the final scheduling risk boundary.

[0173] Furthermore, the system will Var intra With Var inter The variances are accumulated and the predicted variance is output. In practical engineering applications, when reservoirs face complex conditions such as rapid shifts between drought and flood, the samples are located at the intersection of multiple scheduling scenarios. At this time, the variance components between scenarios, Var, are... inter This will increase, leading to a rise in overall prediction variance. When the system determines that the prediction variance exceeds the preset safety alarm threshold, it can switch from automatic scheduling mode to auxiliary analysis mode to ensure the safe operation of the water network.

[0174] This embodiment introduces a continuous soft membership mapping mechanism based on adaptive bandwidth, transforming the originally rigid scenario allocation into a flexible probability distribution. By smoothly weighting the local decisions of multiple scenarios, the system achieves continuous translational transition of dispatch instructions during different hydrological year types and the alternation of flood and dry seasons.

[0175] Simultaneously, by utilizing the full variance decomposition mechanism, the internal variance representing the model fitting discrepancies and the inter-scenario discrepancy variance representing the ambiguous boundaries are separated. This composite prediction variance provides a quantitative benchmark for risk warning and human intervention in cascade reservoirs under complex operating conditions.

[0176] In the embodiments of the present invention, the offline training of the artificial intelligence model and the posterior parameter estimation mechanism are further described in detail.

[0177] In one possible implementation, before training each prediction base model using the optimal scheduling sample set, the method further includes a stratified sampling partition of the optimal scheduling sample set, comprising the following steps:

[0178] Step 501: Perform cluster analysis on the optimal scheduling sample set to divide the overall sample into multiple subsets of scheduling scenario samples with corresponding scenario labels.

[0179] In a further embodiment, step 501 may also be: extracting the multidimensional dimensionless feature vectors corresponding to each sample in the optimal scheduling sample set, performing cluster analysis on the optimal scheduling sample set, dividing the overall sample into multiple scheduling scenario sample subsets with corresponding scenario labels, and recording the cluster center of each scheduling scenario sample subset.

[0180] Specifically, the system acquires a historical set of optimal scheduling samples and extracts the dimensionless feature vectors of all samples. To make the model learning more targeted, the system uses hierarchical clustering to group samples with similar scheduling strategies. When determining the optimal number of clusters, the system constructs a candidate cluster size set containing multiple candidate values, with a preset candidate range of 2 to 10. Further, the system iterates through each candidate value, performs clustering calculations, and extracts the corresponding average silhouette coefficient.

[0181] As an optional implementation, under the premise of ensuring that the sample size under each cluster division meets the pre-configured minimum training sample size constraint, the candidate value that maximizes the average silhouette coefficient is selected from the candidate cluster number set as the target cluster number; the cluster analysis is completed based on the target cluster number, and the sample subset of each scheduling scenario is output.

[0182] Specifically, when selecting the final number of clusters, the system aims to maximize the silhouette coefficient while also ensuring a balanced sample distribution. The system pre-configures a minimum training sample size constraint, stipulating that the number of samples in any cluster subset must not be less than the total sample size. From the candidate values, the system eliminates values ​​that result in small sample categories and selects the value with the largest silhouette coefficient from the remaining candidate values ​​as the target number of clusters. After partitioning, each scheduling scenario sample subset is assigned a discrete scenario label, such as representing a low-water period or a high-water period.

[0183] Step 502: For each subset of scheduling scenarios, determine whether the number of samples in the subset is lower than the pre-configured small sample judgment threshold.

[0184] Correspondingly, since extreme hydrological years are less likely to occur, the sample size of some scheduling scenario subsets may be small. The system configures a small sample determination threshold; in this embodiment, this threshold can be set to 10. Furthermore, the system performs capacity determination on each subset.

[0185] Step 503: For the subset of scheduling scenario samples whose number of subset samples is not lower than the small sample judgment threshold, according to the pre-configured division ratio, randomly extract and split the corresponding scenario training subset and scenario test subset within it.

[0186] Optionally, for a subset of scheduling scenario samples that meets the capacity requirements, the system executes a stratified sampling strategy. In this embodiment, the pre-configured partition ratio can be set to 7:3. The system independently performs random sampling within each subset, allocating 70% of the data to the scenario training subset and the remaining 30% to the scenario testing subset.

[0187] Step 504: For the subset of scheduling scenario samples where the number of subset samples is lower than the small sample judgment threshold, a leave-one-out cross-validation strategy is used to divide the corresponding scenario training subset and scenario test subset.

[0188] Accordingly, for the subset of scheduling scenario samples determined to be small, the system automatically switches to leave-one-out cross-validation mode. In this mode, if the subset contains N samples, the system performs N partitions in a loop. In each partition, one sample is retained as the scenario test subset, and the remaining N-1 samples are used as the scenario training subset.

[0189] Step 505: Summarize the scenario training subsets obtained from the sample subsets of each scheduling scenario for independent training of the prediction base model specific to each scheduling scenario; summarize the test subsets of each scenario for subsequent model parameter verification and bandwidth parameter optimization.

[0190] Through the aforementioned adaptive partitioning mechanism, the system extracts a subset of scene training data that is bound to each scene label. This operation ensures that extreme hydrological conditions are not obscured by the main flow data during model training and validation.

[0191] According to one aspect of this application, training various prediction base models using a scene training subset specifically includes:

[0192] Optionally, multiple pre-defined prediction base models can be assigned to each scheduling scenario.

[0193] In this embodiment, the system instantiates a set of prediction base models in parallel for the extracted scene training subset. To address the complex nonlinear relationships and long-term dependencies in reservoir scheduling, the preset model set may include Random Forest (RF), Gradient Boosting Tree (GB), Long Short-Term Memory (LSTM), and Bidirectional Long Short-Term Memory (BiLSTM).

[0194] Optionally, with the goal of minimizing the prediction error of outbound flow, the Bayesian optimization algorithm is used to search for hyperparameters of each prediction base model on the corresponding scenario training subset to obtain the optimal hyperparameter combination of each model.

[0195] As an optional implementation, the system utilizes Bayesian optimization to replace exhaustive search. For each model, the system establishes a corresponding hyperparameter search space; for example, for the LSTM model, the search space covers parameters such as the number of hidden layer units and the learning rate. The system uses the root mean square error (RMSE) as the objective function and iteratively searches for the hyperparameter combination that minimizes the RMSE on a predetermined subset of training scenarios.

[0196] Optionally, multiple prediction base models can be trained independently using the optimal hyperparameter combination and the corresponding scenario training subset to obtain a pre-trained prediction base model specific to the current scheduling scenario.

[0197] Accordingly, the system extracts the optimal hyperparameter combination and initializes the prediction base model. During the model training phase, the mean squared error (MSE) of outbound traffic is used as the loss function for parameter optimization. For the random forest model and gradient boosting tree model, the tree is constructed using the minimization of the sample mean squared error as the splitting criterion. Those skilled in the art can use conventional optimization methods such as backpropagation and stochastic gradient descent to train the deep learning model. The hyperparameters during training can be determined through the aforementioned Bayesian optimization steps. Then, a subset of the input scenario training data is used to perform backpropagation updates of the model parameters. After training convergence, the system saves the weight parameters, obtaining a prediction base model customized for the current scheduling scenario.

[0198] In another possible implementation, the pre-estimated posterior parameters of the scene are obtained through iterative estimation using an offline expectation-maximization algorithm, specifically including:

[0199] Optionally, for each scheduling scenario, deterministic optimized outbound traffic is extracted from the corresponding scenario training subset as the actual scheduling value, and linear regression is performed between the actual scheduling value and the corresponding pre-trained prediction base model output value to determine the linear correction intercept and the linear correction slope of the prediction base model.

[0200] Further, after acquiring the prediction base models for each scenario, the system enters the posterior parameter estimation stage. The system re-inputs the training subset of the same scenario into the trained prediction base model to obtain the output prediction sequence. Then, using the deterministic optimized outbound flow sequence in the original sample as the true dependent variable and the prediction sequence as the independent variable, least squares linear regression is performed to extract the intercept and slope terms, eliminating the inherent structural biases of each model.

[0201] Optionally, the initial weights and initial prediction variances of each pre-trained prediction base model are initialized.

[0202] In this embodiment, the Expectation-Maximization (EM) algorithm is used to estimate the average posterior parameters of the Bayesian model. The system pre-assigns equal initial weights (1 / k) to the k prediction base models. Simultaneously, the average of the squared residuals of each model on the training set is extracted as the initial prediction variance.

[0203] Optionally, the expectation step and the maximization step are executed alternately. The expectation step is used to calculate the posterior responsibility based on the Gaussian probability density function. The posterior responsibility represents the probability that the true scheduling value is generated by each prediction base model under the current weights and prediction variance. The maximization step is used to update the weights and prediction variance of each model based on the posterior responsibility of each prediction base model.

[0204] The system executes iteratively within a given upper limit on the number of iterations (e.g., 1000). In the expected step, i.e., the E-step, the system calculates the posterior responsibility using the following formula:

[0205] z i_t =(w i *g i_t ) / Σ(w j *g j_t );

[0206] Among them, z i_t For the posterior responsibility of the prediction base model i for sample t, w i To predict the current weights of base model i, g i_t Let be the probability density of the true scheduling value under the Gaussian distribution corresponding to the prediction base model i, and the denominator is the sum of the weighted probability densities of all prediction base models.

[0207] Specifically, in the maximization step, i.e., the M-step, the system updates the global parameters using the calculated posterior responsibility. For the prediction base model i, the posterior responsibility of all its corresponding samples is extracted and averaged, and this value is used as its update weight in the next iteration. Simultaneously, the responsibility is extracted as a weight, and a weighted average of the squared residuals is calculated to obtain the updated prediction variance.

[0208] Optionally, the iteration stops when the relative change of the log-likelihood function meets the preset convergence condition, the weights of each converged model are extracted as the scene prediction base model weights, and the prediction variances of each converged model are extracted as the local baseline prediction variances.

[0209] As an optional implementation, the system calculates the global log-likelihood function after each iteration. The algorithm is considered converged when the relative change in the log-likelihood function value of the current iteration compared to the previous iteration is less than a preset threshold. At this point, the weights and variance generated in the last iteration are extracted.

[0210] Optionally, the linear correction intercept of the prediction base model, the linear correction slope of the prediction base model, the scene prediction base model weights, and the local baseline prediction variance are compiled into pre-estimated scene posterior parameters. Further, the system packages and stores the four parameters extracted above. Since all calculations are performed independently within a predetermined subset of the scheduling scene training, this parameter package constitutes a posterior parameter set specific to the predetermined scheduling scene, which can be extracted and invoked during the online prediction phase.

[0211] Meanwhile, by adaptively dividing massive operational samples into multiple relatively independent scheduling scenarios, and independently executing hyperparameter optimization and posterior weight iteration of the underlying prediction base model within a local environment, this scenario decoupling mechanism breaks the inherent upper limit of the global model's ability to fit extreme droughts or massive flood peaks, improving the model's prediction accuracy and robustness under abnormal conditions such as rapid shifts between drought and flood.

[0212] In an embodiment of the present invention, another method for implementing the prediction base model allocation is provided, namely, using a combination of traditional machine learning models instead of a combination of hybrid deep learning models. The remaining steps are the same as in the foregoing embodiments and can be referred to the foregoing description, and will not be repeated here.

[0213] In one possible implementation, multiple preset prediction base models are assigned to each scheduling scenario.

[0214] In this embodiment, optionally, for a predetermined hardware environment with limited data volume or strict limitations on computing resources, the preset multiple prediction base models can be composed of a combination of artificial neural network models, random forest models, support vector machine models, and extreme gradient boosting tree models. This combination serves as an alternative to the aforementioned deep learning combination including long short-term memory networks, and its input layer data interface and feature dimension remain absolutely consistent, all receiving the aforementioned constructed 8-dimensional input feature set.

[0215] Specifically, artificial neural network models establish a nonlinear mapping between feature inputs and predicted outflow outputs through a multilayer perceptron structure. Support vector machine models utilize kernel functions to map the input multidimensional dimensionless feature vectors to a high-dimensional space, and perform regression calculations by finding the maximum margin hyperplane; this mechanism is suitable for fitting calculations in small-sample scheduling scenarios. Limit gradient boosting tree models, by serially integrating multiple decision trees and optimizing the objective loss function using second-order Taylor expansion, are used for regression prediction of hydrological extreme value data.

[0216] Accordingly, after the aforementioned prediction base model is established, the system executes the Bayesian optimization algorithm described above to perform hyperparameter search. The hyperparameter search space is adjusted accordingly for the model set replaced in this embodiment. For example, for the support vector machine model, the search space includes the penalty coefficient and kernel function bandwidth parameters; for the extreme gradient boosting tree model, the search space includes the learning rate, maximum tree depth, and number of base learners.

[0217] Furthermore, after completing model allocation and hyperparameter search, the system independently trains the aforementioned traditional machine learning models using a subset of scene training data. In the subsequent posterior parameter estimation stage of the Expectation-Maximization (EM) algorithm, the system extracts the predicted and actual values ​​of each model and calculates the initial prediction variance. Because the residual distributions of the Support Vector Machine (SVM) and Limit Gradient Boosting Tree (LTGB) models differ from those of deep learning models under a predetermined hydrological year, the EEM algorithm, during its iterative update phase, automatically assigns higher scene prediction base model weights to traditional models with smaller fitting errors in the current scene subset, based on data-driven principles.

[0218] Furthermore, by providing this alternative, it is demonstrated that the scheduling rule extraction method provided by this invention does not depend on a predetermined underlying physical model architecture. Whether it's an ensemble tree algorithm combination or a deep learning combination including recurrent neural networks, as long as it can provide initial outbound traffic prediction values ​​for a predetermined scenario, it can be integrated into the global framework of soft-scenario membership Bayesian model fusion prediction.

[0219] According to one aspect of this application, the method further includes:

[0220] Based on the scheduling needs and functions of a group of reservoirs in the middle and lower reaches of a certain water area, the optimized scheduling model mainly considers three scheduling objectives: power generation, water supply, and ecology. The power generation objective is represented by cascade power generation, the water supply objective by the water supply guarantee rate, and the ecological objective by the ecological flow guarantee rate. The flood control objective is considered through the constraint of the flood control limit water level. The model mainly consists of two parts: the objective function and the constraints.

[0221] 1) Objective function:

[0222] ① The cascade power generation is the largest (monthly, weekly, and daily scales):

[0223] ;

[0224] In the formula: CPG represents the cascaded power generation; and The generating head and generating flow rate of power station i during time period t are m and m, respectively. 3 / s; The overall output coefficient of power station i; The time period is the length of the period (monthly, weekly, or daily); n is the number of power stations; and T is the number of scheduling periods.

[0225] ② Maximum water supply guarantee rate (monthly, weekly, and daily scales):

[0226] ;

[0227] ;

[0228] In the formula: ARWA is the water supply guarantee rate of the reservoir group; and These represent the water supply and demand at node j during time period t, in billions of m³. 3 J represents the number of water supply nodes.

[0229] ③ The river section with the highest ecological flow guarantee rate (monthly, weekly, and daily scales):

[0230] ;

[0231] ;

[0232] In the formula: AREF is the ecological flow guarantee rate of the river section; and The flow and ecological flow demands at the control section l during time period t are respectively, m 3 / s; L represents the number of ecological flow control sections.

[0233] 2) Constraints:

[0234] Solving the above objective function requires compliance with a series of physical and operational constraints, mainly including water balance constraints, water level constraints, flow constraints, output constraints, water level constraints at the beginning and end of the scheduling period, and the hydraulic connection equations of the cascade reservoirs.

[0235] 3) Solving the optimized scheduling model:

[0236] In this embodiment, the third-generation non-dominated sorting genetic algorithm NSGA-III is selected to solve the optimization scheduling model. This algorithm is based on the NSGA-II algorithm, introduces the reference point method to enhance the diversity of the population, and overcomes the shortcomings of the previous generation algorithm, such as the crowding distance not being suitable for high-dimensional spaces. It has the advantages of fast computation speed, strong robustness, and uniform distribution of non-dominated optimal solutions.

[0237] Based on the actual operational characteristics of reservoir groups and the algorithm's coding method (i.e., the decision variable is water level), the constraints of flow rate and output are transformed into water level constraints. Through trial calculations, the population size of the NSGA-III algorithm was determined to be 1000, and the maximum number of iterations was determined to be 1000. Furthermore, to avoid errors caused by the algorithm's inherent randomness, the algorithm was run independently 10 times.

[0238] 4) Optimal scheduling scheme based on TOPSIS:

[0239] A multi-attribute decision model for scheduling schemes based on the TOPSIS (Topological Solution Approximation) method was established. The Pareto non-dominated solution sets obtained from the power generation-water supply-ecology optimization scheduling model of the middle and lower reaches of the Yalong River at the monthly, weekly, and daily scales were optimized to obtain the optimal scheduling sample sets at the monthly, weekly, and daily scales.

[0240] 4.1) Extraction of optimal scheduling rules for reservoir groups using coupled cluster analysis, prediction base models, and Bayesian theory:

[0241] To address the issue that the accuracy of traditional data mining-based reservoir group scheduling rule extraction is easily affected by the training sample set, model structure, and parameters, a general model for extracting implicit stochastic optimization scheduling rules for reservoir groups is proposed, which couples cluster analysis, prediction base models, and Bayesian theory. This model enhances the representativeness of training samples, reduces the uncertainty of model parameters and structure, and further improves the accuracy of reservoir group scheduling rule extraction methods.

[0242] (1) Determine the input and output variables:

[0243] The most commonly used input variables for reservoir scheduling rules are inflow and water level. Considering the continuous nature of hydrological processes, input factors and scheduling decisions in the preceding periods may influence scheduling decisions in the upcoming period. Furthermore, the hydraulic connections between upstream and downstream reservoirs may also affect scheduling decisions. Therefore, the input variables include the inflow I in the preceding two periods and the upcoming period. t-2 I t-1 and I t The initial water level Z in the first two periods and the period to come. t-2 Z t-1 and Z t Outbound flow in the first two periods O t-2 and O t-1 That is, the input feature set X=[I t-2 ,It-1 ,I t Z t-2 Z t-1 Z t O t-2 O t-1 ].

[0244] Commonly used output variables include the water level at the end of the current period, the reservoir capacity at the end of the current period, the outflow rate during the current period, and the average output during the current period. This project selects the outflow rate O during the current period. t As the output variable, the corresponding decision variable is also the outbound flow rate O during the time period. t .

[0245] (2) Cluster analysis of input variables:

[0246] Hierarchical clustering methods select the distance (such as Euclidean distance) between input sample sets as the clustering criterion. The smaller the distance, the more similar the data objects are, that is, the more similar the implicit scheduling rules are.

[0247] (3) Scheduling rule extraction model construction and hyperparameter optimization based on prediction base model:

[0248] In fact, reservoir group scheduling rule extraction is a fitting problem, that is, using deterministic optimization scheduling results or actual scheduling results as the sample set, and using certain methods to fit the nonlinear relationship between input and output variables. This embodiment uses four representative prediction base models: ANN, RF, SVM, and the extreme gradient boosting algorithm XGBoost, to construct scheduling rule extraction models for different time scales and different modes. Bayesian optimization algorithm is used to optimize the hyperparameters of the prediction base models, where hyperparameter optimization involves finding the hyperparameter combination that best performs on the validation set using a predetermined method. The specific steps are as follows:

[0249] ① Divide the optimal scheduling sample set into a training set and a test set in a 7:3 ratio;

[0250] ② Construct the prediction base model and set the built-in parameters of the model;

[0251] ③ Analyze the sensitivity of different prediction base models to different parameter settings, and determine the parameters to be optimized in the model;

[0252] ④ The Bayesian optimization algorithm was used to optimize the hyperparameters of the four prediction basis models. The Bayesian optimization algorithm is a global optimization algorithm based on Bayesian theory, capable of obtaining near-optimal solutions with minimal evaluation cost. (In the hyperparameter search space...) Internally, the Bayesian optimization algorithm optimizes the objective function through iterative, finite-number trials. In this embodiment, the root mean square error (RMSE) of the model's predicted reservoir outflow is used as the optimization objective function to obtain the hyperparameter combination. : ;

[0253] ⑤ Assign the optimized model parameters to each prediction base model and train the model;

[0254] ⑥ Substitute the sample test set into the trained model to obtain the reservoir outflow predicted by the four models, and calculate and select the correlation coefficient (CC), bias (Bias), and root mean square error (RMSE) to evaluate the extraction results of the reservoir group scheduling rules by each prediction base model.

[0255] (4) Construction of scheduling rule extraction model based on Bayesian model averaging:

[0256] Different prediction base models exhibit significant performance variations when extracting reservoir group scheduling rules across different scales, objects, and dimensions, indicating substantial uncertainty in the model structure. Therefore, this embodiment employs the Bayesian model averaging method, using the scheduling rules extracted by the four prediction base models as input to generate power generation, water supply, and ecological scheduling rules for the Yalong River cascade hydropower reservoir group at monthly, weekly, and daily scales. This reduces the adverse interference of model structure uncertainty on the scheduling decisions of the Yalong River cascade hydropower reservoir group.

[0257] Bayesian Model Averaging (BMA) is a method for synthesizing model results based on Bayesian theory, which can reduce the uncertainty of model structure. In practice, the prediction result of BMA is a weighted average of the predictions from each model, assigning higher weights to models with better performance. In this embodiment, the reservoir outflow O represents the final predicted value based on the Bayesian model averaging. It is the set of simulation results output by k different prediction base models, where k=4 in this embodiment. Let... Let O be the conditional posterior probability distribution of the model simulation value O and the deterministic optimization result set D. Then the Bayesian model-averaged posterior probability distribution of O is:

[0258] ;

[0259] In the formula: Given the set of optimization results D, the result of the i-th prediction base model is the posterior conditional probability of the optimal estimate. If... express ,but The weight As a measure of the probability likelihood that the model results are considered accurate, it reflects the performance of the corresponding model.

[0260] Assumption It is a Gaussian distribution with a standard deviation. The center value is a linear function of the initial model simulation values. ,in and Based on the correction term for the linear fit between the deterministic optimization result value and the simulated value of each prediction base model, the posterior mean and variance of the BMA predicted value of variable O are as follows:

[0261] ;

[0262] ;

[0263] Furthermore, Bayesian models typically require accurate estimation of the weights and variances of each model in the model set. This embodiment employs the expectation-maximization method to determine the probability density function, ultimately obtaining the reservoir group scheduling rules that take into account the uncertainty of the model results. The expectation-maximization method requires initial values ​​for expectation and variance, and iterates between the expectation and maximization steps. In the expectation step, the zi value is estimated given the expectation and variance values. The weights and variances are updated through repeated iterations until convergence.

[0264] In another possible implementation, based on the same inventive concept as the above-described method embodiments, this embodiment provides a reservoir power station group scheduling rule extraction system and a computer-readable storage medium based on multi-model soft scene fusion. This system can be used to execute the methods provided in the above embodiments, and its implementation principle and technical effects are similar; the same parts will not be repeated here. The following focuses on describing the hardware structure and physical implementation.

[0265] In one possible implementation, it specifically includes:

[0266] Memory, used to store computer programs;

[0267] In this embodiment, the memory serves as the core data and instruction storage carrier of the system, used to store all program code and operational status data required for extracting reservoir group scheduling rules in a non-volatile or volatile manner. Specifically, the memory can be implemented using high-speed random access memory or non-volatile memory, specifically at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0268] Furthermore, the physical space of the memory can be divided into different addressing regions, which are used to persistently store the optimal scheduling sample set built offline, the pre-trained prediction base models specific to each scheduling scenario, the posterior parameters of the scenario estimated iteratively through the expectation-maximization algorithm, and the program instruction sequence used for distance mapping and weighted synthesis in the online stage. Through the above hardware-level partitioning and isolation calculation, the reading efficiency of the frequently called online scheduling algorithm module can be guaranteed, while maintaining the storage reliability of historical hydrological data.

[0269] A processor is used to implement a method for extracting scheduling rules for reservoir power station groups based on multi-model soft scene fusion when executing computer programs.

[0270] Specifically, the processor is the control and calculation center of the entire reservoir power station group scheduling rule extraction system. The processor can be implemented using a general-purpose central processing unit (CPU), or using a digital signal processor (DSP), readily available programmable gate arrays (FPGAs), other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can use a microprocessor to execute the computational logic.

[0271] During actual operation, the processor reads program instructions stored in memory via the system data bus and performs intensive matrix and floating-point operations in its computing units. For example, in the online prediction phase, the processor is responsible for calling internal registers to extract the current observation state and executing division and multiplication instructions to calculate a 14-dimensional multidimensional dimensionless feature vector. Further, the processor performs exponential operations based on a Gaussian kernel function to output soft scene membership degrees and performs multiplication-addition summation on the mean and deviation of each local prediction. In the offline training phase, the processor also needs to allocate independent computing threads to perform population evolution calculations for non-dominated sorting genetic algorithms and hyperparameter search for Bayesian optimization algorithms. By employing hardware chips with high concurrency processing capabilities, the real-time computational requirements for grid-level power plant group scheduling instructions can be met.

[0272] In some optional implementations, the system may also be configured with a communication interface component. This communication interface component is used to establish data links with remote watershed automated scheduling systems and sensor networks. It is responsible for receiving analog data such as inflow and outflow rates and reservoir water levels from power stations at all levels in real time, converting them into standard digital signals, and transmitting them to the processor for feature extraction calculations.

[0273] A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of a method for extracting scheduling rules for reservoir power station groups based on multi-model soft scene fusion are disclosed.

[0274] In this embodiment, a computer-readable storage medium serves as the physical entity carrying the software module, providing the code-fixed form of the technical solution of this invention. The computer-readable storage medium can be implemented using any medium capable of storing program code long-term or temporarily, such as a universal serial bus flash drive, portable hard drive, read-only memory, magnetic disk, or optical disk. The computer program is burned into the storage unit of the physical medium in the form of an instruction set.

[0275] When an external server or industrial control computer terminal loads and reads the medium, the microprocessor embedded in the device can extract the instruction set and interpret and execute all the technical processes described in the foregoing embodiments, such as feature representation, soft scene membership mapping, local prediction base model output correction, full variance decomposition calculation, and expectation maximization parameter estimation, line by line in the memory space. By providing a computer-readable storage medium, not only is the offline distribution and cross-regional deployment process of the underlying algorithm firmware of the water conservancy dispatching system simplified, but it can also serve as an independent software product carrier, providing a physical verification basis for subsequent commercial licensing.

[0276] It should be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.

Claims

1. A method for extracting scheduling rules for reservoir power station groups based on multi-model soft scenario fusion, characterized in that, include: Obtain the historical operation sequence and current operational observation status of the reservoir power station group; Based on historical operation sequences and current operational observation status, an input feature set for the prediction base model is constructed, and physical state features reflecting the coupling relationship between the hydrological situation of a single reservoir and the spatial relationship of the cascade are extracted to obtain a multidimensional dimensionless feature vector. Based on the spatial distance mapping between multidimensional dimensionless feature vectors and pre-constructed cluster centers of various scheduling scenarios, the soft scenario membership degree of the current running observation state relative to each scheduling scenario is determined. The input feature set is input into the pre-trained prediction base model specific to each scheduling scenario, and combined with the pre-estimated scenario posterior parameters, to generate the local prediction mean and local prediction bias corresponding to each scheduling scenario. Based on the soft scene membership degree, the local prediction mean and local prediction deviation corresponding to each scheduling scene are weighted and synthesized to obtain the predicted outflow and prediction variance of the target reservoir.

2. The method according to claim 1, characterized in that, Physical state features reflecting the coupling relationship between the hydrological situation of a single reservoir and the spatial relationship of the cascade system are extracted to obtain a multidimensional dimensionless feature vector, including: Based on historical operation sequences and current operational observation status, individual reservoir hydrological situation characteristic indicators are extracted for each regulating reservoir. Based on the hydrological situation characteristics of a single reservoir, we extract the cascade coupling characteristics between adjacent regulating reservoirs. By combining the hydrological situation characteristic indicators of a single reservoir with the characteristic indicators of cascade coupling, a multidimensional dimensionless feature vector is obtained.

3. The method according to claim 2, characterized in that, The hydrological characteristics of a single reservoir include inflow abundance / difference indicators, water storage utilization rate indicators, and inflow trend indicators; among them, the hydrological characteristics of each regulating reservoir are extracted as follows: For each regulating reservoir, the current inflow rate is extracted from the current operating observation status, and the ratio of the current inflow rate to the pre-configured average inflow rate for the same period over many years is calculated to obtain the water abundance / scarcity index. Extract the current water storage from the current operating observation status, and calculate the ratio of the available water after deducting the pre-configured dead storage capacity from the current water storage to the pre-configured total regulating storage capacity to obtain the water storage utilization rate index; Extract the inflow rate of the previous period from the historical operation sequence, calculate the difference between the current inflow rate and the inflow rate of the previous period, and compare this difference with the average inflow rate over the same period to obtain the trend index of water inflow change.

4. The method according to claim 1, characterized in that, Based on the spatial distance mapping between multidimensional dimensionless feature vectors and pre-constructed cluster centers of various scheduling scenarios, the soft scenario membership degree of the current operational observation state relative to each scheduling scenario is determined, including: Calculate the spatial feature distance between the multidimensional dimensionless feature vector and each pre-built cluster center of the scheduling scenario; By using a distance decay function configured based on preset bandwidth parameters, the distance of each spatial feature is mapped to the corresponding assignment weight coefficient of the scheduling scenario. The proportion of the attribution weight coefficient of each scheduling scenario in the total attribution weight coefficient of all scheduling scenarios is determined as the soft scenario membership degree of the current running observation state relative to the corresponding scheduling scenario.

5. The method according to claim 1, characterized in that, The pre-estimated posterior parameters of the scenario include the weights of the scenario prediction base model, the linear correction intercept of the prediction base model, the linear correction slope of the prediction base model, and the local baseline prediction variance. Combining these pre-estimated posterior parameters, the local prediction mean and local prediction bias for each scheduling scenario are generated, including: For each scheduling scenario, obtain the initial outbound flow prediction value output by each prediction base model for the input feature set; By using the linear correction intercept and linear correction slope of the prediction base model, the corresponding initial outflow flow prediction values ​​are linearly corrected to obtain the corrected prediction values ​​of each prediction base model. The corrected prediction values ​​of each prediction base model are weighted and summed using the weights of the scenario prediction base models to obtain the local prediction mean of the current scheduling scenario. By combining the local baseline prediction variance and the dispersion of the corrected prediction values ​​of each prediction base model relative to the local prediction mean, the local prediction bias of the current scheduling scenario is calculated.

6. The method according to claim 1, characterized in that, Based on the soft scene membership degree, the local prediction mean and local prediction deviation corresponding to each scheduling scenario are weighted and synthesized to obtain the predicted outflow and prediction variance of the target reservoir, including: Using the soft scenario membership degree of each scheduling scenario as a weighting coefficient, the local prediction mean of all scheduling scenarios is weighted and summed to obtain the predicted outbound flow. Based on local prediction bias and membership of each soft scene, calculate the intra-scene variance component that reflects the prediction discrepancy of the model within a single scene. Based on the deviation of each local prediction mean from the final predicted outbound flow, and combined with the membership degree of each soft scenario, the inter-scenario variance component reflecting the uncertainty of cross-scenario affiliation is calculated. The variance components within a scenario are added together with the variance components between scenarios to obtain the predicted variance that characterizes the final scheduling risk boundary.

7. The method according to claim 1, characterized in that, Each pre-trained prediction base model is trained using a pre-constructed optimal scheduling sample set; The optimal scheduling sample set is constructed through a multi-timescale nested optimization mechanism, including: Multi-objective optimization scheduling models at monthly, weekly, and daily scales were constructed respectively. The monthly-scale multi-objective optimization scheduling model is solved, and the monthly target water level of each reservoir is extracted. The target water level at the end of the month is used as an upper boundary constraint and injected into the weekly multi-objective optimization scheduling model for solution. The weekend target water level of each reservoir is then extracted from the output. The target water level at the weekend is used as an upper-level boundary constraint and injected into the daily-scale multi-objective optimization scheduling model for solution. The daily-scale optimal scheduling trajectory sequence that satisfies the constraint is selected and compiled into an optimal scheduling sample set.

8. The method according to claim 7, characterized in that, The specific optimization objectives of each multi-objective scheduling model include: The goals are to maximize the cascade power generation, maximize the water supply guarantee rate of the reservoir group, and maximize the ecological flow guarantee rate of the river section. Specifically, the solution for each multi-objective optimization scheduling model is as follows: A non-dominated sorting genetic algorithm with a preset population size is used to iteratively calculate the model considering the optimization objective, generating a non-dominated optimal solution set; A multi-attribute decision-making model based on the approximation ideal solution ranking method is used to optimize the non-dominated optimal solution set and obtain the optimal scheduling trajectory of each reservoir at the corresponding time scale.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for extracting scheduling rules of reservoir power station groups based on multi-model soft scene fusion as described in any one of claims 1 to 8.

10. A reservoir power station group scheduling rule extraction system based on multi-model soft scenario fusion, characterized in that, include: Memory, used to store computer programs; A processor is configured to implement the steps of the method for extracting scheduling rules for reservoir power station groups based on multi-model soft scene fusion as described in any one of claims 1 to 8 when executing a computer program.