Historical similar runoff classification recommendation method based on feature index identification
By combining hydrological mechanisms and deep learning methods, a historical similar runoff classification recommendation system was constructed, which solved the reservoir scheduling errors and the limitations of traditional similarity measurement, achieved efficient long-term runoff forecasting and real-time flood control scheduling support, and improved the scientificity and reliability of water and drought disaster prevention.
Patent Information
- Application Number
- CN202510788315.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
Smart Images

Figure CN120705598A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a runoff classification recommendation method, specifically a historical similar runoff classification recommendation method based on feature index identification, and belongs to the technical field of hydrological forecasting and artificial intelligence intersection. Background Art
[0002] Generally speaking, hydrological forecasting is the core supporting technology for flood and drought prevention, but it is limited by hydrological forecasting capabilities such as meteorological input errors, simplified model structure, and underlying surface changes. As the forecast period increases, the uncertainty of the hydrological forecast process gradually increases.
[0003] In the prior art, 1) a multi-model combined runoff forecasting method based on improved KNN is disclosed in publication number CN117874635A. By coupling the improved KNN method and multi-model integration technology, the runoff simulation capability of the data-driven forecasting model is enhanced from two aspects: feature data selection and multi-model dynamic adaptive weighting. However, this method is still subject to the accumulation of initial field errors and is difficult to solve the runoff morphological distortion problem with a forecast period of more than 7 days; 2) a medium- and long-term runoff forecasting method based on the similarity of previous runoff is disclosed in publication number CN117540906A. This method only relies on the morphological characteristics of the runoff time series, does not consider the multivariate characteristics of the flood formation mechanism (such as fluctuation rate, spatial distribution), and has low computational efficiency, and cannot meet the real-time flood control scheduling needs.
[0004] The existing technologies mentioned above reveal three key technical bottlenecks: 1) Reservoir scheduling causes measured data from hydrological stations to deviate from natural runoff patterns, and direct use introduces systematic errors; 2) Traditional similarity metrics are single and ignore the multidimensional physical characteristics of the runoff process; and 3) Cluster classification relies on artificial empirical thresholds and lacks the ability to adaptively learn complex nonlinear relationships. Summary of the Invention
[0005] The purpose of the present invention is to provide a historically similar runoff classification recommendation method based on feature index identification in order to solve at least one of the above technical problems, and to achieve accurate recommendation of historically similar runoffs by integrating hydrological mechanisms and deep learning into the full-chain runoff process identification framework.
[0006] The present invention achieves the above-mentioned purpose through the following technical solutions: a historical similar runoff classification recommendation method based on feature index identification, the historical similar runoff classification recommendation method comprising the following steps:
[0007] S1. Historical runoff process restoration: restore the measured runoff data of the control hydrological stations affected by reservoir operation to obtain the natural runoff sequence;
[0008] S2. Extraction of historical flooding processes: Based on the super-quantitative sampling method and the maximum flow classification threshold, historical flooding processes are extracted from a long series of runoff processes;
[0009] S3. Characterization of historical runoff characteristics and construction of an archive database, building a runoff characteristic index system including runoff volume, maximum flow, fluctuation rate, amplitude, time and morphology;
[0010] S4. Cluster analysis of runoff processes: principal component analysis is used to reduce the dimension and extract effective features of many rainfall runoff indicators, and the k-means algorithm is used to cluster the runoff processes;
[0011] S5. Recommendation of historical similar runoff processes based on the Siamese network architecture. The Siamese network architecture is used to calculate the similarity between the target runoff process and the historical runoff process, and the historical runoff category with the highest similarity is output to guide flood and drought disaster prevention.
[0012] As a further solution of the present invention: the water balance method is used to analyze the regulating and storing function of the reservoir, and the measured runoff data of the control hydrological station is restored and calculated, and the calculation period is daily;
[0013] According to the reservoir water level, reservoir storage capacity curve and outflow, the restored flow of the controlling hydrological station is calculated based on the water balance method. The formula is as follows:
[0014]
[0015] Where: The average restored flow rate of the controlling hydrological station to be restored during the period; The average actual flow of the controlling hydrological station to be restored during the period; Respectively represent the t j , t j+1 The storage capacity of the i-th reservoir at the beginning of the time period; n represents the number of reservoirs above the basin's controlling hydrological station.
[0016] As a further solution of the present invention: the extraction of historical flooding processes specifically includes:
[0017] Set the maximum flow rate classification threshold;
[0018] The maximum flow rate is selected using the super-quantitative sampling method;
[0019] The starting and ending times of the flooding process are found according to the maximum flow rate, thereby extracting the flooding process.
[0020] As a further solution of the present invention, cluster analysis of runoff processes includes correlation analysis of runoff characteristic indicators. The correlation analysis of runoff characteristic indicators uses the Pearson correlation coefficient to evaluate the relationship between different characteristics in the runoff process. The calculation method is the ratio of the covariance between two variables A and B and the product of their standard deviations:
[0021]
[0022] Where: A is the m characteristic index of a flood [a1, a2, ..., a m ], B is A T , μ A ,μ B is the mean of the characteristic index, σ A ,σ B is the standard deviation of variables A and B, and ρ(A,B) is the correlation coefficient between the two populations A and B.
[0023] As a further solution of the present invention: cluster analysis of runoff processes also includes principal component analysis of runoff processes, which uses principal component analysis to reduce the dimension and extract effective features of many runoff process indicators, specifically including:
[0024] Assume that there are n flooding processes in the data set X, and each flooding process has p variables, then:
[0025]
[0026] Where: x ij (i=1,2…n;j=1,2…p) represents the jth characteristic index of the i-th flood;
[0027] Normalize the data matrix
[0028]
[0029] in, and s j Represent the mean and standard deviation of the jth variable of all samples respectively; X ij is the jth characteristic index of the i-th flood;
[0030] Calculate the correlation coefficient matrix R for the standardized matrix Z
[0031]
[0032] Solve the eigenvalues of the correlation coefficient matrix R, and the eigenvalues are sorted in descending order as λ1, λ2,…,λ p , and calculate the corresponding eigenvectors b1,b2,…,b p , extract the principal component factors;
[0033] According to the cumulative variance contribution rate The criterion of k is used to determine the value of m (usually 85%), thereby establishing the first K principal components and calculating the score value U of the principal component. ij , get the sample matrix U of the principal component;
[0034]
[0035] As a further solution of the present invention: the cluster analysis of the runoff process specifically includes:
[0036] According to the results of principal component analysis, cluster analysis based on the k-means algorithm was carried out: assuming that there are n flood processes, each flood process has q principal components, then n×q data constitute a flood process principal component matrix, that is:
[0037]
[0038] Where: di j (i=1,2,…n;j=1,2…q) represents the jth principal component of the i-th flood;
[0039] Determine the initial number of clustering classes k, divide all samples into k initial classes, and use the centroid of each class as the initial cluster center. Calculate the distance of each object to each cluster center and assign it to the class represented by the nearest center to complete the initial clustering; recalculate the cluster center of each class and iteratively repeat the process of assigning and updating the center until the iterative termination condition is met. Finally, count the categories to which different samples belong and complete the cluster analysis.
[0040] As a further solution of the present invention: the Siamese network architecture is combined with an improved feature extraction network and a contrast loss function to perform similarity analysis on the runoff data. The feature extraction network is responsible for converting the input runoff data into a feature vector, and the Siamese network architecture includes feature extraction network construction, Siamese network construction, training data generation and model training.
[0041] As a further solution of the present invention: the feature extraction network construction specifically includes:
[0042] The bidirectional LSTM layer uses two layers of bidirectional long short-term memory (Bi-LSTM) networks, with 128 and 64 units respectively. Bi-LSTM can simultaneously consider the forward and reverse dependencies of time series data, effectively capturing the temporal characteristics of runoff data.
[0043] The working mechanism of Bi-LSTM is that Bi-LSTM can more comprehensively capture the dependencies in time series by combining information in both forward and reverse directions. Specifically, Bi-LSTM consists of two LSTM layers:
[0044] Forward LSTM: processes data from the beginning to the end of the time series, capturing information from the past to the present; the calculation process of forward LSTM is the same as that of ordinary LSTM, processing the input sequence x1, x2, ..., x in sequence from time step t = 1 to t = T T Output hidden state
[0045] Reverse LSTM: Processes data from the end to the beginning of the time series, capturing information from the future to the present; the calculation process of reverse LSTM is to process the input sequence x from time step t = T to t = 1 T ,x T-1 ,…,x1 output hidden state
[0046] Combining forward and backward information
[0047] The final output of the Bi-LSTM is the concatenation of the hidden states of the forward LSTM and the backward LSTM:
[0048]
[0049] The output at each time step It includes information from the past to the present and from the future to the present, thus being able to more comprehensively capture the bidirectional dependencies in the time series;
[0050] Fully connected layer and activation function
[0051] After the Bi-LSTM layer, two fully connected layers are connected. The number of units in the first layer is set to 256, and the number of units in the second layer is set to 128. Both use the ReLU activation function. A Dropout layer is added after the first fully connected layer to prevent model overfitting. In addition, L2 regularization is added to the fully connected layer to further constrain the complexity of the model and improve the generalization ability of the model.
[0052] As a further solution of the present invention: Siamese network construction specifically includes:
[0053] Network Architecture: Based on the feature extraction network, a Siamese network is constructed. The Siamese network contains two feature extraction network branches with shared weights, which extract features from the two input runoff process data. Then, a custom cosine similarity calculation layer is used to calculate the similarity between the feature vectors of the two runoff processes. The closer the cosine similarity value is to 1, the more similar the two runoff processes are. The specific calculation formula is as follows:
[0054]
[0055] Contrastive loss function: By minimizing the similarity value of positive sample pairs and maximizing the similarity value of negative sample pairs, the model can effectively learn the similarity characteristics of the runoff process;
[0056] The goal of the contrast loss function is to make the feature vector distance of similar sample pairs as small as possible. The specific formula is as follows:
[0057] Loss = y true square_pred+(1-y true )·margin_square
[0058] Where: y true is the label of the sample pair, square_pred is the square of the predicted distance, and margin_square is the square of the maximum distance.
[0059] As a further solution of the present invention, training data generation and model training specifically include:
[0060] Training data generation: Positive and negative sample pairs are generated from historical runoff data through random sampling. A positive sample pair consists of two similar runoff processes, while a negative sample pair consists of two dissimilar runoff processes.
[0061] Model training: The Siamese network was trained using the Adam optimizer in conjunction with an exponentially decaying learning rate scheduler. During training, an early stopping mechanism was implemented. When the loss value stopped decreasing after five consecutive training cycles, training was stopped and the optimal model weights were restored. Finally, after 30 training cycles, the model converged and completed learning the similarity features of the runoff data.
[0062] The beneficial effects of the present invention are:
[0063] 1) The present invention significantly improves the reliability of long-term runoff forecasts by integrating hydrological mechanism models with deep learning technology. First, the runoff restoration technology based on the water balance method eliminates the interference of reservoir scheduling, reconstructs the natural runoff sequence, and provides real benchmark data for similarity analysis; second, an innovative six-dimensional feature index system covering runoff volume, maximum flow, fluctuation rate, amplitude, time and morphology is constructed, breaking through the limitations of traditional single morphological similarity measurement and comprehensively characterizing the flood formation mechanism; further, the combination of principal component analysis and k-means clustering realizes high-dimensional feature dimensionality reduction and objective classification, avoiding the subjectivity of artificial experience thresholds, and generating a runoff pattern category library with clear physical meaning; finally, based on the bidirectional LSTM feature extraction structure of the Siamese network, combined with the improved contrast loss function, the spatiotemporal evolution law of the runoff process is deeply explored, significantly improving the accuracy of similarity matching.
[0064] 2) This method shortens the matching efficiency of historical similar runoff from hours to minutes, supports real-time decision-making in flood control scheduling, and effectively reduces the risk of uncertainty transmission by recommending complete historical runoff processes instead of short-term forecasts. It provides high-confidence data support for joint scheduling of reservoir groups and disaster emergency response, and comprehensively improves the initiative and scientific nature of flood and drought disaster prevention. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a schematic diagram of the process of the present invention;
[0066] Figure 2 This is a schematic diagram of the annual runoff series and its sliding mean process at Xiangjiaba Station from 1960 to 2023;
[0067] Figure 3 This is a schematic diagram of the annual runoff series and its sliding average process from 1960 to 2023 at Gaochang Station;
[0068] Figure 4 This is a schematic diagram of the annual runoff series and its moving average process from 1960 to 2023 at Beibei Station;
[0069] Figure 5 This is a schematic diagram of the annual runoff series and its sliding mean process at Wulong Station from 1960 to 2023;
[0070] Figure 6 This is a schematic diagram of the results of the runoff process selection of four stations in the present invention;
[0071] Figure 7 This is an indication of the flood characteristic index of Xiangjiaba Station in the present invention;
[0072] Figure 8 This is a schematic representation of the flood characteristic index of the high-field station of the present invention;
[0073] Figure 9This is a schematic representation of the flood characteristic index of Beibei Station in the present invention;
[0074] Figure 10 This is a schematic representation of the flood characteristic index of Wulong Station in the present invention;
[0075] Figure 11 This is a schematic diagram of the clustering results of Xiangjiaba Station in the present invention;
[0076] Figure 12 This is a schematic diagram of the high-field station clustering results of the present invention;
[0077] Figure 13 This is a schematic diagram of the clustering results of Beibei Station in the present invention;
[0078] Figure 14 This is a schematic diagram of the clustering results of Wulong Station in the present invention;
[0079] Figure 15 This is a schematic diagram of the runoff process fitting of the present invention. DETAILED DESCRIPTION
[0080] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0081] Example 1, as Figure 1 As shown, a historical similar runoff classification recommendation method based on feature index identification includes the following steps:
[0082] S1. Historical runoff process restoration: restore the measured runoff data of the control hydrological stations affected by reservoir operation to obtain the natural runoff sequence;
[0083] S2. Extraction of historical flooding processes: Based on the super-quantitative sampling method and the maximum flow classification threshold, historical flooding processes are extracted from a long series of runoff processes;
[0084] S3. Characterization of historical runoff characteristics and construction of an archive database, building a runoff characteristic index system including runoff volume, maximum flow, fluctuation rate, amplitude, time and morphology;
[0085] S4. Cluster analysis of runoff processes: principal component analysis is used to reduce the dimension and extract effective features of many rainfall runoff indicators, and the k-means algorithm is used to cluster the runoff processes;
[0086] S5. Recommendation of historical similar runoff processes based on the Siamese network architecture. The Siamese network architecture is used to calculate the similarity between the target runoff process and the historical runoff process, and the historical runoff category with the highest similarity is output to guide flood and drought disaster prevention.
[0087] Example 2. In addition to all the technical features of Example 1, this example also includes: if a group of control reservoirs is built on the basin, affected by the reservoir scheduling, the flood reporting data of the main control hydrological stations in the basin cannot reflect the basin runoff morphology under the natural state, and the hydrological flood reporting data sequence affected by the reservoir scheduling needs to be restored. This study is based on the daily runoff series data of the control hydrological stations in the main and tributary rivers of the upper Yangtze River. According to the water level and storage capacity curves and scheduling operation data of the large reservoirs with strong regulation performance built in the upper Yangtze River, the water balance method is used to analyze the storage function of the reservoir, and the measured runoff data of the control hydrological station is restored and calculated. The calculation period is daily; according to the reservoir water level, reservoir storage capacity curve and outflow flow, the restored flow of the control hydrological station is calculated based on the water balance method. The formula is as follows:
[0088]
[0089] Where: The average restored flow rate of the controlling hydrological station to be restored during the period; The average actual flow of the controlling hydrological station to be restored during the period; Respectively represent the t j , t j+1 The storage capacity of the i-th reservoir at the beginning of the time period; n represents the number of reservoirs above the basin's controlling hydrological station.
[0090] like Figures 2 to 5 As shown, taking Xiangjiaba Station (Jinsha River) in the upper reaches of the Yangtze River as an example:
[0091] Input data: daily runoff series from 1960 to 2023, reservoir water level and storage capacity curves, and operation records (such as Xiluodu and Xiangjiaba reservoirs).
[0092] Reduction calculation: Use the water balance formula to reverse calculate the natural runoff on a daily basis.
[0093] The extraction of historical flooding processes specifically includes: setting a classification threshold for the maximum flow rate; using an over-quantitative sampling method to select the maximum flow rate; and finding the start and end times of the flooding process based on the maximum flow rate, thereby extracting the flooding process.
[0094] like Figure 6 ( Figure 6-1 For Xiangjiaba Station, Figure 6-2 For Gaochang Station, Figure 6-3 For Beipei Station and Figure 6-4(shown as Wulong Station), from June to September 1960 to 2023, the runoff process of the control hydrological stations at the outlets of the Jinsha River, Minjiang River, Jialing River, and Wujiang River in the upper reaches of the Yangtze River was restored and extracted. Among them, there were 220 sessions at the Xiangjiaba Station of the Jinsha River, 315 sessions at the Gaochang Station of the Minjiang River, 401 sessions at the Beibei Station of the Jialing River, and 333 sessions at the Wulong Station of the Wujiang River. Taking the Beibei Station of the Jialing River as an example:
[0095] Parameter setting: The maximum flow rate classification threshold is set to the warning flow rate (historical flood frequency 10%).
[0096] Process Extraction:
[0097] The peak points exceeding the threshold from June to September from 1960 to 2023 were screened by using the super-quantitative sampling method (a total of 401 extreme points were identified);
[0098] With the peak point as the center, search for the inflection point of the flow rate change rate on both sides (for example, when the flow rate change rate is ≤ 5%, it is determined as the starting and ending points).
[0099] like Figures 7 to 10 As shown in the figure, the historical runoff characteristics are characterized by constructing a six-dimensional characteristic index system:
[0100] Runoff: total amount of flood events (e.g., the total amount of flood at Beibei Station in 1981);
[0101] Maximum flow: peak flow (e.g., peak flow at Wulong Station in 2020);
[0102] Rise and fall rate: the change in flow per unit time (e.g., the rise and fall rate of the 1998 flood at Gaochang Station);
[0103] Amplitude: the difference between the start and end flow (e.g., the magnitude of the 2010 flood at Xiangjiaba Station);
[0104] Time: duration of the flood (e.g., average duration 7.2 days);
[0105] Morphology: peak time and symmetry coefficient (such as early peak type, double peak type).
[0106] In addition to all the technical features of the first embodiment, this embodiment also includes:
[0107] Cluster analysis of runoff processes includes correlation analysis of runoff characteristic indicators. The Pearson correlation coefficient is used to evaluate the relationship between different characteristics of the runoff process. The calculation method is the ratio of the covariance between two variables A and B and the product of their standard deviations:
[0108]
[0109] Where: A is the 27 characteristic indicators of a flood [a1, a2, ..., am ], B is A T , μ A ,μ B is the mean of the characteristic index, which is equal in value, σ A ,σ B is the standard deviation of variables A and B, and ρ(A,B) is the correlation coefficient between the two populations A and B.
[0110] Cluster analysis of runoff processes also includes principal component analysis of runoff processes. The principal component analysis method is used to reduce the dimension and extract effective features of many runoff process indicators, including:
[0111] Assume that there are n flooding processes in the data set X, and each flooding process has p variables, then:
[0112]
[0113] Where: x ij (i=1,2…n;j=1,2…p) represents the jth characteristic index of the i-th flood;
[0114] Normalize the data matrix
[0115]
[0116] in, and s j Represent the mean and standard deviation of the jth variable of all samples respectively; X ij is the jth characteristic index of the i-th flood;
[0117] Calculate the correlation coefficient matrix R for the standardized matrix Z
[0118]
[0119] Solve the eigenvalues of the correlation coefficient matrix R, and the eigenvalues are sorted in descending order as λ1, λ2,…,λ p , and calculate the corresponding eigenvectors b1,b2,…,b p , extract the principal component factors;
[0120] According to the cumulative variance contribution rate The criterion of k is used to determine the value of m (usually 85%), thereby establishing the first K principal components and calculating the score value U of the principal component. ij , get the sample matrix U of the principal component;
[0121]
[0122] Cluster analysis of runoff processes specifically includes:
[0123] According to the results of principal component analysis, cluster analysis based on the k-means algorithm was carried out: assuming that there are n flood processes, each flood process has q principal components, and these n×q data constitute a principal component matrix of the flood process, that is:
[0124]
[0125] Where: di j (i=1,2,…n;j=1,2…q) represents the jth principal component of the i-th flood.
[0126] Determine the initial number of clustering classes k, divide all samples into k initial classes, and use the centroid (mean) of each class as the initial cluster center. Calculate the distance of each object to each cluster center and assign it to the class represented by the nearest center to complete the initial clustering; recalculate the cluster center (mean) of each class, and iteratively repeat the process of assigning and updating the center until the iterative termination condition is met. Finally, count the categories to which different samples belong and complete the cluster analysis.
[0127] like Figures 11 to 14 As shown in the figure, the cluster analysis of historical runoff events takes Beibei Station as an example:
[0128] Dimensionality reduction: PCA analysis was performed on the 20 characteristic indicators of 401 floods, and four principal components were retained when the cumulative contribution rate was 85% (explained variance share: PC1=52%, PC2=18%, PC3=9%, PC4=6%).
[0129] k-means clustering:
[0130] Set the initial number of clusters and iterate until the center point movement distance is less than 0.01, forming four typical flood patterns (red / blue / green / purple clusters in the figure):
[0131] Category 1: short-duration steep rise and fall (accounting for 28%);
[0132] Category 2: long duration multi-peak type (35%);
[0133] Category 3: Early peak and slow decline (22%)
[0134] Category 4: Late peak type (accounting for 15%).
[0135] In addition to all the technical features of the first embodiment, this embodiment also includes:
[0136] The Siamese network architecture combines an improved feature extraction network and contrast loss function to perform similarity analysis on runoff data. The feature extraction network is responsible for converting the input runoff data into feature vectors, and the Siamese network architecture includes feature extraction network construction, Siamese network construction, training data generation and model training.
[0137] The feature extraction network construction specifically includes:
[0138] The bidirectional LSTM layer uses two layers of bidirectional long short-term memory (Bi-LSTM) networks, with 128 and 64 units respectively. Bi-LSTM can simultaneously consider the forward and reverse dependencies of time series data, effectively capturing the temporal characteristics of runoff data.
[0139] The working mechanism of Bi-LSTM is that Bi-LSTM can more comprehensively capture the dependencies in time series by combining information in both forward and reverse directions. Specifically, Bi-LSTM consists of two LSTM layers:
[0140] Forward LSTM: processes data from the beginning to the end of the time series, capturing information from the past to the present. The calculation process of forward LSTM is the same as that of ordinary LSTM, processing the input sequence x1, x2, ..., x in sequence from time step t = 1 to t = T. T Output hidden state
[0141] Backward LSTM: processes data from the end to the beginning of the time series, capturing information from the future to the present; the calculation process of the backward LSTM is to process the input sequence x from time step t = T to t = 1. T ,x T-1 ,…,x1 output hidden state
[0142] Combining forward and backward information
[0143] The final output of the Bi-LSTM is the concatenation of the hidden states of the forward LSTM and the backward LSTM:
[0144]
[0145] The output at each time step It includes information from the past to the present (forward LSTM) and information from the future to the present (reverse LSTM), thus being able to more comprehensively capture the bidirectional dependencies in the time series.
[0146] Fully connected layer and activation function
[0147] After the Bi-LSTM layer, two fully connected layers are connected. The number of units in the first layer is set to 256, and the number of units in the second layer is set to 128. Both use the ReLU activation function. A Dropout layer is added after the first fully connected layer to prevent model overfitting. In addition, L2 regularization is added to the fully connected layer to further constrain the complexity of the model and improve the generalization ability of the model.
[0148] The Siamese network construction specifically includes:
[0149] Network Architecture: Based on the feature extraction network, a Siamese network is constructed. The Siamese network contains two feature extraction network branches with shared weights, which extract features from the two input runoff process data. Then, a custom cosine similarity calculation layer is used to calculate the similarity between the feature vectors of the two runoff processes. The closer the cosine similarity value is to 1, the more similar the two runoff processes are. The specific calculation formula is as follows:
[0150]
[0151] Contrastive loss function: To train the Siamese network, an improved contrastive loss function is used. This loss function minimizes the similarity between positive sample pairs (similar runoff processes) and maximizes the similarity between negative sample pairs (dissimilar runoff processes), enabling the model to effectively learn the similarity characteristics of runoff processes.
[0152] The goal of the contrast loss function is to make the distance between the feature vectors of similar sample pairs as small as possible, and the distance between the feature vectors of dissimilar sample pairs as large as possible. The specific formula is as follows:
[0153] Loss = y true square_pred+(1-y true )·margin_square
[0154] Among them, y true is the label of the sample pair (1 for similarity, 0 for dissimilarity), square_pred is the square of the predicted distance, and margin_square is the square of the maximum distance. In this way, the network can learn the similarity characteristics of the runoff process.
[0155] like Figure 15 As shown, the runoff forecast process of Beibei Station on Jialing River in 2024 is taken as an example (forecast period is 10 days).
[0156] Feature extraction:
[0157] Bidirectional LSTM layer: The 128-unit layer captures long-range dependencies (such as the trend 72 hours before the flood peak), and the 64-unit layer extracts local fluctuation features; fully connected layer: The 256-unit layer integrates spatiotemporal features, and the 128-unit layer outputs feature vectors.
[0158] Similarity matching:
[0159] Calculate the cosine similarity between the target runoff and the historical runoff feature vector (for example, the similarity with the 1981 flood is 0.96).
[0160] Output result: The recommended historical similar runoff is the 1981 flood process in Class 2 (long duration multi-peak type).
[0161] Decision support: Generate disaster prevention plans based on the subsequent evolution data of the 1981 flood.
[0162] Example 5: In addition to all the technical features of Example 1, this example also includes:
[0163] Training data generation and model training specifically include:
[0164] Training data generation: Positive and negative sample pairs are generated from historical runoff data through random sampling. A positive sample pair consists of two similar runoff processes, while a negative sample pair consists of two dissimilar runoff processes.
[0165] Model training: The Siamese network was trained using the Adam optimizer in conjunction with an exponentially decaying learning rate scheduler. During training, an early stopping mechanism was implemented. When the loss value stopped decreasing after five consecutive training cycles, training was stopped and the optimal model weights were restored. Finally, after 30 training cycles, the model converged and completed learning the similarity features of the runoff data.
[0166] First, the measured runoff at the control hydrological station was restored and calculated: according to the reservoir water level, storage capacity curve and outflow, the water balance formula was used to reversely infer the natural runoff to eliminate the impact of reservoir storage; secondly, the historical flood process was extracted based on the super-quantitative sampling method, the flow classification threshold was set to locate the extreme point, and the inflection point was searched on both sides of the time axis to determine the start and end time; then, an archive library containing six types of characteristic indicators such as runoff, maximum flow, and fluctuation rate was constructed, and key indicators were selected through Pearson correlation coefficient analysis. Principal component analysis was used to reduce the dimensionality of high-dimensional features, retaining the principal components with a cumulative contribution rate of more than 85%, and eliminating redundant information; based on the principal component matrix, the k-means clustering algorithm was applied to iteratively optimize the classification center, and the historical runoff was divided into several categories according to its morphological characteristics. The Siamese network is used in the core recommendation stage: a dual-branch shared-weight feature extraction network processes the target runoff and historical runoff data respectively. Each branch contains two layers of bidirectional LSTM (128 / 64 units) to capture the forward and reverse dependencies of the temporal sequence. The fully connected layer (256 / 128 units) combines the ReLU activation function and the Dropout mechanism to extract deep features. The similarity score of the two path flow feature vectors is calculated through the cosine similarity layer, and the improved contrast loss function (maximizing the distance between similar sample pairs and minimizing the distance between dissimilar sample pairs) is used to train the network. Finally, the historical runoff category and complete process most similar to the target runoff are output to guide disaster prevention decisions.
[0167] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
[0168] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A historical similar runoff classification recommendation method based on feature index identification, characterized by: The historical similar runoff classification recommendation method comprises the following steps: S1. Historical runoff process restoration: restore the measured runoff data of the control hydrological stations affected by reservoir operation to obtain the natural runoff sequence; S2. Extraction of historical flooding processes: Based on the super-quantitative sampling method and the maximum flow classification threshold, historical flooding processes are extracted from a long series of runoff processes; S3. Characterization of historical runoff characteristics and construction of an archive database, building a runoff characteristic index system including runoff volume, maximum flow, fluctuation rate, amplitude, time and morphology; S4. Cluster analysis of runoff processes: principal component analysis is used to reduce the dimension of rainfall runoff indicators and extract effective features, and the k-means algorithm is used to cluster the runoff processes; S5. Recommendation of historical similar runoff processes based on the Siamese network architecture. The Siamese network architecture is used to calculate the similarity between the target runoff process and the historical runoff process, and the historical runoff category with the highest similarity is output to guide flood and drought disaster prevention.
2. The historically similar runoff classification recommendation method according to claim 1 is characterized by: In S1, the water balance method is used to analyze the regulation and storage function of the reservoir, and the measured runoff data of the control hydrological station is restored and calculated, and the calculation period is daily; According to the reservoir water level, reservoir storage capacity curve and outflow, the restored flow of the controlling hydrological station is calculated based on the water balance method. The formula is as follows: Where: The average restored flow rate of the controlling hydrological station to be restored during the period; The average actual flow of the controlling hydrological station to be restored during the period; Respectively represent the t j , t j+1 The storage capacity of the i-th reservoir at the beginning of the time period; n represents the number of reservoirs above the basin's controlling hydrological station.
3. The historically similar runoff classification recommendation method according to claim 1 is characterized by: In S2, the extraction of historical flooding processes specifically includes: S21, setting the maximum flow rate classification threshold; S22, using the super-quantitative sampling method to select the maximum flow rate; S23. Find the start and end time of the flooding process according to the maximum flow rate, thereby extracting the flooding process.
4. The historically similar runoff classification recommendation method according to claim 1 is characterized by: In S4, the cluster analysis of the runoff process includes a correlation analysis of the runoff characteristic indicators. The runoff characteristic indicator correlation analysis uses the Pearson correlation coefficient to evaluate the relationship between different characteristics in the runoff process. The calculation method is the ratio of the covariance between two variables A and B and the product of their standard deviations: Where: A is the m characteristic index of a flood [a1, a2, ..., a m ], B is A T , μ A ,μ B is the mean of the characteristic index, σ A ,σ B is the standard deviation of variables A and B, and ρ(A,B) is the correlation coefficient between the two populations A and B.
5. The method for recommending historically similar runoff classification according to claim 4, characterized in that: In S4, the cluster analysis of the runoff process also includes principal component analysis of the runoff process, which uses the principal component analysis method to reduce the dimension of the runoff process indicators and extract effective features, specifically including: Assume that there are n flooding processes in the data set X, and each flooding process has p variables, then: Where: x ij (i=1,2…n;j=1,2…p) represents the jth characteristic index of the i-th flood; Normalize the data matrix: in, and s j Represent the mean and standard deviation of the jth variable of all samples respectively; X ij is the jth characteristic index of the i-th flood; To find the correlation coefficient matrix R for the standardized matrix Z, we have: Solve the eigenvalues of the correlation coefficient matrix R, and the eigenvalues are sorted in descending order as λ1, λ2,…,λ p , and calculate the corresponding eigenvectors b1,b2,…,b p , extract the principal component factors; According to the cumulative variance contribution rate The criterion of k is used to determine the first K principal components and calculate the score value U of the principal component. ij , get the sample matrix U of the principal component; 6. The method for recommending historically similar runoff classification according to claim 5, characterized in that: In S4, the cluster analysis of the runoff process specifically includes: According to the results of principal component analysis, cluster analysis based on the k-means algorithm was carried out: assuming that there are n flood processes, each flood process has q principal components, then n×q data constitute a flood process principal component matrix, that is: Where: d ij (i=1,2,…n;j=1,2…q) represents the jth principal component of the i-th flood; Determine the initial number of clustering classes k, divide all samples into k initial classes, and use the centroid of each class as the initial cluster center. Calculate the distance of each object to each cluster center and assign it to the class represented by the nearest center to complete the initial clustering; recalculate the cluster center of each class and iteratively repeat the process of assigning and updating the center until the iterative termination condition is met. Finally, count the categories to which different samples belong and complete the cluster analysis.
7. The historically similar runoff classification recommendation method according to claim 1 is characterized by: In S5, the Siamese network architecture combines an improved feature extraction network and a contrast loss function to perform similarity analysis on the runoff data. The feature extraction network is responsible for converting the input runoff data into a feature vector, and the Siamese network architecture includes feature extraction network construction, Siamese network construction, training data generation and model training.
8. The method for recommending historically similar runoff classification according to claim 7, characterized in that: The feature extraction network construction specifically includes: using a bidirectional LSTM layer; Input layer: Input the runoff number and sample size, time step, and feature dimension to construct the data shape; Embedding layer: Use the Bi-LSTM layer to process the time series. The Bi-LSTM layer contains two LSTM layers with 128 and 64 units respectively. Forward LSTM: Processes the input sequence x1, x2, ..., x in sequence from time step t = 1 to t = T T Output hidden state Reverse LSTM: Process the input sequence x sequentially from time step t=T to t=1 T ,x T-1 ,…,x1 output hidden state Combining the forward and backward information, the final output is the concatenation of the hidden states of the forward LSTM and the backward LSTM: Feature transformation through two fully connected layers After the Bi-LSTM layer, two fully connected layers are connected. The number of units in the first layer is set to 256, and the number of units in the second layer is set to 128. Both use the ReLU activation function and add L2 regularization. A Dropout layer is added after the first fully connected layer to prevent overfitting of the model. Output feature vector, which is used to represent the characteristics of the input runoff time series.
9. The method for recommending historically similar runoff classification according to claim 7, characterized in that: The Siamese network construction specifically includes: Network Architecture: Based on the feature extraction network, a Siamese network is constructed. The Siamese network contains two feature extraction network branches with shared weights, which extract features from the two input runoff process data. Then, a custom cosine similarity calculation layer is used to calculate the similarity between the feature vectors of the two runoff processes. The closer the cosine similarity value is to 1, the more similar the two runoff processes are. The specific calculation formula is as follows: Contrastive loss function: By minimizing the similarity value of positive sample pairs and maximizing the similarity value of negative sample pairs, the model can effectively learn the similarity characteristics of the runoff process; The goal of the contrastive loss function is to make the feature vector distance between similar sample pairs as small as possible: Loss=y true ·square_pred+(1-y true )·margin_square Among them, y true is the label of the sample pair, square_pred is the square of the predicted distance, and margin_square is the square of the maximum distance.
10. The method for recommending historically similar runoff classification according to claim 7, characterized in that: The training data generation and model training specifically include: Training data generation: Positive and negative sample pairs are generated from historical runoff data through random sampling. A positive sample pair consists of two similar runoff processes, while a negative sample pair consists of two dissimilar runoff processes. Model training: The Siamese network was trained using the Adam optimizer in conjunction with an exponentially decaying learning rate scheduler. During training, an early stopping mechanism was implemented. When the loss value stopped decreasing after five consecutive training cycles, training was stopped and the optimal model weights were restored. Finally, after 30 training cycles, the model converged and completed learning the similarity features of the runoff data.
Citation Information
Patent Citations
Medium-and-long-term runoff forecasting method based on early-stage runoff similarity
CN117540906A
Multi-model combination runoff forecasting method based on improved KNN
CN117874635A