A runoff prediction method based on static physical feature enhancement and regional model

CN118673089BActive Publication Date: 2026-09-29HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410816137.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-09-29
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

因此,对于特定流域的洪水预报模型可能无法准确预测极端洪水事件

Benefits of technology

[0044]1、对于现有预测方法中没有考虑静态物理特征对流量的影响可能存在实时变化,本发明采用SFU模块对每个时刻的输入提取不同的静态物理特征,增强模型可解释性的同时,提升预测精度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118673089B_ABST
    Figure CN118673089B_ABST
Patent Text Reader

Abstract

The application discloses a runoff prediction method based on static physical feature enhancement and a regional model, collects static physical features of all small and medium-sized river basins in a region, and monitors time sequence data and climate data of various measuring stations in a typical historical flood process; and the data is processed; a flood prediction model SE-LSTM based on static physical feature enhancement is constructed, including an encoder and a decoder, wherein the encoder Encoder adopts Bi-LSTM as a skeleton, learns the influence of past monitoring values and future prediction values on current time flow; a static feature fusion module SFU in the encoder amplifies the static features of the river basin with the greatest influence on the flow at each time; the decoder Decoder adopts LSTM as a skeleton to process the input, and uses the learned mode in the encoder to perform prediction. The application can effectively utilize the static features of various small river basins, learn the runoff mode in various river basins, so that better prediction accuracy is achieved when facing extreme events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of hydrological forecasting technology, and specifically discloses a runoff prediction method based on static physical feature enhancement and regional model. Background Technology

[0002] In recent years, with global warming, the frequency and intensity of extreme weather events have been increasing. This poses new challenges to flood forecasting research.

[0003] Numerous and complex factors influence flood formation. Traditional forecasting methods using hydrological models struggle to accurately describe the hydrological processes of a watershed and rely heavily on the experience of hydrological experts. With the development of computer technology, deep learning models have been proposed for flood forecasting. Compared to traditional methods, deep learning models are better able to uncover nonlinear relationships between data, which is beneficial for modeling the relationships between flow data and numerous other factors. Currently, the most commonly used method for flood forecasting models is LSTM (Long Short-Term Memory Network), a type of RNN (Recurrent Neural Network) that effectively processes time series data. LSTM internally uses three gate controllers (input gate, forget gate, and output gate) and a cell state to effectively control and regulate the flow of information. However, LSTM can only reason about unidirectional data flows, thus neglecting the impact of future trends in predictable data (such as precipitation) on the current moment's forecast.

[0004] In current deep learning-based flood forecasting models, a mainstream approach is to learn the changing relationships between multiple factors and flow data through feature fusion. This method effectively allows deep learning models to learn the relationship between input features and flow changes, thus making flood forecasts more accurate with the aid of these features. However, most current feature fusion methods focus only on the simple addition of static physical features and time-series data, assuming that static physical features remain completely unchanged throughout the flood process. These methods fail to consider that the impact of different static physical features on flow may vary at different times. The static attributes of a watershed, such as watershed area, topographic slope, soil type, and vegetation cover, are all important components of the watershed's water cycle. The combined effect of these factors, to a certain extent, determines the probability and intensity of floods.

[0005] Existing flood forecasting methods for small and medium-sized river basins mostly focus on specific basins. These methods have good forecasting accuracy for ordinary flood events within the basin, but they are less effective for extreme flood events. A major reason is that extreme flood events occur infrequently, and historical data collected within the basin may not include relevant events. Therefore, flood forecasting models for specific basins may not accurately predict extreme flood events. Accordingly, flood forecasting models can be trained using a regional model approach. A regional model refers to training a flood forecasting model using all river basins within a region. In this way, the model learns flood patterns of extreme flood events from historical data of other similar river basins, and uses this data to guide the forecasting of future extreme flood events. Summary of the Invention

[0006] Purpose of the invention: This invention proposes a runoff prediction method based on enhanced static physical features and regional models, overcoming the impact of existing technologies on flood forecast accuracy due to insufficient utilization of future predictable data or neglect of static watershed data.

[0007] Technical solution: The runoff prediction method based on static physical feature enhancement and regional model described in this invention includes the following steps:

[0008] (1) Collect the static physical characteristics of all small and medium-sized watersheds in the region, as well as the time series data and climate data of each station monitored during typical historical floods;

[0009] (2) The time series data were normalized by standard deviation normalization and the Mahalanobis distance was used to calculate similar watersheds based on time series data, climate data and static physical characteristics, according to a certain threshold.

[0010] (3) Process the static physical features in the selected watershed and embed the static physical features of different dimensions into the same dimension space as the time series data;

[0011] (4) Construct a flood forecasting model SE-LSTM based on static physical feature enhancement, which includes an encoder and a decoder. The encoder uses Bi-LSTM as the skeleton to learn the impact of past monitoring values ​​and future forecast values ​​on the current flow. The static feature fusion module SFU in the encoder amplifies the watershed static features that have the greatest impact on the flow at each time. The decoder uses LSTM as the skeleton to process the input and uses the patterns learned in the encoder to make predictions.

[0012] (5) Divide the data processed by steps (2) and (3) into training set, validation set and test set according to a certain ratio; train the model with training set and validation set and optimize the internal weight parameters of flood forecast model SE-LSTM;

[0013] (6) Test the trained flood forecasting model SE-LSTM with the test set, and evaluate the performance of the flood forecasting model SE-LSTM by comparing the predicted flow value and the actual flow value.

[0014] Furthermore, the static physical characteristics described in step (1) include mean evapotranspiration, mean air temperature, mean surface air temperature, mean elevation, watershed area, slope, and underlying surface type.

[0015] Furthermore, the time-series data mentioned in step (1) includes rainfall, river flow, evapotranspiration, temperature, humidity, and wind speed data.

[0016] Furthermore, the implementation process of step (2) is as follows:

[0017] (21) For dynamic time-series data sequences X = {x1, x2, ... x...} n First, the standard deviation σ and the mean μ are calculated. Then, the following formula is used for preprocessing to obtain the processed result X. * :

[0018]

[0019] The normalized result is segmented according to a preset length of the view window to obtain the input data:

[0020]

[0021] (22) Use the peak finding algorithm to count the peak values ​​of time series data, use the calculated results as features, and integrate all the above features;

[0022] (23) The similarity is calculated using Mahalanobis distance for all features. The Mahalanobis distance formula is:

[0023]

[0024] Where x represents the matrix of each attribute, μ represents the mean matrix of each attribute, S represents the covariance matrix, and W is the weight matrix because static features and time-series features contribute differently to the prediction.

[0025] (24) The calculated Mahalanobis distance is filtered according to the preset threshold. The collected small watershed results are all within the preset value, and similar watersheds are filtered out.

[0026] Furthermore, the implementation process of step (3) is as follows:

[0027] Calculate the common start and end times of all watershed data; segment the watershed data starting from the common start point, set the lookback window length to 120 time steps, and the prediction length to 1 time step; embed each static physical feature and input length into a 120-dimensional vector space.

[0028] Furthermore, the working process of the traffic prediction model SE-LSTM based on static feature enhancement described in step (4) is as follows:

[0029] The Encoder part of the model receives dynamic time-series data and static data. At each time step, the input sequence data is first used with the static data to calculate the weighted input using SFU. Then, the input is fed into Bi-LSTM, and the intermediate states generated during this process are stored in the memory of LSTM. The forward LSTM learns the impact of past rainfall values ​​on the current situation, and the backward LSTM learns the potential impact of future forecast values ​​on the current flow. After all the data in the Encoder has been learned, the memory of the forward and backward LSTMs is merged as the initial input of the Decoder.

[0030] The Decoder part of the model uses the memory and static feature weights learned in the Encoder to set the initial state; the Decoder receives all time-series data except traffic, and after fusing it with the static feature weights, it serves as the input to the model, and outputs the predicted value using the knowledge learned in the memory.

[0031] Furthermore, the workflow of the Static Feature Fusion (SFU) module described in step (4) is as follows:

[0032] The embedded static features are first evaluated using the Sigmoid function to calculate the weight of the current static physical features on the time series data. Then, the time series data is multiplied by this weight. Finally, a residual structure is used to add the data fused with the static features to the original data. The input variable X' is... t and static features The formula for selecting weights using a gating mechanism is as follows:

[0033]

[0034]

[0035] in, Represents the embedding result of the i-th static feature, res i This represents the result of fusing the i-th static feature with the time-series data;

[0036] For the embedded result set res1, res2, ..., res nAggregation is performed along the channel direction to obtain a set of fused results of all static features and input data in their respective channels:

[0037] res=Concat(rea1,res2,…,res n )

[0038] The aggregated data [sequence_len, n] is passed through an attention layer to extract the influence of different static features on the time series data, resulting in a sequence of shape [sequence_len, 1]. Finally, a residual concatenation is performed with the original time series sequence to obtain the output.

[0039] output' = Attention(res)

[0040] output = output' + X' t

[0041] The Sigmoid activation function is used to extract nonlinear relationships in the output data. Then, the data is processed through Dropout layers and fully connected layers to obtain time series data with static feature enhancement. This data is then fed into the Bi-LSTM model for further analysis and processing.

[0042] Furthermore, the preset threshold mentioned in step (24) is the chi-square distribution critical value of 22.36, which is calculated with 13 degrees of freedom and a probability value of 0.05.

[0043] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] 1. Existing prediction methods do not consider the impact of static physical features on traffic flow, which may change in real time. This invention uses the SFU module to extract different static physical features from the input at each time step, thereby enhancing the interpretability of the model and improving prediction accuracy.

[0045] 2. By adopting the training concept of regional models, the resulting model can learn the hydrological characteristics of multiple watersheds. For similar watersheds outside the dataset, this invention can also achieve good prediction results. At the same time, for extreme flood events, this invention solves the problem of inaccurate prediction of target watersheds due to the lack of historical data. The regional model can learn historical extreme flood event patterns from other watersheds and use them for prediction of target watersheds. Attached Figure Description

[0046] Figure 1 This is a flowchart of the method of the present invention;

[0047] Figure 2 This is a schematic diagram of the SE-LSTM model structure proposed in this invention. Detailed Implementation

[0048] The present invention will now be described in further detail with reference to the accompanying drawings.

[0049] like Figure 1 As shown, this invention proposes a runoff prediction method based on static physical feature enhancement and regional models. Taking the Jiangsu-Zhejiang-Anhui region as an example, the Tunxi, Changhua, Pihe, and Jiaojiang river basins are selected as the dataset. The specific steps include:

[0050] Step 1: Collect the static physical characteristics of all small and medium-sized watersheds in the region, as well as the time series data and climate data of each station monitored during typical historical floods.

[0051] In this embodiment, flood-related characteristics of four small watersheds within the region are collected, including flow data from hydrological stations at hourly intervals, precipitation data from rain gauges at hourly intervals, evapotranspiration, air temperature, average surface air temperature, wind speed, and humidity data from meteorological stations, and static characteristic data of each watershed, including average elevation, watershed area, slope, and underlying surface type.

[0052] Step 2: Calculate watershed similarity using the features collected in Step 1. The specific implementation process is as follows:

[0053] The collected data was completed using first-order linear interpolation. Then, the dynamic time-series data sequence X = {x1, x2…x} was processed. n First, the standard deviation σ and the mean μ are calculated. Then, the following formula is used for preprocessing to obtain the processed result X. * :

[0054]

[0055] The normalized result is segmented according to a preset length of the view window to obtain the input data:

[0056]

[0057] The collected time-series data are processed through the above steps to calculate the mean flow rate, variance of flow rate, number of flow rate peaks, mean rainfall, variance of rainfall, average wind speed, average humidity, average evapotranspiration, and average temperature. A peak-finding algorithm is used to count the peak values ​​in the time-series data, and the calculated results are used as features (number of flow rate peaks). Integrating all these features, a total of 14 dimensions are included (mean flow rate, variance of flow rate, number of flow rate peaks, mean rainfall, variance of rainfall, average wind speed, average humidity, average evapotranspiration, average temperature, average surface temperature, average altitude, watershed area, slope, and underlying surface type).

[0058] The similarity is calculated using Mahalanobis distance for all features. The Mahalanobis distance formula is:

[0059]

[0060] Where x represents a matrix of 14 attributes, μ represents the mean matrix of each attribute, and S represents the covariance matrix. Since static features and time-series features contribute differently to the prediction, the covariance matrix S needs to be multiplied by a weight matrix W.

[0061] The calculated Mahalanobis distances were filtered according to a preset threshold. All collected watershed results were within the preset value, thus identifying similar watersheds. In this implementation, with 13 degrees of freedom and a probability of 0.05, the chi-square distribution critical value was calculated to be 22.36. The Mahalanobis distances were then filtered according to this boundary. Since the collected watershed results in step 1 were all within the critical value, it indicates that these watersheds are similar.

[0062] Step 3: Preprocess the selected watersheds. The specific implementation process is as follows:

[0063] Calculate the common start and end times for all watershed data. Divide the watershed data starting from the common start point, setting the lookback window length to 120 time steps and the prediction length to 1 time step.

[0064] Each static physical feature and input length is embedded into a 120-dimensional vector space.

[0065] Step 4: As Figure 2 As shown, a flood forecasting model based on static physical feature enhancement, SE-LSTM, is constructed, including an encoder and a decoder. The encoder uses Bi-LSTM as the skeleton to learn the impact of past monitoring values ​​and future forecast values ​​on the current flow. The static feature fusion module SFU in the encoder amplifies the watershed static features that have the greatest impact on the flow at each time. The decoder uses LSTM as the skeleton to process the input and uses the patterns learned in the encoder to make predictions.

[0066] The Encoder receives both dynamic time-series data and static data. At each time step, the input sequence data is first processed using SFU (Structured Flow Fusion) to calculate a weighted input. This input is then fed into a Bi-LSTM (Bi-LSTM), with intermediate states generated during the process stored in the LSTM's memory. The forward LSTM learns the impact of past rainfall values ​​on the current data, while the backward LSTM learns the potential impact of future forecasts on the current flow. After all data in the Encoder has been learned, the forward and backward memory are fused together as the initial input to the Decoder. The Decoder uses the memory learned in the Encoder and the static feature weights to set the initial state. The Decoder receives all time-series data except for flow data, fuses it with the static feature weights, and uses this fused data as the model's input. It then uses the knowledge learned from the memory to output predicted values.

[0067] The workflow of the Static Feature Fusion (SFU) module in SE-LSTM is as follows:

[0068] The embedded static features are first weighted by the Sigmoid function to calculate their influence on the time series data. Then, the time series data is multiplied by this weight. Finally, a residual structure is used to add the data fused with the static features to the original data. Input variable X' t and static features The formula for selecting weights using a gating mechanism is as follows:

[0069]

[0070] in, Represents the embedding result of the i-th (i∈[1,7]) static feature, res i This represents the result of fusing the i-th static feature and time series data. Linear represents a linear layer, Sigmoid represents an activation function, and GatedLayer represents a gated layer, which is a combination of a linear layer and an activation function.

[0071] The embedded result sets res1, res2, ..., res7 are aggregated using the Concat aggregation function in the channel direction to obtain a set of all static features and input data fused in their respective channels:

[0072] res=Concat(res1,res2,…,res7)

[0073] The aggregated data (size [120, 7]) is passed through an attention layer to extract the influence of different static features on the time series data, resulting in a sequence with the shape [120, 1]. Finally, a residual concatenation is performed with the original time series sequence to obtain an output with the shape [120, 1].

[0074] output' = Attention(res)

[0075] output = output' + X' t

[0076] The output is fed into a Bi-LSTM for further information extraction and prediction, where Attention represents the activation function.

[0077] Step 5: Divide the data processed in Steps 2 and 3 into training set, validation set and test set according to a certain ratio; train the model using the training set and validation set, and optimize the internal weight parameters of the flood forecasting model SE-LSTM.

[0078] Initialize SE-LSTM: Set the hidden layer size to 256; the loss function to NSE; the optimizer to Adam; the input sequence length to 120; the output sequence length to 1; the batch size to 256; the training epochs to 30; and initialize all model parameters to 0.

[0079] In each round, all data is divided into batches of size, and one batch of data is selected and fed into the SE-LSTM model. Each batch randomly contains data from all sub-watersheds, thus achieving the effect of a regional model.

[0080] After each batch of data is trained, the loss function is calculated using the predicted and actual values, and the Adam optimizer is used to perform gradient descent and update the model parameters.

[0081] Step 6: Test the trained flood forecasting model SE-LSTM with the test set, and evaluate the performance of the flood forecasting model SE-LSTM by comparing the predicted flow values ​​and the actual flow values.

[0082] The Yixian County watershed in Anhui Province was selected as the target watershed. Features were extracted to calculate watershed similarity and to verify whether the watershed and the watersheds included in the regional model are similar.

[0083] The model trained in step 5 was validated using historical flow data from the Dongfanghong Reservoir, and the accuracy of the predictions was assessed using MSE and NSE indices.

[0084] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A runoff prediction method based on static physical feature enhancement and regional model, characterized in that, Includes the following steps: (1) Collect the static physical characteristics of all small and medium-sized watersheds in the region, as well as the time series data and climate data of each station monitored during typical historical floods; (2) The time series data were normalized by standard deviation normalization and the Mahalanobis distance was used to calculate the similar watersheds based on the time series data, climate data and static physical characteristics according to a certain threshold. (3) Process the static physical features in the selected watershed and embed the static physical features of different dimensions into the same dimensional space as the time series data; (4) Construct a flood forecasting model SE-LSTM based on static physical feature enhancement, including an encoder and a decoder. The encoder uses Bi-LSTM as the skeleton to learn the impact of past monitoring values ​​and future forecast values ​​on the current flow. The static feature fusion module SFU in the encoder amplifies the watershed static features that have the greatest impact on the flow at each time. The decoder uses LSTM as the skeleton to process the input and uses the patterns learned in the encoder to make predictions. (5) Divide the data processed by steps (2) and (3) into training set, validation set and test set according to a certain ratio; train the model with training set and validation set and optimize the internal weight parameters of flood forecasting model SE-LSTM; (6) Test the trained flood forecasting model SE-LSTM with the test set, and evaluate the performance of the flood forecasting model SE-LSTM by comparing the predicted flow value and the actual flow value; The implementation process of step (2) is as follows: (21) For dynamic time-series data sequences First, calculate the standard deviation. and average Then, the following formula is used for preprocessing to obtain the processing result. : The normalized result is segmented according to a preset length of the view window to obtain the input data: ; (22) Use the peak finding algorithm to count the peak values ​​of the time series data, use the calculated results as features, and integrate all the above features; (23) The similarity is calculated using Mahalanobis distance for all features. The Mahalanobis distance formula is: in, A matrix representing each attribute. This represents the mean matrix of each attribute. Let W represent the covariance matrix. Since static features and time-series features contribute differently to the prediction, W is the weight matrix. (24) The calculated Mahalanobis distance is filtered according to the preset threshold. The collected small watershed results are all within the preset value, and similar watersheds are filtered out. The workflow of the Static Feature Fusion (SFU) module in step (4) is as follows: The embedded static features are first evaluated using a sigmoid function to calculate the weight of the current static physical features on the time series data. Then, the time series data is multiplied by this weight. Finally, a residual structure is used to add the data fused with the static features to the original data. Input variables... and static features The formula for selecting weights using a gating mechanism is as follows: in, Indicates the first Embedding results of static features Indicates the first The result of fusing static features and time-series data; For the embedded result set Aggregation is performed along the channel direction to obtain a set of fused results of all static features and input data in their respective channels: Aggregated data By using an attention layer, the influence of different static features on time-series data is extracted, resulting in a shape like... The sequence is then rejoined with the original time series sequence using residual concatenation to obtain the final sequence. : Use the Sigmoid activation function to extract The nonlinear relationships in the data are then processed through Dropout and fully connected layers to obtain time-series data with static feature enhancement, which is then fed into the Bi-LSTM model for further analysis and processing.

2. The runoff prediction method based on static physical feature enhancement and regional model according to claim 1, characterized in that, The static physical characteristics mentioned in step (1) include mean evapotranspiration, mean air temperature, mean surface air temperature, mean elevation, watershed area, slope, and underlying surface type.

3. The runoff prediction method based on static physical feature enhancement and regional model according to claim 1, characterized in that, The time series data mentioned in step (1) includes rainfall, river flow, evapotranspiration, temperature, humidity, and wind speed data.

4. The runoff prediction method based on static physical feature enhancement and regional model according to claim 1, characterized in that, The implementation process of step (3) is as follows: Calculate the common start and end times of all watershed data; segment the watershed data starting from the common start point, set the lookback window length to 120 time steps, and the prediction length to 1 time step; embed each static physical feature and input length into a 120-dimensional vector space.

5. The runoff prediction method based on static physical feature enhancement and regional model according to claim 1, characterized in that, The working process of the SE-LSTM flood forecasting model based on static physical feature enhancement described in step (4) is as follows: The Encoder part of the model receives dynamic time-series data and static data. At each time step, the input sequence data is first used with the static data to calculate the weighted input using SFU. Then, the input is fed into Bi-LSTM, and the intermediate states generated during this process are stored in the memory of LSTM. The forward LSTM learns the impact of past rainfall values ​​on the current situation, and the backward LSTM learns the potential impact of future forecast values ​​on the current flow. After all the data in the Encoder has been learned, the memory of the forward and backward LSTMs is merged as the initial input of the Decoder. The Decoder part of the model uses the memory and static feature weights learned in the Encoder to set the initial state; the Decoder receives all time-series data except for traffic, and after fusing it with the static feature weights, it serves as the input to the model, and outputs the predicted value using the knowledge learned in the memory.

6. The runoff prediction method based on static physical feature enhancement and regional model according to claim 1, characterized in that, The preset threshold mentioned in step (24) is the chi-square distribution critical value of 22.36, which is calculated with 13 degrees of freedom and a probability value of 0.05.