Data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm

By combining multi-source remote sensing information with machine learning algorithms, the problem of lack of ground monitoring equipment was solved, an intelligent prediction model for river and lake water levels was constructed, the accuracy of river and lake water level predictions was improved, and the prediction needs of river basin flow and water resources were met.

CN118395275BActive Publication Date: 2025-09-16CHINA INST OF WATER RESOURCES & HYDROPOWER RES +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410112992.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-09-16
Estimated Expiration
2044-01-26

AI Technical Summary

Technical Problem

In existing technologies, the lack of ground monitoring equipment and meteorological and hydrological observation data leads to low accuracy in river and lake water level predictions. Traditional physical mechanism models find it difficult to accurately depict the hydrological cycle, resulting in high uncertainty in river and lake water level prediction results, making it difficult to meet business needs.

Method used

By using multi-source remote sensing information and machine learning algorithms, and acquiring satellite remote sensing products for data fusion and processing, an intelligent prediction model for river and lake water levels is constructed. The machine learning model is used to learn the driving mechanism of meteorological factors and river and lake water levels, reducing errors caused by inaccurate characterization of physical processes.

Benefits of technology

It improves the accuracy of river and lake water level predictions, provides key data support, and meets the business needs of basin flow and water resources prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118395275B_ABST
    Figure CN118395275B_ABST
Patent Text Reader

Abstract

The present invention discloses a data-deficient river and lake water level prediction method based on multi-source remote sensing information and a machine learning algorithm. The method utilizes multiple satellite altimetry products to obtain historical river and lake water level data, multiple remotely sensed precipitation products to obtain historical precipitation data for river and lake basins, and multiple remotely sensed evapotranspiration products to obtain historical evapotranspiration data for river and lake basins. A machine learning model is introduced, using remotely sensed inverted river and lake water level data, precipitation data, and evapotranspiration data as the primary training data. With water level as the prediction variable and precipitation and evapotranspiration as the primary predictive factors, an intelligent river and lake water level prediction model is established. The model then uses corrected future precipitation forecast data to predict future river and lake water level changes. The advantages of this method are: By using big data methods to learn the driving mechanisms between basin meteorological factors and river and lake water levels, the method reduces the errors and uncertainties caused by the inaccurate depiction of physical processes by traditional mechanism models, thereby further improving the accuracy of river and lake water level predictions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing hydrological technology, and in particular to a method for predicting water levels of rivers and lakes with insufficient data based on multi-source remote sensing information and a machine learning algorithm. Background Art

[0002] Rivers and lakes are crucial components of a river basin's hydrological cycle. Accurately predicting river and lake water level trends is crucial for forecasting downstream flows and predicting water resources within a river basin. Currently, developing hydrological models tailored to river basin realities is a crucial tool for predicting river and lake water levels. However, over 80% of rivers and lakes worldwide lack ground-based monitoring equipment and corresponding meteorological and hydrological observation data, making it difficult to accurately calibrate hydrological model parameters and resulting in low accuracy in river and lake water level forecasts. Furthermore, because the mechanisms of water transport between land, rivers, lakes, and the vadose zone remain poorly understood, existing physics-based hydrological models often struggle to accurately depict the real-world river-lake hydrological cycle. Even for rivers and lakes with relatively complete monitoring data, water level forecasts are subject to significant uncertainty. Consequently, water level forecasts for data-poor rivers and lakes are relatively inaccurate, making it difficult to fully meet operational needs for basin-wide flow or water resource forecasts. Summary of the Invention

[0003] The purpose of the present invention is to provide a data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithms, thereby solving the aforementioned problems existing in the prior art.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0005] A data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm includes the following steps:

[0006] S1. Acquisition and preliminary processing of satellite remote sensing products:

[0007] Based on the target rivers, lakes and their basin range files, obtain and process satellite altimetry water level products, satellite remote sensing precipitation products and satellite remote sensing evapotranspiration products to obtain satellite altimetry water level product inversion data, satellite remote sensing precipitation product inversion data and satellite remote sensing evapotranspiration product inversion data respectively;

[0008] S2. Optimal fusion of multi-source remote sensing inversion data:

[0009] Based on the intersection of the time ranges of each inversion data, each inversion data is interpolated to the daily scale and sorted. The generalized triangular hat method is used to optimally fuse multiple sets of inversion data to obtain the optimal fused water level of the target river and lake, and the optimal fused precipitation and evapotranspiration of the river and lake basin.

[0010] S3. Construction of intelligent prediction model for river and lake water levels:

[0011] Based on multiple machine learning models, an intelligent prediction model for river and lake water levels was constructed. The optimal fusion of river and lake water levels, precipitation, and evapotranspiration was divided into training and test sets. For each machine learning model, a spatial n-fold cross-validation was performed using the training set, and each cross-validation was tested using the test set. The n predicted values ​​output by the model on the entire training set were superimposed end to end as the predicted value of the training set. The model accuracy was evaluated using evaluation indicators, and the optimal intelligent prediction model for river and lake water levels was obtained through training.

[0012] S4. Future river and lake water level forecast:

[0013] Based on the ratio of the average value of precipitation and evapotranspiration inversion data to the average value of precipitation and evapotranspiration forecast results, a correction factor is obtained; the future precipitation forecast data and evapotranspiration forecast data corrected by the correction factor are input into the optimal river and lake water level intelligent prediction model to obtain the water level changes of rivers and lakes in the corresponding period in the future.

[0014] Preferably, the target rivers, lakes and their basin range files are obtained in the following way:

[0015] Select target rivers and lakes, and obtain the boundary range of rivers and lakes based on publicly available global river and lake datasets; select appropriate digital elevation data, use the river and lake outlet sections as feature points, determine the scope of river and lake basins in geographic information software, and generate target rivers, lakes and their basin scope files.

[0016] Preferably, step S1 specifically includes the following contents:

[0017] S11. Acquisition and processing of satellite altimetry water level products:

[0018] Select the longest possible historical period, download multiple publicly available satellite altimetry water level products, and extract all satellite track points within the target river or lake. For the same revisit of a particular satellite altimetry water level product, calculate the median and standard deviation of the water level elevation corresponding to all satellite track points. Remove data with water level values ​​greater than three times the standard deviation, and use the median of the remaining water level data points as the final water level inversion value for the satellite product. Repeat this process for each satellite altimetry product in sequence.

[0019] S12. Acquisition and processing of satellite remote sensing precipitation products:

[0020] Select the longest possible historical period, download multiple publicly available satellite remote sensing precipitation products, and extract all grid points within the target river or lake. For each satellite remote sensing precipitation product, calculate the average surface rainfall within the river or lake basin at that moment, and use this as the precipitation inversion value for that satellite remote sensing precipitation product. Repeat the above process for each satellite remote sensing precipitation product.

[0021] S13. Acquisition and processing of satellite remote sensing evapotranspiration products:

[0022] Select the longest possible historical period, download multiple publicly available satellite remote sensing evapotranspiration products, and extract all grid points within the target river or lake. For a particular satellite remote sensing evapotranspiration product at the same moment, calculate the surface average evapotranspiration within the river or lake basin at that moment as the evapotranspiration inversion value of the satellite remote sensing evapotranspiration product. Repeat the above process for each satellite remote sensing evapotranspiration product in sequence.

[0023] Preferably, the multiple satellite altimetry water level products are Jason-2, ICESat-2, and Sentinel series satellite altimetry products; the multiple satellite remote sensing precipitation products are GPM, TRMM, and PERSIANN series satellite remote sensing precipitation products; and the multiple remote sensing evapotranspiration products are MODIS, MERRA, and GLEAM series satellite remote sensing evapotranspiration products.

[0024] Preferably, step S2 specifically includes the following contents:

[0025] S21. Time series interpolation of multi-source remote sensing data:

[0026] Determine the time range intersection of the satellite altimetry water level product inversion data, satellite remote sensing precipitation product inversion data, and satellite remote sensing evapotranspiration product inversion data. Use the linear interpolation method within this time range intersection to uniformly interpolate the inversion data to the daily scale and sort them in chronological order.

[0027] S22. Optimal fusion of multi-source remote sensing data:

[0028] The generalized triangular hat method is used to estimate the uncertainty of different data sets and perform optimal fusion based on the estimated variance; the calculation formula is:

[0029]

[0030]

[0031] Among them, X i is the i-th data set, i=1,2,…,N; X is the result of weighted fusion of i data sets; p i For dataset X i The corresponding weight; ω iThe dataset X is estimated by the generalized tricorn hat method i The variance of .

[0032] Preferably, before dividing the training set and the test set in step S3, the optimal fused water level of the target river and lake, the optimal fused precipitation and evapotranspiration of the river and lake basin need to be standardized.

[0033] Preferably, the multiple machine learning models are support vector regression, random forest and artificial neural network.

[0034] Preferably, the evaluation index is one or more of the mean deviation, relative error, root mean square error, and normalized standard deviation.

[0035] Preferably, step S4 specifically includes the following contents:

[0036] S41. Revision of future precipitation forecast data:

[0037] Select precipitation retrieval data and evapotranspiration data for a period of time before the forecast start time and compare them with the forecast data of the meteorological forecast product for the same period. Calculate the ratio of the average precipitation and evapotranspiration retrieval data to the average precipitation and evapotranspiration forecast results for this period, and record it as the correction factor. Multiply the correction factor by the future precipitation forecast results for each period to obtain the revised future precipitation forecast data.

[0038] S42. River and lake water level forecast:

[0039] The revised future precipitation forecast data and evapotranspiration forecast data are input into the optimal river and lake water level intelligent prediction model to obtain the water level changes of rivers and lakes in the corresponding period in the future.

[0040] The beneficial effects of the present invention are as follows: 1. The method of the present invention can overcome the problem of lack of data on rivers and lakes. By acquiring multi-source remote sensing products and performing temporal interpolation and optimal fusion, it generates historical precipitation and evapotranspiration data series for river and lake basins, as well as historical water level change series for rivers and lakes, providing key data support for river and lake water level prediction. 2. Based on the use of remote sensing to obtain relevant meteorological and hydrological data on rivers, lakes and their basins, the method of the present invention uses machine learning methods to establish a prediction model for river and lake water levels. By using big data methods to learn the corresponding driving mechanisms of basin meteorological factors and river and lake water levels, it reduces the errors and uncertainties caused by the inaccurate portrayal of physical processes by traditional mechanism models, thereby further improving the prediction accuracy of river and lake water levels. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0042] Figure 2This is the Lugu Lake water level prediction result obtained by using the method of the present invention in an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0044] Example 1

[0045] like Figure 1 As shown, in this embodiment, in order to address the current difficulty in predicting river and lake water levels due to lack of data, which in turn leads to low accuracy in downstream basin flow prediction, we fully utilize multi-source satellite altimetry products such as Jason, multi-source remote sensing precipitation inversion products such as GPM, and multi-source remote sensing evapotranspiration products such as MODIS. We use historical precipitation and evapotranspiration sequences as prediction factors and historical river and lake altimetry inversion water levels as prediction variables, introduce machine learning models such as random forests and artificial neural networks, establish an intelligent river and lake water level prediction model, and use revised future meteorological forecast data to predict future river and lake water level changes. A data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithms is proposed. The method mainly includes the following parts:

[0046] 1. Acquisition and preliminary processing of satellite remote sensing products

[0047] Based on the target rivers, lakes and their basin range files, obtain and process satellite altimetry water level products, satellite remote sensing precipitation products and satellite remote sensing evapotranspiration products to obtain satellite altimetry water level product inversion data, satellite remote sensing precipitation product inversion data and satellite remote sensing evapotranspiration product inversion data respectively. Specifically including the following contents,

[0048] 1.1 Determination of river and lake water body boundaries and division of river basin boundaries

[0049] Select target rivers and lakes, and obtain the boundary range of rivers and lakes based on publicly available global river and lake datasets; select appropriate digital elevation model (DEM) data, use the river and lake outlet sections as feature points, determine the scope of river and lake basins in geographic information software, and generate target river, lake and basin scope files.

[0050] 1.2. Acquisition and processing of satellite altimetry water level products

[0051] Select the longest possible historical period, download multiple publicly available satellite altimetry water level products such as Jason, and extract all satellite track points within the target river and lake range. For the same revisit of a certain satellite altimetry water level product (such as Jason), calculate the median and standard deviation of the water level elevation corresponding to all satellite track points, eliminate data with water level values ​​greater than three times the standard deviation, and use the median of the remaining water level data points as the final water level inversion value of the satellite product. Repeat the above process for each satellite altimetry product in turn.

[0052] 1.3 Acquisition and processing of satellite remote sensing precipitation products

[0053] Select the longest possible historical period, download multiple publicly available satellite remote sensing precipitation products such as GPM, and extract all grid points within the target river or lake range. For a certain satellite remote sensing precipitation product (such as GPM) at the same moment, calculate the surface average rainfall within the river or lake basin at that moment as the precipitation inversion value of the satellite remote sensing precipitation product. Repeat the above process for each satellite remote sensing precipitation product in turn.

[0054] 1.4 Acquisition and processing of satellite remote sensing evapotranspiration products

[0055] Select the longest possible historical period, download multiple publicly available satellite remote sensing evapotranspiration products such as MODIS, and extract all grid points within the target river or lake. For a particular satellite remote sensing evapotranspiration product (such as MODIS) at the same moment, calculate the surface average evapotranspiration within the river or lake basin at that moment and use it as the evapotranspiration inversion value of the satellite remote sensing evapotranspiration product. Repeat the above process for each satellite remote sensing evapotranspiration product.

[0056] 2. Optimal Fusion of Multi-source Remote Sensing Inversion Data

[0057] Based on the intersection of the time ranges of each inversion data, each inversion data is interpolated to the daily scale and sorted; the generalized triangular hat method is used to optimally fuse multiple sets of inversion data to obtain the optimal fused water level of the target river and lake, the optimal fused precipitation and evapotranspiration of the river and lake basin. Specifically, it includes the following contents:

[0058] 2.1. Time series interpolation of multi-source remote sensing data:

[0059] The intersection of the time ranges of the multi-source river and lake water level inversion data, precipitation inversion data, and evapotranspiration inversion data is taken. The linear interpolation method is used to uniformly interpolate the inversion data to the daily scale at the intersection of the time ranges and sort them in chronological order.

[0060] 2.2 Optimal Fusion of Multi-Source Remote Sensing Data

[0061] The TCH method can be used to evaluate the uncertainty of multiple datasets in the absence of prior information. Therefore, the present invention uses the TCH method to estimate the uncertainty of different datasets (or models) and optimally fuses the estimated variance to obtain more reliable results. i ,i=1,2,…,N}, the weighted fusion result can be expressed as:

[0062]

[0063]

[0064] Among them, X i is the i-th data set, i=1,2,…,N; X is the result of weighted fusion of i data sets; p i For dataset X i The corresponding weight; ω i The dataset X is estimated by the generalized tricorn hat method i The variance of .

[0065] 3. Construction of intelligent prediction model for river and lake water levels

[0066] Based on multiple machine learning models, an intelligent prediction model for river and lake water levels is constructed. The optimal fusion of river and lake water levels, precipitation, and evapotranspiration is divided into training sets and test sets. For each machine learning model, a spatial n-fold cross-validation is performed using the training set and the test set is used for testing during each cross-validation. The n prediction values ​​output by the model on the entire training set are superimposed end to end as the prediction value of the training set. The model accuracy is evaluated through evaluation indicators, and the optimal intelligent prediction model for river and lake water levels is obtained through training. Specifically, it includes the following contents:

[0067] 3.1 Training Data Preprocessing

[0068] After obtaining the optimal fused precipitation, evapotranspiration, and river and lake water levels for river and lake basins, the above data are standardized to avoid subsequent modeling instability and slow model convergence due to excessive data fluctuations:

[0069]

[0070] Where: X is any series of precipitation, evapotranspiration and water level values, X min and X max are the minimum and maximum values ​​in the series, respectively.

[0071] 3.2. Utilize a machine learning model to construct an intelligent prediction model for river and lake water levels. The input variables are precipitation and evapotranspiration in river and lake basins, and the output variable is river and lake water levels. Therefore, the data required for training includes historical precipitation and evapotranspiration data for river and lake basins, as well as river and lake water level data at the corresponding time, all acquired through the aforementioned remote sensing methods. All data is divided into a training set and a test set, and the training set is further divided into n parts. For each artificial intelligence model, these n sub-training sets are used in turn for a spatial n-fold cross-validation, and the test set is tested during each cross-validation. Ultimately, the model outputs n predicted values ​​for the entire training set. The sum of these values ​​represents the predicted value for the entire training set. Model accuracy is calculated using evaluation metrics such as mean deviation, relative error, root mean square error, and normalized standard deviation, ultimately yielding the optimal intelligent prediction model for river and lake water levels.

[0072] IV. Future River and Lake Water Level Forecast

[0073] Based on the ratio of the average value of precipitation and evapotranspiration inversion data to the average value of precipitation and evapotranspiration forecast results, a correction factor is obtained; the future precipitation forecast data and evapotranspiration forecast data corrected by the correction factor are input into the optimal river and lake water level intelligent prediction model to obtain the water level changes of rivers and lakes in the corresponding period in the future. Specifically,

[0074] 4.1. Revision of future precipitation forecast data

[0075] The precipitation inversion data and evapotranspiration data for a period of time before the forecast start time are selected and compared with the forecast data of the meteorological forecast product for the same period. The ratio of the average precipitation and evapotranspiration inversion data of this period to the average precipitation and evapotranspiration forecast results is calculated, and recorded as the correction factor f; the correction factor f is multiplied by the future precipitation forecast results for each period to obtain the revised future precipitation forecast data.

[0076] 4.2 River and Lake Water Level Forecast

[0077] The revised future precipitation forecast data and evapotranspiration forecast data are input into the optimal river and lake water level intelligent prediction model to obtain the water level changes of rivers and lakes in the corresponding period in the future.

[0078] Example 2

[0079] In this embodiment, the method of the present invention is described by taking the water level prediction of Yalong River and Lugu Lake as an example:

[0080] 1. Acquisition and preliminary processing of satellite remote sensing products

[0081] 1.1. Determination of the water body area and division of the watershed of Lugu Lake

[0082] Based on the geographical location of Lugu Lake, the HydroLakes global river and lake dataset (download address: https: / / www.hydrosheds.org / products / hydrolakes) was used to obtain the water body boundary shapefile file of Lugu Lake. The digital elevation model (DEM) data of the Yalong River near Lugu Lake was obtained by selecting STRM90m. The outlet section of Lugu Lake was used as the feature point. The flow direction calculation, river network generation, and watershed delineation steps were performed in the geographic information software ArcGIS to determine the watershed area of ​​Lugu Lake and generate the watershed area shapefile file.

[0083] 1.2. Acquisition and processing of satellite altimetry water level products

[0084] Selecting 2015 to 2021 as a typical period, we downloaded Jason-2, ICESat-2, and Sentinel satellite altimetry products and extracted all satellite track points within the Lugu Lake area. The Jason-2 altimetry product can be downloaded from: https: / / aviso-data-center.cnes.fr / ; the ICESat-2 product from: https: / / nsidc.org / data / icesat-2 / ; and the Sentinel product from: https: / / nsidc.org / data / icesat-2 / . To improve the quality and validity of Lugu Lake water level retrieval data for a given satellite product (such as Jason), we calculated the median and standard deviation of the water level elevations corresponding to all satellite track points. Data with water level values ​​greater than three times the standard deviation were then removed, and the median of the remaining water level data points was used as the final water level retrieval value for that satellite product. This process was repeated for each satellite altimetry product.

[0085] 1.3 Acquisition and processing of remote sensing precipitation products

[0086] Selecting the period from 2015 to 2021 as a representative period, download the GPM, TRMM, and PERSIANN satellite remote sensing precipitation products and extract all grid points within the Lugu Lake area. The GPM download address is: https: / / pmm.nasa.gov / data-access / downloads / gpm; the TRMM download address is: https: / / pmm.nasa.gov / data-access / downloads / trmm; and the PERSIANN download address is: ftp: / / persiann.eng.uci.edu / pub / GCCS. For each remote sensing product (e.g., GPM), use data processing tools such as NCL to calculate the areal average rainfall within the Lugu Lake basin at that time. This serves as the precipitation inversion value for that remote sensing product. Repeat the above steps for each remote sensing precipitation product.

[0087] 1.4 Acquisition and processing of remote sensing evapotranspiration products

[0088] Select the longest possible historical period and download the MODIS, MERRA, and GLEAM satellite remote sensing evapotranspiration products. Extract all grid points within the Lugu Lake water body. The MODIS download address is: https: / / modis.gsfc.nasa.gov / ; the MERRA download address is: https: / / gmao.gsfc.nasa.gov / reanalysis / MERRA-2 / data_access / ; and the GLEAM download address is: https: / / www.gleam.eu / . For a specific remote sensing product (such as GPM) at the same time, use data processing tools such as NCL to calculate the surface average evapotranspiration within the Lugu Lake basin at that time. This serves as the evapotranspiration inversion value for that remote sensing product. Repeat the above steps for each remote sensing evapotranspiration product.

[0089] 2. Optimal Fusion of Multi-source Remote Sensing Inversion Data

[0090] 2.1 Time Series Interpolation of Multi-source Remote Sensing Data

[0091] The intersection of the time range of the multi-source Lugu Lake water level, precipitation, and evapotranspiration data was taken, covering January 1, 2018, to December 31, 2020. For Lugu Lake, the Jason-2, ICESat-2, and Sentinel satellites have temporal resolutions of 10 days, 29 days, and 10 days, respectively. Remotely sensed precipitation and evapotranspiration data for the Lugu Lake basin are all at daily resolution. Linear interpolation was used to uniformly interpolate these remotely sensed data to a daily scale over the 2018-2020 timeframe and arranged in chronological order.

[0092] 2.2 Optimal Fusion of Multi-Source Remote Sensing Data

[0093] The generalized triangular hat method (TCH) is used to optimally fuse multiple sets of precipitation, evapotranspiration and water level inversion data. i ,i=1,2,3}, the weighted fusion result can be expressed as:

[0094]

[0095] Among them, p i For dataset X i The corresponding weight is calculated as follows:

[0096]

[0097] Among them, ω i The data set X estimated by the TCH method i The variance of .

[0098] The results show that for GPM, TRMM and PERSIANN precipitation data, p i 0.44, 0.21 and 0.35 respectively; for MODIS, MERRA and GLEAM evapotranspiration data, p i 0.11, 0.15, and 0.74, respectively; for Jason, ICESat, and Sentinel data, p i Take 0.51, 0.24 and 0.25 respectively.

[0099] 3. Construction of intelligent prediction model for Hugu Lake water level

[0100] 3.1 Training Data Preprocessing

[0101] After obtaining the optimal fused precipitation, evapotranspiration, and water level of the Lugu Lake basin, the above data were standardized to avoid subsequent modeling instability and slow model convergence due to excessive data fluctuations:

[0102]

[0103] Where: X is any series of precipitation, evapotranspiration and water level values, X min and X max are the minimum and maximum values ​​in the series, respectively.

[0104] 3.2 Machine Learning Model Training, Validation, and Testing

[0105] Three artificial intelligence models, including support vector regression (SVR), random forest (RF), and artificial neural network (ANN), were selected to construct an intelligent prediction model for the water level of Lugu Lake. Support vector regression is a typical statistical learning method that predicts future water levels by learning from error samples of historical water level predictions. Its basic idea is to transform the nonlinear problem of water level prediction into a linear problem in a high-dimensional space through a nonlinear kernel function. The random forest model uses ensemble learning to optimize a single weak predictor to improve prediction accuracy. The main idea is to combine multiple weak classifiers, and the final result is voted or averaged, so that the overall model has high accuracy and generalization performance, is less prone to overfitting, and ultimately achieves better water level prediction capabilities. Artificial neural networks can obtain complex nonlinear mapping capabilities by superimposing mathematical operations on each neuron node. They usually include an input layer, an output layer, and an intermediate hidden layer, each with a certain number of neurons. The input layer is mainly used to receive the output features of the previous model and does not participate in the calculation; the hidden layer receives information from the input layer and extracts features; finally, the output layer outputs the final water level prediction result based on the different weights of the hidden layer neural units and its own bias.

[0106] The input variables of the three models are all remote sensing inversion of precipitation and evapotranspiration in the Lugu Lake basin, and the output variables are all satellite inversion of the Lugu Lake water level. All data are divided into training sets and test sets, with the training period from 2018 to 2019 and the test period in 2020. The training set is further divided into five parts. For each artificial intelligence model, these five sub-training sets are used in turn for spatial 5-fold cross-validation, and the test set is tested at the same time as each cross-validation. The specific steps of the 5-fold cross-validation are as follows:

[0107] (1) Divide the dataset into five equal parts, each part is a fold;

[0108] (2) Using the first fold as the test set and the remaining 2 to 5 folds as the training set, a prediction model is obtained by training using the Bayesian optimization algorithm. This embodiment uses the average deviation b as the evaluation index to calculate the prediction accuracy value of the model, and the formula is:

[0109]

[0110] Among them, s i is the predicted value of Lugu Lake water level, i is the measured value of water level, and n is the length of the test set.

[0111] (3) Similarly, the i-th (i=2, 3, 4, 5) fold is used as the test set, and the rest are used as the training set. A prediction model is trained and the prediction accuracy of the model is obtained.

[0112] (4) The average of the prediction accuracy of each time is taken as the final accuracy of the model.

[0113] 4. Future Lugu Lake Water Level Forecast

[0114] 4.1. Revision of future weather forecast data

[0115] This example uses the Global Forecast System (GFS) product's 16-day precipitation and evapotranspiration data, reported starting at 00:00 UTC on May 3, 2022, as the data source for the future weather forecast. Precipitation and evapotranspiration retrieval data for the seven days prior to May 3 are compared with the GFS product's forecast data for the same period. The ratio of the average precipitation and evapotranspiration retrieval data to the average precipitation and evapotranspiration forecast data is calculated, denoted as the correction factor f. The correction factor f is multiplied by the daily weather forecast results for the future (May 3 to May 18) to obtain the revised future weather forecast data.

[0116] 4.2 Lugu Lake Water Level Forecast

[0117] The revised future precipitation forecast data and evapotranspiration forecast data were input into the Lugu Lake water level intelligent prediction model, and the model was used to obtain the water level changes of Lugu Lake in the corresponding period in the future. The results showed that the average error of the predicted water level of Lugu Lake from May 3 to May 18, 2022 was 0.9m (e.g. Figure 2 It meets the relevant accuracy requirements of hydrological forecasting and can provide business support for downstream flow prediction and water resources prediction.

[0118] By adopting the above technical solution disclosed in the present invention, the following beneficial effects are obtained:

[0119] The present invention provides a data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithms. This method overcomes the problem of data scarcity for rivers and lakes. By acquiring multi-source remote sensing products and performing temporal interpolation and optimal fusion, it generates historical precipitation and evapotranspiration data series for river and lake basins, as well as historical river and lake water level change series, providing key data support for river and lake water level prediction. Based on remote sensing-based meteorological and hydrological data related to rivers, lakes, and their basins, the method utilizes machine learning methods to establish a river and lake water level prediction model. Using big data methods, it studies the corresponding driving mechanisms of basin meteorological factors and river and lake water levels, reducing the errors and uncertainties caused by the inaccurate depiction of physical processes by traditional mechanism models, thereby further improving the accuracy of river and lake water level prediction.

[0120] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithms, characterized by: The following steps are included: S1. Acquisition and preliminary processing of satellite remote sensing products: Based on the target rivers, lakes and their basin range files, obtain and process satellite altimetry water level products, satellite remote sensing precipitation products and satellite remote sensing evapotranspiration products to obtain satellite altimetry water level product inversion data, satellite remote sensing precipitation product inversion data and satellite remote sensing evapotranspiration product inversion data respectively; S2. Optimal fusion of multi-source remote sensing inversion data: Based on the intersection of the time ranges of each inversion data, each inversion data is interpolated to the daily scale and sorted. The generalized triangular hat method is used to optimally fuse multiple sets of inversion data to obtain the optimal fused water level of the target river and lake, and the optimal fused precipitation and evapotranspiration of the river and lake basin. S3. Construction of intelligent prediction model for river and lake water levels: Based on multiple machine learning models, an intelligent prediction model for river and lake water levels is constructed. The optimal fusion of river and lake water levels, precipitation, and evapotranspiration is divided into training and test sets. For each machine learning model, a spatial n-fold cross-validation is performed on the training set, and the test set is used for each cross-validation. The n predicted values ​​output by the model on the entire training set are superimposed end to end as the predicted value of the training set. The model accuracy is evaluated using evaluation indicators to train the optimal intelligent prediction model for river and lake water levels. S4. Future river and lake water level forecast: Based on the ratio of the average value of precipitation and evapotranspiration inversion data to the average value of precipitation and evapotranspiration forecast results, a correction factor is obtained; the future precipitation forecast data and evapotranspiration forecast data corrected by the correction factor are input into the optimal river and lake water level intelligent prediction model to obtain the water level changes of rivers and lakes in the corresponding period in the future.

2. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 1 is characterized by: The target rivers, lakes and their basin range files are obtained as follows: Select target rivers and lakes, and obtain the boundary range of rivers and lakes based on publicly available global river and lake datasets; select appropriate digital elevation data, use the river and lake outlet sections as feature points, determine the scope of river and lake basins in geographic information software, and generate target rivers, lakes and their basin scope files.

3. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 1 is characterized by: Step S1 specifically includes the following contents: S11. Acquisition and processing of satellite altimetry water level products: Select the longest possible historical period, download multiple publicly available satellite altimetry water level products, and extract all satellite track points within the target river or lake. For the same revisit of a particular satellite altimetry water level product, calculate the median and standard deviation of the water level elevation corresponding to all satellite track points. Remove data with water level values ​​greater than three times the standard deviation, and use the median of the remaining water level data points as the final water level inversion value for the satellite product. Repeat this process for each satellite altimetry product in sequence. S12. Acquisition and processing of satellite remote sensing precipitation products: Select the longest possible historical period, download multiple publicly available satellite remote sensing precipitation products, and extract all grid points within the target river or lake. For a particular satellite remote sensing precipitation product, calculate the surface average rainfall within the river or lake basin at that moment, and use this as the precipitation inversion value for that satellite remote sensing precipitation product. Repeat the above process for each satellite remote sensing precipitation product in turn; S13. Acquisition and processing of satellite remote sensing evapotranspiration products: Select the longest possible historical period, download multiple publicly available satellite remote sensing evapotranspiration products, and extract all grid points within the target river or lake. For a particular satellite remote sensing evapotranspiration product at the same moment, calculate the surface average evapotranspiration within the river or lake basin at that moment as the evapotranspiration inversion value of the satellite remote sensing evapotranspiration product. Repeat the above process for each satellite remote sensing evapotranspiration product in sequence.

4. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 3 is characterized by: Multiple satellite altimetry water level products include Jason-2, ICESat-2, and Sentinel series satellite altimetry products; multiple satellite remote sensing precipitation products include GPM, TRMM, and PERSIANN series satellite remote sensing precipitation products; and multiple remote sensing evapotranspiration products include MODIS, MERRA, and GLEAM series satellite remote sensing evapotranspiration products.

5. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 1 is characterized by: Step S2 specifically includes the following contents: S21. Time series interpolation of multi-source remote sensing data: Determine the time range intersection of the satellite altimetry water level product inversion data, satellite remote sensing precipitation product inversion data, and satellite remote sensing evapotranspiration product inversion data. Use the linear interpolation method within this time range intersection to uniformly interpolate the inversion data to the daily scale and sort them in chronological order. S22. Optimal fusion of multi-source remote sensing data: The generalized tricorn hat method is used to estimate the uncertainty of different datasets and optimally fuse them based on the estimated variance; The calculation formula is, Among them, X i is the i-th data set, i=1,2,…,N; X is the result of weighted fusion of i data sets; p i For dataset X i The corresponding weight; ω i The dataset X is estimated by the generalized tricorn hat method i The variance of .

6. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 1 is characterized by: Before dividing the training set and the test set in step S3, the optimal fused water level of the target river and lake, the optimal fused precipitation and evapotranspiration of the river and lake basin need to be standardized.

7. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 1 is characterized by: The various machine learning models are support vector regression, random forest, and artificial neural network.

8. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 1 is characterized by: The evaluation indicators are one or more of the mean deviation, relative error, root mean square error, and normalized standard deviation.

9. The data-deficient river and lake water level prediction method based on multi-source remote sensing information and machine learning algorithm according to claim 1 is characterized by: Step S4 specifically includes the following contents: S41. Revision of future precipitation forecast data: Select the precipitation retrieval data and evapotranspiration data for a period of time before the forecast start time and compare them with the forecast data of the meteorological forecast product for the same period. Calculate the ratio of the average precipitation and evapotranspiration retrieval data to the average precipitation and evapotranspiration forecast results for this period, and record it as the correction factor. Multiply the correction factor by the future precipitation and evapotranspiration forecast results for each period to obtain the revised future precipitation forecast data and evapotranspiration forecast data; S42. River and lake water level forecast: The revised future precipitation forecast data and evapotranspiration forecast data are input into the optimal river and lake water level intelligent prediction model to obtain the water level changes of rivers and lakes in the corresponding period in the future.

Citation Information

Patent Citations

  • Runoff prediction system and method based on satellite microwave observation data

    CN112766531A

  • Water depth inversion method for multispectral remote sensing

    CN114117886A