Lake dissolved organic matter remote sensing inversion method based on interpretable ensemble learning
By processing remote sensing image data using an interpretable ensemble learning approach, a model is constructed and optimized to analyze feature contribution and generate a spatial distribution map of dissolved organic matter in lakes. This solves the problems of time-consuming and labor-intensive traditional methods and the lack of model interpretability, achieving high-precision and transparent DOM inversion.
Patent Information
- Application Number
- CN202411978237.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2026-01-23
AI Technical Summary
Traditional methods for monitoring dissolved organic matter concentration in lakes are time-consuming, labor-intensive, and costly, and machine learning models lack interpretability, resulting in insufficient accuracy and reliability in inversion.
An interpretable ensemble learning-based approach was adopted to construct and optimize XGBoost, CatBoost, and NGBoost models by atmospheric correction and masking of remote sensing image data. The BorutaShap algorithm was used to interpret the model features and generate a spatial distribution map of dissolved organic matter abundance.
It improves the accuracy and transparency of remote sensing inversion of dissolved organic matter in lakes, enhances the interpretability of the model, and enables better analysis of the spatiotemporal dynamics of DOM.
Smart Images

Figure CN121389693A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of lake DOM abundance inversion, and particularly relates to a lake dissolved organic matter remote sensing inversion method based on explainable ensemble learning. BACKGROUND
[0002] Dissolved organic matter (DOM) is an important chemical component in the lake ecosystem, rich in carbon, nitrogen, phosphorus and other elements, DOM not only can be decomposed by microorganisms to convert organic carbon into inorganic carbon, affecting the carbon budget and carbon cycle process of the lake. At the same time, it also provides rich energy and nutrients for microorganisms in the lake, affecting the biogeochemical environmental behavior of heavy metals, nutrients and organic pollutants. Understanding and analyzing the distribution and change of DOM concentration in the lake is an important basis for protecting water resources safety, and has practical significance for promoting ecological balance and promoting sustainable economic and social development.
[0003] Traditional DOM concentration monitoring methods mostly use field sampling and laboratory analysis, which is time-consuming, labor-intensive and costly, and cannot spatialize the lake DOM abundance. Remote sensing data is widely used in the monitoring of lake DOM concentration due to its high timeliness and large-scale coverage. The remote sensing inversion algorithm mainly includes empirical algorithms represented by logarithmic, exponential and linear regression, and the precision and reliability still need to be improved in dealing with remote sensing big data and complex water conditions. At present, machine learning technology is widely used in the research of water quality remote sensing inversion. However, machine learning is mostly a black box model, and researchers cannot give a scientific explanation of the internal mechanism and decision-making cause of the model, and cannot fully understand the specific calculation process of the inversion, so that some models with excellent training effect cannot gain people's trust.
[0004] Therefore, it is necessary to propose a lake dissolved organic matter remote sensing inversion method which enhances the model interpretability while improving the inversion accuracy on the basis of the existing lake DOM abundance inversion technology. SUMMARY
[0005] The present application aims to provide a lake dissolved organic matter remote sensing inversion method based on explainable ensemble learning, which aims to enhance the model interpretability while improving the inversion accuracy on the basis of the existing lake DOM abundance inversion technology.
[0006] To achieve the above-mentioned purpose, a lake dissolved organic matter remote sensing inversion method based on explainable ensemble learning is adopted, which comprises the following steps:
[0007] Obtain lake dissolved organic matter abundance data and remote sensing image data, and preprocess the image data; wherein the preprocessing methods include atmospheric correction and mask;
[0008] Divide the training set and test set according to the image data, and build an inversion model;
[0009] Optimize the inversion model, and select the inversion model with the highest precision;
[0010] Interpret the inversion model and analyze the calculation process of the inversion model;
[0011] Based on the inversion model, calculate the spatial distribution of lake dissolved organic matter abundance and analyze the change trend, and output the results.
[0012] Among them, the lake dissolved organic matter abundance data and remote sensing image data are obtained, and the image data is preprocessed; wherein the preprocessing method includes the steps of atmospheric correction, mask:
[0013] The remote sensing image data includes high spatial resolution images from the Sentinel-2 satellite; wherein the image data needs to be corrected by atmospheric correction.
[0014] Among them, the lake dissolved organic matter abundance data and remote sensing image data are obtained, and the image data is preprocessed; wherein the preprocessing method includes the steps of atmospheric correction, mask:
[0015] The mask processing identifies the lake boundary range based on the normalized difference water index.
[0016] Among them, in the step of optimizing the inversion model and selecting the inversion model with the highest precision:
[0017] The process of optimizing the inversion model is: using the Bayesian optimization algorithm to adjust the hyperparameters of XGBoost, CatBoost and NGBoost models; wherein the hyperparameters include learning rate, maximum depth and weak learner number.
[0018] Among them, in the step of interpreting the inversion model and analyzing the calculation process of the inversion model:
[0019] Using BorutaShap algorithm combined with Boruta feature selection and SHAP interpretability technology, the importance of the input features of the inversion model is evaluated, and the key features that contribute most to the inversion of lake dissolved organic matter are identified.
[0020] Among them, in the step of calculating the spatial distribution of lake dissolved organic matter abundance based on the inversion model and analyzing the change trend, and outputting the results:
[0021] The spatial distribution of lake dissolved organic matter abundance includes: using the final optimized inversion model to perform pixel-level prediction on high-resolution remote sensing images in the study area, and generating a spatial distribution map of dissolved organic matter abundance.
[0022] In the step of calculating the spatial distribution of lake dissolved organic matter abundance based on the inversion model, analyzing the change trend, and outputting the results:
[0023] The analysis of the change trend is achieved by comparing the spatial distribution of dissolved organic matter abundance generated in the dry season and the rainy season to reveal the temporal and spatial dynamic change rule of the lake dissolved organic matter.
[0024] The lake dissolved organic matter remote sensing inversion method based on the interpretable ensemble learning provided by the application obtains lake dissolved organic matter abundance data and remote sensing image data, and pre-processes the image data, wherein the pre-processing mode includes atmospheric correction and mask; the image data is divided into a training set and a test set, and an inversion model is constructed; the inversion model is optimized, and the inversion model with the highest precision is selected; the inversion model is interpreted to analyze the calculation process of the inversion model; based on the inversion model, the spatial distribution of lake dissolved organic matter abundance is calculated, the change trend is analyzed, and the results are output; by combining field measurement data with high spatial resolution remote sensing images, the ensemble learning technology is used to estimate the lake DOM abundance, and the model mechanism is explored based on the interpretable technology, which improves the inversion precision, improves the transparency of the model, and is helpful for better carrying out the lake DOM inversion work. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description only represent some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0026] Figure 1 is the step flow chart of the lake dissolved organic matter remote sensing inversion method based on the interpretable ensemble learning of the application. DETAILED DESCRIPTION
[0027] Hereinafter, exemplary embodiments will be described in detail with reference to the accompanying drawings. In the following description, the same numbers refer to the same elements throughout the drawings, unless otherwise represented. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application.
[0028] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used in this application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or," as used herein, refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0029] It is to be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It is to be understood that the term "and / or" as used herein encompasses all possible combinations of one or more of the associated listed items and can be abbreviated as "or". It is further to be understood that the use of "approximately," "substantially," or "about” in describing the contents of items herein can include amounts that are less than or greater than the stated amount due to testing, testing
[0030] Referring now to the drawings Figure 1 The application provides a lake dissolved organic matter remote sensing inversion method based on interpretable ensemble learning, comprising the following steps:
[0031] S100: Obtain lake dissolved organic matter abundance data and remote sensing image data, and pre-process the image data; wherein the pre-processing method includes atmospheric correction and masking.
[0032] In this embodiment, the dissolved organic matter (DOM) abundance data of the lakes in the study area is collected, including the DOC concentration and the absorption coefficients of a CDOM (254), a CDOM (350), and a CDOM (440) representing the humus, lignin and humic acid concentrations, respectively; the sampling time is the dry season and the rainy season. Sentinel-2A satellite images are used as the remote sensing data source, and B2-B8 bands are selected and the ratio features are calculated. The image time resolution is 10 days, and the spatial resolution is 10 meters; Sen2Cor tool is used for atmospheric correction to remove the interference of atmospheric components on image spectral reflection. Based on the normalized difference water index (MNDWI), the lake area is extracted, and the cloud, shadow and non-water features are removed. The MNDWI calculation result greater than 0 is the lake water, and less than or equal to 0 is the non-vegetation, land and other non-water area, and the formula is as follows:
[0033]
[0034] In the formula, MNDWI is the improved normalized difference water index, ρSWIR represents the short-wave infrared reflectivity, ρGreen represents the green band reflectivity.
[0035] S200: Divide the training set and test set according to the image data, and construct the inversion model.
[0036] In this embodiment, the training set and test set are divided according to the image data, wherein the training set and test set are divided: the DOM abundance data and the processed remote sensing image feature data are paired to form a complete data set; and the training set and the test set are randomly divided according to the proportions of 70% and 30%, wherein the training set is used for constructing the inversion model, and the test set is used for verifying and evaluating the inversion model. The inversion model is constructed based on XGBoost, CatBoost and NGBoost algorithms, and the model features are selected: according to the optical properties of dissolved organic matter (DOC, a CDOM (254), a CDOM (350), a CDOM (440), the representative image bands and their ratio features are extracted as model input features. Through correlation analysis, the band features highly correlated with the DOM concentration are selected to optimize the model input variables.
[0037] S300: Optimize the inversion model, and select the inversion model with the highest precision.
[0038] In this embodiment, the Bayesian optimization algorithm is used to optimize the model; the process of optimizing the inversion model is: the hyperparameters of the XGBoost, CatBoost and NGBoost models are adjusted using the Bayesian optimization algorithm; wherein the hyperparameters include learning rate, maximum depth and number of weak learners, the Bayesian optimization iteratively selects the optimal parameter combination in the parameter space by constructing a probability model to maximize the performance of the model. In the optimized model, the coefficient of determination (R 2 ) and the root mean square error (RMSE) are used as evaluation indexes. R 2 represents the goodness of fit of the model, and the closer to 1 indicates that the fitting effect of the model is better. RMSE reflects the difference between the predicted value and the actual value of the model, and the smaller the value indicates that the prediction accuracy is higher. For the inversion model of each dissolved organic matter, the optimization results of the three algorithms (XGBoost, CatBoost, NGBoost) are compared and analyzed on the test set, and the model with the highest R 2 and the lowest RMSE is selected as the final inversion model.
[0039] S400: Interpret the inversion model and analyze the calculation process of the inversion model.
[0040] In this embodiment, the importance of the input features of the inversion model is evaluated using the BorutaShap algorithm. BorutaShap combines the Boruta algorithm and SHAP (Shapley Additive Explanations) values, and can comprehensively analyze the contribution of remote sensing image bands and their ratio features to the prediction results of the inversion model. By comparing the real features of the inversion model with randomly generated shadow features, it is determined whether each input feature has a significant contribution to the prediction of the inversion model. Combined with the importance results of each feature, the key features of the inversion model are identified, for example, it is found that the blue band (B2) has a greater impact on the DOC concentration inversion model, which is a key feature of the inversion model. The BorutaShap analysis results are visualized to intuitively display the calculation process and feature influence of the inversion model. Through the feature importance chart, the importance of each feature to the prediction of the inversion model is displayed, and the influence of key bands and their ratios is highlighted. Through the explanatory analysis of BorutaShap, the decision logic behind the prediction results of the inversion model is further revealed, making the use of the inversion model more transparent and credible, and providing clear basis for scientists and decision makers.
[0041] S500: Based on the inversion model, the spatial distribution of lake dissolved organic matter abundance and the analysis of change trend are calculated, and the results are output.
[0042] In this embodiment, calculating the spatial distribution of lake dissolved organic matter abundance includes using the final optimized inversion model to perform pixel-level prediction on high-resolution remote sensing images in the study area to generate a spatial distribution map of dissolved organic matter abundance. The analysis of change trend is achieved by comparing the spatial distribution maps of dissolved organic matter abundance generated in the dry season and the rainy season to reveal the spatiotemporal dynamic change law of lake dissolved organic matter.
[0043] In the present application, first, lake dissolved organic matter abundance data and remote sensing image data are obtained, and the image data is preprocessed; the preprocessing methods include atmospheric correction and masking; then, the training set and test set are divided according to the image data, and the inversion model is constructed; the inversion model is optimized, and the inversion model with the highest accuracy is selected; the inversion model is then explained to analyze the calculation process of the inversion model; finally, based on the inversion model, the spatial distribution of lake dissolved organic matter abundance and the analysis of change trend are calculated, and the results are output; by combining field measurement data with high spatial resolution remote sensing images, the integrated learning technology is used to estimate the lake DOM abundance, and the model mechanism is explored based on the explainability technology, which improves the inversion accuracy and enhances the transparency of the model, and helps to better carry out the lake DOM inversion work.
[0044] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the application embrace any and all variations of the present application that fall within the scope of the general inventive concept as defined in the claims and that include alterations, modifications and improvements made to the application as disclosed herein.
[0045] It is to be understood that the application is not limited to the precise construction hereinafter described and as shown in the attached drawings, and that various changes in form and detail can be made to the application without departing from the scope thereof.
Claims
1. A method for remote sensing inversion of lake dissolved organic matter based on interpretable ensemble learning, characterized in that, The method comprises the following steps: Obtain lake dissolved organic matter abundance data and remote sensing image data, and preprocess the image data; The preprocessing method includes atmospheric correction and masking; Divide the training set and test set according to the image data, and build an inversion model; Optimize the inversion model, and select the inversion model with the highest accuracy; Interpret the inversion model and analyze the calculation process of the inversion model; Based on the inversion model, calculate the spatial distribution of lake dissolved organic matter abundance and analyze the change trend, and output the results.
2. The lake dissolved organic matter remote sensing retrieval method based on interpretable ensemble learning according to claim 1, characterized in that, In the step of obtaining lake dissolved organic matter abundance data and remote sensing image data, and preprocessing the image data; The preprocessing method includes atmospheric correction and masking; The remote sensing image data includes high spatial resolution images from the Sentinel-2 satellite; the image data needs to be corrected by atmospheric correction.
3. The lake dissolved organic matter remote sensing retrieval method based on interpretable ensemble learning according to claim 1, characterized in that, In the step of obtaining lake dissolved organic matter abundance data and remote sensing image data, and preprocessing the image data; The preprocessing method includes atmospheric correction and masking; The masking process identifies the lake boundary range based on the normalized difference water index.
4. The lake dissolved organic matter remote sensing retrieval method based on interpretable ensemble learning according to claim 1, characterized in that, In the step of optimizing the inversion model and selecting the inversion model with the highest accuracy: The process of optimizing the inversion model is: using the Bayesian optimization algorithm to adjust the hyperparameters of the XGBoost, CatBoost and NGBoost models; the hyperparameters include learning rate, maximum depth and weak learner number.
5. The lake dissolved organic matter remote sensing retrieval method based on interpretable ensemble learning according to claim 1, characterized in that, In the step of interpreting the inversion model and analyzing the calculation process of the inversion model: Use the BorutaShap algorithm to combine Boruta feature selection and SHAP interpretability technology to evaluate the importance of the input features of the inversion model, and identify the key features that contribute most to the inversion of lake dissolved organic matter.
6. The lake dissolved organic matter remote sensing retrieval method based on interpretable ensemble learning according to claim 1, characterized in that, In the step of calculating the spatial distribution of lake dissolved organic matter abundance based on the inversion model and analyzing the change trend, and outputting the results: The calculation of the spatial distribution of lake dissolved organic matter abundance includes: using the finally optimized inversion model to perform pixel-level prediction on high-resolution remote sensing images in the study area, and generating a spatial distribution map of dissolved organic matter abundance.
7. The lake dissolved organic matter remote sensing retrieval method based on interpretable ensemble learning according to claim 1, characterized in that, In the step of calculating the spatial distribution of lake dissolved organic matter abundance based on the inversion model and analyzing the change trend, and outputting the results: The analysis of the change trend is achieved by: comparing the dissolved organic matter abundance spatial distribution maps generated in the dry season and the rainy season, and revealing the spatiotemporal dynamic change law of lake dissolved organic matter.