A multi-parameter water quality remote sensing inversion method based on dynamic weighted ensemble learning

By using a dynamic weighted ensemble learning method, multi-source composite features are constructed and dynamic weight allocation is implemented, which solves the problems of low accuracy and poor stability in water quality remote sensing inversion in existing technologies. It realizes the synchronous and efficient inversion of multi-parameter water quality parameters and is suitable for remote sensing monitoring of water bodies such as rivers, lakes and reservoirs.

CN122470930APending Publication Date: 2026-07-28北京首创大气环境科技股份有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610618233.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing water quality remote sensing inversion technologies suffer from low inversion accuracy, poor stability, and low efficiency in complex water environments, making it difficult to meet the needs of simultaneous monitoring of multiple parameters. Furthermore, a single model cannot adapt to the differences in water quality parameters and complex water scenarios.

Method used

A multi-parameter water quality remote sensing inversion method based on dynamic weighted ensemble learning is adopted. By constructing multi-source composite features, implementing a feature selection mechanism that integrates model importance, linear correlation and mutual information, and combining multiple machine learning models, dynamic weight allocation is achieved to simultaneously invert water quality parameters such as turbidity, permanganate index, dissolved oxygen, total nitrogen and total phosphorus.

Benefits of technology

It enables simultaneous and high-precision inversion of multi-parameter water quality parameters in complex water environments, improves the stability and generalization ability of the inversion results, reduces computational costs and result conflicts, and is applicable to remote sensing monitoring of various water body types such as rivers, lakes and reservoirs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470930A_ABST
    Figure CN122470930A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-parameter water quality remote sensing inversion methods based on dynamic weighted ensemble learning, by spatial and temporal matching and preprocessing to multispectral remote sensing image and water quality monitoring data, construct basic dataset;Further fusion spectrum interaction feature, remote sensing index feature, time period feature and spatial statistical feature, form multi-source composite feature engineering, and obtain optimal feature subset by the feature screening mechanism of model importance, correlation and mutual information combination;On this basis, multi-model pool is constructed, and integrated learning is carried out according to model performance adaptive weight distribution, realize the synchronous inversion of turbidity, permanganate index, dissolved oxygen, total nitrogen and total phosphorus and other various water quality parameters.The present application can improve the stability and generalization ability of inversion result, and is suitable for remote sensing water quality monitoring application of multiple types of water body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing water environment monitoring and intelligent inversion technology, and in particular to a multi-parameter water quality remote sensing inversion method based on dynamic weighted ensemble learning, which is applicable to the regional and continuous monitoring of water quality parameters such as turbidity (TUR), permanganate index (CODMn), dissolved oxygen (DO), total nitrogen (TN), and total phosphorus (TP) in water bodies such as rivers, lakes, and reservoirs. Background Technology

[0002] Water quality monitoring is a crucial foundation for aquatic ecological environment management. Traditional methods relying on manual sampling and fixed monitoring sections suffer from insufficient spatial coverage, limited representativeness, and an inability to reflect the overall distribution characteristics of water bodies, making it difficult to meet the needs of refined monitoring at the basin and regional levels.

[0003] Remote sensing inversion of water quality parameters is a key technology that utilizes the spectral characteristics of satellite or aerial remote sensing imagery, combined with measured water quality data, to construct models for large-scale water quality monitoring. It has significant application value in areas such as river and lake ecological environment management and water resource protection. Currently, mainstream remote sensing inversion methods for water quality mainly include empirical models, semi-analytical models, and machine learning models. With the development of machine learning technology, single machine learning models, such as backpropagation neural networks, long short-term memory networks (LSTM), and random forests, are widely used in water quality parameter inversion. For example, existing technologies include total phosphorus inversion methods based on Sentinel-2 satellite data combined with LSTM algorithms, or single water quality parameter monitoring methods based on multi-source remote sensing data. However, single machine learning models have obvious limitations: the spectral response characteristics of different water quality parameters vary significantly. For example, dissolved oxygen (DO) is a non-optically active parameter with weak correlation of its spectral signals, while turbidity (TUR) is an optically active parameter with strong spectral sensitivity. A single model is difficult to adapt to the differentiated inversion requirements of multiple parameters. At the same time, in complex water body scenarios (such as waters with high turbidity and high nutrient concentration), a single model is prone to overfitting or underfitting, resulting in insufficient inversion accuracy.

[0004] To address the limitations of single-model approaches, some existing technologies employ ensemble learning methods. However, these methods generally use fixed weights, failing to adaptively adjust to the performance differences of different models under varying parameters, time periods, or water body conditions, thus limiting the effectiveness of the ensemble. Furthermore, most methods still rely on an independent inversion strategy of "single parameter—single model," resulting in low computational efficiency and poor consistency among parameters. In addition, existing water quality remote sensing inversion technologies suffer from a lack of diverse feature engineering construction methods: most methods utilize only raw multispectral bands or a limited number of water color indices, failing to systematically incorporate temporal periodicity and spatial statistical features, making it difficult to characterize the spatiotemporal heterogeneity of complex water bodies. Moreover, feature selection lacks a scientific and systematic mechanism, failing to perform multi-dimensional screening of high-dimensional features, leading to high feature redundancy and insufficient model generalization ability and stability. These shortcomings collectively make it difficult for existing technologies to achieve simultaneous, high-precision inversion of multiple water quality parameters, thus failing to meet the needs of refined monitoring in complex water environments.

[0005] Therefore, there is an urgent need for a water quality remote sensing inversion technology that can construct multi-source composite features, implement multi-index fusion feature screening, and achieve multi-parameter synchronous inversion through dynamic weighted ensemble learning, in order to solve the problems of low accuracy, poor stability, and low efficiency in water quality parameter inversion in complex water environments. Summary of the Invention

[0006] The purpose of this invention is to address the technical shortcomings of existing technologies, such as limited generalization ability of single models, insufficient feature utilization, poor adaptability of fixed-weight integration, and low efficiency and poor consistency of multi-parameter inversion. This invention provides a multi-parameter water quality remote sensing inversion method based on dynamic weighted ensemble learning. By constructing a multi-source composite feature engineering that integrates spectral, remote sensing index, temporal, and spatial data, and implementing a feature selection mechanism that integrates model importance, linear correlation, and mutual information, the method adaptively allocates weights based on the model's real-time performance. Ultimately, it achieves simultaneous and high-precision inversion of water quality parameters such as turbidity (TUR), permanganate index (CODMn), dissolved oxygen (DO), total nitrogen (TN), and total phosphorus (TP), thereby improving the stability and generalization ability of inversion results in complex water body scenarios.

[0007] To achieve the above objectives, this invention provides a multi-parameter water quality remote sensing inversion method based on dynamic weighted ensemble learning, comprising the following steps: Step S1: Data acquisition and preprocessing; Acquire multispectral reflectance data from Sentinel-2 satellite or other satellites and measured data from online water quality monitoring points. Perform data registration according to the spatiotemporal matching principle (the difference between satellite transit time and water quality monitoring time ≤ 3 hours). Filter valid data and remove outliers using methods such as box plots to construct a basic dataset. The basic dataset includes reflectance data from multispectral images in bands B1, B2, B3, B4, B5, B6, B7, B8, B8A, B11, and B12, as well as measured data from online water quality monitoring points for parameters such as turbidity (TUR), permanganate index (CODMn), dissolved oxygen (DO), total nitrogen (TN), and total phosphorus (TP).

[0008] In addition to Sentinel-2 satellite, satellite remote sensing images with multispectral observation capabilities, such as Landsat-8, Landsat-9, and Gaofen-6, can also be used as input data, as long as they contain water quality sensitive bands such as blue, green, red, red-edge, or shortwave infrared. When using different satellite data, the spectral interaction features or spectral combination parameters can be adaptively adjusted according to the band configuration and center wavelength relationship of the corresponding sensor, while the multi-source composite feature engineering construction process and dynamic weighted ensemble learning method remain unchanged.

[0009] Step S2: Construction of multi-source composite feature engineering; Based on preprocessed multispectral remote sensing images, a multi-level composite feature system is automatically constructed within the water body area, fusing spectral interactive features, remote sensing index features, temporal periodic features, and spatial statistical features. Specifically, this includes: (1) Construction of spectral interaction features; Based on the band reflectance in the base dataset, various pairwise band interaction features are automatically constructed, including: Band ratio characteristics R i / j and R j / i By constructing the ratio relationship between any two spectral bands, the responsiveness of different water quality parameters to relative spectral changes is enhanced. The calculation method is as follows: ; Band difference and normalized difference characteristics R ij and ND_R ij Used to mitigate the effects of changes in light intensity and highlight spectral differences in water bodies, its calculation method is as follows: ; Log-ratio characteristics (LR) ij The difference in weak signals is enhanced by logarithmic transformation, calculated as follows: ; First-order differential characteristic of the spectrum SD_R i+1When there are at least three available bands, calculate the first-order spectral differential characteristics of adjacent bands based on wavelength order. Based on the center wavelengths of each Sentinel-2 band {'B2': 490, 'B3': 560, 'B4': 665, 'B5': 705, 'B6': 740, 'B7': 783, 'B8': 842, 'B8A': 865, 'B11': 1610, 'B12': 2190}, construct the first-order spectral differential characteristics for adjacent bands. The calculation method is as follows: ; Band product feature (PR): Further constructing the band product and harmonic mean features to characterize the nonlinear coupling relationship between spectra. Band product feature calculation formula: ; Harmonic mean characteristics HR ij Calculation formula: ; In the formula, Bi and Bj represent reflectance at different wavelengths, ε is a small constant to prevent the denominator from being zero, and λi represents the center wavelength of the corresponding wavelength band. These spectral interaction features are automatically generated, significantly improving the model's sensitivity to water quality parameters such as turbidity and chemical oxygen demand.

[0010] (2) Construction of remote sensing index features; Based on green, blue, red-edge, and shortwave infrared bands, key remote sensing indices such as EVI (Enhanced Vegetation Index), SAVI (Soil-Adjusted Vegetation Index), and MSAVI (Modified Soil-Adjusted Vegetation Index) are automatically calculated to characterize water transparency, algal abundance, organic matter absorption characteristics, and water structure. The calculation formulas for each index are as follows: ; In the formula, B4 and B8 are the reflectance of the 4th and 8th bands, respectively. The above indices can be flexibly expanded or combined according to the inversion requirements of target water quality parameters (such as turbidity and permanganate index) to form a targeted subset of index features.

[0011] (3) Construction of time period features Based on the temporal information obtained from remote sensing imagery, temporal correlation features are constructed, and a periodic coding method is introduced to form a time series feature set, including: Basic time features: Generate day_of_year (day number within the year), month (month), and season (season); Sine and cosine periodic characteristics: The periodic time information is converted into continuous numerical characteristics using the following formula.

[0012] Monthly cycle characteristics: ; Annual daily cycle characteristics (daily data scenario): ; Among them, month is the month corresponding to the remote sensing image (value 1~12), and day_of_year is the day of the year corresponding to the remote sensing image (value 1~365); this type of feature can accurately depict the annual / seasonal cycle of water pollution and algal changes (such as the spectral characteristics of summer algal blooms).

[0013] (4) Construction of spatial statistical features Using monitoring sections or preset grid units as statistical units, the local mean, standard deviation, and extreme values ​​of the main spectral bands are calculated, and a multi-band reflectance standard deviation is constructed as a spectral variability feature. The spatial statistical feature construction of this invention does not depend on specific latitude and longitude coordinates; it expresses spatial structural features solely through the differences in spectral distribution at monitoring sections, avoiding the regional adaptability problem caused by coordinate dependence. Specifically, it includes two parts: Grouped statistical features based on monitoring sections: Using "section name" as the grouping unit, long-term statistical features are calculated for the core spectral bands (B2, B3, B4, B5). The specific statistical rules are as follows: For the B4 band, the mean (station_B4_mean), standard deviation (station_B4_std), minimum value (station_B4_min), and maximum value (station_B4_max) are calculated; for the B2, B3, and B5 bands, the mean (station_B2_mean / station_B3_mean / station_B5_mean) and standard deviation (station_B2_std / station_B3_std / station_B5_std) are calculated respectively. All generated features are associated with the original data through "section name" to achieve spatiotemporal aggregation of spectral features under the same section.

[0014] Spectral spatial variability characteristics: If the data simultaneously contains bands B2, B3, B4, and B5, calculate the row dimension standard deviation (spatial_variability) of the reflectance of these four bands in a single sample. The calculation formula is as follows: ; Where n=4 (corresponding to bands B2 / B3 / B4 / B5), and xi is the reflectance value of the i-th band of a single sample. This is the mean of the reflectance of the four bands of the sample; this feature can quantify the dispersion of multi-band reflectance within a single sample and characterize the spectral heterogeneity of the local space.

[0015] Step S3: Multi-indicator fusion feature screening; To address the limitations of single-index selection in the high-dimensional feature set constructed in step S2, this invention proposes a multi-index fusion feature comprehensive evaluation mechanism. This mechanism comprehensively considers the following three types of indicators: (1) Model Importance Index (I) model ): Based on the feature importance of random forest, it captures the non-linear dependency between features and the target; (2) Pearson correlation coefficient (Corr): measures the linear correlation between features and targets, and quickly eliminates irrelevant features; (3) Mutual Information Index (MI): assesses the nonlinear statistical dependence between features and targets, and is applicable to complex water scenarios.

[0016] The three types of indicators are normalized to the [0,1] interval, and then weighted and fused according to adjustable weight coefficients (α,β,γ, satisfying α+β+γ=1) to obtain the comprehensive score of each feature: ; By setting a scoring threshold or selecting TopK features (usually 15-30), an optimal feature subset corresponding to each water quality parameter is formed. This mechanism retains strong linear correlation features while fully exploring nonlinear coupling information, thereby significantly reducing the risk of overfitting while ensuring model accuracy.

[0017] It should be noted that during the multi-indicator fusion feature selection process, the weight coefficients of the model importance index, Pearson correlation coefficient, and mutual information index can be adjusted appropriately according to different water body types, water quality parameter characteristics, or data scale. For example, in scenarios with significant nonlinear relationships or highly turbid water bodies, the weight of the mutual information index can be appropriately increased to enhance the ability to select complex feature relationships without affecting the implementation of the overall feature selection framework of this invention.

[0018] Step S4: Multi-model training; This step aims to build a high-performance base model pool for subsequent dynamic weighted ensemble learning. By training multiple models on the optimal feature subset, diverse inversion patterns are captured for different water quality parameters. The specific process is as follows: (1) Training data preparation and partitioning; Data standardization: RobustScaler is used to standardize the optimal subset of input features to eliminate differences in the units of measurement between features and to enhance the model’s robustness to outliers in the data.

[0019] Dataset partitioning: To ensure the objectivity of model evaluation, the standardized data was strictly divided into training and validation sets in a 7:3 ratio. The training set was used to build the model, while the validation set was used for subsequent model performance evaluation and dynamic weight calculation.

[0020] (2) Model pool construction; Construct a model pool that integrates multiple machine learning methods to cover different inversion mechanisms, including linear, nonlinear, and complex decision boundaries. The core model pool should include at least: Random Forest (RF): Based on ensemble decision trees, it excels at capturing nonlinear relationships and interaction effects.

[0021] Gradient Boosting Decision Tree (GBDT) and Lightweight Gradient Boosting Machine (LightGBM): Efficiently fit complex patterns using the gradient boosting framework.

[0022] Extreme Gradient Boosting Tree (XGBoost): Introduces regularization into the boosting tree framework to optimize generalization performance.

[0023] Ridge regression: A linear model that provides stable baseline predictions and avoids overfitting.

[0024] The model pool can be expanded to include other machine learning models such as Support Vector Regression (SVR), Neural Networks (NN), and Long Short-Term Memory Networks (LSTM) as needed. Expanding the model pool does not affect the weight calculation logic of the dynamic weighted ensemble; each model can participate in weight allocation based on its performance metrics on the validation dataset.

[0025] (3) Independent training Using the aforementioned training set as input, each individual model in the model pool is trained synchronously and independently; the training process is carried out separately for each water quality parameter to ensure that each individual model can learn the optimal mapping relationship between specific parameters and remote sensing features.

[0026] (4) Preliminary performance evaluation After training, the initial performance of each individual model is evaluated on the validation set, with the core evaluation metric being the coefficient of determination (R²). The R² value generated in this step will serve as the direct basis for calculating the weights of each model in the next stage of dynamic weighted ensemble.

[0027] By building and training a diverse pool of models, this step lays a solid foundation for subsequent ensemble learning, avoiding the structural biases of a single model and providing a wealth of performance comparisons for dynamic weight allocation.

[0028] Step S5: Construction of Dynamically Weighted Ensemble Learning Model To overcome the problem of poor adaptability of static weight integration to complex water body scenarios, this invention proposes a dynamic weight allocation strategy based on performance feedback, the specific process of which is as follows: (1) Model performance evaluation: Calculate the coefficient of determination R² and root mean square error (RMSE) of each base model (such as RF, XGBoost, LightGBM, etc.) on the independent validation set.

[0029] (2) Adaptive weight calculation: Based primarily on R², if a model's R² ≤ 0, its weights are reset to zero (i.e., it does not participate in the integration). The weights of the remaining models are allocated according to their R² proportions. The calculation formula is as follows: ; Where wi is the dynamic weight of the i-th sub-model, Ri is the validation set determination coefficient of the corresponding model, and n is the number of sub-models in the model pool.

[0030] (3) Weighted fusion and backoff mechanism: The prediction results of each model are weighted and summed according to the above weights to obtain the integrated prediction value: ; Where Yi is the prediction result of the i-th sub-model. If R² ≤ 0 for all models (extreme case), a backoff mechanism is triggered, and the prediction result is output using a simple averaging method to ensure the robustness of the system.

[0031] This dynamic weighting mechanism can automatically enhance the contribution of high-precision models and suppress the influence of inefficient models, thereby maintaining stable inversion performance under different seasons, water types and pollution levels.

[0032] Step S6: Multi-parameter synchronous inversion output Using the aforementioned dynamic weighted ensemble learning model, the spatial distribution inversion results of turbidity (TUR), permanganate index (CODMn), dissolved oxygen (DO), total nitrogen (TN), and total phosphorus (TP) are simultaneously output to generate a thematic map of water quality across the entire region.

[0033] Based on the above technical solution, the advantages of the present invention are: (1) Combining multi-source composite features with scientific screening mechanisms: By integrating four types of features—spectral interaction, remote sensing index, time period, and spatial statistics—a multi-level composite feature system is constructed to comprehensively characterize the spatiotemporal heterogeneity and spectral response of water bodies. Then, through a screening mechanism that weights and integrates model importance, linear correlation, and mutual information, redundant features are accurately eliminated, key information is retained, the generalization ability of the model is significantly improved, and the problem of insufficient utilization and high redundancy of existing technical features is effectively solved.

[0034] (2) Dynamic weighting improves inversion stability and accuracy: Based on the real-time inversion accuracy (determination coefficient R²) of the model, the weights are adaptively allocated to automatically improve the contribution of high-precision models and eliminate low-performance models, avoiding the limitations of static weights. It can adapt to the changes in spectral characteristics under different water body types and seasonal conditions, and solve the problem of weak generalization ability of single models or static integration.

[0035] (3) Multi-parameter synchronous inversion optimizes efficiency and consistency: Breaking through the isolated inversion mode of "single parameter - single model", multi-parameter synchronous prediction is achieved through a unified feature system and integrated model, which greatly reduces the repetitive modeling process and reduces the computational cost; at the same time, it avoids the result conflict caused by the difference between different models and improves the internal consistency of the inversion results of each parameter.

[0036] (4) Wide range of applications and strong practicality: Spatial statistical feature construction does not depend on specific latitude and longitude coordinates and is applicable to various water body types such as rivers, lakes and reservoirs; the technical process is highly automated and the model pool and feature set can be flexibly expanded according to data sources and monitoring needs, and can directly serve the operational needs of watershed water environment supervision. Attached Figure Description

[0037] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating the overall technical process of the multi-parameter water quality remote sensing inversion method of the present invention. Figure 2 This is a scatter plot showing the accuracy verification of the integrated model for various water quality parameters in this embodiment of the invention. Figure 3 This is a schematic diagram of typical monitoring results of the present invention; Figure 4 This is a table of the first 25 key feature combinations in the embodiments of the present invention; Figure 5 This is a summary of the dynamic weights of each water quality parameter sub-model in the embodiments of the present invention. Detailed Implementation

[0038] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0039] This embodiment takes the rivers and lakes of Huainan City as the research object. Based on Sentinel-2 multispectral remote sensing images and online water quality monitoring data, it employs the multi-parameter water quality remote sensing inversion method based on dynamic weighted ensemble learning proposed in this invention to achieve simultaneous inversion of five water quality parameters: turbidity (TUR), permanganate index (CODMn), dissolved oxygen (DO), total nitrogen (TN), and total phosphorus (TP). Figure 1As shown, the specific steps are as follows: Step S1: Data Acquisition and Preprocessing 1.1 Remote sensing data screening and preprocessing: Sentinel-2 surface reflectance product (S2_SR) images with cloud cover less than 30% and spatial coverage greater than 70% during the period from January 2021 to August 2025 were selected as remote sensing data sources; combined with Sentinel-2 cloud probability data, a threshold method was used to remove clouds, cloud shadows and invalid pixels from the images to ensure image quality.

[0040] 1.2 Water Body Extraction: A modified Normalized Difference Water Index (MNDWI) was used to set a threshold for water body extraction in the study area. The formula is as follows: MNDWI=(B3-B11) / (B3+B11); B3 is the green band (center wavelength 560nm), and B11 is the shortwave infrared band (center wavelength 1610nm). The water body mask of the study area was obtained by threshold segmentation, and the water body dataset of the study area was obtained by cropping, which includes reflectance data of 13 bands: B1, B2, B3, B4, B5, B6, B7, B8, B8A, B11, and B12.

[0041] 1.3 Measured Data Matching and Outlier Removal: Measured data (including five core indicators: TUR, CODMn, DO, TN, and TP) from online water quality monitoring points in the Huainan City river and lake areas were collected simultaneously. Outliers in the measured data were removed using a box plot method (interquartile range IQR = 1.5). Data registration was performed according to the spatiotemporal matching principle of "satellite transit time and monitoring time difference ≤ 3 hours" as specified in the invention. Finally, a basic dataset was constructed, with the following effective sample numbers for each parameter: TUR 1922, CODMn 1829, DO 1986, TP 1881, and TN 1869.

[0042] Step S2: Construct a multi-source composite feature set Based on the basic dataset, and following the requirements of the four-layer feature system of "spectral-index-time-space" in the invention, the four types of features are calculated sequentially, ultimately forming a high-dimensional feature set of 300+ dimensions, as detailed below: Spectral interaction features: including nonlinear interaction features such as band ratio, band difference, normalized difference, logarithmic ratio, first-order spectral derivative, band product, harmonic mean, etc., comprehensively characterizing the linear / nonlinear relationship between water quality parameters and spectral signals; Remote sensing index characteristics: Calculate vegetation-water color correlation indices such as EVI (Enhanced Vegetation Index), SAVI (Soil-Adjusted Vegetation Index), and MSAVI (Modified Soil-Adjusted Vegetation Index) to characterize water transparency, algal abundance, and organic matter absorption characteristics. Time cycle features: Generate basic time features such as month, season, and day of year, as well as periodic continuous features such as month_sin, month_cos, day_sin, and day_cos obtained through sine and cosine encoding, to capture the annual / seasonal cycle patterns of water pollution. Spatial statistical characteristics: First, long-term statistical characteristics such as mean, standard deviation, and extreme values ​​are calculated for the core bands B2, B3, B4, and B5, using "monitoring section name" as the grouping unit; Second, the row dimension standard deviation of reflectance of bands B2, B3, B4, and B5 within a single sample (spectral spatial variability characteristics) is used to quantify local spatial spectral heterogeneity.

[0043] Step S3: Multi-indicator fusion feature screening This step strictly follows the "three-indicator fusion feature comprehensive evaluation mechanism" described in the invention, selecting the optimal feature subset based on the characteristics of different water quality parameters. The specific operation is as follows: 3.1 Screening Indicators and Weighting: Three types of indicators were used: model importance (Imodel), Pearson correlation coefficient (Corr), and mutual information (MI). A comprehensive feature score was calculated with weights of α=0.4, β=0.3, and γ=0.3 (this weighting balances the contributions of linear and nonlinear features, ensuring screening accuracy). Comprehensive score = α × I model +β×Corr+γ×MI (All three indices are first normalized to the [0,1] interval).

[0044] 3.2 Feature Selection Results: Based on the comprehensive score ranking, the top 25 key features were automatically selected, and optimal feature subsets for each water quality parameter were constructed. Specific feature combinations are as follows: Figure 4 As shown in the table below.

[0045] Step S4: Multi-model training According to the requirements of this invention, a diverse model pool is constructed, data standardization and dataset partitioning are completed, and independent training of multiple models is achieved, as detailed below: 4.1 Data Preprocessing: The RobustScaler method described in this invention is used to standardize the optimal feature subset of each parameter to eliminate dimensional differences and enhance the robustness of the model to outliers; the standardized data is divided into a training set (for model building) and a validation set (for performance evaluation and dynamic weight calculation) in a 7:3 ratio to ensure that the partitioning process is random and non-overlapping.

[0046] 4.2 Model Pool Construction and Hyperparameter Optimization: Five types of model pools were constructed, including Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Ridge Regression, Extreme Gradient Boosting Tree (XGBoost), and Lightweight Gradient Boosting Machine (LightGBM), covering various inversion mechanisms such as linear, nonlinear, and gradient boosting. After optimization using 5-fold cross-validation combined with grid search, the hyperparameters of each model were set as follows: RF: Number of decision trees: 300; Maximum depth: 15. GBDT: Learning rate 0.1, number of trees 300; Ridge: Regularization parameter 1.0; XGBoost: Learning rate 0.05, number of trees 300; LightGBM: Learning rate 0.05, number of trees 300.

[0047] 4.3 Model Training and Preliminary Evaluation: Using the training set of each parameter as input, each sub-model in the model pool is trained synchronously and independently. After training, the coefficient of determination R² (core evaluation index) of each sub-model is calculated on the validation set, which serves as the direct basis for subsequent dynamic weight calculation.

[0048] In this embodiment, based on the above model, the accuracy verification of the integrated model for each water quality parameter is as follows: Figure 2 The scatter plot is shown below.

[0049] Step S5: Dynamic Weighted Integration (1) Dynamic weight calculation On the validation set, the dynamic weights of each sub-model are calculated based on its coefficient of determination R². The weight calculation formula is as follows: ; Where: wi is the dynamic weight of the i-th sub-model; Ri is the coefficient of determination of the model on the validation set; n is the number of sub-models in the model pool. The dynamic weights of each water quality parameter sub-model are summarized as follows: Figure 5 As shown in the table below.

[0050] (2) Weighted integration and fusion The ensemble prediction results are obtained through a weighted summation method: ; Where Yi is the prediction result of the i-th sub-model; if all sub-models R²≤0 (extreme case), the backoff mechanism is triggered, and the simple averaging method is used to output the result to ensure the robustness of the inversion system.

[0051] (3) Model performance comparison The test set performance comparison between ensemble models and single optimal machine learning models is shown in the table below:

[0052] Step S6: Output of multi-parameter inversion results Based on the trained dynamic weighted ensemble model, Sentinel-2 rasterized inversion results for TUR, CODMn, TP, TN, and DO are generated simultaneously, and a multi-parameter water quality thematic map of the Huainan City river and lake area is output. The monitoring results are illustrated in the figure below. Figure 3 As shown, it can intuitively present the spatial distribution pattern of various water quality parameters and pollution hotspots.

[0053] Through modeling and verification using multi-temporal Sentinel-2 remote sensing imagery and synchronous water quality monitoring data, this embodiment fully verifies the technical advantages of the method of the present invention, as detailed below: (1) The inversion accuracy has been significantly improved: Through a multi-model dynamic weighted ensemble mechanism, under the data conditions used in this embodiment, the determination coefficient R² of the method of this invention is stably distributed in the range of 0.58–0.73 on the validation dataset, with R² for turbidity (TUR) and total nitrogen (TN) both reaching above 0.72. Compared with the single machine learning model used in the embodiment, the ensemble model has higher overall inversion accuracy and smaller result fluctuations, which to some extent alleviates the instability problem of accuracy caused by structural differences in single models, thereby achieving robust inversion of different water quality parameters.

[0054] (2) High efficiency of multi-parameter synchronous inversion: By using a unified feature engineering construction process and integrated model framework, the process of repeatedly extracting features, training models and adjusting parameters for different water quality parameters is avoided. Compared with the traditional single-parameter independent modeling method, the overall computational complexity and redundant computational overhead are significantly reduced, and the overall processing efficiency of multi-parameter remote sensing inversion is improved.

[0055] (3) Improved spatial distribution consistency and stability: The unified feature system and dynamic weight fusion strategy enable the inversion results of various water quality parameters to exhibit greater consistency and continuity in spatial distribution, effectively reducing spatial noise and abnormal patch phenomena caused by model differences, and improving the stability and interpretability of water quality thematic maps.

[0056] The multi-parameter water quality remote sensing inversion method of the present invention exhibits significant comprehensive technical advantages in terms of inversion accuracy, computational efficiency and result stability, and is applicable to remote sensing water quality parameter inversion and long-term dynamic monitoring scenarios for various types of water bodies.

[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.

Claims

1. A multi-parameter water quality remote sensing inversion method based on dynamic weighted ensemble learning, characterized in that: Includes the following steps: Step S1, Data Preprocessing: Acquire multispectral remote sensing image reflectance data and measured data from water quality monitoring points. Perform data registration and remove outliers according to the spatiotemporal matching of remote sensing image transit time and water quality monitoring time to construct a basic dataset. The multispectral remote sensing image reflectance data includes reflectance data of B1, B2, B3, B4, B5, B6, B7, B8, B8A, B11 and B12 bands of the multispectral image. The measured water quality data includes measured values ​​of turbidity, permanganate index, dissolved oxygen, total nitrogen and total phosphorus. Step S2, Multi-source composite feature engineering construction: Based on the basic dataset, a composite feature system is constructed, which includes: spectral interaction features constructed based on multispectral band reflectance, remote sensing index features calculated based on multispectral bands, time period features constructed based on remote sensing image acquisition time, and spatial statistical features constructed based on monitoring sections or spatial units. Step S3, Multi-index fusion feature screening: Calculate the model importance index I of the composite feature respectively. model Correlation index Corr and mutual information index MI are used to normalize each index and then score it according to preset weights. Based on the scoring results, 15 to 30 key features are selected to obtain the optimal feature subset corresponding to each water quality parameter. Step S4, Multi-model Training: Based on the optimal feature subset, RobustScaler is used for standardization, and the training and validation sets are divided in a 7:3 ratio; a model pool containing multiple machine learning models is constructed, and each individual model in the model pool is trained separately, using the coefficient of determination R0. 2 Evaluate the inversion accuracy of each model on the validation set; Step S5, Dynamic Weighted Integration Learning: Evaluate the inversion performance of each individual model on the validation set, and adaptively calculate the model weights according to the performance indicators that characterize the predictive ability of the models, where the performance indicators include the coefficient of determination R²; when the performance indicator corresponding to a model is lower than a preset threshold, the model does not participate in the weighted fusion, and the prediction results of the remaining models are weighted and fused according to the model weights to obtain the integrated prediction result. Step S6, Multi-parameter synchronous inversion: Based on the integrated prediction results, the remote sensing inversion results of turbidity, permanganate index, dissolved oxygen, total nitrogen and total phosphorus are output synchronously to generate a thematic map of water quality across the entire region.

2. The multi-parameter water quality remote sensing inversion method according to claim 1, characterized in that: The spectral interaction features include band ratio features, band difference features, normalized difference features, logarithmic ratio features, spectral first derivative features, band product features, and harmonic mean features.

3. The multi-parameter water quality remote sensing inversion method according to claim 1, characterized in that: The time cycle features include basic time features and periodic continuous numerical features; the basic time features are month, day of year, and season; the periodic continuous numerical features are month_sin, month_cos, day_sin, and day_cos obtained through sine and cosine encoding, corresponding to the monthly cycle and the daily cycle features within the year, respectively. The periodic time information is converted into periodic continuous numerical features using the following formula: Monthly cycle characteristics: ; Annual daily cycle characteristics: ; Where month is the month corresponding to the remote sensing image, with a value of 1 to 12; day_of_year is the day of the year corresponding to the remote sensing image, with a value of 1 to 365.

4. The multi-parameter water quality remote sensing inversion method according to claim 1, characterized in that: When constructing spatial statistical features, the monitoring section or preset grid unit is used as the statistical unit. The local mean, standard deviation, and extreme values ​​of the core spectral bands B2, B3, B4, and B5 are calculated, and the multi-band reflectance standard deviation is constructed as a spectral variability feature. The steps include: Grouped statistical characteristics based on monitoring sections: Long-term statistical characteristics of core spectral bands B2, B3, B4, and B5 are calculated using "section name" as the grouping unit. All generated features are associated with the original data through the "section name"; Spectral spatial variability characteristics: If the data simultaneously contains bands B2, B3, B4, and B5, calculate the row dimension standard deviation (spatial_variability) of the reflectance of these four bands in a single sample. The calculation formula is as follows: ; Where n=4, corresponding to bands B2 / B3 / B4 / B5, and xi is the reflectance value of the i-th band of a single sample. This is the average reflectance of the sample across its four bands.

5. The multi-parameter water quality remote sensing inversion method according to claim 1, characterized in that: In step S3, the comprehensive score for multi-indicator fusion feature screening is calculated according to the following formula: ; Where α, β, and γ are preset weighting coefficients and α + β + γ = 1, I model Corr and MI have all been normalized to the [0,1] interval.

6. The multi-parameter water quality remote sensing inversion method according to claim 1, characterized in that: The model pool includes Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Ridge Regression (Ridge), Extreme Gradient Boosting Tree (XGBoost), Lightweight Gradient Boosting Machine (LightGBM), Support Vector Regression (SVR), Neural Network (NN), and Long Short-Term Memory Network (LSTM).

7. The multi-parameter water quality remote sensing inversion method according to claim 1, characterized in that: In step S5, the dynamic weighted ensemble learning employs a dynamic weight allocation strategy based on performance feedback, and the steps include: (1) Model performance evaluation: Calculate the coefficient of determination R² and root mean square error RMSE of each base model on the independent validation set; (2) Adaptive weight calculation: Based on the coefficient of determination R², if the model R² ≤ 0, its weights are reset to zero, and the remaining model weights are allocated according to their R² proportions. The calculation formula is as follows: ; Where wi is the dynamic weight of the i-th sub-model, Ri is the validation set determination coefficient of the corresponding model, and n is the number of sub-models in the model pool; (3) Weighted fusion and backoff: The prediction results of each model are weighted and summed according to the model weights mentioned above to obtain the integrated prediction value: ; Where Yi is the prediction result of the i-th sub-model; if R²≤0 for all models, the backoff mechanism is triggered, and the prediction result is output using the simple averaging method.

8. The multi-parameter water quality remote sensing inversion method according to claim 1, characterized in that: The multispectral remote sensing image reflectance data uses Sentinel-2, Landsat-8, Landsat-9 or Gaofen-6 satellite remote sensing images as input data, and the input data includes blue light, green light, red light, red edge or shortwave infrared bands.

9. A water quality remote sensing inversion system for implementing the multi-parameter water quality remote sensing inversion method according to any one of claims 1 to 8, characterized in that: include: The data preprocessing module is used to perform spatiotemporal matching, outlier removal, and construction of a basic dataset for multispectral remote sensing image reflectance data and measured data from water quality monitoring points. The multi-source composite feature engineering module is used to construct a composite feature system that integrates spectral interaction features, remote sensing index features, temporal periodic features, and spatial statistical features based on the aforementioned basic dataset. The multi-indicator fusion feature selection module is used to comprehensively evaluate the model importance, relevance and mutual information of the composite features, and select the optimal feature subset. A multi-model training module is used to construct a model pool based on the optimal feature subset and to complete the training and performance evaluation of each individual model. The dynamic weighted ensemble module is used to adaptively allocate weights and perform ensemble predictions based on the performance metrics of each individual model on the validation dataset. The multi-parameter output module is used to simultaneously output remote sensing inversion results of multiple water quality parameters.