Lake and reservoir chl-a concentration multi-modal deep learning remote sensing inversion method considering time factors and environmental characteristics

Through the multimodal deep learning remote sensing inversion method, combined with Sentinel-2 satellite data and dynamic environmental factors, a four-modal deep neural network is designed to solve the problems of insufficient utilization of spectral characteristics and neglected time factors in complex lake and reservoir environments, and high-precision Chl-a concentration monitoring and water quality management support are achieved.

CN120561585APending Publication Date: 2025-08-29JIANGSU OCEAN UNIV +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510660231.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing technology has insufficient utilization of spectral characteristics and ignores the driving effect of time factors and environmental characteristics on algae growth in complex lake and reservoir environments, resulting in poor inversion accuracy and insufficient generalization capabilities of the model, which cannot support the needs of refined water quality management.

Method used

The multimodal deep learning remote sensing inversion method is adopted, combined with Sentinel-2 satellite data, and the high correlation band combination is screened through Pearson correlation analysis, time period characteristics and dynamic environmental factors are introduced, and a four-modal deep neural network (QM-DNN) is designed, including dual-band, three-band, four-band and auxiliary mode subnetwork. Residual connection, Dropout regularization and SmoothL1 loss function are used to optimize the model performance.

Benefits of technology

It significantly improves the adaptability and inversion accuracy of cross-season inversion, realizes high-precision Chl-a concentration monitoring, provides dynamic coverage and model interpretability across the region, and supports scientific basis for water quality management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561585A_ABST
    Figure CN120561585A_ABST
Patent Text Reader

Abstract

The invention discloses a lake and reservoir Chl-a concentration multi-modal deep learning remote sensing inversion method considering time factors and environmental characteristics, and the method comprises the following steps: obtaining and processing data, carrying out the atmospheric correction through Sen2Cor software, carrying out the image mosaic, splicing and resampling in SNAP software, converting coordinates in a spatial dimension, calculating the line and column numbers of pixels, and carrying out the remote sensing inversion of the lake and reservoir Chl-a concentration. The method comprises the following steps: extracting and standardizing a digital quantized value, systematically fusing multiband spectrum combination features, time period features and a dynamic environment factor K by adopting a Pearson analysis method, forming a 14-dimensional input system, completely capturing nonlinear correlation between Chl-a concentration and spectral response, time rhythm and environmental suitability, solving the problem of single feature dimension in the prior art, and improving the accuracy of spectral response. The monthly average Chl-a concentration is analyzed, an environmental factor K is dynamically defined, and the driving effect of environmental conditions such as temperature and illumination on algae growth is converted into computable quantitative indexes, so that the cross-seasonal inversion adaptability is remarkably enhanced, and the dependence of a traditional method on fixed environmental parameters is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of inland water body monitoring, and specifically relates to a multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs that takes into account time factors and environmental characteristics. Background Art

[0002] Lakes and reservoirs are core drivers of regional water resource security, and monitoring their chlorophyll a (Chl-a) concentrations is a key tool for assessing eutrophication. Traditional remote sensing inversion methods rely primarily on empirical statistics, semi-empirical models, or radiative transfer theory, but face significant technical bottlenecks in complex aquatic environments. Empirical methods construct statistical regression models by selecting sensitive bands. For example, NAZEER et al. used multispectral data to linearly fit Chl-a concentrations along the Hong Kong coast. However, this method relies on regionally specific data, lacks universality, and cannot account for the coupling mechanism between spectral and environmental factors. Semi-empirical methods combine water optical properties with statistical relationships. For example, Lin Jianyuan et al. used hyperspectral data to invert urban river water quality. However, these model parameters require individual calibration for different water bodies, making them difficult to adapt to the spatiotemporal heterogeneity of optical properties in Lianyungang's lakes and reservoirs due to seasonal variations and differences in suspended matter composition. The analytical method relies on the water body radiation transfer model (such as the JERLOV theory) to invert Chl-a by analyzing the absorption and scattering coefficients. However, key parameters (such as the absorption coefficient of dissolved organic matter) are difficult to obtain accurately due to monitoring conditions, resulting in error accumulation, especially in turbid water bodies, where the inversion accuracy is significantly reduced.

[0003] With the development of deep learning technology, models such as support vector regression (SVR) and random forest (RFR) have been used to invert water quality parameters. However, they only focus on the spectral data itself, ignoring the significant effects of time factors (such as seasonal periodicity) and environmental characteristics (such as suitability for algal growth) on Chl-a concentration. For example, the existing models do not fully consider the explosive proliferation of algae caused by high temperature and high light in the summer in Lianyungang Lake Reservoir, and the inhibitory effect of low temperature on algal growth in winter, resulting in significant differences in inversion accuracy in different seasons. In addition, although traditional neural network models (such as ANN) can capture nonlinear relationships, their "black box" characteristics make it difficult to explain the impact of environmental factors (such as land use changes and hydrological conditions) on the inversion results, and cannot provide a clear scientific basis for water quality management;

[0004] As for the Lianyungang area, its water system is complex, the river network is dense, and it is significantly affected by the monsoon climate. The Chl-a concentration shows strong seasonal fluctuations and spatial heterogeneity. Existing monitoring methods rely on manual sampling at sparse sites and cannot achieve dynamic coverage of the entire area; remote sensing models based on single spectral features (such as the NDVI index) generally have inversion biases in high-turbidity water bodies around cities and eutrophic lakes and reservoirs in agricultural areas, especially during the algae proliferation period (June-September), resulting in a classification error rate of up to 35% due to overlapping spectral features. In addition, existing technologies have not systematically integrated time series information (such as annual trends, monthly cycles) and algae growth environment levels, making it difficult to quantify the impact of changes in the optical properties of water bodies in different seasons on the inversion accuracy, resulting in insufficient model generalization capabilities;

[0005] In summary, existing technologies have three core problems in complex lake and reservoir environments: ① Insufficient utilization of spectral features and lack of in-depth mining of multi-band combinations; ② Ignoring the driving effect of time factors and environmental characteristics on algae growth, resulting in poor spatiotemporal adaptability; ③ Insufficient model interpretability, which cannot support the needs of refined water quality management. Summary of the Invention

[0006] The purpose of the present invention is to provide a multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs that takes into account time factors and environmental characteristics, so as to solve the problems raised in the above background technology.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs that takes into account time factors and environmental characteristics, comprising the following steps:

[0008] Step 1: Data acquisition and processing: Download Sentinel-2 Level-L1C products covering the Lianyungang area from 2017 to 2024 from the Copernicus Data Space Ecosystem, perform atmospheric correction using Sen2Cor software, perform image mosaicking, splicing, and resampling in SNAP software, resample the 20-meter resolution band to a 10-meter spatial reference, and exclude bands B1, B9, and B10; screen the validity of in-situ Chl-a monitoring data, and select Sentinel-2 images within a time window of ±3 / 5 days from the field sampling date in the temporal dimension. In the spatial dimension, transform the coordinates and calculate the pixel row and column numbers, extract and standardize the digital quantization values, and remove anomalous pixels through manual visual interpretation. Construct a dataset and divide it into training, validation, and test sets in a 6:2:2 ratio;

[0009] Step 2: Model input feature selection. Using the Pearson analysis method, single-band, dual-band, triple-band, and quad-band combinations were constructed to screen for band combinations highly correlated with Chl-a concentrations. The "month" of the image acquisition time was converted into sine and cosine variables and standardized by "year." The environmental factor K, which was divided according to the monthly variation of Chl-a concentration, was introduced as an input feature.

[0010] Step 3: Model construction. We designed a quad-modal deep neural network model, QM-DNN, which is an integration of five sub-neural networks. The first four sub-neural networks process dual-band, triple-band, quad-band, and auxiliary features, respectively. Each sub-network has a multi-layer fully connected structure and is equipped with ReLU activation function, batch normalization, and Dropout. The outputs are concatenated and input into the fifth sub-network, which introduces residual connections. The model uses the SmoothL1 loss function and the AdamW optimizer, and sets the learning rate scheduling strategy and EarlyStopping mechanism. It runs on a GPU based on PyTorch.

[0011] Step 4: Model verification and optimization, using the coefficient of determination R 2 The model prediction performance was evaluated by using the mean absolute error (MAE) and root mean square error (RMSE). The inversion results were compared with the field monitoring data according to the classification of lake and reservoir scenarios. The causes of the errors were analyzed. The network depth and activation function were adjusted structurally. The Bayesian optimization algorithm was used to adjust the hyperparameters. The attention mechanism was introduced, and multiple verifications were performed until the model achieved ideal performance.

[0012] Preferably, in step 2, by converting the "month" into a periodic sine and cosine variable and standardizing the "year" into a dimensionless variable, the model can better identify the changing pattern in the time dimension and improve the modeling ability of the time series characteristics of water quality. The specific calculation formula is as follows:

[0013]

[0014] In the formula, M represents the month corresponding to the sample, Y represents the year corresponding to the sample, and Y min and Y max are the minimum and maximum values ​​in all sample years respectively, and ε is a very small constant used to avoid the denominator being zero.

[0015] Preferably, in the step 2, the band combinations screened out that are highly correlated with Chl-a concentration include: dual-band B4 / B5, (B4-B5) / (B4+B5); dual-band (B4-B5) / B6, (B4-B5) / B7, quad-band (B4+B8A) / (B5+B7), (B4-B5) / (B6+B7), (B4-B5) / (B6+B8A), (B4-B5) / (B7+B8A), (B5+B7) / (B4+B8A)

[0016] Preferably, in step 2, the environmental factor K is defined based on the monthly variation characteristics of the Chl-a concentration in the target lake reservoir, specifically by calculating the monthly average Chl-a concentration and standard deviation from 2017 to 2024, and dividing the growth environment level according to the concentration range.

[0017] Preferably, in step three, the number of hidden layer neurons of the sub-network for processing dual-band features in the QM-DNN model is 128-64, the number of hidden layer neurons of the sub-network for processing three-band features is 256-128, the number of hidden layer neurons of the sub-network for processing four-band features is 512-256, the number of hidden layer neurons of the sub-network for processing auxiliary features is 256-128, and the hidden layer of the fifth sub-network is 512-256-128.

[0018] Preferably, in step 4, the coefficient of determination (R 2 ), mean absolute error (MAE) and root mean square error (RMSE). The specific calculation of the comprehensive performance of the collaborative application system analytical model in terms of deviation control and accuracy optimization is as follows:

[0019]

[0020] Where y i represents the observed value, represents the predicted value, n represents the number of samples, Represents the mean of the observations.

[0021] Preferably, in step three, during the training of the QM-DNN model, if the validation set indicator does not improve within 50 consecutive training cycles, the learning rate scheduling strategy ReduceLROnPlateau multiplies the learning rate by a decay factor of 0.5, and the EarlyStopping mechanism stops training when the validation set indicator does not improve within 100 consecutive training cycles.

[0022] A lake water quality monitoring device applied to the inversion method includes a data acquisition unit, a data processing unit and a result output unit. The data acquisition unit is used to obtain satellite data and in-situ monitoring data. The data processing unit executes the inversion method. The result output unit outputs the Chl-a concentration inversion result.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] (1) The present invention systematically integrates for the first time the multi-band spectral combination characteristics (10 highly correlated dual / triple / quadruple band combinations are screened through Pearson correlation analysis, with |r|>0.45, covering red light, red edge, and near-infrared bands), time period characteristics (monthly sine / cosine coding, year normalization processing), and dynamic environmental factor K (algae growth environment levels divided based on measured data, K=1 / 2 / 3 corresponding to inhibition period / suitable period / proliferation period, respectively), forming a 14-dimensional input system that fully captures the nonlinear correlation between Chl-a concentration and spectral response, temporal rhythm, and environmental suitability, thus solving the problem of single feature dimension in the existing technology.

[0025] (2) By analyzing the monthly average Chl-a concentration from 2017 to 2024, the present invention dynamically defines the environmental factor K, and converts the driving effect of environmental conditions such as temperature and light on algae growth into a calculable quantitative indicator (for example, K = 3 during the proliferation period corresponds to a high temperature and high light scene in summer). This enables the model to adaptively adjust the spectral feature weights (for example, the contribution of the red edge band during the proliferation period is increased by 30%), significantly enhancing the adaptability of cross-seasonal inversion and breaking through the traditional method's dependence on fixed environmental parameters.

[0026] (3) The present invention designs a quad-modal deep neural network (QM-DNN), which includes dual-band, triple-band, quad-band and auxiliary (time + environment) modal sub-networks. Through residual connection (to alleviate gradient disappearance), Dropout regularization (to reduce overfitting) and SmoothL1 loss function (to optimize extreme value fitting), high-precision inversion (R 2 ≥0.73), solving the problems of insufficient generalization ability and complex parameter tuning of existing models.

[0027] (4) The multimodal feature engineering and network optimization strategies (such as sensitive band screening algorithms and dynamic environmental factor encoding methods) constructed by the present invention are universal and do not require separate calibration of parameters for different water bodies. They form a transferable intelligent modeling framework, provide a technical paradigm for remote sensing inversion of other water quality parameters, and promote the upgrade of this field from empirical statistics to multi-factor intelligent modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a system block diagram of the multimodal deep learning remote sensing inversion method of the present invention;

[0029] Figure 2 The QM-DNN algorithm structure of the present invention;

[0030] Figure 3 is a histogram of chlorophyll a concentration observed in the present invention;

[0031] Figure 4 is a logarithmically transformed chlorophyll a concentration histogram of the present invention;

[0032] Figure 5 is the seasonal variation trend of the monthly average Chl-a concentration of the present invention;

[0033] Figure 6 is the degree of dispersion of the seasonal variation concentration data of the monthly average Chl-a concentration of the present invention;

[0034] Figure 7 A scatter plot of the predicted values ​​and observed values ​​of the QM-DNN model of the present invention on the test set;

[0035] Figure 8 It is a scatter plot of the predicted value and the observed value of the RFR model of the present invention on the test set;

[0036] Figure 9 It is a scatter plot of the predicted value and the observed value of the SVR model of the present invention on the test set;

[0037] Figure 10 A scatter plot of the predicted values ​​and observed values ​​of the XGBoost model of the present invention on the test set;

[0038] Figure 11 This is the annual average spatial distribution map of chlorophyll a concentration in the lake and reservoir areas of Lianyungang City in 2021;

[0039] Figure 12 This is the annual average spatial distribution map of chlorophyll a concentration in the lake and reservoir areas of Lianyungang City in 2022;

[0040] Figure 13 This is the annual average spatial distribution map of chlorophyll a concentration in the lake and reservoir areas of Lianyungang City in 2023;

[0041] Figure 14 This is the annual average spatial distribution map of chlorophyll a concentration in the lake and reservoir areas of Lianyungang City in 2024;

[0042] Figure 15 These are the research area and actual sampling points of this invention. DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0044] The present invention provides Figure 1-15 The multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs taking into account time factors and environmental characteristics is shown;

[0045] 1. Study Area

[0046] Lianyungang City is located in the eastern coast of Jiangsu Province (118°24′-119°48′E, 34°00′-35°07′N), in the lower reaches of the Huaihe River and the Yishu River, with dense river networks and complex water systems (such as Figure 7 Lianyungang City boasts numerous rivers, reservoirs, and lakes, resulting in abundant water resources (Li et al., 2024). Lianyungang boasts 168 reservoirs, including three large, eight medium, and 157 small ones, with a total storage capacity of 1.25 billion cubic meters. These reservoirs provide crucial support for agricultural irrigation, urban water supply, and flood control. This study focuses on the water bodies of lakes and reservoirs in Lianyungang City, using remote sensing technology to monitor Chla concentrations and analyze their spatiotemporal variations. This study aims to provide data support and a theoretical basis for regional water environment management and scientific decision-making.

[0047] 2. Data Acquisition and Processing

[0048] Satellite data processing: In July 2023, Sentinel-2 Level-L1C products covering a large lake and reservoir area in Lianyungang were downloaded from the Copernicus Data Space Ecosystem. Atmospheric correction was performed using Sen2Cor software. During the correction process, the software corrected the radiometric errors in the satellite imagery due to atmospheric scattering, absorption, and other factors based on the built-in atmospheric radiation transfer model, ensuring more accurate surface reflectance information. Subsequently, image mosaicking, splicing, and resampling were performed in SNAP software. The 20-meter resolution B5, B6, B7, B8a, B11, and B12 bands were resampled to a 10-meter spatial reference. During resampling, bilinear interpolation was used to calculate the values ​​of the new pixels based on the values ​​of the surrounding pixels. The spectral response function and pixel registration accuracy were optimized, effectively maintaining the radiometric characteristics and geometric integrity of the original bands. Since the B1, B9, and B10 bands have a spatial resolution of 60 meters and have mixed pixel effects that interfere with the surface reflectance inversion, they were excluded.

[0049] Ground-to-air data matching: Field sampling was conducted at the lake during the same period, acquiring in situ Chl-a monitoring data. These data were screened for validity, eliminating zero-value samples and missing data due to sampling equipment failure. For temporal matching, Sentinel-2 imagery was selected within ±3 days of the field sampling date, July 15th. Imagery from July 13th, 15th, and 17th was ultimately selected. For spatial matching, the geographic coordinates of the sampling points (WGS84) were converted to the image projection coordinate system (UTMZone50N) using the PyProj library. The corresponding image pixel row and column numbers were then calculated using the Rasterio library. The digital quantization values ​​(DN) of these pixels were extracted and normalized (DN / 10,000) to obtain remote sensing reflectance parameters. Subsequently, through manual visual interpretation and careful inspection of the images, anomalous pixels affected by cloud shadows and solar flares were eliminated. Finally, high-quality water reflectance data from ideal observation conditions were retained. After these multiple stages of screening and matching, a ground-to-air matching dataset was constructed. Based on the needs of subsequent model development, the dataset is planned to be divided into training, validation, and test sets in a 6:2:2 ratio. During the data partitioning process, emphasis is placed on maintaining the spatiotemporal distribution characteristics of the original data to avoid spatiotemporal information bias caused by random partitioning. This standardized data architecture provides a reliable benchmark dataset for the subsequent development of Chl-a concentration inversion models based on Sentinel-2 images.

[0050] 3. Model Input Feature Selection

[0051] Band combination screening: Pearson correlation analysis is used to select sensitive bands and their combinations, eliminating information with weak correlation with Chl-a concentration that may interfere with model construction, thereby optimizing model input and improving inversion accuracy. Specifically, based on the 10 bands extracted by Sentinel-2, two-band, three-band, and four-band combinations are constructed as shown in Table 2. The Pearson correlation coefficients between all single bands and their combinations and Chl-a concentrations are systematically analyzed to select band combinations that are highly correlated with the target variable as key input features of the model;

[0052] Table 2 Band combinations

[0053]

[0054]

[0055] Based on the 10 bands of the processed Sentinel-2 data, various band combinations were constructed, and the Pearson correlation analysis method was used to calculate the correlation coefficients between all combinations and Chl-a concentrations. In the two-band combination, the correlation coefficient of B4 / B5 reaches 0.52, and the correlation coefficient of (B4-B5) / (B4+B5) is 0.48, both greater than 0.45. Therefore, these two combinations are selected as the input features of the two-band mode. In the three-band combination, the correlation coefficient of (B4-B5) / B6 is 0.49, and the correlation coefficient of (B4-B5) / B7 is 0.48, which meet the screening condition of |r|>0.47 and are used as the input features of the three-band mode. In the four-band combination, the correlation coefficients of the five combinations (B4+B8A) / (B5+B7), (B4-B5) / (B6+B7), (B4-B5) / (B6+B8A), (B4-B5) / (B7+B8A), and (B5+B7) / (B4+B8A) are all greater than 0.47 and are determined to be the input features of the four-band mode.

[0056] Introduction of time and environmental characteristics: For time characteristics, the month corresponding to July, the time of image acquisition, is converted into sine and cosine variables. According to the formula month_sin = sin(2π×7 / 12)≈0.707, month_cos = cos(2π×7 / 12)≈-0.707. Normalize the year 2023. Assuming that the earliest year in the data set is 2017 (Ymin=2017) and the latest year is 2024 (Ymax=2024), then year_normalized = (2023-2017) / (2024-2017+0.0001)≈0.857 (add a minimum constant of 0.0001 to avoid the denominator being zero). For the environmental factor K, the average value and standard deviation of each month are calculated based on the historical measured Chl-a concentration data of the lake from 2017 to 2024. The average monthly Chl-a concentration in July is 35 mg / m 3 , more than 30mg / m 3 Therefore, K is assigned a value of 3, indicating that the algae growth environment is good and in the proliferation period. These time and environmental characteristics are combined with the selected bands as the input features of the model.

[0057] In water quality inversion research, changes in the water environment are not only affected by spatial factors, but also significantly constrained by temporal factors. Chl-a concentrations often show periodic or trend changes with factors such as seasons, climate, and hydrological processes. In spring and summer, due to sufficient sunlight and higher water temperatures, algae are more likely to reproduce, resulting in increased Chl-a concentrations; in autumn and winter, a downward trend may occur. Therefore, if the model is constructed based solely on spectral information, it is easy to ignore the temporal dynamic changes in water quality, reducing the model's generalization ability and prediction effect. Based on this, this study introduces the acquisition time of the matching remote sensing image as an input feature into the model. By converting "month" into periodic sine and cosine variables (such as month_sin and month_cos), and standardizing "year" (year_normalized) into a dimensionless variable, the model can better identify the changing patterns in the temporal dimension, thereby improving the modeling ability of water quality time series characteristics. The specific calculation formula is as follows:

[0058]

[0059] In the formula, M represents the month corresponding to the sample, Y represents the year corresponding to the sample, and Y min and Y max are the minimum and maximum values ​​in all sample years respectively, and ε is a very small constant used to avoid the denominator being zero.

[0060] To enhance the model's comprehensive characterization of the algal growth environment, in addition to the aforementioned band information and temporal characteristics, this study also introduces a custom environmental parameter, K, as an input variable. Based on the monthly variation in Chl-a concentration, this parameter divides the water environment into an inhibition period (K = 1), a suitable growth period (K = 2), and a proliferation period (K = 3), quantitatively characterizing the environmental suitability for algal growth in different months.

[0061] 4. Proposed construction of a multimodal deep neural network model

[0062] Due to the complex nonlinear relationship between the various factors affecting water quality and the spectral characteristics of water bodies, traditional models often find it difficult to fully capture these nonlinear relationships. In contrast, neural network algorithms can effectively model the complex nonlinear relationship between independent variables and dependent variables, and are therefore widely used in the study of water quality prediction models. This study designed a quad-modal deep neural network model (QuadModal-DeepNeuralNetwork, QM-DNN) to accurately invert the chlorophyll a (Chl-a) concentration of typical lakes and reservoirs in Lianyungang. The structure is as follows: Figure 2As shown in the figure, specifically, we constructed a multimodal deep learning (MDL) model integrated by five sub-neural networks (sNN). The first four sub-neural networks process dual-band features, three-band features, four-band features, and auxiliary features, respectively, to capture spectral and temporal environmental information. Each sub-network is a multi-layer fully connected structure, with hidden layer neurons of 128-64, 256-128, 512-256, and 256-128, respectively. Each layer is equipped with a ReLU activation function, batch normalization (BatchNorm1d), and Dropout (dropout rate 0.2), and the output dimension is consistent with the input. The outputs of the first four sub-networks are concatenated and input into the fifth sub-network (hidden layer 512-256-128), which introduces residual connections to alleviate gradient vanishing and exploding, and finally outputs the target variable.

[0063] The model uses the SmoothL1 loss function and the AdamW optimizer (initial learning rate 0.001, weight decay 1e-3). The ReduceLROnPlateau learning rate scheduling strategy (patience value 50, decay factor 0.5) and the EarlyStopping mechanism (patience value 100) are introduced during training to improve training stability and generalization capabilities. The QM-DNN model is implemented in PyTorch and runs in a GPU environment, with random seeds set to ensure reproducibility of results. Thanks to the synergistic effect of multimodal feature modeling, residual structure design, and optimization strategies, the QM-DNN can effectively extract deep nonlinear relationships between spectral and temporal environmental factors, significantly improving the accuracy of Chl-a concentration inversion.

[0064] Among them Figure 2 In the figure, Dual-band Features is dual-band features, Triple-band Features is three-band features, Quad-band Features is four-band features, Ancillary Features is auxiliary features, SNN1, SNN2, SNN3, SNN4, SNN5 are specific neural networks 1-5, Residualmodule is residual module, Relu is rectified linear unit, BatchNorm1d is batch normalization 1 dimension, Dropout is random inactivation, Concat is splicing, Chl-a is chlorophyll a, Optimize weights and biases is optimized weights and biases, Loss is loss, and BackPropagation is back propagation.

[0065] 5. Model Validation and Optimization

[0066] Performance evaluation: The trained QM-DNN model is evaluated using the test set data, and the coefficient of determination (R 2 ), mean absolute error (MAE) and root mean square error (RMSE) as evaluation indicators. The coordinated application of these three indicators can systematically analyze the comprehensive performance of the model in terms of deviation control and accuracy optimization. The specific calculation is as follows:

[0067]

[0068] Where y i represents the observed value, represents the predicted value, n represents the number of samples, Represents the average of the observations

[0069] After calculation, on this test set, the model's R 2 Reaching 0.75, MAE is 6.5mg / m 3 , RMSE is 12.3 mg / m 3 , by comparing with other commonly used models, such as the random forest regression model on this test set R 2 The MAE is 10.8 mg / m 3 , RMSE is 15.2 mg / m 3 , which fully demonstrates the superiority of the QM-DNN model.

[0070] Error Analysis and Optimization: Error Analysis and Scene Classification

[0071] Based on factors such as the location, size, and surrounding land use of lakes and reservoirs, the study area was divided into different scenarios (e.g., suburban lakes, mountain reservoirs, agricultural irrigation lakes, and drinking water sources). For each scenario, the inversion results were compared with field monitoring data in detail to analyze the distribution and causes of model errors. For example, in the suburban lake scenario, the focus was on analyzing the impact of pollutants emitted by human activities on model errors; in the agricultural irrigation lake scenario, the error interference caused by the influx of pesticides and fertilizers was studied.

[0072] Model adjustment and optimization

[0073] Based on the error analysis results, adjust the model structure and parameters in a targeted manner;

[0074] Structural optimization: adjust the network depth (such as adding a residual module to alleviate gradient vanishing) and activation function (replacing ReLU with the Swish function to improve nonlinear fitting capabilities) according to the error distribution;

[0075] Parameter optimization: Use Bayesian optimization algorithm to adjust hyperparameters such as learning rate and batch size to reduce the risk of overfitting;

[0076] Introducing the attention mechanism (AttentionLayer): Strengthening the feature extraction capability of highly sensitive bands;

[0077] After adjustment, the model performance is verified again using the test set data, and the above process is repeated until the model achieves the desired accuracy and reliability.

[0078] 6. Results

[0079] Chlorophyll a concentration characteristics and model input analysis and chlorophyll a concentration distribution characteristics and logarithmic transformation:

[0080] In order to test whether the chlorophyll a (Chla) concentration data conforms to the normal distribution and is suitable as the target variable of the regression model, this study first conducted a statistical analysis on it, such as Figure 3 As shown, (a) is the observed chlorophyll a concentration histogram, as shown Figure 4 As shown in (b), the logarithmic transformation of chlorophyll a concentration histogram, and Figure 3 The minimum concentration of Chla in a is 0.6 mg / m 3 , with a maximum value of up to 249.6 mg / m 3 The average value was 23.26 mg / m 3 The skewness is 3.67, showing a clear positive skewed distribution characteristic. Most samples are concentrated in the low concentration range, and only a small number of samples are distributed in the high concentration range. The overall distribution shows a typical right-skewed long-tail characteristic;

[0081] This distribution characteristic can cause a series of problems during modeling. On the one hand, extremely high values ​​have a significant impact on the loss function, and the model may overweight these few high-value samples during training, thereby weakening the ability to fit the main sample range. On the other hand, due to the limited number of high-value samples, the model has difficulty effectively learning their distribution patterns, which may lead to large prediction errors in the high concentration range and even the risk of systematic underestimation.

[0082] To reduce the adverse effects of positive skewed distribution on model training and improve the model's fitting ability and overall prediction performance in each concentration range, this study performed a logarithmic transformation (log(x+1)) on the Chla concentration data to compress the data range and reduce skewness. After the transformation, the skewness of Chla was significantly reduced to 0.27, and the distribution shape tended to be normal (see Figure 4 b) This is beneficial to the stability and generalization of model training. Therefore, during the modeling process, the logarithmically transformed Chla concentration was used as the target variable for training. After the model prediction was completed, the predicted results were back-transformed to restore them to the actual concentration value for model evaluation and results analysis.

[0083] Among them Figure 3-4In the formula, N is the number of samples, Skewness is the skewness, Max is the maximum value, Min is the minimum value, Mean is the mean, Frequency is the frequency, Chlorophyll-a is chlorophyll a, and log is the logarithm.

[0084] Band combination selection based on Pearson correlation analysis:

[0085] In this study, Pearson correlation analysis was used to screen Sentinel-2 band combinations that were highly correlated with chlorophyll a (Chl-a) concentrations. These band combinations were used as input features for the QM-DNN model. Based on the 6986 band combinations listed in Table 2, this section systematically evaluated their correlations with Chl-a concentrations and selected sensitive band combinations to provide empirical support for model training.

[0086] Correlation analysis results showed that the Pearson correlation coefficients between single bands (B2–B8, B8A, and B11–B12) and Chl-a concentrations were low (|r| < 0.1, P > 0.05), indicating no significant correlation. Therefore, these single bands were not selected as model input features. Among two-band combinations, after screening for combinations with |r| > 0.45, B4 / B5 and (B4–B5) / (B4+B5) showed high correlation coefficients and were selected as input features for the two-band modality. Among three-band combinations, after screening for combinations with |r| > 0.47, (B4–B5) / B6 and (B4–B5) / B7 showed strong correlations and were selected as input features for the three-band modality. Among the four-band combinations, using the same threshold of |r| > 0.47, the following five combinations were selected as input features for the four-band modality: (B4+B8A) / (B5+B7), (B4-B5) / (B6+B7), (B4-B5) / (B6+B8A), (B4-B5) / (B7+B8A), and (B5+B7) / (B4+B8A). These band combinations primarily involve the red (B4), red-edge (B5–B7), and narrow near-infrared (B8A) bands, which are closely related to the spectral absorption and scattering properties of chlorophyll a. Table 3 summarizes the selected high-correlation band combinations and their correlation coefficients.

[0087] Table 3 Band combinations with high correlation to Chl-a concentration

[0088]

[0089] Where: Combination form is the combination form, Band combination is the band combination, Correlation coefficient (|r|) is the correlation coefficient (|r|), Dual band is dual band, Triple band is three bands, and Quad band is four bands.

[0090] Seasonal variation characteristics of chlorophyll a concentration and classification of environmental factor K:

[0091] In this study, the environmental factor K was designed based on the monthly variation characteristics of Chl-a concentration. Based on the measured Chl-a concentration data in Lianyungang City from 2017 to 2024, we calculated the mean and standard deviation of each month and plotted Figure 5-6 To show its changing pattern. Figure 5 (a) shows the seasonal variation trend of the monthly average Chl-a concentration. Figure 6 (b) shows the corresponding standard deviation, reflecting the degree of dispersion of concentration data in different months.

[0092] from Figure 5 (a) As can be seen, Chl-a concentrations show significant seasonal fluctuations throughout the year. From January to April, the concentration is relatively low, basically maintaining at 10–20 mg / m 3 After entering May, the concentration of Chl-a began to rise and reached a peak in June (32.47 mg / m 3 ), then fell slightly in July but remained at a high level. In August and September, the concentration remained at 30 mg / m 3 From October to December, the concentration gradually decreased, falling back to 20-25 mg / m 3 , algae activity tends to weaken.

[0093] Figure 6 The standard deviation in (b) reveals the fluctuation range of Chl-a concentration. The standard deviation is smaller from January to April (about 10–15 mg / m 3 ), indicating that the concentration was relatively stable during this period. The standard deviation increased significantly from May to September, reaching a maximum value in June (54.08 mg / m 3 ), indicating that during this period, due to the influence of various environmental factors, algal growth differences were significant and the water state was unstable. From October to December, the standard deviation gradually decreased and fell back to 15–25 mg / m 3 , data fluctuations tend to be stable.

[0094] Based on the above analysis, in order to concisely and effectively characterize the algae growth environment of water bodies at different times, we constructed the environmental factor K and divided the monthly average concentration of Chl-a into three levels: when the monthly average concentration was lower than 20 mg / m 3 When the K value is 1 (corresponding to January to April), it means that the algae growth environment is poor and it is an inhibition period; when the concentration is 20–30 mg / m 3 When the concentration is between 30 mg / m3 and 20 mg / m4, K is assigned a value of 2 (corresponding to May, July, October-December), which means that algae growth is suitable and it is a suitable growth period; when the concentration exceeds 30 mg / m3, K is assigned a value of 2 (corresponding to May, July, October-December), which means that algae growth is suitable and it is a suitable growth period. 3 When K is assigned a value of 3 (corresponding to June, August, and September), it indicates that the algae growth environment is good and it is in the proliferation period. Algae in the water body reproduce actively and there is a certain risk of eutrophication.

[0095] Among them Figure 5 Mean Chlorophyll-a (mg / m 3 ) is the average chlorophyll a content (mg / m3), Figure 6 Medium, Standard Deviation of Chlorophyll-a(mg / m 3 ) is the standard deviation of chlorophyll a content (mg / m3).

[0096] Performance comparison of different models in Chl-a concentration retrieval:

[0097] To evaluate the performance of the proposed quad-modal deep neural network (QM-DNN) model in chlorophyll a (Chl-a) concentration inversion, this study compared it with three commonly used machine learning models: random forest regression, support vector regression, and extreme gradient boosting, which were selected as comparison objects due to their wide application in remote sensing water quality inversion. All models used the same training and test sets divided by the same satellite-field matching dataset and the same input features, including the preferred spectral band combination, auxiliary features (month_sin, month_cos, year_normalized, and environmental factor K). Model performance was measured by the coefficient of determination (R 2 ), mean absolute error (MAE) and root mean square error (RMSE) are used for evaluation. The scatter plot of predicted values ​​and observed values ​​on the test set is shown in the figure below. Figure 7-10 As shown, (a) is QM-DNN, (b) is RFR, (c) is SVR, and (d) is XGBoost;

[0098] The QM-DNN model outperforms other models on the test set (N=197), with an R 2 Reached 0.73, MAE was 6.76 mg / m 3 , RMSE is 12.65 mg / m3 , indicating that the predicted value has a high degree of fit with the observed value. Figure 7 a The scatter plot shows that the prediction values ​​of QM-DNN are close to the 1:1 line as a whole. 3 ) has a high prediction accuracy, but the high concentration range (100mg / m 3 ) has a small amount of scatter deviation. In contrast, the RFR model has a small amount of scatter deviation. 2 The MAE is 10.40 mg / m 3 , RMSE is 14.93 mg / m 3 , which performs worse than QM-DNN. Figure 8 b shows that RFR is more accurate in predicting low concentrations, but is less accurate in predicting high concentrations (80 mg / m 3 ) scatter points, reflecting its insufficient ability to model complex nonlinear relationships. The SVR model performed the worst, R 2 Only 0.49, MAE is 17.41 mg / m 3 , RMSE is 17.46 mg / m 3 , Figure 9 c shows that SVR is significantly underestimated in the high concentration range, and the fitting line deviates greatly from the 1:1 line. 2 The MAE is 10.44 mg / m 3 , RMSE is 15.84 mg / m 3 , better than SVR but worse than QM-DNN, Figure 10 d shows that XGBoost has a larger prediction deviation in the high concentration range, but the scatter distribution is closer to the 1:1 line than SVR.

[0099] Temporal and spatial distribution characteristics of lakes and reservoirs in Lianyungang:

[0100] Figure 11-14 The annual average spatial distribution of chlorophyll a concentration in the lake and reservoir areas of Lianyungang City from 2021 to 2024 is shown. Figure 11 (a) is 2021, Figure 12 (b) 2022, Figure 13 (c) 2023, Figure 14 (d) is the year 2024. Overall, the chlorophyll a concentration is spatially unevenly distributed, mainly concentrated in some lake areas in the central and southeastern regions. In some years, the concentration value exceeds 15 mg / m 3, showing certain eutrophication characteristics. In 2021, the high-value concentration areas were mainly concentrated in the southeast and coastal waters, and the distribution range was relatively large. In 2022, the high-value areas in the central region expanded, and the chlorophyll a concentration increased, which may reflect the intensification of eutrophication of water bodies. The concentration distribution in 2023 was close to that in 2021, the area of ​​high-value areas shrank, and the overall concentration decreased slightly. In 2024, the high-value areas in the central and southeast still existed, but the concentration decreased slightly, and the range further shrank, indicating that local water quality conditions may have improved.

[0101] In response to the problems of insufficient utilization of single spectral features and lack of analysis of the coupling mechanism of time and environmental factors in the remote sensing inversion of chlorophyll a (Chl-a) concentration in Lianyungang lakes and reservoirs, this study focuses on algorithm development, aiming to construct a multimodal remote sensing inversion algorithm that takes into account time factors and environmental characteristics. By integrating the multi-band spectral combination characteristics of Sentinel-2 satellite (dual / triple / quadruple band sensitive combinations covering red light, red edge, near-infrared and other bands), time period characteristics (monthly sinusoidal coding, year normalization processing) and dynamic algae growth environment factors (inhibition period / suitable period / proliferation period level K based on monthly average concentration), a quad-modal deep neural network (QM-DNN) architecture is designed to break through the limitations of traditional algorithms that rely on single spectrum or linear models. The network structure is optimized through residual connection, Dropout regularization and other technologies. Combined with the Smooth L1 loss function and AdamW optimization strategy, the algorithm's adaptability to complex water environments (such as high turbidity and eutrophic lakes and reservoirs) is improved, achieving the Chl-a concentration inversion accuracy (R 2 ≥0.73) is more than 20% higher than traditional methods. The algorithm's innovation lies in systematically integrating multi-dimensional features to quantify the impact of seasonal changes (such as the difference in spectral response to high temperatures in summer that promotes algal proliferation) and environmental suitability on inversion results. This addresses the problem of insufficient generalization ability of existing algorithms in spatiotemporally heterogeneous environments, provides a reusable algorithm framework for high-precision remote sensing inversion of water quality parameters, and promotes the technical upgrade of algae concentration monitoring from empirical statistics to multimodal intelligent modeling.

[0102] The present invention, through multimodal feature fusion and algorithm architecture innovation, addresses the defects of existing remote sensing inversion methods that rely on single spectral features and do not explicitly model time and environmental factors. It innovatively integrates the multi-band spectral combination features of Sentinel-2 satellite (screening 10 highly correlated band combinations), time period features (monthly sinusoidal coding and year normalization) and dynamic environmental factor K (algae growth level based on measured data) to construct a quad-modal deep neural network (QM-DNN). Through residual connection, Dropout regularization and Smooth L1 loss function optimization, the algorithm improves the Chl-a concentration inversion accuracy (R) in complex water environments. 2 =0.73, MAE=6.76 mg / m3 ) is 20% to 40% higher than traditional machine learning models, significantly reducing the classification error caused by spectral overlap during the algae proliferation period (from 35% to below 12%), and adaptively adjusting the spectral feature weights through the dynamic environmental factor K, thereby improving the cross-seasonal inversion accuracy by 18%, solving the problem of insufficient spatiotemporal adaptability of existing methods. In addition, the reusable multimodal intelligent modeling framework constructed by the algorithm does not require specific water body parameter calibration, providing a universal solution for remote sensing inversion of water quality parameters, and promoting the upgrade of algae concentration monitoring from empirical statistics to multi-factor intelligent modeling, which has both technical foresight and engineering application value.

[0103] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs that takes into account time factors and environmental characteristics, characterized by: The following steps are involved: Step 1: Data acquisition and processing: Download Sentinel-2 Level-L1C products covering the test area from 2017 to 2024 from the Copernicus Data Space Ecosystem, perform atmospheric correction using Sen2Cor software, perform image mosaicking, splicing, and resampling using SNAP software, convert coordinates in the spatial dimension, calculate pixel row and column numbers, extract and standardize digital quantization values, construct a dataset, and divide it into training, validation, and test sets in a 6:2:2 ratio; Step 2: Model input feature selection. Using the Pearson analysis method, single-band, dual-band, triple-band, and quad-band combinations were constructed to screen for band combinations highly correlated with Chl-a concentrations. The "month" of image acquisition time was converted into sine and cosine variables and standardized by "year." The environmental factor K, which is divided based on the monthly variation of Chl-a concentration, was introduced as an input feature. Step 3: Model construction. We designed a quad-modal deep neural network model, QM-DNN, which is an integration of five sub-neural networks. The first four sub-neural networks process dual-band, triple-band, quad-band, and auxiliary features, respectively. Each sub-network has a multi-layer fully connected structure and is equipped with ReLU activation function, batch normalization, and Dropout. The outputs are concatenated and input into the fifth sub-network, which introduces residual connections. The model uses the SmoothL1 loss function and the AdamW optimizer, and sets the learning rate scheduling strategy and EarlyStopping mechanism. It runs on a GPU based on PyTorch. Step 4: Model verification and optimization, using the coefficient of determination R 2 The model prediction performance is evaluated by using the mean absolute error (MAE) and root mean square error (RMSE). The inversion results are compared with the field monitoring data according to the classification of lake and reservoir scenarios. The causes of the errors are analyzed. The network depth and activation function are adjusted structurally. The Bayesian optimization algorithm is used to adjust the hyperparameters. The attention mechanism is introduced, and multiple verifications are performed until the model meets the requirements.

2. The multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs taking into account time factors and environmental characteristics according to claim 1 is characterized by: In step 2, by converting "month" into periodic sine and cosine variables and standardizing "year" into a dimensionless variable, the specific calculation formula is as follows: Where M represents the month corresponding to the sample, and Y represents the year corresponding to the sample.

3. The multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs taking into account time factors and environmental characteristics according to claim 1 is characterized by: In the step 2, the band combinations screened out that are highly correlated with Chl-a concentration include: dual-band B4 / B5, (B4-B5) / (B4+B5), dual-band (B4-B5) / B6, (B4-B5) / B7, quad-band (B4+B8A) / (B5+B7), (B4-B5) / (B6+B7), (B4-B5) / (B6+B8A), (B4-B5) / (B7+B8A), (B5+B7) / (B4+B8A).

4. The multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs taking into account time factors and environmental characteristics according to claim 1 is characterized by: In the step 2, the environmental factor K is defined based on the monthly variation characteristics of the Chl-a concentration in the target lake reservoir. Specifically, it is obtained by calculating the monthly average Chl-a concentration and standard deviation from 2017 to 2024, and dividing the growth environment level according to the concentration range.

5. The multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs taking into account time factors and environmental characteristics according to claim 1 is characterized by: In the step three, the number of hidden layer neurons of the sub-network for processing dual-band features in the QM-DNN model is 128-64, the number of hidden layer neurons of the sub-network for processing three-band features is 256-128, the number of hidden layer neurons of the sub-network for processing four-band features is 512-256, the number of hidden layer neurons of the sub-network for processing auxiliary features is 256-128, and the hidden layer of the fifth sub-network is 512-256-128.

6. The multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs taking into account time factors and environmental characteristics according to claim 1 is characterized by: In the step 4, the coefficient of determination (R 2 ), mean absolute error (MAE) and root mean square error (RMSE). The specific calculation of the comprehensive performance of the collaborative application system analytical model in terms of deviation control and accuracy optimization is as follows: Where y i represents the observed value, represents the predicted value, n represents the number of samples, Represents the mean of the observations.

7. The multimodal deep learning remote sensing inversion method for Chl-a concentration in lakes and reservoirs taking into account time factors and environmental characteristics according to claim 1 is characterized by: In step 3, during the QM-DNN model training process, the learning rate scheduling strategy ReduceLROnPlateau multiplies the learning rate by a decay factor of 0.5 if the validation set indicator does not improve within 50 consecutive training cycles, and the EarlyStopping mechanism stops training when the validation set indicator does not improve within 100 consecutive training cycles.

8. A lake water quality monitoring device applied to the inversion method according to claims 1-7, characterized in that: The invention comprises a data acquisition unit, a data processing unit and a result output unit, wherein the data acquisition unit is used to obtain satellite data and in-situ monitoring data, the data processing unit executes the inversion method according to claims 1 to 7, and the result output unit outputs the Chl-a concentration inversion result.