Two-stage machine learning algae bloom prediction method based on stationary satellite
By employing geostationary satellite remote sensing and a two-stage machine learning approach, and utilizing geostationary satellite data and environmental parameters, the continuity and accuracy issues in lake algal bloom prediction in existing technologies have been resolved. This approach enables high-precision Chla concentration prediction, improves the timeliness and reliability of algal bloom early warning, and supports lake management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING INST OF GEOGRAPHY & LIMNOLOGY
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies, lacking high-frequency measured data, struggle to achieve continuous and high-precision prediction of lake algal blooms, impacting the timeliness and reliability of algal bloom warnings.
A two-stage machine learning approach based on geostationary satellites is adopted. Using geostationary satellite remote sensing data and multiple environmental parameters, feature variables are selected through a random forest model and combined with an extreme gradient boosting model to predict Chla concentration. This enables the construction of a continuous diurnal dataset and the prediction of future concentration changes by combining environmental parameter feature variables.
It enables high-precision prediction of Chla concentration in lakes even in the absence of measured data, improving the timeliness and reliability of algal bloom early warning and supporting the management and control of eutrophic lakes.
Smart Images

Figure CN121981302A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of satellite remote sensing and water environment analysis technology, specifically involving a two-stage machine learning method for predicting algal blooms based on geostationary satellites. Background Technology
[0002] Harmful algal blooms lead to dissolved oxygen depletion, toxin accumulation, malodorous release, and ecosystem degradation, posing a significant environmental threat to water security and public health. Influenced by human activities and climate change, algal blooms in lakes worldwide are expanding. Due to the significant spatiotemporal heterogeneity of algal blooms over short periods, relying solely on on-site monitoring is insufficient to fully understand the overall situation. Satellite remote sensing, with its advantages of wide coverage, high speed, and periodicity, has been widely used for monitoring lake algal blooms. For algal bloom prediction, existing technologies are mainly divided into mechanism-driven prediction models and data-driven prediction models. Mechanism-driven prediction models have numerous parameters, are complex, and have poor cross-regional applicability; data-driven prediction models typically rely on large amounts of measured data, making it difficult to meet prediction needs under conditions lacking high-frequency observations.
[0003] Chlorophyll a (Chla) is an important indicator for measuring the degree of eutrophication and algal blooms in water bodies. Previous studies have attempted to build predictive models using satellite data. However, existing methods show declining predictive performance in scenarios with discontinuous time series or limited measured data. Therefore, there is an urgent need for a technical solution that can fully utilize long-term satellite observation data and achieve continuous, high-precision Chla prediction even under conditions of scarce measured data, in order to improve the timeliness and reliability of algal bloom early warning. Summary of the Invention
[0004] The purpose of this invention is to provide a two-stage machine learning method for predicting algal blooms based on geostationary satellites.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A two-stage machine learning method for predicting algal blooms based on geostationary satellites, the method comprising: Chla concentration data were retrieved using geostationary satellite remote sensing data; Hourly data of multiple environmental parameters are obtained, and statistical values of environmental parameter data in different time periods before the date to be predicted are calculated as alternative parameters. For each environmental parameter, a correlation analysis was performed between its alternative parameters and the Chla concentration on the date to be predicted, and at least one alternative parameter with the highest correlation was obtained as the characteristic variable of the environmental parameter. Using the characteristic variables of environmental parameters as input, the trained first machine learning model is used to predict the Chla concentration data for missing dates in the Chla concentration data, resulting in a continuous daily-scale Chla dataset. For the daily-scale Chla dataset, the statistical values of environmental parameters in different time periods before the date to be predicted are calculated. These parameters are used as candidate parameters to perform correlation analysis with the Chla concentration on the date to be predicted. At least one candidate parameter with the highest correlation is obtained as the feature variable of the Chla concentration data. Using environmental parameters and Chla concentration data as inputs, a trained second machine learning model is used to predict future Chla concentration changes in lakes.
[0006] In some embodiments of the present invention, the environmental parameters include wind speed, air temperature, solar radiation, air pressure, rainfall, surface runoff, and evaporation.
[0007] In some embodiments of the present invention, the statistical values include the mean, the maximum value, and the cumulative value.
[0008] In some embodiments of the present invention, the average and maximum values of wind speed, temperature, solar radiation and air pressure parameters within different time periods before the date to be predicted are obtained as alternative parameters. For parameters such as rainfall, surface runoff, and evaporation, the mean and cumulative values of these parameters over different time periods before the predicted date are obtained as alternative parameters. For Chla concentration data, the mean value of the concentration over different time periods before the date to be predicted is obtained as an alternative parameter.
[0009] In some embodiments of the present invention, the correlation analysis is performed by ranking the feature importance using a random forest.
[0010] In some embodiments of the present invention, the correlation analysis is performed in the following manner: For a given environmental parameter or Chla concentration data, a random forest model is built using its alternative parameters as input and the Chla concentration data for the predicted date as output. The model is then selected based on feature importance ranking. n The characteristic variable of a bit.
[0011] In some embodiments of the present invention, the Chla concentration values at different times obtained based on geostationary satellite remote sensing data are used as sample values, and each time is used as the prediction date to obtain the corresponding feature variables. The sample values and corresponding feature variables are used to expand the sample database for model training.
[0012] In some embodiments of the present invention, the first machine learning model and the second machine learning model are extreme gradient boosting models.
[0013] In some embodiments of the present invention, the training method of the first machine learning model is as follows: the inverted Chla concentration data is used as the ground truth, and the feature variables of the environmental parameters at the corresponding time are used as inputs to train the first machine learning model. The training method for the second machine learning model is as follows: using the daily-scale Chla dataset as the ground truth, and the feature variables of the environmental parameters and the feature variables of the Chla concentration data at the corresponding time as inputs, the second machine learning model is trained.
[0014] In some embodiments of the present invention, the geostationary satellite remote sensing data is selected from GOCI and GOCI-II satellite data.
[0015] The method of this invention uses a first-stage model to predict Chla concentration by selecting specific characteristic variables from environmental parameters, thereby obtaining continuous Chla concentration data, conducting long-term Chla trend change analysis, and obtaining Chla concentration characteristic variables by combining the continuous Chla concentration dataset. The second-stage model is then trained by combining the characteristic variables of environmental parameters, enabling high-precision prediction of Chla concentration changes across the entire lake several days in advance, providing technical support for the management and precise control of algal blooms in eutrophic lakes. Attached Figure Description
[0016] The accompanying drawings aim to more intuitively and systematically illustrate the overall process of this invention, from model construction to model application, thereby improving the efficiency and readability of the document's information expression. Each component and its label in the drawings is explained accordingly, facilitating the reader's accurate understanding of the specific structure and function of the method. The steps and implementation methods of this invention will be further elaborated below with reference to examples and the accompanying drawings, wherein: Figure 1 This is the annual statistical result of Chla data obtained from the inversion of GOCI and GOCI-II images.
[0017] Figure 2 This is a schematic diagram of an extreme gradient boosting model.
[0018] Figure 3 These are the results of two-stage extreme gradient improvement model accuracy, (a) is the accuracy of the interpolation model, and (b) is the accuracy of the prediction model.
[0019] Figure 4 These are interpolated daily-scale Chla data for Lake A from 2011 to 2024.
[0020] Figure 5 The accuracy of the two-stage extreme gradient boosting model on the independent validation set is shown in (a) Independent validation set results from January to March 2025, and (b) Comparison between the predicted Chla and the Chla retrieved from the image. Detailed Implementation
[0021] To more clearly illustrate the technical solution of the present invention, specific embodiments and accompanying drawings are now described.
[0022] In this disclosure, the description of various parts of the invention will be carried out with reference to the accompanying drawings, but these embodiments and drawings do not exhaust all possible applications of the invention. In fact, the foregoing technical solutions and implementation processes, as well as the specific concepts and steps given below, can be flexibly adjusted and expanded according to needs, and the technical essence of the invention is not limited to specific implementation methods. Furthermore, based on the hydrological and ecological characteristics of different lakes, some steps in the technical process can be appropriately added or deleted to construct a prediction model suitable for specific lake conditions. The technical content disclosed in this invention has strong adaptability and scalability and can be adjusted according to application scenarios.
[0023] Example 1 This embodiment illustrates the two-stage machine learning algal bloom prediction method based on geostationary satellites according to the present invention.
[0024] This embodiment predicts Chla concentration in Lake A based on the extreme gradient boosting model and the GOCI series satellites, as follows: Using Lake A, a typical eutrophic lake, as the study area, we acquired data on air temperature, air pressure, wind speed, and solar radiation in the lake area, as well as rainfall, surface runoff, and evaporation data within the Lake A watershed. Chla data were obtained by inverting GOCI and GOCI-II data from 2011 to 2024. The sample library was expanded by combining satellite and environmental data from different time points. The optimal interval for environmental and Chla features was determined based on feature importance ranking using a random forest model. A two-stage extreme gradient boosting model was constructed for interpolation to obtain a daily-scale Lake A Chla dataset, and Chla prediction was achieved based on this data.
[0025] The implementation process of this invention will be illustrated step by step below with accompanying drawings.
[0026] 1) Obtain Chla data for model training; Due to the difficulty in obtaining long-term measured data and the inability to comprehensively assess the Chla level of the entire lake, this application uses satellite-retrieved Chla data for model training. Chla data is obtained by inverting GOCI and GOCI-II satellite data using an existing trained random forest model. The inversion model used in this embodiment is the Chla inversion model constructed by the applicant based on the random forest method in a previous article (Guo, Y., Wei, X., Huang, Z., Li, H., Ma, R., Cao, Z., Shen, M., & Xue, K. (2023). Retrievals of Chlorophyll-a from GOC I and GOCI-II Data in Optically Complex Lakes). Remote. Sens., 15 , 4886).
[0027] This data will be used as ground truth for training and validation of the two-stage extreme gradient boosting model. Figure 1 The data presents the annual average statistics of Lake A Chla data from 2011 to 2024.
[0028] 2) Based on the feature importance ranking of the random forest model, determine the optimal time interval for environmental features. A random forest model needs to be built for each feature.
[0029] Taking wind speed as an example, the random forest model is constructed as follows: for the multiple calculated wind speed variables (9 in this example), they are input as feature factors into a random forest model. The Chla concentration value obtained in step 1) is used as the ground truth. The top two wind speed variables are selected according to the feature importance ranking built into the model, thus obtaining the WIN 168_mean and WIN 12_mean .
[0030] Table 1 shows the calculation process for the 7 environmental features, and Table 2 shows the ranking of the importance of the features in the random forest model. In the end, a total of 14 features were obtained.
[0031] Chla concentration values at different times in GOCI were used as sample values. Environmental data corresponding to the times in GOCI images were statistically analyzed to expand the sample database for subsequent model training.
[0032] Table 1. Statistical Information on Environmental Variables
[0033] Table 2. The top two most important variables.
[0034] The meanings of the symbols and subscripts shown in the table are as follows: x _mean refers to the front x Hourly average, such as WIN 168_mean This represents the average wind speed over the previous 168 hours. x _sum refers to the preceding text x Hourly cumulative value, such as PRE 168_sum This represents the cumulative rainfall over the previous 168 hours. x _max refers to the preceding text x Hourly maximum value, for example, EVA 12_max This represents the maximum evaporation value in the first 12 hours.
[0035] 3) Construct the first-stage extreme gradient boosting model (XGBoost_Interp); XGBoost_Interp uses only environmental features for interpolation to obtain continuous daily Chla data from 2011 to 2024, and uses grid search to determine the optimal parameter settings for the model. The input features of the XGBoost_Interp model only include environmental variables, i.e., 14 features of 7 environmental variables, and the output is the interpolated Chla concentration. The hyperparameters of the model are set as follows: n_estimators=1200, max_depth=55, min_child_weight=30, gamma=0.7, learning_rate=0.06. Figure 3 Figure (a) shows the accuracy performance of the interpolation model. Figure 4 This presents the daily-scale Chla dataset from 2011 to 2024.
[0036] 4) Determine the optimal environmental time interval for Chla features; Combining the daily-scale Chla dataset obtained through interpolation in section 3), the average Chla concentration for the period from 1 day to 30 days prior to the prediction date is calculated. Following the approach of determining the optimal time interval for environmental features, the Chla concentration for the preceding 26 days is determined by ranking the features according to the importance of the random forest model. 624_mean The mean has the highest importance.
[0037] 5) Construct the second-stage extreme gradient boosting model (XGBoost_Pred); XGBoost_Pred uses environmental features from the optimal interval and Chla features obtained in the first stage to predict subsequent Chla concentrations. The model input features are 14 features from 7 environmental variables (Table 2) and the mean Chla values from the previous 26 days (Chla). 624_mean The output is the predicted Chla concentration. The hyperparameters of the model are set as follows: n_estimators=1250, max_depth=50, min_child_weight=30, gamma=0.8, learning_rate=0.06.
[0038] Figure 3 Figure (b) shows the accuracy performance of the prediction model. Figure 5 Figure (a) shows the model's accuracy performance on the independent validation set. Figure 5 (b) shows a comparison between the predicted Chla and the Chla retrieved from the image.
Claims
1. A two-stage machine learning method for predicting algal blooms based on geostationary satellites, characterized in that, The method includes: Chla concentration data were retrieved using geostationary satellite remote sensing data; Hourly data of multiple environmental parameters are obtained, and statistical values of environmental parameter data in different time periods before the date to be predicted are calculated as alternative parameters. For each environmental parameter, a correlation analysis was performed between its alternative parameters and the Chla concentration on the date to be predicted, and at least one alternative parameter with the highest correlation was obtained as the characteristic variable of the environmental parameter. Using the characteristic variables of environmental parameters as input, the trained first machine learning model is used to predict the Chla concentration data for missing dates in the Chla concentration data, resulting in a continuous daily-scale Chla dataset. For the daily-scale Chla dataset, the statistical values of environmental parameters in different time periods before the date to be predicted are calculated. These parameters are used as candidate parameters to perform correlation analysis with the Chla concentration on the date to be predicted. At least one candidate parameter with the highest correlation is obtained as the feature variable of the Chla concentration data. Using environmental parameters and Chla concentration data as inputs, a trained second machine learning model is used to predict future Chla concentration changes in lakes.
2. The method according to claim 1, characterized in that, The environmental parameters include wind speed, air temperature, solar radiation, air pressure, rainfall, surface runoff, and evaporation.
3. The method according to claim 1, characterized in that, The statistical values include the mean, maximum value, and cumulative value.
4. The method according to claim 2, characterized in that, For wind speed, temperature, solar radiation, and air pressure parameters, obtain their mean and maximum values over different time periods before the predicted date as alternative parameters; For parameters such as rainfall, surface runoff, and evaporation, the mean and cumulative values of these parameters over different time periods before the predicted date are obtained as alternative parameters. For Chla concentration data, the mean values of different time periods before the date to be predicted are obtained as alternative parameters.
5. The method according to claim 1, characterized in that, The correlation analysis was performed by ranking the feature importance using a random forest.
6. The method according to claim 1, characterized in that, The correlation analysis was performed using the following method: For a given environmental parameter or Chla concentration data, a random forest model is built using its alternative parameters as input and the Chla concentration data for the predicted date as output. The model is then selected based on feature importance ranking. n The characteristic variable of a bit.
7. The method according to claim 1, characterized in that, Chla concentration values at different times obtained from geostationary satellite remote sensing data are used as sample values. Each time is used as the prediction date to obtain the corresponding feature variables. The sample database used for model training is expanded using the sample values and the corresponding feature variables.
8. The method according to claim 1, characterized in that, The first and second machine learning models are extreme gradient boosting models.
9. The method according to claim 1, characterized in that, The training method for the first machine learning model is as follows: the inverted Chla concentration data is used as the ground truth, and the feature variables of the environmental parameters at the corresponding time are used as inputs to train the first machine learning model. The training method for the second machine learning model is as follows: using the daily-scale Chla dataset as the ground truth, and the feature variables of the environmental parameters and the feature variables of the Chla concentration data at the corresponding time as inputs, the second machine learning model is trained.
10. The method according to claim 1, characterized in that, The geostationary satellite remote sensing data used are GOCI and GOCI-II satellite data.