Chla concentration prediction method and system based on lagged meteorological data

CN122778080APending Publication Date: 2026-09-18WUXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611255890.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

然而在真实的水华预警工况下,热力积累往往需要约三到四周才充分显现,而降雨带来的营养盐输入脉冲与温度波动则在约一周内即产生响应;当对所有气象因子使用同一固定滞后期时,模型既无法及时捕捉短期降雨脉冲,又会错过长期热力稳定状态的累积信号,致使预测出的Chla浓度峰值与现场实际水华暴发的时间相互错位

Benefits of technology

[0017]The beneficial effects of this invention are as follows: Based on water quality monitoring data of the target water body and spatiotemporally matched meteorological reanalysis data, this invention constructs multiple types of meteorological derived features by configuring lag period parameters to quantify the lag cumulative effect of meteorological conditions on algal growth. Then, a two-stage exhaustive search strategy is employed to objectively determine the optimal lag period combination for each meteorological derived feature under data-driven conditions. Finally, the corresponding meteorological derived features and water quality parameters are merged to construct a feature matrix, which is then input into a machine learning model for training and prediction. This method explicitly encodes the meteorological lag effect and objectively optimizes its time scale, significantly improving the accuracy of Chla prediction and possessing both good ecological interpretability and cross-water body universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122778080A_ABST
    Figure CN122778080A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of water environment monitoring and artificial intelligence, and particularly relates to a Chla concentration prediction method and system based on lagged meteorological data, which comprises obtaining water quality monitoring data and meteorological reanalysis data of a target water body; based on the meteorological reanalysis data, constructing multiple meteorological derived features with clear lag period parameters to quantify the lagged cumulative effect of meteorological conditions on algal growth; adopting a two-stage exhaustive search strategy combining coarse-grained global screening and fine-grained local verification to objectively determine the optimal lag period combination of each meteorological derived feature under data driving; and merging the meteorological derived features corresponding to the optimal lag period combination and the water quality parameters to input a machine learning model for training and prediction. According to the scheme of the present application, the model prediction accuracy and ecological interpretability are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of water environment monitoring and artificial intelligence. More specifically, this invention relates to a method and system for predicting Chla concentration based on lagging meteorological data. Background Technology

[0002] Chlorophyll a (Chla) concentration is a direct proxy for phytoplankton biomass and one of the most important parameters for measuring eutrophication and assessing algal dynamics. It is widely used in lake water quality management, drinking water safety assurance, and aquatic ecosystem health assessment. Predictive modeling of Chla concentration has gradually developed into two main approaches: mechanistic models based on ecological dynamics and data-driven machine learning models. The latter, with its powerful nonlinear fitting capabilities, is finding increasingly widespread application in water quality prediction.

[0003] In early warning scenarios for cyanobacterial blooms in large, shallow, eutrophic lakes (such as Taihu Lake), water intake switching, emergency dredging, and ecological regulation are often arranged based on short-term forecasts of Chla concentration. Cyanobacterial blooms are not triggered immediately by meteorological conditions on the day of the bloom, but rather are the result of a gradual accumulation over a preceding period, influenced by the combined effects of thermal buildup, wind disturbances, and rainfall runoff. These meteorological processes have significant and varying time lags in their impact on algal growth.

[0004] Most existing machine learning-based Chla prediction schemes directly use daily meteorological data or apply the same fixed lag period to all meteorological factors as model input. However, in real-world algal bloom warning scenarios, thermal accumulation often takes about three to four weeks to fully manifest, while nutrient input pulses and temperature fluctuations from rainfall produce a response within about a week. When the same fixed lag period is used for all meteorological factors, the model cannot capture short-term rainfall pulses in time, nor can it miss the cumulative signals of long-term thermal stability, resulting in a misalignment between the predicted Chla concentration peak and the actual time of the algal bloom. The direct consequence is insufficient advance warning or even missed warnings. Management personnel often receive overestimated predictions only when cyanobacteria have already accumulated on a large scale, forcing delays in intake switching and emergency salvage operations, and hindering the achievement of expected results in water plant supply security and lake water quality management.

[0005] Therefore, in the practical scenario of algal bloom early warning, how to match the lag time scale of meteorological data with the actual action patterns of various meteorological processes, so that the Chla prediction results are in line with the actual occurrence rhythm of algal blooms in time, is a technical problem that this type of prediction modeling method urgently needs to solve. Summary of the Invention

[0006] To address the technical problem of discrepancies between the predicted Chla concentration and the actual occurrence rhythm of algal blooms in the field, this invention provides solutions in the following aspects.

[0007] In a first aspect, the present invention provides a Chla concentration prediction method based on lagged meteorological data, comprising: S1, acquiring historical in-situ water quality monitoring data of a target water body and spatiotemporally matched meteorological reanalysis data; wherein the water quality monitoring data includes at least chlorophyll a concentration, and the meteorological reanalysis data includes daily average temperature, daily average wind speed, and daily precipitation; S2, constructing meteorological derivative features at multiple time scales based on the meteorological reanalysis data by configuring different lag period parameters; S3, determining the optimal lag period combination of various meteorological derivative features by employing a data-driven two-stage exhaustive search strategy based on the water quality monitoring data and the constructed meteorological derivative features; S4, extracting the corresponding meteorological derivative features according to the optimal lag period combination, merging them with the water quality monitoring data to construct a feature matrix, inputting the feature matrix into a selected machine learning model for training to construct a Chla concentration prediction model, and using the trained Chla concentration prediction model to predict the Chla concentration at future times.

[0008] Furthermore, step S1 also includes a pre-processing step for cleaning the water quality monitoring data: removing outliers and null values ​​that are outside the physically reasonable range; converting the monitoring timestamp into seasonal and diurnal features, and performing one-hot encoding on the seasonal features, diurnal features, and the site number features corresponding to the monitoring points as categorical features.

[0009] Furthermore, in step S2, the constructed meteorological derivative features include the following five categories: temperature variation coefficient, effective accumulated temperature, effective accumulated wind speed, previous rainfall index, and water body stability cumulative proxy value.

[0010] Further, in step S3, the two-stage exhaustive search strategy includes: a coarse-grained search stage: exhaustively searching all lag combination schemes within a preset full candidate lag range at a set coarse-grained sampling time interval, independently training and evaluating the performance of the prediction model; based on the search results, decomposing the overall performance of multi-feature coupling into independent response curves of single factors, and identifying the preferred lag range where various meteorological derivative features can enable the model performance to reach a high level; a fine-grained search stage: within the preferred lag range of each feature locked by the coarse-grained search, exhaustively verifying the full combination again with a smaller time step, and finally determining the optimal lag combination by comparing the prediction accuracy and cross-validation generalization ability of each combination within the performance saturation region, and comprehensively considering the peak marginal effect and statistical robustness.

[0011] Furthermore, the even smaller time step in the fine-grained search phase is 1 day.

[0012] Further, in step S4, the machine learning model is any one of the following ensemble learning models: extreme random tree model, random forest model, LightGBM model, or XGBoost model.

[0013] Furthermore, step S5 is included: using the SHAP method to perform interpretability analysis on the trained Chla concentration prediction model in order to analyze the marginal contribution of each environmental feature to the Chla concentration prediction and identify its ecological response pattern.

[0014] Further, step S5 specifically includes: S5.1, based on the trained machine learning model and the complete feature matrix, calculating the SHAP value of each feature to quantify the marginal contribution of each feature in each prediction, and obtaining the global importance ranking of each environmental feature by aggregating the absolute values ​​of the SHAP values ​​of all samples; S5.2, verifying and eliminating the information preemption effect for the categorical features that have been processed by one-hot encoding in the input features; S5.3, according to the relationship between feature values ​​and SHAP values, classifying environmental factors into three response modes: positive driving type, negative inhibiting type, and mixed response type, in order to identify the core driving factors of Chla concentration and their direction of influence.

[0015] Further, the steps to verify and eliminate the information preemption effect include: training a first model based on the complete feature matrix; simultaneously constructing a second feature matrix that removes all uniquely coded categorical features, retaining only continuous original water quality parameters and meteorological derived features, and training the second model; performing SHAP analysis on the first and second models respectively, and comparing the feature importance ranking structure of the two; if the SHAP importance of the uniquely coded categorical features in the first model is higher than that of other continuous features, and in the ablation control experiment, the performance of the second model after removing the categorical features decreases by less than a preset threshold compared to the first model, then it is determined that an information preemption effect exists, and the SHAP analysis results of the second model are used as the basis for identifying the ecological driving mechanism.

[0016] In a second aspect, the present invention also provides a Chla concentration prediction system based on lagging meteorological data, comprising: a processor; and a memory storing computer program instructions, wherein when the processor executes the computer program instructions, it implements the Chla concentration prediction method based on lagging meteorological data as described in one or more of the preceding embodiments.

[0017] The beneficial effects of this invention are as follows: Based on water quality monitoring data of the target water body and spatiotemporally matched meteorological reanalysis data, this invention constructs multiple types of meteorological derived features by configuring lag period parameters to quantify the lag cumulative effect of meteorological conditions on algal growth. Then, a two-stage exhaustive search strategy is employed to objectively determine the optimal lag period combination for each meteorological derived feature under data-driven conditions. Finally, the corresponding meteorological derived features and water quality parameters are merged to construct a feature matrix, which is then input into a machine learning model for training and prediction. This method explicitly encodes the meteorological lag effect and objectively optimizes its time scale, significantly improving the accuracy of Chla prediction and possessing both good ecological interpretability and cross-water body universality.

[0018] Furthermore, this invention simultaneously introduces five derived features, each with a clear ecological connotation: temperature variation coefficient, effective accumulated temperature, effective accumulated wind speed, antecedent rainfall index, and water body stability accumulated proxy value. These features quantify the thermal environment stability, available heat accumulation, effective layer-breaking wind energy, non-point source nutrient input pulse, and the continuous accumulation of the water body's thermodynamic-dynamic stable state. Within the same feature system, the various meteorological driving processes involved in algal growth are differentiated, thereby significantly enhancing the model's ability to express meteorological driving mechanisms. This results in more accurate predictions and provides a clear physical basis for ecological mechanism analysis and algal bloom control decisions.

[0019] Furthermore, this invention employs a two-stage strategy in the lag optimization module: first, a coarse-grained exhaustive search of the entire candidate range with a large sampling interval; then, a marginal effect analysis to lock in the optimal interval; and finally, a fine-grained exhaustive search of all combinations within that interval with a smaller step size. This decomposes the overall combination performance into single-factor independent response curves to efficiently filter out a large number of low-performance combinations, ensuring that all fine-grained search solutions fall into the performance saturation region. This significantly reduces computational overhead while taking into account the peak marginal effect and the generalization ability of cross-validation, ensuring that the determined optimal lag combination is both highly accurate and statistically robust. This provides a standardized and operable engineering path for reproducing this optimization process across water bodies. Attached Figure Description

[0020] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein: Figure 1 This is a flowchart of the Chla concentration prediction method based on lagging meteorological data in an embodiment of the present invention; Figure 2 This is a graph showing the trend of R² and cross-validation R² for different lag periods of a single meteorological feature in an embodiment of the present invention. Figure 3This is a graph showing the changing trends of RMSE and MAE for different lag periods of a single meteorological feature in an embodiment of the present invention. Figure 4a This is a marginal effect analysis diagram of R² for the combination of coarse-grained meteorological feature lag periods in an embodiment of the present invention; Figure 4b This is a schematic diagram of the optimal combination set of coarse-grained meteorological feature lag period combinations in an embodiment of the present invention; Figure 5a This is a graph showing the marginal effect of the temperature variation coefficient on model R² under coarse-grained search in an embodiment of the present invention. Figure 5b This is a graph showing the marginal effect of effective accumulated temperature on model R² under coarse-grained search in an embodiment of the present invention. Figure 5c This is a graph showing the marginal effect of effective cumulative wind speed on model R² under coarse-grained search in an embodiment of the present invention. Figure 5d This is a graph showing the marginal effect of the anterior rainfall index on model R² under coarse-grained search in an embodiment of the present invention. Figure 5e This is a graph showing the marginal effect of the cumulative surrogate value of water body stability on model R² under coarse-grained search in an embodiment of the present invention. Figure 6a This is a correlation diagram between the test set R² and CV_R² for the fine-grained meteorological feature lag period combination in this embodiment of the invention; Figure 6b This is a schematic diagram of the optimal combination set of fine-grained meteorological feature lag period combinations in an embodiment of the present invention; Figure 7a This is a graph showing the marginal effect of the temperature variation coefficient on model R² under fine-grained search in an embodiment of the present invention. Figure 7b This is a graph showing the marginal effect of effective accumulated temperature on model R² under fine-grained search in an embodiment of the present invention. Figure 7c This is a graph showing the marginal effect of effective cumulative wind speed on model R² under fine-grained search in an embodiment of the present invention. Figure 7d This is a graph showing the marginal effect of the anterior rainfall index on model R² under fine-grained search in an embodiment of the present invention. Figure 7e This is a graph showing the marginal effect of the cumulative surrogate value of water body stability on model R² under fine-grained search in an embodiment of the present invention. Figure 8a This is an R² distribution curve of the model performance under fine-grained search in an embodiment of the present invention; Figure 8b This is a CV R² distribution curve of the model performance under fine-grained search in an embodiment of the present invention; Figure 8c This is a cumulative R² distribution curve of the model performance under fine-grained search in an embodiment of the present invention; Figure 8d This is a cumulative distribution curve of CV R² for model performance under fine-grained search in an embodiment of the present invention; Figure 9 This is a bar chart showing the importance of SHAP aggregated features of the complete dataset (FD) in this embodiment of the invention. Figure 10 This is a bar chart showing the importance of SHAP features in the dataset (FD-ID & Sea) after removing spatiotemporal proxy variables in this embodiment of the invention. Figure 11 This is a SHAP bee colony graph of the complete dataset (FD) in this embodiment of the invention; Figure 12 This is the SHAP bee colony graph of the spatiotemporal proxy variable dataset (FD-ID&Sea) in an embodiment of the present invention; Figure 13 This is a schematic diagram of the composition principle of the Chla concentration prediction system based on lagging meteorological data in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0023] Figure 1 This is a flowchart of the Chla concentration prediction method based on lagging meteorological data in an embodiment of the present invention.

[0024] like Figure 1 As shown, S1 involves acquiring historical in-situ water quality monitoring data of the target water body and spatiotemporally matched meteorological reanalysis data. The water quality monitoring data includes at least chlorophyll a concentration, and the meteorological reanalysis data includes daily average temperature, daily average wind speed, and daily precipitation.

[0025] Furthermore, the process includes a pre-processing step for cleaning the water quality monitoring data: First, outliers and null values ​​that exceed the physically reasonable range are removed. Then, the monitoring timestamps are converted into seasonal and diurnal features, and the seasonal features, diurnal features, and the station number features corresponding to the monitoring points are subjected to one-hot encoding as categorical features.

[0026] S2. Based on meteorological reanalysis data, meteorological derivative features at multiple time scales are constructed by configuring different lag parameters. The constructed meteorological derivative features include the following five categories: temperature variation coefficient, effective accumulated temperature, effective accumulated wind speed, antecedent precipitation index, and cumulative surrogate value of water body stability.

[0027] Specifically, (1) TCV, Temperature Coefficient of Variation

[0028] in Standard deviation, The temperature over the past nth day, The TCV is the absolute value of the mean. It quantifies the relative fluctuation of temperature over the past n days, reflecting the stability of the thermal environment. Drastic temperature fluctuations can disrupt the stable growth of algal populations; a higher TCV value indicates a more unstable thermal environment and a stronger inhibitory effect on algal growth.

[0029] (2)EAT, Effective Accumulated Temperature

[0030] Using 288.15K (i.e. 15°C) as the baseline temperature, only the temperature contribution exceeding this threshold is counted. It refers to the front The average temperature of that day. 15°C is the empirical lower limit for the effective growth of cyanobacteria; below this temperature, the rate of photosynthesis and cell division of algae decreases significantly. EAT captures "available heat" more accurately than the original temperature average, effectively filtering out noise interference from low-temperature days.

[0031] (3) EAW, Effective Accumulated Wind Speed

[0032] Using 3 m / s as the threshold, only the portion exceeding this wind speed is accumulated. It refers to the front The average wind speed of that day. 3 m / s is the empirical critical wind speed that causes significant vertical mixing of water columns in shallow lakes. Light winds below this threshold have limited impact on the thermal stratification of water bodies. EAW filters out ineffective disturbances, retaining only the wind energy contribution that can truly disrupt thermal stratification and inhibit the upwelling and aggregation of cyanobacteria.

[0033] (4) API, Antecedent Precipitation Index

[0034] API is the exponentially weighted cumulative rainfall, with a decay coefficient K=0.90 to ensure that recent rainfall has a higher weight than long-term rainfall. That is, before The average daily rainfall. Rainfall carries non-point source nitrogen and phosphorus nutrients into the lake via surface runoff, providing a crucial material basis for algal blooms. The exponential decay weighting conforms to the watershed hydrological response pattern and more closely approximates the temporal dynamics of actual nutrient input than simple summation.

[0035] (5)ASP, Accumulated Stability Proxy

[0036] in Convert Kelvin temperature to Celsius. Wind speed represents the intensity of thermal drive. The ASP decreases exponentially with increasing wind speed, representing the inhibitory effect of wind on water body disturbance. The product of these two factors constitutes the daily stability surrogate value: a large positive value indicates stable thermal stratification of the water body, favorable conditions for cyanobacteria to rise and accumulate; a low temperature or high wind indicates a daily ASP value approaching zero or even becoming negative. The cumulative version, ASPn, comprehensively reflects the continuous accumulation effect of the water body's thermodynamic stability over the past n days, directly quantifying the meteorological suitability for algal blooms.

[0037] It should be noted that the baseline temperature for effective accumulated temperature (EAT), the threshold wind speed for effective accumulated wind speed (EAW), and the attenuation coefficient K for the anterior precipitation index (API) can all be adjusted specifically according to the physiological characteristics of different algae species and the hydrological characteristics of the target water body.

[0038] S3. Based on water quality monitoring data and constructed meteorological derivative features, a data-driven two-stage exhaustive search strategy is adopted to determine the optimal lag period combination for various meteorological derivative features.

[0039] Specifically, the two-stage exhaustive search strategy includes a coarse-grained search stage and a fine-grained search stage.

[0040] Coarse-grained search phase: At a set coarse-grained sampling time interval, exhaustively search all lag combination schemes within the preset full candidate lag period range, independently train and evaluate the performance of the prediction model; based on the search results, decompose the overall performance of multi-feature coupling into independent response curves of single factors, and identify the optimal lag period range in which various meteorological derivative features can enable the model performance to reach a high level. Fine-grained search phase: Within the lag period range of each feature selected in the coarse-grained search, a full combinatorial exhaustive verification is performed again with a smaller time step. By comparing the prediction accuracy and cross-validation generalization ability of each combination within the performance saturation region, and considering both the peak marginal effect and statistical robustness, the optimal lag period combination is finally determined. The smaller time step in the fine-grained search phase is 1 day.

[0041] Furthermore, in scenarios with limited computing resources, the fine-grained exhaustive search step can be omitted, and the median of the optimal interval of each feature obtained from the marginal effect analysis in the coarse-grained search can be directly used as the lag parameter, which can still approach the optimal performance.

[0042] S4. Extract the corresponding meteorological features based on the optimal lag period combination, and merge them with water quality monitoring data to construct a feature matrix. Input the feature matrix into the selected machine learning model for training to construct a Chla concentration prediction model, and use the trained Chla concentration prediction model to predict the Chla concentration at future times.

[0043] The aforementioned machine learning model is any one of the following ensemble learning models: Extremely Random Tree Model, Random Forest Model, LightGBM Model, or XGBoost Model.

[0044] Furthermore, step S5 is included: using the SHAP method to perform interpretability analysis on the trained Chla concentration prediction model in order to analyze the marginal contribution of each environmental feature to the Chla concentration prediction and identify its ecological response pattern.

[0045] Specifically, S5.1, based on the trained machine learning model and the complete feature matrix, calculate the SHAP value of each feature to quantify the marginal contribution of each feature in each prediction, and obtain the global importance ranking of each environmental feature by aggregating the absolute values ​​of the SHAP of all samples. S5.2 For the categorical features in the input features that have undergone one-hot encoding, verify and eliminate the information preemption effect.

[0046] The steps to verify and eliminate the information preemption effect include: The first model is trained based on the complete feature matrix; at the same time, a second feature matrix is ​​constructed to remove all uniquely thermally encoded categorical features, retaining only the continuous original water quality parameters and meteorological derived features, and the second model is trained accordingly. SHAP analysis was performed on the first model and the second model respectively to compare the feature importance ranking structures of the two models. If the SHAP importance of the one-hot encoded categorical feature in the first model is higher than that of other continuous features, and in the ablation control experiment, the performance of the second model after removing the categorical feature decreases by less than a preset threshold compared to the first model, then it is determined that there is an information preemption effect, and the SHAP analysis results of the second model are used as the basis for identifying the ecological driving mechanism.

[0047] S5.3 Based on the relationship between characteristic values ​​and SHAP values, environmental factors are classified into three response modes: positive driving, negative inhibiting, and mixed response, in order to identify the core driving factors of Chla concentration and their direction of influence.

[0048] Through the above steps, the core driving factors of Chla concentration and their direction of influence can be objectively identified, providing a basis with clear physical significance for algal bloom early warning and ecological regulation decisions.

[0049] The solution of the present invention will now be described in detail with reference to specific application examples.

[0050] Study Area Overview: Taihu Lake (30°55′N-31°32′N, 119°52′E-120°36′E) is located at the junction of southern Jiangsu Province and northern Zhejiang Province, with a surface area of ​​approximately 2338 km². 2 With an average depth of about 1.9 meters, Taihu Lake is a typical large, shallow, eutrophic lake. The Taihu Lake basin has a subtropical monsoon climate, with an average annual temperature of about 15-17℃ and an average annual precipitation of about 1100-1200 mm. Its shallow water characteristics make Taihu Lake extremely sensitive to meteorological disturbances; the resuspension of bottom sediment and mixing with the water column caused by wind and waves are important physical processes affecting the spatiotemporal distribution of water quality parameters. The dominant species in Taihu Lake's cyanobacterial blooms is *Microcystis*, which typically begins to appear in May and reaches its peak from July to September.

[0051] Water quality monitoring data: Sourced from the National Surface Water Quality Automatic Monitoring Real-time Data Release System, covering six national-level monitoring stations in the Taihu Lake Basin. The original dataset contains monitoring records from June 2021 to December 2024, covering the following parameters: Chla concentration (mg / L), station number (ID), monitoring time (TIME), water temperature (T, ℃), pH, dissolved oxygen (DO, mg / L), conductivity (EC, μS / cm), turbidity (Turb, NTU), permanganate index (PI, mg / L), ammonia nitrogen (NH3-N, mg / L), total phosphorus (TP, mg / L), and total nitrogen (TN, mg / L).

[0052] Meteorological data: sourced from the China Meteorological Forcing Dataset CMFD v2.0 (1951-2024), with a spatial resolution of 0.1°. Daily average temperature (K), daily average wind speed (m / s), and daily precipitation (mm) corresponding to water quality monitoring points and times were extracted.

[0053] Data preprocessing and quality control: After the following quality control process, 14,026 complete data sets were finally retained: (1) outliers that were clearly outside the physical reasonable range were removed; (2) data sets with null values ​​were removed.

[0054] In addition, the monitoring timestamps were converted into two categorical features: Season (divided into four seasons according to climatological standards: spring, summer, autumn, and winter) and Day-Night (defined by seasonal differences for daytime and nighttime periods), and processed using one-hot encoding. The station ID was also processed using one-hot encoding.

[0055] Based on meteorological reanalysis data, five types of meteorological derived features are constructed according to the formula described in step S2 above. The candidate lag period range for each feature is set as follows: TCV is from 1 to 28 days prior (28 candidate values ​​in total), and EAT, EAW, API, and ASP are from the current day (day 0) to 28 days prior (29 candidate values ​​each).

[0056] Before conducting the search for the optimal lag combination, it is first necessary to verify whether the lag effect of meteorological derivative features truly exists and has a substantial impact on model performance. This embodiment systematically evaluates the trend of the predictive performance of each meteorological derivative feature acting alone as a function of lag days.

[0057] The specific method is as follows: For the different lag periods of TCV (lag of 1-28 days), EAT, EAW, API, and ASP (lag of 0-28 days), the corresponding individual meteorological derivative features are extracted one by one, and they are merged with the original water quality parameters to construct feature matrices. The Extra Trees model is trained, and R², cross-validation R², RMSE, and MAE are recorded for each lag period.

[0058] like Figure 2 As shown, the trends of R² and CV_R² for five types of meteorological derivative characteristics under different lag periods are illustrated. Figure 2 It is evident that the impact of meteorological conditions on Chla concentration is not instantaneous, but rather exhibits a clear lag response pattern. Specifically, for cumulative indicators (ASP, EAT, EAW), predictive performance improves with increasing lag time, with the R² curve showing an upward trend and reaching its optimal level after approximately two weeks; for fluctuation-sensitive indicators (TCV, API), R² reaches its peak within a shorter lag window, with the curve showing a pattern of first rising and then falling.

[0059] This result indicates that the timescales of the impact of different ecological processes on algal growth vary significantly. The effects of thermal accumulation and wind disturbance require a longer time to fully manifest, while nutrient input pulses from temperature fluctuations and rainfall produce responses in a shorter period of time.

[0060] Figure 3 The above conclusions were further verified. RMSE and MAE decreased synchronously with the extension of lag days, with the error reduction of ASP being the most significant. RMSE decreased from a relatively high level at lag of 0 days to its lowest level near the optimal lag period, with a change range much larger than that of random fluctuations, proving that the signal of the lag effect is real and not a random phenomenon caused by data noise.

[0061] The above results confirm that explicitly encoding meteorological lag effects can effectively compensate for information gaps caused by monitoring frequency or single-time-section data, enabling the model to capture precursors of algal blooms. Meanwhile, the optimal lag times for different meteorological derivative features differ significantly—ASP, EAT, and EAW require longer lag times (approximately 2-4 weeks), while API and TCV require shorter lag times (approximately 1 week)—providing a direct basis for the necessity of subsequent searches for optimal lag combinations for multiple features. Applying a uniform fixed lag time to all meteorological features would fail to fully capture the true impact of various meteorological factors, leading to impaired model performance.

[0062] The above two-stage exhaustive search determines the optimal lag combination, including a coarse-grained global search stage and a fine-grained global search stage.

[0063] Coarse-grained global search: Based on the results of univariate lag response analysis, a coarse-grained exhaustive search of all combinations is conducted. Specific candidate lag periods are set as follows: EAT, EAW, and ASP: 11 candidate lag periods each (1, 4, 7, 10, 13, 16, 19, 21, 23, 24, 25 days); API: 9 candidate lag periods (1, 3, 5, 7, 9, 12, 15, 18, 22 days); TCV: 6 candidate lag periods (2, 5, 8, 11, 14, 17 days).

[0064] The above settings resulted in a total of 71,874 parameter combinations, achieving balanced sampling across different ecological time scales while balancing computational cost and parameter spatial coverage.

[0065] The evaluation process for each combination is as follows: (1) Extract the corresponding meteorological derivative feature columns according to the specified lag period; (2) Combine with water quality parameters and time-coded features to construct a feature matrix; (3) Randomly divide the training set and test set at an 8:2 ratio; (4) Perform 3-fold cross-validation on the training set; (5) Evaluate R², RMSE and MAE on the test set.

[0066] Figure 4aThe results show a highly significant positive correlation between R² and CV_R² on the test set, with no outliers indicating high R² and low CV_R² on the test set. This strongly suggests that overfitting did not occur during the full search optimization process—the improvement in model performance is based on the simultaneous improvement of generalization ability, rather than being driven by overfitting to the test set. Figure 4b The top 5 optimal combinations are concentrated in the upper right corner of the Pareto front, achieving the lowest RMSE (<0.0049 mg / L) and MAE (<0.0027 mg / L), confirming a synergistic optimization trend in the core evaluation indicators. The top-ranking combinations are not isolated outliers, but rather distributed within a relatively dense band of the high-performance region, indicating a relatively stable optimal structure within a certain range for the optimal lag period.

[0067] To reduce the interference of multi-feature coupling on the interpretation of results, a marginal effect analysis was further performed on the coarse-grained full search results. Specifically, under the global combination context of fixed lag values ​​for other features, the average R², fluctuation range, and theoretical maximum R² corresponding to different candidate lag times for a single feature were statistically analyzed, thereby decomposing the overall combined performance into independent response curves of individual factors.

[0068] Figure 5a to Figure 5e The marginal effect analysis results show that the five meteorological characteristics can be roughly divided into two time scales: Long-term cumulative effects (ASP, EAT, EAW): ASP: The average R² curve peaks at 25 days (R²=0.9130), and an ASP lag of 25 days occurs 138 times (69%) in the Top 200 combinations, thus defining 22-32 days as the fine-grained exhaustive range; EAT: The average R² peaks at 19 days (R²=0.9104), with the optimal range concentrated in 15-21 days. A 19-day lag occurs most frequently (21%) in the Top 200 combinations, defining the fine-grained exhaustive range as 15-20 days; EAW: The average R² peaks at 19 days (R²=0.9110), and lags of 16 and 19 days account for 60% of the cumulative total in the Top 200, defining the fine-grained exhaustive range as 15-20 days.

[0069] Short-term disturbance effects (API, TCV): API: The average maximum R² is 5 days (R²=0.9099), with the optimal lag concentrated between 4-9 days. API lags of 5-9 days account for 50% of the Top 200 combinations, limiting the fine-grained exhaustive range to 4-10 days. TCV: The average peak R² is at 11 days (R²=0.9101), with the optimal range concentrated between 8-11 days. TCV lags of 8 and 11 days account for over 50% of the Top 200 combinations, limiting the fine-grained exhaustive range to 7-12 days.

[0070] Fine-grained local validation: Based on the intervals locked by the coarse-grained search, a second round of exhaustive validation with a step size of 1 day was performed: API: 4-10 days (7 candidate values); TCV: 7-12 days (6 candidate values); EAT: 15-20 days (6 candidate values); EAW: 15-20 days (6 candidate values); ASP: 22-32 days (11 candidate values). A total of 16,632 parameter combinations were formed.

[0071] Figure 6a The test set R for fine-grained exhaustive search is shown. 2 The distribution range is 0.912-0.920, CV_R 2 The distribution range is 0.867–0.876. Compared to the coarse-grained experiment (R... 2 : 0.886-0.920; CV_R 2 (0.829-0.880), significantly improved performance lower limit – test set R 2 The lower bound is improved by approximately 2.9%, CV_R 2 The lower limit is improved by about 4.6% - while the upper limit remains basically the same (both are about 0.920). Figure 6b This also reflects that points with higher R² values ​​have lower RMSE and MAE values, indicating better overall performance. Furthermore, this reflects a general trend; the five points marked in the figure represent the five highest R² values, further demonstrating this result. This result rigorously validates the effectiveness of the coarse-grained interval locking strategy: the high-performance intervals identified through marginal effect analysis successfully filtered out a large number of low-performance combinations, ensuring that all solutions enumerated through fine-grained exhaustive search were within the performance saturation zone.

[0072] More importantly, R of fine-grained point clouds 2 The standard deviation decreased from 0.008 to 0.002, and the distribution bandwidth narrowed significantly. This indicates that within the reduced parameter space, the performance differences between different combinations tend to converge, and the optimal lag interval locked by coarse granularity has statistical robustness, rather than random fluctuations in the data-driven process.

[0073] Figure 7a to Figure 7e This demonstrates the impact of various meteorological features on model R under fine-grained search. 2 Marginal effect analysis. Comprehensive analysis, taking the mean R of each characteristic... 2 The combination of lag periods corresponding to the peak values ​​was determined as the final modeling scheme: API6+TCV9+EAT18+EAW19+ASP25.

[0074] The performance of this combination in fine-grained exhaustive search is as follows: Test set R 2 =0.9196, RMSE=0.00484mg / L, MAE=0.00261mg / L, cross-validation CV_R 2 =0.8760.

[0075] Figure 8a , Figure 8b , Figure 8c and Figure 8d The frequency distribution and cumulative distribution characteristics of this combination are further illustrated. Among them, Figure 8a The corresponding frequency distribution histogram shows that the R of the recommended scheme 2 (0.9196) is located in the high-value dense area of ​​the distribution (0.918-0.920), which contains about 25% of the combinations (about 4000 groups), and is far from an isolated extreme point. Figure 8b The corresponding cumulative distribution curve shows that the recommended solution outperformed 89.3% of the combinations, while the top-ranked combination (R0) was the best. 2 =0.9202) is only slightly higher than 91.7% - the difference between the two in statistical quantiles is only 2.4%, but the recommended scheme has the advantage in cross-validation performance and ecological mechanism interpretability.

[0076] From an ecological mechanism perspective, the timescales of each feature in this optimal combination correspond to well-defined ecological processes: ASP25 represents approximately 3.5 weeks of water stability memory, EAT18 and EAW19 reflect approximately 2.7 weeks of thermodynamic-dynamic cumulative effects, and API6 and TCV9 characterize approximately 1 week of exogenous inputs and thermal environmental disturbances. This scheme, while maintaining high predictive performance, has a strong physical and ecological explanatory foundation.

[0077] Before constructing the complete feature set, six machine learning algorithms were benchmarked based on the original water quality dataset (OD). The results are shown in Table 1: Extra Trees achieved the best performance (R²). 2 =0.8034, CV_R 2 =0.7627, RMSE=0.0076), outperforming Random Forest (R 2 =0.7707) and XGBoost (R 2 =0.7492), and therefore it was selected as the baseline model for subsequent experiments.

[0078] Meteorologically derived features were extracted using the optimal combination API6+TCV9+EAT18+EAW19+ASP25, and a complete feature matrix FD was constructed. To quantify the independent contribution of each feature, a multi-level ablation experiment was designed, including FD, FD-Met (removing meteorological features), FD-ID (removing stations), FD-Sea (removing seasons), FD-DN (removing diurnal variation), and OD (only water quality parameters), which were compared and evaluated on the ExtraTrees model. The results of the ablation experiments have been described above: meteorologically derived features alone contribute 5.6% of the R². 2 Incremental spatial features (ID) contribute the most significantly (R after removal)2 (The decrease was approximately 1.4%), with time features contributing relatively little, validating the actual contribution of each feature group to the model performance.

[0079] Table 1. Comparison of Chl-a prediction performance of Extra Trees model under different feature sets.

[0080] Based on the trained Extra Trees model and the complete feature matrix FD, the SHAP value of each feature is calculated. Figure 9 The importance ranking of SHAP aggregated features in the FD model is presented. The results show that site number and seasonal features account for extremely high attribution weights in the FD model (cumulative contribution of over 27%).

[0081] Following the verification process described in step S5.2 above, a determination is made as to whether an information preemption effect exists.

[0082] Step 1: Construct the comparison model. The first model uses FD (containing all one-hot encoded features), and the second model uses FD-ID&Sea (removing one-hot encoding of site ID and one-hot encoding of season).

[0083] Step 2: Performance Comparison. The performance of the first model (FD) is: R 2 =0.9196, RMSE=0.0048mg / L. The performance of the second model (FD-ID & Sea) is: R 2 =0.9013, RMSE=0.0054mg / L. Both R 2 The difference is only about 1.8 percentage points, indicating that the model's prediction performance does not decrease significantly after removing one-hot encoded features.

[0084] Step 3: SHAP Comparison and Preemption Effect Determination. In the SHAP importance ranking of the first model (FD), the cumulative attribution weight of site ID and season exceeded 27%, while the weights of continuous environmental factors (such as ASP25, EAT18, pH, etc.) were compressed into a narrow range, and directional signals were mixed. One-hot encoded categorical features dominated SHAP attribution, but ablation experiments showed that removing them resulted in only a small decrease in model performance (R0). 2 The asymmetry between the two (which decreased from 0.9196 to 0.9013) confirms the existence of an information preemption effect.

[0085] Step 4: Determine the analysis strategy. Given the information preemption effect, the SHAP analysis results of the second model (FD-ID & Sea) are used as the primary basis for identifying the ecological driving mechanism. Figure 10The importance of SHAP aggregated features for the FD-ID & Sea model is ranked; the SHAP results of the first model are used to evaluate the marginal contribution of each feature to the prediction accuracy.

[0086] Following the method described in step S5.3 above, the SHAP value of the FD-ID&Sea model is visualized and analyzed using a bee colony graph. In the bee colony graph, the horizontal axis represents the SHAP value (positive values ​​increase the Chla prediction value, and negative values ​​decrease it), and the colors from light to dark indicate that the feature value increases from low to high.

[0087] Figure 11 This is the SHAP beehive diagram of the FD model. Under the masking of one-hot encoded features such as ID and Season, continuous environmental factors exhibit a mixed red and blue distribution with compressed intervals. Although values ​​such as pH and ASP25 show a trend of red (high values) slightly to the right and blue (low values) slightly to the left, the distribution interval is compressed to a small range, and the threshold features are not significant. The red and blue dots of values ​​such as EAT18, TN, and TP are highly mixed near the zero value, and the directionality is almost indistinguishable.

[0088] Figure 12 This is the SHAP beehive diagram of the FD-ID&Sea model. After removing one-hot encoded features, all 15 environmental factors are continuous variables, and the color gradients accurately reflect the feature values, presenting the following three clearly identifiable ecological response patterns: (1) Positive-driven type: including pH, PI, ASP25, EAW19, TN, EC and TP. pH and PI show a typical monotonic positive relationship - high values ​​are concentrated in the positive SHAP area (right side) and low values ​​are concentrated in the negative SHAP area (left side), reflecting the strong accompanying effect of algal photosynthetic carbon fixation consuming CO2 leading to pH increase and the high coupling between algal biomass and endogenous organic matter cycle. ASP25 and EC show a significant high-value amplification effect, and the positive SHAP on the right side has a clear long tail, suggesting that when the cumulative water stability index or ion concentration exceeds the critical value, the promoting effect on Chla will be sharply enhanced, which is highly consistent with the threshold burst characteristics of cyanobacterial blooms.

[0089] (2) Reverse suppression type: mainly including EAT18 and API6. High-value samples are mostly concentrated in the negative SHAP region (left side), indicating that excessively high cumulative effective temperature or previous rainfall will suppress the increase of Chla concentration, which may be related to the light suppression caused by high temperature and the dilution effect of rainfall.

[0090] (3) Mixed response type: including EAW19, T, DO, TCV9, NH3-N, Day-Night, and Turb. The SHAP values ​​of these factors are mostly clustered near zero and have mixed red and blue dots, indicating that their effects are highly context-dependent and are synergistically regulated by other environmental factors. For example, EAW19 did not show a simple strong wind suppression effect, which may reflect the dual effect of wind field on water mixing and sediment nutrient release; NH3-N exhibits a nonlinear characteristic of low-promotion and high-inhibition, showing obvious toxic inhibition at high concentrations.

[0091] Based on the SHAP analysis, the five meteorological derivative characteristics contributed a total of 37.06% of the model's explanatory power, exhibiting clear temporal scale differentiation: long-term cumulative effects (ASP25, EAT18, EAW19, lag period 18-25 days) dominated the algal bloom formation process, while short-term disturbance effects (TCV9, API6, lag period 6-9 days) only played a moderating role. Given the nutrient saturation of Taihu Lake, cumulative meteorological conditions represented by ASP25, EAT18, and EAW19 are key triggering factors regulating cyanobacterial blooms—a sustained period of 3-4 weeks of high temperature and weak winds is the key triggering condition for cyanobacterial bloom outbreaks.

[0092] Furthermore, this invention is not limited to the aforementioned application scenarios. The method framework of this invention can be directly extended to Chla concentration prediction modeling for other lakes, reservoirs, rivers, and even nearshore sea areas with the conditions for obtaining water quality monitoring and meteorological reanalysis data. Only the corresponding regional data needs to be replaced, and the two-stage exhaustive search needs to be re-executed to determine the optimal lag period combination applicable to the water body. The method framework of this invention is not limited to Chla concentration prediction; it can also be extended to the prediction of other water quality parameters affected by multiple time-scale meteorological factors, such as the dynamic prediction of dissolved oxygen and turbidity.

[0093] Figure 13 This is a schematic diagram of the composition principle of the Chla concentration prediction system based on lagging meteorological data in an embodiment of the present invention.

[0094] This invention also provides a Chla concentration prediction system based on lagging meteorological data. For example... Figure 13 As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a Chla concentration prediction method based on lagging meteorological data as described above.

[0095] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and therefore will not be described in detail here.

[0096] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.

[0097] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.

Claims

1. A method for predicting Chla concentration based on lagged meteorological data, characterized in that, include: S1. Obtain historical in-situ water quality monitoring data of the target water body and meteorological reanalysis data that matches its temporal and spatial characteristics; The water quality monitoring data includes at least chlorophyll a concentration, and the meteorological reanalysis data includes daily average temperature, daily average wind speed, and daily precipitation. S2. Based on the meteorological reanalysis data, meteorological derivative features at multiple time scales are constructed by configuring different lag parameters; S3. Based on water quality monitoring data and constructed meteorological derivative features, a data-driven two-stage exhaustive search strategy is adopted to determine the optimal lag period combination for various meteorological derivative features. The two-stage exhaustive search strategy includes: coarse-grained search stage: exhaustively search all lag combination schemes within the preset full candidate lag period range at a set coarse-grained sampling time interval, independently train and evaluate the performance of the prediction model; based on the search results, decompose the overall performance of multi-feature coupling into independent response curves of single factors, and identify the optimal lag period range in which various meteorological derivative features can enable the model performance to reach a high level. Fine-grained search stage: Within the lag period range of each feature optimization locked by the coarse-grained search, exhaustive verification of all combinations is performed again with a smaller time step. By comparing the prediction accuracy and cross-validation generalization ability of each combination in the performance saturation region, and taking into account the peak marginal effect and statistical robustness, the optimal lag period combination is finally determined. S4. Extract the corresponding meteorological derivative features based on the optimal lag period combination, and merge them with the water quality monitoring data to construct a feature matrix. Input the feature matrix into the selected machine learning model for training to construct a Chla concentration prediction model, and use the trained Chla concentration prediction model to predict the Chla concentration at future times.

2. The Chla concentration prediction method based on lagged meteorological data according to claim 1, characterized in that, Step S1 also includes a pre-processing step of cleaning the water quality monitoring data: removing outliers and null values ​​that are outside the physically reasonable range; The monitoring timestamps are converted into seasonal and diurnal features, and the seasonal features, diurnal features, and the site number features corresponding to the monitoring points are processed by one-hot encoding to serve as categorical features.

3. The Chla concentration prediction method based on lagged meteorological data according to claim 1, characterized in that, In step S2, the constructed meteorological derived features include the following five categories: Temperature variation coefficient, effective accumulated temperature, effective accumulated wind speed, antecedent rainfall index, and cumulative proxy value of water body stability.

4. The Chla concentration prediction method based on lagging meteorological data according to claim 1, characterized in that, The smaller time step in the fine-grained search phase is 1 day.

5. The Chla concentration prediction method based on lagging meteorological data according to claim 1, characterized in that, In step S4, the machine learning model is any one of the following ensemble learning models: Extremely Random Tree Model, Random Forest Model, LightGBM Model, or XGBoost Model.

6. The Chla concentration prediction method based on lagged meteorological data according to claim 2, characterized in that, It also includes step S5: using the SHAP method to perform interpretability analysis on the trained Chla concentration prediction model, in order to analyze the marginal contribution of each environmental feature to the Chla concentration prediction and identify its ecological response pattern.

7. The Chla concentration prediction method based on lagging meteorological data according to claim 6, characterized in that, Step S5 specifically includes: S5.1 Based on the trained machine learning model and the complete feature matrix, calculate the SHAP value of each feature to quantify the marginal contribution of each feature in each prediction. By aggregating the absolute values ​​of the SHAP values ​​of all samples, obtain the global importance ranking of each environmental feature. S5.

2. For the categorical features in the input features that have undergone one-hot encoding, verify and eliminate the information preemption effect; S5.3 Based on the relationship between characteristic values ​​and SHAP values, environmental factors are classified into three response modes: positive driving, negative inhibiting, and mixed response, in order to identify the core driving factors of Chla concentration and their direction of influence.

8. The Chla concentration prediction method based on lagging meteorological data according to claim 7, characterized in that, The steps to verify and eliminate the information preemption effect include: The first model is trained based on the complete feature matrix; at the same time, a second feature matrix is ​​constructed to remove all uniquely thermally encoded categorical features, retaining only the continuous original water quality parameters and meteorological derived features, and the second model is trained accordingly. SHAP analysis was performed on the first model and the second model respectively to compare the feature importance ranking structures of the two models. If the SHAP importance of the one-hot encoded categorical feature in the first model is higher than that of other continuous features, and in the ablation control experiment, the performance of the second model after removing the categorical feature decreases by less than a preset threshold compared to the first model, then it is determined that there is an information preemption effect, and the SHAP analysis results of the second model are used as the basis for identifying the ecological driving mechanism.

9. A Chla concentration prediction system based on lagged meteorological data, characterized in that, include: processor; A memory storing computer program instructions, which, when executed by the processor, implement the Chla concentration prediction method based on lagging meteorological data as described in any one of claims 1 to 8.