Lake and reservoir chlorophyll concentration prediction method based on SO-KNN model

By using a K-nearest neighbor regression model based on the snake optimization algorithm, the problem of data sparsity and imbalance in the prediction of chlorophyll a concentration in northern reservoirs was solved, achieving a high-precision and highly generalized early warning effect for algal blooms.

CN121834137APending Publication Date: 2026-04-10DALIAN UNIV OF TECH

Patent Information

Application Number
CN202610091328.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing models for predicting chlorophyll a concentration in northern reservoirs suffer from poor fitting, overfitting, weak generalization ability, and insensitivity to predicting critical points for algal bloom risk due to low monitoring data frequency, limited total historical sample size, and highly unbalanced data distribution.

Method used

We adopted a K-Nearest Neighbor Regression (KNN) model based on the Snake Optimization Algorithm (SO). Through multi-source data acquisition and preprocessing, we combined Spearman rank correlation coefficient analysis to screen key feature factors, and used the Snake Optimization Algorithm to optimize the number of nearest neighbors in the KNN model to construct an SO-KNN prediction model, thereby enhancing the sensitivity of identifying algal bloom risks.

Benefits of technology

It significantly improves the high accuracy and strong generalization prediction ability of chlorophyll a concentration under sparse and unbalanced data conditions, effectively overcomes overfitting and insufficient generalization ability, and achieves accurate early warning of algal blooms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834137A_ABST
    Figure CN121834137A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of water environment monitoring and early warning, and discloses a lake and reservoir chlorophyll a concentration prediction method based on an SO-KNN model. According to the invention, multi-time scale meteorological cumulative effect features are introduced to enrich information representation, and an SO-KNN intelligent prediction model is constructed. According to the method, under the conditions of data scarcity and non-equilibrium, the chlorophyll a concentration, especially the high-precision and strong-generalization prediction capability of the water bloom risk critical point, is remarkably improved. The model is simple in structure and efficient in calculation, the common defects of overfitting, insufficient generalization ability and the like of a complex machine learning model in the scene are effectively overcome, and a reliable and practical innovative technical solution is provided for early water bloom warning of northern reservoirs and water areas with similar data conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of water environment monitoring and early warning technology, and relates to a method for predicting chlorophyll a concentration in lakes and reservoirs based on an improved K-nearest neighbor regression (KNN) model using the Snake Optimization (SO) algorithm. Background Technology

[0002] Reservoirs, as important sources of drinking water, industrial and agricultural water, and components of ecosystems, are crucial to national welfare and social stability. Chlorophyll a concentration is a core indicator of phytoplankton biomass in water bodies and a critical parameter for assessing eutrophication and algal bloom risk. Therefore, accurate and timely prediction of reservoir chlorophyll a concentration is of great practical significance for early warning of algal blooms, ensuring water supply security, and implementing effective water environment management.

[0003] The variation of chlorophyll a concentration in water bodies is a complex, nonlinear, dynamic ecological process driven by a combination of water quality and meteorological factors. These influencing factors include not only key water quality parameters such as water temperature, pH, dissolved oxygen, total nitrogen, total phosphorus, ammonia nitrogen, and permanganate index, but also external meteorological conditions such as sunshine duration, wind speed, air temperature, and rainfall. These factors interact intricately, collectively determining the spatiotemporal distribution characteristics of chlorophyll a, making the construction of accurate mechanistic models exceptionally difficult.

[0004] To address the challenge of predicting chlorophyll a concentration, scholars both domestically and internationally have developed and applied various prediction models, which can be broadly categorized into two types: process-based mechanistic models and data-driven models. Process-based mechanistic models, grounded in hydrodynamic and biogeochemical principles, simulate the intrinsic processes of algal growth and reproduction through differential equations. While they can reveal the mechanistic laws governing water quality dynamics, they require a large amount of topographic, hydrological, and biochemical parameters, placing extremely high demands on data integrity. Furthermore, the modeling process is complex and computationally expensive.

[0005] In recent years, with advancements in data acquisition technology and the development of machine learning theory, data-driven models based on historical monitoring data have become the mainstream method for predicting chlorophyll a concentration. These methods bypass complex biochemical processes, directly learning and mining the mapping relationship between input factors and chlorophyll a concentration from the data. Currently, the widely used machine learning models in this field fall into two main categories: single models and fusion models.

[0006] A single model refers to using a single, structurally unified algorithm or model architecture to complete the prediction task. Patent (CN109726857A) discloses an early warning method for algal bloom levels based on a temporal random forest model (Ordinal Forests). This method utilizes time-series data obtained from an online water quality and aquatic ecosystem monitoring system to construct an ordered classification prediction model, enabling high-precision classification and early warning of algal bloom risk levels for the next day. However, this model relies on high-frequency online monitoring data, and hyperparameters need to be manually set. The literature (Huang Cheng, et al. "Research on Chlorophyll a Concentration Prediction Model in Beijiang River Based on BP Artificial Neural Network." Sichuan Environment 44.01(2025): 109-115.) reports a real-time prediction model for chlorophyll a concentration based on a BP neural network. Using real-time monitoring water quality data as input factors, it achieved a good fitting effect (R²=0.892). However, this model also relies on real-time monitoring data, and the network structure is prone to getting trapped in local optima.

[0007] To improve the predictive performance and generalization ability of models, fusion models are widely adopted, which combine multiple single models through strategies to achieve a "1+1>2" effect. For example, the literature Ouyang Tian et al. "Research on Online Algal Temporal Data Prediction Based on LSTM Network: A Case Study of the Three Gorges Reservoir." Lake Science 33.04 (2021):1031-1042. reported on Long Short-Term Memory Network (LSTM) and its WT-LSTM model combined with wavelet transform, using hourly monitored chlorophyll a time-series data for advance prediction. The results showed that the prediction performance was further improved after wavelet transform preprocessing. LSTM models are good at capturing temporal dependencies and are suitable for short-term prediction in high-frequency monitoring scenarios, but they have high requirements for data quality and parameter tuning, and the computational cost is high. Patent (CN114386710A) proposes an STL-RF-LSTM fusion model, which uses temporal decomposition to extract the seasonality, trend term and residual of the sequence, and then uses random forest and LSTM for prediction respectively, and finally synthesizes the results. The model performs well in waters with relatively stable chlorophyll a concentrations (R²≈0.9), but its performance deteriorates significantly (<0.7) when concentrations fluctuate drastically. Furthermore, patent (CN109726857A) proposes a GA-Elman network, which uses a genetic algorithm to optimize the parameters of the Elman recurrent neural network, improving prediction accuracy by 5%~10%. However, its performance still relies on high-frequency monitoring data and converges slowly.

[0008] It should be noted that the above model is mainly applicable to the data conditions of eutrophic reservoirs in southern my country, and its successful application relies on two key foundations: first, high-frequency (hourly / daily) continuous monitoring data; and second, a relatively uniformly distributed dataset within the feature space, i.e., a complete sample containing chlorophyll a concentrations from low to high. However, the ecological background and data conditions of reservoirs in northern my country are different: the water bodies are in a mesotrophic or slightly eutrophic state for a long time, with low frequency and short duration of algal blooms; the water quality monitoring frequency is low, mostly 1-2 times per month by manual sampling; and the data distribution is extremely unbalanced, with the vast majority of samples containing low concentrations of chlorophyll a, and high-concentration samples being rare.

[0009] To address the problems of poor model fitting, overfitting, weak generalization ability, and insensitivity to predicting algal bloom risk thresholds in chlorophyll a concentration prediction in northern reservoirs due to low monitoring data frequency, limited historical sample size, and highly imbalanced data distribution, this invention provides a method for predicting chlorophyll a concentration in lakes and reservoirs based on the snake optimization algorithm-K nearest neighbor model (SO-KNN). This method aims to adaptively determine key parameters in the KNN model under sparse and imbalanced data conditions by utilizing the global optimization capability of the snake optimization algorithm, thereby constructing a regression prediction model based on geometric proximity in the data space. Through the optimized nearest neighbor decision mechanism, the model's sensitivity to minority class samples (algal bloom threshold states) is enhanced, effectively overcoming prediction bias caused by data imbalance and achieving high-precision, highly generalized chlorophyll a concentration prediction, particularly suitable for early warning of algal blooms in northern reservoirs. Summary of the Invention

[0010] The purpose of this invention is to overcome the problems of poor fitting effect, overfitting, poor generalization ability, and insensitivity to the prediction of critical points of algal bloom risk when existing models are applied to scenarios with "low monitoring data frequency, limited total amount of historical samples, and highly unbalanced data distribution". This invention provides a method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model. Through innovative data processing and model construction methods, it achieves accurate and reliable early warning of algal blooms.

[0011] The technical solution of the present invention:

[0012] A method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model includes the following steps:

[0013] Step S1: Acquisition and preprocessing of multi-source data;

[0014] Step S1.1: Water Quality Monitoring Data Acquisition and Preprocessing: Acquire water quality monitoring data from each sampling point during the historical monitoring period of the target lake / reservoir. The water quality monitoring data includes sampling date, water temperature, pH value, dissolved oxygen, total nitrogen, total phosphorus, ammonia nitrogen, permanganate index, and chlorophyll a concentration. Preprocess the water quality monitoring data, including: removing abnormal and duplicate samples to obtain valid samples; using a chain equation-based multiple interpolation method to fill missing values ​​in the valid samples; for data with concentrations below the limit of detection (LOD), using LOD / Substitution is performed to obtain a water quality characteristic set;

[0015] Step S1.2: Acquisition of meteorological monitoring data and construction of meteorological feature sets;

[0016] Step S1.2.1: Meteorological monitoring data acquisition: Acquire the daily average meteorological monitoring data of each sampling point in the valid sample on the sampling day and the previous 14 days. The daily average meteorological monitoring data includes meteorological factors such as daily average temperature, rainfall, and wind speed.

[0017] During the acquisition process, for sampling points with direct monitoring data, actual observation data is used; for meteorological factors or certain dates without direct monitoring data, meteorological data for the corresponding time period of the target sampling point is generated based on observation data from multiple surrounding meteorological stations, ensuring that each sampling point has a complete daily average meteorological data sequence.

[0018] Step S1.2.2: Calculation of lagged meteorological characteristics: For each type of meteorological factor, calculate the moving average of each sampling point 1 day, 2 days, 3 days, 5 days, 7 days, 10 days and 14 days before the sampling date to obtain the lagged meteorological characteristics;

[0019] Step S1.2.3: Construction of meteorological feature set: Integrate the meteorological monitoring data of each sampling point on the sampling day with the lagged meteorological features to obtain the meteorological feature set;

[0020] Step S1.3: Construction of Time Feature Set: Extract the sampling year and sampling month information of each valid sample, and encode the sampling month into a continuous variable through sine-cosine transform. The specific encoding formula is as follows:

[0021] Monthly Sine Characteristics

[0022] Monthly cosine characteristics ;

[0023] The sampling year, the sine feature of the encoded month, and the cosine feature of the month are used as time features to construct a time feature set;

[0024] Step S2: Key feature factor selection and complete dataset construction;

[0025] Step S2.1: Screening of key feature factors: Using Spearman's rank correlation coefficient analysis, the correlation coefficients and corresponding p values ​​between each factor in the water quality feature set and the meteorological feature set and chlorophyll a concentration were calculated respectively. Feature factors with p values ​​< 0.01 (significant correlation) and high absolute values ​​of correlation coefficients were screened as key feature factors.

[0026] Step S2.2: Complete dataset construction: Integrate the key feature factors obtained from the screening with the time feature set constructed in step S1.3 to form the input feature set; integrate the input feature set with the target value of chlorophyll a concentration of the corresponding sample to obtain the complete dataset;

[0027] Step S3: Construction and training of the SO-KNN model;

[0028] Based on the complete dataset obtained in step S2, a K-Nearest Neighbor Regression (KNN) prediction model (SO-KNN) fused with the Snake Optimization Algorithm (SO) is constructed and trained. All key feature factors of the input feature set in step S2 are used as input variables, and chlorophyll a concentration is used as the output variable (prediction target). Specifically, the following steps are included:

[0029] Step S3.1: Data standardization: Normalize the input and output variables to obtain a normalized complete dataset; the normalization process uses the min-max normalization method to map the data to the [0,1] interval, and records the normalization parameters corresponding to each key feature factor and chlorophyll a concentration for subsequent inverse normalization.

[0030] Step S3.2: Time series dataset partitioning: Sort the normalized complete dataset according to the monitoring time, select the samples from the most recent year as the independent time extrapolation test set to simulate the prediction scenario of unknown future data; use all historical samples from before this year as the time training set for the construction and training of the SO-KNN model.

[0031] Step S3.3: Optimal nearest neighbor number K_opt based on SO optimization KNN model, the steps are as follows:

[0032] Step S3.3.1: Parameter Setting and Population Initialization: Preset the search range [K_min, K_max] for the optimal nearest neighbor number K (usually a positive integer between [1, 30]), and initialize the snake population X = {X1, X2, X3, ..., X...} based on this range. n}, where each individual X i This represents a candidate K value, where the individual's position is within the predefined search range of the optimal nearest neighbor number K. The population size N is set to 20-50.

[0033] Step S3.3.2: Fitness assessment: For each individual Xi For the corresponding candidate K values, a K-nearest neighbor regression model is constructed, and time-series cross-validation is used on the training set (the training set is divided into k folds in chronological order, with k-1 folds as the training subset and 1 fold as the validation subset, repeated k times) to calculate the root mean square error (RMSE) of the K-nearest neighbor regression prediction model. The negative value of the RMSE is defined as the fitness F of the individual. i This yields the fitness set F = {F1, F2, F3, ..., F...} n The optimization objective is to maximize fitness (i.e., minimize the verification RMSE).

[0034] Step S3.3.3: Adaptive Iterative Optimization Based on Environmental Awareness: By simulating the adaptive behavior of snake groups to environmental changes, the population is driven to perform iterative optimization. The optimization process is carried out in a loop within a set maximum number of iterations, max_iter. Let the current iteration number be t. The population is divided into two subgroups: male and female.

[0035] Environmental parameter updates: Based on the iteration progress, calculate the current environmental simulation parameters Temperature factor Temp and Food Abundance factor Q. Both are variables that decrease with the iteration process, and their values ​​remain within the range of [0, 1].

[0036] Behavioral pattern switching and location update: The population's behavioral pattern is determined based on environmental parameters, and a new candidate location X_new for each individual is calculated. When the food abundance factor Q is lower than the set food abundance threshold Q_th, a global exploration mode is executed; when the food abundance factor Q is not lower than the set food abundance threshold Q_th, a local development sub-mode is selected based on the relationship between the temperature factor Temp and the set temperature threshold T_th. The specific modes are as follows:

[0037] Global exploration mode (Q < Q_th): Individuals move to randomly selected individuals of the same sex in the population, and the step size is regulated by the exploration ability factor (Am or Af).

[0038] Local development mode (Q ≥ Q_th): When Temp > T_th, all individuals move towards the current optimal solution X_food; when Temp ≤ T_th, they compete with a preset competition probability p (usually 0.5) or reproduce with a probability of (1-p).

[0039] Boundary constraints and updates: Perform boundary checks and rounding on all new positions X_new, and recalculate the fitness using K_new, as shown in the following formula:

[0040] Here, round() is the rounding function, and clamp() is the boundary constraint function (takes K_min when X_new < K_min, takes K_max when X_new > K_max, otherwise takes X_new).

[0041] Reconstruct the K-nearest neighbor regression model using K_new and calculate -RMSE as the new fitness F_new. If F_new is better than the individual's historical best fitness, then update the individual's best position and fitness.

[0042] Global optimal tracking: Compare the updated fitness of all individuals in this iteration, find the individual with the highest fitness, and update its position to the new optimal solution X_food.

[0043] Iteration Termination and Output: Repeat step S3.3.3 until the maximum number of iterations is reached or the optimal solution shows no significant improvement after multiple iterations. The output is the global optimal solution X_best, which is rounded down to be used as the number of K nearest neighbors K_opt for optimization.

[0044] Step S3.4: Optimal Model Training: Using the optimal number of nearest neighbors K_opt obtained in step S3.3.3, initialize and train the K-nearest neighbor regression model, complete the model fitting on the time training set, and obtain the final SO-KNN prediction model;

[0045] Step S4: Model application and performance evaluation, the specific steps are as follows:

[0046] Step S4.1: Extrapolation prediction: Input the independent time extrapolation test set (only input variables) processed in step S3.2 into the SO-KNN prediction model trained in step S3.3 to obtain the normalized chlorophyll a concentration prediction value;

[0047] Step S4.2: Result Inverse Normalization and Comprehensive Performance Evaluation: Using the chlorophyll a concentration normalization parameters recorded in Step S3.1, the predicted values ​​obtained in Step S4.1 are inverse normalized to restore concentration values ​​with actual physical meaning. Multiple statistical indicators are calculated between the inverse normalized predicted values ​​and the actual measured chlorophyll a concentration values ​​in the test set to quantitatively evaluate model performance. Core evaluation indicators include:

[0048] Coefficient of determination (R²): Used to characterize the model's ability to explain changes in the target variable (chlorophyll a concentration). The closer its value is to 1, the higher the model's goodness of fit and the more consistent the predicted results are with the actual value's trend.

[0049]

[0050] Among them, y i For the true value, For predicted values, is the average of the true values, and N is the number of test samples.

[0051] Mean Absolute Error (MAE): The average value of the absolute error between the predicted and the true values. The smaller the MAE value, the higher the average prediction accuracy of the model.

[0052]

[0053] Root mean square error (RMSE): This measures the average deviation between the predicted and actual values. It is more sensitive to larger prediction errors, and the smaller the value, the higher the overall prediction accuracy of the model.

[0054]

[0055] The beneficial effects of this invention are as follows: Addressing the inherent challenges of low data frequency, sparse samples, and highly imbalanced distribution in monitoring data for lakes and reservoirs in northern China, this invention introduces multi-timescale meteorological cumulative effect characteristics to enrich information representation. Furthermore, it innovatively employs the Snake Optimization Algorithm (SO) to automate the core hyperparameters of the K-Nearest Neighbor Regression (KNN) model, constructing an SO-KNN intelligent prediction model. This method significantly improves the high accuracy and strong generalization prediction ability for chlorophyll a concentration, especially the critical point of algal bloom risk, even under conditions of data scarcity and imbalance. The model has a simple structure and high computational efficiency, effectively overcoming the common shortcomings of complex machine learning models in such scenarios, such as overfitting and insufficient generalization ability. It provides a reliable and practical innovative technical solution for early warning of algal blooms in northern reservoirs and similar water bodies with suitable data conditions. Attached Figure Description

[0056] Figure 1 This is the overall flowchart of the lake and reservoir chlorophyll a concentration prediction method based on the SO-KNN model of the present invention.

[0057] Figure 2 This is a heatmap showing the correlation analysis of various water quality and meteorological parameters of Reservoir A in Embodiment 1 of the present invention.

[0058] Figure 3 This is a schematic diagram of the SO-KNN model structure of the present invention.

[0059] Figure 4 This is a comparison curve of the prediction performance of the method of the present invention on the independent time extrapolation test set in Example 1.

[0060] Figure 5 This is a comparison curve of the prediction performance of the method of the present invention on the independent time extrapolation test set in Example 2.

[0061] Figure 6This is a comparison curve of the prediction performance of the method of the present invention on the independent time extrapolation test set in Example 3.

[0062] Figure 7 This is a comparison curve of the prediction performance of the method of the present invention on the independent time extrapolation test set in Example 4. Detailed Implementation

[0063] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.

[0064] Taking one river (River A) and two medium-sized reservoirs (Reservoir B and Reservoir C) and one large reservoir (Reservoir D) in northern my country as examples, these reservoirs share common characteristics such as low monitoring frequency (1-2 times per month), limited data volume, lack of directly monitored meteorological data, and uneven distribution of chlorophyll a concentration, further illustrating the specific implementation of the present invention. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0065] Overall process

[0066] The overall flowchart of the lake and reservoir chlorophyll a concentration prediction method based on the SO-KNN model in this embodiment of the invention is as follows: Figure 1 As shown. Please follow these steps:

[0067] Step S1: Multi-source data acquisition and preprocessing. Acquire water quality data from each sampling point during the historical monitoring period of the target water body, as well as monitoring data from surrounding meteorological stations. Water quality data includes, but is not limited to: water temperature (WT), pH, dissolved oxygen (DO), and permanganate index (COD). Mn The concentrations of ammonia nitrogen (NH3-N), total nitrogen (TN), total phosphorus (TP), and chlorophyll a (Chl-a) were measured. Meteorological data were obtained from surrounding meteorological stations, including three meteorological factors: daily average temperature (TEMP), daily cumulative rainfall (RAIN), and daily average wind speed (WIND). The specific preprocessing steps are as follows:

[0068] Step S1.1: Water Quality Data Preprocessing. Quality control is performed on the water quality data. First, obviously abnormal and duplicate samples are removed. Second, missing values ​​in the data are filled using a multiple imputation method based on chain equations. Finally, for indicators with concentrations below the method detection limit (LOD) (such as TN and TP), LOD / The substitution process is then performed. Following the above steps, valid water quality samples and their corresponding water quality feature sets are obtained.

[0069] Step S1.2: Construction of Meteorological Feature Set. For each valid water quality sample, based on the sampling point and sampling date, obtain the daily average meteorological data (TEMP, RAIN, WIND) from 3-4 surrounding meteorological monitoring stations for the sampling date and the preceding 14 days. Use inverse distance weighted (IDW) spatial interpolation to generate the required data, ensuring the sequence is complete. Based on this, systematically construct the meteorological features:

[0070] Multi-scale lag meteorological characteristics: To characterize the cumulative and lag effects of meteorological conditions, for each type of meteorological factor, the moving average value of the sampling date is calculated for 1, 2, 3, 5, 7, 10, and 14 days prior to the sampling date, forming multi-scale lag characteristics.

[0071] Integration: The meteorological observation values ​​of the sampling day are integrated with the various lag features calculated above to construct a meteorological feature set.

[0072] Step S1.3: Construction of the time feature set. Extract the sampling year and month information for each valid sample. Encode the sampling month into a continuous variable using a sine-cosine transform. The specific encoding formula is as follows:

[0073] Monthly Sine Characteristics

[0074] Monthly cosine characteristics ;

[0075] The sampling year, the sine feature of the encoded month, and the cosine feature of the month are used as time features to construct a time feature set.

[0076] Step S2: Key feature factor screening and complete dataset construction.

[0077] Step S2.1: Screening of key characteristic factors. Spearman's rank correlation coefficient analysis was used to calculate the correlation coefficients and significance (p-value) between each factor in the water quality characteristic set and the chlorophyll a (Chl-a) concentration, respectively. Characteristic factors that were significantly correlated with Chl-a (p-value < 0.01) and had high absolute values ​​of correlation coefficients were selected as key characteristic factors. Figure 2 An example heatmap of correlation analysis for Reservoir A in Example 1 is shown.

[0078] Step S2.2: Complete dataset construction: Integrate the key feature factors obtained from the screening with the time feature set constructed in step S1.3 to form the input feature set; integrate the input feature set with the target value of chlorophyll a concentration of the corresponding sample to obtain the complete dataset;

[0079] Step S3: Construction and Training of the SO-KNN Model. This step aims to construct and train a K-Nearest Neighbor Regression (KNN) prediction model (SO-KNN) that integrates the Snake Optimization Algorithm (SO) based on the complete dataset. Its structural diagram is shown below. Figure 3 As shown. The model uses all the key feature factors of the input feature set in step S2 as input variables and chlorophyll a concentration as the output variable (prediction target). Specifically, it includes the following steps:

[0080] Step S3.1: Data Standardization. For both input and output variables, use the Min-Max normalization method to map the data to the [0,1] interval. Record the normalization parameters for each variable for subsequent inverse normalization of the prediction results.

[0081] Step S3.2: Time Series Dataset Partitioning. The normalized complete dataset is sorted chronologically according to the monitoring time. All samples from the most recent consecutive year are selected as an independent time extrapolation test set to simulate prediction scenarios for unknown future data; all historical samples from before this year are used as the time training set for model construction and parameter optimization.

[0082] Step S3.3: Optimal nearest neighbor number K_opt based on SO optimization KNN model, the steps are as follows:

[0083] Step S3.3.1: Parameter Setting and Population Initialization: Set the search range of the optimal nearest neighbor number K to a positive integer [1, 30], and initialize the snake population X = {X1, X2, X3, ..., X...} based on this range. n}, where each individual X i This represents a candidate K value, where the individual's location is within the predefined search range of the optimal nearest neighbor number K. The population size N is set to 20-50; the maximum number of iterations Max_iter is 50; the food abundance threshold Q_th is set to 0.2; and the temperature threshold T_th is set to 0.5. The population is divided into a male group (top 50%) and a female group (bottom 50%).

[0084] Step S3.3.2: Fitness Evaluation: For each candidate K value, construct a K-nearest neighbor regression model, and calculate the root mean square error (RMSE) using time series cross-validation on the time training set. The negative value of the RMSE is taken as the fitness F. i The optimization objective is to maximize fitness (i.e., minimize validation error).

[0085] Step S3.3.3: Adaptive Iterative Optimization Based on Environmental Awareness. The SO algorithm simulates the foraging and reproductive behaviors of snake populations to drive population iteration. In each iteration, the algorithm adaptively switches between "global exploration" and "local exploitation" modes based on simulated environmental parameters such as the temperature factor Temp and the food abundance factor Q, updating the position and fitness of each individual in the population, and finding the global optimal solution X_best, which is rounded down to become the optimized K-nearest neighbor number K_opt.

[0086] Step S3.4: Optimal Model Training. Using the optimal number of nearest neighbors K_opt obtained in step S3.3.3, initialize the K-nearest neighbor regression model and train it on the entire time training set to obtain the final SO-KNN prediction model.

[0087] Step S4: Model application and performance evaluation, the steps are as follows:

[0088] Step S4.1: Extrapolation Prediction. Input the input variables (already normalized) of the independent time extrapolation test set into the SO-KNN prediction model trained in step S3.3 to obtain the normalized chlorophyll a concentration prediction value.

[0089] Step S4.2: Result Inverse Normalization and Performance Evaluation. Using the chlorophyll a concentration normalization parameters recorded in Step S3.1, the predicted values ​​are inversely normalized to restore them to concentration values ​​(μg / L) with actual physical meaning. Multiple statistical indicators, including the coefficient of determination (R²), mean absolute error (MAE), and root mean square error (RMSE), are calculated between the inversely normalized predicted values ​​and the actual observed values ​​in the test set to quantitatively evaluate the model performance.

[0090] Implementation Examples and Application Effects

[0091] Example 1: Monthly monitoring data from May to October each year (critical period of eutrophication) of River A from 2016 to 2024 were selected, covering 13 sampling points across 5 cross sections, with a total of 695 valid samples. The amount of data is relatively sufficient.

[0092] Correlation analysis identified 10 key input variables: daily average temperature (TEMP), average temperature of the previous 2 days (TEMP_avg2), average temperature of the previous 7 days (TEMP_avg7), average rainfall of the previous 14 days (RAIN_avg14), average wind speed of the previous 2 days (WIND_avg2), water temperature (WT), pH, and permanganate index (COD). Mn Chemical oxygen demand (COD).

[0093] Chl-a prediction was performed using 2024 data (78 samples) as an independent time test set and historical data as a time training set (617 samples).

[0094] like Figure 4 As shown, the prediction result R 2 =0.83, RMSE=5.12μg / L. This indicates that the model can accurately predict the trend of chlorophyll a concentration changes.

[0095] Example 2: Reservoir B is a medium-sized reservoir with two sampling points. From May to October each year from 2018 to 2022, monitoring was conducted twice a month, with a total of only 96 samples. This is the smallest amount of data among the four cases, making it extremely scarce.

[0096] Nine key input variables were selected, with meteorological cumulative characteristics dominating: daily average temperature (TEMP), average temperature of the previous 7 days (TEMP_avg7), average temperature of the previous 14 days (TEMP_avg14), average rainfall of the previous 10 days (RAIN_avg10), average rainfall of the previous 14 days (RAIN_avg14), average wind speed of the previous 2 days (WIND_avg2), average wind speed of the previous 14 days (WIND_avg14), water temperature (WT), and pH.

[0097] The 2022 data (12 samples) was used as the independent time extrapolation test set, and the data from 2018 to 2021 (84 samples) was used as the time training set.

[0098] The model performs well on the independent-time extrapolation test set, such as Figure 5 As shown, the R² is as high as 0.90, and the RMSE is as low as 2.44 μg / L. This fully demonstrates the strong adaptability and high prediction accuracy of the method of this invention under small sample conditions, and makes up for the lack of data by mining effective cumulative meteorological features.

[0099] Example 3: Reservoir C is also a medium-sized reservoir with two sampling points. A total of 106 samples were collected monthly from May to October 2016 to 2024. Its Chl-a concentration remained low for a long period, with only a few months showing abnormally high values, and the data distribution was extremely uneven.

[0100] Eleven key input variables were selected: average wind speed over the previous 3 days (WIND_avg3), average wind speed over the previous 7 days (WIND_avg7), average wind speed over the previous 14 days (WIND_avg14), water temperature (WT), pH, and permanganate index (COD). Mn Chemical oxygen demand (COD), five-day chemical oxygen demand (BOD5), total nitrogen (TN), and sulfate (SO42-) - ), nitrates (NO3-N).

[0101] Chl-a was predicted using 2024 data as an independent time extrapolation test set (12 samples) and historical data as a time training set (94 samples).

[0102] On such an imbalanced dataset, the model successfully captured a single concentration spike in the test set (see [link]). Figure 6 The overall R² reached 0.92, and the RMSE was 4.91 μg / L. This significantly demonstrates the sensitivity and predictive ability of the method of this invention for rare high-concentration events (algal bloom risk points), overcoming the defect of traditional models that tend to ignore high-value signals in such scenarios.

[0103] Example 4: Reservoir D is a large reservoir with a long and narrow area. Five sampling points are located in the reservoir, with a total of 302 samples collected from monthly monitoring data from 2016 to 2024. The spatial heterogeneity is significant, and the model needs to simultaneously handle the spatial location differences and the relationship between water quality and meteorology.

[0104] Fourteen key features were selected, and the sampling point ID was innovatively introduced as a spatial identifier feature, which was input together with other feature variables: sampling point ID, daily average temperature (TEMP), average temperature of the previous 7 days (TEMP_avg7), average temperature of the previous 14 days (TEMP_avg14), average rainfall of the previous 10 days (RAIN_avg10), average wind speed of the previous day (WIND_avg1), average wind speed of the previous 14 days (WIND_avg14), water temperature (WT), pH, and permanganate index (COD). Mn Chemical oxygen demand (COD), nitrate (NO3-N), iron (Fe), and manganese (Mn).

[0105] The 2024 data (35 samples, covering all locations) was used as the independent time extrapolation test set, and the historical data (267 samples) was used as the time training set.

[0106] Independent time extrapolation test set prediction performance as follows Figure 7 As shown, the model's overall prediction R² is 0.84 and RMSE is 4.20 μg / L across test samples from different spatial locations. These results demonstrate that the method of this invention can effectively integrate spatial information, using a unified model to accurately predict chlorophyll a concentrations at different spatial points, simplifying modeling complexity and improving practical efficiency.

[0107] Application Results: The four examples above demonstrate that the method of this invention exhibits good performance in different types of water bodies. Through targeted feature selection, even under conditions of scarce high-concentration samples and limited data, the model maintains a stable prediction accuracy of R²>0.82 in rigorous time extrapolation validation, proving that the method can effectively adapt to the sparse and non-equilibrium data conditions of northern reservoirs and has practical application value.

Claims

1. A method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model, characterized in that, Includes the following steps: Step S1: Acquisition and preprocessing of multi-source data; Step S2: Key feature factor selection and complete dataset construction; Step S3: Construction and training of the SO-KNN model; Step S4: Model application and performance evaluation; Step S4.1: Extrapolation prediction: Input the independent time extrapolation test set processed in step S3.2 into the SO-KNN prediction model trained in step S3.3 to obtain the normalized chlorophyll a concentration prediction value; Step S4.2: Result inverse normalization and comprehensive performance evaluation: Using the chlorophyll a concentration normalization parameter recorded in step S3.1, the predicted value obtained in step S4.1 is inverse normalized to restore the chlorophyll a concentration with actual physical meaning; Multiple statistical indicators were calculated between the inversely normalized predicted values ​​and the actual measured values ​​of chlorophyll a concentration in the test set to quantitatively evaluate the model performance.

2. The method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model according to claim 1, characterized in that, The specific implementation process of step 1 is as follows: Step S1.1: Water Quality Monitoring Data Acquisition and Preprocessing: Acquire water quality monitoring data from each sampling point during the historical monitoring period of the target lake / reservoir. The water quality monitoring data includes sampling date, water temperature, pH value, dissolved oxygen, total nitrogen, total phosphorus, ammonia nitrogen, permanganate index, and chlorophyll a concentration. Preprocess the water quality monitoring data, including: removing abnormal and duplicate samples to obtain valid samples; using a chain equation-based multiple interpolation method to fill missing values ​​in the valid samples; for data with concentrations below the detection limit (LOD), use LOD / Substitution is performed to obtain a water quality characteristic set; Step S1.2: Acquisition of meteorological monitoring data and construction of meteorological feature sets; Step S1.2.1: Meteorological monitoring data acquisition: Acquire the daily average meteorological monitoring data of each sampling point in the valid sample on the sampling day and the previous 14 days. The daily average meteorological monitoring data are meteorological factors, including daily average temperature, rainfall, and wind speed. During the acquisition process, for sampling points with direct monitoring data, actual observation data is used; for meteorological factors or certain dates without direct monitoring data, meteorological data for the corresponding time period of the target sampling point is generated based on observation data from multiple surrounding meteorological stations, ensuring that each sampling point has a complete daily average meteorological data sequence. Step S1.2.2: Calculation of lagged meteorological characteristics: For each type of meteorological factor, calculate the moving average of each sampling point 1 day, 2 days, 3 days, 5 days, 7 days, 10 days and 14 days before the sampling date to obtain the lagged meteorological characteristics; Step S1.2.3: Construction of meteorological feature set: Integrate the meteorological monitoring data of each sampling point on the sampling day with the lagged meteorological features to obtain the meteorological feature set; Step S1.3: Construction of Time Feature Set: Extract the sampling year and sampling month information of each valid sample, and encode the sampling month into a continuous variable through sine-cosine transform. The specific encoding formula is as follows: Monthly Sine Characteristics Monthly cosine characteristics ; The sampling year, the sine feature of the encoded month, and the cosine feature of the month are used as time features to construct a time feature set.

3. The method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model according to claim 2, characterized in that, The specific implementation process of step 2 is as follows: Step S2.1: Screening of key feature factors: Using Spearman's rank correlation coefficient analysis, the correlation coefficients and corresponding p values ​​between each factor in the water quality feature set and the meteorological feature set and chlorophyll a concentration were calculated respectively. Feature factors with p values ​​< 0.01 and high absolute values ​​of correlation coefficients were screened as key feature factors. Step S2.2: Complete dataset construction: Integrate the key feature factors obtained from the screening with the time feature set constructed in step S1.3 to form the input feature set; integrate the input feature set with the target value of chlorophyll a concentration of the corresponding sample to obtain the complete dataset.

4. The method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model according to claim 3, characterized in that, The specific implementation process of step 3 is as follows: Based on the complete dataset obtained in step S2, a K-nearest neighbor regression prediction model with fusion snake optimization algorithm is constructed and trained, with all key feature factors of the input feature set in step S2 as input variables and chlorophyll a concentration as output variable. Step S3.1: Data standardization: Normalize the input and output variables to obtain a normalized complete dataset; The normalization process uses the min-max normalization method to map the data to the [0,1] interval, and records the normalization parameters corresponding to each key feature factor and chlorophyll a concentration for subsequent inverse normalization. Step S3.2: Time series dataset partitioning: Sort the normalized complete dataset according to the monitoring time, select the samples from the most recent year as the independent time extrapolation test set to simulate the prediction scenario of unknown future data; use all historical samples from before this year as the time training set for the construction and training of the K-nearest neighbor regression prediction model that integrates the snake optimization algorithm. Step S3.3: Calculate the optimal number of nearest neighbors K_opt for the K-nearest neighbor regression prediction model based on the fusion snake optimization algorithm; Step S3.3.1: Parameter Setting and Population Initialization: Preset the search range [K_min, K_max] for the optimal nearest neighbor number K, and initialize the snake population X = {X1, X2, X3, ..., X...} based on this search range. n }, where each individual X i This represents a candidate K value, where the individual positions are all within the search range of the preset optimal nearest neighbor number K; Step S3.3.2: Fitness assessment: For each individual X i The corresponding candidate K values ​​are used to construct a K-nearest neighbor regression prediction model based on the fusion snake optimization algorithm. Time series cross-validation is then used on the time training set to calculate the root mean square error (RMSE) of the K-nearest neighbor regression prediction model. The negative value of the RMS error is defined as the fitness F of the individual. i This yields the fitness set F = {F1, F2, F3, ..., F...} n The optimization objective is to maximize fitness and minimize the root mean square error of validation. Step S3.3.3: Adaptive Iterative Optimization Based on Environmental Perception: By simulating the adaptive behavior of snake groups to environmental changes, the population is driven to perform iterative optimization; the optimization process is carried out in a loop within the set maximum number of iterations max_iter. Let the current number of iterations be t. The population is divided into two subgroups: male and female. Environmental parameter update: Based on the iteration progress, calculate the current environmental simulation parameters Temperature factor Temp and food abundance factor Q; both are variables that decrease with the iteration process, and their value range remains within [0, 1]. Behavioral pattern switching and location update: The population's behavioral pattern is determined based on environmental parameters, and the new candidate location X_new for each individual is calculated; When the food abundance factor Q is lower than the set food abundance threshold Q_th, the global exploration mode is executed; when the food abundance factor Q is not lower than the set food abundance threshold Q_th, the local development sub-mode is selected based on the relationship between the temperature factor Temp and the set temperature threshold T_th; the specific modes are as follows: Global exploration mode Q < Q_th: Individuals move to randomly selected individuals of the same sex in the population, and the step size is adjusted by the exploration ability factor; Local development mode Q ≥ Q_th: When Temp > T_th, all individuals move towards the current optimal solution X_food; when Temp ≤ T_th, they compete with a preset competition probability p, or reproduce with a probability of (1-p). Boundary constraints and updates: Perform boundary checks and rounding on all new positions X_new, and recalculate the fitness using K_new, as shown in the following formula: Here, round() is the rounding function; clamp() is the boundary constraint function, which takes K_min when X_new < K_min, takes K_max when X_new > K_max, and otherwise takes X_new; Reconstruct the K-nearest neighbor regression prediction model using K_new and calculate -RMSE as the new fitness F_new. If the new fitness F_new is better than the individual's historical best fitness, then update the individual's best position and fitness. Global optimal tracking: Compare the updated fitness of all individuals in the current iteration, find the individual with the highest fitness, and update its position to the new optimal solution X_food; Iteration Termination and Output: Repeat step S3.3.3 until the maximum number of iterations is reached or the optimal solution shows no significant improvement after multiple iterations. The output is the global optimal solution X_best, which is rounded down and used as the number of K nearest neighbors K_opt for optimization. Step S3.4: Optimal Model Training: Using the optimal number of nearest neighbors K_opt obtained in step S3.3.3, initialize and train the K-nearest neighbor regression model, complete the model fitting on the time training set, and obtain the SO-KNN prediction model.

5. The method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model according to claim 4, characterized in that, The specific implementation process of step 4 is as follows: Step S4.1: Extrapolation prediction: Input the independent time extrapolation test set processed in step S3.2 into the SO-KNN prediction model trained in step S3.3 to obtain the normalized chlorophyll a concentration prediction value; Step S4.2: Result inverse normalization and comprehensive performance evaluation: Using the chlorophyll a concentration normalization parameter recorded in step S3.1, the predicted value obtained in step S4.1 is inverse normalized to restore the chlorophyll a concentration with actual physical meaning; Multiple statistical indicators were calculated between the inversely normalized predicted values ​​and the actual measured values ​​of chlorophyll a concentration in the test set to quantitatively evaluate the model performance.

6. The method for predicting chlorophyll a concentration in lakes and reservoirs based on the SO-KNN model according to claim 5, characterized in that, The core evaluation indicators for quantitative and qualitative assessment of model performance include the coefficient of determination, mean absolute error, and root mean square error.

Citation Information

Patent Citations

  • A cyanobacterial bloom prediction method based on a GA-Elman network

    CN109726857A

  • Lake cyanobacterial bloom long-term forecasting method and system based on STL-RF-LSTM

    CN114386710A

Cited By

  • Multi-time scale prediction method and system for cyanobacterial bloom

    CN122050575A