Blue-green algae bloom extraction method based on spectrum and phenological characteristics

By combining spectral and phenological characteristics and utilizing PCI and LSTM models, we have achieved accurate identification and dynamic analysis of cyanobacterial blooms. This solves the problem of insufficient spectral and temporal information in existing technologies and improves the monitoring accuracy and temporal consistency of cyanobacterial blooms.

CN121904602APending Publication Date: 2026-04-21CHUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHUZHOU UNIV
Filing Date
2025-12-12
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing remote sensing identification methods for cyanobacterial blooms have limited accuracy when dealing with highly dynamic bloom processes, and it is difficult to simultaneously take into account spectral and temporal information, resulting in poor temporal consistency of classification results. Furthermore, traditional methods have limitations in the automatic extraction and dynamic identification of cyanobacterial blooms.

Method used

A fusion method based on spectral and phenological characteristics was adopted, and a localized inversion model was established using the phycocyanin index (PCI). Combined with the gradient method and the long short-term memory neural network (LSTM) model, accurate identification and dynamic analysis of cyanobacterial blooms were achieved.

Benefits of technology

It improved the classification accuracy and temporal stability of cyanobacterial blooms, revealed their spatiotemporal evolution patterns, provided efficient automated monitoring technology, and enhanced the protection and management capabilities of lake ecosystems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904602A_ABST
    Figure CN121904602A_ABST
Patent Text Reader

Abstract

The invention relates to a cyanobacterial bloom extraction method based on spectrum and phenological characteristics, and relates to the field of ecology, and the method comprises the steps: selecting a vegetation index PCI which is excellent in cyanobacterial bloom recognition, building a localized phycocyanobilin inversion model, and carrying out the homogeneity test of pixels around a sampling point; by calculating a PCI gradient image time sequence, cyanobacterial bloom and background water boundaries are identified, and a threshold value for distinguishing cyanobacteria and non-cyanobacteria pixels is determined by using a maximum gradient average value; constructing a cyanobacterial bloom recognition model based on the LSTM neural network; calculating cyanobacterial bloom outbreak intensity and coverage area, identifying a hot spot area and a cold spot area, and depicting distribution characteristics of cyanobacterial bloom in space; a PCI time sequence data set is synthesized, the evolution trend of a cyanobacterial bloom key phenological index time sequence is described, the innovation of the method lies in that spectrum and phenological characteristics are fused, cyanobacterial bloom spatial and temporal distribution is accurately analyzed, and guarantee is provided for water environment treatment and drinking water safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecology, and more particularly to a method for extracting cyanobacterial blooms based on spectral and phenological characteristics. Background Technology

[0002] Cyanobacterial blooms in lakes are a common ecological anomaly in eutrophic water bodies. Frequent outbreaks not only disrupt the ecological balance and affect water quality, but also significantly impact the lake's energy balance and regional climate regulation function. In recent years, due to the combined effects of natural and anthropogenic factors such as global warming, domestic sewage discharge, and agricultural irrigation, the frequency, timing, and spatial extent of cyanobacterial blooms in lakes have been increasing. The massive proliferation of cyanobacteria leads to the depletion of dissolved oxygen and ecological landscape degradation. Furthermore, by altering the albedo and emissivity of the lake surface, it affects the radiation balance and energy equilibrium, weakening the lake's regulatory role in the regional hydrothermal cycle. With the increasing severity of lake ecological and environmental problems, there is an urgent need to establish efficient and stable dynamic identification and monitoring technologies for cyanobacterial blooms to provide scientific support for the protection and management of lake ecosystems.

[0003] Existing remote sensing identification studies of cyanobacterial blooms mainly employ two types of methods: one based on spectral features, using a single vegetation index to invert phycocyanin concentration; the other based on phenological differences, using the growth cycle characteristics of cyanobacteria and aquatic vegetation for classification. However, these methods have limited accuracy when dealing with highly dynamic bloom processes, often failing to simultaneously consider spectral and temporal information, resulting in poor temporal consistency of classification results. Furthermore, due to the complex environment of lakes, significant spectral mixing, and the reliance on experience for threshold setting, traditional methods still have limitations in the automatic extraction and dynamic identification of cyanobacterial blooms.

[0004] To address the aforementioned problems, this invention proposes a method for extracting cyanobacterial blooms based on the fusion of spectral and phenological characteristics. Based on the unique absorption peak characteristics of phycocyanin, which distinguishes cyanobacteria from aquatic vegetation, the phycocyanin index (PCI) is calculated using unique bands from remote sensing satellites, establishing a localized phycocyanin inversion model based on PCI. On this basis, a gradient method is used to determine the threshold for distinguishing cyanobacteria from non-cyanobacteria pixels based on the PCI time series, achieving preliminary extraction of cyanobacterial blooms. Furthermore, a Long Short-Term Memory (LSTM) neural network model is introduced to establish a cyanobacterial bloom identification model based on PCI time series by mining the phenological differences between aquatic vegetation and cyanobacterial blooms in the temporal dimension, thereby effectively improving classification accuracy and temporal stability. Finally, the spatiotemporal evolution characteristics of lake cyanobacterial blooms are analyzed to reveal their dynamic change patterns and ecological environment driving mechanisms.

[0005] In summary, this method integrates spectral information and phenological characteristics, and utilizes temporal phenological features and spectral features to achieve accurate identification and dynamic analysis of cyanobacterial blooms. This has significant scientific and application value for revealing the biophysical feedback mechanism of lake ecosystems and improving the monitoring accuracy of cyanobacterial blooms. Summary of the Invention

[0006] This invention proposes a method for extracting cyanobacterial blooms based on spectral and phenological characteristics, utilizing temporal phenological and spectral features to achieve accurate identification and dynamic analysis of cyanobacterial blooms. Compared with traditional methods, this invention has advantages such as accurate identification of dynamic cyanobacterial blooms and revelation of spatiotemporal evolution patterns.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows: a method for extracting cyanobacterial blooms based on spectral and phenological characteristics, the method comprising the following steps:

[0008] Step S1: Select the vegetation index PCI, which performs well in identifying cyanobacterial blooms, establish a localized phycocyanin inversion model, perform homogeneity testing on the pixels around the sampling points, and ensure the accuracy and reliability of the model through nonlinear fitting methods.

[0009] Step S2: Calculate the PCI gradient image time series, identify the boundary between cyanobacterial blooms and background water bodies, and use the maximum gradient average value to determine the threshold for distinguishing cyanobacteria from non-cyanobacteria pixels, thus achieving preliminary extraction of cyanobacterial blooms.

[0010] Step S3: Combining the phenological differences between aquatic vegetation and cyanobacterial blooms, accurately classify cyanobacterial blooms, and use the PCI sequences determined by the above threshold method to construct a cyanobacterial bloom recognition model based on an LSTM neural network.

[0011] Step S4: Calculate the intensity and coverage area of ​​cyanobacterial blooms based on the PCI threshold, and use the hot spot and cold spot analysis method to identify hot spot areas and cold spot areas to characterize the cyanobacterial blooms from the perspective of spatial distribution.

[0012] Step S5: Based on the 8-day synthetic PCI time series filtered by Savitzky-Golay, extract the phenological characteristic parameters of cyanobacterial blooms (start time, end time, duration, peak time), and combine Sen slope estimation and Mann-Kendall trend test methods to analyze the long-term evolution trend of phenological indicators.

[0013] Further, step S1 includes the following steps:

[0014] Step S11: Phycocyanin Index (PCI) is the value at 620 nm. Rayleigh-corrected reflectance is a reliable indicator for estimating phycocyanin concentration or identifying cyanobacterial blooms. It is used after masking... The dataset and phycocyanin concentration data sampled on the same day constituted a synchronous observation data pair, and a localized phycocyanin model based on PCI was established. The calculation formula for PCI is as follows:

[0015]

[0016] (1),

[0017] in, This indicates the unique absorption peak of cyanobacteria; Indicates spectral differences between baselines. This represents the theoretical reflectance at 620 nm obtained by extrapolation from the spectral baseline.

[0018] Step S12: Calculate the variance or mean deviation of the PCI values ​​within a 3×3 pixel window centered on each in-situ sampling point to perform a homogeneity test, thereby suppressing water mass interference and improving the stability of the inversion model.

[0019] Step S13: The relationship between PCI and measured phycocyanin concentration is fitted using nonlinear least squares fitting, logarithmic fitting, and exponential fitting methods, respectively. Leave-one-out cross-validation is used to evaluate model accuracy, and the model with the highest coefficient of determination and lowest unbiased root mean square error is selected as the optimal fitting model. The nonlinear least squares fitting function is: Its optimal parameters The solution is obtained by minimizing the squared error, as shown in the following formula:

[0020] (2),

[0021] in, Indicates the first PCI values ​​for each sample Indicates the first The measured phycocyanin concentration of each sample Represents a nonlinear fitting function. This represents the optimal parameter vector obtained from the fit. Represents the total number of samples; the formulas for the logarithmic fitting function and the exponential fitting function are as follows:

[0022] (3),

[0023] (4),

[0024] in, and The fitting parameters are obtained by solving using the least squares method or logarithmic transformation; the coefficients of determination The formula used to measure goodness of fit is as follows:

[0025] (5),

[0026] in, This represents the model's predicted value. This represents the mean of the observed values. This represents the total number of samples; the lowest unbiased root mean square error (URMSE) is used to measure prediction accuracy, and the formula is as follows:

[0027] (6),

[0028] in, This represents the model's predicted value. Represents the observed value. This represents the mean of the prediction error to eliminate bias. Indicates the total number of samples;

[0029] Step S14: Compare the fitting results with the remaining observation data, identify abnormal or biased samples, and make necessary corrections or remove them to ensure the reliability of the model.

[0030] Further, step S2 includes the following steps:

[0031] Step S21: Calculate the gradient value for each pixel in the neighboring 3×3 window pixels to determine the boundary of cyanobacterial blooms. The calculation formula is as follows:

[0032] (7),

[0033] in, and This indicates the PCI value and position change of the current pixel relative to its 8 neighboring pixels within a 3x3 window;

[0034] Step S22: Extract the set of pixels with the maximum gradient value from the remaining pixels in the PCI gradient image, and consider these pixels to correspond to the boundary between the cyanobacterial bloom and the background water body;

[0035] Step S23: Apply the above method to the entire PCI time series image to obtain all individual thresholds and generate histograms;

[0036] Step S24: Determine the average value of all individual thresholds in the histogram as the threshold for distinguishing between cyanobacterial blooms and non-cyanobacterial bloom pixels, perform statistical analysis, and take the average or weighted average as the threshold for the entire period to achieve preliminary extraction of cyanobacterial blooms.

[0037] Further, step S3 includes the following steps:

[0038] Step S31: Construct a variable sample set with unified dimensions, units, and scales to unify data units and provide stable input for LSTM model training;

[0039] Step S32: Randomly divide the sample data according to time and space, with 70% as the training set and 30% as the test set to ensure the independence of training and validation;

[0040] Step S33: Build two hidden layers for the LSTM architecture, each with 50 nodes and one output node. Use binary classification to classify pixels into cyanobacterial blooms and non-cyanobacterial blooms.

[0041] Step S34: The hidden layer uses the ReLU activation function, while the output uses the Sigmoid callback function to limit the training accuracy to 98% to prevent the model from overfitting;

[0042] Step S35: Use the standardized n-dimensional PCI sample time series as input to train the LSTM model, and then substitute the n-dimensional cyanobacterial bloom PCI sequence into the trained LSTM model for classification.

[0043] Furthermore, in step S31: to construct a cyanobacterial bloom identification model, building a variable sample set with unified dimensions, units, and scales is fundamental; specifically, it includes the following steps:

[0044] Step S311: Create samples on the remote sensing image based on the latitude and longitude positions of the collected measured data;

[0045] Step S312: For non-geographical variables in the variable set, one-hot encoding is used to represent categorical variables as binary vectors. Specifically, let the categorical variables... The set of values ​​is Then for the first The encoding result of each sample is defined as follows:

[0046] (8),

[0047] The one-hot vector representation of this sample can then be obtained:

[0048] (9),

[0049] Through the above encoding, each categorical variable can be mapped to an independent binary feature dimension, solving the problem that the model has difficulty in handling categorical attributes;

[0050] Step S313: The variables are transformed into dimensionless values ​​between 0 and 1 using the min-max and Z-score normalization methods to improve model accuracy and convergence speed, and prevent gradient explosion. The formula for min-max normalization is as follows:

[0051] (10)

[0052] in, Represents the first of the original input variables Each sample value This represents the minimum value of the variable across all samples. This represents the maximum value of the variable across all samples. This represents the normalized variable value, with a range of values. The standardized Z-score calculation formula is as follows:

[0053] (11),

[0054] in, This represents the mean of the variable, i.e. , This represents the standard deviation of the variable;

[0055] Step S314: Unify the temporal resolution of remote sensing images to 8 days to minimize the problem of missing observations caused by cloud and rain effects, while maintaining a spatial resolution of 300m.

[0056] Further, step S4 includes the following steps:

[0057] Step S41: Based on different PCI thresholds, the outbreak intensity is divided into three levels: severe, moderate, and mild. The PCI intervals corresponding to severe, moderate, and mild are determined by historical observation data or expert experience to quantify the severity of the algal bloom.

[0058] Step S42: Count the number of pixels in cyanobacterial blooms. And multiplied by pixel resolution The formula for calculating the coverage area is as follows:

[0059] (12),

[0060] in, Indicates the coverage area. This indicates the number of pixels representing cyanobacterial blooms. Image resolution;

[0061] Step S43: Determine the hot and cold spots of cyanobacterial blooms using the hot and cold spot analysis method, as shown in the following formula:

[0062] (13)

[0063] in, express Regional cyanobacterial blooms, The spatial weight of cyanobacterial blooms is defined using distance rules. After value standardization, if If the value is positive and significant, it belongs to a high-value aggregation "hotspot" area of ​​cyanobacterial blooms; if If the value is negative and significant, it belongs to the "cold spot" area of ​​low-value aggregation of cyanobacterial blooms;

[0064] Step S44: Calculate the coverage area of ​​cyanobacterial blooms in different hotspot areas at different time scales.

[0065] Further, step S5 includes the following steps:

[0066] Step S51: Based on the 8-day PCI time series dataset of cyanobacterial bloom, the Savitzky-Golay (SG) filtering method is used for smoothing to provide continuous and smooth time series data for phenological feature extraction;

[0067] Step S52: The extraction of phenological characteristic parameters (start time, end time, duration, and peak time) uses a Gaussian fitting model, the formula of which is as follows:

[0068] (14)

[0069] in, This represents the PCI background value (the PCI baseline level determined by the fitted function). The amplitude represents the maximum PCI value during the algal bloom period. This indicates the peak date of the cyanobacterial bloom (the date when the PCI reaches its maximum value). Indicates the width of the algal bloom. Indicates the time step. The duration is defined as the period from the date when the PCI increases to its maximum 50% to the date when the PCI decreases to its maximum 50%.

[0070] Step S53: Describe the evolution trend of the time series of key phenological indicators for cyanobacterial blooms using Sen slope estimation and the Mann-Kendall trend test. The Sen slope is used to estimate the median slope of the time series trend, as shown in the following formula:

[0071] (15)

[0072] in, and These represent the time series at time points. and Phenological characteristic values ​​(such as peak time, duration, etc.). This represents the rate of change (trend rate) of the median of a time series. , This represents the length of the time series. When... When, it indicates that phenological parameters increase with time (e.g., the peak time of algal bloom shifts later); when When, it indicates that the phenological parameter decreases over time (e.g., the peak time of algal blooms arrives earlier); the Mann-Kendall trend test is used to determine whether the trend of a time series is significant, and the trend test is performed using the test statistic Z. The formulas for calculating Z are as follows:

[0073] (16)

[0074] (17)

[0075] in, and These represent the time series at time points. and Phenological characteristic values ​​(such as peak time, duration, etc.). Indicates the length of the time series. Represents the eigenvalue sign function. Represents statistics The variance is used to quantify the impact of repeated values ​​in a data series on trend testing. and The calculation formula is as follows:

[0076] (18)

[0077] (19)

[0078] in, Indicates the length of the time series. This indicates the number of tie groups in the sequence. Indicates the first The number of groups. As can be seen from the above technical solution, the technical effect of the cyanobacterial bloom extraction method based on spectral and phenological characteristics provided by this invention is as follows:

[0079] 1. Utilizing the unique spectral absorption characteristics of phycocyanin, which distinguish it from aquatic vegetation, and combining this with satellite imagery using specific bands to calculate the phycocyanin index (PCI), a localized phycocyanin concentration retrieval model was established. This model can accurately characterize the spectral differences in cyanobacterial blooms, improving the spectral sensitivity and distinguishing ability for cyanobacterial identification.

[0080] 2. By constructing PCI time series and using the gradient method to determine the pixel thresholds for cyanobacteria and non-cyanobacteria, this method achieves automated preliminary extraction of cyanobacterial blooms, avoiding the subjectivity and limitations of traditional empirical threshold methods, thereby significantly improving the stability and consistency of bloom identification.

[0081] 3. This invention further introduces a Long Short-Term Memory (LSTM) neural network model to fully explore the phenological differences between cyanobacterial blooms and aquatic vegetation over time. The LSTM model can capture dynamic dependencies in time series, enhance the accuracy of cyanobacterial bloom time series classification, achieve the identification of cyanobacterial blooms in complex aquatic environments, and quantitatively reveal the spatiotemporal evolution patterns of cyanobacterial blooms.

[0082] 4. The innovation of this invention lies in integrating spectral and phenological characteristics, utilizing temporal phenological features and spectral features to achieve accurate identification and dynamic analysis of cyanobacterial blooms. Compared with traditional methods, this invention has advantages such as accurate identification of dynamic cyanobacterial blooms and revelation of spatiotemporal evolution patterns, providing a new technical approach for the automated monitoring and dynamic extraction of cyanobacterial blooms in lakes. It has significant theoretical and practical application value for aquatic ecosystem monitoring and lake environmental management. Attached Figure Description

[0083] Figure 1 The flowchart illustrates a method for extracting cyanobacterial blooms based on spectral and phenological characteristics, as provided in an embodiment of the present invention.

[0084] Figure 2 The diagram illustrates a specific application of a cyanobacterial bloom extraction method based on spectral and phenological characteristics provided in an embodiment of the present invention.

[0085] Figure 3 A schematic diagram of the LSTM neural network model structure provided in an embodiment of the present invention is shown. Detailed Implementation

[0086] The embodiments of the technical solution of the present invention will now be described and explained in detail with reference to the accompanying drawings.

[0087] Example: According to Figure 1 As shown, a method for extracting cyanobacterial blooms based on spectral and phenological characteristics specifically includes the following steps:

[0088] Step S1: The PCI index is used to construct a local phycocyanin inversion model, and the homogeneity of the sampled pixels is checked. At the same time, nonlinear fitting is used to ensure the accuracy of the model.

[0089] Step S2: Calculate the PCI gradient sequence, identify the boundaries of cyanobacterial blooms, and set a threshold using the maximum gradient average value to achieve preliminary extraction;

[0090] Step S3: Combine the phenological characteristics of aquatic vegetation and cyanobacteria for fine classification, and use the PCI threshold sequence to construct an LSTM network for algal bloom identification;

[0091] Step S4: Calculate the outbreak intensity and coverage area based on the threshold, and use the hot and cold spot analysis method to determine the hot and cold spot areas, and depict the spatial distribution of cyanobacterial blooms;

[0092] Step S5: Based on the PCI time series after SG filtering, extract phenological parameters such as start, end, duration, and peak time, and analyze trend changes using Sen slope and Mann-Kendall test.

[0093] In this embodiment, Chaohu Lake in Anhui Province is used as the research object. Continuous remote sensing image data and some measured water quality parameters are selected to verify the effectiveness of the proposed cyanobacterial bloom extraction method based on spectral and phenological characteristics. The remote sensing data used include Sentinel-2 MSI and Landsat 8 OLI multispectral images, and auxiliary data includes meteorological data and lake boundary vector data, which are used to construct and train a spectral-phenological feature extraction model for cyanobacterial blooms with strong generalization ability.

[0094] The implementation steps of the embodiment are as follows: Figure 2 As shown.

[0095] In this embodiment, the purpose of step S1 is to construct a localized phycocyanin inversion model to provide a reliable spectral data foundation for the subsequent identification and extraction of cyanobacterial blooms. Selecting the PCI spectral index, which is sensitive to cyanobacterial blooms, is crucial because it can effectively capture the unique absorption characteristics of phycocyanin at 620 nm, thereby distinguishing cyanobacteria from other water bodies. Establishing a localized model and subjecting it to rigorous homogeneity testing and nonlinear fitting aims to overcome the differences in optical properties of water bodies in different regions, ensuring that the model inversion results have physical consistency and spatial applicability. This implementation case selected synchronous satellite remote sensing imagery and field-measured phycocyanin concentration data from the Chaohu Lake basin. Through a series of rigorous model construction and verification steps, the accuracy and reliability of the inversion data were ensured. Therefore, step S1 includes the following steps:

[0096] Step S11: Phycocyanin Index (PCI) is the value at 620 nm. Rayleigh-corrected reflectance is a reliable indicator for estimating phycocyanin concentration or identifying cyanobacterial blooms. It is used after masking... The dataset and phycocyanin concentration data sampled on the same day constituted a synchronous observation data pair, and a localized phycocyanin model based on PCI was established. The calculation formula for PCI is as follows:

[0097]

[0098] (1),

[0099] in, This indicates the unique absorption peak of cyanobacteria; Indicates spectral differences between baselines. This represents the theoretical reflectance at 620 nm obtained by extrapolation from the spectral baseline.

[0100] Step S12: Calculate the variance or mean deviation of the PCI values ​​within a 3×3 pixel window centered on each in-situ sampling point to perform a homogeneity test, thereby suppressing water mass interference and improving the stability of the inversion model.

[0101] Step S13: The relationship between PCI and measured phycocyanin concentration is fitted using nonlinear least squares fitting, logarithmic fitting, and exponential fitting methods, respectively. Leave-one-out cross-validation is used to evaluate model accuracy, and the model with the highest coefficient of determination and lowest unbiased root mean square error is selected as the optimal fitting model. The nonlinear least squares fitting function is: Its optimal parameters The solution is obtained by minimizing the squared error, as shown in the following formula:

[0102] (2),

[0103] in, Indicates the first PCI values ​​for each sample Indicates the first The measured phycocyanin concentration of each sample Represents a nonlinear fitting function. This represents the optimal parameter vector obtained from the fit. Represents the total number of samples; the formulas for the logarithmic fitting function and the exponential fitting function are as follows:

[0104] (3),

[0105] (4),

[0106] in, and The fitting parameters are obtained by solving using the least squares method or logarithmic transformation; the coefficients of determination The formula used to measure goodness of fit is as follows:

[0107] (5),

[0108] in, This represents the model's predicted value. This represents the mean of the observed values. This represents the total number of samples; the lowest unbiased root mean square error (URMSE) is used to measure prediction accuracy, and the formula is as follows:

[0109] (6),

[0110] in, This represents the model's predicted value. Represents the observed value. This represents the mean of the prediction error to eliminate bias. This represents the total number of samples.

[0111] Step S14: Compare the fitting results with the remaining observation data, identify abnormal or biased samples, and make necessary corrections or remove them to ensure the reliability of the model.

[0112] In this embodiment, the core of step S2 lies in using the time series of PCI gradient images to accurately identify the boundary between cyanobacterial blooms and background water bodies, and thereby determine a global threshold to achieve preliminary automated extraction of cyanobacterial blooms. Cyanobacterial blooms typically exhibit a clustered distribution in space, and the PCI value at their boundary changes drastically. This spatial gradient characteristic is key to distinguishing bloom areas from clear water areas. By calculating the gradient of each pixel within its neighborhood and statistically analyzing the maximum gradient value over the time series, the pixels with the most significant bloom boundary can be effectively captured. The segmentation threshold determined based on this has both physical significance and statistical robustness, providing high-quality initial samples for subsequent fine classification. Therefore, step S2 includes the following steps:

[0113] Step S21: Calculate the gradient value for each pixel in the neighboring 3×3 window pixels to determine the boundary of cyanobacterial blooms. The calculation formula is as follows:

[0114] (7),

[0115] in, and This represents the PCI value and position change of the current pixel relative to its 8 neighboring pixels within a 3x3 window.

[0116] Step S22: Extract the set of pixels with the largest gradient value from the remaining pixels in the PCI gradient image, and consider these pixels to correspond to the boundary between the cyanobacterial bloom and the background water body.

[0117] Step S23: Apply the above method to the entire PCI time series image to obtain all individual thresholds and generate histograms.

[0118] Step S24: Determine the average value of all individual thresholds in the histogram as the threshold for distinguishing between cyanobacterial blooms and non-cyanobacterial bloom pixels, perform statistical analysis, and take the average or weighted average as the threshold for the entire period to achieve preliminary extraction of cyanobacterial blooms.

[0119] In this embodiment, step S3 is designed primarily to construct a time-series classification model based on an LSTM neural network to achieve accurate classification of cyanobacterial blooms. The model structure diagram is shown below. Figure 3 As shown. The LSTM network, as a deep learning model capable of capturing long-term time dependencies, is based on the principle of learning and memorizing the unique variation patterns of the PCI sequence throughout the entire growth cycle of cyanobacterial blooms. By using the PCI time series initially extracted in step S2 as input, the LSTM model can deeply understand the dynamic characteristics of the entire process of cyanobacterial bloom occurrence, development, outbreak, and decline, thereby effectively distinguishing it from aquatic vegetation with different phenological rhythms. This step, through standardized data preprocessing, reasonable network structure design, and strategies to prevent overfitting, ensures that the model possesses high accuracy and strong generalization ability. Therefore, step S3 includes the following steps:

[0120] Step S31: Construct a variable sample set with unified dimensions, units, and scales to unify data units and provide stable input for LSTM model training.

[0121] Step S32: Randomly divide the sample data according to time and space, with 70% as the training set and 30% as the test set to ensure the independence of training and validation.

[0122] Step S33: Build two hidden layers for the LSTM architecture, each with 50 nodes and one output node. Use binary classification to classify pixels into cyanobacterial blooms and non-cyanobacterial blooms.

[0123] Step S34: The hidden layer uses the ReLU activation function, while the output uses the Sigmoid callback function to limit the training accuracy to 98% to prevent the model from overfitting.

[0124] Step S35: Use the standardized n-dimensional PCI sample time series as input to train the LSTM model, and then substitute the n-dimensional cyanobacterial bloom PCI sequence into the trained LSTM model for classification.

[0125] In this embodiment, the purpose of step S4 is to quantify the outbreak intensity and spatial pattern of cyanobacterial blooms, providing an intuitive and quantitative basis for environmental assessment and risk management. By setting different levels of PCI thresholds to classify bloom intensity, the severity of the bloom is described in detail, and the calculated coverage area is a direct measure of the bloom's impact range. Furthermore, a spatial statistical method of hotspot / coldspot analysis is introduced. Its basic principle is to identify high-value (hotspot) or low-value (coldspot) clusters with significant spatial statistical value, thereby going beyond visual interpretation and statistically revealing the spatial aggregation pattern and core distribution area of ​​cyanobacterial blooms. Therefore, step S4 includes the following steps:

[0126] Step S41: Based on different PCI thresholds, the outbreak intensity is divided into three levels: severe, moderate, and mild. The PCI intervals corresponding to severe, moderate, and mild are determined by historical observation data or expert experience to quantify the severity of the algal bloom.

[0127] Step S42: Count the number of pixels in cyanobacterial blooms. And multiplied by pixel resolution The formula for calculating the coverage area is as follows:

[0128] (12),

[0129] in, Indicates the coverage area. This indicates the number of pixels representing cyanobacterial blooms. This refers to the image resolution.

[0130] Step S43: Determine the hot and cold spots of cyanobacterial blooms using the hot and cold spot analysis method, as shown in the following formula:

[0131] (13)

[0132] in, express Regional cyanobacterial blooms, The spatial weight of cyanobacterial blooms is defined using distance rules. After value standardization, if If the value is positive and significant, it belongs to a high-value aggregation "hotspot" area of ​​cyanobacterial blooms; if If the value is negative and significant, it belongs to the "cold spot" area of ​​low-value aggregation of cyanobacterial blooms.

[0133] Step S44: Calculate the coverage area of ​​cyanobacterial blooms in different hotspot areas at different time scales.

[0134] In this embodiment, the core of step S5 lies in extracting key phenological parameters that characterize cyanobacterial blooms from long-term PCI data and analyzing their changing trends to reveal the long-term evolutionary patterns of cyanobacterial blooms. SG filtering is used to smooth the time series, the basic principle of which is to provide a high-quality data foundation for phenological parameter extraction while preserving the true shape of the time series curve. Subsequently, a Gaussian fitting model is used to extract phenological parameters such as start time and peak time from the smoothed curve. The principle of this process is to approximate the life cycle of cyanobacterial blooms as one or more Gaussian processes, thereby quantifying its phenological characteristics using a mathematical model. Finally, trend analysis is performed by combining Sen slope and Mann-Kendall test, which can robustly assess the direction and significance of changes in phenological parameters over long time scales, providing scientific support for predicting future bloom situations. Therefore, step S5 includes the following steps:

[0135] Step S51: Based on the 8-day PCI time series dataset of cyanobacterial blooms, the Savitzky-Golay (SG) filtering method is used for smoothing to provide continuous and smooth time series data for phenological feature extraction.

[0136] Step S52: The extraction of phenological characteristic parameters (start time, end time, duration, and peak time) uses a Gaussian fitting model, the formula of which is as follows:

[0137] (14)

[0138] in, This represents the PCI background value (the PCI baseline level determined by the fitted function). The amplitude represents the maximum PCI value during the algal bloom period. This indicates the peak date of the cyanobacterial bloom (the date when the PCI reaches its maximum value). Indicates the width of the algal bloom. This indicates the time step. The duration is defined as the period from the date when the PCI increases to its maximum 50% to the date when the PCI decreases to its maximum 50%.

[0139] Step S53: Describe the evolution trend of the time series of key phenological indicators for cyanobacterial blooms using Sen slope estimation and the Mann-Kendall trend test. The Sen slope is used to estimate the median slope of the time series trend, as shown in the following formula:

[0140] (15)

[0141] in, and These represent the time series at time points. and Phenological characteristic values ​​(such as peak time, duration, etc.). This represents the rate of change (trend rate) of the median of a time series. , This represents the length of the time series. When... When, it indicates that phenological parameters increase with time (e.g., the peak time of algal bloom shifts later); when When, it indicates that the phenological parameter decreases over time (e.g., the peak time of algal blooms arrives earlier); the Mann-Kendall trend test is used to determine whether the trend of a time series is significant, and the trend test is performed using the test statistic Z. The formulas for calculating Z are as follows:

[0142] (16)

[0143] (17)

[0144] in, and These represent the time series at time points. and Phenological characteristic values ​​(such as peak time, duration, etc.). Indicates the length of the time series. Represents the eigenvalue sign function. Represents statistics The variance is used to quantify the impact of repeated values ​​in a data series on trend testing. and The calculation formula is as follows:

[0145] (18)

[0146] (19)

[0147] in, Indicates the length of the time series. This indicates the number of tie groups in the sequence. Indicates the first The number of groups / ties.

[0148] The above description is merely a general overview of the steps of this invention and should not be construed as limiting the scope of protection of this invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for extracting cyanobacterial blooms based on spectral and phenological characteristics, characterized in that, The method includes the following steps: Step S1: Select the vegetation index PCI, which performs well in identifying cyanobacterial blooms, establish a localized phycocyanin inversion model, and perform homogeneity testing on the pixels around the sampling points. Step S2: By calculating the PCI gradient image time series, identify the boundary between cyanobacterial blooms and background water bodies, and use the maximum gradient average value to determine the threshold for distinguishing between cyanobacteria and non-cyanobacteria pixels; Step S3: Combine the differences in phenology between aquatic vegetation and cyanobacterial blooms to classify cyanobacterial blooms, and use the cyanobacterial bloom PCI sequence determined by the pixel thresholding method to construct a cyanobacterial bloom recognition model based on an LSTM neural network. Step S4: Calculate the intensity and coverage area of ​​cyanobacterial blooms based on the PCI threshold, and use the hot spot and cold spot analysis method to identify hot spot areas and cold spot areas to characterize the spatial distribution characteristics of cyanobacterial blooms. Step S5: Based on the PCI time series dataset synthesized by Savizky-Golay filtering, extract the phenological characteristic parameters of cyanobacterial blooms, and use Sen slope estimation and Mann-Kendall trend test methods to describe the evolution trend of the time series of key phenological indicators of cyanobacterial blooms.

2. The method for extracting cyanobacterial blooms based on spectral and phenological characteristics according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Utilize the masked... The dataset and phycocyanin concentration data sampled on the same day form a synchronous observation data pair. A localized phycocyanin model based on PCI is established, where the calculation formula for PCI is as follows: , (1), in, This indicates a unique absorption peak in cyanobacteria. Indicates spectral differences between baselines. This represents the theoretical reflectance at 620 nm obtained by extrapolation from the spectral baseline. Step S12: Calculate the variance or mean deviation of the PCI values ​​within a 3×3 pixel window centered on each in-situ sampling point to perform a homogeneity test; Step S13: The relationship between PCI and measured phycocyanin concentration is fitted using nonlinear least squares fitting, logarithmic fitting, and exponential fitting methods, respectively. Leave-one-out cross-validation is used to evaluate model accuracy, and the model with the highest coefficient of determination and lowest unbiased root mean square error is selected as the optimal fitting model. The nonlinear least squares fitting function is: Its optimal parameter a The solution is obtained by minimizing the squared error, as shown in the following formula: (2), in, Indicates the first PCI values ​​for each sample Indicates the first The measured phycocyanin concentration of each sample Represents a nonlinear fitting function. This represents the optimal parameter vector obtained from the fit. Represents the total number of samples; the formulas for the logarithmic fitting function and the exponential fitting function are as follows: (3), (4), in, and The fitting parameters are obtained by solving using the least squares method or logarithmic transformation; the coefficients of determination The formula used to measure goodness of fit is as follows: (5), in, This represents the model's predicted value. This represents the mean of the observed values. The total number of samples is represented; the lowest unbiased root mean square error is used to measure prediction accuracy, and the formula is as follows: (6), in, This represents the model's predicted value. Represents the observed value. This represents the mean of the prediction error to eliminate bias. Indicates the total number of samples; Step S14: Compare the fitting results with the remaining observation data, identify abnormal or biased samples, and make necessary corrections or remove them.

3. The method for extracting cyanobacterial blooms based on spectral and phenological characteristics according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Calculate the gradient value for each pixel in the neighboring 3×3 window pixels to determine the boundary of cyanobacterial blooms. The calculation formula is as follows: (7), in, and This indicates the PCI value and position change of the current pixel relative to its 8 neighboring pixels within a 3x3 window; Step S22: Extract the set of pixels with the maximum gradient value from the remaining pixels in the PCI gradient image, and consider these pixels to correspond to the boundary between the cyanobacterial bloom and the background water body; Step S23: Apply the above method to the entire PCI time series image to obtain all individual thresholds and generate histograms; Step S24: Determine the average value of all individual thresholds in the histogram as the threshold for distinguishing between cyanobacterial blooms and non-cyanobacterial bloom pixels, perform statistical analysis, and take the average or weighted average as the threshold for the entire period.

4. The method for extracting cyanobacterial blooms based on spectral and phenological characteristics according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Construct a variable sample set with unified dimensions, units, and scales to unify data units and provide stable input for LSTM model training; Step S32: Randomly divide the sample data according to time and space, with 70% as the training set and 30% as the test set to ensure the independence of training and validation; Step S33: Build two hidden layers for the LSTM architecture, each with 50 nodes and one output node. Use binary classification to classify pixels into cyanobacterial blooms and non-cyanobacterial blooms. Step S34: The hidden layer uses the ReLU activation function, while the output uses the Sigmoid callback function to limit the training accuracy to 98%; Step S35: Use the standardized n-dimensional PCI sample time series as input to train the LSTM model, and then substitute the n-dimensional cyanobacterial bloom PCI sequence into the trained LSTM model for classification.

5. The method for extracting cyanobacterial blooms based on spectral and phenological characteristics according to claim 1, characterized in that, In step S31, constructing a variable sample set with unified dimensions, units, and scales is fundamental for building a cyanobacterial bloom identification model; specifically... Includes the following steps: Step S311: Create samples on the remote sensing image based on the latitude and longitude positions of the collected measured data; Step S312: For non-geographical variables in the variable set, one-hot encoding is used to represent categorical variables as binary vectors. Specifically, let the categorical variables... The set of values ​​is Then for the first The encoding result of each sample is defined as follows: (8), The one-hot vector representation of this sample can then be obtained: (9); Step S313: Transform the variables into dimensionless values ​​between 0 and 1 using the max-min and Z-score standardization methods. The formula for calculating the max-min standardization is as follows: (10), in, Represents the first of the original input variables Each sample value This represents the minimum value of the variable across all samples. This represents the maximum value of the variable across all samples. This represents the normalized variable value, with a range of values. The standardized Z-score calculation formula is as follows: (11), in, This represents the mean of the variable, i.e. , This represents the standard deviation of the variable; Step S314: Unify the temporal resolution of remote sensing images to 8 days, while maintaining the spatial resolution at 300m.

6. The method for extracting cyanobacterial blooms based on spectral and phenological characteristics according to claim 1, characterized in that, Step S4 includes the following sub-steps: Step S41: Based on different PCI thresholds, the outbreak intensity is divided into three levels: severe, moderate, and mild. The PCI intervals corresponding to severe, moderate, and mild are determined by historical observation data or expert experience to quantify the severity of the algal bloom. Step S42: Count the number of pixels in cyanobacterial blooms. And multiplied by pixel resolution The formula for calculating the coverage area is as follows: (12), in, Indicates the coverage area. This indicates the number of pixels representing cyanobacterial blooms. Image resolution; Step S43: Determine the hot and cold spots of cyanobacterial blooms using the hot and cold spot analysis method, as shown in the following formula: (13), in, express regional cyanobacterial blooms, The spatial weights of cyanobacterial blooms are represented by distance rules. After value standardization, if If the value is positive and significant, it belongs to a high-value clustering "hotspot" area of ​​cyanobacterial blooms; if... If the value is negative and significant, it belongs to the "cold spot" area of ​​low-value aggregation of cyanobacterial blooms; Step S44: Calculate the coverage area of ​​cyanobacterial blooms in different hotspot areas at different time scales.

7. The method for extracting cyanobacterial blooms based on spectral and phenological characteristics according to claim 1, characterized in that, Step S5 includes the following sub-steps: Step S51: Based on the 8-day PCI time series dataset of cyanobacterial bloom, the Savitzky-Golay filtering method is used for smoothing to provide continuous and smooth time series data for phenological feature extraction; Step S52: The extraction of phenological characteristic parameters uses a Gaussian fitting model, the formula of which is as follows: (14), in, Indicates PCI background value, The amplitude represents the maximum PCI value during the algal bloom period. Indicates the peak date of cyanobacterial blooms. Indicates the width of the algal bloom. This indicates the time step, and the duration is defined as the period from the date when the PCI increases to its maximum 50% to the date when the PCI decreases to its maximum 50%. Step S53: Describe the evolution trend of the time series of key phenological indicators for cyanobacterial blooms using Sen slope estimation and the Mann-Kendall trend test. The Sen slope is used to estimate the median slope of the time series trend, as shown in the following formula: (15), in, and These represent the time series at time points. and phenological characteristic values, This represents the rate of change of the median of a time series. , The length of the time series, when When, it indicates that the phenological parameters increase with time; when When, it indicates that the phenological parameter decreases over time; the Mann-Kendall trend test is used to determine whether the trend of the time series is significant, and the trend test is performed using the test statistic Z. The formulas for calculating Z are as follows: (16), (17), in, and These represent the time series at time points. and phenological characteristic values, Indicates the length of the time series. Represents the eigenvalue sign function. Represents statistics The variance is used to quantify the impact of repeated values ​​in a data series on trend testing. and The calculation formula is as follows: (18), (19), in, Indicates the length of the time series. This indicates the number of tie groups in the sequence. Indicates the first The number of groups / ties.