Lake cyanobacterial bloom pixel level prediction method based on multi-source data fusion
By fusing multi-source data and performing partitioned nonlinear fitting, a spatiotemporal dynamic suitability matrix is constructed to identify high-impact factors, achieving pixel-level accurate prediction of cyanobacterial blooms in lakes. This solves the problems of low simulation accuracy and poor applicability of traditional prediction models, and improves the spatiotemporal resolution and accuracy of predictions.
Patent Information
- Application Number
- CN202610069789.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2046-01-20
AI Technical Summary
Existing technologies have low simulation accuracy and poor applicability in predicting cyanobacterial blooms in lakes, making it difficult to cope with the dynamic changes in complex lake ecosystems. Traditional meteorological observation station data has low spatial resolution and fails to effectively combine multi-source data for accurate spatiotemporal matching.
By collecting multi-source basic data such as satellite imagery, wind field, and illumination data, core influencing factors are screened after preprocessing. Dynamic thresholds are determined by partitioning and fitting the cyanobacteria growth mechanism, a spatiotemporal dynamic pixel-level suitability matrix is constructed, high-influence factors are identified, a neighborhood iterative diffusion model is constructed, and the coupling of algal bloom growth and diffusion is realized to perform pixel-level prediction.
It improves the spatiotemporal resolution and accuracy of pixel-level predictions, provides a refined basis for lake algal bloom control decisions, reduces prediction errors caused by homogenization of regional characteristics, dynamically matches the environmental response patterns of cyanobacteria growth, and enhances the scientific validity and rationality of prediction results.
Smart Images

Figure CN121543841A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of algae prediction technology, specifically a pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion. Background Technology
[0002] Currently, technologies such as satellites, drones, real-time monitoring cameras, water quality instruments, and the Internet of Things have enabled dynamic monitoring of algal blooms. However, forecasting and early warning are crucial for targeted prevention and control, providing a basis for pre-emptive measures and emergency responses to avoid ecological disasters. However, the highly dynamic nature and complex outbreak mechanisms of algal blooms pose a challenge to the development of predictive models. Data obtained from traditional meteorological observation stations typically requires interpolation methods to generate the spatial distribution of meteorological factors, resulting in images with low spatial resolution, insufficient to support the construction of pixel-level predictive models. Using a uniform threshold to delineate regions ignores the spatiotemporally heterogeneous physiological mechanisms of cyanobacterial growth and fails to combine multi-source data for accurate spatiotemporal matching, leading to low simulation accuracy, poor applicability, and difficulty in coping with the dynamic changes of complex lake ecosystems. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, this invention proposes a pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion. This invention primarily addresses the problems of low simulation accuracy and poor applicability.
[0004] This invention provides a pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion, comprising: S1: Collect multi-source basic data of lakes at the pixel level from satellite imagery and preprocess it to output a standardized pixel-level dataset.
[0005] S2: Screen core influencing factors from the standardized pixel-level dataset, calculate the algae index to generate a binary distribution product, combine the cyanobacteria growth mechanism to determine the dynamic threshold for partition fitting, and output the partition factor suitability parameter table.
[0006] S3: Based on the wind field and illumination data in the standardized pixel-level dataset, perform spatiotemporal matching with the zoning factor suitability parameter table, invert the cyanobacteria proliferation rate and adjust the factor weights, combine the regional threshold to calculate the pixel comprehensive suitability, and output the spatiotemporal dynamic pixel-level suitability matrix.
[0007] S4: Based on the temporal characteristics of the spatiotemporal dynamic pixel-level fitness matrix, combined with the matched meteorological and algal bloom data, lagged effect factors are identified through correlation analysis and causal tests. High-impact factors are then selected by feature sorting and output as a list of high-impact factors.
[0008] S5: Based on the spatiotemporal dynamic pixel-level suitability matrix quantization and the list of high-impact factors, construct diffusion rules, extract initial algal bloom pixels and determine the diffusion starting point, and output the initial algal bloom pixel distribution results.
[0009] S6: Using the high-precision initial algal bloom pixel distribution as the diffusion starting point, a neighborhood iterative diffusion model is constructed using a spatiotemporal dynamic pixel-level suitability matrix. The algal bloom growth and diffusion are coupled through the proliferation coefficient and migration probability, and the pixel-level algal bloom prediction map is output through iterative simulation.
[0010] According to the pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion provided by the present invention, the specific steps for outputting the zoning factor suitability parameter table in step S2 are as follows: S21: Extract core impact factors from the standardized pixel-level dataset and calculate NDVI and FAI indices using preprocessed satellite imagery from the standardized pixel-level dataset.
[0011] S22: Generate pixel-level binary distribution product and coverage area time series data using threshold segmentation based on NDVI and FAI indices.
[0012] S23: Combining the spatial and temporal heterogeneous physiological mechanisms of cyanobacteria growth, and using time-series data of coverage area to assist in zoning, the lake is divided into nearshore nutrient enrichment zone, lake center hydrodynamic sensitive zone, and transition zone by a nonlinear fitting algorithm to obtain the zoning results.
[0013] S24: Based on the pixel-level binary distribution products and coverage area time series data of the zoning results, analyze the correlation between algal blooms and various factors in different regions, determine the dynamic suitability threshold of the core influencing factors in each region, and output the zoning factor suitability parameter table.
[0014] According to the pixel-level prediction method for lake cyanobacterial blooms based on multi-source data fusion provided by the present invention, the specific steps for obtaining the partitioning results in step S23 are as follows: S231: Based on the core characteristics of the spatial and temporal heterogeneous physiological mechanism of cyanobacteria growth, the physiological response differences of nearshore nutrient-rich areas, lake center hydrodynamic sensitive areas, and transition areas are determined, forming the theoretical basis for zoning.
[0015] S232: Based on the zoning theory, extract the spatial distribution characteristics, expansion and contraction trends and peak occurrence patterns of cyanobacteria coverage area time series data to identify high and low frequency regions of cyanobacteria aggregation.
[0016] S233: Based on the high and low frequency regions, a partitioned nonlinear fitting algorithm is used to construct a fitting target model with the spatial correlation and temporal fluctuation of the coverage area time series data as fitting variables.
[0017] S234: Input the time series data of the coverage area of all pixels in the lake into the fitting model, and divide the initial partitions through iterative calculation.
[0018] S235: By comparing the temporal variation patterns of cyanobacteria coverage area in different zones with the corresponding nutrient and hydrodynamic adjustment fitting parameters, the spatial zoning results of nearshore nutrient enrichment zone, lake center hydrodynamic sensitive zone, and transition zone are output.
[0019] According to the pixel-level prediction method for lake cyanobacterial blooms based on multi-source data fusion provided by the present invention, the specific steps for outputting the spatiotemporal dynamic pixel-level fitness matrix in step S3 are as follows: S31: Perform spatiotemporal alignment preprocessing on the wind field, illumination spatiotemporal data and zoning factor suitability parameters in the standardized pixel-level dataset, unify the temporal granularity and spatial projection, and output the parameter table associated with the dataset.
[0020] S32: Perform factor dimension matching and attribute association based on the associated dataset, extract the fitness benchmark parameters corresponding to each spatiotemporal unit, and output the spatiotemporal matching dataset.
[0021] S33: Substitute the spatiotemporal matching dataset into the cyanobacteria proliferation rate inversion model to calculate the initial proliferation rate, iteratively adjust the partitioning factor weights through error analysis, and output the factor weight set.
[0022] S34: Construct a pixel-level comprehensive suitability calculation model based on the factor weight set and the preset factor thresholds for each region. Sum the wind field and illumination suitability contribution values of each pixel in a weighted manner and output the initial pixel comprehensive suitability dataset.
[0023] S35: Smooth the initial pixel-level fitness dataset, remove outliers, supplement edge pixel interpolation results, and output a pixel-level fitness matrix.
[0024] According to the pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion provided by the present invention, the specific steps for outputting the factor weight set in step S33 are as follows: S331: Extract the wind field frictional wind speed, effective radiation intensity of light, and corresponding zoning factor suitability benchmark parameters of each spatiotemporal unit in the spatiotemporal matching dataset, and output the model input dataset.
[0025] S332: Substitute the model input dataset into the energy conservation-based cyanobacteria proliferation rate inversion model, calculate the initial proliferation rate of each spatiotemporal unit, and output the initial proliferation rate dataset.
[0026] S333: Based on the initial proliferation rate dataset and the measured cyanobacterial biomass growth rate, the model error is calculated using the residual sum of squares to construct the error objective function.
[0027] S334: Use the error objective function to adjust the weights of the wind field factor and the illumination factor until the error is less than the threshold, and output the factor weight set.
[0028] According to the pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion provided by the present invention, the specific steps for outputting the list of high impact factors in step S4 are as follows: S41: Align and calibrate the temporal feature data of the spatiotemporal dynamic pixel-level fitness matrix and the meteorological and algal bloom monitoring data to output a multi-source fusion dataset.
[0029] S42: Analyze the temporal changes of fitness and the fluctuations of meteorological and algal bloom indicators based on the multi-source fusion dataset, calculate the Pearson correlation coefficient, preliminarily screen potential correlation factors, and output a candidate list of correlation factors.
[0030] S43: Based on the candidate list of related factors, use the Granger causality test to verify the causal relationship between the candidate factors and the occurrence of algal blooms, identify key driving factors with lag effects, and output the causal factor set.
[0031] S44: Construct a feature importance ranking model based on the causal factor set, use the random forest algorithm to quantify the contribution of each factor to the impact of algal blooms, and output the factor priority sequence.
[0032] S45: Extract high-impact factors with contributions higher than a preset threshold from the factor priority sequence, calculate the optimal lag time parameter for each factor, and output a list of high-impact factors.
[0033] According to the pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion provided by the present invention, the specific steps for outputting the initial bloom pixel distribution results in step S5 are as follows: S51: Extract the quantitative indicators of the basic conditions for algal bloom occurrence corresponding to each pixel in the spatiotemporal dynamic pixel-level suitability matrix, perform data standardization and value range normalization, and output a pixel-level quantitative dataset.
[0034] S52: Based on the pixel-level quantized dataset, an improved phytoplankton index and phycocyanin index collaborative identification algorithm is introduced to make a preliminary judgment on the probability of pixel-level algal blooms and output a set of suspected algal bloom pixels.
[0035] S53: Analyze the driving mechanism of factor lag effect on the occurrence of algal blooms based on the suspected algal bloom pixel set and the high-impact factor screening list, and output the optimized set of suspected algal bloom pixels.
[0036] S54: Based on the optimized set of suspected algal bloom pixels and the wind-hydrodynamic coupling diffusion rule, simulate the diffusion trend of algal bloom under different wind and hydrodynamic conditions, correct the attribution of edge pixels, and output the algal bloom pixel distribution dataset.
[0037] S55: Perform spatial consistency verification and accuracy verification on the water bloom pixel distribution dataset, remove misjudged pixels and supplement the interpolation results of missed pixels, and output the initial water bloom pixel distribution results.
[0038] According to the pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion provided by the present invention, the specific steps for outputting the suspected bloom pixel set in step S52 are as follows: By retrieving the pixel-level quantized dataset, the basic conditions for algal bloom occurrence corresponding to each pixel are extracted. Combined with the calculation thresholds of PAI and PCI, the input dataset for index calculation is output.
[0039] The input dataset for index calculation is substituted into the improved PAI and PCI collaborative calculation model to calculate the PAI and PCI values of each pixel. Pixels that initially meet the conditions are screened through the double index intersection judgment rule, and the double index preliminary screened pixel set is output.
[0040] Based on the initial screening of the pixel set using the dual-index method, pixel neighborhood correlation analysis is introduced. The single-pixel judgment results are corrected by combining the PAI and PCI distribution characteristics of surrounding pixels, isolated abnormal pixels are removed, and the corrected set of suspected candidate pixels for algal blooms is output.
[0041] The corrected set of suspected algal bloom candidate pixels is compared with the pixel feature data of historical algal blooms to verify the matching degree. Pixels with a matching degree higher than the preset threshold are retained, and finally, the set of suspected algal bloom pixels labeled with the probability of algal bloom occurrence is output.
[0042] According to the pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion provided by the present invention, the specific steps for outputting the pixel-level bloom prediction map in step S6 are as follows: S61: Extract the initial algal bloom pixel distribution results and the spatial coordinates, fitness values and neighborhood topological relationships of the pixels in the spatiotemporal dynamic pixel-level fitness matrix, and output the pixel modeling dataset.
[0043] S62: Construct a neighborhood pixel iterative diffusion model based on the pixel modeling dataset, correlate the proliferation coefficient with the neighborhood spatial migration probability, and output the initial simulation result set of the model.
[0044] S63: Based on the initial simulation result set of the model, select a small number of in-situ observation samples for pixel-level matching, calibrate the accuracy of pixel biomass, call up multi-period remote sensing time series verification data, and output the simulation intermediate dataset.
[0045] S64: Based on the simulated intermediate dataset, reverse the threshold of the spatiotemporal dynamic cell-level suitability matrix and the parameters of the iterative diffusion model, evaluate the simulation effect through the confusion matrix, and output the simulation result set.
[0046] S65: Based on the simulation results set, the core parameters of the model are adjusted using the Bayesian optimization algorithm, and the evolution process of algal blooms at different time periods is simulated iteratively to output a pixel-level algal bloom prediction map.
[0047] According to the pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion provided by the present invention, the specific steps for outputting the initial simulation result set of the model in step S62 are as follows: Extract the fitness value, neighborhood topology, and diffusion parameters of pixels from the pixel modeling dataset, establish a quantitative correlation formula between fitness and proliferation coefficient, and output the model input dataset.
[0048] The core framework of the neighborhood cell iterative diffusion model is constructed based on the model input dataset. The rules for calculating the migration probability of neighborhood cells are defined, and the proliferation coefficient and migration probability are bound together to form a coupled driving logic, which outputs the initial iterative diffusion model.
[0049] Starting with the initial algal bloom pixel, the pixel suitability and neighborhood topology data are substituted into the initial iterative diffusion model to perform the first round of iterative calculation, generate simulation results of algal bloom growth and diffusion, and output the first round of iterative result set of the model.
[0050] Using the first iteration result set of the model, set multiple iteration step sizes for different time periods, repeat the process, simulate the spatiotemporal evolution of algal blooms in different time periods, and output the initial simulation result set of the model.
[0051] This invention provides a pixel-level prediction method for lake cyanobacterial blooms based on multi-source data fusion. By integrating and standardizing multi-source pixel-level data, combining cyanobacterial physiological mechanism partitioning and fitting, constructing a spatiotemporal dynamic fitness matrix, screening high-impact lag factors, and coupling iterative simulation of migration-proliferation processes, it achieves accurate pixel-level prediction of blooms, improves the spatiotemporal resolution and accuracy of prediction, and provides refined technical support for the prevention and control of lake blooms.
[0052] The beneficial effects of this invention are as follows: 1. This invention eliminates noise and cloud obstruction interference from remote sensing images by unifying the spatial resolution and coordinate system of multi-source data and combining preprocessing techniques such as radiometric calibration and atmospheric correction. This constructs a high-quality, standardized pixel-level dataset, laying a reliable data foundation for subsequent analysis. Simultaneously, the method overcomes the limitations of traditional unified thresholds by incorporating the spatiotemporally heterogeneous physiological mechanisms of cyanobacterial growth and employing a partitioned nonlinear fitting algorithm to divide lake functional zones, specifically determining the dynamic suitability thresholds for factors in each region. This partitioned differentiated modeling approach accurately matches the environmental response patterns of cyanobacterial growth in different regions, effectively reducing prediction errors caused by homogenized regional feature processing and significantly improving the accuracy of pixel-level predictions.
[0053] 2. This invention identifies high-impact meteorological factors with significant lag effects on algal blooms, clarifies the time window of these factors' effects, and overcomes the deficiency of traditional models that ignore factor lag. Based on this, a wind-hydrodynamic coupled diffusion rule is constructed, combining the influence of wind speed and direction on cyanobacterial migration with the cyanobacterial proliferation rate corrected for the lag effect of high-impact factors. This achieves a dynamic coupled simulation of both cyanobacterial migration and proliferation processes. This coupling mechanism more closely reflects the real process of natural algal bloom evolution, upgrading the model from simple static statistical prediction to dynamic process simulation, significantly enhancing the scientific rigor and rationality of the prediction results.
[0054] 3. This invention employs a multi-round model calibration and parameter optimization process. Through in-situ sample calibration and remote sensing data correction, combined with confusion matrix evaluation, it continuously optimizes intermediate simulation results and reversely corrects the suitability matrix threshold and iterative diffusion model parameters, effectively reducing the model's systematic error. A Bayesian optimization algorithm is introduced to iteratively adjust core parameters with simulation accuracy as the target, generating a multi-time-period algal bloom evolution simulation dataset. The accuracy and practicality of the results are verified using actual monitoring data. This progressive optimization strategy ensures the model's stability under different spatiotemporal scenarios, and the output pixel-level algal bloom prediction map accurately reflects the spatiotemporal distribution characteristics of algal blooms, providing an intuitive and reliable decision-making basis for lake algal bloom prevention and ecological governance.
[0055] Initial pixels serve as the anchor points for subsequent spatiotemporal evolution simulations of algal blooms. Their accuracy directly determines the reliability of the final pixel-level prediction results. The spread of cyanobacterial blooms gradually originates from localized dominant areas. Blindly predicting the entire lake's pixels directly ignores the ecological pattern of bloom accumulation followed by diffusion. Initial pixels clearly identify the core source of the bloom, providing a realistic spatial starting point for subsequent neighborhood-based iterative diffusion simulations. High-impact factors, lag effect parameters, and zoning suitability thresholds, selected through screening, require initial pixels for accurate implementation. Only by identifying initial pixels can we specifically analyze the factor contribution of that region, avoiding the substitution of local core driving factors with lake-wide average parameters, and ensuring the diffusion rules are more realistic. Without extracting initial pixels, directly performing time-series simulations and diffusion calculations on all lake pixels would result in a large amount of invalid computation. Initial pixels can narrow the simulation scope, focusing on potential bloom areas, significantly improving computational efficiency while maintaining prediction accuracy, and meeting the needs of pixel-level high-resolution prediction. Attached Figure Description
[0056] The invention will now be further described with reference to the accompanying drawings.
[0057] Figure 1 This is a flowchart illustrating the steps of a pixel-level prediction method for lake cyanobacterial blooms based on multi-source data fusion, provided in an embodiment of the present invention. Figure 2This is a flowchart of a pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion, provided in an embodiment of the present invention. Figure 3 This is a flowchart illustrating the steps for outputting the initial algal bloom pixel distribution results provided in an embodiment of the present invention. Detailed Implementation
[0058] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below according to specific embodiments.
[0059] like Figures 1 to 3 As shown in the figure, an embodiment of the present invention provides a pixel-level prediction method for cyanobacterial blooms in lakes based on multi-source data fusion. The system includes: S1: Collect multi-source, pixel-level basic data for lakes, including high spatiotemporal resolution remote sensing images, pixel-level distributed water temperature data, nitrogen-phosphorus ratio spatial interpolation data, and refined meteorological grid data, unifying the spatial resolution to 250m and the coordinate system to WGS-84UTM projection. Perform satellite image data preprocessing: eliminate image noise through radiometric calibration, atmospheric correction, and geometric fine correction; combine cloud pollution pixel restoration and outlier spatiotemporal filtering algorithms to fill in pixel data in areas obscured by optical remote sensing clouds; and output a standardized pixel-level dataset.
[0060] Remote sensing imagery was downloaded from a satellite data sharing platform using optical / radar images. Data with low cloud cover and no extreme weather interference were selected, converted in format, and standardized in coordinate system. Water temperature data was generated as pixel-level data through downscaling and interpolation using measured data from hydrological stations and hydrodynamic model simulations. Nitrogen-phosphorus ratio data was generated as resolution data through gridded in-situ sampling analysis within the lake and combined with spatial interpolation. Measured data from surrounding meteorological stations were collected, simulated using the China Meteorological Administration's refined numerical weather prediction model, and downscaled. All data were standardized to the WGS-84UTM coordinate system and spatial resolution, and timestamp alignment was completed.
[0061] Using standardized optical and radar remote sensing imagery as the processing target, noise is eliminated, data gaps are filled, and a standardized pixel-level dataset is ultimately output. The specific process involves: converting image DN values to radiance values to eliminate sensor errors; eliminating atmospheric interference to obtain the true surface reflectance; correcting geometric distortion to ensure pixel coordinates match actual geographical locations; utilizing radar image penetration to repair pixels in cloud-obscured areas of the optical imagery; removing noisy pixels to ensure data continuity; and finally, matching the preprocessed remote sensing imagery with other data at the pixel level to generate a standardized pixel-level dataset.
[0062] S2: Core influencing factors such as water temperature, nitrogen-phosphorus ratio, wind speed, and light duration are screened from the standardized pixel-level dataset. NDVI (Normalized Difference Vegetation Index) and FAI (Follicular Index) are calculated based on preprocessed satellite imagery. Pixel-level binary distribution products and time-series data of coverage area are generated through threshold segmentation and spatial clustering algorithms. Combining the spatiotemporal heterogeneous physiological mechanisms of cyanobacterial growth, a zonal nonlinear fitting algorithm is used to divide the lake into nearshore nutrient-rich areas, central hydrodynamic sensitive areas, and transition zones. Based on the pixel-level binary distribution products and time-series data of coverage area, dynamic suitability thresholds for each factor in different regions are determined, and a zonal factor suitability parameter table is output. This achieves precise correlation between standardized data and regionally differentiated ecological mechanism thresholds, overcoming the limitations of traditional unified thresholds.
[0063] S21: Extract core influencing factors such as water temperature, nitrogen-phosphorus ratio, wind speed, and sunshine duration from the standardized pixel-level dataset, and simultaneously call the preprocessed satellite images in the dataset to calculate the NDVI and FAI indices.
[0064] The specific steps for calculating the NDVI and FAI indices using multispectral reflectance data from preprocessed satellite imagery are as follows: Extracting the near-infrared band from preprocessed satellite images (denoted as...) ) and red band (denoted as The pixel-level reflectance data ensures that the spatial positions of pixels in the two bands are perfectly matched, without any offset or misalignment.
[0065] The formula for calculating the NDVI index is:
[0066] The NDVI value of the entire pixel is obtained by performing calculations on each pixel individually.
[0067] Pixel-level reflectance data for the blue band (denoted as ρBlue) is extracted from the preprocessed satellite imagery to ensure consistent spatial resolution and corresponding positions for pixels in the blue, red, and near-infrared bands. The FAI index calculation formula is expressed as:
[0068] In the formula, , , The center wavelengths of blue light, red light, and near-infrared bands are respectively used. Based on the matched three-band pixel reflectance data, the FAI value of each pixel is calculated.
[0069] S22: Generate pixel-level binary distribution product and coverage area time series data using threshold segmentation based on NDVI and FAI indices.
[0070] Using the pre-calculated pixel-level NDVI and FAI index datasets, and combining them with a preset threshold, binary classification is performed. The specific steps are as follows: Determine the NDVI threshold range [TNDVI] low TNDVI high ] and FAI threshold range [TFAI low TFAI high ], of which TNDVI low TNDVI high These are the lower and upper thresholds for NDVI classification, respectively, TFAI low TFAI high These are the lower and upper thresholds for FAI classification, respectively.
[0071] Thresholding is applied to the NDVI and FAI indices of each pixel, and binary classification is achieved using a logical AND operation. The binary classification formula is as follows: ;
[0072] In the formula, For time t The binarization result of the location pixel is used, where 1 represents the target land class and 0 represents the non-target land class. The operation is performed pixel by pixel and time by time to generate a pixel-level binary distribution product sequence.
[0073] Extract the total number of pixels at each time step from the time-series binary distribution product, and determine the area S of a single pixel. pixel S pixel The calculation formula is expressed as follows:
[0074] In the formula, Rx and Ry are the horizontal and vertical spatial resolutions of the image, respectively.
[0075] Count the number of pixels with a value of 1 in the binary results at time t. Calculate the coverage area of the target land type at that time. Repeat this statistical process for all time-series images to generate time-series coverage area data. The formula for the coverage area of the target land type is expressed as: ;
[0076] In the formula, The coverage area of the target land type. This represents the number of pixels.
[0077] S23: Combining the spatial and temporal heterogeneous physiological mechanisms of cyanobacteria growth, and using time-series data of coverage area to assist in zoning, the lake is divided into nearshore nutrient enrichment zone, lake center hydrodynamic sensitive zone, and transition zone by a nonlinear fitting algorithm to obtain the zoning results.
[0078] S231: Extract the core characteristics of the spatial and temporal heterogeneous physiological mechanisms of cyanobacterial growth, and clarify the differences in physiological responses in the nearshore nutrient-rich zone (high nutrient salts, weak hydrodynamics, and easy cyanobacterial blooms), the lake center hydrodynamic sensitive zone (strong hydrodynamics, rapid nutrient salt diffusion, and dispersed cyanobacterial distribution), and the transition zone (nutrient salt and hydrodynamic conditions are between the former two), as the theoretical basis for zoning.
[0079] S232: Import the time series data of cyanobacteria coverage area generated in the previous steps, extract the spatial distribution characteristics, expansion / contraction trends and peak occurrence patterns of the coverage area under each time series, and help identify the high-frequency and low-frequency regions of cyanobacteria aggregation.
[0080] S233: Initialize the partitioning algorithm parameters. Based on the core logic of the partitioning nonlinear fitting algorithm, use the spatial correlation and temporal fluctuation of the coverage area time series data as fitting variables, and set the fitting objective function (such as minimizing the coverage area time series variation coefficient within the region and maximizing the variation coefficient between regions).
[0081] S234: Input the time series data of the coverage area of the entire lake pixels into the fitting model, divide the initial partitions through iterative calculation, and then correct the boundaries of the initial partitions by combining the physiological mechanism of cyanobacteria growth to eliminate unreasonable fragmented partitions.
[0082] S235: Verify the rationality of the zoning results. By comparing the temporal variation of cyanobacteria coverage area in different zones with the matching degree of environmental characteristics such as nutrients and hydrodynamics in the corresponding areas, adjust the fitting parameters until the zoning results meet the physiological mechanism expectations. Finally, output the spatial zoning results of nearshore nutrient enrichment zone, lake center hydrodynamic sensitive zone, and transition zone.
[0083] S24: Based on the pixel-level binary distribution products and coverage area time series data of the zoning results, analyze the correlation between algal blooms and various factors in different regions, determine the dynamic suitability threshold of the core influencing factors in each region, and output the zoning factor suitability parameter table.
[0084] S241: Integrate all the data from the preceding steps, including the zoning results, pixel-level binary distribution products for each region, time-series data of coverage area, and previously extracted NDVI, FAI indices and corresponding environmental impact factor data (such as nutrient concentration, water temperature, flow rate, etc.).
[0085] S242: Data is split into zones, and the proportion of cyanobacteria coverage area in the nearshore nutrient-rich zone, the hydrodynamically sensitive zone in the lake center, and the transition zone is calculated for each time series. At the same time, the mean, extreme values and fluctuation characteristics of each influencing factor under the corresponding time series are extracted.
[0086] S243: Construct a correlation analysis model between the proportion of cyanobacteria coverage area and influencing factors in each zone, and screen out the core factors that have a significant impact on the growth of cyanobacteria in each zone through correlation analysis, partial least squares regression and other methods.
[0087] S244: For the core influencing factors of each partition, determine the type of response relationship between them and the proportion of cyanobacteria coverage area (such as positive correlation, negative correlation, non-linear correlation), and use the coverage area proportion range suitable for cyanobacteria growth (such as dominant growth range, inhibited growth range) as a reference to classify the dynamic suitability level of the core factors.
[0088] S245: Calculate the threshold range of each core factor under different fitness levels based on the response relationship, such as the suitable growth threshold and growth inhibition threshold of total phosphorus in nearshore nutrient-rich areas, and the suitable threshold of flow velocity in hydrodynamically sensitive areas in the center of the lake.
[0089] S246: The names of the core influencing factors, suitability levels, threshold ranges, and corresponding cyanobacterial growth response characteristics of each partition should be compiled into a partition factor suitability parameter table in a unified format.
[0090] S3: Based on the refined wind field grid data and the interpolated illumination duration data from the meteorological station in the standardized pixel-level dataset, spatiotemporal matching is performed with the time series data of algal bloom coverage area associated with the zoning factor suitability parameter table. The cyanobacterial proliferation rate is inverted based on the remote sensing time series data of the standardized pixel-level dataset. The cyanobacterial proliferation rate is inverted and the factor weights are adjusted. The pixel comprehensive suitability is calculated in combination with the regional threshold, and the spatiotemporal dynamic pixel-level suitability matrix is output.
[0091] S31: Perform spatiotemporal alignment preprocessing on the wind field, illumination spatiotemporal data and zoning factor suitability parameters in the standardized pixel-level dataset, unify the temporal granularity and spatial projection, and output the parameter table associated with the dataset.
[0092] Based on the wind field, illumination spatiotemporal data, and zoning factor suitability parameter table within the standardized pixel-level dataset, spatiotemporal coordinates and timestamp information of each data point are extracted to determine spatial projection type and temporal granularity differences, outputting a dataset to be aligned with spatiotemporal metadata. A spatial projection transformation algorithm is used to unify the spatial reference of the wind field, illumination data, and parameter table to the same geographic coordinate system, eliminating spatial location bias and outputting an intermediate dataset with a consistent spatial reference. The time series is resampled according to the principle of high-frequency to low-frequency conversion or interpolation frequency compensation, unifying data at different temporal granularities to the target time interval, outputting a transitional dataset with uniform temporal granularity. A spatiotemporal correlation algorithm is used to establish a one-to-one correspondence between wind field, illumination data, and zoning factor suitability parameters, eliminating redundant data with spatiotemporal mismatches, and outputting a parameter table-correlated dataset.
[0093] S32: Perform factor dimension matching and attribute association based on the associated dataset, extract the fitness benchmark parameters corresponding to each spatiotemporal unit, and output the spatiotemporal matching dataset.
[0094] Based on the parameter table association dataset, the spatiotemporal unit codes of wind field and illumination data are parsed, and the dimensional labels of the zoning factor suitability parameters are mapped. Dimensional mapping rules are established, and a dimensional mapping relationship table is output. Attribute data such as wind field speed and illumination duration in the association dataset are matched and bound with factor thresholds and weight coefficients in the suitability parameter table, and an intermediate dataset with attribute binding is output. Grouping and filtering are performed by spatiotemporal unit, and the suitability benchmark parameters corresponding to each spatiotemporal unit are extracted, including core indicators such as factor suitability upper and lower limits and benchmark contribution values, and the parameter extraction result set is output. Spatiotemporal consistency verification is performed, data entries with missing spatiotemporal units or abnormal parameter values are removed, and parameter interpolation results of marginal spatiotemporal units are supplemented. Finally, a spatiotemporal matching dataset with suitability benchmark parameters is output.
[0095] S33: Substitute the spatiotemporal matching dataset into the cyanobacteria proliferation rate inversion model to calculate the initial proliferation rate, iteratively adjust the partitioning factor weights through error analysis, and output the factor weight set.
[0096] S331: By retrieving the spatiotemporal matching dataset, extract the wind field frictional wind speed u, effective radiation intensity I, and corresponding zoning factor suitability benchmark parameter f0 for each spatiotemporal unit, and output the model input dataset with physical quantity labels.
[0097] S332: Based on the model input dataset with physical quantity labels, substitute it into the cyanobacterial proliferation rate inversion model based on energy conservation, calculate the initial proliferation rate r0 of each spatiotemporal unit, and output the initial proliferation rate dataset.
[0098] The formula for cyanobacterial photosynthetic energy absorption is expressed as:
[0099] In the formula, I represents the effective light radiation intensity, and α represents the photosynthetic radiation utilization efficiency of cyanobacteria. This refers to the suitability of the light factor. This represents the equivalent energy flux of wind-driven nutrient diffusion in water bodies. For wind field factor suitability, , The weights are for wind field and light intensity factors.
[0100]
[0101] In the formula, t represents the energy-to-biomass conversion efficiency, t represents the effective light duration per day, and C represents the energy equivalent per unit biomass of cyanobacteria.
[0102] S333: Based on the initial proliferation rate dataset, the measured cyanobacterial biomass growth rate rob is introduced. s The model error is calculated by the sum of squared residuals, and an error objective function is constructed.
[0103] S334: Using the error objective function, the gradient descent method is used to iteratively adjust the wind field factor weights. Light factor weight (constraint (until the error is less than the threshold, then output the optimized factor weight set) .
[0104] S34: Construct a pixel-level comprehensive suitability calculation model based on the factor weight set and the preset factor thresholds for each region. Sum the wind field and illumination suitability contribution values of each pixel in a weighted manner and output the initial pixel comprehensive suitability dataset.
[0105] By retrieving the optimized factor weight set and the preset factor thresholds for each region, the wind field factor weights are determined. Light factor weight The model outputs a set of model parameters with weights and threshold constraints, along with upper and lower limits for the suitability of each factor. Based on this set of parameters, a pixel-level comprehensive suitability calculation model is constructed. The core of the model is a weighted summation formula:
[0106] In the formula, , The system extracts wind field and illumination suitability values for each pixel, and outputs a document explaining the model structure and calculation rules. Based on this document, it extracts the raw wind field and illumination suitability data for each pixel, substitutes them into a weighted summation formula for calculation, and truncates and corrects suitability values exceeding the range based on regional factor thresholds, outputting a single-pixel suitability calculation result set. Using this single-pixel suitability calculation result set, the data is integrated according to spatiotemporal units, linking the pixel's geographic coordinates and timestamp information, removing pixel data with calculation anomalies, supplementing the interpolation calculation results for edge pixels, and finally outputting the initial pixel comprehensive suitability dataset.
[0107] S35: Smooth the initial pixel-level fitness dataset, remove outliers, supplement edge pixel interpolation results, and output a pixel-level fitness matrix.
[0108] By retrieving the initial pixel-level fitness dataset, the fitness value, spatiotemporal coordinates, and neighboring pixel association information of each pixel are extracted. An outlier threshold is set, and a dataset to be corrected with neighboring pixel association labels is output. Based on the dataset to be corrected with neighboring pixel association labels, outlier pixel values exceeding the preset threshold are filtered out. The neighboring pixel mean replacement method is used to remove outlier data, and blank pixel areas without data coverage at the edges are marked, outputting an intermediate dataset after outlier removal. Based on the intermediate dataset after outlier removal, the inverse distance weighted interpolation algorithm is selected. The fitness values of valid edge pixels are used to interpolate the blank edge areas to supplement the fitness data of missing pixels, outputting a complete dataset after interpolation and completion. Using the complete dataset after interpolation and completion, Gaussian filtering is used for spatial smoothing correction to reduce the abrupt differences in fitness between adjacent pixels, improve the spatiotemporal continuity of the data, and output a spatiotemporally and dynamically consistent pixel-level fitness matrix.
[0109] S4: Based on the temporal characteristics of the spatiotemporal dynamic pixel-level fitness matrix, and combined with the spatiotemporally matched meteorological factors and algal bloom status correlation data, conduct a lag effect analysis of meteorological factors: using Pearson correlation analysis and Granger causality test, identify meteorological factors with significant lag effects on algal bloom occurrence, and through random forest feature importance ranking, screen high-impact factors from the core influencing factors and output a high-impact factor screening list.
[0110] S41: By retrieving the temporal feature data of the spatiotemporal dynamic pixel-level fitness matrix and matching the meteorological and algal bloom monitoring data, the alignment and calibration of the data temporal dimension are completed, the time series length and sampling interval are unified, and a multi-source fusion dataset with time series labels is output.
[0111] By retrieving the temporal feature data of the spatiotemporal dynamic pixel-level fitness matrix and matching meteorological and algal bloom monitoring data, timestamps and sampling frequency information of each dataset are extracted to clarify the temporal differences between the data, and a multi-source dataset to be aligned with temporal metadata is output. The maximum time span within the study period is selected as the unified time series length, and the resampling principle of aligning high-frequency data with low-frequency data is determined, outputting an intermediate dataset with time series length normalized. Linear interpolation is used to supplement data at missing time nodes, unifying datasets with different sampling intervals to the target time resolution, and outputting a multi-source dataset with consistent time granularity. A one-to-one mapping relationship between pixel fitness temporal features and corresponding time node meteorological and algal bloom data is established through a spatiotemporal association algorithm, and redundant data with temporal matching failures are eliminated, finally outputting a multi-source fused dataset with time series labels.
[0112] S42: Based on the multi-source fusion dataset with time-series labels, conduct correlation analysis on the time-series changes in fitness and the fluctuations of meteorological and algal bloom indicators, calculate the Pearson correlation coefficient, preliminarily screen potential correlation factors, and output a candidate list of correlation factors.
[0113] By retrieving a multi-source fusion dataset with time-series labels, the time-series sequences of fitness, meteorological factors, and algal bloom indicators for each spatiotemporal unit are extracted. Sequence dimension matching and data cleaning are completed, outputting a time-series analysis dataset with feature labels. For the time-series changes of fitness and each meteorological and algal bloom indicator, Pearson correlation coefficients are calculated group by group to clarify the strength of linear association and positive / negative correlation between variables, outputting a correlation coefficient matrix and a significance test result set. Thresholds for the absolute value of correlation coefficients and significance levels are set, and potential associated factors that pass the threshold test are screened, outputting a candidate list of associated factors including factor name, correlation coefficient, and significance level.
[0114] S43: Based on the candidate list of related factors, the Granger causality test method is introduced to verify the causal relationship between the candidate factors and the occurrence of algal blooms, identify key driving factors with lag effects, and output a set of causal factors with lag attributes.
[0115] By retrieving a candidate list of related factors, the time-series data of each factor in the list and the time-series sequence of algal bloom occurrence are extracted. The stationarity of the data is tested and differencing is performed, outputting a stationary time-series dataset that meets the prerequisites for Granger causality testing. Based on the stationary time-series dataset, a Granger causality test model of the factor time series and algal bloom occurrence time series is constructed. Hypothesis testing is performed with different lag orders, and the F-statistic and significance p-value are calculated. A causality test result table for each factor is output. Based on the causality test result table, factors with p-values less than the significance threshold are selected, the optimal lag order for each factor and algal bloom occurrence is determined, key driving factors with lag effects are marked, and a candidate factor set labeled with lag order is output. Using the candidate factor set labeled with lag order, a secondary screening is performed based on the explanatory power of the factors for algal bloom occurrence, eliminating redundant factors with insignificant lag effects, and outputting a causal factor set with lag attributes.
[0116] S44: Using a set of causal factors with lag attributes, construct a feature importance ranking model, use the random forest algorithm to quantify the contribution of each factor to the impact of algal blooms, and output a priority sequence of factors arranged in descending order of contribution.
[0117] This process extracts the time-series feature values, lag order parameters, and corresponding algal bloom state labels of each factor in a causal factor set with lag attributes. Data feature encoding and label mapping are then performed, outputting a labeled factor feature dataset. The labeled factor feature dataset is divided into training and testing sets. A random forest feature importance ranking model is constructed, setting core parameters such as the number of decision trees and maximum depth, outputting an initialized random forest model. The training set data is input into the initialized random forest model for iterative training. Model accuracy is evaluated using out-of-bag (OOB) data, and the contribution of each factor to the algal bloom is quantified by the reduction in Gini coefficient or mean squared error, outputting a factor contribution score table. Causal factors are sorted from highest to lowest contribution value, and the lag order attribute of each factor is labeled. Weakly influential factors with contribution values below a preset threshold are removed, outputting a factor priority sequence sorted in descending order of contribution.
[0118] S45: Based on the factor priority sequence, extract high-impact factors with contribution values higher than the preset threshold, simultaneously calculate the optimal lag time parameters for each factor, and output a list of high-impact factors and their corresponding lag parameters.
[0119] This process extracts the contribution value, optimal lag order, and corresponding algal bloom impact characteristics of each factor within the factor priority sequence, sets a contribution threshold, and outputs a factor feature dataset with threshold filtering conditions. Based on this dataset, the contribution of each factor is compared with the preset threshold, and high-impact factors with contributions exceeding the threshold are selected. Simultaneously, the optimal lag time parameter corresponding to each high-impact factor is extracted, and a preliminary list of high-impact factors is output. Cross-validation is performed using historical algal bloom event data to confirm the rationality of the lag time parameters. Factors that fail validation are removed, and a calibrated list of high-impact factors and a set of lag parameters are output. Factors are sorted in descending order of contribution, and key information such as the lag time range and intensity of each factor is labeled, resulting in a standardized list of high-impact factors and their corresponding lag parameters.
[0120] S5: Based on the basic conditions for the occurrence of algal blooms quantified by the spatiotemporal dynamic pixel-level suitability matrix, and combined with the improved phytoplankton index (PAI) and phycocyanin index (PCI) collaborative identification algorithm, the initial algal bloom pixels are accurately extracted. A list of high-impact factors and lag effect parameters are used to construct a wind field-hydrodynamic coupled diffusion rule. This rule considers both the influence of wind speed and direction on algal migration and corrects the algal proliferation rate through the lag effect of high-impact factors, thereby achieving the coupling of migration and proliferation processes and outputting high-precision initial algal bloom pixel distribution results.
[0121] S51: By retrieving the spatiotemporal dynamic pixel-level suitability matrix, extract the quantitative indicators of the basic conditions for algal bloom occurrence corresponding to each pixel, complete the data standardization and value range normalization processing, and output the pixel-level quantitative dataset with basic condition labels.
[0122] S52: Based on the pixel-level quantized dataset with basic condition labels, a collaborative identification algorithm of PAI (improved phytoplankton index) and PCI (phycocyanin index) is introduced to make a preliminary judgment on the probability of pixel algal blooms and output the matching results of the suspected algal bloom pixel set and the index.
[0123] By retrieving the pixel-level quantized dataset, the basic conditions for algal blooms corresponding to each pixel are extracted. Combined with the calculation thresholds of PAI and PCI, the input dataset for index calculation is output. The turbidity of the water body is retrieved with the aid of radar imagery, providing background parameters for PCI index calculation.
[0124] The input dataset for index calculation is substituted into the improved PAI and PCI collaborative calculation model to calculate the PAI and PCI values for each pixel. Pixels that initially meet the criteria are screened using the double-index intersection judgment rule, and a preliminary double-index screened pixel set is output. The pixel chlorophyll concentration (Chla) is calculated using the reflectance of specific bands in satellite remote sensing multispectral images, combined with the chlorophyll a inversion model. Based on the reflectance of the red and near-infrared bands of satellite images, turbidity inversion is performed using a turbidity inversion model. The accuracy is then calibrated using turbidimeter measurements at sampling points. If radar imagery is available, its scattering coefficient can also help optimize the turbidity inversion results to obtain the meta-turbidity (T). The PC value at the pixel scale is calculated using the reflectance of the characteristic bands corresponding to phycocyanin in satellite hyperspectral images, combined with the phycocyanin inversion model.
[0125] The formula for calculating the improved phytoplankton index (PAI) is as follows:
[0126] In the formula, T represents the pixel chlorophyll a concentration, and T represents the pixel turbidity. This is the calibration coefficient.
[0127] The modified phycocyanin index (PCI) is calculated using the following formula:
[0128] In the formula, PC represents the pixel phycocyanin concentration. , The minimum and maximum values of PC concentration in the study area are given. Background water reflectance, This is a correction factor.
[0129] Based on the initial screening of the pixel set using the dual-index method, pixel neighborhood correlation analysis is introduced. The single-pixel judgment results are corrected by combining the PAI and PCI distribution characteristics of surrounding pixels, isolated abnormal pixels are removed, and the corrected set of suspected candidate pixels for algal blooms is output.
[0130] The corrected set of suspected algal bloom candidate pixels is compared with the pixel feature data of historical algal blooms to verify the matching degree. Pixels with a matching degree higher than the preset threshold are retained, and finally, the set of suspected algal bloom pixels labeled with the probability of algal bloom occurrence is output.
[0131] S53: Using the matching results of the suspected water bloom pixel set and the index, combined with the high-impact factor screening list and lag effect parameters, analyze the driving mechanism of factor lag effect on water bloom occurrence, and output the optimized set of suspected water bloom pixels driven by factors.
[0132] By retrieving the suspected algal bloom pixel set and index matching results, and associating the high-impact factor screening list and lag effect parameters, the factor type, lag time window, and index matching degree corresponding to each pixel are extracted, outputting a pixel analysis dataset with factor lag attributes. Based on the pixel analysis dataset with factor lag attributes, partial correlation analysis is used to quantify the contribution of each high-impact factor to the probability of algal bloom occurrence under different lag times, clarifying the driving strength and time-dependent characteristics of factor lag effects, and outputting a factor driving strength ranking table. Based on the factor driving strength ranking table, a factor-index coupled driving judgment model is constructed, and a contribution threshold is set to screen strong driving factors. Pixels in the original suspected pixel set are then re-judged, outputting a preliminarily optimized suspected algal bloom pixel set. Using the preliminarily optimized suspected algal bloom pixel set, cross-validation is performed by combining the factor lag patterns of historical algal bloom events, eliminating misjudged pixels with low driving factor matching degrees, and supplementing missed detection pixels corresponding to high contribution factors. Finally, an optimized set of suspected algal bloom pixels driven by factors is output.
[0133] S54: Based on the factor-driven optimization set of suspected algal bloom pixels, combined with the wind-hydrodynamic coupling diffusion rule, simulate the diffusion trend of algal bloom under different wind and hydrodynamic conditions, correct the attribution of edge pixels, and output the corrected algal bloom pixel distribution dataset.
[0134] By retrieving a factor-driven optimized set of suspected algal bloom pixels, the spatial coordinates, algal bloom probability, and neighborhood topological relationships of the pixels are extracted. Combined with the core parameters of wind speed, flow direction, and water velocity from the wind-hydrodynamic coupled diffusion rule, a pixel analysis dataset with diffusion constraints is output. Based on this dataset, a numerical simulation model for algal bloom diffusion is constructed, using wind shear stress and the advection diffusion effect of hydrodynamics as driving terms to simulate the diffusion direction and range of pixels under different spatiotemporal scenarios, outputting a set of algal bloom diffusion trend simulation results. Based on this simulation result set, to address the ambiguity in the attribution of edge pixels, a diffusion probability weighting algorithm for neighborhood pixels is introduced to calculate the algal bloom attribution confidence of edge pixels, correcting misclassified pixels in the original determination results, and outputting an intermediate dataset with edge correction. Using this intermediate dataset, spatial topological consistency is verified, isolated pixels appearing in the diffusion simulation are removed, and missing pixel interpolation results on the diffusion path are supplemented, outputting a corrected algal bloom pixel distribution dataset.
[0135] S55: Using the corrected algal bloom pixel distribution dataset, perform spatial consistency verification and accuracy verification, remove misjudged pixels and supplement the interpolation results of missed pixels, and finally output high-precision initial algal bloom pixel distribution results.
[0136] By retrieving the corrected algal bloom pixel distribution dataset, the spatial coordinates, algal bloom judgment labels, and neighboring pixel association information of each pixel are extracted. Spatial continuity verification rules are set, and a dataset with neighborhood association features is output for verification. Based on the dataset with neighborhood association features, spatial consistency verification is performed to identify and remove isolated misjudged pixels without neighborhood support, and regions of missed pixels due to diffusion simulation bias are marked, outputting a dataset to be completed after removing misjudgments. Based on the dataset to be completed after removing misjudgments, an inverse distance weighted interpolation algorithm is introduced to interpolate using the algal bloom feature values of effective pixels surrounding the missed areas, supplementing the judgment results of missing pixels, and outputting a complete dataset after interpolation. Using the complete dataset after interpolation, accuracy verification is carried out in conjunction with measured algal bloom monitoring point data. Verification indicators such as confusion matrix and Kappa coefficient are calculated, pixel data that meets the accuracy threshold is selected, and finally, high-precision initial algal bloom pixel distribution results are output.
[0137] S6: Using the high-precision initial algal bloom pixel distribution as the diffusion starting point, a neighborhood iterative diffusion model is constructed by combining the spatiotemporal dynamic pixel-level suitability matrix. The algal bloom growth and diffusion are coupled through the proliferation coefficient and migration probability. After in-situ sample calibration, remote sensing data correction and Bayesian optimization of parameters, the pixel-level algal bloom prediction map is output through iterative simulation.
[0138] S61: Extract the initial algal bloom pixel distribution results and the spatial coordinates, fitness values and neighborhood topological relationships of the pixels in the spatiotemporal dynamic pixel-level fitness matrix, and output the pixel modeling dataset with diffusion basic parameters.
[0139] The initial algal bloom pixel distribution results and the spatiotemporal dynamic pixel-level fitness matrix are extracted to obtain the pixel spatial coordinate information, establishing a one-to-one coordinate mapping relationship and outputting a basic pixel dataset with coordinate association labels. The fitness metric value corresponding to each matching pixel is extracted, and the 8-neighborhood topology of the pixel is simultaneously identified and the spatial positional relationships of neighboring pixels are labeled, outputting an intermediate dataset with attribute associations. Based on the basic parameter requirements for algal bloom diffusion, three core types of information—pixel coordinates, fitness values, and neighborhood topology relationships—are integrated. Redundant pixels with coordinate matching failures are removed, and a pixel modeling dataset with diffusion basic parameters is output.
[0140] S62: Based on the pixel modeling dataset with diffusion basic parameters, construct a neighborhood pixel iterative diffusion model, introduce the proliferation coefficient and neighborhood spatial migration probability associated with suitability, realize the dynamic coupling of algal bloom growth and diffusion, and output the initial simulation result set of the model.
[0141] By retrieving a pixel modeling dataset with diffusion parameters, the fitness value, neighborhood topology, and diffusion parameters of pixels are extracted. A quantitative correlation formula between fitness and proliferation coefficient is established, and a model input dataset with proliferation parameters is output. Based on the model input dataset with proliferation parameters, the core framework of the neighborhood pixel iterative diffusion model is constructed. The migration probability calculation rules for neighborhood pixels are defined, and the proliferation coefficient and migration probability are bound together to form a coupled driving logic, outputting an initialized iterative diffusion model. Based on the initialized iterative diffusion model, starting with the initial algal bloom pixel, pixel fitness and neighborhood topology data are substituted to perform the first round of iterative calculation, generating simulation results of algal bloom growth and diffusion, and outputting the model's first round of iteration result set. Using the model's first round of iteration result set, multiple time-period iteration step sizes are set, and the loop logic of proliferation calculation → migration determination → pixel update is repeatedly executed to simulate the spatiotemporal evolution of algal blooms at different time periods, outputting the model's initial simulation result set.
[0142] S63: Based on the initial simulation result set of the model, select a small number of in-situ observation samples for pixel-level matching, calibrate the accuracy of pixel biomass, synchronously call multiple periods of remote sensing time series verification data, and output the simulation intermediate dataset after accuracy calibration.
[0143] The refined wind field grid data, meteorological station interpolated illumination duration data, and algal bloom coverage time series data associated with the zoning factor suitability parameter table in the standardized pixel-level dataset are used to output a temporally consistent and spatially matched pixel-level meteorological-algal bloom fusion dataset through timestamp alignment and spatial grid mapping.
[0144] Based on the remote sensing time-series data of the aforementioned pixel-level meteorological-algal bloom fusion dataset, the cyanobacterial proliferation rate is calculated using an inversion algorithm. The inversion results are then input into a dynamic weight model of factor contribution, and the weight ratios of water temperature and nitrogen-phosphorus ratio are adjusted in real time to output a cyanobacterial proliferation-environmental factor association dataset with dynamic weights.
[0145] Based on the dynamically weighted cyanobacteria proliferation-environmental factor association dataset, and combined with the regional threshold of the zoning factor suitability parameter table, the comprehensive suitability of each pixel is obtained through weighted calculation, and finally the spatiotemporal dynamic pixel-level suitability matrix is output.
[0146] S64: Using the intermediate dataset of the simulation after accuracy calibration, the threshold of the spatiotemporal dynamic cell-level suitability matrix and the parameters of the iterative diffusion model are corrected in reverse. The simulation effect is evaluated by the confusion matrix, and the simulation result set after parameter optimization is output.
[0147] By integrating the intermediate simulation dataset after accuracy calibration with the existing spatiotemporal dynamic pixel-level fitness matrix, the fitness threshold of the matrix is corrected in reverse with reference to the calibration data. At the same time, the core parameters of the iterative diffusion model are adjusted, and the model input dataset after the threshold and parameters are co-corrected is output.
[0148] Based on the model input dataset after threshold and parameter co-correction, it is substituted into the iterative diffusion model for simulation calculation. The degree of fit between the simulation results and the actual monitoring data is compared by quantifying through the confusion matrix. The optimal combination of simulation parameters is evaluated and selected, and a simulation effect evaluation report is output.
[0149] Based on the optimal parameter combination confirmed in the simulation effect evaluation report, the iterative diffusion model is re-driven to execute the simulation process. The entire process is completed by combining the corrected spatiotemporal dynamic cell-level fitness matrix, and the final output is the simulation result set with optimized parameters.
[0150] S65: Based on the simulation results set after parameter optimization, the core parameters of the model are adjusted using the Bayesian optimization algorithm, and the evolution process of algal blooms at different time periods is simulated iteratively to finally output a pixel-level algal bloom prediction map that combines accuracy and practicality.
[0151] Based on the simulation results set after parameter optimization, key feature parameters of algal bloom evolution are extracted as prior information and input into the Bayesian optimization algorithm. The core parameters of the model are adjusted with simulation accuracy as the optimization objective, and the model configuration scheme after parameter iteration optimization is output.
[0152] Based on the model configuration scheme after parameter iteration and optimization, environmental driving data from different time periods are substituted into the model, and multiple rounds of iterative simulation calculations are performed to reconstruct the occurrence, development and diffusion process of algal blooms under different time dimensions, and output a multi-time period algal bloom evolution simulation dataset.
[0153] Based on the multi-period algal bloom evolution simulation dataset, and combined with actual monitoring data, the accuracy and practicality of the simulation results are verified. The simulation results of each period are integrated through pixel-level spatial mapping, and finally, a pixel-level algal bloom prediction map with both accuracy and practicality is output.
[0154] In summary, this embodiment provides a pixel-level prediction method for lake cyanobacterial blooms based on multi-source data fusion. Through in-situ sample calibration, remote sensing data correction, and confusion matrix evaluation, the method continuously optimizes intermediate simulation results and reversely corrects the suitability matrix threshold and iterative diffusion model parameters, effectively reducing the model's systematic error. A Bayesian optimization algorithm is introduced to iteratively adjust core parameters with simulation accuracy as the target, generating a multi-time-period bloom evolution simulation dataset. The accuracy and practicality of the results are verified using actual monitoring data. This progressive optimization strategy ensures the model's stability under different spatiotemporal scenarios, and the output pixel-level bloom prediction map accurately reflects the spatiotemporal distribution characteristics of blooms, providing an intuitive and reliable decision-making basis for lake bloom prevention and ecological governance.
[0155] Initial pixels serve as the anchor points for subsequent spatiotemporal evolution simulations of algal blooms. Their accuracy directly determines the reliability of the final pixel-level prediction results. The spread of cyanobacterial blooms gradually originates from localized dominant areas. Blindly predicting the entire lake's pixels directly ignores the ecological pattern of bloom accumulation followed by diffusion. Initial pixels clearly identify the core source of the bloom, providing a realistic spatial starting point for subsequent neighborhood-based iterative diffusion simulations. High-impact factors, lag effect parameters, and zoning suitability thresholds, selected through screening, require initial pixels for accurate implementation. Only by identifying initial pixels can we specifically analyze the factor contribution of that region, avoiding the substitution of local core driving factors with lake-wide average parameters, and ensuring the diffusion rules are more realistic. Without extracting initial pixels, directly performing time-series simulations and diffusion calculations on all lake pixels would result in a large amount of invalid computation. Initial pixels can narrow the simulation scope, focusing on potential bloom areas, significantly improving computational efficiency while maintaining prediction accuracy, and meeting the needs of pixel-level high-resolution prediction.
[0156] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion, characterized in that, Comprise: S1: Collecting satellite images in the lake pixel level multi-source basic data and pretreatment, output standardization pixel level data set; S2: From the standardization pixel level data set screening core influencing factor, calculating algae index generation binary distribution product, combined with blue algae growth mechanism partition fitting to determine the dynamic threshold, output partition factor suitability parameter table; S3: According to the standardization pixel level data set in the wind field, illumination data, with the partition factor suitability parameter table do space-time matching, inversion blue algae proliferation rate and adjust factor weight, combined with regional threshold calculation pixel comprehensive suitability, output space-time dynamic pixel level suitability matrix; S4: According to the space-time dynamic pixel level suitability matrix time sequence characteristics, combined with the matching meteorological and algal bloom data, through correlation analysis and cause and effect test to identify lag effect factor, through feature sorting filter high influence factor, output high influence factor list; S5: According to the space-time dynamic pixel level suitability matrix quantization and the high influence factor list construct diffusion rule, extract initial algal bloom pixel and determine the diffusion starting point, output initial algal bloom pixel distribution result; S6: With high precision initial algal bloom pixel distribution as the diffusion starting point, using the space-time dynamic pixel level suitability matrix constructs neighborhood iteration diffusion model, through proliferation coefficient and migration probability realizes the growth and diffusion coupling of algal bloom, iteration simulation output pixel level algal bloom prediction map.
2. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S2, the specific steps of outputting the partition factor suitability parameter table are: S21: Extracting core influencing factor from the standardization pixel level data set, using the NDVI and FAI index calculated by the preprocessed satellite image in the standardization pixel level data set; S22: According to the NDVI and FAI index, using threshold segmentation to generate pixel level binary distribution product and coverage area time series data; S23: Combined with the spatio-temporal heterogeneity physiological mechanism of blue algae growth, according to the coverage area time series data auxiliary partition, the lake is divided into nearshore nutrient enrichment area, lake center hydrodynamic sensitive area, transition area by partition nonlinear fitting algorithm, get partition result; S24: According to the partition result pixel level binary distribution product and coverage area time series data, analyze the correlation between algal bloom and each factor in different regions, determine the dynamic suitability threshold of each regional core influencing factor, output partition factor suitability parameter table.
3. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 2, characterized in that: In step S23, the specific steps of obtaining the partition result are: S231: According to the core characteristics of the spatio-temporal heterogeneity physiological mechanism of blue algae growth, determine the physiological response difference of nearshore nutrient enrichment area, lake center hydrodynamic sensitive area, transition area, form partition theoretical basis; S232: According to the partition theoretical basis, extract the spatial distribution characteristics, expansion and contraction trend and peak occurrence law of the blue algae coverage area time series data, identify the high and low frequency areas of blue algae aggregation; S233: According to the high and low frequency areas, using partition nonlinear fitting algorithm to construct fitting target model with the spatial correlation and temporal fluctuation of coverage area time series data as fitting variables; S234: Input the coverage area time series data of lake global pixel into the fitting model, divide the initial partition by iteration calculation; S235: By comparing the time series variation of cyanobacteria coverage area in different subareas with the corresponding adjustment fitting parameters of nutrient salt and water dynamics, the spatial subarea results of nearshore nutrient enrichment area, lake center water dynamics sensitive area and transition area are output.
4. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S3, the specific steps of outputting the spatiotemporal dynamic pixel-level suitability matrix are: S31: Temporally and spatially aligning and preprocessing the wind field, illumination spatiotemporal data and subarea factor suitability parameters in the standardized pixel-level data set, unifying the time granularity and spatial projection, and outputting the parameter table associated data set; S32: According to the associated data set, performing factor dimension matching and attribute association, extracting the corresponding suitability reference parameters of each spatiotemporal unit, and outputting the spatiotemporal matching data set; S33: Substituting the spatiotemporal matching data set into the cyanobacteria proliferation rate inversion model to calculate the initial proliferation rate, iteratively adjusting the subarea factor weight through error analysis, and outputting the factor weight set; S34: According to the factor weight set and the pre-set factor threshold of each region, a pixel-level comprehensive suitability calculation model is constructed, the wind field, illumination suitability contribution values of each pixel are weighted and summed, and an initial pixel-level comprehensive suitability data set is output; S35: The initial pixel-level comprehensive suitability data set is smoothed and corrected, abnormal values are removed, and edge pixel interpolation results are supplemented, and a pixel-level suitability matrix is output.
5. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 4, characterized in that: In step S33, the specific steps of outputting the factor weight set are: S331: Extracting the wind field friction wind speed, illumination effective radiation intensity and corresponding subarea factor suitability reference parameters of each spatiotemporal unit in the spatiotemporal matching data set, and outputting the model input data set; S332: Substituting the model input data set into the cyanobacteria proliferation rate inversion model based on energy conservation to calculate the initial proliferation rate of each spatiotemporal unit, and outputting the initial proliferation rate data set; S333: According to the initial proliferation rate data set and the measured cyanobacteria biomass growth rate, an error objective function is constructed by calculating the model error through residual sum of squares; S334: The wind field factor weight and the illumination factor weight are adjusted using the error objective function until the error is less than the threshold, and the factor weight set is output.
6. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S4, the specific steps of outputting the high impact factor list are: S41: Aligning and calibrating the time series characteristic data of the spatiotemporal dynamic pixel-level suitability matrix, meteorological data and water bloom monitoring data, and outputting a multi-source fusion data set; S42: According to the multi-source fusion data set, the suitability time series variation and the meteorological and water bloom index fluctuation are analyzed, the Pearson correlation coefficient is calculated, the potential associated factors are preliminarily screened, and a candidate list of associated factors is output; S43: According to the candidate list of associated factors, the Granger causality test method is used to verify the causal relationship between the candidate factors and the water bloom, identify the key driving factors with lag effect, and output a causal factor set; S44: According to the causal factor set, a feature importance sorting model is constructed, and the random forest algorithm is used to quantify the contribution degree of each factor to the water bloom, and a factor priority sequence is output; S45: Extracting the high impact factors with a contribution degree higher than a pre-set threshold from the factor priority sequence, and calculating the optimal lag time parameters of each factor, and outputting a high impact factor list.
7. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S5, the specific steps of outputting the initial water bloom pixel distribution result are: S51: Extracting the water bloom occurrence basic condition quantitative index corresponding to each pixel in the spatiotemporal dynamic pixel level suitability matrix, performing data standardization and value range normalization processing, and outputting a pixel level quantitative data set; S52: According to the pixel level quantitative data set, introducing an improved phytoplankton index and phycocyanin index collaborative recognition algorithm to preliminarily judge the water bloom occurrence probability of the pixel, and outputting a water bloom suspected pixel set; S53: According to the water bloom suspected pixel set and the high influence factor screening list, analyzing the driving mechanism of the water bloom occurrence, and outputting a water bloom suspected pixel optimization set; S54: According to the water bloom suspected pixel optimization set and the wind field-water dynamic coupling diffusion rule, simulating the diffusion trend of the water bloom under different wind fields and water dynamic conditions, correcting the attribution determination of the edge pixel, and outputting a water bloom pixel distribution data set; S55: Performing spatial consistency checking and precision verification on the water bloom pixel distribution data set, removing misjudged pixels and supplementing the interpolation results of missed pixels, and outputting the initial water bloom pixel distribution result.
8. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 7, characterized in that: In step S52, the specific steps of outputting the water bloom suspected pixel set are: By calling the pixel level quantitative data set, the water bloom occurrence basic condition index corresponding to each pixel is extracted, the index calculation input data set is outputted by combining the calculation threshold of PAI and PCI; The index calculation input data set is substituted into the collaborative calculation model of the improved PAI and PCI to calculate the PAI value and PCI value of each pixel respectively, the double index intersection judgment rule is used to screen the pixels that preliminarily meet the conditions, and the double index preliminary screening pixel set is outputted; According to the double index preliminary screening pixel set, the pixel neighborhood correlation analysis is introduced, the PAI and PCI distribution characteristics of the surrounding pixels are combined to correct the single pixel determination result, the isolated abnormal pixels are removed, and the corrected water bloom suspected candidate pixel set is outputted; The corrected water bloom suspected candidate pixel set is compared with the pixel feature data of the historical water bloom to perform matching degree verification, the pixels with a matching degree higher than a preset threshold are retained, and finally the water bloom suspected pixel set with a water bloom occurrence probability is outputted.
9. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S6, the specific steps of outputting the pixel level water bloom prediction map are: S61: Extracting the spatial coordinates, suitability values and neighborhood topological relations of the initial water bloom pixel distribution result and the spatiotemporal dynamic pixel level suitability matrix, and outputting a pixel modeling data set; S62: According to the pixel modeling data set, a neighborhood pixel iterative diffusion model is constructed, the proliferation coefficient is associated with the neighborhood space migration probability, and an initial simulation result set of the model is outputted; S63: Based on the initial simulation result set of the model, a small amount of in-situ observation samples are selected for pixel level matching to calibrate the accuracy of the pixel biomass, and a multi-period remote sensing time series verification data is called to output a simulation intermediate data set; S64: According to the simulation intermediate data set, the threshold value of the spatiotemporal dynamic pixel level suitability matrix and the iterative diffusion model parameters are reversely corrected, the simulation effect is evaluated through a confusion matrix, and a simulation result set is outputted; S65: According to the simulation result set, a Bayesian optimization algorithm is used to adjust the core parameters of the model, the water bloom evolution process in different time periods is iteratively simulated, and a pixel level water bloom prediction map is outputted.
10. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 9, characterized in that: In step S62, the specific steps of outputting the initial simulation result set of the model are: Extracting the suitability value, neighborhood topological relationship and diffusion basic parameters of the pixel in the pixel modeling data set, establishing a quantitative correlation formula between the suitability and the multiplication coefficient, and outputting the model input data set; According to the model input data set, constructing the core framework of the neighborhood pixel iterative diffusion model, defining the migration probability calculation rule of the neighborhood pixel, binding the multiplication coefficient and the migration probability to form the coupling driving logic, and outputting the initial iterative diffusion model; Taking the initial algal bloom pixel as the starting point, substituting the pixel suitability and neighborhood topological data into the initial iterative diffusion model, performing the first round of iteration calculation, generating the simulation result of the growth and diffusion of the algal bloom, and outputting the first round of iteration result set of the model; Using the first round of iteration result set of the model, setting the iteration step of multiple time periods, repeating the execution, simulating the spatio-temporal evolution process of the algal bloom at different time periods, and outputting the initial simulation result set of the model.
Citation Information
Patent Citations
Cyanobacterial bloom remote sensing monitoring method based on planktonic algae indexes and deep learning
CN110414488A
Lake cyanobacterial bloom area extraction method
CN116342683A
Lake blue-green algae horizontal and vertical movement rate calculation method based on stationary satellite
CN116542947A
Improved cellular automaton cyanobacterial bloom prediction method
CN117408385A
Blue-green algae bloom weather forecasting and early warning method and system based on hierarchical fusion algorithm
CN118942228A
Cited By
Farmland water-salt evolution rule mining system based on time sequence trajectory clustering
CN122020599A
Agricultural field water-salt evolution rule mining system based on time sequence trajectory clustering
CN122020599B