A lake blue-green algae bloom pixel-level prediction method based on multi-source data fusion
By fusing multi-source data and partitioned nonlinear fitting, a spatiotemporal dynamic fitness matrix is constructed, high-influence factors are identified, and iterative simulation of migration-proliferation processes is coupled. This solves the problems of low simulation accuracy and poor applicability in the prediction of cyanobacterial blooms in lakes, and achieves accurate prediction of blooms at the pixel level.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies have low simulation accuracy and poor applicability in predicting cyanobacterial blooms in lakes, making it difficult to cope with the dynamic changes in complex lake ecosystems. Traditional meteorological observation station data has low spatial resolution, fails to effectively combine multi-source data for accurate spatiotemporal matching, and ignores the spatiotemporal heterogeneous physiological mechanisms of cyanobacterial growth.
By collecting multi-source basic data of lake pixels in satellite imagery, preprocessing and screening core influencing factors, combining the cyanobacterial growth mechanism with zonal fitting to determine dynamic thresholds, constructing a spatiotemporal dynamic pixel-level suitability matrix, identifying high-influence factors, constructing a neighborhood iterative diffusion model, realizing the coupling of algal bloom growth and diffusion, and iteratively simulating and outputting a pixel-level algal bloom prediction map.
It achieves high precision and high resolution in pixel-level algal bloom prediction, improves the spatiotemporal resolution and accuracy of prediction, provides refined technical support for the prevention and control of algal blooms in lakes, eliminates noise interference from remote sensing images, accurately matches the environmental response patterns of cyanobacteria growth, identifies the lag effects of high-impact meteorological factors, and dynamically couples migration and proliferation processes.
Smart Images

Figure CN121543841B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of water algae prediction, and particularly relates to a lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion. BACKGROUND
[0002] At present, satellites, unmanned aerial vehicles, real-time monitoring cameras, water quality instruments and Internet of Things technologies have realized dynamic monitoring of algal blooms, but prediction and early warning are crucial for targeted prevention and control, can provide basis for pre-emptive measures and emergency response, and avoid ecological disasters. However, the high dynamics and complex outbreak mechanism of algal blooms make the development of prediction models a challenge. The data obtained by traditional meteorological observation stations usually need to be interpolated to generate the spatial distribution of meteorological factors, and the image spatial resolution obtained by this method is usually low, which is not enough to support the construction of pixel-level prediction models. The unified threshold is used to divide the region, the spatial and temporal heterogeneity of the physiological mechanism of blue algae growth is ignored, the multi-source data is not combined with the spatial and temporal precise matching, the simulation accuracy is low, the applicability is poor, and it is difficult to cope with the dynamic changes of complex lake ecological systems. SUMMARY
[0003] In order to make up for the shortcomings of the prior art, the application provides a lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion. The application is mainly used to solve the problems of low simulation accuracy and poor applicability.
[0004] The application provides a lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion, which comprises the following steps:
[0005] S1: Collecting lake pixel-level multi-source basic data in satellite images and pre-processing, and outputting standardized pixel-level data set.
[0006] S2: Screening core influencing factors from the standardized pixel-level data set, calculating algal index to generate binary distribution product, combining blue algae growth mechanism to divide and fit to determine dynamic threshold, and outputting partition factor suitability parameter table.
[0007] S3: According to the wind field and light data in the standardized pixel-level data set, and the space-time matching of the partition factor suitability parameter table, the blue algae proliferation rate is inversed and the factor weight is adjusted, the pixel comprehensive suitability is calculated combined with the regional threshold, and the space-time dynamic pixel-level suitability matrix is outputted.
[0008] S4: According to the time sequence characteristics of the space-time dynamic pixel-level suitability matrix, combining the matched meteorological and algal bloom data, identifying the lag effect factor through correlation analysis and causality test, screening high influence factors through feature sorting, and outputting the high influence factor list.
[0009] S5: According to the spatio-temporal dynamic pixel level suitability matrix quantization and the high impact factor list, a diffusion rule is constructed, initial algal bloom pixels are extracted, and an initial algal bloom pixel distribution result is output.
[0010] S6: Taking the high-precision initial algal bloom pixel distribution as a diffusion starting point, a neighborhood iterative diffusion model is constructed using the spatio-temporal dynamic pixel level suitability matrix, the growth and diffusion of the algal bloom are coupled through the proliferation coefficient and the migration probability, and an iterative simulation output pixel level algal bloom prediction map is output.
[0011] According to the lake cyanobacterial bloom pixel level prediction method based on multi-source data fusion provided by the application, in step S2, the specific steps of outputting the partition factor suitability parameter table are:
[0012] S21: Core impact factors are extracted from the standardized pixel level data set, and NDVI and FAI indexes are calculated using the preprocessed satellite image in the standardized pixel level data set.
[0013] S22: According to the NDVI and FAI indexes, a pixel level binary distribution product and coverage area time series data are generated using threshold segmentation.
[0014] S23: Combined with the spatio-temporal heterogeneity physiological mechanism of blue-green algae growth, according to the auxiliary partition of the coverage area time series data, the lake is divided into a nearshore nutrient enrichment area, a lake center hydrodynamic sensitive area and a transition area by a partition nonlinear fitting algorithm to obtain a partition result.
[0015] S24: According to the partition result pixel level binary distribution product and coverage area time series data, the correlation between the algal bloom and each factor in different regions is analyzed, the dynamic suitability threshold of the core impact factor in each region is determined, and a partition factor suitability parameter table is output.
[0016] According to the lake cyanobacterial bloom pixel level prediction method based on multi-source data fusion provided by the application, in step S23, the specific steps of obtaining the partition result are:
[0017] S231: According to the core characteristics of the spatio-temporal heterogeneity physiological mechanism of blue-green algae growth, the physiological response difference of the nearshore nutrient enrichment area, the lake center hydrodynamic sensitive area and the transition area is determined, and a partition theoretical basis is formed.
[0018] S232: According to the spatial distribution characteristics, expansion and contraction trend and peak occurrence law of the blue-green algae coverage area time series data extracted according to the partition theoretical basis, high and low frequency areas of blue-green algae aggregation are identified.
[0019] S233: According to the high and low frequency areas, using the partition nonlinear fitting algorithm, taking the spatial correlation and temporal fluctuation of the coverage area time series data as fitting variables, a fitting target model is constructed.
[0020] S234: input the coverage area time series data of the lake global pixel into the fitting model, and divide the initial partition through iterative calculation.
[0021] S235: by comparing the time series variation law of blue-green algae coverage area in different partitions with the corresponding region of nutrient salt and water dynamic adjustment fitting parameter, output the spatial partition result of nearshore nutrient enrichment area, lake center water dynamic sensitive area and transition area.
[0022] According to the lake blue-green algae bloom pixel level prediction method based on multi-source data fusion provided by the application, in step S3, the specific steps of outputting the spatiotemporal dynamic pixel level suitability matrix are:
[0023] S31: the wind field, illumination spatiotemporal data and partition factor suitability parameter in the standardized pixel level data set are subjected to spatiotemporal alignment preprocessing, the time granularity and space projection are unified, and the parameter table associated data set is output.
[0024] S32: according to the associated data set, the factor dimension matching and attribute association are carried out, the corresponding suitability reference parameter of each spatiotemporal unit is extracted, and the spatiotemporal matching data set is output.
[0025] S33: the spatiotemporal matching data set is substituted into the blue-green algae proliferation rate inversion model to calculate the initial proliferation rate, the partition factor weight is iteratively adjusted through error analysis, and the factor weight set is output.
[0026] S34: according to the factor weight set and the preset factor threshold of each region, a pixel level comprehensive suitability calculation model is constructed, the wind field, illumination suitability contribution value of each pixel is weighted and summed, and an initial pixel comprehensive suitability data set is output.
[0027] S35: the initial pixel comprehensive suitability data set is subjected to smoothing correction, the abnormal value is removed and the edge pixel interpolation result is supplemented, and the pixel level suitability matrix is output.
[0028] According to the lake blue-green algae bloom pixel level prediction method based on multi-source data fusion provided by the application, in step S33, the specific steps of outputting the factor weight set are:
[0029] S331: the wind field friction wind speed, illumination effective radiation intensity and corresponding partition factor suitability reference parameter of each spatiotemporal unit in the spatiotemporal matching data set are extracted, and the model input data set is output.
[0030] S332: the model input data set is substituted into the blue-green algae proliferation rate inversion model based on energy conservation, the initial proliferation rate of each spatiotemporal unit is calculated, and the initial proliferation rate data set is output.
[0031] S333: according to the initial proliferation rate data set and the measured blue-green algae biomass growth rate, the model error is calculated through residual sum of squares, and the error objective function is constructed.
[0032] S334: Adjusting the wind field factor weight and the light factor weight using the error objective function until the error is less than a threshold value, and outputting a factor weight set.
[0033] According to the lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion provided by the application, in step S4, the specific steps of outputting the high impact factor list are:
[0034] S41: Aligning and calibrating the time series feature data of the spatiotemporal dynamic pixel-level suitability matrix, the meteorological data and the water bloom monitoring data, and outputting a multi-source fusion data set.
[0035] S42: According to the multi-source fusion data set, analyzing the suitability time series variation and the meteorological and water bloom index fluctuation, calculating the Pearson correlation coefficient, preliminarily screening the potential correlation factors, and outputting a correlation factor candidate list.
[0036] S43: According to the correlation factor candidate list, using the Granger causality test method to verify the causality between the candidate factors and the water bloom, identifying the key driving factors with lag effects, and outputting a causality factor set.
[0037] S44: According to the causality factor set, constructing a feature importance ranking model, using a random forest algorithm to quantify the contribution of each factor to the water bloom, and outputting a factor priority sequence.
[0038] S45: Extracting high impact factors with a contribution degree higher than a preset threshold from the factor priority sequence, calculating the optimal lag time parameter of each factor, and outputting a high impact factor list.
[0039] According to the lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion provided by the application, in step S5, the specific steps of outputting the initial water bloom pixel distribution result are:
[0040] S51: Extracting the water bloom basic condition quantitative index corresponding to each pixel in the spatiotemporal dynamic pixel-level suitability matrix, performing data standardization and value range normalization processing, and outputting a pixel-level quantitative data set.
[0041] S52: According to the pixel-level quantitative data set, introducing an improved phytoplankton index and algal blue protein index collaborative recognition algorithm to preliminarily judge the water bloom occurrence probability of the pixel, and outputting a water bloom suspected pixel set.
[0042] S53: According to the water bloom suspected pixel set and the high impact factor screening list, analyzing the driving mechanism of the factor lag effect on the water bloom occurrence, and outputting a water bloom suspected pixel optimization set.
[0043] S54: According to the algal bloom suspected pixel optimization set and the wind field-water dynamic coupling diffusion rule, the diffusion trend of the algal bloom under different wind fields and water dynamic conditions is simulated, the attribution determination of the edge pixel is corrected, and an algal bloom pixel distribution data set is output.
[0044] S55: The spatial consistency of the algal bloom pixel distribution data set is checked and the precision is verified, the misjudged pixels are removed, the interpolation results of the missed pixels are supplemented, and an initial algal bloom pixel distribution result is output.
[0045] According to the algal bloom pixel-level prediction method based on multi-source data fusion provided by the application, the specific steps of outputting the algal bloom suspected pixel set in step S52 are:
[0046] By calling the pixel-level quantitative data set, the algal bloom occurrence basic condition index corresponding to each pixel is extracted, the index calculation input data set is output by combining the calculation threshold of PAI and PCI.
[0047] The index calculation input data set is substituted into the improved PAI and PCI cooperative calculation model, and the PAI value and PCI value of each pixel are calculated, respectively. The preliminary qualified pixels are screened through the double-index intersection determination rule, and the double-index preliminary screening pixel set is output.
[0048] According to the double-index preliminary screening pixel set, the pixel neighborhood correlation analysis is introduced, the PAI and PCI distribution characteristics of the surrounding pixels are combined to correct the single-pixel determination result, the isolated abnormal pixels are removed, and the corrected algal bloom suspected candidate pixel set is output.
[0049] The corrected algal bloom suspected candidate pixel set is compared with the pixel feature data of the historical algal bloom occurrence, the matching degree is verified, the pixels with a matching degree higher than a preset threshold are retained, and finally the algal bloom suspected pixel set with labeled algal bloom occurrence probability is output.
[0050] According to the algal bloom pixel-level prediction method based on multi-source data fusion provided by the application, the specific steps of outputting the pixel-level algal bloom prediction map in step S6 are:
[0051] S61: Extract the spatial coordinates, suitability value and neighborhood topological relationship of the initial algal bloom pixel distribution result and the spatiotemporal dynamic pixel-level suitability matrix pixel, and output the pixel modeling data set.
[0052] S62: According to the pixel modeling data set, a neighborhood pixel iterative diffusion model is constructed, the proliferation coefficient is associated with the neighborhood space migration probability, and an initial simulation result set of the model is output.
[0053] S63: Based on the initial simulation result set of the model, a small amount of in-situ observation samples are selected for pixel-level matching, the pixel biomass precision is calibrated, multi-period remote sensing time series verification data is called, and an intermediate simulation data set is output.
[0054] S64: According to the simulation intermediate data set, the spatio-temporal dynamic pixel level suitability matrix threshold and the iterative diffusion model parameter are corrected reversely, the simulation effect is evaluated through the confusion matrix, and a simulation result set is output.
[0055] S65: According to the simulation result set, a Bayesian optimization algorithm is used to adjust the core parameters of the model, and the evolution process of the water bloom in different time periods is iteratively simulated, and a pixel level water bloom prediction map is output.
[0056] According to the lake blue-green algae water bloom pixel level prediction method based on multi-source data fusion provided by the application, in step S62, the specific steps of outputting the initial simulation result set of the model are:
[0057] The suitability value, neighborhood topological relationship and diffusion basic parameter of the pixel in the pixel modeling data set are extracted, a quantitative correlation formula of suitability and proliferation coefficient is established, and a model input data set is output.
[0058] According to the model input data set, the core framework of the neighborhood pixel iterative diffusion model is constructed, the migration probability calculation rule of the neighborhood pixel is defined, the proliferation coefficient and the migration probability are bound to form a coupled driving logic, and an initial iterative diffusion model is output.
[0059] Taking the initial water bloom pixel as the starting point, the pixel suitability and neighborhood topological data are substituted into the initial iterative diffusion model, the first round of iteration calculation is performed, the simulation result of water bloom growth and diffusion is generated, and the first round of iteration result set of the model is output.
[0060] Using the first round of iteration result set of the model, the multi-time period iteration step is set, and the execution is repeated, the spatio-temporal evolution process of the water bloom in different time periods is simulated, and the initial simulation result set of the model is output.
[0061] The lake blue-green algae water bloom pixel level prediction method provided by the application integrates multi-source pixel level data and standardizes the processing, combines the blue-green algae physiological mechanism partition fitting, constructs the spatio-temporal dynamic suitability matrix, selects the high impact lag factor, and iteratively simulates the migration-proliferation process coupling, realizes the accurate prediction of the pixel level water bloom, improves the spatio-temporal resolution and accuracy of the prediction, and provides fine technical support for the prevention and control of the lake water bloom.
[0062] The beneficial effects of the application are as follows:
[0063] 1.The application eliminates the noise and cloud cover interference of remote sensing images by unifying the spatial resolution and coordinate system of multi-source data, combining preprocessing methods such as radiation calibration and atmospheric correction, and constructing a high-quality standardized pixel-level data set, which lays a reliable data foundation for subsequent analysis. At the same time, the method breaks through the limitations of traditional unified threshold, combines the spatio-temporal heterogeneity of cyanobacterial growth, and uses a partitioned nonlinear fitting algorithm to divide the lake functional area, and determines the dynamic suitability threshold of each regional factor. This partitioned differential modeling method accurately matches the environmental response law of cyanobacterial growth in different regions, effectively reduces the prediction error caused by homogeneous treatment of regional characteristics, and significantly improves the accuracy of pixel-level prediction.
[0064] 2.The application identifies high-impact meteorological factors that have a significant lag effect on the occurrence of water bloom, and determines the time window of factor action, which makes up for the defect of traditional models that ignore factor lag. On this basis, the wind field-hydrodynamic coupling diffusion rule is constructed, which combines the influence of wind speed and direction on cyanobacterial migration with the corrected cyanobacterial proliferation rate based on the lag effect of high-impact factors, realizing the dynamic coupling simulation of cyanobacterial migration and proliferation. This coupling mechanism is more in line with the real process of water bloom evolution, making the model upgrade from a simple static statistical prediction to a dynamic process simulation, greatly strengthening the scientificity and rationality of the prediction results.
[0065] 3.The application reduces the systematic error of the model through multiple rounds of model calibration and parameter optimization process, through in-situ sample calibration, remote sensing data correction, and continuous optimization of simulation intermediate results combined with confusion matrix evaluation, and inversely corrects the suitability matrix threshold and iterative diffusion model parameters, effectively reducing the systematic error of the model. The Bayesian optimization algorithm is introduced to iteratively adjust the core parameters based on the simulation accuracy, generate a multi-period water bloom evolution simulation data set, and verify the accuracy and practicality of the results combined with actual monitoring data. This layer-by-layer optimization strategy ensures the stability of the model in different spatio-temporal scenarios, and the output pixel-level water bloom prediction map can accurately reflect the spatio-temporal distribution characteristics of water bloom, providing intuitive and reliable decision-making basis for lake water bloom prevention and ecological management.
[0066] The initial pixel is the anchor point of the subsequent spatio-temporal evolution simulation of the water bloom, and the accuracy thereof directly determines the reliability of the final pixel-level prediction result. The diffusion of the blue-green algae bloom is gradually spread from the local dominant area, and if the whole lake pixel is blindly predicted, the ecological law of the water bloom gathering first and then diffusing will be ignored. The initial pixel determines the core source of the water bloom outbreak, and the subsequent neighborhood iterative diffusion simulation has a real spatial starting point. Through the screening of the high impact factor, the lag effect parameter and the partition suitability threshold, the initial pixel is needed to accurately land. Only by locking the initial pixel, the factor contribution degree of the region can be analyzed, the local core driving factor is replaced by the average parameter of the whole lake, and the diffusion rule is more in line with the actual situation. If the initial pixel is not extracted, the time series simulation and diffusion calculation are directly performed on all pixels of the whole lake, a large amount of invalid operation will be generated. The initial pixel can reduce the simulation range, focus on the potential outbreak area of the water bloom, greatly improve the calculation efficiency while ensuring the prediction accuracy, and adapt to the demand of pixel-level high-resolution prediction. BRIEF DESCRIPTION OF DRAWINGS
[0067] The present application will be further described below according to the drawings.
[0068] Fig. 1 is a step diagram of a lake blue-green algae bloom pixel-level prediction method based on multi-source data fusion provided by an embodiment of the present application;
[0069] Fig. 2 is a flowchart of a lake blue-green algae bloom pixel-level prediction method based on multi-source data fusion provided by an embodiment of the present application;
[0070] Fig. 3 is a step diagram of outputting the initial water bloom pixel distribution result provided by an embodiment of the present application. DETAILED DESCRIPTION
[0071] In order to make the technical means, creative features, purposes and effects achieved by the present application easy to understand, the present application will be further described below according to the specific embodiments.
[0072] As shown in Figs. 1 to 3 , the present application provides a lake blue-green algae bloom pixel-level prediction method based on multi-source data fusion, and the system comprises:
[0073] S1: Collecting lake pixel-level multi-source basic data, including high spatio-temporal resolution remote sensing image, pixel-level distributed water temperature data, nitrogen-phosphorus ratio spatial interpolation data and refined meteorological grid data, and unifying the spatial resolution to 250m and the coordinate system to WGS-84UTM projection. Performing satellite image data preprocessing: eliminating image noise through radiation calibration, atmospheric correction and geometric precision correction, combining cloud pollution pixel repair and abnormal value spatio-temporal filtering algorithm, filling the pixel data of the cloud-shaded area of optical remote sensing, and outputting the standardized pixel-level data set.
[0074] Remote sensing images are downloaded from the satellite data sharing platform, low cloud cover, no extreme weather interference data, format conversion and unified coordinate system. Water temperature data is measured by hydrological station and water power model simulation, and is generated by downscaling interpolation. Nitrogen and phosphorus ratio data are obtained by grid sampling and spatial interpolation. The measured data of the surrounding meteorological station are simulated and down-scaled by the fine numerical weather prediction model of China Meteorological Administration. All data are unified in WGS-84 UTM coordinate system, spatial resolution, and time stamp alignment.
[0075] The unified standard optical and radar remote sensing images are the processing objects, and the noise is eliminated and the data gap is filled. Finally, the standardized pixel-level data set is output. The specific process is as follows: convert the image DN value to radiance value, eliminate sensor error. Eliminate atmospheric interference and obtain the true reflectivity of the ground. Correct geometric distortion to ensure that the pixel coordinates match the actual geographical location. Use the penetrating property of radar image to repair the pixel of optical image cloud cover area. Remove noise pixels to ensure data continuity. Match the pre-processed remote sensing image with other data at the pixel level to generate a standardized pixel-level data set.
[0076] S2: From the standardized pixel-level data set, select water temperature, nitrogen and phosphorus ratio, wind speed, and light duration as core influencing factors. Calculate NDVI (normalized vegetation index) and FAI (phytoplankton index) according to the pre-processed satellite image. Generate pixel-level binary distribution products and coverage area time series data through threshold segmentation and spatial clustering algorithm. According to the spatio-temporal heterogeneity of the physiological mechanism of blue-green algae growth, the lake is divided into nearshore nutrient enrichment area, lake center water power sensitive area, and transition area. According to the pixel-level binary distribution products and coverage area time series data, determine the dynamic suitability threshold of each factor in different regions, output the partition factor suitability parameter table, realize the accurate association of standardized data and regional differentiated ecological mechanism threshold, and solve the limitations of traditional unified threshold.
[0077] S21: Extract water temperature, nitrogen and phosphorus ratio, wind speed, and light duration from the standardized pixel-level data set, and simultaneously call the pre-processed satellite image in the data set to calculate the NDVI and FAI index.
[0078] The specific steps for calculating the NDVI and FAI index through the multi-spectral band reflectance data of the pre-processed satellite image are as follows:
[0079] Extract the pixel-level reflectance data of the near-infrared band (denoted as ) and the red light band (denoted as ) in the pre-processed satellite image, and ensure that the spatial positions of the two bands are completely matched without offset or misplacement.
[0080] The NDVI index calculation formula is:
[0081] Each pixel is operated one by one to obtain the NDVI value of the global pixel.
[0082] The pixel-level reflectivity data of the blue light band (denoted as pBlue) in the preprocessed satellite image is extracted, ensuring that the spatial resolutions of the blue light, red light, and near-infrared three bands are consistent and the positions correspond. The FAI index calculation formula is represented as:
[0083]
[0084] In the formula, , , are the center wavelengths of the blue light, red light, and near-infrared bands, respectively. According to the matched three-band pixel reflectivity data, the FAI value of each pixel is calculated.
[0085] S22: Generate pixel-level binary distribution products and coverage area time series data using threshold segmentation according to the NDVI and FAI indices.
[0086] Through the calculated pixel-level NDVI and FAI index data sets, combined with the preset threshold, the binary classification is completed, and the specific steps are as follows:
[0087] Determine the NDVI threshold interval [TNDVI low , TNDVI high ] and the FAI threshold interval [TFAI low , TFAI high ], where TNDVI low , TNDVI high are the lower and upper threshold values of NDVI classification, respectively, and TFAI low , TFAI high are the lower and upper threshold values of FAI classification, respectively.
[0088] Threshold judgment is performed on the NDVI and FAI indices of each pixel, and logical AND operation is used to realize binary classification. The binary formula is:
[0089] ;
[0090] In the formula, is the binary result of the pixel at position t moment, 1 represents the target land class, and 0 represents the non-target land class. The operation is performed pixel by pixel and time by time to generate a sequence of pixel-level binary distribution products.
[0091] Extract the total number of pixels at each time in the time-series binary distribution product, and determine the area S of a single pixel pixel , S pixel The calculation formula is as follows:
[0092]
[0093] In the formula, Rx and Ry are the horizontal and vertical spatial resolutions of the image respectively.
[0094] Statistical value of 1 in the binary result at time t Calculate the coverage area of the target land class at this time, and repeat the statistical process for all time-series images to generate coverage area time-series data. The formula for the coverage area of the target land class is as follows:
[0095] ;
[0096] In the formula, is the coverage area of the target land class, is the number of pixels.
[0097] S23: Based on the coverage area time-series data, the lake is divided into a nearshore nutrient enrichment zone, a lake center hydrodynamic sensitive zone, and a transition zone according to the spatiotemporal heterogeneity physiological mechanism of cyanobacterial growth, and the division result is obtained through a nonlinear fitting algorithm.
[0098] S231: Retrieve the core characteristics of the spatiotemporal heterogeneity physiological mechanism of cyanobacterial growth, and clarify the physiological response differences of the nearshore nutrient enrichment zone (high nutrient salt, weak hydrodynamic force, and easy cyanobacterial outbreak), the lake center hydrodynamic sensitive zone (strong hydrodynamic force, fast nutrient salt diffusion, and dispersed cyanobacterial distribution), and the transition zone (nutrient salt and hydrodynamic conditions between the former two), which serve as the theoretical basis for the division.
[0099] S232: Import the cyanobacterial coverage area time-series data generated in the previous steps, extract the spatial distribution characteristics, expansion / contraction trend, and peak value occurrence law of the coverage area at each time, and assist in identifying the high-frequency and low-frequency areas of cyanobacterial aggregation.
[0100] S233: Initialize the division algorithm parameters, and according to the core logic of the nonlinear fitting algorithm, take the spatial correlation and temporal fluctuation of the coverage area time-series data as the fitting variables, and set the fitting target function (such as minimizing the coverage area time-series coefficient of variation within the region and maximizing the coefficient of variation between regions).
[0101] S234: Input the coverage area time-series data of the lake global pixels into the fitting model, divide the initial division through iterative calculation, and then modify the boundaries of the initial division according to the physiological mechanism of cyanobacterial growth to eliminate unreasonable fragmented divisions.
[0102] S235: Verify the rationality of the partition results by comparing the temporal variation of cyanobacteria coverage area in different partitions with the matching degree of environmental characteristics such as nutrient salts and hydrodynamic conditions in the corresponding area, adjust the fitting parameters until the partition results meet the expected physiological mechanism, and finally output the spatial partition results of the nearshore nutrient enrichment area, the lake center hydrodynamic sensitive area, and the transition area.
[0103] S24: According to the pixel-level binary distribution product and the temporal data of coverage area of the partition results, analyze the correlation between water bloom and various factors in different areas, determine the dynamic suitability threshold of the core influencing factors of each area, and output the partition factor suitability parameter table.
[0104] S241: Integrate all the data of the previous steps, including the partition results, the pixel-level binary distribution product of each area, the temporal data of coverage area, and the previously extracted NDVI, FAI index and corresponding environmental influencing factor data (such as nutrient salt concentration, water temperature, flow rate, etc.).
[0105] S242: According to the partition, split the data, and calculate the proportion of cyanobacteria coverage area in each time series in the nearshore nutrient enrichment area, the lake center hydrodynamic sensitive area, and the transition area, and extract the mean, extreme value and fluctuation characteristics of each influencing factor at the corresponding time series.
[0106] S243: Construct a correlation analysis model between cyanobacteria coverage area proportion and influencing factors in each partition, and select the core factors that have a significant impact on cyanobacteria growth in each partition through correlation analysis, partial least squares regression, etc.
[0107] S244: For each core influencing factor of each partition, determine its response relationship type (such as positive correlation, negative correlation, nonlinear correlation) with the cyanobacteria coverage area proportion, and divide the dynamic suitability level of the core factor based on the reference of the cyanobacteria coverage area proportion interval suitable for growth (such as the dominant growth interval, the inhibitory growth interval).
[0108] S245: Calculate the threshold range of each core factor at different suitability levels according to the response relationship, such as the suitable growth threshold and inhibitory growth threshold of total phosphorus in the nearshore nutrient enrichment area, and the suitable threshold of flow rate in the lake center hydrodynamic sensitive area.
[0109] S246: The core influencing factor name, suitability level, threshold range and corresponding cyanobacteria growth response characteristics of each partition are compiled into a partition factor suitability parameter table in a unified format.
[0110] S3: According to the refined wind field grid data in the standardized pixel-level data set, the light duration data interpolated by the meteorological station, and the water bloom coverage area time series data associated with the partition factor suitability parameter table, the spatio-temporal matching is performed. According to the remote sensing time series data of the standardized pixel-level data set, the blue-green algae proliferation rate is retrieved. The blue-green algae proliferation rate is retrieved and the factor weight is adjusted. The regional threshold is combined to calculate the pixel comprehensive suitability. The spatio-temporal dynamic pixel-level suitability matrix is output.
[0111] S31: The wind field, light spatio-temporal data and partition factor suitability parameter in the standardized pixel-level data set are preprocessed for spatio-temporal alignment. The time granularity and spatial projection are unified. The parameter table associated data set is output.
[0112] According to the wind field, light spatio-temporal data and partition factor suitability parameter table in the standardized pixel-level data set, the spatio-temporal coordinates and time stamp information of each data are extracted. The spatial projection type and time granularity difference are determined. The to-be-aligned data set with spatio-temporal metadata is output. The spatial reference of the wind field, light data and parameter table is unified to the same geographic coordinate system by using the spatial projection conversion algorithm. The spatial position deviation is eliminated. The intermediate data set with consistent spatial reference is output. According to the high-frequency to low-frequency or interpolation frequency supplement principle, the time series is resampled. The data with different time granularities are unified to the target time interval. The time granularity homogenized transition data set is output. The one-to-one correspondence between the wind field, light data and partition factor suitability parameter is established by using the spatio-temporal correlation algorithm. The redundant data with invalid spatio-temporal matching is removed. The parameter table associated data set is output.
[0113] S32: According to the associated data set, the factor dimension matching and attribute association are performed. The corresponding suitability reference parameter of each spatio-temporal unit is extracted. The spatio-temporal matching data set is output.
[0114] According to the parameter table associated data set, the spatio-temporal unit code of the wind field, light data and the dimension label of the partition factor suitability parameter are parsed. The dimension mapping rule is established. The dimension mapping relationship table is output. The attribute data such as wind speed and light duration in the associated data set are matched and bound with the factor threshold and weight coefficient in the suitability parameter table. The intermediate data set after attribute binding is output. The grouping and screening are performed according to the spatio-temporal unit. The suitability reference parameter corresponding to each spatio-temporal unit is extracted. The core indexes such as the upper and lower limits of the factor suitability, the reference contribution value, etc. are included. The parameter extraction result set is output. The spatio-temporal consistency check is performed. The data entries with missing spatio-temporal units and abnormal parameter values are removed. The parameter interpolation results of the edge spatio-temporal units are supplemented. Finally, the spatio-temporal matching data set with suitability reference parameters is output.
[0115] S33: The spatio-temporal matching data set is substituted into the blue-green algae proliferation rate retrieval model to calculate the initial proliferation rate. The partition factor weight is iteratively adjusted through error analysis. The factor weight set is output.
[0116] S331: Extract the wind field friction velocity u, the effective radiation intensity I of light, and the corresponding partition factor suitability reference parameter f0 of each spatio-temporal unit by calling the spatio-temporal matching data set, and output the model input data set with physical quantity labels.
[0117] S332: According to the model input data set with physical quantity labels, substitute into the cyanobacteria proliferation rate inversion model according to the energy conservation, calculate the initial proliferation rate r0 of each spatio-temporal unit, and output the initial proliferation rate data set.
[0118] The cyanobacteria photosynthetic energy absorption formula is represented as:
[0119]
[0120] In the formula, I is the effective radiation intensity of light, a is the photosynthetic radiation utilization efficiency of cyanobacteria, is the light factor suitability. is the equivalent energy flux of the wind field driving the diffusion of nutrients in the water body, is the wind field factor suitability, , is the wind field and light factor weight.
[0121]
[0122] In the formula, is the energy-biomass conversion efficiency, t is the effective daily light duration, and C is the energy equivalent of cyanobacteria unit biomass.
[0123] S333: According to the initial proliferation rate data set, introduce the measured cyanobacteria biomass growth rate rob s , calculate the model error by residual sum of squares, and construct the error objective function.
[0124] S334: Use the error objective function to iteratively adjust the wind field factor weight , the light factor weight (restricted ), until the error is less than the threshold value, and output the optimized factor weight set .
[0125] S34: According to the factor weight set and the preset factor threshold value of each region, construct a pixel-level comprehensive suitability calculation model, and perform weighted summation on the wind field and light suitability contribution value of each pixel, and output the initial pixel comprehensive suitability data set.
[0126] By calling the optimized factor weight set and the preset factor threshold value of each region, the wind field factor weight , the light factor weight and the upper and lower threshold values of the suitability of each factor, output the model parameter set with weight and threshold constraints. According to the model parameter set with weight and threshold constraints, a pixel-level comprehensive suitability calculation model is constructed, and the core of the model is a weighted summation formula:
[0127]
[0128] wherein, , are the wind field and light suitability values of the pixel, respectively, and the output model structure and calculation rule specification document. According to the model structure and calculation rule specification document, the wind field and light suitability original data of each pixel are extracted and substituted into the weighted summation formula for calculation. At the same time, the suitability values exceeding the range are truncated and corrected according to the regional factor threshold, and the single-pixel suitability calculation result set is output. Using the single-pixel suitability calculation result set, the data is integrated according to the space-time unit, the geographic coordinates and time stamp information of the pixel are associated, the abnormal pixel data is removed, and the interpolation calculation results of the edge pixels are supplemented, and finally the initial pixel comprehensive suitability data set is output.
[0129] S35: The initial pixel comprehensive suitability data set is corrected by smoothing and removing abnormal values and supplementing interpolation results of edge pixels, and a pixel-level suitability matrix is output.
[0130] By calling the initial pixel comprehensive suitability data set, the suitability values, space-time coordinates and neighborhood pixel association information of each pixel are extracted, the abnormal value determination threshold is set, and the data set with neighborhood association labels to be corrected is output. According to the data set with neighborhood association labels to be corrected, the abnormal pixel values exceeding the range are selected by comparing the preset threshold, and the abnormal data is removed by using the neighborhood pixel mean value replacement method, and the blank pixel region without data coverage in the edge is marked, and the intermediate data set after removing abnormal values is output. According to the intermediate data set after removing abnormal values, the inverse distance weighted interpolation algorithm is selected, and the suitability values of the effective pixels in the edge are used for interpolation calculation of the blank edge region, and the suitability data of the missing pixels is supplemented, and the complete data set after interpolation completion is output. Using the complete data set after interpolation completion, Gaussian filtering is used for spatial smoothing correction, the suitability mutation difference of adjacent pixels is weakened, the spatial and temporal continuity of the data is improved, and a spatiotemporal dynamic pixel-level suitability matrix is output.
[0131] S4: According to the time sequence characteristics of the spatiotemporal dynamic pixel-level suitability matrix, combined with the meteorological factor and water bloom state association data after spatiotemporal matching, the meteorological factor lag effect analysis is carried out: Pearson correlation analysis and Granger causality test are used to identify meteorological factors with significant lag effect on water bloom occurrence, high impact factors are selected from the core impact factors by random forest feature importance sorting, and a high impact factor selection list is output.
[0132] S41: Align and calibrate the time sequence dimension of the data by calling the time sequence characteristic data of the spatiotemporal dynamic pixel-level suitability matrix and the matched meteorological and water bloom monitoring data, unify the time sequence length and sampling interval, and output the multi-source fusion dataset with time sequence labels.
[0133] By calling the time sequence characteristic data of the spatiotemporal dynamic pixel-level suitability matrix and the matched meteorological and water bloom monitoring data, the time stamp and sampling frequency information of each dataset are extracted, the time sequence difference between the data is clarified, and the multi-source dataset to be aligned with time sequence metadata is output. The maximum time span in the study period is selected as the unified time sequence length, the resampling principle of aligning high-frequency data to low-frequency data is determined, and the intermediate dataset with normalized time sequence length is output. Linear interpolation method is used to supplement the data of missing time nodes, and the datasets with different sampling intervals are unified to the target time resolution, and the multi-source dataset with consistent time granularity is output. A one-to-one mapping relationship between the pixel suitability time sequence characteristics and the corresponding meteorological and water bloom data is established by the spatiotemporal correlation algorithm, redundant data with invalid time sequence matching are removed, and finally the multi-source fusion dataset with time sequence labels is output.
[0134] S42: According to the multi-source fusion dataset with time sequence labels, the correlation analysis of suitability time sequence change and meteorological and water bloom index fluctuation is carried out, the Pearson correlation coefficient is calculated, the potential correlation factors are preliminarily screened, and the correlation factor candidate list is output.
[0135] By calling the multi-source fusion dataset with time sequence labels, the suitability time sequence, meteorological factor time sequence and water bloom index time sequence of each spatiotemporal unit are extracted, the sequence dimension matching and data cleaning are completed, and the time sequence analysis dataset with feature labels is output. For the time sequence change of suitability and each meteorological and water bloom index, the Pearson correlation coefficient is calculated group by group, the linear correlation strength and positive and negative correlation between variables are clarified, and the correlation coefficient matrix and significance test result set are output. Set the absolute value threshold of the correlation coefficient and the significance level threshold, screen out the potential correlation factors that pass the threshold test, and output the correlation factor candidate list containing the factor name, correlation coefficient and significance level.
[0136] S43: According to the correlation factor candidate list, the Granger causality test method is introduced to verify the causal relationship between the candidate factors and the water bloom, identify the key driving factors with lag effect, and output the causal factor set with lag attribute.
[0137] The time series data of each factor in the list and the time series sequence of the occurrence state of the water bloom are extracted by calling the associated factor candidate list, the data stationarity test and difference processing are completed, and the stationary time series data set meeting the Granger causality test premise is output. According to the stationary time series data set, a Granger causality test model of the factor time series and the water bloom occurrence time series is constructed, different lag orders are set for hypothesis testing, the F statistic and the significance P value are calculated, and the causality test result table of each factor is output. According to the causality test result table, the factors with P values less than the significance threshold are screened, the optimal lag order of each factor and the water bloom occurrence is determined, the key driving factors with lag effect are marked, and the candidate factor set with lag order label is output. The candidate factor set with lag order label is used to perform secondary screening on the explanatory power of the factors on the occurrence of water bloom, and the redundant factors with insignificant lag effect are removed, and the causality factor set with lag attribute is output.
[0138] S44: Using the causality factor set with lag attribute, a feature importance ranking model is constructed, a random forest algorithm is used to quantify the contribution of each factor to the water bloom, and a factor priority sequence ranked in descending order of contribution is output.
[0139] The time series feature values, lag order parameters and corresponding water bloom occurrence state labels of each factor in the causality factor set with lag attribute are extracted, data feature coding and label mapping are performed, and a labeled factor feature data set is output. The labeled factor feature data set is divided into training set and test set, a random forest feature importance ranking model is constructed, core parameters such as the number of decision trees and the maximum depth are set, and an initialized random forest model is output. The training set data is input into the initialized random forest model for iterative training, the out-of-bag (OOB) data is used to evaluate the model accuracy, the contribution of each factor to the water bloom is quantified by the reduction of Gini coefficient or mean square error, and a factor contribution score table is output. The causality factors are sorted according to the contribution value from high to low, the lag order attribute corresponding to each factor is labeled, the weak influence factors with contribution degree lower than the preset threshold are removed, and a factor priority sequence ranked in descending order of contribution is output.
[0140] S45: According to the factor priority sequence, high-impact factors with contribution higher than the preset threshold are extracted, and the optimal lag time parameters of each factor are simultaneously calculated, and a high-impact factor list and corresponding lag parameters are output.
[0141] The contribution degree value of each factor in the factor priority sequence, the optimal lag order and the corresponding water bloom influence characteristics are extracted, a contribution degree threshold is set, and a factor characteristic dataset with threshold screening conditions is output. According to the factor characteristic dataset with threshold screening conditions, the contribution degree of each factor is compared with the preset threshold, and high-impact factors with a contribution degree higher than the threshold are screened out. The optimal lag time parameter corresponding to each high-impact factor is extracted synchronously, and a high-impact factor preliminary screening list is output. Cross-validation is performed in combination with the factor action law of historical water bloom events to confirm the rationality of the lag time parameter, and factors that fail to pass the verification are excluded. A calibrated high-impact factor list and lag parameter set are output. The factor order is arranged in descending order of contribution degree, and key information such as the lag time range and action intensity of each factor is labeled. A standardized high-impact factor list and corresponding lag parameter are output.
[0142] S5: According to the spatio-temporal dynamic pixel-level suitability matrix quantified water bloom occurrence basic condition, combined with the improved phytoplankton index (PAI) and phycocyanin index (PCI) collaborative recognition algorithm, the initial water bloom pixel is accurately extracted, the high-impact factor screening list and the lag effect parameter are constructed, and the wind field-water dynamic coupling diffusion rule is constructed, which considers the influence of wind speed and direction on the migration of blue-green algae, and modifies the blue-green algae proliferation rate through the lag effect of high-impact factors, realizes the coupling of migration and proliferation, and outputs the high-precision initial water bloom pixel distribution result.
[0143] S51: By calling the spatio-temporal dynamic pixel-level suitability matrix, the water bloom occurrence basic condition quantization index corresponding to each pixel is extracted, the data standardization and value range normalization processing is completed, and the pixel-level quantization dataset with basic condition label is output.
[0144] S52: According to the pixel-level quantization dataset with basic condition label, the PAI (improved phytoplankton index) and PCI (phycocyanin index) collaborative recognition algorithm is introduced, the water bloom occurrence probability of each pixel is preliminarily judged, and the water bloom suspected pixel set and index matching result are output.
[0145] By calling the pixel-level quantization dataset, the water bloom occurrence basic condition index corresponding to each pixel is extracted, the index calculation input dataset is output by combining the PAI and PCI calculation threshold, and the radar image assisted inversion of water turbidity is provided. Background parameters for PCI index calculation.
[0146] The exponential calculation input data set is substituted into the improved PAI and PCI collaborative calculation model to calculate the PAI value and the PCI value of each pixel, the preliminary qualified pixels are screened through the double exponential intersection judgment rule, and the double exponential preliminary screening pixel set is output. The chlorophyll concentration (Chla) of the pixel is calculated by the specific band reflectivity of the satellite remote sensing multispectral image and the chlorophyll a inversion model. Based on the reflectivity of the red band and the near-infrared band of the satellite image, the turbidity is inverted through the turbidity inversion model, and the accuracy is calibrated combined with the turbidity instrument measured data of the sampling points. If there is a radar image, the scattering coefficient can also be used to assist the optimization of the turbidity inversion result to obtain the pixel turbidity (T). The PC value of the pixel scale is calculated through the reflectivity of the characteristic band corresponding to the satellite hyperspectral image of the phycocyanobilin and the phycocyanobilin inversion model.
[0147] The calculation formula of the improved phytoplankton algae index (PAI) is:
[0148]
[0149] In the formula, Chla is the chlorophyll a concentration of the pixel, T is the turbidity of the pixel, is the calibration coefficient.
[0150] The calculation formula of the improved phycocyanobilin index (PCI) is:
[0151]
[0152] In the formula, PC is the phycocyanobilin concentration of the pixel, , is the minimum value and the maximum value of the PC concentration of the research area, is the background water reflectivity, is the correction coefficient.
[0153] According to the double exponential preliminary screening pixel set, the pixel neighborhood correlation analysis is introduced, the PAI and PCI distribution characteristics of the surrounding pixels are combined to correct the single pixel judgment result, the isolated abnormal pixels are removed, and the corrected water bloom suspected candidate pixel set is output.
[0154] The corrected water bloom suspected candidate pixel set is compared with the historical water bloom pixel characteristic data to perform matching degree verification, the pixels with a matching degree higher than a preset threshold are reserved, and finally the water bloom suspected pixel set with a labeled water bloom occurrence probability is output.
[0155] S53: Using the water bloom suspected pixel set and the index matching result, combining the high impact factor screening list and the lag effect parameter, analyzing the driving mechanism of the factor lag effect on the water bloom occurrence, and outputting the water bloom suspected pixel optimization set under the factor driving.
[0156] By calling the water bloom suspected pixel set and the index matching result, the high impact factor screening list and the lag effect parameter are associated, the factor type, the lag time window and the index matching degree corresponding to each pixel are extracted, and the pixel analysis data set with factor lag attribute is output. According to the pixel analysis data set with factor lag attribute, the partial correlation analysis method is used to quantify the contribution degree of each high impact factor to the water bloom occurrence probability under different lag time, to clarify the driving strength and time efficiency characteristics of the factor lag effect, and to output the factor driving strength ranking table. According to the factor driving strength ranking table, the factor-index coupling driving judgment model is constructed, the contribution degree threshold is set to screen the strong driving factor, the pixels in the original suspected pixel set are judged again, and the preliminary optimized water bloom suspected pixel set is output. Using the preliminary optimized water bloom suspected pixel set, cross verification is carried out combined with the factor lag rule of historical water bloom events, the misjudged pixels with low driving factor matching degree are removed, the missed pixels corresponding to the high contribution degree factor are supplemented, and finally the water bloom suspected pixel optimization set driven by the factor is output.
[0157] S54: According to the water bloom suspected pixel optimization set driven by the factor, the diffusion trend of water bloom under different wind field and hydrodynamic conditions is simulated combined with the wind field-hydrodynamic coupling diffusion rule, the attribution judgment of edge pixels is corrected, and the corrected water bloom pixel distribution data set is output.
[0158] By calling the water bloom suspected pixel optimization set driven by the factor, the spatial coordinates, the water bloom occurrence probability and the neighborhood topological relationship of the pixels are extracted, combined with the wind speed, flow direction and water flow speed core parameters of the wind field-hydrodynamic coupling diffusion rule, the pixel analysis data set with diffusion constraint condition is output. According to the pixel analysis data set with diffusion constraint condition, the water bloom diffusion numerical simulation model is constructed, the shear stress of wind field and the advection diffusion effect of hydrodynamic are taken as driving items, the diffusion direction and range of pixels under different space-time scenes are simulated, and the water bloom diffusion trend simulation result set is output. According to the water bloom diffusion trend simulation result set, aiming at the attribution fuzzy problem of edge pixels, the diffusion probability weighting algorithm of neighborhood pixels is introduced, the water bloom attribution confidence of edge pixels is calculated, the misjudged pixels in the original judgment result are corrected, and the edge corrected intermediate data set is output. Using the edge corrected intermediate data set, the spatial topological consistency verification is carried out, the isolated pixels appearing in the diffusion simulation are removed, the missing pixel interpolation results on the diffusion path are supplemented, and the corrected water bloom pixel distribution data set is output.
[0159] S55: Using the corrected water bloom pixel distribution data set, spatial consistency verification and precision verification are carried out, misjudged pixels are removed and missing pixel interpolation results are supplemented, and finally the high-precision initial water bloom pixel distribution result is output.
[0160] The corrected water bloom pixel distribution data set is called to extract the spatial coordinates, water bloom determination label and neighborhood pixel correlation information of each pixel, set the spatial continuity verification rule, and output the verification data set with neighborhood correlation characteristics. According to the verification data set with neighborhood correlation characteristics, the spatial consistency verification is performed, the isolated misjudgment pixels without neighborhood support are identified and removed, the missed pixel area caused by the deviation of diffusion simulation is marked, and the post-misjudgment data set is output. According to the post-misjudgment data set to be completed, the inverse distance weighted interpolation algorithm is introduced, the water bloom characteristic value of the effective pixels around the missed area is used for interpolation calculation, the determination result of the missing pixel is supplemented, and the complete data set after interpolation completion is output. The complete data set after interpolation completion is used to carry out precision verification combined with the measured water bloom monitoring point data, the verification indexes such as confusion matrix and Kappa coefficient are calculated, the pixel data meeting the precision threshold are screened, and finally the high-precision initial water bloom pixel distribution result is output.
[0161] S6: Taking the high-precision initial water bloom pixel distribution as the diffusion starting point, constructing a neighborhood iterative diffusion model combined with the time-space dynamic pixel-level suitability matrix, realizing the coupling of water bloom growth and diffusion through the proliferation coefficient and migration probability, and iteratively simulating and outputting the pixel-level water bloom prediction map through in-situ sample calibration, remote sensing data correction and Bayesian optimization parameters.
[0162] S61: Extracting the spatial coordinates, suitability values and neighborhood topological relations of the initial water bloom pixel distribution result and the time-space dynamic pixel-level suitability matrix, and outputting the pixel modeling data set with diffusion basic parameters.
[0163] The spatial coordinate information of the initial water bloom pixel distribution result and the time-space dynamic pixel-level suitability matrix is extracted, the coordinate one-to-one mapping relationship is established, and the pixel basic data set with coordinate correlation label is output. The suitability quantization value corresponding to each matched pixel is extracted, the 8-neighborhood topological structure of the pixel is identified and labeled, and the intermediate data set with attribute correlation is output. Combined with the basic parameter requirements of water bloom diffusion, the three types of core information of pixel coordinates, suitability values and neighborhood topological relations are integrated, the redundant pixels with invalid coordinate matching are removed, and the pixel modeling data set with diffusion basic parameters is output.
[0164] S62: According to the pixel modeling data set with diffusion basic parameters, a neighborhood pixel iterative diffusion model is constructed, the proliferation coefficient associated with suitability and the neighborhood spatial migration probability are introduced, the dynamic coupling of water bloom growth and diffusion is realized, and the model initial simulation result set is output.
[0165] By calling the pixel modeling dataset with diffusion basic parameters, the suitability value, neighborhood topological relationship and diffusion basic parameters of the pixel are extracted, the quantitative correlation formula of suitability and multiplication coefficient is established, and the model input dataset with multiplication parameters is output. According to the model input dataset with multiplication parameters, the core framework of the neighborhood pixel iterative diffusion model is constructed, the migration probability calculation rule of the neighborhood pixel is defined, the multiplication coefficient is bound with the migration probability to form the coupling driving logic, and the initialized iterative diffusion model is output. According to the initialized iterative diffusion model, taking the initial algal bloom pixel as the starting point, the pixel suitability and neighborhood topological data are substituted, the first round of iteration calculation is performed, the simulation results of algal bloom growth and diffusion are generated, and the first round of iteration result set of the model is output. Using the first round of iteration result set of the model, setting the iteration step length of multiple periods, repeating the cycle logic of multiplication calculation-migration determination-pixel update, simulating the spatio-temporal evolution process of algal bloom at different periods, and outputting the initial simulation result set of the model.
[0166] S63: According to the initial simulation result set of the model, a small amount of in-situ observation samples are selected for pixel-level matching, the accuracy of pixel biomass is calibrated, the multi-period remote sensing time series verification data is called synchronously, and the simulation intermediate data set after accuracy calibration is output.
[0167] The fine wind field grid data in the standardized pixel-level data set, the interpolated illumination duration data of the weather station, and the water bloom coverage area time series data associated with the partition factor suitability parameter table are output as the pixel-level meteorological-water bloom fusion data set with consistent time and spatial matching through time stamp alignment and spatial grid mapping.
[0168] According to the remote sensing time series data of the above-mentioned pixel-level meteorological-water bloom fusion data set, the blue-green algae multiplication rate is calculated by using the inversion algorithm, the inversion result is input into the factor contribution degree dynamic weight model, the weight proportion of water temperature and nitrogen-phosphorus ratio is adjusted in real time, and the blue-green algae multiplication-environment factor correlation data set with dynamic weight is output.
[0169] According to the blue-green algae multiplication-environment factor correlation data set with dynamic weight, combined with the regional threshold of the partition factor suitability parameter table, the comprehensive suitability of each pixel is obtained by weighted calculation, and finally the spatio-temporal dynamic pixel-level suitability matrix is output.
[0170] S64: Using the simulation intermediate data set after accuracy calibration, the threshold value of the spatio-temporal dynamic pixel-level suitability matrix and the parameter of the iterative diffusion model are corrected reversely, the simulation effect is evaluated by confusion matrix, and the simulation result set after parameter optimization is output.
[0171] Integrate the simulation intermediate data set after accuracy calibration and the existing spatio-temporal dynamic pixel-level suitability matrix, and use the calibration data as a reference to correct the suitability threshold of the matrix. At the same time, the core parameters of the iterative diffusion model are adjusted, and the model input dataset after threshold and parameter co-correction is output.
[0172] According to the threshold and the parameter cooperative correction model input data set, substitute into the iterative diffusion model for simulation operation, through the confusion matrix quantization comparison simulation results and actual monitoring data of the degree of agreement, evaluate and filter out the optimal simulation parameter combination, output simulation effect evaluation report.
[0173] According to the optimal parameter combination confirmed in the simulation effect evaluation report, re-drive the iterative diffusion model to execute the simulation process, combine the corrected spatiotemporal dynamic pixel level suitability matrix to complete the whole process operation, and finally output the simulation result set after parameter optimization.
[0174] S65: According to the simulation result set after parameter optimization, the Bayesian optimization algorithm is used to adjust the core parameters of the model, and the water bloom evolution process in different time periods is iteratively simulated, and finally the pixel level water bloom prediction map with accuracy and practicality is output.
[0175] According to the simulation result set after parameter optimization, the key feature parameters of water bloom evolution are extracted as prior information, which are input into the Bayesian optimization algorithm to adjust the core parameters of the model with simulation accuracy as the optimization target, and the model configuration scheme after parameter iterative optimization is output.
[0176] According to the model configuration scheme after parameter iterative optimization, substitute into the environmental driving data in different time periods to execute multiple rounds of iterative simulation operation, restore the occurrence, development and diffusion process of water bloom in different time dimensions, and output the multi-period water bloom evolution simulation data set.
[0177] According to the multi-period water bloom evolution simulation data set, the accuracy and practicality of the simulation results are verified combined with the actual monitoring data, and the simulation results of each period are integrated through pixel level spatial mapping, and finally the pixel level water bloom prediction map with accuracy and practicality is output.
[0178] In summary, the embodiment provides a lake cyanobacterial bloom pixel level prediction method based on multi-source data fusion, which calibrates in-situ samples, corrects remote sensing data, continuously optimizes simulation intermediate results combined with confusion matrix evaluation, reversely corrects suitability matrix threshold and iterative diffusion model parameters, and effectively reduces the system error of the model. The Bayesian optimization algorithm is introduced to iteratively adjust the core parameters with simulation accuracy as the target, to generate multi-period water bloom evolution simulation data set, and to verify the accuracy and practicality of the results combined with actual monitoring data. This layer-by-layer progressive optimization strategy ensures the stability of the model in different spatiotemporal scenarios, and the output pixel level water bloom prediction map can accurately reflect the spatiotemporal distribution characteristics of water bloom, which can provide intuitive and reliable decision basis for lake water bloom prevention and ecological management.
[0179] The initial pixel is the anchor point for the subsequent spatio-temporal evolution simulation of water bloom, and its accuracy directly determines the reliability of the final pixel-level prediction result. The diffusion of blue-green algae bloom is gradually spreading from the local dominant area. If the whole lake pixel is blindly predicted, the ecological law of water bloom aggregation and then diffusion will be ignored. The initial pixel determines the core source of water bloom outbreak, and the subsequent neighborhood iterative diffusion simulation has a real spatial starting point. Through the selection of high impact factors, lag effect parameters and partition suitability threshold, the initial pixel is needed to accurately land. Only by locking the initial pixel can the factor contribution of the region be analyzed, avoiding the use of average parameters in the whole lake instead of local core driving factors, and making the diffusion rules more realistic. If the initial pixel is not extracted, the time series simulation and diffusion calculation of all pixels in the whole lake will produce a large amount of invalid operation. The initial pixel can reduce the simulation range and focus on the potential outbreak area of water bloom, greatly improving the calculation efficiency while ensuring the prediction accuracy, and adapting to the demand of high-resolution pixel-level prediction.
[0180] Those skilled in the art can clearly understand from the description of the above embodiments that each embodiment can be realized by means of software and the necessary general hardware platform, and of course, it can also be realized by hardware. Based on such understanding, the above technical solutions or the essential part of the prior art can be embodied in the form of a software product, which can be stored in a computer readable storage medium such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the method of each embodiment or some part of the embodiment.
[0181] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion, characterized in that, The method comprises the following steps: S1: Collecting lake pixel-level multi-source basic data in satellite images and preprocessing, outputting a standardized pixel-level data set; S2: Screening core influencing factors from the standardized pixel-level data set, calculating an algae index to generate a binary distribution product, combining a zoned fitting determination of a dynamic threshold value, and outputting a zoned factor suitability parameter table; S21: Extracting core influencing factors from the standardized pixel-level data set, and calculating NDVI and FAI indexes using the preprocessed satellite images in the standardized pixel-level data set; S22: Generating a pixel-level binary distribution product and coverage area time series data using threshold segmentation according to the NDVI and FAI indexes; S23: According to the coverage area time series data, combining the spatio-temporal heterogeneity physiological mechanism of blue-green algae growth, dividing the lake into a near-shore nutrient enrichment area, a lake center hydrodynamic sensitive area and a transition area through a zoned nonlinear fitting algorithm to obtain a zoned result; S24: According to the zoned result, pixel-level binary distribution product and coverage area time series data, analyzing the correlation between water bloom and each factor in different regions, determining the dynamic suitability threshold of the core influencing factor in each region, and outputting a zoned factor suitability parameter table; S3: According to the wind field and illumination data in the standardized pixel-level data set, and the zoned factor suitability parameter table, performing spatio-temporal matching, inverting the blue-green algae proliferation rate and adjusting the factor weight, combining the regional threshold value to calculate the pixel comprehensive suitability, and outputting a spatio-temporal dynamic pixel-level suitability matrix; S4: According to the time series characteristics of the spatio-temporal dynamic pixel-level suitability matrix, combining the matched meteorological and water bloom data, identifying lag effect factors through correlation analysis and causality test, screening high impact factors through feature sorting, and outputting a high impact factor list; S5: According to the spatio-temporal dynamic pixel-level suitability matrix quantization and the high impact factor list, constructing a diffusion rule, extracting an initial water bloom pixel and determining a diffusion starting point, and outputting an initial water bloom pixel distribution result; S51: Extracting the water bloom occurrence basic condition quantization index corresponding to each pixel in the spatio-temporal dynamic pixel-level suitability matrix, performing data standardization and value range normalization processing, and outputting a pixel-level quantization data set; S52: According to the pixel-level quantization data set, introducing an improved type of plankton algae index and an algal blue protein index cooperative identification algorithm to preliminarily judge the water bloom occurrence probability of the pixel, and outputting a water bloom suspected pixel set; S53: According to the water bloom suspected pixel set and the high impact factor screening list, analyzing the driving mechanism of the factor lag effect on water bloom occurrence, and outputting a water bloom suspected pixel optimization set; S54: According to the water bloom suspected pixel optimization set and the wind field-hydrodynamic coupling diffusion rule, simulating the diffusion trend of water bloom under different wind field and hydrodynamic conditions, correcting the attribution judgment of edge pixels, and outputting a water bloom pixel distribution data set; S55: Spatial consistency checking and precision verification are performed on the water bloom pixel distribution data set, the misjudged pixels are removed, the interpolation results of the missed pixels are supplemented, and an initial water bloom pixel distribution result is output. S6: Taking the high-precision initial algal bloom pixel distribution as the diffusion starting point, using the spatio-temporal dynamic pixel-level suitability matrix to construct a neighborhood iterative diffusion model, and realizing the coupling of algal bloom growth and diffusion through proliferation coefficient and migration probability, an iterative simulation output pixel-level algal bloom prediction map is obtained.
2. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S23, the specific steps for obtaining the partition result are as follows: S231: According to the core characteristics of the spatio-temporal heterogeneity physiological mechanism of cyanobacterial growth, the physiological response differences of the nearshore nutrient enrichment area, the lake center hydrodynamic sensitive area and the transition area are determined to form the theoretical basis for partitioning; S232: According to the partitioning theoretical basis, the spatial distribution characteristics, expansion and contraction trend and peak occurrence law of the cyanobacterial coverage area time series data are extracted to identify high and low frequency areas of cyanobacterial aggregation; S233: According to the high and low frequency areas, using a partitioning nonlinear fitting algorithm, the spatial correlation and temporal fluctuation of the coverage area time series data are used as fitting variables to construct a fitting target model; S234: The coverage area time series data of the lake global pixel are input into the fitting model, and the initial partition is divided through iterative calculation; S235: By comparing the temporal variation law of cyanobacterial coverage area in different partitions with the corresponding region's nutrient salt and water dynamic adjustment fitting parameters, the spatial partition results of the nearshore nutrient enrichment area, the lake center hydrodynamic sensitive area and the transition area are output.
3. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S3, the specific steps for outputting the spatio-temporal dynamic pixel-level suitability matrix are as follows: S31: The spatio-temporal data of wind field, light and partition factor suitability parameters in the standardized pixel-level data set are aligned and preprocessed, the time granularity and space projection are unified, and the parameter table associated data set is output; S32: According to the associated data set, the factor dimension matching and attribute association are carried out, the corresponding suitability reference parameters of each spatio-temporal unit are extracted, and the spatio-temporal matching data set is output; S33: The spatio-temporal matching data set is substituted into the cyanobacterial proliferation rate inversion model to calculate the initial proliferation rate, the partition factor weight is adjusted through error analysis iteration, and the factor weight set is output; S34: According to the factor weight set and the preset factor threshold of each region, a pixel-level comprehensive suitability calculation model is constructed, the wind field and light suitability contribution values of each pixel are weighted and summed, and an initial pixel comprehensive suitability data set is output; S35: The initial pixel comprehensive suitability data set is smoothed and corrected, the abnormal values are removed and the edge pixel interpolation results are supplemented, and the pixel-level suitability matrix is output.
4. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 3, characterized in that: In step S33, the specific steps for outputting the factor weight set are as follows: S331: Extract the wind field friction wind speed, light effective radiation intensity and corresponding partition factor suitability reference parameters of each spatio-temporal unit in the spatio-temporal matching data set, and output the model input data set; S332: Substitute the model input data set into the cyanobacterial proliferation rate inversion model based on energy conservation to calculate the initial proliferation rate of each spatio-temporal unit, and output the initial proliferation rate data set; S333: According to the initial proliferation rate data set and the measured cyanobacterial biomass growth rate, an error objective function is constructed by calculating the model error through residual sum of squares; S334: The wind field factor weight and light factor weight are adjusted using the error objective function until the error is less than the threshold, and the factor weight set is output.
5. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S4, the specific steps for outputting the high impact factor list are: S41: Align and calibrate the temporal feature data of the spatiotemporal dynamic pixel-level suitability matrix, meteorological data, and water bloom monitoring data, and output a multi-source fusion data set; S42: Analyze the suitability temporal variation and meteorological and water bloom index fluctuation based on the multi-source fusion data set, calculate the Pearson correlation coefficient, preliminarily screen the potential associated factors, and output an associated factor candidate list; S43: According to the associated factor candidate list, use the Granger causality test method to verify the causality between the candidate factors and the occurrence of water bloom, identify the key driving factors with lag effects, and output a causal factor set; S44: Construct a feature importance ranking model based on the causal factor set, use the random forest algorithm to quantify the contribution of each factor to water bloom, and output a factor priority sequence; S45: Extract high impact factors with a contribution degree higher than a preset threshold from the factor priority sequence, and calculate the optimal lag time parameter of each factor, and output a high impact factor list.
6. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S52, the specific steps for outputting the water bloom suspected pixel set are: By calling the pixel-level quantitative data set, the corresponding water bloom occurrence basic condition index of each pixel is extracted, and the index calculation input data set is output by combining the calculation threshold of PAI and PCI; The index calculation input data set is substituted into the improved PAI and PCI co-computation model to calculate the PAI value and PCI value of each pixel, respectively. The preliminary qualified pixels are screened by the double-index intersection judgment rule, and the double-index preliminary screening pixel set is output; According to the double-index preliminary screening pixel set, the pixel neighborhood correlation analysis is introduced, the PAI and PCI distribution characteristics of the surrounding pixels are combined to correct the single-pixel judgment result, and the isolated abnormal pixels are removed, and the corrected water bloom suspected candidate pixel set is output; The corrected water bloom suspected candidate pixel set is compared with the historical water bloom pixel feature data to verify the matching degree, and the pixels with a matching degree higher than a preset threshold are retained, and finally the water bloom suspected pixel set with a labeled water bloom occurrence probability is output.
7. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 1, characterized in that: In step S6, the specific steps for outputting the pixel-level water bloom prediction map are: S61: Extract the spatial coordinates, suitability values, and neighborhood topological relationships of the initial water bloom pixel distribution results and the spatiotemporal dynamic pixel-level suitability matrix, and output a pixel modeling data set; S62: Construct a neighborhood pixel iterative diffusion model based on the pixel modeling data set, associate the proliferation coefficient with the neighborhood spatial migration probability, and output a model initial simulation result set; S63: Based on the model initial simulation result set, a small number of in-situ observation samples are selected for pixel-level matching to calibrate the pixel biomass accuracy, and a multi-period remote sensing time series verification data is called to output a simulation intermediate data set; S64: According to the simulation intermediate data set, the threshold value of the spatiotemporal dynamic pixel-level suitability matrix and the iterative diffusion model parameters are corrected in reverse, the simulation effect is evaluated through the confusion matrix, and a simulation result set is output; S65: According to the simulation result set, the Bayesian optimization algorithm is used to adjust the core parameters of the model, and the water bloom evolution process in different time periods is iteratively simulated, and a pixel-level water bloom prediction map is output.
8. The lake cyanobacterial bloom pixel-level prediction method based on multi-source data fusion according to claim 7, characterized in that: In step S62, the specific steps of outputting the initial simulation result set of the model are: Extracting the suitability value, neighborhood topological relationship and diffusion basic parameters of the pixel in the pixel modeling data set, establishing a quantitative correlation formula between the suitability and the multiplication coefficient, and outputting the model input data set; According to the model input data set, constructing the core framework of the neighborhood pixel iterative diffusion model, defining the migration probability calculation rule of the neighborhood pixel, binding the multiplication coefficient and the migration probability to form the coupling driving logic, and outputting the initial iterative diffusion model; Taking the initial algal bloom pixel as the starting point, substituting the pixel suitability and neighborhood topological data into the initial iterative diffusion model, performing the first round of iteration calculation, generating the simulation result of the growth and diffusion of the algal bloom, and outputting the first round of iteration result set of the model; Using the first round of iteration result set of the model, setting the iteration step of multiple time periods, repeating the execution, simulating the spatio-temporal evolution process of the algal bloom at different time periods, and outputting the initial simulation result set of the model.
Citation Information
Patent Citations
Intelligent monitoring method and system for cyanobacterial bloom outbreak
CN121280783A
Method and device for determining extraction model of green tide coverage ratio based on mixed pixels
US20220129674A1