Water bloom prediction method and system based on ecological niche fitness

Through high-time-resolution water sampling and multi-fractal detrended coupled fluctuation analysis combined with a multi-scale deep learning model, the monitoring problem of rapid dynamic changes in algal blooms was solved, and effective early warning of algal bloom outbreaks was achieved.

CN120654889APending Publication Date: 2025-09-16四川省生态环境监测总站
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510789917.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively monitor and warn of the rapid dynamic changes of algal blooms, especially on large spatial scales, making it difficult to prevent and control water pollution.

Method used

A water bloom prediction method based on niche fitness is adopted. Through high-temporal-resolution water sampling, random forest model, multi-fractal detrended coupled fluctuation analysis and multi-scale deep learning model, combined with real-time data from automatic monitoring stations, the concentration of key dominant algae is predicted in real time.

Benefits of technology

It has achieved effective early warning of algal bloom outbreaks, improved the monitoring accuracy of algal growth trends and spatiotemporal heterogeneity, and can predict changes in algal concentration in real time, providing theoretical support for algal bloom early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654889A_ABST
    Figure CN120654889A_ABST
Patent Text Reader

Abstract

The invention provides a water bloom prediction method and system based on ecological niche fitness, and belongs to the technical field of water bloom prediction. In the training stage, a water sample collection method based on high time resolution is adopted, the fluctuation condition of the density of different types of algae along with time can be effectively reflected, the growth trend of the different types of algae can be more clearly reflected, and the prediction precision of the model is improved; in the prediction stage, the ecological niche fitness of competitive growth of different algae is quantitatively analyzed through a multi-fractal detrending coupling fluctuation analysis method, and the competitive growth relation of different types of algae under natural conditions is reflected; water quality data monitored by an automatic monitoring station in real time and ecological niche fitness of competitive growth of different algae are used as input, the concentration of key dominant algae can be predicted in real time, the spatial-temporal heterogeneity and the rapid dynamic change process of the algae are reflected, and an effective theoretical support is provided for early warning of algal bloom outbreak.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of algal bloom prediction, and in particular relates to an algal bloom prediction method and system based on niche fitness. Background Art

[0002] Algal blooms are extremely complex, influenced by a variety of factors. These include anthropogenic emissions of pollutants such as nitrogen and phosphorus, shifts in endogenous pollution in lake sediments, and changes in climatic conditions such as temperature. Most importantly, water bodies often harbor a large number of diverse algae species, which compete for nutrients during their growth. Furthermore, these algae exhibit high spatial and temporal heterogeneity and rapid dynamics under the influence of wind and hydrodynamic forces. Population structure often undergoes significant changes on timescales ranging from minutes to hours, posing significant challenges to the prevention and control of lake water pollution.

[0003] In related technologies, there are many ways to monitor and warn of algal blooms: (1) Satellite remote sensing can achieve synchronous observation of algal blooms on a large spatial scale. However, it is difficult to capture the rapid dynamic changes of different algal blooms on a scale of minutes to hours due to the time and space conversion of satellite operation, the spectral resolution of sensors, and cloud and rain interference. (2) Fixed monitoring stations often have limited observation ranges and still have difficulty observing the rapid dynamic changes of different algal blooms; (3) Large-scale, continuous, and high-frequency manual sampling and monitoring will bring huge manpower costs.

[0004] Therefore, it is necessary to provide a water bloom prediction method and system based on niche fitness to solve the above problems. Summary of the Invention

[0005] The present invention provides a method and system for predicting algal blooms based on niche fitness. In the model training stage, a water sample collection method based on high temporal resolution is adopted, which can effectively reflect the fluctuation of the density of different types of algae over time, more clearly reflect the growth trend of different types of algae, and improve the prediction accuracy of the model. In the prediction stage, a multifractal detrending coupled fluctuation analysis method is used to quantitatively analyze the niche fitness of different algae for competitive growth, reflecting the competitive growth relationship of different types of algae under natural conditions. Using water quality data monitored in real time by an automatic monitoring station and the niche fitness of different algae for competitive growth as input, the concentration of key dominant algae can be predicted in real time, reflecting the spatiotemporal heterogeneity and rapid dynamic change process of algae, providing effective theoretical support for early warning of algal bloom outbreaks, and thus effectively solving at least one technical problem involved in the background technology.

[0006] In order to solve the above-mentioned technical problems, the present invention is achieved as follows: A method for predicting algal blooms based on ecological niche fitness comprises the following steps: Step S1: Collect water samples from the target water area at the same time interval, analyze the water samples to obtain water quality data and the concentrations of different types of algae at different time nodes, calculate the first-order difference of the concentration of each type of algae at two adjacent time nodes, and traverse all time nodes to form a concentration fluctuation sequence for each type of algae; Step S2: Using the concentration fluctuation series of all types of algae and water quality data as the original data set, the synthetic minority class oversampling method is used to enhance the original data set to form a first data set. The first data set is used to construct a random forest model to learn the mapping relationship between water quality data and the concentrations of different types of algae. In the prediction stage, the water quality data of the target water area at the current time node is collected and input into the random forest model. The concentration fluctuation prediction values ​​of different types of algae are output. Combined with the algae concentration at the current time node, the concentration prediction values ​​of different types of algae at the next time node are converted. The algae are sorted in descending order according to the concentration prediction value, and the top 80% of the algae are selected as the key dominant algae. Step S3, analyzing the predicted concentration values ​​of key dominant algae using a multifractal detrended coupled fluctuation analysis method, and quantitatively analyzing the niche fitness of different algae competitive growth by analyzing the intensity changes in the coupled correlation of different types of algae concentrations; Step S4, constructing a multi-scale deep learning prediction model, using the water quality data at the current time node, the niche fitness of different algae competitive growth and the concentration of different types of algae to jointly construct a second data set, using the second data set to train the multi-scale deep learning prediction model, learning the mapping relationship between water quality data and the niche fitness of different algae competitive growth to the concentration of different types of algae, after the training is completed, using the detected water quality data and the niche fitness of different algae competitive growth as input, outputting the predicted value of the concentration of key dominant algae.

[0007] As a preferred improvement, the water sample collection process is as follows: The plankton net with a mesh diameter of 0.064 mm was mounted on an unmanned boat in an area larger than 50m 2 The plankton net was collected in the target waters. During the collection process, the water outlet piston switch at the bottom of the plankton net was closed, and the plankton net was made to do an "∞"-shaped reciprocating motion at a speed of 20-30 cm / s at a depth of 0.5 m below the water surface. After slowly dragging for 1-3 minutes, the unmanned boat was recovered, the plankton net was lifted out of the water surface, and the excess water sample was filtered out, retaining 10-20 ml of the water sample. The bottom outlet of the plankton net was then moved into the sampling bottle, and the water outlet piston switch at the bottom was opened to collect the water sample. The above process was repeated many times to obtain a cumulative water sample of not less than 500 ml. Lugol's iodine solution was added to the water sample for fixation, and the water sample was refrigerated and stored away from light.

[0008] As a preferred improvement, the water quality data include water temperature, dissolved oxygen content, pH value, conductivity, ammonia nitrogen content, total phosphorus content, total nitrogen content, permanganate index and chlorophyll a content.

[0009] As a preferred improvement, in step S1, the concentrations of different types of algae are detected by the following method: 0.1 ml of the obtained water sample is taken and shaken thoroughly, placed in a 0.1 ml counting frame, and microscopically examined using an AlgaeAC plankton automatic classification counter to determine the algae types in the water sample and the concentrations of different types of algae.

[0010] As a preferred improvement, the process of enhancing the original data set using the synthetic minority class oversampling method specifically includes the following steps: Step S211: From the original data set Extract minority class sample subsets respectively and the majority class sample subset , for the minority class sample subset Any minority class sample in , calculate the Euclidean distance between it and other samples, and select K nearest neighbors based on the size of the Euclidean distance to form a neighbor set , expressed as: ; Where, express The K nearest neighbors of Step S212: From the neighbor set Select any neighbor sample , using the interpolation mechanism in the sample and neighbor samples Generate a new data point between , expressed as: ; Where, Represents a random number in uniform distribution U(0,1); Step S213: New sample Add a subset of minority class samples , to achieve the minority class sample subset Enhanced, traversing the minority class sample subset All samples in form the enhanced minority class sample subset ; Step S214: The enhanced minority class sample subset and the majority class sample subset Merge to form the first data set.

[0011] As a preferred improvement, the construction process of the random forest model specifically includes the following steps: Step S221 , selecting multiple samples from the first data set by bootstrapping with replacement as an independent training subset; In step S222, a decision tree is generated using each training subset. For each node in the decision tree, the CART algorithm is used to split the node. The Gini index is used as an indicator to evaluate the purity of the node to find the best splitting feature and splitting point for each split. The above steps are repeated until the preset stopping condition is reached. The Gini index calculation formula is: Where, represents the training subset; Indicates the number of categories in the training subset; Indicates the The proportion of class samples in the training subset; Step S223, repeat the process of building a decision tree, and finally form a random forest model consisting of k decision trees.

[0012] As a preferred improvement, step S3 specifically includes the following steps: Step S31, using a multifractal detrended coupled fluctuation analysis method to analyze the coupling correlation of competitive growth of different types of algae, and obtain the generalized Hurst index of the original CDFA curve; Step S32, randomly shuffling the concentration sequence of algae over time in a manner that the concentration value remains unchanged but the time sequence is shuffled, and reusing the multifractal detrended coupled fluctuation analysis method to analyze the coupling correlation of the competitive growth of different types of algae to obtain the generalized Hurst index of the random CDFA curve; Step S33: quantify the difference between the original CDFA curve and the random CDFA curve by chi-square statistics to obtain the competitive impact factor of the target algae on the overall algae production. i , expressed as: ; Where, Indicates target algae competitive factors affecting overall algal production; Indicates that only the target algae in the original sequence Chi-square statistic obtained by randomization; represents the chi-square statistic obtained by randomizing all algae in the original sequence; and Target algae The impact of competition on overall algal production q The order generalized Hurst exponent and generalized standard deviation; and Represents target algae The q-order generalized Hurst exponent and generalized standard deviation of the CDFA curve after the concentration sequence is randomized; Represents the q-order generalized Hurst exponent of the CDFA curve after all algae concentration series are randomized.

[0013] Step S34, calculating the ecological fitness of different types of algae , the calculation process is expressed as: .

[0014] As a preferred improvement, the multi-scale deep learning prediction model includes a change point detection module, a multi-scale feature extraction module, an adaptive feature fusion module and an output module. The change point detection module uses an iterative cumulative sum of squares algorithm to identify mutation points in the second data set, thereby enhancing the model's ability to capture abnormal changes in data in the second data set; the multi-scale feature extraction module extracts multi-scale features of the second data set through a multi-scale convolution kernel, and captures spatiotemporal dependencies in combination with a gating mechanism; the adaptive feature fusion module dynamically fuses multi-scale spatiotemporal features based on a self-attention mechanism to improve feature utilization efficiency; the output module uses a fully connected layer to output predicted values ​​for the concentration of key dominant algae.

[0015] As a preferred improvement, step S2 further includes the following steps: studying the biodiversity in the water sample by measuring the concentration of different types of algae, wherein the biodiversity is characterized by richness index, diversity index, evenness index, and dominance index, wherein: Richness index is used to measure the number of species in a community. The larger the value, the higher the species richness. The calculation process is expressed as: Where, represents the total number of species in the algal community; represents the total number of individuals in the algal community; The diversity index takes into account the richness and evenness of species. The calculation process is expressed as: Where, Indicates the The ratio of the number of algae individuals of a certain type to the total number of algae individuals in the community; Evenness index, used to assess the balance of species distribution in algal communities, The calculation process is expressed as: Where, is the maximum possible diversity index, that is, the index value when the number of individuals of all species is equal; Dominance index is used to measure the dominance of dominant species in a community. The calculation process is expressed as: .

[0016] A system for executing the above-mentioned algal bloom prediction method based on ecological niche fitness, comprising: The sampling component is used to collect water samples from the target water area at the same time interval, detect and analyze the water samples to obtain water quality data and the concentration of different types of algae at different time nodes, calculate the first-order difference of the concentration of each type of algae at two adjacent time nodes, and traverse all time nodes to form a concentration fluctuation sequence for each type of algae; The control platform is equipped with a SMOTE-RF model, a multi-fractal detrended coupled fluctuation analysis model and a multi-scale deep learning prediction model. The SMOTE-RF model uses the concentration fluctuation series of all types of algae and water quality data as the original data set, and uses the synthetic minority class oversampling method to enhance the original data set to form a first data set. The first data set is used to construct a random forest model to learn the mapping relationship between water quality data and the concentration of different types of algae. In the prediction stage, the water quality data of the target water area at the current time node is collected and input into the random forest model, and the concentration fluctuation prediction value of different types of algae is output. Combined with the algae concentration at the current time node, the concentration prediction value of different types of algae at the next time node is converted, and the algae are sorted in descending order according to the concentration prediction value, and the top 80% of the algae are screened. The multi-fractal detrended coupled fluctuation analysis model analyzes the predicted concentration value of the key dominant algae through the multi-fractal detrended coupled fluctuation analysis method, and quantitatively analyzes the niche fitness of different algae competitive growth through the intensity change of the coupling correlation of different types of algae concentrations; the multi-scale deep learning prediction model constructs a second data set with the water quality data of the current time node, the niche fitness of different algae competitive growth and the concentration of different types of algae. The second data set is used to train the multi-scale deep learning prediction model to learn the mapping relationship between water quality data and the niche fitness of different algae competitive growth to the concentration of different types of algae. After the training is completed, the water quality data obtained by the test and the niche fitness of different algae competitive growth are used as input to output the predicted value of the key dominant algae concentration; Monitoring component, used to monitor water quality data of target water areas in real time; The database is used to store the input and output data of all models during the system construction phase and the daily operation phase of the system.

[0017] The beneficial effects of the present invention are: (1) Based on the high-time-resolution water sampling method, the data on water sample information fluctuations over time are obtained, which can effectively reflect the changes in water quality data and algae concentration at different time points; (2) The multifractal detrended coupled fluctuation analysis method was used to quantitatively analyze the niche fitness of different algae competitive growth, reflecting the competitive growth relationship of different types of algae under natural conditions; (3) Using the water quality data monitored in real time by automatic monitoring stations and the ecological niche fitness of different algae competitive growth as input, the concentration of key dominant algae can be predicted in real time, providing effective theoretical support for early warning of algal blooms. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which: Figure 1 A diagram showing the prediction results of cryptoalgae by the water bloom prediction system provided by the present invention; Figure 2 Graph showing the prediction results of green algae by the water bloom prediction system provided by the present invention Figure 3 Graph showing the prediction results of diatoms by the algal bloom prediction system provided by the present invention Figure 4 A diagram showing the prediction results of cyanobacteria by the water bloom prediction system provided by the present invention. DETAILED DESCRIPTION

[0019] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] This embodiment provides a method for predicting algal blooms based on ecological niche fitness, comprising the following steps: In step S1, water samples are collected from the target water area at the same time intervals, and the water samples are tested and analyzed to obtain water quality data and the concentrations of different types of algae at different time nodes. The first-order difference of the concentration of each type of algae at two adjacent time nodes is calculated, and the concentration fluctuation sequence of each type of algae is formed by traversing all time nodes.

[0021] The water sample collection process is as follows: The plankton net with a mesh diameter of 0.064 mm was mounted on an unmanned boat in an area larger than 50m 2 The plankton net was collected in the target waters. During the collection process, the water outlet piston switch at the bottom of the plankton net was closed, and the plankton net was made to do an "∞"-shaped reciprocating motion at a speed of 20-30 cm / s at a depth of 0.5 m below the water surface. After slowly dragging for 1-3 minutes, the unmanned boat was recovered, the plankton net was lifted out of the water surface, and the excess water sample was filtered out, retaining 10-20 ml of the water sample. The bottom outlet of the plankton net was then moved into the sampling bottle, and the water outlet piston switch at the bottom was opened to collect the water sample. The above process was repeated many times to obtain a cumulative water sample of not less than 500 ml. Lugol's iodine solution was added to the water sample for fixation, and the water sample was refrigerated and stored away from light.

[0022] For the same target water area, a series of unmanned boats are started at the same time interval to collect water samples. In this embodiment, the selected time interval is 5 minutes, that is, an unmanned boat is started every 5 minutes. The sequence of unmanned boats is recorded as A1, A2, A3, ..., A n The sampling time of each unmanned boat is 1 hour. After the collection task is completed, the unmanned boat can be put into use again to perform the next collection task, realizing the recycling of the unmanned boat.

[0023] Water quality data include water temperature, dissolved oxygen, pH value, conductivity, ammonia nitrogen content, total phosphorus content, total nitrogen content, permanganate index and chlorophyll a content.

[0024] Water temperature was measured using a deep-water thermometer; dissolved oxygen, pH, and conductivity were measured using a ProDSS portable multi-parameter water quality meter; transparency was measured using the Secchi disk method; ammonia nitrogen content was measured using the Nessler reagent spectrophotometric method specified in HJ 535-2009; total phosphorus content was measured using the ammonium molybdate spectrophotometric method specified in GB 11893-1989; total nitrogen content was measured using the alkaline potassium persulfate digestion ultraviolet spectrophotometric method specified in HJ 636-2012; the permanganate index was measured using the method specified in GB 11892-1989; and chlorophyll a was measured spectrophotometrically.

[0025] The concentrations of different types of algae were determined by taking 0.1 ml of the water sample, shaking it thoroughly, placing it in a 0.1 ml counting frame, and examining it under a microscope using an AlgaeAC plankton automatic classification counter to determine the type of algae in the water sample and the concentrations of the different types of algae.

[0026] The concentration sequence of algae is B1, B2, B3, ..., B n For example, if the sampling time for unmanned vessel A1 is from 08:00 to 09:00 on January 1, the algae concentration data obtained by this unmanned vessel is recorded as B1; the sampling time for unmanned vessel A2 is from 08:05 to 09:05 on January 1, the algae concentration data obtained by this unmanned vessel is recorded as B2; the sampling time for unmanned vessel A3 is from 08:10 to 09:10 on January 1, the algae concentration data obtained by this unmanned vessel is recorded as B3, and so on. It should be noted that each set of algae concentration data includes concentration data of multiple different types of algae.

[0027] The first-order difference of the concentration of each type of algae at two adjacent time nodes can reflect the concentration fluctuation of this type of algae at a time resolution (5 minutes). By traversing all time nodes, the density fluctuation of different types of algae over time can be effectively reflected, and the growth trend of different types of algae can be more clearly reflected.

[0028] Step S2, using the concentration fluctuation series of all types of algae and water quality data as the original data set, adopting the synthetic minority class oversampling method to enhance the original data set to form a first data set, using the first data set to construct a random forest model, and learning the mapping relationship between water quality data and the concentration of different types of algae; in the prediction stage, collecting the water quality data of the target water area at the current time node and inputting it into the random forest model, outputting the concentration fluctuation prediction value of different types of algae, combining it with the algae concentration at the current time node, and converting it into the concentration prediction value of different types of algae at the next time node, sorting the algae in descending order according to the concentration prediction value, and screening the top 80% of the algae as the key dominant algae.

[0029] Since algae concentration data is completely dependent on actual collection, there is a data imbalance problem in the original dataset. To solve this problem, the present invention adopts the synthetic minority oversampling method (SMOTE) to enhance the algae concentration data. By generating new minority class samples in the feature space, the original data is balanced, thereby effectively improving the model's recognition ability for minority class samples.

[0030] Specifically, the process of enhancing the original dataset using the synthetic minority class oversampling method includes the following steps: Step S211: From the original data set Extract minority class sample subsets respectively and the majority class sample subset , for the minority class sample subset Any minority class sample in , calculate the Euclidean distance between it and other samples, and select K nearest neighbors based on the size of the Euclidean distance to form a neighbor set , expressed as: ; Where, express The K nearest neighbors of Step S212: From the neighbor set Select any neighbor sample , using the interpolation mechanism in the sample and neighbor samples Generate a new data point between , expressed as: ; Where, Represents a random number in uniform distribution U(0,1); Step S213: New sample Add a subset of minority class samples , to achieve the minority class sample subset Enhanced, traversing the minority class sample subset All samples in form the enhanced minority class sample subset ; Step S214: The enhanced minority class sample subset and the majority class sample subset Merge to form the first data set.

[0031] The enhanced minority class samples will form denser local regions in the original feature space, improving the classifier's ability to learn from these samples and significantly reducing the model's sensitivity to majority class shifts. The resulting first dataset maintains a balance between data types. Furthermore, the parameters in the SMOTE algorithm (K value and sample generation ratio) should be appropriately adjusted and optimized based on the specific characteristics and actual needs of the dataset to ensure the quality of the synthesized samples and provide a more balanced data foundation for subsequent model training.

[0032] The Random Forest (RF) model is an algorithm based on ensemble learning. By constructing multiple decision trees and integrating their predictions, it significantly improves the accuracy and robustness of classification and regression tasks. The dataset is divided into a training set and a test set in an 8:2 ratio. The RF model is constructed using the training set. The RF model construction process includes the following steps: Step S221 , selecting multiple samples from the first data set by bootstrapping with replacement as an independent training subset; In step S222, a decision tree is generated using each training subset. For each node in the decision tree, the CART algorithm is used to split the node. The Gini index is used as an indicator to evaluate the purity of the node to find the best splitting feature and splitting point for each split. The above steps are repeated until the preset stopping condition is reached. The Gini index calculation formula is: Where, represents the training subset; Indicates the number of categories in the training subset; Indicates the The proportion of class samples in the training subset; Step S223, repeat the process of building a decision tree, and finally form a random forest model consisting of k decision trees.

[0033] While training the random forest model, we perform hyperparameter optimization, adjusting hyperparameters such as the number of decision trees, maximum depth, and minimum number of samples to improve model performance. We use cross-validation to evaluate performance on the validation set and determine the most optimal hyperparameter combination for the model.

[0034] After the model is built, the test set is used to comprehensively evaluate the RF model and calculate indicators such as accuracy, recall, and F1-Score. Accuracy is used to measure the proportion of samples correctly predicted by the model to the total samples, reflecting the overall prediction accuracy of the model; Recall represents the proportion of positive samples correctly predicted by the model to the actual positive samples, emphasizing the model's ability to identify positive samples; F1-Score is the harmonic mean of accuracy and recall, which comprehensively considers the relationship between the two and can more comprehensively reflect the performance balance of the model under the distribution of different categories of samples. Where TP and FN represent the number of positive samples predicted as positive samples and negative samples respectively; FP and TN represent the number of negative samples predicted as positive samples and negative samples respectively; It indicates the proportion of samples that are actually positive among the samples predicted by the model to be positive.

[0035] In each decision tree, each time a feature is used to split, the reduction in the Gini index resulting from that split is calculated. The reductions in the Gini index resulting from all splits are summed to obtain the feature's importance score within that decision tree. The feature importance scores across all decision trees are averaged to determine the feature's importance within the entire random forest model. All features are ranked based on their calculated importance scores. Features with higher importance scores are selected as key features for subsequent analysis.

[0036] Assessing the relative importance of input variables helps understand the model's decision-making mechanism. To explore the impact of different water quality data on the prediction of algal community structure, feature importance analysis was used to systematically assess the importance of input variables to the final prediction results. By quantifying the weight coefficients of each variable, the contribution of each variable to the prediction results was revealed, thus helping to understand the model's decision-making mechanism.

[0037] Once the SMOTE-RF model is built, monitoring stations can be set up in the target waters to monitor water quality in real time. This water quality data is then fed into the SMOTE-RF model, which outputs predicted algae concentrations at different times. After the predicted concentrations of different algae types are determined, the algae are ranked from highest to lowest according to their predicted concentrations, with the top 80% of the ranked algae identified as the key dominant algae.

[0038] In addition, the concentration of different types of algae can also be used to study the biodiversity in water samples. In order to more comprehensively reflect the community characteristics of algae in water samples, the present invention defines the following index: Richness index is used to measure the number of species in a community. The larger the value, the higher the species richness. The calculation process is expressed as: Where, represents the total number of species in the algal community; represents the total number of individuals in the algal community; The diversity index takes into account the richness and evenness of species. It is affected by both the number of species and the uniformity of the distribution of the number of individuals of each species. The calculation process is expressed as: Where, Indicates the The ratio of the number of algae individuals of a certain type to the total number of algae individuals in the community; The evenness index is used to evaluate the degree of balance in the distribution of species in algal communities and is closely related to the diversity index. The higher the evenness index, the more balanced the distribution of the number of individuals of each species, thereby increasing the value of the diversity index. The calculation process is expressed as: Where, is the maximum possible diversity index, that is, the index value when the number of individuals of all species is equal; The dominance index shows an opposite trend to other indices and is mainly used to measure the dominance of dominant species in a community. When certain dominant species in a community have an absolute advantage in number, the value of this index is high, which means that the overall diversity of the community is low and is negatively correlated with the diversity index. In addition, the presence of dominant species may lead to an uneven distribution of species, thereby reducing the value of the evenness index. The calculation process is expressed as: .

[0039] Since different indices have different sensitivities to changes in community structure, this study comprehensively considered these four indices and analyzed the algal community structure from multiple dimensions in order to comprehensively and accurately evaluate the diversity characteristics of the algal community ecosystem.

[0040] Step S3: Analyze the predicted concentration values ​​of key dominant algae using a multifractal detrended coupled fluctuation analysis method, and quantitatively analyze the niche fitness of different algae competitive growth by comparing the intensity changes of the coupled correlations of different types of algae concentrations.

[0041] Under the influence of complex aquatic environmental factors, different algae species compete with each other for and consume water nutrients, resulting in nonlinear coupling relationships between their growth. Algal growth is reflected in their concentration: higher concentrations indicate better growth and a competitive advantage, while lower concentrations indicate slower growth and a competitive disadvantage. Therefore, the niche fitness of different algal species in competitive growth can be quantitatively analyzed by measuring the strength of the coupling correlations at different algal concentrations.

[0042] Step S3 specifically includes the following steps: Step S31, using a multifractal detrended coupled fluctuation analysis method to analyze the coupling correlation of competitive growth of different types of algae, and obtain the generalized Hurst index of the original CDFA curve; Step S32, randomly shuffling the concentration sequence of algae over time in a manner that the concentration value remains unchanged but the time sequence is shuffled, and reusing the multifractal detrended coupled fluctuation analysis method to analyze the coupling correlation of the competitive growth of different types of algae to obtain the generalized Hurst index of the random CDFA curve; Step S33: quantify the difference between the original CDFA curve and the random CDFA curve by chi-square statistics to obtain the competitive impact factor of the target algae on the overall algae production. i , expressed as: ; Where, Indicates target algae competitive factors affecting overall algal production; Indicates that only the target algae in the original sequence Chi-square statistic obtained by randomization; represents the chi-square statistic obtained by randomizing all algae in the original sequence; and Represents target algae The impact of competition on overall algal production q The order generalized Hurst exponent and generalized standard deviation; and Represents target algae The q-order generalized Hurst exponent and generalized standard deviation of the CDFA curve after the concentration sequence is randomized; Represents the q-order generalized Hurst exponent of the CDFA curve after all algae concentration series are randomized.

[0043] Step S34, calculating the ecological fitness of different types of algae , the calculation process is expressed as: .

[0044] Step S4, constructing a multi-scale deep learning prediction model, using the water quality data at the current time node, the niche fitness of different algae competitive growth and the concentration of different types of algae to jointly construct a second data set, using the second data set to train the multi-scale deep learning prediction model, learning the mapping relationship between water quality data and the niche fitness of different algae competitive growth to the concentration of different types of algae, after the training is completed, using the detected water quality data and the niche fitness of different algae competitive growth as input, outputting the predicted value of the concentration of key dominant algae.

[0045] The multi-scale deep learning prediction model includes a change point detection module, a multi-scale feature extraction module, an adaptive feature fusion module and an output module. The change point detection module uses an iterative cumulative sum of squares algorithm to identify mutation points in the second data set, thereby enhancing the model's ability to capture abnormal changes in data in the second data set. The multi-scale feature extraction module extracts multi-scale features of the second data set through multi-scale convolution kernels (such as 1×4, 1×8, 1×12, 1×24, etc.), and captures spatiotemporal dependencies in combination with a gating mechanism. The adaptive feature fusion module dynamically fuses multi-scale spatiotemporal features based on a self-attention mechanism to improve feature utilization efficiency. The output module uses a fully connected layer to output predicted values ​​for the concentration of key dominant algae.

[0046] This embodiment further provides a system for executing the above-mentioned algal bloom prediction method based on ecological niche fitness, comprising: The sampling component is used to collect water samples from the target water area at the same time interval, detect and analyze the water samples to obtain water quality data and the concentration of different types of algae at different time nodes, calculate the first-order difference of the concentration of each type of algae at two adjacent time nodes, and traverse all time nodes to form a concentration fluctuation sequence for each type of algae; The control platform is equipped with a SMOTE-RF model, a multi-fractal detrended coupled fluctuation analysis model and a multi-scale deep learning prediction model. The SMOTE-RF model uses the concentration fluctuation series of all types of algae and water quality data as the original data set, and uses the synthetic minority class oversampling method to enhance the original data set to form a first data set. The first data set is used to construct a random forest model to learn the mapping relationship between water quality data and the concentration of different types of algae. In the prediction stage, the water quality data of the target water area at the current time node is collected and input into the random forest model, and the concentration fluctuation prediction value of different types of algae is output. Combined with the algae concentration at the current time node, the concentration prediction value of different types of algae at the next time node is converted, and the algae are sorted in descending order according to the concentration prediction value, and the top 80% of the algae are screened. The multi-fractal detrended coupled fluctuation analysis model analyzes the predicted concentration value of the key dominant algae through the multi-fractal detrended coupled fluctuation analysis method, and quantitatively analyzes the niche fitness of different algae competitive growth through the intensity change of the coupling correlation of different types of algae concentrations; the multi-scale deep learning prediction model constructs a second data set with the water quality data of the current time node, the niche fitness of different algae competitive growth and the concentration of different types of algae. The second data set is used to train the multi-scale deep learning prediction model to learn the mapping relationship between water quality data and the niche fitness of different algae competitive growth to the concentration of different types of algae. After the training is completed, the water quality data obtained by the test and the niche fitness of different algae competitive growth are used as input to output the predicted value of the key dominant algae concentration; Monitoring component, used to monitor water quality data of target water areas in real time; The database is used to store the input and output data of all models during the system construction phase and the daily operation phase of the system.

[0047] The algal bloom prediction system based on niche fitness provided in this embodiment includes two stages: First, the system construction stage During the system construction phase, the sampling component is used to generate one or more continuous concentration fluctuation sequences of algae. The concentration fluctuation sequences of all types of algae and water quality data are used as the original data set. The synthetic minority class oversampling method is used to enhance the original data set to form a first data set, and the first data set is used to construct a random forest model; then a multifractal detrended coupled fluctuation analysis model is constructed, and the concentration prediction values ​​of key dominant algae are analyzed by the multifractal detrended coupled fluctuation analysis method, and the ecological niche fitness of different algae competitive growth is quantitatively analyzed by the intensity change of the coupling correlation of different types of algae concentrations; finally, a multi-scale deep learning prediction model is constructed, and a second data set is jointly constructed with the water quality data at the current time node, the ecological niche fitness of different algae competitive growth, and the concentration of different types of algae, and the second data set is used to train the multi-scale deep learning prediction model.

[0048] Second, the daily operation stage of the system During the system's daily operations, the SMOTE-RF model and a multi-scale deep learning prediction model are used to predict algal blooms. Specifically, the system first uses the monitoring component to monitor the target water area's water quality data in real time. This data is then fed into a random forest model, which outputs predicted concentrations of different algae types. The algae are then ranked from highest to lowest by predicted concentration, with the top 80% of the ranked algae selected as the key dominant algae. The system then quantitatively analyzes the ecological niche fitness of different algae for competitive growth by analyzing the intensity of the coupled correlation between the concentrations of different algae types. Using the water quality data obtained by the monitoring component and the ecological niche fitness of different algae for competitive growth as input, the system outputs predicted concentrations of key dominant algae.

[0049] The monitoring component is deployed in the target water area. The monitoring component uses various types of sensors to monitor water quality data in real time, and then uploads it to the control platform. After the control platform processes the monitored water quality data, it outputs the prediction result. When the prediction result exceeds the set threshold, an initial alarm is issued to remind relevant staff to issue an early warning.

[0050] It should be noted that in order to achieve the effect of data transmission, communication is required between the monitoring component and the control platform. The communication method adopts the existing technology in this field, preferably wireless communication to facilitate the remote deployment of the control platform, such as 4G communication, 5G communication, etc. In addition, in order to realize the utilization of monitoring data, steps such as digital-to-analog conversion are required, which also adopt the existing technology in this field.

[0051] The monitoring data uploaded by the monitoring component is stored in a database, which is also deployed on the control platform. As the monitoring time progresses, the data stored in the database is updated in real time. It is understood that the data stored in the database can also be called to generate various statistical charts, which can be done using conventional techniques in the field and will not be described in detail in this embodiment.

[0052] Example 1 This example takes a target water area of ​​100 m2 in the Qiongjiang River Basin of Suining City as the research and development object. From February 27 to March 22, 2025, sampling work was carried out in the target water area multiple times. A total of 10 groups of algae concentration fluctuation sequences with a time resolution of 5 minutes were obtained, totaling 14,400 samples for training related models.

[0053] Then, the water quality data collected by the monitoring component from March 27 to April 2, 2025, was used to predict algal blooms and evaluate the effectiveness of the algal bloom prediction method and system of the present invention. The evaluation indicators include R 2 , RMSE and MAE, the evaluation results are as follows Figure 1-Figure 4 As shown. Figure 1-4 It can be seen that: The prediction system provided by the present invention has a good fitting ability for the outbreak period of key algae concentrations and performs well in predicting key algae concentrations. All indicators have reached a high level. 2 The results were 0.987, 0.923, 0.949 and 0.917 respectively, and the RMSE were 2.928×10 6 , 2.375×10 6 , 2.049×10 6 and 2.656×10 6 , MAE were 2.479×10 6 , 2.145×10 6 , 1.735×10 6 and 1.959×10 6 .

[0054] Among them, compared with the benchmark models (XGBoost, RF, SVD, BPNN, LSTM and GRU), the MSDN model performs better in predicting cryptoalgae in R 2 , RMSE and MAE were improved by 8.968%, 10.477% and 9.33% respectively; in terms of predicting green algae, the MSDN model was superior to the R 2 , RMSE and MAE were improved by 4.145%, 9.461% and 10.56% respectively; in terms of predicting diatoms, the MSDN model was superior to the R 2, RMSE and MAE were improved by 7.751%, 10.885% and 8.214% respectively; in terms of predicting cyanobacteria, the MSDN model was superior to the R 2 , RMSE and MAE are improved by 9.804%, 11.738% and 25.739% respectively, indicating that the performance of the model provided by the present invention is better than other models.

[0055] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. A method for predicting algal blooms based on niche fitness, characterized in that: The steps include: Step S1: Collect water samples from the target water area at the same time interval, analyze the water samples to obtain water quality data and the concentrations of different types of algae at different time nodes, calculate the first-order difference of the concentration of each type of algae at two adjacent time nodes, and traverse all time nodes to form a concentration fluctuation sequence for each type of algae; Step S2: using the concentration fluctuation series of all types of algae and water quality data as the original dataset, enhancing the original dataset using a synthetic minority class oversampling method to form a first dataset, constructing a random forest model using the first dataset to learn the mapping relationship between the water quality data and the concentrations of different types of algae; In the prediction stage, water quality data of the target water area at the current time node is collected and input into the random forest model, which outputs the concentration fluctuation prediction value of different types of algae. Combined with the algae concentration at the current time node, the concentration prediction value of different types of algae at the next time node is converted. The algae are ranked in descending order according to the concentration prediction value, and the top 80% of the algae are selected as the key dominant algae. Step S3, analyzing the predicted concentration values ​​of key dominant algae using a multifractal detrended coupled fluctuation analysis method, and quantitatively analyzing the niche fitness of different algae competitive growth by analyzing the intensity changes in the coupled correlation of different types of algae concentrations; Step S4, constructing a multi-scale deep learning prediction model, using the water quality data at the current time node, the niche fitness of different algae competitive growth and the concentration of different types of algae to jointly construct a second data set, using the second data set to train the multi-scale deep learning prediction model, learning the mapping relationship between water quality data and the niche fitness of different algae competitive growth to the concentration of different types of algae, after the training is completed, using the detected water quality data and the niche fitness of different algae competitive growth as input, outputting the predicted value of the concentration of key dominant algae.

2. The algal bloom prediction method based on niche adaptability according to claim 1, wherein The water sample collection process is as follows: The plankton net with a mesh diameter of 0.064mm was mounted on an unmanned boat in an area larger than 50m 2 The plankton net was collected in the target waters. During the collection process, the water outlet piston switch at the bottom of the plankton net was closed, and the plankton net was made to do an "∞"-shaped reciprocating motion at a speed of 20-30 cm / s at a depth of 0.5 m below the water surface. After slowly dragging for 1-3 minutes, the unmanned boat was recovered, the plankton net was lifted out of the water surface, and the excess water sample was filtered out, retaining 10-20 ml of the water sample. The bottom outlet of the plankton net was then moved into the sampling bottle, and the water outlet piston switch at the bottom was opened to collect the water sample. The above process was repeated many times to obtain a cumulative water sample of not less than 500 ml. Lugol's iodine solution was added to the water sample for fixation, and the water sample was refrigerated and stored away from light.

3. The algal bloom prediction method based on niche adaptability according to claim 1, wherein Water quality data include water temperature, dissolved oxygen, pH value, conductivity, ammonia nitrogen content, total phosphorus content, total nitrogen content, permanganate index and chlorophyll a content.

4. The algal bloom prediction method based on niche adaptability according to claim 1, wherein In step S1, the concentration of different types of algae is detected by the following method: 0.1 ml of the obtained water sample is taken and shaken thoroughly, placed in a 0.1 ml counting frame, and microscopically examined using an AlgaeAC plankton automatic classification counter to determine the type of algae in the water sample and the concentration of different types of algae.

5. The algal bloom prediction method based on niche adaptability according to claim 1, wherein The process of enhancing the original dataset using the synthetic minority class oversampling method specifically includes the following steps: Step S211: From the original data set Extract minority class sample subsets respectively and the majority class sample subset , for the minority class sample subset Any minority class sample in , calculate the Euclidean distance between it and other samples, and select K nearest neighbors based on the size of the Euclidean distance to form a neighbor set , expressed as: ; Where, express The K nearest neighbors of Step S212: From the neighbor set Select any neighbor sample , using the interpolation mechanism in the sample and neighbor samples Generate a new data point between , expressed as: ; Where, Represents a random number in uniform distribution U(0,1); Step S213: New sample Add a subset of minority class samples , to achieve the minority class sample subset Enhanced, traversing the minority class sample subset All samples in form the enhanced minority class sample subset ; Step S214: The enhanced minority class sample subset and the majority class sample subset Merge to form the first data set.

6. The algal bloom prediction method based on niche adaptability according to claim 1, wherein The construction process of the random forest model specifically includes the following steps: Step S221 , selecting multiple samples from the first data set by bootstrapping with replacement as an independent training subset; In step S222, a decision tree is generated using each training subset. For each node in the decision tree, the CART algorithm is used to split the node. The Gini index is used as an indicator to evaluate the purity of the node to find the best splitting feature and splitting point for each split. The above steps are repeated until the preset stopping condition is reached. The Gini index calculation formula is: Where, represents the training subset; Indicates the number of categories in the training subset; Indicates the The proportion of class samples in the training subset; Step S223, repeat the process of building a decision tree, and finally form a random forest model consisting of k decision trees.

7. The algal bloom prediction method based on niche adaptability according to claim 1, wherein Step S3 specifically includes the following steps: Step S31, using a multifractal detrended coupled fluctuation analysis method to analyze the coupling correlation of competitive growth of different types of algae, and obtain the generalized Hurst index of the original CDFA curve; Step S32, randomly shuffling the concentration sequence of algae over time in a manner that the concentration value remains unchanged but the time sequence is shuffled, and reusing the multifractal detrended coupled fluctuation analysis method to analyze the coupling correlation of the competitive growth of different types of algae to obtain the generalized Hurst index of the random CDFA curve; Step S33: quantify the difference between the original CDFA curve and the random CDFA curve by chi-square statistics to obtain the competitive impact factor of the target algae on the overall algae production. i , expressed as: ; Where, Indicates target algae competitive factors affecting overall algal production; Indicates that only the target algae in the original sequence Chi-square statistic obtained by randomization; represents the chi-square statistic obtained by randomizing all algae in the original sequence; and Target algae The impact of competition on overall algal production q The order generalized Hurst exponent and generalized standard deviation; and Target algae The q-order generalized Hurst exponent and generalized standard deviation of the CDFA curve after the concentration sequence is randomized; Represents the q-order generalized Hurst exponent of the CDFA curve after the concentration series of all algae are randomized. Step S34, calculating the ecological fitness of different types of algae , the calculation process is expressed as: 。 8. The algal bloom prediction method based on niche fitness according to claim 1, wherein The multi-scale deep learning prediction model includes a change point detection module, a multi-scale feature extraction module, an adaptive feature fusion module and an output module. The change point detection module uses an iterative cumulative sum of squares algorithm to identify mutation points in the second data set, thereby enhancing the model's ability to capture abnormal changes in data in the second data set; the multi-scale feature extraction module extracts multi-scale features of the second data set through a multi-scale convolution kernel, and captures spatiotemporal dependencies in combination with a gating mechanism; the adaptive feature fusion module dynamically fuses multi-scale spatiotemporal features based on a self-attention mechanism to improve feature utilization efficiency; the output module uses a fully connected layer to output predicted values ​​for the concentration of key dominant algae.

9. The algal bloom prediction method based on niche fitness according to claim 1, wherein After step S2, the following steps are also included: studying the biodiversity in the water sample by measuring the concentration of different types of algae, wherein the biodiversity is characterized by richness index, diversity index, evenness index and dominance index, wherein: Richness index is used to measure the number of species in a community. The larger the value, the higher the species richness. The calculation process is expressed as: Where, represents the total number of species in the algal community; represents the total number of individuals in the algal community; The diversity index takes into account the richness and evenness of species. The calculation process is expressed as: Where, Indicates the The ratio of the number of algae individuals of a certain type to the total number of algae individuals in the community; Evenness index, used to assess the balance of species distribution in algal communities, The calculation process is expressed as: Where, is the maximum possible diversity index, that is, the index value when the number of individuals of all species is equal; Dominance index is used to measure the dominance of dominant species in a community. The calculation process is expressed as: 。 10. A system for executing the algal bloom prediction method based on ecological niche fitness according to any one of claims 1 to 9, characterized in that: include: The sampling component is used to collect water samples from the target water area at the same time interval, detect and analyze the water samples to obtain water quality data and the concentration of different types of algae at different time nodes, calculate the first-order difference of the concentration of each type of algae at two adjacent time nodes, and traverse all time nodes to form a concentration fluctuation sequence for each type of algae; The control platform is equipped with a SMOTE-RF model, a multi-fractal detrended coupled fluctuation analysis model and a multi-scale deep learning prediction model. The SMOTE-RF model uses the concentration fluctuation series of all types of algae and water quality data as the original data set, and uses the synthetic minority class oversampling method to enhance the original data set to form a first data set. The first data set is used to construct a random forest model to learn the mapping relationship between water quality data and the concentration of different types of algae. In the prediction stage, the water quality data of the target water area at the current time node is collected and input into the random forest model, and the concentration fluctuation prediction value of different types of algae is output. Combined with the algae concentration at the current time node, the concentration prediction value of different types of algae at the next time node is converted, and the algae are sorted in descending order according to the concentration prediction value, and the top 80% of the algae are screened. The multi-fractal detrended coupled fluctuation analysis model analyzes the predicted concentration value of the key dominant algae through the multi-fractal detrended coupled fluctuation analysis method, and quantitatively analyzes the niche fitness of different algae competitive growth through the intensity change of the coupling correlation of different types of algae concentrations; the multi-scale deep learning prediction model constructs a second data set with the water quality data of the current time node, the niche fitness of different algae competitive growth and the concentration of different types of algae. The second data set is used to train the multi-scale deep learning prediction model to learn the mapping relationship between water quality data and the niche fitness of different algae competitive growth to the concentration of different types of algae. After the training is completed, the water quality data obtained by the test and the niche fitness of different algae competitive growth are used as input to output the predicted value of the key dominant algae concentration; Monitoring component, used to monitor water quality data of target water areas in real time; The database is used to store the input and output data of all models during the system construction phase and the daily operation phase of the system.

Citation Information

Cited By

  • Algae community structure change prediction algorithm and system based on multi-source data fusion

    CN121502706A