A Reservoir Algal Bloom Prediction Method Based on Convergent Cross Mapping and Machine Learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]综上所述,现有技术在水库水华研究中普遍存在以下不足:一是难以在复杂非线性系统中准确识别水华发生的关键驱动因素及其因果关系;二是驱动因素识别结果与水华预测模型之间缺乏有效衔接,难以同时兼顾预测精度与机理解释能力
(1)采用收敛交叉映射算法分析水库水华暴发与各环境因子的因果关系,克服了传统相关分析方法难以区分相关性与因果性的问题,提高了水库水华暴发关键驱动因素识别的可靠性和准确性。
Smart Images

Figure CN122575503A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of reservoir algal bloom prediction, and in particular to a reservoir algal bloom outbreak prediction method, computer equipment, and storage medium based on convergent cross mapping and machine learning. Background Technology
[0002] Under the combined influence of global warming and human activities, eutrophication in reservoirs is becoming increasingly prominent, with algal blooms exhibiting a trend towards higher frequency, suddenness, and complexity. Reservoir algal blooms not only significantly reduce water transparency and cause dissolved oxygen deficits, but may also produce harmful substances such as odor-causing compounds and algal toxins, seriously threatening drinking water safety and the stability of reservoir ecosystems. Therefore, conducting analysis of reservoir algal bloom mechanisms and effectively predicting algal bloom occurrences are crucial technical issues that urgently need to be addressed in the field of water environment management and risk prevention.
[0003] Analyzing the mechanisms of algal blooms is fundamental to accurate prediction. Existing research largely relies on traditional statistical methods such as Pearson correlation analysis and multiple regression analysis, inferring bloom mechanisms by analyzing the correlation between algal biomass and environmental variables. These methods are simple to implement, computationally inexpensive, and can reflect the linear correlation characteristics between variables to some extent. However, reservoir ecosystems exhibit significant nonlinearity, strong coupling, and dynamic evolutionary characteristics. Algal growth is often driven by multiple factors, and these driving relationships may change over time. Traditional statistical methods struggle to distinguish between correlations and causal relationships and are susceptible to interference from collinearity and time lag effects, leading to insufficient stability and reliability in identifying key drivers of algal blooms.
[0004] Currently, reservoir algal bloom prediction models are mainly divided into two categories: mechanistic models and machine learning models. Mechanistic models explicitly describe reservoir hydrodynamic processes, nutrient cycling, and algal growth, possessing strong physical significance. However, their complex structure, numerous parameters, and high dependence on high-quality long-term monitoring data limit their widespread application in real-time prediction and operational settings. Machine learning models, on the other hand, can predict algal blooms by learning implicit patterns from historical data. They offer advantages such as flexible modeling and high prediction accuracy. However, most existing methods focus on the prediction results themselves, lacking the ability to explain the driving mechanisms of algal blooms. They struggle to clearly define the true driving role of different environmental factors in algal bloom occurrence, and their interpretability and generalization ability remain insufficient.
[0005] In summary, existing technologies generally have the following shortcomings in reservoir algal bloom research: First, it is difficult to accurately identify the key driving factors and their causal relationships in complex nonlinear systems; second, there is a lack of effective connection between the identification results of driving factors and algal bloom prediction models, making it difficult to simultaneously take into account both prediction accuracy and mechanism explanation capabilities. Summary of the Invention
[0006] To address the aforementioned problems in existing technologies, this invention provides a reservoir algal bloom prediction method based on convergent cross-mapping and machine learning. This method can identify key driving factors of reservoir algal blooms within a data-driven framework and, based on this, construct an algal bloom prediction model with a certain lead time, thereby improving the reliability and application value of algal bloom prediction.
[0007] The objective of this invention can be achieved through the following technical solutions: In a first aspect of the present invention, a method for predicting reservoir algal blooms based on convergent cross-mapping and machine learning is provided, comprising the following steps: Long-term monitoring data of the reservoir is acquired and preprocessed to obtain time series of environmental influencing factors and time series of algal biomass. The long-term monitoring data includes algal biomass data and environmental influencing factors. The algal biomass data includes chlorophyll a concentration or algal density. The environmental influencing factors include water temperature, sunshine duration, nutrient content and hydrodynamic conditions. The convergent cross-mapping algorithm is used to analyze the causal relationship between the algal biomass factor sequence and the environmental impact factor sequence. The cross-mapping skill value of algal biomass data to environmental impact factors is calculated. When the cross-mapping skill value converges, it is determined that the environmental impact factors have a causal relationship with the algal biomass data. The environmental impact factors with convergent cross-mapping skill values are selected as the driving factors of algal blooms. The selected driving factors of algal blooms are input into a preset algal bloom prediction model, which outputs the predicted value of algal biomass data in the reservoir within the foreseeable future period. When the predicted value exceeds a preset threshold, it is determined to be a high-risk state of algal bloom and triggers an algal bloom risk warning.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The convergent cross-mapping algorithm was used to analyze the causal relationship between reservoir algal blooms and various environmental factors, which overcame the problem that traditional correlation analysis methods could not distinguish between correlation and causation, and improved the reliability and accuracy of identifying key driving factors of reservoir algal blooms.
[0009] (2) Based on the machine learning model, the accurate prediction of algal blooms in reservoirs has been realized. The data acquisition and model building process is relatively simple, and it has strong versatility and portability. It can provide reliable technical support for algal bloom early warning and reservoir operation management.
[0010] (3) This invention effectively integrates algal bloom driving mechanism analysis with algal bloom prediction. The key algal bloom driving factors identified by the convergent cross-mapping algorithm (CCM) are used as independent variables input into the algal bloom prediction model, which not only improves the interpretability and generalization ability of the prediction model, but also realizes accurate prediction of reservoir algal blooms with a certain prediction period.
[0011] In some embodiments, the preprocessing steps include time-scale unification, missing value processing, outlier removal, and standardization of the acquired long-sequence monitoring data.
[0012] In some embodiments, the method for calculating the cross-mapping skill value includes reconstructing the attractor manifold to obtain the hysteresis coordinate vector, specifically including: Define the time series of environmental influencing factors Time series of algal biomass data The length of each element is L, the dimension of the reconstructed attractor manifold is E, and the sampling interval during the reconstruction of the attractor manifold is τ. The hysteresis coordinate vector of the reconstructed attractor manifold at time t is obtained as follows: ; ; in, Represents the attractor manifold after reconstruction. The lag coordinate vector, Represents the attractor manifold after reconstruction. The lagging coordinate vector, t∈[ ,L].
[0013] In some embodiments, the step of calculating the cross-mapping skill value from algal biomass data to the direction of environmental influencing factors, and determining whether there is a causal relationship between environmental influencing factors and algal biomass data, further includes: After reconstruction of the attractor manifold Find the lagging coordinate vectors of the same period. ,pass express The time indices of the E+1 nearest neighbors are used to determine the reconstructed attractor manifold. Lag coordinate vectors of the same period The nearest neighbor, according to indivual Local weighted average estimation of environmental impact factors Genetic environmental impact factors convergent cross estimate : ; in, Environmental influencing factors Simultaneous value, yes Rather than the attractor manifold after reconstruction Upper The weight of the distance between the nearest neighbors. From the reconstructed attractor manifold The above is generated through cross mapping The estimated value, weight The calculation method is as follows: ; in, It is the reconstructed attractor manifold Lag coordinate vectors of the same period and The Euclidean distance between them It is the reconstructed attractor manifold Lag coordinate vectors of the same period and The Euclidean distance between them; calculate and The cross-mapping skill values between them are used to determine whether there is a causal relationship between environmental factors and algal biomass.
[0014] In some embodiments, calculation and The steps for determining whether there is a causal relationship between environmental influencing factors and algal biomass based on the cross-mapping skill values include: calculate and Cross-mapping skill values : ; in, and They represent and The average value; When the cross-mapping skill value ρ gradually increases and eventually converges to a value greater than 0, it indicates that environmental influencing factors... There is a causal relationship between the algal biomass data Y and environmental influencing factors. It is the driving factor of algal biomass data Y; conversely, if Non-convergence The correlation coefficient ρ does not converge to a value greater than 0, which helps in determining environmental influencing factors. There is no causal relationship between algal biomass data Y and environmental influencing factors. It is not a driver of algal biomass data Y.
[0015] In some embodiments, by introducing a forecast period T, the driving factors of algal blooms at time t are... Algal biomass data at time t+T Matching is performed to construct a prediction sample set with a lead time T.
[0016] In some embodiments, the prediction sample set is divided into a training set and a test set in an 8:2 ratio to construct an algal bloom prediction model based on a random forest model. The specific steps are as follows: Construct a random forest model, wherein the random forest model consists of multiple decision trees, and when training each tree, multiple training subsets are randomly selected from the training set using a bootstrap sampling method; When splitting at each node of the decision tree, k features are randomly selected from all the driving factors of the input, and the optimal split point is selected from them for splitting; The algal bloom prediction model is obtained by optimizing the key hyperparameters of the random forest model through cross-validation, including the number of decision trees, the maximum depth of the trees, and the minimum number of samples per leaf node.
[0017] In some embodiments, after obtaining the algal bloom prediction model, a test set is used to evaluate the preset algal bloom prediction model, and the coefficient of determination index is used to quantitatively evaluate the prediction accuracy of the algal bloom prediction model: ; in, The coefficient of determination is a metric used to measure the goodness of fit between the predicted values of an algal bloom prediction model and the actual observed values. It is the number of test values in the test set. and They represent the first One observed value and one simulated value, It is the average value of the test values in the test set.
[0018] In a second aspect of the invention, a computer device is also provided, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a reservoir algal bloom prediction method based on convergent cross mapping and machine learning.
[0019] In a third aspect of the invention, a storage medium is also provided, the storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting reservoir algal blooms based on convergent cross-mapping and machine learning. Attached Figure Description
[0020] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0021] Figure 1 This is a flowchart illustrating the method for identifying and predicting driving factors of reservoir algal blooms based on convergent cross-mapping and machine learning, as presented in this invention. Figure 2This is a technical flowchart of a method for identifying and predicting driving factors of reservoir algal blooms based on convergent cross mapping and machine learning. Figure 3 This is a schematic diagram illustrating the specific steps involved in constructing the algal bloom prediction model of the present invention. Detailed Implementation
[0022] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.
[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0024] In the first aspect of the invention, please refer to Figure 1 The diagram below illustrates the process of identifying and predicting the driving factors of reservoir algal blooms based on convergent cross-mapping and machine learning, according to the present invention. The method includes the following steps: S1. Acquire long-sequence monitoring data of the reservoir and preprocess the long-sequence monitoring data to obtain time series of environmental influencing factors and time series of algal biomass. The long-sequence monitoring data includes algal biomass data and environmental influencing factors. The algal biomass data includes chlorophyll a concentration or algal density. The environmental influencing factors include water temperature, sunshine duration, nutrient content and hydrodynamic conditions. The study collected long-series monitoring data of the reservoir, with a length of L. The long-series monitoring data included algal biomass and environmental influencing factors. The algal biomass included chlorophyll a concentration or algal density. The environmental influencing factors included water temperature, meteorological conditions, nutrients, and hydrodynamic conditions.
[0025] The acquired data was then preprocessed, including time scale unification, missing value handling, outlier removal, and standardization, to construct a multivariate time series dataset for subsequent analysis, ultimately yielding the time series data of environmental influencing factors. and algal biomass time series .
[0026] In other embodiments, those skilled in the art can make appropriate adjustments to the preprocessing steps and the collected data based on their understanding of the inventive concept of this application.
[0027] S2. The convergent cross-mapping algorithm is used to analyze the causal relationship between the algal biomass factor sequence and the environmental impact factor sequence, and to screen out the driving factors of algal blooms.
[0028] Convergent Cross Mapping (CCM) is a causal inference method based on nonlinear state-space reconstruction, mainly used to determine whether there is a nonlinear causal relationship between two time series variables.
[0029] In this embodiment, to accurately identify the driving factors of algal blooms, the present invention uses a convergent cross-mapping algorithm to calculate the cross-mapping skill value from algal biomass data to environmental influencing factors. When the cross-mapping skill value converges, it is determined that the environmental influencing factors have a causal relationship with the algal biomass data. The environmental influencing factors with convergent cross-mapping skill values are selected as the driving factors of algal blooms, as follows: After preprocessing, S1 yields time series of environmental impact factors, each with a length of L. and algal biomass time series Let the dimension of the reconstructed attractor manifold be . The sampling interval when reconstructing the attractor manifold is So in The hysteresis coordinate vector of the reconstructed attractor manifold at each time step is: ; ; in, Represents the attractor manifold after reconstruction. The lag coordinate vector, Represents the attractor manifold after reconstruction. The lagging coordinate vector, t∈[ ,L].
[0030] Subsequently, information about The convergent cross-estimation values specifically include: definition From The above is generated through cross mapping The estimated value, first in Find the lagging coordinate vectors of the same period. and use express The time indices of the E+1 nearest neighbors are used to determine superior The nearest neighbor, according to indivual Local weighted average estimation of environmental impact factors Genetic environmental impact factors convergent cross estimate : ; in, Environmental influencing factors Simultaneous value, yes Rather than the attractor manifold after reconstruction Upper The weight of the distance between the nearest neighbors. From the reconstructed attractor manifold The above is generated through cross mapping The estimated value, weight The calculation method is as follows: ; in, It is the reconstructed attractor manifold Lag coordinate vectors of the same period and The Euclidean distance between them It is the reconstructed attractor manifold Lag coordinate vectors of the same period and The Euclidean distance between them; Next calculation and Cross-mapping skill values And based on cross-mapping skill values Determine whether there is a causal relationship between environmental factors and algal biomass: ; in, and They represent and The average value; Specifically, with the length of the original time series As the attractor manifold increases, it becomes more compact, and the distance between the E+1 nearest neighbors decreases. Gradually converges to Cross-mapping skill values (i.e., correlation coefficients) The value gradually increases and eventually converges to a value greater than 0, indicating that environmental influencing factors... For algal biomass There is a causal relationship, that is, environmental influencing factors. yes The driving factors. Conversely, if Non-convergence Cross-mapping skill values The fact that the environmental influencing factors do not converge to a value greater than 0 indicates that the factors are not converging. For algal biomass There is no causal relationship, i.e., environmental influencing factors. no The driving factors.
[0031] In other embodiments, those skilled in the art can make appropriate adjustments to the convergent cross-mapping algorithm based on their understanding of the inventive concept of this application.
[0032] S3. Input the selected driving factors of algal blooms into the preset algal bloom prediction model, and output the predicted value of algal biomass data of the reservoir in the future forecast period. When the predicted value exceeds the preset threshold, it is determined to be a high-risk state of algal bloom and triggers an algal bloom risk warning.
[0033] The algal bloom prediction model is based on the random forest model, which is a classifier that uses multiple trees to train and predict samples.
[0034] Based on the identification of the driving factors of algal blooms, an algal bloom prediction model based on a random forest model is constructed to achieve reservoir algal bloom prediction with a predictive period.
[0035] The set of algal bloom drivers identified in step S1 Algal biomass is used as the input variable in the algal bloom prediction model and as the output variable. To achieve early prediction of reservoir algal blooms, this invention introduces a prediction period T, which represents the driving factors of algal bloom outbreaks at time t. Algal biomass at time t+T Matching is performed to construct a predictive sample set with a forward-looking period.
[0036] Please see Figure 3 The constructed sample set is divided into a training set and a test set in an 8:2 ratio to construct the random forest model. The specific steps are as follows: S301. Random forests consist of multiple decision trees. When training each tree, multiple training subsets are randomly selected from the training set using bootstrap sampling. S302. When splitting at each node of the decision tree, randomly select k features from all input driving factors, and select the optimal split point from them for splitting; S303. Optimize the key hyperparameters of the random forest using cross-validation, including the number of decision trees (n_estimators), the maximum depth of the trees (max_depth), and the minimum number of samples per leaf node (min_samples_leaf).
[0037] The trained model was evaluated using a test set, employing the coefficient of determination (R²). 2 The indicators are used to quantitatively evaluate the model's prediction accuracy: ; in, The coefficient of determination is a metric used to measure the goodness of fit between the predicted values of an algal bloom prediction model and the actual observed values. It is the number of test values in the test set. and They represent the first One observed value and one simulated value, It is the average value of the test values in the test set.
[0038] Finally, the data on algal bloom drivers obtained from real-time or historical monitoring are input into the evaluated prediction model, which outputs the predicted value of algal biomass for the foreseeable future. When the predicted value exceeds the preset threshold, it is determined to be a high-risk state for algal bloom and an algal bloom risk warning is automatically triggered.
[0039] In other embodiments, those skilled in the art can make appropriate adjustments to the algal bloom prediction model based on their understanding of the inventive concept of this application.
[0040] In a second aspect of the invention, a computer device is also provided, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a reservoir algal bloom prediction method based on convergent cross mapping and machine learning.
[0041] In a third aspect of the invention, a storage medium is also provided, the storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting reservoir algal blooms based on convergent cross-mapping and machine learning.
[0042] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0043] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the algorithm. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0044] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0045] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0046] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0047] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms.
[0048] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for predicting reservoir algal blooms based on convergent cross-mapping and machine learning, characterized in that, Includes the following steps: Long-term monitoring data of the reservoir is acquired and preprocessed to obtain time series of environmental influencing factors and time series of algal biomass. The long-term monitoring data includes algal biomass data and environmental influencing factors. The algal biomass data includes chlorophyll a concentration or algal density. The environmental influencing factors include water temperature, sunshine duration, nutrient content and hydrodynamic conditions. The convergent cross-mapping algorithm is used to analyze the causal relationship between the algal biomass factor sequence and the environmental impact factor sequence. The cross-mapping skill value of algal biomass data to environmental impact factors is calculated. When the cross-mapping skill value converges, it is determined that the environmental impact factors have a causal relationship with the algal biomass data. The environmental impact factors with convergent cross-mapping skill values are selected as the driving factors of algal blooms. The selected driving factors of algal blooms are input into a preset algal bloom prediction model, which outputs the predicted value of algal biomass data in the reservoir within the foreseeable future period. When the predicted value exceeds a preset threshold, it is determined to be a high-risk state of algal bloom and triggers an algal bloom risk warning.
2. The reservoir algal bloom prediction method based on convergent cross-mapping and machine learning according to claim 1, characterized in that: The preprocessing steps include time-scale unification, missing value processing, outlier removal, and standardization of the acquired long-sequence monitoring data.
3. The reservoir algal bloom prediction method based on convergent cross-mapping and machine learning according to claim 1, characterized in that: The method for calculating the cross-mapping skill value includes reconstructing the attractor manifold to obtain the hysteresis coordinate vector, specifically including: Define the time series of environmental influencing factors Time series of algal biomass data The length of each element is L, the dimension of the reconstructed attractor manifold is E, and the sampling interval during the reconstruction of the attractor manifold is τ. The hysteresis coordinate vector of the reconstructed attractor manifold at time t is obtained as follows: ; ; in, Represents the reconstructed attractor manifold The lag coordinate vector, Represents the reconstructed attractor manifold The lagging coordinate vector, t∈[ ,L].
4. The reservoir algal bloom prediction method based on convergent cross-mapping and machine learning according to claim 3, characterized in that: The steps for calculating the cross-mapping skill value from algal biomass data to environmental influencing factors, and determining whether there is a causal relationship between environmental influencing factors and algal biomass data, also include: After reconstruction of the attractor manifold Find the lagging coordinate vectors of the same period. ,pass express The time indices of the E+1 nearest neighbors are used to determine the reconstructed attractor manifold. Lag coordinate vectors of the same period The nearest neighbor, according to indivual Local weighted average estimation of environmental impact factors Genetic environmental impact factors convergent cross estimate : ; in, Environmental influencing factors Simultaneous value, yes Rather than the attractor manifold after reconstruction Upper The weight of the distance between the nearest neighbors. From the reconstructed attractor manifold The above is generated through cross mapping The estimated value, weight The calculation method is as follows: ; in, It is the reconstructed attractor manifold Lag coordinate vectors of the same period and The Euclidean distance between them It is the reconstructed attractor manifold Lag coordinate vectors of the same period and The Euclidean distance between them; calculate and The cross-mapping skill values between them are used to determine whether there is a causal relationship between environmental factors and algal biomass.
5. The reservoir algal bloom prediction method based on convergent cross-mapping and machine learning according to claim 4, characterized in that: calculate and The steps for determining whether there is a causal relationship between environmental influencing factors and algal biomass based on the cross-mapping skill values include: calculate and Cross-mapping skill values : ; in, and They represent and The average value; When the cross-mapping skill value ρ gradually increases and eventually converges to a value greater than 0, it indicates that environmental influencing factors... There is a causal relationship between the algal biomass data Y and environmental influencing factors. It is the driving factor of algal biomass data Y; conversely, if Non-convergence The correlation coefficient ρ does not converge to a value greater than 0, which helps in determining environmental influencing factors. There is no causal relationship between algal biomass data Y and environmental influencing factors. It is not a driver of algal biomass data Y.
6. The reservoir algal bloom prediction method based on convergent cross-mapping and machine learning according to claim 1, characterized in that: By introducing a forecast period T, the driving factors of algal blooms at time t are... Algal biomass data at time t+T Matching is performed to construct a prediction sample set with a lead time T.
7. The reservoir algal bloom prediction method based on convergent cross-mapping and machine learning according to claim 6, characterized in that: The predicted sample set is divided into a training set and a test set in an 8:2 ratio. An algal bloom prediction model based on a random forest model is then constructed. The specific steps are as follows: Construct a random forest model, wherein the random forest model consists of multiple decision trees, and when training each tree, multiple training subsets are randomly selected from the training set using a bootstrap sampling method; When splitting at each node of the decision tree, k features are randomly selected from all the driving factors of the input, and the optimal split point is selected from them for splitting; The algal bloom prediction model is obtained by optimizing the key hyperparameters of the random forest model through cross-validation, including the number of decision trees, the maximum depth of the trees, and the minimum number of samples per leaf node.
8. The reservoir algal bloom prediction method based on convergent cross-mapping and machine learning according to claim 7, characterized in that: After obtaining the algal bloom prediction model, a test set is used to evaluate the preset algal bloom prediction model, and the coefficient of determination index is used to quantitatively evaluate the prediction accuracy of the algal bloom prediction model: ; in, The coefficient of determination is a metric used to measure the goodness of fit between the predicted values of an algal bloom prediction model and the actual observed values. It is the number of test values in the test set. and They represent the first One observed value and one simulated value, It is the average value of the test values in the test set.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 8.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.