Water source algae early warning method and device based on mechanism map, equipment and medium

CN122548425APending Publication Date: 2026-08-11WATER ENG ECOLOGICAL INST CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本申请的目的在于克服上述技术不足,提出一种基于机理图谱的水源地藻类预警方法、装置、设备及介质,解决现有技术中转折事件样本稀缺、统计扩增样本生态合理性不足以及预警模型对关键风险跃迁识别能力不足的技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548425A_ABST
    Figure CN122548425A_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, equipment, and medium for early warning of algae blooms in water sources based on a mechanistic map. The early warning method includes: firstly, constructing a time-series map of mechanistic knowledge based on historical monitoring data during the training period and aquatic ecology literature; then, extracting algal bloom level inflection event windows from the historical monitoring data during the training period to form a sample set, fitting a multivariate joint distribution model, and sampling to generate candidate synthetic inflection event scenarios; next, performing a knowledge compliance check on the candidate scenarios and screening out synthetic scenarios; then merging them into an enhanced training set to train an early warning model to predict the algal bloom level in a future set time domain; finally, outputting the algal bloom level prediction results and graded warning signals based on real-time online water quality and meteorological and hydrological data. This application can enhance the characterization ability of scarce inflection event samples and improve the accuracy and stability of early warning of abnormal algal proliferation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological environment monitoring and early warning technology, specifically to a method, device, equipment, and medium for early warning of algae in water sources based on mechanism maps. Background Technology

[0002] Enclosed or semi-enclosed water bodies such as reservoirs and lakes, which are drinking water sources, are prone to abnormal algal blooms due to the combined effects of nutrient input, meteorological disturbances, hydrodynamic changes, and biological community succession. Abnormal algal blooms not only affect water transparency, dissolved oxygen, and ecological balance, but may also cause odors, algal toxin risks, and water supply safety issues. Therefore, early warning of algal blooms is of great significance.

[0003] Existing methods for early warning of algal blooms mainly include empirical threshold-based methods, statistical prediction methods, and data-driven methods based on machine learning. Empirical threshold methods rely on manually set rules and have limited adaptability to complex dynamic scenarios. Statistical models often fail to adequately characterize the nonlinearity, time lag, and sudden transitions in variable relationships. While machine learning models possess some nonlinear fitting capabilities, they typically rely on large-scale, evenly distributed training samples. However, the truly valuable early warning events in abnormal algal blooms in water sources are often the transitional events from low-risk to medium-to-high-risk states. Such samples are few in number and unevenly distributed, leading to insufficient model recognition capabilities during critical periods. Furthermore, relying solely on statistical regularities for sample amplification can easily generate samples inconsistent with actual ecological mechanisms. For example, while the direction of change and time lag relationships of certain variables may be mathematically valid, they may not be reasonable from an ecological perspective, potentially interfering with model training. At the same time, the ecological processes of water sources have obvious characteristics of multi-factor coupling, time lag and seasonal stratification. The same driving factor may also have opposite effects at different time scales. For example, heavy rain or high inflow may inhibit algae in the short term, but may promote algae growth through nutrient input in the later stage. Traditional methods are difficult to systematically express such mechanistic constraints.

[0004] Therefore, there is a need for an early warning scheme for algae in water sources that can combine ecological mechanism knowledge, time series representation, targeted enhancement of turning events, and machine learning prediction to improve the early warning capability for abnormal algal proliferation. Summary of the Invention

[0005] The purpose of this application is to overcome the above-mentioned technical deficiencies and propose a method, device, equipment and medium for early warning of algae in water sources based on mechanism maps, which solves the technical problems of scarce transition event samples, insufficient ecological rationality of statistical amplification samples and insufficient ability of early warning models to identify key risk transitions in the existing technology.

[0006] To achieve the above-mentioned technical objectives, the present application adopts the following technical solution: Firstly, this application provides a method for early warning of algae in water sources based on mechanistic maps, including: A time-series graph of mechanistic knowledge is constructed based on historical monitoring data during the training period and aquatic ecology literature. The time-series graph of mechanistic knowledge includes nodes, directed edges, and dual-effect path encoding. Based on the historical monitoring data during the training period, the transition event window of the algal bloom level is extracted, and the monitoring data of each window constitutes the transition event sample set. The sample set is divided into corresponding subsets according to the stratification condition. A multivariate joint distribution model is fitted on each subset, and candidate synthetic transition event scenarios are generated by sampling from the multivariate joint distribution model. Based on the aforementioned mechanism knowledge time series graph, a knowledge compliance check is performed on each of the aforementioned candidate synthesis turning point events, and synthesis scenarios that conform to the graph constraints are retained. The historical monitoring data during the training period is combined with the synthetic scenario data to form an enhanced training set. The enhanced training set is then trained using a machine learning model to obtain an early warning model for predicting the level of algal blooms in a future time domain. The system acquires real-time online water quality monitoring data and meteorological and hydrological data, and outputs algal bloom level prediction results and graded alarm signals through the early warning model.

[0007] In some embodiments of this application, the construction of a time-series graph of mechanistic knowledge based on historical monitoring data during the training period and aquatic ecology literature, wherein the time-series graph of mechanistic knowledge includes nodes, directed edges, and dual-effect path encoding, including: The nodes are defined as environmental factor nodes, biological factor nodes, event nodes, and state nodes. The environmental factor nodes include water temperature, total phosphorus, total nitrogen, ammonia nitrogen, dissolved oxygen, pH, conductivity, and turbidity. The biological factor nodes include cyanobacteria density, diatom density, green algae density, and cryptophyte density. The event nodes include heavy rainfall, high inflow rate, and low water level. The state nodes include algal bloom level: safe level, low risk level, and high risk level. Define directed edges connecting nodes. Each directed edge carries a relation type, a time lag window, a confidence weight, and a source identifier. The relation type includes facilitator, inhibitor, co-occurrence, and predecessor. The time lag window is in hours. The confidence weight ranges from 0 to 1. The source identifier includes literature, data mining, and experts. Encode dual-effect paths to support the opposite effects of the same driving factor within different time lag windows.

[0008] In some embodiments of this application, the step of extracting algal bloom level transition event windows based on the historical monitoring data during the training period, constructing a transition event sample set with monitoring data from each window, dividing the sample set into corresponding subsets according to stratification conditions, fitting a multivariate joint distribution model to each subset, and sampling from the multivariate joint distribution model to generate candidate synthetic transition event scenarios includes: The time step in which the algal bloom level changes from the safe level to the low-risk level or above during the training period is selected as the starting time step. Based on the starting time step, the first preset duration is traced forward and the second preset duration is traced backward as the turning point event window. The sample set was divided into spring / summer and autumn / winter layers according to the season, and a Gaussian copula model was fitted to each layer to generate candidate synthetic turning point scenarios.

[0009] In some embodiments of this application, the step of performing a knowledge compliance check on each candidate synthetic turning point scenario based on the mechanistic knowledge time series graph, and retaining synthetic scenarios that conform to the graph constraints, includes: Obtain the set of applicable edges triggered in each candidate synthetic turning event scenario; Examine whether the direction of change of the target node in the candidate synthetic turning event scenario conforms to the relation type constraint of the corresponding directed edge, and whether the temporal relation falls within the time lag window of the corresponding directed edge; A knowledge compliance score is calculated based on the ratio of the number of applicable edges that pass the test to the total number of applicable edges in the set, and synthetic scenarios with scores not lower than a set threshold are retained.

[0010] In some embodiments of this application, the machine learning model is a gradient boosting decision tree or a long short-term memory network, and the future time domain includes a short-term time domain and a medium-to-long-term time domain. When using a gradient boosting decision tree, the input features include current value features, multi-step time lag features, rolling statistical features, rate of change features, and calendar features; the historical monitoring data during the training period includes online water quality monitoring data, meteorological data, and hydrological data. The online water quality monitoring data includes chlorophyll a and water quality parameters, the meteorological data includes air temperature, precipitation, and wind speed, and the hydrological data includes water level and inflow.

[0011] In some embodiments of this application, the time series map of mechanistic knowledge is constructed from the time-delay cross-correlation scan results of training period data, seasonal stratification analysis results, knowledge from aquatic ecology literature, and expert-confirmed knowledge, and is used to correct counterintuitive relationships in data mining results.

[0012] In some embodiments of this application, the dual-effect path further includes high inflow rate inhibiting algae in a third preset time window via the first edge and promoting algae in a fourth preset time window via the second edge; and silver carp and bighead carp filter feeding inhibiting algae of a preset particle size in a fifth preset time window via the first edge and promoting small cyanobacteria in a sixth preset time window via the second edge.

[0013] Secondly, this application provides a water source algae early warning device based on mechanism mapping, comprising: The graph construction module is used to construct a time-series graph of mechanistic knowledge based on historical monitoring data during the training period and aquatic ecology literature. The time-series graph of mechanistic knowledge includes nodes, directed edges, and dual-effect path encoding. The sample generation module is used to extract the algal bloom level transition event window based on the historical monitoring data of the training period, and to form a transition event sample set with the monitoring data of each window. The sample set is divided into corresponding subsets according to the stratification conditions. A multivariate joint distribution model is fitted to each subset, and candidate synthetic transition event scenarios are generated by sampling from the multivariate joint distribution model. The compliance check module is used to perform knowledge compliance checks on each of the candidate synthetic turning point events based on the time-series graph of the mechanism knowledge, and retain synthetic scenarios that conform to the graph constraints. The model training module is used to merge the historical monitoring data of the training period with the synthetic scenario data to form an enhanced training set, and to train the enhanced training set using a machine learning model to obtain an early warning model for predicting the level of algal blooms in a future set time domain. The early warning output module is used to acquire real-time online water quality monitoring data and meteorological and hydrological data, and output the algal bloom level prediction results and graded alarm signals through the early warning model.

[0014] Thirdly, this application provides an electronic device, including a memory and a processor, wherein, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the water source algae early warning method based on mechanism maps as described in any of the embodiments of the first aspect above.

[0015] Fourthly, this application also provides a computer-readable storage medium for storing a computer-readable program or instructions, which, when executed by a processor, can implement the steps of the water source algae early warning method based on mechanism maps described in any of the embodiments of the first aspect above.

[0016] Compared with the prior art, the beneficial technical effects of the technical solution provided in this application include: This application constructs a sample enhancement process around algal bloom level transition events. Based on multivariate joint distribution sampling, it introduces mechanistic knowledge time series graph constraints to screen candidate synthetic transition event scenarios for knowledge compliance. This ensures that the enhanced samples entering the training phase simultaneously consider statistical distribution characteristics and ecological mechanism rationality. By using the enhanced training set for machine learning modeling, the model's ability to identify low-risk to high-risk transition processes can be improved, thereby enhancing the accuracy and stability of early warning of abnormal algal proliferation in water sources. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the embodiments will be briefly described below: Figure 1 This is a flowchart illustrating the water source algae early warning method based on mechanism maps in an embodiment of this application; Figure 2 This is a flowchart of an embodiment of the early warning method in this application; Figure 3 This is a comparison diagram of the model results of the early warning method embodiments in this application; Figure 4 This is the core pathway diagram of the time-series mechanistic knowledge in the embodiments of this application; Figure 5 This is a pipeline verification diagram of knowledge compliance in a synthetic scenario in the embodiments of this application; Figure 6 This is a schematic block diagram of the early warning device in the embodiments of this application; Figure 7 This is a schematic block diagram of an electronic device in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] Those skilled in the art will understand that, in this specification, the term "comprising" is an open-ended expression, meaning that the stated feature is present but other features are excluded. Directional terms such as "upper," "lower," "left," and "right" refer to exemplary directions based on the accompanying drawings. Features specified as "first" or "second" implicitly include one or more of that feature. Singular expressions can also be used in plural forms. "Multiple" means two or more. The terms "installed," "connected," and "linked" can refer to a fixed connection, a detachable connection, or an integral connection; it can be a direct connection or an indirect connection via an intermediate medium, and it can be a connection within two components. Furthermore, "linked" can include wireless connections.

[0020] The purpose of this application is to overcome the above-mentioned technical deficiencies and propose a method, device, equipment and medium for early warning of algae in water sources based on mechanism maps, which solves the technical problems of scarce transition event samples, insufficient ecological rationality of statistical amplification samples and insufficient ability of early warning models to identify key risk transitions in the existing technology.

[0021] To achieve the above-mentioned technical objectives, the present application adopts the following technical solution: This embodiment provides a method for early warning of algae in water sources based on mechanistic maps. For example... Figure 1 As shown, the method includes the following steps.

[0022] Step S1: Construct a mechanism knowledge time series graph (MT-TKG) based on historical monitoring data during the training period and aquatic ecology literature. The mechanism knowledge time series graph includes nodes, directed edges, and dual-effect path encoding.

[0023] Historical monitoring data for the training period should preferably be continuous time series data, which can come from online monitoring systems, conventional monitoring systems, and supporting meteorological and hydrological stations in water source reservoirs, lakes, or river-type drinking water sources. The data time granularity can be at the hourly level, or it can be set to a shorter or longer time interval depending on the monitoring conditions. The training period should preferably cover a complete year or a multi-year process to include different hydrological and meteorological conditions and algal succession stages.

[0024] Mechanistic knowledge time series graphs are used to express the directed, time-delayed, and conditional associations among multiple factors during abnormal algal proliferation in water sources. Nodes represent environmental factors, biological factors, events, and risk states; directed edges represent the interaction between driving factors and response factors; and dual-effect path encoding represents the opposite effects of the same driving factor on the target node within different time lag windows.

[0025] Step S2: Based on historical monitoring data during the training period, extract the transition event window of algal bloom level, and use the monitoring data of each window to form a transition event sample set. Divide the sample set into corresponding subsets according to the stratification conditions, fit a multivariate joint distribution model on each subset, and sample from the multivariate joint distribution model to generate candidate synthetic transition event scenarios.

[0026] Among them, the algal bloom level transition event is preferably a key event that changes the algal bloom level from a low-risk state to a higher-risk state. The transition event window is used to retain the multivariate dynamic process over a certain period before and after the transition, rather than just capturing a single point value at a certain moment. Stratification conditions can be set according to season, temperature range, hydrological period, typical meteorological event types, etc., to reduce distribution bias caused by mixed modeling of different ecological backgrounds. The multivariate joint distribution model is used to characterize the joint statistical structure among multiple monitoring variables, and based on this, candidate synthetic transition event scenarios are sampled and generated.

[0027] Step S3: Based on the time-series graph of mechanistic knowledge, perform a knowledge compliance check on each candidate synthetic turning point scenario and retain synthetic scenarios that conform to the graph constraints.

[0028] The purpose of knowledge compliance checks is to determine whether candidate synthetic turning point scenarios satisfy the directional and temporal relationships in the ecological mechanism. Samples that meet the graph constraints are retained, while those that do not are removed, thereby reducing the number of statistically feasible but ecologically unreasonable samples entering the training phase.

[0029] Step S4: Combine the historical monitoring data and synthetic scenario data during the training period to form an enhanced training set. Use a machine learning model to train the enhanced training set to obtain an early warning model for predicting the level of algal blooms in the future at a set time domain.

[0030] The enhanced training set includes both real observation sequences and synthetic inflection event scenarios filtered from spectral data. The machine learning model can establish a mapping between multi-source monitoring variables and future algal bloom levels. The future timeframe can be set according to management needs, such as from several hours to several days.

[0031] Step S5: Obtain real-time online water quality monitoring data and meteorological and hydrological data, and output the algal bloom level prediction results and graded alarm signals through the early warning model.

[0032] Real-time online water quality monitoring data and meteorological and hydrological data are preprocessed and then input into the early warning model. The model outputs the predicted value or category of algal bloom level after a set time domain, and further maps it into graded alarm signals for dispatching early warning, increasing patrol density, adjusting water intake process and assisting in emergency response decision-making.

[0033] This embodiment establishes a technical chain around algal bloom level transition events, encompassing "map construction—transition sampling—knowledge filtering—reinforcement training—real-time early warning." First, a time-series map of mechanistic knowledge capable of expressing directionality and time lag is constructed using historical monitoring data and ecological knowledge. Then, real transition event windows are extracted and candidate synthetic transition event scenarios are generated using a hierarchical modeling approach. Next, compliance verification is performed using the mechanistic map, filtering out samples that violate ecological logic. Finally, the enhanced training set is used to train a machine learning model and output real-time early warning results. This allows the model to be more fully trained under key risk transition scenarios. This embodiment uses transition events as the core enhancement object and constrains the quality of synthetic samples through the mechanistic knowledge time-series map, making the enhanced training set closer to the real ecological evolution process, thereby improving the ability to identify key transition processes of abnormal algal proliferation and enhancing the accuracy and stability of early warnings.

[0034] This embodiment describes the construction process of a time-series graph of mechanistic knowledge. The time-series graph of mechanistic knowledge includes nodes, directed edges, and dual-effect path encoding, and its construction includes the following:

[0035] First, define the nodes. Nodes include environmental factor nodes, biological factor nodes, event nodes, and state nodes. Environmental factor nodes include water temperature, total phosphorus, total nitrogen, ammonia nitrogen, dissolved oxygen, pH, conductivity, and turbidity. Biological factor nodes include cyanobacteria density, diatom density, green algae density, and cryptophyte density. Event nodes include heavy rainfall, high inflow, and low water level. State nodes include algal bloom levels: safe, low-risk, and high-risk.

[0036] The nodes can be further expanded based on monitoring capabilities. For example, variable nodes such as chlorophyll a, transparency, light intensity, wind direction, and wind speed components can be added without changing the technical concept. Each node can store basic information such as node name, node type, unit, numerical range, sampling timestamp, and standardization method.

[0037] Next, the directed edges connecting the nodes are defined. Each directed edge carries the relationship type, time lag window, confidence weight, and source identifier. Relationship types include facilitator, inhibitor, co-occurrence, and predecessor. The time lag window is in hours and can be represented as a single time lag point or a start-end interval. The confidence weight ranges from 0 to 1, reflecting the reliability of the interaction. The source identifier includes literature, data mining, and experts, to distinguish the relationship's origin and facilitate subsequent maintenance.

[0038] Furthermore, dual-effect paths are encoded to support the opposite effects of the same driving factor within different time lag windows. Dual-effect paths can be represented by two or more directed edges corresponding to the same starting node but different time windows and different relationship types, or they can be represented as a path chain with conditional labels. For example, a rainstorm node can correspond to edges that inhibit algae growth in a short-term window, while simultaneously corresponding to edges that promote algae growth in a medium-term window.

[0039] The build process may include the following specific operations: First, perform time-delay cross-correlation scans on the training data to identify the possible directions and significant time periods of time-delay correlation between variables; Second, conduct seasonal stratification analysis to distinguish the differences in relationship direction and time lag window under different seasonal conditions; Third, we searched aquatic ecology literature to extract repeatedly confirmed mechanisms of action and their applicable conditions; Fourth, domain experts comprehensively verify the data mining results and literature knowledge, and correct or eliminate relationships that are obviously counterintuitive or beyond ecological common sense. Fifth, the confirmed nodes, edges, and dual-effect paths are uniformly encoded and incorporated into the time series graph of mechanistic knowledge.

[0040] This embodiment describes the process of extracting transition event windows, hierarchical partitioning, and multivariate joint distribution modeling. Specifically, it includes the following:

[0041] The starting time step is selected when the algal bloom level changes from safe to low-risk or above during the training period. A first preset duration is traced backward from this starting time step, and a second preset duration is traced backward as the transition event window. The first preset duration is preferably set to 72 hours, and the second preset duration is preferably set to 24 hours. This covers both the driving accumulation process before the risk escalates and the early evolutionary stage after the risk transition.

[0042] For each turning point event window, the time series data corresponding to each monitored variable within the window are extracted to form a turning point event sample. Multiple turning point event samples together constitute a turning point event sample set. The sample set can be organized in the form of a matrix, tensor, or serialized features.

[0043] The sample set was divided into spring / summer and autumn / winter layers based on the season. The spring / summer layer typically corresponds to the period of active algae growth, while the autumn / winter layer corresponds to the period of relatively low temperatures. If necessary, the start and end times of the spring / summer and autumn / winter layers can be regionalized according to local ecological characteristics, for example, using monthly, ten-day, or sliding temperature thresholds as stratification criteria.

[0044] Gaussian copula models are fitted at each layer to generate candidate synthetic turning point scenarios. Specifically, the marginal distributions of each variable are first estimated, then the correlation structure between variables is fitted using Gaussian copula, and finally sampling is performed under the fitted joint distribution to generate synthetic turning point scenarios with corresponding statistical characteristics. For time series samples, key time lag features, window statistics, or multi-time-point joint variable vectors can be used as the copula input space to preserve the cross-variable correlation structure.

[0045] To improve the quality of synthetic samples, the training data can be imputed for missing values, corrected for outliers, aligned in time, standardized in units, and normalized before sampling; after sampling, variable range truncation, physical feasibility constraints, and extreme noise smoothing can be performed.

[0046] This embodiment describes the knowledge compliance check process, including the following:

[0047] For each candidate synthetic turning point event scenario, the applicable edge set is first identified based on the change process of each variable in the scenario and the event triggering situation. The applicable edge set refers to the set of directed edges that have triggering conditions, determineable directions, and time delay correspondences under the scenario.

[0048] Then, we examine whether the direction of change of the target node in the candidate synthetic turning point event scenario conforms to the relation type constraint of the corresponding directed edge, and whether the time relation falls within the time lag window of the corresponding directed edge.

[0049] When the relationship type is promoting, the changes in the starting node and the changes in the target node are required to have the same gain characteristics in the corresponding lag window. When the relationship type is suppression, the changes in the starting node and the changes in the target node are required to have an inverse relationship within the corresponding hysteresis window; When the relationship type is co-occurrence, it requires that both nodes show significant changes together within the same time period or within the allowable error range; When the relation type is predecessor, the change of the starting node must occur before the change of the target node, and the corresponding time delay window constraint must be satisfied.

[0050] A knowledge compliance score is calculated based on the ratio of the number of applicable edges that pass the test to the total number of applicable edges in the set. Synthetic scenarios with scores not lower than a set threshold are retained. The knowledge compliance score can be expressed as: K = Npass / Nall Where K represents the knowledge compliance score, Npass represents the number of applicable edges that pass the test, and Nall represents the total number of applicable edges. The threshold can be set based on the validation set results, domain experience, or the conservatism of the scenario, with 0.6 being the preferred value. If the set of applicable edges is empty, the candidate scenario is deemed non-compliant and is directly eliminated.

[0051] In practical implementation, the direction of change of the target node can be determined by using difference sign, slope sign, window mean difference, or classification addition / reduction labels; the time relationship can be determined by using fixed time delay window matching or sliding alignment with allowable deviation.

[0052] This embodiment describes the machine learning model, input features, and training data sources. The machine learning model is either a gradient boosting decision tree (LightGBM) or a long short-term memory network (LSTM), and the time domain will be set to include short-term and medium-to-long-term time domains.

[0053] When using gradient boosting decision trees, the input features include current value features, multi-step time lag features, rolling statistics features, rate of change features, and calendar features.

[0054] Current value characteristics refer to the original or standardized values ​​of each monitored variable at the current moment; Multi-step time delay features refer to the variable values ​​at several historical time steps, used to express time-series dependencies; Rolling statistical characteristics refer to statistical measures such as mean, extreme values, standard deviation, skewness, or volatility calculated within a set time window; The rate of change characteristic refers to the magnitude, slope, or relative rate of change within adjacent moments or a set time window; Calendar features refer to features that reflect periodic information, such as hour, day sequence, month, solar term segment, or workday marking.

[0055] Historical monitoring data during the training period includes online water quality monitoring data, meteorological data, and hydrological data. Online water quality monitoring data includes chlorophyll a and water quality parameters; meteorological data includes air temperature, precipitation, and wind speed; and hydrological data includes water level and inflow. All types of data are preferably unified to the same time granularity and time-aligned before modeling.

[0056] When using a Long Short-Term Memory (LSTM) network, a multivariate sequence arranged in chronological order can be directly input, or static auxiliary features or event marker features can be added to the sequence input. The output can be a multi-class classification output to predict the level of algal blooms as safe, low-risk, or high-risk, or it can be a regression output that is then thresholded to the corresponding level.

[0057] The short-term time domain can be set to 4 hours or 12 hours, while the medium-to-long-term time domain can be set to 24 hours, 48 ​​hours, or 72 hours. In practical applications, multiple early warning models can be trained simultaneously, each corresponding to a different forecast duration, to meet different operational needs such as water intake scheduling, chemical dosing control, and emergency response.

[0058] This embodiment explains the knowledge source fusion method of the time series graph of mechanism knowledge.

[0059] The mechanistic knowledge time series map is constructed from the time-delay cross-correlation scan results of the training period data, the seasonal stratification analysis results, the knowledge of aquatic ecology literature, and the knowledge confirmed by experts. It is used to correct counterintuitive relationships in the data mining results.

[0060] Among them, time-delay cross-correlation scan results are used to discover candidate driving relationships and their possible time-delay intervals from the data level; seasonal stratification analysis results are used to determine whether certain relationships have significant seasonal differences; knowledge of aquatic ecology literature is used to provide evidence of existing ecological mechanisms, such as the potential promoting effect of increased nutrient salts, rising water temperature, and intensified stable stratification on algal growth, or the inhibitory effect of high turbidity and low light on specific algal communities; expert confirmation knowledge is used to retain, adjust, or delete candidate relationships based on the watershed characteristics, reservoir scheduling methods, fishery stocking, and historical experience of the study area.

[0061] A counterintuitive relationship refers to a correlation that appears in local statistical data but is clearly inconsistent with known ecological processes, or may be a spurious relationship caused by data noise, sampling bias, extreme missing data, imputation, or other reasons. For example, if a variable shows an abnormally positive correlation with algal changes in a very small number of samples, but lacks ecological evidence and is not reproducible, the relationship can be corrected to a weak relationship or eliminated through expert confirmation and literature evidence.

[0062] This embodiment describes the extended coding content of the dual-effect path.

[0063] The dual-effect pathway also includes high inflow rates inhibiting algae in the third preset time window via the first edge and promoting algae in the fourth preset time window via the second edge; and silver carp and bighead carp filter feeding inhibiting algae of a preset particle size in the fifth preset time window via the first edge and promoting small cyanobacteria in the sixth preset time window via the second edge.

[0064] The dual effect of high inflow can be understood as follows: in the short term, high inflow may inhibit algal accumulation by shortening water retention time, enhancing mixing, and diluting algal concentration; however, in the subsequent stage, high inflow may be accompanied by the input of exogenous nutrients, sediment resuspension, and changes in suitable conditions, thereby promoting algal growth. The third preset time window can be set to 0 to 48 hours, and the fourth preset time window can be set to 72 to 168 hours.

[0065] The dual effect of filter feeding by silver carp and bighead carp can be understood as follows: within a certain particle size range, silver carp and bighead carp can directly consume larger algal populations, thereby reducing the corresponding algal density in a short period of time; however, in subsequent stages, due to particle size-selective feeding, changes in competitive relationships, and niche release, smaller cyanobacteria may gain a competitive advantage and thus increase. The fifth and sixth preset time windows can be set based on the ecological monitoring results of the study area, for example, set to 24 to 72 hours and 72 to 240 hours, respectively.

[0066] In graph coding, the aforementioned dual-effect path can be represented as a combination of multiple time-delayed edges from event nodes or biological regulation nodes to algal community structure nodes, specific algal nodes, and state nodes, or as a composite path formed by connecting intermediate nodes such as nutrients, turbidity, and particle size structure changes.

[0067] This embodiment provides supplementary explanations regarding the classification of algal bloom levels and the early warning output method.

[0068] Algal bloom levels can be classified into safe, low-risk, and high-risk levels based on chlorophyll a concentration, cyanobacterial density, algal cell abundance, fluorescence signal intensity, or a combination thereof. Safe level indicates a low-risk state, low-risk level indicates a state requiring attention, and high-risk level indicates a high-risk state. Level thresholds can be determined based on historical monitoring distribution in the study area, drinking water safety management requirements, and local technical specifications.

[0069] When outputting the grade results for a future time domain, the early warning output module can simultaneously output the confidence score, grade transition probability, or risk index. Graded alarm signals can be set to different alarm levels such as blue, yellow, and orange, or can be presented in text prompt form, such as "Stay alert," "Recommended to increase monitoring," or "Recommended to initiate high-risk response assessment."

[0070] This embodiment provides supplementary explanation of the preprocessing process for real-time input data.

[0071] Before being input into the early warning model, real-time online water quality monitoring data and meteorological and hydrological data can undergo time alignment, outlier identification, missing value imputation, unit conversion, smoothing and denoising, and standardization processing in sequence. Outlier identification can be carried out based on physical boundaries, sensor maintenance records, and sliding window statistical rules; missing value imputation can be achieved using linear interpolation, spline interpolation, backfilling from neighboring time periods, or model estimation methods; the standardized parameters preferably use the transformation parameters saved during the training phase to ensure consistency between training and inference.

[0072] For sudden breakpoints or inconsistencies in clocks from multiple data sources, a data quality flag can be set, and a "low confidence warning" or a manual review prompt can be triggered when the data quality falls below the threshold.

[0073] This embodiment provides supplementary explanation of the evaluation and update mechanism during model training.

[0074] After creating the enhanced training set, it can be further divided into a training set, a validation set, and a test set. The validation set is used to adjust model hyperparameters, knowledge compliance thresholds, and stratified sampling folds, while the test set is used to evaluate the performance of the final early warning model. Evaluation metrics may include accuracy, recall, F1 score, area under the receiver operating characteristic (AUC) curve, and hit rate, false negative rate, and lead time for high-priority events.

[0075] After actual deployment, the model can be incrementally updated monthly, quarterly, or after the algal bloom season. During updates, newly added real monitoring data can be incorporated into the training data, and the turning point event window, sample set, and early warning model parameters can be re-extracted, updated, and updated.

[0076] This embodiment provides a detailed explanation of the dual-effect path of heavy rainfall.

[0077] Heavy rainfall events directly inhibit algae growth within a 0-48 hour window via the first edge, and indirectly promote algae growth via non-point source nutrient input within a 72-168 hour window via the second edge. Specifically, in the short term after a heavy rainfall event, lower light intensity, higher turbulence, and scouring dilution may suppress algal aggregation; in the subsequent stage, the increase in total phosphorus and total nitrogen from watershed non-point sources, under suitable temperature and residence time conditions, transforms into a driver of algal growth.

[0078] In the graph, the rainstorm node can be connected to the turbidity node, total phosphorus node, and cyanobacteria density node, and different relationship types and time lag windows can be set; it can also be further connected to the algal bloom level node to reflect the comprehensive effect.

[0079] This embodiment explains the specific method for setting the knowledge compliance threshold.

[0080] The optimal threshold for knowledge compliance score is 0.6. When setting this threshold, cross-validation can be used to compare the retention rate of synthetic samples, model testing performance, and recall rate of high-level events under different thresholds on historical data. Then, a threshold that balances sample size and ecological rationality can be selected. For water sources where samples are relatively scarce, the threshold can be optimized within the range of 0.5 to 0.7; for scenarios requiring high ecological consistency, the threshold can be increased to strengthen the screening process.

[0081] This embodiment describes the multi-scale output in a future time domain setting.

[0082] The early warning model can output algal bloom severity results after 4 hours, 12 hours, 24 hours, 48 ​​hours, and 72 hours, respectively. For the same moment, multi-time-domain joint early warning results can be generated: the 4-hour and 12-hour results are used for short-term operation scheduling, the 24-hour and 48-hour results are used for monitoring intensification and emergency preparedness, and the 72-hour results are used for medium-term risk assessment.

[0083] In terms of model implementation, multiple models can be trained separately, or a single model with multiple output heads can be used to complete joint prediction across multiple time domains.

[0084] This embodiment further explains the specific composition of LightGBM input features.

[0085] When using a gradient boosting decision tree, multi-step time lag features can be set to the monitoring values ​​of the past 6 steps; rolling statistical features can calculate the mean, maximum, minimum and standard deviation based on the past 48-hour window; rate of change features can include 1-hour rate of change, 6-hour rate of change and 24-hour rate of change; calendar features can include hour, weekday and month codes.

[0086] This embodiment illustrates the application scenario of drinking water source reservoirs.

[0087] This early warning method can be deployed on the online monitoring and early warning platform for drinking water source reservoirs, with the identification and forecasting of high-risk algal bloom levels as the key output. When the model predicts that the risk level will reach a high level within the next 24 hours, it can trigger on-site inspections, algal density verification at water intakes, pre-adjustment of water purification process parameters, and emergency material preparation procedures.

[0088] The synthetic data generation adopts a two-stage quality control: the first stage ensures statistical fidelity through multivariate joint distribution modeling (20 / 20 features pass the KS test, and the relevant structure retention rate is 92.7%); the second stage ensures ecological rationality through MT-TKG compliance filtering (34.3% pass rate, rejecting 65.7% of candidate scenarios that violate ecological constraints).

[0089] like Figures 2-5 As shown, taking a subtropical drinking water source reservoir (Laohutan Reservoir in Huzhou City, Zhejiang Province, monitored at a 4-hour frequency for 10 years) as an example, this reservoir has three monitoring systems: an online automatic water quality station (14 parameters including Chl-a, approximately 24,000 data points every 4 hours, from 2016 to 2025), an automatic weather station / ERA5-Land (temperature, precipitation, wind speed, etc., daily frequency), and a hydrological station (water level, inflow, etc., daily frequency). The data is divided into three periods: training period (April 2016 to December 2020), validation period (January 2021 to December 2022), and testing period (January 2023 to November 2025).

[0090] Step S1 implementation: MT-TKG construction.

[0091] The constructed MT-TKG V1.0 contains 37 nodes and 61 directed edges, including 22 edges from literature sources, 21 edges from data mining sources, and 18 edges from expert confirmation. The average confidence level is 0.73, and there are 43 high-confidence edges (C≥0.7). Typical edges are shown in the table below:

[0092] Among them, D09+D10 constitutes a "suppression followed by enhancement" dual-effect pathway for rainstorms (short-term dilution and inhibition → medium-term non-point source nutrient promotion, with a time lag of 3-4 days, TP increase of +15%, p=0.002). E07+E08b constitutes a similar dual-effect pathway for inflow. E09+E11 encodes the "double-edged sword" effect of filter feeding by silver carp and bighead carp.

[0093] Implementation of steps S2-S3: Synthetic scene generation and filtering.

[0094] 106 turning points (47 in spring / summer and 59 in autumn / winter) were extracted from the training data (10,139 records). Gaussian copula models were fitted to each of the two seasonal layers (spring / summer and autumn / winter), with each layer amplified by 50-fold, generating a total of 5,300 candidate synthetic scenes. Knowledge compliance filtering (K≥0.6 threshold) was performed using 45 available edges in the MT-TKG dataset; 1,817 scenes passed the filter (pass rate 34.3%). Synthetic data quality validation: all 20 core features passed the KS test (p>0.05), and the Frobenius distance preservation rate for the relevant structures was 92.7%.

[0095] Step S4 implementation: Fusion training.

[0096] The LightGBM gradient boosting decision tree was used as the early warning model, employing 168-dimensional features (including current value, 6-step time lag, 48-hour rolling statistics, rate of change, and calendar features). The training set contained 8907 real records, and after KCSS (Knowledge Constrained Synthetic Sampling) enhancement, it contained 10724 records. During the testing period (2023-2025), the L2 F1 score (harmonic mean of precision and recall, used to comprehensively measure the model's ability to identify outbreak-level events) for 24-hour predictions improved from 0.241 to 0.817, and the L2 recall (the proportion of real outbreak-level events successfully predicted by the model) improved from 14.2% to 78.0%.

[0097] The evaluation results during the independent testing period (2023-2025) are as follows:

[0098] In the crucial 24-hour prediction time domain, KCSS enhancement improved the F1 score for high-risk algal blooms (L2) from 0.241 to 0.817 (+0.576), and the L2 recall from 14.2% to 78.0% (+63.8 percentage points). The effect of KCSS enhancement is precisely concentrated in the mid-to-long time domain (24-72h), where the baseline model is weakest, confirming the design goal of "targeted supplementation of scarce training samples".

[0099] This embodiment provides a water source algae early warning device 1 based on mechanism mapping.

[0100] like Figure 6 As shown, the device includes a map construction module 101, a sample generation module 102, a compliance check module 103, a model training module 104, and an early warning output module 105.

[0101] The graph construction module 101 is used to construct a time-series graph of mechanistic knowledge based on historical monitoring data during the training period and aquatic ecology literature. The time-series graph of mechanistic knowledge includes nodes, directed edges, and dual-effect path encoding. The graph construction module 101 may further include a data access unit, a knowledge extraction unit, a relation encoding unit, and a graph storage unit.

[0102] The sample generation module 102 is used to extract algal bloom level transition event windows based on historical monitoring data during the training period, and to construct a transition event sample set using monitoring data from each window. The sample set is then divided into corresponding subsets according to stratification conditions. A multivariate joint distribution model is fitted to each subset, and candidate synthetic transition event scenarios are generated by sampling from the multivariate joint distribution model. The sample generation module 102 may further include a window extraction unit, a sample stratification unit, a joint distribution fitting unit, and a sampling unit.

[0103] The compliance check module 103 is used to perform knowledge compliance checks on each candidate synthetic turning point scenario based on the mechanistic knowledge time series graph, retaining synthetic scenarios that conform to the graph constraints. The compliance check module 103 may include an applicable edge identification unit, a direction determination unit, a time delay determination unit, and a score filtering unit.

[0104] The model training module 104 is used to merge historical monitoring data and synthetic scenario data during the training period to form an enhanced training set. A machine learning model is then used to train the enhanced training set to obtain an early warning model for predicting the level of algal blooms in a future specified time domain. The model training module 104 may include a feature construction unit, a sample balancing unit, a model training unit, and a model evaluation unit.

[0105] The early warning output module 105 is used to acquire real-time online water quality monitoring data and meteorological and hydrological data, and outputs algal bloom level prediction results and graded alarm signals through the early warning model. The early warning output module 5 may further include a data preprocessing unit, a model inference unit, a level mapping unit, and an alarm release unit.

[0106] In some optional embodiments, in order to support the execution logic of the above modules, the device may be further configured with the following specific internal unit architecture: The graph construction module 101 serves as the knowledge foundation of the system. Internally, through the collaborative work of the data access unit, knowledge extraction unit, relation encoding unit, and graph storage unit, it can achieve the structured encapsulation of ecological mechanisms.

[0107] Specifically, the data access unit is responsible for connecting to heterogeneous data sources and cleaning historical monitoring data and aquatic ecology literature during the training period; the knowledge extraction unit is used to define the nodes, including environmental factor nodes, biological factor nodes, event nodes, and state nodes, and to define the directed edges connecting the nodes and their carried relation types, time lag windows, confidence weights, and source identifiers; the relation encoding unit is specifically used to encode dual-effect paths to support the opposite effects of the same driving factor in different time lag windows (e.g., a rainstorm event inhibits algae in [0, 48h] but promotes algae in [72, 168h]); the graph storage unit persistently stores the constructed mechanism knowledge time series graph MT-TKG for subsequent modules to call.

[0108] The sample generation module 102 is responsible for data augmentation, and its internal workflow logic is as follows: The window extraction unit selects the time step during the training period when the algal bloom level changes from the safe level to the low-risk level or above as the starting time step, and uses the starting time step as the reference to trace forward for a first preset time and backward for a second preset time as the turning event window; the sample stratification unit divides the sample set into spring / summer layer and autumn / winter layer according to the season; the joint distribution fitting unit fits Gaussian copula models on the spring / summer layer and autumn / winter layer respectively; the sampling unit samples from the Gaussian copula model to generate candidate synthetic turning event scenarios.

[0109] The compliance inspection module 103 serves as the core checkpoint for filtering fake samples, ensuring the ecological rationality of the synthesized data through multi-layered verification logic.

[0110] For each candidate synthetic turning event scenario, the applicable edge identification unit retrieves the set of triggered applicable edges from the graph storage unit; the direction determination unit checks whether the change direction of the target node in the candidate synthetic turning event scenario conforms to the relationship type constraint of the corresponding directed edge; the time lag determination unit checks whether the time relationship falls within the time lag window of the corresponding directed edge; the score filtering unit calculates the knowledge compliance score based on the ratio of the number of applicable edges that pass the test to the total number of identified applicable edges, and retains synthetic scenarios with scores not lower than a set threshold.

[0111] Model training module 104 is responsible for building predictive capabilities, specifically including: The feature construction unit is used to construct current value features, multi-step time lag features, rolling statistical features, rate of change features, and calendar features based on the augmented training set; the sample balancing unit is used to handle the class imbalance problem in the training set; the model training unit is used to train the augmented training set using machine learning models (such as gradient boosting decision trees or long short-term memory networks); and the model evaluation unit is used to evaluate the model performance on the validation set and output the final early warning model for deployment.

[0112] The early warning output module 105 realizes a closed loop for business applications.

[0113] The data preprocessing unit standardizes the acquired real-time online water quality monitoring data and meteorological and hydrological data; the model inference unit inputs the processed data into the early warning model and outputs the predicted algal bloom level after a set time domain; the level mapping unit maps the predicted values ​​to a safe level, a low-risk level, or a high-risk level; and the alarm release unit triggers the corresponding graded alarm signal according to the final level and pushes it to the user terminal.

[0114] This embodiment describes an electronic device. For example... Figure 7 As shown, the electronic device 2 includes a processor 202, a memory, and a network interface 205 connected via a device bus 201. The memory may include a storage medium 203 and internal memory 204.

[0115] The memory is used to store programs, model parameters, historical running data, scene sample data, device parameters, and optimization results.

[0116] The storage medium 203 may store an operating system 2031 and a computer program 2032. When the computer program 2032 is executed, it enables the processor 202 to execute a water source algae early warning method based on mechanistic maps.

[0117] The processor 202 provides computing and control capabilities to support the operation of the entire electronic device 2.

[0118] The internal memory 204 provides an environment for the operation of the computer program 2032 in the storage medium 203. When the computer program 2032 is executed by the processor 202, the processor 202 can execute a water source algae early warning method based on mechanism maps.

[0119] The network interface 205 is used for network communication, such as providing data information transmission. Those skilled in the art will understand that the structure shown in the figures is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device 2 to which the present invention is applied. The specific electronic device 2 may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.

[0120] The processor 202 is used to run a computer program 2032 stored in a memory to implement the water source algae early warning method based on mechanism map disclosed in the embodiments of the present invention.

[0121] Those skilled in the art will understand that the illustrated embodiments of the computer device do not constitute a limitation on the specific configuration of the computer device. In other embodiments, the computer device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. For example, in some embodiments, the computer device may include only memory and a processor. In such embodiments, the structure and function of the memory and processor are different from those illustrated. Figure 5 The embodiments shown are consistent and will not be described again here.

[0122] It should be understood that, in this embodiment of the invention, the processor 202 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0123] In another embodiment of the present invention, a computer-readable storage medium is provided. This computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein when executed by a processor, the computer program implements the water source algae early warning method based on mechanistic maps disclosed in the embodiments of the present invention.

[0124] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0125] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Units with the same function may be grouped into one unit. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, or may be electrical, mechanical, or other forms of connection.

[0126] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0127] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0128] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, a backend server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks.

[0129] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Any other corresponding changes and modifications made based on the technical concept of this application should be included within the scope of protection of the claims of this application.

Claims

1. A method for early warning of algae in a water source based on a mechanism map, characterized in that, include: A time-series graph of mechanistic knowledge is constructed based on historical monitoring data during the training period and aquatic ecology literature. The time-series graph of mechanistic knowledge includes nodes, directed edges, and dual-effect path encoding. Based on the historical monitoring data during the training period, the transition event window of the algal bloom level is extracted, and the monitoring data of each window constitutes the transition event sample set. The sample set is divided into corresponding subsets according to the stratification condition. A multivariate joint distribution model is fitted on each subset, and candidate synthetic transition event scenarios are generated by sampling from the multivariate joint distribution model. Based on the aforementioned mechanism knowledge time series graph, a knowledge compliance check is performed on each of the aforementioned candidate synthesis turning point events, and synthesis scenarios that conform to the graph constraints are retained. The historical monitoring data during the training period is combined with the synthetic scenario data to form an enhanced training set. The enhanced training set is then trained using a machine learning model to obtain an early warning model for predicting the level of algal blooms in a future time domain. The system acquires real-time online water quality monitoring data and meteorological and hydrological data, and outputs algal bloom level prediction results and graded alarm signals through the early warning model.

2. The method for early warning of algae in water sources based on mechanistic maps according to claim 1, characterized in that, The mechanism knowledge time series graph constructed based on historical monitoring data during the training period and aquatic ecology literature includes nodes, directed edges, and dual-effect path encoding, including: The nodes are defined as environmental factor nodes, biological factor nodes, event nodes, and state nodes. The environmental factor nodes include water temperature, total phosphorus, total nitrogen, ammonia nitrogen, dissolved oxygen, pH, conductivity, and turbidity. The biological factor nodes include cyanobacteria density, diatom density, green algae density, and cryptophyte density. The event nodes include heavy rainfall, high inflow rate, and low water level. The state nodes include algal bloom level: safe level, low risk level, and high risk level. Define directed edges connecting nodes. Each directed edge carries a relation type, a time lag window, a confidence weight, and a source identifier. The relation type includes facilitator, inhibitor, co-occurrence, and predecessor. The time lag window is in hours. The confidence weight ranges from 0 to 1. The source identifier includes literature, data mining, and experts. Encode dual-effect paths to support the opposite effects of the same driving factor within different time lag windows.

3. The method for early warning of algae in water sources based on mechanistic maps according to claim 2, characterized in that, Based on the historical monitoring data during the training period, the process involves extracting algal bloom level transition event windows, constructing a transition event sample set using monitoring data from each window, dividing the sample set into corresponding subsets according to stratification conditions, fitting a multivariate joint distribution model to each subset, and sampling from the multivariate joint distribution model to generate candidate synthetic transition event scenarios, including: The time step in which the algal bloom level changes from the safe level to the low-risk level or above during the training period is selected as the starting time step. Based on the starting time step, the first preset duration is traced forward and the second preset duration is traced backward as the turning point event window. The sample set was divided into spring / summer and autumn / winter layers according to the season, and a Gaussian copula model was fitted to each layer to generate candidate synthetic turning point scenarios. 4.The mechanism map-based water source algae early warning method according to claim 2, characterized in that, Based on the aforementioned mechanism knowledge time series graph, a knowledge compliance check is performed on each of the candidate synthetic turning point scenarios, retaining synthetic scenarios that conform to the graph constraints, including: Obtain the set of applicable edges triggered in each candidate synthetic turning event scenario; Examine whether the direction of change of the target node in the candidate synthetic turning event scenario conforms to the relation type constraint of the corresponding directed edge, and whether the temporal relation falls within the time lag window of the corresponding directed edge; A knowledge compliance score is calculated based on the ratio of the number of applicable edges that pass the test to the total number of applicable edges in the set, and synthetic scenarios with scores not lower than a set threshold are retained.

5. The mechanism map-based water source algae early warning method according to claim 1, characterized in that, The machine learning model is a gradient boosting decision tree or a long short-term memory network, and the future time domain includes short-term time domain and medium-to-long-term time domain. When using gradient boosting decision trees, the input features include current value features, multi-step time lag features, rolling statistics features, rate of change features, and calendar features; The historical monitoring data during the training period includes online water quality monitoring data, meteorological data, and hydrological data. The online water quality monitoring data includes chlorophyll a and water quality parameters. The meteorological data includes air temperature, precipitation, and wind speed. The hydrological data includes water level and inflow.

6. The mechanism map-based water source algae early warning method according to claim 1, characterized in that, The mechanistic knowledge time series map is constructed from the time-delay cross-correlation scan results of the training period data, the seasonal stratification analysis results, the knowledge of aquatic ecology literature, and the knowledge confirmed by experts. It is used to correct counterintuitive relationships in the data mining results.

7. The method for early warning of algae in water sources based on mechanistic maps according to claim 2, characterized in that, The dual-effect path also includes high inflow rate inhibiting algae growth through the first edge within a third preset time window, and promoting algae growth through the second edge within a fourth preset time window; Silver carp and bighead carp filter feed inhibit algae of a preset particle size within the fifth preset time window via the first edge, and promote small cyanobacteria within the sixth preset time window via the second edge.

8. A mechanism map-based water source algae early warning device, characterized in that, include: The graph construction module is used to construct a time-series graph of mechanistic knowledge based on historical monitoring data during the training period and aquatic ecology literature. The time-series graph of mechanistic knowledge includes nodes, directed edges, and dual-effect path encoding. The sample generation module is used to extract the algal bloom level transition event window based on the historical monitoring data of the training period, and to form a transition event sample set with the monitoring data of each window. The sample set is divided into corresponding subsets according to the stratification conditions. A multivariate joint distribution model is fitted to each subset, and candidate synthetic transition event scenarios are generated by sampling from the multivariate joint distribution model. The compliance check module is used to perform knowledge compliance checks on each of the candidate synthetic turning point events based on the time-series graph of the mechanism knowledge, and retain synthetic scenarios that conform to the graph constraints. The model training module is used to merge the historical monitoring data of the training period with the synthetic scenario data to form an enhanced training set, and to train the enhanced training set using a machine learning model to obtain an early warning model for predicting the level of algal blooms in a future set time domain. The early warning output module is used to acquire real-time online water quality monitoring data and meteorological and hydrological data, and output the algal bloom level prediction results and graded alarm signals through the early warning model.

9. An electronic device, comprising: Including memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the water source algae early warning method based on mechanism maps as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the water source algae early warning method based on mechanism maps as described in any one of claims 1 to 7.