A method and system for assessing water environment risks of a river basin through big data analysis

By constructing a structured basic database of multi-source heterogeneous data, dividing risk scenarios into refined categories, calculating the spatiotemporal coupling correlation, setting risk thresholds, and assessing governance measures, the problems of data fragmentation and insufficient identification accuracy in watershed water environment risk assessment have been solved, achieving high-precision risk identification and differentiated governance.

CN122288360APending Publication Date: 2026-06-26CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF GEOSCIENCES (WUHAN)
Filing Date
2026-03-04
Publication Date
2026-06-26

Smart Images

  • Figure CN122288360A_ABST
    Figure CN122288360A_ABST
Patent Text Reader

Abstract

This invention discloses a watershed water environment risk assessment method and system based on big data analysis, belonging to the field of environmental risk monitoring technology. The key technical solutions include the following steps: collecting multi-source heterogeneous water environment data from different sub-watershed units within the watershed; processing the multi-source heterogeneous data to construct a basic watershed water environment database; based on the basic watershed water environment database, classifying watershed risk-sensitive scenario types, extracting characteristic indicator sets of water quality, hydrology, and pollution sources from each watershed risk-sensitive scenario type, and calculating the spatiotemporal coupling correlation between characteristic indicators; setting risk threshold ranges for characteristic indicators, filtering out abnormal indicator data exceeding the threshold range, and matching the risk-sensitive scenarios corresponding to the abnormal indicator data. The effect is to break down the format barriers and spatiotemporal resolution differences of data from different sources, avoiding data fragmentation and insufficient support capabilities, and providing a unified and reliable data foundation for subsequent risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental risk monitoring technology, and more specifically, to a method and system for watershed water environment risk assessment based on big data analysis. Background Technology

[0002] With the increasing demand for water environment governance in my country's river basins, traditional risk assessment methods are no longer sufficient to address the challenges posed by the complex and ever-changing river basin ecosystems. Current water environment management faces prominent problems such as fragmented data, rudimentary risk identification, and delayed governance decisions, severely restricting the accuracy and effectiveness of water environment risk prevention and control.

[0003] Current technologies for watershed water environment risk assessment largely rely on single-dimensional water quality monitoring data, failing to effectively integrate multi-source heterogeneous data from hydrology, pollution sources, meteorology, and land use. Inconsistent data formats and significant differences in spatiotemporal resolution among different sources create numerous data silos, resulting in insufficient basic data support and an inability to comprehensively reflect the dynamic changes in the watershed's water environment. Furthermore, existing methods tend to classify watershed risk scenarios in a coarse manner, failing to refine scenario segmentation based on sub-basins' ecological function positioning, population distribution, and industrial layout. The selection of feature indicators relies heavily on empirical judgment, lacking quantitative analysis of the spatiotemporal coupling relationships among water quality, hydrology, and pollution source indicators, leading to insufficient risk identification accuracy and difficulty in precisely locating high-risk areas and key driving factors. Summary of the Invention

[0004] In view of the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for watershed water environment risk assessment based on big data analysis.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A watershed water environment risk assessment method based on big data analysis, comprising the following steps:

[0007] Collect multi-source heterogeneous water environment data from different sub-basin units within the watershed, process the multi-source heterogeneous data, and construct a basic database of the watershed's water environment.

[0008] Based on the aforementioned watershed water environment database, watershed risk-sensitive scenario types are classified, and feature indicator sets of water quality, hydrology, and pollution sources are extracted from each watershed risk-sensitive scenario type. The spatiotemporal coupling correlation between feature indicators is then calculated.

[0009] Set risk threshold ranges for characteristic indicators, filter out abnormal indicator data that exceed the threshold range, match the risk-sensitive scenarios corresponding to the abnormal indicator data, and determine the high-risk areas and key influencing factor data of the watershed water environment.

[0010] Based on historical risk impact factor data, risk data impact assessment is conducted on key impact factor data of high-risk areas, the decay rate of risk indicators under different governance measures is predicted, and the governance cycle required for the risk level to drop to the safe range is determined to generate a governance cycle prediction dataset.

[0011] Based on the governance cycle and the watershed's ecological carrying capacity, differentiated risk management strategies are formulated, and a priority ranking of management strategies under multiple scenarios is generated.

[0012] Based on the priority ranking of the control strategies, a watershed water environment risk assessment report is output, which includes risk level distribution, estimated treatment cycle, and control strategy recommendations.

[0013] Preferably, the process involves collecting multi-source heterogeneous water environment data from different sub-basin units within the watershed, processing the multi-source heterogeneous data, and constructing a basic watershed water environment database. This specifically includes the following steps:

[0014] The multi-source heterogeneous data is collected, including water quality monitoring data, hydrological runoff data, pollution source emission data, meteorological data, and land use data.

[0015] The multi-source heterogeneous data is normalized, missing values ​​are filled in, and outliers are removed to unify the data format and spatiotemporal resolution.

[0016] A structured watershed water environment database is constructed based on sub-basin units and monitoring time dimensions.

[0017] Preferably, a structured watershed water environment database is constructed according to sub-basin units and monitoring time dimensions, specifically including the following steps:

[0018] Based on the natural confluence boundary of the watershed, the spatial topological relationships of each sub-watershed unit within the watershed are integrated, and the association relationships of each sub-watershed unit are determined to obtain the watershed unit association dataset;

[0019] Based on the frequency and temporal variation characteristics of watershed water environment monitoring, monitoring time levels are divided, and the monitoring behavior of each monitoring time level and the corresponding sub-watershed unit is associated to form a hierarchical watershed correlation coefficient.

[0020] A basic database of watershed water environment is formed by integrating watershed unit association datasets and hierarchical watershed association coefficients.

[0021] Preferably, based on the aforementioned watershed water environment database, the watershed risk-sensitive scenario types are classified, and feature indicator sets of water quality, hydrology, and pollution sources are extracted from each watershed risk-sensitive scenario type. The spatiotemporal coupling correlation between the feature indicators is then calculated. Specifically, this includes the following steps:

[0022] The risk-sensitive scenarios in the watershed are classified into drinking water source protection areas, ecological buffer zones, and pollution prevention and control zones according to ecological functional zones, population density, and industrial and agricultural layout.

[0023] Extract the characteristic indicator set of water quality, hydrology and pollution source in each risk-sensitive scenario type. The water quality indicators include COD, ammonia nitrogen and total phosphorus concentration. The hydrological indicators include runoff and flow velocity. The pollution source indicators include the number of sewage outlets and the amount of pollutants discharged.

[0024] The spatiotemporal coupling correlation degree is obtained by calculating the coupling correlation degree between feature indicators under different spatiotemporal dimensions.

[0025] Preferably, a risk threshold range for characteristic indicators is set, abnormal indicator data exceeding the threshold range is filtered out, and risk-sensitive scenarios corresponding to the abnormal indicator data are matched to determine high-risk areas and key influencing factor data of the watershed water environment. This specifically includes the following steps:

[0026] Based on national water environment quality standards and watershed ecological objectives, risk threshold ranges for each characteristic indicator are set.

[0027] Compare the indicator data in the basic database with the threshold range, and filter out abnormal indicator data that exceed the threshold range;

[0028] Based on the spatiotemporal coupling correlation, risk-sensitive scenarios corresponding to abnormal indicators are matched to locate high-risk areas;

[0029] Key influencing factor data is obtained by identifying the key influencing factors driving abnormal indicators based on high-risk areas.

[0030] Preferably, based on historical risk impact factor data, a risk data impact assessment is conducted on key impact factor data for high-risk areas to predict the decay rate of risk indicators under different governance measures, determine the governance cycle required for the risk level to drop to a safe range, and generate a governance cycle prediction dataset. This specifically includes the following steps:

[0031] Collect risk occurrence history, time series change data of key influencing factors and corresponding risk level evolution data of each sub-basin unit in history, and construct a historical risk factor time series correlation dataset;

[0032] The risk level and indicator decay rate are obtained by comparing the key influencing factor data of high-risk areas with the historical risk factor time series correlation dataset.

[0033] Multiple governance measures were simulated, including pollution source reduction, ecological restoration, and water conservancy regulation.

[0034] By inputting governance parameters for different scenarios, the system predicts the rate of decline of risk indicators, calculates the governance cycle required for the risk level to drop to the safe range, and generates a governance cycle prediction dataset.

[0035] Preferably, by combining the governance cycle with the watershed's ecological carrying capacity, differentiated risk management strategies are formulated, and a priority ranking of management strategies under multiple scenarios is generated, specifically including the following steps:

[0036] Collect watershed ecological carrying capacity data, which includes water body self-purification capacity data, soil adsorption capacity data, and vegetation coverage rate;

[0037] By combining the governance cycle prediction dataset with ecological carrying capacity data, an evaluation index system for management and control strategies is constructed.

[0038] Prioritize the management strategies for different types of risk-sensitive scenarios in different watersheds, and give priority to management strategies with short governance cycles and minimal ecological impact.

[0039] Preferably, based on the priority ranking of the control strategies, a watershed water environment risk assessment report is output. The watershed water environment risk assessment report includes risk level distribution, estimated treatment cycle, and control strategy recommendations, specifically including the following steps:

[0040] Integrate data on the distribution of high-risk areas, data on the estimated governance cycle, and data on the priority of control strategies;

[0041] Based on the sub-basin units, risk level distribution maps are drawn, key influencing factors and governance cycles are marked, and a visualized watershed water environment risk assessment report is output, clarifying the control measures and implementation sequence for each sensitive scenario.

[0042] A watershed water environment risk assessment system based on big data analysis includes:

[0043] Data acquisition module: Collects multi-source heterogeneous water environment data from different sub-basin units within the watershed, processes the multi-source heterogeneous data, and constructs a basic database of the watershed water environment;

[0044] Feature extraction module: Based on the watershed water environment basic database, the watershed risk-sensitive scenario types are divided, and feature index sets of water quality, hydrology, and pollution sources in each watershed risk-sensitive scenario type are extracted, and the spatiotemporal coupling correlation between feature indexes is calculated.

[0045] Anomaly screening module: Sets risk threshold ranges for feature indicators, filters out abnormal indicator data that exceed the threshold range, matches the risk-sensitive scenarios corresponding to the abnormal indicator data, and identifies high-risk areas and key influencing factor data of the watershed water environment;

[0046] Risk assessment module: Based on historical risk impact factor data, it conducts risk data impact assessment on key impact factor data of high-risk areas, predicts the decay rate of risk indicators under different governance measures, determines the governance cycle required for the risk level to drop to the safe range, and generates a governance cycle prediction dataset.

[0047] Strategy formulation module: Combining the governance cycle and the watershed's ecological carrying capacity, formulate differentiated risk management strategies and generate a priority ranking of management strategies in multiple scenarios;

[0048] Report generation module: Based on the priority of the control strategies, outputs a watershed water environment risk assessment report, which includes risk level distribution, estimated treatment cycle, and control strategy recommendations.

[0049] Compared with existing technologies, this invention has the following beneficial effects: By processing multi-source heterogeneous data on water quality, hydrology, pollution sources, meteorology, and land use, a structured basic database covering the entire basin, all time periods, and multiple dimensions is constructed. This breaks down the format barriers and spatiotemporal resolution differences of data from different sources, avoiding data fragmentation and insufficient support capabilities, and providing a unified and reliable data foundation for subsequent risk assessment. By combining the ecological function positioning, population distribution, and industrial layout of sub-basins, refined risk-sensitive scenarios for drinking water source protection areas and pollution control zones are delineated. Furthermore, the spatiotemporal coupling between water quality, hydrology, and pollution source indicators is calculated. The correlation degree quantifies the strength of mutual influence between indicators, accurately identifies high-risk areas and key driving factors, and avoids problems such as coarse scenario division, subjective indicator selection, and low risk identification accuracy, significantly improving the spatial accuracy and factor targeting of risk identification. By using the watershed ecological carrying capacity as a constraint, differentiated management and control strategies are formulated for different risk scenarios and high-risk areas, and the strategy priority is completed through multi-dimensional evaluation, avoiding the problems of strategy mismatch with ecological carrying capacity and poor implementation. This achieves refined management and control by region, scenario, and level, ensuring governance efficiency while avoiding excessive intervention and additional pressure on the ecosystem. Attached Figure Description

[0050] Figure 1 This invention provides a schematic diagram illustrating the steps of a watershed water environment risk assessment method based on big data analysis, as provided in an embodiment of the invention.

[0051] Figure 2 This invention provides a schematic diagram illustrating the steps involved in generating a governance cycle prediction dataset in a watershed water environment risk assessment method based on big data analysis, as described in an embodiment of the present invention.

[0052] Figure 3 This is a schematic diagram of a watershed water environment risk assessment system based on big data analysis, provided as an embodiment of the present invention. Detailed Implementation

[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0055] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0056] Reference Figures 1-3 As shown.

[0057] Example 1 further illustrates the watershed water environment risk assessment method and system based on big data analysis proposed in this invention.

[0058] A watershed water environment risk assessment method based on big data analysis, comprising the following steps:

[0059] Collect multi-source heterogeneous water environment data from different sub-basin units within the watershed, process the multi-source heterogeneous data, and construct a basic database of the watershed's water environment.

[0060] Based on the aforementioned watershed water environment database, watershed risk-sensitive scenario types are classified, and feature indicator sets of water quality, hydrology, and pollution sources are extracted from each watershed risk-sensitive scenario type. The spatiotemporal coupling correlation between feature indicators is then calculated.

[0061] Set risk threshold ranges for characteristic indicators, filter out abnormal indicator data that exceed the threshold range, match the risk-sensitive scenarios corresponding to the abnormal indicator data, and determine the high-risk areas and key influencing factor data of the watershed water environment.

[0062] Based on historical risk impact factor data, risk data impact assessment is conducted on key impact factor data of high-risk areas, the decay rate of risk indicators under different governance measures is predicted, and the governance cycle required for the risk level to drop to the safe range is determined to generate a governance cycle prediction dataset.

[0063] Based on the governance cycle and the watershed's ecological carrying capacity, differentiated risk management strategies are formulated, and a priority ranking of management strategies under multiple scenarios is generated.

[0064] Based on the priority ranking of the control strategies, a watershed water environment risk assessment report is output, which includes risk level distribution, estimated treatment cycle, and control strategy recommendations.

[0065] Collecting multi-source heterogeneous water environment data from different sub-basin units within the watershed, processing the multi-source heterogeneous data, and constructing a basic watershed water environment database specifically includes the following steps:

[0066] The multi-source heterogeneous data is collected, including water quality monitoring data, hydrological runoff data, pollution source emission data, meteorological data, and land use data.

[0067] The multi-source heterogeneous data is normalized, missing values ​​are filled in, and outliers are removed to unify the data format and spatiotemporal resolution.

[0068] A structured watershed water environment database is constructed based on sub-basin units and monitoring time dimensions.

[0069] First, for the target watershed, various types of multi-source heterogeneous water environment data are systematically collected according to pre-divided sub-watershed units. For example, in a plain river network watershed, real-time monitoring data of pH, COD, and ammonia nitrogen from automatic water quality monitoring stations are collected simultaneously; daily runoff and water level data from hydrological stations; daily discharge and pollutant concentration data from industrial wastewater outlets; daily rainfall and temperature data from meteorological stations; and land use type data for cultivated land, urban areas, and water bodies obtained through satellite remote sensing interpretation. These data come from different sources, have different formats, and are collected at different frequencies. For example, water quality data is collected hourly, hydrological data daily, and land use data annually, thus forming a heterogeneous characteristic.

[0070] The collected multi-source heterogeneous data then undergoes standardized preprocessing. First, the data is normalized using existing technologies, such as mapping different magnitudes of water quality concentration, runoff, and pollutant emissions to a 0-1 range to eliminate dimensional differences and prevent significant numerical magnitude discrepancies from affecting subsequent analysis. Next, missing values ​​are imputed. For example, if ammonia nitrogen data for a particular day is missing, the average of data from the same period seven days before and after that monitoring point is used for imputation, or relevant data from adjacent stations are used for extrapolation to ensure the continuity of the data sequence. Finally, outliers are removed, for example, using the 3σ criterion to identify and delete data that significantly deviates from the normal fluctuation range, such as abnormally high values ​​at water quality monitoring points due to equipment malfunction. Finally, the format and spatiotemporal resolution of all data are standardized, for example, all data is converted to CSV format and uniformly adjusted to a daily time resolution to ensure that data from different sources are aligned and matched in both time and space.

[0071] After data preprocessing, the solution constructs a structured watershed water environment database based on two core dimensions: sub-basin units and monitoring time. For example, the system creates an independent data table for each sub-basin, storing data on water quality, hydrology, pollution sources, meteorology, and land use from all monitoring points within that sub-basin on different dates, forming a structured data matrix of sub-basins, monitoring time, and multi-dimensional indicators. This allows for rapid retrieval of complete water environment data for any sub-basin at any time period during subsequent analysis, providing unified and reliable data support for subsequent steps such as risk scenario classification and spatiotemporal correlation analysis.

[0072] A structured watershed water environment database is constructed based on sub-basin units and monitoring time dimensions, specifically including the following steps:

[0073] Based on the natural confluence boundary of the watershed, the spatial topological relationships of each sub-watershed unit within the watershed are integrated, and the association relationships of each sub-watershed unit are determined to obtain the watershed unit association dataset;

[0074] Based on the frequency and temporal variation characteristics of watershed water environment monitoring, monitoring time levels are divided, and the monitoring behavior of each monitoring time level and the corresponding sub-watershed unit is associated to form a hierarchical watershed correlation coefficient.

[0075] A basic database of watershed water environment is formed by integrating watershed unit association datasets and hierarchical watershed association coefficients.

[0076] First, the scheme integrates the spatial topological relationships of each sub-basin unit based on the natural confluence boundaries of the watershed, determining the correlations between sub-basins and forming a watershed unit correlation dataset. Taking a mountainous watershed as an example, the entire watershed is divided into three units: an upstream mountainous sub-basin, a midstream hilly sub-basin, and a downstream plain sub-basin. The system clarifies, based on the natural flow direction of the river, that the water from the upstream sub-basin flows into the midstream sub-basin, and vice versa. This upstream-downstream confluence relationship is transformed into structured topological data, such as using an association matrix to record the water flow transmission relationship from sub-basin A to sub-basin B, and using spatial coordinates to mark the boundaries and confluence locations of each sub-basin. This facilitates subsequent analysis by clearly tracing the migration paths of pollutants between different sub-basins, providing a foundation for spatial risk transmission analysis.

[0077] Secondly, the scheme will divide the monitoring time levels based on the frequency and temporal variation characteristics of watershed water environment monitoring, and associate each monitoring time level with the monitoring behavior of the corresponding sub-watershed unit to form a hierarchical watershed correlation coefficient. For example, for hourly high-frequency water quality automatic monitoring data, daily frequency manual inspection data, and annual frequency land use data, the system will divide the data into five time levels: hourly, daily, weekly, monthly, and annual. Each time level will be bound to the monitoring behavior of the corresponding sub-watershed. For example, if the automatic water quality monitoring station in the upstream sub-watershed generates a set of data every hour, it will be associated with the hourly time level; if the manual inspection in the midstream sub-watershed is carried out once a week, the data will be associated with the weekly time level. To quantify the strength of this correlation, the system will calculate the hierarchical watershed correlation coefficient, for example, using the formula: , where w m,n f represents the correlation coefficient between the m-th time level and the n-th sub-basin. m,n This refers to the monitoring frequency of the sub-basin at this time level. Taking the upstream sub-basin as an example, if the hourly monitoring frequency is 168 times / week, the daily frequency is 7 times / week, and the weekly frequency is 1 time / week, then the correlation coefficient for the hourly time level is 168 / (168+7+1)≈0.955, for the daily level it is 7 / 176≈0.0398, and for the weekly level it is 1 / 176≈0.0057, thus reflecting the monitoring coverage intensity of the sub-basin at different time levels.

[0078] Finally, the solution integrates the watershed unit association dataset with the hierarchical watershed association coefficients to form a complete watershed water environment database. The system integrates the spatial topological data of each sub-watershed, monitoring data at different time levels, and corresponding association coefficients into a unified structured data model. For example, a complete record in the database will include the spatial coordinates of the upstream sub-watershed, hourly water quality monitoring data, the corresponding association coefficient, and the confluence relationship between that sub-watershed and the midstream sub-watershed. This integration not only makes the data spatially traceable and temporally correlated but also reflects the weight of different monitoring behaviors through association coefficients, ensuring that subsequent analysis prioritizes high-frequency, highly correlated monitoring data, thus improving the accuracy and efficiency of risk assessment.

[0079] Based on the aforementioned watershed water environment database, the watershed risk-sensitive scenario types are classified, and feature indicator sets of water quality, hydrology, and pollution sources are extracted from each watershed risk-sensitive scenario type. The spatiotemporal coupling correlation between the feature indicators is calculated, specifically including the following steps:

[0080] The risk-sensitive scenarios in the watershed are classified into drinking water source protection areas, ecological buffer zones, and pollution prevention and control zones according to ecological functional zones, population density, and industrial and agricultural layout.

[0081] Extract the characteristic indicator set of water quality, hydrology and pollution source in each risk-sensitive scenario type. The water quality indicators include COD, ammonia nitrogen and total phosphorus concentration. The hydrological indicators include runoff and flow velocity. The pollution source indicators include the number of sewage outlets and the amount of pollutants discharged.

[0082] The spatiotemporal coupling correlation degree is obtained by calculating the coupling correlation degree between feature indicators under different spatiotemporal dimensions.

[0083] First, the solution will classify the entire watershed into risk-sensitive scenarios based on its ecological function, population density, and industrial and agricultural layout. Taking a typical watershed in southern China as an example, the system will designate the area surrounding centralized urban water intakes as a drinking water source protection zone, the vegetation buffer zone surrounding the water source protection zone as an ecological buffer zone, and industrial parks and large-scale farmland areas as pollution prevention and control zones. This classification allows for more targeted risk analysis and avoids the lack of accuracy that can result from a uniform assessment.

[0084] After scenario segmentation, the solution extracts a set of three-dimensional feature indicators for water quality, hydrology, and pollution sources for each type of risk-sensitive scenario. For drinking water source protection areas, water quality indicators focus on COD, ammonia nitrogen, and total phosphorus concentrations, as these indicators are directly related to drinking water safety. Hydrological indicators select runoff and flow velocity, as these two parameters determine the diffusion and dilution capacity of pollutants at the water source. Pollution source indicators count the number of sewage outlets and pollutant emissions to assess the external pollution pressure. For ecological buffer zones, water quality indicators also include COD, ammonia nitrogen, and total phosphorus concentrations to assess the buffer zone's reduction effect on upstream pollutants. Hydrological indicators such as runoff and flow velocity reflect the hydrodynamic conditions of the buffer zone. Pollution source indicators focus on the number of sewage outlets and pollutant emissions related to non-point source pollution. For pollution control zones, water quality indicators still primarily focus on COD, ammonia nitrogen, and total phosphorus concentrations to measure the pollution intensity within the area. Hydrological indicators such as runoff and flow velocity are used to analyze the migration patterns of pollutants. Pollution source indicators focus on counting the number of industrial sewage outlets and agricultural drainage outlets, as well as the corresponding pollutant emissions.

[0085] Next, the scheme will calculate the coupling correlation between these characteristic indicators under different spatiotemporal dimensions to quantify the strength of their mutual influence. Taking the time dimension as an example, changes in runoff directly affect the concentrations of pollutants such as COD and ammonia nitrogen during dry and wet seasons. During wet seasons, the large runoff dilutes pollutants, typically lowering their concentrations, while during dry seasons, the small runoff makes pollutant concentrations more likely to increase. Spatially, increased discharge from pollution control outlets in pollution prevention zones leads to a corresponding rise in water quality indicators in downstream ecological buffer zones and drinking water source protection areas.

[0086] The spatiotemporal coupling correlation degree can be calculated using an improved grey relational analysis method, with the specific formula as follows:

[0087] Where, r ij (t,s) represents the index x at time t and spatial location s. i With x j The coupling correlation is denoted by ζ, which is the resolution coefficient and is typically set to 0.5. For example, in a pollution control zone, when calculating the correlation between runoff x1 and COD concentration x2, if the runoff is small and the COD concentration is high during the dry season, the absolute value of the difference between the two is large, and the correlation will approach 0.3; while during the wet season, the runoff is large and the COD concentration is low, the absolute value of the difference is small, and the correlation will approach 0.8. Through such calculations, the response relationship between indicators under different time and space conditions can be clearly quantified, providing a quantitative basis for subsequent risk identification.

[0088] The process involves setting risk threshold ranges for characteristic indicators, filtering out abnormal indicator data that exceed these ranges, matching the risk-sensitive scenarios corresponding to the abnormal indicator data, and identifying high-risk areas and key influencing factors for the watershed's water environment. This includes the following steps:

[0089] Based on national water environment quality standards and watershed ecological objectives, risk threshold ranges for each characteristic indicator are set.

[0090] Compare the indicator data in the basic database with the threshold range, and filter out abnormal indicator data that exceed the threshold range;

[0091] Based on the spatiotemporal coupling correlation, risk-sensitive scenarios corresponding to abnormal indicators are matched to locate high-risk areas;

[0092] Key influencing factor data is obtained by identifying the key influencing factors driving abnormal indicators based on high-risk areas.

[0093] First, the plan sets corresponding risk threshold ranges for each characteristic indicator based on national water environmental quality standards and watershed ecological goals. Taking drinking water source protection areas as an example, the national surface water environmental quality standards stipulate that the COD concentration in drinking water sources must be controlled below 15 mg / L, the ammonia nitrogen concentration below 0.5 mg / L, and the total phosphorus concentration below 0.02 mg / L. The plan will further refine these standards into tiered threshold ranges, taking into account the watershed's own ecological goals. For example, COD concentration is divided into a safe range of 0 to 15 mg / L, a warning range of 15 to 20 mg / L, a risk range of 20 to 30 mg / L, and a high-risk range above 30 mg / L. For hydrological indicators such as runoff, the plan will consider the historical flow characteristics of the watershed during dry and wet seasons, setting a safe range of 70% to 130% of the multi-year average runoff. Runoff below 70% or above 130% will be classified as warning or higher. For the number of sewage outlets in the pollution source indicators, a threshold will be set according to the regional environmental capacity. For example, the safe range for the number of industrial sewage outlets in the pollution prevention and control zone is 0 to 5. If it exceeds this range, it will be judged as abnormal.

[0094] The system then compares the real-time and historical indicator data in the watershed's basic water environment database with the set threshold ranges one by one, filtering out all abnormal indicator data that exceed the threshold ranges. For example, in the monitoring data of drinking water source protection areas, if the ammonia nitrogen concentration reaches 0.8 mg / L for three consecutive days, exceeding the safety threshold of 0.5 mg / L, it will be marked as abnormal indicator data; in pollution control areas, if the runoff is only 50% of the multi-year average, and the number of sewage outlets reaches 8, these two indicators will also be identified as abnormal.

[0095] Then, based on the previously calculated spatiotemporal coupling correlation, the system will accurately match abnormal indicator data with corresponding risk-sensitive scenarios to pinpoint high-risk areas. For example, if the number of sewage outlets and the amount of pollutants discharged in a pollution control zone both exceed the threshold, and the spatiotemporal coupling correlation shows that the pollutant discharge in this zone has a correlation of 0.85 with the COD concentration in the downstream ecological buffer zone, the system will determine that the pollution control zone and the downstream ecological buffer zone are high-risk areas. As another example, if the ammonia nitrogen concentration in a drinking water source protection area is abnormal, and the correlation shows that this anomaly is highly correlated with the discharge volume of sewage outlets in the upstream pollution control zone, both the upstream pollution control zone and the drinking water source protection area will be designated as high-risk areas.

[0096] Finally, the system analyzes the spatiotemporal coupling correlation and contribution of various indicators within high-risk areas to identify the key influencing factors driving abnormal indicators. The degree of influence of each factor can typically be quantified using a contribution calculation formula, as follows: , where C k r represents the contribution of the k-th factor. kj(t,s) represents the coupling correlation between factor k and anomaly index j at time t and space s, x k (t,s) represents the actual value of factor k, r ij (t,s) represents the coupling correlation between influencing factor i and anomaly index j at time t and spatial location s. It quantifies the strength of their mutual influence in the spatiotemporal dimensions, and its value typically ranges from 0 to 1. k0 x is the safety threshold for factor k. i0 , is the safety threshold benchmark value for feature index i.

[0097] For example, in high-risk areas, the calculated contribution of pollutant discharge from sewage outlets is 0.62, the contribution of runoff is 0.25, and the contribution of land use type is 0.13. In this case, pollutant discharge from sewage outlets can be identified as the key influencing factor driving water quality anomalies in the area. This facilitates the system's accurate location of high-risk areas and clarifies the core driving factors, providing a clear basis for subsequent risk assessment and control strategy development.

[0098] Based on historical risk impact factor data, a risk data impact assessment is conducted on key impact factor data for high-risk areas. The decay rate of risk indicators under different governance measures is predicted, and the governance cycle required for the risk level to drop to a safe range is determined, generating a governance cycle prediction dataset. The specific steps include:

[0099] Collect risk occurrence history, time series change data of key influencing factors and corresponding risk level evolution data of each sub-basin unit in history, and construct a historical risk factor time series correlation dataset;

[0100] The risk level and indicator decay rate are obtained by comparing the key influencing factor data of high-risk areas with the historical risk factor time series correlation dataset.

[0101] Multiple governance measures were simulated, including pollution source reduction, ecological restoration, and water conservancy regulation.

[0102] By inputting governance parameters for different scenarios, the system predicts the rate of decline of risk indicators, calculates the governance cycle required for the risk level to drop to the safe range, and generates a governance cycle prediction dataset.

[0103] First, the solution systematically collects historical data on the occurrence of risks in each sub-basin unit, the time-series changes of key influencing factors, and the evolution of corresponding risk levels, constructing a historical risk factor time-series correlation dataset. Taking the southern river basin as an example, the system collects records of water quality exceedance events in each sub-basin over the past ten years, monthly changes in key factors such as pollutant discharge from sewage outlets and runoff, and the evolution sequence of risk levels from safe to high risk. This data will be integrated into a structured time-series dataset, for example, by annual, quarterly, and monthly time dimensions, linking the changes in risk factors and risk levels in each sub-basin to form a correlation map between risk factors and risk levels.

[0104] Next, the solution compares the key influencing factor data of the current high-risk area with historical risk factor time-series correlation datasets to obtain the risk level and indicator decay rate under similar scenarios. For example, if the current daily pollutant discharge from the sewage outlets in the pollution control zone is 20 tons and the COD concentration is 35 mg / L, the system will search for historical events in the historical dataset where the discharge volume and COD concentration are in a similar range. It was found that five years ago, a similar risk scenario occurred in this area with a daily discharge volume of 22 tons and a COD concentration of 38 mg / L. At that time, after pollution source reduction measures were implemented, the COD concentration decayed at a rate of 0.8 mg / L per day, eventually reaching the safe threshold of 15 mg / L in 15 days. Through this comparison, the system can quickly obtain the baseline rate of indicator decay under similar scenarios, providing an initial reference for subsequent predictions.

[0105] Subsequently, the plan will set up various simulation scenarios for governance measures, including three core measures: pollution source reduction, ecological restoration, and water conservancy regulation. For the pollution source reduction scenario, different reduction ratios will be set, such as reducing pollutant emissions from sewage outlets by 30%, 50%, and 70%. For the ecological restoration scenario, different restoration scales will be set, such as constructing 10 hectares or 20 hectares of artificial wetlands. For the water conservancy regulation scenario, different ecological water replenishment flows will be set, such as replenishing 50,000 cubic meters or 100,000 cubic meters of water per day. Each scenario corresponds to specific governance parameters, which will be used as input conditions into the prediction model.

[0106] After inputting governance parameters for different scenarios, the system predicts the decay rate of risk indicators based on a machine learning model or a hydrodynamic-water quality coupling model trained on historical data, and calculates the governance cycle required for the risk level to drop to a safe range. The decay rate of risk indicators can be quantified using an exponential decay model, with the following formula: Where x(t) is the concentration of the risk indicator at time t, x0 is the initial concentration of the risk indicator, k is the decay rate coefficient, and t is the treatment time. Taking a high-risk area as an example, the initial COD concentration x0 = 35 mg / L, and the safety threshold is 15 mg / L. Under the scenario of a 50% reduction in pollution sources, the calculated decay rate coefficient k = 0.06 days. −1 Substituting into the formula to solve for t: 15 = 35 × e −0.06t Solving for t, we get t ≈ 14.5 days, meaning that in this scenario, it would take approximately 15 days for the risk level to decrease to a safe range. In an ecological restoration scenario, if a 20-hectare artificial wetland is constructed, the calculated attenuation rate coefficient k = 0.04 days. −1 Substituting into the formula, we get t≈21.6 days, which means it takes approximately 22 days. In a water management scenario, with a daily water replenishment of 100,000 cubic meters, the attenuation rate coefficient k=0.05 days. −1 Solving for t, we get t ≈ 17.3 days, which means it will take approximately 18 days. Through this type of calculation, the system generates a dataset of estimated governance cycles under different governance scenarios, clearly presenting the governance efficiency of each measure and providing a quantitative basis for the formulation of subsequent control strategies.

[0107] Based on the governance cycle and the watershed's ecological carrying capacity, differentiated risk management strategies are formulated, and a priority ranking of management strategies under multiple scenarios is generated. Specifically, this includes the following steps:

[0108] Collect watershed ecological carrying capacity data, which includes water body self-purification capacity data, soil adsorption capacity data, and vegetation coverage rate;

[0109] By combining the governance cycle prediction dataset with ecological carrying capacity data, an evaluation index system for management and control strategies is constructed.

[0110] Prioritize the management strategies for different types of risk-sensitive scenarios in different watersheds, and give priority to management strategies with short governance cycles and minimal ecological impact.

[0111] First, the plan will collect watershed ecological carrying capacity data, which includes three core indicators: water body self-purification capacity, soil adsorption capacity, and vegetation coverage. Taking a typical watershed as an example, water body self-purification capacity data can be obtained by monitoring parameters such as COD degradation rate and ammonia nitrogen nitrification rate in different sub-basins. For instance, the upstream mountainous sub-basin has a strong water body self-purification capacity, with an average daily COD degradation rate of 1.2 mg / L, while the downstream plain sub-basin has a slower water flow rate, with an average daily degradation rate of only 0.5 mg / L. Soil adsorption capacity data is obtained by measuring the adsorption capacity of soil for phosphorus and heavy metal pollutants in different areas. For example, the forest soil in the ecological buffer zone has an adsorption capacity of 12 mg / kg of total phosphorus, while the farmland soil in the pollution control zone has an adsorption capacity of only 5 mg / kg of soil. Vegetation coverage data can be obtained through satellite remote sensing interpretation. The vegetation coverage in the ecological buffer zone can reach 85%, while the vegetation coverage in the sub-basin surrounding the city is only 30%. These data will be integrated into a basic dataset of ecological carrying capacity, reflecting the ecological self-regulation and pollution absorption capacity of different areas in the watershed.

[0112] The plan then combines the governance cycle prediction dataset with ecological carrying capacity data to construct an evaluation index system for management and control strategies. This evaluation index system typically includes five core dimensions: governance cycle, ecological impact intensity, governance cost, implementation difficulty, and risk reduction efficiency. Each dimension is assigned a weight; for example, governance cycle has a weight of 0.3, ecological impact intensity 0.25, governance cost 0.2, implementation difficulty 0.15, and risk reduction efficiency 0.1. The governance cycle is directly calculated using the predicted number of days of governance under different measures. Ecological impact intensity is quantified and assessed using ecological carrying capacity data. For instance, if a governance measure requires the use of a high-vegetation-coverage ecological buffer zone, its ecological impact intensity will be considered high; if the measure relies on the modification of existing sewage outlets, the ecological impact intensity will be low. Taking a pollution control zone as an example, the system will generate three control strategies for addressing COD exceeding standards: 1. Reduce pollution sources by 50%, with an estimated treatment period of 15 days and low ecological impact; 2. Construct a 20-hectare artificial wetland, with an estimated treatment period of 22 days and moderate ecological impact; 3. Replenish ecological water at a rate of 100,000 cubic meters per day, with an estimated treatment period of 18 days and moderate ecological impact. These strategies will be incorporated into the evaluation index system, with each dimension assigned a corresponding weight. For example, the weight for treatment period is 0.3, for ecological impact intensity is 0.25, for treatment cost is 0.2, for implementation difficulty is 0.15, and for risk reduction efficiency is 0.1. The score for each strategy in each dimension will be weighted and calculated using the following formula: In this score, T represents the governance cycle score, with a higher score for a shorter cycle; E represents the ecological impact score, with a higher score for a smaller impact; C represents the governance cost score, with a higher score for a lower cost; D represents the implementation difficulty score, with a higher score for a lower difficulty; and R represents the risk reduction efficiency score, with a higher score for a higher efficiency.

[0113] Finally, the solution prioritizes control strategies for different watershed risk-sensitive scenarios, giving preference to strategies with short treatment cycles and minimal ecological impact. Taking drinking water source protection areas, ecological buffer zones, and pollution control zones as examples, in drinking water source protection areas, the system prioritizes strategies with minimal ecological impact, such as pollution source reduction rather than ecological water replenishment, to avoid disturbing the ecosystem of the water source. In ecological buffer zones, it prioritizes ecological restoration strategies that rely on the absorption capacity of existing vegetation and soil to maximize the use of ecological carrying capacity. In pollution control zones, it prioritizes pollution source reduction strategies with the shortest treatment cycles to quickly reduce pollution input.

[0114] Assuming the pollution source reduction strategy scores 85 points, the constructed wetland strategy scores 72 points, and the ecological water replenishment strategy scores 78 points, the system will prioritize the pollution source reduction strategy as the first priority, ecological water replenishment as the second priority, and constructed wetlands as the third priority. This prioritization not only ensures the efficient implementation of control measures but also avoids excessive intervention that could put additional pressure on the watershed ecosystem, achieving a balance between risk management and ecological protection.

[0115] Based on the priority ranking of the aforementioned control strategies, a watershed water environment risk assessment report is output. This report includes risk level distribution, estimated treatment cycle, and control strategy recommendations, specifically comprising the following steps:

[0116] Integrate data on the distribution of high-risk areas, data on the estimated governance cycle, and data on the priority of control strategies;

[0117] Based on the sub-basin units, risk level distribution maps are drawn, key influencing factors and governance cycles are marked, and a visualized watershed water environment risk assessment report is output, clarifying the control measures and implementation sequence for each sensitive scenario.

[0118] First, high-risk area distribution data, governance cycle prediction data, and control strategy priority data are integrated to form a complete dataset for the assessment report. Taking a certain watershed as an example, the high-risk area distribution data includes the spatial coordinates and risk levels of three high-risk areas: the upstream pollution control zone, the midstream sub-basin buffer zone, and the downstream drinking water source protection zone. The governance cycle prediction data includes the number of days required for governance under different control strategies, such as 15 days for pollution source reduction, 18 days for ecological water replenishment, and 22 days for constructed wetland construction. The control strategy priority data clarifies the optimal strategy for each area, such as prioritizing pollution source reduction in the pollution control zone, prioritizing ecological restoration in the ecological buffer zone, and prioritizing source control in the drinking water source protection zone. These data are then linked to the corresponding sub-basin units, forming a structured correlation matrix of region, risk, cycle, and strategy.

[0119] The system then generates a risk level distribution map based on sub-basin units, marking the key influencing factors and remediation cycles for each high-risk area. For example, on the risk level distribution map, high-risk areas are marked in red, risk areas in orange, warning areas in yellow, and safe areas in green. Key influencing factors are also marked next to each high-risk area; for example, pollutant discharge from discharge outlets is marked in pollution control zones, runoff in ecological buffer zones, and ammonia nitrogen concentration in drinking water source protection zones. The remediation cycle is also marked in the corresponding location, such as 15 days for pollution control zones, 22 days for ecological buffer zones, and 12 days for drinking water source protection zones. This visualization allows managers to intuitively see the spatial distribution of risks, core driving factors, and the time required for remediation.

[0120] Finally, the system outputs a complete, visualized watershed water environment risk assessment report, clearly defining the control measures and implementation timelines for each sensitive scenario. The report includes a spatial distribution map of risk levels, a list of high-risk areas, a table of the contribution of key influencing factors, indicator decay curves under different governance measures, a regional governance cycle forecast table, a list of differentiated control strategies, and a strategy priority ranking table. For example, the report specifies that the first priority measure in pollution control zones is a 50% reduction in pollution sources, to be completed within 15 days; the first priority measure in ecological buffer zones is the construction of 10 hectares of artificial wetlands, to be completed within 22 days; and the first priority measure in drinking water source protection zones is strengthened supervision of upstream pollution sources, to be completed within 12 days. The report also provides implementation support recommendations, such as funding allocation, inter-departmental collaboration, and monitoring plans, to ensure the efficient implementation of control strategies.

[0121] A watershed water environment risk assessment system based on big data analysis includes:

[0122] Data acquisition module: Collects multi-source heterogeneous water environment data from different sub-basin units within the watershed, processes the multi-source heterogeneous data, and constructs a basic database of the watershed water environment;

[0123] Feature extraction module: Based on the watershed water environment basic database, the watershed risk-sensitive scenario types are divided, and feature index sets of water quality, hydrology, and pollution sources in each watershed risk-sensitive scenario type are extracted, and the spatiotemporal coupling correlation between feature indexes is calculated.

[0124] Anomaly screening module: Sets risk threshold ranges for feature indicators, filters out abnormal indicator data that exceed the threshold range, matches the risk-sensitive scenarios corresponding to the abnormal indicator data, and identifies high-risk areas and key influencing factor data of the watershed water environment;

[0125] Risk assessment module: Based on historical risk impact factor data, it conducts risk data impact assessment on key impact factor data of high-risk areas, predicts the decay rate of risk indicators under different governance measures, determines the governance cycle required for the risk level to drop to the safe range, and generates a governance cycle prediction dataset.

[0126] Strategy formulation module: Combining the governance cycle and the watershed's ecological carrying capacity, formulate differentiated risk management strategies and generate a priority ranking of management strategies in multiple scenarios;

[0127] Report generation module: Based on the priority of the control strategies, outputs a watershed water environment risk assessment report, which includes risk level distribution, estimated treatment cycle, and control strategy recommendations.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A watershed water environment risk assessment method based on big data analysis, characterized in that, The method includes the following steps: Collect multi-source heterogeneous water environment data from different sub-basin units within the watershed, process the multi-source heterogeneous data, and construct a basic database of the watershed's water environment. Based on the aforementioned watershed water environment database, watershed risk-sensitive scenario types are classified, and feature indicator sets of water quality, hydrology, and pollution sources are extracted from each watershed risk-sensitive scenario type. The spatiotemporal coupling correlation between feature indicators is then calculated. Set risk threshold ranges for characteristic indicators, filter out abnormal indicator data that exceed the threshold range, match the risk-sensitive scenarios corresponding to the abnormal indicator data, and determine the high-risk areas and key influencing factor data of the watershed water environment. Based on historical risk impact factor data, risk data impact assessment is conducted on key impact factor data of high-risk areas, the decay rate of risk indicators under different governance measures is predicted, and the governance cycle required for the risk level to drop to the safe range is determined to generate a governance cycle prediction dataset. Based on the governance cycle and the watershed's ecological carrying capacity, differentiated risk management strategies are formulated, and a priority ranking of management strategies under multiple scenarios is generated. Based on the priority ranking of the control strategies, a watershed water environment risk assessment report is output, which includes risk level distribution, estimated treatment cycle, and control strategy recommendations.

2. The watershed water environment risk assessment method based on big data analysis according to claim 1, characterized in that, Collecting multi-source heterogeneous water environment data from different sub-basin units within the watershed, processing the multi-source heterogeneous data, and constructing a basic watershed water environment database specifically includes the following steps: The multi-source heterogeneous data is collected, including water quality monitoring data, hydrological runoff data, pollution source emission data, meteorological data, and land use data. The multi-source heterogeneous data is normalized, missing values ​​are filled in, and outliers are removed to unify the data format and spatiotemporal resolution. A structured watershed water environment database is constructed based on sub-basin units and monitoring time dimensions.

3. The watershed water environment risk assessment method based on big data analysis according to claim 2, characterized in that, A structured watershed water environment database is constructed based on sub-basin units and monitoring time dimensions, specifically including the following steps: Based on the natural confluence boundary of the watershed, the spatial topological relationships of each sub-watershed unit within the watershed are integrated, and the association relationships of each sub-watershed unit are determined to obtain the watershed unit association dataset; Based on the frequency and temporal variation characteristics of watershed water environment monitoring, monitoring time levels are divided, and the monitoring behavior of each monitoring time level and the corresponding sub-watershed unit is associated to form a hierarchical watershed correlation coefficient. A basic database of watershed water environment is formed by integrating watershed unit association datasets and hierarchical watershed association coefficients.

4. The watershed water environment risk assessment method based on big data analysis according to claim 3, characterized in that, Based on the aforementioned watershed water environment database, the watershed risk-sensitive scenario types are classified, and feature indicator sets of water quality, hydrology, and pollution sources are extracted from each watershed risk-sensitive scenario type. The spatiotemporal coupling correlation between the feature indicators is calculated, specifically including the following steps: The risk-sensitive scenarios in the watershed are classified into drinking water source protection areas, ecological buffer zones, and pollution prevention and control zones according to ecological functional zones, population density, and industrial and agricultural layout. Extract the characteristic indicator set of water quality, hydrology and pollution source in each risk-sensitive scenario type. The water quality indicators include COD, ammonia nitrogen and total phosphorus concentration. The hydrological indicators include runoff and flow velocity. The pollution source indicators include the number of sewage outlets and the amount of pollutants discharged. The spatiotemporal coupling correlation degree is obtained by calculating the coupling correlation degree between feature indicators under different spatiotemporal dimensions.

5. The watershed water environment risk assessment method based on big data analysis according to claim 4, characterized in that, The process involves setting risk threshold ranges for characteristic indicators, filtering out abnormal indicator data that exceed these ranges, matching the risk-sensitive scenarios corresponding to the abnormal indicator data, and identifying high-risk areas and key influencing factors for the watershed's water environment. This includes the following steps: Based on national water environment quality standards and watershed ecological objectives, risk threshold ranges for each characteristic indicator are set. Compare the indicator data in the basic database with the threshold range, and filter out abnormal indicator data that exceed the threshold range; Based on the spatiotemporal coupling correlation, risk-sensitive scenarios corresponding to abnormal indicators are matched to locate high-risk areas; Key influencing factor data is obtained by identifying the key influencing factors driving abnormal indicators based on high-risk areas.

6. The watershed water environment risk assessment method based on big data analysis according to claim 5, characterized in that, Based on historical risk impact factor data, a risk data impact assessment is conducted on key impact factor data for high-risk areas. The decay rate of risk indicators under different governance measures is predicted, and the governance cycle required for the risk level to drop to a safe range is determined, generating a governance cycle prediction dataset. The specific steps include: Collect risk occurrence history, time series change data of key influencing factors and corresponding risk level evolution data of each sub-basin unit in history, and construct a historical risk factor time series correlation dataset; The risk level and indicator decay rate are obtained by comparing the key influencing factor data of high-risk areas with the historical risk factor time series correlation dataset. Multiple governance measures were simulated, including pollution source reduction, ecological restoration, and water conservancy regulation. By inputting governance parameters for different scenarios, the system predicts the rate of decline of risk indicators, calculates the governance cycle required for the risk level to drop to the safe range, and generates a governance cycle prediction dataset.

7. The watershed water environment risk assessment method based on big data analysis according to claim 6, characterized in that, Based on the governance cycle and the watershed's ecological carrying capacity, differentiated risk management strategies are formulated, and a priority ranking of management strategies under multiple scenarios is generated. Specifically, this includes the following steps: Collect watershed ecological carrying capacity data, which includes water body self-purification capacity data, soil adsorption capacity data, and vegetation coverage rate; By combining the governance cycle prediction dataset with ecological carrying capacity data, an evaluation index system for management and control strategies is constructed. Prioritize the management strategies for different types of risk-sensitive scenarios in different watersheds, and give priority to management strategies with short governance cycles and minimal ecological impact.

8. The watershed water environment risk assessment method based on big data analysis according to claim 7, characterized in that, Based on the priority ranking of the aforementioned control strategies, a watershed water environment risk assessment report is output. This report includes risk level distribution, estimated treatment cycle, and control strategy recommendations, specifically comprising the following steps: Integrate data on the distribution of high-risk areas, data on the estimated governance cycle, and data on the priority of control strategies; Based on the sub-basin units, risk level distribution maps are drawn, key influencing factors and governance cycles are marked, and a visualized watershed water environment risk assessment report is output, clarifying the control measures and implementation sequence for each sensitive scenario.

9. A watershed water environment risk assessment system based on big data analysis, applied to the watershed water environment risk assessment method based on big data analysis as described in claims 1-8, characterized in that, include: Data acquisition module: Collects multi-source heterogeneous water environment data from different sub-basin units within the watershed, processes the multi-source heterogeneous data, and constructs a basic database of the watershed water environment; Feature extraction module: Based on the watershed water environment basic database, the watershed risk-sensitive scenario types are divided, and feature index sets of water quality, hydrology, and pollution sources in each watershed risk-sensitive scenario type are extracted, and the spatiotemporal coupling correlation between feature indexes is calculated. Anomaly screening module: Sets risk threshold ranges for feature indicators, filters out abnormal indicator data that exceed the threshold range, matches the risk-sensitive scenarios corresponding to the abnormal indicator data, and identifies high-risk areas and key influencing factor data of the watershed water environment; Risk assessment module: Based on historical risk impact factor data, it conducts risk data impact assessment on key impact factor data of high-risk areas, predicts the decay rate of risk indicators under different governance measures, determines the governance cycle required for the risk level to drop to the safe range, and generates a governance cycle prediction dataset. Strategy formulation module: Combining the governance cycle and the watershed's ecological carrying capacity, formulate differentiated risk management strategies and generate a priority ranking of management strategies in multiple scenarios; Report generation module: Based on the priority of the control strategies, outputs a watershed water environment risk assessment report, which includes risk level distribution, estimated treatment cycle, and control strategy recommendations.