Machine learning method for predicting dissolved N2O concentration of river

By combining river information, physical and chemical parameters of water and sediments, and microbial composition and functional data through machine learning, a classification and regression model was constructed, which solved the problem of insufficient accuracy in predicting river N2O concentrations and achieved high-precision prediction and the revelation of microbial regulation mechanisms.

CN120688657AActive Publication Date: 2025-09-23HUNAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510567391.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-23
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately predict river N2O concentrations at different scales, and traditional methods are unable to capture the complexity of the local environment and the impact of microorganisms, resulting in insufficient prediction accuracy.

Method used

Using machine learning methods, combined with river information, physical and chemical parameters of water and sediments, and microbial composition and functional data, classification and regression models were constructed to predict N2O concentration using the optimal data set, and to analyze the regulatory mechanism of microorganisms on N2O production.

Benefits of technology

It achieved high-precision prediction of river N2O concentration at different scales, revealed the key regulatory mechanisms of environmental factors and microorganisms on N2O production, and provided more accurate emission reduction strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688657A_ABST
    Figure CN120688657A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning method for predicting the dissolved N2O concentration of a river, and the method comprises the following steps: calculating the N2O dissolved concentration of each region of a target drainage basin, and carrying out the N2O concentration grouping of each region according to the N2O dissolved concentration; respectively acquiring macroscopic features, environmental physical and chemical indexes and microscopic features of a target drainage basin and taking the macroscopic features, the environmental physical and chemical indexes and the microscopic features as explanatory variables, taking N2O concentration grouping results of all regions as response variables, constructing a data set and respectively training corresponding classification models, selecting the explanatory variable corresponding to the classification model with the best classification effect, and taking N2O dissolution concentration as the response variable to obtain a classification model with the optimal classification effect; constructing a new data set training regression model; and screening an optimal parameter from the explanatory variables, training an optimal regression model by using the optimal parameter, and calculating an N2O emission flux predicted value according to the N2O solution concentration predicted by the optimal regression model. According to the method, the N2O concentration prediction precision is improved, and meanwhile, a key regulation and control mechanism of microorganisms on N2O generation is disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of environmental technology, and in particular to a machine learning method for predicting dissolved N2O concentration in a river. Background Art

[0002] Nitrous oxide (N2O) is a major greenhouse gas with a greenhouse effect potential far greater than that of other greenhouse gases, making a significant contribution to global warming. Therefore, N2O emissions have become a key issue in global warming research. N2O is widely present in both natural and anthropogenically altered water bodies. Its generation, transformation, and consumption in water bodies are regulated by multiple factors, including nitrogen input, microbial activity (nitrification / denitrification), environmental conditions (dissolved oxygen, temperature, pH), hydrological characteristics (flow rate, water level), and sediment properties (organic matter, particle size). Nitrogen, as a key substrate, drives N2O generation and emission through microbial metabolism. N2O is primarily produced through microbial nitrification and incomplete denitrification. Riverine N2O emissions are estimated to be 254.9±46.2 Gg N yr⁻¹, accounting for 79.8% of global N2O emissions from inland waters (319.6±58.2 Gg N yr⁻¹). Therefore, monitoring the concentration of N2O in rivers and taking corresponding emission reduction measures in a timely manner are of great significance to ecological and environmental protection. At present, the traditional N2O monitoring method is headspace gas chromatography (HS-GC). This method is based on water sample collection and measurement and is suitable for the accurate detection of trace N2O. However, it has low temporal and spatial resolution and is easily interfered by matrix effects. It can only obtain spatially discrete data and is difficult to meet the real-time monitoring needs of dynamic environments. The in-situ monitoring method membrane injection mass spectrometry (MIMS) that can achieve continuous monitoring uses equipment that is expensive and requires strict maintenance. It is not suitable for large-scale field application and is difficult to cover all emission sources. Therefore, the prediction and simulation of N2O concentrations have become an important supplement to the research on N2O methods. The existing methods for predicting river N2O concentrations are mainly divided into statistical models, mechanism models and hybrid models. The statistical model establishes an empirical formula based on historical data to quickly estimate the concentration of N2O. The International Panel on Climate Change (IPCC) recommends that the N2O concentration be estimated based on the nitrate concentration (NO3 - The concentration of dissolved N2O in rivers is estimated by multiplying the product of [N2O-N] and the default emission factor (EF5r = 0.26%), that is, [N2O-N] = [NO3 - -N]×EF5r. In the formula, [N2O-N] is the dissolved N2O concentration in the river (μg / L), [NO3 - -N] is the nitrate concentration in rivers (μg / L), and EF5r is the default emission factor provided by IPCC.

[0003] This method is suitable for large-scale and large-scale river basin estimation, but N2O concentration has great temporal and spatial variability in global rivers. The default emission factor recommended by IPCC cannot accurately reflect local conditions. Therefore, some scholars use machine learning to identify the key control factors affecting EF5r and propose a method based on dissolved organic carbon (DOC) and nitrate (NO3 - Regionally specific prediction models based on the SWAT-N2O coupler (likely a typo or SWAT-N2O coupler) can more accurately predict N2O concentrations in small areas. Mechanistic models primarily use mathematical equations to describe the physical, chemical, and biological processes of N2O production, transport, and emission, integrating hydrological, microbial, and environmental dynamics. For example, a study developed a riverine N2O greenhouse gas emission model (SWAT-N2O coupler). This model leverages literature data to establish a relationship between DIN (dissolved inorganic nitrogen in rivers) and dissolved N2O, and incorporates an air-water exchange model and N2O coefficients to simulate riverine greenhouse gas emissions. This model, with inputs such as DEM, land use, soil data, daily meteorological data, and tillage management practices, outputs the average daily water nitrogen load, runoff, and nitrogen concentration for a river section. This approach offers high accuracy, but also requires a high level of data and computational complexity. Furthermore, a study has integrated the process-based SWAT (Soil and Water Assessment Tool) model with a linear mixed model (LMM) to predict N2O emissions from headwater agricultural river systems under different fertilization and climate change scenarios, providing information for fertilization management.

[0004] At present, big data-driven machine learning models, such as random forests, support vector machines, and neural networks, have become an important research method due to their simplicity and efficiency. They are applied to the prediction of N2O concentration. Machine learning methods can improve the prediction accuracy and find the key driving factors of river N2O concentration. Some studies have combined the SWAT model and the random forest model to predict the concentration of dissolved inorganic nitrogen (DIN) and dissolved N2O in rivers, revealing the difference in the contribution of human activities (such as sewage treatment policies) and natural factors (such as climate change) to N2O emissions, and thus proposing differentiated emission reduction strategies for urban and rural areas. Some studies have also established a satellite remote sensing-driven machine learning model. By inverting parameters such as sea surface temperature (SST), salinity (SSS), dissolved inorganic nitrogen (DIN), and chlorophyll a (Chl-a) from MODIS and Landsat satellites, combined with indirect indicators such as CDOM absorption coefficient and diffuse attenuation coefficient (Kd), the optimal input feature set (such as SST, SSS, DIN, and aCDOM (400)) was selected to construct an XGBoost model to predict dissolved N2O concentration. In the Pearl River Estuary study, the model test set R 2The results reached 0.74, with an RMSE of 5.16. SHAP analysis revealed that temperature, salinity, and DIN were the core driving factors. Machine learning can capture the nonlinear responses between complex variables through a data-driven approach, quantify the contributions of input variables, and provide a basis for N2O emission reduction strategies. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: In response to the technical problems existing in the prior art, the present invention provides a machine learning method for predicting the dissolved N2O concentration in rivers. It conducts research at the macroscale, mesoscale and microscale respectively, and selects the optimal data set. Then, a machine learning method is used to predict the N2O concentration based on the optimal data set (mesoscale water and sediment physical and chemical parameters). The predicted N2O concentration is combined with microbial community functional analysis. While better predicting the N2O concentration, the key regulatory mechanisms of environmental factors and microorganisms on N2O production are revealed.

[0006] In order to solve the above technical problems, the technical solution proposed by the present invention is:

[0007] A machine learning method for predicting dissolved N2O concentration in rivers includes the following steps:

[0008] Calculate the N2O dissolved concentration in each area of ​​the target watershed and group each area into N2O concentration groups based on the N2O dissolved concentration;

[0009] Obtain the macroscopic characteristics, environmental physicochemical indicators, and microscopic characteristics of each area in the target watershed, and use the grouping results of each area as the macroscopic characteristics. The macroscopic characteristics include river information data of the target watershed, the environmental physicochemical indicators include physicochemical parameter data of water bodies and sediments in the target watershed, and the microscopic characteristics include microbial composition and function data of the target watershed.

[0010] Macroscopic characteristics, environmental physical and chemical indicators, and microscopic characteristics were used as explanatory variables, and the N2O concentration grouping results of each region were used as response variables to construct a classification model for the dataset training.

[0011] The explanatory variables of the classification model with the best classification effect were selected, and the N2O concentration in the target watershed was used as the response variable. A new dataset was constructed and a regression model was trained to predict the N2O concentration.

[0012] The optimal parameters are selected from the explanatory variables, and the optimal parameters are used as the explanatory variables. The N2O dissolved concentration in the target watershed is used as the response variable. The optimal data set is constructed and the optimal regression prediction model is trained. The N2O dissolved concentration predicted by the optimal regression model is obtained, and the N2O emission flux prediction value is calculated based on the N2O dissolved concentration.

[0013] Furthermore, when calculating the dissolved N2O concentration in each area of ​​the target watershed, water samples were collected at sampling points in each area of ​​the target watershed. The water samples were slowly discharged into the bottom of the sample bottle through a silicone tube with minimal turbulence. After the sample bottle was filled with water, the water was allowed to overflow to avoid N2O in the atmosphere from contaminating the water sample. HgCl2 was then injected into the sample bottle to a final concentration of 0.5% v / v to inhibit microbial activity. The sample bottle was stored at low temperature and the headspace equilibrium-gas chromatography method was used to determine the dissolved N2O in the water of the sample bottle.

[0014] Furthermore, when obtaining the macro characteristics of each area in the target basin, the hydrological model of the target basin is constructed, the pollution source database is introduced into the hydrological model, the hydrological model is parameterized and calibrated, and the calibrated hydrological model is output to obtain the river information data of each area.

[0015] Furthermore, when obtaining the environmental physicochemical indicators of each area in the target watershed, the physicochemical parameters of the water body and sediment are measured through in-situ monitoring and calculation at the sampling points in each area of ​​the target watershed. The physicochemical parameters of the water body include one or more of the concentrations of total nitrogen, ammonium nitrogen, nitrate and total phosphorus in the water; the physicochemical parameters of the sediment include one or more of the concentrations of total nitrogen, ammonium nitrogen, nitrate and total phosphorus in the sediment.

[0016] Furthermore, when obtaining the microscopic characteristics of each area in the target watershed, sediment samples are collected at sampling points in each area of ​​the target watershed, the microbial 16S genome of the sediment samples is sequenced, and the sequencing data is matched with a preset database to obtain microbial composition data and functional data.

[0017] Furthermore, the macroscopic characteristics, environmental physical and chemical indicators, and microscopic characteristics were used as explanatory variables, and the N2O concentration grouping results of each region were used as response variables. When constructing the dataset to train the corresponding classification model, the following steps were included:

[0018] Dividing the data set into a training set and a test set;

[0019] Construct a random forest classification model, determine the optimal values ​​of key parameters of the random forest classification model through 5-fold cross-validation, and train the random forest classification model using the training set so that each decision tree in the trained random forest classification model predicts the classification of macroscopic characteristics, environmental physical and chemical indicators, or microscopic characteristics. Vote on the prediction results of all decision trees, and use the category with the most votes as the final prediction result.

[0020] The data of the test set is input into the trained random forest classification model, and the evaluation indicators are calculated based on the classification prediction results of the random forest classification model and the actual classification results of the data of the test set. The evaluation indicators include the accuracy and AUC value of the test set.

[0021] Furthermore, the explanatory variables of the classification model with the best classification effect are environmental physical and chemical indicators, and the optimal parameters are total phosphorus, total nitrogen, ammonium nitrogen, nitrate, water temperature T and sediment nitrate in the physical and chemical parameter data of water and sediment of environmental physical and chemical indicators.

[0022] Furthermore, when training the regression model and training the optimal regression model, both include:

[0023] Dividing the new data set or the optimal data set into a training set and a test set;

[0024] A random forest regression model was constructed and the optimal values ​​of its key parameters were determined through 5-fold cross-validation. The model was then trained using the training set, so that each decision tree in the trained random forest regression model predicted the dissolved N2O concentration based on the physicochemical parameter data of water and sediments. The prediction results of all decision trees were averaged as the final prediction result.

[0025] The test set data is input into the trained random forest regression model, and evaluation indicators are calculated based on the N2O dissolved concentration prediction results of the random forest regression model and the actual N2O dissolved concentration results of the test set data. The evaluation indicators include one or more of the root mean square error, determination coefficient, mean absolute error, mean relative error, and bias deviation.

[0026] Furthermore, when selecting the optimal parameters from the explanatory variables, it includes:

[0027] Input the data set into the random forest regression model and calculate the MDG results of each explanatory variable of the random forest regression model;

[0028] The MDG results were statistically analyzed and the optimal parameters were screened in descending order of ranking.

[0029] Furthermore, after selecting the optimal parameters from the explanatory variables, the method also includes: analyzing the correlation between the optimal parameters and the microbial composition and function data. Specifically, the correlation coefficient between each optimal parameter and the microbial composition data is obtained through the Pearson analysis method, and the pheatmap diagram showing the relationship between each optimal parameter and the microbial function data is obtained through the Spearman analysis method.

[0030] Compared with the prior art, the advantages of the present invention are:

[0031] The present invention uses machine learning to analyze the physical and chemical parameters of water and sediment (such as TN, NH4 + 、NO3 -Machine learning can better exploit the complex nonlinear relationships between physical and chemical parameters and N2O emissions (e.g., etc.). Traditional mechanistic models require a large number of soil, microbial, and environmental parameters, while the machine learning model in this study only requires conventional water quality and sediment monitoring data to achieve high-precision predictions. Furthermore, compared to traditional IPCC prediction methods and other machine learning models, this method adds sediment parameters in addition to water quality parameters, taking into account a more comprehensive range of influencing factors, thereby achieving better N2O concentration prediction results.

[0032] Through correlation analysis between microorganisms and the physicochemical parameters of important water bodies—sediments—this method can identify which environmental factors affect N2O emissions by regulating the structure or function of microbial communities, thereby explaining the importance of features in machine learning models. For example, it can explain how important features such as temperature and nitrogen concentration affect N2O emissions by affecting microbial activity, thereby explaining the nonlinear results in the model at the biological level, so as to more accurately locate critical control points and optimize management strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of a method flow of an embodiment of the present invention.

[0034] Figure 2 3. It is a schematic diagram comparing the evaluation indicators of three random forest classification models trained on watershed macro characteristics, environmental physicochemical indicators and microbial micro indicators in an embodiment of the present invention.

[0035] Figure 3 Schematic diagram of using the optimal regression model to predict the N2O effect in an embodiment of the present invention.

[0036] Figure 4 3 is a schematic diagram comparing the effects of the random forest regression model and the IPCC method on predicting N2O in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the scope of protection of the present invention is not limited thereby.

[0038] Before introducing the specific implementation methods, the relevant concepts are first explained.

[0039] Microbial 16S sequencing is a high-throughput sequencing technology based on 16S rRNA genes. It is widely used to analyze the composition, diversity, and functional potential of bacterial and archaeal communities. By analyzing the variable regions of the 16S gene (such as V1-V9), this technology can provide taxonomic information (from phylum to species level), identify dominant and rare species, and assess alpha diversity (e.g., Chao1 and Shannon indices, which measure species richness and evenness) and beta diversity (e.g., PCoA and NMDS, which reveal inter-sample variability). Furthermore, it can track the dynamics of microbial communities under different environmental conditions, time series, or interventions. Combined with statistical methods (such as LEfSe), it can screen biomarkers associated with disease and environmental stress. Tools such as FAPROTAX and PICRUSt2 can also be used to predict metabolic pathways (e.g., nitrogen cycling and antibiotic synthesis) and ecological functions. In medicine, environmental science, agriculture, and industry, 16S sequencing is being used to explore the relationship between gut microbiota and disease, soil / water microbial function, plant rhizosphere interactions, and fermentation process monitoring. These data provide insights into how microorganisms adapt to their environment, how they respond to environmental changes, and their roles in ecosystems.

[0040] Mean Decrease Gini (MDG): This refers to a feature importance metric typically used in random forest models. In random forests, the Gini index is a measure of node purity, while "Mean Decrease Gini" represents the average reduction in the Gini index by a particular feature.

[0041] When training a random forest, multiple features are considered at each split node, and the "Mean Decrease Gini" metric measures the average contribution of each feature to the overall purity of the model. A higher MDG value generally indicates that the feature is more important in building the decision tree. Calculating this metric involves averaging the performance of many decision trees, as random forests are constructed by combining multiple decision trees. Random forest models in machine learning libraries such as scikit-learn typically provide feature importance scores that include MDG.

[0042] Example

[0043] Traditional IPCC prediction methods use only water quality parameters and may only capture processes in water quality while ignoring the contributions of sediments. Sediments are the primary habitat for microbial survival and metabolism. The organic matter contained in them may serve as a carbon source to promote denitrification, while the anaerobic environment of sediments can promote denitrification to produce N2O. In addition, the hydraulic conductivity of sediments may affect the exchange of substances between water and sediments, thereby affecting the diffusion of N2O. Sediment properties may vary greatly in different river sections, especially in urban and rural areas, where human activities have different impacts on sediments. For example, urban river sediments may contain more pollutants, affecting the structure of microbial communities, while rural areas may be affected by agricultural runoff, resulting in higher nutrient content in sediments. Therefore, using machine learning to build a model based on the physicochemical parameters of river water and sediments can simultaneously capture processes in water quality and sediments.

[0044] Furthermore, while machine learning can predict N2O concentrations and reveal key driving factors, N2O metabolism in real environments is highly complex. Microorganisms are the executors of N2O production, and their community structure and function directly determine nitrogen transformation pathways and N2O yields. Therefore, after identifying the driving factors, analyzing the correlation between river microbial composition and function and water-sediment physicochemical parameters can identify which environmental factors affect N2O emissions by regulating microbial community structure or function. This can explain important features in machine learning models, such as how temperature and nitrogen concentration affect N2O emissions by influencing microbial activity. This can also explore the impact of microbial structure and function on N2O emissions, providing a microbial perspective on the N2O production mechanism and explaining the nonlinear results in the model at a biological level, allowing for more accurate identification of key control points and optimization of river N2O management strategies.

[0045] Therefore, this embodiment proposes a machine learning method for predicting dissolved N2O concentration in rivers, using machine learning to calculate the physical and chemical parameters of water and sediments (such as TN, NH4 + 、NO3 - etc.) to predict N2O concentrations, enabling machine learning to better utilize the complex nonlinear relationships between physical and chemical parameters and N2O emissions. Figure 1 As shown, it mainly includes the following steps:

[0046] S1) Calculate the dissolved N2O concentration in each area of ​​the target watershed;

[0047] S2) Based on the level of dissolved N2O concentration, the target watershed regions are divided into high N2O concentration regions (HA) and low N2O concentration regions (LA), and the N2O concentration groups of each region are obtained;

[0048] S3) obtaining macroscopic characteristics, environmental physicochemical indicators, and microscopic characteristics of each area of ​​the target watershed, wherein the macroscopic characteristics include river information data of the target watershed, the environmental physicochemical indicators include physicochemical parameter data of water bodies and sediments in the target watershed, and the microscopic characteristics include microbial composition and function data of the target watershed;

[0049] S4) Using macroscopic characteristics, environmental physical and chemical indicators, and microscopic characteristics as explanatory variables, and the N2O concentration grouping results of each region as the response variable, a classification model corresponding to the dataset training was constructed;

[0050] S5) Select the explanatory variables of the classification model with the best classification effect, use the N2O dissolved concentration in the target watershed as the response variable, construct a new data set and train a regression model to predict the N2O dissolved concentration;

[0051] S6) using Mean Decrease Gini (MDG) as a feature importance metric, screening optimal parameters from the explanatory variables to form an optimal parameter set, using the optimal parameters as the explanatory variables, taking the N2O dissolved concentration in the target watershed as the response variable, constructing an optimal data set and training an optimal regression model, obtaining the N2O dissolved concentration predicted by the optimal regression model, and calculating the predicted N2O emission flux based on the N2O dissolved concentration;

[0052] S7) analyzing the correlation between microbial composition and function data and optimal parameters as environmental variables to determine parameters with more significant effects on microbial species and functions.

[0053] The above steps first examine the datasets that best classify macroscopic features, environmental physicochemical indicators, and microscopic characteristics. Next, a machine learning model is constructed to simulate N2O concentrations using the best-classified datasets, enabling prediction of dissolved N2O concentrations in rivers. Furthermore, the researchers explain the underlying mechanisms by which the selected optimal data regulates river N2O from a microbial perspective, more accurately revealing the pathways of N2O production and emissions. By identifying parameters that most significantly influence microbial species and functions, this provides a more scientific basis for differentiated emission reduction strategies.

[0054] The following is a detailed description of each step.

[0055] In step S1 of this embodiment, when calculating the dissolved NO concentration in each area of ​​the target watershed, water samples are collected at sampling points in each area of ​​the target watershed, and the dissolved NO concentration is calculated for the water samples at each sampling point, including the following steps:

[0056] S101) Select sampling points.

[0057] Collect land use data, soil properties, topography, climate data, pollution data, etc. within the target area of ​​the designated watershed. Select representative sampling points based on the watershed land use, key provincial and national controlled sections, the location of factories, enterprises, and sewage treatment plants, and the nitrogen pollution status of the river mainstream and tributaries. The specific steps include:

[0058] First, determine the locations of key provincial and national controlled sections and factories, enterprises, and sewage treatment plants:

[0059] Obtain the divided watershed area map, the distribution data of provincial and national controlled sections, and the distribution data of factories, enterprises, and sewage treatment plants; mark the locations of provincial and national controlled sections, factories, enterprises, and sewage treatment plants on the watershed area map; analyze the relationship between these locations and the watershed area, determine the areas where they are located, and obtain the watershed area map with the locations of provincial and national controlled sections, factories, enterprises, and sewage treatment plants marked.

[0060] Then, analyze the nitrogen pollution status of the river mainstream and tributaries:

[0061] Obtain a river basin map with the locations of provincial and national controlled sections, factories, enterprises, and sewage treatment plants, as well as water quality monitoring data (nitrogen pollution status) for the river's main stream and tributaries. Collect and organize nitrogen pollution monitoring data for the river's main stream and tributaries, analyze the spatial distribution of nitrogen pollution, identify areas with more severe pollution, and analyze the relationship between pollution sources and pollution status based on the locations of provincial and national controlled sections, factories, enterprises, and sewage treatment plants. Obtain a nitrogen pollution distribution map for the river's main stream and tributaries, as well as analysis results on the relationship between pollution sources and pollution status.

[0062] Finally, select representative sampling points:

[0063] Obtain nitrogen pollution distribution maps for the river's main stream and tributaries, analysis results on the relationship between pollution sources and pollution status, and a map of the designated watershed regions. Based on the nitrogen pollution distribution maps and pollution source analysis results, identify areas requiring key monitoring. Within each region, select representative sampling points based on land use type, provincially and nationally controlled sections, and the locations of factories, enterprises, and sewage treatment plants. Ensure that the sampling points cover different pollution conditions and land use types to fully reflect the nitrogen pollution status of the watershed. Obtain a representative distribution map of the sampling points, including detailed information on the sampling points (location, region, land use type, etc.).

[0064] S102) Water sampling.

[0065] Use a water sampler to repeatedly collect 4×250mL of river surface water by changing positions around the sampling point into a sterile water sample collection bag.

[0066] S103) N2O analysis.

[0067] The water sample was slowly drained into the bottom of the sample bottle through silicone tubing with minimal turbulence. After the sample bottle was filled, several volumes of water were allowed to overflow, avoiding the formation of bubbles. For nitrous oxide sampling, the water sample was slowly drained into the bottom of the sample bottle through silicone tubing with minimal turbulence. After the sample bottle was filled, several volumes of water were allowed to overflow to prevent atmospheric nitrous oxide contamination of the sample. Next, 200 μL of HgCl₂ was injected into the sample bottle to a final concentration of 0.5% v / v to inhibit microbial activity. The sample bottle was then stored in a freezer. If any signs of bubbles were present, a new sample was collected. The sample was returned to the laboratory, and dissolved nitrous oxide in the sample water was determined using headspace equilibrium gas chromatography (Agilent 7890B, μECD).

[0068] Step S2 of this example aims to group N2O concentrations by region. Specifically, the mclust package in R is used to perform classification based on the distribution of N2O concentration data. All dissolved N2O concentrations are divided into two groups: HA (high dissolved N2O concentration group) and LA (low dissolved N2O concentration group). The corresponding sampling point groups are labeled HA or LA.

[0069] The R language's mclust package is a statistical analysis tool based on Gaussian mixture models (GMM), primarily used for cluster analysis, classification, and density estimation. Specifically, using the R language's mclust package to complete classification involves constructing a mixture model for labeled N2O concentration data and implementing classification predictions to achieve the optimal number of clusters and reasonable classification, resulting in highly accurate and interpretable classification results.

[0070] In step S3 of this embodiment, when obtaining the macro characteristics of each area in the target watershed, specifically, a hydrological model of the target watershed is constructed, a pollution source database is introduced into the hydrological model, the hydrological model is parameterized and calibrated, and the calibrated hydrological model is obtained to output river information data for each area, including the following steps:

[0071] S201) collecting land use data, soil property data, topographic data, climate data, pollution data, etc. within the target watershed, and inputting them into the SWAT hydrological model to establish a hydrological model for the target watershed;

[0072] S202) Parameterizing and calibrating the model, including: comparing the simulation results of the current model with the measured data, adjusting the model parameters to improve the fit; and optimizing the adjusted model parameters using a specified optimization algorithm.

[0073] The parameterization and calibration of hydrological models is an iterative process. The key lies in establishing an initial model based on the target watershed data. Then, by comparing simulation results with measured data, model parameters are adjusted to improve the fit. This involves adjusting parameters such as hydrological processes, soils, and vegetation. The calibration process utilizes the SUFI-2 optimization algorithm to ensure accurate simulations across a wide range of scenarios. This iterative process continuously optimizes the model, ensuring greater adaptability and predictive accuracy.

[0074] In this example, the SWAT model was calibrated and validated using the SWAT calibration and uncertainty program software. After 1000 simulation iterations using the selected parameters, the model accurately simulated the average monthly flow and nitrogenous pollutant loads during the calibration and validation phase (all R2>0.66, NSE>0.65). River-related data for each region was obtained from the SWAT model output and local hydrological datasets, including river width, river length, river depth, and river slope. This involved data such as land use, slope, elevation, flow velocity, and river information.

[0075] In step S3 of this embodiment, when obtaining the environmental physicochemical indicators of each basin in the target watershed, specifically, measuring the physicochemical parameters of water bodies and sediments through in-situ monitoring and calculation at sampling points in each basin in the target watershed, the following steps are included:

[0076] S301) Sediment sampling.

[0077] In this example, the sampling period is divided into summer (July) and winter (February), covering the dry and wet seasons. Sediment samples are collected at each sampling point at the designated sampling time and homogenized. The processing method complies with national standards.

[0078] During sampling, sediment samples were collected using a grab bucket, and three samples were collected at the same location (at a depth of 0 to 5 cm from the uppermost riverbed). The collected samples were then homogenized, sealed in pre-sterilized bags, and stored at -80°C for analysis. In this embodiment, two key sampling periods were clearly defined, namely winter (February) and summer (July). This time selection takes into account the possible impact of seasonal changes on N2O, especially the impact of temperature on N2O production and emissions. At each selected sampling point, a grab bucket sampler was used to collect samples 0-5 cm from the riverbed. Each sample was collected in triplicate to avoid extreme data or operational errors. Finally, the sediment samples were sealed and stored in pre-sterilized bags. All samples were quickly frozen and stored at a low temperature of -80°C, which helped maintain the stability of the samples and provided high-quality sediment samples for subsequent analysis.

[0079] S302) Water sampling.

[0080] Using a water sampler, we repeatedly collected 4 x 250 mL of river surface water at various locations around the sampling point into sterile water sample bags. Two of these samples were preservative-free to measure TP, TN, and NH₄⁺, while the remaining two were preservative-free to measure NO₃⁻. The sterile water sample bags were then stored in a freezer. Upon return to the laboratory, these environmental variables were measured according to standard methods.

[0081] S303) Analyze physicochemical parameters.

[0082] This example analyzes 11 physical and chemical parameters, including 4 sediment variables and 7 water variables. These variables are environmental variables and are measured according to standard methods. The 7 water variables are measured on-site using a multi-parameter instrument (HQ2200, HACH, USA) for water temperature, pH, and dissolved oxygen. The total nitrogen (TN), ammonium nitrogen (NH4) in water are measured according to the Chinese national standards (GB 7480-87, GB 7479-87, GB 11893-89, and GB 11894-89, respectively). + -N), nitrate (NO3 - -N) and total phosphorus (TP) concentrations to obtain the physicochemical parameters of the water body. The four sediment variables were analyzed according to the standards of the Ministry of Ecology and Environment of China (HJ 632-2011, HJ 717-2014 and HJ 634-2012, respectively), to analyze the sediment properties (TN, TP, NO3 - -N and NH4 + -N), thereby obtaining the physical and chemical parameters of the sediment.

[0083] In step S3 of this embodiment, when obtaining the microscopic characteristics of each area in the target watershed, specifically, sediment samples are collected at sampling points in each area of ​​the target watershed, the sediment samples are subjected to microbial 16S genome sequencing, and the sequencing data is matched with a preset database to obtain microbial composition data and functional data. The specific steps include:

[0084] The first step is to collect sediment samples. The specific steps are the same as step S301 and will not be repeated here.

[0085] The collected sediment samples were then sequenced for microbial genomes, including:

[0086] S401) Pre-treating the sediment sample and extracting DNA, the specific steps are as follows:

[0087] A 0.25g fresh sediment sample was weighed and placed in a liquid nitrogen-cooled mortar. Three cycles of freeze-grinding with liquid nitrogen (grinding to a powder after each addition of liquid nitrogen) were performed to thoroughly disrupt the microbial cell walls. The ground powder was transferred to a 2mL sterile centrifuge tube, and 500μL of Solarbio Solution A (containing humic acid adsorbent and 0.5mm glass beads) was added. The suspension was vortexed at 2000rpm for 5 minutes to form a homogenous suspension. Subsequently, 50μL of lysozyme (20mg / mL) was added, and the suspension was incubated at 37°C for 30 minutes to lyse Gram-positive bacteria. 50μL of Solution B (containing 1% SDS and proteinase K) was then added and lysed in a 65°C waterbath for 10 minutes (inverting every 2 minutes to mix). After lysis, 100μL of Solution C (precipitant) was added, the tube was allowed to stand for 5 minutes, and then centrifuged at 12,000×g for 10 minutes to remove impurities. The supernatant was transferred to a fresh tube. The supernatant was mixed with 500 μL of binding buffer and an equal volume of isopropanol, loaded onto the DNA adsorption column in batches, centrifuged at 12,000 × g, and the waste liquid was discarded. After washing twice with rinsing solution (containing ethanol), the adsorption column was dried at room temperature for 5 minutes. Finally, the DNA was eluted with 50 μL of TE buffer preheated at 65°C and stored at -20°C until use;

[0088] S402) The extracted DNA is quality tested and amplified by PCR to obtain a PCR product.

[0089] In this example, after DNA extraction, the purity was first tested using a NanoDrop (OD260 / 280 ratio needed to be between 1.7 and 1.9, OD260 / 230 ≥ 2.0), and the concentration was measured using a Qubit 4.0 fluorometer (required to be ≥ 10 ng / μL). If the concentration was insufficient, vacuum concentration was performed. Subsequently, 2% agarose gel electrophoresis (100 V, 30 minutes) was performed to verify the DNA integrity (the main band needed to be > 10 kb and without diffuse tailing). PCR amplification was performed using primers 338F (10 μM) and 806R (10 μM). A 25 μL reaction system was prepared by adding 5 μL of 5× Q5 buffer, 2 μL of 2.5 mM dNTPs, 1 μL of each primer, 2 μL of DNA template (10 ng / μL), 0.25 μL of Q5 High-Fidelity enzyme, and 0.2 μL of BSA (0.1 mg / mL). The volume was adjusted to 25 μL with ddH2O. The amplification program was as follows: initial denaturation at 98°C for 2 minutes, 25 cycles of (98°C for 15 seconds, 55°C for 30 seconds, and 72°C for 30 seconds), and a final extension at 72°C for 5 minutes. The amplified product was confirmed by 2% agarose gel electrophoresis to identify the 460 bp target band. If primer dimers (<100 bp) were present, the product was purified using AMPure XP magnetic beads (0.8× volume) and stored for future use.

[0090] S403) Using the PCR product, a sequencing library is constructed and sequenced to obtain raw sequencing data.

[0091] Specifically, after PCR product purification, libraries were prepared using the KAPA HyperPrep kit. DNA ends were first blunted using an end-repair reaction (incubation at 20°C for 30 minutes), followed by ligation with Illumina TruSeq dual-index adapters (containing an 8bp specific tag, incubation at 20°C for 15 minutes) to ensure unique identification for each sample. Ligation products were purified using AMPure XP magnetic beads (0.6× volume) to select for target fragments of 450-550bp, removing short fragments (<300bp) and residual primer dimers. Library quantification was performed using Qubit fluorescence quantification and an Agilent 2100 Bioanalyzer to determine fragment distribution. Qualified libraries were diluted to 1.8-2.2pM before loading. Sequencing was performed on the Illumina NovaSeq 6000 platform in PE300 (paired-end 300bp read length) mode, generating at least 50,000 reads per sample. 1% PhiX Control v3 was added to balance base diversity and correct sequencing errors.

[0092] S404) The quality of the original sequencing data is assessed, and then low-quality bases are cut out and primer sequences are removed to obtain high-quality sequences.

[0093] Specifically, the raw sequencing data were quality-assessed using FastQC (Q30 score > 80%), low-quality bases were removed using Trimmomatic (parameters: SLIDINGWINDOW: 4:20, MINLEN: 100), and primer sequences were removed using Cutadapt (allowing 10% mismatches). High-quality sequences were denoised using DADA2 to generate ASVs (Amplicon Sequence Variants), and species annotation was performed using the Silva 138 database (similarity threshold 97%). Alpha diversity (Shannon and Chao1 indices) and beta diversity (PCoA analysis based on Bray-Curtis distance) were calculated using QIIME2, and differential species were screened using LEfSe analysis (LDAScore > 3.0 and p < 0.05). When submitting data to NCBI SRA, the original FASTQ file must be compressed with gzip (naming example: Sample1_R1.fastq.gz), and a metadata file must be attached to describe the sample's geographic coordinates (such as "GPS: 31.23°N, 121.47°E"), sampling depth (0-10cm), pH value and other environmental parameters, and the biological project number PRJNA845821 must be associated to ensure that the data is publicly traceable.

[0094] During each of the above steps, the following key points should be noted: DNA extraction stage: negative control (no sample extraction) to exclude reagent contamination; PCR stage: positive control (standard strain DNA) to verify primer validity; library construction stage: Agilent 2100 detection of fragment distribution (single peak 450-550bp); sequencing stage: PhiX Control to correct base recognition error rate (<0.1%); analysis stage: chimera filtering (UCHIME algorithm) and negative control sequence removal.

[0095] Finally, the sequencing data is matched to a preset database, mainly through microbial bioinformatics analysis, to obtain N2O characteristic bacteria and characteristic bacterial community functional data, including:

[0096] S405) After correcting sequencing errors and removing chimeric sequences for high-quality sequences using the DADA2 algorithm, similar sequences with a similarity greater than a specified degree (similarity ≥ 97%) are clustered into operational taxonomic units (OTUs) to generate a species abundance feature table;

[0097] S406) The operational taxonomic unit OTU is compared with the standard database, and the taxonomic information of the microorganism is annotated by a Bayesian classifier or an exact matching algorithm. According to the annotation results, the relative abundance of species at different classification levels of each sample is counted to form a multidimensional abundance matrix. Specifically, in this embodiment, the OTU sequence obtained by high-throughput sequencing is compared with the standard database (Silva, Greengenes), and the taxonomic information of the microorganism (classification levels such as phylum, class, and genus) is annotated by a Bayesian classifier or an exact matching algorithm. According to the annotation results, the relative abundance of species at different classification levels of each sample is counted to form a multidimensional abundance matrix (such as a phylum level abundance table and a genus level abundance table) for subsequent difference analysis;

[0098] S407) Using the inter-group significance difference test and LEfSe analysis method, based on the multidimensional abundance matrix, the significantly different microbial flora between the sediment samples in the areas grouped by different N2O concentrations are analyzed, and the microbial flora with significant differences in the different N2O concentration groups and related to river nitrogen metabolism is selected as the microbial composition data. Specifically, this embodiment is based on the inter-group significance difference test and LEfSe analysis, i.e., LDA EFfect Size analysis. According to the community abundance data, strict statistical methods are used to perform hypothesis testing on the species between the microbial communities of different groups (or samples), evaluate the significance level of the species abundance difference, obtain the species with significant differences between the groups (or samples), and find biomarkers with statistical differences between the groups, wherein:

[0099] Non-parametric tests (such as the Mann-Whitney U test and the Kruskal-Wallis test) were used to identify species that differed between groups and to compare the statistical differences in species abundance between different experimental groups. Corrections for multiple testing (such as FDR correction) were performed to screen for species that met the significance threshold (such as p < 0.05). LDA (linear discriminant analysis) was used to assess the effect size (LDA Score) of the differential species. Combined with the results of the Kruskal-Wallis test, biomarkers with significant enrichment characteristics between groups were screened (usually LDA Score > 2.0 and p < 0.05). Hierarchical dendrograms or bar charts were generated to visually display the microbial groups that were specifically enriched between different groups. Combining these two classification results, the literature was consulted and bacteria that were significantly different in different N2O classification groups and were related to river nitrogen metabolism, such as nitrifying bacteria and denitrifying bacteria, were selected.

[0100] S408) Using FAPROTAX tools and PICRUSt2 tools, the annotation results are mapped to a reference genome database, and the metabolic functional potential of the microbial flora is inferred as functional data.

[0101] Specifically, FAPROTAX maps the taxonomic annotation results to a preset functional database to predict the ecological functions of microorganisms (such as nitrification and methanogenesis). The results show that the functions of different N2O groupings in the watershed are significantly different, and microbial functions are an important factor in mediating the generation and reduction of N2O. PICRUSt2 predicts the gene family composition (KO genes) of the microbial community based on the phylogenetic association of 16S rRNA gene sequences with reference genome databases (such as KEGG), and then infers the abundance of metabolic pathways (such as carbohydrate metabolism and amino acid synthesis) to obtain relevant functional data. KEGG's metabolic pathway annotation can obtain the activities of microorganisms in different metabolic pathways and reveal their roles in the ecosystem.

[0102] Step S4 of this embodiment is intended to identify the optimal dataset among macroscopic features, environmental physical and chemical indicators, and microscopic features. The macroscopic features, environmental physical and chemical indicators, and microscopic features are used as explanatory variables, and the N2O concentration grouping results of each region are used as the response variable. When constructing a classification model for training the corresponding dataset, the following steps are included:

[0103] S41) Data preprocessing.

[0104] Handling missing values ​​in a dataset:

[0105] While the random forest algorithm has a certain tolerance for missing values, most implementations (such as scikit-learn) require complete input data, so a system must handle missing values. In this example, missing numerical features are padded with their mean. Features with a high missingness rate (over 30%) are simply deleted to avoid introducing noise.

[0106] In this example, the raw data for macroscopic features, environmental physicochemical indicators, and microscopic features serve as explanatory variables to form the feature matrix in the dataset. Each row in the feature matrix represents all features of a sample (sampling point or the area where the sampling point is located), and each column represents the same feature. The N2O concentration grouping corresponding to the raw data serves as the classification label, that is, the grouping result (HA or LA) of the area corresponding to the raw data of macroscopic features, environmental physicochemical indicators, or microscopic features serves as the classification label. These classification labels serve as the target variables in the dataset, with each sample corresponding to a target variable. If some features or target variables have missing values, they can be filled (e.g., filling numerical features with the mean and categorical features with the mode) or the rows containing missing values ​​can be deleted. In this example, the filling parameters (e.g., mean) for the training and test sets must be calculated based on the training set to prevent data leakage.

[0107] For example, for medium-scale data sets, features with missing rates greater than a preset threshold in the physical and chemical parameter data of water bodies and sediments are deleted, and the missing values ​​of other features in the physical and chemical parameter data of water bodies and sediments are filled using the mean filling method.

[0108] Feature encoding: Convert the categorical labels of the target variable into numerical form. Specifically, encoding tools such as one-hot encoding are used to convert high N2O concentration areas (HA) and low N2O concentration areas (LA) into two numerical columns.

[0109] Data Standardization: Random Forest split decisions are based on feature ordering, which is dimensionless, so standardization is generally not necessary. However, given the varying scales of different features, all feature data in the dataset in this example is standardized or normalized to transform the data into a distribution with a mean of 0 and a standard deviation of 1.

[0110] S42) Dataset processing.

[0111] The data set is divided into a training set and a test set, wherein the data in the training set accounts for 80% and is used for training the model, and the data in the test set accounts for 20% and is used for testing the model.

[0112] In random forest modeling, an independent test set is used to evaluate the generalization performance of the model. By dividing the dataset into a training set and a test set, you can more comprehensively evaluate the performance of the model in real-world applications, rather than just its performance during training.

[0113] S43) Model training.

[0114] Construct a random forest classification model, determine the optimal parameters of the random forest classification model, and use the training set to train the random forest classification model. Specifically, randomly extract samples from the data in the training set to train the decision tree of the random forest classification model, so that each decision tree in the trained random forest classification model predicts the classification of macro characteristics or environmental physical and chemical indicators or micro characteristics. Vote the prediction results of all decision trees, and use the category with the most votes as the final prediction result.

[0115] Specifically, in this embodiment, the "scikit-learn" Python package (v1.2.2) is used to construct a random forest classification model, and the 5-fold cross-validation of the GridSearchCV algorithm is used to determine the optimal values ​​of the key parameters 'n_estimators' and 'max_depth'. The random forest classification model is an ensemble model composed of multiple decision trees, and "n_estimators" determines the number of trees in the ensemble. Increasing the number of trees can generally improve the performance of the model because it reduces the risk of overfitting and improves the stability of the model. "max_depth" controls the growth depth of the decision tree, and limiting the depth of the tree helps prevent overfitting. A smaller "max_depth" can generally improve the generalization ability of the model and avoid overfitting of the training data.

[0116] S44) Model evaluation.

[0117] The data of the test set is input into the trained random forest classification model, and the evaluation index is calculated based on the classification prediction results of the random forest classification model and the actual classification results of the test set data.

[0118] In this embodiment, the performance evaluation of the random forest classification model includes two key indicators: the Accuracy (accuracy) and AUC (area under the curve) value of the test set. Accuracy represents the proportion of correctly classified instances in the test set, providing a measure of the overall predictive performance, while the AUC value is the area under the ROC curve (receiver operating characteristic curve), which reflects the model's ability to sort positive and negative samples at different classification thresholds to evaluate the performance of the model. Even if the ratio of positive and negative samples is very different (such as 1:1000), AUC can still stably evaluate the performance of the model. The Accuracy and AUC values ​​obtained from 30 tests were statistically analyzed using box plots to mitigate the impact of data sets characterized by extreme restrictions or severe imbalances.

[0119] Accuracy=(TP+TN) / (TP+TN+FP+FN)

[0120] Among them, TP represents true positive examples (the number of samples correctly predicted by the model as positive), TN represents true negative examples (the number of samples correctly predicted by the model as negative), FP represents false positive examples (the number of samples incorrectly predicted by the model as positive), and FN represents false negative examples (the number of samples incorrectly predicted by the model as negative).

[0121] AUC is the area under the ROC curve. Sort the predicted probabilities of the test set samples from high to low, gradually adjust the classification threshold (from 1.0 to 0.0), calculate the (FPR, TPR) under each threshold, draw the ROC curve and calculate the area under the curve (such as the trapezoidal rule):

[0122] AUC=i=1∑n-12TPRi+TPRi+1×(FPRi+1-FPRi)

[0123] Among them, TPR = TP / (TP+FN): true positive rate (sensitivity), FPR = FP / FP+TN: false positive rate (1-specificity).

[0124] The core advantage of AUC (area under the ROC curve) lies in its ability to comprehensively and robustly evaluate the overall performance of binary classification models. First, it is highly insensitive to class distribution. Even in scenarios with a large disparity in the ratio of positive and negative samples, it can still accurately quantify the model's ability to identify minority classes, avoiding evaluation distortion caused by sample imbalance. Second, AUC directly reflects the model's essential ability to distinguish between positive and negative classes by measuring the quality of the model's ranking of the predicted probabilities of positive samples, making it suitable for tasks that require priority discrimination. In addition, AUC circumvents the limitations of single threshold selection and provides a global perspective on the model's generalization ability. At the same time, it does not rely on the preset weights of FP and FN, and can adapt to both balanced cost scenarios and support high cost-sensitive tasks.

[0125] Comparison of the evaluation indicators of three random forest classification models trained by macroscopic features, environmental physical and chemical indicators and microscopic features, the results are as follows Figure 2 As shown, it can be seen that in this embodiment, the random forest classification model trained by environmental physical and chemical indicators (i.e., physical and chemical parameter data of water bodies and sediments) has the best classification effect.

[0126] In step S5 of this embodiment, the physical and chemical parameter data of water and sediment are used as input variables, and a random forest algorithm is used to train and validate a regression model, aiming to construct a regression model with the best prediction effect to predict the dissolved NO concentration. The specific steps include:

[0127] S51) Preprocessing the physical and chemical parameter data of water and sediment, including:

[0128] The features with missing rates greater than the preset threshold in the physical and chemical parameter data of water bodies and sediments are deleted, and the missing values ​​of other features in the physical and chemical parameter data of water bodies and sediments are filled using the mean filling method.

[0129] The physical and chemical parameter data of water and sediments were standardized and normalized.

[0130] The splitting decision of the random forest is based on feature ranking and is not affected by dimension, so normalization is usually not required. However, considering that the scales of different features vary, the physical and chemical parameter data of the water and sediment in this example are normalized by natural logarithm.

[0131] The physical and chemical parameter data set of the processed water body and sediment is used as the explanatory variable, and the N2O dissolved concentration obtained in S1 is used as the response variable to construct the N2O regression model data set. In this data set, the original data of the physical and chemical parameters of the water body and sediment are used as the explanatory variables to form the feature matrix in the data set. Each row in the feature matrix represents all the features of a sample (sampling point or the area where the sampling point is located), and each column represents the same feature. The N2O dissolved concentration corresponding to the original data is used as the target variable in the data set, and each sample corresponds to a target variable. The data set is divided into a training set and a test set. In this embodiment, 80% of the samples in the data set are used as a training set and 20% of the samples are used as a test set. In random forest modeling, an independent test set is used to evaluate the generalization performance of the model. By dividing the data set into a training set and a test set, the effect of the model in practical applications can be more comprehensively evaluated, not just its performance during the training process.

[0132] Next, train the regression model, including:

[0133] S52) constructing a random forest regression model, and similarly determining optimal values ​​of key parameters of the random forest regression model through 5-fold cross validation, and then using the training set to train the random forest regression model, specifically, randomly selecting samples from the data in the training set to train a decision tree of the random forest classification model, so that each decision tree in the trained random forest regression model predicts the N2O dissolved concentration of the physicochemical parameter data of water bodies and sediments, and taking the average of the prediction results of all decision trees as the final prediction result.

[0134] S53) Evaluate the performance of random forest regression models.

[0135] Specifically, the test set is used to verify the performance of the model on an independent test set. The data of the test set is input into the trained random forest regression model. The evaluation index is calculated based on the N2O dissolved concentration prediction results of the random forest regression model and the actual N2O dissolved concentration results of the test set data. In this embodiment, RMSE and R 2 , MAE, MRE, Bias and other indicators are used to evaluate the performance of the regression model.

[0136] RMSE (Root Mean Squared Error), the square root of MSE, has the same unit as the target variable and more intuitively reflects the magnitude of the error. Where y i True value, Predicted value, sample size.

[0137] R 2 The coefficient of determination (Coefficient of Determination) indicates the proportion of the total variation of the target variable explained by the model. The closer the value is to 1, the stronger the explanatory power of the model. Where, The mean of the true values.

[0138] MAE (Mean Absolute Error) measures the average absolute deviation between the predicted value and the true value, reflecting the error size of the model prediction.

[0139] MRE (Mean Relative Error) calculates the average ratio of the prediction error to the true value and reflects the relative error size in percentage form.

[0140] Bias refers to the systematic deviation between the expected and true values ​​of the model's predicted values, reflecting whether the model has overall overestimation or underestimation. In the formula The expected value of the predicted value, the true value of y.

[0141] In this embodiment, step S6 uses Mean Decrease Gini (MDG) as a feature importance index to evaluate the physical and chemical parameters of water and sediment to form an optimal parameter set, aiming to optimize the performance of the regression model. The specific steps include:

[0142] S61) Filter the optimal parameters from the explanatory variables. In this embodiment, the data set of the physicochemical parameter data of the entire water body and sediment needs to be input into the random forest regression model determined by the method of step S5, and then MeanDecrease Gini (MDG) is used as the importance ranking index of the explanatory variables. Finally, the MDG results are statistically analyzed, and the optimal parameters are filtered in descending order to obtain the optimal parameter set. Feature screening is achieved by recursive feature elimination. Recursive Feature Elimination (RFE) is a model-based feature selection method that iteratively trains the model and gradually eliminates the features that contribute the least to the model, and finally filters out the optimal feature subset. The optimal feature prediction subset is total phosphorus (TP) in water, ammonium nitrogen (NH4 + ), total nitrogen (TN), nitrate (NO3 - ), water temperature T and sediment nitrate (NO3 - ). This is a key factor driving dissolved N2O concentrations in rivers.

[0143] S62) uses the optimal feature subset, that is, the optimal feature prediction subset is total phosphorus (TP), ammonium nitrogen (NH4 + ), total nitrogen (TN), nitrate (NO3 - ), water temperature T and sediment nitrate (NO3 - ) as the explanatory variable, the N2O concentration in the target watershed as the response variable, construct the optimal data set and train the best random forest regression model according to the same process of step S5 to predict the N2O concentration. Figure 3 The more accurate predicted value of N2O dissolved concentration is shown. Then, using the thin boundary layer method and the thin boundary diffusion model, the predicted value of N2O dissolved concentration is used to estimate the N2O emission in the river, and the predicted value of N2O emission flux can be obtained.

[0144] S63) Evaluate the generalizability of the optimal regression model: Cross-validate the N2O prediction model trained in the previous step to evaluate model performance. This helps test the model's generalization ability on different datasets and improve the model's robustness. Based on the cross-validation results, optimize performance (further fine-tune parameters or use other optimization methods). Interpret the optimized model to understand the contribution of each parameter in the model. Verify that the model results are consistent with actual observations to ensure that the model's assessment of the physical and chemical parameters of water sediments is credible. Visualize the results of the final model and parameter set to better convey the evaluation results and model performance.

[0145] Machine learning is a black box model. While it can predict emissions, it cannot directly reveal the production and emission mechanisms of N2O. Microbial information is a sensitive reactor that can more accurately reflect the current state of the river. Microbial processes are the direct driving factor of N2O production. By analyzing microbial data, it is possible to clearly identify which microbial functional genes or community changes lead to changes in N2O concentration. Therefore, in step S7 of this embodiment, Spearman correlation analysis is used to further study the relationship between the optimal parameters in the optimal parameter set and the microbial composition and functional data. The following steps are included:

[0146] The water quality-sediment physicochemical parameters are integrated with the N2O characteristic bacteria and characteristic bacterial community function data to ensure that the data have consistent time and space scales; then the correlation coefficient of each parameter with the microbial composition data is obtained by the Pearson analysis method, and the pheatmap diagram showing the relationship between each optimal parameter and the microbial function data is obtained by the Spearman analysis method, so as to explain the relationship between the physicochemical properties and the N2O characteristic bacteria and characteristic bacterial community function data, determine which physicochemical parameters have a more significant impact on microbial species and functions, and analyze the microbial mechanism of the N2O concentration distribution characteristics. Explain the reasons for the generation, reduction, and distribution of N2O from a microbial perspective. Particular attention is paid to the optimal feature prediction subset, namely total phosphorus (TP) and ammonium nitrogen (NH4 + ), total nitrogen (TN), nitrate (NO3 - ), water temperature T and sediment nitrate (NO3 - ) and its relationship with related bacteria.

[0147] Specifically, there is an intricate and close relationship between the physicochemical parameters of water and sediments as environmental factors and the composition of microbial communities. Through Pearson analysis, the color gradient is used to vividly display the correlation coefficient of environmental factors.

[0148] Behind this rich data lies the KO functional data obtained through PICRUSt2 analysis, which provides detailed information about the microbial community. Spearman analysis allows the construction of a pheatmap, which vividly illustrates the detailed relationships between microbial signatures and physicochemical parameters. This provides deeper insight into the complex interactions between N2O-significant bacterial communities and physicochemical parameters, further providing intuitive recommendations for targeted river N2O control.

[0149] The N2O concentration predicted by the N2O prediction model of this embodiment is compared with the actual N2O concentration observation value and the N2O concentration predicted by the IPCC method. Figure 4 As shown, Figure 4 (a) Comparison of predicted N2O concentrations obtained using the random forest regression model with observed values. Figure 4 (b) Comparison of the predicted N2O concentration obtained using the IPCC method and the observed value. It can be seen that compared with the IPCC, the N2O concentration obtained using the random forest prediction subset using the optimal feature is closer to the true value of N2O, which makes up for the shortcomings of the IPCC prediction method.

[0150] In summary, the machine learning method for predicting dissolved N2O concentration in rivers proposed in this paper has the following advantages:

[0151] (1) Overcoming the shortcomings of traditional prediction methods: This method breaks through the traditional IPCC prediction framework that relies on nitrate concentration and emission factors, combines the prediction of N2O concentration by machine learning models with microbial community functional analysis, and reveals the key regulatory mechanism of N2O production by microbial N2O characteristic bacteria (such as nitrifying bacteria / denitrifying bacteria) while improving the accuracy of N2O concentration prediction.

[0152] (2) The key driving factors of N2O were obtained: by integrating various types of data (water quality parameters, meteorology, land use, etc.) through machine learning modeling (such as random forest), and using water quality and sediment parameters, a good prediction effect was achieved. Through feature importance sorting and recursive feature elimination, the most core driving factors, namely water TP, NH4 + TN, NO3 - , water temperature T and sediment NO3 - , and evaluated the contributions of relevant driving factors to N2O concentrations.

[0153] (3) Deepening the mechanism analysis from a microbial perspective: Through feature importance analysis, combined with the correlation between the abundance and KO function of N2O-specific microorganisms and physicochemical parameters, this study elucidates how environmental factors such as temperature and total nitrogen affect N2O production and emissions by regulating microbial community structure. This biological interpretation of the model's nonlinear results provides actionable intervention targets for river N2O emission reduction, closes the "prediction-attribution-regulation" chain, and provides a theoretical basis for optimizing management strategies.

[0154] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed above with reference to the preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent variations, and modifications to the above embodiment that do not depart from the technical solution of the present invention and are based on the technical essence of the present invention shall fall within the scope of protection of the technical solution of the present invention.

Claims

1. A machine learning method for predicting dissolved N2O concentration in rivers, characterized by: The following steps are involved: Calculate the N2O dissolved concentration in each area of ​​the target watershed and group each area into N2O concentration groups based on the N2O dissolved concentration; Obtain the macroscopic characteristics, environmental physicochemical indicators, and microscopic characteristics of each area in the target watershed. The macroscopic characteristics include river information data of the target watershed, the environmental physicochemical indicators include physicochemical parameter data of water bodies and sediments in the target watershed, and the microscopic characteristics include microbial composition and function data of the target watershed. Macroscopic characteristics, environmental physical and chemical indicators, and microscopic characteristics were used as explanatory variables, and the N2O concentration grouping results of each region were used as response variables to construct a classification model for the dataset training. The explanatory variables of the classification model with the best classification effect were selected, and the N2O concentration in the target watershed was used as the response variable. A new dataset was constructed and a regression model was trained to predict the N2O concentration. The optimal parameters are selected from the explanatory variables, and the optimal parameters are used as the explanatory variables. The N2O dissolved concentration in the target watershed is used as the response variable. The optimal dataset is constructed and the optimal regression model is trained. The N2O dissolved concentration predicted by the optimal regression model is obtained, and the predicted N2O emission flux value is calculated based on the N2O dissolved concentration.

2. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: When calculating the dissolved N2O concentration in each area of ​​the target watershed, water samples are collected at sampling points in each area of ​​the target watershed. The water samples are slowly discharged into the bottom of the sample bottle through a silicone tube with minimal turbulence. After the sample bottle is filled with water, the water is allowed to overflow to avoid contamination of the water sample by N2O in the atmosphere. HgCl2 is then injected into the sample bottle to a final concentration of 0.5% v / v to inhibit microbial activity. The sample bottle is stored at low temperature and the dissolved N2O in the water of the sample bottle is determined by headspace equilibrium-gas chromatography.

3. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: When obtaining the macro characteristics of each area in the target basin, the specific step is to construct a hydrological model of the target basin, introduce the pollution source database into the hydrological model, parameterize and calibrate the hydrological model, and obtain the river information data of each area output by the calibrated hydrological model.

4. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: When obtaining the environmental physicochemical indicators of each area in the target watershed, the physicochemical parameters of the water body and sediment are measured through in-situ monitoring and calculation at the sampling points in each area of ​​the target watershed. The physicochemical parameters of the water body include one or more of the concentrations of total nitrogen, ammonium nitrogen, nitrate and total phosphorus in the water; the physicochemical parameters of the sediment include one or more of the concentrations of total nitrogen, ammonium nitrogen, nitrate and total phosphorus in the sediment.

5. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: When obtaining the microscopic characteristics of each area in the target basin, specifically, sediment samples are collected at sampling points in each area of ​​the target basin, the microbial 16S genome of the sediment samples is sequenced, and the sequencing data is matched with a preset database to obtain microbial composition data and functional data.

6. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: Macroscopic characteristics, environmental physical and chemical indicators, and microscopic characteristics are used as explanatory variables, and the N2O concentration grouping results of each region are used as response variables. When constructing the dataset to train the corresponding classification model, the following steps are included: Dividing the data set into a training set and a test set; Construct a random forest classification model, determine the optimal values ​​of key parameters of the random forest classification model through 5-fold cross-validation, and train the random forest classification model using the training set so that each decision tree in the trained random forest classification model predicts the classification of macroscopic characteristics, environmental physical and chemical indicators, or microscopic characteristics. Vote on the prediction results of all decision trees, and use the category with the most votes as the final prediction result. The data of the test set is input into the trained random forest classification model, and the evaluation indicators are calculated based on the classification prediction results of the random forest classification model and the actual classification results of the data of the test set. The evaluation indicators include the accuracy and AUC value of the test set.

7. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: The explanatory variables of the classification model with the best classification effect are environmental physical and chemical indicators, and the optimal parameters are total phosphorus, total nitrogen, ammonium nitrogen, nitrate, water temperature T and sediment nitrate in the physical and chemical parameter data of water and sediment of the environmental physical and chemical indicators.

8. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: When training a regression model and training an optimal regression model, both include: Dividing the new data set or the optimal data set into a training set and a test set; A random forest regression model was constructed and the optimal values ​​of its key parameters were determined through 5-fold cross-validation. The model was then trained using the training set, so that each decision tree in the trained random forest regression model predicted the dissolved N2O concentration based on the physicochemical parameter data of water and sediments. The prediction results of all decision trees were averaged as the final prediction result. The test set data is input into the trained random forest regression model, and evaluation indicators are calculated based on the N2O dissolved concentration prediction results of the random forest regression model and the actual N2O dissolved concentration results of the test set data. The evaluation indicators include one or more of the root mean square error, determination coefficient, mean absolute error, mean relative error, and bias deviation.

9. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: When selecting the optimal parameters from the explanatory variables, include: Input the data set into the random forest regression model and calculate the MDG results of each explanatory variable of the random forest regression model; The MDG results were statistically analyzed and the optimal parameters were screened in descending order of ranking.

10. The machine learning method for predicting dissolved N2O concentration in a river according to claim 1, characterized in that: After selecting the optimal parameters from the explanatory variables, the process also includes: analyzing the correlation between the optimal parameters and the microbial composition and function data. Specifically, the Pearson analysis method is used to obtain the correlation coefficient between each optimal parameter and the microbial composition data, and the Spearman analysis method is used to obtain a pheatmap diagram showing the relationship between each optimal parameter and the microbial function data.

Citation Information

Patent Citations

  • River dissolved oxygen concentration prediction method and system based on maximum information coefficient

    CN116384541A

  • Machine learning method for tracing nitrogen-containing pollutants in river based on microbial metagenomics

    CN118173178A

  • Method to predict the effluent ammonia-nitrogen concentration based on a recurrent self-organizing neural network

    US20160140437A1

  • Method of Determining River Nitrous Oxide Emission based on Land-River-Atmosphere Simulation

    US20240321403A1

  • AU2020100709A4