A Method for Predicting and Regulating Groundwater Nitrogen Metabolism Function Based on Three-Dimensional Fluorescence Spectroscopy and Automated Machine Learning

By combining three-dimensional fluorescence spectroscopy with automated machine learning, key fluorescence peaks and characteristic components were identified, solving the problem of predicting nitrogen metabolism function in heterogeneous groundwater data from multiple sources. This approach enables low-cost, rapid, and accurate prediction of nitrogen metabolism function, providing efficient technical support for groundwater ecological monitoring.

CN121983143BActive Publication Date: 2026-06-30HOHAI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HOHAI UNIV
Filing Date
2026-04-09
Publication Date
2026-06-30

Smart Images

  • Figure CN121983143B_ABST
    Figure CN121983143B_ABST
Patent Text Reader

Abstract

This invention relates to a method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning, comprising the following steps: collecting groundwater microbial gene sequencing sequence numbers and corresponding water sample fluorescence emission spectrum data from the literature; downloading microbial gene sequencing data based on the acquired gene sequencing data sequence numbers, and normalizing the required emission spectrum data; selecting functional genes reflecting groundwater nitrogen metabolism capacity and calculating their relative abundance; and constructing a machine learning model for groundwater nitrogen metabolism. The optimal model selected in this invention can accurately predict and understand groundwater nitrogen metabolism, reveal the impact of DOM fluorescence properties on groundwater ecosystems, and can be used to predict the nitrogen metabolism function of groundwater microorganisms in unknown areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning, belonging to the technical field of environmental science and technology. Background Technology

[0002] Groundwater is an indispensable natural resource, providing nearly half of the world's drinking water. Groundwater is susceptible to pollution from nitrogen-based fertilizers, and nitrates and nitrites can be converted into carcinogenic nitrosamines. Microbial-mediated nitrogen transformation dominates groundwater nitrogen metabolism. Studies have shown that changes in DOM (microbial activity) properties affect the gene abundance of microorganisms involved in nitrogen metabolism, thereby altering groundwater nitrogen transformation processes. Therefore, establishing a predictive method for the correlation between DOM characteristics and groundwater microbial nitrogen metabolism function is of great significance for groundwater ecological security assessment.

[0003] Three-dimensional fluorescence spectroscopy (3D-EEM) is a commonly used technique for rapidly assessing the properties of groundwater precipitates (DOMs). However, the raw data contains thousands of excitation / emission (Ex / Em) pairs, leading to noise and peak overlap issues. While parallel factor analysis (PARAFAC) can resolve independent fluorescent components, it is limited by stringent split-half test requirements, initial value sensitivity, and high computational cost, making it difficult to meet the needs of rapid screening of large-scale samples. Particularly noteworthy is the inherent mechanism of PARAFAC, which requires all samples to share the same fixed number of components and stable spectral profiles due to the spatial dispersion of groundwater sampling wells, limited sample size per study, and strong heterogeneity of multi-source data. This leads to cross-confusion of component spectra and averaging dilution of region-specific signals when merging EEM data from different aquifers, seasons, or study areas, making it difficult to extract robust feature parameters that can be compared across datasets.

[0004] Principal component analysis (PCA) can achieve rapid dimensionality reduction through singular value decomposition, without the need for pre-defined component numbers and complex validation, and is better suited to multi-source heterogeneous data. However, traditional PCA methods require flattening EEM data into a one-dimensional vector, resulting in the loss of the two-dimensional coupling relationship between Ex and Em; moreover, PCA loadings are abstract weight vectors that cannot be directly mapped to fluorescence feature regions, making it difficult to extract DOM feature parameters with clear ecological significance for subsequent modeling and analysis.

[0005] In recent years, groundwater microbial research has accumulated a large amount of sequencing data, and machine learning has provided an efficient tool for analyzing the relationship between multivariates and functional labels. However, directly using raw high-dimensional EEM data as input poses risks of the curse of dimensionality, noise interference, and overfitting; using PCA abstract scores or PARAFAC components as input loses the physical identifiers of the original Ex / Em, making it impossible to trace the specific fluorescent substances on which the model predictions rely. Therefore, there is an urgent need for a technical solution that combines rapid EEM data dimensionality reduction with machine learning. This solution should be able to extract interpretable and cross-dataset comparable DOM feature parameters through improved PCA methods, and utilize these parameters to accurately predict the nitrogen metabolism function of groundwater microorganisms, providing efficient and reliable technical support for large-scale groundwater ecological monitoring. Summary of the Invention

[0006] To address the aforementioned problems, this invention discloses a method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning. The specific technical solution is as follows:

[0007] A method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning includes the following steps:

[0008] Step (1): Collect groundwater microbial gene sequencing sequence numbers and corresponding three-dimensional fluorescence spectral data of water samples;

[0009] Step (2): Based on the obtained groundwater microbial gene sequencing sequence number, download the groundwater microbial gene sequencing data, process the microbial gene sequencing data using bioinformatics analysis software, classify and annotate the species, obtain the relative abundance of microbial functional genes, and perform logarithmic transformation on all relative abundance data.

[0010] Step (3): Select functional genes that reflect the nitrogen metabolism capacity of groundwater environment. Functional genes include denitrification genes, nitrite respiration genes, nitrate reduction genes and urea decomposition genes. The relative abundance of functional genes is used as a predictive variable. Different feature groups are used as input. Feature group selection: Feature group 1) Spectral information obtained from 3D-EEM based on expanded principal component U-PCA analysis; Feature group 2) Spectral information obtained from 3D-EEM based on PARAFAC analysis; Feature group 3) Water quality data;

[0011] Step (4): Construct 4 ML models and 2 automatic machine learning models. The ML models include K-Nearest Neighbor (KNN), XGBOOST, RF, and GA-SVM optimized by genetic algorithm. The automatic machine learning models include TPOT and H2O. Optimal model for predicting groundwater environmental microbial function based on 3D-EEM spectroscopy is determined, and the accuracy of the optimal model is evaluated.

[0012] Step (5): Sort the importance of features based on Shapley values, identify features that have a direct causal relationship with microbial nitrogen metabolism function based on causal analysis, construct the numerical response relationship between causal relationship pairs, and identify the critical point at which the effect of fluorescence features on nitrogen metabolism function changes from positive to negative.

[0013] Furthermore, in step (1), the gene sequencing sequence number is the 16S rRNA sequencing data sequence number.

[0014] Furthermore, step (1) is performed according to the following steps:

[0015] (1) Search for literature on groundwater microorganisms, high-throughput sequencing, DOM, and mark the sequencing sequence numbers of groundwater microorganism genes to obtain microbial sequencing data points; use Python program to remove Raman and Rayleigh scattering from 3D-EEM spectra and form spectra with the same size;

[0016] (2) Collect the characteristic variables of microbial sequencing data points and use the FAPROTAX database to predict the relative abundance of microbial functional genes.

[0017] Furthermore, in step (3), feature group 1) is: fluorescence intensity of DOM key Ex / Em pairs identified using U-PCA, as well as fluorescence index FI, freshness index β:α, microbial origin index BIX and humification index HIX;

[0018] Feature group 2) consists of DOM components identified using PARAFAC and their maximum fluorescence peak intensities, and DOM spectral indices calculated, including FI, β:α, BIX, and HIX.

[0019] Feature group 3) consists of water quality data, including pH, conductivity, temperature, dissolved oxygen, dissolved organic carbon, chloride ions, nitrate, nitrite, sulfate, magnesium ions, calcium ions, potassium ions, and sodium ions.

[0020] Furthermore, in step (3), U-PCA analysis unfolds the 3D-EEM spectral data into a one-dimensional vector along the excitation-emission plane, extracts the top n principal components with a cumulative variance contribution rate of 95% through singular value decomposition, identifies the excitation / emission wavelength positions with high weights in each principal component load, screens the fluorescence intensity of key fluorescence regions, and organizes the fluorescence intensity of key fluorescence regions of all samples to form input data for predicting nitrogen metabolism; the singular value decomposition formula is as follows:

[0021]

[0022] In the formula, X is the standardized data matrix n×m, where n is the number of samples and m is the number of features; U is the left singular vector matrix n×n, S is the singular value diagonal matrix n×m, and V... T It is the transpose of the right singular vector matrix, m×m;

[0023] To extract a cumulative variance contribution rate of 95%, the sum of squares of the first k singular values ​​of S must account for ≥0.95 of the sum of squares of all singular values. That is, we take the smallest k that satisfies the following formula.

[0024]

[0025] in It is the first singular value matrix S on the diagonal. j There are elements, where m is the number of all singular values.

[0026]

[0027] In the formula, The principal component matrix after dimensionality reduction has dimensions n×k, where k represents the top k principal components with a cumulative variance contribution rate of 95%. Let m × k be the right singular vector matrix, and let the right singular matrix be the loading matrix. The corresponding principal components are superimposed, and the data structure is reconstructed to identify the most critical fluorescence peaks.

[0028] Furthermore, in step (3), the PARAFAC analysis uses the alternating least squares method to decompose the 3D-EEM data, decomposing the EEM datasets of multiple samples into individual fluorescent components with specific excitation emission spectral characteristics.

[0029] Furthermore, in step (4), the model accuracy evaluation index is selected using the coefficient of determination R. 2 and Mean Absolute Error (MSE), R 2 The larger the value, the better the model fits the data and the stronger the explanatory power of the independent variable on the dependent variable. The smaller the value, the closer the model's predicted value is to the true value and the higher the prediction accuracy.

[0030] The R2 The calculation formula is:

[0031]

[0032] Where n is the number of samples, This represents the observed value of groundwater microbial functional abundance. Predicted values ​​for the functional abundance of groundwater microorganisms; This represents the average value of observed microbial functional abundance in groundwater.

[0033] The formula for calculating MSE is:

[0034]

[0035] in, This represents the observed value of groundwater microbial functional abundance. This is a predicted value for the functional abundance of groundwater microorganisms.

[0036] Furthermore, in step (4), the five-fold cross-validation method is used to randomly divide the test set samples into five subsets of similar size, and then the four ML models and two automatic machine learning models are evaluated and trained five times each to evaluate the model accuracy.

[0037] Furthermore, the formula for calculating the Shapley value in step (5) is as follows:

[0038]

[0039] in Let be the Shapley value of feature i in any model f constructed based on dataset x, and M be the total number of all input features. It is the set of all feature combinations containing feature i. It is a combination of features The number of features in They are based on and Train different prediction models.

[0040] Furthermore, in step (5), the causal analysis adopts the dual machine learning method DML, using the optimal model obtained in step (3) as the base learner, to calculate the average causal effect ATE of each fluorescence feature on nitrogen metabolism. When the ATE is statistically significant, it is considered that the feature has a direct causal relationship with the result.

[0041] The process of constructing the numerical response relationship between causal pairs is as follows: it is achieved through the Individual Conditional Expectation Map (ICE), by keeping all other features unchanged and changing only one fluorescence feature, and observing the output of the optimal model to obtain a continuous response curve.

[0042] The beneficial effects of this invention are:

[0043] 1. Microorganisms play a role in nitrogen metabolism through their metabolism, but due to the complexity of groundwater systems, accurate prediction and understanding of groundwater nitrogen metabolism is helpful for the ecological restoration of natural water bodies.

[0044] 2. By collecting 16S rRNA sequencing data and DOM fluorescence spectra of 196 groundwater samples from across continents from literature, and based on the H2O automated machine learning model, we revealed the impact of DOM fluorescence properties on groundwater ecosystems, which can be used to predict the nitrogen metabolism function of groundwater microorganisms in unknown areas.

[0045] 3. Traditional prediction of microbial nitrogen metabolism relies on the simultaneous acquisition of multiple water quality parameters (such as pH, dissolved oxygen, nitrate, etc.), which is costly and difficult, especially in well-hidden groundwater environments. Three-dimensional fluorescence data has been widely used as input for machine learning models in water environment prediction. This invention is the first to verify the feasibility of predicting groundwater nitrogen metabolism using only 3D-EEM data as a single input source, overcoming the technical bottleneck of simultaneous multi-parameter acquisition and significantly reducing monitoring costs and field operation complexity.

[0046] 4. Addressing the challenges of scattered groundwater sampling well locations, limited sample size in single studies, and strong heterogeneity of multi-source data, the U-PCA feature selection method employed in this invention offers significant advantages over PARAFAC: it eliminates the need for pre-setting fixed component numbers, avoids the stringent split-half verification process, is insensitive to initial values, and boasts fast computation speed and stable results. In particular, it can tolerate spectral pattern differences in EEM data from different aquifers, seasons, or study areas, avoiding the cross-confusion of component spectra and the dilution of regionally specific signals caused by forced uniform modeling, effectively extracting robust feature parameters comparable across datasets. Furthermore, this invention introduces automated machine learning technology, replacing the cumbersome manual feature engineering, algorithm selection, and hyperparameter tuning processes of traditional machine learning, achieving full automation from spectral data acquisition to nitrogen metabolism risk prediction. Comparative experiments demonstrate that the model based on U-PCA fluorescence parameters outperforms both multi-water quality parameter models and PARAFAC fluorescence parameter models in prediction accuracy, providing an economical, efficient, and scalable technical means for rapid screening of groundwater nitrogen pollution and ecological safety assessment. Attached Figure Description

[0047] Figure 1 This is a flowchart of the present invention.

[0048] Figure 2 This is a schematic diagram illustrating the principle of U-PCA processing of raw 3D-EEM data and identification of key fluorescence peaks according to the present invention.

[0049] Figure 3 This is a schematic diagram of the key Ex / Em fluorescence peaks identified by U-PCA used in this invention.

[0050] Figure 4 This invention uses three sets of features as input to four classical ML models and two automated ML models to predict nitrate reduction function R. 2 The model performance evaluation graph obtained from MSE.

[0051] Figure 5 This invention uses three sets of features as input to four classical ML models and two automated ML models to predict denitrification function R. 2 The model performance evaluation graph obtained from MSE.

[0052] Figure 6 This invention uses three sets of features as input to four classical ML models and two automated ML models respectively, to predict the R-value of urea decomposition function. 2 The model performance evaluation graph obtained from MSE.

[0053] Figure 7 This invention uses three sets of features as input to four classical ML models and two automated ML models to predict nitrite respiratory function R. 2 The model performance evaluation graph obtained from MSE.

[0054] Figure 8 This is a ranking chart of the importance of Shapley features for predicting four nitrogen metabolism functions using the H2O automatic machine learning model based on fluorescence peaks and fluorescence parameters identified by U-PCA as feature inputs.

[0055] Figure 9 The present invention uses the DML method with the H2O automatic machine learning model as the base learner to estimate the statistically significant ATE plot.

[0056] Figure 10 The diagram shows the ICE sequence between the fluorescence peaks and nitrate reduction, with the Shapley value ranking higher.

[0057] Figure 11 The diagram shows the ICE sequence between the fluorescence peak 3-urea decomposition and the peak ranking high for Shapley values.

[0058] Figure 12 The diagram shows the ICE sequence between the fluorescence peaks 2 and denitrification, with the Shapley value ranking higher.

[0059] Figure 13 The diagram shows the ICE sequence between the fluorescence peak 2-nitrite respiration, with the Shapley value ranking first. Detailed Implementation

[0060] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0061] Combined with appendix Figure 1-13 This section describes a specific embodiment of the invention in practical application; the overall process can be found in [link to full description]. Figure 1 :

[0062] Step 1. Collect water quality parameters, microbial 16S rRNA amplicon sequencing data, and corresponding 3D-EEM data from 193 groundwater samples from different water bodies from publicly published peer-reviewed literature.

[0063] Step 2. The 16S rRNA data were processed using the FAPROTAX tool to obtain 17 nitrogen metabolism-related functional genes of the microbial community. Given that some nitrogen metabolism functional genes were not detected in more than 50% of the samples or had extremely low relative abundance, to avoid the impact of sparse data on model performance, four widely distributed nitrogen metabolism functions (denitrification genes, nitrite respiration genes, nitrate reduction genes, and urea decomposition genes) were ultimately selected as prediction targets. Due to the skewness of the abundance distribution of functional bacteria, a logarithmic transformation was performed on all abundance data.

[0064] Step 3. Preprocess all 3D-EEM data to unify spectral dimensions, remove Rayleigh scattering, and perform Raman normalization. To identify the key fluorescence peaks with the greatest sample discrimination ability, all preprocessed 3D-EEM data are expanded into one-dimensional vectors for principal component analysis (PCA). The top n principal components with a cumulative variance contribution rate of 95% are extracted, and the one-dimensional vector loading data are reconstructed into three dimensions (see [link to U-PCA method] for details). Figure 2 ), Figure 2 S1, S2, and S3 represent 3D-EEM fluorescence spectral data of three different samples. The raw data for each sample is in the form of a three-dimensional matrix: excitation wavelength - emission wavelength - corresponding fluorescence intensity at the excitation and emission wavelengths. PC1 represents the direction of greatest variation in sample dispersion, PC2 is the second largest variation direction orthogonal to PC1, and PC3 represents another variation direction orthogonal to both PC1 and PC2. These three together explain 95% of the variation among the variables. Contour plots were drawn to identify the high-weight excitation / emission wavelength positions in each principal component loading, and several key fluorescence peaks that best distinguish the differences between samples were selected. See [link to relevant documentation]. Figure 3The three-dimensional EEM matrix (Ex × Em) of each sample is flattened into a one-dimensional vector in a row- or column-major order. All one-dimensional vectors from all samples are stacked vertically to construct a two-dimensional data matrix X. PCA decomposition is performed on matrix X to extract the loading matrix V (principal component directions) and the score matrix U. The one-dimensional loading vectors of PC1, PC2, and PC3 are then reconstructed back to the original two-dimensional Ex-Em grid dimensions. Contour plots are drawn by overlaying the reconstructed PC1, PC2, and PC3 loading matrices, and the loading weight distribution is displayed through color gradients or contour lines. The excitation / emission wavelength coordinates corresponding to high absolute value loading weights are identified in the plots. These high-weight peaks are the key fluorescence peaks with the greatest sample discrimination ability, used for subsequent input variable selection in nitrogen metabolism function prediction models. The peak intensities of the key fluorescence peaks corresponding to each sample are also compiled. PARAFAC decomposes all preprocessed 3D-EEM data to identify and quantify five main fluorescent components and their maximum fluorescence intensities. Common DOM spectral indices, including FI, β:α, BIX, and HIX, are calculated.

[0065] Step 4. Construct three sets of input features: Combination 1 consists of the intensity of key fluorescence peaks identified by U-PCA and four DOM spectral indices (fluorescence index FI, freshness index β:α, microbial origin index BIX, and humification index HIX); Combination 2 consists of the intensity of fluorescence components resolved by PARAFAC and four DOM spectral indices (fluorescence index FI, freshness index β:α, microbial origin index BIX, and humification index HIX); Combination 3 consists of 13 collected water quality parameters (pH, conductivity, temperature, dissolved oxygen, dissolved organic carbon, chloride ions, nitrate, nitrite, sulfate, magnesium ions, calcium ions, potassium ions, and sodium ions). To compare the performance of ML models with different input combinations (four traditional ML models including KNN, XGBOOST, RF, and GA-SVM, and two AutoML models including TPOT and H2O), R was selected. 2 MSE and other metrics are used to evaluate model performance. For four traditional ML models, a grid search approach is used to adjust the three most important parameters to optimize model performance. For the TPOT AutoML model framework, performance is automatically optimized by setting the number of generations of genetic evolution and the population size. The H2O AutoML model is configured with training time to automatically explore various algorithms such as generalized linear models, deep neural networks, and gradient boosting machines, and uses a stacked ensemble strategy to automatically select the optimal model.

[0066] Step 5. Figures 4-7 This demonstrates the prediction of four nitrogen metabolism functions using three different combinations of input features. Figure 4 Prediction of Nitrate_reduction function Figure 5For the prediction of denitrification function, Figure 6 For the functional prediction of urea decomposition, Figure 7 Performance comparison of six machine learning models (GA-SVM, KNN, RF, XGB, TPOT, H2O) for predicting nitrite respiration abundance. The radar chart on the left shows the performance of R... 2 The right column shows the MSE comparison. Red represents input combination 1: fluorescence peak intensity identified by U-PCA combined with fluorescence parameters; purple represents input combination 2: maximum fluorescence peak intensity of fluorescent components identified by PARAFAC combined with fluorescence parameters; blue represents input combination 3: water quality parameters. Points further out on the radar chart on the left indicate R... 2 The larger the value, the better the model's performance; conversely, the smaller the MSE (Mean Sequence Size) on the right, the better the model's performance. Besides the TPOT AutoML model with feature combination 1 as input showing better performance in predicting urea decomposition, the H2O AutoML model with feature combination 1 as input also showed the best performance in predicting denitrification, nitrate reduction, and nitrite respiration. When predicting denitrification, the R... 2 The value was 0.83, and the MSE was 0.02; nitrate reduction R 2 The value was 0.83, and the MSE was 0.012; nitrite respiration R 2 The value was 0.79, and the MSE was 0.16; urea decomposition R 2 The value was 0.45, and the MSE was 0.24. The models automatically optimized by AutoML (H2O and TPOT) outperformed those manually optimized by traditional ML models (KNN, XGBOOST, RF, and GA-SVM) networks through hyperparameter search, and fluorescence parameters were better used as inputs to predict nitrogen metabolism than water quality parameters. In particular, using the fluorescence intensity of key fluorescence peaks identified by U-PCA as input features better predicted groundwater nitrogen metabolism than using the fluorescence intensity of fluorescent components separated by PARAFAC.

[0067] Step 6. To gain a deeper understanding of the nonlinear relationship between key fluorescence peaks and nitrogen metabolism functions, the Shapley value combined with the ATE value method was used to perform interpretability analysis on the model results. Based on the Shapley value, the importance of features was ranked to obtain the importance ranking for the prediction of four types of nitrogen metabolism (urea decomposition, nitrite respiration, nitrate reduction, and denitrification). See [link to relevant documentation]. Figure 8Taking the prediction of urea decomposition function as an example, fluorescence peak 4 is the input parameter with the highest special importance, indicating that among the nine input features of feature combination 1, the fluorescence intensity of fluorescence peak 4 has the most significant impact on the accurate prediction of urea decomposition function. DML is used to identify fluorescence peak-nitrogen metabolism function pairs with direct causal relationships, thereby identifying fluorescence parameters with direct causal relationships for the prediction of four nitrogen metabolism functions (…). Figure 9 ), Figure 9 Significance level: · represents p < 0.05; ·· represents p < 0.01, ··· represents p < 0.001. Based on the combined model contribution and causal inference results, the top three features in terms of Shapley importance and fluorescence parameter pairs with a direct causal relationship to nitrogen metabolism were selected. Subsequently, ICE was used to establish the numerical relationship between these DOM fluorescence parameters and the corresponding microbial nitrogen metabolism, see [link to relevant documentation]. Figures 10-13 , Figure 10 The ICE diagram showing the fluorescence peaks between 3-nitrate reduction and the peaks ranked high by Shapley value is shown. Figure 11 The ICE diagram showing the fluorescence peak between 3-urea decomposition and the peak ranked high for Shapley value is provided. Figure 12 The diagram shows the ICE curve between the fluorescence peaks 2 and denitrification, indicating the peaks with higher Shapley values. Figure 13 The diagram shows the ICE sequence between the fluorescence peak 2-nitrite respiration, with the Shapley value ranking first. Figures 10-13 The 0th value represents the predicted response curve corresponding to the sample with the smallest value for that feature among all samples; 10th represents values ​​exceeding 10%; 20th represents values ​​exceeding 20%; and so on up to 100th, which represents 100%. The dashed line represents the average of all ICE curves, reflecting the overall relationship between the feature and the prediction target. The gray area represents the actual sample distribution density of that feature in the training data. The higher the gray bar, the denser the samples in that area, which can then be used to assess the nitrogen metabolism function of groundwater at a specific location and guide the formulation of groundwater ecological restoration policies.

[0068] In summary, this invention develops a complete workflow integrating fluorescence spectroscopy analysis and automated machine learning. This workflow first identifies key fluorescence peaks using the U-PCA method and uses them as feature inputs to the H2O AutoML model, achieving high-precision prediction of nitrogen metabolism functions such as groundwater denitrification and nitrate reduction. Based on this, key fluorescence peaks with direct regulatory effects on specific nitrogen metabolism functions are identified using SHAP and DML causal inference methods; and combined with ICE, the numerical response relationship between DOM fluorescence peaks and nitrogen metabolism functions is further clarified. This invention provides novel technical means and scientific basis for the accurate assessment and targeted regulation of groundwater nitrogen metabolism functions.

[0069] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0070] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning, characterized in that, Includes the following steps: Step (1): Collect groundwater microbial gene sequencing sequence numbers and corresponding three-dimensional fluorescence spectral data of water samples; Step (2): Based on the obtained groundwater microbial gene sequencing sequence number, download the groundwater microbial gene sequencing data, process the microbial gene sequencing data using bioinformatics analysis software, classify and annotate the species, obtain the relative abundance of microbial functional genes, and perform logarithmic transformation on all relative abundance data. Step (3): Select functional genes that reflect the nitrogen metabolism capacity of groundwater environment. Functional genes include denitrification genes, nitrite respiration genes, nitrate reduction genes and urea decomposition genes. The relative abundance of functional genes is used as a predictive variable. Different feature groups are used as input. Feature group selection: Feature group 1) Spectral information obtained from 3D-EEM based on expanded principal component U-PCA analysis; Feature group 2) Spectral information obtained from 3D-EEM based on PARAFAC analysis; Feature group 3) Water quality data; Step (4): Construct 4 ML models and 2 automatic machine learning models. The ML models include K-Nearest Neighbor (KNN), XGBOOST, RF, and GA-SVM optimized by genetic algorithm. The automatic machine learning models include TPOT and H2O. Optimal model for predicting groundwater environmental microbial function based on 3D-EEM spectroscopy is determined, and the accuracy of the optimal model is evaluated. Step (5): Sort the importance of features based on Shapley values, identify features that have a direct causal relationship with microbial nitrogen metabolism function based on causal analysis, construct the numerical response relationship between causal relationship pairs, and identify the critical point at which the effect of fluorescence features on nitrogen metabolism function changes from positive to negative.

2. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 1, characterized in that, In step (1), the gene sequencing sequence number is the 16S rRNA sequencing data sequence number.

3. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 1, characterized in that, Step (1) shall be performed in accordance with the following steps: (1) Search for literature on groundwater microorganisms, high-throughput sequencing, DOM, and mark the sequencing sequence numbers of groundwater microorganism genes to obtain microbial sequencing data points; use Python program to remove Raman and Rayleigh scattering from 3D-EEM spectra and form spectra with the same size; (2) Collect the characteristic variables of microbial sequencing data points and use the FAPROTAX database to predict the relative abundance of microbial functional genes.

4. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 1, characterized in that, In step (3), feature group 1) is: fluorescence intensity of DOM key Ex / Em pairs identified by U-PCA, as well as fluorescence index FI, freshness index β:α, microbial origin index BIX and humification index HIX; Feature group 2) consists of DOM components identified using PARAFAC and their maximum fluorescence peak intensities, and DOM spectral indices calculated, including FI, β:α, BIX, and HIX. Feature group 3) consists of water quality data, including pH, conductivity, temperature, dissolved oxygen, dissolved organic carbon, chloride ions, nitrate, nitrite, sulfate, magnesium ions, calcium ions, potassium ions, and sodium ions.

5. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 4, characterized in that, In step (3), U-PCA analysis unfolds the 3D-EEM spectral data into a one-dimensional vector along the excitation-emission plane. Singular value decomposition extracts the top n principal components with a cumulative variance contribution rate of 95%. It identifies the excitation / emission wavelength positions with high weights in each principal component load, screens the fluorescence intensity of key fluorescence regions, and organizes the fluorescence intensity of key fluorescence regions from all samples to form input data for predicting nitrogen metabolism. The singular value decomposition formula is as follows: ; In the formula, X is a standardized data matrix n x m, n is the number of samples, and m is the number of characteristics; U is a left singular vector matrix n x n, S is a singular value diagonal matrix n x m, V T is the transpose of a right singular vector matrix m x m; To extract a cumulative variance contribution rate of 95%, the sum of squares of the first k singular values ​​of S must account for ≥0.95 of the sum of squares of all singular values. That is, we take the smallest k that satisfies the following formula. ; in It is the first singular value matrix S on the diagonal. j There are elements, where m is the number of all singular values; ; In the formula, The principal component matrix after dimensionality reduction has dimensions n×k, where k represents the top k principal components with a cumulative variance contribution rate of 95%. Let m × k be the right singular vector matrix, and let the right singular matrix be the loading matrix. The corresponding principal components are superimposed, and the data structure is reconstructed to identify the most critical fluorescence peaks.

6. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 4, characterized in that, In step (3), the PARAFAC analysis uses the alternating least squares method to decompose the 3D-EEM data, decomposing the EEM datasets of multiple samples into individual fluorescent components with specific excitation emission spectral characteristics.

7. The method for predicting and regulating groundwater nitrogen metabolism function based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 1, characterized in that, The model precision evaluation index selected in the step (4) is a coefficient R 2 and a mean absolute error MSE, R 2 The greater the coefficient R is, the better the fitting effect of the model is, the stronger the explanatory power of the independent variable to the dependent variable is, the smaller the MSE is, the closer the predicted value of the model to the true value is, and the higher the prediction precision is. The R 2 The calculation formula is: ; Where n is the number of samples, This represents the observed value of groundwater microbial functional abundance. Predicted values ​​for the functional abundance of groundwater microorganisms; This represents the average value of observed microbial functional abundance in groundwater. The formula for calculating MSE is: ; in, This represents the observed value of groundwater microbial functional abundance. This is a predicted value for the functional abundance of groundwater microorganisms.

8. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 1, characterized in that, In step (4), the five-fold cross-validation method is used to randomly divide the test set samples into five subsets of similar size. Then, the four ML models and the two automatic machine learning models are evaluated and trained five times each to evaluate the model accuracy.

9. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 1, characterized in that, The formula for calculating the Shapley value in step (5) is as follows: ; in Let be the Shapley value of feature i in any model f constructed based on dataset x, and M be the total number of all input features. It is the set of all feature combinations containing feature i. It is a combination of features The number of features in They are based on and Train different prediction models.

10. The method for predicting and regulating groundwater nitrogen metabolism based on three-dimensional fluorescence spectroscopy and automated machine learning according to claim 1, characterized in that, In step (5), the causal analysis adopts the dual machine learning method DML, using the optimal model obtained in step (3) as the base learner, to calculate the average causal effect ATE of each fluorescence feature on nitrogen metabolism. When the ATE is statistically significant, it is considered that the feature has a direct causal relationship with the result. The process of constructing the numerical response relationship between causal pairs is as follows: it is achieved through the Individual Conditional Expectation Map (ICE), by keeping all other features unchanged and changing only one fluorescence feature, and observing the output of the optimal model to obtain a continuous response curve.