Method for estimating nucleotide metabolic function intensity of water body based on DOM optical characteristics

By constructing a multiple linear regression model based on the optical properties of DOM, the problems of complex operation and high cost of traditional water nucleotide metabolism detection are solved. This enables rapid and low-cost estimation of water nucleotide metabolism and genetic molecular information, improving the spatiotemporal resolution and real-time monitoring capabilities of the detection.

CN122347993APending Publication Date: 2026-07-07JILIN JIANZHU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN JIANZHU UNIVERSITY
Filing Date
2026-04-01
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Traditional methods for detecting nucleotide metabolism and genetic molecular information in water bodies are cumbersome, costly, and have low spatiotemporal resolution, making it difficult to achieve large-scale water environment monitoring and dynamic assessment.

Method used

A multiple linear regression model was constructed based on the optical properties of DOM. By measuring the optical parameters HIX, BIX, and aCDOM of the water sample, the optimal prediction model was selected by combining stepwise regression and the Akaike information criterion, and the intensity of nucleotide metabolism in the water body was estimated.

Benefits of technology

It enables non-invasive estimation of nucleotide metabolism and genetic molecular information in water bodies, simplifies operation, reduces costs, and improves the spatiotemporal resolution of detection, making it suitable for large-scale water environment monitoring and in-situ real-time monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347993A_ABST
    Figure CN122347993A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of water environment monitoring, and discloses a method for estimating the nucleotide metabolism function intensity of water bodies based on the optical characteristics of DOM (dissolved organic matter), which is specifically as follows: based on the optical characteristics of the dissolved organic matter (DOM), the optical parameters and the functional gene relative abundance data of the modeling samples are acquired, the gene abundance data is subjected to central logarithmic ratio transformation, a multiple linear regression model is constructed, and model verification and evaluation are completed, and finally, the optimal prediction model for rapidly estimating the relative abundance of the key functional genes of the water bodies is obtained. The application solves the defects of the traditional detection methods, such as complicated operation, high cost, low space-time resolution and the like, realizes rapid, non-invasive and in-situ estimation of the nucleotide metabolism and genetic molecular information functional genes of the water bodies, greatly shortens the detection period, and is suitable for large-scale water environment monitoring, water body ecological health evaluation, pollution early warning and carbon cycle process research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water environment monitoring technology, specifically a method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties. Background Technology

[0002] The metabolic function and genetic molecular information of aquatic microorganisms are among the core indicators reflecting the potential ecological health status of the aquatic environment, and are closely related to the self-purification capacity, nutrient cycling efficiency, and pollution response mechanisms of water bodies. Microbial communities in inland water bodies significantly enhance their adaptability to environmental changes through horizontal gene transfer driven by functional genes (such as transposase gene-K07495; ribonucleoside diphosphate reductase-K00526). These genes are closely related to key ecological processes: high abundance of transposase genes may promote the rapid acquisition of enzyme genes by microorganisms to degrade persistent organic matter, thereby enhancing the self-purification capacity of water bodies; ribonucleoside diphosphate reductase corresponds to the functional classification of specific genes or proteins.

[0003] Traditional methods for detecting nucleotide metabolism and genetic molecular information in water rely on metagenomic sequencing or quantitative PCR, which have several drawbacks: (1) cumbersome operation: requiring the collection of large amounts of water samples and the extraction of high-quality DNA / RNA, resulting in long experimental cycles (usually ≥2 weeks); (2) high cost: sequencing costs account for 30%-50% of the total project budget, limiting large-scale application; (3) low spatiotemporal resolution: making it difficult to capture dynamic changes in gene abundance in water (such as diurnal and seasonal fluctuations); (4) sensitive to environmental interference: optical characteristics such as water turbidity and pigment interference are not included in the model, leading to prediction bias. The above methods suffer from problems such as complex operation, long detection cycles, high costs, and difficulty in achieving in-situ real-time monitoring, which cannot meet the needs of large-scale water environment monitoring and dynamic assessment.

[0004] Dissolved organic matter (DOM) is one of the most active organic components in water bodies. Its optical properties (such as fluorescence spectra and ultraviolet absorption characteristics) can sensitively reflect the dynamic changes in the metabolic activities of microorganisms in water bodies. Therefore, developing a rapid, non-invasive method for estimating the gene intensity of nucleotide metabolism and genetic molecular information in water bodies based on the optical properties of DOM has important theoretical and practical application value for water body ecological monitoring and pollution control. Summary of the Invention

[0005] The purpose of this invention is to provide a method for estimating the intensity of nucleotide metabolic function in water bodies based on the optical properties of DOM, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties, comprising the following steps:

[0007] Step S1, Modeling Sample Data Acquisition: Surface water samples were collected from inland lakes and reservoirs and their inflow and outflow rivers. The water samples were filtered through a 0.45 μm filter membrane, and the optical parameters of dissolved organic matter (DOM) were measured. These optical parameters included the humification index (HIX), biogenicity index (BIX), and colored dissolved organic matter absorption coefficient (aCDOM). Separate water samples were filtered through a 0.22 μm filter membrane to enrich microorganisms. DNA was extracted and subjected to metagenomic sequencing. The sequencing reads were aligned to the KEGG database to obtain the relative abundance of functional genes. The gene abundance data were transformed using the central logarithmic ratio and used as the dependent variable for modeling. The optical parameters and the transformed gene abundance data were merged according to sample ID to form a dataset. Samples with missing values ​​were deleted to obtain valid samples.

[0008] Step S2, Construction of multiple linear regression model: The abundance of target genes after central log ratio transformation is used as the dependent variable Y, and the DOM optical parameters HIX, BIX, and aCDOM are used as candidate independent variables X. The optimal prediction model is selected by stepwise regression combined with the Akaike Information Criterion (AIC). The least squares method is used for model fitting.

[0009] Step S3, Model Validation and Evaluation: Randomly divide the effective samples into 70% training set and 30% test set. After building the model on the training set, calculate the coefficient of determination (R²) and root mean square error (RMSE) between the predicted and measured values ​​on the test set. Repeat the random division 10 times and take the average performance as the evaluation index of model stability.

[0010] Step S4, Model Application: Measure the DOM optical parameters HIX, BIX, and aCDOM of the water sample to be tested, and substitute them into the optimal prediction model obtained in step S2 to estimate the relative abundance of the target functional gene in the water sample to be tested.

[0011] Preferably, the formula for calculating the central logarithmic ratio transformation in step S1 is:

[0012]

[0013] Where, x i Let g(x) be the original abundance of the i-th gene in a sample, and g(x) be the geometric mean of the abundance of all target genes in the sample.

[0014] Preferred: The specific algorithm for selecting the optimal prediction model in step S2 using stepwise regression combined with the Akaike Information Criterion (AIC) is as follows: Initialization: Start with an empty model containing only the intercept term; Forward selection: Add the unselected independent variables to the model one by one, calculate the AIC value of the new model, and select the variable that causes the largest decrease in AIC to be introduced into the model; Backward elimination: After each new variable is introduced, check whether there are any variables in the model that cause the AIC to increase due to the introduction of that variable, and if so, remove them; Termination: Repeat the forward selection and backward elimination steps until adding any remaining variable cannot reduce the AIC, and removing any existing variable cannot reduce the AIC, and the model converges to obtain the optimal prediction model.

[0015] Preferably, in step S2, the model fit is evaluated by the coefficient of determination (R²), and the AIC value is used to measure the balance between model complexity and fit.

[0016] Preferably, in step S3, the random forest machine learning method is also used for parallel modeling. The random forest model uses the default parameter ntree=300. The importance of variables is evaluated by %IncMSE. The test set R² of the linear model and the random forest are compared to determine the linear relationship between DOM optical parameters and gene abundance and to verify the reliability of the linear model.

[0017] Preferably, the target functional gene in step S4 is the K07495 putative transposase gene and / or the K00526 ribonucleoside diphosphate reductase β-chain gene.

[0018] Preferably, the optimal prediction model for the K07495 hypothetical transposase gene is: K07495 Abundance =-13.41+0.79×HIX+6.22×BIX; The optimal prediction model for the K00526 ribonucleoside diphosphate reductase β chain gene is: K00526 Abundance =0.27 - 0.42 × HIX + 0.01 × a CDOM Where, Abundance is the relative abundance of genes after central logarithmic ratio transformation.

[0019] Compared with the prior art, the beneficial effects of this invention are as follows:

[0020] Based on the optical properties of DOM, this invention enables non-invasive estimation of functional genes related to nucleotide metabolism and genetic molecular information in water bodies. It eliminates the need to extract DNA / RNA from water samples, making the operation simple and significantly reducing the difficulty of detection.

[0021] This invention avoids expensive detection methods such as metagenomic sequencing, significantly reduces detection costs, and solves the problem that the high cost of traditional methods limits large-scale application, making it suitable for large-scale water environment monitoring;

[0022] This invention only requires measuring the DOM optical parameters of water samples to complete gene abundance estimation, significantly shortening the detection cycle from the traditional several weeks. It can capture the dynamic changes of gene abundance in water bodies during the day and night, and the seasons, improving the spatiotemporal resolution of detection and enabling in-situ real-time monitoring.

[0023] The multiple linear regression model constructed in this invention has an R² greater than 0.5 on the test set, which can effectively explain more than 50% of the gene abundance variation. The model has good fitting effect and high prediction accuracy, and can provide reliable gene abundance data support for water ecological health assessment, pollution early warning and carbon cycle process research. Attached Figure Description

[0024] Figure 1 This is a flowchart of the method of the present invention;

[0025] Figure 2 A scatter plot showing the measured and model-predicted abundance values ​​of the K00526 gene;

[0026] Figure 3 This is a partial regression plot of K07495 gene abundance versus BIX. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Example

[0029] Please see Figures 1-3 The illustration shows a method for estimating the intensity of nucleotide metabolic function in water bodies based on the optical properties of the DOM (Dissolved Organic Matter). The specific steps are as follows: Modeling and sample data acquisition: Taking inland lakes and reservoirs and their inflow and outflow rivers as the research objects, 49 surface water samples were collected. After filtering some of the water samples through a 0.45 μm filter membrane, the optical parameters of the DOM were measured using a fluorescence spectrometer: humification index (HIX), biogenic index (BIX), and colored dissolved organic matter absorption coefficient (aCDOM). Another part of the water samples was filtered through a 0.22 μm filter membrane to enrich microorganisms. Microbial DNA was extracted using a kit, and metagenomic sequencing was performed on the DNA. The reads obtained from the sequencing were aligned to the KEGG database to obtain the relative abundance of genes K07495 and K00526.

[0030] The original abundance of genes K07495 and K00526 was transformed using the central log ratio transformation formula:

[0031]

[0032] Where x i Let g(x) be the original gene abundance and g(x) be the geometric mean of the two gene abundances. The optical parameters and the transformed gene abundance data are merged according to the sample ID. Samples with missing values ​​are deleted to obtain valid modeling samples.

[0033] The multiple linear regression model was constructed with the transformed gene abundances of K07495 and K00526 as dependent variables Y, and HIX, BIX, and aCDOM as candidate independent variables X. Stepwise regression combined with the AIC criterion was used to screen the optimal model: starting with an empty model containing only the intercept term, forward selection was performed, adding independent variables one by one and calculating the AIC value, introducing the variable that caused the largest decrease in AIC; then backward elimination was performed, removing variables that caused the AIC to increase due to the introduction of new variables; forward selection and backward elimination were repeated until the model converged, obtaining the optimal prediction model, and the least squares method was used for model fitting.

[0034] Model validation and evaluation involved randomly dividing the valid samples into a 70% training set and a 30% test set. After building the model on the training set, the R² and RMSE of the predicted values ​​versus the measured values ​​were calculated on the test set. This random division operation was repeated 10 times, and the average R² and RMSE were used as the model stability indicators. Simultaneously, a random forest method was used for parallel modeling (ntree=300), and the importance of variables was evaluated using % IncMSE. The reliability of the linear model was verified by comparing the test set R² of the linear model and the random forest.

[0035] Gene abundance estimation of water samples: Surface water samples were collected from an inland lake and filtered through a 0.45 μm filter membrane. HIX, BIX, and aCDOM values ​​were measured, and these values ​​were substituted into the optimal prediction models K07495 and K00526, respectively: K07495 Abundance =-13.41+0.79×HIX+6.22×BIX; K00526 Abundance =0.27-0.42×HIX+0.01×aCDOM; the relative abundance of genes after central logarithmic ratio transformation is calculated, and the rapid estimation of genes K07495 and K00526 in the water sample is completed.

[0036] Figure 2 Note: The horizontal axis represents the measured CLR abundance, and the vertical axis represents the predicted CLR abundance. The scatter points are evenly distributed around the 1:1 line, indicating that the model has a good fit and prediction effect on the abundance of the K00526 gene. Figure 3 Explanation: The horizontal axis represents the residual of BIX to HIX+FI, and the vertical axis represents the CLR abundance of the K07495 gene. This visually shows a significant positive correlation between BIX and the K07495 transposase gene.

[0037] The method in this embodiment does not require metagenomic sequencing. Gene abundance estimation can be completed simply by measuring the conventional optical parameters of water samples. The detection cycle is shortened from more than two weeks to a few hours. It is simple to operate, low in cost, and the model has high prediction accuracy. It can effectively reflect the abundance status of microbial functional genes in water bodies, providing an efficient and reliable technical means for water environment ecological monitoring.

[0038] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0039] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties, characterized in that, Includes the following steps: Step S1, Modeling Sample Data Acquisition: Surface water samples were collected from inland lakes and reservoirs and their inflow and outflow rivers. The water samples were filtered through a 0.45 μm filter membrane, and the optical parameters of dissolved organic matter (DOM) were measured. These optical parameters included the humification index (HIX), biogenicity index (BIX), and colored dissolved organic matter absorption coefficient (aCDOM). Separate water samples were filtered through a 0.22 μm filter membrane to enrich microorganisms. DNA was extracted and subjected to metagenomic sequencing. The sequencing reads were aligned to the KEGG database to obtain the relative abundance of functional genes. The gene abundance data were transformed using the central logarithmic ratio and used as the dependent variable for modeling. The optical parameters and the transformed gene abundance data were merged according to sample ID to form a dataset. Samples with missing values ​​were deleted to obtain valid samples. Step S2, Construction of multiple linear regression model: The abundance of target genes after central log ratio transformation is used as the dependent variable Y, and the DOM optical parameters HIX, BIX, and aCDOM are used as candidate independent variables X. The optimal prediction model is selected by stepwise regression combined with the Akaike Information Criterion (AIC). The least squares method is used for model fitting. Step S3, Model Validation and Evaluation: Randomly divide the effective samples into 70% training set and 30% test set. After building the model on the training set, calculate the coefficient of determination (R²) and root mean square error (RMSE) between the predicted and measured values ​​on the test set. Repeat the random division 10 times and take the average performance as the evaluation index of model stability. Step S4, Model Application: Measure the DOM optical parameters HIX, BIX, and aCDOM of the water sample to be tested, and substitute them into the optimal prediction model obtained in step S2 to estimate the relative abundance of the target functional gene in the water sample to be tested.

2. The method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties according to claim 1, characterized in that: The formula for calculating the central logarithmic ratio transformation in step S1 is: , Where, x i Let g(x) be the original abundance of the i-th gene in a sample, and g(x) be the geometric mean of the abundance of all target genes in the sample.

3. The method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties according to claim 2, characterized in that: The specific algorithm for selecting the optimal prediction model using stepwise regression combined with the Akaike Information Criterion (AIC) in step S2 is as follows: Initialization: Start with an empty model containing only the intercept term; Forward selection: Add the unselected independent variables to the model one by one, calculate the AIC value of the new model, and select the variable that causes the largest decrease in AIC to be introduced into the model; Backward elimination: After each new variable is introduced, check whether there are any variables in the model that cause the AIC to increase due to the introduction of that variable, and if so, remove them. Termination: Repeat the forward selection and backward elimination steps until adding any remaining variables can no longer reduce AIC, and removing any existing variables can no longer reduce AIC. The model converges to obtain the optimal prediction model.

4. The method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties according to claim 3, characterized in that: In step S2, the model fit is evaluated using the coefficient of determination (R²), and the AIC value is used to measure the balance between model complexity and fit.

5. The method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties according to claim 4, characterized in that: In step S3, the random forest machine learning method is also used for parallel modeling. The random forest model uses the default parameter ntree=300. The importance of variables is evaluated by %IncMSE. The test set R² of the linear model and the random forest is compared to determine the linear relationship between DOM optical parameters and gene abundance and to verify the reliability of the linear model.

6. The method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties according to claim 5, characterized in that: The target functional genes in step S4 are the K07495 putative transposase gene and / or the K00526 ribonucleoside diphosphate reductase β-chain gene.

7. The method for estimating the intensity of nucleotide metabolic function in water bodies based on DOM optical properties according to claim 6, characterized in that: The optimal prediction model for the assumed transposase gene of K07495 is: K07495 Abundance =-13.41+0.79×HIX+6.22×BIX; The optimal prediction model for the K00526 ribonucleoside diphosphate reductase β chain gene is: K00526 Abundance =0.27 - 0.42 × HIX + 0.01 × a CDOM Where, Abundance is the relative abundance of genes after central logarithmic ratio transformation.