A model and construction method for calculating age based on fecal DNA

By constructing a model for calculating age based on mtDNA heteroplasmy sites in fecal DNA, the problems of difficulty and high cost in obtaining materials are solved, and low-cost, damage-free, high-precision age prediction is achieved, which is suitable for age estimation of large or wild animals.

CN119446288BActive Publication Date: 2025-09-30NORTHEAST FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411467387.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-09-30
Estimated Expiration
2044-10-21

AI Technical Summary

Technical Problem

In existing technologies, methods for estimating the age of wild animals are limited by the difficulty and high cost of obtaining materials. Traditional methods cannot accurately determine the age of large or wild animals, especially the high cost of methylation site detection based on fecal DNA.

Method used

A model for calculating age based on mtDNA heteroplasmy sites in fecal DNA was constructed. The elastic regression model, support vector machine model, XGBoost regression model or random forest model was used to extract genomic DNA from feces, construct a library and sequence it, screen for age-related mtDNA heteroplasmy sites, and use machine learning methods to predict age.

Benefits of technology

It achieves low-cost and non-destructive acquisition of materials, improves the precision and accuracy of age prediction, and is suitable for age estimation of large or wild animals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005093763440000081
    Figure BDA0005093763440000081
  • Figure BDA0005093763440000091
    Figure BDA0005093763440000091
Patent Text Reader

Abstract

A model and construction method for calculating age based on fecal DNA, belonging to the field of bioinformatics technology. The present invention solves the problem of the inability to obtain the required materials in the prior art by using feces as the material, and the present invention discloses for the first time a method for obtaining mtDNA heteroplasmy sites, which can more accurately capture the target mtDNA heteroplasmy sites. Based on the target mtDNA heteroplasmy sites and machine learning, a set of methodological processes for screening detection sites related to age was constructed, and a model for calculating age based on mtDNA heteroplasmy sites in fecal DNA was constructed. The model is mainly used to calculate the age of large animals or wild animals. The model solves the problem of high cost in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to an age estimation method based on mtDNA heterogeneity. Background Art

[0002] Age is one of the fundamental parameters in population research and a crucial factor in animal ecology studies. Age plays an important role in determining life history characteristics of animal populations, such as birth and mortality rates. Age typically consists of both chronological age and biological age. Chronological age refers to the total time elapsed since birth, while biological age is inferred based on anatomical and physiological developmental status, reflecting the actual state of an organism's tissue structure and physiological functions. The degree of agreement between biological age and chronological age reflects the interaction and mechanisms between an organism and its environment, and plays a key role in population monitoring, development trend prediction, life history research, and conservation and management decisions. Therefore, accurately determining species age is extremely important.

[0003] However, there are no age records for animals in the wild. Traditionally, the age of wild animals has been estimated mainly based on appearance, morphology, anatomical indicators (growth rings and wear of hard tissues, bone morphology, size of food cuts, growth curves), physiological and biochemical indicators (sex hormone levels, chromosome telomere length, DNA methylation levels), and omics indicators (transcriptome, proteome, metabolome, methylome, and metagenome). The appearance morphology observation method is often limited to long-lived animals, whose body shape does not change much in adulthood, making it impossible to accurately determine their age. However, the other three methods are limited by the selection of species materials and usually rely on animal tissues, organs, tissues, or blood. This type of material is difficult to collect and is subject to various factors. It can only roughly estimate the age of animals (juvenile, subadult, adult, and old). Therefore, traditional methods have obvious shortcomings in counting wild animals or large animals.

[0004] Generally, feces is the most easily obtained material from wild animals. The continuous division and differentiation of stem cells and the shedding of intestinal epithelial cells build a highly dynamic intestinal epithelium. Intestinal epithelial cells are one of the cells with the highest cell turnover rate (they are completely renewed approximately every 3 to 5 days). After completing their mission, these daughter cells fall off into the intestinal cavity and are excreted from the body with feces. The DNA obtained from feces mainly comes from the shed intestinal epithelial cells, and the huge number of intestinal epithelial cells in feces (the number of cells shed by mice every 24 hours can reach 2×10 8 While humans can reach 10 11 The cost of detecting methylation sites in fecal DNA is high in existing methods, so a lower-cost method is needed.

[0005] Mitochondrial DNA (mtDNA) is the genetic material found in mitochondria, which generate energy (ATP) for cells. It is a specialized form of deoxyribonucleic acid found within mitochondria. Mitochondria are organelles that provide cells with energy (ATP). Each mitochondria typically contains multiple DNA molecules. mtDNA in feces primarily comes from shed intestinal epithelial cells. The intestinal epithelium not only has a vast surface area (up to 111 times the body's surface area in humans and 119 times in mice) but also has an extremely rapid cell turnover rate (completely renewed approximately every 3 to 5 days). For example, in humans, each intestinal crypt contains 4 to 6 stem cells, which must produce 300 daughter cells daily to maintain the integrity of the intestinal epithelium. This massive cell proliferation is accompanied by a massive proliferation of mitochondria and their genomes, resulting in the accumulation of mutations in the daughter epithelial cells. After completing their mission, these daughter cells shed into the intestinal lumen and are excreted in feces. Therefore, utilizing the vast number of intestinal epithelial cells in feces (mice shed up to 2×10 cells per 24 hours), it is possible to investigate the effects of mtDNA on the intestinal epithelium. 8 While humans can reach 10 11 ), which can detect the heteroplasmy of mtDNA.

[0006] Changes and accumulation of mtDNA heteroplasmy are inextricably linked to individual age. Previous studies in humans have found that mtDNA heteroplasmy decreases with age, decreasing by an average of 0.4 copies per year. Furthermore, heteroplasmic variants are often more numerous and more frequent in older individuals. On average, mtDNA heteroplasmy in individuals over 70 years old is 58.5% higher than in those under 40 years old. This association between heteroplasmy and increasing age is not limited to humans but has also been observed in a variety of other mammals. However, these studies have not further explored the relationship between mtDNA heteroplasmy and age and used it to infer age. Therefore, to more conveniently, accurately, and cost-effectively calculate the age of large animals (including wild animals), it is imperative to develop an age estimation method based on mtDNA heteroplasmy. Summary of the Invention

[0007] The present invention solves the problem of the difficulty in obtaining the required materials and the high cost in the prior art by constructing a computational model based on the correlation between mtDNA heteroplasmy sites and age in DNA in feces.

[0008] A method for constructing a model for calculating age based on fecal DNA, wherein the method is to construct a model for calculating age based on mtDNA heteroplasmy sites in fecal DNA.

[0009] The model is an elastic regression model, a support vector machine model, an XGBoost regression model or a random forest model.

[0010] The specific steps of the method are as follows:

[0011] Step 1: Extract genomic DNA from feces;

[0012] Step 2: Construct a library based on genomic DNA and sequence it. Then use trimmomatic to process the sequence in the library to obtain the processed library sequence.

[0013] Step 3: Obtain a complete report of mtDNA heterogeneity for each sample through the processed library sequence;

[0014] Step 4: Data processing of the complete report of mtDNA heteroplasmy for each sample;

[0015] Step 5: Screening target heteroplasmic sites;

[0016] Step 6: Construct a machine learning model of age and target heterogeneous sites based on the xgboost package, randomForest package, e1071 package, and glmnet package.

[0017] The nuclear CT value of the DNA in step 1 is less than or equal to 27.

[0018] The bacterial CT value of the DNA in step 1 is greater than 7.

[0019] The sequencing depth in step 2 is greater than or equal to 5000X.

[0020] The condition for screening target heteroplasmic sites in step 5 is that the correlation between age and target mtDNA heteroplasmic sites is limited to be greater than or equal to 0.5, or less than or equal to -0.5.

[0021] The data processing described in step 4 is done by using the dplyr package, setting the outliers to null values ​​according to the 3 sigma principle, and using the KNN model to interpolate the missing values.

[0022] The specific steps for obtaining a complete report of mtDNA heteroplasmy for each sample are:

[0023] Step 31: Use bwa to align the processed library data with the mtgenome genome, and then process it through samtools to obtain the mitochondrial related reading file mt1.bam;

[0024] Step 32: Convert mt1.bam to mt1.fq sequence and use trimmomatic to process the sequence. Then perform a second alignment with the Numt sequence to obtain the secondary alignment file numt-sorted.bam, which is then deduplicated using samtools software.

[0025] Step 33: Convert the deduplicated numt-sorted.bam to numt.fq sequence and perform trimmomatic processing on it, then perform a third alignment with the mitochondrial genome to obtain numttomt.bam;

[0026] Step 3 and 4: Integrate the file mt1.bam in step 1 and the file numttomt.bam in step 3, and remove the file numt-sorted.bam in step 2 to obtain all mitochondrial readings;

[0027] Steps 3 and 5: Use gatk to extract mtDNA heteroplasmy information from all mitochondrial reads, and use mutserve v2.0.1 to extract mtDNA heteroplasmy information from all mitochondrial reads. Integrate the heteroplasmy information obtained from the two to obtain the mtDNA heteroplasmy information report for each sample.

[0028] Step 36: Use bcftools to establish a consensus sequence, construct a circular sequence through Python processing, then build a library based on the circular sequence and sequence it. Then use trimmomatic to process the sequence in the library to obtain the processed library sequence. Repeat steps 1 to 5 to obtain a new heteroplasmy report for each sample. Finally, integrate the new heteroplasmy information of each sample with the mtDNA heteroplasmy information report of each sample obtained in step 5 to obtain a complete report of mtDNA heteroplasmy for each sample.

[0029] A model for calculating age based on mtDNA heteroplasmy sites in fecal DNA, wherein the model is constructed by the above method.

[0030] A model for calculating age based on mtDNA heteroplasmy sites in fecal DNA and its application in calculating the age of large animals or wild animals.

[0031] Beneficial effects

[0032] The invention constructs a model for calculating age based on mtDNA heteroplasmy sites in fecal DNA. Compared with the prior art, the required materials will not cause damage to animals and are relatively easy to obtain.

[0033] The present invention discloses for the first time a method for obtaining mtDNA heteroplasmic sites, which can more accurately capture age-related mtDNA heteroplasmic sites.

[0034] The present invention constructs a model for calculating age based on mtDNA heteroplasmy sites in fecal DNA, which has higher prediction accuracy than the existing technology.

[0035] This invention discloses for the first time a methodological process for screening age-related detection sites based on fecal mtDNA heterogeneity and machine learning. The detection cost is low and the accuracy is high. It can be applied to the screening of age-related heterogeneous sites in more animals, and also provides a reference for the construction of age identification methods for other animals. DETAILED DESCRIPTION

[0036] The present invention will be further described below with reference to embodiments, but the present invention is not limited to the following embodiments.

[0037] Implementation Condition 1: The nuclear CT value of the fecal DNA selected in the present invention does not exceed 27, and the bacterial CT value is greater than 7;

[0038] Implementation condition 2: Screening for target mtDNA heteroplasmic sites must be performed according to the relevant steps described in the detailed method;

[0039] Implementation condition three: The correlation between the screening age and the target mtDNA heteroplasmic site in the present invention is limited to be greater than or equal to 0.5 or less than or equal to -0.5.

[0040] Implementation Condition 4: The present invention constructs a computational model for the correlation between target mtDNA heteroplasmy sites and age based on the elastic regression model, support vector machine model, random forest model, and eXtreme Gradient Boosting (XGBoost) regression model, and determines the optimal parameters and evaluation indicators for each model through leave-one-out cross-validation.

[0041] Implementation condition five: The average mitochondrial sequencing depth of the samples selected in the present invention must be greater than or equal to 5000X.

[0042] The present invention is further described in detail below with reference to specific embodiments.

[0043] 1. Sample DNA extraction and library construction.

[0044] 1. Sample preparation.

[0045] The samples were 33 pieces of Siberian tiger feces from the Siberian Tiger Park and Hengdaohezi in Heilongjiang Province. They were numbered and recorded with the corresponding date of birth, sampling date, and gender, and stored in a -80℃ refrigerator.

[0046] 2. Sample DNA extraction, library construction and WGS sequencing.

[0047] Genomic DNA was extracted from 33 stool samples using the method established by our research group (described in patent document CN113186185A). The extracted DNA was annotated and stored at -20°C until ready for use. Libraries were constructed from qualified DNA from each of the 33 samples (DNA with a nuclear CT value no greater than 27 and a bacterial CT value greater than 7). The constructed libraries were sequenced at a depth of 5000x or greater. Library construction and sequencing were commissioned to Shenzhen BGI. The cost of testing for mtDNA heteroplasmy per sample was 648 yuan (compared to the traditional method of testing methylation sites per sample, which costs 4224 yuan).

[0048] 2. Extract heterogeneous information.

[0049] 1. Find the Numts sequence.

[0050] The nuclear genome and mitochondrial genome sequences were obtained from NCBI, and the NUMTFinder V0.5.4 tool was used to obtain the Numts sequence of the sample through the nuclear genome and mitochondrial genome sequences.

[0051] 2. Library data processing.

[0052] Trimmomatic (version 0.36) was used to process the sequenced data, including adapter trimming, quality filtering, and low-quality base removal, to improve the quality and reliability of the sequence data. The sequencing data were processed using the following criteria: adapter removal, with a maximum of two mismatches allowed and a threshold of 15 for adapter removal; quality filtering using a sliding window of 5 bases, with pruning performed when the average quality value fell below 20; sequences ≥35 bases in length and with an average quality of ≥20 were retained; and no front-end or back-end trimming was performed. The entire process was processed in parallel using two threads, with unpaired sequences discarded and paired sequences saved to a designated file.

[0053] 3. Three comparisons.

[0054] The purpose of the three-way alignment is to ensure that the Numts readings are removed while obtaining the mitochondrial readings, and to recover the mitochondrial readings that are misidentified as Numts.

[0055] 3.1. Obtaining preliminary mitochondrial readings.

[0056] Use bwa (version: 0.7.17-r1188) to align the processed library data with the mtgenome genome (from NCBI), and then process the aligned data through samtools (Version: 1.15.1) to obtain the mitochondrial related reading file mt1.bam.

[0057] 3.2. Obtain Numts readings in mitochondria.

[0058] The file mt1.bam (mitochondrial-related reading file) obtained in 3.1 was converted into the mt1.fq sequence. The mt1.fq sequence was then trimmed by trimmomatic to remove low-quality sequences (trimming was performed when the average quality value was less than 20). It was then aligned with the Numt sequence for a second time to obtain the secondary alignment file nunt-sorted.bam. This file is the Numts reads included in the first alignment, which were sorted and deduplicated by the samtools software.

[0059] 3.3. Recovering mitochondrial reads misidentified as Numts.

[0060] The deduplicated nunt-sorted.bam (i.e., the Numts reads obtained from the second alignment) was converted into the nunt.fq sequence. The nunt.fq sequence was trimmed by trimmomatic to remove low-quality sequences (trimming was performed when the average quality value was less than 20) and aligned with the mitochondrial genome for the third time to obtain numttomt.bam, and the mitochondrial reads that were mistakenly identified as Numts were recovered.

[0061] 3.4. Integrate bam files.

[0062] Integrate the mitochondrial readings mt1.bam obtained in step 3.1 with the recovered mitochondrial readings numttomt.bam obtained in step 3.3, and remove the Numts readings numt-sorted.bam in the second step. At this point, all mitochondrial readings can be obtained.

[0063] 4. Extract heterogeneous information.

[0064] The Mutect2 module in gatk (Version: 4.3.0.0) was used to extract mtDNA heteroplasmy information from all mitochondrial reads obtained in 3.4. At the same time, mutserve v2.0.1 was used to extract mtDNA heteroplasmy information from all mitochondrial reads obtained in 3.4. The heteroplasmy information obtained from the two was integrated to obtain the mtDNA heteroplasmy information report for each sample.

[0065] 5. Create a consensus sequence for each sample.

[0066] Use bcftools to establish a consensus sequence based on the heteroplasmy information of each sample. After Python processing, a circular sequence is constructed. This circular sequence can be used as a reference genome. Repeat steps 2 and 3 to obtain a new heteroplasmy report for each sample. Finally, the new heteroplasmy information for each sample is reintegrated with the heteroplasmy information generated in step 4 to obtain a complete report on mtDNA heteroplasmy for each sample.

[0067] 3. Data processing and screening of age-related mtDNA heteroplasmy.

[0068] 1. Data processing.

[0069] The mtDNA heteroplasmy information of 33 samples was integrated using the dplyr package (R language). Outliers were detected according to the 3 sigma principle and set as null values. The missing values ​​were interpolated using the KNN model.

[0070] 2. Screening for age-related heterogeneity sites.

[0071] Based on Spearman correlation analysis, the conditions for screening the correlation between heterogeneous sites and age were limited to less than or equal to -0.5 or greater than or equal to 0.5.

[0072] 4. Model, evaluate and predict heterogeneous sites and age.

[0073] Four machine learning models were constructed based on the xgboost package, randomForest package, e1071 package, and glmnet package to analyze age and target heterogeneity sites, namely elastic regression model, support vector machine model, XGBoost regression model, and random forest model. The leave-one-out cross-validation method was used to evaluate each model, and the evaluation index of each model (R 2 , MSE, RSME) and the optimal parameters of each model. At the same time, the correlation between the predicted age and the actual age of the four different models needs to be determined to screen the optimal model. 2 It represents correlation, MAE represents the mean of absolute errors (hereinafter referred to as error), and RMSE represents the root mean square error. The evaluation criteria are that the higher the correlation, the smaller the error, and the higher the accuracy of the model.

[0074] 5. Results

[0075] Spearman correlation analysis identified 15 target heteroplasmic sites with correlations with age greater than or equal to 0.5 or less than or equal to -0.5. To further validate the accuracy of age estimation, age prediction was performed on these 15 sites using a leave-one-out cross-validation method combined with four machine learning models: elastic regression, support vector machine, XGBoost regression, and random forest. The results are shown in Table 1. The results indicate that all four models are suitable for constructing a model for calculating age based on mtDNA heteroplasmic sites in fecal DNA.

[0076] When the elastic regression model is used to calculate age, in the training set, the correlation between the calculated age and the actual age can reach 0.859, with an error of only 6.18 months. In the test set, the correlation between the predicted age and the actual age can reach 0.878, with an error of only 13.87 months. When the support vector machine regression model is selected as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.995, with an error of only 1.93 months. In the test set, the correlation between the predicted age and the actual age can reach 0.805, with an error of only 11.01 months. When the random forest regression model was selected as the analysis model, the correlation between predicted and actual age reached 0.956 in the training set, with an error of only 5.41 months; in the test set, the correlation reached 0.811, with an error of only 9.93 months. When the XGBoost regression model was selected as the analysis model, the correlation between predicted and actual age reached 0.987 in the training set, with an error of only 1.75 months; in the test set, the correlation reached 0.905, with an error of only 8.91 months. Because the XGBoost regression model had the highest correlation and the lowest error, it was found to be a superior model for calculating age based on fecal mtDNA heteroplasmy.

[0077] Table 1 Regression evaluation indicators of different machine learning models

[0078]

[0079]

Claims

1. A method for constructing a model for calculating age based on fecal DNA, characterized in that: The method is to construct a model for calculating age based on mtDNA heteroplasmy sites in fecal DNA; The specific steps of the method are as follows: Step 1: Extract genomic DNA from feces; Step 2: construct a library based on genomic DNA and sequence it, then use trimmomatic to process the sequences in the library; Step 3: Obtain a complete report of mtDNA heterogeneity for each sample through the processed library sequence; Step 4: Data processing of the complete report of mtDNA heteroplasmy for each sample; Step 5: Screening target heteroplasmic sites; Step 6: Build a machine learning model for age and target heterogeneous sites based on the xgboost, randomForest, e1071, and glmnet packages; The condition for screening the target heteroplasmic sites in step 5 is that the correlation between age and the target mtDNA heteroplasmic sites is greater than or equal to 0.5, or less than or equal to -0.5; The method for obtaining a complete report of mtDNA heteroplasmy for each sample described in step 3 is as follows: Step 31: Use bwa to align the processed library data with the mtgenome genome, and then process it through samtools to obtain the mitochondrial related reading file mt1.bam; Step 32: After converting mt1.bam to mt1.fq sequence, trimmomatic was used to remove low-quality sequences. The trimming was performed when the average quality value was less than 20. Then, a second alignment was performed with the Numt sequence to obtain the secondary alignment file numt-sorted.bam, which was then deduplicated using samtools software; Step 33: Convert the deduplicated nunt-sorted.bam to the numt.fq sequence, perform trimmomatic trimming to remove low-quality sequences, and perform trimming when the average quality value is lower than 20. Then, perform a third alignment with the mitochondrial genome to obtain numttomt.bam. Step 34: Integrate the file mt1.bam in step 31 and the file numttomt.bam in step 33, and remove the file numt-sorted.bam in step 32 to obtain all mitochondrial readings; Step 35: Use the Mutect2 module in gatk to extract mtDNA heteroplasmy information from all mitochondrial reads. Simultaneously, use mutserve v2.0.1 to extract mtDNA heteroplasmy information from all mitochondrial reads. Integrate the heteroplasmy information obtained from the two to obtain the mtDNA heteroplasmy information report for each sample. Step 36: Use bcftools and Python to construct circular sequences, then build and sequence libraries based on the circular sequences. Then use trimmomatic to process the sequences in the libraries. Repeat steps 31 to 35 to obtain a new heteroplasmy report for each sample. Finally, integrate it with the mtDNA heteroplasmy information report for each sample obtained in step 35 to obtain a complete report on mtDNA heteroplasmy for each sample.

2. The method according to claim 1, characterized in that The model is an elastic regression model, a support vector machine model, an XGBoost regression model or a random forest model.

3. The method according to claim 1, characterized in that The nuclear CT value of the genomic DNA in step 1 is less than or equal to 27.

4. The method according to claim 1, wherein The bacterial CT value of the genomic DNA in step 1 is greater than 7.

5. The method according to claim 1, wherein The sequencing depth in step 2 is greater than or equal to 5000X.

6. The method according to claim 1, characterized in that The data processing described in step 4 is done by using the dplyr package, setting the outliers to null values ​​according to the 3 sigma principle, and using the KNN model to interpolate the missing values.

7. A model for calculating age based on mtDNA heteroplasmy sites in fecal DNA, characterized in that: The model is constructed by the method according to any one of claims 1 to 6.