Model for calculating age based on excrement mtDNA heterogeneity index and construction method
Patent Information
- Application Number
- CN202510034870.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the estimation of the age of wild animals depends on the tissue of difficult to collect animal body, and the traditional methods are costly and have low accuracy, making it difficult to accurately calculate the age of large animals.
By constructing a calculation model based on the age-related mtDNA heterogeneity index of DNA in feces, aging calculation is performed using quantitative information on the inconsistent read ratio of mtDNA heterogeneity sites, base mutation diversity or inconsistent read comparison rates, and aged calculation is performed in combination with machine learning models (such as elastic regression, support vector machine, XGBoost, random forest).
An animal age calculation without causing damage to animals and is easy to obtain samples is achieved, with high prediction accuracy and low detection cost, and is suitable for age identification of more animals.
Smart Images

Figure BDA0005235262340000081 
Figure BDA0005235262340000091 
Figure BDA0005235262340000101
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of bioinformatics, and in particular relates to an age calculation method based on mtDNA heterogeneity indicators. Background Art
[0002] Age is one of the basic parameters of a population and has important indicative significance for life history characteristics such as birth rate and mortality rate of animal populations. Age includes chronological age and biological age. Chronological age refers to the sum of time experienced since birth, while biological age is the age inferred based on the anatomical and physiological development status, which reflects the actual state of the body's tissue structure and physiological function. The consistency between biological age and chronological age reflects the interaction relationship and interaction mechanism between the body and the environment, and also plays a key role in population monitoring, development trend prediction, life history research, and protection and management decision-making. Therefore, it is particularly important to accurately determine the age of species. However, there is no age record for animals in the wild. Traditionally, the age of wild animals is estimated mainly from the aspects of appearance, anatomical indicators, physiological indicators, and omics indicators, but the materials are often limited. These methods usually rely on animal tissues, organs, tissues or blood. Such materials are difficult to collect and are restricted by many factors. They can only roughly estimate the age of animals (juvenile, sub-adult, adult and old). Therefore, traditional methods have obvious shortcomings in calculating wild animals or large animals.
[0003] Mitochondrial DNA (mtDNA) is the genetic material in mitochondria, which is a special form of deoxyribonucleic acid found in the mitochondria of cells. Mitochondria are organelles that provide energy (ATP) to cells. mtDNA has the property of a high mutation rate. Some individuals have two or more types of mtDNA molecules at the same time. This phenomenon is called mtDNA heteroplasmy. The mtDNA in feces mainly comes from desquamated intestinal epithelial cells. The huge surface area of the intestinal epithelium (up to 111 times (human) to 119 times (mice) the surface area of the body) and the extremely fast renewal rate (completely renewed every 3 to 5 days) make it possible to obtain sufficient mtDNA heteroplasmy information in a short time. The mtDNA heteroplasmy information is inextricably linked to the age of the individual. However, there is currently no further exploration of the relationship between mtDNA heteroplasmy and age, nor any report on using it to calculate age.
[0004] Currently, the epigenetic clock is considered to be the most promising molecular method for estimating biological age. By establishing a model of the DNA methylation level of biological tissues and the age of the organism, the age of the animal can be predicted. However, the detection cost of DNA methylation sites is high, so a low-cost detection method is urgently needed. Summary of the invention
[0005] The present invention solves the problem in the prior art that the required materials are difficult to obtain and the cost is high by constructing a calculation model based on the mtDNA heterogeneity index and the correlation with age in the feces.
[0006] A method for constructing a calculation age model, the method is to construct a calculation age model based on mtDNA heteroplasmy indicators extracted from feces, the heteroplasmy indicators include the quantification of inconsistent reading ratios of mtDNA heteroplasmic sites, base mutation diversity of mtDNA heteroplasmic sites or inconsistent reading pair ratios of mtDNA heteroplasmic sites.
[0007] Preferably, the model is an elastic regression model, a support vector machine model, an XGBoost regression model or a random forest model.
[0008] Preferably, the steps of the method are as follows:
[0009] Step 1: Extract genomic DNA from feces;
[0010] Step 2: construct a library based on genomic DNA, sequence and detect mtDNA heterogeneity, and then use trimmomatic to process the sequences in the library;
[0011] Step 3: Align the processed library sequences three times to obtain all mitochondrial reads for each sample;
[0012] Step 4: Process all mitochondrial reads of each sample through Python to obtain mtDNA heteroplasmy index information of each sample;
[0013] Step 5: Process the mtDNA heteroplasmy index information and screen the target sites through the dplyr package;
[0014] Step 6: Construct a machine learning model of age and target site based on the xgboost package, randomForest package, e1071 package and glmnet package.
[0015] Preferably, the sequencing in step 2 is WGS sequencing, and the sequencing depth is greater than or equal to 5000X.
[0016] Preferably, the trimmomatic processing of sequences in the library in step 2 includes removing adapter sequences, quality filtering, and removing low-quality sequences.
[0017] Preferably, the three-way comparison method in step 3 is:
[0018] Step 31: Use bwa to align the processed library sequence with the mtgenome genome, and then process it through samtools to obtain the mitochondrial related reading file mt1.bam;
[0019] Step 32: Convert mt1.bam to mt1.fq sequence and use trimmomatic to process the sequence, and then perform a second alignment with the Numt sequence to obtain the secondary alignment file nunt-sorted.bam, which is then sorted and deduplicated using samtools software;
[0020] Step 33: Convert the deduplicated nunt-sorted.bam to the nunt.fq sequence, remove the low-quality sequences using trimmomatic, and then align it with the mitochondrial genome for the third time to obtain numttomt.bam;
[0021] Step 34: Integrate the file mt1.bam described in step 31 and the numttomt.bam described in step 33, and remove the numt-sorted.bam described in step 32 to obtain all mitochondrial readings.
[0022] Preferably, the mtDNA heteroplasmy index information in step 4 includes the quantification of the inconsistent read ratio of the mtDNA heteroplasmy site, the base mutation diversity of the mtDNA heteroplasmy site, or the inconsistent read pair ratio of the mtDNA heteroplasmy site.
[0023] Preferably, in step 5, the dplyr package processes the mtDNA heteroplasmy index information by respectively integrating the same mtDNA heteroplasmy index information, setting the outliers to null values according to the 3sigma principle, and using the KNN model to interpolate the missing values.
[0024] Preferably, the condition for screening the target site in step five is that the correlation between age and the target site in the mtDNA heterogeneity index is greater than or equal to 0.3, or less than or equal to -0.3.
[0025] A model for calculating age based on mtDNA heteroplasmic sites in fecal DNA, wherein the model is prepared by the above method.
[0026] Beneficial Effects
[0027] The invention constructs a model for calculating age based on mtDNA heterogeneity indicators (MHS-PDR; MHS-BMP; MHR-Entropy; MHS-PDRP; MHS-qPDRP). Compared with the prior art, the required materials will not cause damage to animals and are relatively easy to obtain.
[0028] The invention constructs a model for calculating age based on mtDNA heterogeneity index, which has higher prediction accuracy than the prior art.
[0029] The present invention discloses for the first time a methodological process for calculating age established by using mtDNA heteroplasmy indicators and machine learning. The detection cost is low and the accuracy is high, which can provide a reference for age identification methods of more animals. DETAILED DESCRIPTION
[0030] The present invention is further described below in conjunction with implementation examples, but the present invention is not limited to the following implementation examples.
[0031] Implementation condition 1: The nuclear CT value of the feces DNA selected in the present invention does not exceed 27, and the bacterial CT value is greater than 7;
[0032] Implementation condition 2: Screening of relevant sites of mtDNA heteroplasmy indicators should be carried out according to the relevant steps described in the embodiment;
[0033] Implementation condition three: the correlation between the screening age and the target site in the mtDNA heteroplasmy index of the present invention is greater than or equal to 0.3 or less than or equal to -0.3;
[0034] Implementation condition 4: The present invention constructs a computational model of the correlation between target sites and age in mtDNA heteroplasmy indicators based on elastic regression model, support vector machine regression model, random forest regression model and eXtreme Gradient Boosting (XGBoost) regression model, and determines the optimal parameters and evaluation indicators of each model through leave-one-out cross validation;
[0035] Implementation condition five: The average sequencing depth of the samples in the present invention must be greater than or equal to 5000X.
[0036] The present invention is further described in detail below in conjunction with specific implementation modes.
[0037] Example 1.
[0038] 1. Sample DNA extraction and library construction.
[0039] 1. Sample preparation.
[0040] The samples were 55 pieces of Siberian tiger feces from the Siberian Tiger Park and Hengdaohezi in Heilongjiang Province. They were numbered and recorded with the corresponding date of birth, sampling date, and gender, and stored in a -80℃ refrigerator.
[0041] 2. Sample DNA extraction, library construction and WGS sequencing.
[0042] The genomic DNA was extracted from 55 stool samples according to the method established by our research group (CN113186185A), and the extracted DNA was annotated and stored at -20°C for future use. The qualified DNA in the 55 samples (DNA nuclear CT value not exceeding 27, bacterial CT value greater than 7) was separately constructed, and the constructed library was sequenced by WGS, and the sequencing depth was required to be greater than or equal to 5000X. The library construction and sequencing was entrusted to Shenzhen BGI, and mtDNA heterogeneity was tested. (The cost of detecting mtDNA heterogeneity for each sample library is 648 yuan, and the cost of detecting methylation sites for each sample library in the traditional method is 4224 yuan).
[0043] 2. Extract heterogeneous information.
[0044] 1. Find the Numts sequence.
[0045] The nuclear genome and mitochondrial genome sequences of the Siberian tiger were obtained from NCBI, and the NUMTFinder V0.5.4 tool was used to obtain the Numts sequence of the sample through the nuclear genome and mitochondrial genome sequences.
[0046] 2. Library data processing.
[0047] Use trimmomatic (version: 0.36) to process the sequenced sequences, including removing adapter sequences (adapter trimming), quality filtering (quality filtering), and removing low-quality sequences (trimming low-quality bases) to improve the quality and reliability of the sequence data and obtain the processed library data. The standards for processing sequencing data include:
[0048] The condition for removing the adapter sequence is that a maximum of 2 mismatches are allowed, and the threshold for adapter removal is 15; quality filtering is performed using a sliding window with a window size of 5 bases, and pruning is performed when the average quality value is less than 20; sequences with a length greater than or equal to 35 bases and an average quality greater than or equal to 20 are retained; the front and back ends of the sequence are not trimmed. The entire process is processed in parallel using 2 threads, unpaired sequences are discarded, and paired sequences are saved in the specified file.
[0049] 3. Three comparisons.
[0050] The purpose of the three alignments is to ensure that the Numts reads are removed while obtaining the mitochondrial reads, and to recover the mitochondrial reads that were misidentified as Numts.
[0051] 3.1. Obtaining preliminary mitochondrial readings.
[0052] Use bwa (version: 0.7.17-r1188) to align the processed library data with the mtgenome genome (from NCBI), and then process the aligned data through samtools (Version: 1.15.1) to obtain the mitochondrial related reading file mt1.bam.
[0053] 3.2. Obtain Numts readings in mitochondria.
[0054] The file mt1.bam (mitochondrial related reading file) obtained in 3.1 was converted into the mt1.fq sequence, and then the mt1.fq sequence was trimmed by trimmomatic to remove low-quality sequences (trimming was performed when the average quality value was less than 20), and then aligned with the Numts sequence for the second time to obtain the secondary alignment file nunt-sorted.bam, which is the Numts reading contained in the first alignment file, and was sorted and deduplicated by the samtools software.
[0055] 3.3. Recovering mitochondrial reads misidentified as Numts.
[0056] The nunt-sorted.bam (i.e., the Numts reads obtained from the second alignment) after deduplication was converted into the nunt.fq sequence. The nunt.fq sequence was trimmed by trimmomatic to remove low-quality sequences (trimming was performed when the average quality value was less than 20) and aligned with the mitochondrial genome for the third time to obtain numttomt.bam, and the mitochondrial reads that were misidentified as Numts were restored.
[0057] 3.4. Integrate bam files.
[0058] Integrate the mitochondrial readings mt1.bam obtained in step 3.1 with the recovered mitochondrial readings numttomt.bam obtained in step 3.3, and remove the Numts readings numt-sorted.bam in the second step. All mitochondrial readings can be obtained.
[0059] 4. Extract information of each indicator.
[0060] All mitochondrial reads were processed using relevant Python scripts to obtain the mtDNA heteroplasmy index (MHS-PDR; MHS-BMP; MHR-Entropy; MHS-PDRP; MHS-qPDRP) information report for each sample.
[0061] 4.1mtDNA heyeroplasmy sites'-proportion of Discordant Reads (MHS-PDR).
[0062] MHS-PDR refers to the ratio of discordant bases to the total number of bases in all sequencing reads covering a certain heterogeneous site.
[0063] 4.2mtDNA heyeroplasmy sites'-base mutation Polymorphism (MHS-BMP).
[0064] MHS-BMP refers to the frequency distribution of different base types at a certain heterogeneous site and their variation diversity.
[0065] 4.3 Entropy of mtDNA heteroplasmic region' (mtDNA heyeroplasmy region'-Entropy, MHR-Entropy).
[0066] MHR-Entropy is an indicator based on information entropy, which quantifies the uncertainty of base diversity and distribution in a certain region. The greater the entropy, the higher the base diversity in the region.
[0067] 4.4mtDNA heyeroplasmy sites'-proportion of Discordant Read Pairs (MHS-PDRP).
[0068] MHS-PDRP is the ratio of read pairs with inconsistent mutations in the shared region to the total number of read pairs in all paired sequencing read pairs covering a certain heterogeneous site.
[0069] 4.5 Quantification of discordant read pairs at mtDNA heteroplasmy sites (mtDNA heyeroplasmy sites'-quantitative proportion of Discordant Read Pairs, MHS-qPDRP).
[0070] MHS-qPDRP calculates the mutation degree of heterogeneous sites based on the mutation consistency between paired sequencing reads. The number of mutations and their distribution pattern are converted into quantitative indicators by normalizing the Hamming distance.
[0071] 3. Data processing and screening of target sites in age-related mtDNA heteroplasmy indicators.
[0072] 1. Data processing.
[0073] The dplyr package (R language) was used to integrate the information of five mtDNA heteroplasmy indicators of 55 samples, and outliers were detected according to the 3sigma principle. Outliers were set as null values, and the KNN model was used to interpolate missing values.
[0074] 2. Screening of age-related target sites.
[0075] Based on Spearman correlation analysis, the correlation between the target sites and age in the five mtDNA heteroplasmy indicators was less than or equal to -0.3 or greater than or equal to 0.3.
[0076] 4. Model, evaluate and predict the target sites and ages in different indicators respectively.
[0077] Four machine learning models for age and target sites of each indicator were constructed based on the xgboost package, randomForest package, e1071 package and glmnet package, namely elastic regression model, support vector machine regression model, XGBoost regression model and random forest regression model. The leave-one-out cross-validation method was used to evaluate each model, and the evaluation index of each model (R 2 , MSE, RSME) and the optimal parameters of each model. At the same time, the correlation between the predicted age and the actual age of the four different models needs to be determined to select the optimal model. 2 It represents correlation, MAE represents the mean of absolute errors (hereinafter referred to as errors), and RMSE represents the root mean square error. The evaluation criteria are that the higher the correlation, the smaller the error, and the higher the accuracy of the model.
[0078] 5. Results of target site modeling based on different indicators.
[0079] 1. MHS-PDR modeling results.
[0080] Through Spearman correlation analysis, 37 target sites were screened out (the correlation between the sites and age in the MHS-PDR index was greater than or equal to 0.3 or less than or equal to -0.3), and the age of the 37 target sites was predicted using the leave-one-out cross-validation method combined with four machine learning models: elastic regression model, support vector machine regression model, XGBoost regression model, and random forest regression model. The results are shown in Table 1. The results show that the elastic regression model, XGBoost regression model, and random forest regression model can all be applied to the model for calculating age based on the target sites in the MHS-PDR index.
[0081] When the elastic regression model is used to calculate age, in the training set, the correlation between the calculated age and the actual age can reach 0.9627, with an error of only 5.14 months; in the test set, the correlation between the predicted age and the actual age can reach 0.8558, with an error of only 14.25 months; when the support vector machine regression model is selected as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9969, with an error of only 3.38 months; in the test set, the correlation between the predicted age and the actual age can reach 0.532, with an error of only 32.65 months; when the random regression model is selected as the analysis model, the correlation between the predicted age and the actual age can reach 0.9969, with an error of only 3.38 months; in the test set, the correlation between the predicted age and the actual age can reach 0.532, with an error of only 32.65 months; When the machine forest regression model is used as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9718, with an error of only 9.92 months, and in the test set, the correlation between the predicted age and the actual age can reach 0.8197, with an error of only 22.34 months; when the XGBoost regression model is selected as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9983, with an error of only 1.52 months, and in the test set, the correlation between the predicted age and the actual age can reach 0.7383, with an error of only 19.72 months. According to the analysis of the test set, since the elastic regression model has the highest correlation and the smallest error, the elastic regression model is superior in the model for calculating age based on the MHS-PDR indicator.
[0082] Table 1
[0083]
[0084] 2. MHS-BMP modeling results.
[0085] 67 target sites were screened out by Spearman correlation analysis (the correlation between the sites and age in the MHS-BMP index was greater than or equal to 0.3 or less than or equal to -0.3), and the age of the 67 target sites was predicted by four machine learning models using the leave-one-out cross-validation method combined with the elastic regression model, support vector machine regression model, XGBoost regression model, and random forest regression model. The results are shown in Table 2. The results show that both the support vector machine regression model and the random forest regression model can be applied to the model for calculating age based on the target sites in the MHS-PDR index.
[0086] When the elastic regression model is used to calculate age, in the training set, the correlation between the calculated age and the actual age can reach 0.9956, with an error of only 1.91 months; in the test set, the correlation between the predicted age and the actual age can reach 0.52, with an error of only 22.97 months; when the support vector machine regression model is selected as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9970, with an error of only 3.53 months; in the test set, the correlation between the predicted age and the actual age can reach 0.7029, with an error of only 32.12 months; when the random regression model is selected as the analysis model, the correlation between the predicted age and the actual age can reach 0.9956, with an error of only 1.91 months; in the test set, the correlation between the predicted age and the actual age can reach 0.52, with an error of only 22.97 months; When the random forest regression model is used as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9734, with an error of only 8.89 months, and in the test set, the correlation between the predicted age and the actual age can reach 0.7981, with an error of only 23.87 months; when the XGBoost regression model is selected as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9977, with an error of only 1.25 months, and in the test set, the correlation between the predicted age and the actual age can reach 0.5360, with an error of only 25.44 months. According to the analysis of the test set, the random forest regression model has the highest correlation and the smallest error, so it is superior in the model for calculating age based on the MHS-BMP indicator.
[0087] Table 2
[0088]
[0089]
[0090] 3.MHR-Entropy modeling results.
[0091] Through Spearman correlation analysis, 12 target sites were screened out (the correlation between the site and age in the MHR-Entropy index was greater than or equal to 0.3 or less than or equal to -0.3), and the four machine learning models, elastic regression model, support vector machine regression model, XGBoost regression model and random forest regression model, were used to predict age for the 12 target sites using the leave-one-out cross-validation method. The results are shown in Table 3. The results show that the models are not suitable for the model for calculating age based on the target sites in the MHR-Entropy index.
[0092] When the elastic regression model is used to calculate age, in the training set, the correlation between the calculated age and the actual age can reach 0.7183, with an error of only 15.74 months; in the test set, the correlation between the predicted age and the actual age can reach 0.2762, with an error of only 25.75 months; when the support vector machine regression model is selected as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9924, with an error of only 3.97 months; in the test set, the correlation between the predicted age and the actual age can reach 0.1189, with an error of only 24.55 months; When the random forest regression model was used as the analysis model, the correlation between the predicted age and the actual age in the training set was 0.9512, with an error of only 8.41 months, and in the test set, the correlation between the predicted age and the actual age was 0.3592, with an error of only 22.78 months; when the XGBoost regression model was selected as the analysis model, the correlation between the predicted age and the actual age in the training set was 0.9888, with an error of only 3.29 months, and in the test set, the correlation between the predicted age and the actual age was 0.3975, with an error of only 20.34 months. Due to the low correlation and high error in the test set, it was shown that the models were not suitable for the model for calculating age based on the fecal MHR-Entropy target site.
[0093] Table 3 Regression evaluation indicators of different machine learning modeling in MHR-Entropy
[0094]
[0095]
[0096] 4. MHS-PDRP modeling results
[0097] 34 target sites were screened out by Spearman correlation analysis (the correlation between the sites in the MHS-PDRP index and age was greater than or equal to 0.3 or less than or equal to -0.3). In order to further verify the accuracy of the estimated age, the leave-one-out cross-validation method was combined with four machine learning models: elastic regression, support vector machine, XGBoost regression, and random forest to predict the age of the 34 target sites. The results are shown in Table 4. The results show that none of the models are suitable for the model for calculating age based on the target sites in the MHS-PDRP index.
[0098] When the elastic regression model is used to calculate age, in the training set, the correlation between the calculated age and the actual age can reach 0.9396, with an error of 7.31 months. In the test set, the correlation between the predicted age and the actual age can reach 0.4650, with an error of only 26.69 months. When the support vector machine regression model is selected as the analysis model, in the training set, the correlation between the predicted age and the actual age can reach 0.9973, with an error of 3.48 months. In the test set, the correlation between the predicted age and the actual age can reach 0.4697, with an error of 27.46 months. When the random forest regression model was used as the analysis model, the correlation between the predicted age and the actual age in the training set was 0.9371, with an error of 9.75 months, and in the test set, the correlation between the predicted age and the actual age was 0.2361, with an error of 20.23 months; when the XGBoost regression model was selected as the analysis model, the correlation between the predicted age and the actual age in the training set was 0.9975, with an error of 1.42 months, and in the test set, the correlation between the predicted age and the actual age was 0.3795, with an error of 18.34 months. Due to the low correlation and high error in the test set, it is shown that none of the models are suitable for the model for calculating age based on the target site in the MHS-PDRP indicator.
[0099] Table 4 Regression evaluation indicators of different machine learning modeling of MHS-PDRP
[0100]
[0101]
[0102] 5.MHS-qPDRP modeling results.
[0103] 51 target heterogeneous sites were screened out by Spearman correlation analysis (the correlation between the sites and age in the MHS-qPDRP index was greater than or equal to 0.3 or less than or equal to -0.3), and the four machine learning models, including the elastic regression model, support vector machine model, XGBoost regression model and random forest model, were used to predict age at the 51 target heterogeneous sites using the leave-one-out cross-validation method. The results are shown in Table 5. The results show that the random forest regression model is suitable for constructing a model for calculating age based on the MHS-qPDRP target heterogeneous sites.
[0104] When the elastic regression model is used to calculate age, in the training set, the correlation between the calculated age and the actual age can reach 0.8740, with an error of 14.85 months. In the test set, the correlation between the predicted age and the actual age can reach 0.8983, with an error of only 15.94 months. When the support vector machine regression model is used to calculate age, in the training set, the correlation between the predicted age and the actual age can reach 0.9969, with an error of 3.35 months. In the test set, the correlation between the predicted age and the actual age can reach 0.2618, with an error of 35.14 months. When the random forest regression model is used as the analysis model, the correlation between the predicted age and the actual age can reach 0.9835 in the training set, with an error of 8.62 months, and in the test set, the correlation between the predicted age and the actual age can reach 0.8224, with an error of 23.70 months; when the XGBoost regression model is selected as the analysis model, the correlation between the predicted age and the actual age can reach 0.9982 in the training set, with an error of 1.35 months, and in the test set, the correlation between the predicted age and the actual age can reach 0.6465, with an error of 22.68 months. Since the elastic regression has the highest correlation and the smallest error, the elastic regression model is superior in constructing a model for calculating age based on MHS-qPDRP.
[0105] Table 5
[0106]
[0107]
Claims
1. A method for constructing a calculation age model, characterized in that: The method is to construct a model for calculating age based on mtDNA heteroplasmy indicators extracted from feces, and the heteroplasmy indicators include quantification of inconsistent reading ratios of mtDNA heteroplasmic sites, base mutation diversity of mtDNA heteroplasmic sites, or inconsistent reading pair ratios of mtDNA heteroplasmic sites.
2. The method according to claim 1, characterized in that The model is an elastic regression model, a support vector machine model, an XGBoost regression model or a random forest model.
3. The method according to claim 1, characterized in that The steps of the method are as follows: Step 1: Extract genomic DNA from feces; Step 2: construct a library based on genomic DNA, sequence and detect mtDNA heterogeneity, and then use trimmomatic to process the sequences in the library; Step 3: Align the processed library sequences three times to obtain all mitochondrial reads for each sample; Step 4: Process all mitochondrial reads of each sample through Python to obtain mtDNA heteroplasmy index information of each sample; Step 5: Process the mtDNA heteroplasmy index information and screen the target sites through the dplyr package; Step 6: Construct a machine learning model of age and target site based on the xgboost package, randomForest package, e1071 package and glmnet package.
4. The method according to claim 3, characterized in that: The sequencing in step 2 is WGS sequencing, and the sequencing depth is greater than or equal to 5000X.
5. The method according to claim 3, characterized in that: The trimmomatic processing of sequences in the library in step 2 includes removal of adapter sequences, quality filtering, and removal of low-quality sequences.
6. The method according to claim 3, characterized in that The method of the three comparisons described in step 3 is: Step 31: Use bwa to align the processed library sequence with the mtgenome genome, and then process it through samtools to obtain the mitochondrial related reading file mt1.bam; Step 32: Convert mt1.bam to mt1.fq sequence and use trimmomatic to process the sequence, and then perform a second alignment with the Numts sequence to obtain the secondary alignment file nunt-sorted.bam, which is then sorted and deduplicated using samtools software; Step 33: Convert the deduplicated nunt-sorted.bam to the nunt.fq sequence, remove the low-quality sequences using trimmomatic, and then align it with the mitochondrial genome for the third time to obtain numttomt.bam; Step 34: Integrate the file mt1.bam described in step 31 and the numttomt.bam described in step 33, and remove the numt-sorted.bam described in step 32 to obtain all mitochondrial readings.
7. The method according to claim 3, characterized in that The mtDNA heteroplasmy indicator information in step 4 includes the quantification of the inconsistent read ratio of the mtDNA heteroplasmy site, the base mutation diversity of the mtDNA heteroplasmy site, or the inconsistent read pair ratio of the mtDNA heteroplasmy site.
8. The method according to claim 3, characterized in that In step 5, the dplyr package processes the mtDNA heteroplasmy index information by respectively integrating the same mtDNA heteroplasmy index information, setting the outliers as null values according to the 3sigma principle, and using the KNN model to interpolate the missing values.
9. The method according to claim 3, characterized in that: The condition for screening the target site in step five is that the correlation between age and the target site in the mtDNA heteroplasmy index is greater than or equal to 0.3, or less than or equal to -0.
3.
10. A model for calculating age based on mtDNA heteroplasmic sites in fecal DNA, characterized in that: The model is constructed by the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method for efficiently enriching host DNA (deoxyribonucleic acid) from faeces of mammals
CN113186185A