A method for constructing a biological age prediction model based on DNA methylation
By screening and optimizing DNA methylation sites, a biological age prediction model suitable for the Chinese population was constructed, which solved the problems of high detection cost and significant influence of blood cell components in existing models, and achieved low-cost and accurate biological age prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2022-04-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing DNA methylation biological age prediction models suffer from problems such as high detection costs, excessive number of CpG sites, lack of data on Asian populations, and significant influence from blood cell components, which limits their application in the Chinese population.
By screening 31 candidate CpG sites suitable for the Chinese population, a biological age prediction model was constructed using methods such as elastic network regression and multiple linear regression. The model was optimized to reduce the influence of blood cell components and adjusted in the Zhejiang cohort population to establish a biological age prediction model suitable for the Chinese population.
It achieves low-cost, accurate prediction of biological age, is applicable to the Chinese population, reduces the influence of blood cell components, and provides a simple and effective tool for aging assessment.
Smart Images

Figure CN115240761B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a biological age prediction model based on DNA methylation applicable to the Chinese population and its construction method. Background Technology
[0002] Population aging is a pressing global issue. According to WHO statistics, the global population aged 60 and over will increase from 12% in 2015 to 22% in 2050. my country entered an aging society in 2000 and has now become the country with the largest elderly population and the fastest aging rate. Statistics show that in 2017, the population aged 60 and over in my country accounted for 17.9%, and it is predicted that this will exceed one-third by 2050. Population aging leads to a continuous increase in the incidence of aging and age-related diseases, placing a heavy burden on society and families, and has become a major social problem today. Aging refers to the progressive decline and loss of function in physiological, psychological, and cognitive abilities as age increases, leading to increased susceptibility to disease and ultimately death. Aging is an inevitable phenomenon in living organisms, but delaying aging, reducing the incidence of age-related diseases, and achieving healthy aging are important goals for today's society and core components of the Healthy China 2030 strategy. Conducting research on aging prevention and control has become a priority in the field of population health, and aging assessment is the foundation and key link in aging research and prevention.
[0003] Currently, the most commonly used indicator for assessing aging is chronological age, which is age calculated from the date of birth, also known as time-series age. Chronological age is associated with the decline of organ function, the occurrence of chronic diseases, and the risk of death. However, individuals of the same chronological age may show different levels of aging; some may appear younger in appearance and function, while others exhibit characteristics of older age. Furthermore, the risk of developing age-related diseases varies significantly among individuals of the same chronological age, resulting in high heterogeneity in health outcomes among the elderly. Therefore, chronological age has limitations as an indicator for assessing aging and is not an ideal assessment method.
[0004] Aging refers to the genetic and epigenetic, biochemical, and other functional phenotypic changes that occur in an organism under the influence of genetic and environmental factors. These changes include genomic instability, shortened telomere length, DNA methylation or demethylation, altered protein homeostasis, abnormal nutrient sensing regulation, mitochondrial dysfunction, apoptosis, stem cell exhaustion, and altered intercellular communication. Biological age, established using these aging-related molecular biological markers, can more accurately reflect the degree of aging, damage repair and tissue regeneration capacity, functional status, and assess current (or future) health status and lifespan. Compared to calendar age, biological age is a more suitable indicator for assessing aging. However, there is no gold standard for assessing biological age; therefore, aging-related biomarkers are typically used to predict biological age. Common aging biomarkers fall into two categories: molecular and phenotypic. Molecular biomarkers include genetic susceptibility, characteristic gene expression levels, small molecule metabolites, DNA methylation, and telomere length. Phenotypic biomarkers include physiological and biochemical indicators such as blood pressure, blood lipids, and blood glucose, as well as functional indicators such as cognition, memory, and grip strength. Among these biomarkers, DNA methylation is considered the most promising biomarker for assessing biological age in the human population because it can reflect molecular changes in the body under the influence of genetic and environmental factors, has a high correlation with chronological age, is stable in vivo, and has high-throughput detection methods.
[0005] DNA methylation refers to the covalent methylation modification of cytosine at the 5′ position of CpG islands in the genomic DNA sequence. CpG island methylation can affect and regulate gene expression, and is closely related to the occurrence and development of diseases such as embryonic development, malignant tumors, and immunodeficiency. Since the 1960s, numerous studies have found that age is related to the degree and distribution of genomic DNA methylation, and the methylation level of certain CpG islands increases with age. In the human genome, there are two types of age-related CpG islands: one type is hypomethylated CpG islands that activate or enhance gene expression, and the other type is hypermethylated CpG islands that suppress gene expression. Using these age-related methylated CpG sites, multifactorial statistical methods can be used to establish biological age prediction models, called methylation age models or epigenetic clocks. Currently, the most commonly used epigenetic clocks internationally are the Horvath age clock and the Hannum age clock. The Horvath age clock uses Illumina 27k methylation microarray data, employing 353 CpG sites as markers to build an age prediction model, while the Hannum age clock uses an Illumina 450k microarray, employing 71 CpG sites as markers. Both predicted methylation age and calendar age with high correlations (correlation coefficients of 0.97 and 0.96, respectively). After adjusting for confounding factors such as calendar age, numerous studies have found that methylation age is associated with obesity, insulin resistance, metabolic abnormalities, and the risk of developing and dying from age-related diseases (Alzheimer's disease, malignant tumors, cardiovascular diseases, Parkinson's disease, etc.). Furthermore, DNA methylation is a reversible chemical modification. Recent studies have found that after one year of comprehensive treatment with metformin and other medications, methylation age can be reversed by 1.5 years, and reduced by 2.5 years compared to the control group. Even more excitingly, current research has found that in vitro expression of the Yamanaka factor (which transforms somatic cells into pluripotent stem cells) can completely reset the epigenetic clock. In vivo hematopoietic stem cell therapy or transfusion of blood from young donors can also reverse the epigenetic clock. These findings suggest that biological age based on DNA methylation can be used in aging and anti-aging research.
[0006] However, existing biological age prediction models based on DNA methylation still have limitations. First, both the Horvath and Hannum age clocks include too many CpG sites (353 and 71 sites, respectively). Under current conditions, only whole-genome methylation microarray technology can meet the detection requirements, thus placing high demands on detection technology and costs, limiting the clinical application of the models. Second, existing models are constructed using calendar age as the dependent variable, without considering the heterogeneity of calendar age itself or optimizing for population characteristics. Essentially, these models can only predict calendar age or a "semi-biological age" with partial biological age information, and cannot fully reflect biological age information. Third, existing models are mainly based on data from European and American populations, lacking data from Asian or Chinese populations. Since methylation levels differ between different ethnic groups, these models may not be applicable to the Chinese population. Fourth, the methylation levels of some CpG sites vary among different blood cells, and blood cell composition affects the overall DNA methylation level of whole blood samples. Existing models constructed using whole blood samples do not take into account the influence of blood cell heterogeneity. Therefore, when using these models for biological age assessment, it is necessary to adjust the blood cell composition, which increases the complexity of the models and is not conducive to practical application.
[0007] Against this backdrop, there is an urgent need to establish a DNA methylation-based biological age prediction model suitable for the Chinese population. This model should minimize information redundancy while ensuring prediction accuracy, using the fewest possible CpG sites. Developing this practical, non-invasive, low-cost, and non-blood cell component-related biological age prediction model suitable for the Chinese population will be of great significance in aging assessment and related research in China. Summary of the Invention
[0008] The purpose of this invention is to address the shortcomings of existing technologies by providing a method for constructing a biological age prediction model based on DNA methylation. The prediction model constructed in this invention contains fewer methylation sites, is applicable to the Chinese population, and is unaffected by blood cell components. It can provide an important tool for aging assessment and the research and prevention of age-related diseases in the Chinese population.
[0009] The objective of this invention is achieved through the following technical solution: a method for constructing a biological age prediction model based on DNA methylation, comprising the following steps:
[0010] (1) Acquisition of DNA methylation sample data: Download raw data of 450k methylation chips from whole blood samples of Chinese population, which include calendar age data, from the GEO data website;
[0011] (2) Preprocessing of DNA methylation sample data: The raw data of the 450k methylation chip was preprocessed using the R language package "ChAMP".
[0012] (3) Site selection for the methylation age prediction model: Using calendar age as the dependent variable, 235,021 CpG sites as independent variables, and gender as a covariate, elastic network regression was used to select CpG sites for the prediction model. Elastic network regression was performed using the R language package "glmnet", with parameters set as follows: alpha = 0.5, connection function "Gaussian", and the penalty parameter lambda determined using 10-fold cross-validation, selecting the lambda value that minimizes the root mean square error. Bootstrap resampling was used to resample the training set multiple times. In each resampled dataset, a series of CpG sites were selected using the above elastic network regression. The frequency of each CpG site being selected in the multiple resamplings was then counted, and CpG sites with a selection frequency greater than 50% were selected as candidate methylation sites for model construction. A total of 31 candidate CpG sites were finally obtained: cg168 67657, cg07372824, cg07553761, cg14361627, cg24079702, cg13575925, cg21692159, cg03032497, c g20158366, cg18404041, cg06567855, cg23500537, cg07850154, cg21177396, cg06639320, cg176214 38. cg11847992, cg06515235, cg15059474, cg01620164, cg11423680, cg18507365, cg18933331, cg19 893664, cg15665792, cg17243289, cg03684893, cg16882373, cg25703552, cg00481951, cg27184585;
[0013] (4) Construction and evaluation of methylation age prediction models: Multiple linear regression, support vector machine, random forest, and gradient boosting regression tree were used for initial model construction and evaluation to select the optimal model construction method. Then, optimal subset regression was used to further screen methylation sites, calculating the prediction accuracy of each optimal model with 1 to 31 CpG sites. The prediction effects achievable with different numbers of CpG sites were compared, and the model with the optimal number of CpG sites was selected based on the Bayesian information criterion. Optimal subset regression was performed using the R language package "leaps," selecting 18 CpG sites from 31 CpG sites: cg06567855 The following 18 CpG sites were identified: cg15059474, cg03684893, cg16867657, cg06639320, cg11423680, cg18404041, cg18507365, cg14361627, cg21177396, cg13575925, cg11847992, cg07372824, cg16882373, cg06515235, cg07850154, cg25703552, and cg07553761. A methylation age prediction model was established for these 18 CpG sites using the optimal multiple linear regression method.
[0014] (5) Optimization of methylation age prediction model and establishment of biological age prediction model: The above methylation age prediction model was optimized using natural population data from any provincial cohort. First, people with serious chronic diseases and the 30% of people with the largest prediction bias were excluded. The remaining samples were normal people, whose biological age was considered to be approximately equal to their calendar age. In the normal people, calendar age was used as the dependent variable, and the regression coefficients of 18 CpG sites were fitted and adjusted using multiple linear regression. The optimized model is the biological age prediction model.
[0015] Furthermore, step 2 includes the following sub-steps:
[0016] (2.1) Filtering of probes (CpG sites) specifically includes: ① removing probes with a detection P value > 0.01, ② removing non-CpG probes, ③ removing probes containing SNP (single nucleotide polymorphism) sites, ④ removing probes with multiple reactive sites, and ⑤ removing probes distributed on sex chromosomes.
[0017] (2.2) Sample filtering: Principal component analysis was performed using the R language package "wateRmelon" to exclude outlier samples, i.e., samples whose first two principal components are outside the interquartile range of 2; in addition, samples with missing gender information were excluded.
[0018] (2.3) Calibration of probe signals: The signal values (β values) of each probe in all samples are normalized to correct the bias caused by the inconsistent distribution of type I and type II probe data in the chip design;
[0019] (2.4) Batch effect calibration: Empirical Bayes model was used to calibrate the batch effect of methylated data from different datasets.
[0020] (2.5) Calibration of blood cell heterogeneity: CpG sites with methylation levels that exhibit blood cell heterogeneity were eliminated.
[0021] The beneficial effects of this invention are as follows: Utilizing methylation data from 219 Chinese individuals, this invention employs an elastic network combined with bootstrap resampling to screen for 31 candidate methylation sites for modeling. Multiple linear regression, support vector machines, random forests, and gradient boosting regression trees are used for initial model construction and evaluation. Then, full subset regression is used to further screen methylation sites, resulting in a methylation age prediction model based on 18 methylation sites. Subsequently, the methylation age prediction model is optimized using natural population data from the Zhejiang cohort, ultimately yielding a biological age prediction model based on 18 methylation sites. This model is applicable to the Chinese population, contains a small number of methylation sites, is unaffected by blood cell components, and exhibits good predictive accuracy. Due to its simplicity, cost-effectiveness, and accuracy, this model can be widely promoted in clinical applications, providing an important tool for aging assessment and research and prevention of age-related diseases in the Chinese population. Attached Figure Description
[0022] Figure 1 It is a technology roadmap for the invention;
[0023] Figure 2 This is a graph showing the changes in BIC for models with different numbers of methylation sites in the optimal subset regression of the training set. Detailed Implementation
[0024] This invention relates to a biological age prediction model based on DNA methylation applicable to the Chinese population. The overall technical approach is described below. Figure 1 The present invention is specifically constructed through the following steps:
[0025] Step 1: Obtaining DNA methylation sample data
[0026] Raw 450kDNA methylation microarray data, including chronological age information and derived from blood samples, from a Chinese population were downloaded from the GEO data website (https: / / www.ncbi.nlm.nih.gov / geo / index.cgi). Based on the search results, seven independent databases met the requirements: GSE104812, GSE107737, GSE116379, GSE51388, GSE53740, GSE65638, and GSE84003. Data from healthy samples without age-related diseases were selected from these databases, resulting in 227 raw 450kDNA methylation microarray data sets.
[0027] Step 2: Preprocessing of DNA methylation sample data
[0028] Using the R language package "ChAMP", preprocess the acquired raw 450k DNA methylation microarray data according to the following steps:
[0029] (1) Filtering of probes (CpG sites): ① Remove probes with a detection P value > 0.01, i.e., probes with a signal intensity lower than the negative control probe; ② Remove non-CpG probes, i.e. quality control probes with non-methylation sites; ③ Remove probes containing SNP (single nucleotide polymorphism) sites; ④ Remove probes with multiple-hit sites; ⑤ Remove probes distributed on sex chromosomes.
[0030] (2) Sample filtering: ① Principal component analysis was performed using the R language package "wateRmelon" to exclude outliers. Outliers were defined as samples whose first two principal components were located outside twice the interquartile range. ② Samples with missing gender information were excluded.
[0031] (3) Calibration of probe signals: The signal values (β values) of each probe in all samples are normalized using the “BMIQ (Beta-Mixture Quantile)” method to correct the bias caused by the inconsistent distribution of type I and type II probe data in the chip design.
[0032] (4) Batch effect calibration: Empirical Bayes model was used to calibrate the batch effect of methylated data from different datasets.
[0033] (5) Calibration of blood cell heterogeneity: Multiple studies have shown that the methylation level of some CpG sites varies among different blood cells, and the composition of blood cells in whole blood changes with age. In order to eliminate the influence of blood cell composition on the modeling results, this invention removes CpG sites with blood cell heterogeneity in methylation level by referring to the current research results on blood cell heterogeneity of methylation sites.
[0034] Step 3: Site selection for methylation age prediction models
[0035] After the above data preprocessing, 219 samples remained, with 235,021 CpG loci. The data was split into a training set (155 cases) and a test set (64 cases) using stratified random sampling (5 percentiles) based on age, with 70% for the training set and 30% for the test set. Basic information about the training and test set samples is shown in Table 1.
[0036] Table 1: Basic Information of Training and Test Set Samples
[0037] Training set (n=155) Test set (n=64) p-value Age, median (interquartile range) 32.0(22.0-47.2) 32.0(21.8-47.2) 0.700 Gender, Number of People (Percentage) 0.999 male 76(49.0) 32(50.0) female 79(51.0) 32(50.0)
[0038] Note: Age was compared between groups using the Wilcoxon rank-sum test, and gender was compared between groups using the chi-square test.
[0039] In the training set, calendar age was used as the dependent variable, 235,021 CpG sites as independent variables, and gender as a covariate. Elastic Net regression was used to screen candidate CpG sites for the prediction model. Elastic Net regression was performed using the R package "glmnet" with the following parameters: alpha = 0.5, Gaussian connection function, and lambda penalty parameter determined using 10-fold cross-validation, selecting the lambda value that minimizes the root mean square error. Considering the small number of participants in the training set, to increase the stability of the screened CpG sites, this invention utilizes bootstrap resampling to resample the training set 1000 times. In each resampled dataset, a series of CpG sites were screened using the aforementioned Elastic Net regression method. Then, the frequency of each CpG site in the 1000 resamplings was counted, and CpG sites with a frequency greater than 50% were selected as candidate methylation sites for constructing the prediction model. This step yielded a total of 31 candidate CpG sites, and the basic information of each CpG site is shown in Table 2.
[0040] Table 2: Basic information on 31 candidate CpG sites
[0041]
[0042]
[0043] Note: The screening frequency is the frequency at which each CpG site is screened out in 1000 resamplings using elastic network regression.
[0044] The correlation coefficient is the Pearson correlation coefficient between each CpG locus and calendar age in the training set.
[0045] The p-value is the p-value corresponding to each correlation coefficient.
[0046] Step 4: Construction and Evaluation of Methylation Age Prediction Model
[0047] Based on the 31 CpG loci selected above, models were constructed using different modeling methods in the training set. This invention selected four methods for model construction: multiple linear regression, support vector machine, random forest, and gradient boosting regression tree. For model evaluation, the mean absolute error (MAE) and root mean square error (RMSE) of predicted age and calendar age, as well as the Pearson correlation coefficient (R) between them, were used. Smaller MAE and RMSE, and larger R, indicate better predictive performance of the established model. This invention used bootstrap resampling (1000 resamplings) for internal validation in the training set and performed external validation on the test set based on the model constructed in the training set. The results of internal and external validation for the four modeling methods are shown in Table 3.
[0048] Table 3: Evaluation results of four modeling methods for 31 CpG sites
[0049]
[0050] Note: MAE is the mean absolute error between predicted age and calendar age.
[0051] RMSE is the root mean square error between the predicted age and the calendar age.
[0052] R is the Pearson correlation coefficient between predicted age and calendar age.
[0053] As shown in Table 3, various modeling methods can predict age well using 31 CpG loci. Among them, the multiple linear regression model has the best external validation results (lowest MAE, RMSE, and highest R) in addition to good internal validation results. Furthermore, the multiple linear regression model has a simple structure, good interpretability, and is conducive to practical application. Therefore, this invention chooses multiple linear regression as the model construction method. Based on the multiple linear regression model, this invention uses optimal subset regression to further screen and compress the 31 candidate CpG loci. Optimal subset regression is performed using the R language package "leaps" to calculate the prediction performance of all possible models with CpG loci from 1 to 31, and selects the optimal model for each number of CpG loci, resulting in 31 different models. These 31 models are comprehensively evaluated using three model evaluation metrics: MAE, RMSE, and R, as well as the Bayesian Information Criterion (BIC). The evaluation results of the optimal model under different numbers of CpG sites are shown in Table 4; the changes in the BIC value of the optimal model under different numbers of CpG sites are shown in Table 4. Figure 2 .
[0054] Table 4: Evaluation results of the optimal model under different numbers of CpG sites in whole subset regression.
[0055]
[0056]
[0057] Note: MAE is the mean absolute error between predicted age and calendar age.
[0058] RMSE is the root mean square error between the predicted age and the calendar age.
[0059] R is the Pearson correlation coefficient between predicted age and calendar age.
[0060] As shown in Table 4, with the increase in the number of CpG sites in the model, the MAE and RMSE of both internal and external validation initially decreased rapidly, then the rate of decrease gradually slowed down; the correlation coefficient R initially increased rapidly, then the rate of increase gradually slowed down. When the number of CpG sites in the model reached 18, the MAE and RMSE did not show a significant decrease, and the correlation coefficient R did not show a significant increase. Meanwhile, from... Figure 2The results show that as the number of CpG sites in the model increases, the BIC value exhibits a U-shaped curve, reaching its minimum when the number of CpG sites in the model is 18. Therefore, this invention considers 18 CpG sites to be the optimal number of sites for the current prediction model, and the multiple linear regression model established using these 18 CpG sites is the current optimal prediction model, balancing the inclusion of a smaller number of CpG sites with the accuracy of model prediction. The specific regression coefficients of the 18 CpG sites and the intercept in this multiple linear regression model are shown in Table 5. Considering that there are currently multiple platforms for methylation detection, to ensure the usability of the model across different detection platforms and reduce systematic errors caused by detection platforms, this invention also calculated the standardized regression coefficients of these 18 CpG sites.
[0061] Table 5: Regression coefficients and standardized regression coefficients of 18 CpG sites in the multiple linear regression model
[0062]
[0063]
[0064] Note: The predicted age obtained using standardized regression coefficients is the standardized predicted age. Destandardization is required to obtain the original predicted age. The destandardization method is: Standardized predicted age × Standard deviation of the predicted population's calendrical ages + Mean of the predicted population's calendrical ages.
[0065] Finally, the multiple linear regression model established for these 18 CpG sites was evaluated, and the specific evaluation results are shown in Table 6.
[0066] Table 6: Evaluation results of the multiple linear regression model for 18 CpG sites
[0067]
[0068] Note: MAE is the mean absolute error between predicted age and calendar age.
[0069] RMSE is the root mean square error between the predicted age and the calendar age.
[0070] R is the Pearson correlation coefficient between predicted age and calendar age.
[0071] As shown in Table 6, the model for these 18 CpG sites performed well in both internal validation on the training set and external validation on the test set. In internal validation, the model outperformed the methylation age prediction model established by Horvath and Hannum (Horvath: R = 0.97, MAE = 2.9; Hannum: R = 0.963, RMSE = 3.88). In external validation, the model also outperformed the Horvath and Hannum models (Horvath: R = 0.96, MAE = 3.6; Hannum: R = 0.905, RMSE = 4.89).
[0072] Step 5: Optimization of the methylation age prediction model and establishment of the biological age prediction model
[0073] The methylation age prediction model (GEO model) established using GEO data described above uses calendar age as the dependent variable for fitting. Therefore, the predicted methylation age reflects more of the calendar age, or a "semi-biological age" containing some biological age information. This predicted age does not fully reflect biological age information; therefore, the methylation age prediction model needs to be optimized to establish a biological age prediction model.
[0074] Numerous studies have shown that aging status varies within a natural population. Some individuals, due to genetic, behavioral, and environmental factors, have biological ages that deviate from their chronological age (aging type and young type), and their biological age is not equal to their chronological age. The remaining individuals have smaller deviations in biological age (normal type), and their biological age can be considered equal to their chronological age. Therefore, if we use the normal type population as the dependent variable, and adjust the regression coefficients of the 18 CpG sites in the GEO model, the optimized model predicts the biological age, and this model can be considered a biological age prediction model. In this invention, the population's health information and the deviation between the methylation age predicted by the GEO model and the chronological age are used to reflect population heterogeneity. Based on relevant biological knowledge and statistical experience, this invention defines heterogeneous populations as: (1) those with severe chronic diseases; and (2) the 30% of the population with the largest prediction deviation.
[0075] Based on the above optimization principle, this invention utilizes a sample of the Zhejiang cohort population established in the laboratory and employs random sampling to select 1513 individuals for methylation-targeted sequencing (MethylTarget method) of 18 CpG sites. Among these, individuals with unqualified methylation-targeted sequencing data and those lacking age or gender information were excluded, resulting in 1399 qualified samples. Of these 1399 samples, individuals with heterogeneous aging conditions were excluded: (1) those with cardiovascular disease, tumors, chronic liver disease, or chronic kidney disease (294 individuals); (2) the 30% of individuals whose methylation age predicted by the GEO model deviated most significantly from their chronological age (332 individuals). Since the detection platforms used for methylation-targeted sequencing (MethylTarget method) and the GEO model (450k chip) are different, the regression coefficients used in the GEO model during the deviation calculation are standardized regression coefficients. Finally, in the remaining 773 normal samples, using the calendar age of these samples as the dependent variable, multiple linear regression was employed to fit and adjust the regression coefficients of the 18 CpG sites in the GEO model. The optimized model can be considered a biological age prediction model. The regression coefficients of the 18 CpG sites in this biological age prediction model are shown in Table 7. Similarly, to ensure the usability of the biological age prediction model across different detection platforms and to reduce systematic errors caused by detection platforms, Table 7 also presents the standardized regression coefficients of the 18 CpG sites in the biological age prediction model.
[0076] Table 7: Regression coefficients and standardized regression coefficients of 18 CpG sites in the biological age prediction model
[0077]
[0078]
[0079] Note: The predicted age obtained using standardized regression coefficients is the standardized predicted age. Destandardization is required to obtain the original predicted age. The destandardization method is: Standardized predicted age × Standard deviation of the predicted population's calendrical ages + Mean of the predicted population's calendrical ages.
[0080] The biological age prediction model was evaluated, and the evaluation results of the biological age prediction model in 773 normal samples and 1399 total samples (natural population) are shown in Table 8.
[0081] Table 8: Evaluation results of biological age prediction models for 18 CpG sites
[0082]
[0083] Note: MAE is the mean absolute error between predicted age and calendar age.
[0084] RMSE is the root mean square error between the predicted age and the calendar age.
[0085] R is the Pearson correlation coefficient between predicted age and calendar age.
[0086] As shown in Table 8, the prediction bias of the biological age prediction model was relatively small in the 773 normal samples, but relatively large in the 1399 total samples. This result is consistent with the theoretical characteristics of biological age: in the normal population, biological age is approximately equal to chronological age with a small deviation; however, in the general population, due to heterogeneity, biological age does not equal chronological age, and the deviation is larger. This result supports the conclusion that the optimized model can be used as a biological age prediction model.
[0087] In summary, this invention suggests that the optimized model containing 18 CpG sites can serve as a practical biological age prediction model that is suitable for the Chinese population, non-invasive, low-cost, and not related to blood cell components.
Claims
1. A method for constructing a biological age prediction model based on DNA methylation, characterized in that, Includes the following steps: (1) Acquisition of DNA methylation sample data: Download raw data of 450k methylation chips from whole blood samples of Chinese population, which include calendar age data, from the GEO data website; (2) Preprocessing of DNA methylation sample data: The raw data of the 450k methylation chip was preprocessed using the R language package "ChAMP". (3) Site selection for the methylation age prediction model: Using calendar age as the dependent variable, 235,021 CpG sites as independent variables, and gender as a covariate, elastic network regression was used to select CpG sites for the prediction model. Elastic network regression was performed using the R software package "glmnet", with parameters set as follows: alpha = 0.5, Gaussian connection function, and the penalty parameter lambda determined using 10-fold cross-validation, selecting the lambda value that minimizes the root mean square error. Bootstrap resampling was used to resample the training set multiple times. In each resampled dataset, a series of CpG sites were selected using the above elastic network regression. The frequency of each CpG site being selected in the multiple resamplings was then counted, and CpG sites with a selection frequency greater than 50% were selected as candidate methylation sites for model construction. A total of 31 candidate CpG sites were finally obtained: cg168 67657, cg07372824, cg07553761, cg14361627, cg24079702, cg13575925, cg21692159, cg03032497, c g20158366, cg18404041, cg06567855, cg23500537, cg07850154, cg21177396, cg06639320, cg176214 38. cg11847992, cg06515235, cg15059474, cg01620164, cg11423680, cg18507365, cg18933331, cg19 893664, cg15665792, cg17243289, cg03684893, cg16882373, cg25703552, cg00481951, cg27184585; (4) Construction and evaluation of methylation age prediction models: Multiple linear regression, support vector machine, random forest, and gradient boosting regression tree were used for initial model construction and evaluation to select the optimal model construction method. Then, optimal subset regression was used to further screen methylation sites, calculating the prediction accuracy of each optimal model with 1 to 31 CpG sites. The prediction effects achievable with different numbers of CpG sites were compared, and the model with the optimal number of CpG sites was selected based on the Bayesian information criterion. Optimal subset regression was performed using the R language package "leaps," selecting 18 CpG sites from 31 CpG sites: cg06567855 The following 18 CpG sites were identified: cg15059474, cg03684893, cg16867657, cg06639320, cg11423680, cg18404041, cg18507365, cg14361627, cg21177396, cg13575925, cg11847992, cg07372824, cg16882373, cg06515235, cg07850154, cg25703552, and cg07553761. A methylation age prediction model was established for these 18 CpG sites using the optimal multiple linear regression method. (5) Optimization of methylation age prediction model and establishment of biological age prediction model: The above methylation age prediction model is optimized using natural population data from any provincial cohort. First, people with serious chronic diseases and the 30% of people with the largest prediction bias are excluded. The remaining samples are normal people, and their biological age can be considered to be approximately equal to their calendar age. In the normal population, with calendar age as the dependent variable, multiple linear regression was used to fit and adjust the regression coefficients of 18 CpG sites. The optimized model is the biological age prediction model.
2. The method for constructing a biological age prediction model based on DNA methylation according to claim 1, characterized in that, Step (2) includes the following sub-steps: (2.1) Filtering of probes, specifically including: ① removing probes with a detection P value > 0.01, ② removing non-CpG probes, ③ removing probes containing single nucleotide polymorphism sites, ④ removing probes with multiple reaction sites, and ⑤ removing probes distributed on sex chromosomes. (2.2) Sample filtering: Principal component analysis was performed using the R language package "wateRmelon" to exclude outlier samples, i.e., samples whose first two principal components are outside twice the interquartile range; in addition, samples with missing gender information were excluded. (2.3) Calibration of probe signals: Normalize the signal values of each probe in all samples to correct the bias caused by the inconsistent distribution of Type I and Type II probe data in the chip design; (2.4) Batch effect calibration: An empirical Bayesian model was used to calibrate the batch effect of methylated data from different datasets; (2.5) Calibration of blood cell heterogeneity: CpG sites with methylation levels that exhibit blood cell heterogeneity were eliminated.