DNA Methylation Age Prediction Method and System Based on the uniAge Model

CpG sites compatible with RRBS and methylation chips were screened through the uniAge model, which solved the insufficient data coverage and applicability of the existing DNA methylation age clock, realized cross-platform DNA methylation age prediction, reduced sequencing costs, and improved the flexibility and accuracy of research and application.

CN119207556BActive Publication Date: 2025-07-22SHENZHEN E GENE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411678987.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-07-22
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

The existing DNA methylation age clock model is mainly based on DNA methylation chips. It has insufficient data coverage and limitations in applicability, making it difficult to flexibly apply in different data types and research scenarios, especially in genome-wide research and targeted DNA methylation sequencing.

Method used

Using the uniAge model, the CpG sites covered by RRBS and methylation chips were screened out simultaneously, and the age prediction model was trained using DNAme chip data in human peripheral blood, and the sequencing data was corrected to adapt to different data types, achieving cross-platform data compatibility and flexibility.

Benefits of technology

It improves genome coverage and prediction accuracy, reduces sequencing costs, is suitable for large-scale population research and clinical applications, and provides a wider range of research and application possibilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207556B_ABST
    Figure CN119207556B_ABST
Patent Text Reader

Abstract

The present invention belongs to, but is not limited to, the field of biotechnology, and discloses a method and system for predicting DNA methylation age based on the uniAge model, comprising the following steps: screening out CpG sites that can be covered by both RRBS and methylation chips; using DNAme chip data of human peripheral blood to screen sites and train a prediction model for age in the common CpG sites; correcting the DNAme level obtained from sequencing data to align with the chip data, or separately fitting model coefficients on the sequencing data using the screened CpG sites. The present invention predicts an individual's age using the DNAme levels of several CpG sites in peripheral blood; and trains an age prediction model that is applicable to both methylation chip data and sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, especially the analysis and interpretation of DNA methylation data, and is applicable to the technical field of detecting biological age through DNA methylation. Specifically, it relates to a DNA methylation age prediction method and system based on the uniAge model. Background Art

[0002] Changes in genomic DNA methylation are important driving factors and molecular markers of aging. The biological age modeled by DNA methylation of specific tissues can not only be used to reflect an individual's chronological age, but also reveal the degree of biological aging of an organism and even predict the risk of age-related diseases, which has significant scientific and clinical significance.

[0003] The Horvath Clock in 2013 is one of the most well-known epigenetic clocks, applicable to a variety of tissues and cell types, including human blood, skin, and internal organs. The Horvath Clock uses 353 CpG sites to predict age, and its results are often considered the most widely used and validated biological age indicator currently. The Hannum Clock in the same year was mainly based on the DNA methylation patterns of blood samples and used 71 CpG sites to estimate age. Although the application scope of the Hannum Clock is more limited compared to the Horvath Clock, it has high accuracy in some specific scenarios, such as blood sample analysis. The PhenoAge in 2018 combines methylation data and clinical risk factors (such as inflammatory markers, blood biomarkers) to predict biological age. It not only focuses on biological age but also can predict disease risk and lifespan. This model is considered to have high practicality in health and disease research. The Skin & Blood Clock in the same year was specifically designed for skin and blood samples and is suitable for measuring the biological age of specific tissues, especially widely used in cosmetics and skin aging research. The GrimAge in 2019 is an improved epigenetic clock that combines DNA methylation with multiple clinical risk factors (such as smoking, inflammation, etc.). It is considered to be able to more accurately predict lifespan and the occurrence of age-related diseases than earlier clocks. The DNAmTL (Telomere Length Clock) in 2020 estimates telomere length through methylation patterns and indirectly measures biological age. Telomere length is an important marker of aging, and DNAmTL provides an alternative that does not rely on traditional telomere measurement methods. The PACE (PhenoAge Acceleration Clock) in 2021 focuses on the study of accelerated aging and measures the rate of biological age increase rather than static biological age. This clock is more suitable for evaluating the effects of anti-aging interventions. The ImmuneAge in 2022 is specifically used to predict immune function and the degree of immune system aging, especially has application prospects in vaccine research and geriatric research.

[0004] The existing DNA methylation age clocks have provided a powerful tool for researchers to understand the aging process, evaluate disease risk, and develop personalized health intervention programs. However, the existing methylation age clocks are mainly developed based on the CpG sites measured by DNA methylation chips (such as 450K and EPIC chips). Although this method has been widely used in scientific research and applications, there are still some limitations, especially with the diversification of application scenarios and data types, these limitations become particularly prominent.

[0005] First, the CpG sites used in chip technology are fixed and only cover part of the genome. This means that for the DNA methylation age clocks developed using chip technology, the actual methylation information that can be utilized is limited and it is difficult to comprehensively reflect the genomic methylation status of an individual. Especially when conducting whole-genome or larger-scale studies, this insufficient coverage may lead to an incomplete understanding of the aging process. More importantly, when researchers hope to apply these chip-based models to other types of methylation data, such as studies based on sequencing data like RRBS (Reduced-representation bisulfite sequencing) or WGBS (Whole genome bisulfite sequencing), many key CpG sites often cannot be covered in these sequencing data. This is because the set of CpG sites covered by reduced representation bisulfite sequencing (RRBS) and whole genome bisulfite sequencing (WGBS) technologies is not exactly the same as that of chip technology, thus limiting the applicability of these methylation age clocks in different data types.

[0006] Second, due to the limitations in the number and position of fixed sites on the chip, it is difficult to screen out the most effective sites for more cost-effective targeted DNA methylation sequencing when conducting personalized or specific-scenario methylation detection. In contrast, targeted DNA methylation sequencing technology has significant cost advantages in large-scale applications. Especially when researchers hope to apply the methylation clock technology to clinical or large-scale population studies, the high cost and uncontrollability of the number of fixed sites on the chip make it unsuitable for these more economical and flexible detection schemes. This further indicates that the current methylation clock models lack sufficient flexibility to adapt to scenarios with different research and application requirements. Summary of the Invention

[0007] In view of the problems existing in the prior art, the present invention provides a DNA methylation age prediction method and system based on the uniAge model. Different from traditional DNA methylation age clocks, our model can simultaneously process data from different technical platforms, including data from methylation chips (such as 450K and EPIC) and sequencing technologies (such as RRBS and WGBS). This compatibility ensures that our model can be used in more diverse research and application scenarios. It can not only cover more CpG sites, thereby improving the genomic coverage, but also flexibly switch between different types of methylation data. In this way, we have overcome the limitations of current DNA methylation age clocks in terms of data type and application scenario, and created a more general and adaptable technical platform.

[0008] In addition, our method also lays the foundation for future targeted DNA methylation sequencing technologies. This means that researchers can flexibly screen and design sets of CpG sites that are more suitable for specific scenarios according to different research needs, thereby significantly reducing sequencing costs and improving the accuracy and efficiency of research. The cost advantage of targeted DNA methylation sequencing makes it have great potential in large-scale population studies and clinical applications, and our invention provides the underlying technical support for achieving this goal.

[0009] Through this innovative method, we have not only solved the technical bottlenecks of existing methylation age clocks but also provided more extensive possibilities for future research and applications. The development of this general-purpose technology will promote the application of aging clocks in various scenarios, helping researchers to more deeply explore the aging mechanism, accurately evaluate individual health risks, and provide a more reliable basis for individualized health intervention measures.

[0010] The present invention is implemented as follows. A method for predicting DNA methylation age based on the uniAge model includes the following steps:

[0011] Step 1, screen out CpG sites that can be covered by both RRBS and methylation chips simultaneously;

[0012] Step 2, use the DNAme chip data of human peripheral blood to screen sites and train a prediction model for age in the common CpG sites;

[0013] Step 3, correct the DNAme level obtained from sequencing data, align it with the chip data, or separately fit the model coefficients on the sequencing data using the screened CpG sites.

[0014] Furthermore, the data used includes 450K and EPIC chip data of human whole blood samples, sequencing data of human whole blood samples, sequencing and chip data of the same biological samples, and human RRBS data.

[0015] Furthermore, the initial screening of CpG sites screens out the common and "reliable" CpG sites of RRBS and DNAme chips, meeting the following conditions:

[0016] Core sites of the chip: the intersection of the sites (probes) covered by the 450K, 850K (EPICv1), and 935K (EPICv2) arrays;

[0017] The probe source sequences (50nt) after base conversion (according to the BS sequencing data alignment method) are uniquely aligned to the end-to-end reference genome UCSC hs1 (T2T);

[0018] Core sites of RRBS: Most easily covered by sequencing. RRBS uses the MspI endonuclease, and the cleavage site is C^CGG;

[0019] Consistently methylated CpG sites: The methylation levels of the positive and negative strands have the smallest difference;

[0020] Intersection of chip sites and RRBS sites;

[0021] Located on autosomes.

[0022] Furthermore, training set and test set: Select the training set and test set according to the age distribution of all datasets, so that the training set has a certain number of samples in each age interval, and the training set and test set have a relatively consistent age distribution.

[0023] Furthermore, chip data preprocessing: Only retain the methylation values (beta value) of the pre-screened sites, 65,221 sites;

[0024] Remove samples with missing age and samples with age < 1 or age > 120;

[0025] Remove sites with a missing proportion > 10% in all samples;

[0026] Remove samples with a site missing proportion > 10%;

[0027] For each site, fill in the missing methylation values with random normal distribution values, and calculate the mean and variance from the non-missing values of this site.

[0028] Furthermore, screening of age-related CpG sites:

[0029] Further screen age-related methylation sites from 65,221 pre-screened sites. Let be the methylation level (beta value) of the i th sample at the j th CpG site;

[0030] Calculate the inner quartile range (IQR) of the methylation level of each site, ;

[0031] Calculate the correlation between the methylation of each site and age, Pearsons correlation and Spearmans correlation ;

[0032] Define ;

[0033] According to Sort the CpG sites from largest to smallest and select the 6,000 CpG sites with the largest scores to enter the following model.

[0034] Furthermore, use the LASSO model to screen CpG sites

[0035] The input data is 7,149 microarray training set samples (rows), including variables age, dataset, gender, and methylation beta values of 6,000 CpG sites (columns). Use the R package glmnet to fit the LASSO model

[0036]

[0037] .

[0038] Use the machine learning framework mlr3 in R to call glmnet. The learner parameters are lrn("regr.cv_glmnet", alpha = 1, lambda = 10^(seq(-5, 0, length.out = 100)), s = 'lambda.1se', standardize = T, intercept = T, type.measure ='mae'). You can also directly use glmnet for training without using mlr3.

[0039] Furthermore, Ridge regression age prediction model:

[0040] Use ridge regression to fit age on the CpG sites selected by the LASSO model. The model is:

[0041]

[0042] .

[0043] The learner parameters for using mlr3 to call glmnet are lrn("regr.cv_glmnet", alpha = 0, lambda = 10^(seq(-8, 0, length.out = 100)), s = 'lambda.1se', standardize = F, intercept = T, type.measure ='mae'). Among them, alpha = 0 indicates fitting ridge regression;

[0044] This model is denoted as full model or uniAge - 605, which contains 605 CpG sites, and the estimated value of age is

[0045]

[0046] wherein is the coefficient estimated by ridge regression, is the j -th methylation beta value of the CpG site, .

[0047] Furthermore, when applying the model to the sequencing data, quantile normalization is used to correct the distribution of the methylation levels of 605 (100) sites in the sequencing data to be the same as the methylation distribution of these 605 (100) sites in the microarray data training set. This correction can slightly improve the performance of the model on the sequencing data.

[0048] Another object of the present invention is to provide a DNA methylation age prediction system based on the uniAge model for the DNA methylation age prediction method based on the uniAge model, including:

[0049] A CpG site screening module for screening out CpG sites that can be covered by both RRBS and methylation microarrays;

[0050] An age prediction model training module for screening sites and training an age prediction model in the common CpG sites using the DNAme microarray data of human peripheral blood;

[0051] A sequencing data correction module for correcting the DNAme level obtained from the sequencing data to align with the microarray data, or separately fitting model coefficients on the sequencing data using the screened CpG sites.

[0052] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are as follows:

[0053] First, the present invention predicts an individual's age using the DNA methylation levels of several CpG sites in peripheral blood. Peripheral blood is an easily obtainable human tissue with little (few) contamination, and there is also the most available data on peripheral blood. These conditions ensure that the present invention can obtain a relatively accurate, robust, and easy-to-use DNA methylation age prediction model.

[0054] Second, the present invention uses the common CpG sites of the core sites of DNA methylation chips (the probe intersections of 450K and EPIC chips) and the core sites of RRBS sequencing as the candidate feature set for the prediction model. The 100 CpG sites used in the uniAge model can be captured in both chip samples and sequencing samples. Among them, in actual RRBS samples, an average of 90% of the 100 CpG sites in the uniAge model are covered by sequencing, while only (on average) 30% of the CpG sites are covered by the existing models. This unified CpG site design makes the uniAge model have broad applicability (compared with existing models).

[0055] Third, the present invention uses a step-by-step screening and refinement method for CpG sites, and finally determines the uniAge model with 100 CpG sites. In the evaluation of a large amount of real data, the uniAge model with 100 CpG sites has the same level of prediction accuracy as the complete model that does not limit the number of sites and contains 605 CpG sites. Low-cost targeted DNA methylation sequencing can be conveniently applied to the prediction model with a controllable number of CpG sites (100), which is extremely beneficial for large-scale research and applications. The existing HorvathAge depends on the input of 353 CpG sites, and GrimAge depends on the input of 1030 CpG sites. The dependence on too many site numbers is not conducive to the application of targeted DNA methylation sequencing.

[0056] Fourth, as the creative auxiliary evidence of the claims of the present invention, the technical solution of the present invention has broad commercial prospects in the fields of health management, anti-aging, insurance, and precision medicine after transformation:

[0057] a) Precision health management and personalized medicine

[0058] Health assessment and monitoring: Commercial DNA methylation detection services can provide users with personalized biological age assessments. Compared with traditional health checks, DNA methylation age can detect potential health problems earlier, thus helping users take early intervention measures.

[0059] Anti-aging and longevity intervention: Many anti-aging companies use DNA methylation age data to design personalized lifestyle adjustment and supplement programs for users. By regularly monitoring the changes in methylation age, the effectiveness of anti-aging measures can be evaluated.

[0060] Precision medicine: In the field of precision medicine, DNA methylation age can help doctors formulate treatment plans more suitable for patients, especially in cancer treatment and chronic disease management.

[0061] b) Life insurance and health insurance

[0062] Risk assessment: Insurance companies can use DNA methylation age as part of their health risk assessment. Biological age can better reflect an individual's health status than chronological age, thus helping insurance companies more accurately assess premiums and claims risks.

[0063] Personalized insurance products: By incorporating DNA methylation age, insurance companies can design more flexible personalized insurance products. For example, offering preferential premiums to those with good health and a lower biological age.

[0064] c) Corporate employee health management

[0065] Corporate health programs: Some companies have started to introduce health management programs based on DNA methylation age to help employees understand their biological age and adopt health intervention programs to delay aging, improve work efficiency and well-being.

[0066] Employee benefits and performance improvement: By regularly monitoring biological age, companies can optimize their employee health programs, thereby reducing sick leave and increasing work efficiency.

[0067] d) Anti-aging and beauty industry

[0068] Anti-aging product marketing: Beauty and skincare companies can develop personalized skincare programs based on the results of DNA methylation age testing and promote their products with this as a selling point. For example, a methylation age clock for skin aging can help optimize the design and effectiveness evaluation of anti-aging products.

[0069] Biological age-based personalized advice: By detecting the biological age of the skin, the beauty industry can provide targeted skincare and beauty advice to customers, thereby enhancing the customer experience and loyalty.

[0070] e) Nutrition and supplement market

[0071] Personalized nutrition programs: Based on the results of DNA methylation age testing, nutrition and health companies can customize personalized dietary and supplement programs for users to help delay biological aging.

[0072] Effectiveness evaluation: Supplement companies can use DNA methylation age to prove the effectiveness of their products, enhancing the scientific endorsement and market competitiveness of the products.

[0073] f) Data and analysis services

[0074] Big data analysis and prediction: Gene testing companies and health data analysis companies can use the large-scale DNA methylation data they collect for health trend analysis and disease prediction. Such data is of great value in the fields of pharmaceuticals, public health, and scientific research.

[0075] g) Consumer genetic testing services

[0076] Direct-to-consumer testing services: Genetic testing companies can offer DNA methylation age testing as part of their health and genetic analysis products to provide consumers with more personalized health information.

[0077] Health and aging tracking applications: Some startups have developed DNA methylation age testing products that can be combined with mobile applications to help consumers track their aging process through data visualization. Description of the drawings

[0078] Figure 1 is a flowchart of the DNA methylation age prediction method based on the uniAge model provided by an embodiment of the present invention;

[0079] Figure 2 is a schematic diagram of the age distribution of each dataset provided by an embodiment of the present invention;

[0080] Figure 3 is a schematic diagram of the correlation between methylation and age of CpG sites provided by an embodiment of the present invention;

[0081] Figure 4 is a schematic diagram of the distribution of the IQR of the methylation level of CpG sites provided by an embodiment of the present invention;

[0082] Figure 5 is a structural diagram of the DNA methylation age prediction system based on the uniAge model provided by an embodiment of the present invention;

[0083] Figure 6 is a schematic diagram of the fitting value of uniAge on the training set provided by an embodiment of the present invention;

[0084] Figure 7 is a schematic diagram of the predicted value of uniAge on the chip data test set provided by an embodiment of the present invention;

[0085] Figure 8 is a schematic diagram of the predicted value of the existing model on the chip data test set provided by an embodiment of the present invention;

[0086] Figure 9 is a schematic diagram of the predicted value of different models on the sequencing data test set provided by an embodiment of the present invention;

[0087] Figure 10 is a schematic diagram of the deletion ratio of sites of different models on the sequencing data test set provided by an embodiment of the present invention. Detailed implementation manners

[0088] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0089] As Figure 1 shown, a DNA methylation age prediction method based on the uniAge model provided by an embodiment of the present invention includes the following steps:

[0090] Step 1, screen out CpG sites that can be covered by both RRBS and methylation chips;

[0091] Step 2, use the DNAme chip data of human peripheral blood to screen sites and train the age prediction model among the common CpG sites;

[0092] Step 3, correct the DNAme level obtained from the sequencing data, align it with the chip data, or separately fit the model coefficients on the sequencing data using the screened CpG sites.

[0093] (Public) dataset

[0094] The data involved has 4 parts:

[0095] 1. 450K and EPIC chip data of human whole blood samples, used for CpG site screening and (chip) model training, with a total of 29 datasets and 15,610 samples;

[0096] 2. Sequencing data of human whole blood samples, used to evaluate the performance of the model on sequencing data, with a total of 2 datasets and 102 samples;

[0097] 3. Sequencing and chip data of the same biological samples, used to correct the technical bias between sequencing and chips, with two sources: GEO (Gene Expression Omnibus) and ENCODE (The Encyclopedia of DNA Elements). There are a total of 4 datasets and 92 samples from the GEO source, and a total of 155 samples from the ENCODE source;

[0098] 4. Human RRBS data, used for initial site screening, with a total of 4 datasets and 1,835 samples.

[0099] See "data / DNA methylation age dataset.xlsx" for the data summary information

[0100] See "data / {dataset}.tsv" for the phenotype information such as age, gender, and disease of each dataset sample. The script "data / pheno-data.R" for organizing the phenotype information is provided for reference.

[0101] Processed data of partial (15 datasets) chip data was downloaded for model training and evaluation. For convenient data reading, some files were modified, such as the variable names in the first line. The processed DNAme data (beta value) after modification can be found in "meth-data / *".

[0102] Existing DNAme age models

[0103] Name model Phenotype Platforms Samples Probes / CpGs probe pool PMID HorvathAge Linear model, age transformation Chronological age 450K / 27K 353 21,369, Intersect of 450K and 27K 24138928,25968125 HannumAge Linear model Chronological age 450K 656 71 450K 23177740 SkinBloodAge Linear model, age transformation Chronological age EPIC / 450K 2222 from multiple tissues / cell types 391 Intersect of 450K and EPIC 30048243 ZhangAge Linear model, DNAme normalization Chronological age EPIC / 450K 13,661, most blood samples 514 31443728 PhenoAge Linear model, time to death to age transformation Mortality EPIC / 450K / 27K 513 29676998 GrimAge, (Online only) Mortality EPIC / 450K 2356 1030 450,161 CpGs, Intersect of 450K and EPIC 30669119

[0104] The above models for evaluation were collated from the original text and supplementary materials. The article supplementary materials, etc. can be found in the corresponding folders of the "models / " directory. The CpG sites (chip probes) and corresponding coefficients of each model can be found in "models / {model}.tsv", and the program for calculating each model can be found in "models / DNAme-ages.R". Among them, GrimAge can only be used by online registration.

[0105] uniAge model

[0106] Initial screening of CpG sites

[0107] The purpose of the initial screening of the model is to screen out the common and "reliable" CpG sites of RRBS and DNAme chips, meeting the following conditions

[0108] 1. Core sites of the chip: The intersection of the sites (probes) covered by the 450K, 850K (EPICv1), and 935K (EPICv2) arrays

[0109] 2. The probe source sequences (50nt) after base conversion (according to the BS sequencing data alignment method) are uniquely aligned end-to-end to the reference genome UCSC hs1 (T2T)

[0110] 3. Core sites of RRBS: The most easily sequenced coverage. RRBS uses the MspI endonuclease, and the cleavage site is C^CGG

[0111] 4. CpG sites with consistent methylation: The methylation levels of the positive and negative strands have the smallest difference

[0112] 5. The intersection of chip sites and RRBS sites

[0113] 6. Located on autosomes

[0114] Chip probe screening

[0115] Intersection of covered sites (probes) for 450K, 850K (EPICv1), and 935K (EPICv2) arrays. nProbes = 369,845. Align the 50nt source sequence of the probes to the UCSC hs1 reference genome using bsmap (default parameters), remove probes with non-unique alignments, and the remaining 368,540 probes are denoted as set A.

[0116] RRBS core sites

[0117] Download the processed files of the human multi-tissue RRBS datasets GSE143673 (n = 451), GSE203169 (n = 82), GSE231984 (n = 781), GSE233417 (n = 521), nTotal = 1835.

[0118] According to the screening conditions

[0119] GSE143673: DP >= 10, sample proportion >= 0.4: number = 1285147

[0120] GSE203169: DP >= 5, sample proportion >= 0.9: number = 855454

[0121] GSE231984: DP >= 5, sample proportion >= 0.8: number = 2532617

[0122] GSE233417: DP >= 1, sample proportion >= 0.95: number = 4921971

[0123] CpG sites that appear in at least 3 of the above 4 sets are defined as RRBS core sites (denoted as set B1), nCpGs = 1,448,485

[0124] Inconsistent strand methylation CpG sites

[0125] Use the datasets GSE143673, GSE203169, GSE231984, nTotal = 1314.

[0126] According to the screening conditions, the absolute value of the difference in CpG positive and negative strand methylation >= 0.2, in 3 datasets

[0127] GSE143673: DP >= 10, sample proportion >= 0.05: number = 60204

[0128] GSE203169: DP >= 5, sample proportion >= 0.1: number = 33925

[0129] GSE231984: DP >= 5, sample proportion >= 0.2: number = 89610

[0130] The CpG sites that appear in any of the above sets (the union of the sites) are called inconsistent strand methylated CpG sites (denoted as set B2), nCpG = 174,072

[0131] The available CpG sites screened using RRBS data are set B = B1 - B2, nCpGs = 1,315,096

[0132] Denote the CpG sites on autosomes as set C, then the initially screened CpG sites are A ∩ B ∩ C, nCpGs = 65,221. The information of the initially screened sites can be found in the file "uniAge / final - probe - autosomes.tsv"

[0133] The scripts "uniAge / statCoverage.py" and "uniAge / count - consistent - CpGs.ipynb" for initial site screening are for reference

[0134] Training set and test set

[0135] As Figure 2 shown, the training set and test set are selected according to the age distribution of all datasets, so that the training set has a certain number of samples in each age interval, and the training set and test set have a relatively consistent age distribution

[0136] 1. Training set:

[0137] a) Microarray data: GSE157131, GSE55763, GSE87648, GSE40279, GSE84727, GSE72680, GSE73103, GSE72775, GSE168739, nTotal = 7149

[0138] 2. Test set:

[0139] a) Chip data: GSE125105, GSE42861, GSE80417, GSE147221, GSE72773, GSE61496, nTotal = 3323

[0140] b) Sequencing data: GSE117593, GSE85928, nTotal = 102

[0141] The total number of chip samples used is 10,472, and the number of sequencing samples is 102

[0142] Preprocessing of chip data

[0143] Only retain the methylation values (beta value) of the initially screened sites, 65,221 sites

[0144] 1. Remove samples with missing age and samples with age < 1 or age > 120

[0145] 2. Remove sites with a missing proportion > 10% in all samples

[0146] 3. Remove samples with a site missing proportion > 10%

[0147] 4. For each site, fill in the missing methylation values with randomly generated normal distribution values, where the mean and variance are calculated from the non-missing values of the site

[0148] Screening of age-related CpG sites

[0149] Further screen age-related methylation sites from 65,221 initially screened sites. Let be the methylation level (beta value) of the i th sample at the j th CpG site

[0150] 5. Calculate the inner quartile range (IQR) of the methylation level for each site

[0151] 6. Calculate the correlation between methylation and age for each site, Pearsons correlation and Spearmans correlation

[0152] 7. Define

[0153] 8. Sort the CpG sites in descending order according to and select the top 6000 CpG sites with the highest scores to enter the following model

[0154] As Figure 3 , 4 shown, this step screens for CpG sites with the greatest correlation with age and the largest range of methylation changes. Since the fluorescence values of the chip are continuously distributed, this screening removes those sites whose correlation with age depends on weak methylation changes. Tests show that this step can improve the model performance.

[0155] Use the LASSO model to screen for CpG sites. The input data are 7149 chip training set samples (rows), including variables age, dataset, gender, and methylation beta values of 6000 CpG sites (columns). Use the R package glmnet to fit the LASSO model.

[0156]

[0157] .

[0158] Use the machine learning framework mlr3 in R to call glmnet. The learner parameters are lrn("regr.cv_glmnet", alpha = 1, lambda = 10^(seq(-5, 0, length.out = 100)), s = 'lambda.1se', standardize = T, intercept = T, type.measure ='mae'). It is also possible not to use mlr3 and directly use glmnet for training.

[0159] Cross-validation

[0160] Use 10-fold cross-validation to determine the model hyperparameter lambda. Due to the randomness of cross-validation, repeat 5 times of 10-fold cross-validation to fit the LASSO model. Each fit can select variables (CpG sites), denoted as , and finally take the sites that appear in at least 4 model selections as the model sites, including the CpG site set . The dataset is not the variable of concern.

[0161] Batch effect. Since different datasets come from different studies and use different preprocessing methods, the batch effect between datasets is significant. Using batch effect correction methods limma:removeBatchEffect(), sva:ComBat() (parameter methods), and quantile normalization does not improve the model performance. Therefore, the batch effect of the training data is not corrected in advance, and terms are added to the model to model the batch effect.

[0162] Ridge Regression Age Prediction Model

[0163] Fitting age using ridge regression on the CpG sites selected by the LASSO model. The model is

[0164]

[0165]

[0166] The learner parameters for calling glmnet using mlr3 are lrn("regr.cv_glmnet", alpha = 0,.lambda = 10^(seq(-8, 0, length.out = 100)), s = 'lambda.1se', standardize = F, intercept = T, type.measure ='mae'). Among them, alpha = 0 indicates fitting ridge regression.

[0167] The results show that the differences between genders are very small and can be ignored. Using limma:removeBatchEffect(), sva:ComBat() (parametric methods), and quantile normalization before coefficient estimation did not improve the model performance either.

[0168] This model is denoted as full model or uniAge - 605, which contains 605 CpG sites, and the estimated value of age is

[0169]

[0170] where are the coefficients estimated by ridge regression, is the j th methylation beta value of the CpG site, .

[0171] File list:

[0172] 1. uniAge / clean - meth - data.R: Script for reading methylation level data and phenotype data

[0173] 2. uniAge / uniAge.R: Script for chip data preprocessing, screening of age - related CpG sites, site selection for full model, and parameter estimation

[0174] 3. uniAge / model - uniAge.csv: Sites and coefficients of the full model

[0175] 4. uniAge / uniAge-on-train-datasets.csv: Fitted values of the full model for the training set ages

[0176] 5. uniAge / uniAge-on-test-datasets.csv: Predicted values of the full model for the chip test set ages

[0177] The 100-site model (uniAge-100): To facilitate the application of targeted methylation sequencing, 100 sites were selected from the 605 sites of the full model, which is called the reduced model or uniAge-100. In theory, the performance of the full model without the limitation of the number of sites should be better. Retaining the full model and the reduced model is to facilitate the comparison of the performance loss of the reduced model. For chip and whole-genome sequencing data, there is no need to be restricted by the number of CpG sites, so there is no need to use the reduced model. And the evaluation on the test set shows that the performance of the reduced model and the full model is almost the same.

[0178] Using the same process as the LASSO regression site screening and ridge regression coefficient estimation of the full model, for each CpG site of the full model continue to use LASSO site screening, select the appropriate lambda parameter so that the number of selected sites is exactly 100, and then use ridge regression to fit the age with the methylation beta values of these 100 sites to obtain the reduced model, uniAge-100.

[0179] File list:

[0180] 1. uniAge / reduced-model.R: Script for site selection and parameter estimation to obtain the reduced model from the full model

[0181] 2. uniAge / reduced-model-uniAge.csv: Sites and coefficients of the reduced model

[0182] 3. uniAge / reduced-uniAge-on-train-datasets.csv: Fitted values of the reduced model for the training set ages

[0183] 4. uniAge / reduced-uniAge-on-test-datasets.csv: Predicted values of the reduced model for the age of the chip test set

[0184] Nonlinear effect

[0185] Consider whether adding the nonlinear effect of site methylation on age can significantly improve the model performance. CpG sites may have different effects (slopes) on age at high, medium, and low methylation levels. For the methylation level variable of each site in the full model, construct piecewise linear B-splines using splines:bs(meth, degree = 1, knots = c(0.3, 0.7)), and fit the parameters using ridge regression, that is

[0186]

[0187] where is the methylation beta value of the i th sample at the j th CpG site, are the parameters to be estimated, , .

[0188] The piecewise linear model did not improve the model performance

[0189] Sequencing data correction: When applying the model to sequencing data, use quantile normalization to correct the distribution of methylation levels at 605 (100) sites in the sequencing data to be the same as the methylation distribution of these 605 (100) sites in the chip data training set. This correction can slightly improve the performance of the model on sequencing data. This method is ineffective on the test set of chip data

[0190] For the method of distribution correction, see the function quantile_align() in "models / DNAme-ages.R", and the average methylation level of the model sites on the chip training set can be found in "uniAge / mean-meth-training-set-[reduced-]uniAge.tsv"

[0191] As Figure 5 shown, a DNA methylation age prediction system based on the uniAge model provided by an embodiment of the present invention includes:

[0192] A CpG site screening module for screening out CpG sites that can be covered by both RRBS and methylation chips

[0193] An age prediction model training module for screening sites and training an age prediction model among the common CpG sites by using DNAme chip data of human peripheral blood.

[0194] A sequencing data correction module for correcting the DNAme level obtained from sequencing data, aligning it with the chip data, or separately fitting model coefficients on the sequencing data by using the screened CpG sites.

[0195] The specific application fields and related products of DNA methylation age cover multiple aspects such as health management, anti-aging, insurance, nutrition, and genetic testing. The following are some specific application fields and their representative products:

[0196] a) Health management and biological age monitoring

[0197] Application field: Through DNA methylation age measurement, an individual can obtain their biological age compared to their actual age. This information helps with personalized health management and early intervention.

[0198] Related products:

[0199] MyDNAge: Developed by Zymo Research, it provides an assessment of an individual's biological age by analyzing the DNA methylation status in blood or saliva samples.

[0200] Elysium Health Index: Provided by Elysium Health, the company uses DNA methylation technology to assess biological age and combines lifestyle suggestions to optimize health management.

[0201] Chronomics: This is a service that provides comprehensive DNA methylation age detection, including health status analysis and lifestyle suggestions, mainly for personalized health management.

[0202] b) Anti-aging and beauty

[0203] Application field: In the anti-aging field, DNA methylation age is used to evaluate the effects of various anti-aging therapies, including skin care products, supplements, and lifestyle interventions.

[0204] Related products:

[0205] EpiAging Clock by TruDiagnostic: The test provided by TruDiagnostic can measure multiple epigenetic clocks, helping users understand the biological aging status of the skin and the whole body, and providing verification for the effects of anti-aging products and treatments.

[0206] OneSkin: A company focusing on skin health, whose products combine DNA methylation data and claim to improve skin quality by reversing the biological age of the skin.

[0207] c) Personalized nutrition and dietary management

[0208] Application area: By analyzing DNA methylation age and other biomarkers, personalized dietary and nutritional supplement plans are developed to help slow down the aging process.

[0209] Related products:

[0210] GlycanAge: This product not only analyzes DNA methylation but also combines other biological markers to provide users with health and nutrition advice, aiming to delay biological aging.

[0211] InnerAge 2.0 by InsideTracker: InsideTracker's product provides users with personalized nutrition and lifestyle advice by combining methylation age and other blood indicators.

[0212] d) Insurance and risk assessment

[0213] Application area: Insurance companies use DNA methylation age as part of risk assessment to more accurately customize premiums and products.

[0214] Related products:

[0215] Revolution Insurance by YouSurance: This is a service that incorporates biological age into health insurance assessment and provides personalized insurance plans for customers through DNA methylation age detection.

[0216] e) Scientific research and medical research

[0217] Application area: DNA methylation age is used in research for early disease diagnosis, aging mechanisms, and the effectiveness of health interventions, playing an important role in basic scientific research.

[0218] Related products:

[0219] Illumina EPIC Array: A high-throughput sequencing platform for large-scale epigenetics research, capable of analyzing hundreds of thousands of CpG sites related to DNA methylation, widely used in biological age research.

[0220] Horvaths Clock Toolkits: Multiple research tools and software packages to help researchers analyze DNA methylation age in different tissues.

[0221] f) Consumer Genetic Testing and Lifestyle Optimization

[0222] Application Area: Genetic testing companies offer direct-to-consumer products that combine DNA methylation age analysis to help individuals optimize their lifestyles.

[0223] Related Products:

[0224] 23andMe + DNA Methylation Analysis: Although 23andMe's main business is genetic testing, they also offer methylation age analysis in cooperation with third-party platforms.

[0225] EpigenCare: This product provides personalized skin care advice by analyzing methylation patterns in skin samples.

[0226] g) Dynamic Aging Tracking and Health Optimization

[0227] Application Area: By regularly tracking changes in DNA methylation age, users can dynamically monitor their aging process and evaluate the effectiveness of interventions.

[0228] Related Products:

[0229] TrueAge by TruMe: Provides regular biological age monitoring services, combined with health advice, to help users optimize their lifestyles and adjust their health plans in real time.

[0230] EpiAging Monitoring by Novogenia: Provides long-term methylation age monitoring and trend analysis to support continuous health and anti-aging interventions.

[0231] Datasets for uniAge Performance Evaluation:

[0232] As Figure 6 shown, the training sets are GSE157131, GSE55763, GSE87648, GSE40279, GSE84727, GSE72680, GSE73103, GSE72775, GSE168739, nTotal = 7149

[0233] Distribution of the absolute value of the fitting value deviation of uniAge on the training set

[0234] Mean absolute deviation uniAge-605 uniAge-100 50% (median) 2.095132 2.606567 75% 3.629820 4.580433 90% 5.273679 6.755359 95% 6.603948 8.357867 mean 2.59369 3.27759

[0235] Chip Data Test Set

[0236] As Figure 7As shown, the test sets are GSE125105, GSE42861, GSE80417, GSE147221, GSE72773, GSE61496, nTotal = 3323

[0237] uniAge performs excellently on the microarray dataset:

[0238] As Figure 6 shown, the 100-locus uniAge model has good fitting accuracy on the microarray data training set (n = 7149), with a root mean square error (RMSE) of 3.44 (years) and a mean absolute error of 3.72 (years).

[0239] Distribution of the absolute value of the prediction value deviation of uniAge on the microarray data test set

[0240] Mean absolute deviation uniAge-605 uniAge-100 50% (median) 2.579847 3.023845 75% 4.498416 5.263006 90% 6.519476 7.748378 95% 8.365930 9.358484 mean 3.236782 3.72195

[0241] As Figure 7 shown, the RMSE of the 100-locus uniAge model on the independent microarray data test (n = 3323) is 4.81 (years), which is better than the prediction errors of existing DNA methylation ages ( Figure 8 ), where the RMSE of HorvathAge is 26.15 (years), the RMSE of HannumAge is 26.68 (years), the RMSE of Skip&BloodAge is 22.79 (years), and the systematic bias of ZhangAge is more significant.

[0242] uniAge also performs better than existing methods on the sequencing dataset:

[0243] The performance of uniAge and existing methods was evaluated using bisulfite sequencing RRBS and WGBS samples: RRBS: GSE85928; WGBS: GSE117593, nTotal = 102

[0244] As Figure 9 shown, for bisulfite sequencing samples (RRBS and WGBS, n = 102), the RMSE of uniAge is 15.72 (years), which is better than several existing commonly used methods, where the RMSE of HorvathAge is 21.65 (years), the RMSE of HannumAge is 23.86 (years), and the RMSE of Skip&BloodAge is 17.44 (years).

[0245] On the other hand, Figure 10The missing ratio of CpG sites required by different models in DNA methylation sequencing samples is shown. Through the preliminary screening of RRBS core sites, the site missing of the uniAge model is much lower than that of other models. In the evaluated RBBS samples, the CpG site missing ratio of the uniAge model is about 10% on average, while the average missing ratio of several existing methods is about 70%. Compared with existing methods, the missing ratio of uniAge has been significantly improved, which undoubtedly makes the prediction results of age more reliable.

[0246] Improvement work direction: Ridge model bias correction

[0247] LASSO, Ridge, and Elastic net are all (regularized) biased estimates, which result in larger age predictions for young people and smaller age predictions for older people. Although this bias is a compromise to reduce the prediction variance (minimize the prediction squared loss), sometimes the bias is too large and can be corrected, considering the de-biased estimate of Ridge regression (Zhang and Politis, 2021)

[0248]

[0249] and They are the classical and de-biased Ridge estimates respectively. Compare The bias of is smaller and is not an unbiased estimate. It is not clear which software package implements the de-biased Ridge estimate.

[0250] Correction of sequencing data

[0251] Judging from the current results, the technical deviation of sequencing data and chip data has a great impact on the model. The quantile normalization that has been tried is effective but not sufficient. For each CpG site, the beta value distribution of the chip training set of the site is used as the prior, the read count and beta value in the sequencing sample are used as the likelihood, and the posterior mean of the beta value is used as the correction, so that the beta value of the CpG site obtained by the sequencing data is closer to the chip data.

[0252] Fitting model coefficients separately on sequencing data: Re-estimate the coefficients of the linear model on the sequencing samples using the identified 605 or 100 CpG sites. Since the current sequencing data is relatively small (102 samples), consider an empirical Bayes linear model. The prior distribution of the model coefficients is determined by the microarray data. Specifically, randomly select 102 microarray training sets, fit a Ridge model to obtain the coefficients of the 605 (100) sites, repeat the random selection 1000 times to obtain a distribution for each coefficient, and use this as the prior distribution of the coefficients. Calculate the posterior distribution on the sequencing samples to obtain the model for the sequencing data. If there are approximately 500 sequencing samples, directly fit the model on the sequencing data.

[0253] Example 1: Predicting DNA methylation age using 450K microarray and RRBS data

[0254] 1. Sample preparation and data acquisition:

[0255] Collect human whole blood samples, obtain methylation microarray data using the Illumina HumanMethylation450K microarray, and simultaneously perform RRBS sequencing on the same samples.

[0256] Extract CpG site information from the 450K microarray data and RRBS data.

[0257] 2. Screening CpG sites:

[0258] According to Step 1, screen out the CpG sites that can be covered by both the 450K microarray and RRBS.

[0259] During the screening process, use the intersection of the covered sites of the 450K microarray and RRBS, and ensure that these sites are core sites on the microarray and have a high RRBS sequencing coverage rate.

[0260] Ensure that the methylation levels of the selected CpG sites have the smallest difference between the positive and negative strands and are located on autosomes.

[0261] 3. Model training:

[0262] Use the common CpG sites in the 450K microarray data obtained from peripheral blood samples, and train an age prediction model based on the methylation levels of these sites.

[0263] Adopt linear regression or other suitable machine learning algorithms to fit the model coefficients in the training set to maximize the prediction accuracy.

[0264] 4. Data correction and prediction:

[0265] Correct the DNA methylation levels of RRBS sequencing data to align them with 450K microarray data.

[0266] Use the model trained on 450K microarray data to predict age on RRBS data.

[0267] 5. Result verification:

[0268] Compare the predicted DNA methylation age with the actual age to evaluate the accuracy and stability of the model.

[0269] Repeat the experiment multiple times to ensure the reliability of the results.

[0270] This example is particularly applicable to scientific research experiments that require accurate prediction of the age of biological samples from DNA methylation data from different sources, and is applicable to human genome research and disease prediction.

[0271] Example 2: Predicting DNA methylation age using EPIC v1 microarray and RRBS data

[0272] 1. Sample preparation and data acquisition:

[0273] Perform DNA methylation microarray analysis of human whole blood samples using the Illumina MethylationEPIC v1 microarray (850K), and simultaneously perform RRBS sequencing on the same samples.

[0274] Extract relevant CpG site information from the EPIC microarray and RRBS sequencing data.

[0275] 2. Screening CpG sites:

[0276] According to Step 1 and Step 3, screen out the "reliable" CpG sites jointly covered by the EPIC microarray and RRBS.

[0277] During the screening process, use the intersection of the core sites of the EPIC microarray and the sites covered by RRBS sequencing, and ensure that these sites uniquely map to the reference genome UCSC hs1 (T2T) on the microarray.

[0278] Ensure that the selected sites have consistent methylation levels and are located on autosomes.

[0279] 3. Model training:

[0280] Use the CpG sites screened from the EPIC microarray data. Based on the methylation levels of these sites, train an age prediction model using the microarray data of whole blood samples.

[0281] Apply the uniAge model for fitting to generate the coefficients of the prediction model.

[0282] 4. Data correction and prediction:

[0283] Correct the RRBS data, align it with the EPIC chip data, or directly fit the model coefficients on the selected CpG sites.

[0284] Use the above-trained model to perform age prediction on the RRBS data.

[0285] 5. Result verification:

[0286] Compare the prediction results with the actual age of the samples to evaluate the accuracy of the model.

[0287] Compare the prediction results of different batches of samples to verify the stability of the model on different datasets.

[0288] This embodiment is particularly applicable to the research on DNA methylation age prediction of a large population. Especially in the case where there are both chip and sequencing data in the samples, it is more applicable to precision medicine research and biomarker discovery.

[0289] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made by those skilled in the art within the technical scope disclosed by the present invention and within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A method for predicting DNA methylation age based on the uniAge model, characterized in that, It includes the following steps: Step 1, screen out the CpG sites that can be covered by both RRBS and methylation chips simultaneously; Step 2, use the DNAme chip data of human peripheral blood to screen sites and train the prediction model of age in the common CpG sites; Step 3, correct the DNAme level obtained from the sequencing data, align it with the chip data, or separately fit the model coefficients on the sequencing data using the screened CpG sites; The data used includes the 450K and EPIC chip data of human whole blood samples, the sequencing data of human whole blood samples, the sequencing and chip data of the same biological samples, and human RRBS data; Primary screening of CpG sites: screen the common and "reliable" CpG sites of RRBS and DNAme chips, meeting the following conditions: Core sites of the chip: the intersection of the sites covered by 450K, 850K, and 935K arrays; The probe source sequence after base conversion is uniquely aligned to the end-to-end reference genome UCSChs1; Core sites of RRBS: the most easily covered by sequencing, RRBS uses MspI endonuclease, and the cleavage site is C^CGG; Consistently methylated CpG sites: the methylation levels of the positive and negative strands have the smallest difference; The intersection of chip sites and RRBS sites; Located on autosomes.

2. The DNA methylation age prediction method based on the uniAge model according to claim 1, wherein Training set and test set: Select the training set and test set according to the age distribution of all data sets, so that the training set has a certain number of samples in each age interval, and the training set and test set have a relatively consistent age distribution.

3. The DNA methylation age prediction method based on the uniAge model according to claim 1, wherein Preprocessing of chip data: Only retain the methylation values of the initially screened sites, 65,221 sites; Remove samples with missing age and samples with age < 1 or age > 120; Remove sites with a missing proportion > 10% in all samples; Remove samples with a site missing proportion > 10%; For each site, fill in the missing methylation values with random normal distribution values, and the mean and variance are calculated from the non-missing values of this site.

4. The DNA methylation age prediction method based on the uniAge model according to claim 1, characterized in that Screening of age-related CpG sites: Further screen age-related methylation sites from 65,221 initially screened sites, and let β ij be the methylation level of the j-th CpG site in the i-th sample; Calculate the interquartile range, IQR, of the methylation level at each locus j = q 0.75 (β ij ) - q 0.25 (β ij ); Calculate the correlation between methylation at each locus and age, Pearson correlation coefficient r j = cor Pearson (β ij , Age i ) and Spearman rank correlation coefficient ρ j = cor Spearman (β ij , Age i ); Definition Sort by score j Sort the CpG sites from largest to smallest by score, and select the 6,000 CpG sites with the highest scores to enter the following model.

5. The DNA methylation age prediction method based on the uniAge model according to claim 1, characterized in that, Use the LASSO model to screen CpG sites The input data is 7149 chip training set samples, including variables age, data set, gender, and methylation beta values of 6000 CpG sites. Use the R package glmnet to fit the LASSO model i = 1, 2,..., 7149, p = 6000, Age i ~dataset i +sex i +CpG i1 +…+CpG ip ; Use the machine learning framework mlr3 of R to call glmnet, and the learner parameters are lrn("regr.cv_glmnet", alpha = 1, lambda = 10 ^ (seq(-5, 0, length.out = 100)), s = 'lambda.1se', standardize = T, intercept = T, type.measure ='mae'), or directly use glmnet for training without using mlr3.

6. The DNA methylation age prediction method based on the uniAge model according to claim 1, wherein, Ridge regression age prediction model: Use ridge regression to fit age on the CpG sites selected by the LASSO model. The model is: The learner parameters for calling glmnet using mlr3 are lrn("regr.cv_glmnet", alpha = 0, lambda = 10 ^ (seq(-8, 0, length.out = 100)), s = 'lambda.1se', standardize = F, intercept = T, type.measure ='mae'), where alpha = 0 indicates fitting a ridge regression; This model is denoted as full model or uniAge - 605, which contains 605 CpG sites, and the estimated value of age is where is the coefficient of ridge regression estimation, β j is the methylation betavalue of the j-th CpG site, p f = 605.

7. The DNA methylation age prediction method based on the uniAge model according to claim 1, wherein When applying the model to the sequencing data, quantile normalization is used to correct the distribution of methylation levels at 605 sites in the sequencing data to be the same as the methylation distribution at these 605 sites in the microarray data training set. This correction can slightly improve the performance of the model on the sequencing data.

8. A DNA methylation age prediction system based on the uniAge model for implementing the DNA methylation age prediction method based on the uniAge model according to any one of claims 1 to 7, characterized in that, Including: CpG site screening module, which is used to screen out CpG sites that can be covered by both RRBS and methylation microarray; Age prediction model training module, which is used to screen sites and train the age prediction model in the common CpG sites using the DNAme microarray data of human peripheral blood; Sequencing data correction module, which is used to correct the DNAme level obtained from the sequencing data to align with the microarray data, or to separately fit the model coefficients on the sequencing data using the screened CpG sites.

Citation Information

Patent Citations

  • SE203169C1

  • Method for obtaining age of individual of Chinese population

    CN113373236A