Combinations of methylation biomarkers, prediction methods and devices for biological age prediction
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]有鉴于此,有必要提供一种用于生物年龄预测的甲基化标志物组合、生物年龄预测方法及装置,用以解决现有中生物年龄预测需要依赖细胞组成估算的技术问题
本发明提供的用于生物年龄预测的甲基化标志物组合中,通过223个特定CpG位点的标志物组合,该标志物组合具有年龄相关性显著和细胞组成独立性强的特点,仅需检测223个特定CpG位点的甲基化水平,即可直接计算得到待测样本的生物年龄,无需对待测样本进行细胞分型或细胞组成估算,降低了临床检测的操作复杂度和实施成本。
Smart Images

Figure CN122564128A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biomedical detection and epigenetics, specifically to a combination of methylation markers, a method for predicting biological age, and an apparatus for such prediction. Background Technology
[0002] Biological age is a comprehensive indicator reflecting the actual aging state of an organism's physiological functions. Compared with chronological age, biological age can more accurately assess an individual's health status, aging process, and disease risk, and has significant application value in precision medicine, anti-aging research, and clinical disease risk assessment. DNA methylation, as an important epigenetic modification, shows a regular change in its level with age. Peripheral blood mononuclear cells (PBMCs) have become a commonly used sample type for biological age prediction based on DNA methylation due to their convenient collection and relatively simple detection procedures.
[0003] In related technologies, DNA methylation data is used to assess an individual's age status. However, this type of method has the following problems in practical applications: First, there is a strong correlation between the methylation sites of interest and various immune cell compositions in PBMCs. This means that assessments using this method usually require additional cell typing or cell composition estimation steps to obtain reliable results, increasing the complexity and cost of the detection process. Second, existing methods consider only a single factor when screening methylation sites, typically relying solely on statistical significance. This results in a large number of sites with high redundancy, making it difficult to directly develop simplified detection schemes suitable for convenient clinical testing scenarios.
[0004] Therefore, there is a need for a new combination of methylation biomarkers and a new method for predicting biological age. This method should be able to predict directly based on the methylation level of the sample to be tested, without the need to estimate the cell composition of the sample. Furthermore, the combination of CpG biomarkers on which this method depends should have a clear composition and corresponding model parameters to facilitate clinical testing applications. Summary of the Invention
[0005] In view of this, it is necessary to provide a combination of methylation markers, a method and apparatus for predicting biological age, in order to solve the technical problem that existing biological age predictions rely on cellular composition estimation.
[0006] To address the aforementioned technical problems, in a first aspect, the present invention provides the application of a combination of methylation biomarkers in biological age prediction, wherein the combination of methylation biomarkers comprises: cg23744638、cg01477894、cg01188867、cg27030854、cg04193015、cg04644869、cg16966397、cg18448426、cg13331563、cg25253898、cg00387658、cg25490376、cg03752138、cg06210197、cg22493216、cg15109150、cg25253895、cg01763090、cg15728320、cg09143195、cg16545019、cg03424070、cg19663246、cg03732728、cg07158339、cg17767963、cg16983588、cg16966398、cg26514191、cg23510764、cg10221746、cg18452806、cg00842595、cg11935615、cg04538460、cg08466899、cg23124451、cg00533891、cg05150327、cg03996822、cg06272857、cg06669752、cg07822980、cg14209784、cg19727165、cg16767132、cg27333886、cg15736994、cg27430293、cg02554627、cg14295611、cg18881380、cg09171967、cg14144583、cg22798948、cg22549904、cg18622950、cg13604801、cg04367531、cg06839255、cg00876127、cg23734065、cg03467555、cg13612317、cg18684863、cg20692569、cg15148145、cg00593900、cg14155195、cg10348486、cg14253018、cg01949363、cg13734401、cg15676719、cg20067846、cg14996491、cg04462378、cg07485775、cg06027629、cg24709951、cg09372060、cg15882041、cg00557647、cg18468088、cg24919946、cg04478095、cg08469939、cg18881402、cg07841070、cg24806066、cg16087521、cg23405095、cg15583492、cg17804348、cg17017591、cg08029165、cg27603605、cg16582319、cg21757387、cg14162928、cg04104144、cg09349619、cg15017982、cg14036830、cg01635555、cg09918713、cg13207212、cg26014782、cg24369989、cg22008851、cg04140663、cg11957087、cg21379099、cg26060605、cg23476877、cg26897150、cg00688297、cg06852243、cg18329187、cg05666820、cg14166009、cg18815647、cg09715059、cg01933590、cg06893273、cg27016307、cg14643892、cg20669572、cg27659368、cg26413855、cg21571060、cg11925696、cg17508941、cg05542681、cg07717903、cg01070775、cg01408508、cg12948621、cg21570988、cg01358804、cg03731828、cg17179314、cg23008153、cg18084609、cg26332926、cg15470658、cg00446115、cg23737263、cg11651896、cg21899500、cg17364089、cg20858243、cg07884745、cg10665321、cg17364672、cg17816357、cg13504434、cg17163168、cg06208270、cg21213853、cg26698460、cg08529549、cg20002504、cg08667128、cg13149736、cg01410876、cg00464640、cg20818947、cg11868595、cg19986909、cg24939561、cg18788940、cg22524327、cg11229663、cg21870668、cg26676157、cg06259248、cg13221677、cg07215693、cg22932336、cg24013620、cg00004996、cg11429271, cg05260088, cg00593462, cg01873886, cg01682111, cg07949597, cg24195002, cg13619054, cg20164601, cg04561804, cg 05825244, cg18839637, cg10362475, cg14185918, cg23128634, cg07519847, cg05162523, cg10634619, cg02527881, cg16340422, cg209 77229, cg15380598, cg08046044, cg13185413, cg14880348, cg01127154, cg09570348, cg02601140, cg03025830, cg23114616, cg100122 73. cg13401893, cg12633154, cg24609819, cg26095158, cg02441474, cg19480198, cg06573787, cg02961654, cg04985582, cg03726693. ,
[0007] Secondly, the present invention provides the application of a reagent for detecting the methylation level of a combination of methylation biomarkers in the preparation of a kit for biological age prediction, wherein the combination of methylation biomarkers comprises: cg23744638、cg01477894、cg01188867、cg27030854、cg04193015、cg04644869、cg16966397、cg18448426、cg13331563、cg25253898、cg00387658、cg25490376、cg03752138、cg06210197、cg22493216、cg15109150、cg25253895、cg01763090、cg15728320、cg09143195、cg16545019、cg03424070、cg19663246、cg03732728、cg07158339、cg17767963、cg16983588、cg16966398、cg26514191、cg23510764、cg10221746、cg18452806、cg00842595、cg11935615、cg04538460、cg08466899、cg23124451、cg00533891、cg05150327、cg03996822、cg06272857、cg06669752、cg07822980、cg14209784、cg19727165、cg16767132、cg27333886、cg15736994、cg27430293、cg02554627、cg14295611、cg18881380、cg09171967、cg14144583、cg22798948、cg22549904、cg18622950、cg13604801、cg04367531、cg06839255、cg00876127、cg23734065、cg03467555、cg13612317、cg18684863、cg20692569、cg15148145、cg00593900、cg14155195、cg10348486、cg14253018、cg01949363、cg13734401、cg15676719、cg20067846、cg14996491、cg04462378、cg07485775、cg06027629、cg24709951、cg09372060、cg15882041、cg00557647、cg18468088、cg24919946、cg04478095、cg08469939、cg18881402、cg07841070、cg24806066、cg16087521、cg23405095、cg15583492、cg17804348、cg17017591、cg08029165、cg27603605、cg16582319、cg21757387、cg14162928、cg04104144、cg09349619、cg15017982、cg14036830、cg01635555、cg09918713、cg13207212、cg26014782、cg24369989、cg22008851、cg04140663、cg11957087、cg21379099、cg26060605、cg23476877、cg26897150、cg00688297、cg06852243、cg18329187、cg05666820、cg14166009、cg18815647、cg09715059、cg01933590、cg06893273、cg27016307、cg14643892、cg20669572、cg27659368、cg26413855、cg21571060、cg11925696、cg17508941、cg05542681、cg07717903、cg01070775、cg01408508、cg12948621、cg21570988、cg01358804、cg03731828、cg17179314、cg23008153、cg18084609、cg26332926、cg15470658、cg00446115、cg23737263、cg11651896、cg21899500、cg17364089、cg20858243、cg07884745、cg10665321、cg17364672、cg17816357、cg13504434、cg17163168、cg06208270、cg21213853、cg26698460、cg08529549、cg20002504、cg08667128、cg13149736、cg01410876、cg00464640、cg20818947、cg11868595、cg19986909、cg24939561、cg18788940、cg22524327、cg11229663、cg21870668、cg26676157、cg06259248、cg13221677、cg07215693、cg22932336、cg24013620、cg00004996、cg11429271, cg05260088, cg00593462, cg01873886, cg01682111, cg07949597, cg24195002, cg13619054, cg20164601, cg04561804, cg 05825244, cg18839637, cg10362475, cg14185918, cg23128634, cg07519847, cg05162523, cg10634619, cg02527881, cg16340422, cg209 77229, cg15380598, cg08046044, cg13185413, cg14880348, cg01127154, cg09570348, cg02601140, cg03025830, cg23114616, cg100122 73. cg13401893, cg12633154, cg24609819, cg26095158, cg02441474, cg19480198, cg06573787, cg02961654, cg04985582, cg03726693. ,
[0008] Understandably, reagents for detecting the methylation level of a combination of methylation biomarkers can be conventionally selected based on actual usage needs. For example, they can be reagents from methylation chip detection methods, bisulfite sequencing methods, methylation-specific PCR methods, pyrosequencing methods, or mass spectrometry methods, as long as they can efficiently obtain the methylation level of the combination of methylation biomarkers.
[0009] Thirdly, the present invention provides a method for predicting biological age, comprising: Obtain the methylation level of each CpG site in the above methylation marker combination in the sample to be tested, wherein the methylation level is a value between 0 and 1; The biological age of the sample to be tested is predicted by summing the methylation levels of each CpG site and the corresponding regression coefficients, and by using the baseline biological age value. The biological age baseline value and the regression coefficients corresponding to each CpG site are determined by elastic network regression modeling based on the methylation level of the 223 CpG sites in the PBMC sample and the actual age of each PBMC sample; the biological age baseline value represents the biological age corresponding to when the methylation level of all 223 CpG sites is 0.
[0010] In one possible implementation, the steps for determining the baseline biological age value and the regression coefficients corresponding to each CpG site include: EPIC V2 methylation chip detection data of the PBMC samples were obtained and preprocessed using the QCDPB algorithm in the SeSAMe software package; Epigenome-wide association analysis was performed on the preprocessed data, with age as the independent variable and sex and cell composition as covariates for correction, to screen for candidate CpG sites. The selected candidate CpG sites were filtered for cell composition stability to obtain the target CpG sites; On the PBMC sample, the methylation level of the target CpG site was used as the input variable and the actual age was used as the output variable. Elastic network regression modeling was adopted, and the hyperparameters were optimized by cross-validation to obtain the baseline biological age and the regression coefficients of each CpG site.
[0011] In one possible implementation, the cell composition stability filtering of the selected candidate CpG sites includes: Calculate the maximum correlation coefficient between the methylation level of each candidate CpG site and the proportion of various PBMC cell types; Candidate CpG sites with a maximum correlation coefficient less than a third preset threshold are identified as target CpG sites.
[0012] In one possible implementation, the various PBMC cell types include naive CD4+ T cells, memory CD4+ T cells, naive CD8+ T cells, memory CD8+ T cells, regulatory T cells, memory B cells, naive B cells, natural killer cells, and monocytes.
[0013] In one possible implementation, the step of performing a full epigenome association analysis on the preprocessed data, using age as the independent variable and adjusting for sex and cellular composition as covariates, includes: The methylation level of each CpG site in the preprocessed data is converted into a logarithmic M value; Using the transformed M value as the dependent variable, age as the independent variable, and the proportion of sex and multiple PBMC cell types as covariates, a linear model was constructed and fitted to obtain the fitting results. The fitting results were adjusted using an empirical Bayesian algorithm; Based on the adjusted fitting results, CpG sites with an error detection rate less than the first preset threshold and an absolute value of methylation difference greater than the second preset threshold are selected as candidate CpG sites.
[0014] In one possible implementation, the absolute value of the methylation difference is the mean difference in methylation levels between the older and younger groups.
[0015] In one possible implementation, the preprocessing using the QCDPB algorithm from the SeSAMe software package includes: Quality masking was applied to the EPIC V2 methylation chip detection data to shield against cross-hybridization and non-specific probes. Channel attribution is inferred for the quality-masked data, and the channel attribution of each probe is marked according to the inference results; Nonlinear correction for dye bias is performed on the data after the channel assignment marker; The data after dye bias correction is masked by p-value, and probes that do not meet the significance requirement are marked as missing values based on the detection p-value of each probe. NOOB background correction is performed on the data masked by the p-value.
[0016] In one possible implementation, the sample to be tested is a human peripheral blood mononuclear cell sample.
[0017] In one possible implementation, obtaining the methylation level of each CpG site in the methylation biomarker combination in the sample to be tested includes: Genomic DNA was extracted from the sample to be tested; The genomic DNA was detected using a methylation chip. The methylation level of each CpG site among the 223 CpG sites was determined based on the detection results.
[0018] Fourthly, the present invention also provides a biological age prediction device, comprising: The acquisition unit is used to acquire the methylation level of each CpG site in the above-mentioned methylation marker combination in the sample to be tested, wherein the methylation level is a value between 0 and 1; The prediction unit is used to predict the biological age of the sample to be tested based on the sum of the methylation levels of each CpG site multiplied by the corresponding regression coefficients and the baseline biological age value. The biological age baseline value and the regression coefficients corresponding to each CpG site are determined by elastic network regression modeling based on the methylation level of the 223 CpG sites in the PBMC sample and the actual age of each PBMC sample; the biological age baseline value represents the biological age corresponding to when the methylation level of all 223 CpG sites is 0.
[0019] The beneficial effects of this invention are: The methylation biomarker combination for biological age prediction provided by this invention uses a combination of 223 specific CpG sites. This biomarker combination has the characteristics of significant age correlation and strong cell composition independence. Only the methylation level of 223 specific CpG sites needs to be detected to directly calculate the biological age of the sample to be tested, without the need for cell typing or cell composition estimation of the sample to be tested, thus reducing the operational complexity and implementation cost of clinical testing. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A schematic flowchart of an embodiment of the biological age prediction method provided by the present invention; Figure 2 A schematic diagram illustrating the prediction accuracy of the biological age prediction model provided by this invention; Figure 3 This is a schematic diagram of the biological age prediction device provided by the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0023] In the description of the embodiments of the present invention, unless otherwise stated, "a plurality of" means two or more.
[0024] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.
[0025] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0026] In this invention, the dataset of 500 PBMC samples used is derived from the NCBI GEO database (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE246337).
[0027] This invention provides a combination of methylation biomarkers, a method for predicting biological age, and an apparatus for such prediction, which are described below.
[0028] The subject executing the biological age prediction method in this application embodiment can be the biological age prediction device provided in this application embodiment, or a network device that integrates the biological age prediction device. This application embodiment does not specifically limit this.
[0029] Figure 1 This is a schematic flowchart of an embodiment of the biological age prediction method provided by the present invention, as shown below. Figure 1 As shown, biological age prediction methods include: S101. Obtain the methylation level of each CpG site in a set of CpG methylation markers in the sample to be tested. The marker set consists of 223 CpG sites as shown in Table 1 of the specification. The methylation level is a value between 0 and 1. S102. Based on the sum of the products of the methylation levels of each CpG site and the corresponding regression coefficients, and the baseline biological age value, the biological age of the sample to be tested is predicted; wherein, the baseline biological age value and the regression coefficients corresponding to each CpG site are determined by elastic network regression modeling of the PBMC samples based on the methylation levels of the 223 CpG sites in the PBMC samples and the actual age of each PBMC sample; the baseline biological age value represents the biological age corresponding to when the methylation levels of the 223 CpG sites are all 0.
[0030] The biological age baseline is a constant, which physically represents the biological age output by the model when the methylation level of all 223 CpG sites is 0. This biological age baseline is determined synchronously with each regression coefficient during the elastic network regression modeling process.
[0031] The CpG methylation marker consists of 223 specific CpG sites. The specific CpG site ID (probe ID), the chromosome where each CpG site is located, the position on the chromosome, the corresponding gene name, and the regression coefficient of each CpG site for biological age prediction are shown in Table 1 of the instruction manual.
[0032] Table 1: Information on 223 CpG methylation markers
[0033] In Table 1, CpG sites are identified using probe IDs from the Illumina methylation chip; for example, 'cg23744638' represents a specific CpG site. Those skilled in the art can uniquely determine the corresponding CpG site and its location in the genome based on this probe ID.
[0034] The 223 CpG sites in Table 1 were selected from methylation data of human peripheral blood mononuclear cells. The selection process comprehensively considered the correlation between each CpG site and age, the correlation with immune cell composition, and the regularization selection results of the model. The sites that were ultimately retained were all sites with non-zero regression coefficients in the elastic network regression model.
[0035] This CpG methylation marker is used to predict biological age. The methylation level of each CpG site is multiplied by the regression coefficient corresponding to that site, and all products are summed to obtain a cumulative sum. This cumulative sum is then added to the baseline biological age value to obtain the biological age of the sample.
[0036] It should be understood that the biological age baseline value is a constant, and its physical meaning is the biological age predicted by the model when all 223 CpG sites are in a completely unmethylated state (i.e., the methylation level is 0). This biological age baseline value is determined synchronously with the regression coefficients of each CpG site during the elastic network regression modeling process, reflecting the overall level of the age distribution of the samples in the training set.
[0037] Specifically, after obtaining the methylation levels of the aforementioned 223 CpG sites in the sample to be tested, the biological age of the sample to be tested is calculated according to the following formula: Biological age = β0 + Σ j ·β j × beta j Where j ranges from 1 to 223, β0 is the baseline biological age, and βj Let be the regression coefficient corresponding to the j-th CpG site, and beta j β represents the methylation level of the j-th CpG site. j The specific values are shown in Table 1 of the instruction manual.
[0038] methylation level beta j The value ranges from 0 to 1, where 0 represents that the CpG site is completely unmethylated and 1 represents that it is fully methylated. The regression coefficient β for each CpG site... j This reflects the direction and extent of the influence of changes in methylation levels at this site on biological age. A positive regression coefficient indicates that the higher the methylation level, the older the predicted age, while a negative regression coefficient indicates that the higher the methylation level, the younger the predicted age.
[0039] This CpG methylation marker is derived from EPIC V2 methylation chip detection data of human peripheral blood mononuclear cell samples. It is a combination of age-related and cell composition-independent sites obtained from approximately 750,000 CpG sites through multi-dimensional screening.
[0040] Understandably, the CpG methylation biomarker provided in this embodiment contains 223 specific CpG sites. This biomarker combination has the characteristics of significant age correlation and strong cell composition independence. Only the methylation level of the 223 specific CpG sites needs to be detected to directly calculate the biological age of the sample to be tested. There is no need to perform cell typing or cell composition estimation on the sample to be tested, which significantly reduces the operational complexity and implementation cost of clinical testing.
[0041] In summary, the biological age prediction method provided by this invention uses a combination of 223 specific CpG site markers. This combination of markers has the characteristics of significant age correlation and strong cell composition independence. Only the methylation level of the 223 specific CpG sites needs to be detected to directly calculate the biological age of the sample to be tested. There is no need to perform cell typing or cell composition estimation on the sample to be tested, which reduces the operational complexity and implementation cost of clinical testing.
[0042] In some embodiments of the present invention, the steps for determining the biological age baseline value and the regression coefficients corresponding to each CpG site include: acquiring EPIC V2 methylation chip detection data of the PBMC sample and preprocessing it using the QCDPB algorithm in the SeSAMe software package; performing a whole epigenome association analysis on the preprocessed data, using age as the independent variable and sex and cell composition as covariates for correction, and screening out candidate CpG sites; filtering the screened candidate CpG sites for cell composition stability to obtain target CpG sites; and on the PBMC sample, using the methylation level of the target CpG site as the input variable and actual age as the output variable, using elastic network regression modeling, and optimizing hyperparameters through cross-validation to obtain the biological age baseline value and the regression coefficients of each CpG site.
[0043] The QCDPB algorithm in the SeSAMe software package refers to a preprocessing workflow optimized for EPIC V2 methylated chip data, which includes five sequentially executed steps: quality masking, channel inference, dye bias correction, p-value masking, and NOOB background correction.
[0044] Epigenome-wide association analysis (EPA) is a statistical analysis method that systematically examines the association between methylation levels at various CpG sites and age across the entire genome, using age as the independent variable and sex and cellular composition as covariates. During the analysis, methylation levels are typically converted to M values, calculated using the formula M = log2(beta / (1-beta)).
[0045] Cell composition stability filtering refers to the process of calculating the maximum correlation coefficient between the methylation level of each CpG site and the proportion of various PBMC cell types, and then screening CpG sites based on the magnitude of this correlation coefficient to remove sites that are highly correlated with cell composition.
[0046] Elastic regression is a linear regression method that combines L1 and L2 regularization. L1 regularization shrinks the regression coefficients of some features to zero, achieving feature selection; L2 regularization simultaneously shrinks the coefficients of related features, improving model stability. In the objective function of elastic regression, the alpha parameter controls the weight distribution between L1 and L2 regularization, and the lambda parameter controls the overall strength of regularization.
[0047] The objective function of the elastic network regression model is: min_β (1 / n)Σ( i y i )² + λ[α·Σ|β j | + (1 α) / 2·Σβ j ²]; Where n is the number of samples in the training set. i Let y be the predicted age of the i-th sample. i Let λ be the actual age of the i-th sample, and α be the hyperparameters to be optimized. α is used to control the L1 regularization term Σ|β. j | and L2 regularization term Σβ j The weights are distributed between α and α². The closer α is to 1, the more the model tends to be Lasso regression; the closer α is to 0, the more the model tends to be Ridge regression. λ controls the overall strength of regularization; the larger the value of λ, the stronger the regularization effect.
[0048] Specifically, EPIC V2 methylation chip detection data of PBMC samples were acquired and preprocessed using the QCDPB algorithm in the SeSAMe software package. This preprocessing was performed sequentially: quality masking of the raw data to exclude cross-hybridization probes and non-specific probes; channel assignment inference of the quality-masked data to determine the channel assignment of each probe; nonlinear correction for dye bias to eliminate systematic errors between different fluorescence detection channels; p-value masking, marking probes that do not meet the significance requirement as missing values based on the detection p-value; and finally, NOOB background correction to subtract background signals.
[0049] First, the methylation level of each CpG site is converted into an M value. Then, a linear model is constructed using the M value as the dependent variable for fitting. The standard error is adjusted using the empirical Bayesian method, and sites that meet the statistical significance requirements are selected based on the fitting results.
[0050] The maximum correlation coefficient between the methylation level of each candidate CpG site and the proportion of various PBMC cell types is calculated and denoted as max_cor. Sites with max_cor less than a preset threshold are identified as target CpG sites. In this embodiment, the preset threshold is preferably 0.40, and the corresponding coefficient of determination r² is 0.16, indicating that the portion of methylation variation at the retained sites that can be explained by cell composition does not exceed 16%.
[0051] Stratified sampling by age was used to divide 480 samples into training and test sets. Five-fold cross-validation with five replicates was performed on the training set, and alpha=0.1 and lambda=5.115 were determined as the optimal hyperparameters. Using these hyperparameters, an elastic network regression model was fitted to all training samples, yielding 223 CpG sites with non-zero regression coefficients and their corresponding regression coefficients, as well as the biological age baseline.
[0052] Understandably, this embodiment, through the sequential combination of whole epigenome association analysis, cell composition stability filtering, and elastic network regression modeling, achieves the screening of core site combinations closely related to age and lowly correlated with cell composition from hundreds of thousands of CpG sites. Specifically, the cell composition stability filtering step actively eliminates sites highly correlated with the proportion of immune cells, ensuring that the final model does not require cell composition estimation of the test samples in practical applications. Cross-validation is used to systematically optimize the hyperparameters of the elastic network regression, improving the model's prediction accuracy and generalization ability.
[0053] In one specific implementation, as shown in Table 2, the dataset partitioning table uses stratified sampling by age (5 quantile intervals, each interval divided by 80 / 20) to divide the 480 samples into a training set and a test set, ensuring that the age distribution of the two subsets is highly consistent and avoiding age bias caused by random partitioning.
[0054] The 480 PBMC samples were stratified by age and divided into training and test sets. Specifically, the samples were divided into five age ranges based on the quintiles of age, with 80% and 20% of each range randomly assigned to the training and test sets, respectively. The training set contained 385 samples, ranging in age from 18.1 to 89.0 years, with a median age of 53.6 years; the test set contained 95 samples, ranging in age from 18.6 to 89.0 years, with a median age of 54.5 years. The test set was strictly sealed after partitioning and used only once during the final model evaluation to ensure the unbiasedness of the evaluation results.
[0055] Table 2: Dataset Partitioning Table
[0056] Table 3 shows the results of the alpha search cross-validation. For the 10 candidate alpha values ranging from 0.1 to 1.0, 5 repetitions × 10-fold cross-validation were performed. The cross-validation results show that when alpha is 0.1, the mean absolute error is 4.541 years, the standard deviation is 0.039 years, and the number of selected features is 221. As the alpha value increases, the mean absolute error monotonically increases, while the number of selected features gradually decreases. Based on these results, the optimal alpha value was determined to be 0.1. This value is close to Ridge regression and is suitable for handling the high correlation between adjacent CpG sites in methylation data.
[0057] Table 3: Alpha Search Cross-Validation Results
[0058] As shown in Table 4, the alpha search table is used for 10 candidate values from alpha = 0.1 to 1.0. For each value, 5 repetitions × 10-fold cross-validation (50 CVs in total) are performed in the training set. The mean and standard deviation of the MAE corresponding to lambda.1se are used as evaluation metrics.
[0059] Table 4: Alpha Search Table
[0060] The CV MAE increases monotonically with increasing alpha, with alpha=0.1 (close to Ridge) showing the best performance. This result is consistent with the characteristics of methylation data: adjacent CpG sites are highly correlated, and Ridge-type regularization can synchronously shrink the coefficients of correlated features, which is more reasonable than randomly discarding one of them using Lasso. Five more repetitions of 10-fold CV were performed on the training set with the optimal alpha=0.1, and the lambda.1se of each iteration was recorded. The geometric mean was taken as the final lambda (the geometric mean is more symmetric and robust to the log-scale lambda). The lambda.1se values for each iteration are 3.183, 6.699, 5.309, 5.309, and 5.826. These five lambda.1se values are derived from five independent, repeated cross-validations. Specifically, after determining the optimal alpha = 0.1, this alpha value is fixed, and five more repetitions of 10-fold cross-validation are performed on the training set. Each repetition runs the complete 10-fold cross-validation process independently, and each cross-validation selects a lambda.1se value based on the criterion of minimum error plus one standard deviation (i.e., the one-standard-error rule). The five repetitions produce five lambda.1se values, and the geometric mean of these five values is taken as the final regularization strength parameter, lambda. The geometric mean is used instead of the arithmetic mean because lambda parameters are typically analyzed on a logarithmic scale, and the geometric mean is symmetric and robust to lambda values on a logarithmic scale.
[0061] Final lambda (geometric mean): 5.115 Corresponding CV MAE: 4.54 ± 0.08 years.
[0062] Table 5 shows the data for the linear model. The linear model is constructed using the limma framework, and the design matrix is as follows: M ~ age + gender + CD4Tnv + CD4Tmem + CD8Tnv + CD8Tmem + Treg + Bnv +NK + Mono Table 5: Data Table for Linear Model
[0063] Where M represents the M value after methylation level conversion at each CpG site, age represents the actual age of the individuals from which the sample was obtained, and gender represents sex. CD4Tnv represents the proportion of initial CD4+ T cells, CD4Tmem represents the proportion of memory CD4+ T cells, CD8Tnv represents the proportion of initial CD8+ T cells, CD8Tmem represents the proportion of memory CD8+ T cells, Treg represents the proportion of regulatory T cells, Bnv represents the proportion of initial B cells, NK represents the proportion of natural killer cells, and Mono represents the proportion of monocytes. The sum of the proportions of the above cells is close to 1. To avoid complete collinearity with the intercept term of the model, the proportion of memory B cells was not included in the design matrix of this model. Empirical Bayesian adjustment was performed using eBayes(), with the following parameter settings: trend=TRUE (modeling the mean-variance trend, suitable for methylated data); robust=TRUE (reducing the weight of outliers to reduce the impact of outliers). Multiple test correction was performed using the Benjamini-Hochberg (BH) method to control FDR.
[0064] In some embodiments of the present invention, the step of filtering the selected candidate CpG sites for cell composition stability includes: calculating the maximum correlation coefficient between the methylation level of each candidate CpG site and the proportion of multiple PBMC cell types; and determining the candidate CpG sites whose maximum correlation coefficient is less than a third preset threshold as the target CpG sites.
[0065] The maximum correlation coefficient refers to the maximum absolute value of the correlation coefficient obtained after Pearson correlation analysis of the methylation level of each CpG site with the proportion of multiple PBMC cell types.
[0066] The third preset threshold is a pre-set value used to determine whether the correlation between CpG sites and cell composition is within an acceptable range. The third preset threshold is preferably 0.40.
[0067] Specifically, for candidate CpG sites screened through whole-epigenome association analysis, the Pearson correlation coefficient between the methylation level of each site and the proportion of various PBMC cell types is calculated one by one, resulting in a set of correlation coefficients. The maximum absolute value of this set of correlation coefficients is taken as the maximum correlation coefficient for that site. This maximum correlation coefficient is compared with a third preset threshold. If the maximum correlation coefficient is less than the third preset threshold, the candidate CpG site is identified as the target CpG site; if the maximum correlation coefficient is greater than or equal to the third preset threshold, the candidate CpG site is eliminated.
[0068] Table 6 shows the max_cor quantile distribution. The core constraint of the production environment is that PBMC cell composition does not need to be estimated when predicting new samples (additional cell typing in clinical practice increases complexity and cost). Therefore, CpGs with low correlation to cell composition are actively screened from EWAS significant sites as modeling features.
[0069] For each significant EWAS CpG, calculate its beta value and the Pearson correlation coefficient between the proportions of the nine cell types. Take the maximum absolute value of the correlation coefficient for each cell type (max_cor) as the cell composition dependence index of that CpG.
[0070] Table 6: Quantile Distribution of max_cor
[0071] After calculating the maximum cor values for each CpG site and the proportions of the nine PBMC cell types, their quantile distributions were statistically analyzed. The results showed that the 25th percentile of maximum cor was 0.332, the 50th percentile was 0.401, the 75th percentile was 0.482, the 90th percentile was 0.554, and the 95th percentile was 0.595. Based on these distributions, a third preset threshold was set to 0.40. This threshold is close to the 50th percentile, with a corresponding coefficient of determination r² of 0.16, indicating that the portion of methylation variations at the retained CpG sites that can be explained by cellular composition does not exceed 16%.
[0072] Table 7 shows the feature selection table. After initial screening using whole epigenome association analysis, a total of 4,854 candidate CpG loci were obtained. Cellular composition stability filtering was performed on these 4,854 candidate CpG loci, using a max_cor < 0.40 criterion. A total of 2,409 CpG loci were retained, accounting for 49.6% of the total number of candidate loci before filtering. Finally, through elastic network regression modeling and hyperparameter optimization, 223 CpG loci with non-zero regression coefficients were selected from the 2,409 candidate loci as the final core CpG biomarker combination.
[0073] Table 7: Feature Filtering Table
[0074] Understandably, this embodiment eliminates CpG sites highly correlated with PBMC cell composition by setting a screening condition where the maximum correlation coefficient is less than a third preset threshold. The coefficient of determination corresponding to the third preset threshold is 0.16, meaning that of the methylation variations of the retained CpG sites, the portion that can be explained by cell composition does not exceed 16%, and cell composition-independent signals dominate. The target CpG sites obtained in this way can be directly used for biological age prediction without the need for additional estimation of sample cell composition in clinical applications.
[0075] In some embodiments of the present invention, the various PBMC cell types include naive CD4+ T cells, memory CD4+ T cells, naive CD8+ T cells, memory CD8+ T cells, regulatory T cells, memory B cells, naive B cells, natural killer cells, and monocytes.
[0076] Among them, multiple PBMC cell types refer to different immune cell subsets isolated or identified from human peripheral blood mononuclear cells, and each subset differs in immune response and epigenetic characteristics.
[0077] Naïve CD4+ T cells are precursors of helper T cells that have not yet encountered the antigen, while memory CD4+ T cells are helper T cells that have encountered the antigen and formed immune memory. Naïve CD8+ T cells are precursors of cytotoxic T cells that have not yet encountered the antigen, while memory CD8+ T cells are cytotoxic T cells that have encountered the antigen and formed immune memory. Regulatory T cells are a subset of T cells with immunosuppressive functions.
[0078] Memory B cells are B cells that have been exposed to antigens and formed immune memory, while naive B cells are B cell precursors that have not been exposed to antigens. Natural killer cells are lymphocytes involved in innate immunity, and monocytes are mononuclear phagocytes involved in innate immunity and antigen presentation.
[0079] Specifically, in this embodiment, when performing whole epigenome association analysis, the proportions of the aforementioned nine PBMC cell types were used as covariates in the linear model. Using the EpiDISH software package, based on the reference map cent12CT.m, the proportions of the nine cell types in each sample were estimated using the RPC algorithm, and the proportions of each cell type were added as correction factors to the EWAS model.
[0080] Understandably, by incorporating nine key PBMC immune cell subsets into covariate correction, this embodiment can effectively separate the age effect from the cell composition effect, avoid false positive screening of age-related CpG sites due to differences in the proportion of immune cells between samples, thereby improving the cell composition independence of the screened markers and the generalization ability of the prediction model.
[0081] Table 8 shows nine PBMC cell types.
[0082] Table 8: Nine PBMC cell types
[0083] In some embodiments of the present invention, the step of performing a full epigenome association analysis on the preprocessed data, using age as the independent variable and sex and cell composition as covariates for correction, includes: converting the methylation level of each CpG site in the preprocessed data into a logarithmic M value; using the converted M value as the dependent variable, age as the independent variable, and the proportion of sex and multiple PBMC cell types as covariates, constructing a linear model for fitting, and obtaining the fitting result; adjusting the fitting result using an empirical Bayesian algorithm; and, based on the adjusted fitting result, selecting CpG sites with an error detection rate less than a first preset threshold and an absolute value of methylation difference greater than a second preset threshold as candidate CpG sites.
[0084] In some embodiments of the present invention, the absolute value of the methylation difference is the mean difference in methylation levels between the older and younger groups.
[0085] Among them, whole epigenome association analysis refers to the analytical method of systematically screening CpG sites that are significantly associated with the target phenotype (age) across the entire genome.
[0086] The M value is obtained by converting the methylation beta value using the formula M = log2(beta / (1-beta)). This conversion can make the data more consistent with the homogeneity of variance assumption of the linear model.
[0087] Covariates are factors other than independent variables that may affect the dependent variable in regression analysis. Including them in the model can eliminate the interference of these factors.
[0088] Empirical Bayesian algorithms are statistical methods that stabilize the variance estimate of a single gene by borrowing information from all genes, and are suitable for high-dimensional data analysis. The false positive rate refers to the proportion of false positives out of all positive results, and is used to control for Type I errors in multiple hypothesis testing.
[0089] Δbeta refers to the mean difference in methylation levels between the older and younger age groups. The formula is Δbeta = mean(beta_older group) - mean(beta_younger group). A positive value indicates that methylation increases with age, while a negative value indicates that methylation decreases with age.
[0090] Specifically, after performing SeSAMe QCDPB preprocessing on the EPIC V2 chip data, the methylation beta value of each CpG site is converted to an M value according to M = log2(beta / (1-beta)). Using the converted M value as the dependent variable, age as the independent variable, and the proportions of sex and various PBMC cell types as covariates, a linear model is constructed for fitting, yielding a fitting result containing statistics, p-values, and logFC information for each CpG site. Further, an empirical Bayesian algorithm is used to adjust the standard error of the fitting result to stabilize statistical inference under small sample conditions. Based on the adjusted fitting result, the false discovery rate is controlled using the Benjamini-Hochberg method, and CpG sites with a false discovery rate less than a first preset threshold and an absolute methylation difference Δbeta greater than a second preset threshold are selected as candidate CpG sites. In this embodiment, the first preset threshold is preferably 0.05, and the second preset threshold is preferably 0.05.
[0091] Understandably, this embodiment converts methylation beta values into M values to ensure the data meets the application conditions of a linear model; it eliminates the interference of sex and the proportion of multiple PBMC cell types as covariates in the model; it improves the stability of statistical inference under small sample conditions by adjusting the standard error using an empirical Bayesian algorithm; and it ensures that the selected candidate CpG sites are both statistically significant and have substantial biological changes by simultaneously controlling the false discovery rate and effect size threshold.
[0092] In one specific implementation, as shown in Table 9, the EWAS screening results are presented. The screening criteria and results employ a dual filtering standard, taking into account both statistical significance and biological meaning. Statistical significance is defined as an FDR-corrected p-value (adj.P.Val) < 0.05; the effect size is |Δbeta| > 0.05 (beta value difference > 5%).
[0093] Table 9: EWAS Screening Results
[0094] The EWAS screening results table, using a dual criterion of a false discovery rate of less than 0.05 and an absolute methylation difference Δbeta greater than 0.05, yielded 4,854 candidate CpG sites significantly associated with age from approximately 751,238 CpG sites. Sites with increasing methylation levels with age were categorized as hypertypes, while sites with decreasing methylation levels with age were categorized as hypotypes.
[0095] Table 10 shows representative CpG site examples, specifically examples of the most significant age-related CpG sites (sorted by original p-values): Among the 4,854 candidate CpG sites screened, some sites showed extremely significant age correlation. For example, the Δbeta of site cg16867657 is +0.22, with a very small adjusted p-value, indicating that the methylation level of this site increases significantly with age; the Δbeta of site cg10501210 is -0.36, indicating that the methylation level of this site decreases significantly with age. The above examples are used to illustrate the effectiveness of the screening method in this embodiment, but the biological age prediction method in this embodiment does not rely solely on the above example sites, but rather on all 223 CpG sites finally screened and determined.
[0096] Table 10: Examples of Representative CpG Sites
[0097] In some embodiments of the present invention, the preprocessing using the QCDPB algorithm in the SeSAMe software package includes: quality masking of the EPIC V2 methylation chip detection data to mask cross-hybridization and non-specific probes; channel attribution inference of the quality-masked data, and labeling the channel attribution of each probe according to the inference results; nonlinear correction of dye bias for the channel-attributed data; p-value masking of the dye bias-corrected data, and marking probes that do not meet the significance requirements as missing values according to the detection p-values of each probe; and NOOB background correction for the p-value-masked data.
[0098] Quality masking refers to identifying and masking probes that may cross-hybridize with other genomic regions or have non-specific binding based on the probe's design quality indicators, in order to avoid these probes introducing errors into subsequent analyses.
[0099] Channel assignment inference refers to determining the fluorescence detection channel to which each probe belongs on the EPIC V2 chip. Because the EPIC V2 chip uses a dual-color channel detection system, different probes may be assigned to the Cy3 or Cy5 channel. Correctly inferring the channel assignment is a prerequisite for subsequent dye bias correction.
[0100] Nonlinear correction of dye bias is a correction process for the systematic signal difference between two color channels. Since the optical characteristics of different channels are different, nonlinear methods are needed to eliminate this channel bias so that the signals from different channels are comparable.
[0101] p-value masking refers to marking probes whose signal strength does not reach the significance threshold as missing values based on the detection p-value of each probe, instead of retaining their noisy data. This processing method can avoid low-quality data from interfering with subsequent analysis.
[0102] NOOB background correction is a background subtraction method based on a normal exponential distribution. This method uses the mismatch signal of the probe to estimate the background distribution, subtracts the background component from the original signal, and obtains a more accurate estimate of the methylation signal.
[0103] Specifically, the raw data is quality-masked to exclude cross-hybridization probes and non-specific probes. Then, channel assignment inference is performed on the quality-masked data to determine the fluorescence detection channel to which each probe belongs. Next, non-linear correction for dye bias is applied to the channel-assigned data to eliminate systematic errors between the two-color channels. Afterward, p-value masking is performed on the dye bias-corrected data, marking probes whose detection p-values do not meet the significance requirement as missing values. Finally, NOOB background correction is performed on the p-value-masked data to subtract background signals.
[0104] Understandably, this embodiment, through the preprocessing process executed sequentially in the above five steps, can effectively reduce probe non-specific signals, channel bias, and background noise in EPICV2 chip data, improve the accuracy of methylation level estimation, and provide a more reliable data foundation for subsequent CpG site screening and model construction.
[0105] In this embodiment, a sample quality control step was also performed during the preprocessing of EPIC V2 methylation chip data. The probe missing rate (NA) for each sample was used as the quality control indicator. Samples with an NA rate greater than 10% were judged as low-quality samples and removed. After this quality control, all 480 samples met the requirements, and no samples were removed due to quality issues. Table 11 shows the sample quality control indicator table. SeSAMe's pval masking mechanism will generate 5%~8% probe NA in normal samples, which is expected behavior. Based on the NA rate distribution of 500 samples (median approximately 5.2%), low-quality samples were filtered out with an NA rate <10%, ultimately retaining 480 samples and filtering out 20.
[0106] Table 11: Sample Quality Control Indicators
[0107] In some embodiments of the present invention, the sample to be tested is a human peripheral blood mononuclear cell sample.
[0108] Human peripheral blood mononuclear cells refer to a population of single nucleated cells isolated from human peripheral blood, mainly including lymphocytes, monocytes, and natural killer cells. These cells are commonly used as a sample type in biological age prediction studies because they are easy to obtain, relatively simple to detect, and their DNA methylation modifications show regular changes with age.
[0109] Specifically, the test sample is collected from the subject's peripheral venous blood. After collection, the blood sample is separated into a peripheral blood mononuclear cell layer by density gradient centrifugation, thus obtaining the test sample. After obtaining the test sample, genomic DNA is extracted from the sample, and the methylation level of the aforementioned 223 CpG sites is detected. Substituting these values into the prediction formula, the subject's biological age can be calculated.
[0110] Understandably, this embodiment fully utilizes the advantages of convenient cell collection and simple detection operation by using human peripheral blood mononuclear cells as the test sample. The biological sample required for the test can be obtained without complicated invasive operations, thus improving the clinical operability of the detection method.
[0111] In some embodiments of the present invention, obtaining the methylation level of each CpG site in the combination of methylation markers in the sample to be tested includes: extracting genomic DNA from the sample to be tested; detecting the genomic DNA using a methylation chip; and determining the methylation level of each CpG site among the 223 CpG sites based on the detection results.
[0112] Specifically, after extracting genomic DNA from the sample to be tested, the methylation level of 223 CpG sites was detected according to the following steps: first, methylation microarray was used to detect the genomic DNA, and then the methylation level of each CpG site was determined based on the detection results.
[0113] For example, the Illumina EPIC V2 methylation chip can be used for detection. This chip can simultaneously detect the methylation status of approximately 750,000 CpG sites across the entire genome, from which the detection results of 223 target CpG sites can be extracted.
[0114] Understandably, this embodiment defines the methylation level as a beta value, thus clarifying the quantitative standard for the degree of methylation and facilitating subsequent formula calculations. Furthermore, by listing various methylation detection methods, it provides a flexible implementation method for those skilled in the art, allowing them to select appropriate detection methods based on actual experimental conditions and equipment, ensuring that the methylation levels of the 223 CpG sites can be accurately obtained.
[0115] Table 12 shows the model evaluation results. On the training set, the mean absolute error (MAO) was 2.84 years, the root mean square error (RMSE) was 3.75 years, and the correlation coefficient was 0.985. In cross-validation, the MAO was 4.54 ± 0.08 years. On the archived test set, the MAO was 4.80 years, the RMSE was 6.14 years, and the correlation coefficient was 0.962. The difference between the MAO of the cross-validation and the MAO of the test set was only 0.26 years, indicating that the prediction model constructed by this method has extremely low overfitting and good generalization ability. The samples with larger prediction errors in the test set had a maximum deviation of approximately ±20 years, resulting in an RMSE greater than the MAO.
[0116] Table 12: Model Evaluation Results
[0117] In one specific implementation, such as Figure 2 The figure shows a schematic diagram illustrating the prediction accuracy of the biological age prediction model. The horizontal axis represents the actual age of the sample to be tested, in years, and the vertical axis represents the biological age calculated according to the method provided in this embodiment of the invention, in years. Each circular marker in the figure represents a validation set sample. Figure 2 The model also includes a diagonal line, where points on this diagonal represent biological ages equal to actual ages. The closer a sample point is to this diagonal line, the more accurate the prediction. Validated on an independent validation set, the model's mean absolute error (MAE) is 4.80 years, the root mean square error (RMSE) is 6.14 years, and the correlation coefficient (r) is 0.962. The mean absolute error obtained from repeated cross-validation on the training set is 4.54 ± 0.08 years. These results demonstrate that the model possesses high prediction accuracy and good generalization ability.
[0118] Table 13 shows the distribution of age acceleration. Based on the biological age predicted by this method, age acceleration can be further calculated. The formula for calculating age acceleration is: age acceleration equals biological age minus chronological age. Age acceleration was calculated for 95 samples in the test set. The minimum value was -15.82 years, the first quartile was -3.04 years, the median was 0.58 years, the mean was 0.57 years, the third quartile was 5.18 years, and the maximum value was 19.48 years. A positive age acceleration indicates that the biological age of the sample is greater than the chronological age, i.e., accelerated aging; a negative age acceleration indicates that the biological age is less than the chronological age, i.e., delayed aging.
[0119] Table 13: Age Acceleration Distribution Table
[0120] A positive value indicates that the biological age is greater than the actual age (accelerated aging), while a negative value indicates that the biological age is less than the actual age (delayed aging). Age acceleration distribution in the test set.
[0121] To better implement the biological age prediction method in the embodiments of the present invention, based on the biological age prediction method, such as... Figure 3 As shown, this embodiment of the invention also provides a biological age prediction device, the biological age prediction device 300 including: The acquisition unit 301 is used to acquire the methylation level of each CpG site in a set of CpG methylation markers in the sample to be tested. The marker set consists of 223 CpG sites as shown in Table 1 of the specification, and the methylation level is a value between 0 and 1. The prediction unit 302 is used to predict the biological age of the sample to be tested based on the sum of the methylation levels of each CpG site multiplied by the corresponding regression coefficients and the baseline biological age value. The biological age baseline value and the regression coefficients corresponding to each CpG site are determined by elastic network regression modeling based on the methylation level of the 223 CpG sites in the PBMC sample and the actual age of each PBMC sample; the biological age baseline value represents the biological age corresponding to when the methylation level of all 223 CpG sites is 0.
[0122] The biological age prediction device 300 provided in the above embodiments can realize the technical solutions described in the above biological age prediction method embodiments. The specific implementation principles of each module or unit can be found in the corresponding content in the above biological age prediction method embodiments, and will not be repeated here.
[0123] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.), and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0124] The foregoing has provided a detailed description of the methylation biomarker combination, biological age prediction method, and apparatus provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. The application of a combination of methylation biomarkers in biological age prediction, characterized in that, The methylation marker combination includes: cg23744638、cg01477894、cg01188867、cg27030854、cg04193015、cg04644869、cg16966397、cg18448426、cg13331563、cg25253898、cg00387658、cg25490376、cg03752138、cg06210197、cg22493216、cg15109150、cg25253895、cg01763090、cg15728320、cg09143195、cg16545019、cg03424070、cg19663246、cg03732728、cg07158339、cg17767963、cg16983588、cg16966398、cg26514191、cg23510764、cg10221746、cg18452806、cg00842595、cg11935615、cg04538460、cg08466899、cg23124451、cg00533891、cg05150327、cg03996822、cg06272857、cg06669752、cg07822980、cg14209784、cg19727165、cg16767132、cg27333886、cg15736994、cg27430293、cg02554627、cg14295611、cg18881380、cg09171967、cg14144583、cg22798948、cg22549904、cg18622950、cg13604801、cg04367531、cg06839255、cg00876127、cg23734065、cg03467555、cg13612317、cg18684863、cg20692569、cg15148145、cg00593900、cg14155195、cg10348486、cg14253018、cg01949363、cg13734401、cg15676719、cg20067846、cg14996491、cg04462378、cg07485775、cg06027629、cg24709951、cg09372060、cg15882041、cg00557647、cg18468088、cg24919946、cg04478095、cg08469939、cg18881402、cg07841070、cg24806066、cg16087521、cg23405095、cg15583492、cg17804348、cg17017591、cg08029165、cg27603605、cg16582319、cg21757387、cg14162928、cg04104144、cg09349619、cg15017982、cg14036830、cg01635555、cg09918713、cg13207212、cg26014782、cg24369989、cg22008851、cg04140663、cg11957087、cg21379099、cg26060605、cg23476877、cg26897150、cg00688297、cg06852243、cg18329187、cg05666820、cg14166009、cg18815647、cg09715059、cg01933590、cg06893273、cg27016307、cg14643892、cg20669572、cg27659368、cg26413855、cg21571060、cg11925696、cg17508941、cg05542681、cg07717903、cg01070775、cg01408508、cg12948621、cg21570988、cg01358804、cg03731828、cg17179314、cg23008153、cg18084609、cg26332926、cg15470658、cg00446115、cg23737263、cg11651896、cg21899500、cg17364089、cg20858243、cg07884745、cg10665321、cg17364672、cg17816357、cg13504434、cg17163168、cg06208270、cg21213853、cg26698460、cg08529549、cg20002504、cg08667128、cg13149736、cg01410876、cg00464640、cg20818947、cg11868595、cg19986909、cg24939561、cg18788940、cg22524327、cg11229663、cg21870668、cg26676157、cg06259248、cg13221677、cg07215693、cg22932336、cg24013620、cg00004996、cg11429271、cg05260088、cg00593462、cg01873886、cg01682111、cg07949597、cg24195002、cg13619054、cg20164601、cg04561804、cg05825244、cg18839637、cg10362475、cg14185918、cg23128634、cg07519847、cg05162523、cg10634619、cg02527881、cg16340422、cg20977229、cg15380598、cg08046044、cg13185413、cg14880348、cg01127154、cg09570348、cg02601140、cg03025830、cg23114616、cg10012273、cg13401893、cg12633154、cg24609819、cg26095158、cg02441474、cg19480198、cg06573787、cg02961654、cg04985582、cg03726693。、 2. The application of a reagent for detecting the methylation level of a combination of methylation markers in the preparation of a kit for biological age prediction, characterized in that, The methylation marker combination includes: cg23744638、cg01477894、cg01188867、cg27030854、cg04193015、cg04644869、cg16966397、cg18448426、cg13331563、cg25253898、cg00387658、cg25490376、cg03752138、cg06210197、cg22493216、cg15109150、cg25253895、cg01763090、cg15728320、cg09143195、cg16545019、cg03424070、cg19663246、cg03732728、cg07158339、cg17767963、cg16983588、cg16966398、cg26514191、cg23510764、cg10221746、cg18452806、cg00842595、cg11935615、cg04538460、cg08466899、cg23124451、cg00533891、cg05150327、cg03996822、cg06272857、cg06669752、cg07822980、cg14209784、cg19727165、cg16767132、cg27333886、cg15736994、cg27430293、cg02554627、cg14295611、cg18881380、cg09171967、cg14144583、cg22798948、cg22549904、cg18622950、cg13604801、cg04367531、cg06839255、cg00876127、cg23734065、cg03467555、cg13612317、cg18684863、cg20692569、cg15148145、cg00593900、cg14155195、cg10348486、cg14253018、cg01949363、cg13734401、cg15676719、cg20067846、cg14996491、cg04462378、cg07485775、cg06027629、cg24709951、cg09372060、cg15882041、cg00557647、cg18468088、cg24919946、cg04478095、cg08469939、cg18881402、cg07841070、cg24806066、cg16087521、cg23405095、cg15583492、cg17804348、cg17017591、cg08029165、cg27603605、cg16582319、cg21757387、cg14162928、cg04104144、cg09349619、cg15017982、cg14036830、cg01635555、cg09918713、cg13207212、cg26014782、cg24369989、cg22008851、cg04140663、cg11957087、cg21379099、cg26060605、cg23476877、cg26897150、cg00688297、cg06852243、cg18329187、cg05666820、cg14166009、cg18815647、cg09715059、cg01933590、cg06893273、cg27016307、cg14643892、cg20669572、cg27659368、cg26413855、cg21571060、cg11925696、cg17508941、cg05542681、cg07717903、cg01070775、cg01408508、cg12948621、cg21570988、cg01358804、cg03731828、cg17179314、cg23008153、cg18084609、cg26332926、cg15470658、cg00446115、cg23737263、cg11651896、cg21899500、cg17364089、cg20858243、cg07884745、cg10665321、cg17364672、cg17816357、cg13504434、cg17163168、cg06208270、cg21213853、cg26698460、cg08529549、cg20002504、cg08667128、cg13149736、cg01410876、cg00464640、cg20818947、cg11868595、cg19986909、cg24939561、cg18788940、cg22524327、cg11229663、cg21870668、cg26676157、cg06259248、cg13221677、cg07215693、cg22932336、cg24013620、cg00004996、cg11429271、cg05260088、cg00593462、cg01873886、cg01682111、cg07949597、cg24195002、cg13619054、cg20164601、cg04561804、cg05825244、cg18839637、cg10362475、cg14185918、cg23128634、cg07519847、cg05162523、cg10634619、cg02527881、cg16340422、cg20977229、cg15380598、cg08046044、cg13185413、cg14880348、cg01127154、cg09570348、cg02601140、cg03025830、cg23114616、cg10012273、cg13401893、cg12633154、cg24609819、cg26095158、cg02441474、cg19480198、cg06573787、cg02961654、cg04985582、cg03726693。、 3. A method for predicting biological age, characterized in that, include: Obtain the methylation level of each CpG site in the combination of methylation markers according to claim 1 or 2 in the sample to be tested, wherein the methylation level is a value between 0 and 1; The biological age of the sample to be tested is predicted by summing the methylation levels of each CpG site and the corresponding regression coefficients, and by using the baseline biological age value. The baseline biological age value and the regression coefficients corresponding to each CpG site are determined by elastic network regression modeling of the PBMC samples based on the methylation level of the 223 CpG sites in the PBMC samples and the actual age of each PBMC sample. The biological age baseline value represents the biological age corresponding to the methylation level of all 223 CpG sites being 0.
4. The biological age prediction method according to claim 3, characterized in that, The steps for determining the baseline biological age value and the regression coefficients corresponding to each CpG site include: The EPIC V2 methylation chip detection data of the PBMC samples were obtained and preprocessed using the QCDPB algorithm in the SeSAMe software package; We performed a whole-epigenome association analysis on the preprocessed data, using age as the independent variable and sex and cell composition as covariates for correction, and screened out candidate CpG sites. The selected candidate CpG sites were filtered for cell composition stability to obtain the target CpG sites; On the PBMC sample, the methylation level of the target CpG site was used as the input variable and the actual age was used as the output variable. Elastic network regression modeling was adopted, and the hyperparameters were optimized by cross-validation to obtain the baseline biological age and the regression coefficients of each CpG site.
5. The biological age prediction method according to claim 4, characterized in that, The process of filtering candidate CpG sites for cellular compositional stability includes: Calculate the maximum correlation coefficient between the methylation level of each candidate CpG site and the proportion of various PBMC cell types; Candidate CpG sites with a maximum correlation coefficient less than a third preset threshold are identified as target CpG sites.
6. The biological age prediction method according to claim 5, characterized in that, The various PBMC cell types include naive CD4+ T cells, memory CD4+ T cells, naive CD8+ T cells, memory CD8+ T cells, regulatory T cells, memory B cells, naive B cells, natural killer cells, and monocytes.
7. The biological age prediction method according to claim 4, characterized in that, The preprocessed data underwent a whole-epigenome association analysis, with age as the independent variable and sex and cellular composition as covariates for correction, including: The methylation level of each CpG site in the preprocessed data is converted into a logarithmic M value; Using the transformed M value as the dependent variable, age as the independent variable, and the proportion of sex and multiple PBMC cell types as covariates, a linear model was constructed and fitted to obtain the fitting results. The fitting results were adjusted using an empirical Bayesian algorithm; Based on the adjusted fitting results, CpG sites with an error detection rate less than the first preset threshold and an absolute value of methylation difference greater than the second preset threshold are selected as candidate CpG sites.
8. The biological age prediction method according to claim 7, characterized in that, The absolute value of the methylation difference is the mean difference in methylation levels between the older and younger groups.
9. The biological age prediction method according to claim 3, characterized in that, The acquisition of the methylation level of each CpG site in the methylation biomarker combination in the sample to be tested includes: Genomic DNA was extracted from the sample to be tested; The genomic DNA was detected using a methylation chip. The methylation level of each CpG site among the 223 CpG sites was determined based on the detection results.
10. A biological age prediction device, characterized in that, include: The acquisition unit is used to acquire the methylation level of each CpG site in the combination of methylation markers according to claim 1 or 2 in the sample to be tested, wherein the methylation level is a value between 0 and 1; The prediction unit is used to predict the biological age of the sample to be tested based on the sum of the methylation levels of each CpG site multiplied by the corresponding regression coefficients and the baseline biological age value. The baseline biological age value and the regression coefficients corresponding to each CpG site are determined by elastic network regression modeling of the PBMC samples based on the methylation level of the 223 CpG sites in the PBMC samples and the actual age of each PBMC sample. The biological age baseline value represents the biological age corresponding to the methylation level of all 223 CpG sites being 0.