A method for constructing a biological age prediction model based on DNA methylation
By screening age-related CpG sites using whole-genome methylation and single-cell RNA sequencing data, a cell type-weighted elastic network regression model was constructed. This model overcomes the shortcomings of traditional models in distinguishing between normal and pathological aging and reflecting real-time changes in biological age, thus achieving more accurate prediction of biological age.
Patent Information
- Application Number
- CN202510503663.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Traditional biological age prediction models based on DNA methylation struggle to distinguish between normal and pathological aging, and cannot reflect changes in an individual's biological age in real time after interventions. Furthermore, DNA methylation data is highly dimensional and noisy, making it difficult to identify the methylation sites most relevant to age.
By acquiring whole-genome methylation sequencing data and single-cell RNA sequencing from human peripheral blood samples, cell subpopulation classification was performed, and CpG sites with significant changes in methylation levels and related to aging gene expression in the 20-80 age range were screened. A biological age prediction model was constructed using a cell type-weighted elastic network regression algorithm.
It improves the accuracy and stability of the model, enabling more precise prediction of an individual's biological age, reflecting the trend of methylation changes at different age stages, and providing a tool for assessing an individual's aging level and health status.
Smart Images

Figure CN120032713B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, and particularly relates to a method for constructing a biological age prediction model based on DNA methylation. BACKGROUND
[0002] The biological age prediction model is a model for predicting the biological age of an individual by analyzing biomarker data, so as to evaluate the health status and aging degree of the individual. The traditional age measurement mainly relies on the actual age (i.e. the date of birth) calculated according to the calendar, however, the actual age cannot accurately reflect the biological state and health level of the individual. For example, there can be significant differences in physical function, disease risk and life span among people of the same age, which is mainly due to various factors such as genetic factors, lifestyle, environmental exposure, etc. Therefore, the scientific community urgently needs a more accurate biological age evaluation method to better reflect the health status and aging rate of the individual. DNA methylation is an epigenetic modification that changes with age. Horvath's Clock is a classic DNA methylation age prediction model that calculates the DNA methylation age by analyzing the methylation level of specific CpG sites. In addition, DNAm PhenoAge is also a commonly used model, which combines multiple age-related DNA methylation markers and blood indicators to more accurately reflect the biological age.
[0003] However, the construction of the traditional biological age prediction model based on DNA methylation often has the following problems: some models are difficult to distinguish between normal aging and pathological aging, such as the biological age evaluation of premature aging, chronic disease, etc. is not accurate enough; the DNA methylation data has high dimension and large noise, how to screen out the methylation sites most related to age from it and construct a stable and efficient prediction model has been a technical difficulty; the traditional model is mostly static prediction, which cannot reflect the dynamic changes of the individual's biological age after intervention measures (such as lifestyle changes, drug treatment). SUMMARY
[0004] Therefore, it is necessary to provide a method for constructing a biological age prediction model based on DNA methylation to solve at least one of the above technical problems.
[0005] To achieve the above-mentioned purpose, a method for constructing a biological age prediction model based on DNA methylation comprises the following steps:
[0006] Step S1: obtaining whole genome methylation sequencing data of a human peripheral blood sample; classifying the human peripheral blood sample into cell subpopulations, and performing single-cell RNA sequencing processing on each cell subpopulation to obtain single-cell sequencing data including lymphocytes, neutrophils and monocytes;
[0007] Step S2: Organize specific analysis of CpG sites within 2000 base pairs upstream and downstream of each cell subpopulation-specific transcription factor binding site according to single-cell sequencing data and whole-genome methylation sequencing data, and obtain candidate marker site data;
[0008] Step S3: Perform dynamic evolution analysis on the candidate marker site data, and screen out CpG sites with a monotonous increasing or decreasing trend in methylation level within the age range of 20-80 years old and a change amplitude greater than 30%, and obtain age-related site data;
[0009] Step S4: According to the pre-acquired promoter region information of the aging gene and the age-related site data, screen the CpG sites located within 1000 base pairs upstream and downstream of the aging gene promoter and with an absolute value of the correlation coefficient between the methylation level and the gene expression greater than 0.6, and obtain core age marker site data;
[0010] Step S5: Construct a biological age prediction model according to the core age marker site data using an elastic network regression algorithm based on cell type weighting.
[0011] The application lays a solid data foundation for subsequent analysis by obtaining whole genome methylation sequencing data of human peripheral blood samples and processing cell subpopulations by single cell RNA sequencing. This not only comprehensively understands the genomic methylation state of individuals, but also deeply explores the gene expression characteristics of different cell subpopulations, so as to more accurately grasp the biological information at the cell level, provide rich data sources for screening age-related marker sites, and effectively improve the starting quality of model construction. Based on the single cell sequencing data and the whole genome methylation sequencing data, the CpG sites within the range of 2000 base pairs upstream and downstream of the cell subpopulation specific transcription factor binding site are analyzed for tissue specificity. This operation accurately locks the specific region that may be related to age change. Such analysis helps to screen candidate marker sites with potential value, greatly reduces the scope of subsequent research, improves the pertinence and efficiency of research, makes the subsequent screening of age-related sites more targeted, and provides key intermediate data for constructing a precise biological age prediction model. Further dynamic evolution analysis of the candidate marker site data selects the CpG sites with a methylation level change of more than 30% in the 20-80 age range. This screening process can effectively eliminate sites with unstable methylation level changes or small change amplitudes, and the remaining age-related site data is more representative and reliable, and can more accurately reflect the methylation change trend of individuals at different ages. The model provides high-quality core data for subsequent construction, enhancing the accuracy and stability of the model for age prediction. According to the promoter region information of the pre-obtained aging genes and the age-related site data, the CpG sites located within the range of 1000 base pairs upstream and downstream of the promoter of the aging gene and with an absolute value of the correlation coefficient between methylation level and gene expression greater than 0.6 are screened. This step closely links the methylation level and the gene expression, and the core age marker site data screened is not only closely related to the age change, but also directly related to the expression regulation of the aging gene. Such a screening mechanism makes the final core age marker site more biologically meaningful, providing a strong biological basis for the constructed model and further improving the scientificity and reliability of the model. A biological age prediction model is constructed according to the core age marker site data by using a cell type weighted elastic network regression algorithm. This algorithm fully considers the weight difference of different cell types in age prediction, can more accurately integrate the core age marker site data, and the constructed model not only has high prediction accuracy, but also can effectively avoid overfitting and other problems. The optimal regularization parameter is determined by cross-validation of the model, further optimizing the performance of the model, so that it can more accurately predict the biological age of individuals in practical application, providing a powerful tool for evaluating the aging degree and health status of individuals, and having important scientific value and application prospect. BRIEF DESCRIPTION OF DRAWINGS
[0012] Other features, objects, and advantages of the application will become more apparent from the following detailed description when read in connection with the following drawings:
[0013] Figure 1 A schematic diagram of the step flow for the method for constructing a DNA methylation-based biological age prediction model of the present application;
[0014] Figure 2 A schematic diagram of the step flow for the method for constructing a DNA methylation-based biological age prediction model of the present application; Figure 1 A schematic diagram of the step flow for the method for constructing a DNA methylation-based biological age prediction model of the present application;
[0015] Figure 3 A schematic diagram of the step flow for the method for constructing a DNA methylation-based biological age prediction model of the present application; Figure 1 A schematic diagram of the step flow for the method for constructing a DNA methylation-based biological age prediction model of the present application. DETAILED DESCRIPTION
[0016] The technical method of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0017] In addition, the accompanying drawings are only schematic illustrations of the present application, and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0018] It should be understood that although the terms "first", "second", etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the example embodiments, a first element can be called a second element, and similarly a second element can be called a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0019] To achieve the above-mentioned object, in accordance with the present application Figures 1 to 3 The present application provides a method for constructing a DNA methylation-based biological age prediction model, which comprises the following steps:
[0020] Step S1: obtaining whole genome methylation sequencing data of a human peripheral blood sample; classifying cell subpopulations of the human peripheral blood sample, and performing single-cell RNA sequencing processing on each cell subpopulation to obtain single-cell sequencing data including lymphocytes, neutrophils and monocytes;
[0021] In the embodiment of the present application, peripheral blood samples of 1000 healthy subjects are selected, each sample has a volume of 5 mL, collected using an EDTA anticoagulant tube, and stored at 4℃ for no more than 4 hours before processing. First, single nucleated cells are separated by density gradient centrifugation, and red blood cells are removed using a red blood cell lysis solution. After centrifugation, white blood cells are collected for DNA extraction and RNA extraction. For DNA extraction, genomic DNA is extracted using a silica gel membrane column method, and is treated using a bisulfite conversion method for whole genome methylation sequencing. For RNA extraction, total RNA is extracted using TRIzol reagent, followed by mRNA enrichment and library construction, and single-cell RNA sequencing of cell subpopulations is performed using a 10X Genomics single-cell sequencing platform. The Seurat software package is used to cluster the single-cell sequencing data, and the cells are divided into three types of lymphocytes, neutrophils and monocytes according to the known characteristic gene expression patterns of the cell types, and a single-cell expression matrix is constructed for each type of cell to obtain a high-resolution single-cell sequencing data set.
[0022] Step S2: performing tissue-specific analysis on CpG sites within a range of 2000 base pairs upstream and downstream of cell subpopulation-specific transcription factor binding sites according to single-cell sequencing data and whole genome methylation sequencing data to obtain candidate marker site data;
[0023] In the embodiment of the present application, transcription factors specifically expressed in different cell types are extracted from single-cell RNA sequencing data, transcription factors with expression significantly higher than that in other cell subpopulations (p value < 0.05, log2FC > 1.5) in at least one cell subpopulation are screened, and the gene promoter region information thereof is recorded. Then, the methylation level data of CpG sites within a range of 2000 base pairs upstream and downstream of the transcription factor binding sites are extracted from the whole genome methylation sequencing data, and combined with the transcription factor binding site information of the ENCODE database, the CpG sites are subjected to tissue-specific analysis. The tSNE dimension reduction method is used to analyze the CpG methylation patterns of different cell types, and sites with methylation levels significantly higher than those in other cell subpopulations (p value < 0.05, methylation difference > 20%) in a specific cell subpopulation are screened as candidate marker site data, and stored in a database for subsequent analysis.
[0024] Step S3: Dynamic evolution analysis is performed on the candidate marker site data, CpG sites with a monotonically increasing or monotonically decreasing trend in the methylation level and a change amplitude greater than 30% in the 20-80 age group are screened out, and age-related site data is obtained;
[0025] In the embodiment of the application, candidate marker site data of healthy individuals with balanced age distribution (age range of 20-80 years old, at least 100 people per 10-year age group) are selected, the mean methylation level of each site in different age groups is calculated, and linear regression analysis is performed. The correlation between the methylation level and the age is calculated by Spearman rank correlation analysis, and CpG sites with an absolute value of Spearman correlation coefficient greater than 0.6 and a p value less than 0.05 are screened out. The change amplitude of the methylation level of these sites between 20 years old and 80 years old is further calculated, and CpG sites with a monotonically increasing or monotonically decreasing trend and a change amplitude greater than 30% in 60 years are selected as age-related site data.
[0026] Step S4: According to the pre-acquired promoter region information of the aging gene and the age-related site data, CpG sites located within 1000 base pairs upstream and downstream of the promoter of the aging gene and with an absolute value of the correlation coefficient between the methylation level and the gene expression greater than 0.6 are screened out, and core age marker site data is obtained.
[0027] In the embodiment of the application, aging-related gene information in the aging gene database (such as GenAge database) is pre-acquired, and the promoter region (within 1000 base pairs upstream and downstream of the gene transcription start site) is extracted. CpG sites located in these regions are screened out from the age-related site data, and the Pearson correlation coefficient between the methylation level of these sites and the expression of the corresponding gene is calculated combined with single-cell RNA sequencing data. Only CpG sites with an absolute value of the correlation coefficient greater than 0.6 and a p value less than 0.05 are retained as core age marker site data. For the finally screened CpG sites, MeDIP-qPCR experiment is used to further verify the methylation level, and high-correlation sites verified by experiment are selected for subsequent analysis.
[0028] Step S5: A biological age prediction model is constructed according to the core age marker site data by using an elastic network regression algorithm based on cell type weighting.
[0029] The embodiment of the present application adopts an elastic network regression algorithm based on cell type weighting to construct a biological age prediction model. First, the core age marker site data is normalized, and the weighted methylation level of each CpG site in different age groups is calculated. Then, the cell type information is taken as a weighting factor, and the mapping relationship between the methylation level of the CpG site and the actual age of the individual is learned through a weighted elastic network regression model. The prediction performance of the model is evaluated by 10-fold cross-validation, and the mean square error (MSE) and the determination coefficient (R²) are used as the model evaluation indicators. Finally, the generalization ability of the model is verified on an independent test set, and the model parameters (such as L1 / L2 regularization coefficients) are optimized to improve the prediction accuracy.
[0030] The application lays a solid data foundation for subsequent analysis by obtaining whole genome methylation sequencing data of human peripheral blood samples and processing cell subpopulations by single cell RNA sequencing. This not only comprehensively understands the genomic methylation state of individuals, but also deeply explores the gene expression characteristics of different cell subpopulations, so as to more accurately grasp the biological information at the cell level, provide rich data sources for screening age-related marker sites, and effectively improve the starting quality of model construction. Based on the single cell sequencing data and the whole genome methylation sequencing data, the CpG sites within the range of 2000 base pairs upstream and downstream of the cell subpopulation specific transcription factor binding site are analyzed for tissue specificity. This operation accurately locks the specific region that may be related to age change. Such analysis helps to screen candidate marker sites with potential value, greatly reduces the scope of subsequent research, improves the pertinence and efficiency of research, makes the subsequent screening of age-related sites more targeted, and provides key intermediate data for constructing a precise biological age prediction model. Further dynamic evolution analysis of the candidate marker site data selects the CpG sites with a methylation level change of more than 30% in the 20-80 age range. This screening process can effectively eliminate sites with unstable methylation level changes or small change amplitudes, and the remaining age-related site data is more representative and reliable, and can more accurately reflect the methylation change trend of individuals at different ages. The model provides high-quality core data, enhances the accuracy and stability of the model for age prediction. According to the promoter region information of the pre-obtained aging genes and the age-related site data, the CpG sites located within 1000 base pairs upstream and downstream of the promoter of the aging gene and with an absolute value of the correlation coefficient between methylation level and gene expression greater than 0.6 are screened. This step closely links the methylation level and the gene expression, and the core age marker site data screened not only closely related to age change, but also directly related to the expression regulation of aging genes. Such a screening mechanism makes the final core age marker site more biologically meaningful, providing a strong biological basis for the constructed model, further improving the scientificity and reliability of the model. A biological age prediction model is constructed according to the core age marker site data by using a cell type weighted elastic network regression algorithm. The algorithm fully considers the weight difference of different cell types in age prediction, can more accurately integrate the core age marker site data, and the constructed model not only has high prediction accuracy, but also can effectively avoid overfitting and other problems. The optimal regularization parameter is determined by cross-validation of the model, further optimizing the performance of the model, so that it can more accurately predict the biological age of individuals in practical application, providing a powerful tool for evaluating the aging degree and health status of individuals, and having important scientific value and application prospect.
[0031] Preferably, step S1 comprises the following steps:
[0032] Step S11: obtaining a human peripheral blood sample, and separating and processing the human peripheral blood sample by using a density gradient centrifugation method to obtain mononuclear cell layer separation data;
[0033] In the embodiment of the present application, peripheral blood samples of healthy volunteers are selected, each sample has a volume of 10 mL, is collected by using an EDTA anticoagulation tube, and is stored at 4℃ for no more than 4 hours to ensure cell activity. The blood sample is diluted to PBS buffer (without calcium and magnesium ions) at a ratio of 1:1, an equal volume of Ficoll-Paque Plus density gradient separation liquid (density 1.077 g / mL) is slowly added, the liquid interface is kept clear, and then centrifugation is performed at 800xg for 30 minutes (temperature 4℃), without braking during the centrifugation to avoid disturbance of the interface layer. After the centrifugation is completed, the white cell cloud layer, i.e., the mononuclear cell layer, located at the junction of the plasma and the density gradient separation liquid is carefully collected by using a Pasteur pipette, and is washed twice (400xg, 10 minutes) by using PBS buffer to remove residual platelets and density separation liquid, and is finally resuspended in 1 mL of cold PBS buffer to obtain mononuclear cell layer separation data, and cell counting and viability detection (trypan blue staining method, cell viability should be higher than 90%) are performed.
[0034] Step S12: performing whole genome methylation level sequencing processing on DNA according to the mononuclear cell layer separation data by using a bisulfite sequencing method to obtain whole genome methylation sequencing data;
[0035] In the embodiment of the present application, 10 6 mononuclear cell samples are taken, genomic DNA is extracted by using QIAGEN DNeasy Blood & Tissue Kit according to the instructions, and the DNA concentration is measured by using NanoDrop 2000 to ensure that the DNA integrity OD260 / 280 ratio is between 1.8-2.0. Bisulfite treatment is performed by using Zymo Research EZ DNA Methylation-Gold Kit, and the specific method includes dissolving 500 ng of genomic DNA in a bisulfite reagent, performing denaturation (98℃ for 10 minutes) and bisulfite conversion (64℃ for 2.5 hours), and then performing purification treatment by using a silica gel membrane column and library construction. Whole genome methylation sequencing is performed by using an Illumina NovaSeq 6000 platform, 150 bp double-end sequencing strategy (PE150) is selected, the target sequencing depth is 30x, and the generated data is converted to obtain whole genome methylation sequencing data after base recognition.
[0036] Step S13: According to the mononuclear cell layer separation data, the different cell subgroups are subjected to enrichment separation treatment based on magnetic bead sorting, so as to obtain preliminary sorting subgroup data;
[0037] In the embodiment of the present application, CD14 magnetic beads are used to sort monocyte macrophage subgroups in mononuclear cells, CD3 magnetic beads are used to sort T lymphocyte subgroups, CD19 magnetic beads are used to sort B lymphocyte subgroups, and CD16 / CD66b magnetic beads are used to sort neutrophil subgroups. During the experiment, the mononuclear cells are suspended in PBS (containing 0.5% bovine serum albumin BSA), and the corresponding magnetic bead labeling antibodies (concentration 5 μL / 10 6 cells) are added, and incubated at 4°C for 15 minutes. Then, magnetic sorting columns are used for sorting, and the effluent and adsorbed cells are collected, and after excluding dead cells using trypan blue, cell counting is performed. Each subgroup of cells is suspended in RPMI-1640 medium (containing 10% FBS) to obtain preliminary sorting subgroup data.
[0038] Step S14: The preliminary sorting subgroup data is subjected to flow cytometry purity verification processing, and a cell population with a purity greater than 95% is screened, so as to obtain high-purity cell subgroup data;
[0039] In the embodiment of the present application, 10 6 cells of each cell subgroup are taken for flow cytometry (FACS) detection. A BD FACSCanto II flow cytometer is used, and fluorescently labeled antibodies (such as CD14-PE, CD3-FITC, CD19-APC, CD16-PerCP, etc.) consistent with the magnetic bead sorting are used for cell staining. The staining conditions are 4°C, dark incubation for 30 minutes, and then washed twice with PBS. The purity of each cell subgroup is determined using a flow cytometer, and the cell purity is analyzed using FlowJo software. A cell population with a purity greater than 95% is selected, and the cell suspension is centrifuged (300xg, 5 minutes), the supernatant is discarded, and the cells are resuspended in an RNA protection solution to obtain high-purity cell subgroup data.
[0040] Step S15: According to the high-purity cell subgroup data, the transcriptional sequencing of each cell subgroup is performed, so as to obtain single-cell transcriptional original data;
[0041] In the embodiment of the present application, 10 5Total RNA was extracted from each cell using QIAGEN RNeasy Micro Kit, and the RNA integrity was measured using Agilent 2100 Bioanalyzer, with the requirement that the RNA integrity number (RIN value) be greater than 7.0. Subsequently, cDNA synthesis was performed using the Smart-seq2 technology, and the 10X Genomics Chromium single-cell transcriptome sequencing platform was selected to prepare single-cell libraries and perform sequencing. Sequencing was performed using Illumina NovaSeq 6000 in 150bp paired-end sequencing mode, with a target data output of 50G per sample, and finally single-cell transcriptome raw data was obtained.
[0042] Step S16: Data quality control was performed on the single-cell transcriptome raw data to obtain single-cell sequencing data, wherein the data quality control included removing sequencing adapter sequences, filtering low-quality reads, and correcting PCR amplification bias.
[0043] In the present embodiment, fastp software was used to perform preliminary quality control on the single-cell transcriptome raw data, remove sequencing adapter sequences (adapter trimming), remove reads with base quality lower than Q20, and remove low-quality short fragment sequences shorter than 30bp. Next, UMI-tools was used to remove PCR repeat sequences to reduce the influence of amplification bias on data analysis. Then, STAR software was used to align the cleaned data to the human reference genome GRCh38, and the alignment rate was calculated, with the requirement that the alignment rate be greater than 85%. Finally, Cell Ranger software was used for cell barcode recognition and gene expression matrix construction, cells with mitochondrial gene expression ratio greater than 10% were removed, and low-quality cells with gene expression less than 500 genes were filtered, and finally high-quality single-cell sequencing data was obtained.
[0044] The present application successfully obtains mononuclear cell layer separation data by obtaining human peripheral blood samples and using density gradient centrifugation method for separation processing, lays a foundation for subsequent cell subpopulation analysis and sequencing work, ensures the accuracy and reliability of the research object, and enables the subsequent steps to be smoothly carried out and focused on the target cell components. With the help of bisulfite sequencing method, the DNA is sequenced for whole genome methylation level, and comprehensive and detailed whole genome methylation sequencing data is obtained. This operation provides key data support for subsequent analysis of the relationship between methylation and age, enabling researchers to fully understand the methylation state of the sample, providing rich and accurate data sources for screening age-related marker sites, and helping to deeply explore the mechanism of methylation in the aging process. According to the mononuclear cell layer separation data, the enrichment separation processing based on magnetic bead sorting is used to separate different cell subpopulations, and the preliminary sorting subpopulation data is obtained. This process effectively enriches the target cell subpopulation, improves the specificity and accuracy of subsequent analysis, enables researchers to focus more on the characteristics of specific cell subpopulations, and helps to deeply explore the unique role and contribution of different cell subpopulations in the aging process, providing important cell level data for constructing a precise biological age prediction model. After the flow cytometry purity verification process of step S14, cell populations with a purity greater than 95% are screened, and high-purity cell subpopulation data is obtained. High-purity cell subpopulation data is crucial for subsequent transcriptome sequencing, which can effectively reduce data interference and errors caused by cell impurities, ensure the accuracy and reliability of transcriptome sequencing results, and further improve the analysis accuracy of cell subpopulation-specific gene expression, providing a cleaner and more representative data basis for screening age-related genes and marker sites. Based on the high-purity cell subpopulation data, the transcriptome of each cell subpopulation is sequenced, and single-cell transcriptome raw data is obtained. This operation enables researchers to deeply understand the characteristics and differences of different cell subpopulations in the gene expression level, provides direct data support for exploring the relationship between gene expression and age change, helps to discover specific gene expression patterns related to aging, and further enriches the data dimension and biological basis for constructing a biological age prediction model. Data quality control is performed on the single-cell transcriptome raw data, including removing sequencing adapter sequences, filtering low-quality reads, correcting PCR amplification bias, etc., and finally obtaining high-quality single-cell sequencing data. Data quality control is an important link to ensure data accuracy and reliability, and the data after quality control can more truly reflect the gene expression state of the cell, avoid errors and biases in analysis and conclusions caused by data quality problems, provide a solid data guarantee for subsequent analysis and model construction based on single-cell sequencing data, and improve the scientificity and credibility of the whole research.
[0045] Preferably, step S13 comprises the following steps:
[0046] Based on the mononuclear cell layer separation data, different cell subpopulations were enriched and separated using magnetic bead sorting to obtain preliminary sorted subpopulation data. The enrichment and separation process specifically included enriching lymphocytes with CD3 magnetic beads, enriching mononuclear cells with CD14 magnetic beads, and enriching neutrophils with CD66b magnetic beads.
[0047] In this embodiment of the invention, the mononuclear cell layer separation data obtained in step S11 is used, with a total cell count of approximately 1 × 10⁻⁶. 7 Cells were resuspended in PBS (containing 2 mM EDTA and 0.5% bovine serum albumin BSA) and the concentration was adjusted to 1 × 10⁶ cells / day. 7 cells / mL. First, T lymphocytes were enriched using CD3 magnetic beads. The cell suspension was added to CD3 microbeads (manufactured by Miltenyi Biotec) at a ratio of 5 μL microbeads / 10 mL. 6 The cells were added at a ratio of 1:1 and incubated at 4°C in the dark for 15 minutes. After incubation, the cell suspension was transferred to an MS magnetic sorting column (placed within a MACS sorting magnetic field). The cells were washed three times with 1 mL of sorting buffer (PBS + 2 mM EDTA + 0.5% BSA), each time using gravity flow to elute unbound cells and collecting the effluent. The magnetic field was then removed, and the cells were washed with 1 mL of sorting buffer and the T lymphocyte population adsorbed to the magnetic beads was collected. Next, the remaining cells in the effluent were used for CD14 bead enrichment of monocytes at a ratio of 5 μL CD14 beads / 10 cells. 6 Magnetic beads were added at a specific ratio to the cells, and monocytes were separated using the same magnetic sorting process. The enriched CD14+ cells were then collected. Finally, the unbound cell effluent was collected, and CD66b magnetic beads (5 μL / 1000 ml) were added. 6 Neutrophils were enriched and separated using the same magnetic sorting method (100 cells). Each enriched cell subpopulation was suspended in 1 mL of RPMI-1640 medium (containing 10% fetal bovine serum FBS), counted, and their viability was assessed. The viability of each cell was greater than 90%. Preliminary sorted subpopulation data were obtained, in which T lymphocytes, monocytes, and neutrophils accounted for approximately 60%, 30%, and 10% of the total cell count, respectively.
[0048] The present application can accurately obtain preliminary sorting data of different cell subpopulations by in-depth analysis of mononuclear cell layer separation data and enrichment separation processing based on magnetic bead sorting. Specifically, lymphocytes are enriched by CD3 magnetic beads, monocytes are enriched by CD14 magnetic beads, and neutrophils are enriched by CD66b magnetic beads, which significantly improves the purity and quantity of each cell subpopulation and provides high-quality samples for subsequent single-cell sequencing. This enrichment separation method not only effectively removes impurity cells and reduces background noise, but also ensures the integrity and activity of target cell subpopulations, thereby improving the accuracy and reliability of single-cell sequencing data. This is of great significance for in-depth study of the gene expression characteristics and biological functions of each cell subpopulation and lays a solid foundation for constructing an accurate biological age prediction model.
[0049] Preferably, step S2 comprises the following steps:
[0050] Step S21: According to the single-cell sequencing data, differential expression gene analysis is performed on each cell subpopulation, and differential expression genes with enrichment fold greater than 2 and P value less than 0.01 are screened, thereby obtaining cell subpopulation-specific gene data;
[0051] The single-cell sequencing data obtained in step S16 is first standardized in the embodiment of the present application, including using the log2(TPM+1) method to normalize the transcript expression data to reduce the influence of sequencing depth. Subsequently, cell clustering analysis is performed using the Seurat software package, t-SNE or UMAP method is used for dimensionality reduction of the cell population, and cell subpopulation annotation is performed according to cell marker genes. For each cell subpopulation, differential expression analysis is performed using the DESeq2 software package, taking other cell subpopulations as controls, and the enrichment fold (Fold Change, FC) and significance P value of the genes are calculated. Genes that meet FC>2 and P value<0.01 are screened as the differential expression gene set. To improve data reliability, the screened genes are subjected to GO function annotation and KEGG pathway enrichment analysis to ensure their biological rationality, and finally cell subpopulation-specific gene data is obtained.
[0052] Step S22: Based on the JASPAR database, transcription factor binding site prediction processing is performed on the cell subpopulation-specific gene data, thereby obtaining potential transcription factor binding site data;
[0053] The embodiment of the present application takes the cell subpopulation-specific gene data obtained in step S21, and first acquires the sequence of the promoter region of these genes (defined as the range of 1000 base pairs upstream and downstream of the transcription start site). Using the JASPAR database containing known transcription factor binding site (TFBS) information, the target sequence is scanned based on the FIMO (Find Individual Motif Occurrences) tool to predict possible transcription factor binding sites. The specific parameter setting is P-value threshold <1e-4 to ensure the reliability of the prediction results. For each gene, the distribution of binding sites is compared between multiple cell subpopulations, and the binding sites enriched in a specific cell subpopulation are screened, and finally the potential transcription factor binding site data is obtained.
[0054] Step S23: ChIP-seq data verification processing is performed according to the potential transcription factor binding site data, and sites with a binding strength score greater than 0.8 are screened, thereby obtaining confirmed transcription factor binding site data.
[0055] The embodiment of the present application takes the potential transcription factor binding site data obtained in step S22, and uses the ENCODE database or other public ChIP-seq experimental data for verification. First, the ChIP-seq data most matched to the target cell type is downloaded, and Bowtie2 software is used for alignment, and the original sequencing data is aligned to the reference genome, and then MACS2 (Model-based Analysis for ChIP-Seq) is used for peak calling to identify significant transcription factor binding regions. For each potential binding site, calculate its binding strength score (Signal Score), which is calculated based on the enrichment signal of ChIP-seq data, and is usually represented by peak height (Peak Intensity) normalization. Sites with a binding strength score greater than 0.8 are screened to ensure the reliability of the binding sites, and finally the confirmed transcription factor binding site data is obtained.
[0056] Step S24: According to the confirmed transcription factor binding site data and whole genome methylation sequencing data, the CpG sites within the range of 2000 base pairs upstream and downstream of each cell subpopulation-specific transcription factor binding site are analyzed for tissue specificity, and candidate marker site data is obtained.
[0057] The embodiment of the present application takes the verified transcription factor binding site data obtained in step S23 and the whole genome methylation sequencing data obtained in step S12, extracts CpG sites within a range of 2000 base pairs upstream and downstream of the verified transcription factor binding sites, and calculates the methylation levels of these CpG sites in different cell subpopulations. The HOMER software is used to perform tissue-specific analysis on these regions, and CpG sites with a methylation level significantly lower than that in other cell subpopulations in a certain cell subpopulation are defined as tissue-specific CpG sites. The specific screening criteria are that the methylation level difference is greater than 30% (Δβ>0.3), and the P value is less than 0.01 (based on t test). Finally, the CpG sites meeting the criteria are selected as candidate marker site data, and visual analysis is performed, such as using heatmap to display the methylation differences of CpG sites in different cell subpopulations, to ensure the biological specificity of the candidate marker sites.
[0058] Through in-depth analysis of single-cell sequencing data, these steps can accurately identify specific genes of each cell subpopulation, laying a foundation for subsequent research. First, the differentially expressed gene analysis and screening of genes with an enrichment fold greater than 2 and a P value less than 0.01 ensure that the screened genes have significant expression differences, thereby obtaining cell subpopulation-specific gene data. Then, based on the JASPAR database, transcription factor binding site prediction is performed, further expanding the depth of research and providing important clues for understanding gene regulation mechanisms. Subsequently, ChIP-seq data is used to verify and screen sites with a binding strength score greater than 0.8, which not only improves the reliability of the data but also ensures the accuracy of subsequent analysis. Finally, through tissue-specific analysis of CpG sites within a range of 2000 base pairs upstream and downstream of the verified transcription factor binding sites, candidate marker site data is obtained, providing key information for subsequent construction of a biological age prediction model. The organic combination of these steps significantly improves the accuracy and reliability of the research and provides strong support for in-depth exploration of the role of cell subpopulations in the aging process.
[0059] Preferably, step S24 comprises the following steps:
[0060] Step S241: performing positioning mapping processing on the CpG sites within a range of 2000 base pairs upstream and downstream of the transcription factor binding sites according to the verified transcription factor binding site data and the whole genome methylation sequencing data, thereby obtaining preliminary marker site data;
[0061] The embodiment of the application takes the verified transcription factor binding site data obtained in step S23 and the whole genome methylation sequencing data obtained in step S12, first converts the genomic coordinates of the verified transcription factor binding sites to be consistent with the genomic coordinates of the methylation sequencing data. Then, using the intersectBed tool of the BEDTools software package, the transcription factor binding sites are extended by 2000 base pairs upstream and downstream, and the CpG sites in the methylation sequencing data are searched within the range. The CpG sites overlapping the transcription factor binding site interval are screened, and their methylation levels in different cell subpopulations are recorded to form a preliminary marker site data set.
[0062] Step S242: Organizing specific analysis processing of the preliminary marker site data, calculating the variation coefficient of the methylation level of each CpG site between different cell subpopulations, to obtain site-specific score data;
[0063] The embodiment of the application takes the preliminary marker site data obtained in step S241, first classifies the cell subpopulations, calculates the methylation level of each CpG site in different cell subpopulations, and uses the coefficient of variation (Coefficient of Variation, CV) to measure the variation degree of the methylation level. The specific calculation method is to first calculate the methylation mean of each CpG site in all cell subpopulations, then calculate the standard deviation, and obtain the coefficient of variation by the ratio of the standard deviation to the mean. The dplyr and tidyverse packages of R language are used to calculate the coefficients of variation of all preliminary marker sites in batches, and the sites with measurement data in at least 3 different cell subpopulations are screened out to ensure the integrity of the data, and finally the site-specific score data of each CpG site is obtained.
[0064] Step S243: Screening CpG sites with a variation coefficient greater than 0.5 according to the site-specific score data, to obtain candidate marker site data.
[0065] The embodiment of the application takes the site-specific score data obtained in step S242, and screens according to the size of the coefficient of variation, and only retains the CpG sites with a coefficient of variation greater than 0.5. First, the filter function of R language is used to screen the CpG sites meeting the threshold condition, and the average methylation level of these CpG sites in each cell subpopulation is calculated to ensure that there is a significant methylation difference between different cell subpopulations. For the screened CpG sites, the PCA (Principal Component Analysis) method is used to evaluate their discrimination in cell subpopulation classification, and a heatmap is drawn to visualize the methylation level distribution of these sites, and finally the candidate marker site data is obtained to provide data support for subsequent function verification.
[0066] The present application can accurately identify the CpG sites located within the range of 2000 base pairs upstream and downstream of the transcription factor binding sites through in-depth analysis of transcription factor binding site data and whole genome methylation sequencing data, laying a foundation for subsequent tissue-specific analysis. First, the positioning mapping process enables researchers to determine the specific location of these CpG sites, thereby obtaining preliminary marker site data. Next, tissue-specific analysis further evaluates the specificity of these sites by calculating the coefficient of variation of methylation levels of each CpG site among different cell subpopulations, thereby obtaining site-specific score data. Finally, candidate marker site data is obtained by screening CpG sites with a coefficient of variation greater than 0.5 according to the site-specific score data. This process not only improves the accuracy and reliability of the data, but also provides key information for subsequent biological age prediction model construction. The organic combination of these steps significantly improves the accuracy and reliability of the research, and provides strong support for in-depth exploration of the role of cell subpopulations in the aging process.
[0067] Preferably, step S3 comprises the following steps:
[0068] Step S31: Obtain peripheral blood samples of healthy people in each 5-year interval within the age range of 20-80 years old, with a sample size of not less than 50 in each age interval, thereby obtaining stratified age sample data;
[0069] In the embodiment of the present application, healthy people in the age range of 20-80 years old are selected, stratified according to every 5-year age interval, and not less than 50 healthy volunteers are recruited in each age interval to ensure balanced and representative sample size. 5-10 mL of peripheral blood of the volunteers is collected using an EDTA anticoagulant tube and stored briefly at 4°C. After sample collection, density gradient centrifugation is performed within 2 hours, using Ficoll solution with a density of 1.077 g / mL as the separation medium, and centrifuging at 800g for 20 minutes to obtain the mononuclear cell layer, which is then washed twice with PBS buffer to remove residual plasma proteins. After collecting the obtained cell precipitate, DNA extraction is immediately performed, or it is stored at -80°C for standby, thereby forming a stratified age sample data set.
[0070] Step S32: Determine the methylation level of the samples in each age interval according to the stratified age sample data and candidate marker site data, thereby obtaining age interval methylation profile data;
[0071] The methylation level is determined by adopting bisulfite conversion combined with pyrosequencing, using the layered age sample data obtained in step S31 and the candidate marker site data obtained in step S243. The specific operation includes: taking 500 ng of genomic DNA, using EZ DNA Methylation-Gold Kit for bisulfite conversion, and then using PCR to amplify the target CpG site region, and at least 30 CpG sites are amplified for each sample. The product after PCR amplification is subjected to pyrosequencing, and the methylation level (methylation cytosine proportion) of each CpG site is calculated. All data are analyzed using Bio-Rad CFX Manager software, and samples with low sequencing quality (such as signal peak lower than 2 times background noise) are removed, and finally the methylation spectrum data of each age interval is obtained.
[0072] Step S33: The methylation spectrum data of the age interval is subjected to time series analysis processing, and the slope of the methylation level of each CpG site with respect to age is calculated, thereby obtaining methylation dynamic change data;
[0073] The methylation spectrum data of the age interval obtained in step S32 is used to evaluate the methylation level change trend of the CpG site by adopting a time series analysis method. First, the LOESS (local weighted regression) method is used to smooth the methylation data of different age intervals to reduce the influence of experimental errors. Then, a linear regression model is used to calculate the change slope of the methylation level of each CpG site with respect to age increase, wherein the change slope is calculated in the following manner: taking the age interval as the independent variable and the methylation level as the dependent variable, a straight line is fitted by adopting the least square method, and the regression coefficient is calculated. A positive slope indicates that the methylation level increases with age, and a negative slope indicates that the methylation level decreases with age, and finally the methylation dynamic change data is obtained.
[0074] Step S34: According to the methylation dynamic change data, CpG sites with a monotonically increasing or monotonically decreasing methylation level change trend are screened, thereby obtaining monotonous change site data;
[0075] The methylation dynamic change data obtained in step S33 is used to screen CpG sites with a monotonically increasing or monotonically decreasing methylation level change trend. First, Spearman rank correlation analysis is used to calculate the correlation coefficient of the methylation level of each CpG site with respect to age, and sites with a correlation coefficient greater than 0.6 or less than -0.6 are screened to ensure the stability of the trend. Subsequently, Mann-Kendall trend test is used to statistically test the monotonicity of the methylation level, and the significance level is set to P<0.05, and only CpG sites with a significant monotonous change trend are retained, and finally a monotonous change site data set is formed.
[0076] Step S35: Change amplitude calculation processing is performed on the monotonous change site data to calculate the difference percentage between the maximum methylation level and the minimum methylation level in the 20-80 age group, thereby obtaining site change amplitude data.
[0077] The embodiment of the present application takes the monotonous change site data obtained in step S34 to calculate the methylation level change amplitude of each CpG site in the 20-80 age group. First, the average methylation level of each CpG site in the 20-year-old group and the 80-year-old group is calculated respectively, then the absolute difference between the two is calculated, and the difference is divided by the methylation level of the 20-year-old group to obtain the percentage change amplitude. In order to ensure the reliability of the data, the sites with large fluctuations in the measured values (such as standard deviation greater than 50% of the average value) are removed, and finally the site change amplitude data is obtained.
[0078] Step S36: According to the site change amplitude data, CpG sites with a change amplitude greater than 30% are screened, thereby obtaining age-related site data.
[0079] The embodiment of the present application takes the site change amplitude data obtained in step S35 to screen CpG sites with a methylation level change amplitude greater than 30% to determine age-related characteristic sites. The screening method includes setting the threshold to 30%, using the dplyr package of R language to screen the data, and drawing a density distribution graph to verify the rationality of the screening result. For the screened CpG sites, further regression analysis is performed using a linear mixed effects model to ensure that the methylation level change of these sites is dominated by age, rather than caused by individual differences or technical errors. Finally, an age-related site data set is formed to provide a basis for subsequent biological function research.
[0080] The application can accurately determine the CpG sites closely related to age changes, and provide key data for constructing a biological age prediction model. First, peripheral blood samples of healthy people in each 5-year interval within the age range of 20 to 80 years old are obtained, and the sample size of each age interval is not less than 50, which ensures the representativeness of the samples and the reliability of the data, so that the subsequent analysis can accurately reflect the methylation characteristics of different age groups. Then, the methylation level of each sample in each age interval is determined to obtain age interval methylation spectrum data, which provides basic data for subsequent analysis. Then, the slope of the methylation level of each CpG site with age change is calculated by time series analysis to obtain methylation dynamic change data, which helps to identify the law of methylation level change with age. Subsequently, the CpG sites with monotonically increasing or monotonically decreasing methylation level change trend are screened out, which further narrows down the research scope and focuses on those sites consistent with the age change trend. Then, the difference percentage between the maximum methylation level and the minimum methylation level within the age range of 20 to 80 years old is calculated to obtain site change amplitude data, which quantifies the change amplitude of the methylation level. Finally, the CpG sites with a change amplitude greater than 30% are screened out to obtain age-related site data, which ensures that the selected sites have significant changes in methylation level, so they are more likely to be related to age. Through rigorous data analysis and screening throughout the process, the accuracy and reliability of the research are improved, and strong support is provided for in-depth exploration of the relationship between methylation and age.
[0081] Preferably, step S32 comprises the following steps:
[0082] Step S321: performing genome DNA separation and purification based on the phenol-chloroform method on each sample in the stratified age sample data, to obtain high-quality DNA sample data;
[0083] The peripheral blood sample in the stratified age sample data obtained in step S31 is taken in the embodiment of the application, 1-2 mL of whole blood is taken from each sample, an equal volume of cell lysis solution (containing 1% SDS and 10 mM Tris-HCl) is added, and incubation is performed in a 37°C water bath for 30 minutes to fully lyse the cells. Subsequently, an equal volume of phenol-chloroform-isoamyl alcohol (volume ratio of 25:24:1) is added, and after vigorous shaking for 1 minute, centrifugation is performed at 12,000 g for 10 minutes, and the supernatant is taken to repeat the step 2 times to remove proteins and other impurities. Finally, 2.5 times the volume of anhydrous ethanol and 0.1 times the volume of 3M sodium acetate are added to precipitate the DNA, and after standing at -20°C for 1 hour, the DNA precipitate is recovered by centrifugation at 12,000 g for 15 minutes, and after washing twice with 70% ethanol, the DNA is dissolved in a TE buffer. The DNA concentration is detected using a Nanodrop spectrophotometer, and the A260 / A280 ratio is between 1.8 and 2.0. The DNA sample is used for subsequent experiments, and finally high-quality DNA sample data is obtained.
[0084] Step S322: Bisulfite conversion is performed on the high-quality DNA sample data, and unmethylated cytosine is converted into uracil in an alkaline environment, so as to obtain converted DNA data;
[0085] In the embodiment of the present application, 500 ng of DNA is taken from the high-quality DNA sample data obtained in step S321, and bisulfite conversion is performed using EZ DNA Methylation-Gold Kit. First, CT conversion reagent is added, incubated at 55°C for 10 minutes, denatured at 95°C for 30 seconds, and the cycle is repeated 16 times to ensure that the unmethylated cytosine is fully converted into uracil. Subsequently, the sample is placed on a Zymo-Spin IC column for washing and purification, and DNA is eluted using M-Elution Buffer, and finally converted DNA data is obtained. To verify the conversion efficiency, a plurality of CpG sites are selected for pyrosequencing to ensure that the conversion rate of unmethylated cytosine is greater than 98%.
[0086] Step S323: The converted DNA data is selectively amplified by methylation-specific primers to obtain target region amplification product data;
[0087] In the embodiment of the present application, the converted DNA data obtained in step S322 is selectively amplified by methylation-specific PCR (MSP) method. First, specific primers for methylated and unmethylated sequences are designed to ensure that the amplicon length is within the range of 150-300 bp to meet the requirements of high-throughput sequencing. The PCR reaction system includes: 12.5 μL 2×Taq PCR Master Mix, 1 μL 10 μM primer, 2 μL converted DNA, 9.5 μL ddH2O, total volume 25 μL, and the PCR cycle conditions are set as: 95°C pre-denaturation for 5 minutes, 95°C denaturation for 30 seconds, 55°C annealing for 30 seconds, 72°C extension for 30 seconds, a total of 35 cycles, and finally 72°C extension for 5 minutes. The amplification product is detected by 1.5% agarose gel electrophoresis, and after confirming that the band is clear and there is no non-specific amplification, the AMPure XP magnetic beads are used for purification and quantification, and finally the target region amplification product data is obtained.
[0088] Step S324: High-throughput sequencing library construction is performed on the target region amplification product data, so as to obtain sequencing library data, wherein the high-throughput sequencing library construction includes end repair, adapter ligation and fragment selection;
[0089] The embodiment of the present application takes the target region amplification product data obtained in step S323 to construct a high-throughput sequencing library. First, end repair is performed, and the PCR amplification product is incubated at 37 DEG C for 30 minutes to phosphorylate the 5' end and fill in the 3' end protruding bases. Then, the adapter ligase and specific adapter sequence are added, and the DNA fragments are incubated at 16 DEG C for 16 hours to connect the sequencing adapter to the DNA fragments for sequencing. After the adapter ligation, the DNA fragments are purified using AMPure XP magnetic beads, and the fragment screening is performed by a DNA fragment analyzer, and the DNA fragments of 150-350 bp are retained to improve the sequencing uniformity, and finally the sequencing library data is obtained.
[0090] Step S325: The library concentration determination and fragment size distribution evaluation are performed on the sequencing library data, and the library meeting the preset quality control standard is screened to obtain qualified library data;
[0091] The embodiment of the present application takes the sequencing library data obtained in step S324 to perform library concentration determination and fragment size distribution evaluation. First, the Qubit dsDNA HS Assay Kit is used to determine the DNA concentration, and the library concentration is ensured to be in the range of 10-50 ng / μL to meet the sequencing requirements. Then, the Agilent 2100 bioanalyzer is used for fragment size detection to confirm that the library fragments are concentrated in the range of 150-350 bp, and there is no excessive adapter dimer contamination. For the samples with low library concentration or abnormal fragment distribution, the library needs to be reconstructed or subjected to library enrichment treatment, and only the library data meeting the preset quality control standard is retained to finally obtain qualified library data.
[0092] Step S326: The qualified library data is subjected to depth sequencing processing to ensure that the sequencing depth of each candidate marker site is not less than 30X to obtain raw sequencing data;
[0093] The embodiment of the present application takes the qualified library data obtained in step S325 to perform sequencing using the Illumina NovaSeq 6000 high-throughput sequencing platform. First, the library is pooled according to the library concentration requirement, and the Illumina NextSeq500 / 550 High Output Kit is used to prepare the sequencing reaction system. The sequencing parameter is set to PE150, i.e. double-end sequencing, and the sequencing depth of each candidate marker site is not less than 30X to ensure the data quality. During the sequencing process, the sequencing error rate of each base needs to be controlled below 1%, and the base quality value (Q30) needs to be greater than 85%. After the sequencing is completed, the data is split and the low-quality bases are removed to finally obtain the raw sequencing data.
[0094] Step S327: Bioinformatics analysis is performed on the raw sequencing data to obtain age interval methylation profile data, wherein the bioinformatics analysis includes sequence alignment, methylation level quantification and data standardization.
[0095] The raw sequencing data obtained in step S326 is subjected to bioinformatics analysis in the embodiment of the present application. First, the data quality is evaluated using the FastQC tool, and low-quality sequences and adapter contamination are removed using Trim Galore. Then, the sequencing data is aligned to the reference genome GRCh38 using Bismark software, and the matching rate needs to be greater than 80%. After alignment, the CpG site methylation information is extracted, and the methylation proportion, i.e., the ratio of C / C+T, is calculated to obtain methylation level quantification data. Finally, the data is normalized using the Beta distribution standardization method to eliminate experimental batch effects and sequencing errors, and finally age interval methylation profile data is obtained.
[0096] The application plays a key role in determining age-related CpG sites and provides important data support for constructing a biological age prediction model. First, the genomic DNA of each sample in the stratified age sample data is separated and purified based on the phenol-chloroform method, which can effectively remove impurities and obtain high-quality DNA sample data, ensuring the accuracy and reliability of subsequent analysis. Then, the high-quality DNA sample data is treated with bisulfite conversion, which converts unmethylated cytosine to uracil in an alkaline environment, obtaining converted DNA data. This step is a key step in detecting DNA methylation status and lays a foundation for subsequent methylation-specific amplification and sequencing. Then, the methylation-specific primer is used to selectively amplify the candidate marker site region, obtaining the target region amplification product data, which enables the study to focus on specific gene regions that may be related to age changes, improving the specificity and efficiency of the analysis. Next, high-throughput sequencing library construction processing is performed, including end repair, adapter ligation and fragment selection steps, obtaining sequencing library data. This step ensures the quality of the sequencing library and meets the requirements of high-throughput sequencing, providing high-quality templates for subsequent deep sequencing. The library concentration and fragment size distribution of the sequencing library data are determined, and the library that meets the preset quality control standards is selected to obtain qualified library data. This step further ensures the quality of the sequencing data and avoids sequencing errors and data bias caused by library quality problems. Finally, the qualified library data is subjected to deep sequencing processing to ensure that the sequencing depth of each candidate marker site is not less than 30X, obtaining the original sequencing data. Then, through bioinformatics analysis processing, including sequence alignment, methylation level quantification and data standardization steps, age interval methylation spectrum data is obtained, which enables the study to accurately determine the methylation level of each CpG site in different age intervals, providing accurate data basis for screening age-related CpG sites. The entire process is through a series of rigorous operations and quality control measures to ensure the high quality and reliability of the data, providing a strong guarantee for in-depth study of the relationship between DNA methylation and age change.
[0097] Preferably, step S4 comprises the following steps:
[0098] Step S41: Obtain known human aging-related gene database information, and perform annotation analysis processing on the promoter region of the human aging-related gene database information, to obtain aging gene promoter region positioning data;
[0099] The embodiment of the present application obtains gene information closely related to the human aging process from a public human aging-related gene database (such as GenAge, Human Ageing Genomic Resources, etc.), including gene name, gene function description, and reported aging regulation mechanism. Then, the promoter region of each aging-related gene is functionally annotated using a genome annotation tool (such as Homer, UCSC Genome Browser, or Ensembl), and the transcription start site (TSS) is usually selected as the promoter region, and the distribution of CpG islands in the region, transcription factor binding sites, and other key features are analyzed. Finally, the annotation data is integrated to construct precise positioning information of the aging gene promoter region, and the aging gene promoter region positioning data for subsequent analysis is output.
[0100] Step S42: According to the aging gene promoter region positioning data and the age-related site data, the CpG sites located within the range of 1000 base pairs upstream and downstream of the aging gene promoter are screened and processed, thereby obtaining the aging gene-related site data;
[0101] Based on the aging gene promoter region positioning data obtained in step S41, the embodiment of the present application combines the published age-related DNA methylation site data set (such as the results of EWAS research or the DNA methylation data of different age groups of samples in the TCGA database), and uses a genome coordinate alignment algorithm (such as BEDTools intersect or liftOver tool) to map the age-related sites to the range of 1000 base pairs upstream and downstream of the aging gene promoter region. Then, the CpG sites located in the range are screened out, and further exclude sites with sequencing coverage less than 30X or large methylation level variability between samples to ensure the reliability of the data. Finally, the aging gene-related CpG sites that meet the screening criteria are obtained, and the aging gene-related site data is output.
[0102] Step S43: Perform Pearson correlation analysis on the methylation level of each CpG site in the aging gene-related site data and the expression amount of the corresponding gene, thereby obtaining the methylation-expression correlation data;
[0103] The embodiment of the application obtains the aging gene related site data obtained in step S42, acquires gene expression data matched with the corresponding sample from a public database (such as GTEx, TCGA or GEO), and ensures that the sample source, age information and tissue type are consistent. Then, the methylation level of each CpG site and the expression amount of the gene adjacent to the CpG site are subjected to Pearson correlation analysis, the correlation coefficient is calculated, and it is judged whether the change in the methylation level is significantly correlated with the gene expression level. In the specific calculation, the number of samples is not less than 50, and a significance test threshold (such as p value < 0.05) is set. Finally, the calculated CpG methylation-gene expression correlation coefficient is sorted into methylation-expression correlation data, and output for subsequent analysis.
[0104] Step S44: According to the methylation-expression correlation data, the CpG sites with an absolute value of the correlation coefficient greater than 0.6 are screened, and thus functional related site data is obtained.
[0105] Based on the methylation-expression correlation data obtained in step S43, the embodiment of the application screens the CpG sites with an absolute value of the correlation coefficient greater than 0.6, and ensures that there is a strong correlation between the methylation change of these sites and the gene expression. In the specific operation, the correlation data is sorted first, and the sites with no significant correlation (such as sites with a p value greater than 0.05) are removed, and then the sites with an absolute value of the correlation coefficient greater than or equal to 0.6 are extracted as the functional related site data. This screening step helps to remove weakly correlated or irrelevant CpG sites, and improves the accuracy of subsequent analysis. Finally, the functional related site data meeting the screening standard is output.
[0106] Step S45: The functional related site data is subjected to principal component analysis processing, the contribution of each CpG site to the sample age variation is calculated, and thus site weight data is obtained.
[0107] The embodiment of the application uses the functional related site data obtained in step S44, and uses the principal component analysis (PCA) method to calculate the contribution of each CpG site to the sample age variation. First, a samplexCpG methylation level matrix is constructed, and the data is subjected to standardization processing (such as Z-score normalization), so as to ensure that the methylation levels of different sites are compared on the same scale. Then, the PCA method is used to extract the main components, and the contribution of each CpG site to the first principal component (PC1) is calculated as the influence degree of the site on the age variation. Generally, the variance explanation rate of the first two principal components reaches more than 80%, so as to ensure that the selected sites can effectively represent the age variation. Finally, the contribution data of each CpG site, i.e. the site weight data, is output.
[0108] Step S46: According to the site weight data, the CpG sites with a contribution degree in the top 50% are selected, and thus core age marker site data is obtained.
[0109] The embodiment of the present application sorts the site weight data obtained in step S45 according to the contribution degree from high to low, and selects the top 50% CpG sites as the core age marker sites. In the specific screening process, it is ensured that the selected CpG sites have relatively stable methylation variation trends in different age group samples, and the applicability thereof in an independent data set is verified (such as cross-validation using another independent sample). Finally, the screened core age marker site data can be used for subsequent construction of an age prediction model and applied to biological age assessment or research on aging-related diseases.
[0110] The present application plays a key role in the process of screening core age marker sites, and provides important support for constructing a precise biological age prediction model. First, the information of a known human aging-related gene database is obtained, and the promoter region is annotated and analyzed to obtain aging gene promoter region positioning data, which provides precise positioning information for subsequent screening of gene sites closely related to aging. Then, according to the aging gene promoter region positioning data and the age-related site data, the CpG sites located within the range of 1000 base pairs upstream and downstream of the aging gene promoter are positioned and screened to obtain aging gene related site data. This step focuses on the key regulatory region of the aging gene, improving the pertinence and effectiveness of the screening. Then, the methylation level of each CpG site in the aging gene related site data and the expression amount of the corresponding gene are subjected to Pearson correlation analysis to obtain methylation-expression correlation data, which helps to reveal the relationship between methylation level change and gene expression, and provides a basis for screening functionally related sites. Subsequently, according to the methylation-expression correlation data, CpG sites with an absolute value of correlation coefficient greater than 0.6 are screened to obtain functionally related site data, which ensures that the selected sites have significant correlation between methylation level and gene expression, further improving the functionality and prediction value of the sites. Then, the functionally related site data is subjected to principal component analysis to calculate the contribution degree of each CpG site to the age variation of the sample to obtain site weight data, which enables the study to quantify the importance of each site in age prediction and provides a key parameter for constructing a prediction model. Finally, according to the site weight data, the top 50% CpG sites in terms of contribution degree are selected to obtain core age marker site data, which selects the most representative and predictive sites, lays a solid foundation for constructing an accurate and reliable biological age prediction model, and the entire process ensures that the final core age marker site data has high accuracy and reliability through a series of rigorous analysis and screening steps, providing a powerful tool for in-depth study of aging mechanisms and prediction of individual biological age.
[0111] Preferably, step S45 comprises the following steps:
[0112] Step S451: transforming the methylation levels of each CpG site in the function-related site data into a standard normal distribution to obtain standardized methylation data;
[0113] The function-related site data obtained in step S44 of the embodiment of the present application contains the CpG site methylation level data of multiple samples. Since the methylation levels of different CpG sites have different value ranges, which may affect subsequent analysis, it is necessary to standardize the methylation levels. The specific method is as follows: first, calculate the mean and standard deviation of each CpG site in all samples, then standardize the data of all samples of the site, that is, subtract the mean from the original methylation level and divide by the standard deviation, so that the methylation level of each CpG site conforms to the standard normal distribution with a mean of 0 and a standard deviation of 1. This process can be implemented using statistical analysis tools (such as the scale function of R or the StandardScaler of Python). Finally, standardized methylation data is obtained to ensure that the data of different CpG sites are in the same scale range for principal component analysis.
[0114] Step S452: constructing a sample covariance matrix according to the standardized methylation data, and performing eigenvalue decomposition processing on the covariance matrix to obtain eigenvector data;
[0115] The embodiment of the present application constructs a sample covariance matrix based on the standardized methylation data obtained in step S451 to characterize the linear correlation of CpG site methylation levels between samples. The specific method is as follows: first, represent the standardized data as a matrix, where the rows represent samples and the columns represent CpG sites; then calculate the covariance matrix, where each matrix element represents the covariance between two CpG sites, i.e., whether their methylation levels have correlation. The covariance matrix is usually symmetric, and the diagonal elements of the matrix represent the variances of the sites themselves, and the non-diagonal elements represent the covariances between different sites. Next, perform eigenvalue decomposition on the covariance matrix, i.e., calculate the eigenvalues and eigenvectors of the matrix, where the eigenvalues represent the variance information of the data in the corresponding eigenvector direction, and the eigenvectors define the direction of the new coordinate axis. Eigenvalue decomposition can be implemented through numerical calculation library (such as linalg.eig of NumPy or eigen function of R). Finally, eigenvector data is obtained for subsequent principal component transformation.
[0116] Step S453: performing principal component transformation processing on the standardized methylation data according to the eigenvector data, projecting the original data into a new feature space to obtain principal component score data;
[0117] The embodiment of the application performs principal component conversion on the standardized methylation data based on the feature vector data obtained in step S452, that is, projects the original data into a new feature space. The specific operation is as follows: first, the standardized methylation data matrix is multiplied by the feature vector matrix to obtain the projection value of the sample in the principal component direction, that is, the principal component score. This projection process is equivalent to re-describing the original CpG methylation data as a weighted linear combination of different principal components, so that the new coordinate system can better explain the variance of the data. The calculation of the principal component score can be realized by using a matrix operation tool (such as the dot function of Python or the % * % operator of R). Finally, the score of each sample on each principal component, that is, the principal component score data, is obtained, which provides a basis for subsequent calculation of the principal component importance.
[0118] Step S454: Perform variance contribution rate calculation processing on the principal component score data to determine the degree of explanation of each principal component to the overall variation, thereby obtaining the principal component importance data;
[0119] The embodiment of the application calculates the variance contribution rate of each principal component based on the principal component score data obtained in step S453 to determine the degree of explanation of different principal components to the overall data variation. The specific method is as follows: first, the variance of all principal component scores is calculated, and the proportion of the variance of each principal component to the total variance is calculated, which is the variance contribution rate of the principal component. The higher the variance contribution rate, the more the principal component can explain the overall variability of the data. Then, the variance contribution rates are sorted in descending order, and the cumulative variance contribution rate is calculated to observe whether the first few principal components can explain most of the variation of the data. Generally, when the cumulative variance contribution rate reaches more than 85%, it is considered that the selected principal components are sufficient to represent the main features of the data. Finally, the principal component importance data is obtained, which provides a basis for subsequent selection of key principal components.
[0120] Step S455: Select the first N principal components with a cumulative variance contribution rate of 85% according to the principal component importance data, thereby obtaining the key principal component data;
[0121] The embodiment of the application sorts the principal component importance data obtained in step S454 according to the variance contribution rate from high to low, and selects the first N principal components with a cumulative variance contribution rate of 85% as the key principal components. The specific method is as follows: first, the variance contribution rate of each principal component is calculated and arranged in descending order; then, the variance contribution rates of the first few principal components are accumulated until the sum reaches 85%; finally, the number N of principal components included when the sum of the contribution rates reaches 85% is selected, and the data of the N principal components is extracted. This process can be realized by using a programming language (such as the cumsum function of Python or the cumsum of R). Finally, the key principal component data is obtained, which provides a basis for subsequent calculation of the CpG site weight.
[0122] Step S456: Calculate the loading coefficients of each CpG site on the principal components according to the key principal component data, and take the sum of the squares of the loading coefficients as the contribution of the site to the sample age variation, thereby obtaining site weight data.
[0123] Based on the key principal component data obtained in step S455, the loading coefficients of each CpG site on the principal components are calculated, and the contribution of the site to the sample age variation is measured according to the sum of the squares of the loading coefficients. The specific method is as follows: first, extract the principal component loadings of each CpG site from the key principal component data, that is, the projection coefficients of each CpG site on the key principal components. The loading coefficient represents the influence degree of the site on the corresponding principal component; then, square the loading coefficient of each CpG site, and calculate the sum of the squares on all key principal components to obtain the comprehensive contribution of the CpG site. The larger the value, the stronger the site's ability to explain the age variation. Finally, normalize the contribution of all sites to facilitate subsequent screening and sorting. Finally, the site weight data of each CpG site is obtained, which can be used for subsequent selection of the most representative core age marker site.
[0124] The present application plays a key role in determining the core age marker site and provides important data support for constructing a biological age prediction model. First, the methylation levels of each CpG site in the functionally related site data are converted into standard normal distribution to obtain standardized methylation data. This step eliminates the differences in methylation levels between different samples, making the data more comparable. Next, a sample covariance matrix is constructed from the standardized methylation data, and eigenvalue decomposition is performed to obtain eigenvector data, which helps to reveal the main variation direction in the data. Then, the principal component transformation is performed on the standardized methylation data using the eigenvector data, and the original data is projected into a new feature space to obtain principal component score data. This step further simplifies the data structure and highlights the main variation information. After that, the variance contribution rate of the principal component score data is calculated to determine the explanation degree of each principal component to the overall variation, and the principal component importance data is obtained, which enables the study to identify the principal components with the largest contribution to the data variation. According to the principal component importance data, the first N principal components with a cumulative variance contribution rate of 85% are selected to obtain key principal component data. This step ensures that the selected principal components can explain most of the data variation. Finally, the loading coefficients of each CpG site on the principal components are calculated according to the key principal component data, and the sum of the squares of the loading coefficients is taken as the contribution of the site to the sample age variation to obtain site weight data. This step quantifies the importance of each CpG site in age prediction and provides a basis for screening core age marker sites. Through a series of rigorous mathematical and statistical methods, the key information in the data is effectively extracted, and the prediction ability and interpretability of the model are improved.
[0125] Preferably, step S5 comprises the following steps:
[0126] Step S51: randomly divide the samples into a training set and a test set according to the core age marker site data, wherein the training set accounts for 70% and the test set accounts for 30%, so as to obtain model training data and model validation data;
[0127] According to the core age marker site data obtained in step S50, the samples are randomly divided into a training set and a test set for model training and validation, respectively. In actual operation, first, the core age marker site data of all samples are stored in a matrix by rows, wherein each row represents a sample and the columns are different CpG sites. When the data set is divided, 70% of the samples are used as the training set and the remaining 30% are used as the test set. To ensure the randomness of data allocation, a random number generator (such as the train_test_split function in Python) can be used to randomly allocate the samples. The training set is used to train the model, and the test set is used to evaluate the performance of the model after the model training is completed. To ensure that the samples in the training set and the test set have similar age distribution, stratified sampling method can be used to ensure consistent age distribution, so as to avoid the influence of uneven age distribution of the training set and the test set on the model effect.
[0128] Step S52: weighting processing is performed on the methylation level in the core age marker site data, so as to obtain cell type weight data, wherein the weighting processing is specifically to determine the optimal weight coefficient of each cell type by minimizing the mean square error of the predicted age and the actual age;
[0129] In the core age marker site data obtained in step S51, the methylation level of each sample is associated with its corresponding cell type, so weighting processing needs to be performed on the methylation data to adjust the influence of the model according to the difference of the cell types. Specifically, first, the weight coefficient of each cell type is determined by minimizing the mean square error (MSE) between the predicted age and the actual age. In this process, the methylation data of each cell type is used as the input feature, and the weight coefficient of each cell type is calculated by linear regression method. These weight coefficients can reflect the contribution of each cell type to the predicted age, and the cell type weight data is obtained after calculation. Through this method, the influence of the cell types is effectively quantified, and the influence degree of each cell type can be automatically optimized according to the actual situation of the data.
[0130] Step S53: using elastic network regression algorithm to construct a prediction model according to the cell type weight data and the model training data, and determining the optimal regularization parameter through cross-validation, so as to obtain initial prediction model data;
[0131] The embodiment of the present application utilizes the cell type weight data obtained in step S52 and the model training data obtained in step S51 to construct a prediction model by using an elastic network regression algorithm (Elastic Net Regression). The elastic network regression combines the advantages of Lasso regression and Ridge regression, and can perform variable selection and prevent overfitting when facing high-dimensional data. In the implementation process, first, the cell type weight data is taken as a feature, and the predicted age in the model training data is taken as a label, and then input into the elastic network regression algorithm. By optimizing the regularization parameter, the model can automatically select important features and eliminate redundant features. In order to determine the optimal regularization parameter, the method of cross-validation is adopted, and K-fold cross-validation is performed on the training set. Different parameter combinations are selected during each training, the performance index of the model is calculated, and the parameters that make the performance of the model optimal are selected. Finally, the initial prediction model data is obtained, and the model contains the importance and weight of each cell type in age prediction.
[0132] Step S54: According to the model verification data, the performance evaluation process of the initial prediction model data on the test set is performed, the correlation coefficient, mean square error and consistency index of the predicted age and the actual age are calculated, and the model evaluation data is obtained.
[0133] After obtaining the initial prediction model in step S53, the embodiment of the present application uses the test set data segmented in step S51 to evaluate the performance of the model. During the evaluation, first, the samples in the test set are input into the initial prediction model to obtain the predicted age. Then, the correlation coefficient, mean square error (MSE) and consistency index (such as Kappa coefficient or other statistical consistency index) between the predicted age and the actual age are calculated to measure the accuracy and consistency of the model. The closer the correlation coefficient is to 1, the stronger the linear relationship between the predicted age and the actual age; the smaller the mean square error, the smaller the error of the prediction result; and the consistency index represents the prediction consistency of the model between different samples. Through these indexes, the model evaluation data is obtained, and the effect and reliability of the model in actual application are evaluated.
[0134] Step S55: According to the model evaluation data, the biological age acceleration is calculated and standardized, and the age acceleration data is obtained, wherein the biological age acceleration is specifically the difference between the predicted age and the actual age.
[0135] The embodiment of the present application calculates the biological age acceleration according to the model evaluation data obtained in step S54, and performs standardization processing. The biological age acceleration is represented as the difference between the predicted age and the actual age. The specific steps are as follows: for each sample, calculate the difference between the predicted age and the actual age, that is, the biological age acceleration. Then, the biological age acceleration data of all samples is standardized, usually using Z-score standardization, that is, the biological age acceleration of each sample is subtracted from the average of all samples, and then divided by the standard deviation, so that the biological age acceleration conforms to the standard normal distribution. The standardized data can help identify samples that deviate from the normal aging rate, thereby providing a basis for subsequent abnormal aging markers.
[0136] Step S56: samples deviating from the mean by more than two standard deviations in the age acceleration data are marked as abnormal aging, thereby obtaining a biological age prediction model.
[0137] The embodiment of the present application marks samples deviating from the mean by more than two standard deviations in the standardized age acceleration data obtained in step S55 as abnormal aging. The specific method is as follows: first, calculate the mean and standard deviation of the standardized age acceleration of all samples, then judge whether the difference between the standardized age acceleration value of each sample and the mean is greater than two standard deviations. If the age acceleration of a certain sample exceeds this range, it indicates that the aging rate of the sample is abnormal, and there may be a risk of biological abnormal aging, so it is marked as abnormal aging. Finally, a biological age prediction model is obtained, which can assist medical researchers or clinicians to identify individuals who may have accelerated aging according to the age acceleration and its abnormal markers, so as to perform early intervention and treatment.
[0138] The application plays a crucial role in constructing a biological age prediction model, providing a strong guarantee for accurate age prediction. First, the samples are randomly divided into a training set and a test set, with the training set accounting for 70% and the test set accounting for 30%, which enables the model to be trained on a larger data set while retaining a portion of the data for validating the performance of the model, ensuring the generalization ability of the model. Next, the methylation level in the core age marker site data is weighted, and the optimal weight coefficient of each cell type is determined by minimizing the mean square error of the predicted age and the actual age, which enables the model to fully consider the contribution of different cell types to age prediction and improves the accuracy of prediction. Then, the elastic network regression algorithm is used to construct a prediction model according to the cell type weight data and the model training data, and the optimal regularization parameter is determined through cross-validation, which further optimizes the performance of the model and avoids problems such as overfitting, so that the performance of the model on the training set can be better generalized to the test set. Subsequently, the performance of the initial prediction model data on the test set is evaluated according to the model validation data, and the correlation coefficient, mean square error and consistency index of the predicted age and the actual age are calculated, which comprehensively evaluates the prediction performance of the model and provides an important reference for subsequent model optimization. According to the model evaluation data, the biological age acceleration is calculated and standardized to obtain the age acceleration data, which provides a quantitative indicator for evaluating the degree of aging of individuals. Finally, the samples deviating from the mean by more than two standard deviations in the age acceleration data are marked as abnormal aging, thereby obtaining the biological age prediction model, which enables the model to identify individuals with abnormal aging speed and provides an important tool for medical research and health management. The entire process ensures the accuracy and reliability of the biological age prediction model through rigorous data division, weight assignment, model construction, performance evaluation and abnormal identification, and provides strong support for in-depth research on aging mechanisms and prediction of individual biological age.
[0139] Therefore, the embodiments should be considered in all respects as illustrative and not restrictive, the scope of the application being defined by the appended claims rather than the description presented above, and all changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.
[0140] The above description is merely that of a specific implementation of the present application, which enables those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application shall not be limited to the embodiments shown herein, but shall conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a biological age prediction model based on DNA methylation, characterized in that, Includes the following steps: Step S1: Obtain whole-genome methylation sequencing data from human peripheral blood samples; Human peripheral blood samples were classified into cell subpopulations, and single-cell RNA sequencing was performed on each cell subpopulation to obtain single-cell sequencing data including lymphocytes, neutrophils, and monocytes. Step S2: Based on single-cell sequencing data and whole-genome methylation sequencing data, tissue-specific analysis was performed on CpG sites within 2000 base pairs upstream and downstream of the transcription factor binding sites of each cell subpopulation to obtain candidate marker site data. Step S3: Perform dynamic evolution analysis on the candidate marker site data to screen out CpG sites whose methylation levels show a monotonically increasing or monotonically decreasing trend within the 20-80 age group and whose change is greater than 30%, thus obtaining age-related site data. Step S3 includes the following steps: Step S31: Obtain peripheral blood samples from healthy individuals aged 20-80 years, with each age range divided into 5-year intervals. The sample size for each age interval should be no less than 50 cases, thereby obtaining stratified age sample data. Step S32: Based on the stratified age sample data and candidate marker site data, the methylation level of the samples in each age range is determined to obtain the methylation spectrum data of the age range. Step S33: Perform time series analysis on the methylation spectrum data of the age range, calculate the slope of the methylation level of each CpG site with age, and thus obtain the dynamic change data of methylation. Step S34: Based on the dynamic changes in methylation, screen out CpG sites whose methylation levels show a monotonically increasing or monotonically decreasing trend, thereby obtaining monotonically changing site data; Step S35: Perform variation amplitude calculation on the monotonic variation site data, calculate the percentage difference between the maximum and minimum methylation levels in the 20-80 age group, and thus obtain the site variation amplitude data; Step S36: Based on the locus change magnitude data, filter CpG loci with a change magnitude greater than 30% to obtain age-related locus data; Step S4: Based on the pre-acquired promoter region information of aging genes and age-related site data, CpG sites located within 1000 base pairs upstream and downstream of the aging gene promoter and with an absolute correlation coefficient between methylation level and gene expression level greater than 0.6 are screened to obtain core age marker site data. Step S4 includes the following steps: Step S41: Obtain known information from the human aging-related gene database and perform annotation analysis on the promoter regions of the human aging-related gene database information to obtain the promoter region localization data of aging genes. Step S42: Based on the aging gene promoter region localization data and age-related site data, CpG sites located within 1000 base pairs upstream and downstream of the aging gene promoter are localized and screened to obtain aging gene-related site data. Step S43: Perform Pearson correlation analysis on the methylation level of each CpG site in the aging gene-related site data and the expression level of the corresponding gene to obtain methylation-expression correlation data; Step S44: Based on the methylation-expression correlation data, CpG sites with an absolute correlation coefficient greater than 0.6 are screened to obtain functionally related site data; Step S45: Perform principal component analysis on the functionally relevant locus data to calculate the contribution of each CpG locus to the age variation of the sample, thereby obtaining the locus weight data; Step S46: Select the top 50% of CpG sites based on the site weight data to obtain the core age marker site data; Step S5: Construct a biological age prediction model based on core age marker loci data using a cell type-weighted elastic network regression algorithm.
2. The method for constructing a biological age prediction model based on DNA methylation according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain human peripheral blood samples and separate the human peripheral blood samples using density gradient centrifugation to obtain mononuclear cell layer separation data; Step S12: Using bisulfite sequencing, the DNA is sequenced at the whole genome methylation level based on the mononuclear cell layer isolation data to obtain whole genome methylation sequencing data; Step S13: Based on the mononuclear cell layer separation data, different cell subpopulations are enriched and separated using magnetic bead sorting to obtain preliminary sorted subpopulation data; Step S14: Perform flow cytometry purity verification on the preliminary sorted subpopulation data and screen cell populations with a purity greater than 95% to obtain high-purity cell subpopulation data. Step S15: Perform transcriptome sequencing on each cell subpopulation based on the high-purity cell subpopulation data to obtain raw single-cell transcriptome data; Step S16: Perform data quality control on the raw single-cell transcriptome data to obtain single-cell sequencing data. The data quality control includes removing sequencing adapter sequences, filtering low-quality reads, and correcting PCR amplification bias.
3. The method for constructing a biological age prediction model based on DNA methylation according to claim 2, characterized in that, Step S13 includes the following steps: Based on the mononuclear cell layer separation data, different cell subpopulations were enriched and separated using magnetic bead sorting to obtain preliminary sorted subpopulation data. The enrichment and separation process specifically included enriching lymphocytes with CD3 magnetic beads, enriching mononuclear cells with CD14 magnetic beads, and enriching neutrophils with CD66b magnetic beads.
4. The method for constructing a biological age prediction model based on DNA methylation according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Based on single-cell sequencing data, differentially expressed genes of each cell subpopulation are analyzed, and differentially expressed genes with an enrichment fold greater than 2 and a P value less than 0.01 are screened to obtain cell subpopulation-specific gene data. Step S22: Based on the JASPAR database, predict transcription factor binding sites for cell subpopulation-specific gene data to obtain potential transcription factor binding site data; Step S23: Perform ChIP-seq data validation based on potential transcription factor binding site data, and screen sites with binding strength scores greater than 0.8 to obtain confirmed transcription factor binding site data. Step S24: Based on the confirmed transcription factor binding site data and whole-genome methylation sequencing data, tissue-specific analysis was performed on CpG sites within 2000 base pairs upstream and downstream of the transcription factor binding sites of each cell subpopulation to obtain candidate marker site data.
5. The method for constructing a biological age prediction model based on DNA methylation according to claim 4, characterized in that, Step S24 includes the following steps: Step S241: Based on the confirmed transcription factor binding site data and whole-genome methylation sequencing data, CpG sites located within 2000 base pairs upstream and downstream of the transcription factor binding site are mapped to obtain preliminary labeled site data. Step S242: Perform tissue-specific analysis on the preliminary labeled site data, calculate the coefficient of variation of methylation level of each CpG site among different cell subpopulations, and thus obtain site-specific score data; Step S243: Based on the site-specificity score data, CpG sites with a coefficient of variation greater than 0.5 are screened to obtain candidate marker site data.
6. The method for constructing a biological age prediction model based on DNA methylation according to claim 1, characterized in that, Step S32 includes the following steps: Step S321: Perform genomic DNA isolation and purification on each sample in the stratified age sample data using the phenol-chloroform method to obtain high-quality DNA sample data; Step S322: Based on the high-quality DNA sample data, perform bisulfite conversion treatment to convert unmethylated cytosine into uracil under alkaline conditions, thereby obtaining the converted DNA data; Step S323: Selectively amplify candidate marker site regions using methylation-specific primers on the transformed DNA data to obtain amplified product data of the target region; Step S324: Perform high-throughput sequencing library construction processing based on the target region amplification product data to obtain sequencing library data. The high-throughput sequencing library construction processing includes end repair, adapter ligation, and fragment selection. Step S325: Determine the library concentration and evaluate the fragment size distribution of the sequencing library data, and screen libraries that meet the preset quality control standards to obtain qualified library data; Step S326: Perform deep sequencing on the qualified library data to ensure that the sequencing depth of each candidate marker site is not less than 30X, thereby obtaining the raw sequencing data; Step S327: Perform bioinformatics analysis on the raw sequencing data to obtain age-range methylation profile data. The bioinformatics analysis includes sequence alignment, methylation level quantification, and data standardization.
7. The method for constructing a biological age prediction model based on DNA methylation according to claim 1, characterized in that, Step S45 includes the following steps: Step S451: Convert the methylation level of each CpG site in the functionally related site data into a standard normal distribution to obtain standardized methylation data; Step S452: Construct a sample covariance matrix based on the standardized methylation data, and perform eigenvalue decomposition on the covariance matrix to obtain eigenvector data; Step S453: Perform principal component transformation on the standardized methylated data based on the feature vector data, projecting the original data onto the new feature space to obtain the principal component score data; Step S454: Calculate the variance contribution rate of the principal component score data to determine the extent to which each principal component explains the overall variation, thereby obtaining the principal component importance data; Step S455: Select the top N principal components with a cumulative variance contribution rate of 85% based on the principal component importance data to obtain the key principal component data; Step S456: Calculate the loading coefficient of each CpG site on the principal components based on the key principal component data, and use the sum of squares of the loading coefficients as the contribution of the site to the age variation of the sample, thereby obtaining the site weight data.
8. The method for constructing a biological age prediction model based on DNA methylation according to claim 1, characterized in that, Step S5 includes the following steps: Step S51: Based on the core age marker locus data, the samples are randomly divided into a training set and a test set, with the training set accounting for 70% and the test set accounting for 30%, thereby obtaining model training data and model validation data; Step S52: Weight the methylation level in the core age marker site data to obtain cell type weight data. Specifically, the weighting process involves determining the optimal weight coefficient for each cell type by minimizing the mean square error between the predicted age and the actual age. Step S53: Use the elastic network regression algorithm to build a prediction model based on cell type weight data and model training data, and determine the optimal regularization parameter through cross-validation to obtain the initial prediction model data; Step S54: Based on the model validation data, perform performance evaluation processing on the test set for the initial prediction model data, calculate the correlation coefficient, mean square error, and consistency index between the predicted age and the actual age, and thus obtain the model evaluation data; Step S55: Calculate the biological age acceleration based on the model evaluation data and perform standardization to obtain age acceleration data, where the biological age acceleration is specifically the difference between the predicted age and the actual age; Step S56: Mark samples in the age acceleration data that deviate from the mean by more than two standard deviations as abnormal aging, thereby obtaining the biological age prediction model.
Citation Information
Patent Citations
Methylation clock for detecting human biological age and detection method thereof
CN118957042A
DNA methylation biomarker of aging for human ex VIVO and in VIVO studies
US20210381051A1