Peripheral blood DNA methylation model reflecting the thickness of the central zone structure
By constructing a peripheral blood DNA methylation model and using various machine learning algorithms to screen for stable methylation biomarkers, the problem of the inability to accurately predict the thickness of the central region structure in existing technologies has been solved. This has enabled predictions with high sensitivity and high specificity, improving the accuracy and early identification capabilities of non-invasive detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUDAN UNIVERSITY
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-31
AI Technical Summary
Current technologies have not been able to effectively utilize DNA methylation to predict the thickness of the central region structure, making it difficult to accurately assess an individual's brain development and aging trajectory and the risk of neurological diseases in non-invasive testing.
A peripheral blood DNA methylation model was constructed. Through rigorous data preprocessing and feature stability screening, combined with multiple machine learning algorithms, a dual-track prediction system was established to screen out stable combinations of methylation biomarkers to reflect the thickness of the central region structure.
It achieves high sensitivity and high specificity in prediction, reduces false positive and false negative rates, improves the early identification of individual brain development and aging trajectories and the ability to warn of neurological disease risks, and is applicable to accurate prediction of multiple phenotypes.
Smart Images

Figure CN122493941A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to a model for predicting peripheral blood DNA methylation that reflects the thickness of the central region structure. Background Technology
[0002] This 2006 study, published in Cerebral Cortex, used magnetic resonance imaging (MRI) to investigate the anatomical dynamics of the cerebral sulci cortex during aging. The study included 35 healthy elderly individuals aged 59-84 years (mean age 70.9 years), who underwent three imaging scans over four years, with cross-sectional and longitudinal analyses performed. The study focused on the buried cortex of eight sulci: left and right central sulci, frontal sulci, cingulate sulci, and parieto-occipital sulci, analyzing six geometric features including surface area, gray matter thickness, gray matter volume, and sulcus depth. Results showed that cross-sectional analysis confirmed a significant impact of age on mean gray matter thickness, gray matter volume, sulcus depth, and sulcus curvature; longitudinal analysis further confirmed a clear downward trend in these indicators with aging. This study provides direct longitudinal imaging evidence for understanding the degenerative changes in cerebral cortex structure during old age. (Rettmann ME, Kraut MA, Prince JL, Resnick SM. Cross-sectional and longitudinal analyses of anatomical sulcal changes associated with aging.[J]) Cerebralcortex,2006,16(11)).
[0003] The gray matter thickness in the central cortex (including the precentral gyrus, postcentral gyrus, and the area surrounding the central sulcus) shows a longitudinal decreasing trend with age, especially in the elderly population aged 59-84 years, where the average thickness, gray matter volume, and sulcus depth of the central cortex all decrease significantly. (Mousley, A., Bethlehem, RAI, Yeh, FC. et al. Topological turningpoints across the human lifespan. Nat Commun 16, 10055 (2025). https: / / doi.org / 10.1038 / s41467-025-65974-8).
[0004] Objective To explore the value of peripheral blood DNA methylation (DNAm) biological age as a non-invasive molecular biomarker of brain aging. Methods Seventy-nine healthy adults were included. DNAm biological age and aging acceleration were measured. Whole-brain cortical thickness was measured by structural magnetic resonance imaging. The correlation between the two was analyzed and confounding factors were corrected. Results DNAm biological age was highly correlated with chronological age (R²=0.90, P<0.001). After controlling for confounding factors, aging acceleration was significantly negatively correlated with thinning of cortical areas in multiple brain regions. For each year of aging acceleration, the thickness of key brain regions decreased by 0.011–0.024 mm (partial correlation coefficients -0.34 to -0.43, all P<0.01). Conclusion DNAm biological age can independently reflect cerebral cortical aging and provide a non-invasive quantitative biomarker for early warning of neurodegenerative diseases. Proskovec AL, Rezich MT, O'Neill J, et al. Association of EpigeneticMetrics of Biological Age With Cortical Thickness[J]. JAMA Network Open, 2020, 3(9):e2015428. DOI:10.1001 / jamanetworkopen.2020.15428.
[0005] The mechanisms underlying changes in central gyrus thickness are not fully understood, but current research identifies DNA and other components as potential biomarkers. Studies have reported a significant correlation between BDNF methylation (CpG4) and postcentral gyrus thickness in individuals with depression. (Na KS, Won E, Kang J, Chang HS, Yoon HK, Tae WS, Kim YK, Lee MS, JoeSH, Kim H, Ham BJ. Brain-derived neurotrophic factor promoter methylation and cortical thickness in recurrent major depressive disorder. Sci Rep. 2016 Feb15;6:21089. doi: 10.1038 / srep21089. PMID: 26876488; PMCID: PMC4753411.)
[0006] In conclusion, central structural thickness serves as a crucial bridge connecting psychological issues, clinical symptoms, and individual differences in aging trajectories, providing key evidence for assessing brain health and preventing potential diseases.
[0007] DNA methylation is an important component of systems biology and a key technology in precision medicine. It can reveal the physiological and pathological state of an organism in real time, reflecting changes related to its phenotype. Currently, numerous DNA methylation biomarker prediction models based on blood, urine, feces, and tissue samples are used to predict brain structure-related phenotypes.
[0008] However, no studies have yet used DNA methylation to predict individual differences in central region structural thickness. Therefore, developing a highly sensitive and specific integrated DNA methylation biomarker prediction model to accurately assess central region structural thickness is of great significance for improving the early identification of individual brain development and aging trajectories and the ability to warn of neurological disease risks, while ensuring non-invasive and convenient detection. Summary of the Invention
[0009] The purpose of this invention is to provide a highly sensitive and specific peripheral blood DNA methylation model that predicts the thickness of the central region structure, so as to improve the early identification of individual brain development and aging trajectory and the ability to warn of neurological disease risks.
[0010] The peripheral blood DNA methylation model provided by this invention predicts the thickness of the central region structure, wherein the thickness of the central region structure is one or more of the following: central sulcus thickness (S_central_thickness), precentral gyrus thickness (G_precentral_thickness), postcentral gyrus thickness (G_postcentral_thickness), and postcentral sulcus thickness (S_postcentral_thickness).
[0011] The DNA methylation model for predicting the thickness of the central region structure provided by this invention is constructed by the following steps:
[0012] S1. Sample set construction: Obtain a sample set containing complete phenotypic and methylation data; specifically including:
[0013] (a) Sample preparation:
[0014] EDTA or sodium citrate anticoagulant tubes are preferred for venous blood collection. The collected venous blood is then separated from the whole blood using density gradient centrifugation to isolate the target cell population, followed by DNA extraction. Aseptic techniques and sample information recording are required throughout the process to maximize the integrity of the genomic DNA and prevent degradation and unnatural changes in DNA methylation patterns.
[0015] (b) Data Acquisition:
[0016] First, DNA was extracted from frozen PBMC cell pellets, and its purity, concentration, and integrity were rigorously controlled. Then, genomic DNA was subjected to bisulfite treatment, converting unmethylated cytosine to uracil while leaving methylated cytosine unchanged, thus creating a difference in methylation status at the sequence level. The treated DNA was then hybridized on an EPIC microarray, ultimately generating a raw data file (IDAT file) containing the fluorescence intensity values of each probe. These data will lay a solid foundation for subsequent bioinformatics analyses, including background correction, normalization, and screening of differentially methylated sites / regions.
[0017] S2. Data preprocessing: The above sample set is preprocessed to obtain a dataset that can be used for subsequent analysis; specifically, this includes: performing linear correlation analysis for each phenotype and screening out methylation sites that are significantly correlated with the phenotype (p value < 0.05) for the next step of machine learning modeling.
[0018] S3. Prediction Model Construction: In the above dataset, linear correlation analysis is performed for each phenotype to screen CpG sites that are significantly related to the phenotype as candidate features. Then, a dual-track analysis system is established. The preprocessed samples are randomly divided into training set and validation set in a 7:3 ratio. Continuous variable prediction model and binary classification prediction model are constructed respectively to establish the dual-track analysis system.
[0019] (I) Construction of continuous prediction model:
[0020] (1) Data preprocessing: Prioritize the use of samples from the training set; remove low variance features and reduce noise / redundant features.
[0021] (2) 10-fold cross-validation: At least 100 rounds of Bootstrap sampling analysis were performed on the thickness of the posterior central gyrus, the central groove, the posterior central groove, and the anterior central gyrus; in each round, 50% of the thickness was randomly selected as the training set and 50% was used to validate the model.
[0022] (3) Bootstrap sampling is used to perform single CpG linear regression screening on each round of subsamples. The top 20 significant CpGs are selected using single CpG linear regression on 50% of the samples. The stability of the feature ranking is checked every 10 iterations. Then, based on the screened features, multiple models such as LASSO, multiple linear regression, random forest, Elastic Net, and Kernel Ridge are used to calculate Ri. 2(This is an indicator that measures the goodness of fit of a regression model, representing the proportion of the variance in the dependent variable that can be explained by the independent variables in the model. The value ranges from 0 to 1; a higher value indicates a better model fit.) Statistically analyze the frequency of each CpG selection and its corresponding performance. Every 10 iterations, calculate a comprehensive score for each feature using the following formula:
[0023] Composite_Score = 0.7 × Selected Frequency + 0.3 × Average R 2 ;
[0024] (4) Sort by comprehensive score to obtain the current ranking; perform Spearman correlation with the previous batch ranking + average rank change of the top 20; if the correlation > 0.95 and the average rank change < 2, it is considered that "feature ranking is stable", and the Bootstrap is stopped in advance to obtain the original result, and the result is processed according to R 2 The criteria were selected based on a value >0.5 and an AUC (area under the receiver operating characteristic curve, used to evaluate the performance of the classification model; AUC >0.75 indicates that the model's discrimination ability has reached a good level). The specific results were obtained by screening for values >0.5 and AUC (area under the receiver operating characteristic curve, used to evaluate the performance of the classification model; AUC >0.75 indicates that the model's discrimination ability has reached a good level).
[0025] (5) Finally, construct multiple final models:
[0026] Multiple linear regression (using only the 'feature ranking stable' features selected in step (4) as input variables), LASSO (for continuous dependent variables) (selecting the lambda value that minimizes the model error in cross-validation from all preprocessed features as the final regularization strength), random forest (using only the 'feature ranking stable' features selected in step (4) as input variables), Elastic Net (all preprocessed features), Kernel Ridge (using only the 'feature ranking stable' features selected in step (4) as input variables, and the number of features < the number of samples), Stacking ensemble uses the predictions of the three base models, linear model / random forest / kernel ridge regression, as machine learning output, and then uses linear regression as a second-layer model.
[0027] (II) Construction of Binary Classification Prediction Model
[0028] (1) Model building data preprocessing: Prioritize the use of training set samples; remove low variance features and reduce noise / redundant features;
[0029] (2) 10-fold cross-validation, performing at least 100 bootstrap sampling analyses for each phenotype. In each round, 50% is randomly selected as the training set, and 50% is used to validate the model;
[0030] (3) Constructing binary classification labels and automatic rollback mechanism
[0031] For continuous phenotypic values, two labels, "high" or "low," are constructed based on a pre-configured binary classification strategy (binary_threshold_type), including:
[0032] (a) Median segmentation: Using the median of all samples as the threshold, samples with phenotypic values greater than the median are marked as high, and the rest are marked as low;
[0033] (b) Quantile extreme group segmentation (q30_70): Calculate the 30th percentile (Q30) and 70th percentile (Q70) of the phenotype in all samples, and only retain the "extreme samples" outside these two quantiles, that is: samples with phenotype values ≤ Q30 are retained and marked as low; samples with phenotype values ≥ Q70 are retained and marked as high; samples between Q30 and Q70 (the middle 40% interval) are all discarded and do not participate in binary classification modeling.
[0034] This creates a "low" and a "high" extreme group, used to train a clearer binary classification model. The preferred split is q30_70. If this split results in only one class or extreme imbalance (e.g., too few samples in one class), it automatically reverts to a robust split based on the median of the entire sample to ensure sufficient samples for modeling in both classes.
[0035] (4) For each round of bootstrapping samples, fit a univariate logistic regression (Ventricle_Class ~ CpG) for each CpG feature, sort them according to the p-value, and select the top 20 most significant CpGs as candidate features for that round; record the number of times these features are selected in the current round. In the same round, use all candidate CpG features as independent variables, train the model on the bootstrapping samples using LASSO logistic regression, and calculate the area under the ROC curve (AUC) on the validation set samples. Allocate the AUC contribution of this round to the CpGs obtained in the pre-screening of this round, and accumulate it into the average AUC statistics of the corresponding features. Every 10 iterations, calculate the reconstruction comprehensive score of each CpG based on the statistical results so far:
[0036] .
[0037] Sort the composite features from highest to lowest to obtain the current feature ranking, and compare it with the ranking of the previous batch (first 10 rounds) using Spearman correlation and the average change of the top 20 rankings: if the ranking correlation coefficient is > 0.95 and the average change of the top 20 rankings is < 2, then the feature ranking is considered to have stabilized and the Bootstrap is stopped early; otherwise, continue iterating until the maximum number of iterations is reached.
[0038] S4. Comprehensive evaluation of multiphenotypic prediction models;
[0039] After completing the continuous and binary classification task modeling for each phenotype, all model results are summarized and output in a unified manner; specifically:
[0040] (1) Single-type result management:
[0041] For each phenotype, save it in a separate directory named after the phenotype:
[0042] (a) Continuous task related files: including feature frequency and performance analysis, Bootstrap iteration results, final model performance evaluation, stable feature list, feature ranking history trajectory, and RDS files for LASSO, linear regression, random forest, ElasticNet, kernel ridge regression and Stacking models;
[0043] (b) Documents related to the binary classification task: including the comprehensive score of binary classification features, Bootstrap performance records, test set prediction results, final LASSO-logistic model and feature summary, etc.
[0044] (2) Summary of cross-phenotypic model performance:
[0045] The script iterates through the result list of all phenotypes and compares the key performance indicators (such as R) of continuous and binary classification tasks. 2 The results of multi-phenotype modeling (such as RMSE, MAE, AUC, etc.) are combined into a summary table to provide an overview and comparison of the results, which facilitates the rapid identification of phenotypes with better predictive performance and their corresponding combinations of methylation features.
[0046] This invention uses a model to predict peripheral blood DNA methylation and screens out methylation biomarkers with high reproducibility. These biomarkers have a reproducibility rate of over 85% in other independent datasets.
[0047] The present invention also provides a diagnostic product based on the peripheral blood DNA methylation markers, which can non-invasively predict the thickness of four structures in the central brain region (central sulcus, precentral gyrus, postcentral gyrus, and postcentral sulcus) by detecting the methylation level of target CpG sites in the peripheral blood of subjects. This can help assess brain development status and brain aging process, and provide molecular evidence for risk warning of neurodegenerative diseases and brain development-related diseases.
[0048] The diagnostic products are reagent kits, methylation panels, or methylation chips (EPIC chips).
[0049] I. Diagnostic reagent kits:
[0050] (I) Reagent kit preparation (core components and dosage)
[0051] The kit contains the following core components, and the specifications of each component can be adjusted according to the detection throughput, as shown in the example below:
[0052] (1) Target CpG site specific amplification primer pair: 1× working solution (20μM / tube), used for specific amplification of target methylation site regions in peripheral blood genomic DNA;
[0053] (2) Bisulfite conversion reagent: 1 set / box, used for bisulfite conversion of genomic DNA to distinguish between methylated and unmethylated cytosine;
[0054] (3) PCR amplification premix (containing dNTPs and Taq enzyme): 1× reaction system / box, to provide the reaction environment and enzyme raw materials for PCR amplification;
[0055] (4) Positive / negative control DNA: 1 tube each (10 ng / μL), used for experimental quality control and to verify the accuracy and specificity of the detection system;
[0056] (5) Supporting consumables: elution buffer, centrifuge column, etc., for DNA extraction and purification.
[0057] (II) Reagent Kit Operation
[0058] (1) Sample collection: Peripheral venous blood of the subjects was collected using EDTA anticoagulant tubes or sodium citrate anticoagulant tubes. The blood was processed within 2 hours after collection and stored at 4°C for no more than 24 hours. For long-term storage, the sample should be placed in a -80°C freezer to avoid hemolysis and repeated freeze-thaw cycles.
[0059] (2) DNA extraction: Peripheral blood genomic DNA was extracted using a standard genomic DNA extraction kit. After passing quality inspection (OD260 / 280=1.8~2.0), it was used for subsequent experiments.
[0060] (3) Bisulfite conversion: DNA was converted to bisulfite according to the kit instructions, and the converted DNA was purified and recovered;
[0061] (4) PCR amplification: PCR amplification was performed on the transformed DNA using target site-specific primers. After the amplification products passed quality control, the methylation level was detected by pyrosequencing, next-generation sequencing or Sanger sequencing.
[0062] (5) Results analysis: The detected methylation level data is input into the prediction model constructed in this invention to calculate the predicted thickness values of each structure in the central brain region of the subject, and the diagnostic assessment is completed in combination with the reference threshold.
[0063] II. Methylation Chip:
[0064] (a) Chip Structure
[0065] Based on the Infinium MethylationEPIC chip platform, this invention adds 81 specific probes for stable CpG sites selected in this invention to the original chip probe set. The probes cover the target methylation region and can simultaneously detect the whole genome methylation level and the precise methylation status of the target site. The chip package includes probe array, hybridization buffer, washing buffer, labeling reagent and other components.
[0066] (II) Chip fabrication
[0067] Probe arrays were prepared using solid-phase in-situ synthesis technology. The probes were then specifically modified by methylation to optimize the hybridization efficiency and specificity of the probes, ensuring the accuracy and repeatability of target site detection.
[0068] (III) Chip Usage
[0069] (1) Sample pretreatment: Same as the kit procedure, complete peripheral blood DNA extraction and bisulfite conversion;
[0070] (2) Chip hybridization: The transformed DNA is hybridized with chip probes to capture target methylation sites through specific binding;
[0071] (3) Washing and scanning: Wash away unbound nonspecific DNA and use a chip scanner to read the probe fluorescence signal;
[0072] (4) Data standardization: The original signal was processed using the Illumina official standardization process to obtain the β value (methylation level) of the target CpG site.
[0073] (5) Model prediction: Input the standardized methylation data into the prediction model of this invention, and output the prediction result of the central region structure thickness.
[0074] The peripheral blood was obtained by venous blood collection using EDTA anticoagulant tubes or sodium citrate anticoagulant tubes.
[0075] The technical features and performance advantages of this invention mainly include:
[0076] (1) The blood sample collection used in this invention is simple, and the DNA methylation detection technology is mature and stable. In addition, the methylation data can be stored for a long time and analyzed in batches, which is suitable for large-scale cohort studies. The acceptance of the tested population is high, and it is also convenient for clinical and scientific researchers to promote its application.
[0077] (2) Compared with traditional single phenotype or single model analysis methods, this invention first corrects the interference of potential confounding factors such as sample heterogeneity and technical batch effects through rigorous data preprocessing and feature stability screening, which can accurately reflect the association between multiple phenotypes and methylation markers and significantly reduce the false positive rate and false negative rate.
[0078] (3) This invention is the first to construct a dual-track prediction system for continuous variables and binary classification tasks. Through the integration of various machine learning algorithms (including LASSO, random forest, Elastic Net, etc.) and Bootstrap stability evaluation, it selects combinations of methylation biomarkers that can stably associate with the target phenotype in different datasets and different detection batches, and will not fail due to sample fluctuations, thus avoiding the limitations of single algorithm or single phenotype analysis. At the same time, it combines Stacking ensemble learning to construct a robust multi-phenotype prediction model, and performs cross-validation and independent testing. In practical applications, multiple phenotypes (such as EAA-related indicators) can be accurately predicted using only blood DNA methylation data. This helps researchers and clinicians identify high-risk individuals or explore molecular mechanisms based on the prediction results, providing support for early intervention and personalized medicine, and has broad scientific research and clinical translational value.
[0079] (4) The multi-phenotypic prediction model provided by this invention performs well in performance evaluation and is validated using multiple metrics (applicable to continuous and binary classification tasks). In continuous prediction tasks, the average R-value is [missing information]. 2The ROC curve for binary classification tasks achieved a score above 0.876, with the area under the ROC curve exceeding 0.719 on both the training and validation sets. Accuracy, sensitivity, and specificity all exceeded 0.599. More importantly, this invention not only constructed a high-performance prediction model but also screened highly reproducible methylation biomarkers through rigorous bootstrap stability analysis. These biomarkers achieved a reproducibility rate of over 85% on other independent datasets, providing a reliable evidence base for feature selection for research teams in related fields. This feature selection method based on stability assessment can effectively reduce the trial-and-error costs for other research teams in the biomarker validation stage, shorten the research cycle, and provide methodological reference and practical basis for the discovery and validation of disease biomarkers. The prediction model combines excellent predictive performance and reliability, while the selected methylation biomarker combinations also provide important clues for subsequent mechanism research, offering a powerful tool for the application of DNA methylation in complex phenotypic studies. Attached Figure Description
[0080] Figure 1 This is a flowchart of the prediction model construction process of the present invention.
[0081] Figure 2 For Test ROC - G_precentral_thickness.
[0082] Figure 3 For Test ROC - S_central_thickness.
[0083] Figure 4 For Test ROC - G_postcentral_thickness.
[0084] Figure 5 For Test ROC - S_postcentral_thickness.
[0085] Figures 2-5 This is an image showing the effect of the peripheral blood DNA methylation marker of the present invention. Detailed Implementation
[0086] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0087] Example 1,
[0088] This embodiment provides a peripheral blood DNA methylation model for predicting the thickness of brain regions reflecting central brain structures. The predicted brain regions are one or more of the following: central sulcus thickness, precentral gyrus thickness, postcentral gyrus thickness, and postcentral sulcus thickness. The model construction process is as follows: Figure 1 As shown, it includes the following steps:
[0089] S1: Construct a sample set by obtaining complete phenotypic and methylation data; specifically including:
[0090] (a) Sample preparation:
[0091] To ensure the accuracy and reliability of subsequent DNA methylation detection results, the collection of fresh blood samples must follow strict standard procedures. EDTA anticoagulant tubes (purple caps) or sodium citrate anticoagulant tubes (blue caps) are preferred for venous blood collection, as these two anticoagulants cause minimal interference with subsequent DNA extraction and methylation analysis. Heparin anticoagulant tubes should be avoided, as heparin severely inhibits subsequent enzymatic reactions. Collected samples should be temporarily stored and transported at 4°C and must be processed within 24 hours of collection. The core processing step is to separate the target cell population, such as peripheral blood mononuclear cells (PBMCs), from whole blood using density gradient centrifugation (commonly using Ficoll reagent) to eliminate the dilution effect of highly heterogeneous blood cells on the methylation signal. The separated cell pellet should be immediately placed in an ultra-low temperature freezer at -80°C or liquid nitrogen for long-term storage until the DNA extraction stage. Aseptic techniques must be used throughout the process, and sample information must be recorded to maximize the integrity of genomic DNA and prevent degradation and unnatural changes in DNA methylation patterns.
[0092] (b) Data Acquisition:
[0093] The EPIC chip-based data acquisition process begins with the acquisition of high-quality genomes. First, DNA is extracted from frozen PBMC cell pellets, and its purity, concentration, and integrity are rigorously controlled using instruments such as Nanodrop or Qubit. After passing quality control, approximately 500 ng of genomic DNA undergoes bisulfite treatment. This step is crucial for the success of the experiment; it converts unmethylated cytosine to uracil, while methylated cytosine remains unchanged, thus creating differences in methylation status at the sequence level. The treated DNA is then hybridized on an EPIC chip, which covers over 850,000 CpG sites in the human genome, encompassing multiple important functional regions such as enhancers, gene promoters, and coding regions. The chip reads the hybridization signals via fluorescence scanning, and the resulting raw data file (IDAT file) contains the fluorescence intensity values for each probe. This data will lay a solid foundation for subsequent bioinformatics analyses, including background correction, normalization, and screening of differentially methylated sites / regions.
[0094] S2: Data preprocessing, which involves preprocessing the above sample set to obtain a dataset that can be used for subsequent analysis; specifically including:
[0095] Spearmen's analysis method was used to find all CPG features that were linearly and significantly correlated with the thickness of the central region structure.
[0096] S3: Predictive Model Building
[0097] 3.1: Construction of Continuous Prediction Model
[0098] 1. Model building data preprocessing: Prioritize the use of samples from the training set; remove low-variance features and reduce noisy / redundant features.
[0099] 2. Ten-fold cross-validation: Perform at least 100 bootstrap sampling analyses on the thicknesses of the posterior central gyrus, central groove, posterior central groove, and anterior central gyrus. In each round, 50% of the sample is randomly selected as the training set, and the remaining 50% is used to validate the model.
[0100] 3. Using Bootstrap sampling, single CpG linear regression is performed on each round of subsamples for selection. The top 20 significant CpGs are selected using single CpG linear regression on 50% of the samples. The stability of the feature ranking is checked every 10 iterations. Then, multiple models, including LASSO, linear regression, random forest, Elastic Net, and Kernel Ridge, are used to calculate R-squared on the selected features. 2 The frequency of each CpG selection and its corresponding performance are statistically analyzed. Every 10 times, a comprehensive score is calculated for each feature using the following formula:
[0101] Composite_Score = 0.7 × Selected Frequency + 0.3 × Average R 2 .
[0102] 4. Sort by comprehensive score to obtain the current ranking; calculate Spearman correlation with the previous batch rankings and add the average change in rank of the top 20; if the correlation > 0.95 and the average rank change < 2, it is considered that the "feature ranking is stable", and the bootstrap is stopped early to obtain the original results, and the results are processed according to R. 2 The results obtained by filtering for AUC >0.5 and AUC >0.75 are shown in Table 1.
[0103] Table 1 Overview of the prediction results of structural thickness-related features in the central area
[0104] .
[0105] 5. Finally, construct multiple final models:
[0106] Multiple linear regression (using only stable features), LASSO (continuous) (all features, lambda.min), random forest (stable features), Elastic Net (all features), Kernel Ridge (stable features, and the number of features < the number of samples), Stacking ensemble uses the predictions of LM / RF / KR as machine learning output, and then uses linear regression to make a two-layer model.
[0107] For linear regression and stacking, construct explicit prediction equation strings (convert the trained model into mathematical formula strings that can be directly substituted into calculations, clearly showing the quantitative relationship between the target phenotype and each CpG feature), for example:
[0108] Target phenotype (S_central_thickness) = 4.085 + 0.6924 * cg18848366 + -0.6462 * cg01794867 + -0.6914 * cg11081694 + -1.155 * cg07665929 + 0.3466 * cg26767214 + -0.7083 * cg02267698.
[0109] 3.2: Construction of Binary Classification Prediction Model
[0110] 1. Model building data preprocessing: Prioritize the use of training set samples; remove low-variance features and reduce noisy / redundant features.
[0111] 2. 10-fold cross-validation, performing at least 100 bootstrap sampling analyses for each phenotype. In each round, 50% is randomly selected as the training set, and 50% is used to validate the model.
[0112] 3. Constructing binary classification labels and an automatic rollback mechanism
[0113] For continuous phenotypic values, two categories of labels, "high" and "low," are constructed according to a pre-configured binary classification strategy, including:
[0114] (a) Median segmentation: Using the median of all samples as the threshold, samples with phenotypic values greater than the median are marked as high, and the rest are marked as low;
[0115] (b) Quantile extreme group segmentation (q30_70): Calculate the 30th percentile (Q30) and 70th percentile (Q70) of the phenotype in all samples, and only retain the "extreme samples" outside these two quantiles, that is: samples with phenotype values ≤ Q30 are retained and marked as low; samples with phenotype values ≥ Q70 are retained and marked as high; samples between Q30 and Q70 (the middle 40% interval) are all discarded and do not participate in binary classification modeling.
[0116] This creates a "low" and a "high" extreme group, used to train a clearer binary classification model. The preferred split is q30_70. If this split results in only one class or extreme imbalance (e.g., too few samples in one class), it automatically reverts to a robust split based on the median of the entire sample to ensure sufficient samples for modeling in both classes.
[0117] 4. For each round of bootstrapping samples, fit a univariate logistic regression (Ventricle_Class ~ CpG) for each CpG feature, sort them according to the p-value, and select the top 20 most significant CpGs as candidate features for that round; record the number of times these features are selected in the current round. In the same round, use all candidate CpG features as independent variables, train the model on the bootstrapping samples using LASSO logistic regression, and calculate the area under the ROC curve (AUC) on the validation set samples. Distribute the AUC contribution of this round to the CpGs pre-screened in this round, and accumulate it in the average AUC statistics of the corresponding features. Every 10 iterations, calculate the reconstruction comprehensive score for each CpG based on the statistical results so far:
[0118] .
[0119] Sort the composite features from highest to lowest to obtain the current feature ranking, and compare it with the ranking of the previous batch (first 10 rounds) using Spearman correlation and the average change of the top 20 rankings: if the ranking correlation coefficient is > 0.95 and the average change of the top 20 rankings is < 2, then the feature ranking is considered to have stabilized and the Bootstrap is stopped early; otherwise, continue iterating until the maximum number of iterations is reached.
[0120] S4: Comprehensive Evaluation of Multiphenotypic Prediction Models
[0121] After completing the continuous and binary classification tasks for each phenotype, all model results are summarized and output in a unified manner:
[0122] 1. Single-table result management:
[0123] For each phenotype, save it in a separate directory named after the phenotype:
[0124] (a) Continuous task related files: including feature frequency and performance analysis, Bootstrap iteration results, final model performance evaluation, stable feature list, feature ranking history trajectory, and RDS files for LASSO, linear regression, random forest, ElasticNet, kernel ridge regression and Stacking models;
[0125] (b) Documents related to the binary classification task: including the comprehensive score of binary classification features, Bootstrap performance records, test set prediction results, final LASSO-logistic model and feature summary, etc.
[0126] 2. Performance summary of cross-phenotype models:
[0127] The script iterates through the result list of all phenotypes and compares the key performance indicators (such as R) of continuous and binary classification tasks. 2 The results of multi-phenotype modeling (such as RMSE, MAE, AUC, etc.) are combined into a summary table to provide an overview and comparison of the results, which facilitates the rapid identification of phenotypes with better predictive performance and their corresponding combinations of methylation features.
[0128] Specific experiments and results regarding the method of this invention:
[0129] ① Characteristics of the study population:
[0130] This invention included data from 134 adult crew members on long-distance voyages. Participants were recruited through community recruitment and health check-up centers, and none had a history of clinical diagnosis of neurological diseases. Exclusion criteria included: a history of stroke, traumatic brain injury, or neurodegenerative diseases (such as Alzheimer's disease or Parkinson's disease); severe mental illness (such as schizophrenia or a major depressive episode); contraindications to MRI scans (such as implanted metal devices or claustrophobia); recent (within 6 months) use of medications that may significantly affect DNA methylation patterns (such as chemotherapy drugs or immunosuppressants); and a known systemic disease that may affect epigenetics (such as autoimmune diseases or advanced liver or kidney disease).
[0131] All participants underwent high-resolution brain MRI scans (for precise measurement of parietal cortex thickness) and peripheral venous blood collection (e.g., for whole-genome DNA methylation microarray analysis). Plasma was separated within 2 hours of blood sample collection, and the leukocyte layer was cryopreserved at -80°C until DNA extraction. This data was used to construct and validate a predictive model of "peripheral blood characteristics-brain structural phenotype." All participants signed written informed consent forms prior to enrollment, agreeing to the use of their imaging data, blood samples, and molecular data in this study.
[0132] The central tendency and interquartile distribution of the thickness of the four brain regions in the population are as follows (units are based on the original brain image processing output): Central sulcus thickness: mean 2.3566, median 2.3420, Q1=2.2515, Q3=2.4215, minimum 2.0220, maximum 2.7460; Precentral gyrus thickness: mean 3.1703, median 3.2080, Q1=3.0930, Q3=3.3135. Minimum value 2.3340, maximum value 3.9030; Central rear groove thickness: mean 2.4321, median 2.4100, Q1=2.3135, Q3=2.5160, minimum value 2.0130, maximum value 2.9280; Central rear groove thickness: mean 2.381, median 2.3900, Q1=2.3035, Q3=2.4615, minimum value 2.0000, maximum value 2.7500.
[0133] Furthermore, the differences between individuals within the same brain region in the population were significant and could be quantified by specific numerical values: the range of thickness in each brain region was 0.7240 (2.7460-2.0220) for the central sulcus, 1.5690 (3.9030-2.3340) for the precentral gyrus, 0.9150 (2.9280-2.0130) for the postcentral gyrus, and 0.7500 (2.7500-2.0000) for the postcentral sulcus; the interquartile range (IQR) was 0.1700 (2.4215-2.2515) for the central sulcus, 0.2205 (3.3135-3.0930) for the precentral gyrus, 0.2025 (2.5160-2.3135) for the postcentral gyrus, and 0.1580 (2.4615-2.3035) for the postcentral sulcus. The above results indicate that not only is there a significant gap between extreme individuals (minimum and maximum values), but the middle 50% of the population (Q1 to Q3) also exhibits stable dispersion.
[0134] Therefore, the central region structural thickness phenotype predicted by this invention has sufficient variability and distinguishability within the population (Range approximately 0.7240–1.5690, IQR approximately 0.1580–0.2205), providing a clear and quantifiable basis for the distribution of brain structure in the research population for subsequent screening of brain region structures based on peripheral blood DNA methylation characteristics and construction of a central region structural thickness prediction model.
[0135] Using four cortical thickness phenotypes as prediction targets (Pheno), we compared models such as LASSO regression, ElasticNet, random forest, and Stacking ensemble learning in continuous tasks; in binary classification tasks, we used the LASSO-logistic final model for validation, and uniformly adopted the q30_70 extreme grouping strategy (segmentation at 30% and 70% quantiles, with Cut_Info recording the threshold of each phenotype).
[0136] Table 2 shows the overall results for the continuous tasks: Stacking ensemble learning performs best overall across the four phenotypes, with R0 being the highest. 2 The mean is 0.9774 (range 0.9723–0.9819); Random Forest (stable feature) R 2 The mean is 0.9693 (range 0.9559–0.9800). In comparison, Elastic Net and LASSO have lower R... 2 The mean values were 0.9045 and 0.8805, respectively, showing significant fluctuations. Furthermore, based on the statistics of the "optimal model for each phenotype," the optimal R-values for the four phenotypes were... 2The mean was 0.9781 (range 0.9723–0.9846), corresponding to RMSE of 0.0156–0.1972 and MAE of 0.0126–0.1441. The optimal results for each phenotype are as follows:
[0137] (1)G_precentral_thickness: Stacking, R 2 =0.9723;
[0138] (2) S_central_thickness: Stacking, R 2 =0.9811;
[0139] (3)G_postcentral_thickness: Elastic Net, R 2 =0.9846;
[0140] (4) S_postcentral_thickness: Stacking, R 2 =0.9744;
[0141] Table 3 shows the overall results of the binary classification task: After using the q30_70 extreme grouping (retaining only low and high quantile samples), the number of training set samples N_train=28 and the number of test set samples N_test=10 for each phenotype; the overall low / high sample count is 19. The final LASSO-logistic model achieved a mean AUC of 0.84 (range 0.76–0.88), a mean Accuracy of 0.7333 (range 0.70–0.80), a mean Sensitivity of 0.7333 (range 0.60–0.80), and a mean Specificity of 0.7333 (range 0.60–0.8) for the four phenotypes. Among these, S_central_thickness showed the best discriminative performance (AUC=0.88, Accuracy=0.80, Specificity=0.80).
[0142] Meanwhile, to ensure the reproducibility of methylation marker screening, a Bootstrap stability assessment was conducted for each target phenotype: the frequency_Ratio of repeated inclusion of CpG sites was calculated during repeated sampling, and a comprehensive score was constructed by combining the model performance indicators.
[0143] In continuous prediction tasks, the composite score is set to 0.7 × Frequency Ratio + 0.3 × Avg R. 2This refers to a model with a high frequency of repeated inclusion (Frequency_Ratio) and a goodness of fit (Avg_R) corresponding to the inclusion round. 2 (Good CpG sites)
[0144] In the binary classification task, Composite = 0.7 × Frequency_Ratio + 0.3 × Avg_AUC is used, which refers to CpG sites with a high repeated inclusion rate (Frequency_Ratio) and good model discrimination ability (Avg_AUC) in the corresponding selection round.
[0145] Each phenotype was selected using the 70th percentile of the comprehensive score as a threshold. The top 30% of the CpGs in terms of comprehensive score were used to obtain a stable feature set, resulting in stable feature tables: continuous model markers.xlsx and binary classification model markers.xlsx.
[0146] In the continuous model feature selection results, the stable feature set corresponding to each phenotype is of the same size, containing 81 methylation features (CpG sites). Based on FR≥0.3, the number of superior features corresponding to the four phenotypes is as follows:
[0147] (1) S_central_thickness: 13;
[0148] (2) S_postcentral_thickness: 19;
[0149] (3) G_postcentral_thickness: 11;
[0150] (4) G_precentral_thickness: 8;
[0151] In the feature selection results of the binary classification model, the stable feature set corresponding to each phenotype has a consistent size, containing 81 methylation features (CpG sites). Based on FR≥0.3, the number of superior features corresponding to the four phenotypes are as follows:
[0152] (1) S_central_thickness: 14;
[0153] (2) G_precentral_thickness: 17;
[0154] (3) G_postcentral_thickness: 19;
[0155] (4) S_postcentral_thickness: 16;
[0156] Meanwhile, the overall distribution of stable features shows a very consistent "dominant distribution pattern" across the four phenotypes. In gene structural regions (features): Body and IGR dominate, accounting for approximately 60%–75% of all phenotypes in the continuous model and approximately 53.3%–65.9% in the binary model. This indicates that candidate CpGs are not concentrated solely in promoters, but rather fall more heavily within the gene body / intergenic regions, consistent with the common characteristic in genome-wide methylation screening where "stable signals originate from broader genomic regions."
[0157] On the CpG island related annotations (CGI), opensea is dominant, with opensea accounting for the highest proportion (approximately 54%–65%) across all phenotypes. This indicates that candidate sites are more biased towards regions far from the CpG island (potentially containing enhancers, distal regulatory elements, etc.), rather than the traditional single pathway of the "CpG island promoter".
[0158] The superiority of the "preferred features" described in this invention is reflected in at least the following two aspects:
[0159] (1) High repeatability: It is stably and repeatedly selected in multiple rounds of Bootstrap resampling iterations, which reduces the sensitivity of feature selection to random sample fluctuations, thereby improving the reproducibility of marker screening results;
[0160] (2) More robust prediction contribution: Avg_R_squared is at a high level, indicating that it is not only "easily selected repeatedly", but also has a more stable contribution to the fitting / interpretation ability of continuous prediction models.
[0161] This embodiment focuses on four cortical thickness phenotypes, establishing continuous prediction models and binary classification discriminant models respectively, and conducting system evaluations. In the continuous task, various algorithms were compared. As shown in Table 2, the ensemble learning model represented by Stacking exhibits the best overall performance. The average R² of the optimal model for the four phenotypes reached 0.9794 (0.9667–0.9988), and both RMSE and MAE were at low levels, demonstrating strong prediction accuracy and stability. In the binary classification task, under extreme grouping conditions of q30_70 and small sample size, LASSO-logistic still achieved good discriminant results (mean AUC 0.8562, highest AUC = 0.96). Figures 2-5 As shown in Table 3, the accuracy, sensitivity, and specificity metrics are also presented, verifying the effectiveness of the model in classification scenarios.
[0162] Table 2. Comparison of the effects of various algorithms for the continuous thickness model of the central area structure.
[0163] .
[0164] Table 3 Performance Indicators of the Binary Classification Model for Structural Thickness in the Central Area
[0165] .
[0166] More importantly, this embodiment not only reports model performance but also introduces Bootstrap stability analysis to screen CpG features for reproducibility: reproducibility is represented by the inclusion rate (FR), and FR is correlated with model performance (Avg_R for continuous tasks). 2 The binary classification task (Avg_AUC) is used to weight a composite score, which is then used to form a stable feature set with a 70th percentile threshold. This achieves "high-performance prediction / discrimination" and "reproducible biomarker screening based on stability assessment," providing a reliable basis and feasible process for subsequent biomarker validation, panel construction, and mechanism research.
Claims
1. A model for predicting peripheral blood DNA methylation reflecting the thickness of the central region structure, wherein the thickness of the central region structure is one or more of the following: central sulcus thickness, precentral gyrus thickness, postcentral gyrus thickness, and postcentral sulcus thickness; the model is constructed by the following steps: S1. Sample set construction: Obtain a sample set containing complete phenotypic and methylation data; specifically including: The target cell population was separated from the collected venous blood using density gradient centrifugation, and then DNA was extracted and sample information was recorded. The purity, concentration, and integrity of the DNA were strictly controlled. Genomic DNA was then subjected to bisulfite treatment to convert unmethylated cytosine to uracil, while methylated cytosine remained unchanged, creating a difference in methylation status at the sequence level. The treated DNA was hybridized on an EPIC chip to generate a raw data file containing the fluorescence intensity value of each probe. S2. Data preprocessing: The above sample set is preprocessed to obtain a dataset that can be used for subsequent analysis. Specifically, this includes: performing linear correlation analysis for each phenotype and screening out methylation sites that are significantly correlated with the phenotype for the next step of machine learning modeling; S3. Prediction Model Construction: The preprocessed samples are randomly stratified into training and validation sets in a 7:3 ratio, and continuous variable prediction models and binary classification prediction models are constructed respectively to establish a dual-track analysis system. (I) Construction of Continuous Prediction Model (1) Data preprocessing: Use samples from the training set; remove low-variance features and reduce noise / redundant features; (2) 10-fold cross-validation: At least 100 rounds of Bootstrap sampling analysis were performed on the thickness of the posterior central gyrus, the central groove, the posterior central groove, and the anterior central gyrus; in each round, 50% of the thickness was randomly selected as the training set and 50% was used to validate the model. (3) By Bootstrap sampling, do single CpG linear regression for each round of sub-sample, use 50% sample, single CpG linear regression selects the top 20 significant CpGs; check "whether the feature ranking is stable" every 10 times; then on the selected features, combine LASSO, multiple linear regression, random forest, Elastic Net and Kernel Ridge multiple models to calculate R 2 , count the frequency of each CpG selected and the corresponding performance, every 10 times, calculate a comprehensive score for each feature according to the following formula: Composite_Score = 0.7 x Selected Frequency + 0.3 x Average R 2 ; (4) Rank by the comprehensive score to get the current ranking; and Spearman correlation of the ranking of the previous batch + average rank change of the top 20; if the correlation > 0.95 and the average rank change < 2, consider "feature ranking stable", stop Bootstrap in advance, get the original result, and filter the results according to R 2 > 0.5 and AUC greater than 0.75 to get specific results; (5) Finally, construct multiple final models: For multiple linear regression, LASSO, random forest, Elastic Net, Kernel Ridge, and Stacking integration, the predictions of three base models (linear model / random forest / kernel ridge regression) are used as machine learning outputs, and then linear regression is used as a second-layer model. (II) Construction of Binary Classification Prediction Model (1) Model building data preprocessing: Prioritize the use of training set samples; remove low variance features and reduce noise / redundant features; (2) 10-fold cross-validation, with at least 100 Bootstrap sampling analyses performed for each phenotype; in each round, 50% of the sample is randomly selected as the training set and 50% is used to validate the model. (3) Constructing binary classification labels and an automatic rollback mechanism: For continuous phenotypic values, two labels, "high" or "low," are constructed according to a pre-configured binary classification strategy, including: (a) Median segmentation: Using the median of all samples as the threshold, samples with phenotypic values greater than the median are marked as high, and the rest are marked as low; (b) Quantile extreme group splitting, denoted as q30_70: Calculate the 30th percentile of the phenotype in the entire sample, denoted as Q30, and the 70th percentile, denoted as Q70. Only "extreme samples" outside these two quantiles are retained, that is: samples with phenotype values ≤ Q30 are retained and marked as low; samples with phenotype values ≥ Q70 are retained and marked as high; samples between Q30 and Q70 are all discarded and do not participate in binary classification modeling; forming a "low extreme group" and a "high extreme group" to train a clearer binary classification model; q30_70 splitting is preferred. If this splitting method results in only a single class or extreme imbalance, it will automatically fall back to a robust splitting method based on the "median >" of the entire sample to ensure that there are enough samples for modeling in both classes; (4) For each round of bootstrapping samples, fit a univariate logistic regression for each CpG feature, sort them according to the p-value, and select the top 20 most significant CpGs as candidate features for that round; record the number of times these features are selected in the current round; in the same round, use all candidate CpG features as independent variables, train the model on the bootstrapping samples using LASSO logistic regression, and calculate the area under the ROC curve (AUC) on the validation set samples; allocate the AUC contribution of this round to the CpGs obtained in the pre-screening of this round, and accumulate it into the average AUC statistics of the corresponding features; every 10 iterations, calculate the reconstruction comprehensive score of each CpG based on the statistical results so far: ; Sort the features from highest to lowest according to the Composite to obtain the current feature ranking, and compare it with the Spearman correlation and the average change of the top 20 rankings in the previous batch: if the ranking correlation coefficient is > 0.95 and the average change of the top 20 rankings is < 2, then the feature ranking is considered to have stabilized and the Bootstrap is stopped early; otherwise, continue iterating until the maximum number of iterations is reached. S4. Comprehensive evaluation of multiphenotypic prediction models; After completing the continuous and binary classification task modeling for each phenotype, all model results are summarized and output in a unified manner; specifically: (1) Single-type result management: For each phenotype, save it in a separate directory named after the phenotype: (a) Continuous task related files: including feature frequency and performance analysis, Bootstrap iteration results, final model performance evaluation, stable feature list, feature ranking history trajectory, and RDS files for LASSO, linear regression, random forest, Elastic Net, kernel ridge regression and Stacking models; (b) Documents related to the binary classification task: including the comprehensive score of binary classification features, Bootstrap performance records, test set prediction results, and the final LASSO-logistic model and feature summary; (2) Summary of cross-phenotypic model performance: The script iterates through the results list of all phenotypes, merges the key performance indicators of continuous tasks and binary classification tasks into a summary table, and serves as an overview and comparison basis for multi-phenotype modeling results, which facilitates the rapid identification of phenotypes with better prediction performance and their corresponding methylation feature combinations.
2. The peripheral blood DNA methylation model described in claim 1, which predicts the thickness of the central region structure, yields highly reproducible methylation markers, which have a reproducibility rate of over 85% in other independent datasets.
3. The diagnostic product for peripheral blood DNA methylation markers according to claim 2, comprising a diagnostic product kit and a methylation chip; wherein: (a) Diagnostic kits are used to non-invasively predict the thickness of four structures in the central brain region by detecting the methylation level of target CpG sites in the peripheral blood of subjects, to assist in assessing brain development status and brain aging process, and to provide molecular evidence for risk warning of neurodegenerative diseases and brain development-related diseases. A. Reagent Kit Configuration: Contains the following components, with the specifications of each component adjusted according to the detection throughput: (1) Target CpG site specific amplification primer pair: 1× working solution, 20 μM / tube, used for specific amplification of target methylation site regions in peripheral blood genomic DNA; (2) Bisulfite conversion reagent: 1 set / box, used for bisulfite conversion of genomic DNA to distinguish between methylated and unmethylated cytosine; (3) PCR amplification premix, containing dNTPs and Taq enzyme: 1× reaction system / box, to provide the reaction environment and enzyme raw materials for PCR amplification; (4) Positive / negative control DNA: 1 tube each, 10 ng / μL, used for experimental quality control and to verify the accuracy and specificity of the detection system; (5) Supporting consumables: elution buffer and centrifuge column, used for DNA extraction and purification; B. Reagent kit operation: (1) Sample collection: Peripheral venous blood of the subjects was collected using EDTA anticoagulant tubes or sodium citrate anticoagulant tubes. The blood was processed within 2 hours after collection and stored at 4°C for no more than 24 hours. For long-term storage, the sample should be placed in a -80°C freezer to avoid hemolysis and repeated freeze-thaw cycles. (2) DNA extraction: Genomic DNA was extracted from peripheral blood using a standard genomic DNA extraction kit and used in subsequent experiments after passing quality inspection; (3) Bisulfite conversion: DNA was converted to bisulfite according to the kit instructions, and the converted DNA was purified and recovered; (4) PCR amplification: PCR amplification was performed on the transformed DNA using target site-specific primers. After the amplification products passed quality control, the methylation level was detected by pyrosequencing, next-generation sequencing or Sanger sequencing. (5) Results analysis: The detected methylation level data is input into the prediction model constructed in claim 1 to calculate the predicted thickness values of each structure in the central brain region of the subject, and the diagnostic assessment is completed in combination with the reference threshold; (ii) Methylated chip A. Chip Structure: Based on the Infinium MethylationEPIC chip platform, 81 new specific probes for stable CpG sites were added to the original chip probe set. The probes cover the target methylation region and can detect the whole genome methylation level and the precise methylation status of the target site. The chip kit includes components such as probe array, hybridization buffer, washing buffer, and labeling reagents; B. Chip fabrication: Probe arrays were prepared using solid-phase in-situ synthesis technology, and the probes were specifically modified by methylation to optimize the probe hybridization efficiency and specificity, ensuring the accuracy and repeatability of target site detection. C. Chip Usage: (1) Sample pretreatment: Same as the kit procedure, complete peripheral blood DNA extraction and bisulfite conversion; (2) Chip hybridization: The transformed DNA is hybridized with chip probes to capture target methylation sites through specific binding; (3) Washing and scanning: Wash away unbound nonspecific DNA and use a chip scanner to read the probe fluorescence signal; (4) Data standardization: The original signal was processed using the Illumina official standardization process to obtain the β value of the target CpG site; (5) Model prediction: Input the standardized methylation data into the prediction model described in claim 1, and output the prediction result of the central region structure thickness.