Biological age prediction method, device, electronic device, and storage medium
By comprehensively predicting biological age using multi-omics data, the problem of inaccurate assessment in existing technologies has been solved, enabling more accurate assessment of aging status and association analysis of chronic diseases in the elderly.
Patent Information
- Application Number
- CN202511205011.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing technologies for biological age assessment are inaccurate; single-mathematical data models cannot comprehensively assess an individual's aging status and cannot capture the multi-system collaborative degeneration mechanisms.
By acquiring multi-omics data of the individual to be tested, and inputting them into trained genomic, proteomic, metabolomic and imaging aging clock models respectively for prediction, the target biological age is obtained by combining the weighting coefficients.
It improves the accuracy of biological age assessment, can more comprehensively reflect an individual's aging status, identify individuals who are aging faster or slower, and analyze its correlation with chronic diseases of old age.
Smart Images

Figure CN120727299B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, and in particular to biological age prediction methods, devices, electronic devices, and storage media. Background Technology
[0002] In traditional medicine, the measurement of an individual's aging process primarily relies on chronological age, determined by their birth date. However, this method has significant limitations. The rate of aging varies greatly among individuals, and chronological age cannot adequately reflect an individual's health status and degree of aging. For example, significant differences can exist among peers in physical function, disease risk, and lifespan. These differences stem mainly from the combined effects of various factors, including genetics, lifestyle, and environmental exposure. Therefore, the scientific community urgently needs a more precise method for assessing biological age to better reflect an individual's health status and rate of aging.
[0003] Against the backdrop of an aging population, developing an "aging clock" that can quantify the degree of biological aging has become a hot research topic in the biomedical field. The concept of biological age has emerged, assessing the process of functional decline in the body through various biomarkers and considered an indicator of the actual rate of "aging" in an individual. Currently, researchers mainly rely on single-agent genome data to predict biological age when constructing aging clocks.
[0004] However, aging clock models built upon single-atom mathematical data have many limitations. They suffer from one-dimensional bias, and existing aging clock models can only reflect local features of biological aging, failing to capture multi-system collaborative degradation mechanisms and making it difficult to comprehensively assess an individual's aging status.
[0005] There is currently no effective solution to the problem of inaccurate biological age assessment in related technologies. Summary of the Invention
[0006] This embodiment provides a biological age prediction method, apparatus, electronic device, and storage medium to address the problem of inaccurate biological age assessment in related technologies.
[0007] Firstly, this embodiment provides a biological age prediction method, including:
[0008] Acquire the target multi-omics data of the individuals to be tested;
[0009] Each omics data in the target multi-omics data is input into the corresponding trained aging clock model for prediction, and the predicted biological age corresponding to each omics data is obtained.
[0010] The predicted biological ages are combined to obtain the target biological age of the individual to be tested.
[0011] In some embodiments, acquiring the target multi-omics data of the individual to be detected further includes:
[0012] Obtain the raw multi-omics data of the individual to be tested;
[0013] The original multi-omics data is preprocessed and features are extracted to obtain pre-processed multi-omics data;
[0014] The pre-processed multi-omics data is subjected to feature dimensionality reduction and time alignment to obtain the target multi-omics data.
[0015] In some embodiments, the target multi-omics data includes genomic data, proteomic data, metabolomic data, and radiomic data; each omics data in the target multi-omics data is input into a corresponding trained aging clock model for prediction, obtaining the predicted biological age corresponding to each omics data, including:
[0016] The genomic data of the individual to be tested is input into a trained genomic data aging clock model for prediction, and the predicted biological age corresponding to the genomic data is obtained.
[0017] The proteomic data of the individual to be tested is input into a trained proteomic data aging clock model for prediction, and the predicted biological age corresponding to the proteomic data is obtained.
[0018] The metabolomics data of the individual to be tested are input into a trained metabolomics data aging clock model for prediction, and the predicted biological age corresponding to the metabolomics data is obtained.
[0019] The image group data of the individual to be tested is input into a trained image group data aging clock model for prediction, and the predicted biological age corresponding to the image group data is obtained.
[0020] In some embodiments, combining the predicted biological ages to obtain the target biological age of the individual to be tested includes:
[0021] Based on the predictive power of each model in the aging clock model for biological age, corresponding weighting coefficients are set;
[0022] The predicted biological ages are combined according to the weighting coefficients to obtain the target biological age of the individual to be tested.
[0023] In some embodiments, combining the predicted biological ages according to the weighting coefficients to obtain the target biological age of the individual to be tested includes:
[0024] The predicted biological ages are linearly regressed based on the weighting coefficients to obtain the target biological age of the individual to be tested.
[0025] In some of these embodiments, the genomic data includes whole-genome sequencing, DNA methylation, or telomere length detection data;
[0026] The proteomic data includes cardiac metabolic proteins, inflammatory proteins, and oncological proteins;
[0027] The metabolomics data include blood glucose, blood lipids, liver and kidney function, blood cell count, and liver enzymes;
[0028] The image group data includes structural magnetic resonance imaging, diffusion magnetic resonance imaging, and resting-state magnetic resonance imaging.
[0029] In some of these embodiments, the difference between the target biological age of the individual to be tested and the actual age of the individual to be tested is determined;
[0030] Determine whether the difference is greater than a preset difference threshold; if so, perform an aging-related causal analysis on the individual to be tested based on the difference.
[0031] Secondly, this embodiment provides a biological age prediction device, including: an acquisition module, a prediction module, and a combination module, wherein...
[0032] The acquisition module is used to acquire multi-omics data of the individual to be detected;
[0033] The prediction module is used to input the multi-omics data into a trained multi-omics data aging clock model for prediction, and obtain the predicted biological ages corresponding to the multi-omics data.
[0034] The combination module is used to combine the predicted biological ages to obtain the target biological age of the individual to be tested.
[0035] Thirdly, this embodiment provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the biological age prediction method described in the first aspect above.
[0036] Fourthly, this embodiment provides a storage medium storing a computer program that, when executed by a processor, implements the biological age prediction method described in the first aspect above.
[0037] Compared with related technologies, the biological age prediction method provided in this embodiment improves the accuracy of biological age assessment by acquiring target multi-omics data of the individual to be tested; inputting each omics data in the target multi-omics data into the corresponding trained aging clock model for prediction to obtain each predicted biological age corresponding to each omics data; and combining the predicted biological ages to obtain the target biological age of the individual to be tested.
[0038] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0039] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0040] Figure 1 This is a hardware structure block diagram of the terminal of the biological age prediction method in this embodiment.
[0041] Figure 2 This is a flowchart of the biological age prediction method in this embodiment.
[0042] Figure 3 This is a flowchart of another biological age prediction method in this embodiment.
[0043] Figure 4 This is a structural block diagram of the biological age prediction device in this embodiment. Detailed Implementation
[0044] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0045] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.
[0046] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of the terminal for the biological age prediction method in this embodiment. (See diagram for example.) Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.
[0047] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the biological age prediction method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0048] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0049] This embodiment provides a biological age prediction method. Figure 2 This is a flowchart of the biological age prediction method in this embodiment, such as... Figure 2 As shown, the process includes the following steps:
[0050] Step S201: Obtain the target multi-omics data of the individual to be tested.
[0051] Specifically, multi-omics data of the individual to be tested is acquired. This multi-omics data includes genomic data, proteomic data, metabolomic data, and radiomic data, which are used to comprehensively reveal the biological characteristics of an organism and their dynamic changes under different conditions. In this embodiment, the biological age is predicted by acquiring multi-omics data of the individual to be tested, thereby overcoming the inaccuracies caused by single-data detection in existing technologies. After acquiring the multi-omics data of the individual to be tested, preprocessing such as cleaning can be performed on the multi-omics data to remove interfering data and improve the accuracy of the prediction.
[0052] Step S202: Input each omics data in the target multi-omics data into the corresponding trained aging clock model for prediction, and obtain the predicted biological age corresponding to each omics data.
[0053] Specifically, in this embodiment, aging clock models corresponding to each omics data are pre-constructed. The specific construction process is as follows:
[0054] First, multi-omics data of the same subject are acquired through a data collection module. This can be done by requesting and downloading publicly available population cohort datasets, or by recruiting several subjects and collecting relevant cohort data from them. Healthy subjects are selected based on self-reports and standards such as the International Classification of Diseases (ICD). Subjects with common chronic diseases in the elderly, such as diabetes, hypertension, hyperlipidemia, coronary heart disease, stroke, gout, chronic renal failure, and chronic bronchitis, are considered non-healthy subjects.
[0055] Specifically, for publicly available population cohort datasets, the fields containing each omics data are located based on the experimental data description. The data is then decoded using the corresponding encoding methods to extract the omics data. This multi-omics data includes, but is not limited to, genomic data, proteomic data, metabolomic data, and radiomic data. The integrity of the multi-omics data for all participants is ensured, and incomplete data is deleted. Genomic data includes whole-genome sequencing, DNA methylation, or telomere length detection data; proteomic data includes several plasma proteins detected by a protein detection platform; radiomic data includes structural magnetic resonance imaging (MRI), diffusion-weighted MRI, and resting-state MRI of the participants; and metabolomic data includes blood glucose, blood lipids, liver and kidney function, blood cell counts, and liver enzymes.
[0056] After recruiting a number of participants and collecting their cohort data, and following ethical review by relevant institutions, multi-omics data will also be collected from the participants, including but not limited to genomic data, proteomic data, metabolomic data, and radiomic data. This embodiment does not specifically limit the types of multi-omics data; they can be selected and set according to the actual situation.
[0057] Secondly, through the data processing module, the data of each omics in the multi-omics data are preprocessed and feature extracted, and the LASSO linear regression analysis method (Least Absolute Shrinkage and Selection Operator) is used to reduce the dimensionality of the features and align the features in time.
[0058] Specifically, the process involves preprocessing and feature extraction of genomic data, obtaining DNA methylation data from the population cohort using methylation microarrays, and performing quality control on the DNA methylation data.
[0059] (1) Filter the DNA methylation data to remove samples with a genotype detection rate of <98%, inconsistent sex, and close relatives. The genotype detection rate refers to the ratio of the number of samples that successfully detect a specific genotype to the total number of samples tested during the gene testing process, and is used to evaluate the efficiency and reliability of gene testing.
[0060] (2) SNP filtering to remove sites with minor allele frequency (MAF) < 0.01, Hardy-Weinberg equilibrium test p < 1e-6 and deletion rate > 5%.
[0061] (3) Population stratification correction: The top 20 principal components (PCAs) are generated based on PLINK to eliminate differences between individuals. For methylation data, low-quality probes (detection p-value > 0.01), cross-reactive probes, and sex-related chromosome sites need to be filtered out. Batch effects are corrected using ComBat, and the Houseman algorithm is used to estimate and correct the proportion of leukocyte subtypes. Feature extraction includes methylation features, and age-related CpG sites are screened using elastic network regression.
[0062] The proteomic data was processed, including missing value detection, log transformation of the raw concentration values, and normalization.
[0063] Specifically, protein data preprocessing aims to eliminate technical variations and improve the signal-to-noise ratio. Preprocessing steps include:
[0064] (1) Data quality control: remove proteins and abnormal samples with a detection rate of <80%.
[0065] (2) Data standardization processing: the NPX value output by the Olink platform is transformed by log2 and normalization is used to eliminate the deviation between platforms.
[0066] (3) Batch calibration to eliminate batch effects in experiments.
[0067] (4) Outlier handling: extreme values are removed based on the MAD method (median absolute deviation > 3 times).
[0068] Proteomics data feature extraction stage:
[0069] (1) Variance filtering: retain the top 30% of highly variable proteins by variance.
[0070] (2) Age-related screening: Spearman analysis (FDR corrected p < 0.05) was used to identify protein biomarkers that are significantly associated with age.
[0071] The processing of the imaging data includes preprocessing of structural magnetic resonance imaging, resting-state magnetic resonance imaging, and diffusion magnetic resonance imaging. Based on brain atlases, the volume, cortical thickness, and surface area of each brain region in the structural images are extracted; functional connectivity indicators are extracted from the resting-state functional images; and anisotropy fraction, average diffusion rate, and other indicators are extracted from the diffusion images.
[0072] Specifically, the preprocessing procedure for structural magnetic resonance imaging includes merging structural images, removing non-brain structures, segmenting brain images into three different components based on the structure of gray matter, white matter, and cerebrospinal fluid, standardizing them to the MNI152 space, and performing cortical reconstruction.
[0073] The resting-state magnetic resonance imaging (MRI) data preprocessing process includes: using rigid body transformation to estimate head motion parameters and time-layer correction, spatially aligning and registering functional MRI images with structural images, standardizing different brain atlas data to the MNI152 space using a standardization process, and finally performing spatial smoothing using a Gaussian kernel with a full width at half maximum (FWHM) of 6 mm.
[0074] The diffusion magnetic resonance data preprocessing first uses the DWIdenoise method to denoise the diffusion magnetic resonance data, then performs head motion correction, eddy current correction based on the Eddy method, and distortion correction, then uses the singleShell method to reconstruct the model, and finally performs fiber tracking.
[0075] Finally, brain imaging indicators for different modalities were extracted based on the BNA atlas.
[0076] The processing of metabolomics data includes extracting hundreds of metabolites and lipoprotein particle indicators from plasma or serum, as well as indicators from routine clinical tests such as blood glucose, blood lipids, liver and kidney function, and red blood cell count. These indicators are then subjected to difference completion, extreme value removal, and normalization.
[0077] After extracting the feature metrics for each modality, the LASSO model was used to reduce the dimensionality of the multimodal features, removing features with a coefficient value of 0, while ensuring that all features were acquired in a single data collection, rather than from follow-up data.
[0078] Finally, aging clock models were established for each type of omics data. The dimensionality-reduced biological features of the multiple omics data were used as prediction features, and the actual age of the subjects was used as the prediction label. XGBoost, random forest or deep neural network were used for regression training to obtain the trained aging clock models for each omics data.
[0079] Specifically, taking radiomics data as an example, an aging clock model for radiomics data is constructed. First, all healthy subjects are divided into a training set and a test set in an 8:2 ratio. Indicators such as brain map volume, cortical thickness, surface area, functional connectivity, anisotropy score, and average diffusion rate for each subject are integrated as predictive features, with the subject's actual age as the predictive label. Any method such as XGBoost, random forest, or deep neural network is used for prediction. Finally, the radiomics data aging clock model is tested on the test set to obtain the optimal radiomics data aging clock model, i.e., the trained radiomics data aging clock model.
[0080] Following the same approach described above, aging clock models based on any of the methods such as XGBoost, random forest, or deep neural networks were constructed using genomic data, proteomic data, and metabolomic data, respectively, resulting in aging clock models based on genomic data, proteomic data, and metabolomic data.
[0081] For example, the target multi-omics data includes genomic data, proteomic data, metabolomic data, and radiomic data; each omics data in the target multi-omics data is input into the corresponding trained aging clock model for prediction, obtaining the predicted biological age corresponding to each omics data, including:
[0082] The genomic data of the individual to be tested is input into a trained genomic data aging clock model for prediction, and the predicted biological age corresponding to the genomic data is obtained.
[0083] The proteomic data of the individual to be tested is input into a trained proteomic data aging clock model for prediction, and the predicted biological age corresponding to the proteomic data is obtained.
[0084] The metabolomics data of the individual to be tested is input into a trained metabolomics data aging clock model for prediction, and the predicted biological age corresponding to the metabolomics data is obtained.
[0085] The image group data of the individual to be tested is input into the trained image group data aging clock model for prediction, and the predicted biological age corresponding to the image group data is obtained.
[0086] That is, by using the trained aging clock model, predictions are made on each omics data in the target multi-omics data of the individual to be tested, and the DNA methylation age, protein age, metabolic age and brain age of the individual to be tested are obtained.
[0087] Step S203: Combine the predicted biological ages to obtain the target biological age of the individual to be tested.
[0088] For example, combining the predicted biological ages to obtain the target biological age of the individual to be tested includes: setting corresponding weight coefficients based on the predictive power of each model in the aging clock model for biological age; and combining the predicted biological ages based on the weight coefficients to obtain the target biological age of the individual to be tested.
[0089] Specifically, based on the accuracy and reliability of each aging clock model, different weights are assigned to the biological ages predicted by each model. A weighted average is then calculated to obtain the target biological age of the individual being tested. For example, if a genomics-based prediction model outperforms other models on the validation dataset, it can be given a higher weight. Alternatively, a fusion model can be constructed using a fusion algorithm, taking the outputs of multiple aging clock models as input and further training to optimize the final prediction results.
[0090] Alternatively, a linear regression model can be used to combine the predicted biological ages to obtain the target biological age for the individual being tested. Specifically, this includes:
[0091] First, Z-score standardization is performed on each predicted biological age. The specific formula is as follows:
[0092] ;
[0093] Where, μ mod σ represents the mean of all predicted biological ages. mod represents the standard deviation of each predicted biological age, and Age represents each predicted biological age.
[0094] Next, a linear regression model is constructed based on the standardized predicted biological ages:
[0095] ;
[0096] The coefficients β1, β2, β3, and β4 were solved using the least squares method, and these coefficients represent the contribution weights of each predicted biological age. This is the error term.
[0097] Using a linear regression model Age chrono The predicted biological ages are combined to obtain the target biological age of the individual being tested.
[0098] After obtaining the target biological age of the individual being tested, this embodiment can also calculate the age difference between the target biological age and the actual age of the individual being tested, and based on the pre-constructed causal relationship between the age difference and chronic diseases of the elderly, correlate the individual being tested with chronic diseases of the elderly, thereby obtaining aging factors related to chronic diseases of the elderly, and generating an individual aging report related to chronic diseases of the elderly.
[0099] All of the steps S201 to S203 described above can be executed on a computer or other devices with data processing capabilities.
[0100] Through steps S201 to S203, target multi-omics data for the individual to be tested are obtained. Each omics data point in the target multi-omics data is then input into a corresponding trained aging clock model for prediction, yielding the predicted biological age for each omics data point. These predicted biological ages are then combined to obtain the target biological age for the individual to be tested. Compared to existing technologies that rely on single-omics data to predict biological age, this embodiment first establishes a corresponding aging clock model for each omics data point based on healthy subjects. Each aging clock model then predicts the multi-omics data of the individual to be tested, obtaining the biological age for each omics data point. Finally, these predicted biological ages are combined to obtain the target biological age for the individual to be tested. By using multi-omics data for comprehensive prediction of biological age, the accuracy of biological age assessment is improved.
[0101] In some embodiments, acquiring the target multi-omics data of the individual to be detected further includes:
[0102] Obtain the original multi-omics data of the individuals to be tested; preprocess and extract features from the original multi-omics data to obtain the pre-processed multi-omics data; perform feature dimensionality reduction and time alignment on the pre-processed multi-omics data to obtain the target multi-omics data.
[0103] Specifically, before predicting the biological age of the individual under test using multi-omics data, the multi-omics data is preprocessed and features are extracted. The types of multi-omics data for the individual under test are the same as those used in the aging clock model construction process corresponding to each omics data in step S202 above, and the data processing methods are also the same, which will not be elaborated further here.
[0104] In another embodiment, genomic data includes whole-genome sequencing, DNA methylation, or telomere length detection data; proteomic data includes cardiac metabolic proteins, inflammatory proteins, and oncological proteins; metabolomic data includes blood glucose, blood lipids, liver and kidney function, blood cell counts, and liver enzymes; and radiomic data includes structural magnetic resonance imaging, diffusion magnetic resonance imaging, and resting-state magnetic resonance imaging.
[0105] Specifically, the multi-omics data of the individuals being tested are the same data types used in the training process of each aging clock model. Genomic data includes whole-genome sequencing, DNA methylation, or telomere length detection data. Whole-genome sequencing refers to sequencing an individual's entire genome to obtain complete DNA sequence information, including the coding and non-coding regions of all genes, and can detect single nucleotide polymorphisms (SNPs), insertions / deletions (InDels), copy number variations (CNVs), and structural variations (SVs). DNA methylation refers to the addition of methyl groups to DNA molecules, typically occurring on the CpG dinucleotide of cytosine. By detecting DNA methylation levels, we can understand gene expression regulation, cell differentiation, and disease mechanisms. Telomeres are repetitive sequences at the ends of chromosomes, and their length is closely related to cellular aging and carcinogenesis. By detecting telomere length, we can assess the aging state of cells and the biological age of an individual.
[0106] The proteomic data included cardiac metabolic proteins, inflammatory proteins, and oncological proteins. Cardiac metabolic proteins are those related to cardiac metabolic processes, including energy metabolism, oxidative stress, and signal transduction, such as cardiac troponin, brain natriuretic peptide (BNP), and creatine kinase isoenzyme (CK-MB). Inflammatory proteins are those that play a key role in the inflammatory response, including cytokines, chemokines, and acute-phase proteins, such as C-reactive protein (CRP), interleukin-6 (IL-6), and tumor necrosis factor-α (TNF-α). Oncological proteins are those related to tumor development, progression, and metastasis, including tumor markers and cell proliferation-related proteins, such as carcinoembryonic antigen (CEA), alpha-fetoprotein (AFP), and prostate-specific antigen (PSA).
[0107] Metabolomics data include blood glucose, blood lipids, liver and kidney function, blood cell count, and liver enzymes. Blood glucose refers to the glucose content in the blood and is the body's main energy source. It includes fasting blood glucose, postprandial blood glucose, and glycated hemoglobin (HbA1c). Blood lipids refer to lipids in the blood, including cholesterol, triglycerides, high-density lipoprotein (HDL), and low-density lipoprotein (LDL). Examples include total cholesterol (TC), triglycerides (TG), HDL-C, and LDL-C. Liver and kidney function indicators reflect the metabolic and excretory functions of the liver and kidneys. Examples include liver function indicators (ALT, AST, ALP, GGT, etc.) and kidney function indicators (creatinine, blood urea nitrogen, uric acid, etc.). Blood cell count includes the number of red blood cells, white blood cells, and platelets and their related indicators. Examples include red blood cell count (RBC), hemoglobin (HGB), white blood cell count (WBC), and platelet count (PLT). Liver enzymes are a class of enzymes in the liver; changes in their levels can reflect liver damage and functional status. Examples include alanine aminotransferase (ALT), aspartate aminotransferase (AST), alkaline phosphatase (ALP), and gamma-glutamyl transferase (GGT).
[0108] The imaging data included structural magnetic resonance imaging (SMRI), diffusion-weighted magnetic resonance imaging (DMRI), and resting-state MRI. SMRI is a technique for acquiring high-resolution structural images of the brain and other organs, including gray and white matter volume, cortical thickness, and ventricular size. DMRI is a technique for measuring the diffusion of water molecules in tissues, including diffusion tensor imaging (DTI) parameters such as apparent diffusion coefficient (ADC) and fractional anisotropy (FA). Resting-state MRI measures functional connectivity between different brain regions while the subject is at rest, including functional connectivity strength, amplitude at low frequencies (ALFF), and local coherence (ReHo).
[0109] In summary, genomic data, proteomic data, metabolomic data, and radiomic data together constitute the components of multi-omics data. By integrating these multi-omics data, we can more accurately predict biological age, diagnose diseases, evaluate treatment effects, and provide a scientific basis for personalized medicine.
[0110] In some embodiments, the difference between the target biological age of the individual to be tested and the actual age of the individual to be tested is determined; it is determined whether the difference is greater than a preset difference threshold; if so, an aging-related causal analysis is performed on the individual to be tested based on the difference.
[0111] Specifically, currently, due to insufficient clinical translation, existing aging clock models lack a framework for association analysis with chronic diseases of the elderly (such as Alzheimer's disease and cardiovascular diseases), limiting their application value in clinical diagnosis and treatment. Based on this, this embodiment acquires multi-omics data from subjects with chronic diseases, uses a trained aging clock model to predict the biological age of these subjects, identifies individuals experiencing accelerated aging based on the deviation between biological age and chronological age, and analyzes its correlation with chronic diseases and clinical phenotypes of the elderly. This allows for the analysis of aging-related causes based on the target biological age of the individual being tested. The specific process includes:
[0112] The data acquisition module collects non-healthy multi-omics data from subjects with chronic diseases, including diabetes, hypertension, hyperlipidemia, coronary heart disease, stroke, gout, chronic renal failure, and chronic bronchitis.
[0113] Based on the aging clock model trained in step S202 above, biological age prediction is performed on non-healthy multi-omics data to obtain the biological age corresponding to each omics data. The predicted biological ages corresponding to each non-healthy multi-omics data are combined to obtain the target biological age of subjects with chronic diseases.
[0114] The study calculated the difference between the target biological age and the chronological age (ΔAge) of participants with chronic diseases. Participants with ΔAge > 0 were classified as experiencing accelerated aging, while those with ΔAge < 0 were classified as experiencing slowed aging. Participants exceeding their chronological age by 10% were identified as the significantly accelerated aging group, and those below 10% were identified as the slowed aging group. Longitudinal follow-up data from the population cohort were used to calculate the incidence of chronic diseases in these participants using survival analysis, validating the predictive power of ΔAge for chronic diseases. Secondly, a mediation analysis was conducted to construct the pathway of "multi-omics factors—age difference—chronic diseases in the elderly." Multi-omics factors refer to the set of features with high predictive weights in the aging clock model. This analytical framework reveals how abnormalities in certain key factors trigger the difference between biological age and chronological age, further promoting the occurrence of chronic diseases. Based on this mediation effect, earlier behavioral or pharmacological intervention recommendations can be provided to individuals, thereby achieving precise aging and disease prevention and control.
[0115] Next, Mendelian randomization was used to verify the causal relationship between accelerated aging and the target disease. Publicly available GWAS datasets on multiple groups of biological age differences and chronic diseases were searched, with biological age difference as the exposure factor and chronic disease as the outcome factor. First, a p-value < 5 × 10⁻⁶ was set. -8SNPs strongly correlated with the exposure factor were identified as potential instrumental variables. Then, linkage-disequilibrium genetic variants were removed, retaining only the more significant SNPs. An F<10 threshold was set to remove instrumental variables weakly correlated with the exposure factor. Confounding factors leading to chronic disease, such as economic status and gender, were also removed. Finally, Mendelian randomization analysis was performed using inverse variance weighting to determine the causal relationship between biological age difference and chronic disease.
[0116] If the difference between the target biological age and the actual age of the individual being tested exceeds a preset difference threshold, such as ΔAge>0, the individual is determined to be aging rapidly. Then, based on the causal relationship between the biological age difference and chronic diseases in the above analysis process, the causes of aging of the individual being tested are matched and analyzed to generate an individual aging report, which includes biological age, risk of accelerated aging, and disease warning suggestions.
[0117] This embodiment also provides a method for predicting biological age. Figure 3 This is a flowchart of another biological age prediction method in this embodiment, such as... Figure 3 As shown, the process includes the following steps:
[0118] Step S301: Obtain the raw multi-omics data of the individual to be tested; preprocess and extract features from the raw multi-omics data to obtain pre-processed multi-omics data; perform feature dimensionality reduction and time alignment on the pre-processed multi-omics data to obtain target multi-omics data; the target multi-omics data includes genomic data, proteomic data, metabolomic data and radiomic data.
[0119] Step S302: Input the genomic data of the individual to be tested into the trained genomic data aging clock model for prediction, and obtain the predicted biological age corresponding to the genomic data;
[0120] Step S303: Input the proteomic data of the individual to be tested into the trained proteomic data aging clock model for prediction, and obtain the predicted biological age corresponding to the proteomic data.
[0121] Step S304: Input the metabolomics data of the individual to be tested into the trained metabolomics data aging clock model for prediction, and obtain the predicted biological age corresponding to the metabolomics data.
[0122] Step S305: Input the image group data of the individual to be tested into the trained image group data aging clock model for prediction, and obtain the predicted biological age corresponding to the image group data.
[0123] Step S306: Based on the predictive power of each model in the aging clock model for biological age, set the corresponding weight coefficients; perform linear regression on each predicted biological age based on the weight coefficients to obtain the target biological age of the individual to be tested.
[0124] Step S307: Calculate the age difference between the target biological age and the actual age of the individual to be tested; based on the pre-constructed causal relationship between the age difference and chronic diseases of the elderly, match the causes of aging of the individual to be tested, and generate an individual aging report.
[0125] Through steps S301 to S307, compared with the prior art which relies on a single omics data to predict biological age, this embodiment pre-trains a multi-omics data clock aging model and the causal relationship between biological age and chronic diseases, and predicts the biological age of the test individual using multi-omics data respectively, obtaining the biological age corresponding to each omics data; then, a linear regression model is used to combine the biological ages to predict the target biological age, improving the accuracy of biological age assessment; then, individuals with accelerated aging are identified based on the deviation between biological age and actual age, and their correlation with chronic diseases of the elderly is analyzed.
[0126] This embodiment also provides a biological age prediction device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below refer to combinations of software and / or hardware that perform a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0127] Figure 4 This is a structural block diagram of the biological age prediction device in this embodiment, as shown below. Figure 4 As shown, the device 40 includes: an acquisition module 41, a prediction module 42, a combination module 43, and an aging report generation module 44, wherein,
[0128] Module 41 is used to acquire multi-omics data of the individual to be tested;
[0129] Prediction module 42 is used to input multi-omics data into a trained multi-omics data aging clock model for prediction, and obtain the predicted biological ages corresponding to the multi-omics data.
[0130] The combination module 43 is used to combine the predicted biological ages to obtain the target biological age of the individual to be tested;
[0131] The aging report generation module 44 is used to calculate the age difference between the target biological age and the actual age of the individual being tested; based on the pre-constructed causal relationship between the age difference and chronic diseases of the elderly, the aging causes of the individual being tested are matched and an individual aging report is generated.
[0132] This embodiment also provides an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0133] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0134] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0135] S1, acquire the target multi-omics data of the individual to be tested;
[0136] S2, input each omics data in the target multi-omics data into the corresponding trained aging clock model for prediction, and obtain the predicted biological age corresponding to each omics data respectively;
[0137] S3, combine the predicted biological ages to obtain the target biological age of the individual to be tested.
[0138] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.
[0139] Furthermore, in conjunction with the biological age prediction methods provided in the above embodiments, this embodiment can also provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the biological age prediction methods described in the above embodiments.
[0140] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0142] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.
[0143] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0144] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0145] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.
Claims
1. A method of biological age prediction, characterized in that, The method comprises the following steps: obtaining target multi-omics data of a to-be-detected individual; inputting each omics data in the target multi-omics data into a corresponding trained aging clock model for prediction to obtain each predicted biological age corresponding to each omics data; combining the predicted biological ages to obtain a target biological age of the to-be-detected individual; calculating an age difference between the target biological age of the to-be-detected individual and an actual age of the to-be-detected individual; generating an individual aging report by matching the aging causes of the to-be-detected individual according to a pre-constructed causal relationship between the age difference and an old chronic disease; wherein the target multi-omics data comprises genomic data, proteomic data, metabolomic data and imaging data; wherein the imaging data comprises structural magnetic resonance imaging, diffusion magnetic resonance imaging and resting-state magnetic resonance imaging; combining the predicted biological ages to obtain the target biological age of the to-be-detected individual comprises: setting a corresponding weight coefficient according to the prediction ability of each model in the aging clock model to each biological age; and performing linear regression on the predicted biological ages according to the weight coefficient to obtain the target biological age of the to-be-detected individual; the construction process of the causal relationship between the age difference and the old chronic disease comprises: based on the collected non-healthy multi-omics data of a subject suffering from an old chronic disease, performing biological age prediction on the non-healthy multi-omics data by using a trained aging clock model to obtain biological ages corresponding to each non-healthy multi-omics data; combining the biological ages corresponding to each non-healthy multi-omics data to obtain a target biological age of the subject suffering from the old chronic disease, and calculating a difference ΔAge between the target biological age of the subject suffering from the old chronic disease and an actual age; combining longitudinal follow-up data of a population cohort, using survival analysis to calculate the incidence rate of the old chronic disease of the subject, and verifying the prediction performance of the difference ΔAge on the old chronic disease; constructing the action path of "multi-omics factors - age difference - old chronic disease" through mediation analysis; verifying the causal relationship between accelerated aging and the target disease by using Mendelian randomization method; based on the causal relationship, constructing the causal relationship between the age difference and the old chronic disease.
2. The biological age prediction method according to claim 1, characterized in that, The method further comprises the following steps: obtaining original multi-omics data of the to-be-detected individual; performing preprocessing and feature extraction on the original multi-omics data to obtain preliminarily processed multi-omics data; performing feature dimension reduction and time alignment processing on the preliminarily processed multi-omics data to obtain the target multi-omics data.
3. The biological age prediction method according to claim 1, characterized in that, the step of inputting each omics data in the target multi-omics data into a corresponding trained aging clock model for prediction to obtain each predicted biological age corresponding to each omics data comprises: inputting the genomic data of the to-be-detected individual into a trained genomic data aging clock model for prediction to obtain a predicted biological age corresponding to the genomic data; inputting the proteome data of the individual to be detected into the trained proteome data aging clock model for prediction to obtain a predicted biological age corresponding to the proteome data; inputting the metabolome data of the individual to be detected into the trained metabolome data aging clock model for prediction to obtain a predicted biological age corresponding to the metabolome data; inputting the image data of the individual to be detected into the trained image data aging clock model for prediction to obtain a predicted biological age corresponding to the image data.
4. The biological age prediction method according to claim 3, characterized in that, the genomic data comprises whole genome sequencing, DNA methylation or telomere length detection data; the proteome data comprises cardiometabolic proteins, inflammatory proteins and oncology proteins; the metabolome data comprises blood glucose, blood lipids, liver and kidney function, blood cell count and liver enzymes.
5. A biological age prediction device characterized by comprising: comprises: an acquisition module, a prediction module, a combination module and an aging report generation module, wherein the acquisition module is configured to acquire target multi-omics data of an individual to be detected; the prediction module is configured to input each omics data in the target multi-omics data into a corresponding trained aging clock model for prediction to obtain each predicted biological age corresponding to the each omics data; the combination module is configured to combine the each predicted biological age to obtain a target biological age of the individual to be detected; the combination of the each predicted biological age to obtain the target biological age of the individual to be detected comprises: setting a corresponding weight coefficient according to the prediction ability of each model in the aging clock model for biological age; performing linear regression on the each predicted biological age according to the weight coefficient to obtain the target biological age of the individual to be detected; the aging report generation module is configured to calculate an age difference between the target biological age of the individual to be detected and an actual age of the individual to be detected; after matching the aging cause of the individual to be detected according to a pre-constructed causal relationship between the age difference and the old chronic disease, an individual aging report is generated; wherein the target multi-omics data comprises genomic data, proteome data, metabolome data and image data; wherein the image data comprises structural magnetic resonance imaging, diffusion magnetic resonance imaging and resting state magnetic resonance imaging; the construction process of the causal relationship between the age difference and the old chronic disease comprises: based on the collected non-healthy multi-omics data of subjects with old chronic diseases, the trained aging clock model is used to predict the biological age of the non-healthy multi-omics data to obtain the biological age corresponding to each non-healthy multi-omics data; combining the biological age corresponding to each non-healthy multi-omics data to obtain the target biological age of the subjects with old chronic diseases, calculating the difference ΔAge between the target biological age of the subjects with old chronic diseases and the actual age; In combination with longitudinal follow-up data of the cohort, the survival analysis is used to calculate the incidence of chronic diseases in the elderly of the subjects, and the predictive performance of the age difference ΔAge on the chronic diseases in the elderly is verified; The mediation analysis is used to construct the action path of "multi-omics factors - age difference - chronic diseases in the elderly"; The Mendelian randomization method is used to verify the causal relationship between accelerated aging and target diseases; Based on the causal relationship, the causal relationship between the age difference and the chronic diseases in the elderly is constructed. 6.An electronic device comprising a memory and a processor, the electronic device comprising: The memory stores a computer program, and the processor is configured to run the computer, and the processor is configured to run the computer program to execute the biological age prediction method in any one of claims 1-4.
7. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the biological age prediction method in any one of claims 1-4.
Citation Information
Patent Citations
Aging clock construction method based on causal machine joint model
CN120183731A
Biological data signatures of aging and methods of determining a biological aging clock
US20200286625A1
Systems and methods for a novel image-based multi-omics aging clock for the prediction of remaining lifespan
WO2024006917A1