Method for constructing multiple-group biological age metabolic aging scores and risk prediction and application thereof
Patent Information
- Application Number
- CN202611328961.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
然而,目前相关研究仍处于探索阶段,多组学信息的整合方式尚不成熟,对不同组学之间复杂非线性关系的挖掘仍然不足,难以充分发挥多维生物学数据的互补优势
[0040]本发明方法通过构建融合多组学生物年龄特征的代谢疾病衰老评分体系,可以有效解决现有技术中代谢健康评价维度单一、难以量化生物学衰老本质以及缺乏精准风险分层能力的突出问题。本发明首先通过整合临床表型、蛋白质组及代谢组多维数据源,突破传统评价体系仅依赖体格测量或单一生化指标的局限,为代谢衰老的量化提供更为全面的数据基础。在此基础上,本发明创新性地引入临床、蛋白组及代谢组三个层面的生物年龄预测模型,将各模型输出的预测生物学年龄作为核心特征输入,这一设计不仅实现了对机体不同层级衰老状态的独立刻画,更通过多源生物年龄特征的融合,精准捕捉了代谢系统生物学衰老的协同变化规律,相较于单一组学生物标志物模型,可以显著提升特征空间的生物学代表性,实现对代谢衰老状态的全面量化评价。
Smart Images

Figure CN122842950A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of healthcare information technology, and in particular to the construction and risk prediction methods and applications of multi-group biological age metabolic aging scores. Background Technology
[0002] With the accelerating aging of the population and changes in lifestyle and dietary structure, the incidence of obesity and related metabolic diseases (such as type 2 diabetes, hypertension, coronary heart disease, and chronic kidney disease) continues to rise, and the metabolic burden is constantly increasing. Metabolic diseases have become a major public health problem affecting human health worldwide. Numerous studies have shown that obesity is not only characterized by excessive accumulation of adipose tissue, but is also accompanied by chronic low-grade inflammation, immune imbalance, cellular senescence, and decline in the function of multiple organs, and is considered an important risk factor for biological aging.
[0003] Currently, body mass index (BMI) is the primary clinical measure used to assess obesity and assess disease risk. However, a growing body of research reveals significant heterogeneity in metabolic health among individuals with the same BMI. For instance, some obese individuals have high BMIs but normal glucose and lipid metabolism, termed metabolically healthy obesity (MHO); others exhibit significant insulin resistance, dyslipidemia, hypertension, and other metabolic abnormalities, termed metabolically unhealthy obesity (MUO). Simultaneously, some individuals with normal BMIs also exhibit significant metabolic abnormalities, known as normal weight metabolically abnormal obesity (NWMO). Increasing evidence-based medicine research indicates that the risk of cardiovascular metabolic diseases depends not only on the degree of obesity but also on the individual's overall metabolic health. Therefore, relying solely on BMI to assess obesity is insufficient to accurately reflect an individual's true metabolic health status and fails to meet the needs of precision medicine for early identification and risk stratification of high-risk groups.
[0004] To compensate for the limitations of BMI, current clinical practice relies on traditional metabolic indicators such as blood glucose, blood lipids, blood pressure, and waist circumference to assess health risks. While these indicators have significant clinical value in the diagnosis and treatment of metabolic diseases, they primarily reflect metabolic abnormalities after disease onset, belonging to the phenotypic level of disease detection and failing to capture early biological changes in disease development. Furthermore, these indicators are independent, reflecting only one aspect of the body's metabolic state, and cannot comprehensively integrate multidimensional biological information such as inflammatory responses, immune regulation, proteomics, metabolomics, and organ function, thus failing to fully reveal the body's overall metabolic health level and its dynamic changes. In addition, traditional evaluation systems mainly focus on clinical metabolic abnormalities, lacking a systematic evaluation of the body's biological aging process, and cannot accurately quantify the degree of metabolic aging, nor can they achieve early identification and precise risk stratification of high-risk groups for metabolic abnormalities.
[0005] In recent years, biological age, as an important indicator for measuring the aging state of the body, has become a significant research direction in the fields of aging medicine and precision medicine. Compared with chronological age, biological age can comprehensively reflect the degree of functional decline of human tissues and organs and the risk of future diseases, and is considered to be a better representation of an individual's true aging state. Currently, biological age models based on different levels such as clinical indicators, biochemical indicators, proteomics, and metabolomics have been established both domestically and internationally, such as PhenoAge, BioAge, Proteomic Aging Clock, and MetaboAge. These models provide new technical means for the quantitative evaluation of biological aging.
[0006] However, most existing biological age models primarily aim to assess overall aging levels or all-cause mortality risk. Their design is mainly based on the general population and has not been optimized for the heterogeneity of metabolic health status. Furthermore, they struggle to accurately identify different metabolic phenotypes, such as metabolically healthy obesity, metabolically unhealthy obesity, and metabolic abnormalities in individuals with normal weight. According to the Developmental Origins of Health and Disease (DOHaD) theory, the onset of metabolic diseases in adulthood does not entirely begin at the clinical disease stage. Instead, it may be influenced by long-term nutritional, metabolic, and environmental exposures in early life, accumulating through continuous biological changes, ultimately manifesting as adult obesity, insulin resistance, type 2 diabetes, and cardiovascular metabolic diseases. Therefore, metabolic abnormalities are not only clinical manifestations of disease but may also continuously drive the body's biological aging process, causing individuals to experience varying degrees of advanced metabolic-related biological age before a clear clinical disease endpoint is reached. Thus, shifting the focus from traditional disease risk assessment to the early identification and intervention of "metabolic abnormality-driven biological aging" may become an important new strategy for the early prevention and control of metabolic diseases.
[0007] Furthermore, existing biological age models are typically built based on a single data dimension, reflecting only a certain level of aging characteristics. They lack comprehensive integration of multidimensional biological information such as clinical indicators, proteomics, and metabolomics, and cannot fully characterize the systemic biological aging process caused by metabolic abnormalities. Especially for individuals with similar ages or BMIs but different metabolic health statuses, traditional biological age indicators struggle to accurately reflect their potential metabolic aging degree and disease risk. Therefore, there is an urgent need to establish a comprehensive evaluation method that focuses on metabolic abnormality-related biological changes and integrates clinical phenotypes and multidimensional omics information. This method would integrate metabolic abnormalities with biological aging status, enabling early identification, risk assessment, and precise intervention of metabolic-driven biological aging, providing new technical means for the early prevention, control, and individualized management of obesity and related metabolic diseases. In recent years, with the rapid development of high-throughput omics technologies and artificial intelligence algorithms, multi-omics data fusion analysis has gradually become an important research direction for the precise diagnosis and risk prediction of metabolic diseases. Existing studies have begun to attempt to combine multidimensional data such as clinical indicators, proteomics, and metabolomics to construct disease prediction models, which has improved the ability to identify metabolic abnormalities to some extent. However, current research is still in the exploratory stage. The integration methods for multi-omics information are not yet mature, and the exploration of complex nonlinear relationships between different omics remains insufficient, making it difficult to fully leverage the complementary advantages of multidimensional biological data. Furthermore, most existing studies remain at the model construction and cross-sectional evaluation stage, lacking systematic validation of model predictive performance and clinical application value based on large-scale population long-term follow-up data. In particular, there is still insufficient evidence regarding the predictive ability of important clinical outcomes such as mortality, type 2 diabetes, chronic kidney disease, and cardiovascular disease. At the same time, existing evaluation systems rarely incorporate genetic methods such as genome-wide association studies (GWAS), Mendelian randomization (MR), and colocalization analysis, lacking systematic validation of the biological mechanisms and causal relationships reflected by the models. Therefore, their scientific rigor and reliability need further improvement.
[0008] Therefore, there is still a lack of a comprehensive metabolic aging assessment system that can effectively integrate multidimensional omics information, make full use of machine learning algorithms to mine complex biological characteristics, and combine long-term clinical follow-up and genetic evidence for systematic validation, so as to achieve accurate identification of metabolic health status, risk stratification, and early prediction of cardiovascular metabolic diseases. Summary of the Invention
[0009] In view of the technical problems of existing metabolic health assessment systems, such as difficulty in comprehensively reflecting the degree of biological aging related to metabolism, lack of integration of multidimensional omics information, and insufficient ability to accurately identify metabolic health status, this invention provides a method and application for constructing and predicting the risk of multi-omics biological age metabolic aging scores, which can achieve accurate identification of metabolic health status, risk stratification, and early prediction of cardiovascular metabolic diseases.
[0010] To achieve the above and related objectives, the present invention adopts the following technical solution:
[0011] The first aspect of this invention provides a method for constructing and predicting the risk of metabolic aging scores for multiple biological ages, comprising the following steps:
[0012] Step S100: Obtain clinical phenotype data, proteomics data and metabolomics data of the target object, and construct a multidimensional feature dataset;
[0013] Step S200: Based on the multidimensional feature dataset, construct biological age prediction models at the clinical, proteomic, and metabolomic levels respectively, and use the predicted biological age output by each model as biological age features.
[0014] Step S300: The biological age characteristics are fused with the aging-related features in the multidimensional feature dataset. The preset metabolic health phenotype classification is used as a supervision signal to train a multi-omics fusion model to identify aging-related factors driven by metabolic disorders based on the distinction of different metabolic types.
[0015] Step S400: Based on the importance contribution and effect direction of each input feature in the multi-omics fusion model, determine the feature weights, and generate a metabolic disease aging score by weighted summation, wherein the metabolic disease aging score reflects the contribution of factors driving aging by metabolic disorders.
[0016] Step S500: Based on the metabolic disease aging score, predict and stratify the risk of metabolic-related diseases and / or the risk of aging progression in the target subjects.
[0017] Furthermore, in step S300, the metabolic health phenotype classification includes normal weight metabolic health, normal weight metabolic abnormality, metabolic health obesity, and metabolic abnormality obesity.
[0018] Furthermore, before training the multi-omics fusion model, a sample equalization step is also included: the training samples are processed using synthetic minority oversampling technology, and the similarity between samples is measured by heterogeneous Euclidean overlap distance to balance the differences in the number of samples between different metabolic health phenotypes.
[0019] Furthermore, in step S400, the importance contribution is quantified using the feature gain value during model training, and the effect direction is quantified using the SHapley additive interpretation value; the weight coefficients of each input feature are determined by integrating the feature gain value and the additive interpretation value, so as to identify features that make a significant contribution to metabolic disorders driving aging.
[0020] Furthermore, step S500 also includes: using the metabolic disease aging score as a continuous variable, and using a survival analysis model to assess its dose-response relationship with the risk of all-cause mortality and multiple age-related chronic diseases;
[0021] Based on the distribution characteristics of metabolic disease aging scores and the distribution differences within different metabolic health phenotypes, an overall risk stratification system and a stratified risk assessment system were established to classify the target population into high, medium and low risk levels.
[0022] Furthermore, the method also includes a genetic attribution step:
[0023] Genome-wide association analysis was performed using metabolic disease aging scores as a phenotype to identify significantly associated genetic loci;
[0024] We constructed a weighted genetic risk score based on significantly associated genetic loci to quantify the explanatory power of genetic factors on metabolic disease aging scores and to help identify the genetic basis associated with metabolic disorders driving aging.
[0025] Furthermore, the method also includes a causal inference step:
[0026] Using independent genetic variations obtained from genome-wide association analysis as instrumental variables, we employed two-sample Mendelian randomization analysis and colocalization analysis to verify the causal association between metabolic disease aging scores and metabolic-related diseases and / or aging-related outcomes.
[0027] Furthermore, the method also includes cross-cohort validation and intervention steps:
[0028] In independent cohorts of children and adolescents, validate the consistency of significant associations of genetic loci and / or metabolic disease aging score-related features;
[0029] Analyze the correlation between significant associated genetic loci and / or metabolic disease aging score-related features and serum amino acid metabolites, adipokines, and liver factors;
[0030] To analyze gene-environment interactions and assess the protective effect of lifestyle interventions on individuals carrying high-risk genetic traits or metabolic aging risk characteristics.
[0031] The second aspect of this invention provides a multi-group biological age metabolic aging score construction and risk prediction system, including:
[0032] The acquisition module is used to acquire clinical phenotypic data, proteomics data, and metabolomics data of the target object and construct a multidimensional feature dataset.
[0033] The module is used to build biological age prediction models at the clinical, proteomic, and metabolomic levels based on multidimensional feature datasets, and the predicted biological age output by each model is used as biological age feature.
[0034] The training module is used to fuse biological age characteristics with aging-related features in a multidimensional feature dataset. Using a pre-defined metabolic health phenotype classification as a supervision signal, it trains a multi-omics fusion model to identify aging-related factors driven by metabolic disorders based on the distinction between different metabolic types.
[0035] The calculation module is used to determine the feature weights based on the importance contribution and effect direction of each input feature in the multi-omics fusion model, and generate a metabolic disease aging score by weighted summation, wherein the metabolic disease aging score reflects the contribution of factors driving aging by metabolic disorders.
[0036] The prediction and stratification module is used to predict and stratify the risk of metabolic-related diseases and / or the risk of aging progression in target subjects based on the metabolic disease aging score.
[0037] A third aspect of the present invention provides a computer-readable storage medium storing computer-readable instructions, which, when executed by a computer processor, cause the computer to execute the aforementioned method for constructing and predicting the metabolic aging scores of multiple biological ages.
[0038] A fourth aspect of the present invention provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-mentioned method for constructing and predicting the metabolic aging scores of multiple biological ages.
[0039] The beneficial technical effects of this invention are as follows:
[0040] This invention addresses the prominent problems of existing metabolic health assessment methods, such as limited dimensions, difficulty in quantifying the biological nature of aging, and lack of precise risk stratification, by constructing a metabolic disease aging scoring system that integrates multiple biomarker characteristics. Firstly, by integrating multi-dimensional data sources including clinical phenotypes, proteomics, and metabolomics, this invention overcomes the limitations of traditional assessment systems that rely solely on anthropometric measurements or single biochemical indicators, providing a more comprehensive data foundation for the quantification of metabolic aging. Secondly, this invention innovatively introduces biological age prediction models at three levels: clinical, proteomics, and metabolomics. The predicted biological age output by each model is used as the core feature input. This design not only independently characterizes different levels of aging but also accurately captures the synergistic changes in the biological aging of the metabolic system through the fusion of multi-source biological age features. Compared to a single biomarker model, this significantly improves the biological representativeness of the feature space, enabling a comprehensive quantitative evaluation of metabolic aging.
[0041] Furthermore, this invention uses a pre-defined metabolic health phenotype classification (including four categories: normal weight metabolic health (NWMH), metabolic health obesity (MHO), normal weight metabolic abnormality (NWMO), and metabolic abnormality obesity (MUO)) as a supervisory signal to train a multi-omics fusion model. This allows the model's learning objective to be directly anchored to the heterogeneous characteristics of metabolic health, rather than generalized aging phenotypes. This effectively solves the problem of traditional methods struggling to distinguish between complex phenotypes such as metabolic health obesity and metabolic abnormality obesity, achieving accurate identification of metabolic-related aging in different metabolic health phenotypes. In model construction, this invention employs machine learning algorithms to fully explore the complex nonlinear relationships between multi-omics data and innovatively combines feature importance (Gain) and SHAP interpretation methods to determine the weights of each feature to construct the MDAS comprehensive score. This strategy not only improves the model's classification ability and stability but also endows the scoring system with clear biological interpretability, overcoming the difficulty of tracing the origins of black-box models in clinical applications and significantly enhancing the model's generalization ability.
[0042] The metabolic disease aging score constructed in this invention, as a continuous variable, can quantify the cumulative degree of metabolic-related biological aging in an individual and possesses significant predictive ability for long-term cardiovascular metabolic disease risk. This invention demonstrates that the MDAS can not only evaluate the current degree of metabolic aging but also predict the risk of all-cause mortality and various age-related chronic diseases such as type 2 diabetes, hypertension, heart failure, myocardial infarction, stroke, and chronic kidney disease. It can be used for long-term health risk assessment and dynamic monitoring, providing a basis for early disease warning and individualized intervention.
[0043] To solidify the scientific foundation of the evaluation system, this invention further combines genome-wide association analysis (GWAS), weighted genetic risk score (wGRS), Mendelian randomization (MR), and colocalization analysis to systematically elucidate the genetic basis of MDAS and its potential causal relationship with age-related diseases. It establishes a genetically based method for evaluating metabolic aging, improves the scientific rigor and reliability of the evaluation system, and provides new technical support for research on metabolic aging-related mechanisms and the discovery of potential intervention targets.
[0044] In summary, compared with the prior art, the present invention has achieved substantial improvements in evaluation depth, phenotypic resolution, and clinical guidance value. The constructed MDAS can be widely used in multiple scenarios such as health check-ups, chronic disease screening, disease risk assessment, and health management, providing an objective and accurate quantitative tool for the early identification, risk stratification, precise intervention, and personalized health management of metabolic diseases. It has extremely high scientific value and broad application prospects.
[0045] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0046] The accompanying drawings, incorporated in and forming part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without inventive effort. In the drawings:
[0047] Figure 1 A flowchart illustrating the method for constructing and predicting the metabolic aging scores of multiple groups of biological ages in this application;
[0048] Figure 2 This is a flowchart of another exemplary method for constructing and predicting the metabolic aging score of multiple groups of biological ages in this application;
[0049] Figure 3 A flowchart illustrating the process of constructing an MDAS based on machine learning algorithms to select the optimal model;
[0050] Figure 4 This is a schematic diagram illustrating the recognition effect of MDAS on different metabolic health phenotypes.
[0051] Figure 5 This is a diagram illustrating the clinical relevance of MDAS.
[0052] Figure 6 A schematic diagram illustrating the association and prognostic value of MDAS with age-related diseases;
[0053] Figure 7 Kaplan-Meier estimate of the cumulative incidence of nine age-related diseases over 14 years, stratified by overall MDAS risk level;
[0054] Figure 8 This is a schematic diagram of the genome-wide association analysis results of MDAS;
[0055] Figure 9 This is a schematic diagram illustrating the causal effects of MDAS on cardiovascular disease.
[0056] Figure 10 A schematic diagram of the genome association and gene-environment interaction analysis results validated for the BCAMS cohort;
[0057] Figure 11 A schematic diagram of gene-environment interaction analysis results validated for the BCAMS cohort;
[0058] Figure 12A framework diagram for constructing a risk prediction system for metabolic aging scores of multiple groups of biological ages;
[0059] Figure 13 A schematic diagram of the structure of a computer system suitable for an embodiment of this application is shown. Detailed Implementation
[0060] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should be understood that certain features of the invention (described in the context of separate embodiments for clarity) may also be provided in a single embodiment. Conversely, multiple features of the invention (described in the context of a single embodiment for brevity) may also be provided separately or in any suitable combination or, where appropriate, in any other described embodiment of the invention. Certain features described in the context of various embodiments will not be considered essential features of those embodiments unless the embodiment is inoperable without those elements. The invention is further illustrated below by specific examples; however, it should be noted that the specific process conditions and results described in the embodiments of the invention are merely illustrative and should not be construed as limiting the scope of protection of the invention. All equivalent changes or modifications made in accordance with the spirit and essence of the invention should be covered within the scope of protection of the invention.
[0061] This application understands the context of metabolic-related aging based on the developmental origins of health and disease (DOHaD / Doha Theory). It posits that early life factors such as nutrition, growth, environmental exposure, and lifestyle can have long-term effects on metabolic programming, leading some individuals to exhibit metabolic abnormalities, obesity-related metabolic heterogeneity, and accelerated biological aging in adulthood or even adolescence. Therefore, assessing health risk solely based on weight or obesity status is insufficient. Further research is needed to identify aging factors driven by metabolic disorders based on different metabolic health phenotypes and their multi-omics aging characteristics, enabling early risk stratification and intervention.
[0062] Therefore, please refer to Figure 1 The flowchart for the construction and risk prediction method of multiple biological age metabolic aging scores in this application is detailed below:
[0063] Step S100: Obtain clinical phenotype data, proteomics data, and metabolomics data of the target object, and construct a multidimensional feature dataset.
[0064] Specifically, in combination Figure 2This application first establishes a cross-population research cohort system: using a large-scale longitudinal population cohort from the UK Biobank (UKB) as the model development set, it collects demographic data (age, sex, race, education level, smoking and drinking status), lifestyle information, anthropometric indicators (height, weight, body mass index BMI, waist circumference, blood pressure), clinical biochemical indicators (fasting blood glucose FBG, glycated hemoglobin HbA1c, triglycerides TG, high-density lipoprotein cholesterol HDL-C, etc.), proteomics data detected on the Olink platform, metabolomics data detected on the Nightingale platform, and whole-genome genotype data, and correlates them with long-term health outcome information with a median follow-up of 13.8 years; at the same time, it uses the Beijing Children and Adolescents Metabolic Syndrome Cohort (BCAMS) as an independent external validation set, and supplements it with data on targeted amino acid metabolomics profiles detected by liquid chromatography-tandem mass spectrometry (LC-MS / MS) and specific biomarkers such as adipokines and liver factors detected by enzyme-linked immunosorbent assay (ELISA) to verify the model's cross-age-group generalization ability. After obtaining the raw data, a unified data cleaning process is executed: missing value checking, outlier identification based on the 3-fold standard deviation method, Z-score standardization of continuous variables, and dummy variable encoding of categorical variables are performed in sequence to eliminate dimensional differences and measurement biases of data from different sources. Finally, a multidimensional feature dataset containing clinical phenotype, proteomics, metabolomics and genetic information is constructed.
[0065] Step S200: Based on the multidimensional feature dataset, construct biological age prediction models at the clinical, proteomic, and metabolomic levels respectively, and use the predicted biological age output by each model as biological age features.
[0066] Specifically, this application establishes three levels of prediction models based on the multidimensional feature dataset constructed in step S100: First, routine clinical indicators such as fasting blood glucose, blood pressure, and blood lipids are screened to construct a clinical-dimensional biological age model (preferably using the PhenoAge model framework); second, the expression levels of over a thousand plasma proteins detected by the Olink platform are used to construct a proteomic-dimensional biological age model (preferably using the Proteomic Aging Clock model framework); and finally, the concentrations of over a hundred metabolites detected by the Nightingale platform are used to construct a metabolomic-dimensional biological age model (preferably using the MetaboAgeMort model framework). During model training, the subject's actual chronological age is used as the regression label, and the model parameters are optimized by minimizing the deviation between the predicted age and the actual age, ultimately outputting three core biological age features: clinical predicted age, proteomic predicted age, and metabolomic predicted age. These features not only quantify the aging degree of various biological systems but also serve as key aging-related features, providing a high-value feature foundation for the input of subsequent multi-omics fusion models.
[0067] Step S300: The biological age characteristics are fused with the aging-related features in the multidimensional feature dataset. The preset metabolic health phenotype classification is used as a supervision signal to train a multi-omics fusion model to identify aging-related factors driven by metabolic disorders based on the differentiation of different metabolic types.
[0068] Specifically, this application's metabolic health phenotype classification includes normal-weight metabolic health, normal-weight metabolic abnormality, metabolically healthy obesity, and metabolically abnormal obesity. Furthermore, before training the multi-omics fusion model, a sample equalization step is included: the training samples are processed using synthetic minority class oversampling technology, and the similarity between samples is measured using heterogeneous Euclidean overlap distance to balance the differences in sample numbers among different metabolic health phenotypes. In addition, this application's metabolic health phenotype classification not only distinguishes current metabolic states but also reflects differences in metabolic regulation formed under the combined effects of genetic background, early developmental programming, and acquired environment / lifestyle. Based on the Doha theory, different metabolic phenotypes can be considered as explicit groupings of metabolic-related aging trajectories; this application aims to identify the metabolic disorder drivers and accelerated aging characteristics hidden behind these phenotypes by distinguishing between normal-weight metabolic health, normal-weight metabolic abnormality, metabolically healthy obesity, and metabolically abnormal obesity, rather than simply classifying obesity phenotypes.
[0069] Based on the differentiation of different metabolic types, this application identifies aging factors driven by metabolic disorders and constructs a multi-omics fusion model and a metabolic disease aging score accordingly.
[0070] More specifically, before model construction, this application first systematically evaluates different biological age models based on a multi-level biological age assessment system encompassing clinical, proteomics, and metabolomics perspectives, using the coefficient of determination (R²) of each model to predict actual age. 2 As an evaluation metric, representative models with the best predictive performance (e.g., PhenoAge, Proteomic Aging Clock (PAC), and MetaboAgeMort) were selected from clinical biological age models, proteomic biological age models, and metabolomic biological age models, respectively, as the basis for subsequent multi-omics fusion models. Since a certain proportion of high-dimensional omics data contains missing values, a multivariate imputation by chained equations (MICE, e.g., using the missRanger software package) based on random forest prediction mean matching was used to impute missing values, thereby improving data integrity and reducing the impact of missing data on model training. After data preprocessing, all research subjects were randomly divided into training and test sets (approximately 70% training set and approximately 30% test set). Considering the imbalance in the number of samples from different metabolic health phenotypes, a sample equalization step is included before training the multi-omics fusion model: the synthetic minority oversampling technique (SMOTE) is used to process the training samples, in which samples of normal weight metabolic abnormality (NWMO) and normal weight metabolic health (NWMH) are oversampled, samples of metabolically healthy obesity (MHO) and metabolically abnormal obesity (MUO) are appropriately undersampled, and the heterogeneous Euclidean overlap distance (HEOM) is used to measure the similarity between samples to calculate the distance between samples, generating new synthetic samples, thereby balancing the difference in the number of samples between different metabolic health phenotypes and improving the balance between categories during model training. Subsequently, aging-related features from clinical, proteomics, and metabolomics sources (e.g., 172 items) and features output from the aforementioned biological age model were used as input variables for the model. Four metabolic health phenotypes—Normal Weight Metabolic Health (NWMH), Metabolic Health Obesity (MHO), Normal Weight Metabolic Disorder (NWMO), and Metabolic Disorder Obesity (MUO)—were used as supervised learning classification labels. Six supervised learning algorithms—Extreme Gradient Boosting (XGBoost), RandomForest, Linear Discriminant Analysis (LDA), Naive Bayes, Decision Tree, and K-Nearest Neighbor (KNN)—were used for model training. The overall classification accuracy, Kappa statistic, and area under the receiver operating characteristic curve (AUC) were used to comprehensively evaluate the classification performance of each model (e.g., overall classification accuracy, Kappa statistic, and area under the receiver operating characteristic curve (AUC)). Figure 3 As shown, where Figure 3 (A) Six supervised learning algorithms—XGBoost, Random Forest, LDA, Naive Bayes, Decision Tree, and KNN—were used to classify metabolic health phenotypes, and the classification accuracy of each model was compared to select the optimal model. Figure 3 (B) This section compares the predictive performance of different machine learning models for four metabolic phenotypes: NWMH, NWMO, MHO, and MUO. This evaluation assesses the ability of each model to identify different metabolic health states, providing a model foundation for the subsequent construction of the MDAS. Preferably, XGBoost, with the best classification performance, is used as the final modeling algorithm. The MDAS obtained after training and selection demonstrates superior AUC and score distribution discrimination (e.g., metabolically healthy vs. metabolically abnormal individuals, MHO vs. MUO, NWMH vs. NWMO) across different classification scenarios, compared to traditional biological age models (PhenoAge, PAC, and MetaboAgeMort). Figure 4 As shown, where Figure 4 (A) ROC curves were used to compare the ability of MDAS and traditional biological age models to identify different metabolic health phenotypes. Figure 4 (B) Compare the score distributions of MDAS and traditional biological age models across different metabolic health phenotypes to evaluate the ability of each model to distinguish between different metabolic states.
[0071] Step S400: Based on the importance contribution and effect direction of each input feature in the multi-omics fusion model, determine the feature weights, and generate a metabolic disease aging score by weighted summation. The metabolic disease aging score reflects the contribution of metabolic disorders to the factors driving aging.
[0072] Specifically, the importance contribution of this application is quantified by the feature gain value during model training, and the effect direction is quantified by the SHapley additive interpretation value. The weight coefficient of each input feature is determined by integrating the feature gain value and the additive interpretation value in order to identify features that make a significant contribution to metabolic disorders driving aging.
[0073] More specifically, this application utilizes the Gain value output by the XGBoost model to evaluate the contribution of each aging-related feature to the classification of metabolic health phenotypes. This contribution is quantified using the feature gain value (Gain) during model training. Simultaneously, it combines SHapley Additive exPlanations (SHAP) analysis to determine the direction of each feature's effect on the model's prediction results. This effect direction is quantified using the SHapley Additive Explanation Value (SHAP value). By integrating the Gain value and SHAP direction information, the final weight coefficient W of each aging-related feature is obtained. iFinally, based on the weights corresponding to each feature, the various aging-related features are evaluated. i A weighted summation is performed to construct the Metabolic Disease Senescence Score (MDAS), whose calculation formula satisfies the following formula 1:
[0074] (Formula 1),
[0075] In Formula 1, i = 1, 2, ..., n, for example, n is 172; This represents the i-th aging-related feature. This represents the weights calculated by combining XGBoost Gain values and SHAP direction information. The resulting MDAS is a continuous variable; a higher MDAS value indicates a higher degree of metabolic-related biological aging in the study subjects. This score can more accurately identify different metabolic health phenotypes (including NWMH, MHO, NWMO, and MUO) and can be used for subsequent cardiovascular metabolic disease risk assessment and long-term clinical outcome prediction.
[0076] More specifically, the MDAS described in this application can be understood as a continuous indicator reflecting the degree of biological aging driven by metabolic disorders. Based on the Doha theory and the longitudinal / cross-cohort observations in this application, metabolic programming abnormalities can gradually transform into chronic low-grade inflammation, lipid / amino acid metabolism abnormalities, altered protein homeostasis, and the accumulation of cardiovascular metabolic risks in early life or adulthood, ultimately manifesting as higher MDAS and a higher risk of subsequent adverse clinical outcomes.
[0077] Step S500: Based on the metabolic disease aging score, predict and stratify the risk of metabolic-related diseases and / or the risk of aging progression in the target subjects.
[0078] Specifically, after completing the MDAS construction, this invention uses an independent test set to validate the model performance and further evaluate its predictive ability for different metabolic health states and long-term clinical outcomes. First, the trained MDAS model is applied to the independent test set, the MDAS value for each study subject is calculated, and the distribution differences among four metabolic phenotypes—normal weight metabolic health (NWMH), metabolically healthy obesity (MHO), normal weight metabolic abnormality (NWMO), and metabolic abnormality obesity (MUO)—are compared (e.g., Figure 5 As shown in Figure A, MDAS scores showed significant differences across different metabolic types, including NWMH, MHO, NWMO, and MUO. Furthermore, the study participants were divided into three age groups: young adults (<45 years), middle-aged adults (45-60 years), and elderly adults (≥60 years) to evaluate the changing patterns of MDAS across different age groups (e.g., ...). Figure 5 As shown in B, MDAS exhibits age-related changes in youth, middle-aged, and elderly populations, and further compares its score changes under different numbers of chronic diseases (e.g., ...). Figure 5(as shown in C).
[0079] Furthermore, this invention utilizes long-term follow-up data from the UKB (Universal Kidney Health Board) to evaluate the clinical predictive value of MDAS. Using a cohort with a median follow-up of 13.8 years as the study subjects, a Cox proportional hazards regression model was employed to analyze the association between MDAS and the risk of all-cause mortality and age-related chronic diseases. These age-related diseases include hypertension, type 2 diabetes, heart failure, myocardial infarction, stroke, chronic obstructive pulmonary disease, dementia, chronic kidney disease, and malignant tumors. After adjusting for confounding factors such as age, sex, race, education level, smoking, and alcohol consumption, the model calculated the hazard ratio (HR) and 95% confidence interval (95% CI) to evaluate the disease risk corresponding to different MDAS levels (e.g., ...). Figure 6 A forest plot illustrates the association between MDAS and various age-related health outcomes. Figure 6 The bar chart (B-bar) shows the statistical significance and effect size of the association between MDAS and various clinical outcomes. To further evaluate the application value of MDAS in early disease identification, this invention classifies diseases into early-onset and late-onset types based on the age of diagnosis, and analyzes the predictive ability of MDAS for disease risk at different stages of onset (e.g., Figure 6 As shown in C, MDAS has different predictive value for early-onset and late-onset diseases. Simultaneously, based on the overall distribution of MDAS and the distribution within different metabolic phenotypes, overall risk stratification and stratified risk assessment systems were established. Study subjects were stratified into high, medium, and low risk groups, and Kaplan-Meier survival analysis and Log-rank tests were used to compare the cumulative mortality rate and disease incidence rate (e.g., ...) between different risk levels. Figure 7 As shown, the Kaplan-Meier estimates of the cumulative incidence rates of nine age-related diseases over 14 years, stratified by overall MDAS risk levels, are presented. These diseases include chronic obstructive pulmonary disease, chronic kidney disease, cancer, type 2 diabetes, stroke, myocardial infarction, hypertension, heart failure, and dementia. Statistical differences in cumulative incidence rates among the three risk levels were assessed using the Log-rank test (p-values for all comparisons were <0.0001), thus achieving long-term risk prediction and accurate risk stratification based on MDAS.
[0080] This application uses the metabolic disease aging score as a continuous variable and assesses its dose-response relationship with the risk of all-cause mortality and various age-related chronic diseases through a survival analysis model. Based on the distribution characteristics of the metabolic disease aging score and the distribution differences within different metabolic health phenotypes, an overall risk stratification system and a stratified risk assessment system are established to classify the target subjects into high, medium and low risk levels.
[0081] Furthermore, the method of this application also includes a genetic attribution step: performing genome-wide association analysis using metabolic disease aging score as a phenotype to identify significantly associated genetic loci; constructing a weighted genetic risk score based on significantly associated genetic loci to quantify the explanatory power of genetic factors on metabolic disease aging score and to help identify the genetic basis associated with metabolic disorders driving aging.
[0082] Specifically, this application uses the metabolic disease aging score as the phenotype for genome-wide association analysis (GWAS). A generalized linear model (GLM) is employed to analyze the association between genetic variations and MDAS, adjusting for confounding factors such as age, sex, race, smoking, alcohol consumption, genotyping batch, and principal components of population structure. Quality control of genetic variations is performed, including minor allele frequency (MAF) ≥ 0.01 and Hardy–Weinberg equilibrium test. and the genotyping detection rate is ≥95%, and with As a genome-wide significance threshold (e.g.) Figure 8 A Manhattan map and Figure 8 The B-quantile-quantile plot shows the results of the genome-wide association analysis (GWAS) and the assessment of genome expansion in MDAS. Subsequently, functional annotation and candidate gene analysis were performed on the identified significant associated loci. LocusZoom was used to construct the associated region map, and the FUMA platform was employed to identify independent lead SNPs, locate candidate genes, and conduct tissue expression enrichment and functional pathway analysis. A weighted genetic risk score (wGRS) was constructed based on the significant associated genetic loci to quantify the explanatory power of genetic factors on metabolic disease aging scores. The correlation between wGRS and clinical indicators, proteomic and metabolomic characteristics was analyzed. Simultaneously, the differences in lead SNP allele frequencies among different metabolic health phenotypes were compared to evaluate the association between genetic variation and metabolic health status.
[0083] Furthermore, the method of this application also includes a causal inference step: using independent genetic variations obtained from genome-wide association analysis as instrumental variables, and employing two-sample Mendelian randomization analysis and colocalization analysis, the causal association between metabolic disease aging scores and metabolic-related diseases and / or aging-related outcomes is verified.
[0084] Specifically, this application uses independent genetic variants obtained from genome-wide association analysis (MDAS) as instrumental variables, employs two-sample Mendelian randomization (MR) analysis and colocalization analysis to verify the causal association between metabolic disease aging score and metabolic-related diseases and / or aging-related outcomes, evaluate whether MDAS and diseases are driven by common causal variants, and verify the biological rationale of MDAS from a genetic perspective (e.g., Figure 9As shown, the Mendelian randomization analysis results of MDAS for cardiovascular diseases and the co-localization analysis results with various cardiovascular diseases / outcomes show a high posterior probability PP.H4, indicating the existence of shared causal variants mainly mapped to the rs964184 locus.
[0085] Furthermore, the method of this application also includes cross-cohort validation and intervention steps: in independent cohorts of children and adolescents, the consistency of significant associated genetic loci and / or metabolic disease aging score-related characteristics is validated; the correlation between significant associated genetic loci and / or metabolic disease aging score-related characteristics and serum amino acid metabolites, adipokines, and liver factors is analyzed; gene-environment interactions are analyzed, and the protective effect of lifestyle interventions on individuals carrying high-risk genetic traits or metabolic aging risk characteristics is assessed.
[0086] Specifically, this application validated the stability of significantly associated genetic loci in an independent child and adolescent cohort (BCAMS), genotyped the child subjects, and compared the allele frequencies of lead SNPs (such as rs964184, rs174548) in four metabolic phenotypes: NWMH, MHO, NWMO, and MUO. Logistic regression analysis was used to evaluate the impact of genetic variations on the outcomes of different metabolic health states (e.g., Figure 10 As shown in A-10C, the results of genome association and gene-environment interaction analysis validated by the BCAMS cohort reveal the risk allele frequencies and OR values of specific gene variants in different metabolic types. Further analysis was conducted on multidimensional biomarkers such as serum amino acid metabolites, adipokines, and liver factors in children, and the correlation between lead SNPs and these biomarkers was analyzed to validate the impact of MDAS-related genetic loci on metabolic health from childhood. In addition, gene-environment interactions were analyzed to assess the protective effect of lifestyle interventions on individuals carrying high-risk genetic loci. Using the significantly associated locus rs964184 as a representative, the interaction between rs964184 and physical activity levels was analyzed, and the effects of different exercise frequencies on key metabolic indicators such as triglycerides were evaluated. After adjusting for confounding factors such as age, sex, BMI, diet score, sleep duration, birth weight, and urban / rural factors, a multivariate regression model (e.g., Figure 11 As shown, the gene-environment interaction analysis results of the BCAMS cohort reveal a hierarchical analysis of the effect of rs964184-G on triglyceride levels under different physical activity frequencies, as well as a gene-environment interaction framework diagram. This indicates that exercising ≥3 times per week and exercising daily can significantly reduce the genetic risk effect of this allele on elevated triglyceride levels by 36.5% and 52.4%, respectively, thus verifying that a healthy lifestyle can reduce the risk of metabolic aging in genetically susceptible individuals.
[0087] Therefore, metabolic-related aging risks do not only develop in adulthood, but are traceable to metabolic programming, genetic susceptibility, and the interaction of environmental and lifestyle factors during childhood and adolescence. By validating the association between MDAS-related genetic loci and early serum amino acid metabolites, adipokines, and liver factors in the BCAMS cohort, and by assessing the protective effects of lifestyle factors such as physical activity on high-risk allele carriers, this application further demonstrates the feasibility of early identification and intervention of drivers of metabolic aging.
[0088] Please see Figure 12 A framework diagram for constructing a risk prediction system for metabolic aging scores of multiple groups of students based on biological age is shown, including:
[0089] The acquisition module 1201 is used to acquire clinical phenotypic data, proteomics data and metabolomics data of the target object, and construct a multidimensional feature dataset;
[0090] Module 1202 is used to construct biological age prediction models at the clinical, proteomic, and metabolomic levels based on a multidimensional feature dataset, and to use the predicted biological age output by each model as a biological age feature.
[0091] Training module 1203 is used to fuse biological age characteristics with aging-related features in a multidimensional feature dataset, and to train a multi-omics fusion model using a preset metabolic health phenotype classification as a supervision signal, so as to identify aging-related factors driven by metabolic disorders based on the distinction of different metabolic types.
[0092] The calculation module 1204 is used to determine the feature weights based on the importance contribution and effect direction of each input feature in the multi-omics fusion model, and generate a metabolic disease aging score by weighted summation, wherein the metabolic disease aging score reflects the contribution of factors driving aging by metabolic disorders.
[0093] The prediction and stratification module 1205 is used to predict and stratify the risk of metabolic-related diseases and / or the risk of aging progression in target subjects based on the metabolic disease aging score.
[0094] Specifically, the system proposed in this application can effectively integrate multidimensional biological information, improve the accuracy and reliability of metabolic health status assessment, and provide new technical approaches for the early identification, risk stratification and precise intervention of metabolic diseases. It has high scientific value and broad application prospects.
[0095] It should be noted that the multi-group biological age metabolic aging score construction and risk prediction system provided in the above embodiments and the multi-group biological age metabolic aging score construction and risk prediction method provided in the above embodiments belong to the same concept. The specific operation methods of each module and unit have been described in detail in the method embodiments and will not be repeated here. In practical applications, the multi-group biological age metabolic aging score construction and risk prediction system provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.
[0096] Embodiments of this application also provide a computer device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, enable the computer device to implement the multi-group biological age metabolic aging score construction and risk prediction method provided in the above embodiments.
[0097] Figure 13 A schematic diagram of the structure of a computer system suitable for an embodiment of this application is shown. It should be noted that... Figure 13 The computer system 1300 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0098] like Figure 13As shown, the computer system 1300 includes a central processing unit (CPU) 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage section 1308 into a random access memory (RAM) 1303, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in the RAM 1303. The CPU 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304. The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, mouse, etc.; an output section 1307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (local area network) card, modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. Drive 1310 is also connected to I / O interface 1305 as needed. Removable media 1311, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1310 as needed so that computer programs read from them can be installed into storage section 1308 as needed.
[0099] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer tool programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1309, and / or installed from removable medium 1311. When the computer program is executed by central processing unit (CPU) 1301, it performs various functions defined in the system of this application.
[0100] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, flash memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. Computer programs contained on computer-readable media can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0102] The units described in the embodiments of this application can be implemented by tools or by hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the unit itself.
[0103] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a computer's processor, causes the computer to perform the aforementioned method for constructing and predicting the metabolic aging scores of multiple biological ages. This computer-readable storage medium may be included in the computer device described in the above embodiments, or it may exist independently and not incorporated into the computer device.
[0104] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods for constructing and predicting multiple sets of biological age metabolic aging scores provided in the various embodiments above.
[0105] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A method for constructing and predicting the metabolic aging score of multiple biological ages, characterized in that... Includes the following steps: Step S100: Obtain clinical phenotype data, proteomics data and metabolomics data of the target object, and construct a multidimensional feature dataset; Step S200: Based on the multidimensional feature dataset, construct biological age prediction models at the clinical, proteomic, and metabolomic levels respectively, and use the predicted biological age output by each model as biological age features. Step S300: The biological age characteristics are fused with the aging-related features in the multidimensional feature dataset. A preset metabolic health phenotype classification is used as a supervision signal to train a multi-omics fusion model to identify aging-related factors driven by metabolic disorders based on the distinction of different metabolic types. Step S400: Based on the importance contribution and effect direction of each input feature in the multi-omics fusion model, determine the feature weights, and generate a metabolic disease aging score by weighted summation, wherein the metabolic disease aging score reflects the contribution of factors driving aging by metabolic disorders. Step S500: Based on the metabolic disease aging score, predict and stratify the risk of metabolic-related diseases and / or the risk of aging progression of the target subjects.
2. The method for constructing and predicting the metabolic aging score of multiple groups of biological ages according to claim 1, characterized in that, In step S300, the metabolic health phenotype classification includes normal weight metabolic health, normal weight metabolic abnormality, metabolic health obesity, and metabolic abnormality obesity. Furthermore, before training the multi-omics fusion model, a sample equalization step is also included: the training samples are processed using synthetic minority oversampling technology, and the similarity between samples is measured by heterogeneous Euclidean overlap distance to balance the differences in the number of samples between different metabolic health phenotypes.
3. The method for constructing and predicting the metabolic aging score of multiple groups of biological ages according to claim 1, characterized in that, In step S400, the importance contribution is quantified using the feature gain value during model training, and the effect direction is quantified using the SHapley additive interpretation value. The weight coefficients of each input feature are determined by integrating the feature gain value and the additive interpretation value to identify features that significantly contribute to metabolic disorders driving aging.
4. The method for constructing and predicting the metabolic aging score of multiple groups of biological ages according to claim 1, characterized in that, Step S500 further includes: using the metabolic disease aging score as a continuous variable, and using a survival analysis model to assess its dose-response relationship with the risk of all-cause mortality and multiple age-related chronic diseases; Based on the distribution characteristics of the metabolic disease aging score and the distribution differences within different metabolic health phenotypes, an overall risk stratification system and a stratified risk assessment system are established to classify the target population into high, medium and low risk levels.
5. The method for constructing and predicting the metabolic aging score of multiple groups of biological ages according to claim 1, characterized in that, The method also includes a genetic attribution step: Genome-wide association analysis was performed using the metabolic disease aging score as a phenotype to identify significantly associated genetic loci; A weighted genetic risk score was constructed based on the significant associated genetic loci to quantify the explanatory power of genetic factors on metabolic disease aging scores and to help identify the genetic basis associated with metabolic disorders driving aging.
6. The method for constructing and predicting the metabolic aging score of multiple groups of biological ages according to claim 5, characterized in that, The method also includes a causal inference step: Using independent genetic variations obtained from genome-wide association analysis as instrumental variables, two-sample Mendelian randomization analysis and colocalization analysis were employed to verify the causal association between the metabolic disease aging score and metabolic-related diseases and / or aging-related outcomes.
7. The method for constructing and predicting the metabolic aging score of multiple groups of biological ages according to claim 6, characterized in that, The method also includes cross-queue verification and intervention steps: In independent cohorts of children and adolescents, the consistency of the significant associated genetic loci and / or the metabolic disease aging score-related features was verified; Analyze the correlation between the significant associated genetic loci and / or the metabolic disease aging score-related features and serum amino acid metabolites, adipokines, and liver factors; Analyze gene-environment interactions and assess the protective effect of lifestyle interventions on individuals carrying high-risk genetic traits or metabolic aging risk characteristics.
8. A multi-group biological age metabolic aging score construction and risk prediction system, characterized in that, include: The acquisition module is used to acquire clinical phenotypic data, proteomics data, and metabolomics data of the target object and construct a multidimensional feature dataset. The construction module is used to construct biological age prediction models at the clinical, proteomic, and metabolomic levels based on the multidimensional feature dataset, and to use the predicted biological age output by each model as biological age features. The training module is used to fuse the biological age characteristics with the aging-related features in the multidimensional feature dataset, and use a preset metabolic health phenotype classification as a supervision signal to train a multi-omics fusion model to identify aging-related factors driven by metabolic disorders based on the distinction of different metabolic types. The calculation module is used to determine the feature weights based on the importance contribution and effect direction of each input feature in the multi-omics fusion model, and generate a metabolic disease aging score by weighted summation, wherein the metabolic disease aging score reflects the contribution of factors driving aging by metabolic disorders. The prediction and stratification module is used to predict and stratify the risk of metabolic-related diseases and / or the risk of aging progression of the target subjects based on the metabolic disease aging score.
9. A computer-readable storage medium, characterized in that, It stores computer-readable instructions, which, when executed by the computer's processor, cause the computer to perform the method for constructing and predicting the metabolic aging score of multiple biological ages according to any one of claims 1 to 7.
10. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the steps of the method for constructing and predicting the risk of multiple sets of biological age metabolic aging scores as described in any one of claims 1 to 7.