Personalized biological age construction method based on deep learning and causal inference
By building a multi-dimensional biological age evaluation index system through deep learning and causal inference and integrating multi-source data, we have solved the problem of low-cost and high-precision construction of biological age in primary health management, and achieved accurate prediction and personalized intervention of personalized biological age.
Patent Information
- Application Number
- CN202510534460.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-09-09
AI Technical Summary
Existing biological age construction methods are difficult to effectively integrate the low cost of clinical phenotypic biological age and the high accuracy of multi-omic biological age in the primary health management system, making it difficult to promote.
A method based on deep learning and causal inference was used to construct a multidimensional biological age evaluation indicator system. Multi-source data (clinical phenotype, DNA methylation and metabolomics data) were integrated, and a multimodal biological age model was constructed through deep neural networks. Causal inference was performed to optimize risk assessment, and a comprehensive evaluation was performed in combination with the hierarchical weighted TOPSIS method.
It has achieved the precise construction of personalized biological age in the primary health management system, improved the scalability and accuracy of biological age prediction, and supported personalized intervention and intelligent early warning in primary health management.
Smart Images

Figure CN120613012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of health assessment technology, and in particular to a method for constructing personalized biological age based on deep learning and causal inference. Background Art
[0002] Aging is a complex biological process driven by the interplay of multiple dysregulated cellular and biochemical processes, manifesting as a gradual decline in tissue and organ function. Biological age can be used to measure aging, but biomarkers must: 1) be correlated with age, with higher correlations increasing sensitivity; 2) be reproducible and non-invasive; 3) be applicable to both humans and laboratory animals; and 4) monitor the underlying aging process, rather than the effects of disease.
[0003] Currently, the screening of aging biomarkers in research is typically based on correlation with chronological age, disease incidence, or mortality. A variety of biomarkers are used to estimate biological age, ranging from obvious phenotypic characteristics such as skin manifestations to molecular changes such as telomere length. The recent development of high-throughput sequencing technologies has enabled the characterization and quantification of a large number of epigenetic marks, transcriptomes, proteins, and metabolites, revealing how complex organisms undergo global changes with age at the molecular level.
[0004] Depending on the different biomarkers used in biological age, biological age can be divided into clinical phenotypic group age, molecular biological age, multi-omics age, etc. The biomarkers of clinical phenotypic group age mainly include physical characteristics and functional abilities. Frailty phenotype and frailty index are the most widely used frailty assessment tools in community populations and can be used as predictors of health and longevity. In addition to these functional evaluation indicators, biological age constructed based on anthropometric indicators and blood biochemical indicators is also associated with age-related diseases or death, such as phenotypic age and EMR age. The high accessibility, standardized protocols and affordable prices make clinical indicator phenotypes such as blood biochemistry and physical examinations an effective source of grassroots aging biomarkers.
[0005] In addition to clinical phenotypic age, a growing number of molecular age models have been constructed in recent years to investigate the molecular and cellular mechanisms of aging. Telomeres are DNA-protein complexes located at the ends of chromosomes, and telomere length is a biological determinant of age-related diseases. Epigenetic changes are also a hallmark of aging, with cellular and functional recovery achieved through transient epigenetic reprogramming. In 2013, the Horvath clock and Hannum clock, based on DNA methylation data, were proposed, clarifying and popularizing epigenetic clocks. Levine et al. developed DNAmPhenoAge by regressing phenotype age and DNA methylation levels. DNAmGrimAge overlaps with the Horvath and Hannum clocks at 41 and 6 CpG sites, respectively. Subsequently, DNAmGrimAge, constructed through a two-stage modeling approach, outperformed the Horvath, Hannum, and DNAmPhenoAge in predicting mortality and coronary heart disease risk. Epigenetic clocks can provide a more accurate representation of cellular aging (CA) and have the potential to reflect individual health status, disease, and even mortality. In addition, metabolites are the end products of cellular metabolism and can provide a more complete picture of biological processes and stronger phenotypic characterization than other "genomic maps", and have strong predictive value in the development of various chronic diseases.
[0006] As can be seen from the above, different omics data types provide unique characteristic windows to quantify the aging process. Integrating multiple omics information to optimize the construction of multi-omics clocks will be more clinically valuable. OMICmAge is the first multi-omics aging biomarker based on DNA methylation that uses electronic medical data, and its predictive performance is better than that of current biomarkers. However, the current data for constructing biological age basically comes from large-scale cohort surveys or medical data. Most survey items do not exist in primary health record data, making it difficult to promote existing aging biomarkers to primary health management systems. How to fully integrate the low cost of clinical phenotypic biological age and the high accuracy of multi-omics biological age remains an urgent problem to be solved. Summary of the Invention
[0007] The purpose of the present invention is to provide a personalized biological age construction method based on deep learning and causal inference. This method predicts personalized biological age by integrating deep learning and causal inference, and is suitable for grassroots health management systems.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A method for constructing personalized biological age based on deep learning and causal inference, comprising the following steps:
[0010] Step 1: Construct a multi-dimensional biological age evaluation indicator system;
[0011] Step 2: Data collection and preprocessing;
[0012] First, multi-source data collection; multi-source data includes clinical phenotype data, DNA methylation data, and metabolomics data; then, Z-score is used for standardization, and finally, nonlinear dimensionality reduction is performed;
[0013] Step 3: Screening of aging biomarkers;
[0014] First, we conduct a feature importance assessment and calculate the contribution of each marker to aging prediction based on a random forest model. Then, we construct a two-dimensional map based on economic cost analysis to screen for high-importance, low-cost markers.
[0015] Step 4: Construct a multi-level and multi-modal biological age model;
[0016] A biological age model is constructed using a deep neural network (DNN), comprising an input layer, a hidden layer, and an output layer. The biological age model is first trained using screened clinical phenotypic markers, and then, using an intermediate fusion framework for multimodal data fusion, DNA methylation data and metabolomics data are simultaneously integrated into the biological age model to construct a multimodal fusion biological age model.
[0017] Step 5: Causal inference optimizes risk assessment;
[0018] First, based on previous literature and expert consultation, we developed DAGs for accelerated aging and various adverse health outcomes to identify confounding variables. We used a fixed-effects model to generate dummy variables for each group, assigning each group a specific intercept and calculating multilevel propensity scores. We also used inverse probability weighting to assess accelerated aging.
[0019] Step 6: Use the hierarchical weighted TOPSIS method to conduct a multi-dimensional comprehensive evaluation with the help of a multi-dimensional biological age evaluation index system and a multi-level multi-modal biological age model.
[0020] Furthermore, the construction of a multi-dimensional biological age evaluation indicator system in step 1 includes:
[0021] First, qualitatively integrate indicators and preliminarily screen indicators through literature-based evidence-based methods, covering key indicators of the biological age evaluation index system in multiple dimensions, such as generalizability, accuracy, aging acceleration, and predictive ability;
[0022] Then, in the first round of expert consultation, the preliminary screening indicators were scored on a five-level scale of importance, feasibility, and sensitivity to screen out key indicators. The Delphi method was used to calculate the rank sum, full score ratio, arithmetic mean, and coefficient of variation of each indicator score.
[0023] Finally, the second round of expert consultation and weight determination was conducted; the second round of expert consultation compared the importance of key indicators, constructed a judgment matrix, and used the hierarchical analysis method to calculate the weights of indicators at all levels to form a final systematic, complete and comprehensive biological age evaluation indicator system.
[0024] Furthermore, the aging biomarker screening in step 3 includes:
[0025] The contribution of each marker to the aging phenotype was assessed using the permutation feature importance method. The aging prediction model was trained using a random forest model, and the mean square error was calculated. Bootstrap resampling was used to calculate the 95% confidence interval of the importance score.
[0026] Combining the evaluation of biomarker importance and economic cost, a two-dimensional map of aging biomarker screening is formed to show the distribution of each biomarker in terms of characteristic importance and economic cost; key biomarkers with high importance and low economic cost are identified.
[0027] Furthermore, during training, the biological age model calculates the gradient of the mean absolute error (MAE) loss function through the back-propagation algorithm; the gradient optimization adopts the Adam algorithm, integrates the momentum term with the adaptive learning rate mechanism, introduces L2 regularization to constrain the weight amplitude, and combines the Dropout layer to randomly shield neurons; the Bayesian search strategy is used for hyperparameter optimization, and the objective function is the MAE mean of 5-fold cross-validation, and the combination of learning rate, number of hidden layers, and number of nodes is explored in the continuous space.
[0028] This paper, balancing the needs of healthy aging with the actual situation of aging work at the grassroots level, has developed a personalized aging recognition technology that can be promoted at the grassroots level; and constructed a systematic, complete, and comprehensive biological age evaluation index system. Next, based on deep learning, a multi-level, multimodal biological age that can be promoted under multimodal, high-dimensional data is constructed, including clinical phenotypic biological age (RHRAge), DNA methylation biological age (DNAmRHRAge), and multi-omics biological age (OMICmRHRAge). Causal inference methods such as multi-level propensity scores are also introduced to fully consider high-dimensional confounding variables and population heterogeneity, optimizing the risk assessment of aging and adverse health outcomes. Finally, based on hierarchical weighted TOPSIS, weights are assigned to different dimensions of the multi-level, multimodal biological age, achieving scoring and ranking across multiple dimensions including scalability, accuracy, aging acceleration, and predictive ability, forming targeted biological age recommendations and providing a reference basis for grassroots health management services. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of the process of the present invention.
[0030] Figure 2Screening two-dimensional profiles for aging biomarkers.
[0031] Figure 3 Schematic diagram of multi-level and multi-modal biological age. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0033] like Figure 1 As shown, the method for constructing personalized biological age based on deep learning and causal inference disclosed in this embodiment includes the following steps:
[0034] Step 1: Construct a multi-dimensional biological age evaluation indicator system;
[0035] First, indicators were qualitatively integrated and preliminarily screened through literature-based evidence-based methods, covering key indicators of the biological age evaluation index system in multiple dimensions, including generalizability, accuracy, aging acceleration, and predictive ability.
[0036] Then, in the first round of expert consultation, the preliminary screening indicators were rated on a five-level scale of importance, feasibility, and sensitivity to screen out key indicators. The Delphi method was used to evaluate the scoring results of each indicator, and the grade, full score ratio, arithmetic mean, and coefficient of variation of each indicator score were calculated.
[0037] Finally, a second round of expert consultation and weighting was conducted. This second round of expert consultation compared the importance of key indicators, constructed a judgment matrix, and used the analytic hierarchy process to calculate the weights of each level of indicators, ultimately forming a systematic, complete, and comprehensive biological age evaluation indicator system. This step, by assigning weights to the multi-dimensional evaluation indicators, will establish a quantitative evaluation indicator system for biological age.
[0038] Step 2: Data collection and preprocessing;
[0039] First, multi-source data are collected; multi-source data mainly include clinical phenotype data (as shown in Table 1), DNA methylation data and metabolomics data.
[0040] Table 1 Age composition indicators of major clinical phenotype groups
[0041]
[0042]
[0043] The “Specific component indicators” column in bold indicates variables that are not included in the health records of community elderly people.
[0044] Then, the data was standardized; Z-score was used for standardization to eliminate the difference in feature dimensions and enhance comparability; after standardization, the data obeyed a normal distribution with a mean of 0 and a variance of 1.
[0045] Finally, nonlinear dimensionality reduction is performed using t-SNE to address the "crowding problem" of high-dimensional data and reduce redundant features. Conditional probabilities between high-dimensional samples are calculated based on a Gaussian distribution, similarity is defined in low-dimensional space using the t-distribution, and the mapping from high-dimensional to low-dimensional data is optimized by minimizing the KL divergence. The t-SNE algorithm removes redundant features, addresses multicollinearity, and improves the usability of subsequent algorithms.
[0046] Step 3: Screening of aging biomarkers;
[0047] First, feature importance was assessed, and the contribution of each marker to aging prediction was calculated using a random forest model. Then, combined with economic cost analysis, a two-dimensional (importance-cost) map was constructed to screen for high-importance, low-cost markers.
[0048] The contribution of each marker to the aging phenotype was evaluated using the permutation feature importance method; the aging prediction model was trained using the random forest model and the mean square error was calculated; for example, for the jth feature X j Generate X by random permutation j ′, calculate the mean square error, and then calculate the importance score is the jth feature X j Mean square error, MSE original is the initial mean square error. To verify the reliability of the results, Bootstrap resampling was used to calculate the 95% confidence interval of the importance score.
[0049] Combining the biomarker importance evaluation and economic cost evaluation to form a two-dimensional map for aging biomarker screening, such as Figure 2 The figure shows the distribution of each biomarker in terms of characteristic importance and economic cost. Key biomarkers with both high importance and low acquisition cost are identified, thereby further extracting important and generalizable aging factors for constructing biological age.
[0050] Step 4: Construct a multi-level and multi-modal biological age model;
[0051] A deep neural network (DNN) was used to construct a biological age model consisting of an input layer, hidden layers, and an output layer. The model was first trained using selected clinical phenotypic markers. During training, the gradient of the mean absolute error (MAE) loss function was calculated using a backpropagation algorithm. The gradient was optimized using the Adam algorithm, integrating a momentum term with an adaptive learning rate mechanism. To prevent overfitting, L2 regularization was introduced to constrain the weight amplitude, and a dropout layer was used to randomly mask neurons to enhance generalization. Hyperparameter optimization was performed using a Bayesian search strategy, with the objective function being the mean MAE of 5-fold cross-validation. Combinations of learning rate, number of hidden layers, and number of nodes were explored in a continuous space. Finally, the contribution of each aging marker was quantified using Shapley values, and cross-validation was used to ensure the model's generalization ability.
[0052] Next, RHRAge and DNA methylation data were fused to train a DNN-based biological age model. Using an intermediate fusion framework for multimodal data fusion, DNA methylation and metabolomics data were simultaneously integrated into the DNN model to construct a multimodal fusion biological age model. This method not only preserves the unique information of each modality during the fusion process but also allows for interaction between different modalities, thereby improving the effectiveness of data fusion. The predictive performance of each biological age was evaluated by calculating the Pearson correlation coefficient, standard coefficient of determination, and mean absolute error.
[0053] Step 5: Causal inference optimizes risk assessment;
[0054] Based on the described biological age model, individual aging scores were calculated to assess aging rates and determine whether individuals experienced accelerated aging. Taking into account population heterogeneity, the causal effects of accelerated aging on various health outcomes were investigated. First, based on previous literature and expert consultation, DAG plots were constructed for accelerated aging and various adverse health outcomes to identify confounding variables. To account for population heterogeneity across cohorts, multilevel propensity scores were used. A fixed-effects model was used to generate dummy variables for each cohort, assigning each cohort a specific intercept to account for unobserved heterogeneity between cohorts.
[0055] After multilevel propensity scoring, inverse probability weighting is used to estimate accuracy; to avoid infinite variance that occurs when unstable weights are used, stable weights are selected, which can achieve better effect estimation performance. In this case, the numerator of the weight is the probability of the mean of the group-specific intervention variable.
[0056] However, when the data has a hierarchical structure, the problem of homogeneity within the complete layer will limit the use of methods such as traditional multi-level propensity scores. Especially when small clusters appear, the problem of homogeneity within the complete layer will cause traditional methods to face the dilemma of a large number of model failures, which will lead to the inability to obtain parameters of interest such as the average treatment effect (ATE). In view of this, this embodiment intends to adopt an improved cluster propensity score model. First, the unbiasedness of the improved cluster propensity score model for the estimated ATE is explored from a theoretical level. Furthermore, it is demonstrated that compared with the traditional method, the improvement strategy proposed in this study can effectively alleviate the extreme weight problem and reduce the interference of outliers on the overall estimation, thereby significantly improving the robustness and reliability of the estimation results. Secondly, in order to comprehensively and objectively evaluate the estimation performance of the improved model, different simulation scenarios are set. The estimation performance of the improved model is systematically evaluated from different aspects such as simulation convergence, estimation bias, extreme weight ratio, and computational efficiency.
[0057] Finally, we used proportional hazards regression models to estimate the causal effect of accelerated aging on the incidence of various diseases and mortality, and calculated cluster-robust standard errors as an approximation of the estimated variability of the intervention effect. This was then validated in an external validation cohort.
[0058] Step 6: Comprehensive evaluation and application;
[0059] Using the hierarchical weighted TOPSIS method, with the help of the multidimensional biological age evaluation index system and the multi-level multimodal biological age model evaluation verification results, a comprehensive evaluation of the constructed biological age model and classic biological age assessment indicators, such as frailty phenotype and DNAm PhenoAge, was conducted in terms of generalizability, accuracy, aging acceleration, and predictive ability.
[0060] Assume that the BA evaluation index system has a total of k-level indicators, and the minimum level evaluation index has a total of m indicators. The normalized weights are W = [W1, W2, ..., W m ]. At the same time, assume that there are n biological ages to be evaluated. Therefore, the original score matrix X can be expressed as: Negative data were transformed into positive data using the reciprocal method or the difference method to ensure the consistency of the comparison data. Due to the different dimensions of different indicators, the original score matrix was standardized to eliminate the influence of different dimensions and obtain the data matrix Z.
[0061]
[0062] According to the above normalized matrix Z, the positive ideal solution is defined as and negative ideal solutions in i=1,2,…,m.
[0063] Combined with the weight W of each indicator, calculate the weighted Euclidean distance from the indicator value of each biological age to be evaluated to the positive and negative ideal solutions and The larger the value, the farther the biological age to be evaluated is from the optimal solution or the worst solution. The ideal value is The smaller the value, the Larger values:
[0064]
[0065] Therefore, the pth biological age composite score is C p , between 0 and 1. Sort biological ages by the size of the comprehensive score, with larger values representing better biological ages.
[0066]
[0067] To achieve a categorized evaluation based on generalizability, accuracy, aging acceleration, and predictive ability, we will determine the comprehensive scores of each biological age across these different dimensions. This will then lead to a ranking of aging models based on their performance across these dimensions. Based on the sub-dimensional and overall comprehensive evaluation results, we will formulate recommendations for biological age, specifically those that are applicable to grassroots organizations. This method will also be used to assess the individual data collection needs of older adults and provide feedback guidance for accurate biological age prediction.
[0068] Therefore, this embodiment intends to combine clinical indicators such as blood biochemistry and physical examination, as well as omics data, to develop a set of personalized biological age technologies that can be promoted at the grassroots level. By associating or inputting individual data, it is possible to assist in biological age selection, guide data collection, evaluate aging conditions, and output personalized precision interventions based on the data input situation; through secondary development, it can quickly adapt to grassroots health management service scenarios and achieve multi-dimensional results transformation: for example, for community health service centers, hospital physical examination departments and other institutions, it can integrate dynamic management of health records, multimodal data analysis and intelligent early warning functions, and integrate IoT device data collection (such as blood pressure monitors, blood glucose meters, etc.) to achieve automatic collection and real-time updating of residents' health data, generate personalized aging intervention plans, and significantly improve grassroots service efficiency. For individuals, lightweight health management applications can be developed to support automated interpretation of physical examination reports, synchronization of smart device data, and personalized intervention recommendations, to achieve closed-loop management of residents' health monitoring-assessment-intervention.
[0069] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A personalized biological age construction method based on deep learning and causal inference, characterized by: The steps include: Step 1: Construct a multi-dimensional biological age evaluation indicator system; Step 2: Data collection and preprocessing; First, multi-source data collection; multi-source data includes clinical phenotype data, DNA methylation data, and metabolomics data; then, Z-score is used for standardization, and finally, nonlinear dimensionality reduction is performed; Step 3: Screening of aging biomarkers; First, we conduct a feature importance assessment and calculate the contribution of each marker to aging prediction based on a random forest model. Then, we construct a two-dimensional map based on economic cost analysis to screen for high-importance, low-cost markers. Step 4: Construct a multi-level and multi-modal biological age model; A biological age model is constructed using a deep neural network (DNN), comprising an input layer, a hidden layer, and an output layer. The biological age model is first trained using screened clinical phenotypic markers, and then, using an intermediate fusion framework for multimodal data fusion, DNA methylation data and metabolomics data are simultaneously integrated into the biological age model to construct a multimodal fusion biological age model. Step 5: Causal inference optimizes risk assessment; First, based on previous literature and expert consultation, DAG diagrams of accelerated aging and multiple adverse health outcomes will be drawn separately to identify confounding variables; A fixed-effects model was used to generate dummy variables for each group, assigning each group a specific intercept and calculating multilevel propensity scores. Inverse probability weighting was used to assess accelerated aging. Step 6: Use the hierarchical weighted TOPSIS method to conduct a multi-dimensional comprehensive evaluation with the help of a multi-dimensional biological age evaluation index system and a multi-level multi-modal biological age model.
2. The method for constructing personalized biological age based on deep learning and causal inference according to claim 1, characterized in that: The construction of a multi-dimensional biological age evaluation indicator system described in step 1 includes: First, qualitatively integrate indicators and preliminarily screen indicators through literature-based evidence-based methods, covering key indicators of the biological age evaluation index system in multiple dimensions, such as generalizability, accuracy, aging acceleration, and predictive ability; Then, in the first round of expert consultation, the preliminary screening indicators were scored on a five-level scale of importance, feasibility, and sensitivity to screen out key indicators. The Delphi method was used to calculate the rank sum, full score ratio, arithmetic mean, and coefficient of variation of each indicator score. Finally, the second round of expert consultation and weight determination was conducted; the second round of expert consultation compared the importance of key indicators, constructed a judgment matrix, and used the hierarchical analysis method to calculate the weights of indicators at all levels to form a final systematic, complete and comprehensive biological age evaluation indicator system.
3. The method for constructing personalized biological age based on deep learning and causal inference according to claim 1, characterized in that: The aging biomarker screening in step 3 includes: The contribution of each marker to the aging phenotype was assessed using the permutation feature importance method. The aging prediction model was trained using a random forest model, and the mean square error was calculated. Bootstrap resampling was used to calculate the 95% confidence interval of the importance score. Combining the evaluation of biomarker importance and economic cost, a two-dimensional map of aging biomarker screening is formed to show the distribution of each biomarker in terms of characteristic importance and economic cost; key biomarkers with high importance and low economic cost are identified.
4. The method for constructing personalized biological age based on deep learning and causal inference according to claim 1, characterized in that: During training, the biological age model calculates the gradient of the mean absolute error (MAE) loss function using a backpropagation algorithm. Gradient optimization uses the Adam algorithm, which integrates a momentum term with an adaptive learning rate mechanism, introduces L2 regularization to constrain weight amplitudes, and combines a Dropout layer to randomly mask neurons. Hyperparameter optimization is performed using a Bayesian search strategy, with the objective function being the MAE mean of 5-fold cross-validation. Combinations of learning rate, number of hidden layers, and number of nodes are explored in a continuous space.