Aging clock construction method based on causal machine joint model
Through the joint model of causal machine, a hybrid causal effect graph model is constructed, which solves the problem of insufficient integration capabilities of high-dimensional multimodal data in the existing technology, and realizes a high-precision and interpretable aging clock model, which is suitable for primary medical care and personalized health management.
Patent Information
- Application Number
- CN202510336852.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
The existing aging clock construction method has insufficient integration capabilities when processing high-dimensional multimodal data, making it difficult to fully capture key variables and complexities in the aging process, resulting in low prediction accuracy and low applicability.
Using a method based on the joint model of causal machine, feature screening and data processing are performed through multimodal data fusion, saddle point approximation method and proportional advantage logistic regression hybrid model, a mixed causal effect graph model is constructed to evaluate multidimensional causal relationships and improve prediction accuracy and explanatory.
It significantly improves the prediction accuracy and explanatory power of the aging clock model, is suitable for primary medical institutions, and supports personalized health management and early disease intervention.
Smart Images

Figure BDA0005321984650000071 
Figure BDA0005321984650000081 
Figure BDA0005321984650000091
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of health assessment, and particularly relates to a method for constructing an aging clock based on a causal machine joint model. Background Art
[0002] Aging is a complex biological process, involving the gradual decline of body functions and changes in multiple genetic and epigenetic mechanisms. In recent years, epigenetics has gradually become a research hotspot due to its important role in aging regulation. In particular, DNA methylation, this reversible epigenetic modification, can regulate the expression pattern of genes, and its level changes significantly with age. Therefore, the epigenetic clock based on DNA methylation data has become an important tool for predicting biological age (BA). The epigenetic clock establishes a mathematical model through the methylation status of specific CpG sites, which can more accurately reflect the physiological state and health status of an individual than chronological age, providing a new perspective for aging research and health management.
[0003] The integration of multi-omics data provides new opportunities for aging research. Multi-omics combines information such as epigenome, transcriptome, proteome, and metabolome, and can reveal the biological process of aging from multiple perspectives. By integrating these data, not only can aging markers be captured more comprehensively, but also the accuracy of the biological age prediction model can be improved. However, the high-dimensionality and complexity of the data also pose higher requirements for model construction.
[0004] Although DNA methylation shows significant potential in biological age detection, traditional methods for detecting human biological age still have some limitations; traditional methods are usually based on single-omics data, such as telomere length measurement or transcriptome analysis, and have problems of low accuracy and insufficient specificity. At the same time, these methods often ignore the complex interactions between multi-omics data and are difficult to comprehensively reveal the multi-dimensional characteristics of the aging process. The wide application of existing methods for constructing aging clocks still faces the problem of insufficient ability to integrate high-dimensional multi-modal data, making it difficult to comprehensively capture key variables in the aging process, the imbalance and complexity of data distribution, resulting in biases in traditional statistical methods and low applicability in primary medical scenarios, and it is difficult to achieve low-cost and high-efficiency popularization and application.
[0005] Causal Machine Learning (CausalML), as an emerging technology, can formalize the data generation process into Structural Causal Models (SCMs). By revealing the causal relationships between variables, it can solve the non-linear, multi-dimensional, and complex interaction problems that cannot be handled by traditional statistical methods. In the medical field, Causal Machine Learning combined with medical clinical data can more accurately analyze the causal relationships between variables through causal inference techniques, providing a scientific basis for health management and personalized medical decision-making.
[0006] In aging research, the introduction of Causal Machine Learning makes it possible to construct a high-performance multi-modal aging clock. By integrating multi-modal data through a causal graph model, it can simultaneously reveal the mechanism of action of aging biomarkers and the causal relationships of health outcomes. In addition, causal inference techniques can reduce model bias caused by confounding variables, thereby improving the accuracy and interpretability of biological age prediction. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for constructing an aging clock based on a causal machine joint model. By introducing Causal Machine Learning and multi-modal data fusion technology, a new type of aging clock model is developed; in feature screening, the saddle point approximation method and the proportional odds logistic regression mixed model (POLMM) are applied to effectively solve the problem of unbalanced data distribution; by analyzing the interaction of multi-dimensional variables through a causal graph model, the interpretability and prediction accuracy of the model are significantly improved; it not only has high accuracy and interpretability, but also can adapt to the data needs of primary medical institutions, providing scientific support for health management and early disease intervention.
[0008] To achieve the above purpose, the present invention adopts the following technical solutions:
[0009] A method for constructing an aging clock based on a causal machine joint model, comprising the following steps:
[0010] Step 1: Obtain a multi-modal data set of the target population based on multi-modal data fusion and standardized preprocessing; the multi-modal data set includes clinical phenotype data, DNA methylation data, transcriptome data, and metabolome data;
[0011] Step 2: Perform two-stage screening on the multi-modal data set to obtain core features, wherein the two-stage screening includes a screening method for removing low-correlation features based on the Pearson correlation coefficient, and a screening method for removing low-importance features based on recursive feature elimination combined with a machine learning model;
[0012] Step 3: Analyze the interactions among clinical phenotypes, DNA methylation, transcriptomics, and metabolomics based on the extended proportional odds logistic regression mixture model; use the Score test to analyze the association between each single genetic variant and the multi-class phenotype and generate P-values; then calibrate the P-values by the saddlepoint approximation method;
[0013] Among them, the expression of the extended proportional odds logistic regression mixture model is: Y k = α0 + θG k + γ T X k + β T Z k + e k + b k , where G k represents the genotype variable, X k represents the nutrigenomics variable, Z k represents the metabolomics variable, e k represents the interaction between omics, b k represents the description of genetic correlation, θ, γ, β are model parameters, and α0 is the intercept term;
[0014] Step 4: Construct an aging clock based on the mixed causal effect graph model to evaluate multi-dimensional causal relationships, evaluate the biological age of an individual, and output a personalized health assessment report and aging acceleration warning information;
[0015] Constructing an aging clock based on the mixed causal effect graph model includes: First, extract the causal relationships related to aging from medical literature and expert knowledge bases, and use fuzzy theory to process the causal relationships to output a fuzzy causal relationship matrix;
[0016] Then, estimate the strength of the causal relationships based on the Bayesian network and the fuzzy causal relationship matrix; use the partitioned MCMC method to explore and optimize the Bayesian network;
[0017] Finally, use the degree and rate of aging as the outcome indicators for risk assessment, and select the optimal parameters of the Bayesian network through grid search and 5-fold cross-validation to obtain the aging clock.
[0018] Furthermore, the standardized preprocessing refers to: processing the data of continuous variables using the Z-score normalization method; processing the data of categorical variables using one-hot encoding; processing the missing value data using the K-nearest neighbor interpolation or multiple imputation method; processing the systematic missing data using the method based on Bayesian inference.
[0019] Furthermore, in Step 2, the screening method of removing low-correlation features based on the Pearson correlation coefficient means that when the Pearson correlation coefficient |r| between the features in the multi-modal dataset and the aging clock is less than 0.1, the feature is a low-correlation feature and should be removed.
[0020] Further, the screening method for eliminating low-importance features based on recursive feature elimination combined with a machine learning model means: selecting a machine learning model, training it with a dataset from which low-correlation features have been eliminated, calculating the importance scores of the features in the dataset, eliminating the features with importance scores lower than the threshold, and repeating the training until the preset number of features is reached or the model performance no longer improves significantly;
[0021] Among them, the machine learning model is a random forest or LASSO regression; the random forest calculates the importance scores based on Gini importance; LASSO regression calculates the importance scores using the L1 regularization coefficient.
[0022] Further, the parameters θ, γ, and β of the extended proportional odds logistic regression mixture model are optimized and estimated for the extended proportional odds logistic regression mixture model using the penalized quasi-likelihood method and the average information restricted maximum likelihood algorithm.
[0023] Further, calibrating the P-value by the saddlepoint approximation method means: using the cumulant generating function to estimate the probability density function or cumulative distribution function of the distribution; when the statistic T adj < 2, calculating the P-value based on the normal approximation; when the statistic T adj ≥ 2, using the saddlepoint approximation to calculate the P-value.
[0024] Further, the biological age of an individual is estimated and obtained through multi-omics variables and regression models.
[0025] Further, the accelerated aging warning information is obtained by calculating the individual's aging risk index using the hierarchical weighted TOPSIS method and then evaluating the individual's aging risk.
[0026] A lightweight aging clock system based on a causal machine joint model, the system includes one or more processors and a memory, the memory stores a program, and the program is executed by the processor to perform the aging clock construction method based on the causal machine joint model; the processor executes the program using the edge computing mode and model quantization technology.
[0027] Further, the processor is connected to the display of the visual health management platform.
[0028] The present invention has the following beneficial effects:
[0029] (1) Multi-modal data fusion improves accuracy. By integrating clinical phenotype data, DNA methylation data, transcriptome data, and metabolome data, it comprehensively reflects the individual aging process and significantly improves the prediction accuracy.
[0030] (2) Innovative algorithms optimize data processing. By introducing the saddle point approximation method and the POLMM model, the problem of unbalanced data distribution is solved, and the model stability and feature screening effect are improved.
[0031] (3) Causal inference enhances interpretability. Based on the hybrid causal effect graph model, the complex causal relationships between variables are revealed, providing a scientific basis for personalized health management.
[0032] (4) Adapt to primary healthcare and personalized applications. The model is designed with low cost and simple operation, generating health assessment reports and risk warnings, and is applicable to primary healthcare institutions and individual health management. Specific implementation manners
[0033] A method for constructing an aging clock based on a causal machine joint model provided in this embodiment includes the following steps:
[0034] Step 1: Obtain a multi-modal data set of the target population based on multi-modal data fusion and standardized preprocessing. Among them, the multi-modal data set includes clinical phenotype data, DNA methylation data, transcriptome data, and metabolome data.
[0035] Step 2: Perform two-stage screening on the multi-modal data set to obtain the core features of the multi-modal data set. Among them, the two-stage screening includes a screening method for removing low-correlation features based on the Pearson correlation coefficient, and a screening method for removing low-importance features based on recursive feature elimination combined with a machine learning model.
[0036] Step 3: Analyze the multi-omics interaction based on the extended proportional odds logistic regression mixture model, and then perform P-value calibration based on the fusion saddle point approximation method; solve the problem of unbalanced data distribution.
[0037] Step 4: Construct an aging clock based on the hybrid causal effect graph model to evaluate multi-dimensional causal relationships and calculate the biological age of an individual; then verify the performance of the aging clock in combination with the hierarchical weighted TOPSIS method; finally, output a personalized health assessment report and aging acceleration warning information.
[0038] In Step 1, collect the multi-modal data of the target population and construct a user clinical data table; the multi-modal data includes, but is not limited to, clinical phenotype data, DNA methylation data, transcriptome data, and metabolome data; perform standardization and missing value filling on the multi-modal data to generate a complete multi-modal data set.
[0039] The so-called standardized preprocessing refers to: for continuous variable data (such as age, blood glucose, cholesterol, etc.), the Z-score normalization method is used for standardization to eliminate the differences in feature dimensions and enhance comparability; for categorical variable data (such as gender, smoking status, etc.), one-hot encoding is used for processing to make it suitable for subsequent machine learning modeling; for data with missing values, the K-nearest neighbor interpolation (KNN) or multiple imputation method (MICE) is used to fill in the randomly missing data; for systematically missing data, a method based on Bayesian inference is used for reasonable filling to ensure the integrity of the data. As shown in Table 1, it is an example of data preprocessing.
[0040] Table 1 Dataset Information
[0041]
[0042]
[0043]
[0044] In step 2, two-stage feature screening is adopted to effectively improve the generalization ability of the aging prediction model, while optimizing the computational efficiency to ensure that the model can accurately extract key variables in the high-dimensional feature space.
[0045] First, based on the Pearson correlation coefficient, low-correlation features are removed for preliminary screening; for any given feature X i and the target variable (aging clock) Y i , calculate their Pearson correlation coefficient r,
[0046]
[0047] In the formula, and respectively represent the means of X and Y, and n is the number of samples.
[0048] When |r| < 0.1, the correlation between feature X i and the target variable Y i is weak, then it is considered that the contribution of this feature to the prediction result is small and should be removed; thereby reducing the data dimension, excluding features that are irrelevant or have small contributions to the target variable, thus reducing the computational complexity of the model, while reducing noise interference and improving the robustness of the model.
[0049] Then, deep screening is carried out using Recursive Feature Elimination (RFE); RFE is a feature selection method based on model weights, which uses a machine learning model (such as a random forest or LASSO regression) to evaluate the importance of features and recursively removes features with lower contributions. For example: in an initial dataset of 500 subjects, there are a total of 80 features. First, 30 features with a Pearson correlation below 0.1 are screened out, leaving 50 features. Subsequently, recursive feature elimination (RFE) is combined with a random forest model, and finally 20 core features are retained.
[0050] Specifically, the steps of the RFE method: Select a machine learning model (such as a random forest or LASSO) and train it on the current feature set; calculate the importance scores of each feature. If LASSO is used, the L1 regularization coefficient is used; if a random forest is used, the feature weights based on Gini importance are used; finally, the feature with the lowest weight is removed, and training is repeated on the remaining features; until the preset number of features is reached or the model performance no longer improves significantly.
[0051] During the RFE process, if LASSO regression is used as the base model, its objective function is: In the formula, λ is the regularization hyperparameter, and β j is the regression coefficient corresponding to feature j. When λ is appropriately adjusted, the regression coefficients can be effectively sparsified, making the coefficients of unimportant features tend to zero, thus automatically completing feature selection.
[0052] If a random forest is used, the feature weights are calculated based on Gini importance:
[0053]
[0054] where T represents all nodes of the decision tree, and P t is the proportion of samples at node t, and p tk is the probability of class k at node t.
[0055] The importance of a feature is determined by its contribution to the Gini index. A higher Gini importance means that the feature plays a more crucial role in the tree splitting process. For example, the Gini importance of smoking is 0.64 > the importance of a history of hypertension is 0.60, so it can be considered that the feature importance of smoking is greater than that of a history of hypertension.
[0056] The preliminary screening process effectively reduces the data dimension, eliminates irrelevant variables with low correlation, and makes the subsequent modeling more efficient; RFE combines random forest or LASSO to further optimize feature selection, ensuring that the finally selected feature set not only has good predictive ability but also can minimize redundancy to the greatest extent, improving the interpretability and generalization performance of the model. Therefore, this method is not only applicable to the aging prediction model but can also be extended to other machine learning modeling tasks with high-dimensional features, providing an efficient and robust solution for complex biological data analysis.
[0057] In step 3, the POLMM model is Y k = α0 + θG k + γ T X k + β T Z k + e k + b k , where G k represents the genotype variable, X k represents the nutrigenomics variable, Z k represents the metabolomics variable, e k characterizes the interaction between omics, b k describes the genetic correlation, and θ, γ, β are model parameters; α0 is the intercept term, representing the baseline aging level.
[0058] To optimize the parameter estimation of the POLMM model, this embodiment uses the penalized quasi-likelihood method and the average information restricted maximum likelihood algorithm to optimize the POLMM model.
[0059] In the POLMM framework, assume that the phenotypic variable Y k of an individual is affected by multiple omics data, and at the same time, considering the genetic correlation between individuals, a penalty term P(β) is defined to control the complexity of the model;
[0060]
[0061] In the formula, P(β) can use L1 or L2 penalty to improve the sparsity or stability of the model. When using the average information restricted maximum likelihood method for parameter estimation, iterative optimization is performed based on the Fisher information matrix;
[0062]
[0063] In the formula, W represents the weight matrix. Considering the dependence relationship between different omics levels, W is further decomposed into: Through this optimization framework, it is possible to better control the weights of different omics features and improve the predictive ability of the model.
[0064] Test the association between each single genetic variant and the multi-class phenotype, and use the Score test to determine the association between each single genetic variant and the multi-class phenotype. Since the conventional Score test assumes that the test statistic asymptotically follows a normal distribution and only uses the first two moments. However, when the sample size distribution among different classes is highly imbalanced, the potential distribution of the test statistic may be significantly different from the normal distribution.
[0065] To calculate the P-value more accurately, the saddlepoint approximation method is used to characterize the null hypothesis distribution when the phenotypic data distribution is imbalanced. The saddlepoint approximation method uses the cumulant generating function (CGP) to estimate the probability density function or cumulative distribution function of the distribution, and its statistic
[0066] When T adj < 2, then calculate the P-value based on the normal approximation; if T adj ≥ 2, then use the saddlepoint approximation to calculate the P-value.
[0067] To further improve the application ability of the POLMM model in constructing the aging clock, this embodiment introduces a biological age calculation method; the biological age BA is deduced from multi-omics variables and estimated using a regression model;
[0068]
[0069] Among them, X i represents biomarkers at different omics levels (such as DNA methylation level, telomere length, inflammatory factors, etc.), w i represents the weight of each variable, and ∈ is the error term.
[0070] This model can optimize the parameters through the POLMM model to ensure the accuracy of predicting the biological age.
[0071] In the example of predicting biological age by DNA methylation, select the key CpG site M j for Lasso regression and establish a biological age calculation formula: Among them, λ j represents the regression coefficient.
[0072] To further optimize the calculation, this embodiment combines the TOPSIS method to perform weighted adjustment on the biological age to make the individual aging assessment more accurate.
[0073] For the case where telomere length affects aging, if the telomere length (TL) is used as a biomarker, construct an aging clock model:
[0074] BA TL = α + βTL + γH + ∈;
[0075] Among them, TL represents telomere length, H represents health score, and α, β, and γ are regression coefficients. If an individual has a telomere length of 8.5 kb and a health score of 12, the biological age is calculated using the model: BA TL = 80 - (3×8.5) - (0.7×12) = 80 - 25.5 - 8.4 = 46.1; This method combines multi-omics data, improves the prediction accuracy and interpretability of the model, and makes the aging clock more personalized and scientifically based.
[0076] Based on the traditional POLMM modeling, the method of fusing saddle point approximation (SPA) and multi-view learning enables the POLMM model to show higher robustness and prediction ability in small-sample and multi-modal data; Through the comprehensive analysis of multi-omics data, potential biological mechanisms can be effectively mined, providing new methodological support for precision medicine research.
[0077] Suppose in an elderly cohort study, the impact of genetic variation of a certain genotype (rs123456) on the degree of aging is evaluated. The analysis results of the POLMM model are as follows: rs123456 genotype: AA vs. AG vs. GG; IL-6 level: 3.1 pg / mL, 4.2 pg / mL, 5.8 pg / mL; Biological age (predicted value): 60.5 years old, 62.8 years old, 65.2 years old. The results show that the impact of this genotype on biological age is significant (p = 0.002), and the IL-6 level shows a gradient upward trend among different genotypes.
[0078] In addition, in terms of computational efficiency, this embodiment uses the stochastic gradient descent method (SGD) to train the model and combines the Adam optimization algorithm to improve the convergence speed. For high-dimensional data, in order to further improve computational efficiency, sparse matrix operations are used to accelerate the processing of large-scale data, and at the same time, GPU parallel computing methods are used to improve the computational efficiency and scalability of the algorithm.
[0079] In step 4, based on the mixed causal effect graph (CausalEffectGraph, CEG), the scientificity and interpretability of aging assessment are ensured. This embodiment integrates fuzzy theory, Bayesian network, MCMC optimization, biological age modeling, and multi-model integrated analysis to form a complete aging assessment system.
[0080] In the process of extracting and modeling causal relationships, it is first necessary to extract aging-related causal relationships from medical literature and expert knowledge bases. Since the literature data and expert knowledge may be uncertain, fuzzy theory is introduced to process these causal relationships to ensure the reliability of causal inference. Fuzzy causal relationship matrix C = [c ij , c ij represents variable Xi For X j the causal strength, which is jointly determined by expert experience and literature data and calculated through fuzzy numbers;
[0081]
[0082] wherein, I k (X i → X j ) represents the causal relationship score of the k-th data source (literature or expert opinion) for X i causing X j , and μ k is its fuzzy weight.
[0083] Take the fuzzy causal relationship matrix C as the prior information of the Bayesian network to construct the initial causal structure; subsequently, use the Markov Chain Monte Carlo (MCMC) method to optimize the Bayesian network structure and estimate the strength of the causal relationship.
[0084] The optimization objective function of the Bayesian network: wherein, Pa(X i ) represents the set of parent nodes of X i in the graph G, and Ω(G) is the structural regularization term used to avoid overfitting.
[0085] The specific process of MCMC sampling is as follows. First, initialize the causal graph G0 and perform random initialization according to C; then use the Metropolis-Hastings algorithm to sample in the parameter space, and the probability of accepting the new graph G′ is: Through continuous iterative updates until convergence, finally obtain the optimized Bayesian network, which can be used to infer the individual's aging influencing factors and calculate the causal weights of each factor.
[0086] For example, if the likelihood of the current graph G, P(D|G) = 0.3, the prior probability of the current graph G, P(D|G) = 0.4, while the likelihood of the new graph G′, P(D|G′ = 0.5), and the prior probability of the new graph G′, P(D|G′) = 0.3, then the probability of accepting the new graph G′ can be calculated:
[0087] 1. Calculate the numerator: P(D|G′)P(G′) = 0.5 × 0.3 = 0.15;
[0088] 2. Calculate the denominator: P(D|G)P(G) = 0.3 × 0.4 = 0.12;
[0089] 3. Calculate the ratio:
[0090] 4. Calculate the acceptance probability: P(G′|G) = min(1, 1.25) = 1.
[0091] The probability of accepting the new graph G' is 1, which means the new graph G' will be accepted. For example: in a causal analysis based on a Bayesian network, smoking (pack-year > 20) → CRP↑ (p = 0.001); CRP↑ → IL-6↑ (p = 0.003); IL-6↑ → accelerated biological age (p = 0.002). By modeling with a mixed causal effect graph (CEG), the mediating role of inflammatory factors between smoking and accelerated aging is determined, and it is calculated that each unit increase in CRP level will increase the biological age by 0.8 years.
[0092] Biological age estimation is an important part of the aging clock; in this embodiment, the integrated K-modes (EKM) method is used to cluster individual samples, and regression modeling is performed in combination with models such as support vector machine (SVM) and random forest (RF). In the EKM clustering method, the K-modes clustering centers are randomly initialized multiple times and weighted averages are taken to improve the stability of the clustering.
[0093] Objective function Among them, X i is the sample data, C k is the clustering center, w k is the clustering weight, and d(X i , C k ) is the distance from the sample to the clustering center. For each clustering cluster, in this embodiment, SVM and RF are used for biological age regression, and the regression objective is: y i = f(X i ) + ∈, where f(X i ) is learned by SVM or RF, and ∈ is the error term. Finally, the predicted values of multiple models are fused by weighted voting to obtain a more accurate biological age.
[0094] To enhance the applicability of CEG in the aging clock, in this embodiment, a personalized aging risk assessment system is constructed by combining the hierarchical weighted TOPSIS method.
[0095] Define the accelerated aging factor Among them, S ij represents the standardized score of individual i on the jth aging dimension, and w j is the weight of different aging dimensions, which is dynamically adjusted by the Bayesian optimization method.
[0096] In addition, TOPSIS evaluates the individual aging risk by calculating the relative closeness of the positive ideal solution and the negative ideal solution. Define the positive ideal solution A + = max i S ij and the negative ideal solution A - = min iS ij , calculate the distance from individual i to the positive ideal solution and the negative ideal solution Finally, calculate the aging risk index of the individual
[0097] For example, the distance from a certain individual i to the positive ideal solution The distance to the negative ideal solution is Then the calculated aging risk index of the individual is R i The closer it is to 1, the higher the aging risk indicates.
[0098] This embodiment can not only accurately evaluate the individual aging process, but also has interpretability and clinical applicability, providing important technical support for personalized health management. In practical applications, the aging risk of an individual can be evaluated by regularly measuring individual health data and combining with CEG, and a personalized intervention plan can be provided based on the TOPSIS evaluation results. For example, for individuals with a higher aging acceleration factor, specific lifestyle adjustments can be recommended, such as optimizing diet, increasing physical activity, or adopting precision drug interventions to reduce the aging rate. In addition, the system can be integrated into an intelligent health management platform to automatically collect physiological indicators and perform real-time analysis, enabling individuals to keep track of their own aging status at any time and take corresponding measures, further improving the scientific and intelligent level of health management.
[0099] This embodiment takes the healthy aging score as the core and establishes a set of personalized health management systems with dynamic adjustment; the system accurately evaluates the individual health status, predicts future health trends, and provides targeted lifestyle intervention suggestions by integrating multi-omics data, biomarker analysis, and individual health behavior information, so as to achieve individualized health management and early warning.
[0100] The healthy aging score is calculated based on multiple biological indicators (such as DNA methylation level, telomere length, protein expression profile, metabolomics data, etc.); the system uses a weighted linear regression model combined with a Bayesian update mechanism to optimize the score, making it more predictive and individualized adaptable.
[0101] Suppose the healthy aging score of an individual at time t X i is the measured value of the i-th biomarker, w i is its weight, and ∈ is the random error term. The system calculates the aging rate through time series data Δt is the measurement time interval. When the individual's AR continuously exceeds the average value of the same age group (such as more than 1 standard deviation), the system will prompt an increase in health risk and generate a personalized intervention plan.
[0102] For individuals with a relatively low healthy aging score and a relatively fast aging rate, an intervention plan consisting of five major modules is provided: lifestyle adjustment, nutritional intervention, exercise management, sleep optimization, and psychological regulation.
[0103] Lifestyle adjustment: The system conducts a health risk assessment by integrating an individual's daily behavior data (such as the number of steps, sedentary time, eating habits, etc.) and provides scientific and reasonable suggestions for optimizing the lifestyle. For example, if an individual's daily activity level is insufficient and the sedentary time exceeds 8 hours, the system will send a reminder and recommend methods to increase the daily activity level, such as "take a short stand or walk every 30 minutes of work". In addition, for those with an abnormal decline in the healthy aging score, the system may recommend optimizing the work-rest balance, such as increasing meditation, social interaction, or outdoor activities, to reduce the negative impact of long-term chronic stress on health.
[0104] Nutritional intervention: The system combines an individual's metabolomic characteristics and microbiome data to provide precise dietary guidance. For example, for individuals with a relatively high level of chronic inflammation (such as elevated CRP and IL-6), the system may recommend a low-inflammatory diet, reduce the intake of refined carbohydrates and saturated fatty acids, and at the same time increase the intake of foods rich in antioxidant components (such as resveratrol, quercetin, catechins). For those with abnormal glucose metabolism, a low-GI diet or a ketogenic diet is recommended to improve insulin sensitivity. In terms of nutritional supplements, the system can calculate the optimal intake of dietary supplements according to an individual's nutritional needs. For example: supp D target = k(T current ),D supp is the recommended supplementary dose, T target is the target level, T current is the current detected value, and k is an individual adjustment factor.
[0105] Exercise management: The system analyzes an individual's exercise pattern using data from smart wearable devices and physiological parameters, and identifies the risk of insufficient exercise through the TimeSeriesClustering method. If it is detected that an individual's daily activity level is below the healthy threshold, the system will dynamically adjust the exercise goal. For example: E t represents the current exercise recommendation amount (such as the number of steps per day or the training intensity), is the population mean, and α is an individual adaptation coefficient. For individuals with a relatively high risk of muscle decline, the system will recommend resistance training (such as weight training), while for those with a decline in cardiovascular health, high-intensity interval training (HIIT) or moderate-intensity aerobic exercise (such as brisk walking, swimming) is recommended.
[0106] Sleep Optimization: The system evaluates an individual's sleep quality based on polysomnography (PSG) and wearable device data, and makes adjustments in combination with the melatonin secretion rhythm. For example, if the individual's sleep quality score (based on the Pittsburgh Sleep Quality Index) is below 75 points and the proportion of deep sleep is too low, the system may recommend adjusting melatonin secretion, such as: D melatonin = k(T bedtime - T optimal ), D melatonin is the recommended melatonin supplementation dose, T bedtime is the current bedtime, T optimal is the optimal bedtime, and k is the individual adaptation coefficient. In addition, the system will adjust the light exposure strategy to optimize melatonin synthesis and increase the proportion of deep sleep. For example, the system may recommend "reducing blue light exposure after 8 pm and getting at least 30 minutes of natural light exposure in the morning to optimize the circadian rhythm."
[0107] Psychological Regulation: The system combines heart rate variability (HRV) monitoring and stress index to evaluate an individual's emotional state, and optimizes the intervention plan through Reinforcement Learning (RL). Suppose the individual's intervention strategy set is A = {a1, a2,..., a n}, and its cumulative reward function is: r i represents the effect feedback of the i-th intervention, and γ is the discount factor (0 < γ ≤ 1).
[0108] For example, for an individual with an intervention strategy set A = {a1, a2, a3}, the system obtains the following effect feedback r i in three interventions: For the first intervention: r1 = 0.8, for the second intervention: r2 = 0.6, for the third intervention: r3 = 0.9. Suppose the discount factor y = 0.9, and the calculation of its discounted reward can be obtained by the following method: ① Discounted reward for each intervention: For the first intervention (i = 1) y t- i r1 = y 3-1 × 0.8 = 0.9 2 × 0.8 = 0.81 × 0.8 = 0.648 For the second intervention (i = 2) y t-i r2 = y 3-2 × 0.6 = 0.9 1 × 0.6 = 0.9 × 0.6 = 0.54; For the third intervention (i = 3) y t-i r3 = y 3-3 × 0.9 = 0.9 0 × 0.9 = 1 × 0.9 = 0.9 ② Calculate the cumulative reward R t : Add the discounted rewards of each intervention: Rt = 0.648 + 0.54 + 0.9 = 2.088.
[0109] The system adjusts the strategy according to real-time feedback. For example, for individuals with an anxiety score (GAD-7) exceeding 10 points, the system may recommend HRV biofeedback training, mindfulness meditation, or short-term cognitive behavioral therapy (CBT), and optimize the intervention plan in combination with neurotransmitter metabolism data (such as serotonin and dopamine levels). Another example: for individuals with an accelerated aging index (AAI) higher than 1.2, the system recommends: ① increasing at least 30 minutes of moderate-to-high-intensity exercise per day, such as brisk walking or swimming; ② restricting the intake of refined sugar, with daily sugar intake controlled below 25g; ③ increasing foods rich in antioxidant components, such as blueberries (50g per day) and green tea (1-2 cups per day); ④ for those with a vitamin D level below 20 ng / mL, it is recommended to supplement 1000 IU of vitamin D3 per day.
[0110] In view of the limited computing resources of primary medical institutions, a lightweight computing framework is designed to ensure that even in a low-computing-power environment, the processing and analysis of health data can still be efficiently completed. Traditional health management systems rely on cloud computing. However, cloud data processing not only increases communication latency and bandwidth pressure but also involves user privacy and security issues. Therefore, the present invention adopts the edge computing mode, enabling data processing, feature extraction, and model inference to be completed locally, thereby reducing the dependence on cloud computing resources and improving data security and response speed at the same time.
[0111] In this computing framework, the system first uses model distillation technology to enable complex deep learning models to run on low-computing-power devices. By training a complex teacher model (TeacherModel) and then having a smaller student model (StudentModel) learn its output, the amount of computation is reduced while still maintaining high accuracy.
[0112] For example, in the health data prediction task, assume that the prediction function f T (x) = σ(W T x + b T ) of the teacher model, where x is the input health data, W T and b T are model parameters, and σ is a non-linear activation function. The student model learns the soft labels of the teacher model instead of directly learning from the data, reducing the computational complexity while keeping the accuracy within an acceptable range. The optimization process uses a weighted combination of cross-entropy loss and Kullback-Leibler (KL) divergence loss: L = αL CE + (1 - α)L KL, L CE Measure the error between the student model and the true label, L KL Measure the error between the student model and the teacher model, and α controls the weight of the two. In this way, a more lightweight model can be run on the edge device to achieve efficient inference.
[0113] In addition, the system adopts model quantization technology to reduce the computational load and storage requirements. The 32-bit floating-point model trained in the cloud will be converted into an 8-bit or 16-bit fixed-point model to adapt to the low-computing-power environment. For example, if the storage precision of a user's continuous heart rate data is relatively high (32-bit floating point), when calculating on the edge device, the system will convert it to: w is the original weight value, w min and w max are the minimum and maximum values of this weight, and b is the quantization bit number (such as 8-bit). This can greatly reduce the storage and computational requirements without significantly affecting the prediction results.
[0114] In terms of optimizing data transmission, the sliding window data interception method is adopted to reduce the data transmission volume while maintaining information integrity. For example, when collecting blood oxygen data by a wearable device, the system will not upload the complete original waveform data, but calculate key feature values, such as the mean, standard deviation, and Fourier transform features F(x) = {μ(x), σ(x), FFT(x)}. Only transmit the necessary data features, which not only reduces the bandwidth occupancy but also protects user privacy.
[0115] To make the analysis of health data more intuitive, the present invention also develops a set of visual health management platforms, which can display the individual's health trends, aging acceleration warnings, and personalized intervention suggestions through the Web side or the mobile side. For example, when the health aging score of a certain user shows an abnormal downward trend, the system will automatically draw a time series analysis chart and calculate the deviation from the health baseline: S t is the current health score, S baseline is the population mean, and σ is the standard deviation. The system will also combine the individual's living habits to propose specific intervention measures. For example, if a user's aging rate is higher than 1.5 standard deviations of the mean and the average daily exercise volume is less than 6000 steps, the system may suggest that the user increase the daily walking time and give specific goals: Among them, E t represents the recommended exercise volume (such as the number of steps per day or the training intensity), is the mean of the healthy population, and α is the individual adaptation coefficient. Users can view the exercise recommendations on the mobile side and synchronize them to the smart bracelet, and the system can remind the user to exercise regularly through the bracelet vibration.
[0116] In terms of sleep optimization, the platform combines the data from wearable devices, analyzes the user's sleep cycle, and proposes improvement plans. Suppose the polysomnography (PSG) data of a certain user shows that the proportion of deep sleep is less than 15%, and the user wakes up frequently at night. The system may recommend adjusting the melatonin supplement dosage: D melatonin = k(T bedtime - T optimal ), D melatonin is the recommended melatonin supplement dosage, T bedtime is the current sleep onset time, T optimal is the optimal sleep onset time, and k is an individual adjustment factor. The system may also recommend that the user reduce nighttime blue light exposure, such as reducing the use of electronic devices after 8 pm, to promote melatonin secretion.
[0117] This embodiment makes full use of wearable devices (such as smart bracelets and blood glucose monitors) to collect physiological data and performs dynamic analysis through cloud AI. The system uses a long short-term memory network (LSTM) + attention mechanism to model long-term health trends. For example, if a user's blood glucose level has been fluctuating for a long time, the LSTM will predict its future blood glucose trend: h t = σ(W h h t-1 + W x x t + b h ), h t is the hidden state, x t is the current input, and W h and W x are weight matrices. Through the attention mechanism, the system will automatically focus on the key factors affecting blood glucose fluctuations, such as too rapid a decrease in blood glucose at night or abnormal blood glucose response after eating, and provide personalized diet recommendations.
[0118] The above are only the preferred embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications and substitutions based on the technical solutions and inventive concepts provided by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for constructing an aging clock based on a causal machine joint model, characterized in that: The steps include: Step 1: Obtain a multimodal data set of the target population based on multimodal data fusion and standardized preprocessing; the multimodal data set includes clinical phenotype data, DNA methylation data, transcriptome data, and metabolome data; Step 2: Perform two-stage screening on the multimodal dataset to obtain core features, wherein the two-stage screening includes a screening method for eliminating low-correlation features based on the Pearson correlation coefficient, and a screening method for eliminating low-importance features based on recursive feature elimination combined with a machine learning model; Step 3: The interaction between genotype group, nutritional group and metabolomics was analyzed based on the extended proportional odds logistic regression mixed model; the Score test was used to analyze the association between each single genetic variant and multi-classification phenotype and generate P value; then the P value was calibrated by saddle point approximation method; Among them, the expression of the extended proportional advantage logistic regression mixed model is: k =α0+θG k +γ T X k +β T Z k +e k +b k , G k represents the genotype variable, X k represents the nutrigenomic variable, Z k represents the metabolomics variable, e k represents the interaction between omics, b k represents the description of genetic correlation, θ, γ, β are model parameters, and α0 is the intercept term; Step 4: Construct an aging clock based on a mixed causal effect graph model to evaluate multidimensional causal relationships, assess the biological age of individuals, and output personalized health assessment reports and aging acceleration warning information; The construction of an aging clock based on a hybrid causal effect graph model includes: first, extracting aging-related causal relationships from medical literature and expert knowledge base, and using fuzzy theory to process causal relationships to output a fuzzy causal relationship matrix; Then, the strength of causal relationship is estimated based on Bayesian network and fuzzy causal relationship matrix; the partitioned MCMC method is used to explore and optimize the Bayesian network; Finally, the degree and rate of aging were used as outcome indicators for risk assessment, and the optimal parameters of the Bayesian network were selected through grid search and 5-fold cross validation to obtain the aging clock.
2. The method for constructing an aging clock based on a causal machine joint model according to claim 1, characterized in that: The standardized preprocessing refers to: using the Z-score normalization method to process the data of continuous quantitative variables; using one-hot encoding to process the data of categorical variables; using the K-nearest neighbor interpolation or multiple interpolation method to process the data of missing values; and using the Bayesian inference-based method to process the systematic missing data.
3. The method for constructing an aging clock based on a causal machine joint model according to claim 1, characterized in that: In step 2, the screening method for eliminating low-correlation features based on the Pearson correlation coefficient means that when the Pearson correlation coefficient |r| between the feature in the multimodal dataset and the aging clock is less than 0.1, the feature is a low-correlation feature and should be eliminated.
4. The method for constructing an aging clock based on a causal machine joint model according to claim 3, characterized in that: The screening method based on recursive feature elimination combined with a machine learning model to remove low-importance features refers to: selecting a machine learning model, and using a data set with low-relevance features removed for training, and calculating the importance scores of the features in the data set, removing features with importance scores below a threshold, and repeating the training until a preset number of features is reached or the model performance is no longer significantly improved; Among them, the machine learning model is random forest or LASSO regression; random forest uses Gini importance to calculate the importance score; LASSO regression uses L1 regularization coefficient to calculate the importance score.
5. The method for constructing an aging clock based on a causal machine joint model according to claim 1, characterized in that: The parameters θ, γ, and β of the extended proportional odds logistic regression mixed model were optimized and estimated using the penalized quasi-likelihood method and the average information restricted maximum likelihood algorithm.
6. The method for constructing an aging clock based on a causal machine joint model according to claim 1, characterized in that: Calibrating the P value by the saddle point approximation method means: using the cumulant generating function to estimate the probability density function or cumulative distribution function of the distribution; when the statistic T adj <2, the P value is calculated based on the normal approximation; when the statistic T adj ≥2, the P value is calculated using the saddle point approximation.
7. The method for constructing an aging clock based on a causal machine joint model according to claim 1, characterized in that: The biological age of an individual is estimated through multi-omics variable extrapolation and regression modeling.
8. The method for constructing an aging clock based on a causal machine joint model according to claim 1, characterized in that: Accelerated aging warning information is obtained by calculating the individual aging risk index using the stratified weighted TOPSIS method, thereby assessing the individual aging risk.
9. A lightweight aging clock system based on a causal machine joint model, characterized by: The system includes one or more processors and a memory, the memory stores a program, and the program is executed by the processor to implement the aging clock construction method based on the causal machine joint model described in claims 1 to 8; the processor uses edge computing mode and model quantization technology to execute the program.
10. The method for constructing an aging clock based on a causal machine joint model according to claim 1, characterized in that: The processor is connected to a display of a visual health management platform.
Citation Information
Cited By
Biological age prediction method and device, electronic device and storage medium
CN120727299A
Biological age prediction method, device, electronic device, and storage medium
CN120727299B
Kidney biological age calculation method based on glycosylation marker and integrated machine learning
CN122050801A