Method for constructing biological age assessment model based on body composition and proteome data'two-step method 'and application
Through the ‘two-step’ method of combining body composition and proteomic data, a biological age assessment model was constructed, which solved the limitations of the existing technology in capturing the body and muscle aging status, and achieved a more accurate assessment of the muscle aging situation in Chinese population and the formulation of personalized anti-aging solutions.
Patent Information
- Application Number
- CN202510098634.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing biological age models have limitations in capturing the body and muscle aging status, especially when reflecting the muscle aging status of the Chinese population, lack stability and accuracy.
A biological age assessment model was constructed using a ‘two-step’ method based on body composition and proteomic data. First, the body composition age of an individual is calculated by a biological age model constructed based on body composition data; second, the biological age evaluation model is constructed using proteomic data and elastic network regression algorithm.
This method is not only significantly related to temporal age, but also has a significant correlation with the number of individual diseases after excluding the impact of temporal age. It can more accurately reflect the individual's muscle aging status and provide personalized anti-aging solutions.
Smart Images

Figure SMS_2 
Figure SMS_3 
Figure SMS_4
Abstract
Description
Technical Field
[0001] The present invention is applicable to the technical field of human biological age assessment. Specifically, it relates to a method and application for constructing a biological age assessment model based on body composition and proteomic data, which is applicable to the Chinese population. Background Art
[0002] The attenuation of muscle function is one of the typical changes in aging, which will increase the health risks of the elderly such as falls, disability, disability, and cognitive impairment. Accurately assessing the muscle mass and function of the elderly will help improve the intrinsic ability of the elderly to maintain motor function and achieve active health.
[0003] Chronological Age (CA) is regarded as an important parameter for evaluating the degree of individual aging and predicting the risk of death. However, there are significant heterogeneities in the aging process among individuals. A single chronological age cannot comprehensively capture the aging differences at the biological level of individuals, and these differences are crucial for deeply understanding the essence of aging and its impact on health. Correspondingly, Biological Age (BA) can more accurately evaluate the degree of individual aging through different-dimensional biomarkers, and has been proven to have great application potential in the accurate prediction of life expectancy, disease risk, and disease prognosis. Among them, the body composition age based on muscle-related body composition data can accurately reflect the attenuation state of muscle function with age. Fermín-Martínez et al. used the US NHANES database to establish the biological age prediction models "AnthropoAge" for men and women respectively based on the waist-to-height ratio (WHtR), arm circumference, and thigh circumference of men, and the weight, WHtR, thigh circumference, subscapular, and triceps skin folds of women. At the same time, a simplified prediction model "S-AnthropoAge" based on BMI and WHtR was developed (Aging Cell. 2023 Jan; 22(1): e13756.). We previously established a biological age prediction model based on the body composition data of the Chinese population, which can more accurately evaluate the true muscle health status of individuals. The above biological age assessment models based on body composition data are characterized by being fast and non-invasive, but are vulnerable to many factors such as short-term changes in individual nutritional diets, lack stability, and cannot reflect the biological change process inside the body.
[0004] Skeletal muscle is mainly composed of structural proteins and contractile proteins, and protein integrity directly affects muscle strength, mass, and function. Muscle homeostasis is maintained by the balance between protein synthesis and degradation signals. When muscle protein degradation exceeds the synthetic capacity, there is a trend of muscle attenuation. Ceereena Ubaida-Mohien et al. used tandem mass tag (TMT) protein quantification based on mass spectrometry and SOMAscan proteomics methods to analyze proteins in plasma and skeletal muscle of the Baltimore Longitudinal Study of Aging and the Geriatric Evaluation by Laboratory Tests and Epidemiology (GESTALT), and applied an epidemiological model to study the data. They found that mitochondrial proteins in skeletal muscle decreased with age, while spliceosome complex proteins showed an increasing trend (Methods Mol Biol. 2022;2399:173-192.). However, current proteomics studies on muscle function mainly focus on European and American populations, and there are significant differences in the degree of muscle aging among different ethnic groups. These research results are not applicable to the Chinese population. Secondly, current studies mainly focus on the elderly, so the results of these studies lack accuracy when applied to young people.
[0005] Based on the above background, we innovatively constructed a biological age clock through a "two-step method" based on the body composition and proteomics data of the Chinese adult population, accurately assessing the true muscle health status of individuals, which is conducive to the stratified management and preventive care of high-risk populations of muscle attenuation in the early stage of life (such as middle-aged and young people). This is of great significance for identifying individuals with accelerated aging and establishing personalized and precise anti-aging programs. Summary of the Invention
[0006] The present invention aims to overcome the limitations of existing biological age models in capturing the aging status of the body and muscles, and proposes a method - the two-step method - for constructing a biological age assessment model for the Chinese population based on body composition and proteomics data. The biological age evaluated by using this model is not only significantly correlated with chronological age, but also significantly associated with the number of individual diseases after excluding the influence of chronological age. The present invention also provides a biological age assessment method applicable to the Chinese population, which can help users scientifically evaluate their own muscle aging status, so as to take effective measures for prevention and intervention in a timely manner to delay the aging process and reduce the risk of occurrence of geriatric diseases.
[0007] A biological age assessment method, which obtains the key protein indicators of the person to be evaluated and inputs them into the improved biological age assessment model I to obtain the biological age of the person to be evaluated; the improved biological age assessment model I is constructed based on the elastic net regression algorithm using the proteomics data and body composition age of an individual as training data.
[0008] Further, the body composition age is obtained from the biological age evaluation model II constructed based on body composition data.
[0009] Further, the body composition data is obtained from muscle fat analysis, body composition analysis, obesity analysis, muscle balance analysis, segmental water analysis, segmental extracellular water ratio analysis, extracellular water ratio analysis, segmental fat analysis, and body measurement information.
[0010] Furthermore, the biological age evaluation model II constructed based on body composition data is the model structure described in the publication number CN117711623A. Specifically as follows:
[0011] The biological age prediction model for women is:
[0012] Biological age
[0013] = (ln(BMI) - 3.101961) * 0.274298
[0014] + (BCWtoTBW - 0.383087) * 23.957711
[0015] - (ln(BCWtoTBW_LA) + 0.971392) * 22.645593
[0016] + (BCWtoTBW_LL - 0.384664) * 30.010663
[0017] + (ln(ECWtoTBW_RA) + 0.971933) * 11.754412
[0018] + (ECWtoTBW_RL - 0.38359) * 15.232615
[0019] + (ECWtoTBW_TR - 0.383359) * 13.887273
[0020] + (ln(FFM_LA_p) - 4.559868) * 0.486808
[0021] + (ln(FFM_RA_p) - 4.586863) * 2.527928
[0022] - (FFM_RL_p - 96.404821) * 0.031468
[0023] + (ln(FFMI) - 2.733266) * 0.663589
[0024] -(ln(Height) - 5.068349) * 3.15674
[0025] +(ln(Obesitydegree) - 4.662253) * 0.265601
[0026] -(FFM_LL_p - 96.148192) * 0.014774
[0027] +(ln(TBWtoFFM) - 4.296985) * 22.027408 - (ln(VFL)
[0028] -1.975382) * 0.224676
[0029] Where: BMI is the body mass index, ECWtoTBW is the extracellular water ratio, ECWtoTBW_LA is the extracellular water ratio of the left upper limb, ECWtoTBW_LL is the extracellular water ratio of the left lower limb, ECWtoTBW_RA is the extracellular water ratio of the right upper limb, ECWtoTBW_RL is the extracellular water ratio of the right lower limb, ECWtoTBW_TR is the extracellular water ratio of the left upper limb, FFM_LA_p is the fat-free mass of the left upper limb (%), FFM_RA_p is the fat-free mass of the right upper limb (%), FFM_RL_p is the fat-free mass of the right lower limb (%), FFMI is the fat-free mass index, Height is the height, Obesity_degree is the obesity degree, FFM_LL_p is the fat-free mass of the left lower limb (%), TBWtoFFM is the total body water / fat-free mass, and VFL is the visceral fat level.
[0030] The biological age prediction model for men is:
[0031] Biological age
[0032] =(Circ LT -52.119365) * 0.000078
[0033] +(Circ_RA - 32.108621) * 0.031951
[0034] +(ln(ECWtoTBW_TR)+0.974133) * 36.644066
[0035] -(FFM_LA_p - 95.154181) * 0.008383
[0036] -(ln(FFM_RL_p)-4.569205) * 1.207248
[0037] -(FFM_TR_p - 98.489964) * 0.005927
[0038] -(ln(Height) - 5.134496) * 6.316220
[0039] +(ln(Inbodyscore) - 4.244677) * 0.917546
[0040] -(ln(FFM_LL_p) - 4.562414) * 0.601829 + (SMI
[0041] - 7.828159) * 0.212899 + (ln(TBWtoFFM) - 4.298732)
[0042] * 24.649412
[0043] Where: Circ_LT is the left thigh circumference, Circ_RA is the right arm circumference, ECWtoTBW_TR is the extracellular water ratio of the left upper limb, FFM_LA_p is the fat-free mass of the left upper limb (%), FFM_RL_p is the fat-free mass of the right lower limb (%), FFM_TR_p is the relative percentage of the trunk muscle mass compared to the standard muscle mass value of the actual body weight, Height is the height, Inbody score is the InBody score, FFM_LL_p is the fat-free mass of the left lower limb (%), SMI is the skeletal muscle index, and TBWtoFFM is the total body water / fat-free mass.
[0044] Another object of the present invention is achieved by the following technical solution: A "two-step" method for constructing a biological age assessment model based on body composition and proteomic data, comprising the following steps
[0045] (1) Calculate the body composition age of an individual through a biological age model constructed based on body composition data, denoted as Muscle Age (MA);
[0046] (2) Collect the proteomic data of the individual, and construct a biological age assessment model based on the body composition age using the elastic net regression algorithm.
[0047] In step (1), the biological age model constructed based on body composition data is constructed by gender among more than 800 Chinese young, middle-aged, and elderly individuals, with an age coverage range of 18 - 87 years.
[0048] Furthermore, the construction method of the improved biological age assessment model I is as follows:
[0049] (2 - 1) Collect human proteomic indicators and perform preprocessing;
[0050] (2-2) Using Spearman partial correlation analysis to screen candidate protein indicators;
[0051] (2-3) Using candidate protein indicators and body composition age as the dependent variable, the biological age assessment model I is constructed using the elastic network regression algorithm.
[0052] In step (2-1), physical examination data and blood samples of healthy people are collected, and DIA (Data-independent acquisition) mass spectrometry technology is used to measure the proteomic data in the blood samples, and the proteomic data is cleaned, quality controlled and standardized;
[0053] In step (2-2), Spearman partial correlation analysis is used to screen candidate protein indicators for constructing a biological age assessment model;
[0054] In step (2-3), candidate protein indicators were used, body composition age was taken as the dependent variable, and the elastic network regression algorithm was used for further feature screening to construct a biological age assessment model, which was denoted as Proteomics-inferred MuscleAge (ProteomicMA).
[0055] After the model is constructed, the effectiveness of the constructed biological age assessment model can be further evaluated.
[0056] Furthermore, the preprocessing includes:
[0057] (a) Protein markers missing in more than 40% of samples were eliminated;
[0058] (b) Use k-nearest neighbor (KNN) interpolation method to fill missing values in the remaining data;
[0059] (c) Natural logarithm transformation of protein indexes was performed to make them more consistent with normal distribution;
[0060] (d) Z-score normalization of protein indicators.
[0061] Furthermore, when applying Spearman partial correlation analysis to screen candidate protein indicators, Spearman partial correlation analysis was performed on each protein indicator and body composition age, and the control variables of the Spearman partial correlation analysis were gender, drinking status, smoking status, education level, income level, physical activity level, marital status and number of diseases; the correlation P value between each protein indicator and body composition age was obtained, and protein indicators with correlation P values less than the target value (for example, P values less than 0.05) were selected as candidate protein combinations M for constructing a biological age assessment model.
[0062] More specifically, the controlled variables are as follows: gender (male; female), drinking status (currently still drinking; has quit drinking; never drinks), smoking status (currently still smoking; has quit smoking; never smokes), education level (primary school or below; junior high school, high school, technical secondary school; junior college or above), income level (less than 10,000 yuan; 10,000 yuan or above), physical activity level (low activity level; medium and high activity levels), marital status (married; unmarried), and number of diseases (the number of the following diseases currently suffered from: diabetes, hypertension, heart disease, cancer, stroke, gout, lung disease, liver disease, kidney disease, and stomach disease).
[0063] How to select representative biomarkers for muscle aging from a large number of protein indicators is an important technical highlight of the present invention. The present invention applies statistical thinking to screen protein indicators, fully considering the association between protein indicators and body composition age, so as to select a protein indicator combination that is representative of the muscle aging state to the greatest extent. As an implementation method, first, the biological age model constructed based on the existing body composition data is used to calculate the body composition age of an individual, and then Spearman partial correlation analysis is performed between each protein indicator and the body composition age. The influences of gender, drinking status, smoking status, education level, income level, physical activity level, and marital status are adjusted in the model to obtain the correlation P value between each protein and the body composition age. Proteins with a correlation P value less than 0.05 are selected as the candidate protein combination M for constructing the biological age assessment model.
[0064] In an embodiment, the selected candidate protein combination M includes the following 78 proteins: A0A0C4DH35, A0A0J9YXX1, O00300, O00429, O14791, O75915, O75955, P00387, P01743, P01772, P01861, P04632, P05089, P05107, P05387, P05556, P07585, P08575, P11413, P12236, P12814, P12883, P13667, P15170, P15880P20340, P23229, P24941, P26927, P28906, P30040, P32121, P36871, P36959, P40121, P43487, P50148, P51888, P53041, P53621, P60842, P60953, P61981, P62241, P62942, P68363, Q02809, Q07507, Q08431, Q13177, Q13586, Q14165, Q14677, Q16674, Q16762, Q641Q3, Q6Q788, Q6WN34, Q86YW5, Q8IWY4, Q8IZP0, Q8N5C6, Q8N699, Q92520, Q92765, Q96AX2, Q96P63, Q96QR1, Q99719, Q99988, Q9BUN1, Q9BXJ1, Q9BZE9, Q9H0B8, Q9H4G4, Q9P270, Q9UGI8, Q9UIB8. These protein indicators are significantly correlated with body composition age and can better reflect the muscle aging state of an individual.
[0065] In addition, the selection of the algorithm is a difficult point that the present invention needs to overcome. Multiple algorithms have been used by scholars to construct biological age assessment models based on biomarkers, such as multiple linear regression, Lasso regression, Ridge regression, Elastic Net Regression, and machine learning algorithms. Elastic Net Regression combines the advantages of Lasso regression and Ridge regression, and simultaneously introduces L1 and L2 norm regularization terms in the loss function. This combination enables Elastic Net Regression to flexibly handle the problems of multicollinearity and feature redundancy, and has strong stability and robustness to noise on complex datasets. Therefore, the present invention finally selects the Elastic Net Regression algorithm to construct the biological age assessment model. The Elastic Net Regression algorithm was proposed by Hui Zou and Trevor Hastie in 2005, and the derivation details can be found at the following website: https: / / scikit-learn.org / stable / modules / generated / sklearn.linear_model.ElasticNet.html.
[0066] Briefly, the working process of the Elastic Net Regression algorithm is as follows: First, assign initial weight coefficients to m protein features to form an initialized prediction model (m-dimensional hyperplane), and calculate the initial loss function value based on this. Then, the algorithm gradually updates these weight coefficients through the coordinate descent method until the loss function converges to the minimum value. Specifically, in each iteration process, the Elastic Net Regression model will fix all other weights and only optimize the weight of a specific feature to reduce the value of the loss function. This process will continue until the loss function completely converges. At this time, the weights of each protein feature reach the optimal state, that is, the constructed biological age prediction model achieves the highest prediction accuracy under the set hyperparameter conditions. The loss function formula of Elastic Net Regression is as follows:
[0067] ElasticNetLoss = MSE + α * [λ * L1 norm + (1 - α) * 0.5 * λ * L2 norm
[0068] Among them, MSE represents the mean square error between the model prediction value and the actual value, which is used to quantify the difference between the model prediction value and the actual value; λ is the regularization parameter, which controls the strength of regularization; L1norm is the L1 regularization term, that is, the sum of the absolute values of the model coefficients, which helps to sparse feature selection; L2norm is the L2 regularization term, that is, the sum of the squares of the model coefficients, which helps to reduce overfitting; α is a parameter between 0 and 1, which is used to balance the effects of L1 and L2 regularization. In particular, when α = 0, the model degenerates into Ridge regression; when α = 1, the model is equivalent to Lasso regression.
[0069] As an implementation, the present invention uses the R language software package "caret" to implement the automatic parameter tuning process of the elastic net regression algorithm. The parameter settings are: alpha = 0.5, the link function is selected as "Gaussian", and the penalty parameter lambda is determined by the method of 10-fold cross-validation, and the lambda value that minimizes the root mean square error is selected. Finally, the elastic net regression model is retrained with the optimal parameters to obtain the final biological age assessment model.
[0070] Through training, the parameter values of the final biological age model are alpha = 0.5 and lambda = 1.6.
[0071] The final elastic net regression model contains 21 proteins (i.e., key protein indicators), including O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, Q9H4G4.
[0072] Furthermore, the structure of the improved biological age assessment model I is as follows:
[0073] ProteomicMA
[0074] = [ln(O75915) - 3.744371] * 0.999284
[0075] + [ln(P00387) - 3.910639] * 0.121916
[0076] + [ln(P01772) - 7.429088] * -0.7755
[0077] + [ln(P05089) - 4.153092] * -0.985205
[0078] + [ln(P05387) - 3.179462] * 1.021923
[0079] + [ln(P08575) - 3.784147] * -2.114094
[0080] + [ln(P12883) - 4.816762] * -0.060567
[0081] + [ln(P15170) - 2.358900] * -1.183973
[0082] +[ln(P15880) - 2.694445] * 0.062588
[0083] +[ln(P23229) - 4.548532] * 0.891459
[0084] +[ln(P28906) - 4.880842] * 0.671088
[0085] +[ln(P50148) - 4.106702] * -1.231852
[0086] +[ln(P53041) - 2.730823] * 1.545147
[0087] +[ln(P61981) - 4.036439] * -0.083734
[0088] +[ln(P62241) - 4.054173] * -0.315841
[0089] +[ln(Q07507) - 5.315179] * 0.097354
[0090] +[ln(Q14677) - 3.687083] * 1.163949
[0091] +[ln(Q16762) - 3.254478] * -0.003892
[0092] +[ln(Q92765) - 4.763237] * -0.189882
[0093] +[ln(Q96QR1) - 3.884713] * -0.6148500 + [ln(Q9H4G4)
[0094] -4.626910] * 3.118311 + 52.453141
[0095] Where: ProteomicMA is the biological age; O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, Q9H4G4 are the key protein indicators retained by the model.
[0096] A construction method for constructing a biological age assessment model based on body composition and proteomic data. The biological age assessment model constructed based on body composition and proteomic data is the improved biological age assessment model I described in any of the above, and is constructed by using the method for constructing the improved biological age assessment model I described in any of the above.
[0097] The present invention uses two methods to evaluate the obtained biological age assessment model I, which are as follows:
[0098] (1) Based on the obtained biological age assessment model, calculate the biological age of an individual, and analyze the association between this biological age and the number of diseases, grip strength, and walking speed of the individual;
[0099] Furthermore, in the present invention, based on this biological age evaluation model, calculate the biological age of an individual, and use a linear regression model and a Logistic regression model to analyze the association between the biological age evaluation index and the number of diseases, grip strength, and walking speed, and record the regression coefficient value and the P value. If the regression coefficient is greater than 0 and the P value is less than 0.05, it is considered that the biological age evaluation index is significantly correlated with the number of diseases. The significant correlation between the biological age assessment index and the number of diseases, grip strength, and walking speed confirms its predictive efficacy and application value. The evaluation results show that the biological age obtained from the improved biological age assessment model I constructed by the present invention is positively correlated with the number of diseases, and negatively correlated with grip strength and walking speed, indicating that ProteomicMA can effectively reflect the physiological and muscle aging status of an individual.
[0100] (2) Use biological age acceleration to evaluate aging:
[0101] To eliminate the influence of chronological age on the aging process, the present invention defines biological age acceleration to evaluate aging. Biological age acceleration is defined as the residual generated when biological age is linearly regressed against chronological age, which means that an individual is physiologically older (positive value) or younger (negative value) than expected; perform a linear regression of MA and ProteomicMA against chronological age respectively, take the residuals, and obtain MA acceleration (denoted as MA_ACC) and ProteomicMA acceleration (denoted as ProteomicMA_ACC).
[0102] Measuring whether the biological age model based on protein construction can capture more individual muscle aging-related information is a major highlight of our technology. As an implementation, the present invention classifies each individual according to MA_ACC and ProteomicMA_ACC (aging acceleration: MA_ACC>0 or ProteomicMA_ACC>0; aging deceleration: MA_ACC<0 or ProteomicMA_ACC<0), and classifies each individual into a matching or non-matching group according to whether the MA_ACC and ProteomicMA_ACC categories match or not, and compares the differences in clinical indicators between the matching group and the non-matching group. If there are differences in clinical indicators between the matching group and the non-matching group and they are of practical significance, it is considered that the biological age constructed based on body composition and proteomic data can capture individual muscle aging information better than the biological age constructed solely based on body composition data. The evaluation results show that ProteomicMA, as an evaluation index of biological age, can capture the individual's health status and its potential muscle function aging trend more accurately than the biological age constructed based on body composition data.
[0103] The present invention discloses a method for constructing a biological age evaluation model based on body composition and proteomic data, and uses this model to evaluate the biological age of the Chinese population. Calculate the body composition age of an individual through a biological age model constructed based on body composition indicators; use proteomic data, and based on the body composition age, construct a biological age evaluation model using the elastic net regression algorithm; evaluate the aging status by calculating the biological age acceleration. The present invention allows users to calculate the body composition age according to their proteomic data, realizing personalized monitoring of the muscle aging status, and has great guiding significance in identifying individuals with premature muscle function decline and formulating personalized and precise anti-aging programs. The present invention relies on the body composition data and proteomic data in the established Zhejiang Longitudinal Study of Healthy Aging in China (JASHA). First, calculate the individual body composition age through a biological age model constructed based on body composition data, and screen 78 candidate protein indicators that play a role in muscle aging from more than two thousand protein indicators through statistical criteria, and then apply the elastic net regression algorithm to construct a biological age evaluation model. Data prove that there is a correlation between the model ProteomicMA constructed by the present invention and the number of individual diseases, grip strength, and walking speed, and it can effectively reflect the muscle aging status of an individual. In addition, the present invention also compares the reflection ability of the traditional body composition age model and ProteomicMA on the body state, and the results show that ProteomicMA has higher accuracy and sensitivity. Users can use the present invention to calculate their biological age more accurately based on proteomic information, evaluate their own muscle aging status, without the need for complex special tests, and the operation is simple and fast.
[0104] The method for the biological age evaluation model provided by the present invention is applicable to the Chinese population, and has important value especially in the early identification and prevention of premature muscle function decline and the evaluation of the intervention effect of geriatric diseases. This method can provide a scientific basis for formulating personalized and precise anti-aging programs and the research of aging-related diseases, and has guiding significance for promoting secondary and tertiary aging prevention in the population. Description of the Drawings
[0105] Figure 1 It is the process framework for constructing a biological age evaluation model based on body composition and proteomic data.
[0106] Figure 2 It is the correlation between MA and chronological age: among them, MA is the body composition age calculated through the biological age model (existing) constructed based on body composition data, CA is the chronological age, and R is the correlation coefficient.
[0107] Figure 3The correlation between ProteomicMA, MA, and chronological age is as follows: Among them, MA is the body composition age calculated by a biological age model (existing) constructed based on body composition data, ProteomicMA is a new biological age evaluation index constructed based on MA, CA is the chronological age, and the color of the heatmap represents the correlation magnitude between ProteomicMA, MA, and chronological age. The darker the color, the greater the correlation.
[0108] Figure 4 The correlation between ProteomicMA_ACC and MA_ACC is as follows: Among them, ProteomicMA_ACC is the residual obtained by regressing ProteomicMA on chronological age, representing ProteomicMA acceleration; MA_ACC is the residual obtained by regressing MA on chronological age, representing MA acceleration.
[0109] Figure 5 This is a comparison of the grip strength and walking speed of the matched and unmatched groups of ProteomicMA_ACC and MA_ACC. Among them, the abscissa decel represents aging deceleration determined based on MA_ACC, and accel represents aging acceleration determined based on MA_ACC; the color of the box plot is gray for the matched group (Matched), that is, the determination results of ProteomicMA_ACC and MA_ACC are consistent, and blue for the unmatched group (Mismatched), that is, the determination results of ProteomicMA_ACC and MA_ACC are inconsistent. Detailed implementation mode
[0110] The present invention will be further described below in conjunction with the embodiments and the drawings.
[0111] The specific content of the embodiment is as follows:
[0112] 1. Through statistical criteria, select protein indicators that can reflect the individual muscle aging state, and apply the elastic net regression algorithm to construct a biological age assessment model in the JASHA cohort population.
[0113] This invention collected the baseline data (from 2022 to 2023) of the JASHA cohort. JASHA is a prospective, longitudinal cohort study on multi-dimensional aging phenotypes (including physical function, cognitive function, brain health, and mental health) in the general population of Zhejiang Province, aiming to explore the influencing factors of age-related health outcomes in the Chinese population across time and space. The samples of this study were sourced from the Health Management Center of Dongyang People's Hospital, the largest and most comprehensive physical examination center in Dongyang City. The study subjects were the people who underwent annual routine physical examinations at this center. Abundant biological samples (such as blood, urine, etc.) were collected, and information on demographics, socioeconomic status, lifestyle, medical history, and mental health was obtained by trained investigators through face-to-face questionnaires. Participants were recruited starting from June 2022, and finally, a total of 876 participants agreed to participate in the study and provided written informed consent. The research protocol of this study was approved by the Ethics Committee of the School of Public Health, Zhejiang University.
[0114] In this example, we selected 79 individuals in the JASHA cohort with complete body composition measurement data and proteomic data, aged from 45.2 to 83.5 years. Body composition data was detected by a high-end medical body composition analyzer (model: InBody 770); blood proteomics was detected by DIA mass spectrometry technology. After obtaining the proteomic data, the following data cleaning and quality control steps were carried out: Proteins missing in more than 40% of the samples were excluded; the KNN imputation method was used to fill the remaining missing values; the natural logarithm transformation was performed on the protein expression levels to make them more conform to the normal distribution; the Z-score normalization was performed on the protein expression levels to eliminate the influence of dimensions. After the above preprocessing steps, 1774 proteins were finally retained.
[0115] The MA of an individual was calculated through a biological age model constructed based on body composition indicators (using the biological age evaluation model constructed in patent document CN 117711623A). As Figure 2 shown, MA was significantly correlated with chronological age (correlation coefficient r = 0.71, P < 0.05); the range of MA was from 32.1 to 75.7 years (median 52.1 years, interquartile range 10.3 years).
[0116] Spearman partial correlation analysis was performed between each protein indicator and body composition age, and the effects of gender, drinking status, smoking status, education level, income level, physical activity level, and marital status were adjusted in the model (specifically: gender (1 = male; 2 = female), drinking status (1 = still drinking; 2 = quit drinking; 3 = never drink), smoking status (1 = never smoke; 2 = still smoking; 3 = quit smoking), education level (1 = primary school or below; 2 = junior high school, high school, technical secondary school; 3 = junior college or above), income level (1 = less than 10,000 yuan; 2 = 10,000 yuan and above), physical activity level (0 = medium and high activity levels; 1 = low activity level), marital status (0 = unmarried; 1 = married) and the number of diseases (the number of the following diseases currently suffered: diabetes, hypertension, heart disease, cancer, stroke, gout, lung disease, liver disease, kidney disease, and stomach disease)), and the correlation P - values between each protein and body composition age were obtained. The results showed that 78 proteins were significantly correlated with body composition age (P < 0.05, Table 1), and these proteins were selected as the candidate protein combination M for constructing the biological age assessment model.
[0117] Table 1. 78 candidate protein indicators for constructing the biological age model
[0118]
[0119]
[0120]
[0121]
[0122] Note: The protein indicators were all subjected to natural logarithm transformation and Z - score standardization, and the model adjusted for gender, drinking status, smoking status, education level, income level, physical activity level, marital status, and the number of diseases.
[0123] The biological age model ProteomicMA (MA as the dependent variable) was constructed based on the above - mentioned candidate protein combination using the elastic net regression algorithm. The automatic parameter tuning process was implemented using the R language software package "caret", and the parameter settings were: alpha = 0.5, the link function was "Gaussian", and the penalty parameter lambda was determined by the method of 10 - fold cross - validation. The lambda value that minimized the root mean square error was selected, and the final biological age model parameter values were alpha = 0.5, lambda = 1.6. The final model included 21 proteins, and the model coefficients are shown in Table 2. The model calculation formula is:
[0124] ProteomicMA
[0125] = [ln(O75915) - 3.744371] * 0.999284
[0126] + [ln(P00387) - 3.910639] * 0.121916
[0127] + [ln(P01772) - 7.429088] * -0.7755
[0128] + [ln(P05089) - 4.153092] * -0.985205
[0129] + [ln(P05387) - 3.179462] * 1.021923
[0130] + [ln(P08575) - 3.784147] * -2.114094
[0131] + [ln(P12883) - 4.816762] * -0.060567
[0132] + [ln(P15170) - 2.358900] * -1.183973
[0133] + [ln(P15880) - 2.694445] * 0.062588
[0134] + [ln(P23229) - 4.548532] * 0.891459
[0135] + [ln(P28906) - 4.880842] * 0.671088
[0136] + [ln(P50148) - 4.106702] * -1.231852
[0137] + [ln(P53041) - 2.730823] * 1.545147
[0138] + [ln(P61981) - 4.036439] * -0.083734
[0139] + [ln(P62241) - 4.054173] * -0.315841
[0140] + [ln(Q07507) - 5.315179] * 0.097354
[0141] + [ln(Q14677) - 3.687083] * 1.163949
[0142] +[ln(Q16762) - 3.254478] * -0.003892
[0143] +[ln(Q92765) - 4.763237] * -0.189882
[0144] +[ln(Q96QR1) - 3.884713] * -0.6148500 + [ln(Q9H4G4)
[0145] - 4.626910] * 3.118311 + 52.453141
[0146] Where: ProteomicMA is the biological age; O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, Q9H4G4 are the key protein indicators retained by the model.
[0147] Table 2. Coefficients of the ProteomicMA Model
[0148]
[0149]
[0150] The range of ProteomicMA is from 42.0 to 64.2 years old (median 52.6 years old, interquartile range 5.5 years old). ProteomicMA has a strong correlation with both MA and chronological age ( Figure 3 ), and the correlation coefficients are 0.81 (P < 0.05) and 0.52 (P < 0.05) respectively. To eliminate the influence of chronological age, we regressed the calculated ProteomicMA and MA on chronological age to obtain the residuals, which are ProteomicMA_ACC and MA_ACC, and the two are significantly correlated ( Figure 4 ).
[0151] 2. Analyze the associations between the biological age evaluation index ProteomicMA constructed above and the number of individual diseases, grip strength, and walking speed; compare the differences in grip strength and walking speed between the matched and unmatched groups of MA_ACC and ProteomicMA_ACC
[0152] The JASHA cohort collected information on the prevalence of major diseases such as hypertension, diabetes, stroke, and cancer in participants at baseline. By quantifying and accumulating the information on these diseases, we constructed an index reflecting the co - morbidity status of individuals - the number of diseases, where a higher value indicates that the individual has more diseases. Further, based on this disease number index, we defined a binary morbidity variable, that is, whether the individual has one or more diseases. To evaluate the physical function status of individuals, the study used walking speed measurement as an important indicator. The measurement method was to use a stopwatch to record the time required for participants to complete a 4 - meter walk at a normal speed, and calculate the walking speed accordingly. In addition, the study measured the grip strength of participants using a spring dynamometer. The measurement method was to stand upright with feet naturally separated, bend the arm at 90 degrees, hold the dynamometer with the hand and squeeze it as hard as possible, and the number displayed on the dynamometer was the grip strength value. Each hand was measured twice, and the maximum value of the two measurements was used as the grip strength value of that arm. The average value of the maximum grip strengths of both hands was used as the grip strength value of the individual.
[0153] By constructing a Logistic regression model (relationship between ProteomicMA and morbidity) or a linear regression model (relationship between ProteomicMA and grip strength and walking speed), we explored the relationships between the biological age evaluation index and the number of diseases, grip strength, and walking speed. As shown in Table 3, ProteomicMA was positively associated with the number of diseases and negatively associated with grip strength and walking speed, suggesting that ProteomicMA can effectively reflect the physiological and muscle aging status of individuals.
[0154] Table 3. Association analysis between ProteomicMA and morbidity, grip strength, and walking speed
[0155]
[0156] Note: The association between ProteomicMA and morbidity was calculated by a Logistic regression model; the associations between ProteomicMA and grip strength and walking speed were calculated by a linear regression model. The above models were adjusted for chronological age and gender.
[0157] Based on the differences between biological age and actual age (MA_ACC and ProteomicMA_ACC), we divided the study subjects into two categories: "accelerated aging" (MA_ACC > 0 or ProteomicMA_ACC > 0) and "decelerated aging" (MA_ACC < 0 or ProteomicMA_ACC < 0). Further, according to the consistency of MA_ACC and ProteomicMA_ACC, the study subjects were divided into a matched group and a non - matched group. As Figure 5As shown, in the non-matching group where MA_ACC indicates accelerated aging while ProteomicMA_ACC indicates decelerated aging, individuals showed a faster walking speed compared to those in whom both MA_ACC and ProteomicMA_ACC indicate accelerated aging; in the non-matching group where MA_ACC indicates decelerated aging while ProteomicMA_ACC indicates accelerated aging, individuals showed a smaller grip strength and a slower walking speed compared to those in whom both MA_ACC and ProteomicMA_ACC indicate decelerated aging. The above findings suggest that, compared to the biological age constructed based on body composition data, ProteomicMA, as an evaluation index of biological age, can more accurately capture an individual's health status and its potential trend of muscle function aging.
[0158] In summary, the present invention proposes a "two-step" method for constructing a biological age evaluation model by combining body composition and proteomic data, which is particularly applicable to the Chinese population. Based on the JASHA cohort study, 78 proteins related to body composition aging were screened, and the ProteomicMA model was constructed through the elastic net regression algorithm. This model can effectively evaluate an individual's muscle aging status, is significantly correlated with health indicators such as the number of diseases, grip strength, and walking speed, and has higher accuracy and sensitivity than traditional models. The present invention simplifies the calculation process of biological age, facilitating individuals to evaluate their own aging status without complex detections. This model has important application values in aspects such as early identification of muscle function decline and evaluation of the intervention effects of geriatric diseases, provides a scientific basis for the formulation of personalized anti-aging strategies, and helps to promote the development of aging prevention measures.
[0159] The above are only the preferred embodiments of one or more embodiments of this specification, and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.
Claims
1. A biological age assessment method, characterized in that: include: Obtain key protein indicators of the person to be evaluated, input them into the improved biological age evaluation model I, and obtain the biological age of the person to be evaluated; The improved biological age assessment model I is constructed based on an elastic network regression algorithm using individual proteome data and body composition age as training data.
2. The biological age assessment method according to claim 1, characterized in that: The body composition age is obtained by a biological age assessment model II constructed based on body composition data.
3. The biological age assessment method according to claim 2, characterized in that: The body composition data is obtained by muscle fat analysis, body composition analysis, obesity analysis, muscle balance analysis, segment water analysis, segment extracellular water ratio analysis, extracellular water ratio analysis, segment fat analysis, and body measurement information.
4. The biological age assessment method according to claim 1, characterized in that: The improved biological age assessment model I is constructed as follows: (2-1) Collect human proteome data and perform preprocessing; (2-2) Using Spearman partial correlation analysis to screen candidate protein indicators; (2-3) Using candidate protein indicators and body composition age as the dependent variable, the biological age assessment model I is constructed using the elastic network regression algorithm.
5. The biological age assessment method according to claim 4, characterized in that: The pre-processing comprises: (a) Protein markers missing in more than 40% of samples were eliminated; (b) Use k-nearest neighbor (KNN) interpolation method to fill missing values in the remaining data; (c) Natural logarithm transformation of protein indexes was performed to make them more consistent with normal distribution; (d) Z-score normalization of protein indicators.
6. The biological age assessment method according to claim 4, characterized in that: When applying Spearman partial correlation analysis to screen candidate protein indicators, Spearman partial correlation analysis was performed on each protein indicator and body composition age. The control variables of the Spearman partial correlation analysis were gender, drinking status, smoking status, education level, income level, physical activity level, marital status and number of diseases. The correlation P value between each protein indicator and body composition age was obtained, and protein indicators with correlation P values less than the target value were selected as candidate protein combinations for constructing the biological age assessment model.
7. The biological age assessment method according to claim 1 or 4, characterized in that: When constructing the improved biological age assessment model I based on the elastic network regression algorithm, the Caret software package is used to adjust parameters.
8. The biological age assessment method according to claim 1 or 4, characterized in that: The key protein indicators contained in the improved biological age assessment model I include O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, and Q9H4G4.
9. The biological age assessment method according to claim 1 or 4, characterized in that: The structure of the improved biological age assessment model I is as follows: ProteomicMA =[ln(O75915)-3.744371]*0.999284 +[ln(P00387)-3.910639]*0.121916 +[ln(P01772)-7.429088]*-0.7755 +[ln(P05089)-4.153092]*-0.985205 +[ln(P05387)-3.179462]*1.021923 +[ln(P08575)-3.784147]*-2.114094 +[ln(P12883)-4.816762]*-0.060567 +[ln(P15170)-2.358900]*-1.183973 +[ln(P15880)-2.694445]*0.062588 +[ln(P23229)-4.548532]*0.891459 +[ln(P28906)-4.880842]*0.671088 +[ln(P50148)-4.106702]*-1.231852 +[ln(P53041)-2.730823]*1.545147 +[ln(P61981)-4.036439]*-0.083734 +[ln(P62241)-4.054173]*-0.315841 +[ln(Q07507)-5.315179]*0.097354 +[ln(Q14677)-3.687083]*1.163949 +[ln(Q16762)-3.254478]*-0.003892 +[ln(Q92765)-4.763237]*-0.189882 +[ln(Q96QR1)-3.884713]*-0.6148500+[ln(Q9H4G4) -4.626910]*3.118311+52.453141 Among them: ProteomicMA is the biological age; O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, and Q9H4G4 are the key protein indicators retained by the model.
10. A method for constructing a biological age assessment model I based on body composition and proteomic data, characterized in that: include: (1) Calculate the individual's body composition age using the biological age model II constructed based on body composition data; (2) Proteomic data of individuals are collected, and based on the body composition age, an elastic network regression algorithm is used to construct a biological age assessment model I.
Citation Information
Patent Citations
Noninvasive rapid human body biological age prediction model, method and detection system
CN117711623A
Protein marker for lung cancer detection and lung cancer prediction method
CN118584117A
Age estimation method based on proteome data
CN119324052A
Age estimation method based on multi-dimensional data
CN119339933A
Phenotypic age and DNA methylation based biomarkers for life expectancy and morbidity
US20200347461A1
Cited By
Aging prediction method and device based on age enhancement, equipment and storage medium
CN120998505A
Disease risk prediction method and device based on plasma proteomics and waistline
CN122177477A