Method for constructing a biological age evaluation model based on a two-step method of body composition and proteomic data and application thereof

By combining body composition and proteomics data in a "two-step" approach, the ProteomicMA model was constructed, which solved the problem of insufficient applicability of existing biological age models to the Chinese population. This enabled accurate assessment and early intervention of muscle aging status, and improved the precision of health management.

CN120164612BActive Publication Date: 2025-12-30ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510098634.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-12-30
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing biological age models have limitations in capturing the aging state of the body and muscles, especially in their insufficient applicability to the Chinese population. They are also easily affected by short-term changes in nutrition and diet, and cannot accurately reflect the biological changes inside the body.

Method used

A two-step approach based on body composition and proteomics data was adopted to construct a biological age assessment model. First, body composition age was calculated using body composition data. Then, the elastic network regression algorithm was used to screen and construct proteomics data. Combined with Spearman partial correlation analysis, key protein indicators were selected to construct the ProteomicMA model.

Benefits of technology

This model can more accurately assess an individual's muscle aging status, and is significantly correlated with the number of diseases, grip strength, and walking speed. It can identify high-risk individuals for muscle function decline in the early stages, provide personalized anti-aging solutions, and improve the accuracy of health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_4
    Figure SMS_4
Patent Text Reader

Abstract

The application discloses a biological age evaluation method, comprising the following steps: obtaining a key protein index of a person to be evaluated, inputting an improved biological age evaluation model I, and obtaining the biological age of the person to be evaluated; the improved biological age evaluation model I is obtained based on an elastic network regression algorithm by taking proteomic data and body composition age of an individual as training data. The application also discloses a construction method of the model, i.e., a two-step method. The application allows a user to calculate the biological age according to the proteomic data, realizes personalized monitoring of muscle aging state, and has great guiding significance in identifying individuals with early muscle function decline and formulating personalized and accurate anti-aging schemes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention is applicable to the field of human biological age assessment technology, specifically relating to a method and application for constructing a biological age assessment model based on body composition and proteome data, and is suitable for the Chinese population. Background Technology

[0002] Muscle function decline is a typical change of aging, increasing health risks such as falls, disability, and cognitive impairment in older adults. Accurately assessing the muscle mass and function of older adults can help improve their intrinsic ability to maintain motor function and achieve proactive health.

[0003] Chronological age (CA) is considered an important parameter for assessing an individual's aging process and predicting mortality risk. However, the aging process exhibits significant heterogeneity among individuals, and a single chronological age cannot fully capture the biological differences in aging among individuals, which are crucial for a deeper understanding of the nature of aging and its impact on health. Correspondingly, biological age (BA), through various biomarkers, can more accurately assess an individual's aging process and has proven to have great application potential in the accurate prediction of life expectancy, disease risk, and disease outcomes. Among them, body composition age based on muscle-related body composition data can accurately reflect the decline in muscle function with age. Fermín-Martínez et al. used the US NHANES database to establish "AnthropoAge" predictive models for male and female biological ages, based on waist-to-height ratio (WHtR), arm circumference, and thigh circumference for men, and weight, WHtR, thigh circumference, subscapularis, and triceps skin folds for women. They also developed a simplified prediction model, "S-AnthropoAge," based on BMI and WHtR (Aging Cell. 2023 Jan; 22(1):e13756.). Previously, we established a biological age prediction model based on the body composition data of the Chinese population, which can accurately assess an individual's true muscle health status. The above biological age assessment models based on body composition data are characterized by being rapid and non-invasive, but they are easily affected by many factors such as short-term changes in individual nutrition and diet, lack stability, and cannot reflect the internal biological changes within the body.

[0004] Skeletal muscle is mainly composed of structural and contractile proteins, and protein integrity directly affects muscle strength, mass, and function. Muscle homeostasis is maintained through a balance between protein synthesis and degradation signals; when muscle protein degradation exceeds synthesis capacity, muscle decline occurs. Ceereena Ubaida-Mohien et al. used tandem mass spectrometry-based labeling (TMT) protein quantification and SOMAscan proteomics to analyze plasma and skeletal muscle proteins from the Baltimore Longitudinal Study of Aging and the Genetic Transitional Aging Laboratory Test (GESTALT). They applied an epidemiological model to the data and found that mitochondrial proteins in skeletal muscle decrease with age, while spliceosome complex proteins tend to increase (Methods Mol Biol. 2022; 2399:173-192.). However, current proteomics studies related to muscle function primarily focus on European and American populations, and the degree of muscle aging varies significantly among different ethnic groups, making these findings inapplicable to the Chinese population. Furthermore, current research mainly focuses on the elderly, thus the results lack accuracy when applied to younger individuals.

[0005] Against this backdrop, we have innovatively constructed a biological age clock using a two-step method based on the composition and proteomics data of the Chinese adult population. This method accurately assesses an individual's true muscle health status and is beneficial for stratified management and preventive healthcare of high-risk groups of muscle loss in early life (such as youth and middle age). This is of great significance for identifying individuals with accelerated aging and establishing personalized and precise anti-aging programs. Summary of the Invention

[0006] This invention aims to overcome the limitations of existing biological age models in capturing the aging state of the body and muscles. It proposes a two-step method for constructing a biological age assessment model for the Chinese population based on body composition and proteome data. The biological age assessed using this model is not only significantly correlated with chronological age, but also, after excluding the influence of chronological age, shows a significant correlation with the number of individual diseases. This invention also provides a biological age assessment method applicable to the Chinese population. This method helps users scientifically assess their own muscle aging status, thereby enabling timely and effective preventative and intervention measures to slow the aging process and reduce the risk of age-related diseases.

[0007] A biological age assessment method involves obtaining key protein indicators of the person to be assessed, inputting them into an improved biological age assessment model I, and obtaining the biological age of the person to be assessed; the improved biological age assessment model I is constructed based on an elastic network regression algorithm using the individual's proteomic data and body composition age as training data.

[0008] Furthermore, the body composition age is obtained by a biological age assessment model II constructed based on body composition data.

[0009] Furthermore, the body composition data is obtained from muscle-fat analysis, body composition analysis, obesity analysis, muscle balance analysis, segmental water content analysis, segmental extracellular water ratio analysis, extracellular water ratio analysis, segmental fat analysis, and body measurement information.

[0010] Furthermore, the biological age assessment model II constructed based on body composition data has the model structure described in publication number CN117711623A. Specifically:

[0011] The biological age prediction model for women is as follows:

[0012] Biological age

[0013] =(ln(BMI)-3.101961)*0.274298

[0014] +(BCWtoTBW-0.383087)*23.957711

[0015] -(ln(BCWtoTBW_LA)+0.971392)*22.645593

[0016] +(BCWtoTBW_LL-0.384664)*30.010663

[0017] +(ln(ECWtoTBW_RA)+0.971933)*11.754412

[0018] +(ECWtoTBW_RL-0.38359)*15.232615

[0019] +(ECWtoTBW_TR-0.383359)*13.887273

[0020] +(ln(FFM_LA_p)-4.559868)*0.486808

[0021] +(ln(FFM_RA_p)-4.586863)*2.527928

[0022] -(FFM_RL_p-96.404821)*0.031468

[0023] +(ln(FFMI)-2.733266)*0.663589

[0024] -(ln(Height)-5.068349)*3.15674

[0025] +(ln(Obesitydegree)-4.662253)*0.265601

[0026] -(FFM_LL_p-96.148192)*0.014774

[0027] +(ln(TBWtoFFM)-4.296985)*22.027408-(ln(VFL)

[0028] -1.975382)*0.224676

[0029] Where: BMI is Body Mass Index, ECWtoTBW is Extracellular Water Ratio, ECWtoTBW_LA is Extracellular Water Ratio of Left Upper Limb, ECWtoTBW_LL is Extracellular Water Ratio of Left Lower Limb, ECWtoTBW_RA is Extracellular Water Ratio of Right Upper Limb, ECWtoTBW_RL is Extracellular Water Ratio of Right Lower Limb, ECWtoTBW_TR is Extracellular Water Ratio of Left Upper Limb, FFM_LA_p is Fat-Free Body Mass (%) of Left Upper Limb, FFM_RA_p is Fat-Free Body Mass (%) of Right Upper Limb, FFM_RL_p is Fat-Free Body Mass (%) of Right Lower Limb, FFMI is Fat-Free Body Mass Index, Height is Height, Obesity_degree is Obesity Degree, FFM_LL_p is Fat-Free Body Mass (%) of Left Lower Limb, TBWtoFFM is Total Body Water / Fat-Free Body Mass, and VFL is Visceral Fat Grade.

[0030] The biological age prediction model for men is as follows:

[0031] Biological age

[0032] =(Circ LT -52.119365)*0.000078

[0033] +(Circ_RA-32.108621)*0.031951

[0034] +(ln(ECWtoTBW_TR)+0.974133)*36.644066

[0035] -(FFM_LA_p-95.154181)*0.008383

[0036] -(ln(FFM_RL_p)-4.569205)*1.207248

[0037] -(FFM_TR_p-98.489964)*0.005927

[0038] -(ln(Height)-5.134496)*6.316220

[0039] +(ln(Inbodyscore)-4.244677)*0.917546

[0040] -(ln(FFM_LL_p)-4.562414)*0.601829+(SMI

[0041] -7.828159)*0.212899+(ln(TBWtoFFM)-4.298732)

[0042] *24.649412

[0043] Where: Circ_LT is the circumference of the left thigh, Circ_RA is the circumference of the right arm, ECWtoTBW_TR is the extracellular water ratio of the left upper limb, FFM_LA_p is the lean body mass (%) of the left upper limb, FFM_RL_p is the lean body mass (%) of the right lower limb, FFM_TR_p is the relative percentage of trunk muscle mass compared to the standard muscle mass value of actual body weight, Height is the height, Inbody score is the InBody score, FFM_LL_p is the lean body mass (%) of the left lower limb, SMI is the skeletal muscle index, and TBWtoFFM is the total body water / lean body mass.

[0044] Another objective of this invention is achieved through the following technical solution: a two-step method for constructing a biological age assessment model based on body composition and proteome data, comprising the following steps.

[0045] (1) The body composition age of an individual is calculated by using a biological age model based on body composition data, and is denoted as Muscle Age (MA).

[0046] (2) Collect individual proteome data and construct a biological age assessment model based on the body composition age using an elastic network regression algorithm.

[0047] The biological age model constructed based on body composition data in step (1) was constructed by gender from more than 800 Chinese young, middle-aged and elderly individuals, with an age range of 18-87 years.

[0048] Furthermore, the improved biological age assessment model I is constructed as follows:

[0049] (2-1) Collect human proteomic indicators and perform preprocessing;

[0050] (2-2) Spearman partial correlation analysis was used to screen candidate protein indicators;

[0051] (2-3) Using candidate protein indicators, with body composition age as the dependent variable, the biological age assessment model I is constructed using the elastic network regression algorithm.

[0052] In step (2-1), health checkup data and blood samples are collected from healthy individuals undergoing physical examinations. Data-independent acquisition (DIA) mass spectrometry is used to determine the proteomic data in the blood samples. The proteomic data is then cleaned, quality controlled, and standardized.

[0053] In step (2-2), Spearman partial correlation analysis is used to screen candidate protein indicators for constructing a biological age assessment model;

[0054] In steps (2-3), candidate protein indicators are used, with body composition age as the dependent variable, and elastic network regression algorithm is used for further feature screening to construct a biological age assessment model, denoted as Proteomics-inferred MuscleAge (ProteomicMA).

[0055] Once the model is built, the effectiveness of the constructed biological age assessment model can be further evaluated.

[0056] Furthermore, the preprocessing includes:

[0057] (a) Remove protein markers that are missing in more than 40% of the samples;

[0058] (b) Use k-nearest neighbor (KNN) imputation to fill in the missing values ​​in the remaining data;

[0059] (c) Perform a natural logarithmic transformation on protein indicators to make them more consistent with a normal distribution;

[0060] (d) Standardize the protein metrics using Z-score.

[0061] Furthermore, when using Spearman partial correlation analysis to screen candidate protein indicators, each protein indicator is subjected to Spearman partial correlation analysis with body composition age. The control variables for Spearman partial correlation analysis are gender, alcohol consumption, smoking status, education level, income level, physical activity level, marital status, and number of diseases. The correlation p-value between each protein indicator and body composition age is obtained, and protein indicators with correlation p-values ​​less than the target value (e.g., p-value less than 0.05) are selected as candidate protein combinations M for constructing a biological age assessment model.

[0062] More specifically, the control variables are as follows: gender (male; female), alcohol consumption status (currently drinking; quit drinking; never drinking), smoking status (currently smoking; quit smoking; never smoking), education level (primary school and below; junior high school, high school, vocational school; junior college and above), income level (below 10,000 yuan; above 10,000 yuan), physical activity level (low activity level; moderate to high activity level), marital status (married; unmarried), and number of diseases (currently suffering from the following diseases: diabetes, hypertension, heart disease, cancer, stroke, gout, lung disease, liver disease, kidney disease, and stomach disease).

[0063] A key technical highlight of this invention is the selection of representative biomarkers for muscle aging from a large number of protein indicators. This invention applies statistical methods to screen protein indicators, fully considering the correlation between protein indicators and body composition age, to maximize the selection of protein indicator combinations that are representative of muscle aging status. As one implementation scheme, firstly, an individual's body composition age is calculated using a biological age model constructed based on existing body composition data. Then, Spearman partial correlation analysis is performed on each protein indicator and body composition age. The model adjusts for the effects of gender, alcohol consumption, smoking, education level, income level, physical activity level, and marital status to obtain the correlation p-value between each protein and body composition age. Proteins with correlation p-values ​​less than 0.05 are selected as candidate protein combinations M for constructing the biological age assessment model.

[0064] In the implementation plan, the selected candidate protein combination M includes the following 78 proteins: A0A0C4DH35, A0A0J9YXX1, O00300, O00429, O14791, O75915, O75955, P00387, P01743, P01772, P01861, P04632, P05089, P05107, P05387, P0 5556, P07585, P08575, P11413, P12236, P12814, P12883, P13667, P15170, P15880P20340 , P23229, P24941, P26927, P28906, P30040, P32121, P36871, P36959, P40121, P43487, P5 0148, P51888, P53041, P53621, P60842, P60953, P61981, P62241, P62942, P68363, Q0280 9. Q07507, Q08431, Q13177, Q13586, Q14165, Q14677, Q16674, Q16762, Q641Q3, Q6Q788, Q The protein levels listed are: 6WN34, Q86YW5, Q8IWY4, Q8IZP0, Q8N5C6, Q8N699, Q92520, Q92765, Q96AX2, Q96P63, Q96QR1, Q99719, Q99988, Q9BUN1, Q9BXJ1, Q9BZE9, Q9H0B8, Q9H4G4, Q9P270, Q9UGI8, and Q9UIB8. These protein indicators are significantly correlated with body composition age and can effectively reflect an individual's muscle aging status.

[0065] Another challenge this invention faces is the selection of the appropriate algorithm. Various algorithms have been used by researchers to construct biological age assessment models based on biomarkers, such as multiple linear regression, Lasso regression, Ridge regression, Elastic Net Regression, and machine learning algorithms. Elastic Net Regression combines the advantages of Lasso and Ridge regression, introducing both L1 and L2 norm regularization terms into the loss function. This combination allows Elastic Net Regression to flexibly address multicollinearity and feature redundancy issues, exhibiting strong stability and robustness to noise on complex datasets. Therefore, this invention ultimately chooses the Elastic Net Regression algorithm to construct the biological age assessment model. The Elastic Net Regression algorithm was proposed by Hui Zou and Trevor Hastie in 2005; its derivation details can be found at: https: / / scikit-learn.org / stable / modules / generated / sklearn.linear_model.ElasticNet.html.

[0066] In short, the workflow of the Elastic Regression Network algorithm is as follows: First, initial weight coefficients are assigned to m protein features to form an initial prediction model (m-dimensional hyperplane), and an initial loss function value is calculated based on this. Next, the algorithm iteratively updates these weight coefficients using coordinate descent until the loss function converges to its minimum. Specifically, in each iteration, the Elastic Regression Network model fixes all other weights and optimizes only the weights of a specific feature to reduce the value of the loss function. This process continues until the loss function fully converges, at which point the weights of each protein feature reach their optimal state, meaning the constructed biological age prediction model achieves the highest prediction accuracy under the set hyperparameter conditions. The formula for the loss function of Elastic Regression Network is as follows:

[0067] ElasticNetLoss=MSE+α*[λ*L1 norm +(1-α)*0.5*λ*L2 norm ]

[0068] Here, MSE represents the mean squared error between the model's predicted and actual values, used to quantify the difference between them; λ is the regularization parameter, controlling the strength of regularization; L1norm is the L1 regularization term, the sum of the absolute values ​​of the model coefficients, which helps to sparsify feature selection; L2norm is the L2 regularization term, the sum of the squares of the model coefficients, which helps to mitigate overfitting; α is a parameter between 0 and 1, used to balance the effects of L1 and L2 regularization. Specifically, when α = 0, the model degenerates into Ridge regression; when α = 1, the model is equivalent to Lasso regression.

[0069] As one implementation scheme, this invention utilizes the R language package "caret" to automate the parameter tuning process of the resilient regression algorithm. The parameters are set as follows: alpha = 0.5, the connection function is "Gaussian", and the penalty parameter lambda is determined using 10-fold cross-validation, selecting the lambda value that minimizes the root mean square error. Finally, the resilient regression model is retrained using the optimal parameters to obtain the final biological age assessment model.

[0070] Through training, the final biological age model parameters were alpha = 0.5 and lambda = 1.6.

[0071] The final elasticity network regression model includes 21 proteins (i.e., key protein indicators), including O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, and Q9H4G4.

[0072] Furthermore, the structure of the improved biological age assessment model I is as follows:

[0073] ProteomicMA

[0074] =[ln(O75915)-3.744371]*0.999284

[0075] +[ln(P00387)-3.910639]*0.121916

[0076] +[ln(P01772)-7.429088]*-0.7755

[0077] +[ln(P05089)-4.153092]*-0.985205

[0078] +[ln(P05387)-3.179462]*1.021923

[0079] +[ln(P08575)-3.784147]*-2.114094

[0080] +[ln(P12883)-4.816762]*-0.060567

[0081] +[ln(P15170)-2.358900]*-1.183973

[0082] +[ln(P15880)-2.694445]*0.062588

[0083] +[ln(P23229)-4.548532]*0.891459

[0084] +[ln(P28906)-4.880842]*0.671088

[0085] +[ln(P50148)-4.106702]*-1.231852

[0086] +[ln(P53041)-2.730823]*1.545147

[0087] +[ln(P61981)-4.036439]*-0.083734

[0088] +[ln(P62241)-4.054173]*-0.315841

[0089] +[ln(Q07507)-5.315179]*0.097354

[0090] +[ln(Q14677)-3.687083]*1.163949

[0091] +[ln(Q16762)-3.254478]*-0.003892

[0092] +[ln(Q92765)-4.763237]*-0.189882

[0093] +[ln(Q96QR1)-3.884713]*-0.6148500+[ln(Q9H4G4)

[0094] -4.626910]*3.118311+52.453141

[0095] Wherein: ProteomicMA represents biological age; O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, and Q9H4G4 are key protein indicators retained by the model.

[0096] A method for constructing a biological age assessment model based on body composition and proteome data, wherein the biological age assessment model based on body composition and proteome data is an improved biological age assessment model I as described above, and is constructed using any of the methods described above for constructing the improved biological age assessment model I.

[0097] This invention uses two methods to evaluate the obtained biological age assessment model I, as follows:

[0098] (1) Based on the obtained biological age assessment model, calculate the biological age of an individual and analyze the correlation between the biological age and the number of diseases, grip strength and walking speed of the individual.

[0099] Furthermore, this invention calculates an individual's biological age based on the biological age assessment model, and uses linear regression and logistic regression models to analyze the correlation between biological age assessment indicators and the number of diseases, grip strength, and gait speed, recording regression coefficients and p-values. If the regression coefficient is greater than 0 and the p-value is less than 0.05, the biological age assessment indicators are considered to be significantly correlated with the number of diseases. The significant correlation between biological age assessment indicators and the number of diseases, grip strength, and gait speed confirms its predictive power and application value. Evaluation results show that the biological age obtained by the improved biological age assessment model I constructed in this invention is positively correlated with the number of diseases and negatively correlated with grip strength and gait speed, suggesting that ProteomicMA can effectively reflect an individual's physiological and muscle aging state.

[0100] (2) Using biological age acceleration to evaluate aging:

[0101] To eliminate the influence of chronological age on the aging process, this invention defines biological age acceleration (BAA) for evaluating aging. BAA is defined as the residual produced when biological age is linearly regressed against chronological age; it indicates whether an individual is physiologically older (positive) or younger (negative) than expected. BAA and ProteomicMA are linearly regressed against chronological age respectively, and the residuals are used to obtain the BAA acceleration (denoted as MA_ACC) and the ProteomicMA acceleration (denoted as ProteomicMA_ACC).

[0102] A key highlight of our technology is assessing whether protein-based biological age models can capture more information related to individual muscle aging. As one implementation method, this invention categorizes each individual based on MA_ACC and ProteomicMA_ACC (accelerated aging: MA_ACC>0 or ProteomicMA_ACC>0; decelerated aging: MA_ACC<0 or ProteomicMA_ACC<0). Based on whether MA_ACC and ProteomicMA_ACC categories match or not, each individual is further classified into a matching or non-matching group, and the differences in clinical indicators between the matching and non-matching groups are compared. If there are significant differences in clinical indicators between the matching and non-matching groups, then the biological age constructed based on body composition and proteomics data is considered to capture more information about individual muscle aging than the biological age constructed solely based on body composition data. Evaluation results show that ProteomicMA, as an indicator of biological age, can more accurately capture an individual's health status and potential muscle function aging trends compared to biological age constructed based on body composition data.

[0103] This invention discloses a method for constructing a biological age assessment model based on body composition and proteomics data, and uses this model to assess the biological age of the Chinese population. The method calculates an individual's body composition age using a biological age model constructed based on body composition indicators; then, using proteomics data and the body composition age, a biological age assessment model is constructed using an elastic network regression algorithm; finally, the aging state is evaluated by calculating the biological age acceleration. This invention allows users to calculate their body composition age based on their proteomics data, enabling personalized monitoring of muscle aging status, and has significant guiding significance in identifying individuals with premature muscle function decline and developing personalized and precise anti-aging plans. This invention relies on body composition and proteomics data from the established Zhejiang Longitudinal Study of Healthy Aging (JASHA). First, it uses a biological age model constructed based on body composition data to calculate an individual's body composition age. Then, through statistical criteria, 78 candidate protein indicators that play a role in muscle aging are screened from over two thousand protein indicators. Finally, an elastic network regression algorithm is applied to construct a biological age assessment model. Data demonstrates that the ProteomicMA model constructed in this invention is correlated with the number of individual diseases, grip strength, and walking speed, effectively reflecting an individual's muscle aging state. Furthermore, this invention compares the ability of traditional body composition age models and ProteomicMA to reflect the body's state, showing that ProteomicMA has higher accuracy and sensitivity. Users can use proteomic information to accurately calculate their biological age and assess their muscle aging state using this invention, without the need for complex special tests; the operation is simple and quick.

[0104] The biological age assessment model method provided by this invention is applicable to the Chinese population, and is particularly valuable in the early identification and prevention of premature muscle function decline and the evaluation of intervention effects for age-related diseases. This method can provide a scientific basis for developing personalized and precise anti-aging programs and research on age-related diseases, and has guiding significance for promoting the prevention of secondary and tertiary aging in the population. Attached Figure Description

[0105] Figure 1 This provides a workflow framework for constructing biological age assessment models based on body composition and proteome data.

[0106] Figure 2 The correlation between MA and time-series age is shown below: where MA is the body composition age calculated using a biological age model (existing) based on body composition data, CA is the time-series age, and R is the correlation coefficient.

[0107] Figure 3The correlation between ProteomicMA, MA, and time-series age is shown below: MA is the body composition age calculated using a biological age model (existing) built based on body composition data, ProteomicMA is a new biological age assessment index built based on MA, CA is time-series age, and the color of the heatmap represents the magnitude of the correlation between ProteomicMA, MA, and time-series age. The darker the color, the greater the correlation.

[0108] Figure 4 The correlation between ProteomicMA_ACC and MA_ACC is as follows: ProteomicMA_ACC is the residual obtained by regressing ProteomicMA on time-series age, representing the acceleration of ProteomicMA; MA_ACC is the residual obtained by regressing MA on time-series age, representing the acceleration of MA.

[0109] Figure 5 This section compares grip strength and gait speed in the ProteomicMA_ACC and MA_ACC matched and unmatched groups. The horizontal axis, decel, represents slowing aging as determined by MA_ACC, while accel represents accelerating aging as determined by MA_ACC. In the box plot, gray represents the matched group (where the results of ProteomicMA_ACC and MA_ACC are consistent), and blue represents the unmatched group (where the results of ProteomicMA_ACC and MA_ACC are inconsistent). Detailed Implementation

[0110] The present invention will be further described below with reference to the embodiments and accompanying drawings.

[0111] The specific details of the embodiment are as follows:

[0112] 1. By using statistical criteria, protein indicators that can reflect an individual's muscle aging status were selected, and a biological age assessment model was constructed in the JASHA cohort population using an elastic network regression algorithm.

[0113] This invention collected baseline data (2022-2023) from the JASHA cohort. JASHA is a prospective, longitudinal cohort study of multidimensional aging phenotypes (including physical function, cognitive function, brain health, and mental health) in the general population of Zhejiang Province, aiming to explore the influencing factors of age-related health outcomes in the Chinese population across time and space. The sample for this study came from the Health Management Center of Dongyang People's Hospital, the largest and most comprehensive physical examination center in Dongyang City. The study participants were individuals undergoing annual routine physical examinations at the center. A rich collection of biological samples (such as blood and urine) was collected, and trained investigators obtained information on demographics, socioeconomic status, lifestyle, medical history, and mental health through face-to-face questionnaires. Recruitment of participants began in June 2022, and a total of 876 participants ultimately agreed to participate in the study and provided written informed consent. This study protocol was approved by the Ethics Committee of the School of Public Health, Zhejiang University.

[0114] In this embodiment, we selected 79 individuals from the JASHA cohort with complete body composition measurement data and proteomic data, ranging in age from 45.2 to 83.5 years. Body composition data were analyzed using a high-end medical body composition analyzer (model: InBody 770); blood proteomic analysis was performed using DIA mass spectrometry. After acquiring the proteomic data, the following data cleaning and quality control steps were performed: proteins missing in more than 40% of the samples were removed; the remaining missing values ​​were filled using KNN imputation; protein expression levels were transformed using the natural logarithm to make them more consistent with a normal distribution; and protein expression levels were Z-score standardized to eliminate the influence of dimensions. After the above preprocessing steps, 1774 proteins were ultimately retained.

[0115] The individual's MA is calculated using a biological age model based on body composition indicators (using the biological age evaluation model constructed in patent document CN 117711623A). For example... Figure 2 As shown, MA was significantly associated with time-series age (correlation coefficient r = 0.71, P < 0.05); the range of MA was 32.1 to 75.7 years (median 52.1 years, interquartile range 10.3 years).

[0116] Spearman partial correlation analysis was performed on each protein indicator and body composition age. The model was adjusted for the effects of gender, alcohol consumption, smoking status, education level, income level, physical activity level, and marital status (specifically: gender (1=male; 2=female), alcohol consumption (1=currently drinking; 2=quit drinking; 3=never drinking), smoking status (1=never smoking; 2=currently smoking; 3=quit smoking), education level (1=primary school or below; 2=junior high school, high school, vocational school; 3=college or above), income level (1=less than 10,000 yuan; 2=10,000 yuan or above), physical activity level (0=moderate to high activity level; 1=low activity level), marital status (0=unmarried; 1=married), and number of diseases (current number of the following diseases: diabetes, hypertension, heart disease, cancer, stroke, gout, lung disease, liver disease, kidney disease, and stomach disease)). The p-value of the correlation between each protein and body composition age was obtained. The results showed that 78 proteins were significantly associated with body composition age (P<0.05, Table 1), and these proteins were selected as candidate protein combinations M for constructing a biological age assessment model.

[0117] Table 1. 78 candidate protein indicators for constructing biological age models

[0118]

[0119]

[0120]

[0121]

[0122] Note: Protein indicators were all subjected to natural logarithmic transformation and Z-score standardization. The model was adjusted for gender, alcohol consumption, smoking status, education level, income level, physical activity level, marital status, and number of diseases.

[0123] A biological age model, ProteomicMA (MA as the dependent variable), was constructed based on the aforementioned candidate protein combinations using the elastic network regression algorithm. The parameter tuning process was automated using the R language package "caret". The parameters were set as follows: alpha = 0.5, the connection function was "Gaussian", and the penalty parameter lambda was determined using 10-fold cross-validation. The lambda value that minimized the root mean square error was selected, resulting in the final biological age model parameters of alpha = 0.5 and lambda = 1.6. The final model includes 21 proteins, and the model coefficients are shown in Table 2. The model calculation formula is as follows:

[0124] ProteomicMA

[0125] =[ln(O75915)-3.744371]*0.999284

[0126] +[ln(P00387)-3.910639]*0.121916

[0127] +[ln(P01772)-7.429088]*-0.7755

[0128] +[ln(P05089)-4.153092]*-0.985205

[0129] +[ln(P05387)-3.179462]*1.021923

[0130] +[ln(P08575)-3.784147]*-2.114094

[0131] +[ln(P12883)-4.816762]*-0.060567

[0132] +[ln(P15170)-2.358900]*-1.183973

[0133] +[ln(P15880)-2.694445]*0.062588

[0134] +[ln(P23229)-4.548532]*0.891459

[0135] +[ln(P28906)-4.880842]*0.671088

[0136] +[ln(P50148)-4.106702]*-1.231852

[0137] +[ln(P53041)-2.730823]*1.545147

[0138] +[ln(P61981)-4.036439]*-0.083734

[0139] +[ln(P62241)-4.054173]*-0.315841

[0140] +[ln(Q07507)-5.315179]*0.097354

[0141] +[ln(Q14677)-3.687083]*1.163949

[0142] +[ln(Q16762)-3.254478]*-0.003892

[0143] +[ln(Q92765)-4.763237]*-0.189882

[0144] +[ln(Q96QR1)-3.884713]*-0.6148500+[ln(Q9H4G4)

[0145] -4.626910]*3.118311+52.453141

[0146] Wherein: ProteomicMA represents biological age; O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, and Q9H4G4 are key protein indicators retained by the model.

[0147] Table 2. Coefficients of the ProteomicMA model

[0148]

[0149]

[0150] The proteomic MA ranged from 42.0 to 64.2 years (median 52.6 years, interquartile range 5.5 years). The proteomic MA showed a strong correlation with both the MA and time-series age. Figure 3 The correlation coefficients were 0.81 (P<0.05) and 0.52 (P<0.05), respectively. To eliminate the influence of time-series age, we regressed the calculated ProteomicMA and MA against the time-series age, obtaining the residuals, namely ProteomicMA_ACC and MA_ACC, which were significantly correlated. Figure 4 ).

[0151] 2. Analyze the association between the constructed biological age assessment index ProteomicMA and the number of individual diseases, grip strength, and gait speed; compare the differences in grip strength and gait speed between the MA_ACC and ProteomicMA_ACC matched groups and the unmatched groups.

[0152] At baseline, the JASHA cohort collected data on participants' prevalence of major diseases such as hypertension, diabetes, stroke, and cancer. By quantifying and aggregating this disease information, we constructed an indicator reflecting an individual's co-occurrence of multiple diseases—the number of diseases. A higher number indicates that an individual has more diseases. Further, based on this disease number indicator, we defined a binary disease status variable: whether an individual has one or more diseases. To assess an individual's physical function, the study used gait measurement as an important indicator. The method involved using a stopwatch to record the time it took for participants to walk 4 meters at a normal pace, and calculating the walking speed accordingly. In addition, the study measured participants' grip strength using a spring grip dynamometer. The measurement method involved standing with feet naturally apart, arms bent at 90 degrees, and gripping the dynamometer with full force. The reading on the dynamometer was the grip strength value. Each hand was gripped twice, and the higher value of the two measurements was taken as the grip strength value for that arm. The average of the two maximum grip strength values ​​was taken as the individual's grip strength value.

[0153] By constructing a logistic regression model (the relationship between ProteomicMA and disease status) or a linear regression model (the relationship between ProteomicMA and grip strength and gait speed), we explored the relationship between biological age assessment indicators and the number of diseases, grip strength, and gait speed. As shown in Table 3, ProteomicMA is positively correlated with the number of diseases and negatively correlated with grip strength and gait speed, suggesting that ProteomicMA can effectively reflect an individual's physiological and muscle aging status.

[0154] Table 3. Association analysis between ProteomicMA and disease status, grip strength, and gait speed

[0155]

[0156] Note: The association between ProteomicMA and disease status was calculated using a logistic regression model; the association between ProteomicMA and grip strength and gait speed was calculated using a linear regression model. All models were adjusted for time-series age and gender.

[0157] Based on the difference between biological age and chronological age (MA_ACC and ProteomicMA_ACC), we divided the study subjects into two categories: "accelerated aging" (MA_ACC>0 or ProteomicMA_ACC>0) and "decelerated aging" (MA_ACC<0 or ProteomicMA_ACC<0). Furthermore, based on the consistency between MA_ACC and ProteomicMA_ACC, the study subjects were further divided into matched and unmatched groups. Figure 5As shown, in the unmatched group where MA_ACC indicates accelerated aging while ProteomicMA_ACC indicates decelerated aging, individuals exhibited faster gait compared to those whose MA_ACC and ProteomicMA_ACC both indicated accelerated aging. Conversely, in the unmatched group where MA_ACC indicated decelerated aging while ProteomicMA_ACC indicated accelerated aging, individuals exhibited weaker grip strength and slower gait compared to those whose MA_ACC and ProteomicMA_ACC both indicated decelerated aging. These findings suggest that, compared to biological age constructed based on body composition data, ProteomicMA, as an indicator of biological age, can more accurately capture an individual's health status and potential trends in muscle function aging.

[0158] In summary, this invention proposes a two-step method for constructing a biological age assessment model by combining body composition and proteomics data, which is particularly suitable for the Chinese population. Based on the JASHA cohort study, 78 proteins related to body composition aging were screened, and a ProteomicMA model was constructed using an elastic network regression algorithm. This model can effectively assess an individual's muscle aging status, showing significant correlations with health indicators such as disease quantity, grip strength, and gait speed, and exhibits higher accuracy and sensitivity than traditional models. This invention simplifies the calculation of biological age, facilitating individual assessment of their own aging status without the need for complex tests. This model has significant application value in the early identification of muscle function decline and the evaluation of the effectiveness of interventions for age-related diseases, providing a scientific basis for the development of personalized anti-aging strategies and contributing to the advancement of aging prevention measures.

[0159] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit the scope of one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the protection scope of one or more embodiments of this specification.

Claims

1. A method for biological age assessment, characterized in that, The application relates to a method for evaluating biological age of a person, comprising the following steps: obtaining key protein indicators of a person to be evaluated, inputting an improved biological age evaluation model I, and obtaining the biological age of the person to be evaluated; the improved biological age evaluation model I is obtained by taking proteomic data of an individual and body composition age as training data and based on an elastic network regression algorithm; the body composition age is obtained by a biological age evaluation model II constructed based on body composition data; the key protein indicators contained in the improved biological age evaluation model I include O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1 and Q9H4G4.

2. The biological age assessment method according to claim 1, characterized in that, the body composition data is obtained by muscle fat analysis, human body composition analysis, obesity analysis, muscle balance analysis, segment water analysis, segment extracellular water ratio analysis, extracellular water ratio analysis, segment fat analysis and body measurement information.

3. The biological age assessment method according to claim 1, characterized in that, The improved biological age evaluation model I is constructed by the following method: (2-1) collecting human proteomic data and performing pretreatment; (2-2) screening candidate protein indicators by applying Spearman partial correlation analysis; (2-3) using the candidate protein indicators and taking the body composition age as a dependent variable, and using an elastic network regression algorithm to construct the biological age evaluation model I.

4. The biological age assessment method according to claim 3, characterized in that, The pretreatment comprises the following steps: (a) removing protein indicators missing in more than 40% of samples; (b) filling in the missing values in the remaining data by using a k-nearest neighbor imputation method; (c) performing natural logarithmic conversion on the protein indicators to make them more conform to normal distribution; (d) performing Z-score standardization on the protein indicators.

5. The method of biological age assessment according to claim 3, characterized in that, When the candidate protein indicators are screened by applying Spearman partial correlation analysis, Spearman partial correlation analysis is performed on each protein indicator and the body composition age, the control variables of the Spearman partial correlation analysis are gender, drinking status, smoking status, education level, income level, physical activity level, marital status and disease quantity; the correlation P value of each protein indicator and the body composition age is obtained, and the protein indicators with a correlation P value less than a target value are selected as candidate protein combinations for constructing the biological age evaluation model.

6. The biological age assessment method according to claim 1 or 3, characterized in that, When the improved biological age evaluation model I is constructed based on the elastic network regression algorithm, a Caret software package is used for parameter adjustment.

7. The biological age assessment method according to claim 1 or 3, characterized in that, The structure of the improved biological age evaluation model I is as follows: ; wherein: is the biological age; O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, Q9H4G4 are the key protein indicators reserved for the model.

8. A method for constructing an improved biological age assessment model I based on body composition and proteome data, characterized in that, The application further relates to a method for evaluating biological age of a person, comprising the following steps: (1) calculating the body composition age of an individual by a biological age model II constructed based on body composition data; (2) collecting proteomic data of the individual, constructing an improved biological age evaluation model I based on the proteomic data and the body composition age by using an elastic network regression algorithm. The improved biological age assessment model I includes key protein indicators including O75915, P00387, P01772, P05089, P05387, P08575, P12883, P15170, P15880, P23229, P28906, P50148, P53041, P61981, P62241, Q07507, Q14677, Q16762, Q92765, Q96QR1, and Q9H4G4.

Citation Information

Patent Citations

  • Noninvasive rapid human body biological age prediction model, method and detection system

    CN117711623A

  • Age estimation method based on proteome data

    CN119324052A