Method for evaluating healthy life and disease risk based on phenotypic age acceleration
By calculating the phenotype age acceleration value and combining survival analysis models to evaluate health lifespan and disease risk, the problem of inability to comprehensively quantify health lifespan and disease risk in the existing technology is solved, and scientific guidance for individualized health management and disease prevention is achieved.
Patent Information
- Application Number
- CN202510877569.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
Smart Images

Figure CN120388745A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for accelerating the assessment of healthspan and disease risk based on phenotypic age, and belongs to the field of information technology. Background Art
[0002] Aging is characterized by the progressive deterioration of organismal homeostasis with increasing chronological age. Notably, the rate of biological aging exhibits significant inter-individual variability, which is closely associated with different risks of death and life expectancy. To quantify this inter-individual difference, a recent study proposed PhenoAge Acceleration (PhenoAgeAccel for short), which is defined as the residual of the linear regression of phenotypic age on chronological age, taking into account chronological age. This innovative metric is superior to traditional measures of phenotypic age in predicting an individual's relative and biological aging, and can more precisely distinguish individuals showing accelerated or decelerated aging patterns from the normal standards of their age peers. Emerging evidence suggests that PhenoAge Acceleration is an independent risk factor for morbidity and mortality in chronic respiratory diseases, cancer, and neurological diseases. However, current studies have mainly focused on single disease endpoints, ignoring the cumulative health burden across major disease categories (type 2 diabetes, respiratory diseases, cancer, dementia, cardiovascular diseases). Comprehensive quantification of healthspan, i.e., measuring the disease-free years lost due to cumulative disease burden, has become crucial because isolated disease studies cannot capture the synergistic effects of accelerated aging on the development of multimorbidity. In addition, healthspan metrics (such as disease-free survival years) provide key insights into differences in the incidence of aging-related diseases at the population level. These absolute metrics enable comparative analysis of disease incidence patterns across different degrees of PhenoAge Acceleration, thereby informing public health strategies and guiding clinical interventions for the prevention of aging-related diseases.
[0003] In view of the above problems, the present application is specifically proposed. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present application is to provide a method for accelerating the assessment of healthspan and disease risk based on phenotypic age.
[0005] To achieve the above object, the present application is implemented by the following technical solutions: A method for accelerating the assessment of healthspan and disease risk based on phenotypic age, comprising the following steps: S101: Obtain the baseline data of the target population, where the baseline data includes the actual age and multi-system clinical chemical biomarker data; S102: Calculate the phenotypic age based on the baseline data, and determine the phenotypic age acceleration value through regression analysis of the phenotypic age and the actual age; S103: Stratify the target population according to the phenotypic age acceleration value; S104: Evaluate the association between the stratification of the degree of aging and the disease-free survival years and disease risk through a preset survival analysis model; S105: Output the evaluation results, and the evaluation results characterize the differences in healthspan and changes in disease risk of the target population.
[0006] Preferably, step S102 is specifically as follows: S1021: Obtain the multi-system clinical chemical biomarker data, and perform extreme value adjustment on each biomarker data to reduce distribution skewness; S1022: Calculate the phenotypic age through a preset double Gompertz proportional hazards model combined with the actual age and the adjusted biomarker data; S1023: Use linear regression method to analyze the residuals between the phenotypic age and the actual age to obtain the phenotypic age acceleration value; S1024: Determine the degree of aging acceleration of the target population according to the magnitude of the phenotypic age acceleration value.
[0007] Preferably, step S103 is specifically as follows: S1031: Obtain the distribution range of the phenotypic age acceleration value; S1032: Divide the phenotypic age acceleration value into multiple aging degree categories through a preset stratification standard, and the aging degree categories include severe aging, mild aging and non-aging; S1033: For each aging degree category, determine the corresponding proportion of the number of people in the target population; S1034: Generate a stratification result according to the proportion of the number of people and the aging degree category, and the stratification result is used for subsequent healthspan and disease risk assessment.
[0008] Preferably, step S103 is specifically as follows: S1031: Obtain multiple stratification frameworks for the phenotypic age acceleration value, where the stratification frameworks include three-category stratification and two-category stratification; S1032: Perform multi-dimensional classification on the phenotypic age acceleration value through the stratification framework to obtain multiple aging degree stratification results; S1033: For each aging degree stratification result, analyze its dose-dependent relationship with the disease risk; S1034: Determine the final aging degree stratification scheme according to the dose-dependent relationship, and the final aging degree stratification scheme is used for subsequent survival analysis model evaluation.
[0009] Preferably, step S104 is specifically as follows: S1041: Obtain distribution data of aging degree stratification; S1042: Analyze the relationship between aging degree stratification and the risk of major diseases through a preset proportional hazard regression model to obtain the risk ratio; S1043: Use a flexible parameter survival model to calculate the difference in disease-free survival years corresponding to the aging degree stratification; S1044: Generate association analysis results based on the risk ratio and the difference in disease-free survival years. The association analysis results are used to characterize the impact of different aging degrees on healthy life expectancy.
[0010] Preferably, step S104 is specifically as follows: S1041: Obtain the lifestyle and socioeconomic factor data of the target population; S1042: Combine the lifestyle and socioeconomic factor data with the aging level stratification through stratified analysis to analyze the differences in healthy life expectancy in different subgroups; S1043: Use the multivariate adjustment method to correct the analysis results to obtain the corrected correlation data; S1044: Determine the consistency effect of the aging level stratification in different subgroups based on the corrected correlation data, and the consistency effect is used for the subsequent evaluation result output.
[0011] Preferably, step S104 is specifically as follows: S1041: Obtain follow-up data of the target population, including disease diagnosis records and death records; S1042: Analyze the relationship between follow-up data and aging level stratification through a competing risk regression model to obtain the association results after adjusting for competing risks; S1043: Use sensitivity analysis methods to verify the robustness of the association results; S1044: Update the output data of the survival analysis model based on the association results after robustness verification, and the output data will be used for subsequent healthy life expectancy difference assessment.
[0012] Preferably, step S105 is specifically as follows: S1051: Obtain the association analysis data generated by the survival analysis model; S1052: Organize the association analysis data using a preset output format, which includes the difference in disease-free survival years and the disease risk ratio; S1053: Generate corresponding healthy life expectancy prediction values for different aging levels of the target population; S1054: Output a final evaluation report based on the healthy life expectancy prediction values and the disease risk ratios. The final evaluation report is used to characterize the changes in the health status of the target population at different aging levels.
[0013] Beneficial effects of this application: This application obtains the biological data of the target population, calculates the phenotypic age acceleration value, and classifies it into different aging acceleration categories. Subsequently, a risk regression model is constructed to analyze the association between each category and disease risk, and predict the number of disease-free survival years. By extracting the difference values of healthy life expectancy between different categories and verifying their consistency in different populations, the phenotypic age acceleration value is finally determined as a reliable indicator for predicting healthy life expectancy. This application also combines a competing risk regression model to generate an evaluation report including individual aging risk and prediction of disease-free years. This method reliably evaluates the degree of biological aging, provides a scientific basis for individualized health management and disease prevention, helps to extend healthy life expectancy, and improve the quality of life. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flowchart of a method for evaluating healthy life expectancy and disease risk based on phenotypic age acceleration according to this application.
[0015] Figure 2 It is a curve graph of the increase in the number of disease-free years at 45 and 65 years old according to this application.
[0016] Figure 3 It is a K-M survival curve graph according to this application.
[0017] Figure 4 It is the life expectancy (male) at 45 years old of different subgroups of the UK Biobank cohort according to this application.
[0018] Figure 5 It is the life expectancy (female) at 45 years old of different subgroups of the UK Biobank cohort according to this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the technical means, creative features, achieved purposes and functions of this application easy to understand, the following further elaborates this application in combination with specific embodiments.
[0020] Embodiment S101: Obtain the baseline data of the target population, where the baseline data includes at least the actual age and multi-system clinical chemistry biomarker data; Specifically, the actual age data of the target population is obtained through database query to determine the record set containing the age field. According to the actual age data, multi-system clinical chemical biomarker data is extracted from the clinical database to obtain a biomarker data set containing albumin, creatinine, glucose, CRP, lymphocyte percentage, mean cell volume, and red blood cell distribution width. If there are missing values in the biomarker data, the missing values are filled by the mean imputation method to obtain a complete biomarker data set. The extreme value adjustment method is used to limit the value of each biomarker between the 1st percentile and the 99th percentile to determine the adjusted biomarker data set. The phenotypic age of each individual is calculated through the PhenoAge formula to obtain a phenotypic age data set. According to the linear regression model of phenotypic age and actual age, the age-adjusted residuals are calculated to determine the set of phenotypic age acceleration values. If the phenotypic age acceleration value is greater than 0, the first 50% of it is classified into the severe aging category, and the remaining is classified into the mild aging and non-aging categories to obtain the stratified set of phenotypic age acceleration categories. Through association analysis, the stratified phenotypic age acceleration categories are matched with the disease-free survival years data to determine the disease-free survival years statistical results for each category. The double Gompertz proportional hazards model is used to estimate the risks of the phenotypic age acceleration categories and the disease-free survival years data to obtain the association analysis results of the aging-related physiological indicators and the survival years.
[0021] Exemplarily, the actual age data of the target population is obtained through database query. For example, the actual age fields of 88,382.4 person-years of severely aged population and 93,179.6 person-years of mildly aged population are extracted from the Chinese cohort. According to the actual age data, multi-system clinical chemical biomarker data is extracted from the clinical database. For example, the specific values of albumin, creatinine, glucose, CRP, lymphocyte percentage, mean cell volume, and red blood cell distribution width are obtained. The range of albumin is 3.5 g / dL - 5.0 g / dL, and the range of creatinine is 0.6 mg / dL - 1.2 mg / dL. If there are missing values in the biomarker data, the missing values are filled by the mean imputation method. For example, the missing glucose value is replaced with the average value of 6.5 mmol / L. The extreme value adjustment method is used to limit the value of each biomarker between the 1st percentile and the 99th percentile. For example, the CRP value is limited between 0.1 mg / L and 10 mg / L. The phenotypic age of each individual is calculated through the phenotypic age formula. For example, the formula PhenoAge = 141.50 + ln(-0.00553×ln(1 - xb) / 0.09165) is used, where xb is the adjusted biomarker data, and xb = -19.907 - 0.0336×albumin + 0.0095×creatinine + 0.0195×glucose + 0.0954×ln(CRP) - 0.0120×lymphocyte percentage + 0.0268×mean cell volume + 0.3356×red blood cell distribution width + 0.00188. According to the linear regression model of phenotypic age and actual age, the residuals after age correction are calculated. For example, the distribution of phenotypic age acceleration values between -2.5 and 3.5 is obtained. If the phenotypic age acceleration value is greater than 0, the top 50% of it is classified into the severely aged category, and the remaining is classified into the mildly aged and non-aged categories. For example, individuals with a phenotypic age acceleration value greater than 1.5 are classified as severely aged. Through association analysis, the stratified phenotypic age acceleration categories are matched with the disease-free survival years data. For example, the average disease-free survival years of the severely aged group is 45 years, and that of the mildly aged group is 47.96 years. The double Gompertz proportional hazards model is used to estimate the risks of the phenotypic age acceleration categories and the disease-free survival years data.
[0022] S102: Calculate the phenotypic age based on the baseline data, and determine the phenotypic age acceleration value through the regression analysis of phenotypic age and actual age. The phenotypic age value is derived by combining the actual age and biomarker data through a preset double Gompertz proportional hazards model to obtain the phenotypic age calculation result; Specifically, baseline data of 308,592 participants were obtained, and the actual age and data of nine multi-system clinical chemical biomarkers were extracted to obtain the original biological data set. The biomarker data distribution skewness was processed by extreme value adjustment by setting each biomarker to the lowest 1st percentile and the highest 99th percentile, and the adjusted biomarker data were obtained. According to the adjusted biomarker data, an elastic net regularization Cox proportional hazards framework was used to optimize for death prediction in 10-fold cross-validation, and the regression coefficients of the nine biomarkers were determined. Through the double Gompertz proportional hazards model, combined with the actual age and the regression coefficients of the nine biomarkers, the xb of each individual was calculated, where xb was the adjusted biomarker data, and the xb calculation result was obtained. Using the formula PhenoAge = 141.50 + ln(-0.00553 × ln(1 - xb) / 0.09165), based on the xb calculation result, the phenotypic age value of each individual was calculated, and the phenotypic age calculation result was obtained. Through linear regression of the phenotypic age value and the actual age, the regression residuals were calculated to obtain the phenotypic age acceleration value. If the phenotypic age acceleration value was greater than 0 and in the top 50%, it was classified into the severe aging category; if the phenotypic age acceleration value was less than or equal to 0, it was classified into the non-aging or mild aging category, and the phenotypic age acceleration classification result was obtained. According to the phenotypic age acceleration classification result, a flexible curve model was used to estimate the number of years without major diseases (type 2 diabetes, respiratory diseases, cancer, dementia) in non-aging and mild aging individuals, and the disease-free years estimate value was obtained. Through the Cox proportional hazards regression model, the association between the phenotypic age acceleration classification result and the chronic disease incidence risk was analyzed to obtain the risk assessment result.
[0023] Exemplarily, baseline data of 308,592 participants were extracted from the UK Biobank, including actual age (such as 56.88 ± 8.19 years old) and nine biomarkers such as albumin, creatinine, glucose, C-reactive protein (CRP), percentage of lymphocytes, mean cell volume, and red blood cell distribution width, which constituted the original dataset. Extreme value adjustment was performed on each biomarker. For example, the outliers of CRP were restricted between the 1st percentile (0.1 mg / L) and the 99th percentile (10 mg / L) to eliminate data skewness. An elastic net regularized Cox proportional hazards regression model was used to optimize death prediction through 10-fold cross-validation, and the regression coefficients of each biomarker were determined (such as albumin coefficient -0.0336, creatinine coefficient 0.0095). Based on the double Gompertz proportional hazards model, xb was calculated by combining age and biomarker coefficients. For example, for an individual, xb = -19.907 - 0.0336 × 40 (albumin) + 0.0095 × 80 (creatinine). Substituting into the formula PhenoAge = 141.50 + ln(-0.00553 × ln(1 - xb) / 0.09165), if xb is 0.2, the phenotypic age is 141.50 + ln(-0.00553 × ln(0.8) / 0.09165) ≈ 65 years old. Linear regression was performed between the phenotypic age and the actual age, and the residual was the phenotypic age acceleration. For example, for a 60-year-old individual with a phenotypic age of 65 years old, the residual was +5 years. If the residual > 0 and is in the top 50% (such as +5 years > threshold +3 years), it was classified into the severe aging group; otherwise, it was classified into the non- / mild aging group. A flexible curve model was used to estimate the disease-free years. For example, the average disease-free years of the non-aging group was 20 years, and that of the mild aging group was 15 years. The Cox proportional hazards regression model was used to analyze the phenotypic age acceleration and the risk of chronic diseases. For example, the risk ratio of the severe aging group was 1.5 (95% CI 1.2 - 1.8).
[0024] S103: Stratify the target population according to the phenotypic age acceleration value; S1031: Obtain the distribution range of the phenotypic age acceleration value; Specifically, obtain the individual's actual age data and nine multi-system clinical chemical biomarker data, and determine the biomarker values such as albumin, creatinine, glucose, C-reactive protein, lymphocyte percentage, mean cell volume, and red blood cell distribution width. By setting the lowest 1st percentile and the highest 99th percentile, extreme value adjustment is performed on the biomarker data to obtain the biomarker values after distribution skewness correction. Use the phenotypic age formula PhenoAge = 141.50 + ln(-0.00553 × ln(1 - xb)) / 0.09165 to calculate the individual's phenotypic age, where xb = -19.907 - 0.0336× albumin + 0.0095 × creatinine + 0.0195 × glucose + 0.0954 × ln(CRP) - 0.0120 × lymphocyte percentage + 0.0268 × mean cell volume + 0.3356 × red blood cell distribution width + 0.00188, to obtain the phenotypic age value. According to the individual's actual age and phenotypic age, construct a linear regression model to determine the regression coefficient and intercept between the phenotypic age and the actual age. Through the linear regression model, calculate the predicted values of the phenotypic age and the actual age to obtain the predicted phenotypic age value after age correction. Subtract the predicted phenotypic age value from the phenotypic age value to calculate the regression residual and obtain the phenotypic age acceleration value. If the phenotypic age acceleration value is greater than 0, compare it with the 50th percentile to determine whether it belongs to the severe aging category; if the phenotypic age acceleration value is less than or equal to 0, classify it as non / mild aging to obtain the phenotypic age acceleration category. According to the phenotypic age acceleration value and category, combined with the double Gompertz proportional hazard model, estimate the individual's death risk to obtain the death risk prediction value. Through the phenotypic age acceleration value and the death risk prediction value, analyze their association with the number of disease-free life years to determine the quantitative impact of phenotypic age acceleration on the number of disease-free survival years.
[0025] Exemplarily, actual age data of an individual and data of nine biomarkers such as albumin, creatinine, glucose, C-reactive protein, lymphocyte percentage, mean cell volume, and red blood cell distribution width are extracted from a clinical database. For example, the albumin concentration of a certain individual is 4.2 g / dL, creatinine is 0.8 mg / dL, and glucose is 95 mg / dL. Extreme value adjustment is performed on the biomarker data. If the original value of C-reactive protein of an individual is 12 mg / L, which is higher than the 99th percentile threshold of 10 mg / L, it is truncated to 10 mg / L. Using the phenotypic age calculation formula, the adjusted biomarker data is input. For example, calculate xb = -19.907 - 0.0336×4.2 + 0.0095×0.8 + 0.0195×95 + 0.0954×ln(10) - 0.0120×25 + 0.0268×90 + 0.3356×14 + 0.00188, and get xb = -16.42. Substitute it into the phenotypic age formula to output a phenotypic age of 65 years old. A linear regression model is constructed. If the regression equation is PhenoAge = 1.2×actual age + 15, the predicted phenotypic age of a 50-year-old individual is 75 years old. Calculate the residual. If the actual phenotypic age is 80 years old, then the phenotypic age acceleration value = 80 - 75 = 5. If the quantile of the phenotypic age acceleration value > 0 in the population is 3, then this individual is classified as severely aged. Based on the double Gompertz proportional hazard model, input the phenotypic age acceleration value = 5 and clinical data, and output a 5-year death risk probability of 8%. Analyze the association between phenotypic age acceleration and the number of disease-free survival years. For example, the average number of disease-free survival years in the severely aged group is reduced by 2.3 years compared to the non / lightly aged group.
[0026] S1032: Divide the phenotypic age acceleration value into multiple aging degree categories according to a preset stratification standard. The aging degree categories at least include severely aged, lightly aged, and non-aged; Specifically, obtain the phenotypic age acceleration value dataset, and perform stratification processing on the phenotypic age acceleration values using a preset classification standard to obtain three aging acceleration categories: severe aging, mild aging, and non-aging. According to the stratification results, calculate the person-year cumulative value for each aging acceleration category, and determine the person-year data corresponding to the severe aging, mild aging, and non-aging categories. Based on the person-year data, construct a baseline feature description for each aging acceleration category, and use descriptive statistical methods to calculate the mean and standard deviation of continuous variables as well as the number and percentage of categorical variables to obtain the baseline feature statistical results. Use the chi-square test to compare categorical variables between groups, and at the same time use the Mann-Whitney-Wilcoxon test to compare continuous variables between groups to determine the statistical differences between different aging acceleration categories. According to the statistical difference results, construct a Royston-Parmar flexible parametric proportional hazards survival model, with age as the time scale, to predict the average disease-free survival period for each aging acceleration category and obtain the survival model parameters. Through the survival model parameters, calculate the survival curves at 1-year intervals between 45 and 100 years old, and use numerical integration methods to construct individual survival curves to determine the survival probabilities for each aging acceleration category. According to the survival probabilities, calculate the area under the survival curve, and use the area difference method to compare the differences in population healthy life expectancy between the mild aging group and the severe aging group to obtain the difference value of disease-free survival years. If there is statistical uncertainty in the difference value, use 1000 independent bootstrap methods to repeatedly generate 95% confidence intervals to judge the statistical robustness of the difference in disease-free survival years. Through the stratification analysis method, combined with demographic characteristics, lifestyle, and socioeconomic status data, evaluate the consistency of healthy life expectancy prediction and potential modification effects of phenotypic age acceleration values to obtain the stratification analysis results.
[0027] Exemplarily, a phenotypic age acceleration value dataset is obtained from the UK Biobank cohort and stratified using binary classification to obtain three categories: severe aging, mild aging, and non-aging. Among them, the UK Biobank cohort has accumulated 755,084.3 (severe aging), 855,722.2 (mild aging), and 2,238,140 (non-aging) person-years. According to the stratification results, the person-year cumulative values of each aging category are calculated. For example, in the Chinese cohort, the severe aging group has 883,824 person-years, the mild aging group has 931,796 person-years, and the non-aging group has 10,559,934 person-years. Baseline characteristics are described using person-year data, and descriptive statistical methods are used to calculate the mean of continuous variables (such as age 56.84 ± 18.45 years) and the percentage of categorical variables (such as the proportion of 13,797 females). The chi-square test is used to compare categorical variables (such as smoking status), and the Mann-Whitney-Wilcoxon test is used to compare continuous variables (such as body mass index) to determine differences between groups. A Royston-Parmar flexible parametric proportional hazards survival model is constructed to predict disease-free survival with age as the time scale, and survival curves are calculated at 1-year intervals between 45 and 100 years. The survival probability is obtained through numerical integration (such as the number of disease-free survival years in the male mild aging group increases by 2.96 years). The area under the curve is calculated based on the survival probability, and the difference in healthy life expectancy between the mild aging group and the severe aging group is compared using the area difference method (such as a difference of 2.97 years in females). If the difference is uncertain, a 95% confidence interval is generated using 1000 bootstrap resamples (such as 2.45 - 3.47 years). The prediction consistency of the phenotypic age acceleration value is evaluated in combination with stratification variables (such as education level, sleep duration), and potential modifying effects are analyzed.
[0028] S1033: For each aging degree category, determine the corresponding proportion of the number of people in the target population; Specifically, baseline data of 308,592 participants were extracted from the UK Biobank database, including phenotypic age, chronological age, gender, demographic characteristics, lifestyle, and socioeconomic status variables. The phenotypic age acceleration value was calculated to obtain the phenotypic age acceleration value. According to the phenotypic age acceleration value, the participants were divided into three categories: severe aging (> severe threshold), mild aging (> 0 but below the severe threshold), and non-aging (≤ 0), resulting in a classified dataset. Disease diagnosis data during the follow-up period were obtained, and disease-free year records of those not diagnosed with the four major age-related diseases between 45 and 100 years old were screened to determine the disease-free year time period dataset. Descriptive statistical methods were used to calculate the baseline characteristics of each phenotypic age acceleration category, including the mean and standard deviation of continuous variables, and the number and percentage of categorical variables, to obtain the baseline characteristics statistical table. The between-group differences of categorical variables were compared by chi-square test, and the between-group differences of continuous variables were compared by Mann-Whitney-Wilcoxon test to determine the results of between-group statistical differences. A multivariate Cox proportional hazards regression model was constructed, with the classified phenotypic age acceleration category and disease-free year time period data input, and demographic characteristics, lifestyle, and socioeconomic status variables adjusted to obtain the hazard ratio (HR) and 95% confidence interval (CI). If the confidence interval of the hazard ratio contains 1, 1000 bootstrap replicates were calculated to generate a new 95% confidence interval to judge statistical uncertainty. Based on the hazard ratio and confidence interval output by the Cox proportional hazards regression model, combined with the Royston-Parmar flexible parametric proportional hazards survival model, the average disease-free survival period of each phenotypic age acceleration category was predicted to obtain the disease-free survival period estimate. By numerical integration method, the area under the survival curve at 1-year intervals between 45 and 100 years old was calculated, and the difference in the area under the survival curve between different phenotypic age acceleration categories was compared to obtain the population healthy life expectancy difference value.
[0029] Exemplarily, baseline data of 308,592 participants were extracted from the UK Biobank database, including variables such as phenotypic age, chronological age, gender, BMI, smoking status, sleep duration, educational level, etc. A linear regression model was used to calculate the phenotypic age acceleration value. According to the distribution of the phenotypic age acceleration value, classification thresholds for the severe aging group (>1.5 standard deviations), mild aging group (0 - 1.5 standard deviations), and non-aging group (≤0) were set, and grouping was completed using SQL query statements. Disease diagnosis data coded with ICD-10 were extracted through the medical record system, and the confirmed ages of four major diseases, namely E11 (diabetes), I25 (coronary heart disease), J18 (pneumonia), and G30 (Alzheimer's disease), were screened. The disease-free years were calculated for the interval from 45 to 100 years old (confirmed age, 100 years old). The mean ± standard deviation (such as age 56.88 ± 8.19 years) and frequency percentage (such as the proportion of males being 46.1%) were calculated for the baseline data of each group. The chi-square test was used to compare categorical variables such as the smoking rate, and the Wilcoxon rank-sum test was used to compare continuous variables such as BMI. When constructing the Cox proportional hazards regression model, the three-category phenotypic age acceleration variable was input, covariates such as gender and BMI were adjusted, and the partial likelihood estimation method was used to calculate the hazard ratio (such as the hazard ratio of the severe aging group = 2.34, 95% CI 2.12 - 2.58). The baseline hazard function output by the Cox proportional hazards regression model was input into the Royston-Parmar flexible parametric proportional hazards survival model, and the survival curve was fitted using restricted cubic splines (3 knots), and the trapezoidal integral of the annual survival probability for the interval from 45 to 100 years old was calculated (such as the area of the non-aging group = 38.7 years). Finally, the difference in the area under the curve was calculated to obtain the difference in population healthy life expectancy.
[0030] S1034: Generate a stratification result according to the proportion of people and the category of aging degree, and the stratification result is used for subsequent assessment of healthy life expectancy and disease risk; According to the risk assessment result, a flexible parametric survival model is used to predict the disease-free survival years of different aging acceleration categories within the preset age range, and the disease-free survival years are calculated through the area integral under the survival curve.
[0031] S104: Evaluate the association between the stratification of aging degree and the disease-free survival years and disease risk through a preset survival analysis model; Specifically, based on the extracted disease-free survival data, the difference in disease-free years between the mild aging group and the severe aging group was calculated to determine the healthy lifespan gain of mild aging relative to severe aging. Based on the extracted disease-free survival data, the difference in disease-free years between the non-aging group and the severe aging group was calculated to determine the healthy lifespan gain of non-aging relative to severe aging. Using the Royston-Parmar flexible parameter proportional hazards survival model, survival curves for individuals in each aging category were generated over a time scale of 45 to 100 years, and the numerical integration of survival probabilities was obtained. By numerically integrating the areas under the survival curves, the remaining disease-free years for each aging category between 45 and 100 years were estimated, and the healthy lifespan area values for the mild, non-aging, and severe aging groups were obtained. If the difference in the survival curve areas between the mild and severe aging groups was greater than 0, the difference between the two curve areas was calculated to determine the difference in healthy lifespan for the mild aging group relative to the severe aging group. If the difference in the survival curve area between the non-aging group and the severe aging group was greater than 0, the difference in the areas of the two curves was calculated to determine the difference in healthy lifespan between the non-aging group and the severe aging group. 1000 independent bootstrap replicates were used to generate 95% confidence intervals to assess the statistical uncertainty of the difference in healthy lifespan between the mildly aging group and the non-aging group relative to the severe aging group and to determine the confidence interval range. Based on the results of the stratified analysis, baseline data on demographic characteristics, lifestyle, and socioeconomic status were extracted and integrated into a multivariable Cox proportional hazards regression model to calculate the hazard ratio for the difference in healthy lifespan and to determine the predictive consistency of the stratified analysis.
[0032] Exemplarily, for instance, the number of person - years in the mild senescence group is 93,179.6, in the severe senescence group is 88,382.4, and in the non - senescence group is 1,055,993.4. According to the extracted disease - free survival person - year data, calculate the difference in disease - free person - years between the mild senescence group and the severe senescence group, and determine the healthy life - span gain value of mild senescence relative to severe senescence. For example, in the UK Biobank cohort, for men, the mild senescence group has a 2.96 - year increase compared to the severe senescence group, and for women, it has a 2.97 - year increase. According to the extracted disease - free survival person - year data, calculate the difference in disease - free person - years between the non - senescence group and the severe senescence group, and determine the healthy life - span gain value of non - senescence relative to severe senescence. For example, the non - senescence group has a 4.67 - year increase compared to the severe senescence group. Using the Royston - Parmar flexible parametric proportional hazards survival model, with the time scale from 45 to 100 years old, generate the survival curves of individuals in each senescence category, and obtain the numerical integration results of the survival probabilities. For example, predict the survival probability distribution between 45 and 100 years old through the model. By numerically integrating the area under the survival curve, estimate the remaining disease - free person - years of each senescence category between 45 and 100 years old, and obtain the healthy life - span area values of the mild senescence, non - senescence, and severe senescence groups. For example, calculate the areas under the survival curves as X for the mild senescence group, Y for the non - senescence group, and Z for the severe senescence group. If the difference in the areas under the survival curves between the mild senescence group and the severe senescence group is greater than 0, calculate the difference between the two curve areas to obtain the population healthy life - span difference value of mild senescence relative to severe senescence. For example, the area difference is A. If the difference in the areas under the survival curves between the non - senescence group and the severe senescence group is greater than 0, calculate the difference between the two curve areas to obtain the population healthy life - span difference value of non - senescence relative to severe senescence. For example, the area difference is B. Use 1000 independent bootstrap resampling methods to repeatedly generate 95% confidence intervals, evaluate the statistical uncertainty of the healthy life - span difference values of the mild senescence group and the non - senescence group relative to the severe senescence group, and determine the confidence interval range. For example, the 95% confidence interval of the difference value in the mild senescence group is 2.45 to 3.47 years. According to the stratified analysis results, extract the baseline characteristic data of demographic characteristics, lifestyle, and socioeconomic status, integrate the multivariate Cox proportional hazards regression model, calculate the hazard ratio of the healthy life - span difference value, and obtain the prediction consistency results of the stratified analysis. For example, the hazard ratio is 1.2, and the 95% confidence interval is 1.1 to 1.3.
[0033] Through the stratified analysis method, verify the consistency of the healthy life - span difference values in different demographic characteristics and lifestyle factors. The stratified analysis method is based on preset socioeconomic and behavioral variables for grouped calculations.
[0034] Specifically, baseline data of 308,592 participants in the UK Biobank were obtained, and phenotypic age acceleration values were calculated. The residuals after age correction were obtained through a linear regression model of phenotypic age and actual age. According to the phenotypic age acceleration values, participants were divided into an accelerated aging group (phenotypic age acceleration value > 0) and a non-accelerated aging group (phenotypic age acceleration value ≤ 0) using a binary classification method, and the baseline characteristics of the two groups were determined. Through a stratified analysis method, based on demographic characteristics, lifestyle, and socioeconomic status variables, the accelerated aging group and the non-accelerated aging group were further grouped to obtain stratified subpopulations. Using the Royston-Parmar flexible parametric proportional hazards survival model, with age as the time scale, the survival curves of each subpopulation between 45 and 100 years old were calculated to obtain individualized survival probabilities. Through the numerical integration of the annual survival probabilities, the survival curves of each subpopulation were constructed, and the area under the survival curve was calculated to obtain the remaining disease-free years of each subpopulation. If the calculation of the survival curve area of a certain subpopulation was completed, the area difference between the corresponding subpopulations of the accelerated aging group and the non-accelerated aging group was compared to obtain the healthy life expectancy difference value. Through 1000 independent bootstrap replicates, the statistical uncertainty of the healthy life expectancy difference value of each subpopulation was evaluated to generate a 95% confidence interval and determine the robustness of the difference value. Descriptive statistical methods were used to calculate the mean and standard deviation of continuous variables of the baseline characteristics of each subpopulation, as well as the number and percentage of categorical variables, to obtain the statistical distribution of stratified characteristics. The differences in categorical variables between stratified subpopulations were compared using the chi-square test, and the differences in continuous variables were compared using the Mann-Whitney-Wilcoxon test to judge the predictive consistency of phenotypic age acceleration in different stratified variables.
[0035] Exemplarily, baseline data of 308,592 participants were extracted from the UK Biobank, including biomarkers such as age, albumin, creatinine, glucose, C-reactive protein (CRP), percentage of lymphocytes, mean cell volume, red blood cell distribution width, etc. The phenotypic age (PhenoAge) was calculated using an elastic net regularized Cox proportional hazards framework, with the formula PhenoAge = 141.50 + ln(-0.00553×ln(1 - xb) / 0.09165), where xb is the adjusted biomarker data, and xb = -19.907 - 0.0336×albumin + 0.0095×creatinine + 0.0195×glucose + 0.0954×ln(CRP) - 0.0120×percentage of lymphocytes + 0.0268×mean cell volume + 0.3356×red blood cell distribution width + 0.00188×alkaline phosphatase. Subsequently, the phenotypic age acceleration residuals were calculated through a linear regression model. The participants were divided into an accelerated aging group (>0) and a non-accelerated aging group (≤0) according to the phenotypic age acceleration value, with the severely aging group accounting for 23.8%. Using a stratified analysis method, further grouping was performed according to variables such as body mass index (BMI≥30 or <30), smoking status (never, past, current), sleep duration (<6 hours, 6 - 8 hours, >8 hours), educational level (below high school, high school, college and above), etc. The Royston-Parmar flexible parametric proportional hazards survival model was used to calculate the survival probability starting from 45 years old at 1-year intervals until 100 years old, and the area under the survival curve was estimated through numerical integration to obtain the disease-free survival years. If the area under the survival curve of a certain subgroup (such as smokers with BMI≥30) is 25.3 years and the corresponding subgroup of the non-accelerated aging group is 32.7 years, then the healthy life expectancy difference value is 7.4 years. Using 1000 bootstrap samplings, the 95% confidence interval of the difference value was calculated as [6.8, 8.0]. The baseline characteristics of each subgroup were statistically analyzed. For example, the average age of the accelerated aging group was 58.2 years (standard deviation 8.5), and that of the non-accelerated aging group was 55.6 years (standard deviation 7.9). The distribution of smoking status was compared through a chi-square test (18.5% of current smokers in the accelerated aging group, 12.3% in the non-accelerated aging group, p<0.001), and the Mann-Whitney-Wilcoxon test was used to compare the BMI differences to analyze the prediction consistency of phenotypic age acceleration.
[0036] If the healthy life expectancy difference value maintains a consistent direction and magnitude in each stratified group, then the phenotypic age acceleration value is determined as a reliable indicator for predicting healthy life expectancy, and the output result of the prediction model is generated.
[0037] S105: Output the evaluation result, where the evaluation result characterizes the healthy life expectancy difference and disease risk change of the target population; Specifically, baseline data of 308,592 participants in the UK Biobank were obtained. Phenotypic age acceleration was calculated by performing linear regression on phenotypic age and chronological age to obtain the phenotypic age acceleration value. A double Gompertz proportional hazards model was used, combined with chronological age and nine multi-system clinical chemistry biomarkers, to calculate the phenotypic age and determine the phenotypic age value for each participant. Through an elastic net regularized Cox proportional hazards framework, the biomarkers were optimized by 10-fold cross-validation, and the biomarker extreme values were adjusted to the 1st and 99th percentiles to obtain the calculation result of the standardized phenotypic age formula. According to the phenotypic age acceleration value, the participants were divided into three categories: severe aging (the top 50% with phenotypic age acceleration value > 0), mild aging, and non-aging. The baseline characteristics between groups were compared by chi-square test and Mann-Whitney-Wilcoxon test to determine the statistical significance of the classification. The Royston-Parmar flexible parametric proportional hazards survival model was used, with age as the time scale, to construct an individual survival curve and obtain the numerical integration result of the annual survival probability between 45 and 100 years old. By calculating the area under the survival curve, the remaining disease-free years were estimated, and the disease diagnosis and death endpoints were adjusted in combination with the competing risks regression model to determine the individual disease-free years prediction value. If there is statistical uncertainty in the disease-free years prediction value after adjustment by the competing risks regression model, a 95% confidence interval was repeatedly generated through 1000 independent bootstrap methods to obtain a robust estimate of the disease-free years. According to the stratified analysis, combined with demographic characteristics, lifestyle, and socioeconomic status, the association between phenotypic age acceleration and disease-free years was evaluated to determine the potential modifying effects in the biological and behavioral fields. Through a multivariable Cox proportional hazards regression model, the hazard ratio and 95% confidence interval were estimated to generate the final biological aging assessment report and obtain the individual aging risk and disease-free years prediction value.
[0038] Exemplarily, baseline data of 308,592 participants were extracted from the UK Biobank, including biomarkers such as age, albumin (mean 42.3 g / L), creatinine (mean 80.5 μmol / L), glucose (mean 5.2 mmol / L), etc. The relationship between phenotypic age and actual age was fitted by a linear regression model, and the residual was calculated to obtain the phenotypic age acceleration value (range -10.2 to 15.3). Using the double Gompertz proportional hazards model, the actual age (mean 56.88 years) and standardized biomarker data were input, such as logarithmically transformed CRP (median 1.8 mg / L), lymphocyte percentage (mean 28.5%), and the individual PhenoAge value was output (formula: PhenoAge = 141.50 + ln(-0.00553×ln(1 - xb) / 0.09165)). Through the elastic net regularized Cox proportional hazards regression model (α = 0.5, λ = 0.01) for 10-fold cross-validation, extreme value truncation was performed on indicators such as red blood cell distribution width (RDW, mean 13.2%) (1st percentile 11.5%, 99th percentile 15.1%), and the standardized phenotypic age equation was obtained after optimization. Based on the median of the phenotypic age acceleration value of 0, the participants were divided into a severe aging group (phenotypic age acceleration value > 0, n = 33,900), a mild aging group, and a non-aging group. The chi-square test was used to compare the smoking rates (severe aging group 32.1% vs non-aging group 18.7%, p < 0.001), and the Mann-Whitney test was used to compare the BMI (severe aging group 28.4 vs non-aging group 26.1, p < 0.001). Using the Royston-Parmar model (degrees of freedom 3) with age as the time scale, the annual survival probabilities from 45 to 100 years old were calculated (such as the survival probability at 70 years old was 0.892), and the area under the curve (mean 38.7 years) was obtained by numerical integration using the trapezoidal method (step size 1 year). Combining with the Fine-Gray competing risks regression model (sub-distribution hazard ratio = 1.21, 95% CI 1.15 - 1.28), the competing risks of death and disease diagnosis were adjusted, and the corrected disease-free years were output (such as severe aging group 31.2 years vs non-aging group 42.5 years). If the bootstrap method detected that the standard error > 0.5, 1000 repeated samplings (seed number 1234) were performed to generate the 95% CI (such as the CI for 31.2 years was 30.8 - 31.6). In the stratified analysis, the Cox proportional hazards regression model was fitted by grouping according to educational level (below high school / college) (hazard ratio for college = 0.87, 95% CI 0.82 - 0.93), and finally the multivariate Cox regression results (adjusted hazard ratio of model 4 = 1.34, 95% CI 1.27 - 1.41) and the predicted values of disease-free years were integrated to generate a JSON format report.
[0039] Experimental Example (1)Study participants of this application This study analyzed two large-scale, longitudinal, and prospective cohorts: the UK Biobank cohort, which contains anonymized data of 502,386 adults aged between 37 and 73 years from 22 assessment centers in England, Wales, and Scotland during March 2006 to October 2010; and the Chinese cohort from the Second Affiliated Hospital of Nanchang University, which contains 400,441 participants aged between 0 and 120 years during July 2006 to September 2024. In the UK Biobank cohort, participants were excluded if they had any missing data in these 9 clinical biomarkers (n = 147,063, 29.3%), a history of type 2 diabetes, dementia, cancer, or respiratory disease before baseline assessment (n = 46,484, 9.3%), and died before the age of 45 (n = 29, 0.1%). Finally, the remaining 308,810 (61.5%) participants were included in the analysis. Similar to the Chinese Biobank cohort, the Chinese cohort excluded 243,360 (60.8%) participants with missing phenotypic age biomarker data, 531 (0.1%) participants with missing diagnosis dates for the four major diseases, and 1,034 (0.2%) participants with a history of the four major diseases before baseline assessment. Finally, 155,516 (38.8%) participants were included in the analysis. Both cohorts used a standardized data collection protocol for longitudinal follow-up to obtain clinical measurements and outcome confirmation information.
[0040] In the Chinese cohort, longitudinal biomarker stability assessment involved repeated biosample collection from 300,345 participants. After applying the exclusion criteria to this subsample, 83,888 (27.9%) participants with missing examination or diagnosis data, 137,127 (45.7%) participants with missing phenotypic age biomarker data, and 1,427 (0.05%) participants with a history of the four major diseases before baseline assessment were excluded. Finally, 77,903 (25.9%) participants met the criteria for repeated analysis.
[0041] Baseline data collection adopted a multi - mode assessment, including questionnaire surveys implemented by interviewers, self - reported digital surveys, standardized anthropometric measurements, and multi - analyte biological sample collection. All eligible participants completed a written informed consent form through a document approved by the institutional review board before study enrollment. The UK Biobank cohort study has been approved by the North West - Multicentre Research Ethics Committee (Ethics Committee reference number: 16 / NW / 0274), and the Chinese cohort study has been approved by the Research Ethics Committee of the Second Affiliated Hospital of Nanchang University. After data collation, the data was analyzed from May 13, 2024, to September 1, 2024.
[0042] (2)Main study The primary endpoint of this study was the first record of any of the four major diseases (type 2 diabetes, respiratory diseases, cancer, dementia). Information on disease diagnosis and diagnosis date was coded according to the 10th Revision of the International Classification of Diseases (ICD - 10) terms in UK Biobank data fields 41270 and 41280, respectively. The date of death and its underlying cause were obtained from the death certificate. Follow - up data was as of May 13, 2024, and the death outcome was censored at the last follow - up data or the date of death, or at the earlier date if the death data occurred earlier.
[0043] (2.1)Covariates Covariates included: age, gender (male; female), body mass index (BMI) categorized by the obesity threshold (≥30 kg / m²; <30 kg / m²), average annual household income (≥31000£; <31000£), employment status (employed / self - employed; unemployed / retired), education level (university degree; non - university degree), aspirin use (non - user; user), smoking history (never; former / current smoker), sleep duration (≤6 hours; ≥7 hours), frequency of dietary change (never; occasionally / often). Specific information on covariates is shown in Table 1.
[0044] Table 1 Field ID of covariates and exposure factors used in this study in the UK Biobank cohort (2.2)Phenotypic age acceleration assessment Phenotypic age was derived using a double-Gompertz proportional hazards model to estimate the risk of death, combining chronological age with nine multisystem clinical chemistry biomarkers. These biomarkers were initially identified by optimizing for mortality prediction using an elastic-net regularized Cox proportional hazards framework with 10-fold cross-validation. To reduce the skewness of the biomarker distributions, extreme value adjustment was performed by setting each biomarker to the lowest 1st and 99th percentiles. Phenotypic age (PhenoAge) was calculated as follows: PhenoAge = 141.50 + ln(-0.00553×ln(1-xb) / 0.09165), where xb = -19.907-0.0336×albumin + 0.0095×creatinine + 0.0195×glucose + 0.0954×ln(CRP)-0.0120×lymphocyte percentage + 0.0268×mean cellular volume + 0.3356×red cell distribution width + 0.00188. Subsequently, phenotypic age acceleration, which is the age-adjusted residual from the linear regression model of phenotypic age against chronological age, was calculated. Analyses were restricted to participants with complications.
[0045] (2.3) Definition of phenotypic age acceleration categories This application implemented multiple phenotypic age acceleration stratification methods to comprehensively assess its association with disease-free life years. The primary classification divided phenotypic age acceleration into three categories (i.e., tertiary classification): (1) severe aging (the first 50% with phenotypic age acceleration values > 0); (2) mild aging (the second 50% with phenotypic age acceleration values > 0); and (3) no aging (phenotypic age acceleration values ≤ 0). Secondary analysis frameworks included: (1) population-based tertiles (three equal groups); (2) quartile-based classification (four equal groups); and (3) dichotomy (accelerated aging: phenotypic age acceleration values > 0 vs. non-accelerated aging: phenotypic age acceleration values ≤ 0).
[0046] (2.4) Disease-free years This application defines a disease - free year as the time period between 45 and 100 years old without being diagnosed with four major age - related diseases. This age range is selected because the disease incidence rate increases exponentially after 45 years old. The analysis used the Royston - Parmar flexible parametric proportional hazards survival model to predict the average disease - free survival time with age as the time scale. The process of calculating life years (the difference in average disease - free years) includes the following analysis stages. The number of remaining disease - free years is estimated by the area under the survival curve at 1 - year intervals between 45 and 100 years old. The individual survival curve is constructed by numerical integration of the annual survival probability, and then the difference in population healthy life expectancy between different accelerated aging groups is calculated, that is, the difference in the area under the two survival curves (such as mild and severe aging groups). The statistical uncertainty is evaluated by generating 95% confidence intervals through 1000 independent bootstrap replicates.
[0047] (2.5) Stratified analysis This application conducted a stratified analysis to evaluate the robustness and potential heterogeneity of phenotypic age - accelerated healthy life expectancy. The analysis was based on demographic characteristics, lifestyle (body mass index, smoking status, sleep duration, dietary changes, and aspirin use), and socioeconomic status (educational level, employment status, average household income). This comprehensive stratified cohort assessment not only verified the predictive consistency of phenotypic age acceleration but also identified its potential modifying effects in biological and behavioral domains. <id
[0048] (2.6) Statistical analysis This application used descriptive statistics to summarize the baseline characteristics of phenotypic age - acceleration categories, reporting the mean (standard deviation, std) of continuous variables and the number (percentage) of categorical variables. The chi - square test was used to compare categorical variables between groups, and the Mann - Whitney - Wilcoxon test was used to compare continuous variables. The multivariable Cox proportional hazards regression model estimated the hazard ratio (HR) and 95% confidence interval (CI). Four sequential adjustment schemes were conducted: (1) Model 0 - unadjusted; (2) Model 1 - adjusted for educational level and BMI; (3) Model 2 - further adjusted for employment status and average household income; (4) Model 3 - further adjusted for sleep duration, smoking status, dietary changes, and aspirin use. In addition, to verify the model stability, supplementary analysis further adjusted for cardiovascular disease and mental disorders (anxiety / depressive disorders). The Kaplan - Meier survival curve was used to estimate the cumulative survival probability. The log - rank test was used to compare the differences in survival curves between strata.
[0049] (3) Trial results (3.1) Baseline characteristics Tables 2 and 3 show the basic demographic characteristics of the study population, including 308,810 participants in the UK Biobank cohort (142,282 men, age, mean [standard deviation], 56.88 [8.19] years; 166,528 women, age, 56.69 [7.99] years) and 155,516 participants in the China cohort (89,641 men, age, 56.72 [19.21] years; 65,875 women, age, 56.75 [18.15] years), with a median follow-up time of 15.22 (interquartile range [IQR], 14.47–15.93) years and 8.33 (IQR, 5.45–10.96) years, respectively. Disease prevalence analysis identified 104,663 (33.9%) UK participants (53,257 men, age 59.78 [7.40] years; 51,406 women, age 58.96 [7.60] years) and 34,303 (22.1%) Chinese participants (20,506 men, age 57.86 [19.66] years; 13,797 women, age 56.84 [18.45] years) diagnosed with ≥1 major disease.
[0050] Calculations of person-years showed that there were differences in risk exposure between different categories of accelerated aging: the UK Biobank cohort accumulated 755,084.3 (severe aging), 855,722.2 (mild aging) and 223,814 (no aging) person-years; the Chinese cohort accumulated 88,382.4 (severe aging), 93,179.6 (mild aging) and 1,055,993.4 (no aging) person-years.
[0051] Table 2 Baseline characteristics of the UK Biobank cohort Table 3 Baseline characteristics of nine phenotypic age clinical chemistry biomarkers in the Chinese cohort (3.2) Phenotypic age acceleration and disease-free years The analysis showed that within the four age strata of accelerated aging, the number of years of disease-free survival at age 45 years increased in a stepwise manner, with a consistent pattern observed across sex and cohort ( Figure 2). The UK Biobank cohort showed that, compared with the severely senescent group, the average disease - free survival years of men in the mildly senescent group increased by 2.96 (95% CI: 2.45, 3.47) years, and those of women increased by 2.97 (2.23, 3.18) years; while the Chinese cohort showed a more significant difference, with an increase of 8.58 (7.65, 9.51) years for men and 7.87 (6.86, 8.87) years for women. Non - senescent individuals showed a better healthy life expectancy advantage. UK men gained 4.23 (3.75, 4.70) years more than the severely senescent group, and women gained 4.48 (4.07, 4.90) years more; in contrast, Chinese cohort men benefited more, with 13.68 (12.67, 14.70) years, and women with 14.31 (13.11, 15.52) years. When comparing the mildly senescent group, this non - senescent advantage still existed. In the UK Biobank cohort, men increased by 1.27 (0.81, 1.73) years and women increased by 1.78 (1.37, 2.19) years; while in the Chinese cohort, men increased by 5.10 (3.78, 6.43) years and women increased by 6.45 (4.95, 7.95) years.
[0052] (3.3) Non - linear association between phenotypic age acceleration and disease risk Multivariable Cox regression analysis showed a significant dose-response association between phenotypic age acceleration and disease risk across all cohorts (Tables 4 and 5). For every 1 standard deviation increase in phenotypic age acceleration, the risk of major disease increased significantly in UK men (hazard ratio, 1.13; 95% confidence interval, 1.12 to 1.14) and in UK women (1.11, 1.10 to 1.12). Compared with those with severe ageing, UK men in the mild ageing group had a 17% lower risk (hazard ratio, 0.83; 95% confidence interval, 0.81-0.85), rising to a 25% reduction in the non-ageing group (hazard ratio, 0.75; 95% confidence interval, 0.74-0.77). UK women showed a similar protective pattern, with a 14% lower risk in the mild ageing group (hazard ratio, 0.86; 95% confidence interval, 0.83-0.88), rising to a 23% reduction in the non-ageing group (hazard ratio, 0.77; 95% confidence interval, 0.75-0.79). Analysis of the Chinese cohort showed an 11% lower risk in men with mild ageing (hazard ratio, 0.89; 95% confidence interval, 0.84-0.94). Secondary analyses of the two distinct scenarios of accelerated versus non-accelerated aging showed a 17% reduced risk in UK men (hazard ratio, 0.83; 95% confidence interval, 0.82-0.84) and a 16% reduced risk in women (hazard ratio, 0.84; 95% confidence interval, 0.82-0.85), while in Chinese individuals, the reduced risks were 6% (hazard ratio, 0.94; 95% confidence interval, 0.90-0.97) and 4% (hazard ratio, 0.96; 95% confidence interval, 0.91-1.00), respectively. Stratified analyses by quartiles and tertiles of phenotypic age acceleration showed consistent dose-dependent associations, indicating a decreasing risk with increasing phenotypic age category.
[0053] Kaplan-Meier survival curves stratified by phenotypic age acceleration categories are shown in Figure 2. Figure 3 Adjusted Fine-Gray competing risk regression models taking into account the mortality endpoint maintained the protective association of the mild / no aging group in both men and women (Table 7).
[0054] Table 4 Relationship between biological aging and major disease risks in the UK Biobank cohort Table 5 The relationship between biological aging and major disease risks in the Chinese population Table 6 Phenotypic acceleration in participants stratified by lifestyle factors in the UK Biobank cohort Table 7 Fine and Gray competing risks regression models in the UK Biobank cohort (3.4)Stratified analysis Multivariable-adjusted stratified analysis showed that the association between phenotypic age acceleration and healthspan was consistent across different socioeconomic levels (educational level, employment status, average household income) and lifestyle levels (body mass index, dietary changes, sleep duration, smoking status) ( Figure 4 and Figure 5 ). Regardless of the changes in the stratification variables, the direction and magnitude of the observed association between phenotypic age acceleration and the increase in disease-free years remained stable, indicating the consistency of the effect.
[0055] (3.5)Mediating effect of phenotypic age acceleration Higher levels of phenotypic age acceleration were significantly associated with modifiable risk factors in both men and women: shorter sleep duration, higher body mass index, occasional or frequent dietary changes, lower educational level, lower household income, past or current smoking, non-regular employment status, and aspirin use (Table 6). Mediation analysis showed that phenotypic age acceleration mediated to some extent the causal pathways of disease development by socioeconomic and lifestyle factors. The lifestyle pathway showed a mediation proportion of 8.91% for shorter sleep duration (95% confidence interval 7.18 - 10.88%, p<0.001), 11.79% for BMI≥30 (11.14 - 12.52%, p<0.001), 8.36% for occasional or frequent dietary changes (6.16 - 11.19%, p<0.001), 6.56% for past or current smoking (6.09 - 7.03%, p<0.001), and 3.74% for aspirin use (3.41 - 4.07%, p<0.001). The mediating factors in the socioeconomic pathway explained 6.49% (5.94 - 7.05%, p<0.001) for lower educational level, 5.18% (4.79 - 5.60%, p<0.001) for lower household income, and 1.24% (1.04 - 1.46%, p<0.001) for non-regular employment status (Table 8).
[0056] Table 8 Mediation of phenotypic age acceleration in the associations between lifestyle, social class, BMI, and aspirin use and major diseases The foregoing has shown and described the basic principles, main features and advantages of the present application. For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic features of the present application, the present application can be implemented in other specific forms. Therefore, in any regard, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present application is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes that fall within the meaning and scope of the equivalent elements of the claims in the present application.
[0057] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for accelerating the assessment of healthspan and disease risk based on phenotypic age, characterized in that, It includes the following steps: S101: Obtain the baseline data of the target population, where the baseline data includes the actual age and multi-system clinical chemical biomarker data; S102: Calculate the phenotypic age based on the baseline data, and determine the phenotypic age acceleration value through the regression analysis of the phenotypic age and the actual age; S103: Stratify the target population according to the phenotypic age acceleration value; S104: Evaluate the association between the aging degree stratification and the disease-free survival years and disease risk through a preset survival analysis model; S105: Output the evaluation result, where the evaluation result characterizes the difference in healthy life expectancy and the change in disease risk of the target population.
2. The method for accelerating the assessment of healthy lifespan and disease risk based on phenotypic age as claimed in claim 1, wherein Step S102 is specifically as follows: S1021: Obtain the multi-system clinical chemical biomarker data, and perform extreme value adjustment on each biomarker data to reduce the distribution skew; S1022: Calculate the phenotypic age through a preset double Gompertz proportional hazards model combined with the actual age and the adjusted biomarker data; S1023: Analyze the residual between the phenotypic age and the actual age by using the linear regression method to obtain the phenotypic age acceleration value; S1024: Determine the aging acceleration degree of the target population according to the size of the phenotypic age acceleration value.
3. The method for accelerating the assessment of healthspan and disease risk based on phenotypic age as claimed in claim 1, wherein Step S103 is specifically as follows: S1031: Obtain the distribution range of the phenotypic age acceleration value; S1032: Divide the phenotypic age acceleration value into multiple aging degree categories through a preset stratification standard, where the aging degree categories include severe aging, mild aging, and non-aging; S1033: For each aging degree category, determine the corresponding proportion of the number of people in the target population; S1034: Generate a stratification result according to the proportion of the number of people and the aging degree category, and the stratification result is used for subsequent healthy life expectancy and disease risk assessment.
4. The method for accelerating the assessment of healthspan and disease risk based on phenotypic age as claimed in claim 1, wherein Step S103 is specifically as follows: S1031: Obtain multiple stratification frameworks for the phenotypic age acceleration value, where the stratification frameworks include three-category stratification and two-category stratification; S1032: Perform multi-dimensional classification on the phenotypic age acceleration value through the stratification framework to obtain multiple aging degree stratification results; S1033: For each aging degree stratification result, analyze its dose-dependent relationship with the disease risk; S1034: Determine the final aging degree stratification scheme according to the dose-dependent relationship, and the final aging degree stratification scheme is used for subsequent survival analysis model evaluation.
5. The method for accelerating the assessment of healthspan and disease risk based on phenotypic age as claimed in claim 1, wherein Step S104 is specifically as follows: S1041: Obtain the distribution data of the aging degree stratification; S1042: Analyze the relationship between the aging degree stratification and the incidence risk of major diseases through a preset proportional hazards regression model to obtain the risk ratio; S1043: Calculate the difference in disease-free survival years corresponding to the aging degree stratification by using a flexible parametric survival model; S1044: Generate an association analysis result according to the risk ratio and the difference in disease-free survival years, and the association analysis result is used to characterize the impact of different aging degrees on healthy life expectancy.
6. The method for accelerating the assessment of healthspan and disease risk based on phenotypic age as claimed in claim 1, wherein Step S104 is specifically as follows: S1041: Obtain data on the lifestyle and socioeconomic factors of the target population; S1042: Through the stratified analysis method, combine the lifestyle and socioeconomic factor data with the stratification of the degree of aging, and analyze the differences in healthy life expectancy among different subgroups; S1043: Use the multivariate adjustment method to correct the analysis results to obtain the corrected correlation data; S1044: Determine the consistent impact of the aging degree stratification in different subgroups according to the corrected correlation data, and the consistent impact is used for the subsequent output of the evaluation results.
7. The method for accelerating the assessment of healthspan and disease risk based on phenotypic age according to claim 1, wherein, Step S104 is specifically as follows: S1041: Obtain the follow-up data of the target population, and the follow-up data includes disease diagnosis records and death records; S1042: Analyze the relationship between the follow-up data and the aging degree stratification through a competing risks regression model to obtain the correlation results after competing risks adjustment; S1043: Use the sensitivity analysis method to verify the robustness of the correlation results; S1044: Update the output data of the survival analysis model according to the correlation results after robustness verification, and the output data is used for the subsequent evaluation of the differences in healthy life expectancy.
8. The method for accelerating the assessment of healthspan and disease risk based on phenotypic age as claimed in claim 1, wherein Step S105 is specifically as follows: S1051: Obtain the correlation analysis data generated by the survival analysis model; S1052: Organize the correlation analysis data through a preset output format, and the output format includes the differences in disease-free survival years and the disease risk ratio; S1053: Generate corresponding healthy life expectancy prediction values for different aging degree categories of the target population; S1054: Output the final evaluation report according to the healthy life expectancy prediction values and the disease risk ratio, and the final evaluation report is used to characterize the changes in the health status of the target population at different aging degrees.
Citation Information
Patent Citations
Method and system for constructing biological age and senescence evaluation based on physical examination markers
CN116779077A
Physiological aging speed evaluation system based on Chinese population, construction method and application
CN118448041A
Biological age evaluation system based on clinical indexes, construction method and application
CN119601208A
Cited By
Physiological age calculation and fracture risk data processing method based on multiple modes
CN122266745A
A pilot driving risk prediction method and device based on big data, and a medium
CN122494264A