Prediction model construction method and system for sarcopenia of old people
By constructing a prediction model based on age, gender and calf circumference, the problem of different diagnostic accuracy of sarcopenia screening tools is solved, and high accuracy and low-cost sarcopenia risk prediction for the elderly in the community are achieved, which is suitable for widespread promotion in the elderly community in China.
Patent Information
- Application Number
- CN202510268693.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the diagnostic accuracy of sarcopenia screening tools is uneven and require professional equipment and personnel, making it difficult to make early diagnosis in economically underdeveloped areas, resulting in a large number of sarcopenia patients missing.
A prediction model based on age, gender, BMI and calf circumference was constructed. A high-accuracy prediction model was established through simple measurement tools such as a tape measure and a height scale, combined with single-factor and multi-factor logistic regression, and used for sarcopenia screening in the elderly in the community.
It has achieved high accuracy and low cost sarcopenia risk prediction among the elderly in the community, avoided dependence on professional equipment and personnel, and was suitable for widespread promotion in the elderly community in China.
Smart Images

Figure CN120376123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of prediction models, and particularly to a method and system for constructing a prediction model for sarcopenia in the elderly population. Background Art
[0002] Sarcopenia is a disease related to age-related disorders of muscle mass, strength, and function, which is highly prevalent in the elderly population and can lead to a series of complications, bringing huge health problems and economic burdens to individuals and society. According to the "Statistical Bulletin on the Development of China's Civil Affairs in 2023", the proportion of the population aged 60 and above in China is 21.1%, exceeding 296 million, and the proportion of the population aged 65 and above is 15.4%, exceeding 216 million. China has currently entered a moderately aging society and will enter a severely aging society around 2035. According to previous studies, the prevalence of sarcopenia in the community elderly population ranges from 5.5% to 25.7%. Considering the huge population base, there are a large number of undetected sarcopenia cases among the elderly in China, and it will continue to increase in the next few decades. Based on the intervenability of sarcopenia, early screening and early diagnosis are of great practical significance for the prevention and treatment of sarcopenia and its complications in China.
[0003] According to the consensus proposed by the Asian Working Group for Sarcopenia (AWGS) and the European Working Group on Sarcopenia in Older People (EWGSOP2) in 2019, it is recommended to use SARC-F, SARC-CalF, or calf circumference (CC) to screen for community sarcopenia. However, the sensitivities of SARC-F and SARC-CalF are relatively low, only 29.0%-33.3% and 58.0%-66.7% respectively, which will lead to a large number of missed diagnoses of sarcopenia patients. Moreover, the critical value of calf circumference (CC) is controversial, and different studies have found different optimal critical values for the population. The AUROCs of the three tools are between 0.50 and 0.82. There are also some other community screening tools, such as SARC-FEBM, SARC-F+AC, SARC-CalF+AC, or Ishii test. According to the diagnostic criteria of AWGS2019, the AUROCs are approximately 0.82, 0.61-0.80, 0.71-0.85, and 0.82-0.85 respectively. The diagnostic accuracies of the screening tools are uneven, and the stabilities are also different, and more studies with larger sample sizes are needed for further verification.
[0004] The diagnosis of sarcopenia requires the measurement of muscle mass and quantity simultaneously. AWGS2019 recommends evaluating muscle mass through muscle strength (grip strength test, <28 kg for men and <18 kg for women) or physical function (6-meter walking test <1.0 m / s, 5-time sit-to-stand time ≥12 s, or Short Physical Performance Battery (SPPB) ≤9 points), and estimating muscle quantity through appendicular skeletal muscle mass (ASM) (measurement methods: DXA (dual-energy X-ray absorptiometry) or BIA (bioelectrical impedance analysis)). Based on the diagnostic criteria, the diagnosis of sarcopenia requires professional equipment and personnel simultaneously, which greatly limits the early diagnosis of sarcopenia, especially in less developed regions. Therefore, a simple and easy-to-use early self-screening tool for sarcopenia is particularly important to screen out high-risk and suspected cases early and promote the early diagnosis and treatment of sarcopenia. Summary of the Invention
[0005] The purpose of the present invention is to provide a new, highly accurate and easy-to-use prediction model, which can accurately estimate the likelihood of sarcopenia in community-dwelling elderly people without the need for complex professional tools and personnel, and can be widely promoted and used in the community-dwelling elderly population to solve the problems raised in the above background technology.
[0006] To achieve the above invention purpose, one aspect of the present invention provides a method for constructing a prediction model for sarcopenia in the elderly population, including the following steps:
[0007] Step S1, determining the population scope and criteria included in the prediction model;
[0008] Step S2, determining the diagnostic criteria for sarcopenia;
[0009] Step S3, collecting sarcopenia-related factors and related physical measurement data from the selected population;
[0010] Step S4, using univariate and forward, backward, and stepwise multivariate logistic regression methods to select the best sarcopenia prediction model factors and establish a prediction model;
[0011] Step S5, selecting the best model by comparing model performance parameters and visualizing it as a nomogram;
[0012] Step S6, using the nomogram to predict sarcopenia in the elderly.
[0013] Further, the inclusion criteria in step S1 are age ≥60 years, and those who cannot provide information, cannot walk, and those who cannot cooperate with body measurements are excluded.
[0014] Further, the diagnostic criteria for sarcopenia are as follows:
[0015] Low ASM, M: <7.0 kg / m 2 , F: <5.7 kg / m2 ;
[0016] And the combined muscle strength is low, M: < 28 kg, F: < 18 kg;
[0017] Where M represents male and F represents female;
[0018] Or those with poor physical function and a 6 - meter walking speed of < 1.0 m / s are diagnosed with sarcopenia.
[0019] Furthermore, the sarcopenia - related factors and relevant physical measurement data in step S3 include: age, gender, marital status, education level, chronic diseases, surgeries, drug use, alcohol consumption, smoking, sleep, physical activity, and eating habits, as well as estimating physical activity through the short - form International Physical Activity Questionnaire and measuring the calf circumference when sitting at rest.
[0020] Furthermore, in step S4, all the collected data is randomly divided into a training set and a validation set in a ratio of 7:3, and an optimized model based on the training set is constructed respectively. The formula is expressed as the following formula:
[0021] logit(p) = 0.1227age - 0.7862sex - 0.1970BMI - 0.3476CC - 0.0158dbp + 0.1061sitting_time,
[0022] The simplified model formula is expressed as: logit(p) = 0.1263age - 0.7775sex - 0.1943BMI - 0.3472CC,
[0023] Where: logit(p) is the dependent variable y, and the variables on the right - hand side of the equation are all independent variables x. age is the age, with the unit of years; sex is the gender, 1 = male, 2 = female, BMI is the BMI index, with the unit of kg / m 2 ; CC is the calf circumference, with the unit of cm; dbp is the diastolic blood pressure, with the unit of mmHg; sitting_time is the daily sedentary time, with the unit of hours;
[0024] Based on the generalizable population, applicable scenarios, and accuracy of the model, the best sarcopenia prediction model factors are determined to be age, gender, BMI, and calf circumference, and the validation set is used for model validation.
[0025] Furthermore, in step S5, by comparing the AUROC, sensitivity, and specificity of the model, it is determined that the best model is the simplified model based on the training set. All the data is incorporated into the best model to train a prediction model for sarcopenia in the elderly population. The formula is expressed as: logit(p) = 0.1287age - 0.7152sex - 0.2036BMI - 0.2955CC,
[0026] Among them, logit(p) is the dependent variable y, and the variables on the right side of the equation are all independent variables x. age represents age in years; sex represents gender, where 1 = male and 2 = female, and BMI is the BMI index in kg / m 2 ; CC is the calf circumference in cm.
[0027] Furthermore, the method for predicting sarcopenia in the elderly using the nomogram in step S6 is to measure the calf circumference and BMI, find the scores corresponding to each variable on the nomogram of the prediction model according to the age and gender of the self-tester, add up the scores to obtain the total score, and query the prevalence rate under the total score of the nomogram to obtain the risk of sarcopenia.
[0028] Another aspect of the present invention provides a prediction model construction system for sarcopenia in the elderly population, including a population standard module, a diagnostic standard module, a data collection module, a model establishment module, a model determination module, and a model usage module, where:
[0029] The population standard module is used to determine the population range and criteria included in the prediction model;
[0030] The diagnostic standard module is used to determine the diagnostic criteria for sarcopenia;
[0031] The data collection module is used to collect sarcopenia-related factors and related physical measurement data from the selected population;
[0032] The model establishment module is used to select the best sarcopenia prediction model factors using univariate and forward, backward, and stepwise multivariate logistic regression methods and establish a prediction model;
[0033] The model determination module is used to select the best model by comparing model performance parameters and visualize it as a nomogram;
[0034] The model usage module is used to predict sarcopenia in the elderly using the nomogram.
[0035] Due to the adoption of this system and method, compared with the prior art, it has the following advantages:
[0036] Based on the present invention, in the elderly population aged 60 and above in the community, only a tape measure and a height and weight scale are needed to predict the risk of sarcopenia, without the need for professional medical equipment such as a grip strength meter, X-ray, bioelectrical impedance measurement and other related instruments, nor a professional examining doctor, and without high medical expenses, which will bring great convenience to the prediction of the risk of sarcopenia in the elderly population in the community. Moreover, the model has extremely high accuracy and stability, facilitating popularization and application in Chinese elderly communities. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1It is a flow chart of a method for constructing a prediction model for sarcopenia in the elderly population.
[0038] Figure 2 It is a nomogram of the sarcopenia prediction model in the elderly population for Model 2.
[0039] Figure 3 It is an ROC curve graph predicted by the data of 1000 bootstrap repeated samplings for both the training set and the validation set in the simplified model, where the left graph is the ROC curve graph of the training set and the right graph is the ROC curve graph of the validation set.
[0040] Figure 4 It is a nomogram of the sarcopenia prediction model in the elderly population for the final model.
[0041] Figure 5 It is a graph of sensitivity and specificity of 4 models at different cut-off points.
[0042] Figure 6 It is a calibration curve graph of 4 models.
[0043] Figure 7 It is a DCA curve graph of 4 models. Specific implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] As Figure 1 shown in the flow chart of the method of the present invention, the embodiments of the present invention provide a method, and the specific steps are as follows:
[0046] Step S1, determine the population range and criteria for including in the prediction model.
[0047] Two-stage random sampling was carried out in Pudong New Area, Shanghai according to the nature of the suburbs and communities. In the first stage, 2 suburbs and 2 urban areas were randomly selected according to the regional location (suburbs and urban areas). In the second stage, 1 elderly community or 1 ordinary community was randomly selected from the urban areas or suburbs selected in the first stage. A total of 2528 community elderly people were recruited. The inclusion criteria were: age ≥ 60 years old. The exclusion criteria were: (1) unable to provide information; (2) unable to walk, and other situations where it was impossible to cooperate with physical measurements.
[0048] Step S2, determine the diagnostic criteria for sarcopenia.
[0049] The diagnosis of sarcopenia was performed according to the AWGS 2019 criteria. Handgrip strength (HGS) was measured using a digital handgrip dynamometer (EH106, Kangdu Medical Equipment Co., Ltd., Beijing, China). The dominant hand was tested twice, with a 30-second interval between tests, and the higher value was taken as the handgrip strength value of the participant. Physical function was measured by the 6-meter gait speed. The participant walked 6 meters at a normal speed, and the walking time was measured and converted into a walking speed. The participant was required to walk twice, and the higher value was taken. Appendicular skeletal muscle mass (ASM) was evaluated using a bioelectrical impedance analysis (BIA) device (Inbody270, Biospace Co., Ltd., Korea). People with low ASM (M: < 7.0 kg / m 2 , F: < 5.7 kg / m 2 ), low muscle strength (M: < 28 kg, F: < 18 kg), or poor physical function (6-meter walking speed: < 1.0 m / s) were diagnosed with sarcopenia.
[0050] Step S3, collect sarcopenia-related factors and related physical measurement data from the selected population.
[0051] From May 2023 to October 2023, 2453 participants aged ≥ 60 years were consecutively enrolled. All participants were informed, and questionnaires and physical indicators were measured by doctors and clinical research coordinators. According to the unique participant item code, the questionnaires and anthropometric information of the 2453 participants were successfully matched and included in the statistical dataset.
[0052] The sarcopenia prediction model factors were selected from sarcopenia-related factors and related physical measurement data. The questionnaires covered the general demographic characteristics, clinical characteristics, and lifestyle habits of the participants. Information such as age, gender, marital status, education level, chronic diseases, surgeries, drug use, alcohol consumption, smoking, sleep, physical activity, and eating habits was collected. Physical activity was estimated using the International Physical Activity Questionnaire Short Form (IPAQ-SF). The calf circumference (CC) at rest while sitting was measured twice, and the larger value was taken.
[0053] Step S4, use univariate and forward, backward, and stepwise multivariate logistic regression methods to select the best sarcopenia prediction model factors and establish a prediction model.
[0054] Participants were randomly assigned to the training set (N = 1718) and the validation set (N = 735) according to sarcopenia status in a 7:3 ratio, and a random seed was set to ensure the reproducibility of randomization. The maximum missing value in all datasets did not exceed 5%. Assuming that the missingness was completely random, the data was not imputed, and the analysis was based on the complete data. Univariate and forward, backward, and stepwise multivariate logistic regression were used to select the best predictors and build a prediction model. Among the data containing variables in the multivariate logistic regression model, there was only 1 missing value, so no data imputation or sensitivity analysis was performed on the model.
[0055] The prevalence of sarcopenia in all participants was 14.1% (345 / 2453). In the training and validation datasets, the prevalence rates were 14.1% and 13.9% respectively. The characteristics of the participants in the training set and the validation set are shown in Table 1. The average age of the participants was 72.55 ± 6.23 years, and the average BMI was 24.57 ± 3.43 kg / m 2 There were no statistically significant differences in other general information and clinical characteristics between the two datasets except for educational level.
[0056] Table 1. General demographics and clinical characteristics of participants in the training and validation sets
[0057] Feature N <![CDATA[Total, N = 2,453 1 > <![CDATA[Training set, N = 1,718 1 > <![CDATA[Validation set, N = 735 1 > <![CDATA[p-value 2 > Street 2,453 0.166 Beicai 327 (13.3%) 212 (12.3%) 115 (15.6%) Hangtou 1,900 (77.5%) 1,343 (78.2%) 557 (75.8%) Kangqiao 160 (6.5%) 115 (6.7%) 45 (6.1%) Zhangjiang 66 (2.7%) 48 (2.8%) 18 (2.4%) Age, mean(SD), y 2,453 72.55(6.23) 72.68(6.24) 72.22(6.18) 0.094 Gender 2,453 0.285 Male 1,038 (42.3%) 715 (41.6%) 323 (43.9%) Female 1,415 (57.7%) 1,003 (58.4%) 412 (56.1%) Marital status 2,453 0.559 Married 2,450 (99.9%) 1,715 (99.8%) 735 (100.0%) Unmarried 3 (0.1%) 3 (0.2%) 0 (0.0%) Height, mean(SD), cm 2,453 157.19(8.32) 157.01(8.28) 157.59(8.42) 0.248 Weight, mean(SD), kg 2,453 60.81(10.21) 60.69(10.00) 61.07(10.70) 0.355 <![CDATA[BMI, mean(SD), kg / m 2 > 2,453 24.57(3.43) 24.59(3.43) 24.52(3.43) 0.925 Systolic blood pressure, mean(SD), mmHg 2,453 140.21(17.00) 140.48(17.33) 139.57(16.21) 0.698 Diastolic blood pressure, mean(SD), mmHg 2,453 77.94(10.10) 77.96(10.19) 77.89(9.87) 0.742 Heart rate, mean(SD), bpm 2,453 77.61(10.63) 77.76(10.66) 77.27(10.54) 0.276 Educational level 2,402 0.037 Below middle school 1,396 (58.1%) 1,003 (59.5%) 393 (54.9%) Middle school or above 1,006 (41.9%) 683 (40.5%) 323 (45.1%) History of chronic diseases 2,453 0.338 No 616 (25.1%) 422 (24.6%) 194 (26.4%) Yes 1,837 (74.9%) 1,296 (75.4%) 541 (73.6%) Osteoporosis 2,453 0.821 No 2,380 (97.0%) 1,666 (97.0%) 714 (97.1%) Yes 73 (3.0%) 52 (3.0%) 21 (2.9%) CKD 2,453 0.311 No 2,436 (99.3%) 1,708 (99.4%) 728 (99.0%) Yes 17 (0.7%) 10 (0.6%) 7 (1.0%) COPD 2,453 0.627 No 2,433 (99.2%) 1,703 (99.1%) 730 (99.3%) Yes 20 (0.8%) 15 (0.9%) 5 (0.7%) Diabetes 2,453 0.697 No 2,028 (82.7%) 1,417 (82.5%) 611 (83.1%) Yes 425 (17.3%) 301 (17.5%) 124 (16.9%) Hypertension 2,453 0.140 No 1,137 (46.4%) 813 (47.3%) 324 (44.1%) Yes 1,316 (53.6%) 905 (52.7%) 411 (55.9%) Heart failure 2,453 0.639 No 2,448 (99.8%) 1,715 (99.8%) 733 (99.7%) Yes 5 (0.2%) 3 (0.2%) 2 (0.3%) Malignant tumor 2,453 0.696 No 2,378 (96.9%) 1,667 (97.0%) 711 (96.7%) Yes 75 (3.1%) 51 (3.0%) 24 (3.3%) Surgical history 2,453 0.729 No 1,566 (63.8%) 1,093 (63.6%) 473 (64.4%) Yes 887 (36.2%) 625 (36.4%) 262 (35.6%) History of drug use (specific drugs) 2,452 0.480 No 762 (31.1%) 541 (31.5%) 221 (30.1%) Yes 1,690 (68.9%) 1,176 (68.5%) 514 (69.9%) Smoking 2,453 0.821 No 2,019 (82.3%) 1,416 (82.4%) 603 (82.0%) Yes 434 (17.7%) 302 (17.6%) 132 (18.0%) Drinking 2,452 0.403 No 2,133 (87.0%) 1,500 (87.4%) 633 (86.1%) Yes 319 (13.0%) 217 (12.6%) 102 (13.9%) Sleep duration mean(SD), h 2,453 6.79(1.52) 6.82(1.53) 6.74(1.51) 0.486 Dietary habits 2,453 0.472 No special 2,273 (92.7%) 1,597 (93.0%) 676 (92.0%) Vegetarian 141 (5.7%) 97 (5.6%) 44 (6.0%) High-protein diet 39 (1.6%) 24 (1.4%) 15 (2.0%) Calf circumference, mean(SD), cm 2,453 33.35(3.00) 33.34(2.91) 33.37(3.19) 0.791 Sarcopenia 2,453 0.862 No 2,108 (85.9%) 1,475 (85.9%) 633 (86.1%) Yes 345 (14.1%) 243 (14.1%) 102 (13.9%)
[0058] In the training set, the prevalence of sarcopenia was 14.1%. There were statistically significant differences in age, BMI, diastolic blood pressure, presence of osteoporosis, eating habits, sitting time, physical activity level, and calf circumference between sarcopenic and non-sarcopenic participants (p < 0.05). Medication use and surgical history were not included as factors in the analysis because they were highly correlated with the disease history. Since BMI is a combined measure of height and weight, height and weight were also not included in the analysis. More detailed information can be found in Table 2. In the univariate logistic regression analysis, variables with a P-value less than 0.1 were considered for inclusion in the multivariate logistics regression model.
[0059] Table 2 Results of univariate logistics regression analysis in the training set
[0060]
[0061] In the training set, multivariate logistic regression analysis was performed using forward, backward, and stepwise regression methods. Age, BMI, diastolic blood pressure, osteoporosis, eating habits, sedentary duration, physical activity level, calf circumference, gender, and hypertension were included in the optimized model (Model 1). Considering the final application scenario of the model and its prediction accuracy, a simplified model (Model 2) was developed, which only included four simple and objective predictive factors: age, gender, BMI, and calf circumference.
[0062] The variance inflation factor (VIF) was used to test the multicollinearity of variables. VIF tests were performed in the model analysis, and the VIF values of all variables were less than 1.60, indicating no multicollinearity problem. As shown in Table 3, the regression coefficients of age, gender, BMI, and calf circumference in Model 1 were 0.1227 (0.0128), -0.7862 (0.1806), -0.1970 (0.0365), and -0.3476 (0.0448), respectively. The formula is expressed as: logit(p) = 0.1227age - 0.7862sex - 0.1970BMI - 0.3476CC - 0.0158dbp + 0.1061sitting_time. In Model 2, they were 0.1263 (0.0126), -0.7775 (0.1795), -0.1943 (0.0362), and -0.3472 (0.0444), respectively. The formula is expressed as: logit(p) = 0.1263age - 0.7775sex - 0.1943BMI - 0.3472CC. Where: logit(p) is the dependent variable y, and the variables on the right side of the equation are all independent variables x. age is age in years; sex is gender, 1 = male, 2 = female, BMI is the BMI index in kg / m 2 ; CC is the calf circumference in cm; dbp is the diastolic blood pressure in mmHg; sitting_time is the daily sedentary time in hours. This shows a high degree of consistency between Model 1 and Model 2. For more information on model evaluation, please refer to Table 4. The nomogram model visualization of the simplified model (Model 2) in the training set is as Figure 2 shown.
[0063] Table 3 Results of the optimized and simplified multivariate logistic regression models based on the training set and the simplified model based on all data
[0064]
[0065] Among them, Model 1 is the optimized model based on the training set, Model 2 is the simplified model based on the training set, and Model 4 is the simplified model based on all data.
[0066] Table 4. Evaluation indicators of the multi-factor logistic regression model
[0067]
[0068] Step S5, select the best model by comparing the model performance parameters and visualize it as a nomogram.
[0069] The area under the receiver operating characteristic curve (AUROC) was used to determine the discrimination ability of the model. The Hosmer-Lemeshow goodness-of-fit test was used to evaluate the model fit. To evaluate the stability of the prediction model, internal validation was performed. Using the 1000-times unrestricted random sampling method, the adjusted AUROC values of the training set and the validation set were obtained. The calibration curve was used to describe the consistency between the predicted probability and the observed results. Decision curve analysis (DCA) was used to evaluate the clinical utility. The nomogram was used to describe the risk of sarcopenia in Chinese elderly people.
[0070] The model performances of all models are shown in Table 5. The AUROC and 95% CI of Model 1 and Model 2 were 0.8720 (0.8510 - 0.8930) and 0.8688 (0.8476 - 0.8900), respectively. Delong's test found that there was no statistical significance between the AUROCs of Model 1 and Model 2 (-0.0032 (95% CI, -0.0077 - 0.001, p = 0.166)). At the threshold probability of 0.111, the sensitivity of Model 1 was 0.8724 (95% CI, 0.8305 - 0.9144), and the specificity was 0.7119 (95% CI: 0.6888 - 0.7350), while the sensitivity of Model 2 was 0.8383 (95% CIC: 0.8258 - 0.9108), and the specificity was 0.7186 (95% CID: 0.6957, 0.7416). In the internal validation of the training set, the AUROC was 0.8700 (95% CI, 0.8488 - 0.8894), the sensitivity was 0.8706 (95% CI, 0.8289 - 0.9127), and the specificity was 0.7190 (95% CI: 0.6963 - 0.7409) (as Figure 3 shown on the left). In the validation set, the AUROC was 0.8465 (95% CI, 0.8102 - 0.8828), the sensitivity was 0.7941 (95% CI, 0.7156 - 0.8726), and the specificity was 0.7062 (95% CI, 0.6707 - 0.7416).
[0071] In addition, for the validation dataset, the AUROC obtained by internal validation using the bootstrap method was 0.8463 (95% CI, 0.8103 - 0.8814), the sensitivity was 0.7944 (95% CI: 0.7130 - 0.8684), and the specificity was 0.7057 (95% CI - 0.6719 - 0.7402) (as shown in Figure 3 the right). These validation results indicate that at a specific threshold, the AUROC, sensitivity, and specificity are relatively high, demonstrating the stability and generalizability of the model. Therefore, we used all the datasets to incorporate the variables in the simplified model to fit the final prediction model (Model 4), and the resulting formula is: logit(p) = 0.1287age - 0.7152sex - 0.2036BMI - 0.2955CC, where logit(p) is the dependent variable y, and the variables on the right side of the equation are all independent variables x. Age is in years, sex is gender, 1 = male, 2 = female, BMI is the BMI index in kg / m 2 ; CC is the calf circumference in cm. The results show that Model 4 exhibits highly consistent regression coefficients, AUROC, sensitivity, and specificity with the simplified model at a threshold probability of 0.111, as shown in Table 5. The visualization of the final model is as shown in Figure 4 the nomogram.
[0072] Table 5. Performance of the Multivariate Logistic Regression Model for Prediction (Threshold = 0.111)
[0073] Model AUROC(95%CI) Sen(95%CI) Spe(95%CI) PPV(95%CI) NPV(95%CI) Model 1 0.8720(0.8510-0.8930) 0.8724(0.8305-0.9144) 0.7119(0.6888-0.7350) 0.3328(0.2962-0.3694) 0.9713(0.9614-0.9813) Model 2 0.8688(0.8476-0.8900) 0.8683(0.8258-0.9108) 0.7186(0.6957-0.7416) 0.3371(0.3000-0.3741) 0.9707(0.9607-0.9807) Model 2 Bootstrap internal validation 0.8700(0.8488-0.8894) 0.8706(0.8289-0.9127) 0.7190(0.6963-0.7409) 0.3377(0.3002-0.3764) 0.9712(0.9618-0.9808) Model 3 0.8465(0.8102-0.8828) 0.7941(0.7156-0.8726) 0.7062(0.6707-0.7416) 0.3034(0.2482-0.3585) 0.9551(0.9364-0.9739) Model 3 Bootstrap validation 0.8463(0.8103-0.8814) 0.7944(0.7130-0.8684) 0.7057(0.6719-0.7402) 0.3021(0.2481-0.3521) 0.9554(0.9353-0.9735) Model 4 0.8617(0.8433-0.8801) 0.8464(0.8083-0.8844) 0.7102(0.6908-0.7295) 0.3234(0.2929-0.3539) 0.9658(0.9568-0.9749)
[0074] Among them, Model 3 is the internal validation model for the validation set.
[0075] The risk of sarcopenia is higher in the elderly, males, those with lower BMI, and those with lower calf circumference. For example, as shown in Figure 4 a male aged 87 with a body mass index of 20.2 kg / m 2 and a calf circumference of 31 cm has a total score of 286 in the prediction model, which means his risk of sarcopenia is 77.0%.
[0076] As shown in Figure 5 are the sensitivity and specificity of each model at different threshold probabilities. A threshold of 0.111 was determined as the optimal cut-off value for this study, generally making the sensitivity of the model higher than 0.80 and the specificity higher than 0.70, making it suitable for the purpose of sarcopenia screening.
[0077] As shown in Figure 6The calibration curves of all models are shown as follows. The calibration slope and Spiegelhalter z-score of the optimized model (Model 1) are 1.000 and 0.584 (p = 0.560), respectively. The calibration slope and Spiegelhaller z-score of the simplified model (Model 2) are 1.000 and 0.533 (p = 0.573), respectively. The calibration slope and Spigelhalter z-score of the internal validation model (Model 3) are 0.856 and 1.175 (p = 0.240), respectively. The calibration slope and Spiegelhalter z-score of the final model (Model 4) are 1.000 and 0.615 (p = 0.538). In addition, the Hosmer-Lemeshow test did not reach statistical significance in the optimized model (χ² = 10.02, p = 0.264), the simplified model (χ² = 6.79, p = 0.559), the internal validation model (χ² = 9.42, p = 0.309), and the final model (χ² = 12.89, p = 0.116). The above results indicate that the prediction results of the model are in good agreement with the actual results.
[0078] The results of the decision curve analysis (DCA) on the clinical application value of the prediction model are as Figure 7 shown. The results show that the net benefit (NB) threshold probability range of the optimized model (Model 1) is 0.5% to 75.9%, the net benefit threshold probability of the simplified model (Model 2) is 1.0% to 77.4%, the net benefit threshold probability of the validation model (Model 3) is 0.9% to 93.4%, and the net benefit threshold probability of the final model (Model 4) is 0.9% - 77.2%. These results indicate that compared with non-model methods, the model has a larger range of net benefits and has clinical practicability in predicting sarcopenia in community-dwelling elderly people. At a threshold probability of 0.111, the NB of Model 1 is 0.0925, the NB of Model 2 is 0.0927, the NB of Model 3 is 0.0787, and the NB of Model 4 is 0.0879. This means that to detect one case of sarcopenia, Model 1 requires approximately 11 people (1 / 0.0925 = 10.81), Model 2 requires approximately 11 people (1 / 0.0927 = 10.79), Model 3 requires approximately 13 people (1 / 0.0787 = 12.71), and Model 4 requires approximately 12 people (1 / 0.0879 = 11.38), and all of these can be achieved only by the model without a complex diagnostic process. The difference in NB between Model 2 and Model 1 is 0.0002, the difference in NB between Model 2 and Model 3 is 0.0141, and the difference in NB between Model 2 and Model 4 is 0.0048, indicating that the NB of the optimized model, the simplified model, the validation model, and the final model are comparable.
[0079] Step S6, use the nomogram to predict sarcopenia in the elderly.
[0080] In the actual application scenario of the prediction model, the elderly person or their family member uses a tape measure to measure the calf circumference (the thickest part of the calf), measures height and weight. It is more convenient if there is a BMI converter or paper plate at home. Without conversion, the BMI value can be directly obtained based on height and weight. If not, the conversion can be done according to the formula BMI = weight (kg) / height (m) 2 After conversion, find the scores corresponding to each variable on the prediction model nomogram according to the age and gender of the self-tester, calculate the total score, and the corresponding prevalence rate under the total score can indicate the risk of sarcopenia. For example, as Figure 4 shown, an 87-year-old male with a body mass index of 20.2 kg / m 2 , a calf circumference of 31 cm, the total score of the prediction model is 286, and the risk of sarcopenia is 77.0%, which is extremely high. It is recommended to immediately go to the hospital for the diagnosis and treatment of sarcopenia. This prediction model is most stable when the threshold probability is 0.111. Therefore, it is recommended that when the self-tested disease probability is greater than 0.111, go to the relevant department of sarcopenia in the hospital for medical treatment to achieve early detection and early treatment to prevent complications.
[0081] Based on the present invention, in the elderly population aged 60 and above in the community, only a tape measure and a height and weight scale are needed to predict the risk of sarcopenia. There is no need for professional medical equipment such as a grip strength meter, X-ray, bioelectrical impedance measurement and other related instruments, nor professional examining doctors, and no high medical expenses, which will bring great convenience to the prediction of the risk of sarcopenia in the elderly population in the community. Moreover, the model has extremely high accuracy and stability, which is convenient for popularization and application in Chinese elderly communities.
[0082] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a prediction model for sarcopenia in the elderly population, characterized in that, It includes the following steps: Step S1: Determine the population scope and criteria for inclusion in the prediction model; Step S2: Determine the diagnostic criteria for sarcopenia; Step S3: Collect sarcopenia-related factors and relevant physical measurement data from the selected population; Step S4: Use univariate and forward, backward, and stepwise multivariate logistic regression methods to select the best sarcopenia prediction model factors and establish a prediction model; Step S5: Select the best model by comparing model performance parameters and visualize it as a nomogram; Step S6: Use the nomogram to predict sarcopenia in the elderly.
2. The method for constructing a prediction model for sarcopenia in the elderly population according to claim 1, wherein In step S1, the inclusion criteria are age ≥ 60 years, and those who are unable to provide information, unable to walk, or unable to cooperate with body measurements are excluded.
3. A method for constructing a prediction model for sarcopenia in the elderly population according to claim 1, wherein, The diagnostic criteria for sarcopenia are as follows: Low ASM, M: < 7.0 kg / m 2 , F: < 5.7 kg / m 2 ; And combined with low muscle strength, M < 28 kg, F < 18 kg; Where M represents male and F represents female; Or those with poor physical function, 6-meter walking speed: < 1.0 m / s are diagnosed with sarcopenia.
4. A method for constructing a prediction model for sarcopenia in the elderly population according to claim 1, wherein In step S3, the sarcopenia-related factors and relevant physical measurement data include: age, gender, marital status, education level, chronic diseases, surgeries, drug use, alcohol consumption, smoking, sleep, physical activity, and eating habits, as well as estimating physical activity through the short form of the International Physical Activity Questionnaire and measuring the calf circumference when sitting at rest.
5. A method for constructing a prediction model for sarcopenia in the elderly population according to claim 1, characterized in that, In step S4, all the collected data are randomly divided into a training set and a validation set in a ratio of 7:3, and an optimized model based on the training set is constructed respectively. The formula is expressed as the following formula: logit(p) = 0.1227age - 0.7862sex - 0.1970BMI - 0.3476CC - 0.0158dbp + 0.1061sitting_time, The simplified model formula is expressed as: logit(p) = 0.1263age - 0.7775sex - 0.1943BMI - 0.3472CC, where: logit(p) is the dependent variable y, and the variables on the right side of the equation are all independent variables x, age is age in years; sex is gender, 1 = male, 2 = female, BMI is the BMI index in kg / m 2 ; CC is the calf circumference in cm; dbp is the diastolic blood pressure in mmHg; sitting_time is the daily sedentary time in hours; Based on the generalizable population, applicable scenarios, and accuracy of the model, the best sarcopenia prediction model factors are determined to be age, gender, BMI, and calf circumference, and the validation set is used for model validation.
6. The method for constructing a prediction model for sarcopenia in the elderly population according to claim 1, wherein In step S5, by comparing the AUROC, sensitivity, and specificity of the model, the best model is determined to be the simplified model based on the training set. All the data are incorporated into the best model for training to obtain a prediction model for sarcopenia in the elderly population. The formula is expressed as: logit(p) = 0.1287age - 0.7152sex - 0.2036BMI - 0.2955CC, where logit(p) is the dependent variable y, and the variables on the right side of the equation are all independent variables x, age is the age in years; sex is the gender, 1 = male, 2 = female, BMI is the BMI index in kg / m 2 ; CC is the calf circumference in cm.
7. A method for constructing a prediction model for sarcopenia in the elderly population according to claim 1, characterized in that, In step S6, the method of using the nomogram to predict sarcopenia in the elderly is to measure the calf circumference and BMI, find the scores corresponding to each variable on the prediction model nomogram according to the age and gender of the self-tester, add up the scores of each item to get the total score, and query the prevalence rate under the total score of the nomogram to obtain the risk of sarcopenia.
8. A prediction model construction system for sarcopenia in the elderly population, characterized in that, It includes a population criteria module, a diagnostic criteria module, a data collection module, a model establishment module, a model determination module, and a model usage module, where: The population criteria module is used to determine the population scope and criteria for inclusion in the prediction model; The diagnostic criteria module is used to determine the diagnostic criteria for sarcopenia; The data collection module is used to collect sarcopenia-related factors and relevant physical measurement data from the selected population; The model establishment module is used to select the best sarcopenia prediction model factors and establish a prediction model using univariate and forward, backward, and stepwise multivariate logistic regression methods; The model determination module is used to select the best model by comparing model performance parameters and visualize it as a nomogram; The model usage module is used to predict sarcopenia in the elderly using the nomogram.