Community old people falling risk identification and intervention method based on machine learning

By combining offline questionnaires and machine learning models with the SHAP method, the risk of falls among the elderly in the community was identified and targeted intervention strategies were developed. This solved the systemic gap in fall prevention for the elderly at the community level and enabled the refined identification of risk factors and fall prevention measures.

CN120974281APending Publication Date: 2025-11-18SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511192322.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify fall risks among the elderly at the community level and develop targeted intervention strategies, resulting in a lack of systematic solutions.

Method used

Data was collected through offline questionnaires, a fall risk identification model was built using machine learning algorithms, and the model was interpreted using the SHAP method to identify risk factors and thresholds, and to develop targeted fall prevention measures.

Benefits of technology

This paper presents a fall prevention framework for older adults applicable to communities in different regions. It can identify risk characteristics and develop targeted strategies, and combines machine learning and SHAP methods to explore nonlinear relationships, thereby improving prevention effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974281A_ABST
    Figure CN120974281A_ABST
Patent Text Reader

Abstract

The invention discloses a community old people falling risk identification and intervention method based on machine learning, and relates to the technical field of old people falling risk identification, and the method comprises the steps: collecting the social economy, health condition, perceived environment and falling information data of old people by a system through offline questionnaire survey, and carrying out the data processing and statistics; based on the processed data, taking the tumble information data as a dependent variable, and utilizing a machine learning algorithm to build a tumble risk identification model; explaining the tumble risk identification model through an SHAP method to obtain a risk factor identification result; and based on a risk factor identification result, formulating a targeted anti-falling intervention measure for a corresponding community. Therefore, the method for identifying and intervening the falling risk of the old people in the community based on the machine learning is beneficial to finding the nonlinear relationship and the threshold value between the risk factor and the falling, so that different communities output targeted anti-falling measures for the old people.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of fall risk identification for the elderly, and in particular to a community elderly fall risk identification and intervention method based on machine learning. BACKGROUND

[0003] Existing fall prediction methods for the elderly mainly focus on practical devices such as wearable devices, AI intelligent detection, and fall alarm devices, and realize individual fall risk identification and real-time warning through sensors and machine learning algorithms. Although individual protection technology is developing rapidly, not every elderly person can benefit from it, and most community-dwelling elderly people cannot access more advanced fall prevention or risk identification devices. In addition, the building conditions, environmental characteristics, and living habits of different regions are not the same, resulting in differences in risk factors for falls among community-dwelling elderly people.

[0004] Therefore, it is necessary to identify the fall risk of elderly people in a specific community from a community perspective and develop targeted intervention strategies, which is a more effective means of preventing falls in the elderly population. Today, there is still a gap in the above-mentioned comprehensive fall prevention framework for community elderly people, and there is a lack of systematic solutions from the perspectives of community planning, public facility optimization, and environmental risk assessment. SUMMARY

[0005] The purpose of the present application is to provide a community elderly fall risk identification and intervention method based on machine learning, which can be applied to communities in different regions and output fall prevention strategies for the elderly population in that community.

[0006] To achieve the above-mentioned purpose, the present application provides a community elderly fall risk identification and intervention method based on machine learning, comprising the following steps:

[0007] S1, through offline questionnaire survey, the system collects the social and economic, health status, perceived environment, and fall information data of the elderly, and performs data processing and statistics;

[0008] S2, based on the processed data, taking the fall information data as the dependent variable, a fall risk identification model is built using machine learning algorithms;

[0009] S3, the fall risk identification model is explained by SHAP method, and the risk factor identification result is obtained, including the ranking of risk factors and the critical value of corresponding independent variables promoting or inhibiting falls;

[0010] S4, based on the risk factor identification result, targeted fall prevention intervention measures are developed for the corresponding community.

[0011] Further, in S1, the data processing comprises converting the collected questionnaire results into categorical variables or continuous variables, and calculating the body mass index of each elderly person and the weekly physical activity metabolic expenditure of each elderly person.

[0012] Further, the categorical variables comprise binary variables and multi-component variables.

[0013] The binary variables comprise gender, education level, monthly income, living alone, self-care ability, accessibility of leisure facilities, accessibility of service facilities, accessibility of public transport facilities, community safety, and falling.

[0014] The multi-component variables comprise sleep duration and fear of falling.

[0015] Further, the continuous variables comprise age, grip strength, balance ability, residence floor, residence duration, number of diseases, number of home environment defects, number of building environment defects, and number of road environment defects.

[0016] Further, S2 comprises dividing the training set and the test set, training and testing the falling risk identification model, and selecting the machine learning algorithm by the area under the curve, the accuracy, the sensitivity, or the specificity.

[0017] Further, S3 comprises:

[0018] S31, attributing the changes of the falling risk identification model to each independent variable by the SHAP method, and estimating the marginal contribution effect of each feature by the SHAP value;

[0019] S32, sorting the importance of the independent variables according to the SHAP value, and determining the falling risk factors.

[0020] S33, for the continuous variables, outputting the feature dependence graph of the SHAP value and fitting the curve, and identifying the threshold of falling.

[0021] Further, in S33, one continuous variable comprises one or more thresholds.

[0022] Further, S4 comprises formulating the targeted prevention measures according to the falling risk factors or the threshold of falling.

[0023] Therefore, the community elderly falling risk identification and intervention method based on machine learning has the following technical effects:

[0024] (1) The application provides an elderly falling prevention framework method that can be universally applied in different regions of the community, which can help the community level to master the falling risk characteristics of the elderly in the community and the environment defects that cause falling, and is helpful for the community to formulate targeted anti-falling strategies and measures.

[0025] (2) The application combines machine learning algorithms and SHAP methods to identify risk factors and explore variable thresholds, which helps to explore the non-linear relationship between risk factors and falls, and is superior to traditional binary logistic regression models.

[0026] The technical solutions of the application will be further described in detail below with the help of drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 is a flowchart of a community elderly fall risk identification and intervention method based on machine learning;

[0028] Figure 2 is a significance ranking diagram of different variables in an embodiment of the community elderly fall risk identification and intervention method based on machine learning, wherein a is a SHAP value scatter plot of the top 15 variables, and b is a mean value ranking diagram of the absolute value of the SHAP of the top 15 variables;

[0029] Figure 3 is a feature dependence diagram of continuous variables in an embodiment of the community elderly fall risk identification and intervention method based on machine learning, wherein a is balance ability, b is grip strength, c is the number of diseases, d is physical activity, e is body mass index BMI, and f is age. DETAILED DESCRIPTION

[0030] The application can be explained in more detail through the following examples, and the purpose of the disclosure is to protect all changes and improvements within the scope of the application, and the application is not limited to the following examples.

[0031] As shown in Figure 1 , the application provides a community elderly fall risk identification and intervention method based on machine learning, which specifically includes the following steps:

[0032] S1, data collection and processing: offline questionnaire survey is adopted, and the system collects social and economic, health status, perceived environment, and fall information data of the elderly.

[0033] In the process of offline sampling survey, the sample size is determined by the size of the community, the number of elderly population, the survey budget, etc., generally between 200-1000. At the same time, paper questionnaire is used to collect data, and auxiliary old people fill in to ensure the authenticity of the data. The questionnaire should include relevant topics: (1) Social economy: age, gender, education, monthly income, living alone; (2) Health status: height, weight, number of diseases, sleep duration, walking aid use, lower limb function, fear of falling, self-care ability (measured by Barthel index measurement table), 3 intensity physical activity days and duration; In addition, after filling out the questionnaire, the grip strength is tested using a grip strength meter, and the balance ability test scale in the Technical Guidelines for Elderly Fall Intervention published by the National Health Commission is used to score the balance ability of the elderly; (3) Perceived environment: home environment defects, building environment defects, road environment defects, accessibility of leisure facilities, accessibility of service facilities, accessibility of public transportation facilities, community safety, residence floor, residence length; (4) Fall information: whether falling in the past year, falling place, falling reason.

[0034] In the data processing process, the questionnaire results collected are input into Excel, and the initial content is converted into categorical variables or continuous variables during input to facilitate subsequent analysis and calculation. The binary variables, multiple classification variables, continuous variables, body mass index, and weekly physical activity metabolic expenditure of each elderly person are calculated, and a data set for training machine learning algorithms is constructed.

[0035] The binary variables of each old person include: gender (0 female, 1 male), education level (0 junior high school and below, 1 high school and above), monthly income (0 less than 5000 yuan per month, 1 5000 yuan and above per month), living alone (0 living alone, 1 not living alone), walking aid use (0 use, 1 not use), lower limb function (0 weak, 1 normal), self-care ability (0 need help, 1 no need help), accessibility of leisure facilities (0 poor, 1 good), accessibility of service facilities (0 poor, 1 good), accessibility of public transportation facilities (0 poor, 1 good), community safety (0 poor, 1 good), fall (0 no fall in the past year, 1 fall in the past year).

[0036] The multiple classification variables of each elderly person include: sleep duration (1 less than 6h, 2 6-9h per day, 3 more than 9h per day), whether afraid of falling (1 afraid, 2 a little afraid, 3 not afraid).

[0037] Continuous variables of each elderly person include: age (unit: years), grip strength (unit: kg), balance ability (unit: score), residence floor (unit: number of floors), residence length (unit: years), number of diseases (unit: number), number of home environment defects (unit: number), number of building environment defects (unit: number), and number of road environment defects (unit: number).

[0038] Body mass index (BMI) of each elderly person is calculated by the following formula:

[0039] BMI = W / H 2 ;

[0040] wherein, BMI is body mass index, W is body weight (kg), and H is height (m).

[0041] Weekly physical activity metabolic expenditure of each elderly person is calculated by the following formula:

[0042] PA = 8*t1*f1+4*t2*f2+3.3*t3*f3

[0043] wherein, PA is weekly physical activity metabolic expenditure (MET-min / week), t1, t2, and t3 are average daily exercise time of high intensity, moderate intensity, and low intensity (min / d) respectively, and f1, f2, and f3 are weekly exercise days of high intensity, moderate intensity, and low intensity (d / week) respectively.

[0044] S2, machine learning analysis: based on the collected data, a fall risk identification model is built by using machine learning algorithm.

[0045] Specifically, the processed data is divided into two parts for model testing: 70% of the training set and 30% of the test set; falls in the statistical data of S1 are taken as the dependent variable, and six machine learning algorithms Logistic Regression (LR), Support Vector Machine (SVM), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Categorical Boosting (CatBoost), and Random Forest (RF) are used to build classification models, the performance of the six algorithms is compared through four indicators of area under the curve (AUC), accuracy (Accuracy), sensitivity (Sensitivity), and specificity (Specificity), and the algorithm with the optimal AUC is preferentially selected for further analysis.

[0046] S3, ranking of fall risk factors: using the SHAP framework to interpret the built machine learning model, and obtaining the risk factor identification result.

[0047] S31, attributing the changes of the optimal model to each independent variable through the SHAP method, the difference in prediction of this independent variable in an independent variable set is the marginal contribution, the SHAP value can be used to estimate the marginal contribution effect of each feature, and the importance of the independent variable is quantified by the size of the SHAP value of each independent variable.

[0048] Among them, SHAP is a method for quantifying the contribution of each feature to the result of the machine learning model using game theory, which makes the black box model interpretable. The core of SHAP method is to quantify the "individual contribution" of each independent variable to the model prediction, which calculates the change of prediction value before and after adding an independent variable under different feature combinations (i.e. marginal contribution), and then weighs all possible combinations to get the SHAP value of the independent variable, which directly reflects the importance of the independent variable to the model output, and the calculation method is as follows:

[0049]

[0050] In the formula, φ p represents the contribution of feature p, x' represents a specific instance of feature vector set; Z' represents the non-zero subset in x';! represents factorial; P represents the total number of features; |Z'| represents the size of Z'; f x (Z') represents the output of the machine learning model when given the feature subset Z'; f x (Z' / P) represents the output of the machine learning model when given the feature subset Z' and removing the feature p.

[0051] It is worth noting that the same independent variable has its specific SHAP value in each sample, which represents the importance of the independent variable to the model output in each sample. SHAP value greater than 0 represents that the independent variable produces positive contribution to the result in a specific sample, and less than 0 represents negative contribution.

[0052] S32, according to the SHAP value, the relationship between the value of each independent variable and its influence on classification is identified, and the importance of the independent variable is ranked, the independent variable with higher ranking indicates that it plays a greater role in the classification of fall risk of the elderly in the community, that is, it is a more important fall risk factor.

[0053] For example Figure 2As shown in Figure a, taking the variable "balance ability" as an example, the figure shows a scatter plot of the SHAP value of this variable for each sample. The specific interpretation method is as follows: the color of the scatter points represents the magnitude of the variable value, with red points representing higher balance ability and blue points representing lower balance ability. Among them, the red points are mostly SHAP values ​​< 0, and the blue points are mostly SHAP values ​​> 0, indicating that the older people with higher balance ability are less likely to fall (that is, the higher the balance ability, the negative contribution to the fall prevention, i.e., less likely to fall); the rest can be inferred by analogy. Figure 2 Figure b shows the average absolute value of the SHAP values ​​of each variable in each sample, and the variables are ranked according to their importance based on this average value. The higher the value, the greater the role it plays in the fall classification process.

[0054] S33. Output the feature dependency graph of the SHAP value of each continuous variable and fit the curve to perform threshold identification. When the SHAP value is 0, the corresponding independent variable value is the critical value that the independent variable promotes or inhibits falls. An independent variable may contain multiple thresholds, which can reflect the non-linear relationship between the variable and falls (e.g., U-shaped relationship).

[0055] Figure 3 The diagram illustrates the correspondence between the specific values ​​of a particular independent variable and the SHAP values ​​in each sample, and curve fitting is performed. For example... Figure 3 As shown in Figure b, the independent variable is "grip strength": when the grip strength is 15 kg, the SHAP value is 0, indicating that its contribution to falls is zero, meaning it neither promotes nor inhibits falls; when the grip strength is less than 15 kg, the SHAP value is greater than 0, indicating that falls are promoted, and vice versa. Therefore, it can be concluded that a grip strength of 15 kg is the critical value for classifying falls in older adults. Furthermore, when two or more SHAP values ​​are 0, it indicates a significant non-linear relationship between the corresponding variable and falls.

[0056] In summary, compared to using machine learning alone, the SHAP method can effectively explain the relationship between independent and output variables, derive the importance ranking of independent variables, and explore potential nonlinear relationships between them. Traditional machine learning models, on the other hand, are typically black-box methods, only able to obtain outputs from inputs, but struggling to explore the specific connections between input and output variables.

[0057] S4. Analyze the risk factors and threshold identification results in S3, and propose fall prevention strategies or specific prevention recommendations based on the risk factors and thresholds. For example, if the results in S3 show that a grip strength of less than 15 kg is a factor that increases the risk of falls, the corresponding measure recommendation could be: It is recommended that older adults increase strength training to increase muscle mass and improve grip strength to more than 15 kg.

[0058] Therefore, the application adopts the above-mentioned community elderly fall risk identification and intervention method based on machine learning, and through the framework method of "investigation-analysis-intervention", different communities can output targeted fall prevention measures for the elderly population; at the same time, combined with machine learning and SHAP method, it is helpful to find the nonlinear relationship and threshold between risk factors and falls, and to provide fine response measures.

[0059] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, but not to limit them, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that: it can still modify or equivalently replace the technical solutions of the present application, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for identifying and intervening in the fall risk of elderly people in the community based on machine learning, characterized in that, Includes the following steps: S1. Through offline questionnaire surveys, the system collects data on the socioeconomic status, health status, perceived environment, and fall information of the elderly, and performs data processing and statistics. S2. Based on the processed data, using fall information data as the dependent variable, a fall risk identification model is built using machine learning algorithms. S3. The fall risk identification model is interpreted using the SHAP method to obtain the risk factor identification results, including the ranking of risk factors and the critical values ​​of corresponding independent variables that promote or inhibit falls. S4. Based on the risk factor identification results, develop targeted fall prevention intervention measures for the corresponding communities.

2. The method for identifying and intervening in the risk of falls among elderly people in the community based on machine learning, as described in claim 1, is characterized in that... In S1, data processing includes converting the collected questionnaire results into categorical or continuous variables, and calculating the body mass index and weekly physical activity metabolic expenditure for each elderly person.

3. The method for identifying and intervening in the risk of falls among elderly people in the community based on machine learning, as described in claim 2, is characterized in that... Categorical variables include binary variables and multi-component variables; Binary variables include gender, education level, monthly income, living alone status, self-care ability, accessibility of recreational facilities, accessibility of service facilities, accessibility of public transportation facilities, community safety, and falls; Multicategorical variables include sleep duration and fear of falling.

4. A method for identifying and intervening in the risk of falls among elderly people in the community based on machine learning, as described in claim 2, is characterized in that... Continuous variables include age, grip strength, balance ability, floor of residence, length of residence, number of illnesses, number of defects in home environment, number of defects in building environment, and number of defects in road environment.

5. A method for identifying and intervening in the risk of falls among elderly people in the community based on machine learning, as described in claim 1, is characterized in that... S2 involves dividing the data into training and testing sets, training and testing the fall risk identification model, and selecting the best machine learning algorithm based on metrics such as area under the curve, accuracy, sensitivity, or specificity.

6. A method for identifying and intervening in the risk of falls among elderly people in the community based on machine learning, as described in claim 1, characterized in that, S3 include: S31. Using the SHAP method, the changes in the fall risk identification model are attributed to each independent variable, and the marginal contribution effect of each feature is estimated using the SHAP value. S32. Rank the independent variables according to their importance based on their SHAP values ​​to determine the fall risk factors; S33. For continuous variables, output the feature dependency graph of SHAP values ​​and fit the curve to identify the threshold for falling.

7. A method for identifying and intervening in the risk of falls among elderly people in the community based on machine learning, as described in claim 6, is characterized in that... In S33, a continuous variable contains one or more thresholds.

8. A method for identifying and intervening in the risk of falls among elderly people in the community based on machine learning, as described in claim 6, is characterized in that... S4 includes developing targeted preventative measures based on fall risk factors or fall thresholds.

Citation Information

Patent Citations

  • Old people nutrition and health state assessment and risk prediction system based on machine learning

    CN114974570A

  • Method for identifying influence depression level by using machine learning

    CN118430822A

  • Fall risk assessment method and system

    CN118629644A