Method for constructing prediction model for subcutaneous fat hyperplasia caused by insulin injection
By constructing a machine learning model and combining it with ultrasound examination data to screen key predictive variables, the problem of early identification of lipomatosis in diabetic patients was solved, enabling accurate identification and timely intervention of high-risk patients, and improving the accuracy and interpretability of diagnosis.
Patent Information
- Application Number
- CN202510946458.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
In the current technology, lipomatosis is difficult to identify in diabetic patients at an early stage, traditional diagnostic methods underestimate the prevalence, and there is a lack of effective predictive models to guide prevention and intervention.
A machine learning-based model for predicting subcutaneous fat hyperplasia induced by insulin injection was constructed. Random forest, extreme gradient boosting, and logistic regression algorithms were used, combined with ultrasound examination data, to screen out key predictive variables and interpret the model's impact using the SHAP method. The best-performing model was then selected for prediction.
It enables accurate identification of high-risk liposuction, allowing for timely intervention, reducing the burden on patients, and improving the accuracy of diagnosis and the interpretability of predictive models.
Smart Images

Figure CN120809209A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and in particular to a method for constructing a subcutaneous fat hyperplasia prediction model caused by insulin injection. BACKGROUND
[0002] Fat hyperplasia is the most common skin complication in patients receiving insulin therapy, which is mainly manifested as hypertrophy of adipocytes, leading to nodular swelling and tissue hardening at the injection site. Early identification and intervention of fat hyperplasia is of great significance for diabetic patients receiving insulin therapy, which can help optimize blood glucose control and reduce the economic burden of this population. Therefore, timely detection of fat hyperplasia by a reliable method is of great significance for effective blood glucose management.
[0003] Currently, in clinical practice, fat hyperplasia is usually identified by visual inspection and palpation. However, larger flat fat hyperplasia lesions may be missed even by experienced clinicians. Our team's previous study showed that the prevalence of subclinical fat hyperplasia undetected by visual or palpation was 19.9%. In recent years, ultrasonography has been recommended as a reliable diagnostic tool for fat hyperplasia, which can show detailed tissue features, including size, depth, and morphological structure. Although a systematic review of 26865 diabetic patients reported a global prevalence of fat hyperplasia of 41.8%, studies using ultrasonography showed a prevalence as high as 86.5%, indicating that traditional diagnostic methods may severely underestimate this condition. Although the exact pathogenesis of fat hyperplasia is not yet clear, most of it is considered to be preventable, so identifying high-risk patients or situations is crucial for guiding the prevention of fat hyperplasia. SUMMARY
[0004] To solve the above problems, the present application aims to provide a method for constructing a subcutaneous fat hyperplasia prediction model caused by insulin injection. By using three machine learning algorithms, random forest, extreme gradient boosting, and logistic regression, a model for predicting fat hyperplasia and identifying risk factors is constructed. The best performance of the extreme gradient boosting machine learning model is selected, and the constructed extreme gradient boosting machine learning model shows high efficiency in predicting the occurrence of fat hyperplasia in diabetic patients, which can accurately identify high-risk diabetic patients with fat hyperplasia, thereby achieving timely and targeted intervention on the occurrence of fat hyperplasia to reduce its impact.
[0005] The above technical problems are solved by the following technical solutions: The application discloses a method for constructing an insulin injection subcutaneous fat hyperplasia prediction model, which comprises the following steps: S1, selecting a plurality of participants according to inclusion criteria to conduct a questionnaire survey, the inclusion criteria including: patients who have been diagnosed with diabetes, are 18 years old or above, and have continuously used insulin for more than three months; the questionnaire survey including three parts of social demographic information, diabetes-related information and insulin injection behavior; S2, performing ultrasonic examination on each injection site of insulin of all participants and non-injection sites symmetrical to the injection sites, and performing detailed evaluation on subcutaneous tissue thickness, echo intensity, vascularization and dermal and subcutaneous tissue boundary according to the ultrasonic examination results, then performing fat hyperplasia diagnosis according to the evaluation results, and obtaining modeling group data; S3, selecting 22 prediction variables related to fat hyperplasia from a plurality of prediction factors, and classifying the prediction variables into three categories; then randomly dividing the modeling group data into a training group and an internal validation group at a ratio of 8:2, continuously performing five times, and performing feature screening on the 22 prediction variables by using LASSO regression on the training group in each time; S4, developing fat hyperplasia prediction models by using three algorithms of random forest, extreme gradient boosting and logistic regression on the training groups randomly divided for five times, respectively, then evaluating model performance by using average area under the curve, selecting a fat hyperplasia prediction model with the best performance for research, and verifying external validity of the model by using an external validation group; and S5, explaining the influence of each prediction variable screened out on the fat hyperplasia prediction model by using a SHAP method based on the fat hyperplasia prediction model with the best performance, and visually displaying the mutual relationship between the plurality of prediction variables screened out.
[0006] The method can obtain the prediction variable most significantly affecting the occurrence of fat hyperplasia in the diabetic patients by selecting the fat hyperplasia prediction model with the best performance from the three prediction fat hyperplasia models, so that the occurrence of fat hyperplasia in the diabetic patients can be effectively prevented, and the influence can be reduced.
[0007] Preferably, in step S1, the social demographic information at least includes age, gender, height, weight, education level, residence and working state; the diabetes-related information at least includes diabetes type, disease duration, insulin use type, insulin treatment duration and insulin dose; and the insulin injection behavior at least includes needle replacement frequency, injection area, needle length and insulin storage mode. The questionnaire survey on the social demographic information, the diabetes-related information and the insulin injection behavior of the participants can obtain a plurality of factors that may affect fat hyperplasia, so that the accuracy of the fat hyperplasia prediction model can be improved.
[0008] Preferably, in step S2, the basis for the diagnosis of fat proliferation is that the lesion is located in the subcutaneous tissue and meets at least four of the following five characteristics: high echo point or nodular morphology, low echo halo, echo texture heterogeneity, surrounding connective tissue deformation, no vascularization and no capsule structure. By diagnosing the participants with fat proliferation based on multiple diagnostic criteria, the accuracy of the diagnosis can be ensured, and reliable data basis is provided for subsequent evaluation of the fat proliferation prediction model.
[0009] Preferably, in step S3, the three types of prediction variables include: demographic data, diabetes-related risk factors and insulin injection behavior. By classifying the prediction variables into three categories, the effects of prediction variables of different dimensions on fat proliferation can be determined, and confounding analysis can be avoided.
[0010] Preferably, in step S3, the plurality of prediction factors at least include: needle reuse frequency, diabetes type, insulin continuous treatment time, insulin dose and insulin injection area. By selecting 22 prediction variables related to fat proliferation from a plurality of prediction factors, the performance of the prediction model can be optimized, and the interpretability of the results can be enhanced.
[0011] Preferably, in step S3, when performing LASSO regression, a normalization method is used to preprocess the data of the 22 prediction variables, and the influence of each prediction variable on fat proliferation is scaled to the range of [0, 1]. By normalizing the prediction variables, the scale difference between different prediction variables can be reduced, and the feature selection can be accelerated.
[0012] Preferably, in step S4, the external validation group data collected is different from the hospital and time period of the modeling group data, and there is no overlapping patient between the external validation group and the modeling group data. By collecting the external validation group data, the performance of the finally selected fat proliferation prediction model can be further tested, the applicability of the model in a specific population can be verified, and overfitting of the training data can be avoided.
[0013] Preferably, in step S5, the SHAP method calculates the marginal contribution of each prediction variable to the fat proliferation prediction model to obtain the SHAP value of each prediction variable, and quantifies the positive and negative effects of each prediction variable on the fat proliferation prediction model according to the SHAP value. By using the SHAP method, the prediction results of the fat proliferation prediction model can be visually explained, which helps to better understand the significance of the fat proliferation prediction model.
[0014] The present application has the following beneficial effects compared with the prior art: The technical solution of the present invention uses LASSO regression to perform feature screening on multiple predictor variables, and adopts three machine learning algorithms, random forest, extreme gradient boosting, and logistic regression, to construct a model for predicting adipose hyperplasia and identify risk factors; then the model performance is evaluated by the average area under the curve, and for the best-performing adipose hyperplasia prediction model, its SHAP value is calculated to illustrate the positive and negative impact of each feature on the model prediction results, thereby accurately identifying diabetic patients with high risk of adipose hyperplasia, and implementing timely and targeted intervention in the occurrence of adipose hyperplasia to mitigate its impact. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A flowchart of a method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection; Figure 2 is the correlation heat map of the predictor variables; Figure 3 A plot of the area under the curve for fat hyperplasia in the extreme gradient boosting learning model test dataset. Figure 4 This is the importance ranking diagram of the seven predictive variables and SHAP values of fat hyperplasia; Figure 5 Importance ranking diagram of the seven predictive variables of lipohypertrophy. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0017] like Figure 1 Figure 2 is a flowchart of a method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection. First, multiple participants were selected for a questionnaire survey based on the inclusion criteria, and ultrasound doctors were arranged to perform ultrasound examinations on all participants. LASSO regression was then used to perform feature screening on multiple predictor variables. Three algorithms, random forest, extreme gradient boosting, and logistic regression, were used to develop a fat hyperplasia prediction model, and the fat hyperplasia prediction model with the best performance was selected for research. Finally, the SHAP method was used to explain the influence of the predictor variables on the fat hyperplasia prediction model, so that the best-performing fat hyperplasia prediction model could be used to accurately identify diabetic patients with high risk of fat hyperplasia.
[0018] The method specifically comprises the following steps: S1, selecting a plurality of participants according to the inclusion criteria to conduct a questionnaire survey, wherein 1095 participants are selected in the present application, the inclusion criteria include: patients who have been diagnosed with diabetes, are 18 years old or older, and have been using insulin continuously for more than three months, and are excluded from patients who have dermatitis, skin diseases, and scars or wounds at the insulin injection site; the questionnaire survey includes three parts of social demographic information, diabetes-related information and insulin injection behavior, and the social demographic information at least includes: age, gender, height, weight, education level, residence and working status; the diabetes-related information at least includes: diabetes type, disease duration, insulin use type, insulin treatment duration and insulin dose; the insulin injection behavior at least includes: needle replacement frequency, injection area, needle length and insulin storage method; by conducting a questionnaire survey on the participants' social demographic information, diabetes-related information and insulin injection behavior, a variety of factors that may affect lipohypertrophy can be obtained, thereby improving the accuracy of the lipohypertrophy prediction model.
[0019] S2, performing ultrasonic examination on each injection site of insulin and non-injection site symmetrical to the injection site of all participants, and performing detailed evaluation of subcutaneous tissue thickness, echo intensity, vascularization and dermal and subcutaneous tissue boundary according to the results of ultrasonic examination, then performing lipohypertrophy diagnosis according to the evaluation results, and obtaining modeling group data.
[0020] Specifically, the non-injection site symmetrical to the injection site is subjected to ultrasonic examination, and the "symmetrical" means that: if a diabetic patient injects insulin on the left lateral thigh, the corresponding part of the right lateral thigh is the symmetrical part thereof, and the ultrasonic examination is used as an individual control for evaluating the injection site of insulin, so as to increase the accuracy of the lipohypertrophy diagnosis.
[0021] Among them, the basis for diagnosing lipohypertrophy is that the lesion needs to be located in the subcutaneous tissue and meet at least four of the following five characteristics: high echo point or nodular morphology with low echo halo, which means that the density of the subcutaneous tissue of the lesion is higher than that of the surrounding normal tissue or lower than that of the surrounding normal tissue; echo texture heterogeneity, which means that the echo texture of the subcutaneous tissue of the lesion is not uniform compared with the surrounding tissue; surrounding connective tissue deformation, which means that the surrounding connective tissue of the subcutaneous tissue of the lesion is deformed; no vascularization, which means that the subcutaneous tissue of the lesion has no blood vessel formation and blood flow signal; no envelope structure, which means that the subcutaneous tissue of the lesion has no boundary or envelope; by diagnosing lipohypertrophy of the participants according to a plurality of diagnostic bases, the accuracy of the diagnosis can be ensured, and reliable data basis is provided for subsequent evaluation of the lipohypertrophy prediction model.
[0022] S3, according to the questionnaire results, fat hyperplasia diagnosis results and previous research data, 22 prediction variables related to fat hyperplasia are selected from multiple prediction factors, and the 22 prediction variables are risk factors for causing fat hyperplasia in diabetic patients, and the prediction variables are divided into three categories: demographic data, diabetes-related risk factors and insulin injection behavior.
[0023] As shown in FIG. 4, it is a correlation heat map of prediction variables, which shows that there is a significant correlation between multiple prediction variables, so the data of 22 prediction variables is preprocessed by normalization method, and the influence of each prediction variable on fat hyperplasia is scaled to the range of [0, 1], then the modeling group data is randomly divided into training group and internal validation group according to the ratio of 8:2, and the process is repeated for 5 times, and LASSO regression is used for feature selection of 22 prediction variables in each training group, and 7 prediction variables with the strongest prediction ability are selected at least 3 times in 5 times of feature selection, including: insulin duration, body mass index, needle reuse frequency, correct site rotation, insulin injection area, insulin dose and needle length. Figure 2
[0024] Among them, the multiple prediction factors include needle reuse frequency, diabetes type, insulin continuous treatment time, insulin dose and insulin injection area; the demographic data includes age, gender and body mass index; the diabetes-related risk factors include duration, diabetes type and insulin dose; the insulin injection behavior includes needle replacement frequency, injection area, needle length and insulin storage method, so that the specific influencing factors of fat hyperplasia can be obtained, so that the diabetic patients can prevent in time and targeted.
[0025] In step S3 of the embodiment, LASSO regression, i.e. Least Absolute Shrinkage and Selection Operator Regression, is a linear regression model, which adds an L1 regularization term, i.e. the sum of model coefficients, to the objective function to compress the model coefficients, so that some coefficients are reduced to 0, so that multiple prediction variables with the strongest prediction ability are selected from 22 prediction variables, and the complexity of the fat hyperplasia prediction model is reduced, the risk of overfitting of the fat hyperplasia prediction model is reduced, and the accuracy of the fat hyperplasia prediction model is improved.
[0026] S4, the modeling group data is randomly divided into a training group and an internal validation group in a ratio of 8:2, which are respectively used for model training and unknown data testing, and three algorithms of random forest, extreme gradient boosting and logistic regression are used to develop a fat hyperplasia prediction model respectively, and the performance of the model in the training group, the internal validation group and the external validation group is evaluated by the average area under the curve, and the best performance fat hyperplasia prediction model is selected for research.
[0027] Among them, the external validation group data collected is different from the hospital and time period of the modeling group data, and there is no overlapping patient between the external validation group and the modeling group data; through the collected external validation group data, the performance of the finally selected fat hyperplasia prediction model can be further tested, and the applicability of the model in a specific population is verified, so as to avoid overfitting the training data.
[0028] As shown in the following table, it is the test data of three kinds of machine learning models, from the area under the curve value, in the training group, the prediction performance of the extreme gradient boosting machine learning model and the random forest machine learning model is better, the average area under the curve value is 0.9943, and the average area under the curve value of the logistic regression machine learning model is 0.8087; in the internal validation group, the prediction performance of the logistic regression machine learning model is the best, the average area under the curve value is 0.7873, and the average area under the curve values of the random forest machine learning model and the extreme gradient boosting machine learning model are 0.7767 and 0.7683 respectively; in the external validation group, the prediction performance of the extreme gradient boosting machine learning model is the best, the average area under the curve value is 0.7542, and the average area under the curve values of the logistic regression machine learning model and the random forest machine learning model are 0.7521 and 0.7471 respectively.
[0029]
[0030] As Figure 3 As shown in the following table, it is the test data of three kinds of machine learning models, from the area under the curve value, in the training group, the prediction performance of the extreme gradient boosting machine learning model and the random forest machine learning model is better, the average area under the curve value is 0.9943, and the average area under the curve value of the logistic regression machine learning model is 0.8087; in the internal validation group, the prediction performance of the logistic regression machine learning model is the best, the average area under the curve value is 0.7873, and the average area under the curve values of the random forest machine learning model and the extreme gradient boosting machine learning model are 0.7767 and 0.7683 respectively; in the external validation group, the prediction performance of the extreme gradient boosting machine learning model is the best, the average area under the curve value is 0.7542, and the average area under the curve values of the logistic regression machine learning model and the random forest machine learning model are 0.7521 and 0.7471 respectively.
[0031] In the present application, by calculating the average area under the curve value of the three kinds of machine learning models, that is, the average value of the area under the curve value of the training group and the validation group, it is found that the average area under the curve value of the external validation group of the extreme gradient boosting machine learning model is the highest, therefore the extreme gradient boosting machine learning model is selected as the prediction model of fat hyperplasia of insulin-treated diabetic patients.
[0032] S5, based on the extreme gradient boosting machine learning model, the SHAP method is used to calculate the SHAP value corresponding to each prediction variable, and the influence of each prediction variable screened out on the extreme gradient boosting machine learning model is explained through the SHAP value, and the mutual relationship between the plurality of prediction variables screened out is visualized.
[0033] As shown in Figure 4 , it is the importance ranking diagram of 7 prediction variables and SHAP values of fat hyperplasia, as shown in Figure 5 , it is the importance ranking diagram of 7 prediction variables of fat hyperplasia, combined with Figure 4 and Figure 5 , the influence of 7 prediction variables on the prediction model of fat hyperplasia is shown, and 7 prediction variables are arranged in order of importance according to the absolute value of their average SHAP value from high to low, wherein, from high to low, in order: correct site rotation, insulin dose, body mass index, insulin duration, insulin injection area, needle reuse frequency and needle length.
[0034] In this embodiment, in step S5, the SHAP method, i.e. SHapley Additive exPlanations, is a method for explaining the prediction results of machine learning model, by calculating the marginal contribution of each prediction variable to the extreme gradient boosting machine learning model, obtaining the SHAP value of each prediction variable, and quantifying the positive and negative influence of each prediction variable on the extreme gradient boosting machine learning model through the SHAP value, wherein the higher the SHAP value of the prediction variable, the greater the possibility of fat hyperplasia, so that the SHAP value can be used to explain the prediction results of the extreme gradient boosting machine learning model, which helps to better understand the significance of the extreme gradient boosting machine learning model.
[0035] In summary, the present application uses LASSO regression to screen a plurality of prediction variables, and uses random forest, extreme gradient boosting and logistic regression three kinds of machine learning algorithms to construct the model for predicting fat hyperplasia and identify risk factors; then the average area under the curve is used to evaluate the model performance, and for the fat hyperplasia prediction model with the best performance, the SHAP value is calculated to clarify the positive and negative influence degree of each feature on the model prediction result, so that the high-risk fat hyperplasia diabetic patients can be accurately identified, and timely and targeted intervention can be realized for the occurrence of fat hyperplasia to reduce its influence, which has significant progress.
[0036] The above embodiments only illustrate the technical idea of the present application, and cannot limit the protection scope of the present application. Any modification made on the basis of the technical solution according to the technical idea of the present application falls within the protection scope of the present application.
Claims
1. A method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection, characterized in that: The steps include: S1. We selected multiple participants for a questionnaire survey based on the inclusion criteria: patients diagnosed with diabetes, aged ≥18 years, and taking insulin for more than three months. The questionnaire included sociodemographic information, diabetes-related information, and insulin injection behavior. S2. Perform ultrasound examinations on all participants at each insulin injection site and at non-injection sites symmetrical to the injection site. A detailed assessment of subcutaneous tissue thickness, echogenicity, vascularization, and the dermal-subcutaneous tissue boundary will be performed based on the ultrasound examination results. A diagnosis of adipose hyperplasia will be made based on the assessment results, and modeling group data will be obtained. S3. Select 22 predictive variables associated with lipohypertrophy from multiple predictive factors and categorize these predictive variables into three categories. Then, randomly divide the modeling data into a training group and an internal validation group in an 8:2 ratio for five consecutive times. LASSO regression is used to screen the 22 predictive variables for features in each training group. S4. 5 randomly divided training groups were used to develop three fat hyperplasia prediction models, namely random forest, extreme gradient boosting, and logistic regression. The model performance was evaluated by the mean area under the curve. The best performing fat hyperplasia prediction model was selected for research and the external validity of the model was verified by an external validation group. S5. Based on the best-performing adipose hyperplasia prediction model, the SHAP method was used to explain the influence of each screened predictive variable on the adipose hyperplasia prediction model, and the relationships between the multiple screened predictive variables were visualized.
2. The method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection according to claim 1, characterized in that: In step S1, sociodemographic information includes at least: age, gender, height, weight, education level, place of residence and work status; diabetes-related information includes at least: diabetes type, course of disease, type of insulin used, duration of insulin treatment and insulin dosage; insulin injection behavior includes at least: needle replacement frequency, injection area, needle length and insulin storage method.
3. The method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection according to claim 1, characterized in that: In step S2, the diagnosis of fat hyperplasia is based on the following: the lesion must be located in the subcutaneous tissue and meet at least four of the five characteristics of high echo points or nodules with low echo halos, heterogeneous echo texture, deformation of surrounding connective tissue, lack of vascularization, and lack of capsule structure.
4. The method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection according to claim 1, characterized in that: In step S3, the three types of predictor variables include: demographic data, diabetes-related risk factors, and insulin injection behavior.
5. The method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection according to claim 1, characterized in that: In step S3, the multiple prediction factors include at least: needle reuse frequency, diabetes type, insulin treatment duration, insulin dosage, and insulin injection area.
6. The method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection according to claim 1, characterized in that: In step S3, when performing LASSO regression, the 22 predictor variables are first preprocessed using a normalization method to scale the effect of each predictor variable on fat hyperplasia to the range of [0, 1].
7. The method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection according to claim 1, characterized in that: In step S4, the hospitals and time periods of the collected external validation group data and the modeling group data are different, and there are no overlapping patients in the external validation group and the modeling group data.
8. The method for constructing a prediction model for subcutaneous fat hyperplasia induced by insulin injection according to claim 1, characterized in that: In step S5, the SHAP method calculates the marginal contribution of each predictor variable to the fat hyperplasia prediction model to obtain the SHAP value of each predictor variable, and quantifies the positive and negative effects of each predictor variable on the fat hyperplasia prediction model according to the SHAP value.