Diabetic complication typing system and method based on clustering pseudo labels

Through a system based on clustered pseudo-label, the problem of diabetic complication evaluation and classification is solved, and the accurate typing and grading of diabetic complications is achieved, providing a basis for early prevention and treatment.

CN120089339APending Publication Date: 2025-06-03BEIJING HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510170915.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate and type complications of diabetes, making it difficult to achieve early prevention and treatment.

Method used

A system based on clustering pseudo-labeling is adopted to realize the classification and evaluation of diabetes complications through data collection, outcome attribute clustering, pseudo-label generation, classification model training and classification modules.

Benefits of technology

Accurate typing and grading of diabetes complications has been achieved, providing a scientific basis for the early prevention and treatment of diabetes complications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089339A_ABST
    Figure CN120089339A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical data processing, in particular to a diabetic complication typing system and method based on clustering pseudo tags. According to the method, physical examination attribute data related to diabetic complication typing is obtained through weighted kmeans clustering, K-means hierarchical clustering is carried out in combination with outcome attribute data of diabetic complication of coronary heart disease, cerebral apoplexy and fatty liver to obtain a complication typing clustering result set, and each patient in the data set is endowed with a pseudo tag. And training the classification model to obtain a diabetic complication typing model, so as to realize typing and grading of the diabetic complication and provide a basis for early prevention and treatment of the diabetic complication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical data processing, and particularly relates to a diabetes complication classification system and method based on clustering pseudo-labels. Background Art

[0002] Diabetes is a chronic disease marked by hyperglycemia, which is caused by absolute or relative insulin secretion deficiency and utilization disorders. This disease is mainly divided into three types: type 1, type 2, and gestational diabetes. The causes are mainly attributed to the combined effects of genetic factors and environmental factors, including the decline in insulin secretion caused by pancreatic islet cell dysfunction, or the body's insensitivity to insulin action or both, resulting in the ineffective utilization and storage of glucose in the blood. There is a disease aggregation phenomenon in some diabetic patients and their families. In addition, the incidence and prevalence of diabetes are on the rise globally.

[0003] Diabetes complications are various complications that may occur after a long-term diabetes, including cardiovascular diseases, neuropathy, eye diseases, kidney diseases, etc. Due to the long-term instability of blood sugar in diabetic patients, various tissues and organs will be damaged, leading to the occurrence of complications. Therefore, it is very important to prevent and treat diabetes complications early. The main symptoms of diabetes complications depend on the affected tissues or organs. The symptoms of cardiovascular diseases include chest pain, arrhythmia, and fatigue; neuropathy may cause paresthesia, pain, and muscle weakness; eye diseases are mainly manifested as visual impairment and eye infections; the typical symptoms of kidney diseases are increased urine output and proteinuria.

[0004] About 10 years after the onset of diabetes, 30% - 40% of patients will develop at least one complication. More than 50% of diabetic patients die from cardiovascular complications; for those with type 1 diabetes with a history of more than 15 years, the prevalence of retinopathy is 98%; for those with type 2 diabetes with a history of more than 15 years, the prevalence of retinopathy is 78%; about 10% of diabetic patients die from kidney lesions; the incidence of diabetic foot is about 2.6% in type 1 diabetes and about 5.2% in type 2 diabetes; the number of amputated patients is 10 - 20 times that of non-diabetic patients. Hyperosmolar hyperglycemic state is more common in elderly type 2 diabetic patients over 60 years old: more than 50% of patients may have diabetic neuropathy.

[0005] Therefore, the evaluation of diabetes complications can prevent and treat related complications as early as possible. Summary of the Invention

[0006] Aiming at the deficiencies of the above-mentioned existing technologies, the present invention aims to provide a diabetes complication classification system and method based on clustering pseudo-labels to achieve the classification and evaluation of diabetes complications and provide a basis for the prevention and treatment of related complications as early as possible.

[0007] To solve the above problems, the present invention adopts the following technical solutions: On the one hand, the present invention provides a diabetes complication typing system based on clustering pseudo-labels, including a data collection module, an outcome attribute clustering module, a pseudo-label generation module, a classification model training module, and a typing module; The data collection module is used to collect the physical examination attribute data related to the diabetes complication typing of patients and the outcome attribute data of whether they have diabetes complications such as coronary heart disease, stroke, and fatty liver, and form a data set; The outcome attribute clustering module is used to use the K-means hierarchical clustering algorithm to cluster the physical examination attribute data and the outcome attribute data to obtain a set of complication typing clustering results; The pseudo-label generation module is used to construct a consensus matrix according to the set of complication typing clustering results. Each element in the consensus matrix represents the frequency that two data points are assigned to the same cluster in multiple base clusterers, and cluster the consensus matrix to assign pseudo-labels of outcome attributes to each piece of data in the data set; The classification model training module is used to train based on the data set with pseudo-labels, using the pseudo-labels as supervision information, and adopt a classification algorithm to obtain a diabetes complication typing model; The typing module is used to predict the typing of the diabetes complications of the patient to be estimated according to the physical examination attribute data of the patient to be estimated, using the diabetes complication typing model.

[0008] As an implementable manner, the physical examination attribute data is obtained by weighted kmeans clustering of the data of patients without diabetes at baseline.

[0009] As an implementable manner, the physical examination attribute data includes alanine aminotransferase, triglyceride, waist circumference, body mass index, total cholesterol, uric acid, resting systolic blood pressure, heart rate, high-density lipoprotein cholesterol, age, blood urea nitrogen, gender, and fasting blood glucose.

[0010] As an implementable manner, it further includes a risk prediction module; the risk prediction module is used to calculate the risk curves under different categories of each diabetes complication according to the data set with pseudo-labels through the Nelson-Aalen model; calculate the risk of the patient to be estimated through the risk curves under different categories of each diabetes complication.

[0011] As an implementable manner, the classification algorithm includes support vector machine, random forest, or gradient boosting decision tree.

[0012] On the other hand, the present invention provides a typing method for non-disease diagnosis purposes of diabetes complications based on clustering pseudo-labels, including: Collect the physical examination attribute data related to the classification of diabetic complications of the patient and the outcome attribute data of whether the patient has coronary heart disease, stroke, and fatty liver diabetic complications to form a data set; Use the K-means hierarchical clustering algorithm to cluster according to the physical examination attribute data and outcome attribute data to obtain a set of clustering results for complication classification; Construct a consensus matrix based on the set of clustering results for complication classification. Each element in the consensus matrix represents the frequency that two data points are assigned to the same cluster in multiple base clusterers. Cluster the consensus matrix and assign pseudo-labels of outcome attributes to each piece of data in the data set; Based on the data set with pseudo-labels assigned, use the pseudo-labels as supervised information and adopt a classification algorithm for training to obtain a diabetic complication classification model; According to the physical examination attribute data of the patient to be estimated, use the diabetic complication classification model to predict and classify the diabetic complications of the patient to be estimated.

[0013] As an implementable manner, the physical examination attribute data is obtained by performing weighted kmeans clustering on the data of patients without diabetes at the baseline.

[0014] As an implementable manner, the physical examination attribute data includes alanine aminotransferase, triglyceride, waist circumference, body mass index, total cholesterol, uric acid, resting systolic blood pressure, heart rate, high-density lipoprotein cholesterol, age, blood urea nitrogen, gender, and fasting blood glucose.

[0015] As an implementable manner, it further includes a risk prediction module; the risk prediction module is used to calculate the risk curves under different categories of each diabetic complication according to the data set with pseudo-labels assigned through the Nelson-Aalen model; calculate the risk of the patient to be estimated through the risk curves under different categories of each diabetic complication.

[0016] As an implementable manner, the classification algorithm includes support vector machine, random forest, or gradient boosting decision tree.

[0017] The beneficial effects of the present invention are as follows: The present invention obtains the physical examination attribute data related to the classification of diabetic complications through weighted kmeans clustering, combines the outcome attribute data of whether the patient has coronary heart disease, stroke, and fatty liver diabetic complications, performs K-means hierarchical clustering to obtain a set of clustering results for complication classification, assigns pseudo-labels to each patient in the data set, and trains a classification model to obtain a diabetic complication classification model, realizing the classification and grading of diabetic complications, and providing a basis for the early prevention and treatment of diabetic complications. Description of the Drawings

[0018] Figure 1Schematic diagram of a diabetes complication classification system based on clustering pseudo-labels of the present invention.

[0019] Figure 2 Schematic diagram of a clustering graph with a low risk of the present invention developing into diabetes in the future.

[0020] Figure 3 Schematic diagram of a clustering with a high risk of the present invention developing into diabetes in the future and a relatively high risk of having FLD.

[0021] Figure 4 Schematic diagram of a clustering with a medium risk of the present invention developing into diabetes in the future and relatively high risks of having Stroke and CVD.

[0022] Figure 5 Violin plot of the SGPT1 physical examination attribute data of the present invention in three classes.

[0023] Figure 6 Violin plot of the TC1 physical examination attribute data of the present invention in three classes.

[0024] Figure 7 Violin plot of the XL physical examination attribute data of the present invention in three classes.

[0025] Figure 8 Violin plot of the BMI physical examination attribute data of the present invention in three classes.

[0026] Figure 9 Violin plot of the GENDERCODE physical examination attribute data of the present invention in three classes.

[0027] Figure 10 Violin plot of the RSBP physical examination attribute data of the present invention in three classes.

[0028] Figure 11 Violin plot of the TG physical examination attribute data of the present invention in three classes.

[0029] Figure 12 Violin plot of the age physical examination attribute data of the present invention in three classes.

[0030] Figure 13 Violin plot of the BUN1 physical examination attribute data of the present invention in three classes.

[0031] Figure 14 Violin plot of the FBGDL11 physical examination attribute data of the present invention in three classes.

[0032] Figure 15 Violin plot of the HDLC1 physical examination attribute data of the present invention in three classes.

[0033] Figure 16Violin plot of the NIAOSUAN physical examination attribute data of the present invention in three categories.

[0034] Figure 17 Violin plot of the WAIST physical examination attribute data of the present invention in three categories.

[0035] Figure 18 Nelson-Aalen risk curve graph of coronary heart disease of the present invention.

[0036] Figure 19 Nelson-Aalen risk curve graph of stroke of the present invention.

[0037] Figure 20 Nelson-Aalen risk curve graph of fatty liver of the present invention. Detailed implementation manners

[0038] The present invention will be further described in detail below in conjunction with specific embodiments.

[0039] It should be noted that these embodiments are only used to illustrate the present invention, rather than limiting the present invention. Any simple improvement of the method under the premise of the concept of the present invention belongs to the scope protected by the present invention.

[0040] See Figure 1 , a diabetes complication classification system based on clustering pseudo-labels, including a data acquisition module 100, an outcome attribute clustering module 200, a pseudo-label generation module 300, a classification model training module 400, and a classification module 500.

[0041] The data acquisition module 100 is used to collect the physical examination attribute data related to the diabetes complication classification of patients and the outcome attribute data of whether they have diabetes complications such as coronary heart disease, stroke, and fatty liver, and form a data set.

[0042] The physical examination attribute data is obtained by weighted kmeans clustering of the data of patients without diabetes at baseline. Specifically, it includes: (1) Data acquisition: Collect the clinical data of the patient cohort without diabetes at baseline (using the patient's first physical examination record as the baseline, and screening patients with a pre-meal blood glucose not exceeding 7 at the baseline physical examination).

[0043] (2) Data screening: Index screening: Remove the ID class attributes, index attributes with a missing value exceeding 50%, eliminate the nominal attributes, and retain the numerical and categorical attributes.

[0044] (3) Patient sample processing: For each index, remove the outliers outside 3 standard deviations and the extreme values below 0.1% and above 99.9% in the sample.

[0045] Normalization of continuous variables and encoding of categorical variables.

[0046] After index screening and patient sample processing, a dataset containing 13,609 patients without diabetes at baseline was obtained. After weighted kmeans clustering, more than 13 clustering metrics / physical examination attributes were obtained.

[0047] Physical examination attribute data: SGPT1: Alanine aminotransferase, TG: Triglyceride, WAIST: Waist circumference, BMI: Body mass index, TC1: Total cholesterol, NIAOSUAN: Uric acid, RSBP: Resting systolic blood pressure, XL: Heart rate, HDLC1: High-density lipoprotein cholesterol, age: Age, BUN1: Blood urea nitrogen, GENDERCODE: Gender, FBGDL11: Fasting blood glucose.

[0048] Alanine aminotransferase (ALT): Mainly present in the liver, when the liver is damaged, the ALT level in the blood will rise. It is an important indicator for detecting liver function.

[0049] Triglyceride: (TG): A lipid in the blood, high TG levels are associated with an increased risk of cardiovascular disease.

[0050] Waist circumference (WAIST): Used to evaluate the degree of abdominal obesity and is associated with the risk of multiple metabolic diseases.

[0051] Body mass index: (BMI): A commonly used indicator to measure the degree of body fatness and health, calculated as weight (kg) divided by the square of height (m).

[0052] Total cholesterol (TC): Refers to the total amount of cholesterol carried by all lipoproteins in the blood. High levels of total cholesterol are associated with an increased risk of cardiovascular disease.

[0053] Uric acid (UA): Excessively high levels of uric acid in the blood can cause problems such as gout and kidney stones.

[0054] Resting systolic blood pressure (RSBP): The pressure in the blood vessels during heart contraction, which is one of the important indicators for measuring hypertension.

[0055] Heart rate (HR): Refers to the number of times the heart beats per minute, and the resting heart rate is the heart rate measured when a person is in a quiet state.

[0056] High-density lipoprotein cholesterol (HDLC): Higher HDL-C levels help remove low-density lipoproteins (bad cholesterol) from the blood, thereby reducing the risk of cardiovascular disease.

[0057] Age (AGE): A basic demographic variable that is associated with many health conditions and disease risks.

[0058] Blood Urea Nitrogen (BUN): The end product of protein metabolism, mainly excreted from the body through the kidneys. An elevated BUN level may indicate renal insufficiency.

[0059] Gender code, used to distinguish between males and females, and gender also plays an important role in the risk assessment of some diseases.

[0060] Fasting Blood Glucose (FBG): Refers to the blood glucose level measured after at least 8 hours without consuming any calorie-containing food or drink. The fasting blood glucose level is a key indicator for diagnosing diabetes and evaluating blood glucose control.

[0061] The outcome attribute data of the diabetes complications studied: Coronary heart disease, stroke, fatty liver.

[0062] Due to the instability of the clustering algorithm, even minor differences in different category outcome attributes may change in the clustering results of different rounds. And for different outcome attributes, we may attach different importance.

[0063] In response to this, here we use the method of weighted kmeans ensemble clustering to pay different degrees of attention to different outcome attributes, and reduce the clustering instability and improve the overall clustering quality through the method of integrating multiple clusterings. See Figures 2 - 4 .

[0064] The outcome attribute clustering module 200 is used to cluster the physical examination attribute data and the outcome attribute data by using the K-means hierarchical clustering algorithm to obtain a set of clustering results for complication classification types.

[0065] General clustering methods belong to unsupervised learning and do not effectively utilize the subsequent disease development of patients (outcome attributes) to assist in category division. For the classification of disease subtypes, only indicators can be selected through prior knowledge for clustering. The different categories clustered may have no difference or only minor differences in outcome attributes.

[0066] In response to this, we use the method of generating pseudo-labels based on outcome attributes and then classifying with physical examination indicators to effectively generate different pseudo-labels according to outcome attributes, assist in the division of disease subtypes, and make the different disease subtypes have greater differences in outcome attributes.

[0067] The pseudo-label generation module 300 is used to construct a consensus matrix based on the set of clustering results for complication classification types. Each element in the consensus matrix represents the frequency of two data points being assigned to the same cluster in multiple base clusterers, and cluster the consensus matrix to assign pseudo-labels of outcome attributes to each data point in the dataset.

[0068] See Figures 5 - 17 For the 13 physical examination indicators of the present invention, there are good differences in the distributions among the three categories, and all of them pass the Kruskal-Wallis test (p < 0.05).

[0069] The classification model training module 400 is used to train based on the dataset with pseudo-labels assigned, taking the pseudo-labels as supervision information, and adopting a support vector machine, a random forest, or a gradient boosting decision tree to obtain a diabetes complication classification model.

[0070] Evaluate the model performance through methods such as cross-validation, and adjust the model parameters or reselect features according to the evaluation results to improve the accuracy and generalization ability of the model.

[0071] The classification module 500 is used to predict and classify the diabetes complications of the patient to be evaluated by using the diabetes complication classification model according to the physical examination attribute data of the patient to be evaluated.

[0072] The system of the present invention further includes a risk prediction module 600; the risk prediction module 600 is used to calculate the risk curves of different categories of each diabetes complication through the Nelson-Aalen model according to the dataset with pseudo-labels assigned; calculate the risk of the patient to be evaluated through the risk curves of different categories of each diabetes complication.

[0073] The estimation of the Nelson-Aalen model of the cumulative risk rate function is calculated by the following formula:

[0074] where is the number of individuals who died at time (j = 1,..., r), is the number of individuals who survived before time .

[0075] See Figures 18 - 20 For the present invention, the regression cumulative risk curves of different categories of each diabetes complication are calculated through the Nelson-Aalen risk function model.

[0076] The present invention also provides a classification method for diabetes complications for non-disease diagnosis purposes based on clustering pseudo-labels, including: S100. Collect the physical examination attribute data related to the diabetes complication classification of the patient and the outcome attribute data of whether the patient has coronary heart disease, stroke, and fatty liver diabetes complications to form a dataset; S200. Adopt the K-means hierarchical clustering algorithm to cluster according to the physical examination attribute data and the outcome attribute data to obtain a set of complication classification clustering results; S300. Construct a consensus matrix based on the set of clustering results of the complication classification. Each element in the consensus matrix represents the frequency that two data points are assigned to the same cluster in multiple base clusterers. Cluster the consensus matrix to assign a pseudo-label of the outcome attribute to each data in the dataset. S400. Based on the dataset with pseudo-labels assigned, using the pseudo-labels as supervised information, train with a support vector machine, random forest, or gradient boosting decision tree to obtain a diabetes complication classification model. S500. According to the physical examination attribute data of the patient to be estimated, use the diabetes complication classification model to predict and classify the diabetes complications of the patient to be estimated.

[0077] The method of the present invention further includes S600. Calculate the cumulative regression risk curve for different categories of each diabetes complication based on the dataset with pseudo-labels assigned through the Nelson-Aalen model; calculate the risk of the patient to be estimated through the cumulative regression risk curves for different categories of each diabetes complication.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described by referring to the preferred embodiments of the present invention, those of ordinary skill in the art should understand that various changes can be made in form and details without departing from the spirit and scope of the present invention defined by the appended claims.

Claims

1. A diabetes complication classification system based on clustering pseudo-labels, characterized in that: It includes data collection module, outcome attribute clustering module, pseudo-label generation module, classification model training module and typing module; The data collection module is used to collect the patient's physical examination attribute data related to the classification of diabetes complications and the outcome attribute data of whether the patient suffers from coronary heart disease, stroke and fatty liver diabetes complications to form a data set; The outcome attribute clustering module is used to obtain a set of complication classification clustering results by clustering the physical examination attribute data and the outcome attribute data using a K-means hierarchical clustering algorithm; The pseudo-label generation module is used to construct a consensus matrix based on the complication classification clustering result set, each element in the consensus matrix represents the frequency of two data points being assigned to the same cluster in multiple base clusterers, clustering the consensus matrix, and assigning a pseudo-label of the outcome attribute to each data in the data set; The classification model training module is used to obtain a diabetes complication classification model based on a dataset assigned with pseudo labels and using pseudo labels as supervision information and a classification algorithm for training; The classification module is used to predict and classify the diabetic complications of the patient to be evaluated based on the physical examination attribute data of the patient to be evaluated and using the diabetic complications classification model.

2. The diabetes complication classification system based on clustering pseudo labels according to claim 1, characterized in that: The physical examination attribute data is obtained by performing weighted kmeans clustering on the data of patients without diabetes at baseline.

3. The diabetes complication classification system based on clustering pseudo labels according to claim 2, characterized in that: The physical examination attribute data include alanine aminotransferase, triglycerides, waist circumference, body mass index, total cholesterol, uric acid, resting systolic blood pressure, heart rate, high-density lipoprotein cholesterol, age, blood urea nitrogen, gender and fasting blood glucose.

4. The diabetes complication classification system based on clustering pseudo labels according to claim 1, characterized in that: It also includes a risk prediction module; the risk prediction module is used to calculate the risk curves of different categories of diabetic complications through the Nelson-Aalen model according to the data set after the pseudo labels are assigned; and the risk of the patient to be assessed is calculated through the risk curves of different categories of diabetic complications.

5. The diabetes complication classification system based on clustering pseudo labels according to claim 1, characterized in that: The classification algorithms include support vector machines, random forests, or gradient boosted decision trees.

6. A classification method for diabetes complications based on clustered pseudo-labels for non-disease diagnosis purposes, characterized in that: include: Collect the patient's physical examination attribute data related to the classification of diabetes complications and the outcome attribute data of coronary heart disease, stroke and fatty liver complications to form a data set; Using K-means hierarchical clustering algorithm to cluster the physical examination attribute data and outcome attribute data to obtain a set of complication classification clustering results; A consensus matrix is ​​constructed according to the complication classification clustering result set, wherein each element in the consensus matrix represents the frequency of two data points being assigned to the same cluster in multiple base clusterers, the consensus matrix is ​​clustered, and a pseudo label of the outcome attribute is assigned to each data in the data set; Based on the dataset with pseudo labels, the pseudo labels are used as supervision information and the classification algorithm is used for training to obtain a diabetes complication classification model. The diabetes complications classification model is used to predict and classify the diabetes complications of the patient to be evaluated based on the physical examination attribute data of the patient to be evaluated.

7. The method for typing diabetes complications for non-disease diagnosis purposes based on clustered pseudo-labels according to claim 6, characterized in that: The physical examination attribute data is obtained by performing weighted kmeans clustering on the data of patients without diabetes at baseline.

8. The method for typing diabetes complications for non-disease diagnosis purposes based on clustered pseudo-labels according to claim 7, characterized in that: The physical examination attribute data include alanine aminotransferase, triglycerides, waist circumference, body mass index, total cholesterol, uric acid, resting systolic blood pressure, heart rate, high-density lipoprotein cholesterol, age, blood urea nitrogen, gender and fasting blood glucose.

9. The method for typing diabetes complications for non-disease diagnosis purposes based on clustered pseudo-labels according to claim 6, characterized in that: It also includes a risk prediction module; the risk prediction module is used to calculate the risk curves of different categories of diabetic complications through the Nelson-Aalen model according to the data set after the pseudo labels are assigned; and the risk of the patient to be assessed is calculated through the risk curves of different categories of diabetic complications.

10. The method for typing diabetes complications for non-disease diagnosis purposes based on clustered pseudo-labels according to claim 6, characterized in that: The classification algorithms include support vector machines, random forests, or gradient boosted decision trees.