Risk assessment method and system for type 2 diabetes complications and storage medium

By constructing a two-dimensional clustering tree and using a competitive risk model, the problem that a single clinical indicator in the prior art cannot accurately evaluate the risk of type 2 diabetes complications is solved, and higher evaluation accuracy and computing efficiency are achieved.

CN120183705AActive Publication Date: 2025-06-20DATA SPACE RES INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510645724.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The prior art cannot accurately assess the risk of complications in patients with type 2 diabetes through a single clinical indicator, and cannot capture the potential link between multiple clinical indicators.

Method used

The DDRTree algorithm is used to construct a two-dimensional clustering tree, and the position coordinates of the patient to be evaluated are obtained based on the clinical index data of the patient to be evaluated, and the probability of him suffering from complications in the next r years is calculated using a competitive risk model.

Benefits of technology

It improves the accuracy of risk assessment of type 2 diabetes complications, simplifies the expression of clinical indicator data, improves computational efficiency, and takes into account time factors to make the probability of disease assessment more scientific.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183705A_ABST
    Figure CN120183705A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of risk assessment of diabetic complications, and particularly relates to a risk assessment method and system for type 2 diabetic complications and a storage medium. The risk assessment method comprises the following steps: S1, based on clinical index data of a patient in a first training set, constructing a two-dimensional clustering tree by using a DDRTree algorithm; the two-dimensional clustering tree comprises a plurality of clustering areas, each clustering area comprises a plurality of nodes, and each node represents a patient; s2, based on clinical index data of a to-be-evaluated patient, using a mapping function to obtain position coordinates of the to-be-evaluated patient in the two-dimensional clustering tree; and S3, on the basis of the position coordinates of the patient to be assessed, calculating the probability that the patient to be assessed suffers from various complications in the next r years by using a competitive risk model. According to the method, the risk that the type 2 diabetes patients suffer from complications can be accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of risk assessment of diabetes complications, and particularly relates to a method, system and storage medium for risk assessment of type 2 diabetes complications. Background Art

[0002] Type 2 diabetes is a chronic metabolic disease, and its complications include diabetic retinopathy, heart failure, stroke, myocardial infarction, end-stage renal disease, chronic kidney disease, metabolism-related fatty liver, liver cirrhosis, and diabetic peripheral vascular disease, etc. The development process of these complications is intricate, seriously affecting the quality of life and prognosis of patients. Therefore, risk assessment of type 2 diabetes complications can not only perform preventive interventions on the living habits of patients, but also provide targeted prognosis for type 2 diabetes patients, improving the prognosis effect.

[0003] The existing technology usually evaluates the risk of complications in type 2 diabetes patients based on a single clinical index, such as blood glucose or blood pressure, etc., using simple statistical methods or a single machine learning algorithm.

[0004] However, type 2 diabetes complications often affect multiple clinical indexes, and there are also some potential connections between the clinical indexes under type 2 diabetes that technicians cannot directly capture; moreover, different complications may make the same clinical index show the same change trend. For example, after long-term follow-up of type 2 diabetes patients with high blood pressure, it is found that only a part of the patients develop diabetic foot within ten years. Therefore, the existing technology cannot accurately evaluate the risk of complications in type 2 diabetes patients through a single clinical index.

[0005] Therefore, how to improve the accuracy of risk assessment of type 2 diabetes complications has become an urgent problem to be solved. Summary of the Invention

[0006] The purpose of the present invention is to overcome the above-mentioned deficiencies of the existing technology and provide a method for risk assessment of type 2 diabetes complications, which can accurately evaluate the risk of complications in type 2 diabetes patients.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions: A method for risk assessment of type 2 diabetes complications, comprising the following steps: S1, based on the clinical index data of patients in the first training set, use the DDRTree algorithm to construct a two-dimensional clustering tree; the two-dimensional clustering tree contains several clustering regions, each clustering region contains several nodes, and each node represents a patient; S2, based on the clinical index data of the patient to be evaluated, use a mapping function to obtain the position coordinates of the patient to be evaluated in the two-dimensional clustering tree; S3. Based on the position coordinates of the patient to be evaluated, use the competing risks model to calculate the probabilities of the patient to be evaluated having various complications within the next r years.

[0008] Preferably, after S1 and before S2, it further includes S1': In S1', after the technical staff annotate the common characteristics of the clinical indicators of each clustering region, use the mapping function to obtain the position coordinates of the patients in the validation set in the two-dimensional clustering tree. If the consistency rate between the patients in the validation set and the common characteristics of the clinical indicators of the corresponding clustering region is greater than the first threshold, it is determined that the current two-dimensional clustering tree is successfully verified; otherwise, the verification fails. After discarding the current two-dimensional clustering tree, increase the number of patients in the first training set and return to S1 again.

[0009] Preferably, in S1, it further includes the following sub-steps: S11. Obtain the feature matrix X based on the first training set. The first training set contains the data of m clinical indicators of n patients; the n rows of the feature matrix X respectively represent n patients, and the m columns respectively represent m clinical indicators; both n and m are positive integers. S12. After constructing the optimization objective function F of the DDRTree algorithm based on the feature matrix X and solving it, obtain the key parameters of the two-dimensional clustering tree: ; Among them, the key parameters are the optimal coordinate matrix and the optimal set of central nodes; Y represents the coordinate matrix. The coordinate matrix Y contains n rows and 2 columns. The n rows of the coordinate matrix Y respectively represent n patients, and are the same as the patients represented by the corresponding rows in the feature matrix X. The first column of the coordinate matrix Y represents the abscissa, and the second column represents the ordinate; Y(f) represents the position coordinates corresponding to the f-th row in the coordinate matrix Y; W represents the linear transformation parameter matrix in the DDRTree algorithm; represents the transpose of the linear transformation parameter matrix W; λ represents the first hyperparameter; represents the position coordinates of the central node i in the current two-dimensional clustering tree; represents the position coordinates of the central node j in the current two-dimensional clustering tree; and are both position coordinates in the coordinate matrix Y; ε represents the edge set of the minimum spanning tree in the current two-dimensional clustering tree; Z represents the set of central node position coordinates; α represents the second hyperparameter; represents the position coordinates of the central node closest to the position coordinates Y(f); represents the coordinate matrix Y and the set of central nodes Z when taking the minimum value. At this time, the coordinate matrix Y and the set of central nodes Z are the optimal coordinate matrix and the optimal set of central nodes; represents the square of the 2-norm; S13. Construct a two-dimensional clustering tree according to the key parameters.

[0010] Preferably, in S1´, the following content is further included: S11´, technicians label the common characteristics of the clinical indicators of each clustering region according to the clinical indicators of the patients in each clustering region in the current two-dimensional clustering tree; S12´, use the mapping function to obtain the position coordinates of each patient in the two-dimensional clustering tree in the validation set; S13´, determine the clustering region where each patient in the validation set is located according to the position coordinates of each patient in the two-dimensional clustering tree in the validation set; S14´, after determining whether the clinical indicator characteristics of each patient in the validation set are consistent with the common characteristics of the corresponding clustering region, calculate the consistency rate CR = NUM(C) / NUM(All); where NUM(·) represents the quantity; NUM(C) represents the number of patients whose clinical indicator characteristics in the validation set are consistent with the common characteristics of the corresponding clustering region, and NUM(All) represents the total number of patients in the validation set; If the consistency rate CR is greater than the first threshold, it is determined that the current two-dimensional clustering tree is successfully verified; if the consistency rate CR is below the first threshold, it is determined that the current two-dimensional clustering tree verification fails. After discarding the current two-dimensional clustering tree, construct a new first training set or increase the number of patients in the first training set, and then return to S1 again.

[0011] Preferably, use a machine learning model to obtain the position coordinates of patient G in the two-dimensional clustering tree based on the mapping function ; The mapping function is: ; ; Among them, represents the ordinate of the position coordinates of patient G in the two-dimensional clustering tree; represents the abscissa of the position coordinates of patient G in the two-dimensional clustering tree; represents the first global bias parameter; represents the second global bias parameter; represents the age regression coefficient; represents the gender regression coefficient; represents the age of patient G; represents the gender parameter of patient G; represents the data of the p-th clinical indicator of patient G; represents the k-th power of; represents the intermediate quantity of the p-th clinical indicator of patient G, and max(·) represents taking the maximum value, Data representing the p-th clinical index of the patient corresponding to node d in the two-dimensional clustering tree. There are n nodes in total in the two-dimensional clustering tree.

[0012] Preferably, in S3, u undiagnosed complications of the patient to be evaluated are evaluated, where u is a positive integer, including the following: Based on the position coordinates of the patient to be evaluated R, the competing risk model calculates the probability that the patient to be evaluated will develop the V-th complication within the next r years of , where r ≥ 0 and V is a positive integer less than or equal to u: ; ; where t represents the t-th year in the future for the patient to be evaluated R starting from the current moment; represents the instantaneous probability that the patient to be evaluated R will develop a complication in the t-th year in the future of represents the first regression coefficient; represents the second regression coefficient; represents the ordinate of the position coordinates of the patient to be evaluated R in the two-dimensional clustering tree; represents the abscissa of the position coordinates of the patient to be evaluated R in the two-dimensional clustering tree; represents the complication of the baseline subdistribution hazard function varying with the future time t.

[0013] Preferably, before using the competing risk model to calculate the probabilities of various complications that the patient to be evaluated will develop within the next r years, the competing risk model is trained using the second training set; the second training set includes e pieces of training data, and each piece of training data corresponds to 1 patient who has already had more than one complication. Each piece of training data contains the time when the patient was diagnosed with type 2 diabetes, the times when the patient was diagnosed with various complications, and the clinical index data at the time of the most recent diagnosis of a complication; e is a positive integer.

[0014] Preferably, the clinical index data of the patients in the first training set, the validation set, and the second training set are all cleaned clinical index data. Cleaning the clinical index data also includes the following: Step 1, extract the clinical index data of several patients. If a certain patient has more than one piece of dirty data, then all the clinical index data of the current patient are excluded; dirty data refers to: among the same clinical index data of several patients, the clinical index data located outside 5 standard deviations. Step 2, transform the remaining clinical index data in Step 1 into clinical index data within the range of 0 to 1 through the rank normalization method. At this time, the cleaning of the clinical index data is completed.

[0015] The present invention also provides a risk assessment system for type 2 diabetes complications, including: a two-dimensional clustering tree construction module, a position coordinate mapping module, and a risk assessment module; the two-dimensional clustering tree construction module contains the DDRTree algorithm, which constructs a two-dimensional clustering tree based on the clinical index data of patients in the first training set and sends the constructed two-dimensional clustering tree into the position coordinate mapping module; the position coordinate mapping module is used to convert the clinical index data of the patient to be evaluated into position coordinates on the two-dimensional clustering tree and send the position coordinates of the patient to be evaluated into the risk assessment module; the risk assessment module contains a competing risks model, which calculates the probabilities of the patient to be evaluated having various complications within the next r years based on the position coordinates of the patient to be evaluated and outputs them; each module is configured to execute the steps of a risk assessment method for type 2 diabetes complications as described above.

[0016] The present invention also provides a computer-readable storage medium: the computer-readable storage medium stores a computer program programmed or configured to execute a risk assessment method for type 2 diabetes complications as described above.

[0017] The beneficial effects of the present invention are as follows: (1) The risk assessment method for type 2 diabetes complications of the present invention comprehensively evaluates the probability of a patient having a certain complication in the future based on various clinical characteristic data of the patient to be evaluated, improving the accuracy of the assessment.

[0018] (2) The two-dimensional clustering tree constructed by the present invention greatly simplifies the representation form of the patient's clinical index data without losing the information represented by the patient's clinical index data, improving the calculation efficiency during the assessment; and the position coordinates of the nodes in the two-dimensional clustering tree also contain the potential connections between various clinical index data of the patient himself; and the distribution of the nodes in the two-dimensional clustering tree also contains the potential connections between various clinical index data of patients.

[0019] (3) After constructing the two-dimensional clustering tree, the present invention also verifies the current two-dimensional clustering tree using a validation set, ensuring that only the two-dimensional clustering tree with sufficient generalization and accuracy will be used for subsequent complication probability assessment, further improving the accuracy of the complication probability assessment.

[0020] (4) The present invention uses a competing risks model to calculate the probabilities of various complications that a patient to be evaluated will develop within the next r years based on a two-dimensional clustering tree. The two-dimensional clustering tree constructed by the present invention presents this information in the form of simple two-dimensional position coordinates while ensuring the integrity of the patient's clinical index information and covering potential connection information. For those skilled in the art: Those skilled in the art do not care about the two-dimensional position coordinates of the two-dimensional clustering tree and the vast amount of information represented in the clustering tree structure, and only expect to know the final probability of developing complications. Therefore, whether those skilled in the art understand all the information covered by the two-dimensional clustering tree has no impact on the final evaluation of the probability of developing complications by the competing risks model. For the competing risks model: It simplifies the input of the competing risks model. By only using the two-dimensional position coordinates of the patient to be evaluated as the input of the model, the competing risks model can receive various clinical index information and potential connection information of the patient to be evaluated represented by the two-dimensional position coordinates. While improving the evaluation accuracy of the competing risks model, it greatly reduces the computational cost of the competing risks model and improves the evaluation efficiency.

[0021] (5) When the present invention evaluates the probability that a patient will develop a certain complication within the next r years, considering that the probability of developing the disease will increase over time, time is integrated during the evaluation process, making the evaluation result of the probability of developing the disease more scientific.

[0022] (6) The present invention can evaluate the probabilities of various complications that a patient to be evaluated will develop within the next r years. Therefore, based on the patient's current various clinical index data, the patient can evaluate the probabilities of developing different complications at different times in the future, which is very flexible and convenient. And regularly conducting risk assessments of complications is also beneficial for providing feedback to doctors and patients on the intervention and prognosis effects of type 2 diabetes. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of a method for risk assessment of type 2 diabetes complications of the present invention; Figure 2 is a schematic diagram of the two-dimensional clustering tree constructed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0024] To make the technical solutions of the present invention clearer and more definite, the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All solutions obtained by equivalent substitution of the technical features of the technical solutions of the present invention by those of ordinary skill in the art and conventional reasoning fall within the protection scope of the present invention.

[0025] The "patients" mentioned in the present invention are all patients who have been diagnosed with type 2 diabetes, and will not be repeatedly explained hereinafter.

[0026] Such asFigure 1 As shown in the figure, it is a flowchart of a risk assessment method for type 2 diabetes complications of the present invention, including the following steps: S1. Based on the clinical index data of patients in the first training set, use the DDRTree algorithm to construct a two-dimensional clustering tree; the two-dimensional clustering tree contains several clustering regions, each clustering region contains several nodes, and each node represents a patient; S2. Based on the clinical index data of the patient to be evaluated, use the mapping function to obtain the position coordinates of the patient to be evaluated in the two-dimensional clustering tree; S3. Based on the position coordinates of the patient to be evaluated, use the competing risks model to calculate the probabilities of the patient to be evaluated having various complications within the next r years.

[0027] In S1, the following sub-steps are further included: S11. Obtain the feature matrix X based on the first training set. The first training set contains the data of m clinical indexes of n patients; each row of the feature matrix X represents the m clinical indexes of a patient, and each column represents a clinical index; both n and m are positive integers.

[0028] The feature matrix X is an n×m matrix, and the elements in the matrix represent the data of a certain clinical index of a certain patient.

[0029] The clinical index data of the patients in the first training set comes from the Chinese Nephropathy Database and is the clinical index data of the patients in the first year after being diagnosed with type 2 diabetes.

[0030] In this embodiment, n = 26000; m = 10, and they are respectively high-density lipoprotein cholesterol, triglyceride, systolic blood pressure, alanine aminotransferase, glycated hemoglobin, low-density lipoprotein cholesterol, creatinine, heart rate, body mass index, and diastolic blood pressure.

[0031] The patient age and age are known data, but we do not need to use them when constructing the training set and the feature matrix X.

[0032] S12. After constructing the optimization objective function F of the DDRTree algorithm based on the feature matrix X and solving it, obtain the key parameters of the two-dimensional clustering tree: ; Among them, the key parameters are the optimal coordinate matrix and the optimal central node set; Y represents the coordinate matrix. The coordinate matrix Y contains n rows and 2 columns. Each row of the coordinate matrix Y represents a patient in the first training set and is the same as the patient represented by the corresponding row in the feature matrix X. The first column of the coordinate matrix Y represents the abscissa, and the second column represents the ordinate; Y(f) represents the f-th row in the coordinate matrix Y, which is a position coordinate; W represents the linear transformation parameter matrix in the DDRTree algorithm; represents the transpose of the linear transformation parameter matrix W; λ represents the first hyperparameter, which is set by technicians according to experience and is used to control the compactness between nodes in the two-dimensional clustering tree; represents the position coordinate of the central node of i in the current two-dimensional clustering tree, and is also a position coordinate in the coordinate matrix Y; represents the position coordinate of the central node of j in the current two-dimensional clustering tree, and is also a position coordinate in the coordinate matrix Y; ε represents the edge set of the minimum spanning tree in the current two-dimensional clustering tree, which is automatically generated by the DDRTree algorithm in the process of learning the potential relationships between the clinical indicators of patients in the first dataset; Z represents the set of central node position coordinates; α represents the second hyperparameter, which is set by technicians according to experience and is used to control the degree to which non-central nodes in the two-dimensional clustering tree approach the central nodes; represents the position coordinate of the central node closest to the position coordinate Y(f); represents the coordinate matrix Y and the central node set Z corresponding to taking the minimum value. At this time, the coordinate matrix Y and the central node set Z are the optimal coordinate matrix and the optimal central node set; represents the square of the 2-norm.

[0033] S13. Construct a two-dimensional clustering tree according to the key parameters.

[0034] The schematic diagram of the two-dimensional clustering tree is as Figure 2 shown, Figure 2 Each gray dot in it represents a patient in the first training set and is also a node in the two-dimensional clustering tree.

[0035] In the DDRTree algorithm, constructing a two-dimensional clustering tree according to the key parameters is a prior art and will not be elaborated here.

[0036] Optionally, after S1 and before S2, it further includes S1': S1'. After technicians mark the common characteristics of the clinical indicators in each clustering region, use the mapping function to obtain the position coordinates of patients in the validation set in the two-dimensional clustering tree. If the consistency rate between the patients in the validation set and the common characteristics of the clinical indicators in the corresponding clustering region is greater than the first threshold, it is determined that the current two-dimensional clustering tree is successfully verified; otherwise, the verification fails. After discarding the current two-dimensional clustering tree, increase the number of patients in the first training set and return to S1 again.

[0037] The set of m clinical indicators of patient G is , represents the data of the p-th clinical indicator of patient G, where 1 ≤ p ≤ m and p is a positive integer.

[0038] Use a machine learning model based on the mapping function to obtain the position coordinates of patient G in the two-dimensional clustering tree ; The mapping function is as follows: ; ; wherein, represents the ordinate of the position coordinate of patient G in the two-dimensional clustering tree; represents the abscissa of the position coordinate of patient G in the two-dimensional clustering tree; represents the first global bias parameter; represents the second global bias parameter, and are known quantities within the machine learning model and are generated by the machine learning model itself; represents the age regression coefficient; represents the gender regression coefficient; represents the age of patient G; represents the gender parameter of patient G. Different gender parameters correspond to male and female respectively, and the gender parameter is set by the technical staff; represents the data of the p-th clinical index of patient G; represents to the k-th power of; represents the intermediate quantity of the p-th clinical index of patient G. max(·) represents taking the maximum value, represents the data of the p-th clinical index of the patient corresponding to node d in the two-dimensional clustering tree. There are n nodes in total in the two-dimensional clustering tree.

[0039] In S1´, the following content is also included: S11´, the technical staff mark the common characteristics of the clinical indexes of each clustering region according to the clinical indexes of the patients in each clustering region in the current two-dimensional clustering tree.

[0040] For example, the common characteristics of the clinical indexes in the first clustering region are high body mass index (body mass index > 30) and high blood sugar (fasting blood sugar > 9.8 mmol / L), and the common characteristics of the clinical indexes in the second clustering region are high high-density lipoprotein cholesterol (> 2.6 mmol / L) and high heart rate (resting heart rate > 100 beats / minute).

[0041] After the two-dimensional clustering tree is generated, the range of each clustering region in the current two-dimensional clustering tree is determined.

[0042] S12´, use the mapping function to obtain the position coordinates of each patient in the validation set in the two-dimensional clustering tree.

[0043] S13´, determine the clustering region where each patient in the validation set is located according to the position coordinates of each patient in the validation set in the two-dimensional clustering tree.

[0044] As long as the position coordinates of a patient fall within a certain clustering region of the current two-dimensional clustering tree, this region is the clustering region where the current patient is located.

[0045] S14´. After determining whether the clinical index characteristics of each patient in the validation set are consistent with the common characteristics of the corresponding clustering region, calculate the consistency rate CR = NUM(C) / NUM(All), where NUM(·) represents the quantity, NUM(C) represents the number of patients in the validation set whose clinical index characteristics are consistent with the common characteristics of the corresponding clustering region, and NUM(All) represents the total number of patients in the validation set. If the consistency rate CR is greater than the first threshold, it is determined that the current two-dimensional clustering tree is successfully verified; if the consistency rate CR is below the first threshold, it is determined that the current two-dimensional clustering tree verification fails. After discarding the current two-dimensional clustering tree, construct a new first training set or increase the number of patients in the first training set, and then return to S1.

[0046] The clinical index data of the patients in the validation set also come from the Chinese Nephropathy Database, and they are also the clinical index data of the patients in the first year after the diagnosis of type 2 diabetes, but there is no overlap between the patients in the validation set and those in the first training set. In this embodiment, the number of patients in the validation set is 6,501.

[0047] In the present invention, S1 and its sub-steps construct a two-dimensional clustering tree; while S1´ and its sub-steps are used to verify whether the structure of the current two-dimensional clustering tree has sufficient generalization and accuracy. If the structure of the two-dimensional clustering tree has sufficient generalization and accuracy, the clustering region where the patients in the validation set fall through the mapping function must have consistent clinical index characteristics with the patients in the validation set. Even if there are inconsistencies, they are extremely small. Therefore, in S14´, the consistency rate is finally calculated to determine whether the structure of the current two-dimensional clustering tree has sufficient generalization and accuracy. If the consistency rate is too low, it means that the generalization and accuracy of the current two-dimensional clustering tree are insufficient and cannot be used for subsequent calculations. Therefore, in the present invention, in this case, the first training set is reconstructed or the number of patients in the original first training set is increased (i.e., the sample size is increased), and then the two-dimensional clustering tree is reconstructed based on the new first training set until the consistency rate is greater than the first threshold.

[0048] In S1´, S12´ and S2, the position coordinates of the patients in the two-dimensional clustering tree are obtained using the mapping function, which has been described above and will not be elaborated here.

[0049] Figure 2 The red dots in represent the positions of the patients to be evaluated in the two-dimensional clustering tree.

[0050] Before using the competing risks model to calculate the probabilities of various complications that the patient to be evaluated may develop within the next r years, the competing risks model is trained using a second training set, which further includes the following: The second training set includes e pieces of training data. Each piece of training data corresponds to a patient who has already developed more than one complication. Each piece of training data contains the time when the patient was diagnosed with type 2 diabetes, the times when various complications were diagnosed, and the clinical index data at the time of the most recent diagnosis of a complication; e is a positive integer.

[0051] In this embodiment, e = 48000.

[0052] The clinical index data and complication data of the patients in the second training set also come from the Chinese Nephrology Database.

[0053] The complications of type 2 diabetes mainly include: myocardial infarction (MI), stroke (including ischemic and hemorrhagic), heart failure (HF), metabolic associated fatty liver disease (MAFLD), liver cirrhosis, diabetic retinopathy (DR), chronic kidney disease (CKD), end-stage renal disease (ESRD), and diabetic peripheral vascular disease (DPD).

[0054] The clinical index data of the patients in the first training set, the validation set, and the second training set are all cleaned clinical index data. Cleaning the clinical index data further includes the following: Step 1: Extract the clinical index data of a number of patients; if a patient has more than one piece of dirty data, all the clinical index data of the current patient will be excluded; dirty data refers to the clinical index data that is outside 5 standard deviations among the same type of clinical index data of a number of patients. Step 2: Transform the remaining clinical index data in Step 1 into clinical index data within the range of 0 to 1 through the rank normalization method, and at this time the cleaning of the clinical index data is completed.

[0055] Cleaning the clinical index data can exclude significantly abnormal clinical index data. Even if a very small amount of clinical index data with a relatively low degree of abnormality is retained, the impact on other normal clinical index data in the future will be reduced due to rank normalization.

[0056] In this embodiment, the clinical index data of the patient to be evaluated is the clinical index data within the last 3 months.

[0057] In S3, only u types of complications that the patient to be evaluated has not been diagnosed with are evaluated, where u is a positive integer, and it further includes the following: The competing risks model calculates the probability that the patient to be evaluated will develop the Vth complication within the next r years based on the position coordinates of the patient to be evaluated R of , where r ≥ 0 and V is a positive integer less than or equal to u: ; ; Where t represents the future t-year of the patient R to be evaluated starting from the current moment; Indicates that the patient R to be evaluated will have complications in the next t years The instantaneous probability of represents the first regression coefficient; represents the second regression coefficient; The ordinate represents the position coordinate of the patient R to be evaluated in the two-dimensional clustering tree; The horizontal coordinate represents the position coordinate of the patient R to be evaluated in the two-dimensional clustering tree; Indicates complications The benchmark sub-distribution risk function that changes with future time t; , as well as It is obtained automatically by the competing risk model during the training and optimization process.

[0058] In this embodiment, r=10.

[0059] The risk assessment method for type 2 diabetes complications of the present invention is based on multiple clinical characteristic data of the patient to be assessed to comprehensively assess the probability of the patient suffering from certain complications in the future, thereby improving the accuracy of the assessment.

[0060] The present invention first uses the DDRTree algorithm to construct a two-dimensional clustering tree based on various clinical indicator data of patients in the first training set, converts patients carrying multiple clinical indicator data into nodes in the two-dimensional clustering tree, and then we only need to process the position coordinates of the nodes, so as to achieve the dimensionality reduction of the patient's multidimensional information. The dimensionality reduction of this information is not to directly delete the information of certain dimensions, but to integrate and compress the multidimensional information, so that the DDRTree algorithm not only excavates the potential connection between the multiple clinical indicator data of each patient in the training process, but also learns the potential connection between the various clinical indicator data between patients, and then uses the two-dimensional node clustering method to reflect the potential connection between the multiple clinical indicators between patients in the two-dimensional plane through the angle of node distribution. That is, the two-dimensional clustering tree constructed by the present invention greatly simplifies the expression form of the patient's clinical indicator data without losing the information represented by the patient's clinical indicator data, and improves the calculation efficiency in the evaluation process; and the node position coordinates in the two-dimensional clustering tree also include the potential connection between the patient's own multiple clinical indicator data excavated; and the distribution of nodes in the two-dimensional clustering tree also includes the potential connection between the various clinical indicator data between patients.

[0061] After constructing the two-dimensional clustering tree, the present invention also verifies the current two-dimensional clustering tree using a validation set, ensuring that only two-dimensional clustering trees with sufficient generalization and accuracy are used for subsequent complication probability assessment, further improving the accuracy of complication probability assessment.

[0062] The present invention uses a competing risks model to calculate the probabilities of a patient to be evaluated having various complications within the next r years based on the two-dimensional clustering tree. While ensuring the integrity of the patient's clinical index information and covering potential connection information, the two-dimensional clustering tree constructed by the present invention presents this information in the form of simple two-dimensional position coordinates. For technicians: Technicians are not concerned with the two-dimensional position coordinates of the two-dimensional clustering tree and the vast amount of information represented in the clustering tree structure, and only expect to know the final complication probability of getting sick. Therefore, whether technicians understand all the information covered by the two-dimensional clustering tree has no impact on the final competing risks model's assessment of the complication probability of getting sick. For the competing risks model: It simplifies the input of the competing risks model. By only using the two-dimensional position coordinates of the patient to be evaluated as the input of the model, the competing risks model can receive various clinical index information and potential connection information of the patient to be evaluated represented by the two-dimensional position coordinates. While improving the accuracy of the competing risks model's assessment, it greatly reduces the computational cost of the competing risks model and improves the assessment efficiency.

[0063] When the present invention evaluates the probability of a patient having a certain complication within the next r years, considering that the probability of getting sick will increase over time, the time is integrated during the evaluation process, making the evaluation result of the probability of getting sick more scientific.

[0064] The present invention can evaluate the probabilities of a patient to be evaluated having various complications within the next r years. Therefore, the patient can evaluate the probabilities of getting sick with different complications at different times in the future based on their current various clinical index data, which is very flexible and convenient. Doctors can perform preventive interventions on the patient's living habits and targeted prognosis based on the evaluation results, improving the patient's treatment effect. And regularly conducting risk assessments of complications is also beneficial for feeding back the effects of interventions and prognosis to doctors and patients, facilitating doctors and patients to make timely adjustments.

[0065] Technicians use the various clinical index data of 5000 patients with diagnosed complications at the time of diagnosis of type 2 diabetes, and use the complication risk assessment method of the present invention to evaluate the probabilities of having various complications within 3 years from the time of diagnosis of type 2 diabetes. For these 5000 patients, for the complication with the highest individual evaluation probability and a probability greater than 85%, the accuracy rate that the corresponding patient actually developed within 3 years from the time of diagnosis of type 2 diabetes is as high as 73.2%, while the average accuracy rate of the existing technology for risk assessment of type 2 diabetes complications is only 46.9%.

[0066] Technicians used the data of type 2 diabetes patients with diagnosed complications and adopted the complication risk assessment method of the present invention to verify the model performance. Based on the two-dimensional clustering tree structure and the competing risks model, the present invention demonstrated good prediction performance in the prediction of diabetes complications within 10 years. Specifically: (1) Model discrimination ability: When the method of the present invention predicts major complications (including myocardial infarction, stroke, heart failure, metabolic associated fatty liver disease, liver cirrhosis, diabetic retinopathy, chronic kidney disease, end-stage renal disease, and diabetic peripheral vascular disease), the ROC AUC (area under the receiver operating characteristic curve) reaches 0.82 - 0.91, indicating that the model has excellent discrimination.

[0067] (2) External validation performance: In the external validation of the independent dataset, the ROC AUC remains at 0.79 - 0.86, proving the strong generalization ability of the model.

[0068] Compared with the traditional methods that only use single clinical indicators or simple statistical methods, the present invention integrates multi-dimensional clinical data and uses dimensionality reduction technology, significantly improving the prediction efficiency and accuracy while retaining complete information, providing an efficient and reliable technical solution for the personalized risk assessment of type 2 diabetes complications.

[0069] The present invention also provides a risk assessment system for type 2 diabetes complications, including: A two-dimensional clustering tree construction module, a position coordinate mapping module, and a risk assessment module, The two-dimensional clustering tree construction module contains the DDRTree algorithm. The DDRTree algorithm constructs a two-dimensional clustering tree based on the clinical indicator data of patients in the first training set and sends the constructed two-dimensional clustering tree into the position coordinate mapping module; The position coordinate mapping module is used to convert the clinical indicator data of the patient to be evaluated into position coordinates on the two-dimensional clustering tree and send the position coordinates of the patient to be evaluated into the risk assessment module; The risk assessment module contains a competing risks model. The competing risks model calculates the probabilities of the patient to be evaluated having various complications within the next r years based on the position coordinates of the patient to be evaluated and outputs them; Each module is configured to execute the steps of a risk assessment method for type 2 diabetes complications as described above.

[0070] The present invention also provides a computer-readable storage medium: The computer-readable storage medium stores a computer program programmed or configured to execute a risk assessment method for type 2 diabetes complications as described above.

[0071] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies. It should also be noted that the above are only preferred embodiments of the present invention, and are not intended to limit the present invention. Each component or each step in the embodiments of the present invention can be decomposed and / or recombined, and these decompositions and / or recombinations should be regarded as equivalent solutions of this application and should all fall within the protection scope of the present invention.

Claims

1. A method for risk assessment of complications of type 2 diabetes, characterized in that: The following steps are involved: S1, based on the clinical indicator data of the patients in the first training set, a two-dimensional clustering tree is constructed using the DDRTree algorithm; the two-dimensional clustering tree contains a number of clustering regions, each clustering region contains a number of nodes, and each node represents a patient; S2, based on the clinical indicator data of the patient to be evaluated, using a mapping function to obtain the position coordinates of the patient to be evaluated in the two-dimensional clustering tree; S3, based on the location coordinates of the patient to be evaluated, uses the competing risk model to calculate the probability that the patient to be evaluated will suffer from various complications in the next r years.

2. A method for risk assessment of type 2 diabetes complications according to claim 1, characterized in that: After S1 and before S2, it also includes S1´: S1´, after the technicians mark the common features of the clinical indicators of each clustering area, they use the mapping function to obtain the position coordinates of the patients in the verification set in the two-dimensional clustering tree. If the consistency rate of the common features of the clinical indicators of the patients in the verification set and the corresponding clustering areas is greater than the first threshold, the current two-dimensional clustering tree is judged to be successfully verified; otherwise, the verification fails, and after discarding the current two-dimensional clustering tree, the number of patients in the first training set is increased and returns to S1.

3. A method for risk assessment of type 2 diabetes complications according to claim 1 or 2, characterized in that: In S1, the following sub-steps are also included: S11, obtaining a feature matrix X based on the first training set, wherein the first training set contains data of m clinical indicators of n patients; the n rows of the feature matrix X represent the n patients respectively, and the m columns represent the m clinical indicators respectively; n and m are both positive integers; S12, construct the optimization objective function F of the DDRTree algorithm based on the feature matrix X and then solve it to obtain the key parameters of the two-dimensional clustering tree: ; Among them, the key parameters are the optimal coordinate matrix and the optimal central node set; Y represents the coordinate matrix, which contains n rows and 2 columns. The n rows of the coordinate matrix Y represent n patients respectively, and are the same as the patients represented by the corresponding rows in the feature matrix X. The first column of the coordinate matrix Y represents the horizontal coordinate, and the second column of the coordinate matrix Y represents the vertical coordinate; Y(f) represents the position coordinate corresponding to the fth row in the coordinate matrix Y; W represents the linear transformation parameter matrix in the DDRTree algorithm; represents the transpose of the linear transformation parameter matrix W; λ represents the first hyperparameter; Indicates the position coordinates of the central node i in the current two-dimensional clustering tree; Indicates the position coordinates of the central node j in the current two-dimensional clustering tree; and are all position coordinates in the coordinate matrix Y; ε represents the edge set of the minimum spanning tree in the current two-dimensional clustering tree; Z represents the set of central node position coordinates; α represents the second hyperparameter; Indicates the position coordinates of the central node closest to the position coordinate Y(f); Indicates the coordinate matrix Y and the central node set Z corresponding to the minimum value. At this time, the coordinate matrix Y and the central node set Z are the optimal coordinate matrix and the optimal central node set; represents the square of the 2-norm; S13, construct a two-dimensional clustering tree based on key parameters.

4. A method for risk assessment of type 2 diabetes complications according to claim 2, characterized in that: In S1´, the following are also included: S11´, the technicians mark the common characteristics of the clinical indicators of each cluster area according to the clinical indicators of the patients in each cluster area in the current two-dimensional cluster tree; S12´, use the mapping function to obtain the position coordinates of each patient in the validation set in the two-dimensional clustering tree; S13´, determining the clustering region where each patient in the validation set is located according to the position coordinates of each patient in the validation set in the two-dimensional clustering tree; S14´, after determining whether the clinical indicator characteristics of each patient in the validation set are consistent with the common characteristics of the corresponding clustering region, calculate the consistency rate CR=NUM(C) / NUM(All); wherein NUM(·) represents the number; NUM(C) represents the number of patients whose clinical indicator characteristics of the patients in the validation set are consistent with the common characteristics of the corresponding clustering region, and NUM(All) represents the total number of patients in the validation set; If the consistency rate CR is greater than the first threshold, the current two-dimensional clustering tree is determined to be successfully verified; if the consistency rate CR is below the first threshold, the current two-dimensional clustering tree is determined to have failed verification, and after discarding the current two-dimensional clustering tree, a new first training set is constructed or the number of patients in the first training set is increased, and then return to S1.

5. The method for risk assessment of type 2 diabetes complications according to claim 1, characterized in that: The machine learning model is used to obtain the position coordinates of patient G in the two-dimensional clustering tree based on the mapping function. ; The mapping function is: ; ; in, The ordinate represents the position coordinate of patient G in the two-dimensional clustering tree; The horizontal coordinate represents the position coordinate of patient G in the two-dimensional clustering tree; represents the first global bias parameter; represents the second global bias parameter; represents the age regression coefficient; represents the gender regression coefficient; represents the age of patient G; represents the gender parameter of patient G; represents the data of the pth clinical indicator of patient G; express to the kth power; represents the pth clinical index intermediate value of patient G, max(·) represents the maximum value, It indicates that the node d in the two-dimensional clustering tree corresponds to the data of the pth clinical indicator of the patient. The two-dimensional clustering tree includes n nodes in total.

6. A method for risk assessment of type 2 diabetes complications according to claim 1, characterized in that: In S3, the undiagnosed complications of the patient to be evaluated are evaluated, where u is a positive integer and includes the following: The competing risk model is based on the location coordinates of the patient R to be evaluated, and calculates the probability that the patient will suffer from the Vth complication in the next r years. Probability , where r ≥ 0 and V is a positive integer less than or equal to u: ; ; Where t represents the future t-year of the patient R to be evaluated starting from the current moment; Indicates that the patient R to be evaluated will have complications in the next t years The instantaneous probability of represents the first regression coefficient; represents the second regression coefficient; The ordinate represents the position coordinate of the patient R to be evaluated in the two-dimensional clustering tree; The horizontal coordinate represents the position coordinate of the patient R to be evaluated in the two-dimensional clustering tree; Indicates complications Benchmark subdistribution risk function varying with future time t.

7. A method for risk assessment of type 2 diabetes complications according to claim 2, characterized in that: Before using the competing risk model to evaluate the probability of a patient suffering from various complications in the next r years, the competing risk model is trained using the second training set; the second training set includes e training data, each training data corresponds to a patient who has suffered from more than one complication, and each training data includes the time when the patient was diagnosed with type 2 diabetes, the time when various complications were diagnosed, and the clinical indicator data when the complication was most recently diagnosed; e is a positive integer.

8. A method for risk assessment of type 2 diabetes complications according to claim 7, characterized in that: The clinical indicator data of patients in the first training set, validation set, and second training set are all cleaned clinical indicator data, and the cleaned clinical indicator data also includes the following: Step 1: extract clinical indicator data of several patients. If a patient has more than one type of dirty data, all clinical indicator data of the current patient will be removed; dirty data refers to clinical indicator data that is outside 5 standard deviations among the same clinical indicator data of several patients; Step 2: The remaining clinical indicator data in step 1 are converted into clinical indicator data within the range of 0 to 1 by using a rank normalization method. At this time, the clinical indicator data cleaning is completed.

9. A risk assessment system for type 2 diabetes complications, characterized in that: include: Two-dimensional clustering tree construction module, location coordinate mapping module and risk assessment module, The two-dimensional clustering tree construction module includes a DDRTree algorithm, which constructs a two-dimensional clustering tree based on the clinical indicator data of the patients in the first training set, and sends the constructed two-dimensional clustering tree to the position coordinate mapping module; The position coordinate mapping module is used to convert the clinical indicator data of the patient to be evaluated into the position coordinates on the two-dimensional clustering tree, and send the position coordinates of the patient to be evaluated to the risk assessment module; The risk assessment module includes a competing risk model, which calculates the probability of the patient suffering from various complications in the next r years based on the location coordinates of the patient to be assessed and outputs it; Each module is configured to execute the steps of a method for risk assessment of complications of type 2 diabetes as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program programmed or configured to execute a method for assessing the risk of complications of type 2 diabetes as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for judging 2 type diabetes mellitus risk state

    CN102930163A

  • Method and system for predicting relapse risk of cerebral apoplexy

    CN117877727A

  • Diabetic complication prediction method based on k-means clustering analysis

    CN118942703A

  • Information processing method for disease stratification and assessment of disease progressing

    US20040243362A1