Method and system for constructing a diabetic kidney disease knowledge graph based on clinical data of DKD
By constructing a knowledge graph of DKD clinical data and combining traditional Chinese medicine syndromes with modern medical indicators, the problem of difficulty in judging the speed of DKD progression has been solved. This has enabled the identification of individuals prone to early DKD progression and the accurate prediction of the risk of proteinuria in the middle and late stages, thus improving the accuracy and timeliness of DKD progression assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies are insufficient to effectively analyze and predict the correlation between TCM symptoms and physicochemical indicators in different stages of diabetic nephropathy (DKD), making it difficult to determine the rate of DKD progression. Furthermore, traditional progression indicators lack strong evidence and are difficult to identify individuals prone to early progression and risk factors for proteinuria in the middle and late stages.
We constructed a knowledge graph of diabetic nephropathy based on clinical data of DKD, combined with traditional Chinese medicine syndromes and modern medical indicators, and analyzed the correlation between TCM symptoms and physicochemical indicators of DKD staging through visualization to determine the strength of influencing factors.
It enables accurate assessment of DKD progression and identification of individuals at an early stage of progression, improving the accuracy and timeliness of DKD progression assessment and allowing for timely intervention to slow kidney damage.
Smart Images

Figure CN118070895B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of databases, and more particularly, to a method and system for constructing a diabetic kidney disease (DKD) knowledge graph based on DKD clinical data. BACKGROUND
[0002] Diabetic kidney disease (DKD) is one of the most serious complications of diabetes, and about 40-70% of the causes of kidney dialysis are due to diabetes, which is the main cause of end-stage kidney disease in developed countries and regions. According to the ADA 2020 guidelines, the progression of diabetic kidney disease is as follows: early changes in glomerular capillary perfusion, gradual increase in urinary protein excretion rate, further to the microalbuminuria stage, clinical proteinuria stage (massive proteinuria accompanied by decreased glomerular filtration rate), and finally to end-stage kidney. Studies have shown that patients in the clinical proteinuria stage are irreversible, and will eventually progress to end-stage kidney, and the mortality rate of patients who have reached end-stage kidney will be greatly increased, with a ten-year survival rate of only 8.5%, causing a huge burden on public medical resources and human health.
[0003] A large number of research teams have constructed prediction models for the occurrence, development, progression risk, and even kidney replacement therapy risk of DKD. Such models can help clinicians assess the risk of patients and classify the risk, and thus provide precise and effective intervention measures to achieve the rational use of medical resources. However, such models only predict the progression risk outcome, but do not take into account the time factor, making the model evidence not strong, and thus the model cannot meet the comprehensive and continuous risk assessment needs of DKD patients. Evidence shows that DKD patients in the clinical proteinuria stage generally progress to end-stage kidney at a rate of 2.3% per year, but there are still many related factors that can accelerate the progression of diabetic kidney disease. Currently, there are very limited means to monitor and assess the progression rate of DKD patients from the clinical proteinuria stage to end-stage kidney, and indicators such as urinary protein excretion rate and glomerular filtration rate (eGFR) have certain limitations. According to research statistics, not all DKD patients have elevated albuminuria levels, and the prevalence of normal proteinuria DKD is nearly 9.7%. Therefore, traditional DKD progression indicators are difficult to determine the speed of DKD progression, which brings certain difficulties to clinicians in assessing the condition.
[0004] In addition, the characteristics of the medical syndrome have a clear diagnostic significance in different stages of diabetic kidney disease. Diabetic kidney disease is developed from diabetes. In the early stage, it is mainly yin deficiency and internal heat. Internal heat gradually damages yin and consumes qi. In the middle stage, it further develops yin and yang deficiency, forming pathological products such as blood stasis and phlegm dampness. After a long time, it forms "micro-tumor" to block the kidney collaterals, causing kidney damage, leading to the occurrence of diabetic kidney disease. The kidney stores essence, and if the kidney qi is insufficient, the kidney essence will not be fixed, and the subtle substances will leak out to form proteinuria, which corresponds to the "microalbuminuria period" and "clinical proteinuria period" in modern medicine. With further development of the disease, endogenous turbidity toxin, five-organ failure, and qi movement disorder, it eventually becomes oliguria, vomiting, and "guan ge" syndrome, and enters the "end stage of kidney". The "kidney collaterals tumor" stage forms slowly and often has no symptoms, making it difficult to find. The "kidney essence loss" stage is accompanied by fatigue, proteinuria and other clinical symptoms, which is the key syndrome stage for early diagnosis and treatment of diabetic kidney disease. Current studies have also shown that the kidney essence deficiency syndrome in traditional Chinese medicine is closely related to proteinuria, fibrosis and other pathological processes of chronic kidney disease. However, the current method cannot effectively analyze and predict the correlation between the TCM syndrome, TCM syndrome elements and physicochemical indicators in different DKD stages.
[0005] Therefore, it is necessary to introduce a new method and system to compare the TCM syndrome distribution characteristics of patients with DKD clinical proteinuria period and end-stage kidney, qualitatively clarify the correlation between TCM syndrome elements and modern medical indicators in DKD clinical proteinuria period and end-stage kidney, construct a knowledge graph to clarify the relationship between the pathogenesis characteristics, physicochemical indicators and TCM syndrome elements in DKD stages, and use a visual method to describe the relationship between various factors, in order to solve the technical problems that the traditional DKD progression indicators in the prior art cannot determine the speed of DKD progression, and cannot effectively analyze and predict the correlation between the TCM syndrome, TCM syndrome elements and physicochemical indicators in different DKD stages, so as to master the differences in the distribution characteristics of patients with clinical proteinuria period and end-stage kidney, and the quantitative relationship between TCM syndrome elements and diabetic kidney disease in the middle and late stages, and then accurately identify the early DKD progression population and the risk factors for rapid increase of proteinuria in the middle and late stages, and intervene and slow down the kidney damage of patients in time. SUMMARY
[0006] To solve the technical problems mentioned in the background art, the present application provides a method and system for constructing a diabetic kidney disease (DKD) knowledge graph based on clinical data, which combines traditional Chinese medicine (TCM) syndromes, TCM syndrome elements, and physicochemical indicators based on the differences and information of human information, modern medical indicators, and TCM syndromes at different DKD stages. The method can accurately analyze and determine the correlation between TCM syndromes, TCM syndrome elements, and physicochemical indicators at different DKD stages in a visual manner, thereby solving the technical problems of traditional DKD progression indicators being difficult to determine the speed of DKD progression and being difficult to effectively analyze and predict the correlation between TCM syndromes, TCM syndrome elements, and physicochemical indicators at different DKD stages. Furthermore, the method can quickly and accurately find the TCM syndrome elements and physicochemical indicators that affect TCM syndromes and determine the strength of the relationship, thereby improving the accuracy and timeliness of determining the speed of DKD progression.
[0007] The present application provides a method for constructing a diabetic kidney disease (DKD) knowledge graph based on clinical data, which includes:
[0008] S101, data collection, preprocessing, and storage: Collect clinical data according to the diabetic kidney disease (DKD) diagnosis standard, clinical staging, patient age, and sample size estimation model F, group and sample estimation process the clinical data, obtain basic clinical information data, and preprocess and store the basic clinical information data; S102, data export and annotation: Export data from the stored basic clinical information data, and perform annotation classification and co-occurrence analysis on the exported data to obtain structured basic clinical information data, determine the variable factors related to the DKD clinical proteinuria stage and end-stage kidney, and the correlation and co-occurrence values between the variable factors; S103, constructing a knowledge graph based on the correlation between each variable factor: Determine nodes based on the variable factors and add node identifiers and node labels, calculate the weights of each node and connect the edges between each node based on the correlation and co-occurrence values between each node, and determine the types and weights of the edges, construct a DKD knowledge graph based on the nodes and edges; S104, graph visualization and dynamic adjustment: Based on the constructed DKD knowledge graph, draw a visual DKD knowledge graph, and dynamically adjust the visual DKD knowledge graph based on the weights of the edges;
[0009] In step S101, the step of grouping and estimating the clinical data includes: S101-1, determining the patient sample size range based on common risk factors for the progression of diabetic nephropathy X, TCM symptom Y, dropout rate α, and the sample size estimation model F; S101-2, selecting the patient sample size HZ_Sum, and dividing the clinical data into a clinical proteinuria group and a renal failure group according to the DKD stage ratio and the patient sample size HZ_Sum, wherein the number of patients in the clinical proteinuria group is LK_Sum, and the number of patients in the renal failure group is SK_Sum.
[0010] As described above, the sample size estimation model is F = [5 × (X + Y) × α, 10 × (X + Y) × α];
[0011] The patient sample size range is [5×(X+Y)×α, 10×(X+Y)×α];
[0012] The number of patient samples HZ_Sum is greater than or equal to {5×(X+Y)×α}, and the number of patient samples HZ_Sum is less than or equal to {10×(X+Y)×α}.
[0013] The DKD phase ratio is LK_Sum:SK_Sum = 360:128;
[0014]
[0015]
[0016] As described above, step S101, the step of collecting clinical data, includes:
[0017] 1) Collect basic information data, baseline data, and clinical information data based on the data acquisition MDRD model;
[0018] 2) Collect TCM syndrome data according to the TCM clinical syndrome level, and calculate and obtain TCM syndrome score based on the TCM syndrome data;
[0019] in,
[0020] The clinical data includes the basic information data, the baseline data, the clinical information data, and the TCM syndrome data;
[0021] The glomerular filtration rate is GFR, serum creatinine is Cr, and age is Age;
[0022] When the patient being collected is male, the data collection MDRD model is as follows:
[0023] GFR = 175 × Cr -1.234 ×Age-0.179 ;
[0024] When the patient being collected is female, the data collection MDRD model is as follows:
[0025] GFR = 175 × Cr -1.234 ×Age -0.179 ×0.79;
[0026] The TCM clinical syndrome levels include mild, intermediate, and advanced levels. Each level corresponds to a different score, with higher levels corresponding to higher scores. When calculating the TCM syndrome score, the score for each TCM syndrome data is determined based on the TCM clinical syndrome level of the TCM syndrome data. Then, the scores of all TCM syndrome data corresponding to the syndrome element are summed to obtain the TCM syndrome score.
[0027] As described above, step S101, the step of preprocessing and storing the basic clinical information data, includes:
[0028] Based on the composition and category of the basic clinical information data, as well as the collection time and number of follow-up visits of the basic clinical information data, the storage structure of the basic clinical information data is determined, and a basic clinical information data table is constructed.
[0029] Based on the collection time and number of follow-up visits of the basic clinical information data, field description information is added to the collected basic clinical information data and stored in the basic clinical information data table;
[0030] in,
[0031] The storage structure of the basic clinical information data is {number, field description information, field name, field type};
[0032] The field description information includes: the number of follow-up visits and the collection time of the basic clinical information data;
[0033] The storage of basic clinical information data tables includes a basic information data table, a baseline data table, a clinical information data table, and a traditional Chinese medicine syndrome data table.
[0034] As described above, step S101, the step of preprocessing and storing the basic clinical information data, further includes:
[0035] Descriptive analysis and normality tests were performed on the basic clinical information data, and data characteristics were presented based on whether the quantitative basic clinical information data met a normal distribution.
[0036] If the basic clinical information data does not conform to the normal distribution, the data characteristics are presented as the median and quartiles.
[0037] If two or more of the basic clinical information data are compared, for the basic clinical information data that conforms to the normal distribution and has homogeneous variances, an independent samples test is used; for the basic clinical information data that does not conform to the normal distribution or has inhomogeneous variances, a non-parametric rank sum test is used. For count data, an X 2 test is performed, and all continuous variable characteristic data in the basic clinical information data are standardized, and the basic clinical information data is scaled to between [0, 1] in percentage form.
[0038] As described above, in step S102, the step of annotating, classifying, and co-occurrence analyzing the exported data further includes:
[0039] The data exported from the basic clinical information data is classified and annotated according to number, status judgment, basic information, physical and chemical indicators, traditional Chinese medicine symptoms, and traditional Chinese medicine syndrome scores to obtain structured basic clinical information data, including basic information data, clinical information data, and traditional Chinese medicine syndrome data. Among them, the annotation method for the basic information data is field name and field type, the annotation method for the clinical information data is field name, field type, and number of follow-up visits, and the annotation method for the traditional Chinese medicine syndrome data is traditional Chinese medicine syndrome label, traditional Chinese medicine syndrome element, traditional Chinese medicine symptoms, and traditional Chinese medicine syndrome score;
[0040] Based on the clinical information of the same patient during the same visit, co-occurrence analysis is performed on the structured basic clinical information data to determine the correlation between traditional Chinese medicine syndrome elements, traditional Chinese medicine symptoms, and baseline data, and the co-occurrence values between the traditional Chinese medicine syndrome elements, the traditional Chinese medicine symptoms, and the baseline data with correlations are recorded. Among them, the greater the co-occurrence value, the stronger the correlation between the traditional Chinese medicine syndrome elements, the traditional Chinese medicine symptoms, and the baseline data;
[0041] According to the correlation and the co-occurrence value, variable factors related to the clinical proteinuria stage and the end-stage kidney of diabetic nephropathy are determined;
[0042] Among them,
[0043] The coefficient value of the correlation is r, -1 < r < 1r. When r > 0, it indicates a positive correlation; when r < 0, it indicates a negative correlation; when r = 0, it indicates a zero correlation.
[0044] As described above, in step S103, the step of constructing a knowledge graph based on the correlation between variable factors further includes:
[0045] Node identification: Based on the variables related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy, the TCM syndrome elements, TCM symptoms and baseline data corresponding to the variables are used as nodes, and the node identifier and node label are added.
[0046] Calculate node weights: Based on the correlation between the TCM syndrome elements and TCM symptoms corresponding to the variable factors and the baseline data, calculate and determine the number of other nodes connected to each node, determine the node weight of the node based on the number of nodes, and determine the edges connecting the node to other nodes with which there is a correlation.
[0047] Calculate the weight of each edge: Calculate and determine the weight of each edge based on the co-occurrence value between the TCM syndrome elements and TCM symptoms corresponding to the variable factors and the baseline data;
[0048] Graph Construction: Construct a knowledge graph of diabetic nephropathy based on the nodes, node weights, edges, and edge weights.
[0049] As described above, in step S104, the step of drawing a visualized knowledge graph of diabetic nephropathy based on the constructed knowledge graph further includes:
[0050] Based on the node weights of the nodes constituting the diabetic nephropathy knowledge graph, the nodes are sorted, and the diameter of the nodes and the color of the node labels are determined according to the node sorting results. The diameter of the nodes and the intensity of the node label colors are both proportional to the node weights.
[0051] The edges are sorted according to their weights. Each edge is labeled according to its sorting result, and the rendering degree of each edge is determined. The rendering degree includes the thickness, length, and color of the edge. The thickness, length, and color intensity of the edge are all proportional to the weight of the edge.
[0052] Based on the diameter of the nodes and the color of the node labels, as well as the labels and rendering levels of the edges, the visualized knowledge graph of diabetic nephropathy is generated and drawn.
[0053] As described above, the basic clinical information data includes: basic information data, baseline data, clinical information data, and traditional Chinese medicine syndrome data;
[0054] The basic information data includes: name, gender, age, height, weight, body mass index, waist-to-hip ratio, occupation, and ethnicity;
[0055] The baseline data includes: past medical history data, personal data, physicochemical indicators, and functional evaluation data;
[0056] The functional evaluation data include urinary protein excretion rate and glomerular filtration rate.
[0057] Accordingly, the present invention also provides a system for constructing a knowledge graph of diabetic nephropathy based on DKD clinical data. The system includes a data preprocessing module, a data classification and annotation module, a knowledge graph construction module, and a graph dynamic processing module. Specifically, the data preprocessing module (S101) collects, preprocesses, and stores data: clinical data is collected according to the diagnostic criteria for diabetic nephropathy, clinical stages, patient age, and a sample size estimation model F; the clinical data is grouped and sample estimated to obtain basic clinical information data; and the basic clinical information data is preprocessed and stored. The data classification and annotation module is used for data export and annotation: data is exported from the stored basic clinical information data; the exported data is annotated, classified, and co-occurrence analyzed to obtain structured basic clinical information data; and variables related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy, as well as the correlation and co-occurrence values among these variables, are identified. The knowledge graph construction module is used to construct a knowledge graph based on various variables... The knowledge graph is constructed based on the correlations between quantitative factors: Nodes are determined based on the variable factors, and node identifiers and labels are added. The weights of each node and the edges connecting each node are calculated according to the correlations and co-occurrence values between nodes, as well as the type and weight of the edges. A knowledge graph for diabetic nephropathy is constructed based on the nodes and edges. The graph dynamic processing module is used for graph visualization and dynamic adjustment: Based on the constructed knowledge graph for diabetic nephropathy, a visualized knowledge graph for diabetic nephropathy is drawn, and the visualized knowledge graph for diabetic nephropathy is dynamically adjusted according to the weights of the edges. Specifically, when dynamically adjusting the visualized knowledge graph for diabetic nephropathy, the number of visualized nodes and the number of edges connecting the nodes are inversely proportional to the weight values of the edges connecting the nodes. The smaller the weight value of the edge connecting a node, the more visualized edges the node has, and the more uniform the distribution of correlations between nodes. Conversely, the larger the weight value of the edge connecting a node, the fewer visualized edges the node has, and the stronger the correlations between nodes.
[0058] This invention, by applying the above technical solutions, realizes the differences and information in the distribution of human information, modern medical indicators, and traditional Chinese medicine (TCM) syndromes based on different DKD stages. It combines TCM symptoms, TCM syndrome elements, and physicochemical indicators, allowing for precise analysis and judgment of the correlation between TCM symptoms, TCM syndrome elements, and physicochemical indicators at different DKD stages in a visualized manner. This solves the technical problems in existing technologies where traditional DKD progression indicators are difficult to use to determine the speed of DKD progression, and where it is difficult to effectively analyze and predict the correlation between TCM symptoms, TCM syndrome elements, and physicochemical indicators at different DKD stages. This allows for the differentiation of disease distribution characteristics between patients in the clinical proteinuria stage and those in the end-stage of renal disease, as well as the quantitative relationship between TCM syndrome elements and the middle and late stages of diabetic nephropathy. It accurately identifies individuals prone to early DKD progression and risk factors for rapid increases in proteinuria in the middle and late stages, enabling timely intervention and mitigation of kidney damage. Furthermore, it quickly and accurately identifies TCM syndrome elements and physicochemical indicators affecting TCM symptoms and determines the strength of their influence, improving the accuracy and timeliness of judging the speed of DKD progression. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 A flowchart illustrating the method for constructing a knowledge graph of diabetic nephropathy based on DKD clinical data, as proposed in an embodiment of the present invention, is shown.
[0061] Figure 2 This diagram illustrates the node relationships in the knowledge graph construction method for diabetic nephropathy based on DKD clinical data proposed in this embodiment of the invention.
[0062] Figure 3 This diagram illustrates the overall structure of the knowledge graph for diabetic nephropathy based on DKD clinical data proposed in this embodiment of the invention.
[0063] Figure 4 This diagram illustrates the poor correlation between nodes in the knowledge graph of diabetic nephropathy based on DKD clinical data proposed in this embodiment of the invention.
[0064] Figure 5 This illustration shows a schematic diagram of weak correlations between nodes in a knowledge graph of diabetic nephropathy based on DKD clinical data, as proposed in an embodiment of the present invention.
[0065] Figure 6 This illustration shows a diagram illustrating the strong correlation between nodes in the knowledge graph of diabetic nephropathy based on DKD clinical data proposed in this embodiment of the invention.
[0066] Figure 7 This illustration shows a strong correlation diagram of nodes in a knowledge graph of diabetic nephropathy based on DKD clinical data, as proposed in an embodiment of the present invention.
[0067] Figure 8 The diagram shows a schematic of the system for constructing a knowledge graph of diabetic nephropathy based on DKD clinical data, as proposed in an embodiment of the present invention. Detailed Implementation
[0068] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0069] This invention provides a method for constructing a knowledge graph of diabetic nephropathy based on clinical data of DKD, such as... Figure 1 As shown, the method includes the following steps:
[0070] S101, Data collection, preprocessing and storage: Clinical data is collected according to the diagnostic criteria for diabetic nephropathy, clinical stage, patient age and sample size estimation model F. The clinical data is grouped and sample estimation is performed to obtain basic clinical information data. The basic clinical information data is then preprocessed and stored.
[0071] in,
[0072] In step S101, the step of grouping and estimating the clinical data includes:
[0073] S101-1, Based on the common risk factors for the progression of diabetic nephropathy X, TCM symptom Y, dropout rate α, and the sample size estimation model F, determine the patient sample size range;
[0074] S101-2, Select the number of patient samples HZ_Sum, and divide the clinical data into a clinical proteinuria group and a renal failure group according to the DKD stage ratio and the number of patient samples HZ_Sum, wherein the number of patients in the clinical proteinuria group is LK_Sum and the number of patients in the renal failure group is SK_Sum.
[0075] In this embodiment, the sample size estimation model is F = [5 × (X + Y) × α, 10 × (X + Y) × α];
[0076] The patient sample size range is [5×(X+Y)×α, 10×(X+Y)×α];
[0077] The number of patient samples HZ_Sum is greater than or equal to {5×(X+Y)×α}, and the number of patient samples HZ_Sum is less than or equal to {10×(X+Y)×α}.
[0078] The DKD phase ratio is LK_Sum:SK_Sum = 360:128;
[0079]
[0080]
[0081] To better illustrate the technical solution provided by this invention, this step is further explained using a sample of 488 patients as an example. Based on the diagnostic criteria and data collection standards for diabetic nephropathy, the clinical data of the final 488 patients were divided into two groups according to DKD stage: Clinical proteinuria group (360 cases): ACR > 300 mg / g or 24-hour urinary protein excretion rate > 300 mg / 24H, with or without decreased renal function, eGFR > 15 ml / min / 1.73 m2. Renal failure group (128 cases): Glomerular filtration rate < 15 ml / (min·1.73 m2), often accompanied by clinical manifestations related to end-stage renal disease.
[0082] The sample size was calculated using Kendall's principle of using 5 to 10 times the number of variables. Twenty-six common risk factors for the progression of diabetic nephropathy were selected, along with 11 TCM syndrome elements. The minimum sample size was between 37 × 5 = 185 and 37 × 10 = 370 cases. Assuming a dropout rate of 20%, the total number of cases needed was between 185 × 120% = 222 and 370 × 120% = 444 cases. The actual number of cases included was 488, consistent with the estimated sample size.
[0083] In this embodiment, the basic clinical information data includes: basic information data, baseline data, clinical information data, and traditional Chinese medicine syndrome data;
[0084] The basic information data includes: name, gender, age, height, weight, body mass index, waist-to-hip ratio, occupation, and ethnicity;
[0085] The baseline data includes: past medical history data, personal data, physicochemical indicators, and functional evaluation data;
[0086] The functional evaluation data include urinary protein excretion rate and glomerular filtration rate.
[0087] To help professionals in this field better understand the diagnostic criteria and data collection standards for diabetic nephropathy, the ADA 2020 Guidelines for the Diagnosis and Treatment of Diabetes and the Mogensen staging system can be referenced before collecting clinical data. The specific details are as follows:
[0088] (1) Meets the diagnostic criteria for diabetes: fasting blood glucose ≥7.0 mmol / L; random blood glucose ≥11.1 mmol / L or oral glucose tolerance test (OGTT) ≥11.1 mmol / L in 2 hours; glycated hemoglobin ≥6.5%; if there are symptoms of diabetes, a single blood glucose value that meets the diagnostic criteria for diabetes can be diagnosed as diabetes.
[0089] (2) Individuals meeting one of the following criteria may be diagnosed with clinical proteinuria stage or above of DKD (see Table 1 for details):
[0090] a. Urinary albumin / creatinine ratio (UACR) ≥ 300 mg / g or urinary albumin excretion ratio (UAER) ≥ 300 mg / 24h;
[0091] b. Estimated glomerular filtration rate (eGFR) <60 ml / (min·1.73m2) for more than 3 months;
[0092] c. Age ≥ 18 years old, gender not limited;
[0093] d. Sign an informed consent form.
[0094] Table 1
[0095]
[0096]
[0097] Exclusion criteria:
[0098] (1) Patients with type 1 diabetes or type 2 diabetes without kidney damage; patients with diabetic nephropathy whose urinary protein excretion rate does not meet the requirements;
[0099] (2) Non-diabetic kidney diseases, such as primary glomerular diseases, drug-induced kidney damage, nephrotic syndrome secondary to other causes, autoimmune diseases and connective tissue diseases, tumors, etc., have not been clearly excluded;
[0100] (3) All or important parts of the medical records and related clinical data are missing;
[0101] (4) Patients with serious primary diseases of the respiratory, digestive, or hematologic systems, as well as those currently suffering from concurrent infections or mental illness, or those with congestive heart failure; or those diagnosed with malignant tumors, or pregnant or breastfeeding patients; or those with a history of critical illnesses such as malignant hypertension, myocardial infarction, cerebrovascular accident, or diabetic ketoacidosis within the past 6 months.
[0102] Exclusion criteria:
[0103] (1) The subject voluntarily withdrew;
[0104] (2) Unable to continue follow-up observation or missing important follow-up data.
[0105] To collect more accurate clinical data that meets the requirements, in this embodiment, step S101, the step of collecting clinical data, includes:
[0106] 1) Collect basic information data, baseline data, and clinical information data based on the data acquisition MDRD model;
[0107] 2) Collect TCM syndrome data according to the TCM clinical syndrome level, and calculate and obtain TCM syndrome score based on the TCM syndrome data;
[0108] in,
[0109] The clinical data includes the basic information data, the baseline data, the clinical information data, and the TCM syndrome data;
[0110] The glomerular filtration rate is GFR, serum creatinine is Cr, and age is Age;
[0111] When the patient being collected is male, the data collection MDRD model is as follows:
[0112] GFR = 175 × Cr -1.234 ×Age -0.179 ;
[0113] When the patient being collected is female, the data collection MDRD model is as follows:
[0114] GFR = 175 × Cr -1.234 ×Age -0.179 ×0.79;
[0115] The TCM clinical syndrome levels include mild, intermediate, and advanced levels. Each level corresponds to a different score, with higher levels corresponding to higher scores. When calculating the TCM syndrome score, the score for each TCM syndrome data is determined based on the TCM clinical syndrome level of the TCM syndrome data. Then, the scores of all TCM syndrome data corresponding to the syndrome element are summed to obtain the TCM syndrome score.
[0116] To better illustrate the technical solution provided by this invention, a sample size of 488 patients will be used as an example for further explanation. The details are as follows:
[0117] I. Collection of basic information data, baseline data, and clinical information data.
[0118] Patients' medical history and clinical information relevant to this study were collected, and patients diagnosed with diabetic nephropathy in the clinical proteinuria stage were included in the follow-up observation list. Specific clinical information included in the study is shown in Table 2. The aforementioned MDRD model was used in the data collection. The MDRD model can also be written as:
[0119] GFR (glomerular filtration rate) = 175 × serum creatinine (SCr, mg / dl) - 1.234 × age (Age, years) - 0.179 × 0.79 (if female), or
[0120] GFR (glomerular filtration rate) = 175 × serum creatinine (Cr, mg / dl) -1.234 × Age (in years) - 0.179 ×0.79 (if female).
[0121] Table 2
[0122]
[0123] II. Collection of Clinical Syndromes in Traditional Chinese Medicine
[0124] Based on the Clinical Respiratory Scale (CRF) for Diabetic Nephropathy with Kidney Essence Deficiency (hereinafter referred to as Essence Deficiency), a Traditional Chinese Medicine (TCM) syndrome assessment form for diabetic nephropathy patients was developed (see Table 3, TCM Syndrome Assessment Form for Diabetic Nephropathy). TCM syndrome types were determined for enrolled patients, and syndrome data were recorded. Scoring criteria were used to calculate scores of 2, 4, and 6 points for mild, moderate, and severe cases, respectively. The scores were then summed to obtain the total score for each syndrome element.
[0125] Table 3
[0126]
[0127]
[0128] In this embodiment, step S101, the step of preprocessing and storing the basic clinical information data, includes:
[0129] Based on the composition and category of the basic clinical information data, as well as the collection time and number of follow-up visits of the basic clinical information data, the storage structure of the basic clinical information data is determined, and a basic clinical information data table is constructed.
[0130] Based on the collection time and number of follow-up visits of the basic clinical information data, field description information is added to the collected basic clinical information data and stored in the basic clinical information data table;
[0131] in,
[0132] The storage structure of the basic clinical information data is {number, field description information, field name, field type};
[0133] The field description information includes: the number of follow-up visits and the collection time of the basic clinical information data;
[0134] The storage of basic clinical information data tables includes a basic information data table, a baseline data table, a clinical information data table, and a traditional Chinese medicine syndrome data table.
[0135] It is worth noting that in some practical applications, data storage methods can also be "field name", "field name + type", or "field name + type + number of follow-up visits". For example, "24-hour urine protein quantification at the second follow-up visit" can be named as "24-hour urine protein value 2". After the raw data is entered and verified as required, it is archived and stored in sequence, and a retrieval catalog is filled in. Electronic data files include databases, analysis results, and explanatory documents, which are stored in categories, and multiple backups are saved on different disks or recording media, and properly stored to prevent damage.
[0136] To ensure that clinical data better meets the accuracy requirements for constructing a knowledge graph of diabetic nephropathy, in this embodiment, step S101, the step of preprocessing and storing the basic clinical information data, further includes:
[0137] Descriptive analysis and normality tests were performed on the basic clinical information data, and data characteristics were presented based on whether the quantitative basic clinical information data met a normal distribution.
[0138] If the basic clinical information data does not conform to a normal distribution, the data characteristics are presented as the median and quartiles;
[0139] When comparing two or more sets of basic clinical information data, an independent samples test is used for basic clinical information data that conforms to a normal distribution and has homogeneous variance; a nonparametric rank-sum test is used for basic clinical information data that does not conform to a normal distribution or has heterogeneous variance. Count data are compared using the X-ray discrepancy test. 2 The basic clinical information data was tested and standardized, and then scaled to the range [0, 1] by percentage.
[0140] To better illustrate the technical solution provided by this invention, a sample size of 488 patients will be used as an example for further explanation. During data processing, descriptive analysis is performed on basic clinical data; normality tests are conducted on quantitative basic clinical information data, and data characteristics are presented based on whether they conform to a normal distribution. If they do not conform to a normal distribution, the median and interquartile range (IQR) are used to present the data characteristics. For comparisons between two independent basic clinical information data samples, independent samples t-tests are used for continuous data that conform to a normal distribution and have homogeneous variances; nonparametric rank-sum tests are used for data that do not conform to a normal distribution or have heterogeneous variances; and chi-square tests are used for categorical data. 2 Test. Count data are expressed as frequency n and percentage (%). Further standardization is performed on all continuous variable characteristic data by processing the data using MinMax Scaler to scale the data to the range [0,1].
[0141] Regarding the handling of missing data and imputation, for variables with a missing rate of less than 20%, median imputation and adjacent imputation methods were used for data imputation. For non-observed indicators and those with unclear relationships to observed indicators, multiple imputation methods had no significant impact on the final accuracy of the model. Questionnaires with more than 20% missing content were excluded; outlier data were replaced using the median method.
[0142] S102, Data Export and Labeling: Export data from the stored basic clinical information data, and perform labeling, classification and co-occurrence analysis on the exported data to obtain structured basic clinical information data, and determine the variable factors related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy, as well as the correlation and co-occurrence values among the variable factors.
[0143] To better illustrate the technical solution provided by this invention, a sample size of 488 patients will be used as an example for further explanation. In this example, clinical data is exported from the database and divided into six aspects: "patient number, status assessment, basic information, physicochemical indicators, TCM symptoms, and TCM syndrome score". Each patient is identified using a patient code, and these six data points are then linked together.
[0144] Then, the clinical data is annotated, specifically including the following:
[0145] (1) Basic information data annotation
[0146] Patient basic information is stored and labeled in the database using field names (Incase Fieldname) and field types (Incase Field Type). For example, the first visit date is labeled as: First Visit Date.
[0147] (2) Clinical information data annotation
[0148] Patient clinical information data is labeled as "field name", "field name + type", or "field name + type + number of follow-up visits". For example, "first follow-up blood routine hemoglobin value" is named as: "blood routine, hemoglobin value, 1".
[0149] (3) TCM syndrome data annotation
[0150] Based on the data annotation results of the TCM syndrome element diagnostic scale collected from the same clinical information collection of the patient, the relationship between adjacent elements at two levels is recorded in the form of "Level 1 element α, Level 2 element β". For example, the patient's "Essence deficiency syndrome with tinnitus 4 points" at the first baseline enrollment can be recorded as: "TCM syndrome, essence deficiency, tinnitus, 4 points".
[0151] Co-occurrence analysis was performed on the labeled clinical data. Based on the correlation analysis of clinical data from the same patient visit, the co-occurrence of element records that showed correlation after verification (P<0.05) was defined as "element 1, 1". For example, "sperm deficiency syndrome, 1".
[0152] To effectively utilize the preprocessed clinical data, in this embodiment, step S102, the step of labeling, classifying, and performing co-occurrence analysis on the exported data, further includes:
[0153] The data exported from the basic clinical information data is classified and labeled according to number, status judgment, basic information, physicochemical indicators, TCM symptoms, and TCM syndrome scores to obtain structured basic clinical information data, including basic information data, clinical information data, and TCM syndrome data. The labeling method of the basic information data is field name and field type, the labeling method of the clinical information data is field name, field type, and number of follow-up visits, and the labeling method of the TCM syndrome data is TCM syndrome label, TCM syndrome element, TCM symptoms, and TCM syndrome score.
[0154] Based on the clinical information of the same patient during the same visit, co-occurrence analysis is performed on the structured basic clinical information data to determine the correlation between TCM syndrome elements, TCM symptoms and baseline data, and the co-occurrence value of the TCM syndrome elements, TCM symptoms and baseline data that have a correlation is recorded. The larger the co-occurrence value, the stronger the correlation between the TCM syndrome elements, TCM symptoms and baseline data.
[0155] Based on the correlation and co-occurrence value, determine the variable factors associated with the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy;
[0156] in,
[0157] The coefficient value of the correlation is r, where -1 < r < 1. When r > 0, it indicates a positive correlation; when r < 0, it indicates a negative correlation; and when r = 0, it indicates a zero correlation.
[0158] It should be noted that in some practical examples, for continuous variables, the Spearman rank correlation coefficient is used to obtain the relationship between the two variables in question. The correlation coefficient takes values in the range -1 < r < 1. |r| indicates the degree of correlation between the two variables. r > 0 indicates a positive correlation, r < 0 indicates a negative correlation, and r = 0 indicates a zero correlation. For two unordered categorical variables, the φ coefficient (Phi coefficient) is used to measure the degree of association between the two binary variables.
[0159] S103. Construct a knowledge graph based on the correlations between various variable factors: Determine nodes based on the variable factors, add node identifiers and node labels, calculate the weights of each node and the edges connecting each node, as well as the type and weight of the edges, based on the correlations and co-occurrence values between each node, and construct a diabetic nephropathy knowledge graph based on the nodes and the edges.
[0160] In this embodiment, in step S103, the step of constructing a knowledge graph based on the correlations between various variable factors further includes:
[0161] Determine nodes: Based on the variable factors related to the clinical proteinuria stage and the end-stage of the kidney in diabetic nephropathy, take the traditional Chinese medicine syndrome elements, traditional Chinese medicine symptoms corresponding to the variable factors, and baseline data as the nodes, and add the node identifier and the node label;
[0162] Calculate the node weights: According to the correlations between the traditional Chinese medicine syndrome elements, traditional Chinese medicine symptoms corresponding to the variable factors, and the baseline data, calculate and determine the number of other nodes connected to each node, determine the node weight of the node according to the number of nodes, and the edges connecting the node and other nodes with correlations;
[0163] Calculate the weights of the edges: According to the co-occurrence values between the traditional Chinese medicine syndrome elements, traditional Chinese medicine symptoms corresponding to the variable factors, and the baseline data, calculate and determine the weights of the edges connecting each edge;
[0164] Construct the graph: Construct a diabetic nephropathy knowledge graph based on the nodes, the node weights, as well as the edges and the edge weights.
[0165] It should be noted that the nodes and edges in constructing the diabetic nephropathy knowledge graph can be understood as the mapping between "entity", "relationship", and "entity", such as Figure 2As shown, entities (pathological characteristics of DKD patients, TCM syndrome elements, etc.) serve as "nodes" in the knowledge graph, and the correlations between each entity serve as "edges." In a narrow sense, the measurement of edges and nodes is related to the frequency of data occurrence and their hierarchical level. However, since some instances use pre-labeled data rather than literature or textual data, in addition to using hierarchical relationships as the measure of edges, the number of edges entering and leaving a node is used as the "weight" or "quality" of that node, and the correlation coefficient between independent variables is used as the "measure" of the edge. Data was imported using Gephi knowledge graph software, and the parameter settings for nodes and edges were adjusted to create a knowledge graph of DKD progression, visualizing the risk factors for DKD progression.
[0166] S104, Graph Visualization and Dynamic Adjustment: Based on the constructed diabetic nephropathy knowledge graph, a visualized diabetic nephropathy knowledge graph is drawn, and the visualized diabetic nephropathy knowledge graph is dynamically adjusted according to the weight of the edges.
[0167] In this embodiment, step S104, the step of drawing a visualized knowledge graph of diabetic nephropathy based on the constructed knowledge graph, further includes:
[0168] Based on the node weights of the nodes constituting the diabetic nephropathy knowledge graph, the nodes are sorted, and the diameter of the nodes and the color of the node labels are determined according to the node sorting results. The diameter of the nodes and the intensity of the node label colors are both proportional to the node weights.
[0169] The edges are sorted according to their weights. Each edge is labeled according to its sorting result, and the rendering degree of each edge is determined. The rendering degree includes the thickness, length, and color of the edge. The thickness, length, and color intensity of the edge are all proportional to the weight of the edge.
[0170] Based on the diameter of the nodes and the color of the node labels, as well as the labels and rendering levels of the edges, the visualized knowledge graph of diabetic nephropathy is generated and drawn.
[0171] To better illustrate the technical solution provided by this invention, a sample of 488 patients will be used as an example for further explanation. In this example, when constructing the visualized knowledge graph of diabetic nephropathy,
[0172] (1) Parameter design of the points: Based on the structured basic clinical information data obtained after previous processing, the variables related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy identified by correlation analysis were used as nodes, and the main parameters of the points were set and statistically analyzed. See Table 4 for details.
[0173] Node ID and Label: Each variable name serves as the identifier and label for each node.
[0174] Node weight: The weight of a node is determined by the number of edges it is associated with. The number of edges is determined by a prior relevance list; the more connections a node has as a related element, the greater its weight.
[0175] Table 4
[0176]
[0177] (2) Edge parameter settings: Based on the data from the above correlation analysis, the main parameters of the edges are set and statistically analyzed. See Table 5 for details.
[0178] Edge endpoints (Nodes): The variable elements in the correlation analysis are used as the source and target of the edge, respectively.
[0179] Edge Type: Considering that the relationships between nodes in this study are correlational, i.e., the correlation between clinical physiochemical indicators, basic information, and TCM syndrome elements, and that this relationship is not hierarchical or subordinate, but bidirectional, the edge type is set to undirected.
[0180] Edge weight: The weight of an edge is the correlation between two nodes. When data is imported into Gephi 0.10.1 software, the weights of the same edges are automatically merged and summed.
[0181] Table 5
[0182]
[0183]
[0184] Based on the constructed knowledge graph of diabetic nephropathy, the parameters are set after importing the parameters of the nodes and edges.
[0185] 1) Node display parameter settings
[0186] Configure node display parameters. In the node size settings, set the rendering method for "Ranking" to "Weight," with the ranking order based on the node's weight. Set the minimum size to 5 and the maximum size to 40; that is, the greater the node's weight, the darker the node's label color and the larger the node's diameter. Continue configuring the font, size, color, and other parameters for the node labels on this interface.
[0187] 2) Setting the display parameters of the edges
[0188] Similarly, continue setting parameters such as edge label, color appearance, and thickness. For the edge's "ranking" rendering method, select "degree." The ranking order is based on the edge's weight; that is, the thickness of the edge is displayed according to its weight. The greater the edge's weight, the thicker it appears, the longer the distance, and the darker the color; conversely, the smaller the weight, the thinner the edge, the longer the length, and the lighter the color.
[0189] 3) Overall layout settings
[0190] Configure the layout. Select "ForceAtlas2", adjust the scaling to 20.0, select "Prevent Overlap", and adjust the graphic structure.
[0191] After setting up, draw a knowledge graph of diabetic nephropathy. For example... Figure 3 As shown, according to the knowledge graph, the major points are, in descending order: sperm deficiency syndrome, 24-hour urine protein quantification, serum creatinine, serum urea nitrogen, estimated glomerular filtration rate, urine protein qualitative analysis, and urine microalbumin, etc.; followed by urine protein quantification and fasting blood glucose, etc. Symptom nodes are evenly distributed around syndrome elements, but their weights are relatively small.
[0192] After drawing the knowledge graph of diabetic nephropathy, the range of edge weights can be adjusted to dynamically adjust the knowledge graph of diabetic nephropathy.
[0193] like Figure 4 As shown. When the weight range is 0.2-0.4, adjusting the weight range of the "edges" can yield a knowledge graph with related nodes within that range. Filtering the knowledge graph with weight relationships of 0.2-0.4 reveals weakly correlated disease-symptom factors, including sperm deficiency-fasting blood glucose, 24-hour urine protein quantification-lower back and knee weakness, 24-hour urine protein quantification-yin deficiency syndrome, and glomerular filtration rate-yin deficiency syndrome. These disease-symptom factors have weak correlations and cannot be strongly correlated with disease progression. The correlation between symptoms and modern medical indicators is also poor.
[0194] like Figure 5 As shown, when the weight range is 0.4-0.6, further filtering of the knowledge graph with a weight of 0.4-0.6 reveals new weighted edges. These include: deficiency of essence syndrome - serum urea nitrogen, 24-hour urine protein quantification - lethargy, 24-hour urine protein quantification - blood stasis syndrome, deficiency of essence syndrome - serum uric acid, damp-heat syndrome - qualitative urine protein, etc. The weights of these nodes are increased, and the correlation between nodes is relatively high. Figure 4 It's even more obvious.
[0195] like Figure 6As shown, when the weight range is 0.6-0.8, further filtering of the knowledge graph with weight ratios of 0.6-0.8 reveals a decrease in the number of lines, indicating the following: Essence Deficiency Syndrome - ACR, Essence Deficiency Syndrome - Glomerular Filtration Rate, Dampness Syndrome - 24-hour Urine Protein, Dampness Syndrome - Qualitative Urine Protein, Dampness Syndrome - Quantitative Urine Protein, Blood Deficiency Syndrome - Glomerular Filtration Rate, Foamy Urine - ACR, Dampness and Turbidity Syndrome - Glomerular Filtration Rate, Dampness and Turbidity Syndrome - Serum Creatinine, etc. This weight ratio range is quite clear, indicating a strong correlation between TCM syndrome types and modern medical indicators, and between symptoms and modern medical indicators.
[0196] like Figure 7 As shown, when the weight ranges from 0.8 to 0.8, filtering the knowledge graph with a weight ratio greater than 0.8 reveals that only two sets of nodes are correlated: sperm deficiency syndrome - biochemical albumin, and serum creatinine - water retention syndrome.
[0197] By applying the above technical solutions, the differences and information of human information, modern medical indicators, and traditional Chinese medicine (TCM) syndrome distribution based on different DKD stages are realized. By combining TCM symptoms, TCM syndrome elements, and physicochemical indicators, the correlation between TCM symptoms, TCM syndrome elements, and physicochemical indicators at different DKD stages can be accurately analyzed and judged in a visualized manner. This solves the technical problems of existing technologies where traditional DKD progression indicators are difficult to use to determine the speed of DKD progression, and where it is difficult to effectively analyze and predict the correlation between TCM symptoms, TCM syndrome elements, and physicochemical indicators at different DKD stages. This allows for the differentiation of disease distribution characteristics between patients in the clinical proteinuria stage and end-stage renal disease, as well as the quantitative relationship between TCM syndrome elements and the middle and late stages of diabetic nephropathy. It accurately identifies individuals prone to early DKD progression and risk factors for rapid increases in proteinuria in the middle and late stages, enabling timely intervention and slowing down kidney damage. Furthermore, it quickly and accurately identifies TCM syndrome elements and physicochemical indicators affecting TCM symptoms and determines the strength of their influence, improving the accuracy and timeliness of judging the speed of DKD progression.
[0198] Corresponding to the method for constructing a knowledge graph of diabetic nephropathy based on DKD clinical data described in the embodiments of the present invention, the present invention also discloses a system for constructing a knowledge graph of diabetic nephropathy based on DKD clinical data, such as... Figure 8 As shown, the system includes a data preprocessing module, a data classification and labeling module, a knowledge graph construction module, and a knowledge graph dynamic processing module;
[0199] in,
[0200] The data preprocessing module is used for S101, data collection, preprocessing and storage: collecting clinical data according to the diagnostic criteria for diabetic nephropathy, clinical stage, patient age and sample size estimation model F, grouping and sample estimation processing of the clinical data to obtain basic clinical information data, and preprocessing and storing the basic clinical information data.
[0201] The data classification and annotation module is used for data export and annotation: exporting data from the stored basic clinical information data, and performing annotation, classification and co-occurrence analysis on the exported data to obtain structured basic clinical information data, and determining the variable factors related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy, as well as the correlation and co-occurrence values between the variable factors;
[0202] The knowledge graph construction module is used to construct a knowledge graph based on the correlation between various variable factors: determine nodes based on the variable factors and add node identifiers and node labels; calculate the weight of each node and the edges connecting each node according to the correlation and co-occurrence value between each node, as well as the type and weight of the edges; and construct a knowledge graph of diabetic nephropathy based on the nodes and the edges.
[0203] The graph dynamic processing module is used for graph visualization and dynamic adjustment: based on the constructed diabetic nephropathy knowledge graph, a visualized diabetic nephropathy knowledge graph is drawn, and the visualized diabetic nephropathy knowledge graph is dynamically adjusted according to the weight of the edges;
[0204] Specifically, when dynamically adjusting the visualized diabetic nephropathy knowledge graph, the number of visualized nodes and the weight values of the edges connecting the nodes are inversely proportional. The smaller the weight value of the edge of a node, the more visualized edges the node has, and the more uniform the distribution of correlations among the nodes. Conversely, the larger the weight value of the edge of a node, the fewer visualized edges the node has, and the stronger the correlations among the nodes.
[0205] The various embodiments in this specification are described in a related manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0206] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for constructing a knowledge graph of diabetic nephropathy based on DKD clinical data, characterized in that, The method includes: S101, Data collection, preprocessing and storage: Clinical data is collected according to the diagnostic criteria for diabetic nephropathy, clinical stage, patient age and sample size estimation model F. The clinical data is grouped and sample estimation is performed to obtain basic clinical information data. The basic clinical information data is then preprocessed and stored. S102, Data export and annotation: Export data from the stored basic clinical information data, and perform annotation, classification and co-occurrence analysis on the exported data to obtain structured basic clinical information data, and determine the variable factors related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy, as well as the correlation and co-occurrence values among the variable factors; S103, Construct a knowledge graph based on the correlation between various variable factors: Determine nodes based on the variable factors and add node identifiers and node labels. Calculate the weight of each node and the edges connecting each node, as well as the type and weight of the edges, based on the correlation and co-occurrence value between each node. Construct a knowledge graph of diabetic nephropathy based on the nodes and the edges. S104, Graph Visualization and Dynamic Adjustment: Based on the constructed diabetic nephropathy knowledge graph, a visualized diabetic nephropathy knowledge graph is drawn, and the visualized diabetic nephropathy knowledge graph is dynamically adjusted according to the weight of the edges; In step S101, the step of grouping and estimating the clinical data includes: S101-1, based on common risk factors for the progression of diabetic nephropathy X, traditional Chinese medicine symptom Y, and dropout rate. Using the sample size estimation model F, determine the patient sample size range; S101-2, Select the number of patient samples HZ_Sum, and divide the clinical data into a clinical proteinuria group and a renal failure group according to the DKD stage ratio and the number of patient samples HZ_Sum. The number of patients in the clinical proteinuria group is LK_Sum, and the number of patients in the renal failure group is SK_Sum. In step S103, the step of constructing a knowledge graph based on the correlation between various variable factors further includes: Node identification: Based on the variables related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy, the TCM syndrome elements, TCM symptoms and baseline data corresponding to the variables are used as nodes, and the node identifier and node label are added. Calculate node weights: Based on the correlation between the TCM syndrome elements and TCM symptoms corresponding to the variable factors and the baseline data, calculate and determine the number of other nodes connected to each node, determine the node weight of the node based on the number of nodes, and determine the edges connecting the node to other nodes with which there is a correlation. Calculate the weight of each edge: Calculate and determine the weight of each edge based on the co-occurrence value between the TCM syndrome elements and TCM symptoms corresponding to the variable factors and the baseline data; Graph Construction: Construct a knowledge graph of diabetic nephropathy based on the nodes, node weights, edges, and edge weights.
2. The method as described in claim 1, characterized in that, The sample size estimation model is as follows: ; The patient sample size range is: ; The number of patient samples, HZ_Sum, is greater than or equal to { }, and the number of patient samples HZ_Sum is less than or equal to { }; The DKD phase ratio is LK_Sum:SK_Sum = 360:128; ; 。 3. The method as described in claim 1, characterized in that, Step S101, the step of collecting clinical data, includes: 1) Collect basic information data, baseline data, and clinical information data based on the data acquisition MDRD model; 2) Collect TCM syndrome data according to the TCM clinical syndrome level, and calculate and obtain TCM syndrome score based on the TCM syndrome data; in, The clinical data includes the basic information data, the baseline data, the clinical information data, and the TCM syndrome data; The glomerular filtration rate (GFR) is [value], serum creatinine (Cr) is [value], and age is [age]. When the patient being collected is male, the data collection MDRD model is as follows: ; When the patient being collected is female, the data collection MDRD model is as follows: ; The TCM clinical syndrome levels include mild, intermediate, and advanced levels. Each level corresponds to a different score, with higher levels corresponding to higher scores. When calculating the TCM syndrome score, the score for each TCM syndrome data is determined based on the TCM clinical syndrome level of the TCM syndrome data. Then, the scores of all TCM syndrome data corresponding to the syndrome element are summed to obtain the TCM syndrome score.
4. The method as described in claim 1, characterized in that, In step S101, the step of preprocessing and storing the basic clinical information data includes: Based on the composition and category of the basic clinical information data, as well as the collection time and number of follow-up visits of the basic clinical information data, the storage structure of the basic clinical information data is determined, and a basic clinical information data table is constructed. Based on the collection time and number of follow-up visits of the basic clinical information data, field description information is added to the collected basic clinical information data and stored in the basic clinical information data table; in, The storage structure of the basic clinical information data is {number, field description information, field name, field type}; The field description information includes: the number of follow-up visits and the collection time of the basic clinical information data; The storage of basic clinical information data tables includes a basic information data table, a baseline data table, a clinical information data table, and a traditional Chinese medicine syndrome data table.
5. The method as described in claim 1, characterized in that, In step S101, the step of preprocessing and storing the basic clinical information data further includes: Descriptive analysis and normality tests were performed on the basic clinical information data, and data characteristics were presented based on whether the quantitative basic clinical information data met a normal distribution. If the basic clinical information data does not conform to a normal distribution, the data characteristics are presented as the median and quartiles; When comparing two or more sets of basic clinical information data, an independent samples test is used for basic clinical information data that conforms to a normal distribution and has homogeneous variance; a nonparametric rank-sum test is used for basic clinical information data that does not conform to a normal distribution or has heterogeneous variance; and a count data test is used. The basic clinical information data was tested and standardized, and then scaled to the range [0, 1] by percentage.
6. The method as described in claim 1, characterized in that, In step S102, the step of labeling, classifying, and performing co-occurrence analysis on the exported data further includes: The data exported from the basic clinical information data is classified and labeled according to number, status judgment, basic information, physicochemical indicators, TCM symptoms, and TCM syndrome scores to obtain structured basic clinical information data, including basic information data, clinical information data, and TCM syndrome data. The labeling method of the basic information data is field name and field type, the labeling method of the clinical information data is field name, field type, and number of follow-up visits, and the labeling method of the TCM syndrome data is TCM syndrome label, TCM syndrome element, TCM symptoms, and TCM syndrome score. Based on the clinical information of the same patient during the same visit, co-occurrence analysis is performed on the structured basic clinical information data to determine the correlation between TCM syndrome elements, TCM symptoms and baseline data, and the co-occurrence value of the TCM syndrome elements, TCM symptoms and baseline data that have a correlation is recorded. The larger the co-occurrence value, the stronger the correlation between the TCM syndrome elements, TCM symptoms and baseline data. Based on the correlation and co-occurrence value, determine the variable factors associated with the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy; in, The correlation coefficient is... , When r>0, it indicates a positive correlation; when r<0, it indicates a negative correlation; and when r=0, it indicates zero correlation.
7. The method as described in claim 1, characterized in that, In step S104, the step of drawing a visualized knowledge graph of diabetic nephropathy based on the constructed knowledge graph of diabetic nephropathy further includes: Based on the node weights of the nodes constituting the diabetic nephropathy knowledge graph, the nodes are sorted, and the diameter of the nodes and the color of the node labels are determined according to the node sorting results. The diameter of the nodes and the intensity of the node label colors are both proportional to the node weights. The edges are sorted according to their weights. Each edge is labeled according to its sorting result, and the rendering degree of each edge is determined. The rendering degree includes the thickness, length, and color of the edge. The thickness, length, and color intensity of the edge are all proportional to the weight of the edge. Based on the diameter of the nodes and the color of the node labels, as well as the labels and rendering levels of the edges, the visualized knowledge graph of diabetic nephropathy is generated and drawn.
8. The method as described in claim 1, characterized in that, The basic clinical information data includes: basic information data, baseline data, clinical information data, and traditional Chinese medicine syndrome data; The basic information data includes: name, gender, age, height, weight, body mass index, waist-to-hip ratio, occupation, and ethnicity; The baseline data includes: past medical history data, personal data, physiochemical indicators, and functional evaluation data; The functional evaluation data include urinary protein excretion rate and glomerular filtration rate.
9. A system for implementing the method of constructing a knowledge graph of diabetic nephropathy based on DKD clinical data as described in claim 1, characterized in that, The system includes a data preprocessing module, a data classification and labeling module, a knowledge graph construction module, and a knowledge graph dynamic processing module; in, The data preprocessing module is used for S101, data collection, preprocessing and storage: collecting clinical data according to the diagnostic criteria for diabetic nephropathy, clinical stage, patient age and sample size estimation model F, grouping and sample estimation processing of the clinical data to obtain basic clinical information data, and preprocessing and storing the basic clinical information data. The data classification and annotation module is used for data export and annotation: exporting data from the stored basic clinical information data, and performing annotation, classification and co-occurrence analysis on the exported data to obtain structured basic clinical information data, and determining the variable factors related to the clinical proteinuria stage and end-stage renal disease of diabetic nephropathy, as well as the correlation and co-occurrence values between the variable factors; The knowledge graph construction module is used to construct a knowledge graph based on the correlation between various variable factors: determine nodes based on the variable factors and add node identifiers and node labels; calculate the weight of each node and the edges connecting each node according to the correlation and co-occurrence value between each node, as well as the type and weight of the edges; and construct a knowledge graph of diabetic nephropathy based on the nodes and the edges. The graph dynamic processing module is used for graph visualization and dynamic adjustment: based on the constructed diabetic nephropathy knowledge graph, a visualized diabetic nephropathy knowledge graph is drawn, and the visualized diabetic nephropathy knowledge graph is dynamically adjusted according to the weight of the edges; Specifically, when dynamically adjusting the visualized diabetic nephropathy knowledge graph, the number of visualized nodes and the weight values of the edges connecting the nodes are inversely proportional. The smaller the weight value of the edge of a node, the more visualized edges the node has, and the more uniform the distribution of correlations among the nodes. Conversely, the larger the weight value of the edge of a node, the fewer visualized edges the node has, and the stronger the correlations among the nodes.