Method and system for constructing and optimizing kidney essence deficiency diagnosis model

By constructing a quantitative relationship model of diabetic nephropathy based on the kidney essence deficiency syndrome score, the deficiency of diagnosis and prediction of kidney essence deficiency syndrome in the existing technology is solved, and a higher accuracy diagnosis and prediction of kidney essence deficiency syndrome is achieved, supporting the quantitative evaluation of disease progression of traditional Chinese medicine syndromes.

CN120376158APending Publication Date: 2025-07-25DONGZHIMEN HOSPITAL OF BEIJING UNIV OF CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410225595.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology lacks lower-specific, simple and inexpensive technical means for diagnosis and prediction of kidney essence deficiency syndrome, which leads to difficulties in the early diagnosis and treatment of diabetic nephropathy, especially the influencing factors and the lack of effective methods for the treatment of kidney essence deficiency syndrome.

Method used

A quantitative relationship model of diabetic nephropathy based on renal essence deficiency syndrome score was constructed. By grouping, standard balanced processing and grouping analysis of the basic clinical information data of patients with trace proteinuria, the optimal model training set was generated, the initial diagnostic model was constructed and trained, and the optimal diagnostic model of renal essence deficiency was finally obtained.

Benefits of technology

It has improved the diagnosis and prediction level and accuracy of kidney essence deficiency syndrome, clarified the qualitative quantitative relationship between kidney essence deficiency syndrome and the overall progress of diabetic nephropathy, and provided scientific support for quantitative evaluation of disease progression in traditional Chinese medicine syndromes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376158A_ABST
    Figure CN120376158A_ABST
Patent Text Reader

Abstract

The invention discloses a kidney essence deficiency diagnosis model construction and optimization construction method and system, and the method comprises the steps: carrying out the grouping processing of basic clinical information data, and obtaining the baseline data of kidney essence deficiency syndromes; performing standard equalization processing on the kidney essence deficiency syndrome baseline data to obtain kidney essence deficiency syndrome standard data, performing grouping analysis on the kidney essence deficiency syndrome standard data to generate an optimal model training set, and calculating and determining a kidney essence deficiency syndrome score of each group of kidney essence deficiency syndrome standard data after grouping; according to a diagnosis model structure, the kidney essence deficiency syndrome score and the quantification interval of the kidney essence deficiency syndrome score, constructing an initial kidney essence deficiency diagnosis model; and training the initial kidney essence deficiency diagnosis model based on each group of kidney essence deficiency syndrome standard data and the optimal model training set, and obtaining an optimal kidney essence deficiency diagnosis model according to a training result. By applying the method and the system provided by the invention, the diagnosis and prediction level and accuracy of the kidney essence deficiency syndrome are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of database technology, and more specifically, to a method and system for constructing and optimizing a kidney essence deficiency diagnosis model. Background Art

[0002] Diabetic kidney disease (DKD) is a complex multi-factorial disease involving genetic and environmental factors. Current understanding of the risk factors for diabetic kidney disease mainly comes from cohort studies with sample sizes ranging from a few hundred to several thousand subjects. Traditional Chinese medicine believes that the occurrence and development of diabetic kidney disease are roughly affected by multiple aspects such as congenital deficiency, weak viscera, improper diet, emotional depression, long-term use of drugs, and exogenous pathogenic factors of the six climatic factors. In terms of pathogenesis, most medical experts believe it is a condition of deficiency in the root and excess in the branch, which later develops into deficiency of the spleen and kidney, blood stasis blocking the collaterals, and even decline of yin and yang, with dampness-toxicity flooding with water. In addition, diabetic kidney disease is a long-term chronic disease with dynamic changes, and there are also different pathogenesis characteristics at different pathological stages of the disease.

[0003] At the same time, diabetic kidney disease is also a serious complication of diabetes and is the main cause of global chronic kidney disease (CKD) and end-stage renal disease (ESRD), and current treatment methods are limited. Therefore, early diagnosis and treatment are of great significance for preventing the progression of the disease. Currently, urinary albumin excretion rate and estimated glomerular filtration rate (eGFR) are widely accepted as indicators for evaluating the diagnosis and staging and progression of diabetic kidney disease, and microalbuminuria is also commonly recommended as the earliest clinically identifiable sign of diabetic kidney disease. However, urinary albumin has certain limitations and heterogeneity. Approximately 20%-40% of type 2 diabetes patients have declining renal function but no albuminuria, and the increase and decrease of glomerular filtration rate are not absolutely proportional to the progression of diabetic kidney disease. Therefore, there is a need to find a diagnostic and predictive method with lower idiosyncrasy, simplicity, affordability. In addition, through practice, the syndrome of kidney essence deficiency and other traditional Chinese medicine syndrome elements have been verified to be related to the progression of diabetic kidney disease. Currently, there is a lack of effective technical means to accurately determine the influencing factors and treatment of diabetic kidney disease, especially the syndrome of kidney essence deficiency.

[0004] To further explore the quantitative relationship of each stage of diabetic kidney disease based on the syndrome score of kidney essence deficiency, it is necessary to introduce a new method and system that can construct and train the best kidney essence deficiency diagnosis model based on the syndrome score of kidney essence deficiency, so as to solve the technical problem in the prior art of lacking technical means for diagnosing and predicting the syndrome of kidney essence deficiency with lower idiosyncrasy, simplicity, and affordability, assist in improving the accuracy and stability of the diagnosis and prognosis of urinary protein and serum creatinine, provide reference for the quantitative evaluation of disease progression by traditional Chinese medicine syndrome elements, and thus improve the level and accuracy of the diagnosis and prediction of the syndrome of kidney essence deficiency. Summary of the Invention

[0005] In view of the above-mentioned technical problems, the present invention provides a method and system for constructing and optimizing a kidney essence deficiency diagnosis model. According to the differences in the distribution of kidney essence deficiency scores among people in different disease stages of diabetic kidney disease, a quantitative relationship model between the clinical proteinuria stage and the end-stage kidney of diabetic kidney disease based on the kidney essence deficiency syndrome is constructed, that is, a kidney essence deficiency diagnosis model, so as to solve the technical problem in the prior art of lacking technical means for diagnosing and predicting the kidney essence deficiency syndrome with lower idiosyncrasy, simpler, easier, cheaper, and further clarify the qualitative and quantitative relationship between the kidney essence deficiency syndrome and the whole-stage progression of DKD, and better provide scientific technical support and high-precision model support for the quantitative evaluation of disease progression by traditional Chinese medicine syndrome elements and the diagnosis and prediction of the kidney essence deficiency syndrome, thereby improving the level and accuracy of the diagnosis and prediction of the kidney essence deficiency syndrome.

[0006] The present invention provides a method for constructing and optimizing a kidney essence deficiency diagnosis model, and the method includes:

[0007] S101, based on the number of patients in the microalbuminuria stage, group the basic clinical information data to obtain the baseline data of the kidney essence deficiency syndrome; S102, perform standard equalization processing on the baseline data of the kidney essence deficiency syndrome to obtain the standard data of the kidney essence deficiency syndrome, and perform grouped analysis on the standard data of the kidney essence deficiency syndrome according to the data classification analysis standard to generate an optimal model training set, and calculate and determine the kidney essence deficiency syndrome scores of each group of the grouped standard data of the kidney essence deficiency syndrome; S103, construct an initial kidney essence deficiency diagnosis model according to the diagnostic model structure, the kidney essence deficiency syndrome scores, and the quantization interval of the kidney essence deficiency syndrome scores; S104, train the initial kidney essence deficiency diagnosis model based on the standard data of each group of kidney essence deficiency syndromes and the optimal model training set, and obtain the optimal kidney essence deficiency diagnosis model according to the training results.

[0008] As described above, in step S101, the step of grouping the basic clinical information data further includes: accessing the basic clinical information data according to the factors affecting diabetic kidney disease and the diabetic kidney disease diagnosis criteria; grouping the basic clinical information data to obtain the baseline data of the kidney essence deficiency syndrome, where the baseline data of the kidney essence deficiency syndrome includes three groups, namely, the baseline data of the clinical proteinuria group, the baseline data of the renal failure group, and the baseline data of the microalbuminuria group; storing the baseline data of the kidney essence deficiency syndrome according to a preset data storage format; where the preset data storage format includes patient number, field description information, field name, field type, kidney essence deficiency syndrome score, total kidney essence deficiency syndrome score, and kidney essence deficiency syndrome score baseline; the kidney essence deficiency syndrome score baseline is determined based on the maximum value in the kidney essence deficiency syndrome scores.

[0009] As described above, in step S102, the steps of performing standard equalization processing on the baseline data of the kidney essence deficiency syndrome include: 1) data extraction: extracting the baseline data of the kidney essence deficiency syndrome according to the traditional Chinese medicine symptom assessment structure; 2) data standardization conversion: scaling the extracted baseline data of the kidney essence deficiency syndrome according to the data standardization interval [0, 1] to obtain the standardized baseline data of the kidney essence deficiency syndrome; 3) sample size equalization and test set: respectively extracting standardized baseline data of the kidney essence deficiency syndrome with the quantity of the preset test set fixed value from the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group included in the standardized baseline data of the kidney essence deficiency syndrome, and generating a full-scale data set; wherein, the traditional Chinese medicine symptom assessment structure includes the traditional Chinese medicine symptom grade and the traditional Chinese medicine symptom score, the traditional Chinese medicine symptom grade corresponds one-to-one with the traditional Chinese medicine symptom score, and the higher the traditional Chinese medicine symptom grade, the higher the traditional Chinese medicine symptom score; the full-scale data set is T = {t1, t2, t3, …… t n-1 , t n}, where n is the patient number.

[0010] As described above, in step S102, the steps of performing grouped analysis on the standard data of the kidney essence deficiency syndrome according to the data classification analysis standard to generate an optimal model training set and calculating and determining the kidney essence deficiency syndrome scores of each group of the grouped standard data of the kidney essence deficiency syndrome further include: respectively calculating the kidney essence deficiency syndrome scores, the total kidney essence deficiency syndrome scores, the kidney essence deficiency syndrome baseline, and the trend test value P corresponding to the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group based on the data classification analysis standard and the traditional Chinese medicine symptom score; sorting the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group in different disease stages from large to small according to the size of the trend test value P; respectively constructing confusion matrices corresponding to the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group according to the sorting results and the total kidney essence deficiency syndrome scores, and calculating and determining the optimal model training set according to the precision and recall rates corresponding to different disease stages in the confusion matrices; wherein, the data classification analysis standard includes disease stage, number of neighbors, test set ratio, verification method, and number of cross-validation folds; the disease stage includes the microalbuminuria stage, the clinical proteinuria stage, and the renal failure stage; the kidney essence deficiency syndrome score baseline is determined based on the maximum value in the kidney essence deficiency syndrome scores.

[0011] As described above, in step S103, the step of constructing the initial kidney essence deficiency diagnosis model according to the diagnosis model structure, the kidney essence deficiency syndrome score, and the quantization interval of the kidney essence deficiency syndrome score further includes: defining the structure of the quantization interval according to the kidney essence deficiency syndrome score, including the starting point a, window length b, and ending point c of the quantization interval; taking the point where the value of the kidney essence deficiency syndrome score is equal to zero as the origin of the kidney essence deficiency diagnosis model according to the diagnosis model structure, and based on the origin, determining the monotonic influence relationship of kidney essence deficiency staging according to the kidney essence deficiency syndrome score baseline, the kidney essence deficiency syndrome score, the quantization interval, the kidney essence deficiency syndrome score X, and the disease staging variable Y, and constructing the initial kidney essence deficiency diagnosis model.

[0012] As described above, the diagnosis model structure includes an origin, the kidney essence deficiency syndrome score X, and the disease staging variable Y; the quantization interval is [a, c], c = a + b, and the starting point a, window length b, and ending point c of the quantization interval are all integers equal to or greater than zero; the monotonic influence relationship of kidney essence deficiency staging is: Y = f(x), where X is an integer equal to or greater than zero; the initial kidney essence deficiency diagnosis model is: P = f(a, b).

[0013] As described above, after determining the monotonic influence relationship of kidney essence deficiency staging, it further includes the step of dividing the sampling data subset, specifically: calculating and obtaining the values of the corresponding disease staging variable Y according to different kidney essence deficiency syndrome scores X, and dividing the sampling data subset according to the comparison result of the values of the two disease staging variables Y: if f(x2) - f(x1) > 0 or f(x2) - f(x1) < 0, determining the sampling data subset according to the minimum value of x1 and x2; where x1 and x2 are both different values of the kidney essence deficiency syndrome score X; the sampling data subset is τ ab = T(a < x < a + b), a is the starting point of the quantization interval, b is the window length of the quantization interval, and T is the full dataset including all clinical proteinuria group baseline data, renal failure group baseline data, and microalbuminuria group baseline data.

[0014] As described above, in step S103, the step of constructing the initial kidney essence deficiency diagnosis model according to the diagnosis model structure, the kidney essence deficiency syndrome score, and the quantization interval of the kidney essence deficiency syndrome score further includes: constructing a curve AOC matrix according to the initial kidney essence deficiency diagnosis model P = f(a, b); calculating and obtaining the area under the curve (AUC) value of the microalbuminuria stage, the AUC value of the clinical proteinuria stage, and the AUC value of the renal failure stage respectively according to the starting point a and window length b of the quantization interval corresponding to the kidney essence deficiency syndrome score of each group of data of the clinical proteinuria group baseline data, renal failure group baseline data, and microalbuminuria group baseline data.

[0015] As described above, in step S104, the step of training the initial kidney essence deficiency diagnosis model based on the standard data of each group of kidney essence deficiency syndrome and the optimal model training set, and obtaining the optimal kidney essence deficiency diagnosis model according to the training results further includes: 1) According to the kidney essence deficiency syndrome score X, sample the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group respectively to obtain a sampling data subset τ ab ; 2) Based on the sampling data subset τ ab , train the initial kidney essence deficiency diagnosis model and obtain the corresponding training results; 3) Repeat step 2) until all the sampling data subsets τ ab are trained. Compare the area under the curve (AUC) values in all the training results, and determine the optimal kidney essence deficiency diagnosis model according to the comparison results; among them, when the AUC value of the curve is the largest, the initial kidney essence deficiency diagnosis model corresponding to the AUC value of the curve is the optimal kidney essence deficiency diagnosis model; the training results include the AUC value of the curve, precision, and recall rate.

[0016] Correspondingly, the present invention also provides a system for constructing and optimizing a kidney essence deficiency diagnosis model. The system includes a baseline data acquisition unit, a baseline data processing unit, an initial model construction unit, and an optimal model training unit; wherein, the baseline data acquisition unit is used to group and process the basic clinical information data based on the number of patients in the microalbuminuria stage to obtain the baseline data of the kidney essence deficiency syndrome; the baseline data processing unit is used to perform standard equalization processing on the baseline data of the kidney essence deficiency syndrome to obtain the standard data of the kidney essence deficiency syndrome, and perform grouped analysis on the standard data of the kidney essence deficiency syndrome according to the data classification analysis criteria to generate an optimal model training set, and calculate and determine the kidney essence deficiency syndrome score of each group of the grouped standard data of the kidney essence deficiency syndrome; the initial model construction unit is used to construct an initial kidney essence deficiency diagnosis model according to the diagnosis model structure, the kidney essence deficiency syndrome score, and the quantization interval of the kidney essence deficiency syndrome score; the optimal model training unit is used to train the initial kidney essence deficiency diagnosis model based on the standard data of each group of kidney essence deficiency syndrome and the optimal model training set, and obtain the optimal kidney essence deficiency diagnosis model according to the training results; the standard data of the kidney essence deficiency syndrome includes the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group; the data classification analysis criteria include disease stage, number of neighbors, test set ratio, verification method, and number of cross-validation folds, and the disease stage includes the microalbuminuria stage, the clinical proteinuria stage, and the renal failure stage; the diagnosis model structure includes the origin, the kidney essence deficiency syndrome score X, and the disease stage variable Y.

[0017] By applying the above technical solutions, the present invention realizes the construction of a quantitative relationship model for the clinical proteinuria stage and the end-stage kidney of diabetic kidney disease based on the kidney essence deficiency syndrome, that is, a kidney essence deficiency diagnosis model, according to the distribution differences of the kidney essence deficiency scores of people in different disease stages of diabetic kidney disease. It solves the technical problem in the prior art of lacking technical means for diagnosing and predicting the kidney essence deficiency syndrome with lower idiosyncrasy, simplicity, easy affordability, etc. It not only further clarifies the qualitative and quantitative relationship between the kidney essence deficiency syndrome and the whole-stage progression of DKD, but also can better provide scientific technical support and high-precision model support for the quantitative evaluation of disease progression by TCM syndrome elements and the diagnosis and prediction of the kidney essence deficiency syndrome, thereby improving the level and accuracy of the diagnosis and prediction of the kidney essence deficiency syndrome. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 The flowchart showing a method for constructing and optimizing a kidney essence deficiency diagnosis model proposed in an embodiment of the present invention;

[0020] Figure 2 The schematic diagram of the quantitative relationship of the kidney essence deficiency syndrome showing a method for constructing and optimizing a kidney essence deficiency diagnosis model proposed in an embodiment of the present invention;

[0021] Figure 3 The schematic diagram of the distribution of kidney essence deficiency scores in each stage of DKD showing a method for constructing and optimizing a kidney essence deficiency diagnosis model proposed in an embodiment of the present invention;

[0022] Figure 4 The schematic diagram of the error of kidney essence deficiency scores in each stage showing a method for constructing and optimizing a kidney essence deficiency diagnosis model proposed in an embodiment of the present invention;

[0023] Figure 5 The schematic diagram of the model learning curve showing a method for constructing and optimizing a kidney essence deficiency diagnosis model proposed in an embodiment of the present invention;

[0024] Figure 6 The schematic diagram of the confusion matrix showing a method for constructing and optimizing a kidney essence deficiency diagnosis model proposed in an embodiment of the present invention;

[0025] Figure 7 The schematic diagram of the principle of the diagnosis model showing a method for constructing and optimizing a kidney essence deficiency diagnosis model proposed in an embodiment of the present invention;

[0026] Figure 8Shows the schematic diagram of the AUC set of the DKD clinical proteinuria stage diagnosis model for the method of constructing and optimizing the kidney essence deficiency diagnosis model proposed in the embodiment of the present invention;

[0027] Figure 9 Shows the schematic diagram of the AUC set of the DKD end-stage kidney disease diagnosis model for the method of constructing and optimizing the kidney essence deficiency diagnosis model proposed in the embodiment of the present invention;

[0028] Figure 10 Shows the schematic diagram of the ROC curve of the optimal diagnosis model in the clinical proteinuria stage for the method of constructing and optimizing the kidney essence deficiency diagnosis model proposed in the embodiment of the present invention;

[0029] Figure 11 Shows the schematic diagram of the ROC curve of the optimal diagnosis model in the DKD end-stage kidney disease for the method of constructing and optimizing the kidney essence deficiency diagnosis model proposed in the embodiment of the present invention;

[0030] Figure 12 Shows the schematic diagram of the structure of the system for constructing and optimizing the kidney essence deficiency diagnosis model proposed in the embodiment of the present invention. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] The present invention provides a method for constructing and optimizing a kidney essence deficiency diagnosis model, as Figure 1 shown, the method includes the following steps:

[0033] S101, based on the number of patients in the microalbuminuria stage, group the basic clinical information data to obtain the baseline data of the kidney essence deficiency syndrome.

[0034] In this embodiment, in step S101, the step of grouping the basic clinical information data further includes:

[0035] According to the factors affecting diabetic kidney disease and the diabetic kidney disease diagnosis criteria, access the basic clinical information data;

[0036] Group the basic clinical information data to obtain the baseline data of the kidney essence deficiency syndrome, where the baseline data of the kidney essence deficiency syndrome includes three groups, namely the baseline data of the clinical proteinuria group, the baseline data of the renal failure group, and the baseline data of the microalbuminuria group;

[0037] Store the baseline data of kidney essence deficiency syndrome according to the preset data storage format;

[0038] Among them,

[0039] The preset data storage format includes patient number, field description information, field name, field type, kidney essence deficiency syndrome score, total kidney essence deficiency syndrome score, and baseline of kidney essence deficiency syndrome score;

[0040] The baseline of the kidney essence deficiency syndrome score is determined based on the maximum value in the kidney essence deficiency syndrome score.

[0041] To enable those skilled in the art to better understand the technical solution provided in this step, this step will be further described below by taking 384 patients as an example.

[0042] The diagnostic criteria for diabetic kidney disease can be formulated according to the "ADA 2020 Diabetes Diagnosis and Treatment Guidelines" and Mogensen staging. The specific formulated criteria can be as follows:

[0043] (1) Patients who meet the diagnostic criteria for diabetes; that is, fasting blood glucose ≥ 7.0 mmol / l for 8 hours; random blood glucose ≥ 11.1 mmol / l or oral glucose tolerance test (OGTT) ≥ 11.1 mmol / l for two hours; glycated hemoglobin ≥ 6.5%; if there are diabetes symptoms and a blood glucose value reaches the diabetes diagnostic criteria once, diabetes can be diagnosed.

[0044] (2) Patients with a clear history of diabetes and meeting one of the following conditions can be diagnosed with DKD: urine albumin / creatinine ratio ≥ 30 mg / g or urine albumin excretion rate ≥ 30 mg / 24 h;

[0045] (3) Age ≥ 18 years old, regardless of gender.

[0046] S102. Perform standard equalization processing on the baseline data of kidney essence deficiency syndrome to obtain standard data of kidney essence deficiency syndrome, and perform grouped analysis on the standard data of kidney essence deficiency syndrome according to the data classification analysis standard to generate an optimal model training set, and calculate and determine the kidney essence deficiency syndrome scores of each group of standard data of kidney essence deficiency syndrome after grouping.

[0047] In this embodiment, in step S102, the step of performing standard equalization processing on the baseline data of kidney essence deficiency syndrome includes:

[0048] 1) Data extraction: Extract the baseline data of kidney essence deficiency syndrome according to the traditional Chinese medicine symptom assessment structure;

[0049] 2) Data standardization conversion: Scale the extracted baseline data of kidney essence deficiency syndrome according to the data standardization interval [0, 1] to obtain standardized baseline data of kidney essence deficiency syndrome;

[0050] 3) Sample size balance and test set: According to the preset fixed value of the test set, standardized baseline data of kidney essence deficiency syndrome are respectively extracted from the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group included in the standardized baseline data of kidney essence deficiency syndrome, and the quantity of each extraction is the preset fixed value of the test set, and a full dataset is generated;

[0051] Among them,

[0052] The traditional Chinese medicine syndrome assessment structure includes the traditional Chinese medicine syndrome grade and the traditional Chinese medicine syndrome score. The traditional Chinese medicine syndrome grade corresponds one-to-one with the traditional Chinese medicine syndrome score, and the higher the traditional Chinese medicine syndrome grade, the higher the traditional Chinese medicine syndrome score;

[0053] The full dataset is T = {t1, t2, t3, …… t n-1 , t n} where n is the patient number.

[0054] In this embodiment, in step S102, the step of grouping and analyzing the standard data of kidney essence deficiency syndrome according to the data classification and analysis standard to generate an optimal model training set and calculating and determining the kidney essence deficiency syndrome score of each group of standardized data of kidney essence deficiency syndrome after grouping further includes:

[0055] Based on the data classification and analysis standard and the traditional Chinese medicine syndrome score, calculate the kidney essence deficiency syndrome score, the total kidney essence deficiency syndrome score, the baseline of kidney essence deficiency syndrome, and the trend test value P corresponding to the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group respectively;

[0056] According to the magnitude of the trend test value P, sort the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group in different disease stages from large to small;

[0057] According to the sorting result and the total kidney essence deficiency syndrome score, construct confusion matrices corresponding to the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group respectively, calculate, and determine the optimal model training set according to the precision and recall corresponding to different disease stages in the confusion matrix;

[0058] Among them,

[0059] The data classification and analysis standard includes disease stage, number of neighbors, test set ratio, verification method, and number of cross-validation folds;

[0060] The disease stage includes the microalbuminuria stage, the clinical proteinuria stage, and the renal failure stage;

[0061] The scoring baseline for the kidney essence deficiency syndrome is determined based on the maximum value in the scoring for the kidney essence deficiency syndrome.

[0062] To enable those skilled in the art to better understand the technical solution provided in this step, this step will still be further described by taking 384 patients as an example of the sample size.

[0063] Data cleaning and data standardization: Rearrange the collected clinical information and data, sort out the traditional Chinese medicine syndrome assessment form and calculate the scores to complete data cleaning. Considering the differences in the total scores of each syndrome element in the traditional Chinese medicine scoring scale, the Minmax Scaler standardization method is used for the data to scale the data information into the interval [0,1].

[0064] Sample size balance: Since there are large deviations in the number of patients in each stage of diabetic nephropathy, before training for further classification problems, the data must be balanced in terms of sample size. Because the focus of this invention is on the traditional Chinese medicine diagnosis model for each stage of diabetic nephropathy, and considering finding the factors for the progression of diabetic nephropathy as much as possible, a sampling method is adopted.

[0065] Fixing the test set: To avoid the overlap between the test set and the training set caused by the oversampling method, the method of fixing the test set is adopted. The preset test set fixed value is selected as 90, that is, a total of 90 patients' data are randomly fixed in advance, that is, 30 test samples in each of the 3 groups for each stage, to ensure that the test set does not coincide with the training set to achieve the best test purpose.

[0066] In addition, in some embodiments, it is also necessary to perform missing value imputation on the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group obtained after grouping. Specifically, programs can be written using R language version 4.0 and python3.7 version. For the blank data with a vacancy rate less than 10% in the collected data information, data filling is performed through the median imputation method and the adjacent imputation method, and the questionnaires with more than 20% content missing are excluded; the outlier data is removed, and the categorical variables are standardized and transformed.

[0067] In addition, in some embodiments, before the grouped analysis and processing of the baseline data of the kidney essence deficiency syndrome, it is also necessary to perform descriptive analysis on the baseline data of the kidney essence deficiency syndrome; perform a normality test on the quantitative baseline data of the kidney essence deficiency syndrome, and present the data characteristics according to whether it meets the normal distribution. If it does not conform to the normal distribution, the data characteristics are presented with the median and the interquartile range (IQR). For the comparison between two independent samples, for measurement data that conform to the normal distribution and have homogeneous variances, an independent samples t-test is used; for those that do not conform to the normal distribution or have inhomogeneous variances, a non-parametric rank sum test is used; for count data, X 2Examination. Count data is expressed as frequency n and percentage (%). Continuously standardize all continuous variable characteristic data, process the data through MinMax Scaler, and scale the data to between [0,1].

[0068] S103. Construct an initial kidney essence deficiency diagnosis model based on the diagnostic model structure, the kidney essence deficiency syndrome score, and the quantization interval of the kidney essence deficiency syndrome score.

[0069] In this embodiment, in step S103, the step of constructing an initial kidney essence deficiency diagnosis model based on the diagnostic model structure, the kidney essence deficiency syndrome score, and the quantization interval of the kidney essence deficiency syndrome score further includes:

[0070] Define the structure of the quantization interval according to the kidney essence deficiency syndrome score, including the starting point a, window length b, and ending point c of the quantization interval;

[0071] Take the point where the value of the kidney essence deficiency syndrome score is equal to zero as the origin of the kidney essence deficiency diagnosis model according to the diagnostic model structure. Based on the origin, determine the monotonic influence relationship of kidney essence deficiency staging according to the kidney essence deficiency syndrome score baseline, the kidney essence deficiency syndrome score, the quantization interval, the kidney essence deficiency syndrome score X, and the disease staging variable Y, and construct the initial kidney essence deficiency diagnosis model.

[0072] The diagnostic model structure includes the origin, the kidney essence deficiency syndrome score X, and the disease staging variable Y;

[0073] The quantization interval is [a, c], c = a + b, and the starting point a, window length b, and ending point c of the quantization interval are all integers equal to or greater than zero;

[0074] The monotonic influence relationship of kidney essence deficiency staging is: Y = f(x), where X is an integer equal to or greater than zero;

[0075] The initial kidney essence deficiency diagnosis model is: P = f(a, b).

[0076] In this embodiment, after determining the monotonic influence relationship of kidney essence deficiency staging, it further includes the step of dividing the sampling data subset, specifically:

[0077] Calculate and obtain the value of the corresponding disease staging variable Y according to different kidney essence deficiency syndrome scores X, and divide the sampling data subset according to the comparison result of the values of the two disease staging variables Y:

[0078] If f(x2) - f(x1) > 0 or f(x2) - f(x1) < 0, determine the sampling data subset according to the minimum value of x1 and x2;

[0079] Among them, both x1 and x2 are different values of the Kidney Essence Deficiency Syndrome Score X;

[0080] The sampled data subset is τ ab = T(a < x < a + b), a is the starting point of the quantization interval, b is the window length of the quantization interval, and T is the full dataset containing the baseline data of all clinical proteinuria groups, renal failure groups, and microalbuminuria groups.

[0081] In this embodiment, in step S103, the step of constructing the initial Kidney Essence Deficiency diagnosis model according to the diagnosis model structure, the Kidney Essence Deficiency Syndrome Score, and the quantization interval of the Kidney Essence Deficiency Syndrome Score further includes:

[0082] Construct a curve AOC matrix according to the initial Kidney Essence Deficiency diagnosis model P = f(a, b);

[0083] According to the starting point a and the window length b of the quantization intervals corresponding to the Kidney Essence Deficiency Syndrome Scores of each group of data of the clinical proteinuria group baseline data, renal failure group baseline data, and microalbuminuria group baseline data, calculate and obtain the curve area AUC value in the microalbuminuria stage, the curve area AUC value in the clinical proteinuria stage, and the curve area AUC value in the renal failure stage respectively.

[0084] To enable those skilled in the art to better understand the technical solution provided in this step, this step will be further described below by taking 384 patients as an example of the sample size.

[0085] When constructing the initial Kidney Essence Deficiency diagnosis model, the K-Nearest Neighbor (KNN) classification algorithm can be used. The KNN algorithm is a method of classifying the K nearest (i.e., the closest in the feature space) samples near a certain sample in the feature space into the same category, and it is a commonly used and effective classification algorithm in machine learning algorithms.

[0086] Such as Figure 2 shown, to improve the simplicity of the model in clinical applications, the KNN algorithm is further improved, and the target attribute is changed to the Kidney Essence Deficiency Syndrome Score interval. The score sets of Kidney Essence Deficiency of the three groups of patients are regarded as the feature space, and the center origin is defined as 0 points of Kidney Essence Deficiency. The farther the sample is from the center origin, the higher the score. According to the different distributions of the Kidney Essence Deficiency Syndrome Scores in each stage of DKD of the three groups, calculate and set the Kidney Essence Deficiency Syndrome Score intervals (Windows) in each stage of DKD as the initial Kidney Essence Deficiency diagnosis model.

[0087] The starting point and ending point of the interval are the start value and end value of the distinguish value for the diagnosis of kidney essence deficiency between different stages; the difference between the end value and the start value is named the windows size of the diagnosis. Next, taking the diagnosis stage as the dependent variable and the scoring interval of kidney essence deficiency as the independent variable, a diagnostic stage equation can be obtained using Logistic regression. Therefore, each scoring interval of kidney essence deficiency can correspond to a prediction result. By setting the start value of the scoring interval of kidney essence deficiency and the interval size as the independent variables, and the AUC value of the curve of the corresponding diagnostic model as the dependent variable, the best combination of independent variables (i.e., the starting point of the interval and the scale of the interval size) with the best AUC score can be calculated.

[0088] Model verification and evaluation: When constructing the model, taking the AUC value of the curve as the dependent variable, the model with a higher AUC value of the curve is selected as the optimal diagnostic model for kidney essence deficiency.

[0089] S104, training the initial diagnostic model for kidney essence deficiency based on the standard data of kidney essence deficiency syndrome for each group and the optimal model training set, and obtaining the optimal diagnostic model for kidney essence deficiency according to the training results.

[0090] In this embodiment, in step S104, the step of training the initial diagnostic model for kidney essence deficiency based on the standard data of kidney essence deficiency syndrome for each group and the optimal model training set, and obtaining the optimal diagnostic model for kidney essence deficiency according to the training results further includes:

[0091] 1) According to the score X of kidney essence deficiency syndrome, sampling the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group respectively to obtain the sampling data subset τ ab ;

[0092] 2) Based on the sampling data subset τ ab , training the initial diagnostic model for kidney essence deficiency and obtaining the corresponding training results;

[0093] 3) Repeating step 2) until all the sampling data subsets τ ab are trained, comparing the AUC values of the curves in all the training results, and determining the optimal diagnostic model for kidney essence deficiency according to the comparison results;

[0094] Among them,

[0095] when the AUC value of the curve is the largest, the initial diagnostic model for kidney essence deficiency corresponding to the AUC value of the curve is the optimal diagnostic model for kidney essence deficiency;

[0096] The training results include the AUC value of the curve area, precision, and recall rate.

[0097] To enable those skilled in the art to better understand the technical solution provided by the present invention, this step will be further described below by taking 384 patients as an example of the sample size.

[0098] a. Baseline analysis of kidney essence deficiency syndrome

[0099] As shown in Table 1 and Figure 3 shown, the baseline distribution of kidney essence deficiency syndrome in each stage of DKD: Univariate logistic regression analysis was performed on the kidney essence deficiency syndrome of the baseline data of the three groups of kidney essence deficiency syndrome, and it was presented in median [IQR]. The results showed that there were significant differences in the scores of kidney essence deficiency syndrome among the three groups (P<0.001). The total scores of kidney essence deficiency syndrome in the baseline data of each group of kidney essence deficiency syndrome were: renal failure group > clinical proteinuria group > microalbuminuria group.

[0100] Table 1

[0101]

[0102] b. The distribution of the scores of kidney essence deficiency syndrome among the three groups was grouped by the score range of kidney essence deficiency, and the distribution of the scores of kidney essence deficiency in each group was further statistically analyzed.

[0103] As shown in Table 2 and Figure 4 shown, in the microalbuminuria stage, the number of people with a score of less than 6 for kidney essence deficiency syndrome was 290, accounting for 75.5% of all patients in the microalbuminuria stage; the scores of kidney essence deficiency syndrome in patients with clinical proteinuria were relatively uniform, with a large degree of dispersion; in the end-stage of the kidney, the number of people with a score of more than 12 for kidney essence deficiency was 57, accounting for 44.5% of the total number of patients in the end-stage of the kidney.

[0104] Table 2

[0105]

[0106]

[0107] As Figure 5 shown, the KNN classification model.

[0108] Classification was performed using the KNN machine learning method. The classification variable was the disease stage, the independent variable in the model was the total score of kidney essence deficiency, and the model parameters were adjusted as follows: number of neighbors: 5; proportion of the test set: a fixed test set of 88 cases; verification method: cross-validation; number of cross-validation folds: 10 times.

[0109] As shown in Table 3 and Figure 6As shown, the optimal model training set was obtained, with an accuracy of 61.4% in the test set. The confusion matrix of the test set was used to evaluate the model, and the performance of the diagnosis model for kidney essence deficiency in each stage of DKD was obtained. It can be seen from the chart that the diagnostic performance of the model is poor and cannot classify well. The precision of the clinical proteinuria group is 56.4%, and the recall rate is only 58.7%; the precision of the end-stage kidney is as high as 100%, but the recall rate is as low as 40.0%, indicating that there is a certain bias in the model.

[0110] Table 3

[0111] name precision recall F1-score support Microalbuminuria stage 0.638 0.682 0.659 44 Clinical proteinuria stage 0.564 0.611 0.587 36 End-stage kidney 0.1 0.25 0.4 8.0 macro avg 0.734 0.514 0.549 88.0 weighted avg 0.641 0.614 0.606 88.0

[0112] c. Quantification model of "score window" for kidney essence deficiency

[0113] In order to thoroughly study the impact of essence deficiency on staging, considering that different groups with different essence deficiency scores may have different impacts on staging. The dataset T = {t1, t2... t872} was designed and used, and 872 patients were grouped according to different stages.

[0114] As Figure 7 shown, the essence deficiency score was defined as the independent variable X = {2, 4, 6...} as ordinal data; the staging was the dependent variable Y = {3, 4, 5}, also ordinal data. Assuming that the essence deficiency score x has a monotonic effect on the staging y = f(x), there must be a relationship: f(x2) - f(x1) > 0 or f(x2) - f(x1) < 0, and only one of the two situations exists. The data subset was divided τab = T(a < x < a + b).

[0115] where a is the starting position of the window interval (Window start), that is, the training set was divided according to the essence deficiency scores of each group. Starting from the essence deficiency score a of the patient, the window size (Window size) is b. Therefore, the quantification interval of the staging of this model is [a, a + b]. The AUC matrix was constructed with this model. The construction of the AUC matrix mainly considered that the calculation method of AUC can balance sensitivity and specificity and is a general balance measure of accuracy. The pseudocode of the calculation process is as follows:

[0116] AUC matrix = all 0.5 (baseline) matrix; for all starting positions a: 0, 2, 4, 6, 8,......;

[0117] for all window sizes b: 2, 4, 6, 8,......;

[0118] Each group of datasets was sampled according to the essence deficiency score τab;

[0119] The same Logistic regression parameters were used for training LR(τ

[0120] If the size of the dataset LEN(τab) > 30 and the classification balance ratio is between 0.1 and 0.9

[0121] AUC(a,b) = r ab ;

[0122] Otherwise

[0123] AUC(a,b) = r ab ;

[0124] As Figure 8 and Figure 9 shown, according to the distribution of the kidney essence deficiency score in the actual research, the quantitative relationship models between microalbuminuria and clinical albuminuria, and between clinical albuminuria and end-stage kidney disease are calculated respectively. According to the above calculation principle and method, two AUC Tables of "3∣4" and "4∣5" are obtained, with AUC as the dependent variable, the starting point a of the interval, and the interval size b as the independent variables.

[0125] It can be seen that the quantitative relationship model with the best AUC value is the lightest-colored grid area in the figure, that is:[[]]

[0126] The starting score of the quantitative window in the clinical albuminuria stage: a = 8 points, the window size b = 10 points, and the quantitative interval score is [8, 18]; the accuracy rate is 75.8%, and the AUC of the ROC curve: 0.79. The result of drawing the ROC curve of this optimal diagnostic model is as Figure 10 shown.

[0127] As Figure 11 shown, the starting score of the quantitative window for end-stage kidney disease a = 24 points, the window size b = 16 points. Therefore, the quantitative scoring interval for end-stage kidney disease is: [24, 40], the accuracy rate: 87.5%, and the AUC: 0.87. The result of drawing the ROC curve of this optimal diagnostic model.

[0128] By applying the above technical solutions, a quantitative relationship model for the clinical albuminuria stage and end-stage kidney disease of diabetic nephropathy based on the kidney essence deficiency syndrome, that is, the kidney essence deficiency diagnosis model, is constructed according to the differences in the distribution of the kidney essence deficiency scores of people in different disease stages of diabetic nephropathy. This solves the technical problem in the existing technology of lacking technical means for diagnosing and predicting the kidney essence deficiency syndrome with lower idiosyncrasy, simplicity, affordability, and convenience. It not only further clarifies the qualitative and quantitative relationship between the kidney essence deficiency syndrome and the full-stage progression of DKD, but also can better provide scientific technical support and a high-precision model support for the quantitative evaluation of disease progression by traditional Chinese medicine syndrome elements and the diagnosis and prediction of the kidney essence deficiency syndrome, thereby improving the level and accuracy of the diagnosis and prediction of the kidney essence deficiency syndrome.

[0129] In addition, by applying the above technical solutions, it was also found that there were significant statistical differences in the scores of kidney essence deficiency syndrome between the clinical proteinuria stage and the end-stage of diabetic kidney disease (P<0.05), which proved that there was also a linear positive correlation between kidney essence deficiency syndrome and the whole-course progression of DKD, further clarified the specific quantitative values of this qualitative relationship, and made the quantitative relationship model constructed by the Logistic regression analysis method (i.e., the kidney essence deficiency diagnosis model) have a high recognition accuracy. It can be used as an effective auxiliary tool for judging the stage of DKD by the score of kidney essence deficiency syndrome in clinical practice, further improving the accuracy of the score judgment of kidney essence deficiency syndrome, further enriching the diagnostic methods of diabetic kidney disease, and reducing the heterogeneity and instability of traditional modern medical indicators.

[0130] Corresponding to the method for constructing and optimizing a kidney essence deficiency diagnosis model in an embodiment of the present invention, the present invention also discloses a system for constructing and optimizing a kidney essence deficiency diagnosis model, as Figure 12 shown. The system includes a baseline data acquisition unit, a baseline data processing unit, an initial model construction unit, and an optimal model training unit;

[0131] Among them,

[0132] The baseline data acquisition unit is used to group and process the basic clinical information data based on the number of patients in the microalbuminuria stage to obtain the baseline data of kidney essence deficiency syndrome;

[0133] The baseline data processing unit is used to perform standard equalization processing on the baseline data of kidney essence deficiency syndrome to obtain the standard data of kidney essence deficiency syndrome, group and analyze the standard data of kidney essence deficiency syndrome according to the data classification analysis standard to generate an optimal model training set, and calculate and determine the scores of kidney essence deficiency syndrome for each group of the grouped standard data of kidney essence deficiency syndrome;

[0134] The initial model construction unit is used to construct an initial kidney essence deficiency diagnosis model according to the diagnostic model structure, the score of kidney essence deficiency syndrome, and the quantitative interval of the score of kidney essence deficiency syndrome;

[0135] The optimal model training unit is used to train the initial kidney essence deficiency diagnosis model based on the standard data of kidney essence deficiency syndrome for each group and the optimal model training set, and obtain the optimal kidney essence deficiency diagnosis model according to the training results;

[0136] The standard data of kidney essence deficiency syndrome includes the baseline data of the clinical proteinuria group, the baseline data of the renal failure group, and the baseline data of the microalbuminuria group;

[0137] The data classification analysis standard includes disease stage, number of neighbors, test set ratio, verification method, and number of cross-validation folds. The disease stage includes the microalbuminuria stage, the clinical proteinuria stage, and the renal failure stage;

[0138] The diagnostic model structure includes the origin, the score X of kidney essence deficiency syndrome, and the disease staging variable Y.

[0139] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A method for constructing and optimizing a diagnosis model of kidney essence deficiency, characterized in that The method includes: S101. Based on the number of patients in the microalbuminuria stage, group the basic clinical information data to obtain the baseline data of the kidney essence deficiency syndrome. S102. Perform standard equalization processing on the baseline data of the kidney essence deficiency syndrome to obtain the standard data of the kidney essence deficiency syndrome. Then, group and analyze the standard data of the kidney essence deficiency syndrome according to the data classification and analysis criteria to generate an optimal model training set, and calculate and determine the kidney essence deficiency syndrome scores of each group of the grouped standard data of the kidney essence deficiency syndrome. S103. Construct an initial kidney essence deficiency diagnosis model based on the diagnostic model structure, the kidney essence deficiency syndrome scores, and the quantization interval of the kidney essence deficiency syndrome scores. S104. Train the initial kidney essence deficiency diagnosis model based on the standard data of the kidney essence deficiency syndrome in each group and the optimal model training set, and obtain the optimal kidney essence deficiency diagnosis model according to the training results.

2. The method according to claim 1, wherein In step S101, the step of grouping the basic clinical information data further includes: Access the basic clinical information data according to the factors affecting diabetic nephropathy and the diabetic nephropathy diagnosis criteria. Group the basic clinical information data to obtain the baseline data of the kidney essence deficiency syndrome. Among them, the baseline data of the kidney essence deficiency syndrome includes three groups, namely the baseline data of the clinical proteinuria group, the baseline data of the renal failure group, and the baseline data of the microalbuminuria group. Store the baseline data of the kidney essence deficiency syndrome according to the preset data storage format. Wherein, The preset data storage format includes patient number, field description information, field name, field type, kidney essence deficiency syndrome score, total kidney essence deficiency syndrome score, and kidney essence deficiency syndrome score baseline. The kidney essence deficiency syndrome score baseline is determined based on the maximum value in the kidney essence deficiency syndrome scores.

3. The method according to claim 1, wherein In step S102, the step of performing standard equalization processing on the baseline data of the kidney essence deficiency syndrome includes: 1) Data extraction: Extract the baseline data of the kidney essence deficiency syndrome according to the traditional Chinese medicine syndrome assessment structure. 2) Data standardization conversion: Scale the extracted baseline data of the kidney essence deficiency syndrome according to the data standardization interval [0, 1] to obtain the standardized baseline data of the kidney essence deficiency syndrome. 3) Sample size equalization and test set: According to the preset fixed value of the test set, extract the standardized baseline data of the kidney essence deficiency syndrome with the same quantity of the preset fixed value of the test set from the three groups of data, namely the baseline data of the clinical proteinuria group, the baseline data of the renal failure group, and the baseline data of the microalbuminuria group, included in the standardized baseline data of the kidney essence deficiency syndrome, and generate a full-scale data set. Wherein, The traditional Chinese medicine syndrome assessment structure includes traditional Chinese medicine syndrome grades and traditional Chinese medicine syndrome scores. The traditional Chinese medicine syndrome grades correspond one-to-one with the traditional Chinese medicine syndrome scores, and the higher the traditional Chinese medicine syndrome grade, the higher the traditional Chinese medicine syndrome score. The full dataset is T = {t1, t2, t3, …… t n-1 , t n}, where n is the patient number.

4. The method according to claim 1, wherein In step S102, the step of grouping and analyzing the standard data of the kidney essence deficiency syndrome according to the data classification and analysis criteria to generate an optimal model training set, and calculating and determining the kidney essence deficiency syndrome scores of each group of the grouped standard data of the kidney essence deficiency syndrome further includes: Based on the above data classification and analysis criteria and the TCM syndrome score, calculate the Kidney Essence Deficiency syndrome scores, total Kidney Essence Deficiency syndrome scores, baselines of Kidney Essence Deficiency syndrome, and trend test value P corresponding one by one to the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group respectively; According to the magnitude of the trend test value P, sort the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group at different disease stages from large to small; According to the sorting results and the total Kidney Essence Deficiency syndrome score, construct confusion matrices corresponding to the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group respectively, calculate, and determine the optimal model training set according to the precision and recall corresponding to different disease stages in the confusion matrix; Among them, the data classification and analysis criteria include disease stage, number of neighbors, test set ratio, verification method, and number of cross-validation folds; the disease stage includes the microalbuminuria stage, the clinical proteinuria stage, and the renal failure stage; the baseline of the Kidney Essence Deficiency syndrome score is determined based on the maximum value in the Kidney Essence Deficiency syndrome score.

5. The method according to claim 1, characterized in that, In step S103, the step of constructing the initial Kidney Essence Deficiency diagnosis model according to the diagnostic model structure, the Kidney Essence Deficiency syndrome score, and the quantization interval of the Kidney Essence Deficiency syndrome score further includes: According to the Kidney Essence Deficiency syndrome score, define the structure of the quantization interval, including the starting point a, window length b, and ending point c of the quantization interval; According to the diagnostic model structure, take the point where the value of the Kidney Essence Deficiency syndrome score is equal to zero as the origin of the Kidney Essence Deficiency diagnosis model. Based on this origin, according to the baseline of the Kidney Essence Deficiency syndrome score, the Kidney Essence Deficiency syndrome score, the quantization interval, the Kidney Essence Deficiency syndrome score X, and the disease stage variable Y, determine the monotonic influence relationship of the Kidney Essence Deficiency stage, and construct the initial Kidney Essence Deficiency diagnosis model.

6. The method according to claim 5, characterized in that The diagnostic model structure includes the origin, the Kidney Essence Deficiency syndrome score X, and the disease stage variable Y; The quantization interval is [a, c], c = a + b, and the starting point a, window length b, and ending point c of the quantization interval are all integers equal to or greater than zero; The monotonic influence relationship of the Kidney Essence Deficiency stage is: Y = f(x), where X is an integer equal to or greater than zero; The initial Kidney Essence Deficiency diagnosis model is: P = f(a, b).

7. The method according to claim 5, characterized in that, After determining the monotonic influence relationship of the Kidney Essence Deficiency stage, it further includes the step of dividing the sampling data subset, specifically: According to different Kidney Essence Deficiency syndrome scores X, calculate and obtain the corresponding values of the disease stage variable Y, and divide the sampling data subset according to the comparison results of the two values of the disease stage variable Y: If f(x2) - f(x1) > 0 or f(x2) - f(x1) < 0, determine the sampling data subset according to the minimum value of x1 and x2; where x1 and x2 are both different values of the Kidney Essence Deficiency syndrome score X; The sampled data subset is τ ab = T(a < x < a + b), where a is the starting point of the quantization interval, b is the window length of the quantization interval, and T is the full dataset containing all the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group.

8. The method according to claim 1, wherein In step S103, the step of constructing the initial Kidney Essence Deficiency diagnosis model according to the diagnostic model structure, the Kidney Essence Deficiency syndrome score, and the quantization interval of the Kidney Essence Deficiency syndrome score further includes: Construct a curve AOC matrix according to the initial kidney essence deficiency diagnosis model P = f(a, b). According to the starting points a of the quantization intervals corresponding to the kidney essence deficiency syndrome scores of each group of data in the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group, and the window lengths b of the quantization intervals, calculate and obtain the area under the curve (AUC) values in the microalbuminuria stage, the clinical proteinuria stage, and the renal failure stage respectively.

9. The method according to claim 1, wherein In step S104, the step of training the initial kidney essence deficiency diagnosis model based on the standard data of kidney essence deficiency syndrome for each group and the optimal model training set, and obtaining the optimal kidney essence deficiency diagnosis model according to the training results further includes: 1) According to the score X of kidney essence deficiency syndrome, sample the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group respectively to obtain the sampling data subset τ ab ; 2) Based on the sampled data subset τ ab , train the initial diagnosis model for kidney essence deficiency and obtain the corresponding training results; 3) Repeat step 2) until all of the sampled data subsets τ ab are trained. Compare the AUC values of the curve areas in all of the training results, and determine the optimal kidney essence deficiency diagnosis model according to the comparison results; Wherein,[[]] When the AUC value of the curve area is the largest, the initial kidney essence deficiency diagnosis model corresponding to the AUC value of the curve area is the optimal kidney essence deficiency diagnosis model. The training results include the AUC value of the curve area, precision, and recall rate.

10. A system for implementing the method of constructing and optimizing the kidney essence deficiency diagnosis model according to claim 1, characterized in that, The system includes a baseline data acquisition unit, a baseline data processing unit, an initial model construction unit, and an optimal model training unit. Wherein,[[]] The baseline data acquisition unit is used to group and process the basic clinical information data based on the number of patients in the microalbuminuria stage to obtain the baseline data of the kidney essence deficiency syndrome. The baseline data processing unit is used to perform standard equalization processing on the baseline data of the kidney essence deficiency syndrome to obtain the standard data of the kidney essence deficiency syndrome, group and analyze the standard data of the kidney essence deficiency syndrome according to the data classification and analysis criteria to generate an optimal model training set, and calculate and determine the kidney essence deficiency syndrome scores of the standard data of the kidney essence deficiency syndrome for each group after grouping. The initial model construction unit is used to construct an initial kidney essence deficiency diagnosis model according to the diagnosis model structure, the kidney essence deficiency syndrome score, and the quantization interval of the kidney essence deficiency syndrome score. The optimal model training unit is used to train the initial kidney essence deficiency diagnosis model based on the standard data of kidney essence deficiency syndrome for each group and the optimal model training set, and obtain the optimal kidney essence deficiency diagnosis model according to the training results. The standard data of the kidney essence deficiency syndrome includes the baseline data of the clinical proteinuria group, the renal failure group, and the microalbuminuria group. The data classification and analysis criteria include disease stage, number of neighbors, test set ratio, verification method, and number of cross-validation folds. The disease stage includes the microalbuminuria stage, the clinical proteinuria stage, and the renal failure stage. The diagnosis model structure includes an origin, a kidney essence deficiency syndrome score X, and a disease stage variable Y.