A method and device for constructing and evaluating a highly reliable glomerular filtration rate prediction model

By building multiple local models and building global models, the problem of insufficient reliability of existing GFR prediction models in small samples is solved, and high-reliability GFR prediction and evaluation is achieved.

CN115206534BActive Publication Date: 2025-05-06XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210915183.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-01
Publication Date
2025-05-06
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

The existing glomerular filtration rate (GFR) prediction models are ineffective in small samples and fail to effectively evaluate the reliability of the model and predicted results.

Method used

By constructing multiple local models and using the median values ​​of these models to build a global model, combining data preprocessing, feature selection and log transformation techniques, a highly reliable GFR prediction model is proposed. The model is evaluated through model stability and predicted results stability to ensure the reliability of predicted results.

Benefits of technology

Highly reliable GFR prediction in small samples is achieved, which improves the reliability of the model and prediction results, and can predict GFR values ​​more accurately and stably.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115206534B_ABST
    Figure CN115206534B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing and evaluating a highly reliable glomerular filtration rate prediction model, comprising: step (1): data preprocessing; step (2): local model setting; step (3): local model construction; step (4): global model construction; step (5): model stability calculation of the global model; step (6): accuracy calculation of GFR prediction of the global model; step (7): stability calculation of the GFR prediction result; step (8): model selection; the method constructs multiple local models, and a technical solution of constructing a global model through the median model of these models can be used to obtain a highly reliable GFR prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of glomerular filtration rate prediction models, and in particular relates to a method and device for constructing and evaluating a highly reliable glomerular filtration rate prediction model. Background Art

[0002] Direct measurement of GFR (Glomerular Filtration Rate) has poor operability. GFR formula prediction methods based on serum creatinine concentration (SCr) and serum cystatin C are widely used in clinical practice. The GFR calculated according to the formula is called eGFR (Estimated Glomerular Filtration Rate), which has become the main basis for diagnosing stage 3 to 5 CKD. This method is low-cost and less harmful to the patient's body. However, there are the following problems:

[0003] (1) Among the various existing GFR prediction models, each model is established based on a specific data set, its generalization ability is unknown, and the quality of the model is difficult to determine;

[0004] (2) Due to the small amount of data on the GFR gold standard, the reliability of the GFR prediction model constructed based on it is questionable. Obviously, the reliability of the prediction results of GFR values ​​predicted by using a GFR prediction model with low reliability is questionable.

[0005] Therefore, there is an urgent need for a technology that can build a highly reliable GFR prediction model for the GFR prediction problem in small sample conditions and evaluate the reliability of the GFR prediction results given by the model.

[0006] The prediction of GFR usually involves building a regression model of the relationship between the patient's physical indicators and renal metabolite indicators and the GFR value. These indicators include the patient's age, gender, race, and serum creatinine value. Existing models are constructed based on different data sets and are all linear regression models based on the linear model assumption.

[0007] There are currently no studies on whether a model is reliable; however, this question is particularly important, especially for small sample sizes of GFR in kidney disease, where it is more challenging to construct a reliable model.

[0008] The existing methods have the following main shortcomings:

[0009] (1) The reliability of the GFR prediction model is not examined. The evaluation of the GFR prediction model uses evaluation indicators such as prediction accuracy, mean square error, and correlation coefficient (R2) based on existing data, but does not examine the reliability of the model. Unreliable models are unusable.

[0010] (2) The reliability of the GFR prediction results is not examined. The model only gives the prediction results of GFR, but does not examine the reliability of this result, which may lead to the problem that the prediction results may seem good but are not reliable and therefore unreliable. Summary of the invention

[0011] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method and device for constructing and evaluating a highly reliable glomerular filtration rate prediction model. The method proposes a technical solution of constructing multiple local models and constructing a global model through the median model of these models. Using this solution, a highly reliable GFR prediction model can be obtained.

[0012] In order to achieve the above object, the technical solution adopted by the present invention is:

[0013] A method for constructing and evaluating a highly reliable glomerular filtration rate prediction model, comprising:

[0014] Step (1): data preprocessing;

[0015] In the data used for modeling, each sample contains the numerical values ​​of its clinical characteristics and its GFR measurement value, both of which are numerical.

[0016] Data preprocessing is divided into two steps: feature selection and log transformation;

[0017] Step (2): local model setting;

[0018] The local model is set as a linear regression model, as follows:

[0019] y=w T x+b (4)

[0020] The independent variable x is the model input, i.e., the characteristic value used to predict GFR; y is the model output, i.e., the predicted GFR value calculated by the model based on the input; w and b are model parameters. After learning w and b through training, the model is determined;

[0021] Formula (4) gives the general form of the local model. By setting different powers of the independent variables, various forms of local models can be obtained.

[0022] Step (3): local model construction;

[0023] Data sampling and deduplication processing and local model training;

[0024] Step (4): Global model construction;

[0025] The constructed GFR prediction model f(x), i.e. the global model, is the median model of all local models. The median of the N local model parameters is taken as the parameter of the global model, i.e.:

[0026]

[0027] Step (5): Model stability calculation of the global model;

[0028] Model stability is used to characterize the reliability of a model. For the obtained global model f(x), its model stability is calculated:

[0029]

[0030] Here f i (x) is the local model;

[0031] Step (6): Calculation of the accuracy of the global model GFR prediction;

[0032] ①Data sampling and deduplication processing;

[0033] The basic data set is subjected to sampling with replacement to remove duplicates. This process is repeated M times to obtain M sampling data sets, denoted by D k It is a sampled data set obtained by performing the k-th sampling and deduplication processing on the basic data set, where the number of sampling times M can be as large as possible.

[0034] ②GFR prediction of resampled data;

[0035] For the dataset D k Any sample x in (k=1, 2, ..., M) i , use the global model f(x) to predict the value of GFR, that is, calculate f(x i );

[0036] ③Accuracy of GFR prediction results;

[0037] For each sample data set D k , calculate the predicted GFR value f(x i ) and the true value y i The MSE between is calculated as:

[0038]

[0039] Here, n is the dataset D k The number of samples in y i For the dataset D k The label value of the i-th sample in (i.e., the true value of GFR), f(x i ) is the predicted value of the i-th sample given by the model;

[0040] The average value of MSE of M samples is used to characterize the accuracy of the global model in predicting GFR. The smaller the value, the more accurate the global model is in predicting GFR.

[0041] Step (7): stability calculation of GFR prediction results;

[0042] The stability of the global model for GFR prediction results is calculated according to the following formula, where M is the number of sampling sets, D k (k=1, 2, ..., M) is the kth sampling set, |D k | is the sampling set D k The number of samples in x i represents the i-th sample in the sampling set, y i is the true value of the GFR label of the i-th sample, f(x i ) is the GFR prediction value of the i-th sample given by the model:

[0043]

[0044] Step (8): Model selection;

[0045] For the various forms of local models set in step (2), the model stability of the corresponding global model, the accuracy of the GFR prediction results, and the stability of the prediction results are comprehensively evaluated, and the global model corresponding to the model form with the highest three values ​​and its parameters is selected as the GFR prediction model; when the three cannot be guaranteed to be the highest at the same time, the global model with relatively high values ​​of all three values ​​is selected after comprehensive evaluation.

[0046] The original data set in step (1) is collected from the clinic, including GFR measurement values ​​of 22 chronic kidney disease patients and corresponding 91 features, and the data type is numerical.

[0047] The feature selection screens out the top two features of these candidate features: cystatin C and age. The present invention does not limit the specific method of feature selection and the number of selected features. If the present invention is applied to other data sets for modeling, the feature selection method can be set according to actual conditions.

[0048] The log transformation takes the logarithm with base e for the GFR measurement value and the selected characteristic independent variables, cystatin C and age, and uses the data set after feature selection and log transformation as the basic data set for subsequent modeling.

[0049] In step (2), for the GFR prediction problem, the value y of the dependent variable GFR and the feature independent variable x=(x 1 , x 2 ) T——The value of age x 1 and the value of cystatin C x 2 Take the logarithm with base e, and we get x′ 1 =log e x 1 , x′ 2 =log e x 2 , y′=log e y, and use the data to build four linear regression models:

[0050] ①y′=a 1 x′ 1 +a 2 x′ 2 +a 3 (Type 0+1)

[0051] ②y′=c 1 (x′ 1 ) 2 +c 2 (x′ 2 ) 2 +c 3 x′ 1 x′ 2 +c 4 (Type 0+2)

[0052] ③y′=a 1 (x′ 1 ) 2 +a 2 (x′ 2 ) 2 +a 3 x′ 1 x′ 2 +a 4 x′ 1 +a 5 x′ 2 +b(0+1+2 type)

[0053] ④y′=a 1 (x′ 1 ) 3 +a 2 (x′ 2 ) 2 +b(2+3 type)

[0054] Since formula ① contains only constant terms (zero-order terms) and first-order terms, it is named 0+1 type; formula ② contains only constant terms and second-order terms, it is named 0+2 type;

[0055] Formula ③ contains a constant term, a linear term, and a quadratic term, so it is named the 0+1+2 type; Formula ④ contains a quadratic term and a cubic term, so it is named the 2+3 type.

[0056] The data sampling deduplication processing in step (3) includes sampling with replacement on the basic data set (the number of samples in each sampling data set is the same as that in the basic data set), and only one duplicate sample is retained. This operation is called data sampling deduplication processing, and the data set after such sampling deduplication processing is called a sampling data set. The basic data set is subjected to N independent sampling deduplication processing to obtain N sampling data sets. N is as large as possible. In the following modeling example, N is 1000.

[0057] The local model training trains the set model on the N sampled data sets after sampling to obtain its parameters: let the model trained with the i-th data set be f i (x) = w i T x+b i , get N models It is called a local model.

[0058] A highly reliable glomerular filtration rate prediction device, comprising:

[0059] A collection unit, the collection unit comprising a cystatin C collection device and an age collection device;

[0060] A processing unit, running the high-reliability glomerular filtration rate prediction model, receiving data from the cystatin C collection device and the age collection device in a wired or wireless manner, and estimating a high-reliability glomerular filtration rate;

[0061] a display unit, receiving and displaying the estimated high-reliability glomerular filtration rate by wired or wireless means;

[0062] The hardware form of the cystatin C collection device is a smart phone or a computer, and the data transmission method with the processing unit is a Socket method;

[0063] The hardware form of the age collection device is a smart phone or a computer, and the data transmission method with the processing unit is Socket;

[0064] The hardware form of the processing unit is a computer.

[0065] Beneficial effects of the present invention:

[0066] (1) The technology involved in the present invention can be used for but not limited to the prediction of GFR in kidney disease;

[0067] (2) The model stability is proposed to characterize the reliability of the model; the stability measure of the prediction results is proposed to characterize the reliability of the prediction results given by the model;

[0068] (3) A highly reliable model is constructed. The model established in the present invention, namely the median model obtained by integrating local models, has the characteristics of high reliability of the model and high robustness to noise in the data;

[0069] (4) It was applied to the prediction of GFR in chronic kidney disease, and a more reliable prediction model and more reliable prediction results of GFR were developed;

[0070] (5) Different from the previously submitted patent application (CN202210581688.1), the present invention examines the reliability of the model and the reliability of the model's GFR prediction results. Only reliable ones are trustworthy and usable. In this sense, the present invention is an extension of the previous patent application in terms of technical solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 Schematic diagram of the technical route of the present invention.

[0072] Figure 2 Schematic diagram comparing the model stability of the median model obtained for each model form.

[0073] Figure 3 Schematic diagram of the accuracy of the prediction results of each model.

[0074] Figure 4 Schematic diagram of the stability of the prediction results of each model. DETAILED DESCRIPTION

[0075] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0076] Example:

[0077] The following uses kidney disease data to compare various GFR prediction models, which are shown in Table 1. The first 8 are models reported in the literature, and the last 4 are median models obtained under the 4 model forms established by the present application technology. The purpose of model selection is achieved by comparing the model stability, accuracy of GFR prediction results and stability of prediction results of these models.

[0078] Table 1 All models used in the experiment

[0079]

[0080]

[0081] (1) Comparison of model stability

[0082] Among the 12 models in Table 1, the model with number 3 needs to be removed first, because the model has different formulas for different genders, and our data set has only 22 samples, which is too small. If the gender division is performed again, the amount of data for each gender will be even smaller, which will make the numerical value of model stability very unreliable. Therefore, the model with number 3 is removed from the comparison here. For the remaining models, the models with numbers 1, 2, 4, and 8 have the same model form, the models with numbers 5 and 6 have the same model form, and the model with number 7 is a separate model form. In addition to these three model forms, model numbers 8 to 12 are each a model form. Therefore, there are a total of 7 model forms in the 12 formulas in Table 1, which are shown in Table 2.

[0083] Table 2 Classification of model forms of GFR prediction models

[0084]

[0085]

[0086] For each of the seven model forms in Table 2, the following operations are performed to compare their stability:

[0087] ① Perform sampling with replacement to remove duplicates on the basic data set of kidney disease, perform linear regression modeling on the sampled samples, and obtain a local model;

[0088] ② Repeat step ① 1000 times, and use the 1000 local models obtained by 1000 samplings to build a global model;

[0089] ③Calculate the model stability of this global model.

[0090] From this, the model stability of each of the seven model forms can be obtained, as shown in Figure 2 (Numbers 1 to 7 correspond to the models represented by numbers 1 to 7 in Table 2). Figure 2 It can be seen that the model corresponding to No. 4 in Table 2 (i.e., 2+3 type) has the highest model stability.

[0091] In fact, the models numbered 9, 10, 11, and 12 in Table 1, wherein the parameters are global models in the corresponding model form obtained by the technology of the present application.

[0092] (2) Comparison of the accuracy of GFR prediction results

[0093] The accuracy of the GFR prediction results refers to the difference between the GFR prediction results given by the model and the actual GFR value given in the data, usually expressed as the mean square error (MSE) between them. The following is a comparison of the accuracy of the 12 models in Table 1. The specific steps are as follows: For each of the 12 models,

[0094] ① Perform sampling and deduplication processing on the basic data set of kidney disease, use the 12 models in Table 1 to predict GFR for the sampled data set, and record the predicted values;

[0095] ② Compare the predicted value with the true value of the GFR label in the sampling data set, and use the MSE calculation formula to calculate the MSE of this sampling;

[0096] ③ Repeat ①② 1000 times and take the average of 1000 MSEs as the MSE of the final global model prediction result.

[0097] Here, the smaller the average MSE value is, the smaller the gap between the predicted value and the true value is, and the better the accuracy of the prediction result is.

[0098] The results of the above approach are shown in Figure 3 It can be seen that the prediction results of model 12 (i.e., type 2+3) in Table 1 have the highest accuracy, followed by model 10 (i.e., type 0+2) in Table 1.

[0099] (3) Comparison of the stability of GFR prediction results

[0100] For each of the 12 GFR prediction models in Table 1, the calculation process of its stability is as follows:

[0101] ① Perform sampling and deduplication processing on the basic data of kidney disease, and use the GFR prediction model to predict the GFR of the sampled samples;

[0102] ② Repeat ①, total repetition M = 1000 times;

[0103] ③ Using formula (8), calculate the stability of the prediction results of the GFR prediction model.

[0104] Figure 4 The stability of the results of GFR prediction by each model is given. Figure 4 Among them, the model No. 8 in Table 1 has the highest stability in its GFR prediction results, followed by the model No. 10 in Table 1, and then the model No. 12 and the model No. 11 in Table 1.

[0105] Figure 4 Among them, the model No. 8 in Table 1 has the highest stability in its GFR prediction results, followed by the model No. 10 in Table 1, and then the model No. 12 and the model No. 11 in Table 1.

[0106] (4) Model selection

[0107] The results of the above experiments are summarized in Table 3.

[0108] Table 3 Performance comparison of each model

[0109]

[0110] As can be seen from Table 3, Model 12 (2+3 type) in Table 1 has the strongest model stability, the highest accuracy in GFR prediction, and this accuracy is relatively stable. Other models either have poor model stability, poor accuracy in GFR prediction, poor stability in GFR prediction results, or both (for example, Model 8 in Table 1, although its GFR prediction stability is the best, its model stability and GFR prediction result accuracy are not good, that is, its GFR prediction results are stable at an unsatisfactory poor accuracy). Therefore, Model 12 in Table 1 was finally selected as the final kidney disease GFR prediction model.

Claims

1. A method for constructing and evaluating a highly reliable glomerular filtration rate prediction model, characterized in that: include; Step (1): data preprocessing; In the data used for modeling, each sample contains the numerical values ​​of its clinical characteristics and its GFR measurement value, both of which are numerical. Data preprocessing is divided into two steps: feature selection and log transformation; Step (2): local model setting; The local model is set as a linear regression model, as follows: y=w T x+b (4) The independent variable x is the model input, i.e., the characteristic value used to predict GFR; y is the model output, i.e., the predicted GFR value calculated by the model based on the input; w and b are model parameters. After learning w and b through training, the model is determined; Step (3): local model construction; Data sampling, deduplication processing and local model training; Step (4): global model construction; The constructed GFR prediction model f(x); Step (5): Model stability calculation of the global model; Model stability is used to characterize the reliability of a model. For the obtained global model f(x), its model stability is calculated: Step (6): Calculation of the accuracy of the global model GFR prediction; Including data sampling deduplication processing, GFR prediction of resampled data and the accuracy of GFR prediction results; Step (7): stability calculation of GFR prediction results; The stability of the global model for GFR prediction results is calculated according to the following formula, where M is the number of sampling sets, D k (k=1,2,...,M) is the kth sampling set, |D k | is the sampling set D k The number of samples in x i represents the i-th sample in the sampling set, y i is the true value of the GFR label of the i-th sample, f(x i ) is the GFR prediction value of the i-th sample given by the model: Step (8): Model selection; For the various forms of local models set in step (2), the model stability of the corresponding global model, the accuracy of the GFR prediction results, and the stability of the prediction results are comprehensively evaluated, and the global model corresponding to the model form with the highest three values ​​and its parameters is selected as the GFR prediction model; when the three cannot be guaranteed to be the highest at the same time, the global model with relatively high values ​​of all three values ​​is selected after comprehensive evaluation.

2. The method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: The original data set in step (1) is collected from the clinic, including GFR measurement values ​​of 22 patients with chronic kidney disease and corresponding 91 features, and the data type is numerical; The feature selection screens out the top two features of these candidate features: cystatin C and age; The log transformation takes the logarithm with base e for the GFR measurement value and the selected characteristic independent variables, cystatin C and age, and uses the data set after feature selection and log transformation as the basic data set for subsequent modeling.

3. The method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: In step (2), for the GFR prediction problem, the value y of the dependent variable GFR and the feature independent variable x = (x1, x2) T ——The value of age x1 and the value of cystatin C x2 are taken with the logarithm of base e, and we get x′1=log e x1,x′2=log e x2,y′=loge y, and use this data to build four linear regression models: ①y′=a1x′1+a2x′2+a3(0+1 type) ②y′=c1(x′1) 2 +c2(x′2) 2 +c3x′1x′2+c4(0+2 type) ③y′=a1(x′1) 2 +a2(x′2) 2 +a3x′1x′2+a4x′1+a5x′2+b(0+1+2 type) ④y′=a1(x′1) 3 +a2(x′2) 2 +b(2+3 type) Since formula ① contains only constant terms (zero-order terms) and first-order terms, it is named 0+1 type; formula ② contains only constant terms and second-order terms, it is named 0+2 type; Formula ③ contains a constant term, a linear term, and a quadratic term, so it is named the 0+1+2 type; Formula ④ contains a quadratic term and a cubic term, so it is named the 2+3 type.

4. The method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: The data sampling and duplication removal processing in step (3) includes sampling with replacement on the basic data set, retaining only one repeated sample. This operation is called data sampling and duplication removal processing. The data set after such sampling and duplication removal processing is called a sampling data set. The basic data set is subjected to N independent sampling and duplication removal processing to obtain N sampling data sets, where N is as large as possible. In the following modeling example, N is 1000.

5. The method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: In the step (3), the local model training is to train the set model on the N sampled data sets after sampling to obtain its parameters: let the model trained with the i-th data set be f i (x) = w i T x+b i , get N models It is called a local model.

6. The method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: The GFR prediction model f(x) in step (4) is the median model of all local models, that is, the median of the N local model parameters is taken as the parameter of the global model, that is:

7. The method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: The model stability calculation formula of the global model in step (5) is: Here f i (x) is the local model.

8. The method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: The step (6) of data sampling deduplication processing includes performing sampling deduplication processing with replacement on the basic data set; this process is repeated M times to obtain M sampling data sets, denoted by D k is the sampled data set obtained by performing the k-th sampling and deduplication processing on the basic data set, where the number of sampling times M can be as large as possible; GFR prediction of resampled data For dataset D k Any sample x in (k=1,2,...,M) i , use the global model f(x) to predict the value of GFR, that is, calculate f(x i ); The accuracy of GFR prediction results is for each sample data set D k , calculate the predicted GFR value f(x i ) and the true value y i The MSE between is calculated as: Here, n is the dataset D k The number of samples in y i For the dataset D k The GFR label of the i-th sample (i.e., the true value of GFR), f(x i ) is the predicted value of the i-th sample given by the model; The average value of the MSE of M samples is used to characterize the accuracy of the global model in predicting GFR. The smaller the value, the more accurate the global model is in predicting GFR.

9. A device for executing the method for constructing and evaluating a highly reliable glomerular filtration rate prediction model according to claim 1, characterized in that: include: A collection unit, the collection unit comprising a cystatin C collection device and an age collection device; A processing unit, running the high-reliability glomerular filtration rate prediction model, receiving data from the cystatin C collection device and the age collection device in a wired or wireless manner, and estimating a high-reliability glomerular filtration rate; a display unit, receiving and displaying the estimated high-reliability glomerular filtration rate by wired or wireless means; The hardware form of the cystatin C collection device is a smart phone or a computer, and the data transmission method with the processing unit is a Socket method; The hardware form of the age collection device is a smart phone or a computer, and the data transmission method with the processing unit is Socket; The hardware form of the processing unit is a computer.

Citation Information

Patent Citations

  • A method and system for constructing a glomerular filtration rate estimation model

    CN115081190B

  • Detecting system for glomerular filtration rate

    CN105277723A

  • Method for establishing model for obtaining glomerular filtration rate and application

    CN109545377A