Method for predicting the occurrence of postoperative acute kidney injury and system thereof

KR102999118B1Active Publication Date: 2026-08-03THE CATHOLIC UNIV OF KOREA IND ACADEMIC COOP FOUND
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
THE CATHOLIC UNIV OF KOREA IND ACADEMIC COOP FOUND
Filing Date
2022-08-16
Publication Date
2026-08-03

Smart Images

  • Figure 112022085396559-PAT00004_ABST
    Figure 112022085396559-PAT00004_ABST
Patent Text Reader

Abstract

A method and system for predicting the occurrence of acute kidney injury are provided. A method for predicting the occurrence of acute kidney injury according to some embodiments of the present disclosure trains a model that predicts the risk of acute kidney injury after surgery using a dataset of multiple patients, and can accurately predict the risk of acute kidney injury after surgery in a specific patient at an early stage using the trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to a method and system for predicting the occurrence of acute kidney injury after surgery, and more specifically, to a method for predicting the risk of occurrence of acute kidney injury after surgery using machine learning / deep learning technology and a system for performing the method. Background Technology

[0002] Acute kidney injury (AKI) is statistically known to occur in about 7% of all hospitalized patients, and up to 20% of patients receiving treatment in the Intensive Care Unit (ICU). Furthermore, it is known that AKI occurs in as many as 40% of patients who have undergone surgery.

[0003] In particular, acute kidney injury occurring after surgery drastically reduces the patient's survival rate, and if a patient fails to recover from acute kidney injury and undergoes renal replacement therapy, even if they survive, their quality of life is bound to be significantly reduced.

[0004] Therefore, research on methods to predict the risk of acute kidney injury (API) after surgery at an early stage is continuously being conducted in the medical field. However, since postoperative ATI is caused by complex factors, early prediction is considerably difficult, and research results to date remain minimal. Prior art literature

[0005] Korean Patent Publication No. 10-2022-0075046 (Published June 7, 2022) The problem to be solved

[0006] The technical problem to be solved through some embodiments of the present disclosure is to provide a method for accurately predicting the risk of acute kidney injury occurring after surgery at an early stage, and a system for performing the method.

[0007] Another technical problem to be solved through some embodiments of the present disclosure is to provide a method for constructing a model that predicts the risk of acute kidney injury occurring after surgery and a system for performing the method.

[0008] Another technical problem to be solved through some embodiments of the present disclosure is to provide a method for generating a high-quality dataset for a model that predicts the risk of acute kidney injury after surgery, and a system for performing the method.

[0009] Another technical problem to be solved through some embodiments of the present disclosure is to provide important variables (features) that can ensure the performance of a model for predicting the risk of acute kidney injury after surgery.

[0010] The technical problems of the present disclosure are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by a person skilled in the art of the present disclosure from the description below. means of solving the problem

[0011] A method for predicting the occurrence of acute kidney injury after surgery according to some embodiments of the present disclosure, for solving the above technical problem, is a method performed by at least one computing device and may include the steps of preparing a dataset for a plurality of patients—wherein the dependent variable of the dataset relates to the occurrence of acute kidney injury after surgery, and the independent variables of the dataset include variables related to preoperative examination items of said patients—and constructing a model that predicts the risk of acute kidney injury after surgery using said prepared dataset.

[0012] In some embodiments, the preoperative test items may include albumin, creatinine (Cr), potassium, protein, and urine specific gravity.

[0013] In some embodiments, the independent variables of the dataset further include variables regarding the patients' disease history and medication history, and the disease history may include a history regarding chronic kidney disease (CKD), hypertension (HTN), cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), and liver cirrhosis (LC). In this case, the medication history may be regarding antihypertensive drugs.

[0014] In some embodiments, the independent variables of the dataset may further include variables regarding the type and duration of the surgery received by the patients.

[0015] In some embodiments, the model may be based on at least one of a neural network, logistic regression, and LGBM (Light Gradient Boosting Machine).

[0016] In some embodiments, the step of preparing the dataset may include removing data of patients who meet a predetermined kidney-related condition from the original patient dataset.

[0017] In some embodiments, the step of preparing the dataset includes removing data of patients who meet certain surgery-related conditions from the original patient dataset, and the certain surgery-related conditions may be defined based on the time required for surgery or the type of surgery.

[0018] A method for predicting the occurrence of acute kidney injury after surgery according to some other embodiments of the present disclosure, for solving the above technical problem, may include the step of obtaining a model trained to predict the risk of acute kidney injury occurring after surgery—the model being trained using a dataset of a plurality of patients, wherein the dependent variable of the dataset relates to the occurrence of acute kidney injury after surgery, and the independent variables of the dataset include variables related to pre-operative examination items of said patients—and the step of predicting the risk of acute kidney injury occurring in a specific patient after a target surgery using said trained model.

[0019] A system for predicting the occurrence of acute kidney injury after surgery according to some embodiments of the present disclosure, for solving the technical problem described above, comprises one or more processors and a memory for storing one or more instructions, wherein the one or more processors can perform the operation of preparing a dataset for a plurality of patients by executing the one or more stored instructions—wherein the dependent variable of the dataset relates to the occurrence of acute kidney injury after surgery and the independent variable of the dataset includes a variable related to pre-operative examination items of the patients—and the operation of constructing a model for predicting the risk of acute kidney injury after surgery using the prepared dataset.

[0020] A system for predicting the occurrence of acute kidney injury according to several other embodiments of the present disclosure, for solving the technical problem described above, comprises one or more processors and a memory for storing one or more instructions, wherein the one or more processors can perform the operation of obtaining a model trained to predict the risk of acute kidney injury occurring after surgery by executing the one or more stored instructions—the model being trained using a dataset of multiple patients, wherein the dependent variable of the dataset is related to the occurrence of acute kidney injury after surgery, and the independent variable of the dataset includes variables related to pre-operative examination items of said patients—and the operation of predicting the risk of acute kidney injury occurring in a patient after a target surgery using said trained model.

[0021] A computer program according to some embodiments of the present disclosure for solving the technical problem described above may be combined with a computing device and stored on a computer-readable recording medium to execute the steps of preparing a dataset for a plurality of patients—wherein the dependent variable of the dataset relates to the occurrence of acute kidney injury after surgery, and the independent variables of the dataset include variables related to preoperative examination items of said patients—and constructing a model that predicts the risk of acute kidney injury occurring after surgery using said prepared dataset.

[0022] A computer program according to some other embodiments of the present disclosure for solving the technical problem described above may be combined with a computing device and stored on a computer-readable recording medium to execute the steps of: obtaining a model trained to predict the risk of acute kidney injury occurring after surgery—said that the model is trained using a dataset of multiple patients, said dataset having dependent variables related to the occurrence of acute kidney injury after surgery, and said dataset having independent variables related to pre-operative examination items of said patients—and predicting the risk of acute kidney injury occurring in a specific patient after a target surgery using said trained model. Effects of the invention

[0023] According to some embodiments of the present disclosure, the risk of acute kidney injury after surgery can be accurately predicted early through a machine learning / deep learning model (hereinafter referred to as the 'prediction model'). For example, the risk of acute kidney injury occurring in a patient after surgery can be accurately predicted early, and the incidence of acute kidney injury can be reduced through preemptive measures based on the prediction results.

[0024] Furthermore, by removing unnecessary patient data from the original patient dataset, a high-quality training dataset for the prediction model can be generated. Consequently, a high-performance prediction model can be easily constructed.

[0025] In addition, the independent variables of the patient dataset may consist of variables related to the patient's disease history, medication history, surgery-related information, and pre- and post-operative examination items. Accordingly, the prediction model can be trained to predict the risk of acute kidney injury by considering various factors in combination, and the performance of the prediction model can be further improved.

[0026] The effects according to the technical concept of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0027] FIG. 1 is an exemplary drawing for schematically illustrating a system for predicting the occurrence of acute kidney injury after surgery and its inputs and outputs according to some embodiments of the present disclosure. FIGS. 2 and 3 are exemplary drawings for explaining the process of a system for predicting the occurrence of acute kidney injury after surgery according to some embodiments of the present disclosure providing a prediction service. FIGS. 4 and 5 are exemplary flowcharts illustrating a method for predicting the occurrence of acute kidney injury after surgery according to some embodiments of the present disclosure. FIG. 6 is an exemplary drawing illustrating an artificial neural network-based prediction model according to some embodiments of the present disclosure. FIG. 7 is an exemplary drawing for illustrating a method for selecting major independent variables according to some embodiments of the present disclosure. FIG. 8 is an exemplary drawing for illustrating a method for selecting major independent variables according to several other embodiments of the present disclosure. FIGS. 9 and FIGS. 10 are exemplary drawings for illustrating a data augmentation method according to some embodiments of the present disclosure. Figure 11 is a diagram illustrating the process of preparing a patient dataset for performance evaluation of a prediction model. Figures 12 to 14 illustrate the performance evaluation results for various types of prediction models. FIG. 15 illustrates an exemplary computing device capable of implementing a system for predicting the occurrence of acute kidney injury after surgery according to some embodiments of the present disclosure. Specific details for implementing the invention

[0028] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the attached drawings. The advantages and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the attached drawings. However, the technical concept of the present disclosure is not limited to the following embodiments but can be implemented in various different forms. The following embodiments are provided merely to complete the technical concept of the present disclosure and to fully inform those skilled in the art of the scope of the present disclosure, and the technical concept of the present disclosure is defined only by the scope of the claims.

[0029] It should be noted that when assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the present disclosure, if it is determined that a detailed description of related known components or functions could obscure the essence of the present disclosure, such detailed description is omitted.

[0030] Unless otherwise defined, all terms used herein (including technical and scientific terms) may be used in a meaning commonly understood by those skilled in the art to which this disclosure pertains. Additionally, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. The terms used herein are for describing the embodiments and are not intended to limit this disclosure. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text.

[0031] Additionally, terms such as first, second, A, B, (a), (b), etc., may be used to describe the components of the present disclosure. These terms are intended merely to distinguish the components from other components, and the nature, order, or sequence of the components is not limited by such terms. Where it is stated that a component is "connected," "combined," or "connected" to another component, it should be understood that the component may be directly connected or connected to the other component, but that another component may also be "connected," "combined," or "connected" between each component.

[0032] As used in this disclosure, "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components, steps, actions, and / or elements to the mentioned components, steps, actions, and / or elements.

[0033] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0034] FIG. 1 is an exemplary drawing for explaining a system (10) for predicting the occurrence of acute kidney injury after surgery and its inputs and outputs according to some embodiments of the present disclosure. In FIG. 1 and below, the system for predicting the occurrence of acute kidney injury after surgery (10) is depicted as a ‘prediction system (10)’, and in the following description, the system for predicting the occurrence of acute kidney injury after surgery (10) will also be abbreviated as a ‘prediction system (10)’.

[0035] As illustrated in FIG. 1, the prediction system (10) may be a computing device / system that predicts and outputs the risk of acute kidney injury (AKI) occurring after surgery in a patient based on input patient data. For example, the prediction system (10) can predict the risk of acute kidney injury (e.g., risk of AKI occurring within about 30 days after surgery) after surgery in a patient who has undergone (or is scheduled to undergo) surgery through a learned prediction model (11).

[0036] More specifically, the prediction system (10) trains a prediction model (11) using a dataset of multiple patients and can predict the risk of acute kidney injury occurring after surgery in a specific patient through the trained prediction model (11). The prediction results may include, for example, whether acute kidney injury occurs after surgery, the risk of acute kidney injury (i.e., the probability of acute kidney injury occurring), the progression / risk stage of acute kidney injury (e.g., acute kidney injury stage according to KDIGO criteria), the risk level at each stage, but are not limited thereto. The specific method by which the prediction system (10) trains the prediction model (11) and predicts the risk of acute kidney injury occurring after surgery will be explained in detail with reference to the drawings from Fig. 4 onwards.

[0037] The patient dataset (or data) used for training (or prediction) the prediction model (11) may consist of at least one dependent variable and a number of independent variables, which will be described later. For reference, in the field of the art, the term "variable" may be used interchangeably with terms such as "feature," "attribute," "element," "item," and "field." Additionally, each individual data constituting the patient dataset may be used interchangeably with terms such as "sample," "example," "record," "instance," "entry," "data point," and "observation."

[0038] In some embodiments, as illustrated in FIG. 2, the prediction system (10) may provide a prediction service regarding the occurrence of acute kidney injury after surgery to a user. For example, the prediction system (10) may receive patient data from a user terminal (20), predict the risk of acute kidney injury after surgery based on the received patient data, and provide the prediction result to the user terminal (20). The user may be a patient or medical staff, but the scope of the present disclosure is not limited thereto. As a more specific example, the prediction system (10) may provide the prediction service through a web interface (or app interface). For example, the prediction system (10) may provide a web page (30) as illustrated in FIG. 3 to the user terminal (20), receive patient data through the web page (30), and provide the predicted result based on the input patient data through the web page (30).

[0039] The prediction system (10) may be implemented with at least one computing device. For example, the prediction system (10) may be implemented with one computing device. As another example, the prediction system (10) may be implemented with multiple computing devices, wherein the first function of the prediction system (10) may be implemented in the first computing device and the second function may be implemented in the second computing device. Alternatively, a specific function of the prediction system (10) may be implemented in multiple computing devices.

[0040] A computing device may encompass any device equipped with computing (processing) functions, and for an example of such a device, refer to FIG. 15. Since a computing device is a collection of multiple components (e.g., processor, memory, etc.) interacting with each other, it may be referred to as a 'computing system' depending on the case. Additionally, a computing system may refer to a collection of multiple computing devices interacting for the same purpose.

[0042] Up to now, a prediction system (10) according to some embodiments of the present disclosure has been schematically described with reference to FIGS. 1 to 3. Below, various methods that can be performed in the prediction system (10) illustrated in FIG. 1 will be described in detail.

[0043] For convenience of understanding, the following description will continue by assuming that all steps (operations) of the methods to be described below are performed in the prediction system (10) exemplified in FIG. 1. Therefore, if the subject of a specific step (operation) is omitted, it can be understood that it is performed by the prediction system (10). Of course, in an actual environment, some steps of the methods to be described below may be performed on a different computing device. For example, the training of the prediction model (e.g., 11 in FIG. 1) may be performed on a different computing device depending on the case.

[0044] FIG. 4 is an exemplary flowchart schematically illustrating a method for predicting the occurrence of acute kidney injury after surgery according to some embodiments of the present disclosure. However, this is merely a preferred embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as necessary.

[0045] As illustrated in FIG. 4, the prediction method according to the embodiments may begin with step S41 of preparing a patient dataset to be used for training a prediction model. As described above, the patient dataset may consist of one or more dependent variables and a plurality of independent variables, and may include a plurality of patient data (i.e., data samples).

[0046] One or more dependent variables (e.g., correct answer labels) may be related to the occurrence of acute kidney injury after surgery (e.g., occurrence of acute kidney injury within 30 days after surgery), such as whether acute kidney injury occurred after surgery, or the stage of progression / risk of acute kidney injury (e.g., acute kidney injury stage according to KDIGO criteria), but are not limited thereto.

[0047] Multiple independent variables may include, for example, variables regarding the patient's demographic characteristics, disease history, medication history, surgeries performed, and tests performed before and after surgery (i.e., items indicating the patient's health status). However, the scope of the present disclosure is not limited thereto. For more detailed examples of independent variables, refer to Table 1 below.

[0048] division Detailed variables Demography Age, gender, BMI, height, weight, blood pressure (SBP, DBP), etc. History of disease Chronic kidney disease (CKD), diabetes mellitus (DM), hypertension (HTN), cardiovascular disease (CVD), coronary artery disease (CAD), chronic obstructive pulmonary disease (COPD), liver cirrhosis (LC), smoking status, duration of smoking, duration of illness, etc. Medication history Antihypertensive drugs (e.g., angiotensin receptor blockers (ARBs), angiotensin-converting enzyme inhibitors (ACEi)), anti-inflammatory drugs (e.g., non-steroidal anti-inflammatory drugs (NSAIDs)), duration of medication, dosage, etc. surgery Surgery department, duration of surgery, whether the surgery is on a weekday or weekend, etc. Inspection items Before surgery White blood cell count (WBC), hemoglobin, CRP (C-reactive protein), glucose, BUN (Blood Urea Nitrogen), creatinine (Cr), eGFR, protein, total protein, albumin, AST, ALT, sodium (Na), potassium (K), chloride (Cl), calcium (Ca), uric acid, CPK, LDH, urine specific gravity (SG), urine protein

[0049] By utilizing the various independent variables mentioned above to train the prediction model, the model can accurately predict the risk of acute kidney injury after surgery by considering various factors in combination. For example, a prediction model trained using the independent variables exemplified in Table 1 can accurately predict the risk of acute kidney injury after surgery (e.g., risk of acute kidney injury occurring within 30 days after surgery) by considering the patient's disease history, information on surgeries received (or scheduled to be received), and the patient's pre- and post-operative condition in combination.

[0050] The detailed process of this step S41 is illustrated in Fig. 5.

[0051] As illustrated in FIG. 5, the patient dataset preparation step (S41) may include a step of cleaning the original patient dataset (S51) and a step of removing some patient data (i.e., data samples) (S52). FIG. 5 illustrates, as an example, that step S52 is performed after step S51, but the order of execution of step S51 and step S52 can be changed at any time. Below, each step will be described in detail.

[0052] In step S51, the original patient dataset can be refined in various ways.

[0053] As an example, the prediction system (10) can correct outliers in the original patient dataset. For instance, the prediction system (10) can determine the values ​​of the top n% (e.g., 1%, 5%, etc.) and / or bottom k% (e.g., 1%, 5%, etc.) for each variable as outliers and remove patient data containing outliers.

[0054] As another example, the prediction system (10) can correct missing values ​​in the original patient dataset. For instance, the prediction system (10) can correct (i.e. fill) missing values ​​in the original patient dataset using Multiple Imputation by Chained Equations (MICE). Since those skilled in the art are likely already familiar with the MICE technique, a description thereof will be omitted.

[0055] As another example, the prediction system (10) can convert the values ​​of non-numerical variables of the original patient dataset into numerical values. For instance, the prediction system (10) can convert the values ​​of non-numerical variables into numerical values ​​using a one-hot encoding technique.

[0056] As another example, the prediction system (10) can normalize the original patient dataset (or values ​​of numeric variables). For instance, the prediction system (10) can normalize the original patient dataset (or values ​​of numeric variables) using a min-max normalization technique.

[0057] As another example, the prediction system (10) can refine the original patient dataset based on various combinations of the examples described above.

[0058] In step S52, patient data meeting certain conditions may be removed from the original patient dataset. However, the specific method may vary depending on the embodiment.

[0059] In some embodiments, data of patients meeting certain renal-related conditions may be removed. Here, certain renal-related conditions may be conditions defined based, for example, a history of renal replacement therapy, preoperative eGFR levels, preoperative creatinine (Cr) levels, or the degree to which creatinine (Cr) levels have increased within a certain period prior to surgery. However, the scope of the present disclosure is not limited by these examples. As a more specific example, the prediction system (10) may remove data from the original patient dataset of patients who have a history of renal replacement therapy, patients whose preoperative eGFR levels are below a reference value (e.g., about 15 ml / min), patients whose preoperative creatinine (Cr) levels (concentrations) are above a reference value (e.g., about 4.0 mg / dL), or patients whose preoperative creatinine (Cr) levels have increased above a reference value (e.g., about 1.5 times the previous value or 0.3 mg / dL) within a certain period (e.g., about 2 weeks). The reason for excluding data from these patients is that the exemplified patients can be seen as having severe kidney problems with chronic kidney disease stage 5 or having recently suffered acute kidney injury. In other words, it can be understood that the data from the exemplified patients is excluded because it is important to accurately predict acute kidney injury that occurs suddenly in general patients.

[0060] In some other embodiments, patient data that meets certain surgery-related conditions may be removed. Here, certain surgery-related conditions may be conditions defined, for example, based on the duration of the surgery or the type of surgery. However, the scope of the present disclosure is not limited by these examples. As a more specific example, the prediction system (10) may remove data of patients whose surgery duration is less than a threshold (e.g., about 1 hour). This is because the association between a relatively simple surgery performed in a short time and acute kidney injury is typically very low. As another example, the prediction system (10) may remove data of patients whose surgery type corresponds to heart surgery, nephrectomy, or kidney transplantation. This is also understood to be because it is important to accurately predict acute kidney injury that occurs suddenly in general patients.

[0061] In some other embodiments, some patient data may be removed from the original patient dataset based on various combinations of the embodiments described above.

[0062] Meanwhile, in some embodiments, a major independent variable may be selected from among the numerous independent variables constituting the original patient dataset (or patient dataset). Then, the patient dataset for the major independent variable may be used as a training dataset for the prediction model. By doing so, the performance of the prediction model may be further improved, and this embodiment will be described in more detail later with reference to FIGS. 7 and 8.

[0063] In addition, in some embodiments, a process of augmenting the dataset of the acute kidney injury class (i.e., the group of patients who developed acute kidney injury after surgery) may be performed to alleviate the class imbalance problem of the original patient dataset (or patient dataset). The present embodiment will be described in more detail later with reference to FIGS. 9 and 10.

[0064] Referring again to Fig. 4, the explanation will be provided.

[0065] In step S42, a model for predicting the risk of acute kidney injury after surgery can be built (trained) using a prepared patient dataset. For example, the prediction system (10) can input each patient data (i.e., values ​​of independent variables) into the prediction model to obtain a prediction result, and train the prediction model in a direction that minimizes the difference (i.e., prediction error) between the prediction result and the correct answer (i.e., value of dependent variable).

[0066] Predictive models can be designed and implemented based on various types of models. For example, predictive models can be designed and implemented based on deep learning / machine learning models such as artificial neural networks (see 60 in FIG. 6), logistic regression, LGBM (Light Gradient Boosting Machine), naive bayes, support vector machines, decision trees, and random forests. However, the scope of the present disclosure is not limited by these examples, and predictive models may also be implemented based on other types of models (e.g., deep learning models such as convolutional neural networks, recurrent neural networks, Transformers, etc.). Additionally, predictive models may be designed in the form of classification models (e.g., decision trees, naive bayes, etc.) or in the form of regression models.

[0067] In some embodiments, multiple prediction models may be constructed. For example, the prediction system (10) may construct a first prediction model (e.g., an artificial neural network-based prediction model) using a prepared patient dataset and further construct a second prediction model of a different type from the first prediction model (e.g., a logistic regression-based prediction model). Alternatively, the prediction system (10) may construct a first prediction model using a patient dataset for first independent variables and construct a second prediction model using a patient dataset for second independent variables that are at least partially different from the first independent variables. In this case, the prediction system (10) can predict the risk of acute kidney injury after surgery by comprehensively considering the prediction results of the two prediction models.

[0068] In step S43, the risk of acute kidney injury occurring after surgery in a specific patient can be predicted through a learned prediction model. For example, the prediction system (10) can predict the risk of acute kidney injury occurring after surgery in a specific patient who has undergone (or is scheduled to undergo) the target surgery through the learned prediction model. Specifically, the prediction system (10) can construct input data for the prediction model based on the patient's data (e.g., type of target surgery, duration (or estimated duration), pre-operative test results, disease history, medication history, etc.) and perform the prediction by inputting the input data into the prediction model. As described above, the prediction results may include, for example, whether acute kidney injury occurs after surgery, the risk of acute kidney injury (e.g., confidence score for the acute kidney injury class), the stage of progression / risk of acute kidney injury (e.g., acute kidney injury stage according to KDIGO criteria), the risk level for each stage, but are not limited thereto.

[0069] Up to this point, a method for predicting the occurrence of acute kidney injury after surgery according to several embodiments of the present disclosure has been described with reference to FIGS. 4 to 6. As described above, the risk of acute kidney injury after surgery can be accurately predicted early through a machine learning / deep learning model. For example, the risk of acute kidney injury occurring in a patient after surgery can be accurately predicted early, and the incidence of acute kidney injury can be reduced through preemptive measures based on the prediction results.

[0070] Hereinafter, embodiments regarding a method for selecting major independent variables will be described with reference to FIGS. 7 and FIGS. 8.

[0071] First, with reference to FIG. 7, a method for selecting major independent variables according to some embodiments of the present disclosure will be described.

[0072] As illustrated in FIG. 7, the embodiments relate to a method for selecting key independent variables (73-1 to 73-k) based on the performance evaluation results for a model (72).

[0073] Specifically, the prediction system (10) can train a model (72) using independent variables (71-1 to 71-n) that constitute a patient dataset and evaluate the performance of the trained model (72). For example, the prediction system (10) can train a first model using a first independent variable (e.g., 71-1) and train a second model using a second independent variable (e.g., 71-2). Then, the prediction system (10) can evaluate the performance of each of the first model and the second model. Of course, the prediction system (10) can train a model (72) using two or more independent variables (e.g., 71-1, 71-2).

[0074] The model (72) illustrated in FIG. 7 can be understood as an abstraction of all models used for selecting major independent variables. The model (72) is a learnable model (i.e., a machine learning / deep learning model) and may be of the same type as the prediction model described above, or may be of a different type.

[0075] Next, the prediction system (10) can select major independent variables (73-1 to 73-k) based on the performance evaluation results of the model (72). For example, the prediction system (10) can select K independent variables (73-1 to 73-k) used in training a model whose performance evaluation score (e.g., accuracy) is above a threshold value (where K is a value smaller than the total number of independent variables N) as major independent variables.

[0076] When the major independent variables (73-1 to 73-k) are selected, the prediction system (10) can build a prediction model using a patient dataset composed of the major independent variables (73-1 to 73-k). By doing so, a higher performance prediction model can be built.

[0077] Hereinafter, a method for selecting major independent variables according to several other embodiments of the present disclosure will be described with reference to FIG. 8.

[0078] As illustrated in FIG. 8, the embodiments relate to a method for selecting key independent variables by using the degree of influence that a change in the value of a specific independent variable (e.g., variable 2) has on the prediction results (e.g., 82, 83) of a model (81).

[0079] Specifically, the prediction system (10) can train a model (81) using a patient dataset. The model (81) is a trainable model (i.e., a machine learning / deep learning model) and may be a model of the same type as the prediction model described above, or a model of a different type. Then, the prediction system (10) can input the first patient data (84) into the trained model (81) to obtain a first prediction result (82). FIG. 8 illustrates an example where the model (81) is a classification model that outputs a confidence score for an acute kidney injury class (AKI) and a normal class (No-AKI).

[0080] Next, the prediction system (10) can generate second patient data (86) by changing the value (85) of a specific independent variable (e.g., variable 2) in the first patient data (84), and input the second patient data (86) into a retrained model (81) to obtain a second prediction result (83). For example, the prediction system (10) may change the value (85) of a specific independent variable (e.g., variable 2) to '0' or change it to the average value of a dataset belonging to the normal class (or acute kidney injury class).

[0081] Next, the prediction system (10) can calculate the difference between two prediction results (82, 83). The prediction system (10) can measure the influence (e.g., average of the difference values) of a specific independent variable (e.g., variable 2) on the prediction results (e.g., 82, 83) of the model (81) by repeating these processes on other patient data. For example, if the confidence score of the acute kidney injury class (AKI) decreases significantly overall as a result of changing the value of a specific independent variable (e.g., variable 2) to '0', the prediction system (10) can determine that the influence of the independent variable on the prediction results (or dependent variable) of the model (81) is high.

[0082] The prediction system (10) can measure the influence of each independent variable constituting the patient dataset and select independent variables whose measured influence is greater than or equal to a threshold as major independent variables. Then, the prediction system (10) can build a prediction model using the patient dataset composed of major independent variables. By doing so, a higher-performance prediction model can be built.

[0083] Meanwhile, according to some other embodiments of the present disclosure, the prediction system (10) may select key independent variables based on the odds ratio of independent variables. Specifically, the prediction system (10) may train a logistic regression model using a patient dataset and calculate the odds ratio of each independent variable through the trained logistic regression model. Then, the prediction system (10) may select independent variables having odds ratios that differ from '1' by more than a threshold value as key independent variables. For the odds ratios of the independent variables exemplified in Table 1, refer to Tables 3 through 7 below.

[0084] So far, embodiments regarding a method for selecting key independent variables have been described with reference to FIGS. 7 and 8. Hereinafter, a dataset augmentation method according to some embodiments of the present disclosure will be described with reference to FIGS. 9 and 10.

[0085] As illustrated in FIG. 9, the embodiments relate to a method for augmenting a dataset of patients of the acute kidney injury (AKI) class.

[0086] Specifically, the prediction system (10) can sample a latent vector (102) within a data region or latent region (101) in which a patient dataset of the acute kidney injury (AKI) class is encoded. Any method of mapping (or encoding method) the patient dataset of the acute kidney injury (AKI) class to the data region or latent region is acceptable. Then, the prediction system (10) can generate virtual patient data belonging to the acute kidney injury (AKI) class by decoding the sampled latent vector (102). As this sampling and decoding process is repeated, the dataset of the acute kidney injury (AKI) class can be easily augmented.

[0087] Up to now, with reference to FIGS. 9 and FIGS. 10, a dataset augmentation method according to some embodiments of the present disclosure has been described. As described above, by augmenting the patient dataset of the acute kidney injury (AKI) class, the class imbalance problem can be significantly reduced, and accordingly, the performance of the prediction model can be significantly improved.

[0088] Hereinafter, the experimental results performed by the inventors of the present disclosure will be briefly introduced with reference to FIGS. 11 to 14.

[0089] To demonstrate the effectiveness of the technical concept according to the present disclosure, the inventors constructed a model that predicts the risk of acute kidney injury occurring after surgery (specifically, the risk of acute kidney injury occurring within 30 days after surgery) using a dataset of actual patients, and evaluated the performance of the constructed model.

[0090] More specifically, as illustrated in FIG. 11, the inventors prepared a final dataset to be used for training and evaluating the performance of a prediction model by refining a patient dataset (i.e., a cohort dataset) and removing some patient data that met the conditions (see description in FIG. 5). The final dataset consisted of a total of 239,267 data points (i.e., data samples), of which the number of data points corresponding to the acute kidney injury class was 7,935 (i.e., whether acute kidney injury occurred was used as the dependent variable). In addition, the final dataset consisted of patient data for the independent variables exemplified in Table 1.

[0091] Next, the inventors set approximately 80% of the final dataset as the training dataset and the remaining approximately 20% as the test dataset. Then, the inventors constructed prediction models based on artificial neural networks, logistic regression, decision trees, random forests, LGBM, and Naive Bayes using the training dataset, and evaluated the performance of each prediction model using the test dataset. The evaluation results are illustrated in Table 2 below and Figures 12 to 14. Figures 12 to 14 illustrate the Area Under the Curve (AUC) evaluation results for logistic regression, artificial neural networks, and LGBM, respectively. As those skilled in the art would already be familiar with the evaluation metrics illustrated in Table 2 and Figures 12 to 14, a detailed explanation thereof will be omitted.

[0092] division AUC Accuracy Precision Specificity Sensitivity (Recall) F1-score artificial neural networks 0.824 0.727 0.088 0.725 0.775 0.159 Logistic regression 0.818 0.735 0.089 0.735 0.754 0.159 Decision tree 0.747 0.771 0.085 0.777 0.599 0.148 Random Forest 0.803 0.704 0.081 0.702 0.764 0.146 LGBM 0.828 0.712 0.085 0.709 0.793 0.154

[0094] Referring to Table 2 and Figures 12 to 14, it can be seen that the performance of the artificial neural network-based prediction model is slightly superior, and the performance of other types of prediction models is also generally superior. This is believed to be because there are many variables among the independent variables exemplified in Table 1 that are closely associated with the occurrence of acute kidney injury, and because the prediction model performs predictions by comprehensively considering various factors (independent variables).

[0095] In addition, the inventors calculated the odds ratios of each independent variable using a logistic regression-based prediction model to analyze the relationship between the independent variables and the dependent variable. The results of the calculation are listed in Tables 3 to 7 below. Tables 3 and 4 show the odds ratios of independent variables regarding demographic characteristics and disease history, respectively; Table 5 shows the odds ratios of independent variables regarding medication history and time required for surgery; and Tables 6 and 7 show the odds ratios of independent variables regarding pre-operative examination items.

[0096] division Odds 95% confidence interval p-value age 1.021 1.019, 1.023 <0.0001* gender 0.690 0.652, 0.730 <0.0001* BMI 1.011 1.006, 1.017 0.0001*

[0098] division Odds 95% confidence interval p-value Chronic kidney disease (CKD) 2.248 1.728, 2.925 <0.0001* Diabetes mellitus (DM) 1.161 1.050, 1.284 0.0037* High blood pressure (HTN) 1.210 1.080, 1.357 0.0011* Cardiovascular disease (CVD) 1.217 1.118, 1.326 <0.0001* Coronary artery disease (CAD) 1.049 0.917, 1.199 0.4873 Chronic Obstructive Pulmonary Disease (COPD) 1.136 0.896, 1.441 0.2934 Liver cirrhosis (LC) 1.316 1.086, 1.595 0.0051*

[0100] division Odds 95% confidence interval p-value Surgery duration 1.164 1.150, 1.178 <0.0001* Taking ARB / ACEi 1.326 1.216, 1.447 <0.0001* Taking NSAIDs 1.000 0.941, 1.062 0.9997

[0102] division Odds 95% confidence interval p-value SBP 1.013 1.011, 1.014 <0.0001* DBP 0.983 0.981, 0.985 <0.0001* albumin 0.524 0.489, 0.561 <0.0001* ALT 1.000 0.999, 1.001 0.6259 AST 1.000 1.000, 1.001 0.5229 BUN 1.001 0.997, 1.005 0.6289 Calcium (Ca) 1.003 0.959, 1.049 0.8826 Chlorine (Cl) 1.005 0.997, 1.013 0.2040 CPK 1.000 1.000, 1.000 0.1227 Creatinine (Cr) 3.218 2.871, 3.607 <0.0001*

[0104] division Odds 95% confidence interval p-value CRP 0.998 0.997, 0.999 0.0001* eGFR 1.012 1.011, 1.013 <0.0001* Sugar (Glucose) 1.002 1.002, 1.002 <0.0001* hemoglobin 1.000 0.958, 1.043 0.9925 Hematocrit 0.963 0.948, 0.978 <0.0001* Potassium (K) 0.755 0.715, 0.798 <0.0001* LDH 1.000 1.000, 1.000 <0.0001* Sodium (Na) 0.980 0.971, 0.990 <0.0001* Total Protein 1.042 0.998, 1.089 0.0644 Uric acid 1.078 1.062, 1.095 <0.0001* protein in urine 1.369 1.325, 1.416 <0.0001* Urine specific gravity (SG) 0.042 0.009, 0.196 <0.0001* Urine white blood cell count (WBC) 1.005 1.003, 1.008 <0.0001*

[0106] Referring to Tables 3, 5 through 7, it can be seen that independent variables regarding gender, chronic kidney disease (CKD), hypertension (HTN), cardiovascular disease (CVD), liver cirrhosis (LC), use of angiotensin receptor blockers (ARBs), angiotensin-converting enzyme inhibitors (ACEi), and non-steroidal anti-inflammatory drugs (NSAIDs) have a relatively close association with the dependent variable (e.g., occurrence of acute kidney injury after surgery). In addition, independent variables regarding albumin, creatinine (Cr), potassium (K), protein, and urine specific gravity (SG) also have a relatively close association with the dependent variable.

[0107] The results of experiments performed by the inventors so far have been briefly introduced. Below, with reference to FIG. 15, an exemplary computing device (150) capable of implementing a prediction system (10) according to some embodiments of the present disclosure will be described.

[0108] FIG. 15 is an exemplary hardware configuration diagram showing a computing device (150).

[0109] As illustrated in FIG. 15, a computing device (150) may include one or more processors (151), a bus (153), a communication interface (154), a memory (152) for loading a computer program executed by the processor (151), and a storage (155) for storing a computer program (156). However, FIG. 15 illustrates only the components related to the embodiments of the present disclosure. Therefore, a person skilled in the art to which the present disclosure belongs will understand that other general-purpose components may be included in addition to the components illustrated in FIG. 15. That is, the computing device (150) may include various additional components in addition to the components illustrated in FIG. 15. Furthermore, depending on the case, the computing device (150) may be configured in a form in which some of the components illustrated in FIG. 15 are omitted. Each component of the computing device (150) will be described below.

[0110] The processor (151) can control the overall operation of each component of the computing device (150). The processor (151) may be configured to include at least one of a CPU (Central Processing Unit), MPU (Micro Processor Unit), MCU (Micro Controller Unit), GPU (Graphic Processing Unit), or any form of processor well known in the art of the present disclosure. Additionally, the processor (151) may perform operations for at least one application or program for executing operations / methods according to embodiments of the present disclosure. The computing device (150) may have one or more processors.

[0111] Next, the memory (152) may store various data, commands and / or information. The memory (152) may load a computer program (156) from storage (155) to execute an operation / method according to embodiments of the present disclosure. The memory (152) may be implemented as a volatile memory such as RAM, but the technical scope of the present disclosure is not limited thereto.

[0112] Next, the bus (153) can provide communication functions between components of the computing device (150). The bus (153) can be implemented as various types of buses, such as an address bus, a data bus, and a control bus.

[0113] Next, the communication interface (154) may support wired and wireless internet communication of the computing device (150). Additionally, the communication interface (154) may support various communication methods other than internet communication. To this end, the communication interface (154) may be configured to include a communication module well known in the art of the present disclosure.

[0114] Next, the storage (155) may store one or more computer programs (156) non-temporarily. The storage (155) may be configured to include non-volatile memory such as ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), flash memory, a hard disk, a removable disk, or any form of computer-readable recording medium well known in the art to which this disclosure belongs.

[0115] Next, the computer program (156) may include one or more instructions that cause the processor (151) to perform an operation / method according to various embodiments of the present disclosure when loaded into memory (152). That is, the processor (151) may perform an operation / method according to various embodiments of the present disclosure by executing the one or more instructions.

[0116] For example, the computer program (156) may include one or more instructions to perform the operation of acquiring a model learned to predict the risk of acute kidney injury occurring after surgery and the operation of predicting the risk of acute kidney injury occurring in a patient after the target surgery using the learned model. In such a case, the prediction system (10) according to some embodiments of the present disclosure may be implemented through the computing device (150).

[0117] Up to now, with reference to FIG. 15, an exemplary computing device (150) capable of implementing a prediction system (10) according to some embodiments of the present disclosure has been described.

[0118] The technical concept of the present disclosure, as described so far with reference to FIGS. 1 through 15, may be implemented as computer-readable code on a computer-readable medium. The computer-readable recording medium may be, for example, a removable recording medium (CD, DVD, Blu-ray disc, USB storage device, removable hard disk) or a fixed recording medium (ROM, RAM, computer-equipped hard disk). The computer program recorded on the computer-readable recording medium may be transmitted to another computing device via a network such as the Internet and installed on the other computing device, thereby being used on the other computing device.

[0119] In the foregoing, although all components constituting the embodiments of the present disclosure have been described as being combined or operating together, the technical concept of the present disclosure is not necessarily limited to such embodiments. That is, within the scope of the purpose of the present disclosure, all components may be selectively combined and operated in one or more ways.

[0120] Although operations are depicted in a specific order in the drawings, it should not be understood that the operations must be executed in the specific order depicted or in a sequential order, or that all depicted operations must be executed to obtain the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various configurations in the embodiments described above should not be understood as a necessary separation, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.

[0121] Although embodiments of the present disclosure have been described above with reference to the attached drawings, those skilled in the art will understand that the present disclosure may be practiced in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. The scope of protection of the present disclosure shall be interpreted by the claims below, and all technical concepts within the equivalent scope shall be interpreted as being included within the scope of rights of the technical concepts defined by the present disclosure.

Claims

Claim 1 A method performed by at least one computing device, comprising the step of preparing a dataset for a plurality of patients— wherein the dependent variable of the dataset relates to the occurrence of acute kidney injury after surgery, and the independent variables of the dataset include variables related to preoperative examination items of the patients; A method for predicting the occurrence of acute kidney injury, comprising the step of constructing a model that predicts the risk of acute kidney injury occurring after surgery using the prepared dataset, wherein the pre-operative examination items include albumin, creatinine (Cr), potassium, protein, and urine specific gravity, and the independent variables of the dataset include variables regarding the patients' disease history, medication history, and the type and duration of surgery received by the patients, wherein the disease history includes a history regarding chronic kidney disease (CKD), hypertension (HTN), cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), and liver cirrhosis (LC), and the medication history is regarding antihypertensive drugs, and the model is based on at least one deep learning or machine learning model among a neural network, logistic regression, and LGBM (Light Gradient Boosting Machine), and constructing a first prediction model using the prepared dataset, and further constructing a second prediction model of a different type from the first prediction model. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 A method for predicting the occurrence of acute kidney injury according to claim 1, wherein the step of preparing the dataset includes the step of removing data of patients who meet a predetermined kidney-related condition from the original patient dataset. Claim 7 A method for predicting the occurrence of acute kidney injury according to claim 6, wherein the aforementioned specified kidney-related conditions are defined based on a history of renal replacement therapy or preoperative eGFR values. Claim 8 A method for predicting the occurrence of acute kidney injury according to claim 6, wherein the above-mentioned predetermined kidney function condition is defined based on the creatinine (Cr) level before surgery or the degree to which the creatinine (Cr) level has increased within a certain period prior to surgery. Claim 9 A method for predicting the occurrence of acute kidney injury according to claim 1, wherein the step of preparing the dataset includes the step of removing data of patients who meet a predetermined surgery-related condition from the original patient dataset, and the predetermined surgery-related condition is defined based on the time required for surgery or the type of surgery. Claim 10 A method for predicting the occurrence of acute kidney injury according to claim 1, wherein the step of preparing the dataset comprises: a step of correcting outliers in the original patient dataset; a step of correcting missing values ​​in the original patient dataset using Multiple Imputation by Chained Equations; and a step of normalizing the original patient dataset in which the outliers and missing values ​​have been corrected. Claim 11 A method for predicting the occurrence of acute kidney injury according to claim 1, wherein the step of preparing the dataset comprises: a step of acquiring an original patient dataset—the original patient dataset includes a first dataset for a group of patients in whom acute kidney injury occurred after surgery and a second dataset for a group of patients in whom it did not occur—; and a step of augmenting the first dataset. Claim 12 A method performed by at least one computing device, comprising the step of obtaining a model trained to predict the risk of acute kidney injury occurring after surgery— said model being trained using a dataset of multiple patients, said dataset having dependent variables relating to the occurrence of acute kidney injury after surgery, and said dataset having independent variables relating to preoperative examination items of said patients— ; A method for predicting the occurrence of acute kidney injury, comprising the step of predicting the risk of acute kidney injury occurring in a specific patient after a target surgery using the above-mentioned learned model, wherein the pre-operative examination items include albumin, creatinine (Cr), potassium, protein, and urine specific gravity, and the independent variables of the above-mentioned dataset include variables regarding the patients' disease history, medication history, and the type and duration of surgery received by the patients, wherein the disease history includes a history regarding chronic kidney disease (CKD), hypertension (HTN), cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), and liver cirrhosis (LC), and the medication history is regarding antihypertensive drugs, and wherein the model is based on at least one deep learning or machine learning model among a neural network, logistic regression, and LGBM (Light Gradient Boosting Machine), and wherein a first prediction model is constructed using the above-mentioned prepared dataset, and a second prediction model of a different type from the first prediction model is further constructed. Claim 13 A method for predicting the occurrence of acute kidney injury according to claim 12, wherein the predicting step comprises: a step of configuring input data based on the type and duration of the target surgery and the test results of the specific patient regarding the pre-operative test items; and a step of inputting the input data into the learned model to predict the risk. Claim 14 One or more processors; and includes a memory for storing one or more instructions, and the one or more processors perform the operation of obtaining a model trained to predict the risk of acute kidney injury occurring after surgery by executing the one or more stored instructions—the model is trained using a dataset of multiple patients, the dependent variable of the dataset is related to the occurrence of acute kidney injury after surgery, and the independent variables of the dataset include variables related to preoperative examination items of the patients—and perform the operation of predicting the risk of acute kidney injury occurring in a patient after a target surgery using the trained model, wherein the preoperative examination items include albumin, creatinine (Cr), potassium, protein, and urine specific gravity, and the independent variables of the dataset include variables related to the patients' disease history, medication history, and the type and duration of surgery received by the patients, wherein the disease history includes a history regarding chronic kidney disease (CKD), hypertension (HTN), cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), and liver cirrhosis (LC), and the medication history An acute kidney injury occurrence prediction system relating to antihypertensive drugs, wherein the model is based on at least one deep learning or machine learning model among a neural network, logistic regression, and LGBM (Light Gradient Boosting Machine), and wherein a first prediction model is constructed using the prepared dataset, and a second prediction model of a different type from the first prediction model is further constructed.