Prediction of Albuminuria Using Machine Learning
By employing a machine learning model to predict UACR levels based on patient data, the challenge of delayed albuminuria detection is addressed, enabling early intervention and improved health outcomes.
Patent Information
- Application Number
- JP2024555968
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-04
- Filing Date
- 2023-05-12
- Publication Date
- 2025-06-17
AI Technical Summary
Current methods for detecting albuminuria, a marker of chronic kidney disease, are inadequate as they often require urine samples, which are not commonly collected from patients without diabetes or kidney issues, leading to delayed diagnosis and treatment.
A machine learning model is developed using patient demographics, vital signs, and blood test data to predict urine albumin-to-creatinine ratio (UACR) levels, enabling early identification of undiagnosed albuminuria without the need for urine samples.
This approach allows for earlier detection of albuminuria, reducing the risk of kidney function deterioration and associated cardiovascular events, and facilitates timely medical intervention and potential participation in clinical trials.
Smart Images

Figure 2025518440000007 
Figure 2025518440000008 
Figure 2025518440000009
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 456,855, filed on April 4, 2023, and U.S. Provisional Patent Application No. 63 / 343,778, filed on May 19, 2022, and the entire contents of both are hereby incorporated by reference into this application.
[0002] This disclosure relates to the prediction of albuminuria using machine learning.
Background Art
[0003] Albumin is a protein secreted by the liver and found in the blood. A properly functioning kidney does not permit albumin to permeate from the blood into the urine. Albuminuria is a pathological condition in which albumin abnormally present in the urine serves as an indicator of chronic kidney disease (CKD). If left untreated, CKD can lead to kidney failure and death.
Prior Art Documents
Non - Patent Documents
[0004]
Non - Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The measured value of albumin in urine can be used to evaluate the progression of CKD, serves as an indicator that a patient is a candidate for treatment, and also becomes an alternative marker for treatment effect. To detect the presence of other markers of CKD, generally blood tests are used, but relatively few patients are screened for albuminuria. In particular, the albuminuria test requires the collection of a urine sample, which is generally not indicated for patients without diabetes or other kidney-related conditions. Therefore, some patients may be affected by albuminuria over a long time period (e.g., months or years) without receiving appropriate diagnosis or necessary treatment.
Means for Solving the Problems
[0006] The urine albumin level can be determined from a sample using the urine albumin-to-creatinine ratio (UACR) test. Results in the range of 30 - 300 mg / g are called microalbuminuria, and results exceeding 300 mg / g are called macroalbuminuria. Nevertheless, this cutoff at 300 mg / g is somewhat arbitrary and does not show a clinically significant inflection point. In some clinical scenarios, the cutoff may vary and can be any value between 150 mg / g and 700 mg / g.
[0007] Early detection of which category albuminuria is in, and subsequent medical intervention, can reduce the risk of worsening kidney function, including progression of CKD and progression to end-stage renal disease (ESRD). Early detection may also lead to a reduced risk of incidental cardiovascular events, even when albuminuria is detected in patients with normal kidney function. If UACR results are not available for many patients, it is beneficial to improve albuminuria screening by identifying patients suspected of undiagnosed albuminuria based on more widely available biomarkers. Data on these biomarkers already exist in the electronic health records of patients (e.g., from demographic data, health status data, blood tests, etc.). Therefore, based on these records and / or other available medical information, it is possible to predict which patients may have undiagnosed albuminuria. Once these patients are identified, they can be contacted prospectively for UACR testing and treatment that may indicate albuminuria.
[0008] Accordingly, the applications of the embodiments of the present application can identify albuminuria earlier than would otherwise be detected without the present application. Early detection and subsequent early medical intervention can reduce the impact of CKD and prevent ultimate renal failure. Furthermore, the identified patients may also be recruited for medical research or clinical trials of new treatments and / or medications that may lead to improved treatment options and health outcomes.
[0009] To address these issues, these embodiments use a machine learning model learned with respect to a combination of patient demographics, vital signs, blood tests, and / or other medical information. The learning results in a classifier that predicts UACR levels. In some cases, the machine learning model is based on gradient boosting techniques, but other underlying techniques (e.g., artificial neural networks or expert systems) may be used instead of, or in connection with, gradient boosting. Such models have been shown to be effective when applied to clinical data to predict UACR levels for a wide range of patients. In some embodiments, urine protein-to-creatinine ratio (UPR) levels may be used instead of, or to derive, UACR values. These UACR and UPR levels may be calculated from separate measurements of urinary albumin, urinary creatinine, and / or urinary protein.
[0010] Accordingly, a first exemplary embodiment includes obtaining, by a computing system, a training dataset and applying, by the computing system, a machine learning trainer to the training dataset. The training dataset includes observed values of corresponding demographic values, vital sign values, blood test values, and either UACR values or UPR values for a plurality of individuals. The machine learning trainer generates a machine learning model. The machine learning model is configured to receive, as input, new observed values of new demographic values, new vital sign values, and new blood test values and provide a prediction as to whether an individual presenting the new observed values has undiagnosed albuminuria or proteinuria.
[0011] A second exemplary embodiment includes obtaining, by a computing system, observed values of an individual's demographic values, an individual's vital sign values, and an individual's blood test values; applying, by the computing system, a machine learning model to the observed values; and providing, by the computing system, a prediction as to whether the individual exhibits undiagnosed albuminuria or proteinuria based on the observed values. The machine learning model is learned by a training dataset. The training dataset includes observed values of corresponding demographic values, vital sign values, blood test values, and either a UACR value or a UPR value for a plurality of individuals. The machine learning model is configured to provide a prediction as to whether further observed values indicate undiagnosed albuminuria or proteinuria.
[0012] A third exemplary embodiment includes obtaining, by a computing system, a training dataset; and applying, by the computing system, a quantile regression machine learning trainer to the training dataset. The training dataset includes observed values of corresponding demographic values, vital sign values, blood test values, and either a UACR value or a UPR value for a plurality of individuals. The quantile regression machine learning trainer generates a quantile regression machine learning model. The quantile regression machine learning model is configured to receive, as input, a quantile of new demographic values, new vital sign values, and new blood test values and new observed values, and provide a prediction of a UACR or UPR value at the quantile for an individual presenting the new observed values.
[0013] A fourth exemplary embodiment includes obtaining, by a computing system, percentile and observed values of an individual's demographic values, an individual's vital sign values, and an individual's blood test values; applying, by the computing system, a quantile regression machine learning model to the observed values; and providing, by the computing system, a prediction of a UACR or UPR value at a percentile, based on the observed values and with respect to the individual. The quantile regression machine learning model is learned by a learning dataset. The learning dataset includes observed values of corresponding demographic values, vital sign values, blood test values, and either a UACR value or a UPR value for a plurality of individuals. The quantile regression machine learning model is configured to provide a prediction of a UACR or UPR value at one or more percentiles with respect to further observed values.
[0014] In a fifth exemplary embodiment, a product includes a non-transitory computer-readable medium storing program instructions that, when executed by a computing system, cause the computing system to perform operations according to the first, second, third, and / or fourth exemplary embodiments.
[0015] In a sixth exemplary embodiment, a computing system includes at least one processor, memory, and program instructions. The program instructions may be stored in the memory and, when executed by the at least one processor, may cause the computing system to perform operations according to the first, second, third, and / or fourth exemplary embodiments.
[0016] In a seventh exemplary embodiment, a system includes various means for performing each of the operations according to the first, second, third, and / or fourth exemplary embodiments.
[0017] Together with other embodiments, these aspects, advantages, and alternatives will become apparent to those skilled in the art by reading the following detailed description, with appropriate reference to the accompanying drawings. Further, this summary and other descriptions and drawings provided herein are intended to illustrate embodiments by way of example only, and thus, numerous variations are possible. For example, structural elements and processing steps may be rearranged, combined, distributed, removed, or otherwise changed within the scope of the embodiments described in the claims.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 12C
Figure 12D
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19A
Figure 19B
MODE FOR CARRYING OUT THE INVENTION
[0019] In this application, exemplary methods, apparatuses, and systems will be described. The terms "embodiment" and "exemplary" should be understood in this application to mean "provided as an example, instance, or illustration". Any embodiment or feature described as "embodiment" or "exemplary" in this application need not be construed as being preferred or advantageous over other embodiments or features, unless such reference is made. Accordingly, other embodiments may be utilized and other changes may be made without departing from the scope of the subject matter presented in this application.
[0020] Accordingly, the exemplary embodiments described in this application are not meant to be limiting. It will be readily understood that the aspects of the present disclosure generally outlined herein and shown in the drawings may be arranged, replaced, combined, separated, and designed in a variety of different configurations. For example, the separation of features into "client" and "server" components may be done in numerous ways.
[0021] Furthermore, unless the context otherwise proposes, the features shown in each figure may be used in combination with each other. Accordingly, the drawings should generally be seen as component aspects of one or more overall embodiments, understanding that not all of the features illustrated are required in each embodiment.
[0022] Furthermore, any listing of components, blocks, or steps in this specification or the claims is for the purpose of clarity. Accordingly, such listing should not be construed as requiring or implying that these components, blocks, or steps are fixed in a particular configuration or are to be performed in a particular order.
[0023] I. Exemplary Computing Devices and Cloud-Based Computing Environments FIG. 1 is a simplified block diagram illustrating a computing device 100 and shows some of the components that may be included in a computing device configured to operate in accordance with embodiments of the present application. The computing device 100 may be a client device (e.g., a device actively operated by a user), a server device (e.g., a device that provides computing services to client devices), or some other type of computing platform. Some server devices may sometimes operate as client devices to perform certain operations, and some client devices may incorporate server features.
[0024] In this embodiment, the computing device 100 includes a processor 102, a memory 104, a network interface 106, and an input / output device 108, all of which may be connected by a system bus 110 or a similar mechanism. In some embodiments, the computing device 100 may include other components and / or peripherals (e.g., a removable storage device, a printer, etc.).
[0025] The processor 102 may be any type of computer processing component, such as a central processing unit (CPU), a coprocessor (e.g., a math, graphics, or cryptographic coprocessor), a digital signal processor (DSP), a network processor, and / or an integrated circuit or controller in a predetermined form that executes processor operations. In some cases, the processor 102 may be one or more single-core processors. In other cases, the processor 102 may be one or more multi-core processors having multiple independent processing units. The processor 102 may also include register memory for temporarily storing executed instructions and related data, and cache memory for temporarily storing recently used instructions and data.
[0026] Memory 104 may be any form of memory usable by a computer, including but not limited to random access memory (RAM), read-only memory (ROM), and non-volatile memory (e.g., flash memory, hard disk drive, solid state drive, and / or tape storage device). Thus, memory 104 represents both a main memory device and a long-term storage device. Other types of memory may include biological memory.
[0027] Memory 104 may store program instructions and / or data on which the program instructions may act. By way of example, memory 104 may store these program instructions on a non-transitory computer-readable medium. It is such that the instructions are executable by processor 102 to perform any of the methods, processes, or operations disclosed in this specification or the accompanying drawings.
[0028] As shown in FIG. 1, memory 104 may include firmware 104A, kernel 104B, and / or application 104C. Firmware 104A may be program code used to boot or otherwise initiate some or all of computing device 100. Kernel 104B may be an operating system and may include modules for memory management, scheduling, and management of processing, input / output, and communication. Kernel 104B may also include device drivers that enable the operating system to communicate with hardware modules of computing device 100 (e.g., memory devices, networking interfaces, ports, and buses). Application 104C may be one or more user-space software programs such as a web browser or an email client, or any software library used by these programs. Memory 104 may also store data used by these and other programs and applications.
[0029] The network interface 106 may take the form of one or more wired interfaces, such as Ethernet (registered trademark) (e.g., Fast Ethernet, Gigabit Ethernet, etc.). The network interface 106 may also support communication via one or more non-Ethernet media, such as coaxial cables or power lines, or may support communication via wide-area media, such as SONET (Synchronous Optical Networking) or software-defined wide-area networking (SD-WAN) technology. The network interface 106 may further take the form of one or more wireless interfaces, such as IEEE802.11 (Wifi (registered trademark)), BLUETOOTH (registered trademark), global positioning system (GPS), or wide-area wireless interfaces. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used via the network interface 106. Further, the network interface 106 may comprise a plurality of physical interfaces. For example, some embodiments of the computing device 100 may include Ethernet, BLUETOOTH, and Wifi interfaces.
[0030] The input / output device 108 can facilitate the interaction between the user and the computing device 100 of the peripheral device. The input / output device 108 may include one or more types of input devices such as a keyboard, a mouse, a touch screen, and the like. Similarly, the input / output device 108 may include one or more types of output devices such as a screen, a monitor, a printer, and / or one or more light emitting diodes (LEDs). Additionally or alternatively, the computing device 100 may communicate with other devices, for example, using a USB (universal serial bus) or HDMI (high-definition multimedia interface: registered trademark) port interface.
[0031] One or more computing devices such as the computing device 100 may be deployed to support the embodiments of the present application. The exact physical location, connectivity, and configuration of these computing devices may be unknown and / or unimportant to the client device. Thus, the computing devices may sometimes be referred to as "cloud-based" devices, which may be housed at various remote data center locations.
[0032] FIG. 2 shows a cloud-based server cluster 200 according to an exemplary embodiment. In FIG. 2, the operation of the computing device (e.g., the computing device 100) may be distributed among a plurality of server devices 202, a data storage device 204, and a plurality of routers 206. All of them may be connected by a local cluster network 208. The number of server devices 202, data storage devices 204, and routers 206 in the server cluster 200 may depend on one or more computing tasks and / or applications assigned to the server cluster 200.
[0033] For example, the server device 202 may be configured to execute various computing tasks of the computing device 100. Accordingly, the computing tasks may be distributed across one or more of the server devices 202. Such a distribution of tasks can reduce the total time to complete these tasks and return the results to the extent that these computing tasks can be executed in parallel. For the purpose of simplification, both the server cluster 200 and the individual server devices 202 may sometimes be referred to as "server devices". This nomenclature should be understood to mean that one or more separate server devices, data storage devices, and cluster routers may be involved in the operation of the server device.
[0034] The data storage device 204 may be an array of data storage devices including a drive array controller configured to manage read and write access to a group of hard disk drives and / or solid state drives. The drive array controller may also manage backups or redundant copies of the data stored in the data storage device 204, either alone or together with the server device 202, to protect one or more of the server devices 202 from drive failures or other types of failures that would prevent access to the devices of the data storage device 204. Other types of memory may be used in addition to the drives.
[0035] The router 206 may include networking equipment configured to provide internal and external communication for the server cluster 200. For example, the router 206 may include one or more packet switching devices and / or routing devices (including switches and / or gateways) configured to provide (i) network communication between the server device 202 and the data storage device 204 via the local cluster network 208, and / or (ii) network communication between the server cluster 200 and other devices via the communication link 210 to the network 212.
[0036] Furthermore, the configuration of router 206 may be at least partially based on the data communication requirements of server device 202 and data storage device 204, the latency and throughput of local cluster network 208, the latency, throughput, and cost of communication link 210, and / or other factors that may contribute to cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the system architecture.
[0037] As a possible example, data storage device 204 may include any form of database, such as a structured query language (SQL) database. Various types of data structures may store information including, but not limited to, tables, arrays, lists, trees, and tuples in such a database. Furthermore, any database in data storage device 204 may be monolithic or distributed across multiple physical devices.
[0038] Server device 202 may be configured to transmit and receive data to and from data storage device 204. This transmission and retrieval may take the form of SQL queries or other types of database queries and the output of such queries, respectively. Additional text, images, video, and / or audio may be included as well. Furthermore, server device 202 may organize the received data into a web page or web application representation or organize it in some other way for use by a software application. Such a representation may take the form of a markup language, such as HTML, XML (eXtensible Markup Language), or some other standardized or proprietary format.
[0039] In addition, the server device 202 may have the ability to execute various types of computerized scripting languages, such as, but not limited to, Perl, Python, PHP (PHP Hypertext Preprocessor), ASP (Active Server Pages), JAVASCRIPT (registered trademark), and the like. The computer program code written in these languages can facilitate the provision of web pages to the client device and the interaction of the client device with the web pages. Alternatively or additionally, JAVA (registered trademark) may be used to facilitate the generation of web pages and / or to provide web application functionality.
[0040] II. Exemplary Gradient Boosting Model The gradient boosting algorithm is a machine learning technique that can be used to develop a prediction model for a multi-dimensional dataset. For convenience, these sets are often represented in matrix form using rows and columns. One or more columns represent input variables, and an additional column represents an output variable. The output variable is an unknown function of one or more of the input variables. The rows represent observations of the input variables and their corresponding output variables, usually based on real-world data. In many cases, the number of rows may be very large, hundreds, thousands, or more. The machine learning process involves learning a gradient boosting model so that it can predict the output variable for new observations of the input variables. In other words, the model attempts to learn or at least approximate the unknown function from the existing instances of the input variables and their corresponding output variables.
[0041] Figure 3 further illustrates these concepts. The training dataset 300 includes a set of training observations (rows), each consisting of input variables X1, X2, and X3 and their corresponding output variable Y. These input and output variables are related by some unknown function f. Here, Y = f(X1, X2, X3). The output variable Y may take various forms, such as an integer or real number, text, or a boolean value.
[0042] The learning dataset 300 may be collected from medical data of actual patients, for example, from medical experts, hospitals, clinical trials, or other sources. In such a learning dataset, the value of the output variable for each observation is expected to be known, but not all values of the input variables need to be present. For example, the learning dataset may be sparse.
[0043] The learning dataset 300 is provided to a gradient boosting trainer 302, which applies one or more learning techniques to generate a gradient boosting model 304. The gradient boosting model 304 may be an algorithm that can be used to apply an approximation of an unknown function f to new observations of the input variables, or it may be a set of parameters that control the behavior of the algorithm.
[0044] Accordingly, the gradient boosting model 304 may receive new observations 306 and generate a predicted output variable 308. The accuracy of such predictions may vary depending on the operation of the gradient boosting trainer 302 and the quality of the learning dataset 300. The goal is to make the gradient boosting model 304 as accurate as reasonably possible given a sufficiently rich learning dataset and a reasonable amount of time spent on learning. As will be described in more detail later, this accuracy can be measured in various ways.
[0045] The operation of the gradient boosting trainer 302 may include learning a set of decision trees (colloquially referred to as a "forest") each having a limited depth or a limited number of leaves. Thus, these trees are weak learners in that they generally do not consider all the available information in the training dataset, and thus, their individual predictions may or may not have high accuracy. However, gradient boosting makes an overall prediction based on the weighting of the predictions from the individual trees. These overall predictions consider all or most of the training dataset and thus may be more accurate than the predictions from any of the individual trees.
[0046] However, unlike a random forest where each tree is independent of the others, the construction of subsequent trees in a gradient boosting model may be based on the error (or residual) associated with one or more of the previously constructed trees. In some cases, subsequent trees that compensate well for the errors of the previous trees are given a greater weight towards the overall prediction, while in other cases, all trees may be equally weighted. Gradient boosting continues to construct trees in this way until a predetermined number of trees are constructed or until new trees no longer improve the accuracy of the prediction beyond a predetermined margin.
[0047] When the output variable takes on continuous values, e.g., integers within some range, the trees are constructed based on the magnitude of the residual between the actual value of the output variable in the training data and the associated predicted value. This is sometimes referred to as gradient boosting for regression. In some cases, these residuals are referred to as "pseudo-residuals" to distinguish gradient boosting from linear regression, but the terms "residual" and "pseudo-residual" are used interchangeably in this application. The initial prediction p i,0 for each observation i may have the same value p0, such as the average of some or all of the output variables in the training dataset. In other words, p0 = p i,0 ∀ i is true.
[0048] Here, the notation p i,j represents a prediction regarding the observed value i made using trees 0 to n (more specifically, see the following regarding the method of calculating the prediction using multiple trees). Since the initial prediction p0 is generally based only on the value of the output variable, it may take the form of a single node rather than a tree.
[0049] In any case, the tree is configured to predict the value of the residual. The non-leaf nodes of the tree represent the state of the input variables. For example, the root node in a tree constructed from the learning dataset 300 may represent the state X2 > 5, such that if this state is true, the left branch of the node is followed, and if this state is false, the right branch of the node is followed. Either of these branches may lead to another non-leaf node representing a state or a leaf node representing a residual. There may be more than two branches, but for the sake of convenience, binary trees are used in the embodiments of the present application.
[0050] The tree construction may be based on various algorithms used for decision trees. In some cases, this may include selecting the input variables and, possibly, associated cutoff values based on entropy or Gini impurity. The cutoff values are selected to divide the values of the input variables in a way that rationally helps the input variables in predicting the output variable. The input variables are then arranged as nodes in the tree such that the input variables that are more helpful for prediction are generally placed higher in the tree (e.g., closer to the root node). In some cases, randomness may be added to the process of determining where to place the input variables in the tree.
[0051] Since the number of observed values is usually much larger than the number of leaves in a tree of limited size, the residuals of each observed value that result in the same leaf are typically averaged and then placed in the leaf. Thus, a leaf may represent the aggregated residual r i,0 for a number of observed values. Here, the notation r i,jshows the residuals for the observed value i made using trees 0 to n (more specifically, see below regarding the method of calculating the residuals using multiple trees).
[0052] Subsequently, iterative predictions are made for the observed values in the training dataset. Each prediction p in the first iteration i,1 traverses tree 1 with respect to the observed value until reaching a leaf, and then adds the residual r of that leaf i,0 to the initial prediction according to the learning rate 0 < α < 1. In other words, p i,1 = p i,0 + αr i,0 The learning rate helps prevent overfitting of the training dataset and enables taking small steps for higher prediction accuracy.
[0053] From the prediction p i,1 a new residual r i,1 is calculated. Again, the residual is based on the difference between the actual output variable value in the training dataset and the associated predicted value with possible aggregations as described above. These new residuals are generally expected to be smaller than those of r i,0 , but this does not apply to all cases of residuals.
[0054] Based on these new residuals, a next tree, tree 2, may be constructed. This tree may have the same structure as tree 1, or may have a different structure such that the nodes representing the input variables appear at different locations (e.g., using randomness).
[0055] Next, each prediction p in the second iteration i,2 traverses both tree 1 and tree 2 for the observed value until reaching their respective leaves, and then adds the associated residuals r i,0 and r i,1 to the initial prediction according to the learning rate. In other words, p i,2 = p i,0 + αr i,0 + αr i,1It is.
[0056] Prediction p i,2 From this, a new residual r i,2 is calculated. As the number of trees increases, these residuals are expected to continue to decrease.
[0057] The process of constructing a new tree based on the residuals and making a new prediction continues, as described above, until a predetermined number of trees are constructed or until adding a new tree no longer reduces the size of the residuals beyond a predetermined margin. At this point, the learning is complete, and the learned gradient boosting model is ready to make predictions regarding new observations of the input variables.
[0058] When such a new observation is given, the model starts from the initial prediction p0, traverses all of the trees according to the values of the input variables, and adds the resulting residuals. Thus, assuming n trees in addition to the initial node, the predicted value of the output variable for the new observation is given by the following formula.
[0059]
Equation
[0060] Here, r0 is the initial residual, and r j , 1 ≦ j ≦ n is the residual of the j-th tree for this observation.
[0061] Gradient boosting may be used to predict the value of the output variable from a limited number of possible values. For example, if the output variable is a boolean value, gradient boosting may be used to learn a binary classifier. This is sometimes called gradient boosting for classification.
[0062] In this case, the prediction may be based on, for all observations in the training dataset, (i) the natural logarithm of the odds that the output variable is true, and (ii) the probability that the output variable is true. For example, assume that there are 100 observations including 70 "true" and 30 "false". The natural logarithm of the odds that the observation is true is ln(70 / 30) = 0.847, while the probability that the output variable is true is 0.7.
[0063] The natural logarithm of the odds (0.847) is used as the initial prediction for all observations, and the probability (0.7) is used to calculate the residuals. Since 0.7 is greater than 0.5, the initial prediction is "true" for all observations (note that values other than 0.5 can be used as the cutoff in this process). Clearly, these initial predictions are not accurate as indicated by their residuals. Assigning a value of 1.0 to true and 0.0 to false, the residuals are 0.3 for each observation with an output variable that is "true" and -0.7 for each observation with an output variable that is "false".
[0064] Furthermore, the residuals are typically transformed to produce the output value of the leaf. This is because they are a function of probability, while the prediction is a function of the natural logarithm of the odds. For a leaf with residual r k an exemplary transformation is given by the following equation when the corresponding predicted probability is ρ k as follows.
[0065]
Equation
[0066] These output values for the leaf are then scaled by the learning rate and added to the initial prediction. These predictions still have the form of the natural logarithm of the odds and can then be converted to the probability form used by the residuals via the application, so they are logistic functions. For example, the predicted value of p has the probability of the following equation.
[0067]
Number
[0068] These residuals are then determined as the difference between the value of the output value from the training dataset and the probability. As described above, a new tree can be calculated based on these residuals. This process continues until a predetermined number of trees are constructed or until adding a new tree can no longer reduce the size of the residuals beyond a predetermined margin.
[0069] The learned gradient boosting model is then applied to new observations in a similar manner. The process adds the initial prediction and the transformed output value of each leaf associated with the new observation to find the natural logarithm of the predicted odds. Then, a logistic function is applied to this prediction to provide a probability. If the probability is greater than 0.5, the final prediction for this new observation is "true"; otherwise, it is "false".
[0070] Note that the above description merely provides a general overview of some methods for performing gradient boosting. Other techniques are also possible. Further, instead of or in connection with gradient boosting techniques, other types of machine learning models such as artificial neural networks or expert systems may be used. Nevertheless, gradient boosting generally works well for a given type of data and produces an explainable model, i.e., a model that can be analyzed to understand how and why it generates its answers. This may not be the case for artificial neural networks. Also, gradient boosting models also tend to work well when the training dataset is sparse (e.g., a large number of observations containing at least one missing input variable), whereas artificial neural networks tend to be less efficient for sparse data.
[0071] There are two popular gradient boosting frameworks, XGBoost and LightGBM, that can be used to learn and deploy gradient boosting models. Each is briefly described below for illustrative purposes. Nevertheless, other frameworks may be used.
[0072] A. XGBoost When the output variable takes on continuous values, XGBoost constructs a tree by placing all the residuals at the root node and then calculates a similarity score for the residuals. The similarity score may be calculated as the squared sum of the residuals divided by the number of residuals. In some cases, a regularization constant is also added to the denominator to reduce sensitivity to outliers and overfitting in the training dataset. Regardless, as the similarity score gets higher, the residuals become more similar.
[0073] Next, various ways of splitting the residuals into multiple groups are considered, and it is ascertained whether any of these splits results in a higher overall similarity score. This may be determined by calculating the gain of the split with respect to the original grouping of all the residuals at the root node. The gain may be the similarity score of the root node subtracted from the sum of the similarity scores for each node of the split. Next, the split that generates the maximum gain among all splits is selected so as to generate a branch from the root node (i.e., each group of residuals in the selected split becomes a child node of the root node).
[0074] Next, the same process is executed for each of the new child nodes. If a child node contains only one residual, it cannot be further split and becomes a leaf. Also, the tree may be limited to a maximum number of levels (e.g., 4, 6, or 8), and nodes at this maximum depth are also not further split. In some cases, XGBoost may require that a minimum number of residuals (e.g., 2, 3, 4, …) be represented at each node.
[0075] Once the XGBoost tree is constructed in this way, some branches that generate a threshold gain less than a certain value may be pruned. The pruning process reconnects them to their respective parent nodes.
[0076] Predictions are made by traversing the tree with the values of the input variables until a leaf is reached. The output value of a leaf is the sum of the residuals in that leaf divided by the number of residuals in that leaf. Again, a normalization constant may also be added to the denominator.
[0077] Like standard gradient boosting, this tree is then used to make predictions that are scaled by a learning rate. The residuals from these predictions are then used to construct the next tree and so on. The tree construction ends when a predetermined maximum number of trees is constructed or when the residuals become smaller than a predetermined threshold.
[0078] When the output variable takes one of a discrete number of values (e.g., for classification), XGBoost maps these to a numerical range. For example, each value of a boolean output variable would be mapped to 1.0 or 0.0. The splits are then performed as described above.
[0079] However, a different similar score calculation is used for each node, and this one is the sum of the squares of the residuals in the node divided by the sum over all observations. The product of (i) the previous probability and (ii) one minus the previous probability. A normalization constant may also be added to the denominator. The tree construction is also performed as described above, but this different similarity score calculation is used to determine the gain.
[0080] As described above, XGBoost may require that the minimum number of residuals is represented at each leaf. However, in this version of XGBoost, instead of counting the residuals, a value called "cover" is used. The cover is the value obtained by subtracting the normalization constant from the denominator of the similarity score. Leaves with a cover below a threshold may be removed from the tree, effectively pruning the tree. Other pruning techniques described above may also be used.
[0081] For prediction, the output value of a leaf is the sum of the residuals in the leaf divided by the sum over all observations. The product of (i) the previous probability and (ii) one minus the previous probability. A normalization constant may also be added to the denominator. Across multiple trees, the prediction for an observation is the natural logarithm of the odds of the initial output value, added to the output values for each tree scaled by the learning rate.
[0082] The logistic function may be applied to this result to convert it back to a probability. Based on this probability value (e.g., above or below 0.5 for a binary output variable), the value of the output variable may be selected.
[0083] XGBoost also uses a number of techniques to improve its processing speed with respect to large learning datasets. These techniques include an approximate greedy algorithm for selecting splits, a weighted sketch algorithm for focusing on hard-to-predict observations, using distributed learning across multiple processors or computers, and / or maintaining commonly used variables and constants in the processor cache. Other techniques may also be applied.
[0084] B. LightGBM LightGBM also uses gradient boosting, but generally does so in a way that increases the learning speed, reduces memory usage, and provides improved accuracy. In particular, rather than considering all values of the input variables, LightGBM bins these values to form a histogram and operates on the bins rather than the values. Also, LightGBM uses exclusive feature bundling to reduce the dimensionality of the feature space when two or more features tend to take mutually exclusive values. Further, LightGBM uses gradient-based one-sided sampling to identify the observations with the largest residuals and operate only on those observations, and also uses random sampling of observations with lower residuals.
[0085] As a result, LightGBM focuses its computations on the most needed cases, i.e., computations involving input variables that are dissimilar to each other and observations related to the most error-prone initial trees. Thus, in practice, LightGBM can execute about 10 times faster than other gradient boosting implementations with similar accuracy.
[0086] III. Prediction of Albuminuria As described above, albuminuria is a condition in which albumin is detected in a patient's urine and indicates chronic kidney disease (CKD). Generally, since urine tests are not ordered for most patients, CKD can progress undiagnosed and ultimately lead to kidney failure. Urine albumin-to-creatinine ratio (UACR) tests with results in the range of 30 - 300 mg / g are called microalbuminuria. Results greater than 300 mg / g are called macroalbuminuria. Albuminuria is also indicative of other health conditions, including but not limited to cardiovascular disease, hypertension, systemic vasculitis, and / or diabetes. Non-limiting examples of cardiovascular disease include heart failure such as heart failure with preserved ejection fraction (HFpEF), heart failure with mid-range ejection fraction (HFmrEF), and heart failure with reduced ejection fraction, and major adverse cardiac events (MACE) including myocardial infarction, stroke, and cardiovascular death, but are not limited thereto. Additionally, glomerular diseases may also be similarly associated. Generally, proteinuria or albuminuria is associated with an increased risk of progression to end-stage renal disease (ESRD) and overall mortality.
[0087] In addition to UACR testing, estimated glomerular filtration rate (EGFR), as indicated by a serum creatinine blood test, is also a biomarker for CKD (i.e., EGFR can be derived from serum / plasma creatinine). Thus, EGFR is generally tested regularly for most patients, usually in a basic metabolic panel. EGFR calculation is also based on a patient's age, sex, height, weight, and ethnicity, although other factors may be included. EGFR may be measured in milliliters per minute of purified blood per body surface area (mL / min / 1.73m2), and higher measurements generally indicate healthier kidney function.
[0088] When a patient exhibits a low EGFR over a predetermined time period, typically, a urine test to determine UACR is indicated. However, the interaction between EGFR and UACR is complex, and typically only 30% - 40% of patients screened for albuminuria actually exhibit the condition. For example, as shown in FIG. 4, patients with an EGFR in the range of 60 - 90 mL / min / 1.73 m2 are at low risk (defined as an EGFR of less than 60 mL / min / 2.73 m2) for the progression of CKD if their UACR is less than 30 mg / g, but are at high risk if their UACR is 300 mg / g or more. As a result, considering EGFR alone does not accurately predict the risk of CKD.
[0089] Identifying patients with undiagnosed albuminuria based on readily available information (e.g., patient demographics, vital signs, blood tests, and / or other medical information) who do not have recent or any available UACR results has numerous advantages. First, screening for albuminuria using urine tests can be made more effective by selecting patients who are likely to exhibit the condition. Further, these patients, if ultimately diagnosed with albuminuria, can receive treatment at an earlier stage and before CKD progresses to renal failure. Additionally, diagnosed patients are natural candidates to participate in clinical trials of new treatments and / or medications.
[0090] For example, clinical trials may be designed around patients who exhibit (i) an EFGR between 20 - 30 mL / min / 1.73 m2 and a UACR between 30 - 5000 mg / g, or (ii) an EFGR of at least 30 mL / min / 1.73 m2 and a UACR between 200 - 5000 mg / g. However, other criteria for participation may be used that include different UACR and EFGR ranges.
[0091] Accordingly, embodiments of the present application may include developing a machine learning model that can identify patients who may have undiagnosed albuminuria from an electronic health record corpus or other data sources. This model may then be validated, for example, based on clinical data. Once validated, the model may be used to identify patients in hospitals and other facilities who have a high risk of CKD, an increased rate of declining eGFR, cardiovascular events, and kidney events. These patients may be recommended for further testing, treatment, and / or participation in clinical trials.
[0092] Figure 5 shows how the model makes predictions. Demographic data 500 (e.g., age, gender, ethnicity), vital sign data 502 (e.g., body mass index, blood pressure, heart rate), and blood test data 504 (see below) are provided as inputs to the machine learning model 506. Other data such as comorbidities and administered drug therapies may be included in the input.
[0093] The machine learning model 506 may be a classification model (e.g., classifying patients into categories such as those who may have undiagnosed albuminuria and those who may not have the condition) as shown, or it may be a regression model (e.g., predicting a specific UACR value for each patient). Also, the machine learning model 506 may be based on LightGBM as shown, or it may be based on XGBoost or some other gradient boosting technique. In an alternative embodiment, the machine learning model 506 may be at least partially based on an artificial neural network, an expert system, or some ensemble combination of any of these models. Accordingly, the machine learning model 506 may generate a prediction 508, which may be a classification or value determined by regression.
[0094] Data for training and validating the machine learning model 506 may come from a variety of sources, including hospitals, healthcare providers, clinical sources, and / or insurance claims. For example, insurance claim data may include demographics, vital signs, and blood test results for patients with and without albuminuria, so the machine learning model 506 may be trained on insurance claim data. This training data may be preprocessed in various ways, such as removing outliers, removing skewness, and / or normalizing.
[0095] The trained machine learning model 506 may then be validated against clinical data to determine how accurately it predicts a patient's albuminuria. Once validated, the machine learning model 506 may be applied to identify patients in a hospital using hospital services or patients in other clinical or primary care facilities who are candidates for further testing, treatment, or participation in a clinical trial.
[0096] In some cases, training data may be collected from multiple geographic regions. However, the data may be segmented into regions to develop region-specific models. In some situations, information from identified patients may be checked for novelty, e.g., with respect to whether the input variables for these patients are consistent with those in the training data. To do so, a similarity model may be applied to information from the identified patients and the training data, and dissimilar patients are identified for further processing before the prediction is completed. This can be useful when a machine learning model is trained on data from a given population of individuals (e.g., located in North America) but applied to other given populations of individuals (e.g., located in Europe).
[0097] Exemplary candidate features are shown in FIG. 6. Demographics and vital signs 600 include age, gender, race, body mass index (BMI), blood pressure, and heart rate. Blood tests 602 include a number of possible measurements, including but not limited to albumin, calcium, cholesterol, EGFR, glucose, hemoglobin, magnesium, sodium, triglycerides, white blood cell (WBC) count, etc. Other demographics, vital signs, and blood tests may be used. The output variable (not shown) for this training dataset will be the UACR test result (e.g., either a numerical UACR value or an indication of whether the patient from whom the observed value was derived was diagnosed with albuminuria). For example, in order to be included in the training dataset, each observed value may require a UACR test recorded in the past five years.
[0098] Also, some of this data may be sparse in that 20% to 50% of the input variable values are missing across the dataset. As described above, gradient boosting models typically work well on sparse datasets.
[0099] In particular, techniques the same as or similar to those described in the present application may also be used to predict urinary proteinuria (UPR). Proteinuria is an increased level of protein in the urine and can also be an indicator of CKD. Normal urinary protein values in healthy adults are lower than 150 mg / 24 hours. Mild proteinuria is typically in the range of 150 mg / 24 hours to 2000 mg / 24 hours, and severe proteinuria is typically in the range of 2000 mg / 24 hours to 4000 mg / 24 hours. Similar to albuminuria, these cutoffs may vary or change. UACR is usually better than UPR in predicting CKD, but UPR is still useful and in some areas is a more commonly ordered test.
[0100] Also, the UPR value may be converted to an estimated UACR value using an established equation (see Non-Patent Document 1). Therefore, the UACR value used in the present application may be derived from the UPR value, or the UPR value may be used instead of the UACR value.
[0101] IV. Exemplary Learning Data For illustrative purposes, two sources of learning data are considered. The first is the Limited Claims and Electronic Health Record (LCED) dataset, and the second is the Optum Clinformatics Data Mart (Optum) dataset. Both are from patients based in the United States and consist of at least hundreds of thousands of observations. Nevertheless, different learning datasets may be used.
[0102] FIG. 7 presents an overview of these datasets. The overview 700 of the LCED dataset indicates that it contains 268,605 observations from 104,272 patients. Approximately 19.7% of these observations show microalbuminuria, 8.2% show macroalbuminuria, and the rest show normal values. Similarly, the overview 702 of the Optum dataset indicates that it contains 7.7 million observations from 1.7 million patients. Approximately 22.5% of these observations show microalbuminuria, 11.2% show macroalbuminuria, and the rest show normal values.
[0103] Figure 8 shows the distribution of UACR test results for the patient's LCED and Optum cohorts. As expected, most patients exhibit low UACR within the normal range. Since the ends of the distribution are heavy, Figure 8 also provides each log transformation to provide a better sense of the entire range of the distribution. Since the distortion of the log transformation of these datasets is smaller, the machine learning model may be trained on the log transformation rather than the raw data. Thus, the machine learning model may predict the log transformation of the UACR value, which may then be returned to the UACR value using exponentiation.
[0104] As noted, the machine learning model of embodiments of the present application may be based on regression (predicting a specific UACR value) or may be based on classification (predicting whether microalbuminuria, macroalbuminuria, or either is present). In particular, a regression-based model can be easily used for classification by mapping the predicted UACR values to "true" or "false" based on the side of the cutoff value (e.g., 30 mg / g) to which they apply.
[0105] V. Exemplary Model Results Figure 9 shows a receiver operating characteristic (ROC) curve 900 for the learned model assuming that the classification of macroalbuminuria was used (when at least 300 UACRs indicate macroalbuminuria). The ROC curve 900 plots the true positive rate (the number of true positives divided by the sum of true positives and false negatives) against the false positive rate (the number of false positives divided by the sum of false positives and true negatives) of the model.
[0106] In a diagnostic design, if the goal is to identify all patients with a given condition, the number of false negatives should be small, and thus the true positive rate should be high. However, if the goal is to identify only patients with a given condition, the number of false positives should be small, and thus the false positive rate should also be low. Setting various model parameters (including the cut-off value of UACR that divides the "true" and "false" classifications) can affect both the true positive rate and the false positive rate. Note that the dashed line represents a model that behaves equivalently to random guessing.
[0107] The ROC curve visually explains the trade-off between the true positive rate and the false positive rate. One measure of model quality is the area under the curve (AUC) with respect to the ROC curve. This value is typically between 0.5 and 1.0, and a higher value indicates better model performance across various parameter settings. For example, as shown in ROC curve 900, the AUC is approximately 0.801. This indicates a model that operates moderately well.
[0108] Figure 9 also shows the precision-recall curve 902 for the model, again assuming that the classification of macroalbuminuria was used. The precision-recall curve 902 plots the precision (the number of true positives divided by the sum of true positives and false positives) against the recall (the number of true positives divided by the sum of true positives and false negatives) of the model. Setting various model parameters (including the cut-off value of UACR that divides the "true" and "false" classifications) can affect both the precision and the recall. Note that the dashed line represents a model that behaves equivalently to random guessing.
[0109] The precision-recall curve visually explains the trade-off between precision and recall and is particularly useful for datasets where there are significantly more observations in one class than the other. Again, the AUC may be used to evaluate model quality, with higher values indicating better quality. For example, as shown in the precision-recall curve 902, the AUC is approximately 0.369. This indicates a model that performs significantly better than random guessing, which has an AUC of approximately 0.08.
[0110] Similar to FIG. 9, FIG. 10 shows the ROC curve 1000 for the model, assuming that the classification of microalbuminuria was used (where at least 30 UACRs indicate microalbuminuria). FIG. 10 also shows the precision-recall curve 1002 for the model, again assuming that the classification of macroalbuminuria was used. Again, the AUCs of these curves (approximately 0.740 and 0.562) indicate moderately good model performance compared to random guessing.
[0111] To facilitate further understanding of the learned model, FIGS. 11A and 11B provide the Shapley values 1100 and 1102 of the model learned with respect to macroalbuminuria and microalbuminuria, respectively. Briefly, given a particular output of the model (e.g., the predicted UACR) for a particular set of input variable values and the mean of this output across all input values, the Shapley data assigns a contribution to each of the input variables. This contribution quantifies how much each input variable contributed to the difference between the predicted output value and the mean output value. The Shapley data also accounts for possible interdependencies between input variables such that the Shapley data is independent with respect to the order in which the input variables are applied (if the model is sensitive to such ordering).
[0112] Regarding the UACR prediction model described in the present application, the Shapley value for each input variable may be used to quantify the influence of that input variable on the difference between the predicted UACR of a particular observation and the average UACR of all observations. Some input variables may correlate more highly with the predicted UACR of a particular observation, while others may correlate more weakly with the predicted UACR of a particular observation.
[0113] At Shapley values 1100 and 1102, scatter plots are shown along the corresponding X-axis for each input variable. Each value in the scatter plot corresponds to one of the observations considered by the model. The X-axis for each input variable represents the difference between the predicted UACR and the average UACR, centered at 0. The input variables at the top of the list generally have a greater influence on the predicted UACR than the input variables lower in the list (e.g., have a greater positive and / or negative correlation). In some embodiments, the input variables may be sorted top to bottom in descending order with respect to the magnitude of this influence.
[0114] For example, the three most influential input variables in both Shapley value 1100 (relating to macroalbuminuria) and Shapley value 1102 (relating to microalbuminuria) are creatinine, systolic blood pressure, and HbA1c (glycated hemoglobin), and are shown to have a greater influence on the predicted UACR than any other input variable. On the other hand, age has a large influence when predicting microalbuminuria (the fourth most influential input variable), while having insignificantly little influence when predicting macroalbuminuria (the sixteenth most influential input variable). Similar scatter plots for other input variables show how, and approximately to what extent, they have a positive or negative influence on UACR. These results also show that EGFR alone is not a reliable predictor of albuminuria, which further justifies this more sophisticated modeling approach.
[0115] The advantage of representing SHAP values in this way is that it can provide a detailed understanding, at a quick glance, of why the model predicted a particular UACR for a particular observation. For example, the fact that age is significantly more useful for predicting microalbuminuria than macroalbuminuria may initially seem counterintuitive, and such results can be further investigated. Nevertheless, other types of graphs and visualizations of SHAP data may also be possible.
[0116] In some cases, in addition to albuminuria prediction, selected SHAP values and other data may be provided to healthcare professionals or clinicians. These values can help explain which input variables had the most impact on the prediction. For example, in initial use, the model described in this application determined that the top 10 most influential features (in descending order) were creatinine (in blood / serum), systolic blood pressure, hemoglobin A1C (HbA1C), BMI, EGFR, albumin (in blood / serum), triglyceride, glucose, age, and bilirubin. These features were identified by a combination of SHAP analysis, frequency in the dataset, and importance with respect to splits / gains when constructing decision trees.
[0117] Figures 12A and 12B provide ROC graphs 1200 and 1202 for at least 30 UACRs and at least 200 UACRs, respectively, which can be used as indicators of microalbuminuria and macroalbuminuria. Each of graphs 1202 and 1204 is a model trained using a learning dataset derived from the United States, and plots data regarding each model used to predict albuminuria for the US datasets (LCED and Optum) and anonymized non-US datasets (non-US1, non-US2, and non-US3). As shown in Figure 12A, learning on US data results in a model that works well for predicting microalbuminuria for both the US datasets (AUC of 0.71 - 0.73) and the non-US datasets (AUC of 0.65 - 0.68). As shown in Figure 12B, learning on US data results in a model that works well for predicting macroalbuminuria for both the US datasets (AUC of 0.78 - 0.81) and the non-US datasets (AUC of 0.69 - 0.75).
[0118] Figure 12C shows table 1204 which provides AUC values for among other factors, each race, age cohort, and gender. For the latter, M indicates male, F indicates female, and U indicates unknown. All of these AUC values indicate that the trained model works well across multiple races, multiple age cohorts, and multiple genders, although there is some variability in terms of AUC within each. Figure 12C also evaluates the model with respect to the positive prediction value (PPV), which is another term for the precision (number of true positives divided by the sum of true and false positives). As shown, the model also works well with respect to PPV.
[0119] FIG. 12D shows Table 1206 for providing AUC and PPV factors for different ranges of EGFR and HbA1C. As shown, the model works very well, even when the EGFR is unknown in particular. This suggests that the model has important clinical utility for predicting albuminuria in populations where EFGR measurements are not available. Further, the model operates with the same level of accuracy regardless of the HbA1C level.
[0120] VI. Deployment Scenarios Embodiments of the present application may be deployed in a number of configurations. With respect to the training of one or more machine learning models, this may be performed on one or more computing devices within a server cluster, such as in a public cloud network (e.g., Amazon AWS® or Microsoft Azure®) or a private system. With respect to the execution of these models on new observations, the models may be hosted in various locations and environments.
[0121] In one possible example, the trained model may be hosted in a public cloud network or a private network and may provide results to a client device via a web or application interface. For example, the client device may send a request containing a new set of input variables including new observations to the remotely hosted model. The model may take these as input and then generate a corresponding output, which is then sent to the client device in response to the request. These trained models may be operated by various entities, such as a hospital, a hospital network, a physician, a physician network, a university, a pharmaceutical company, or some consortium involving one or more of these or other entities.
[0122] Alternatively, the learned model may be packaged into a client application that can be downloaded and installed on a desktop, laptop, or mobile computing device. Thus, the client application will include a user interface that enables the user to input or otherwise indicate input variables regarding new observations. The client application will then apply the model to these new observations and generate corresponding outputs that are displayed and / or stored by the client device. This scenario has the advantage that a live network connection is not required to use the model.
[0123] In another option, the learned model may be used to develop simple clinical prediction rules, such as decision trees, that medical practitioners may follow to predict whether a patient may have albuminuria.
[0124] Other deployment scenarios may exist. Thus, embodiments of the present application are not limited to these scenarios.
[0125] VII. Exemplary Operations FIGS. 13 and 14 are flowcharts showing exemplary embodiments. The operations shown in FIGS. 13 and 14 may be performed by a computing system or computing device that includes a software application configured to perform any of the embodiments of the present application. Non-limiting examples of the computing system or computing device include, for example, computing device 100 or server cluster 200. However, the operations may be performed by other types of devices or device subsystems. For example, the operations may be performed by a portable computer such as a laptop or tablet device.
[0126] The embodiments of FIGS. 13 and 14 may be simplified by removing any one or more of the features shown therein. Further, these embodiments may be combined with any of the previous drawings or otherwise with the features, aspects, and / or implementations described in this application. Such embodiments may include instructions executable by one or more processors of one or more computing devices of a system or virtual machine or container. For example, the instructions may take the form of software, hardware, and / or firmware instructions. In an exemplary embodiment, the instructions may be stored on a non-transitory computer-readable medium. When executed by one or more processors of one or more computing devices, the instructions may cause the one or more computing devices to perform various operations of the embodiment.
[0127] Block 1300 of FIG. 13 includes obtaining a training dataset by a computing system. The training dataset includes corresponding demographic values for a plurality of individuals, vital sign values, blood test values, and observed values of either a UACR value or a UPR value.
[0128] Block 1302 of FIG. 13 also includes applying a machine learning trainer to the training dataset by the computing system. The machine learning trainer generates a machine learning model. The machine learning model is configured to receive new observed values of new demographic values, new vital sign values, and new blood test values as inputs and provide a prediction as to whether an individual presenting the new observed values has undiagnosed albuminuria or proteinuria.
[0129] In some embodiments, the demographic values include the age, gender, or ethnicity of a plurality of individuals.
[0130] In some embodiments, the vital sign values include the body mass index, blood pressure readings, or heart rate of a plurality of individuals.
[0131] In some embodiments, the blood test values include creatinine levels, glycated hemoglobin levels, triglycerides, blood albumin levels, or white blood cell counts of multiple individuals.
[0132] In some embodiments, the values within the training dataset include 20% to 50%.
[0133] In some embodiments, the machine learning model is based on gradient boosting.
[0134] In some embodiments, predicting whether an individual presenting with new observation values has undiagnosed albuminuria includes predicting whether the individual has microalbuminuria.
[0135] In some embodiments, predicting whether an individual presenting with new observation values has undiagnosed albuminuria includes predicting whether the individual has macroalbuminuria.
[0136] In some embodiments, predicting whether an individual presenting with new observation values has undiagnosed albuminuria or proteinuria includes predicting the UACR value or UPR value for the individual.
[0137] In some embodiments, the training dataset includes at least 100,000 observations collected from medical claims records or electronic health records.
[0138] In some embodiments, the training dataset includes at least 1,000,000 observations collected from medical claims records or electronic health records.
[0139] In some embodiments, 5% to 25% of the observations have a UACR value indicating albuminuria or a UPR value indicating proteinuria.
[0140] In some embodiments, the UACR value is mathematically derived from the UPR value.
[0141] Block 1400 in FIG. 14 includes obtaining, by a computing system, observed values of an individual's demographic values, an individual's vital sign values, and an individual's blood test values.
[0142] Block 1402 in FIG. 14 also includes applying, by a computing system, a machine learning model to the observed values. The machine learning model is trained by a training dataset. The training dataset includes observed values of corresponding demographic values, vital sign values, blood test values, and UACR for a plurality of individuals. The machine learning model is configured to provide a prediction as to whether additional observed values indicate undiagnosed albuminuria or proteinuria.
[0143] Block 1404 in FIG. 14 also includes providing, by a computing system, a prediction as to whether an individual exhibits undiagnosed albuminuria or proteinuria based on the observed values.
[0144] In some embodiments, providing the prediction includes displaying the prediction in a graphical user interface.
[0145] In some embodiments, obtaining the observed values includes receiving the observed values from a client device that communicates with the computing system via a network. Providing the prediction includes transmitting the prediction to the client device.
[0146] In some embodiments, the demographic values include the age, gender, or ethnicity of a plurality of individuals.
[0147] In some embodiments, the vital sign values include the body mass index, blood pressure readings, or heart rate of a plurality of individuals.
[0148] In some embodiments, the blood test values include the creatinine level, hemoglobin A1c level, triglycerides, blood albumin level, or white blood cell count of a plurality of individuals.
[0149] In some embodiments, the values in the training dataset include those ranging from 20% to 50%.
[0150] In some embodiments, the machine learning model is based on gradient boosting.
[0151] In some embodiments, predicting whether an individual presenting with observed values has undiagnosed albuminuria includes predicting whether the individual has microalbuminuria.
[0152] In some embodiments, predicting whether an individual presenting with observed values has undiagnosed albuminuria includes predicting whether the individual has macroalbuminuria.
[0153] In some embodiments, predicting whether an individual presenting with observed values has undiagnosed albuminuria or proteinuria includes predicting the UACR value or UPR value for the individual.
[0154] In some embodiments, the training dataset includes at least 100,000 observed values collected from medical claim records or electronic health records.
[0155] In some embodiments, the training dataset includes at least 1,000,000 observed values collected from medical claim records or electronic health records.
[0156] In some embodiments, 5% to 25% of the observed values have a UACR value indicating albuminuria or a UPR value indicating proteinuria.
[0157] In some embodiments, the UACR value is mathematically derived from the UPR value.
[0158] VIII. Quantile Regression Model for Predicting Albuminuria Instead of, or in addition to, the above-described embodiments, a quantile-based regression model may be used to predict albuminuria. Unlike a classification model that predicts whether a patient has microalbuminuria or macroalbuminuria, this regression model predicts the 25th, 50th, and / or 75th quantiles of UACR or UPR. Other quantiles, such as the 10th and 90th, may be predicted.
[0159] Figure 15 shows a regression-based model. This model is similar to that of Figure 5 but has some notable differences.
[0160] Demographic data 1500 (here, simply age and gender), vital sign data 1502 (here, simply body mass index and systolic blood pressure), and blood test data 1504 are provided as inputs to a machine learning model 1506. Other data, such as comorbidities and administered drug therapies, may be included in the input. Additionally, additional data, such as that described in the context of Figure 5, may be included.
[0161] In the case of this model, the blood test data 1504 focuses on the levels of albumin, bilirubin, creatinine, HbA1C, triglycerides, glucose, white blood cell count, and ALT. Other combinations of markers may be used.
[0162] The machine learning model 1506 may be a quantile regression model as illustrated. Also, the machine learning model 1506 may be based on LightGBM as illustrated, or it may be based on XGBoost or some other gradient boosting technique. In alternative embodiments, the machine learning model 1506 may be at least partially based on an artificial neural network, an expert system, or some ensemble combination of any of these models. Thus, the machine learning model 1506 may generate a prediction 1508, which may be one or more quantiles of UACR or UPR.
[0163] Quantile regression is a statistical technique used to estimate the relationship between one or more input variables and an output variable across one or more quantiles of the output variable. Unlike ordinary least squares regression, which estimates the mean of the output variable as a function of the input variables, quantile regression estimates the conditional quantiles of the output variable. For example, the model may be configured to estimate the 25th, 50th, and 75th percentiles of the output variable for different values of the input variables. This makes it possible to determine how the relationship between the input and output variables changes across different parts of the distribution of the output variable.
[0164] Such a model may operate at least in part by minimizing a loss function that penalizes the difference between the predicted and observed quantiles. For example, a quantile regression model for quantile τ is given by the following equation.
[0165]
Equation
[0166] In this equation, the β j (τ) coefficients are functions of the quantiles, not constants. Finding the values of these coefficients for a particular quantile is similar to that in linear regression, except that the median absolute deviation (MAD) is minimized. In particular, the following equation holds.
[0167]
Equation
[0168] The function ρ τ is defined by the following equation.
[0169]
Equation
[0170] Therefore, ρ τDepending on the quantile and the overall sign of the error, it gives an asymmetric weight to the error. This means that when the error is positive, ρ τ multiplies the error by τ, and when the error is negative, ρ τ multiplies the error by (1 - τ). For example, to determine the median of the 25th quantile, 75% of the errors should be positive and 25% should be negative. Weights are added to the errors to find the minimum MAD for which this property is true. For the 25th quantile, a weight of 0.75 is added to the negative errors and a weight of 0.25 is added to the positive errors.
[0171] This quantile regression technique acts to explain the relationship between the input and output variables at each quantile and can be used to make predictions about the output variable at a specific quantile for a new value of the input variable. Thus, quantile regression is a useful tool for examining how the relationship between the input and output variables changes across the distribution of the output variable.
[0172] Similar to the case of the machine learning model 506, the data for training and validating the machine learning model 1506 may come from various sources including hospitals, healthcare providers, clinical sources, and / or insurance claims. For example, since insurance claim data can include demographics, vital signs, and blood test results for patients with and without albuminuria, the machine learning model 1506 may be trained on insurance claim data. This training data may be pre - processed in various ways, such as removing outliers, removing skewness, and / or normalizing.
[0173] The trained machine learning model 1506 may then be validated against clinical data to determine how accurately it predicts the quantiles of a patient's UACR or UPR. Once validated, the machine learning model 1506 may be applied to identify patients in a hospital using hospital services or in other clinical or primary care facilities who are candidates for further testing, treatment, or participation in clinical trials.
[0174] Compared with the machine learning model 506, the machine learning model 1506 predicts the quantiles of UACR or UPR rather than classifying patients with or without microalbuminuria or macroalbuminuria. Thus, the user may select the quantiles and thresholds to be used to classify patients into multiple risk groups. For example, when using the 25th quantile, there is a 75% probability that the actual value will be higher than the predicted value, which means high precision. Next, the predicted UACR or UPR value at the quantile is compared to 300 macroalbuminuria thresholds to determine whether the patient is eligible for treatment or research participation. In contrast, using the 50th quantile decreases the precision but increases the recall, and thus it may be more appropriate for different purposes, such as in the case of very high UACR or UPR values. The difference between the quantile predictions also provides an estimate of how reliable it is that the model is in the process of making a prediction (e.g., a smaller difference means a higher confidence level).
[0175] FIG. 16 provides an example of the results that can be generated by such a system. The graphs in this figure show the number of survival days on the X-axis and the probability on the Y-axis for various situations. The dashed lines relate to individuals who do not actually have albuminuria (e.g., UACR ≤ 30), and the solid lines relate to individuals predicted to be without albuminuria (e.g., the median predicted to have UACR ≤ 30). The dashed lines with circles relate to individuals who actually have microalbuminuria (e.g., 30 < UACR ≤ 300), and the solid lines with circles relate to individuals predicted to have microalbuminuria (e.g., the median predicted to have 30 < UACR ≤ 300). The dashed lines with triangles relate to individuals who actually have macroalbuminuria (e.g., 300 < UACR), and the solid lines with triangles relate to individuals predicted to have macroalbuminuria (e.g., the median predicted to have 300 < UACR).
[0176] As can be seen from FIG. 16, the predicted median UACR generally agrees with the survival rate with respect to the corresponding true UACR value. In these curves, the estimated survival (e.g., Kaplan-Meier estimate) is plotted for patients who had a urine test (these are the dashed lines of the "true values"). The estimated survival for patients identified by the model as having no albuminuria, microalbuminuria, or macroalbuminuria is also plotted based on the predicted quantiles. However, due to the nature of quantile regression, predictions are provided for cases where the model has the most certain values (e.g., for the 25% or 50% quantiles), and thus patients with uncertain values are not included in the estimates. The model then identifies patients with the worst renal function, which becomes clearer in clinical tests, and these patients also have the worst survival rates. This is why the predicted survival for patients differs from the true survival rate. The reason is that they simply represent a subset of all patients in that category of albuminuria who may have the worst renal function. Nevertheless, for the main use cases, these are the desired survival curves.
[0177] IX. Additional Exemplary Operations FIGS. 17 and 18 are flowcharts showing exemplary embodiments. The operations shown in FIGS. 17 and 18 may be performed by a computing system or computing device including a software application configured to perform any of the embodiments of the present application. Non-limiting examples of the computing system or computing device include, for example, computing device 100 or server cluster 200. However, the operations may be performed by other types of devices or device subsystems. For example, the operations may be performed by a portable computer such as a laptop or tablet device.
[0178] The embodiments of FIGS. 17 and 18 may be simplified by removing any one or more of the features shown therein. Further, these embodiments may be combined with any of the previous drawings or otherwise with the features, aspects, and / or implementations described in this application. In particular, these embodiments may be used with any relevant features described in the context of FIGS. 13 and 14.
[0179] Such embodiments may include instructions executable by one or more processors of one or more computing devices of a system or virtual machine or container. For example, the instructions may take the form of software, hardware, and / or firmware instructions. In an exemplary embodiment, the instructions may be stored on a non-transitory computer-readable medium. When executed by one or more processors of one or more computing devices, the instructions may cause the one or more computing devices to perform various operations of the embodiment.
[0180] Block 1700 of FIG. 17 includes obtaining a training dataset by a computing system. The training dataset includes corresponding demographic values, vital sign values, blood test values, and observed values of either urinary albumin to creatinine ratio (UACR) values and urinary protein to creatinine ratio (UPR) values for a plurality of individuals.
[0181] Block 1702 of FIG. 17 also includes applying a quantile regression machine learning trainer to the training dataset by the computing system. The quantile regression machine learning trainer generates a quantile regression machine learning model. The quantile regression machine learning model is configured to receive as input quantiles and new observed values of new demographic values, new vital sign values, and new blood test values, and provide a prediction of UACR or UPR values at the quantiles for an individual presenting the new observed values.
[0182] Block 1800 of FIG. 18 includes obtaining, by a computing system, demographic values of an individual, vital sign values of the individual, and quantiles and observed values of blood test values of the individual.
[0183] Block 1802 of FIG. 18 also includes applying, by a computing system, a quantile regression machine learning model to the observed values. The quantile regression machine learning model is trained by a training dataset. The training dataset includes corresponding demographic values, vital sign values, blood test values, and observed values of either urinary albumin to creatinine ratio (UACR) values and urinary protein to creatinine ratio (UPR) values for a plurality of individuals. The quantile regression machine learning model is configured to provide a prediction of UACR or UPR values at one or more quantiles for further observed values.
[0184] Block 1804 of FIG. 18 also includes providing, by a computing system, a prediction of UACR or UPR values at a quantile based on the observed values and for an individual.
[0185] X. Exemplary User Interfaces and Workflows FIGS. 19A and 19B show exemplary workflows for software that executes a model for predicting factors related to albuminuria using inputs from a health care provider (HCP) or other user. The software includes a user interface and a front end of an application, which may include a web-based or stand-alone application (e.g., a mobile app). The software also includes a back-end service (e.g., located on a remote server) that executes authentication and model execution functions. Nevertheless, the model may also include other features and functions not shown in these drawings.
[0186] Referring to FIG. 19A, at step 1900, the HCP navigates to the web page or application of the albuminuria screening tool. The HCP is presented with screen 1902 in the user interface. Screen 1902 includes a text box where the HCP can enter their security credentials (e.g., user id and password).
[0187] After activating the "Sign On" button on screen 1902, the HCP's credentials (or some representation thereof) are sent to the back-end service. In response, the back-end service executes step 1904 of authenticating the HCP.
[0188] Assuming this authentication is successful, the HCP is presented with screen 1906 in the user interface. Screen 1906 requests that the HCP confirm their data entry to verify that the patient has provided informed consent. Assuming this is the case, the HCP activates the "OK" button and control proceeds to screen 1912 in FIG. 19B.
[0189] Referring to FIG. 19B, at step 1910, the HCP enters parameters from the patient's data into the text boxes on screen 1912. As shown, these parameters may include the patient's age, BMI, systolic blood pressure, HbA1C, creatinine, triglyceride, glucose, albumin, and WBC value. Nevertheless, a greater or lesser number of parameters may be entered, and other parameters may be used.
[0190] After the HCP activates the "OK" button, step 1914 is executed by the back-end service. In particular, one or more of the albuminuria prediction models described in the present application are executed against the parameters. As described above, these models may be used to predict whether the patient has microalbuminuria or macroalbuminuria, or neither.
[0191] Accordingly, the output of the model is made to occupy screen 1916. This screen may display, for example, whether the patient is predicted to have microalbuminuria or macroalbuminuria. The screen may also display the patient's predicted UACR.
[0192] In addition to the features and functions described, the front-end and back-end services may also support the HCP to log out of the application, help screens, etc.
[0193] XI. Conclusion The present disclosure should not be limited to the specific embodiments described in this application, which are intended to be examples of various aspects. As will be apparent to those skilled in the art, numerous modifications and changes may be made without departing from its scope. In addition to what is described in this application, functionally equivalent methods and apparatuses within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and changes are intended to be included within the scope of the appended claims.
[0194] The detailed description above explains various features and operations of the disclosed systems, apparatuses, and methods with reference to the accompanying drawings. The exemplary embodiments described in this specification and the drawings are not meant to be limiting. Other embodiments may be utilized and other changes may be made without departing from the scope of the subject matter presented in this application. It will be readily understood that the aspects of the present disclosure generally described and illustrated in the drawings may be arranged, replaced, combined, separated, and designed in a variety of different configurations.
[0195] Regarding any or all of the message flow diagrams, scenarios, and flowcharts in the drawings and described in this application, each step, block m, and / or communication may represent the processing and / or transmission of information according to an exemplary embodiment. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, operations described as a plurality of steps, a plurality of blocks, a plurality of transmissions, a plurality of communications, a plurality of requests, a plurality of responses, and / or a plurality of messages may be performed in an order different from that illustrated or described, including substantially simultaneously or in a reverse order, depending on the relevant functions. Further, with any of the message flow diagrams, scenarios, and flowcharts described in this application, more or fewer blocks and / or operations may be used, and these message flow diagrams, scenarios, and flowcharts may be combined, in part or in whole, with each other.
[0196] A step or block representing the processing of information may correspond to a circuit configured to perform a specific logical function of the method or technology described in this application. Alternatively or additionally, a step or block representing the processing of information may correspond to a module, segment, or portion of program code (including related data). The program code may include one or more instructions executable by a processor to perform a specific logical operation or action in the method or technology. The program code and / or related data may be stored on any type of computer-readable medium, such as a storage device including RAM, a disk drive, a solid-state drive, or other storage media.
[0197] A computer-readable medium may also include non-transitory computer-readable media, such as register memory and processor caches, that store data for short periods of time. The non-transitory computer-readable media may further include non-transitory computer-readable media that store program code and / or data for longer periods of time. Thus, non-transitory computer-readable media may include, for example, secondary or persistent long-term storage devices such as ROM, optical or magnetic disks, solid state drives, or CD-ROM (compact disc read only memory). The non-transitory computer-readable media may also be any other volatile or non-volatile memory system. The non-transitory computer-readable media may be regarded, for example, as a computer-readable storage medium or a tangible storage device.
[0198] Also, one or more steps or blocks representing information transmission may correspond to information transmission between multiple software and / or hardware modules in the same physical device. However, other information transmissions may be between software modules and / or hardware modules in multiple different physical devices.
[0199] The specific configurations shown in the drawings should not be regarded as limitations. It should be understood that other embodiments may include more or fewer of each of the components shown in a given drawing. Further, some of the illustrated components may be combined or omitted. Still further, exemplary embodiments may include components not shown in the drawings.
[0200] Although various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and not limitation, and the true scope is indicated by the claims.
Claims
1. obtaining a training dataset by a computing system; applying a machine learning trainer to the training dataset by the computing system; the training dataset includes corresponding demographic values, vital sign values, blood test values, and observed values of either urinary albumin to creatinine ratio (UACR) values and urinary protein to creatinine ratio (UPR) values for a plurality of individuals; the machine learning trainer generates a machine learning model; the machine learning model is configured to receive, as input, new observed values of new demographic values, new vital sign values, and new blood test values, and provide a prediction as to whether an individual presenting the new observed values has undiagnosed albuminuria or proteinuria; A method.
2. the demographic values include the ages, genders, or ethnicities of the plurality of individuals; The method according to claim 1.
3. the vital sign values include the body mass index, blood pressure readings, or heart rates of the plurality of individuals; The method according to claim 1.
4. the blood test values include creatinine levels, glycated hemoglobin levels, triglycerides, blood albumin levels, or white blood cell counts of the plurality of individuals; The method according to claim 1.
5. the values within the training dataset are included in the range of 20% to 50%; The method according to claim 1.
6. the machine learning model is based on gradient boosting; The method according to claim 1.
7. Predicting whether an individual presenting the new observation value has undiagnosed albuminuria includes predicting whether the individual has microalbuminuria, The method according to claim 1.
8. Predicting whether an individual presenting the new observation value has undiagnosed albuminuria includes predicting whether the individual has macroalbuminuria, The method according to claim 1.
9. Predicting whether an individual presenting the new observation value has undiagnosed albuminuria or proteinuria includes predicting the UACR value or UPR value for the individual, The method according to claim 1.
10. The training dataset includes at least 100,000 observation values collected from medical claim records or electronic health records, The method according to claim 1.
11. The training dataset includes at least 1,000,000 observation values collected from medical claim records or electronic health records, The method according to claim 1.
12. 5% to 25% of the observation values have a UACR value indicating albuminuria or a UPR value indicating proteinuria, The method according to claim 1.
13. The UACR value is mathematically derived from the UPR value, The method according to claim 1.
14. A product including a non-transitory computer-readable medium storing program instructions that cause a computing system to perform any of the operations of claims 1 to 13 when executed by the computing system.
15. A computing system comprising one or more processors and a memory including program instructions, When the above program instructions are executed by the one or more processors, the computing system is caused to perform any of the operations of claims 1 to 13. Computing system.
16. Obtaining, by a computing system, observed values of an individual's demographic values, the individual's vital sign values, and the individual's blood test values; Applying, by the computing system, a machine learning model to the observed values; Providing, by the computing system, a prediction as to whether the individual exhibits undiagnosed albuminuria or proteinuria based on the observed values, The machine learning model is trained by a training dataset, The training dataset includes corresponding demographic values, vital sign values, blood test values, and observed values of either urinary albumin-to-creatinine ratio (UACR) values and urinary protein-to-creatinine ratio (UPR) values for a plurality of individuals, The machine learning model is configured to provide a prediction as to whether additional observed values indicate undiagnosed albuminuria or proteinuria. Method.
17. Providing the prediction includes displaying the prediction in a graphical user interface. The method according to claim 16.
18. Obtaining the observed values includes receiving the observed values from a client device that communicates with the computing system via a network, Providing the prediction includes transmitting the prediction to the client device. The method according to claim 16.
19. The demographic values include the ages, genders, or ethnicities of the plurality of individuals. The method according to claim 16.
20. The vital sign values include the obesity index, blood pressure readings, or heart rates of the plurality of individuals. The method according to claim 16.
21. The blood test values include creatinine levels, glycated hemoglobin levels, triglycerides, blood albumin levels, or white blood cell counts of the plurality of individuals. The method according to claim 16.
22. The values in the training dataset are included in the range of 20% to 50%. The method according to claim 16.
23. The machine learning model is based on gradient boosting. The method according to claim 16.
24. Predicting whether an individual presenting the observed values has undiagnosed albuminuria includes predicting whether the individual has microalbuminuria. The method according to claim 16.
25. Predicting whether an individual presenting the observed values has undiagnosed albuminuria includes predicting whether the individual has macroalbuminuria. The method according to claim 16.
26. Predicting whether an individual presenting the observed values has undiagnosed albuminuria or proteinuria includes predicting the UACR value or UPR value for the individual. The method according to claim 16.
27. The training dataset includes at least 100,000 observed values collected from medical claim records or electronic health records. The method according to claim 16.
28. The training dataset includes at least 1,000,000 observed values collected from medical claim records or electronic health records. The method according to claim 16.
29. Five percent to twenty-five percent of the above-mentioned observed values have a UACR value indicating albuminuria or a UPR value indicating proteinuria, The method according to claim 16.
30. Recommending that the individual receive treatment for albuminuria or proteinuria based on a prediction indicating that the individual exhibits undiagnosed albuminuria or proteinuria Further comprising The method according to claim 16.
31. Recommending that the individual enroll in a clinical trial related to albuminuria or proteinuria based on a prediction indicating that the individual exhibits undiagnosed albuminuria or proteinuria Further comprising The method according to claim 16.
32. The UACR value is mathematically derived from the UPR value, The method according to claim 16.
33. Obtaining a training dataset by a computing system, and Applying a quantile regression machine learning trainer to the training dataset by the computing system, The training dataset includes corresponding demographic values, vital sign values, blood test values, and observed values of either the urinary albumin-to-creatinine ratio (UACR) value or the urinary protein-to-creatinine ratio (UPR) value for a plurality of individuals, The quantile regression machine learning trainer generates a quantile regression machine learning model, The quantile regression machine learning model is configured to receive as input the quantiles and new observed values of new demographic values, new vital sign values, and new blood test values, and provide a prediction of the UACR or UPR value at the quantile for an individual presenting the new observed values. Method.
34. The computing system obtains the demographic values of an individual, the vital sign values of the individual, and the quantiles and observed values of the blood test values of the individual, applying a quantile regression machine learning model to the observed values by the computing system, providing, by the computing system, a prediction of the UACR or UPR value at the quantile based on the observed values and for the individual, wherein the quantile regression machine learning model is trained by a training dataset, the training dataset includes corresponding demographic values, vital sign values, blood test values, and observed values of either the urinary albumin-to-creatinine ratio (UACR) value or the urinary protein-to-creatinine ratio (UPR) value for a plurality of individuals, the quantile regression machine learning model is configured to provide a prediction of the UACR or UPR value at one or more quantiles for additional observed values, A method.
35. A non-transitory computer-readable medium storing program instructions that cause the computing system to perform any of the operations of claims 16 to 34 when executed by the computing system.
36. A computing system comprising one or more processors and a memory including program instructions, wherein the program instructions cause the computing system to perform any of the operations of claims 16 to 34 when executed by the one or more processors, A computing system.
Citation Information
Patent Citations
Modeling method for diabetic complication prediction model
CN111968748A
Diagnosis of human kidney disease
JP2002543386A
Method for determining treatment for cancer patients
JP2022505266A
Computer-based systems, computing components and computing objects configured to implement dynamic outlier bias reduction in machine learning models
US20210110313A1