Prediction of disease progression in portal hypertension using machine learning
A machine learning model predicts portal hypertension progression using demographic and health data, addressing the variability in disease progression and enabling timely, patient-specific treatment recommendations.
Patent Information
- Application Number
- JP2024555967
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-26
- Filing Date
- 2023-05-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-05-18
AI Technical Summary
Current medical knowledge on portal hypertension is limited, and existing treatments are ineffective for many patients, as the progression of the disease varies significantly among individuals, making it difficult for healthcare providers to recommend patient-specific treatments.
A machine learning model is developed using demographic, comorbidity, vital sign, and blood test values to predict the progression of portal hypertension, including the likelihood and timeframe for complications such as varices, ascites, and hepatic encephalopathy, enabling earlier identification of at-risk patients.
The model accurately predicts disease progression, allowing healthcare providers to make informed treatment recommendations, potentially improving patient outcomes by identifying high-risk individuals earlier.
Smart Images

Figure 2025519998000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 346,189, filed May 26, 2022, the entire disclosure of which is incorporated herein by reference.
Background Art
[0002] The portal system supplies blood to the liver. Portal hypertension is an increase in blood pressure within these veins and is typically caused by cirrhosis (scarring of the liver) or thrombosis (clotting). In some cases, portal hypertension can lead to varices (aneurysms) in the esophagus and / or stomach that are at risk of bleeding, resulting in life - threatening hemorrhage. Portal hypertension can also cause fluid retention (ascites) in the abdomen. Beta - blockers are a preferred treatment for portal hypertension but are effective in controlling the condition in less than half of patients. Other treatments such as liver shunts and liver transplants are more invasive.
[0003] For over 20 years, no large - scale epidemiological studies on portal hypertension have been published. Thus, the medical community's knowledge regarding the prognosis and clinical outcomes of patients with portal hypertension is limited.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Portal hypertension typically progresses from mild cases to clinically significant cases, further to severe cases, and ultimately to death. However, different patients with portal hypertension progress differently at different rates. Furthermore, the types of treatments that can help a patient can be very sensitive not only to the patient's current diagnosis but also to how the disease progresses. Thus, a diagnosis of portal hypertension alone does not provide a healthcare provider with sufficient information to successfully recommend patient - specific treatments.
[0005] Embodiments of this specification include the development, training, and use of a machine learning model for predicting disease progression in patients with portal hypertension. This predicted progression may indicate an outcome (e.g., varices, variceal bleeding, recurrent variceal bleeding, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, and / or jaundice, to name a few). The predicted progression may also provide an expected time until that outcome is reached. Thus, patients at risk of complications and / or death can be identified earlier and more accurately. Healthcare providers can then make more informed recommendations regarding treatment plans and their timing.
Means for Solving the Problems
[0006] Accordingly, a first exemplary embodiment is for a computing system to obtain a training dataset, where the training dataset includes corresponding demographic values, comorbidity values, vital sign values, blood test values, and / or observed values of disease progression values for a plurality of individuals diagnosed with portal hypertension and / or cirrhosis, and for the computing system to apply a machine learning trainer to the training dataset, where the machine learning trainer generates a plurality of machine learning models, and each of the machine learning models obtains new observed values of new demographic values, new comorbidity values, new vital sign values, and / or new blood test values as input, and (i) a hazard ratio of whether an individual diagnosed with portal hypertension and / or cirrhosis with the new observed values is expected to show progression to each state associated with portal hypertension or cirrhosis, and / or (ii) a prediction of the period between the time of the new observed values and a further diagnosis of each state.
[0007] A second exemplary embodiment is for a computing system to obtain observed values of an individual's demographic values, comorbidity values, vital sign values, and / or blood test values, where the individual has been diagnosed with portal hypertension and / or cirrhosis, and for the computing system to apply a machine learning model to the observed values, where the machine learning model has been trained on a training dataset that includes observed values of corresponding demographic values, comorbidity values, vital sign values, blood test values, and / or disease progression values for a plurality of individuals diagnosed with portal hypertension and / or cirrhosis, and where the machine learning model is configured to provide a prediction of (i) a hazard ratio of whether the individual is expected to show progression to a condition associated with portal hypertension or cirrhosis, and / or (ii) a period between the time of the observed values and a further diagnosis of the condition, and for the computing system to provide a prediction based on the observed values.
[0008] In a third exemplary embodiment, a manufactured article includes a non-transitory computer-readable medium that, when executed by a computing system, stores program instructions that cause the computing system to perform operations according to the first exemplary embodiment and / or the second exemplary embodiment.
[0009] In a fourth exemplary embodiment, a computing system includes at least one processor, a memory, and program instructions. The program instructions are stored in the memory and, when executed by the at least one processor, can cause the computing system to perform operations according to the first and / or second exemplary embodiments.
[0010] In a fifth exemplary embodiment, a system includes various means for performing each of the operations of the first and / or second exemplary embodiments.
[0011] The above other embodiments, aspects, advantages, and alternatives will become apparent to those skilled in the art by reading the following detailed description with appropriate reference to the accompanying drawings. Further, this summary and other descriptions and figures provided herein are intended only to illustrate embodiments by way of example, and thus, numerous variations are possible. For example, structural elements and process steps can be rearranged, combined, distributed, eliminated, or changed while remaining within the scope of the claimed embodiments.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Best Mode for Carrying Out the Invention
[0013] This specification describes exemplary methods, devices, and systems. It should be understood that the terms "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "exemplary" or "an example" is not necessarily to be construed as preferred or advantageous over other embodiments or features, unless so stated. Thus, other embodiments may be utilized and other changes may be made without departing from the scope of the subject matter presented herein.
[0014] Accordingly, the exemplary embodiments described herein are not meant to be limiting. It will be readily understood that the aspects of the present disclosure generally described herein and illustrated in the figures can be arranged, substituted, combined, separated, and designed in a variety of different configurations. For example, the separation of the features into "client" and "server" components can be done in several ways.
[0015] Furthermore, unless the context clearly indicates otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should generally be regarded as depicting the structural aspects of one or more overall embodiments, with the understanding that not all illustrated features are necessary for each embodiment.
[0016]
[0017] I. Exemplary Computing Devices and Cloud-Based Computing Environments FIG. 1 is a simplified block diagram illustrating a computing device 100 that exemplifies some of the components that can be included in a computing device arranged to operate in accordance with an embodiment of the present specification. The computing device 100 can be a client device (e.g., a device actively operated by a user), a server device (e.g., a device that provides computing services to a client device), or some other type of computing platform. Some server devices may sometimes operate as client devices to perform certain operations, and some client devices may incorporate server functions.
[0018] In this example, the computing device 100 includes a processor 102, a memory 104, a network interface 106, and an input / output unit 108, all of which can be coupled by a system bus 110 or a similar mechanism. In some embodiments, the computing device 100 may include other components and / or peripheral devices (e.g., removable storage, printers, etc.).
[0019] The processor 102 can be one or more of any type of computer processing element, such as a central processing unit (CPU), a coprocessor (e.g., a computing, graphics, or encryption coprocessor), a digital signal processor (DSP), a network processor, and / or an integrated circuit or controller that performs processor operations. In some cases, the processor 102 may be one or more single-core processors. In other cases, the processor 102 may be one or more multi-core processors having multiple independent processing units. The processor 102 may also include register memory for temporarily storing executed instructions and related data, as well as cache memory for temporarily storing recently used instructions and data.
[0020] Memory 104 may be any form of computer-usable memory including, but not limited to, random access memory (RAM), read-only memory (ROM), and non-volatile memory (e.g., flash memory, hard disk drive, solid state drive, and / or tape storage). Thus, memory 104 represents both main memory units and long-term storage. Other types of memory may include biological memory.
[0021] Memory 104 may store program instructions and / or data that the program instructions may operate on. By way of example, memory 104 may store these program instructions on a non-transitory computer-readable medium such that the instructions are executable by processor 102 to perform any of the methods, processes, or operations disclosed herein or in the accompanying drawings.
[0022] As shown in FIG. 1, memory 104 may include firmware 104A, kernel 104B, and / or application 104C. Firmware 104A may be program code used to boot or otherwise initiate some or all of computing device 100. Kernel 104B may be an operating system including modules for memory management, scheduling, and management of processes, input / output, and communication. Kernel 104B may also include device drivers that enable the operating system to communicate with hardware modules of computing device 100 (e.g., memory units, network interfaces, ports, and buses). Application 104C may be one or more user space software programs such as a web browser or email client, and any software libraries used by these programs. Memory 104 may also store data used by these and other programs and applications.
[0023] The network interface 106 may take the form of one or more wired interfaces such as Ethernet (e.g., Fast Ethernet, Gigabit Ethernet, etc.). The network interface 106 may also support communication via one or more non-Ethernet media such as coaxial cables or power lines, or via wide area media such as Synchronous Optical Networking (SONET) technology or Software-Defined Wide Area Networking (SD-WAN) technology. The network interface 106 may further take the form of one or more wireless interfaces such as IEEE 802.11 (Wifi), BLUETOOTH (R), Global Positioning System (GPS), or wide area wireless interfaces. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used via the network interface 106. Additionally, the network interface 106 may comprise multiple physical interfaces. For example, some embodiments of the computing device 100 may include Ethernet, BLUETOOTH (R), and Wifi interfaces.
[0024] The input / output unit 108 may facilitate interaction between the user and peripheral devices with the computing device 100. The input / output unit 108 may include one or more types of input devices such as a keyboard, mouse, touch screen, etc. Similarly, the input / output unit 108 may include one or more types of output devices such as a screen, monitor, printer, and / or one or more light-emitting diodes (LEDs). Alternatively or additionally, the computing device 100 may communicate with other devices using, for example, a Universal Serial Bus (USB) port interface or a High-Definition Multimedia Interface (HDMI (R)) port interface.
[0025] To support the embodiments of this specification, one or more computing devices, such as computing device 100, may be deployed. The exact physical location, connectivity, and configuration of these computing devices may be unknown and / or unimportant to the client device. Thus, the computing devices may be referred to as "cloud-based" devices that may be housed at various remote data center locations.
[0026] FIG. 2 illustrates a cloud-based server cluster 200 according to an exemplary embodiment. In FIG. 2, the operations of a computing device (e.g., computing device 100) may be distributed among server device 202, data storage 204, and router 206, all of which may be connected by a local cluster network 208. The number of server devices 202, data storage 204, and routers 206 within server cluster 200 may depend on the (one or more) computing tasks and / or applications assigned to server cluster 200.
[0027] For example, server device 202 may be configured to perform various computing tasks of computing device 100. Thus, computing tasks may be distributed among one or more of server devices 202. To the extent that these computing tasks can be performed in parallel, such distribution of tasks may reduce the total time to complete these tasks and return results. For simplicity, both server cluster 200 and individual server devices 202 may be referred to as "server devices." This name is to be understood to mean that one or more separate server devices, data storage devices, and cluster routers may be involved in server device operation.
[0028] Data storage 204 may be a data storage array that includes a drive array controller configured to manage read and write access to a group of hard disk drives and / or solid state drives. The drive array controller may also be configured to manage backups or redundant copies of data stored in data storage 204, either alone or in cooperation with server device 202, to protect against drive failures or other types of failures that prevent one or more of server devices 202 from accessing units of data storage 204. Other types of memory other than drives may be used.
[0029] Router 206 may include network equipment configured to provide internal and external communications for server cluster 200. For example, router 206 may include one or more packet switching devices and / or routing devices (including switches and / or gateways) configured to provide (i) network communications between server device 202 and data storage 204 via local cluster network 208, and / or (ii) network communications between server cluster 200 and other devices via communication link 210 to network 212.
[0030] Furthermore, the configuration of router 206 can be at least partially based on the data communication requirements of server device 202 and data storage 204, the latency and throughput of local cluster network 208, the latency, throughput, and cost of communication link 210, and / or other factors that may contribute to the cost, speed, fault tolerance, elasticity, efficiency, and / or other design goals of the system architecture.
[0031] As a possible example, data storage 204 may include any form of database, such as a Structured Query Language (SQL) database. Various types of data structures, including but not limited to tables, arrays, lists, trees, and tuples, can store information in such a database. Further, any database within data storage 204 may be monolithic or distributed across multiple physical devices.
[0032] Server device 202 may be configured to send data to and receive data from data storage 204. This sending and retrieval may take the form of, respectively, an SQL query or other type of database query and the output of such a query. Additional text, images, videos, and / or audio may also be included. Further, server device 202 may compile the received data into a web page or web application representation or in some other way for use by a software application. Such a representation may take the form of a markup language such as HTML, XML (eXtensible Markup Language), or some other standardized or proprietary format.
[0033] Further, server device 202 may have the ability to execute various types of computerized scripting languages, including but not limited to Perl, Python, PHP (PHP Hypertext Preprocessor), ASP (Active Server Page), JAVASCRIPT (registered trademark), etc. Computer program code written in these languages can facilitate the provision of web pages to client devices and the interaction between client devices and web pages. Alternatively or additionally, JAVA (registered trademark) may be used to facilitate the generation of web pages and / or provide web application functionality.
[0034] II. Examples of Gradient Boosting Models The gradient boosting algorithm is a machine learning technique that can be used to develop predictive models for multi-dimensional datasets. For convenience, these sets are often represented in matrix form using columns and rows. One or more columns represent input variables, and an additional column represents the output variable. The output variable is an unknown function of one or more of the input variables. Rows typically represent observations of the input variables and their corresponding output variables based on real-world data. Often, the number of rows can be very large, in the hundreds, thousands, or more. The machine learning process involves training a gradient boosting model so that it can predict the output variable for new observations of the input variables. In other words, the model attempts to learn or at least approximate the unknown function from the existing instances of the input variables and their corresponding output variables.
[0035] Figure 3 further illustrates these concepts. The training dataset 300 includes a set of training observations (rows) consisting of input variables TIFF2025519998000002.tif6150, TIFF2025519998000003.tif6150, and TIFF2025519998000004.tif6150, and their corresponding output variables TIFF2025519998000005.tif6150. These input variables and output variables are related by some unknown function TIFF2025519998000006.tif6150, where TIFF2025519998000007.tif6150. The output variable TIFF2025519998000008.tif6150 can take various forms, such as an integer or real number, text, or a boolean value.
[0036] The training dataset 300 can be collected, for example, from medical data of actual patients from medical experts, hospitals, clinical trials, or other information sources. In such a training dataset, the value of the output variable for each observation is expected to be known, but not all values of the input variables need to be present. For example, the training dataset may be sparsely loaded.
[0037] The training dataset 300 is provided to the gradient boosting trainer 302, and the gradient boosting trainer 302 applies one or more training techniques to generate a gradient boosting model 304. The gradient boosting model 304 is an algorithm that can be used to apply an approximation of an unknown function to new observations of the input variables, or a set of parameters for controlling the behavior of the algorithm, such as TIFF2025519998000009.tif6150.
[0038] Thus, the gradient boosting model 304 can receive new observations 306 and generate predicted output variables 308. The accuracy of such predictions can vary based on the operation of the gradient boosting trainer 302 and the quality of the training dataset 300. The goal is for the gradient boosting model 304 to be as accurate as reasonably possible, given a sufficiently rich training dataset and a reasonable amount of time spent on training. This accuracy can be measured in various ways, as will be explained in more detail below.
[0039] The operation of the gradient boosting trainer 302 can involve training a set of decision trees (collectively referred to colloquially as a "forest"), each having a limited depth or a limited number of leaves. Thus, these trees are weak learners in that they generally do not take into account all the available information in the training data set, and thus their individual predictions may or may not have high accuracy. However, gradient boosting makes an overall prediction based on the weighting of the predictions from the individual trees. These overall predictions take into account most, if not all, of the training data set and thus are likely to be more accurate than the predictions from any of the individual trees.
[0040] However, unlike a random forest where each tree is independent of the others, the construction of subsequent trees in a gradient boosting model can be based on the error (or residual) of one or more of the previously constructed trees. In some cases, subsequent trees that sufficiently compensate for the error of the previous tree are given more weight towards the overall prediction, while in other cases, all trees can be weighted equally. Gradient boosting continues to build trees in this way until a predetermined number of trees have been constructed or until a new tree can no longer improve the accuracy of the prediction by more than a predetermined margin.
[0041] When the output variable takes on continuous values, such as integers within a certain range, the trees are constructed based on the magnitude of the residual between the actual value of the training data output variable and the predicted value associated with it. This can be referred to as gradient boosting for regression. In some cases, these residuals are referred to as "pseudo-residuals" to distinguish gradient boosting from linear regression, but the terms "residual" and "pseudo-residual" are used interchangeably herein. Each observation The initial prediction for TIFF2025519998000010.tif6150 is TIFF2025519998000011.tif6150, the same value such as the average of some or all of the output variables in the training data set TIFF2025519998000012.tif6150 is possible. In other words, it is TIFF2025519998000013.tif6150.
[0042] Here, the notation TIFF2025519998000014.tif6150 refers to tree 0 to the observed values obtained using TIFF2025519998000015.tif6150 and the prediction of TIFF2025519998000016.tif6150 (see below for further details on how the prediction is calculated using multiple trees). The initial prediction is TIFF2025519998000017.tif6150, and since it is generally based only on the value of the output variable, it can take the form of a single node rather than a tree.
[0043] In any case, the tree is constructed to predict the value of the residual. The non-leaf nodes of the tree represent conditions on the input variables. For example, the root node within a tree constructed from the training dataset 300 represents a condition such that the left branch of the node is followed when this condition is true and the right branch of the node is followed when this condition is false, which can be represented by TIFF2025519998000018.tif6150. Either of these branches can lead to another non-leaf node representing a condition or a leaf node representing a residual. There may be more than three branches, but for the sake of convenience, binary trees are used in the examples in this specification.
[0044] The tree construction may be based on various algorithms used in decision trees. In some cases, this may involve selecting input variables and, optionally, associated cutoff values based on entropy or Gini impurity. The cutoff values are selected to split the values of the input variables in such a way that the output variables can be reasonably predicted from the input variables. The input variables are then placed as nodes within the tree, with more predictive input variables generally placed higher up in the tree (e.g., closer to the root node). In some cases, randomness may be added to the process of determining where to place the input variables within the tree.
[0045] Since the number of observations is usually much larger than the number of leaves of a tree with a limited size, the residuals of each observation leading to the same leaf are typically averaged and then placed at the leaf. Thus, the leaf represents the aggregated residuals of several observations It can represent TIFF2025519998000019.tif6150. Here, the notation TIFF2025519998000020.tif6150 is Tree 0 to Observations made using TIFF2025519998000021.tif6150 Refers to the residuals of TIFF2025519998000022.tif6150 (see below for further details on how residuals are calculated using multiple trees).
[0046] Subsequently, iterative predictions are made for the observations within the training dataset. Each prediction in the first iteration TIFF2025519998000023.tif6150 involves traversing tree 1 for the observation until reaching a leaf, and then adding the residual of that leaf TIFF2025519998000024.tif6150 to the initial prediction according to the learning rate TIFF2025519998000025.tif6150. In other words, TIFF2025519998000026.tif is 6150. The learning rate helps prevent overfitting of the training dataset and enables small steps towards higher prediction accuracy.
[0047] Prediction From TIFF2025519998000027.tif 6150, new residuals TIFF2025519998000028.tif 6150 are calculated. Again, the residuals are based on the difference between the predicted values associated with the actual output variable values in the training dataset and involve the aggregations possible as described above. These new residuals are generally Expected to be smaller than the residuals of TIFF2025519998000029.tif 6150, although this is not always true for all residuals.
[0048] The next tree, Tree 2, can be constructed based on these new residuals. This tree may have the same structure as Tree 1 or may be structured differently by nodes representing input variables that appear at different positions (e.g., using randomness).
[0049] Then, each prediction TIFF2025519998000030.tif 6150 traverses both Tree 1 and Tree 2 for the observed values until reaching each leaf and then adds the associated residuals TIFF2025519998000031.tif 6150 and TIFF2025519998000032.tif 6150 to the initial prediction according to the learning rate. In other words, TIFF2025519998000033.tif is 6150.
[0050] Prediction From TIFF2025519998000034.tif 6150, new residuals TIFF2025519998000035.tif6150 is calculated. These residuals are expected to continue to decrease as the number of trees increases.
[0051] The process of constructing a new tree based on the residuals and making a new prediction, as described above, continues until a predetermined number of trees are constructed or adding a new tree can no longer reduce the size of the residuals by more than a predetermined margin. At this point, training is complete, and the trained gradient boosting model is ready to make predictions for new observations of the input variables.
[0052] Given such a new observation, the model starts with an initial prediction TIFF2025519998000036.tif6150, traverses all the trees according to the values of the input variables, and adds the resulting residuals. Thus, assuming TIFF2025519998000037.tif6150 trees are added to the initial node, the predicted value of the output variable for the new observation is TIFF2025519998000038.tif10150, where TIFF2025519998000039.tif6150 is the initial residual, and TIFF2025519998000040.tif6150 is TIFF2025519998000041.tif6150, and TIFF2025519998000042.tif6150 is the residual of the 6150th tree for this observation.
[0053] Gradient boosting can also be used to predict the value of an output variable from a limited number of possible values. For example, if the output variable is a boolean value, gradient boosting can be used to train a binary classifier. This can be called gradient boosting for classification.
[0054] In this case, the prediction can be based on (i) the natural logarithm of the odds that the output variable is true, and (ii) the probability that the output variable is true, across all observations in the training dataset. For example, assume there are 100 observations, 70 of which are "true" and 30 are "false". The natural logarithm of the odds that an observation is true is TIFF2025519998000043.tif6150, and the probability that the output variable is true is 0.7.
[0055] The natural logarithm of the odds ( TIFF2025519998000044.tif6150) is used as the initial prediction for all observations, and the probability (0.7) is used to calculate the residuals. Since 0.7 is greater than 0.5, the initial prediction is "true" for all observations (note that values other than 0.5 can be used as the cutoff in this process). Clearly, these initial predictions are not accurate as indicated by their residuals. Assigning values of 1.0 for true and 0.0 for false, the residuals will be 0.3 for each observation with an output variable that is "true" and -0.7 for each observation with an output variable that is "false".
[0056] Furthermore, the residuals are typically transformed to produce the output values of the leaves. This is because the residuals are in terms of probability while the predictions are in terms of the natural logarithm of the odds. The corresponding predicted probability TIFF2025519998000045.tif6150 and the residuals TIFF2025519998000046.tif6150 for an exemplary transformation for a leaf are as follows. TIFF2025519998000047.tif10150
[0057] These output values of the leaves are then scaled by the learning rate and added to the initial predictions. Since these predictions are still in the form of the natural logarithm of the odds, they can be transformed to the probability form used by the residuals by applying the logistic function. For example, The predicted value of TIFF2025519998000048.tif6150 should have the following probabilities. TIFF2025519998000049.tif9150
[0058] These residuals are then determined as the difference between the value and probability of the output value from the training dataset. Similarly to before, new trees can be calculated based on these residuals. This process continues until a predetermined number of trees are built or adding new trees can no longer reduce the size of the residuals by more than a predetermined margin.
[0059] The trained gradient boosting model is then applied similarly to new observations. The process adds the initial prediction and the transformed output value of each leaf associated with the new observation to find the natural logarithm of the predicted odds. A logistic function is then applied to this prediction to provide the probability. If the probability is greater than 0.5, the final prediction for this new observation is "true", otherwise it is "false".
[0060] Note that the above description only provides a general overview of some ways to perform gradient boosting. Other techniques are possible. Additionally, other types of machine learning models, such as artificial neural networks and expert systems, can be used instead of, or in combination with, gradient boosting techniques. That being said, gradient boosting generally works well with certain types of data and produces a model that is explainable, i.e., one can analyze the model to understand how and why it is generating its answers, which may not be the case for artificial neural networks. Furthermore, gradient boosting models also tend to work well when the training dataset is sparse (e.g., many observations missing at least one input variable), while artificial neural networks tend to be less efficient with sparse data.
[0061] There are two common gradient boosting frameworks, XGBoost and LightGBM, that can be used to train and deploy gradient boosting models. Each will be briefly described below for illustrative purposes. That being said, other frameworks can also be used.
[0062] A.XGBoost When the output variable takes continuous values, XGBoost constructs a tree by placing all the residuals at the root node and then calculates the similarity score of the residuals. The similarity score can be calculated as the sum of the squares of the residuals divided by the number of residuals. In some cases, a regularization constant is also added to the denominator to reduce sensitivity to outliers and overfitting of the training dataset. In any case, the higher the similarity score, the more similar the residuals are.
[0063] Next, various ways of splitting the residuals into groups are considered to check whether any of these splits results in a higher overall similarity score. This can be determined by calculating the gain of the split with respect to the original grouping of all the residuals within the root node. The gain can be calculated as the sum of the similarity scores of each node of the split minus the similarity score of the root node. Then, the split that produces the maximum gain among all splits is selected to create a branch from the root node (i.e., each residual group within the selected split becomes a child node of the root node).
[0064] Next, the same process is performed for each of the new child nodes. If a child node contains only one residual, the child node cannot be split further and becomes a leaf. Also, the tree may be limited to a maximum number of levels (e.g., 4, 6, or 8), and nodes at this maximum depth are not split further. In some cases, XGBoost may require that a minimum number of residuals (e.g., 2, 3, 4…) be represented at each node.
[0065] Once the XGBoost tree is constructed in this way, some branches that generate gains below the threshold can be pruned. The pruning process reconnects those branches to their respective parent nodes.
[0066] Predictions are made by traversing the tree with the values of the input variables until a leaf is reached. The output value of a leaf is the sum of the residuals of that leaf divided by the number of residuals in that leaf. Again, a regularization constant can also be added to the denominator.
[0067] Similar to standard gradient boosting, this tree is then used to make predictions scaled by the learning rate. The residuals from these predictions are then used to construct the next tree, and so on. Tree construction ends when a predetermined maximum number of trees have been constructed or when the residuals become smaller than a predetermined threshold.
[0068] When the output variable takes one of the values of a discrete number (e.g., for classification), XGBoost maps these to a numerical range. For example, each value of a boolean output variable is mapped to 1.0 or 0.0. The splits are then done as described above.
[0069] However, a different similar score calculation per node is used, which is the sum of the squares of the residuals within the node divided by the sum of (i) the previous probability over all observations and (ii) the product of 1 minus the previous probability. A regularization constant can also be added to the denominator. Tree construction is also done as above, but this different similarity score calculation is used to determine the gain.
[0070] As mentioned above, XGBoost may require that the minimum number of residuals be represented at each leaf. However, in this version of XGBoost, a value called "cover" is used instead of the count of residuals. Cover is the denominator obtained by subtracting the regularization constant from the similarity score. Leaves with a cover value below the threshold may be removed from the tree, effectively pruning the tree. Other pruning techniques described above can also be used.
[0071] In the case of prediction, the output value of a leaf is the sum of the leaf residuals divided by the sum of the products of (i) the previous probability over all observations and (ii) one minus the previous probability. A regularization constant can also be added to the denominator. Across multiple trees, the prediction of an observation is the natural logarithm of the odds of the initial output value, plus the output value of each tree scaled by the learning rate.
[0072] To convert this result back to a probability, the logistic function can be applied to this result. Based on the value of this probability (e.g., above or below 0.5 for a binary output variable), the value of the output variable can be selected.
[0073] XGBoost also uses several techniques to speed up the processing of its large training datasets. These techniques include an approximate greedy algorithm for selecting splits, a weighted sketch algorithm for focusing on observations that are difficult to predict, using distributed training across multiple processors or computers, and / or holding commonly used variables and constants in the processor cache. Other techniques can also be applied.
[0074] B. LightGBM LightGBM also uses gradient boosting, but generally uses gradient boosting to increase training speed, reduce memory usage, and improve accuracy. Specifically, LightGBM does not consider all values of the input variables, but bins these values to form a histogram and operates on the bins rather than the values. Also, LightGBM uses exclusive feature bundling to reduce the dimensionality of the feature space when two or more features tend to take mutually exclusive values. Additionally, LightGBM uses gradient-based one-sided sampling to identify observations with the largest residuals and operates as a random sampling of observations with lower residuals only for those observations.
[0075] As a result, LightGBM focuses its calculations where they are most needed, that is, on input variables that are not similar to each other and on observations with the most errors in the initial trees. Thus, in practice, LightGBM can perform about 10 times faster than other gradient boosting implementations with similar accuracy.
[0076] III. Survival Time Analysis Survival time analysis is a set of statistical methods that provide estimates of the amount of time (e.g., number of days, weeks, months, years) until an event occurs. An example of the use of such an analysis would be to predict the number of days from when a patient is diagnosed with a condition (e.g., portal hypertension) until the patient is expected to show an outcome such as the onset of varices, ascites, or death. This analysis can be performed on both treated and untreated patients. Often, the goal of survival time analysis is to determine whether a particular treatment affects the predicted survival time.
[0077] In particular, survival time analysis is not limited to predicting the time until a patient dies and is used to predict the time between two events. However, it has such a name because it is often used to predict a patient's survival time.
[0078] Survival time analysis is typically performed in the form of regression using a training dataset of observations. Each observation can indicate whether the patient showed an outcome (a binary true or false), and if so, the time between some initial state and that outcome. For example, the initial state could be a diagnosis of portal hypertension and the outcome could be varices. If the patient is diagnosed with varices during the observation period, the time between the two diagnoses is the survival time. Some conditions (such as portal hypertension) have more than one possible progression, so each could be modeled separately.
[0079] Most survival time analyses must take into account censored data. For example, in a clinical trial, some patients may drop out of the study before the outcome for that patient can be observed. This can be due to various reasons, such as the study ending before an outcome is observed in the patient, or the patient leaving the study early for some reason (e.g., loss of interest, moving to another location, or death). Such data are considered to be "right-censored" and are typically assumed to be uninformative. In other words, right-censored observations are not taken into account by the model.
[0080] Cox regression is a survival time analysis technique that allows multiple input variables to be used to predict an outcome (the output variable). The Cox model assumes that the log hazard of an observation is a linear function of its covariates and a population-level baseline hazard function that varies over time. Embodiments of the present specification include a training data set having multiple input variables, and since the respective effects on various outcomes are unknown, Cox regression is a good candidate for predicting survival time.
[0081] Another method of survival time analysis is Accelerated Failure Time (AFT). AFT is also a regression-based approach that supports multiple input variables. Unlike Cox regression, AFT allows for a fully parametric specification of the hazard function. That said, other parametric, semi-parametric, or non-parametric survival functions can also be used.
[0082] XGBoost embodiments support Cox regression and AFT and can thus be used for survival analysis. Other gradient boosting embodiments may provide similar functionality. Another software package, NGBoost, includes an AFT functionality integrated into a gradient boosting framework that allows the gradient to be characterized as a probability distribution (e.g., as a random survival forest).
[0083] IV. Prediction of the Progression of Portal Hypertension As described above, portal hypertension is a condition in which liver damage (e.g., cirrhosis) or portal vein occlusion (endogenous or exogenous) leads to an increase in the venous blood pressure of the portal venous system that carries blood from the gastrointestinal organs to the liver. Untreated portal hypertension can lead to several conditions including, but not limited to, varices, variceal bleeding, recurrent variceal bleeding, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, jaundice, lower extremity swelling, coagulation disorders, pulmonary complications, portosystemic shunts, and death.
[0084] Figure 4 illustrates an exemplary progression pathway of portal hypertension. These pathways generally progress from the observed values of compensated cirrhosis to the observed values of decompensated cirrhosis and further to the observed values of further decompensation.
[0085] Patients with compensated cirrhosis may mostly be asymptomatic and may not have ascites, variceal bleeding, hepatic encephalopathy, or jaundice. Nevertheless, such patients can also be diagnosed with mild portal hypertension (usually, the condition is not observed). Mild portal hypertension can progress to clinically significant portal hypertension (where one or more conditions such as varices are observed).
[0086] Clinically significant portal hypertension can progress to decompensated cirrhosis. Patients with decompensated cirrhosis may exhibit one or more conditions (e.g., ascites, variceal bleeding, hepatic encephalopathy, and / or jaundice). Patients with late-stage decompensated cirrhosis may have more severe conditions such as recurrent variceal bleeding, refractory ascites, hepatic encephalopathy, portosystemic shunts, and / or jaundice. The average life expectancy of patients with decompensated cirrhosis can be counted in months or a few years.
[0087] Therefore, the possible progression pathways of portal hypertension lead from compensated cirrhosis without varices to compensated cirrhosis with varices, decompensated cirrhosis with variceal bleeding, and death through recurrent variceal bleeding. Other progression pathways are also possible.
[0088] Current treatments for portal hypertension are either ineffective or invasive for many patients. Thus, it is desirable to have a framework for assessing whether a particular patient is likely to progress to one or more of these conditions, as well as a prediction of the time frame for progression. If such a framework were available, patients who are more likely to progress more rapidly towards an undesirable state could be identified early in that progression. In that case, more aggressive treatment could be considered for these patients to slow that progression. In other situations, patients with any predicted rate of progression could be selected for clinical trials of new treatments (e.g., diet, supplements, and / or pharmaceuticals). Embodiments herein can include, for example, software that analyzes a patient database and provides a ranking of these patients for inclusion in clinical trials in order of risk of progression (trials in patients with a higher risk of progression can shorten the duration of the trial and reduce the placebo effect). Thus, embodiments herein can potentially lead to an improvement in the lifespan and quality of life of patients with portal hypertension.
[0089] Thus, in order to predict whether a state is likely to occur and the length of time until the state is expected to become observable, for each state of interest, one by one, a series of TIFF2025519998000050.tif6150 machine learning models can be trained on patient data. For example, there can be one model for each of varices, ascites, variceal bleeding, hepatic encephalopathy, and death. Thus, in this example, TIFF2025519998000051.tif6150. To predict whether any of these states are likely to occur and the length of time until that state is expected to be observable, additional overall models can be trained on patient data. Thus, in total There can be 6150 machine learning models. In some cases, the model can predict the progression from early complications to more severe later complications (e.g., from portal hypertension with aneurysms without bleeding to aneurysms with bleeding).
[0090] These machine learning models can operate based on the observed progression timeline. A general framework for such a timeline is illustrated in FIG. 5. Specifically, timeline 500 includes four main points of interest for modeling possible progressions. These progressions include cases where an outcome is observed, where an unknown outcome is observed, or where no outcome is observed. The progression of portal hypertension for each patient can be mapped to timeline 500.
[0091] Time point 504 is the reference date for prediction, the starting date. The starting date may be the first recorded diagnosis of the patient's portal hypertension and / or cirrhosis, or it may be the point in time from which predictions are made for patients who have already been diagnosed with portal hypertension and / or cirrhosis. In some cases, these patients are selected such that the complications of portal hypertension or the causes of non-cirrhotic portal hypertension have not been previously recorded. Patients who exhibit certain characteristics or do not observe certain features, for example, patients who have had a liver transplant or severe cirrhosis prior to this starting date, may be excluded from the training data.
[0092] Thus, time point 504 is the point in time from which the outcome (if any) is measured. In practice, it is desirable to have at least six months of observational values prior to time point 504 to increase the likelihood of identifying incident patients. Thus, the time between time point 502 and time point 504 should be at least six months, although a minimum period other than six months may be used.
[0093] Time point 506 is the time point at which an outcome is observed. As noted above, in some cases, no outcome is observed. In these situations, time point 506 does not exist. The outcome may be known (e.g., aneurysm, aneurysm hemorrhage, recurrent aneurysm hemorrhage, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, and / or jaundice). In other situations, the outcome may be unknown. For example, a condition may have been observed, but the exact nature of the condition cannot be determined from the available data (e.g., it may be known that the patient has an aneurysm, but since the patient has not yet had an endoscopy to make that determination, it may be unknown whether the aneurysm is bleeding).
[0094] Time point 508 represents the end of the data for a patient in whom no outcome was observed. For example, the patient has died, dropped out of the medical system from which data was being collected, or the study period has ended. The data for such a patient can be right-censored.
[0095] Figure 6 illustrates this patient data in another way. Graph 600 includes two exemplary timelines of the progression of portal hypertension. In both cases, it is assumed that the patient has been diagnosed with cirrhosis and portal hypertension. Also in both cases, it includes health data (e.g., test, procedure, vital sign measurement, and / or clinical examination results) aggregated monthly and represented by asterisks. For simplicity, only three asterisks are shown per timeline in Figure 6, but these health data entries can continue for several months. In any case, the health data can be considered a sparse data set with some months missing at least some values. There may be a possibility that 20% to 60% of all expected health data entries do not exist.
[0096] Since the timeline 602 shows the progression of liver disease to an outcome (e.g., aneurysm, aneurysm bleeding, recurrent aneurysm bleeding, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, and / or jaundice), it is called a "Type 1" timeline. The timeline 604 is called a "Type 2" timeline because no progression of liver disease is observed and the associated data is right-censored.
[0097] The health records of patients can be formed into a training dataset according to such timelines. As a possible example, a starting date is determined for each patient. This can be the date considering that the patient's data is incorporated into the training dataset. Patients with portal hypertension observed for at least 6 months before the starting date may incorporate their data into the training dataset. Then, for patients with known or unknown outcomes, each outcome and the associated outcome time are determined. Patients without an outcome are right-censored. Alternatively, the progression may be represented as the amount of time (e.g., in days) between time point 504 and time point 506, along with an indication of the outcome observed at time point 506.
[0098] Since some specific unknown outcomes can define a lower bound for the time to progression to a certain state, some data can also be left-censored. For example, when a patient did not show an aneurysm before a given date, currently shows an aneurysm, and the patient's bleeding status is unknown. [Table 1]
[0099] Table 1 provides a simple example of possible training data. In this table, Patient A was observed to have a venous aneurysm 432 days after the starting date of May 22, 2019, Patient B was observed to have intractable ascites 325 days after the starting date of June 1, 2019, and Patient C was observed to have an unknown condition 517 days after the starting date of May 27, 2019. In some cases, when multiple conditions are observed, multiple entries may be possible for each patient. For example, Table 1 may also have a second entry for Patient A indicating that Patient A was also observed to have a venous aneurysm with bleeding 489 days later. Alternatively or additionally, there may be entries for monthly aggregated health data (not shown). The monthly test results are aggregated over the 6 months prior to the starting date or observation, and the last observed value can be carried forward to the starting date.
[0100] In any case, the format and content of Table 1 are for illustrative purposes, and other formats and / or content of the training data may be possible. As described above, the training data may also include, for each patient, demographic data, comorbidities, test results, vital signs, and / or other health data (e.g., medication, prescription, treatment, exercise, nutrition, mental health).
[0101] Figure 7 illustrates the training and prediction stages of a series of machine learning models. The training dataset 700 may include observed values from patients related to the progression of portal hypertension. As described above, these may include input variables (e.g., demographics, comorbidities, test results, and / or vital signs) as well as output variables (an indication of the diagnosis of one or more additional conditions and the number of days from the starting date until one or more of these additional conditions are diagnosed).
[0102] A series of To generate 6,150 trained machine learning models, a machine learning trainer 702 can be applied to a training dataset 700. The machine learning trainer 702 can use gradient boosting and survival time aspects in a robust manner when the training dataset 700 is sparse.
[0103] Each trained model can predict, given new input variables from a patient (not shown), whether the patient is expected to be diagnosed with a particular condition and how many days it is expected to take until such a diagnosis can be made. For example, model 704A may be trained to predict 706A whether a diagnosis of aneurysm will be made and, if a diagnosis of aneurysm is made, how many days it will take from a starting date until the diagnosis. Similarly, model 704B may be trained to predict 706B whether a diagnosis of ascites will be made and, if a diagnosis of ascites is made, how many days it will take from a starting date until the diagnosis. Similarly, model 704C may be trained to predict 706C whether a diagnosis of hepatic encephalopathy will be made and, if a diagnosis of hepatic encephalopathy is made, how many days it will take from a starting date until the diagnosis. Each of these models may be trained independently and may also operate independently when making predictions.
[0104] The prediction of whether a patient is expected to be diagnosed with a particular condition may be in the form of a hazard ratio of whether an individual is expected to progress to a condition associated with portal hypertension or cirrhosis. The hazard ratio may be relative to the general population, i.e., whether the patient is expected to progress faster or slower. This hazard ratio can be a non - negative value used to compare a particular cohort to the general population. Thus, a hazard ratio of 0.1 means that a particular cohort is 10 times less likely to exhibit the condition, and a hazard ratio of 2 means that a particular cohort is 2 times more likely to exhibit the condition. The hazard ratio can be thresholded to a boolean true or false for inclusion or exclusion (e.g., if the hazard ratio is above 0.1, the boolean value is "true", and if the hazard ratio is 0.1 or below, the boolean value is "false").
[0105] Alternatively, in the context of FIG. 3, the output variable within the training data set 300 may be a time range, and the upper and lower limits of this range are the same if the event is observed at an exact point in time. Further, the predicted output variable 308 can be a hazard ratio and / or a predicted time value (e.g., number of days).
[0106] Furthermore, a general model is further generated to predict whether any such condition will be diagnosed and how many days it is expected to take to make such a general diagnosis (to create a total of 6150 trained models of TIFF2025519998000055.tif), the machine learning trainer 702 can be applied to the training data set 700. In the case of a patient predicted to be diagnosed with two or more conditions (e.g., varices and ascites), the predicted number of days can be until the earliest diagnosis is expected to be made from the starting date. Thus, the model 704D may be trained to predict whether any one or more conditions associated with varices will be diagnosed and, if a diagnosis is made, how many days it will take from the starting date to the diagnosis 706D.
[0107] As shown by the omission of FIG. 7, TIFF2025519998000056.tif6150 can take various values. In some cases, TIFF2025519998000057.tif6150 can be 1 or 2, and in other cases, TIFF2025519998000058.tif6150 can be 7, 8, or more. Thus, the embodiments of this specification can use TIFF2025519998000059.tif6150 with any value. These embodiments may or may not include a general model.
[0108] FIG. 8 illustrates how any one of these trained models makes predictions. Demographic data 800 (e.g., age, gender, race), comorbidities 802 (e.g., diabetes, obesity), and test results and vital sign data 804 (e.g., body mass index, blood pressure, heart rate, blood test data) are provided as inputs to the trained machine learning model 806. Other data such as medications taken may also be included in the input.
[0109] As described above, the training dataset can indicate the starting date of a patient representing the time point when the patient is known or at least suspected of having portal hypertension and / or cirrhosis. Further, the test results and vital sign data 804 may optionally include a plurality of measurements taken over time, aggregated monthly.
[0110] The machine learning model 806 may use some form of survival time technique (e.g., Cox proportional hazards and / or accelerated failure time) to predict the number of days from the starting date until a further condition is diagnosed. The machine learning model 806 may be based on LightGBM or XGBoost as shown, or may be based on NGBoost or some other gradient boosting technique that supports survival analysis. In an alternative embodiment, the machine learning model 806 may be at least partially based on an artificial neural network, an expert system, or some ensemble combination of these or other models.
[0111] The machine learning model 806 may generate a per-patient prediction 508, which may indicate whether a condition is expected to be observed and the number of data points between the starting date and when the condition is expected to be observed.
[0112] Data for training and validating the machine learning model 806 may be obtained from various information sources including hospitals, healthcare providers, clinical information sources, and / or insurance claims. For example, the machine learning model 806 can be trained on insurance claim data because this data can include patient demographics, vital signs, and blood test results regardless of the presence or absence of portal hypertension and / or cirrhosis. This training data may be preprocessed in various ways, e.g., for outlier removal, deskewing, and / or normalization.
[0113] Next, the trained machine learning model 806 can be validated with clinical data to determine how accurately it predicts the progression of a patient's portal hypertension. Once validated, the machine learning model 806 can be applied in a hospital to identify patients or for use in hospital services that are candidates for further testing, treatment, or inclusion in a clinical trial.
[0114] In some cases, the training data can be collected from multiple geographic regions. However, the data may be segmented by region to develop region-specific models. Depending on the situation, information from identified patients may be checked for novelty, for example, whether the input variables of these patients match the input variables of the training data. For this purpose, a similarity model may be applied to the information from the identified patients and the training data, and dissimilar patients may be identified for further processing before the prediction is finalized. This can be useful when a machine learning model is trained on data from one population (e.g., located in North America) but applied to another population (e.g., located in Europe).
[0115] V. Experimental Results The embodiments herein were applied to training data containing a cohort of 10,429 patients that met the inclusion and exclusion criteria, and a severe condition was observed in 41% of these patients. The median progression time for these patients was 190 days. Patient data was collected from the Optum Clinformatics Data Mart (Optum) dataset (de-identified US electronic health records, 2007 - 2021). A machine learning survival time model robust to sparse data was used to model the hazard rate of progression. Model performance was evaluated using three-way cross-validation of the area under the cumulative dynamic curve (AUC) and compared to four established liver disease scores (fib-4, ALBI, PALBI, and MELD), as well as patient age as a baseline.
[0116] The AUC metric can be based on the Receiver Operating Characteristic (ROC) curve or the precision-recall curve generated from the model results. The ROC curve plots the true positive rate of the model (the number of true positives divided by the sum of the number of true positives and false negatives) against the false positive rate (the number of false positives divided by the sum of the number of false positives and true negatives). The ROC curve visually represents the trade-off between the true positive rate and the false positive rate. One measure of model quality is the AUC of the ROC curve. This value is typically between 0.5 and 1.0, and higher values indicate better model performance across various parameter settings. The precision-recall curve plots the precision of the model (the number of true positives divided by the sum of the number of true positives and false positives) against the recall (the number of true positives divided by the sum of the number of true positives and false negatives). Again, the AUC can be used to evaluate model quality, and higher values indicate higher quality.
[0117] Figure 9 illustrates the evaluation of the model within the graph using a general model for predicting progression to any of the states. As shown, the embodiments herein generate a model that outperforms all of the fib-4, ALBI, PALBI, and MELD techniques in terms of the ability to predict at least the disease progression of patients with portal hypertension. As a baseline for comparison, age was also used as a predictor, which yielded an AUC of 0.49 and was tantamount to a random guess.
[0118] Thus, the prediction model described herein provides an improvement in AUC compared to the prior art for predicting the progression of portal hypertension. This model also functions well for partially observed data, and the state-specific model can distinguish the risk of progression due to individual complications, which is an improvement over the existing techniques.
[0119] VI. Deployment Scenario The embodiments of this specification can be deployed in several arrangements. With respect to the training of one or more machine learning models, this can be done on one or more computing devices within a server cluster, for example, within a public cloud network (such as Amazon AWS or Microsoft Azure) or on a private system. With respect to the execution of these models on new observations, the models can be hosted in various locations and environments.
[0120] In one possible example, the trained model is hosted on a public cloud network or a private network and can provide results to a client device via a web or application interface. For example, the client device may send a request containing a new set of input variables including new observations to the remotely hosted model. The model can take these as inputs and then generate corresponding outputs that are sent to the client device in response to the request. These trained models can be deployed by various entities such as hospitals, hospital networks, physicians, physician networks, universities, pharmaceutical companies, or some consortiums of one or more of these to other entities.
[0121] Alternatively, the trained model may be packaged with a client application that can be downloaded and installed on a desktop, laptop, or mobile computing device. Thus, the client application will include a user interface that allows the user to input or otherwise indicate input variables for new observations. The client application will then apply the model to this new observation and generate the corresponding output that is displayed and / or stored by the client device. This scenario has the advantage that a live network connection is not required to use the model.
[0122] In other alternative forms, the trained model can be used to develop simple clinical prediction rules, such as decision trees, that a healthcare provider can follow to predict the progression and / or prognosis of portal hypertension.
[0123] Other deployment scenarios may exist. Thus, the embodiments herein are not limited to these scenarios.
[0124] VII. Exemplary Operations Figures 10 and 11 are flowcharts illustrating exemplary embodiments. The operations illustrated in Figures 10 and 11 can be performed by a computing system or computing device that includes a software application configured to perform any of the embodiments herein. Non-limiting examples of a computing system or computing device include, for example, computing device 100 and server cluster 200. However, the operations can be performed by other types of devices or device subsystems. For example, the operations can also be performed by a portable computer such as a laptop or tablet device.
[0125] The embodiments of Figures 10 and 11 may be simplified by removing any one or more of the illustrated features. Further, these embodiments may be combined with any of the features, aspects, and / or embodiments described in any of the previous figures or elsewhere herein. Such embodiments may include instructions executable by one or more processors of one or more computing devices of a system or virtual machine or container. For example, the instructions may take the form of software instructions and / or hardware instructions and / or firmware instructions. In an exemplary embodiment, the instructions may be stored on a non-transitory computer-readable medium. When executed by one or more processors of one or more computing devices, the instructions may cause the one or more computing devices to perform the various operations of the embodiments.
[0126] In these embodiments, an individual diagnosed with portal hypertension and / or cirrhosis may include an individual diagnosed with compensated cirrhosis, an individual with compensated cirrhosis, or an individual showing symptoms of compensated cirrhosis. Thus, the term "portal hypertension and / or cirrhosis" may include compensated cirrhosis without a specific diagnosis of portal hypertension.
[0127] Block 1000 of FIG. 10 is to obtain a training dataset by a computing system, wherein the training dataset includes corresponding demographic values, comorbidity values, vital sign values, blood test values, and / or observed values of disease progression values for a plurality of individuals diagnosed with portal hypertension and / or cirrhosis.
[0128] Block 1002 of FIG. 10 is to apply a machine learning trainer to the training dataset by a computing system, wherein the machine learning trainer generates a plurality of machine learning models, and each of the machine learning models obtains new observed values of new demographic values, new comorbidity values, new vital sign values, and / or new blood test values as inputs, and (i) a hazard ratio of whether an individual diagnosed with portal hypertension and / or cirrhosis showing the new observed values is expected to show progression to each state related to portal hypertension or cirrhosis, and / or (ii) a prediction of the period between the time of the new observed values and a further diagnosis of each state.
[0129] In some embodiments, the machine learning trainer also obtains new observed values as inputs and generates additional machine learning models configured to provide further predictions of (i) a further hazard ratio of whether an individual is expected to show progression to any state related to portal hypertension or cirrhosis, and / or (ii) a further period between the time of the new observed values and an additional diagnosis of any state related to portal hypertension or cirrhosis.
[0130] In some embodiments, the disease progression value for a particular individual among a plurality of individuals includes a starting date and one or more outcomes, and each of the one or more outcomes indicates a particular state and the observation period between the starting date and when the particular state was diagnosed.
[0131] In some embodiments, the disease progression value also includes one or more additional outcomes, and each of the one or more additional outcomes indicates an unknown state and an additional observation period between the starting date and when the unknown state was identified. For example, in a patient with an aneurysm, the bleeding status of the state may be unknown.
[0132] In some embodiments, there are vital sign values or blood test values for at least six months before the starting date in the disease progression values for a plurality of individuals.
[0133] In some embodiments, the particular state is one of an aneurysm, aneurysm bleeding, recurrent aneurysm bleeding, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, portosystemic shunt, or jaundice.
[0134] In some embodiments, the demographic values include the age, gender, race, or ethnicity of a plurality of individuals.
[0135] In some embodiments, the vital sign values include the body mass index, blood pressure readings, or heart rate of a plurality of individuals.
[0136] In some embodiments, the comorbidity values include signs of diabetes or obesity.
[0137] In some embodiments, 20% - 60% of the values in the training dataset are populated ( populated).
[0138] In some embodiments, the machine learning model is based on gradient boosting.
[0139] In some embodiments, the machine learning model is based on gradient boosting and survival analysis.
[0140] In some embodiments, the training dataset includes at least 10,000 observations collected from medical claims records or electronic health records.
[0141] In some embodiments, the hazard ratio is provided as a boolean indication of progression to each state (e.g., the hazard ratio can be thresholded to either a "true" value indicating that progression to each state is predicted, or a "false" value indicating that progression to each state is not predicted).
[0142] In some embodiments, the observations in the training dataset also include instructions for dosing, prescribing, or treating related to multiple individuals, and the new observations also include instructions for dosing, prescribing, and / or treating related to an individual.
[0143] Block 1100 of FIG. 11 is to obtain observations of an individual's demographic values, comorbidity values, vital sign values, and / or blood test values by a computing system, where the individual is diagnosed with portal hypertension and / or cirrhosis.
[0144] Block 1102 of FIG. 11 is to apply a machine learning model to the observations by a computing system, where the machine learning model is trained with a training dataset, the training dataset includes corresponding demographic values, comorbidity values, vital sign values, blood test values, and / or disease progression values for a plurality of individuals diagnosed with portal hypertension and / or cirrhosis, and the machine learning model is configured to provide a prediction of (i) the hazard ratio of whether an individual is expected to show progression to a state related to portal hypertension or cirrhosis, and / or (ii) the period between the time of the observations and a further diagnosis of the state.
[0145] Block 1104 in FIG. 11 includes providing a prediction based on observed values by a computing system.
[0146] Some embodiments may further include applying, by a computing system, a second machine learning model to observed values, where the second machine learning model is trained on at least a portion of a training dataset and is configured to provide a second prediction of (i) a second hazard ratio of whether an individual is expected to exhibit progression to a second state associated with portal hypertension or cirrhosis and / or (ii) a second period between the time of the observed values and a second further diagnosis of the second state, and providing, by the computing system, a second prediction based on the observed values.
[0147] Some embodiments may further include applying, by a computing system, a further machine learning model to observed values, where the further machine learning model is trained on at least a portion of a training dataset and is configured to provide a further prediction of (i) a further hazard ratio of whether an individual is expected to exhibit progression to any state associated with portal hypertension or cirrhosis and / or (ii) a further period between the time of the observed values and a further diagnosis of any state associated with portal hypertension or cirrhosis, and providing, by the computing system, a further prediction based on the observed values.
[0148] In some embodiments, providing the prediction includes displaying the prediction on a graphical user interface.
[0149] In some embodiments, obtaining the observed values includes receiving the observed values from a client device that communicates with the computing system via a network, and providing the prediction includes transmitting the prediction to the client device.
[0150] In some embodiments, the disease progression value for a particular individual among a plurality of individuals includes a starting date and one or more outcomes, and each of the one or more outcomes indicates a particular condition and the observation period between the starting date and when the particular condition was diagnosed.
[0151] In some embodiments, the disease progression value also includes one or more additional outcomes, and each of the one or more additional outcomes indicates an unknown condition and an additional observation period between the starting date and when the unknown condition was identified.
[0152] In some embodiments, there are vital sign values or blood test values for at least six months prior to the starting date in the disease progression values for a plurality of individuals.
[0153] In some embodiments, the particular condition is one of aneurysm, aneurysm hemorrhage, recurrent aneurysm hemorrhage, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, portal systemic shunt, or jaundice.
[0154] In some embodiments, the demographic values include the age, gender, race, or ethnicity of a plurality of individuals.
[0155] In some embodiments, the vital sign values include the body mass index, blood pressure readings, or heart rate of a plurality of individuals.
[0156] In some embodiments, the comorbidity values include signs of diabetes or obesity.
[0157] In some embodiments, 20% - 60% of the values in the training dataset are loaded.
[0158] In some embodiments, the machine learning model is based on gradient boosting.
[0159] In some embodiments, the machine learning model is based on gradient boosting and survival time analysis.
[0160] In some embodiments, the training dataset includes at least 10,000 observations collected from medical claim records or electronic health records.
[0161] In some embodiments, the hazard ratio is provided as a boolean indication of progression to each state (e.g., the hazard ratio can be thresholded to either a "true" value indicating that progression to each state is predicted, or a "false" value indicating that progression to each state is not predicted).
[0162] In some embodiments, the observations in the training dataset also include dosing, prescription, or treatment instructions related to multiple individuals, and the observations also include dosing, prescription, or treatment instructions related to an individual.
[0163] VIII. Conclusion The present disclosure is not limited to the specific embodiments described in this application, which are intended as examples of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from its scope. In addition to those described herein, functionally equivalent methods and apparatuses within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to be included within the scope of the appended claims.
[0164] The above detailed description has described various features and operations of the disclosed systems, devices, and methods with reference to the accompanying drawings. The exemplary embodiments described herein and in the drawings are not meant to be limiting. Other embodiments can be utilized and other changes can be made without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure generally described herein and illustrated in the figures can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.
[0165] For any or all of the message flow diagrams, scenarios, and flowcharts of the figures, as described herein, each step, block, and / or communication can represent the processing of information and / or the transmission of information according to an exemplary embodiment. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, the operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages can be performed in an order different from the illustrated or described order, including substantially simultaneously or in a reverse order, depending on the functions involved. Further, more or fewer blocks and / or operations can be used with any of the message flow diagrams, scenarios, and flowcharts described herein, and these message flow diagrams, scenarios, and flowcharts can be combined with each other, either partially or wholly.
[0166] A step or block representing the processing of information can correspond to a circuit configured to perform a specific logical function of the method or technique described herein. Alternatively or additionally, a step or block representing the processing of information can also correspond to a module, segment, or part of program code (including related data). The program code can include one or more instructions executable by a processor to perform specific logical operations or actions in a method or technique. The program code and / or related data can be stored in any type of computer-readable medium, such as a storage device including RAM, disk drive, solid-state drive, or another storage medium.
[0167] A computer-readable medium can also include non-transitory computer-readable media such as register memory and processor caches that store data for a short period of time. The non-transitory computer-readable media can further include non-transitory computer-readable media that store program code and / or data over a longer period of time. Thus, the non-transitory computer-readable media can include, for example, secondary or persistent long-term storage such as ROM, optical or magnetic disks, solid state drives, or compact disc read-only memory (CD-ROM). The non-transitory computer-readable media can also be any other volatile or non-volatile memory system. The non-transitory computer-readable media can be regarded as, for example, a computer-readable storage medium or a tangible storage device.
[0168] Furthermore, one or more steps or blocks representing information transmission can correspond to information transmission between software modules and / or between hardware modules within the same physical device. However, other information transmissions can be between software modules and / or between hardware modules in different physical devices.
[0169] The specific arrangements shown in the figures should not be considered limiting. It should be understood that other embodiments can include more or fewer of each element shown in a given figure. Additionally, some of the illustrated elements can be combined or omitted. Furthermore, exemplary embodiments can include elements not illustrated in the figures.
[0170] Although various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and not intended to be limiting, and the true scope is indicated by the following claims.
Claims
1. Obtaining, by a computing system, a training dataset, wherein the training dataset includes corresponding demographic values, comorbidity values, vital sign values, blood test values, or disease progression values for a plurality of individuals diagnosed with portal hypertension or cirrhosis; obtaining; Applying, by the computing system, a machine learning trainer to the training dataset, wherein the machine learning trainer generates a plurality of machine learning models, and each of the machine learning models obtains, as input, new observation values of new demographic values, new comorbidity values, new vital sign values, or new blood test values, and (i) a hazard ratio indicating whether an individual diagnosed with portal hypertension or cirrhosis showing the new observation values is expected to show progression to each state associated with portal hypertension or cirrhosis, or (ii) a prediction of the period between the time of the new observation values and a further diagnosis of each state. Applying; A method comprising.
2. The machine learning trainer also obtains, as input, the new observation values and generates a further machine learning model configured to provide a further prediction of (i) a further hazard ratio indicating whether the individual is expected to show progression to any state associated with portal hypertension or cirrhosis, or (ii) a further period between the time of the new observation values and a further diagnosis of any state associated with portal hypertension or cirrhosis. The method according to claim 1.
3. The disease progression value for a particular individual among the plurality of individuals includes a starting date and one or more outcomes, and each of the one or more outcomes indicates a particular state and an observation period between the starting date and the time when the particular state was diagnosed. The method according to claim 1.
4. The disease progression value also includes one or more additional outcomes, and each of the one or more additional outcomes indicates an unknown state and an additional observation period between the starting date and the time when the unknown state was identified. The method according to claim 3.
5. There are vital sign values or blood test values for at least six months prior to the starting date in the disease progression value for the plurality of individuals. The method according to claim 3.
6. The method according to claim 3, wherein the specific state is one of aneurysm, aneurysm bleeding, recurrent aneurysm bleeding, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, portosystemic shunt, or jaundice.
7. The method according to claim 1, wherein the demographic values include the ages, genders, races, or ethnicities of the plurality of individuals.
8. The method according to claim 1, wherein the vital sign values include the body mass index, blood pressure readings, or heart rates of the plurality of individuals.
9. The method according to claim 1, wherein the comorbidity values include signs of diabetes or obesity.
10. The method according to claim 1, wherein 20% to 60% of the values in the training data set are loaded.
11. The method according to claim 1, wherein the machine learning model is based on gradient boosting.
12. The method according to claim 1, wherein the machine learning model is based on gradient boosting and survival time analysis.
13. The method according to claim 1, wherein the training data set includes at least 10,000 observations collected from medical claim records or electronic health records.
14. The method according to claim 1, wherein the hazard ratio is provided as a boolean indication of progression to each respective state.
15. The method according to claim 1, wherein the observations in the training data set also include instructions for medications, prescriptions, or treatments related to the plurality of individuals, and the new observations also include instructions for medications, prescriptions, or treatments related to the individual.
16. A manufactured article including a non-transitory computer-readable medium storing program instructions that, when executed by a computing system, cause the computing system to perform the operations according to any one of claims 1 to 15.
17. One or more processors and a memory including program instructions that, when executed by the one or more processors, cause a computing system to perform the operations according to any one of claims 1 to 15 A computing system comprising.
18. Obtaining, by a computing system, observed values of an individual's demographic values, the individual's comorbidity values, the individual's vital sign values, or the individual's blood test values, wherein the individual is diagnosed with portal hypertension or cirrhosis, the obtaining and Applying a machine learning model to the observed values by the computing system, wherein the machine learning model is trained with a training dataset, the training dataset includes observed values of corresponding demographic values, comorbidity values, vital sign values, blood test values, or disease progression values for a plurality of individuals diagnosed with portal hypertension or cirrhosis, and the machine learning model is configured to provide a prediction of (i) a hazard ratio of whether the individual is expected to show progression to a state related to portal hypertension or cirrhosis, or (ii) a period between the time of the observed values and a further diagnosis of the state. Providing the prediction based on the observed values by the computing system A method comprising.
19. Applying a second machine learning model to the observed values by the computing system, wherein the second machine learning model is trained with at least a part of the training dataset, and the second machine learning model is configured to provide a second prediction of (i) a second hazard ratio of whether the individual is expected to show progression to a second state related to portal hypertension or cirrhosis, or (ii) a second period between the time of the observed values and a second further diagnosis of the second state. Providing the second prediction based on the observed values by the computing system The method according to claim 18, further comprising.
20. Applying a further machine learning model to the observed values by the computing system, wherein the further machine learning model is trained with at least a part of the training dataset, and the further machine learning model is configured to provide a further prediction of (i) a further hazard ratio of whether the individual is expected to show progression to any state related to portal hypertension or cirrhosis, and (ii) a further period between the time of the observed values and a further diagnosis of any state related to portal hypertension or cirrhosis. Providing the further prediction based on the observed values by the computing system The method according to claim 19, further comprising.
21. The method of claim 18, wherein providing the prediction includes displaying the prediction on a graphical user interface.
22. The method of claim 18, wherein obtaining the observed values includes receiving the observed values from a client device that communicates with the computing system via a network, and providing the prediction includes transmitting the prediction to the client device.
23. The disease progression value for a particular individual among the plurality of individuals includes a starting date and one or more outcomes, each of the one or more outcomes indicating a particular state and an observation period between the starting date and when the particular state was diagnosed, according to the method of claim 18.
24. The method of claim 23, wherein the disease progression value also includes one or more additional outcomes, each of the one or more additional outcomes indicating an unknown state and an additional observation period between the starting date and when the unknown state was identified.
25. The method of claim 23, wherein there are vital sign values or blood test values for at least six months prior to the starting date in the disease progression values for the plurality of individuals.
26. The method of claim 23, wherein the particular state is one of aneurysm, aneurysm hemorrhage, recurrent aneurysm hemorrhage, ascites, refractory ascites, hepatic encephalopathy, recurrent hepatic encephalopathy, portosystemic shunt, or jaundice.
27. The method of claim 18, wherein the demographic values include the age, gender, race, or ethnicity of the plurality of individuals.
28. The method of claim 18, wherein the vital sign values include the body mass index, blood pressure readings, or heart rate of the plurality of individuals.
29. The method of claim 18, wherein the comorbidity values include signs of diabetes or obesity.
30. The method of claim 18, wherein 20% to 60% of the values in the training data set are loaded.
31. The method of claim 18, wherein the machine learning model is based on gradient boosting.
32. The method of claim 18, wherein the machine learning model is based on gradient boosting and survival time analysis.
33. The method of claim 18, wherein the training data set includes at least 10,000 observed values collected from medical claim records or electronic health records.
34. The method according to claim 18, wherein the hazard ratio is provided as a boolean representation of the progression to each of the respective states.
35. The method according to claim 18, wherein the observations in the training data set also include instructions for dosing, prescribing, or treating related to the plurality of individuals, and the observations also include instructions for dosing, prescribing, or treating related to the individuals.
36. A manufactured article including a non-transitory computer-readable medium storing program instructions that, when executed by a computing system, cause the computing system to perform the operations according to any one of claims 18 to 35.
37. One or more processors, a memory including program instructions that, when executed by the one or more processors, cause a computing system to perform the operations according to any one of claims 18 to 35 A computing system comprising.
Citation Information
Patent Citations
Method for determining treatment for cancer patients
JP2022505266A
Methods for predicting or detecting disease
US20190108912A1