Bayesian causal relationship network model for healthcare diagnosis and treatment based on patient data

By constructing a causal network model through the Bayesian network algorithm, the problem of variable pre-selection limitations in health care data analysis is solved, the discovery of unknown variables and improved treatment strategies are achieved, and the comprehensiveness and accuracy of data analysis are improved.

CN114203296BActive Publication Date: 2025-10-17BUPUG BIOPHARMACEUTICALS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111477970.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-09-11
Filing Date
2015-09-11
Publication Date
2025-10-17
Estimated Expiration
2035-09-11

AI Technical Summary

Technical Problem

In existing technologies for healthcare data analysis, pre-selecting variables limits the ability to discover new or unknown relationships and cannot fully mine the potential information in the data.

Method used

The Bayesian network algorithm is used to construct a causal network model based on a large amount of patient data. The network model is constructed through an unbiased data-driven method to depict the associations supported by the data set and identify the interactions between unknown variables.

Benefits of technology

It can discover new interactions between variables that were previously unknown to the healthcare community, provide improved treatment strategies and experimental protocols, and increase the comprehensiveness and accuracy of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114203296B_ABST
    Figure CN114203296B_ABST
Patent Text Reader

Abstract

Systems, methods, and computer readable media for healthcare analytics are provided. Data corresponding to a plurality of patients is received. The data is parsed to produce standardized data for a plurality of variables, the standardized data produced for more than one variable for each patient. A causal relationship network model involving the plurality of variables is produced based on the produced standardized data using a Bayesian network algorithm. The causal relationship network model includes variables related to a plurality of medical conditions or medical drugs. In another aspect, a selection of a medical condition or drug is received. A subnetwork is determined from the causal relationship network model. The subnetwork includes one or more variables related to the selected medical condition or drug. One or more predictors for the selected medical condition or drug are identified.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 201580058049.5, filed on September 11, 2015, and entitled "Bayesian Causal Relationship Network Model for Health Care Diagnosis and Treatment Based on Patient Data."

[0002] Related applications

[0003] This application is related to and claims priority from U.S. Provisional Patent Application No. 62 / 049,148, filed on September 11, 2014, the entire disclosure of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0004] The present disclosure relates generally to systems and methods for data analysis, and in particular to generating causal relationship network models using health care data. BACKGROUND

[0005] Many systems analyze data to gain insight into various aspects of health care. Insight can be gained by determining relationships among the data. Conventional methods pre-determine several relevant variables to extract from health care data for processing and analysis. Relationships are established among various factors, such as medical drugs, diseases, symptoms, etc., based on a small number of pre-selected variables. Pre-selecting variables of interest limits the ability to discover new or unknown relationships. Pre-selecting variables also limits the ability to discover other relevant variables. For example, if variables are pre-selected when considering an analysis of diabetes, one would be limited to those variables and would not recognize that data analysis supports another variable related to diabetes that was not known to the health care community before. SUMMARY

[0006] In one aspect, the present invention relates to a computer-implemented method for generating a causal relationship network model based on patient data. The method includes receiving data corresponding to a plurality of patients, wherein the data includes diagnostic information and / or treatment information for each patient, parsing the data to generate standardized data for a plurality of variables, wherein for each patient, the standardized data is generated for more than one variable, generating a causal relationship network model involving the plurality of variables based on the generated standardized data using a Bayesian network algorithm, the causal relationship network model including variables related to a plurality of medical conditions, and generating the causal relationship network using a programmed computing system, the programmed computing system including a memory containing network model building code and one or more processors configured to execute the network model building code.

[0007] In certain embodiments, the causal network model includes relationships indicating one or more predictors for each of a plurality of medical conditions. In certain embodiments, the received data is not preselected as being relevant to one or more of the plurality of medical conditions. In some embodiments, the method further comprises receiving additional data corresponding to one or more additional patients and updating the causal network model based on the additional data. In certain embodiments, the causal network model is generated based solely on the generated standardized data.

[0008] In some embodiments, the method further includes determining a subnetwork from the causal network model, one or more variables in the subnetwork associated with the selected medical condition, and exploring relationships in the subnetwork to determine one or more predictors for the selected medical condition. In certain embodiments, the one or more predictors for the selected medical condition indicate medical conditions that co-occur with the selected medical condition. In certain embodiments, the scope of the subnetwork is determined based on the one or more variables associated with the selected medical condition and the strength of the relationship between the one or more variables and other variables in the causal network model. In certain embodiments, the subnetwork includes one or more variables associated with the selected medical condition, a first set of other variables each having a first degree relationship with the one or more variables, and a second set of other variables each having a second degree relationship with the one or more variables. In some embodiments, at least one of the one or more predictors is previously unknown. In some embodiments, at least one of the one or more predictors is newly identified as a predictor for the medical condition. In certain embodiments, the number of predictors is less than the number of variables.

[0009] In some embodiments, the method further comprises displaying the one or more predictors in a user interface, the display comprising a graphical representation of the one or more variables, the one or more predictors, and the relationships between the one or more variables and the one or more predictors. In some embodiments, the method further comprises displaying a graphical representation of the subnetwork in the user interface. In some embodiments, the method further comprises ranking the one or more predictors based on the strength of the relationship between the one or more variables and the one or more predictors.

[0010] In certain embodiments, the method further comprises determining a subnetwork from the causal relationship network model, one or more variables in the subnetwork being associated with a selected drug, and probing the subnetwork to determine one or more predictors associated with the selected drug. In some embodiments, the one or more predictors associated with the selected drug are indicative of a drug to be administered in conjunction with the selected drug. In some embodiments, the one or more predictors are indicative of an adverse drug interaction between the selected drug and one or more other drugs. In some embodiments, the scope of the subnetwork is determined based on the one or more variables associated with the selected drug and the strength of the relationship between the one or more variables and other variables in the causal relationship network model. In some embodiments, the subnetwork comprises one or more variables associated with the selected drug, each having a first set of other variables having a first degree of relationship to the one or more variables and each having a second set of other variables having a second degree of relationship to the one or more variables. In other embodiments, at least one of the one or more predictors is previously unknown. In some embodiments, at least one of the one or more predictors is newly identified as a predictor of the medical condition. In certain embodiments, the number of predictors is less than the number of variables.

[0011] In certain embodiments, the causal relationship network model is generated based on at least 50 variables.

[0012] In certain embodiments, the causal relationship network model is generated based on at least 100 variables.

[0013] In certain embodiments, the causal relationship network model is generated based on at least 1000 variables.

[0014] In certain embodiments, the causal relationship network model is generated based on at least 100,000 variables.

[0015] In certain embodiments, the causal relationship network model is generated based on at least 50 variables to 1,000,000 variables.

[0016] In other embodiments, the causal relationship network is generated based on data from 50 patients to 1,000,000 patients.

[0017] In other embodiments, the data comprises information from a patient electronic health record.

[0018] In certain embodiments, the received data further includes at least one of the following information for at least some of the plurality of patients: patient demographic data, medical history, patient family medical history, active medication information, past medication information that is not active, allergy information, immune status information, laboratory test results, radiological images, vital sign information, patient weight, billing information, lifestyle information, habit information, insurance claim information, and pharmacy information. In some embodiments, the patient demographic data includes at least one of patient age, patient race, and patient ethnicity.

[0019] In certain embodiments, the received data includes information from patient charts. In some embodiments, the information from patient charts includes at least one of health care professional notes, health care professional observations, administration of medications and treatments, order of administration of medications and treatments, test results, and x-rays.

[0020] In certain embodiments, the received data includes patient discharge information. In some embodiments, the patient discharge information includes at least one of diagnosis codes, treatment codes, insurance codes, diagnosis related group codes, and international classification of diseases codes.

[0021] In certain embodiments, the received data relates to a plurality of patients from a selected hospital. In certain embodiments, the received data relates to a plurality of patients from a selected geographic region.

[0022] In some embodiments, generating a causal relationship network model involving the variables for the plurality of patients based on the generated normalized data using a Bayesian network algorithm includes forming a library of network fragments based on the variables by a Bayesian fragment enumeration process, forming a population of trial networks, each trial network constructed from a different subset of network fragments from the library, and evolving each trial network by local transformations via simulated annealing to globally optimize the population of trial networks to produce a consistent causal relationship network model. In some embodiments, generating a causal relationship network model involving the variables for the plurality of patients based on the generated normalized data using a Bayesian network algorithm further includes in silico simulation of the consistent causal relationship network model based on the input data to provide a predictive confidence level for one or more causal relationships within the resulting causal relationship network model.

[0023] In another aspect, the present application relates to a computer-implemented method of using a causal relationship network model. The method includes receiving a selection of a medical condition from a plurality of medical conditions, determining a subnetwork from a computer-generated causal relationship network model, the causal relationship network model generated from patient data using a Bayesian network algorithm and comprising a plurality of variables, the variables including variables related to the plurality of medical conditions, the causal relationship network model based on the selected medical condition, the subnetwork including one or more variables related to the selected medical condition, traversing the subnetwork to identify one or more predictors for the selected medical condition, and storing the one or more predictors for the selected medical condition.

[0024] In certain embodiments, the selection of the medical condition is received from a user through a user interface. In some embodiments, at least one of the one or more predictors is previously unknown. In some embodiments, at least one of the one or more predictors is newly identified as a predictor of the medical condition. In some embodiments, the number of predictors is less than the number of variables.

[0025] In some embodiments, the method further includes applying a regression algorithm to the predictors to determine a relationship of each predictor to the selected medical condition or a drug. In some embodiments, the method further includes displaying the predictors on the user interface, the display including one or more selected variables, one or more predictors, and a graphical representation of the relationship among the one or more selected variables and the one or more predictors. In some embodiments, the method further includes displaying a graphical representation of the subnetwork on the user interface. In some embodiments, the method further includes ranking the one or more predictors based on the strength of the relationship between the one or more selected variables and the one or more predictors.

[0026] In certain embodiments, the one or more predictors are related to one or more medical drugs. In certain embodiments, the predictors are related to one or more medical conditions.

[0027] In another aspect, the present application relates to a computer-implemented method using a causal relationship network model. The method includes receiving an inquiry relating to a medical condition from a plurality of medical conditions, determining a subnetwork from a computer-generated causal relationship network model, the causal relationship network model generated from patient data using a Bayesian network algorithm and comprising a plurality of variables, the variables comprising variables relating to a plurality of medical conditions, the causal relationship network model based on the medical condition of the inquiry, the subnetwork comprising one or more variables relating to the medical condition of the inquiry, exhaustively investigating the subnetwork to identify one or more predictors for the medical condition of the inquiry, and storing the one or more predictors for the medical condition of the inquiry. In certain embodiments, the inquiry received from the user comprises information relating to a medical condition and / or a medical drug.

[0028] In another aspect, the present application relates to a computer-implemented method using a causal relationship network model. The method includes receiving an inquiry relating to a medical drug, determining a subnetwork from a computer-generated causal relationship network model, the causal relationship network model generated from patient data using a Bayesian network algorithm and comprising a plurality of variables, the variables comprising variables relating to a plurality of medical drugs, the causal relationship network model based on the medical drug, the subnetwork comprising one or more variables relating to the medical drug, exhaustively investigating the subnetwork to identify one or more predictors for the medical drug, and storing the one or more predictors for the medical drug.

[0029] In yet another aspect, the present application relates to a system for generating a causal relationship network model based on patient data. The system includes a data receiving module configured to receive data relating to a plurality of patients, the data comprising diagnostic information and / or treatment information for each patient, a parsing module configured to parse the data to generate standardized data for a plurality of variables, wherein, for each patient, the standardized data is generated for more than one variable, and a processor-implemented relationship network module configured to generate a causal relationship network model involving the plurality of variables based on the generated standardized data using a Bayesian network algorithm, the causal relationship network model comprising variables relating to a plurality of medical conditions. In certain embodiments, the causal relationship network model comprises relationships indicative of one or more predictors for each of the plurality of medical conditions.

[0030] In yet another aspect, the present application provides a system for using a causal relationship network model based on patient data. The system includes a data receiving module configured to receive information related to a medical condition, a subnetwork module configured to determine a subnetwork from a computer generated causal relationship network model, the causal relationship network model generated from patient data using a Bayesian network algorithm and comprising a plurality of variables, the variables including variables related to a plurality of medical conditions, the causal relationship network model based on a medical condition, the subnetwork comprising one or more variables related to the medical condition, and a variable identification module configured to study the subnetwork comprehensively and identify one or more predictors of the medical condition. In certain embodiments, the computer generated causal relationship network model is generated using the methods disclosed in certain embodiments.

[0031] In yet another aspect, the present application provides a non-transitory machine readable storage medium storing at least one program which, when executed by at least one processor, causes the at least one processor to carry out any of the methods disclosed as part of certain embodiments.

[0032] In another aspect, the present application provides a system for generating predictors of a medical condition. The system includes a causal relationship network model generator configured to receive data corresponding to a plurality of patients, the data including diagnostic information and / or treatment information for each patient, parse the data to generate standardized data for a plurality of variables, wherein for each patient, the standardized data is generated for more than one variable, and generate a causal relationship network model associating each variable with one or more of the plurality of variables based on the generated standardized data using a Bayesian network algorithm, the causal relationship network model including variables related to a plurality of medical conditions. In certain embodiments, the causal relationship network model includes relationships indicative of one or more predictors for each of the plurality of medical conditions. In some embodiments, the system further includes a subnetwork selection module configured to receive information related to a medical condition from a user via a user interface, determine a subnetwork from the causal relationship network model, the subnetwork including one or more variables related to the medical condition, study the subnetwork comprehensively to identify one or more predictors for the medical condition, and store the one or more predictors for the medical condition.

[0033] In another aspect, the present application provides a system for generating a causal relationship network model based on patient data. The system includes a data receiving module executed by a first processor configured to receive data corresponding to a plurality of patients, the data including diagnostic information and / or treatment information for each patient, and to parse the data to generate standardized data for a plurality of variables, wherein the standardized data is generated for more than one variable for each patient, and a causal relationship network module executed by one or more other processors configured to generate a causal relationship network model involving the plurality of variables based on the generated standardized data using a Bayesian network algorithm, the causal relationship network model including variables related to a plurality of medical conditions.

[0034] Throughout this application, all numerical values presented in lists of values, such as those above, can also be intended as upper or lower limits of a range as part of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0035] The disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings in which like reference numerals refer to similar elements unless otherwise specified.

[0036] Figure 1 is a schematic network diagram depicting a healthcare analytics system according to an embodiment.

[0037] Figure 2 is a block diagram schematically depicting a healthcare analytics system according to an embodiment in terms of modules.

[0038] Figure 3 is a flowchart of a healthcare analytics method according to an embodiment by generating a relationship network model.

[0039] Figure 4 is a flowchart of a method of using a relationship network model according to an embodiment.

[0040] Figure 5 is a flowchart of a healthcare analytics method in Example 1.

[0041] Figure 6 is an example dataset from a healthcare entity used in Example 1.

[0042] Figure 7 schematically depicts a relationship network model generated in Example 1.

[0043] Figure 8 schematically depicts a subnetwork of the relationship network model focusing on heart failure & shock and renal failure. Figure 7

[0044] Figure 9A ​A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted. Figure 7 A subnetwork of the relationship network model.

[0045] Figure 9B A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0046] Figure 10A A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted. Figure 9A A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0047] Figure 10B A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0048] Figure 11 A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0049] Figure 12A A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0050] Figure 12B A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0051] Figure 13 A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0052] Figure 14A A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0053] Figure 14B A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0054] Figure 15A A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0055] Figure 15B A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0056] Figure 16 A Kidney Failure & Heart Failure & Shock subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted. Figure 8 A second part of the Kidney Failure subnetwork focusing on the association between Kidney Failure and Heart Failure & Shock is schematically depicted.

[0057] Figure 17A The kidney failure subnetwork is schematically depicted, focusing on kidney failure and highlighting the association between bronchitis & asthma and kidney failure. Figure 7 Subnetworks of the relational network model.

[0058] Figure 17B The association between bronchitis & asthma and kidney failure in the kidney failure subnetwork is schematically depicted, along with a regression model for the relationship between bronchitis & asthma and kidney failure.

[0059] Figure 18A A graph of the association between kidney failure and bronchitis and asthma, generated from the relational network model in Example 1.

[0060] Figure 18B A fourfold graph of the association between kidney failure and bronchitis and asthma, generated from the relational network model in Example 1.

[0061] Figure 19A A diagram schematically depicting a pathway based on clinical studies, explaining a new association between bronchitis & asthma and kidney failure determined from the kidney failure subnetwork of Example 1.

[0062] Figure 19B A diagram schematically depicting the potential molecular mechanisms underlying the new association between bronchitis & asthma and kidney failure determined from the kidney failure subnetwork of Example 1.

[0063] Figure 20 The relational network model of Example 1 and the selected red blood cell disorder subnetwork of Example 2 are schematically depicted.

[0064] Figure 21 A block diagram of a computing device that can be used for some embodiments of the health care analytics systems and methods described herein. DETAILED DESCRIPTION

[0065] In recent years, there has been an explosion in the amount of health care data due to the drive to adopt electronic health records. This health care data can be leveraged to open new avenues to advance health care by improving patient care and creating new efficiencies in the delivery of care. For example, understanding the variation in treatment outcomes due to patient-specific molecular and clinical factors enables the creation of precise models in medicine.

[0066] The primary focus of big data efforts in health care has been on better management and curation of health care data, and data mining to test hypotheses. Conventional analysis of health data is limited by reliance on long-held assumptions of biological or clinical phenotypes, or other assumptions that underlie the analysis.

[0067] Embodiments described herein include systems, methods, and computer readable media for health care analytics. Some embodiments use a Bayesian network algorithm to produce a causal relationship network model based on data related to various areas of health care, such as patient care. The data used to produce the relationship network model is a large set of data that is not pre-selected or pre-filtered for relevance. Further, the production of the causal relationship network model does not rely on assumptions about which variables are relevant or irrelevant, or prior knowledge about relationships between variables. This unbiased approach enables embodiments of the methods and systems to construct network models that depict associations supported by the data set and are unbiased by known clinical studies. Thus, in contrast to conventional approaches where data is pre-selected for relevance and involves prior knowledge about relationships between variables, the networks resulting from some embodiments are more likely to include new interactions between variables that were not previously known to the health care community, or that were not previously studied or explored by the health community. Some embodiments relate to methods and systems for patient data modeling that are fully data driven and unbiased by current knowledge. Such data driven and unbiased models can be used to discover new and often surprising trends in disease outcomes. Some embodiments can be used to identify unsuspected comorbidities, and to develop improved treatment strategies and experimental protocols.

[0068] In some embodiments, data comprising health care related information is obtained or received. The received data is processed and parsed to produce standardized data for a plurality of variables. A causal relationship network model is produced based on the variables using a Bayesian network algorithm. In some cases, the relationship network model includes a plurality of medical conditions and a plurality of medical drugs, and indicates relationships between the medical conditions and the drugs.

[0069] Some embodiments include methods that use the produced causal relationship network model. For example, some embodiments include receiving information from a user related to a medical condition or a medical drug. Based on the received information, a sub-network is determined from the produced causal relationship network model. The sub-network is thoroughly studied or explored to determine a predictive factor for the medical condition or drug of interest. The predictive factor can be a significant factor that has an impact on the medical condition or drug.

[0070] definition

[0071] Certain terms used herein are intended to be specifically defined as set forth below, but not elsewhere in the specification.

[0072] The term "diagnostic / treatment information" refers to any information that encodes a diagnosis made or information that describes which treatment was provided.

[0073] The term "medical condition" refers to any pathological condition, disease, and / or illness that can present with symptoms or signs that affect a person.

[0074] The term "medical drug" or "drug" refers to any pharmaceutical, medicinal, therapeutic, and / or chemical substance that can be used to cure, treat, and / or prevent a medical condition and / or to diagnose a medical condition.

[0075] The term "predictor" refers to a variable in a mathematical formula, algorithm, or decision support tool that can be used to predict an outcome. A mathematical formula, algorithm, or decision support tool can use multiple predictors to predict an outcome.

[0076] The following description presents systems and methods for health care analytics in a manner such that any person skilled in the art can practice and use the systems and methods. Various changes to the embodiments will be apparent to those skilled in the art and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the application. Further, in the following description, numerous specific details are set forth for the purpose of explanation. However, one of ordinary skill in the art will recognize that the application can be practiced without the use of these specific details. In other instances, well-known structures and processes have not been described in detail in order not to unnecessarily obscure the description of the application. Thus, the present disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0077] Figure 1 A network diagram is illustrated that depicts an example system 100 that can be included in part or in whole in a health care analytics offering in accordance with embodiments of the application. The system 100 can include a network 105, a client device 110, a client device 115, a client device 120, a client device 125, a server 130, a server 135, a database 140, and a database server 145. The client devices 110, 115, 120, 125, the servers 130, 135, the database 140, and the database server 145 are each in communication with the network 105.

[0078] In one embodiment, one or more portions of the network 105 can be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a cellular telephone network, a wireless network, a WiFi network, a WiMax network, any other type of network, or a combination of two or more such networks.

[0079] Examples of client devices include, but are not limited to, workstations, personal computers, general purpose computers, Internet appliances, notebooks, desktops, multiprocessor systems, set top boxes, network PCs, wireless devices, portable devices, wearable computers, cellular or mobile phones, personal digital assistants (PDAs), smart phones, tablet computers, ultrabooks, netbooks, multi-processor systems, microprocessor-based or programmable consumer electronics, minicomputers, and the like. Each of the client devices 110, 115, 120, 125 can be connected to the network 105 through a wired or wireless connection.

[0080] In exemplary embodiments, the health care analytics system included on the client devices 110, 115, 120, 125 can be configured to locally perform some of the functions described herein, while the server 130, 135 performs other functions described herein. For example, the client devices 110, 115, 120, 125 can receive patient data and parse the patient data, while the server 135 can generate a causal relationship network. In another example, the client devices 110, 115, 120, 125 can receive a selection of a criterion, while the server 135 can determine a subnetwork from the causal relationship network, thoroughly investigate the subnetwork to identify a predictive factor for the selected criterion, and store the identified predictive factor. In yet another example, the client devices 110, 115, 120, 125 can receive a selection of a criterion, the server 135 can determine a subnetwork from the causal relationship network and thoroughly investigate the subnetwork to identify a predictive factor for the selected criterion, and the client devices 110, 115, 120, 125 can store the identified predictive factor.

[0081] In alternative embodiments, the client devices 110, 115, 120, 125 can perform all of the functions described herein. For example, the client devices 110, 115, 120, 125 can receive patient data, parse the patient data, and generate a causal relationship network based on the patient data. In another example, the client devices 110, 115, 120, 125 can receive a selection of a criterion, such as a medical condition or a medical drug, determine a subnetwork from the causal relationship network based on the selected criterion, thoroughly investigate the subnetwork to identify a predictive factor for the selected criterion, and store the identified predictive factor.

[0082] In another alternative embodiment, the healthcare analytics system can be included on the client devices 110, 115, 120, 125 and the server 135 performs the functions described herein. For example, the server 135 can receive patient data, parse the patient data, and generate a causal relationship network based on the patient data. In another example, the server 135 can receive a selection of a criterion, e.g., a medical condition or a medical drug, determine a subnetwork from the causal relationship network based on the selected criterion, thoroughly investigate the subnetwork to identify a predictive factor for the selected criterion, and store the identified predictive factor.

[0083] In some embodiments, the servers 130 and 135 can be part of a distributed computing environment in which some tasks / functions are distributed between the servers 130 and 135. In some embodiments, the servers 130 and 135 are part of a parallel computing environment in which the servers 130 and 135 perform tasks / functions in parallel to provide the computational and processing resources needed to generate the causal relationship network models described herein.

[0084] In some embodiments, the servers 130, 135, the database 140, and the database server 145 are each connected to the network 105 by a wired connection. Alternatively, one or more of the servers 130, 135, the database 140, or the database server 145 can be connected to the network 105 by a wireless connection. Although not shown, the database server 145 can be directly connected to the database 140, or the servers 130, 135 can be directly connected to the database server 145 and / or the database 140. The servers 130, 135 include one or more computers or processors configured to communicate with the client devices 110, 115, 120, 125 over the network. The servers 130, 135 host one or more applications or websites that are accessed by the client devices 110, 115, 120, and 125 and / or facilitate access to the contents of the database 140. The database server 145 includes one or more computers or processors configured to facilitate access to the contents of the database 140. The database 140 includes one or more storage devices for storing data and / or instructions for use by the servers 130, 135, the database server 145, and / or the client devices 110, 115, 120, 125. The database 140, the servers 130, 135, and / or the database server 145 can be located in one or more locations that are geographically distributed from each other and from the client devices 110, 115, 120, 125. Alternatively, the database 140 can be included within the servers 130 or 135, or the database server 145.

[0085] Figure 2is a block diagram 200 showing a health care analytics system implemented in modules according to an example embodiment. In some embodiments, the modules include a data module 210, a parsing module 220, a relationship-network module 230, a subnetwork module 240, and a predictor module 250. In example embodiments, one or more of the modules 210, 220, 230, 240, and 250 are included in the server 130 and / or the server 135, while other of the modules 210, 220, 230, 240, and 250 can be provided in the client devices 110, 115, 120, 125. For example, the data module 210 can be included in the client devices 110, 115, 120, 125, while the parsing module 220, the relationship-network module 230, the subnetwork module 240, and the predictor module 250 are provided in the server 130 or the server 135. In another example, the data module 210 can be included in the client devices 110, 115, 120, 125, while the parsing module 220 and the relationship-network module 230 are provided in the server 130, and the subnetwork module 240 and the predictor module 250 are provided in the server 135. In yet another example, part of the functionality of the relationship-network module 230 can be performed by the server 130, and other parts of the functionality of the relationship-network module 230 can be performed by the server 135.

[0086] In alternative embodiments, the modules can be implemented in any of the client devices 110, 115, 120, 125. The modules can include one or more units of software components, programs, applications, apps, or other code bases or instructions configured to be executed by one or more processors included in the client devices 110, 115, 120, 125. In some embodiments, the modules 210, 220, 230, 240, and 250 can be downloaded from a website. In other embodiments, the modules 210, 220, 230, 240, and 250 can be installed from an external hardware component, such as an external storage component (e.g., a USB drive, a thumb drive, a CD, a DVD, etc.).

[0087] Although the modules 210, 220, 230, 240, and 250 are shown as distinct modules in Figure 2 it should be understood that the modules 210, 220, 230, 240, and 250 can be executed as fewer or more modules than shown. It should be understood that any of the modules 210, 220, 230, 240, and 250 can communicate with one or more external components, such as a database, a server, a database server, or other client devices.

[0088] The data module 210 may be a hardware-implemented module configured to receive and manage data. The parsing module 220 may be a hardware-implemented module configured to process, parse, and analyze received data for multiple variables. The relationship-network module 230 may be a hardware-implemented module configured to generate a causal network model involving multiple variables from the received data using a Bayesian network algorithm. Generating a causal network model may require considerable processor power; therefore, in some embodiments, the functionality of the relationship-network module may be implemented by the server 130 and the server 135. The subnetwork module 240 may be a hardware-implemented module configured to manage the causal network module and determine, from the causal network model, a subnetwork related to the information received from the user. The predictor module 250 may be a hardware-implemented module configured to comprehensively explore the subnetwork to identify one or more predictors corresponding to the information received from the user.

[0089] Figure 3 An example flow chart 300 of a method for generating a causal network model according to an embodiment is illustrated. In block 302, data corresponding to a plurality of patients is received. In some embodiments, the data is received by the data module 210 (see Figure 2 ). In some embodiments, the data includes diagnostic information and / or treatment information for each patient. The data may include information such as any of patient demographic data, medical history, patient family medical history, active medication information, inactive past medication information, allergy information, immune status information, laboratory test results, radiological images, vital signs information, patient weight, billing information, lifestyle information, habit information, insurance claim information, pharmacy information, and the like. Patient demographic data may include patient age, patient race, and patient ethnicity. The data may also include or alternatively include information from the patient's medical record, such as notes from a healthcare professional, observations from a healthcare professional, administration of medications and treatments, the order in which medications and treatments were administered, test results, x-rays, and the like. The data may also include or alternatively include patient discharge information, such as diagnosis codes, treatment codes, insurance charge codes, diagnosis-related group codes, International Classification of Disease codes, and the like. Data module 210 (see Figure 2 ) can extract or obtain data from entities that manage and utilize various healthcare data, information, and / or statistics. The data can be obtained from a variety of sources, such as publicly available resources, commercial entities that collect data, healthcare providers, etc. In some embodiments, the data may not be pre-selected or pre-determined to be relevant to multiple medical conditions or drugs. For illustrative examples of input data, see the following references Figure 6 discussed Example 1 .

[0090] In block 304, the data received in block 302 is parsed to produce standardized data for a plurality of variables. In some embodiments, the data is parsed by parsing module 220 (see Figure 2 ). Standardized data is produced for more than one variable for each patient. Standardization of the data can include reducing the data to its canonical form, and / or organizing the data into a form that facilitates further use. In some embodiments, parsing the data further includes filtering the data and imputation of the data. Filtering the data can include removing data points based on criteria such as completeness and accuracy of the data points. Imputation of the data can include replacing missing data points with suitable substitute values.

[0091] In block 306, a causal relationship network model is produced for the plurality of variables based on the produced standardized data. In some embodiments, the causal relationship network model is produced using causal relationship network module 230 (see Figure 2 ). In some embodiments, a Bayesian network algorithm is used to produce the causal relationship network model involving the plurality of variables. In some embodiments, the produced causal relationship network model includes variables involving a plurality of medical conditions and / or drugs. In some cases, the causal relationship network model is produced using a programmed computing system including memory for network model construction code and one or more processors for executing the network model construction code. The causal relationship network model can include relationships indicating one or more predictors for each of a plurality of medical conditions or drugs. The causal relationship network model can be produced based on the standardized data alone. In some embodiments, the relationship network model is an artificial intelligence based network model.

[0092] The causal relationship network model can be produced based on as many variables as desired or appropriate for meaningful data analysis. For example, in some embodiments, the model can be produced based on at least 50 variables. In other embodiments, the model can be produced based on at least 100 variables, at least 1000 variables, at least 10,000 variables, at least 100,000 variables, or at least 1,000,000 variables. As discussed above, the variables correspond to data determined / extracted from the unprocessed / raw input data set. In some embodiments, the variables become nodes in the causal relationship network model.

[0093] The causal relationship network model can be generated based on data from as many patients as desired or appropriate for meaningful data analysis. For example, in some embodiments, the model can be generated based on data from at least 50 patients. In other embodiments, the model can be generated based on data from at least 100 patients, from at least 1000 patients, from at least 10,000 patients, from at least 100,000 patients, or from at least 1,000,000 patients.

[0094] The method can include further steps. For example, additional data corresponding to one or more other patients can be received, at which time the causal relationship network model can be updated or regenerated based on the additional data.

[0095] In some embodiments, a graphical representation of some or all of the generated causal relationship network model can be presented to a user. In some embodiments, the generated relationship network model is stored for later use.

[0096] It should be noted that many different artificial intelligence-based platforms or systems can be used to generate the causal relationship network model using a Bayesian network algorithm. Some example embodiments use a system commercially available from GNS (Cambridge, MA) called REFS TM (Reverse Engineering / Forward Simulation). AI-based systems or platforms suitable for performing some embodiments use data algorithms to build a network of causal relationships among variables based on an input data set alone, without taking into account existing knowledge about any potential, established, and / or validated relationships.

[0097] For example, the REFS TM AI-based informatics platform can take standardized input data and perform trillions of calculations quickly to determine how data points interact in a system. Based on the REFS TM AI-based informatics platform performs a reverse engineering process to attempt to form a computer-implemented relationship network model based on the input data that quantitatively represents relationships among various health conditions and predictors. In addition, hypotheses can be generated and simulated quickly based on the computer-implemented relationship network model to obtain predictions about the hypothesis with associated confidence levels. More details about an example of generating a causal relationship model using the REFS platform are provided below.

[0098] Figure 4An example flowchart 400 of a method of using a causal relationship network model according to embodiments is illustrated. The method 400 can be implemented in an exemplary system and can be described as a method of using a health care analytics system. The method of using a causal relationship network model can also be referred to as a method of interpreting results from a causal relationship network model.

[0099] In block 402, information is received from a user. In some embodiments, the information is received by the data module 210 (see FIG. 2). In some embodiments, the data module used to receive a data set for generating a relationship network model is different or separate from the data module used to receive information from a user for using the generated relationship network model. Figure 2 ) In some embodiments, the data module used to receive a data set for generating a relationship network model is different or separate from the data module used to receive information from a user for using the generated relationship network model.

[0100] The information received from the user can include a selection of one or more medical conditions or one or more medical drugs. In some embodiments, the user can be presented with a list of medical conditions and / or medical drugs from which to select.

[0101] The information received from the user can be in the form of a query regarding one or more medical conditions or one or more medical drugs. In some embodiments, the user can enter text in a search / query field.

[0102] In some embodiments, a graphical representation of a portion or all of the causal relationship network model can be displayed. In some embodiments, the information received from the user can be a selection of one or more nodes displayed in the graphical representation of a portion or all of the causal relationship network model.

[0103] In block 404, a subnetwork is determined from the causal relationship network model based on the information received from the user in block 402. In some embodiments, the subnetwork is determined by the subnetwork module 240. The causal relationship network model from which the subnetwork is determined has been generated from patient data using a Bayesian network algorithm as described above and can include a plurality of variables, including variables related to a plurality of medical conditions. One or more variables included in the determined subnetwork correspond to the information received from the user. For example, if the information received relates to one or more medical conditions, one or more variables in the subnetwork relate to the one or more medical conditions. As another example, if the information received relates to one or more drugs, one or more variables in the subnetwork relate to the one or more drugs.

[0104] The extent of the subnetwork can be determined based on one or more variables associated with the selected one or more medical conditions or one or more drugs and the strength of the relationship between the one or more variables and other variables in the causal network model. The subnetwork can include one or more variables associated with the medical condition or drug of interest and a first set of other variables each having a first degree of relationship with the one or more variables. In some embodiments, the subnetwork can further include a second set of other variables each having a second degree of relationship with the one or more variables.

[0105] In block 406, the subnetwork is fully explored to identify predictors. In some embodiments, the subnetwork is fully explored by the predictor module 250 (see Figure 2 ). Predictors are identified as relevant to the information received from the user. A predictor is a factor, data point, or node that has a causal relationship with the medical condition or medical drug of interest. For example, renal failure can be a predictor of heart failure. After one or more predictors are identified from the causal network model, the identified predictors can be used in traditional statistical or regression analysis to determine the significance of the predictors relative to the medical condition or drug of interest. One or more predictors for the medical condition of interest can indicate medical conditions that co-occur with the medical condition of interest. One or more predictors can indicate drugs that are administered in combination with the drug of interest, or indicate adverse drug interactions. The identified predictors can be previously unknown or new predictors for the medical condition or drug. The identified predictors can be newly identified as predictors for the medical condition or drug. The number of predictors can be less than the number of variables or nodes in the subnetwork. In box 408, the identified predictors are stored.

[0106] The above method may include further steps. For example, the identified predictors may be displayed in a user interface through a graphical representation of the variables, predictors, and relationships between the variables and predictors. The determined subnetwork may be displayed in a user interface.

[0107] The Generation of Causal Network Models

[0108] The following is a more detailed explanation of the steps for generating a causal network model involving multiple variables based on the generated standardized data using the Bayesian network algorithm for the REFS AI-based informatics system (for illustrative purposes only): Figure 3 However, one of ordinary skill in the art will recognize that other systems employing Bayesian analysis may be used.

[0109] Normalized data for multiple variables are inputted into the REFS system as an input dataset. The REFS system forms a library of "network fragments" that include variables that drive associations and relationships in a healthcare system (e.g., medical conditions, medical drugs, discharge codes). The REFS system selects a subset of network fragments from the library and constructs an initial trial network from the selected subset. The AI-based system also selects a different subset of network fragments from the library to construct another initial trial network. Ultimately, a whole set of initial trial networks (e.g., 1000 networks) are formed from different subsets of network fragments from the library. This process is referred to as parallel whole-set sampling. Each trial network in the whole set of trial networks is evolved or optimized by adding, subtracting, and / or replacing additional network fragments from the library. More details about the formation of the network fragment library, the formation of trial networks, and network evolution are provided below. If additional data is obtained, the additional data can be incorporated into network fragments in the library and into the whole set of trial networks through evolution of each trial network. Upon completion of the optimization / evolution process, the whole set of trial networks can be described as a resulting relational network model.

[0110] The whole set of resulting relational network models can be used to simulate the behavior of associations between various medical conditions and / or drugs. The simulations can be used to predict interactions between medical conditions and / or medical drugs, which can be validated using clinical studies and experiments. In addition, quantitative parameters of relationships in the resulting relational network model can be extracted using simulation functions by applying perturbations to each node individually while observing the effects on other nodes in the resulting relational network model.

[0111] The building blocks of REFS incorporate multiple data types from an unlimited number of data forms, e.g., continuous, discrete, Boolean. The generation of a whole set of models requires considerable processing power. In some embodiments, a parallel IBM Blue Gene machine with 30,000+ processors is used to generate a whole set of models. The resulting networks are capable of high-throughput computer simulation testing of hypotheses. The networks also include a hierarchical order that provides verifiable hypotheses and a prediction confidence metric.

[0112] As described above, pre-processed data is used to construct a library of network fragments. A network fragment defines a quantitative, continuous relationship among a small set of all possible measured variables (e.g., a set of 2-3 members or a set of 2-4 members) (input data). The relationship between variables in a fragment can be linear, logistic, multinomial, explicit or implicit homozygous, etc. Each relationship in a fragment is assigned a Bayesian probability score that reflects how likely the input data is to give the candidate relationship and also penalizes the relationship for its mathematical complexity. By scoring all possible pairings and three-way relationships (and in some embodiments, four-way relationships) inferred from the input data, the most likely fragments (possible fragments) in the library can be identified. Quantitative parameters of the relationships are also computed based on the input data and stored for each fragment. Various model types can be used in fragment counting, including but not limited to linear regression, logistic regression, (analysis of variance) ANOVA models, (analysis of covariance) ANCOVA models, non-linear / multinomial regression models, and even non-parametric regressions. Previous assumptions about model parameters can assume a Gull distribution or a Bayesian Information Criterion (BIC) penalty related to the number of parameters used in the model. During network inference, each network in the full set of initial trial networks is constructed from a subset of fragments in the fragment library. Each initial trial network in the full set of initial trial networks is constructed with a different subset of fragments from the fragment library.

[0113] The model is evolved or optimized by determining the most likely factorization and the most likely parameters given the input data. This can be described as "learning a Bayesian network," or in other words, finding the network that best matches the input data given a training set of input data. This can be done by using a scoring function that evaluates each network with respect to the input data.

[0114] The likelihood of a factorization given the input data is determined using a Bayesian framework. Bayes' rule states that the posterior probability of a model M given data D, P(D|M), is proportional to the product of the posterior probability of data given the model assumption, P(D|M), multiplied by the prior probability of the model, P(M), given that the probability of data, P(D), is constant across models. This is expressed in the following equation:

[0115]

[0116] The posterior probability of data given the model is assumed to be the integral of the likelihood of data over the previous parameter distribution:

[0117] P(D|M) = ∫ P(D|M(Θ)) P(Θ|M) dΘ.

[0118] Assuming all model possibilities are equally likely (i.e., P(M) is constant), the posterior probability of a model M given data D can factorize into a product of integrals over the parameters of each local network fragment M i , as follows:

[0119]

[0120] Note that in the above equation, the leading constant term has been omitted. In some embodiments, the Bayesian Information Criterion (BIC), which takes the negative log of the model posterior probability P(D|M), can be used to "score" each model, as follows:

[0121]

[0122] where the total score S tot for a model M is the sum of the local scores S i for each local network fragment. BIC further gives an expression for determining each individual network fragment score:

[0123]

[0124] where K(M i ) is the number of fitted parameters in model Mj, and N is the number of samples (data points). S MLE (M i ) is the negative log of the likelihood function for the network fragment, which can be computed from the functional relationship for each network fragment. For the BIC score, the lower the score, the more likely the model is to fit the input data.

[0125] Optimizing the full set of trial networks, which can be described as optimizing or evolving the network. For example, the trial networks can be evolved and optimized according to a Metropolis-Monte Carlo sampling algorithm. Simulated annealing can be used to optimize or evolve each trial network in the set by local transformations. In an example simulated annealing method, each trial network is changed by adding a network fragment from the library, by deleting a network fragment from the trial network, by subtracting a network fragment, or by otherwise changing the network topology, and then the new score of the network is computed. Generally, if the score improves, the change is kept, and if the score worsens, the change is discarded. A "temperature" parameter allows some locally worsening score changes to be retained, which helps the optimization process to avoid some local minima. The "temperature" parameter is lowered over time to allow the optimization / evolution process to converge.

[0126] All or part of the network inference process can be performed in parallel for different trial networks. Each network can be optimized in parallel on separate processors and / or separate computing devices. In some embodiments, the optimization process can be performed on a supercomputer that integrates hundreds to thousands of processors running in parallel. Information can be shared among the optimization processes performed on the parallel processors. In some embodiments, the optimization process can be performed on one or more quantum computers, which have the potential to perform certain calculations significantly faster than silicon-based computers.

[0127] The optimization process can include a network filter that removes any network from the ensemble of networks that fails to meet a threshold criterion for the overall score. The removed network can be replaced by a new initial network. In addition, any network that is not "scale free" can be removed from the ensemble of networks. After the ensemble of networks is optimized or evolved, the result can be referred to as an ensemble of generated relational network models, which can be collectively referred to as generated consensus networks.

[0128] Simulations can be used to extract quantitative parameter information about each relationship in the generated relational network models. For example, a simulation for quantitative information extraction can involve perturbing (increasing or decreasing) each node in the network by ten-fold and calculating the subsequent distribution for the other nodes in the model. Endpoints are compared by t-test, with an assumption of 100 samples / group and a significance cutoff of 0.01. The t-test statistic is the median of 100 t-tests. By using this simulation technique, an area under the curve (AUC) representing the strength of the prediction and a fold change in the computer simulated node amplitude representing the driving endpoint are generated for each relationship in the ensemble of networks.

[0129] A relationship quantification module of a local computer system can be used to direct the AI-based system to perform the perturbations and extract the AUC information and fold information. The extracted quantitative information can include the fold change and AUC for each edge connecting a parent node to a child node. In some embodiments, a custom R program can be used to extract the quantitative information.

[0130] In some embodiments, the ensemble of generated relational network models can be used by simulation for predicting responses to changes in conditions, which can subsequently be validated by clinical studies and experiments.

[0131] The output of the AI-based system can be quantitative relationship parameters and / or other simulation predictions.

[0132] Some example embodiments incorporate the Berg Interrogative Biology TMmethods performed by the Informatics Suite, which is a tool for understanding a broad range of biological processes (such as disease pathophysiology) and the key molecular drivers underlying these biological processes, including factors that enable disease processes. Some example embodiments use the Berg Interrogative Biology TM Informatics Suite to gain new insights about the interactions of diseases with other diseases, medical drugs, biological processes, and the like. Some example embodiments include systems that can incorporate at least a portion or all of the Berg Interrogative Biology TM Informatics Suite.

[0133] Example

[0134] Example 1 - Relational Network Model Generated from CMS Data: Subnetworks of Heart Failure & Shock and Renal Failure

[0135] Mathematical and statistical learning tools developed in artificial intelligence (AI) are well suited to decipher complex interaction patterns in big data. The Berg Interrogative Biology TM Informatics Suite uses Bayesian networks (BNs) in a purely data-driven fashion for integration and inference of causal effects for various data forms. This embodiment relates to the use of BNs in health care big data analytics that has a significant impact on enhancing patient care and improving health care and hospital efficiency.

[0136] New findings and new hypotheses directly informing care using the resulting causal relationship network models are formed using low resolution, publicly available data. Data relationship networks of diagnostic codes are generated based on publicly available billing data from the Centers for Medicare & Medicaid Services (CMS).

[0137] Data is extracted from CMS releases, as illustrated in the method 500 of Figure 5 Data is pre-processed by filtering, standardization, and imputation. A.I.-based model building is used to generate models from the pre-processed data, forming relationship networks in the form of diagnostic code networks.

[0138] A data set is obtained from CMS data releases. CMS releases (i.e., made publicly available on their website) involve data on patient care, insurance information, diagnostic codes, discharge codes, and fee codes for various procedures by health care providers. Figure 6A sample of CMS data for a portion of a single medical institution is shown. In this example, the data obtained included a number of health care providers and the top 100 diagnosis codes for 2011.

[0139] Berg Interrogative Biology TM Informatics Suite was used for data pre-processing and model building. The collected data was processed to extract columns from the data set containing information about diagnosis related group (DRG) codes and total number of discharges. The discharge count information was organized into a matrix of DRG codes versus hospitals. DRG codes that were missing in more than 70% of the hospitals were removed from further analysis. Hospitals that had information missing for more than 25 DRG codes were filtered out from the data set. After filtering, the data set contained 100 DRG codes and 1618 hospitals. The data matrix was median polished normalized and missing data was imputed using a "zero procedure." A causal relationship network model was built using the REFS technique. Associations in the resulting relationship network were filtered using an area under the curve (AUC) cutoff of 0.65. The resulting network was visualized using appropriate graphical environment / interface software, in particular Cytoscape, a software environment for integrated models of molecular interaction networks from the Cytoscape Consortium. Within the Cytoscape environment, network visualization is by "nodes" connected by "edges," each edge graphically depicting a relationship between the two nodes connected by the edge.

[0140] A schematic graphical depiction of the resulting relationship network based on DRG codes is shown in Figure 7 This particular relationship network contains 60 diagnoses and 88 associations / connectivity between them. Each node in the network represents the number of discharges for a particular diagnosis, and the edges (e.g., connections between nodes) represent interactions between the number of discharges associated with various diagnoses. In Figure 7 the size of the nodes corresponds to the number of discharge codes for a particular diagnosis. The interpretation of the edges and associations in the network depends on the particular diagnoses involved. For example, when the source node is diabetes and the target node is hypertension, the association represents co-morbidity. When the source node is diabetes and the target node is neuropathy, the association represents a complication of the disease. In this way, the relationship network represents interactions between a variety of DRG codes.

[0141] Sub-networks were selected to obtain relevant information from the resulting relationship network. From the causal relationship network model shown in Figure 7 a sub-network centered on "heart failure & shock" and "renal failure" was determined (shown in Figure 8In this example, "heart failure & shock" and "kidney failure" were selected due to their prominence in the mortality index compiled by the Centers for Disease Control and Prevention (CDC). According to the CDC, heart disease was the leading cause of death in 2011. Several kidney-related conditions ranked 9th, 12th, and 13th; however, all kidney-related conditions combined hold a significant place among the causes of death.

[0142] exist Figure 8 In the graphical representation of the subnetwork of , an association in the form of an arrow indicates that the number of diagnoses in the first condition and the number of diagnoses in the second condition are positively correlated. Clinically, associations can be interpreted in different ways depending on the conditions involved (e.g., an arrow pointing from heart failure & shock to simple pneumonia & pleurisy indicates that heart failure & shock causes or is secondary to simple pneumonia & pleurisy in a statistically significant proportion of patients). Each association between the annotations displayed as lines can also be described as an edge between the nodes. The width of the line forming the association between the nodes (also described as the "weight" of the line) provides a graphical indication of the strength of the relationship between the nodes. For example, the arrow pointing from heart failure & shock to simple pneumonia & pleurisy is wider than the arrow connecting heart failure & shock to respiratory infection & inflammation. This indicates that the model predicts that the relationship between heart failure & shock and simple pneumonia & pleurisy is stronger than the relationship between heart failure & shock and respiratory infection & inflammation.

[0143] As discussed below, the associations in the heart failure subnetwork and in the renal failure subnetwork used as validation of the associations appearing in the relational network model were reliable.

[0144] Heart failure is caused by conditions that reduce the heart's ability to pump blood effectively. These conditions include congenital heart defects, irregular heartbeats, coronary artery disease that narrows the arteries over time, and high blood pressure that can make the heart too weak or too stiff to pump blood effectively. The relationship between heart failure and other conditions is discussed below, as reflected in the associations between heart failure and other nodes in the subnetwork.

[0145] Respiratory Infections and Inflammation / Uncomplicated Pneumonia and Pleurisy: Heart failure is known to cause blood to flow through the body more slowly and to cause fluid retention in the kidneys. Fluid retention begins in the lower body but progresses to the lungs, causing pneumonia. Fluid accumulation in the lungs leads to increased rates of respiratory infections and respiratory tract infections. The relationship network model predicts that a diagnosis of heart failure & shock leads to or follows a diagnosis of uncomplicated pneumonia and pleurisy in a statistically significant proportion of patients, as determined by Figure 9A The relationship network model predicts that the diagnosis of heart failure & shock leads to or follows the diagnosis of respiratory infection & inflammation in a statistically significant proportion of patients, as indicated by the arrows pointing from heart failure & shock to simple pneumonia and pleurisy. Figure 9AIndicated by the arrow pointing from heart failure & shock to respiratory infection & inflammation. Figure 9B Diagram depicting the pathways by which heart failure and shock can lead to uncomplicated pneumonia and pleurisy, as well as respiratory tract infection and inflammation.

[0146] Chronic Obstructive Pulmonary Disease (COPD): COPD causes pulmonary hypertension when the lungs attempt to compensate for low oxygen concentrations in the blood by increasing blood pressure inside the lungs. Increased blood pressure inside the lungs leads to pulmonary hypertension, which strains the right ventricle and causes heart failure. Therefore, a diagnosis of heart failure can lead to a previously undiagnosed diagnosis of COPD. Therefore, an increase in the number of heart failure diagnoses will increase the number of COPD diagnoses. The relationship network model predicts that an increase in the number of diagnoses of heart failure & shock leads to an increase in the number of COPD diagnoses, as shown by Figure 10A This is indicated by the arrow pointing from heart failure & shock to COPD. This can be directly interpreted as heart failure & shock causing COPD in a statistically significant number of patients. However, Figure 10B The diagram depicts the pathway by which COPD leads to heart failure and shock, rather than the other way around. The apparent temporal reversal between this model's predictions and the known relationship between heart failure and COPD is likely due to previously undiagnosed COPD being diagnosed after heart failure. Clinicians may discover that COPD only after heart failure leads to a causal investigation. Thus, although COPD was present before heart failure, it was not diagnosed until after heart failure was diagnosed.

[0147] Arrhythmias and conduction disorders: Arrhythmias and conduction disorders can directly cause heart failure. A diagnosis of heart failure can lead to the diagnosis of these causal conditions. Therefore, an increase in the number of heart failure diagnoses will increase the number of arrhythmias and conduction disorders diagnoses. The network model predicts that an increase in the number of heart failure and shock diagnoses will lead to or follow an increase in the number of arrhythmias and conduction disorders diagnoses in a statistically significant proportion of patients, as measured by Figure 11 Indicated by the arrow pointing from heart failure and shock to arrhythmias and conduction disorders.

[0148] Gastrointestinal (GI) bleeding: Heart failure can be a result of hypovolemic shock. Hypovolemic shock occurs due to a rapid loss of circulating blood volume. The most common causes of hemorrhagic shock (a precursor to hypovolemic shock) are trauma, GI bleeding, and organ damage. The relationship network model predicts that the diagnosis of GI bleeding leads to or follows the diagnosis of heart failure & shock in a statistically significant proportion of patients, as measured by Figure 12A Indicated by the arrows pointing from GI bleeding to heart failure and shock. Figure 12B Diagram depicting the pathways by which GI bleeding leads to heart failure.

[0149] Renal failure: The association between anemia, heart problems, and kidney disease is well known, and the challenges of treating patients with cardio-renal dysfunction have been demonstrated. About one fourth of patients with kidney disease have congestive heart problems. As kidney disease worsens, the fraction of patients with heart disease increases to about 65-70%. Large studies have shown that worsening kidney disease in patients previously diagnosed with heart failure is associated with higher mortality and hospitalization rates. Thus, this association is well known. The relational network model predicts that the diagnosis of renal failure leads or follows the diagnosis of heart failure and shock in a statistically significant proportion of patients, as indicated by the arrows pointing from renal failure to heart failure & shock. Figure 13

[0150] Figure 14A and 14B More details are provided about the relational network model for the prediction of the strength of the relationship between heart failure & shock and renal failure. Figure 14A is a contingency plot of heart failure and renal failure derived from the data used to build the relational network model. The contingency plot, which can be represented as a 2x2 table, shows the deviation of the observed frequencies from the expected frequencies for the variables in the dataset. The Pearson residuals represent the distance between the expected and observed frequencies and thus allow the identification of the categories that contribute to the deviation from the expected values. In Figure 14A the contingency plot, the width of each rectangle represents the number of data points and the height and color represent the Pearson residuals. In this case, the expected values are calculated based on the null hypothesis that heart failure and renal failure are independently distributed. If this were true, the plot would be mostly light gray, indicating that the Pearson residuals are close to zero (i.e., less than 2 and greater than -2). Instead, a medium gray portion representing Pearson residuals less than -2 and a dark gray portion representing Pearson residuals greater than 2 are observed, indicating that heart failure and renal failure are not independent and thus deviate from the null hypothesis. As indicated by the Pearson residuals, the largest component of the change is for renal failure and heart failure in the high categories. Overall, in this case, Figure 14A the contingency plot indicates an interaction between heart failure and renal failure. This interaction is primarily caused by patients who have both renal failure and heart failure.

[0151] Figure 14B is a fourfold plot of heart failure and renal failure derived from the data used to build the relational network model. The fourfold plot is another representation of the deviation from independence between the target categories. In the fourfold plot, the numbers in each square represent the number of data points in that category. Figure 14B ​The impact of kidney failure on heart failure is shown. The left half of the figure indicates that high kidney failure rates have a stronger impact on high heart failure rates compared to low kidney failure rates based on comparing the size of the quarter circles. The right half of the figure indicates that there is less difference in the impact of high or low kidney failure on low heart failure rates. Overall, this figure reinforces the interaction between the target condition heart failure and kidney failure. Based on the relational network model, the relative risk of heart failure with kidney failure is calculated to be 2.57.

[0152] Figure 15A and 15B More details are provided regarding the relational network model for the strength of the relationship between heart failure & shock and G.I. bleeding. Figure 15A is a correlation plot of heart failure and G.I. bleeding. In this figure, the dark gray rectangles representing Pearson residuals greater than 4.0 and the medium gray rectangles representing Pearson residuals less than -2 indicate that G.I. bleeding and heart failure are not independent. In particular, the dark gray rectangles indicate that high rates of both conditions have the strongest interaction. Figure 15B is a four-fold plot of heart failure and G.I. bleeding. This figure indicates that high rates of heart failure are strongly related to high rates of G.I. bleeding. As shown, the relative risk of heart failure with G.I. bleeding is calculated to be 3.22.

[0153] The relational network model was generated based only on DRGs and discharge codes without any assumptions and other information regarding the relationship between heart disease & shock and various conditions (i.e., simple pneumonia and pleurisy, respiratory infections and inflammation, COPD, cardiac dysrhythmia & conduction disorders, G.I. bleeding, and kidney failure). Regardless, the heart failure and shock centered subnetwork of the generated relational network model reflects existing knowledge in the medical field regarding the relationship between heart failure & shock and other conditions, which supports the validity of the relational network model.

[0154] The diagnosis codes corresponding to kidney failure in this analysis include chronic / acute kidney failure and other renal disorders. Kidney failure / dysfunction refers to a decrease in the ability of the kidneys to remove waste from the blood. More than 10% of adults 20 years of age or older have CKD, and the cost of treating CKD is very high due to costs associated with comorbidities and quality of life factors. One study indicates that the cost for treating end-stage renal disease (ESRD) continues to increase and that the medical cost for this condition reached $30 billion in 2009. Figure 16 is depicted schematically. The associations in the subnetwork with kidney failure are studied as follows:

[0155] Kidney and urinary tract infections: The relational network model predicts a statistically significant relationship between kidney failure and kidney and urinary tract infections as indicated by Figure 16The arrow from kidney failure to urinary tract infection is shown. It is known that untreated urinary tract and kidney infections can lead to kidney failure. Although the reversal of the direction is still unexplained, the association between the conditions is represented in the relational network model. It is therefore not surprising that this association is identified.

[0156] Disorders of nutrition, metabolism and fluid / electrolytes: The kidneys play an important role in maintaining fluid and electrolyte balance. Therefore, the diagnosis of kidney failure makes it possible to track tests and diagnoses of problems of nutrition, metabolism, fluid and electrolyte balance. The relational network model predicts that the diagnosis of kidney failure leads to the diagnosis of disorders of nutrition, metabolism and fluid / electrolytes in a statistically significant proportion of patients, as by Figure 16 The arrow from kidney failure to disorders of nutrition, metabolism and fluid / electrolytes is shown.

[0157] Simple pneumonia and pleurisy: It is established that chronic kidney disease increases susceptibility to infections, and pneumonia has been shown to be an infectious complication of kidney disease. The relational network model predicts a statistically significant relationship between kidney failure and simple pneumonia & pleurisy, as by Figure 16 The association between kidney failure and simple pneumonia & pleurisy is shown in the graphical representation of the subnetwork of Figure 8 In the graphical representation of the subnetwork of, the T-shaped form of the association indicates that the number of diagnoses in the first condition and the number of diagnoses in the second condition are negatively correlated. In particular, the number of diagnoses of simple pneumonia & pleurisy and the number of diagnoses of kidney failure are negatively correlated.

[0158] With regard to the kidney failure centred subnetwork, all but one of the selected associations in the subnetwork (in particular, all but the association between bronchitis & asthma and kidney failure) are supported by existing knowledge in the medical field. This serves as further validation that the new predictions from the generated relational network model are reliable and worthy of further investigation.

[0159] The results from the generated network relational model also lead to the identification of new interactions. For example, based on the kidney failure subnetwork, a new interaction between kidney failure and bronchitis & asthma is identified. As Figure 17AThe relationship network model predicts that the diagnosis of bronchitis & asthma leads to or follows the diagnosis of renal failure in a statistically significant proportion of patients, as shown in FIG. 1. The thickness of the arrow indicates that this is a stronger relationship or association between renal failure and any of heart failure & shock, kidney & urinary tract infections, simple pneumonia & pleurisy, and various disorders of nutrition, metabolism, and fluid / electrolytes. Because the literature on the relationship between renal failure and bronchitis & asthma is not widely known or available, this interaction has potential for new discoveries in terms of drug therapy, disease causality, and / or diagnostic sequencing. Regression models were constructed to identify the strength of the interaction, as depicted in FIG. 2. The regression model indicates that bronchitis and asthma account for ~2.5% of the renal failure data (p-value 1.8 x 10 Figure 17B -10 ). The p-value is the probability of obtaining a test statistic at least as extreme as the one actually observed, given the null hypothesis. In this case, if bronchitis and asthma were completely independent, the probability of obtaining this data would be 1.8 x 10 -10

[0160] Figure 18A is a scatter plot of renal failure versus bronchitis & asthma. This plot also shows that a high rate of bronchitis & asthma increases the rate of renal failure. The relative risk of patients with asthma & bronchitis leading to renal failure is calculated to be 5.13. Figure 18A -16 The data represented in FIG. 3 has a p-value of less than 2.22 x 10 Figure 18B is a four-fold plot of renal failure versus bronchitis & asthma. This plot also shows that a high rate of bronchitis & asthma increases the rate of renal failure. The relative risk of patients with asthma & bronchitis leading to renal failure is calculated to be 5.13.

[0161] A new hypothesis was generated to link these conditions. Specifically, the hypothesis is that renal failure and / or dysfunction is caused as a side effect of treating asthma and bronchitis. Thus, as the number of cases of bronchitis and asthma increases, the number of cases of renal failure increases. This hypothesis is discussed further below.

[0162] Bronchitis and asthma are respiratory diseases in which airway narrowing is caused by inflammation. To control symptoms and reduce airway swelling, a number of drug options are available, as shown in Table 1 below. The main ingredient in many of these drugs is a long-acting β2-adrenergic agonist, and a contraindication includes hypokalemia. In fact, albuterol (a β2-agonist) is widely used to treat hyperkalemia in patients with renal failure / dysfunction. Thus, long-term use of drugs containing potent β2-agonists can lower potassium levels, leading to hypokalemia. Electrolyte imbalances have been noted in patients treated for asthma with β2-agonists.

[0163] ​​​Table 1: Clinical pharmacology of the top 10 asthma medications selected in 2011-12. Drug information obtained from Drug Index of Prescription Drugs. Prescription count information obtained from IMS Health's Survey of the Prescriptions.

[0164]

[0165]

[0166] Studies have shown that hypokalemia induces kidney injury in rats and hypokalemia has been observed to cause kidney failure in humans. In a study of 55 patients, chronic hypokalemia was associated with kidney cyst formation, which led to scarring and kidney injury resulting in renal insufficiency. Other studies have also shown that hypokalemia in patients with kidney disease increases the rate of progression to end-stage renal disease and increases mortality. Beta2-agonists can increase aldosterone levels, which in turn is associated with kidney dysfunction. Blocking the function of aldosterone results in improvement in kidney function.

[0167] Figure 19A and 19B The pathways that depict the treatment of bronchitis and asthma can cause kidney failure and / or insufficiency are presented, which support the hypothesis of a new association between bronchitis and asthma and kidney failure and explain this new association. The pathways in Figure 19A were constructed by correlating published clinical studies. Figure 19B The underlying molecular mechanisms that form the basis of the pathways are presented, which were identified to support the clinical findings in the published clinical studies. Figure 19B The protein abbreviations used in the pathways in Figure 19A The hypothesis that correlates bronchitis & asthma with kidney failure is schematically depicted. The treatment for bronchitis & asthma most often involves the use of medications containing long-acting beta2-adrenergic agonists. Long-acting beta2-adrenergic agonists are known to increase aldosterone levels, which have been shown to cause kidney failure. According to the Federal Drug Administration (FDA) label, hypokalemia is a contraindication for long-acting beta2-adrenergic agonists, and hypokalemia is a known marker for an increased rate of kidney failure. Figure 19BThe proposed molecular mechanism linking asthma & bronchitis treatment to kidney failure is schematically depicted. Long-acting beta2-adrenergic agonists increase the activity of Gs, which is involved in the production of cAMP, which in turn increases PKA. This leads to increased renin secretion from the juxtaglomerular cells. Renin catalyzes the formation of angiotensin I, which is subsequently converted to angiotensin II by ACE activity in the lungs. Elevated angiotensin II levels cause aldosterone to increase. In the adrenal cortex of the kidney, increased aldosterone leads to increased potassium excretion, water reabsorption, and sodium reabsorption, ultimately causing hypokalemia and kidney failure. Table 2 includes a list of published clinical studies supporting the pathways and molecular mechanisms shown in Figure 19A and 19B

[0168] Table 2: Published clinical studies supporting the pathways and molecular mechanisms about the link between bronchitis and asthma and kidney failure.

[0169]

[0170]

[0171]

[0172] In this CMS dataset, which includes 3000 hospitals enrolled in the Inpatient Prospective Payment System (IPPS), the total cost of treating patients with kidney disease was $2.3 billion. This CMS dataset represents 60% of total Medicare expenditures. Therefore, the total Medicare cost for 2011 for this diagnosis code alone would be $38.3 billion. In this reference, 2.5% would be attributed to side effects of medications used to treat bronchitis and asthma. Further research and refinement of treatment guidelines for asthma / bronchitis based on the relationship network model determined using Example 1 between kidney failure and bronchitis and asthma not only leads to better patient care, but also saves considerable costs. For example, using patient-level data, it is possible to identify patients at high risk for kidney side effects based on medical history and genetic factors. For high-risk patients, alternative treatment strategies or kidney function monitoring can lead to better outcomes for the patient and lower costs for payers such as Medicare. By integrating patient electronic health records with the knowledge base, such care improvements can be incorporated into the clinic through a clinical decision support system to provide personalized patient care guidance.

[0173] Example 2 - Relational Network Model Generated from CMS Data: Subnetwork of Red Blood Cell (RBC) Disorders

[0174] A subnetwork centered on red blood cell (RBC) disorders was also selected from the generated relationship network model described in Example 1. Figure 20 ​The relational network model and selected RBC disorder subnetworks are schematically depicted. Another novel interaction was identified from the RBC disorder subnetwork.

[0175] DRG codes for RBC disorders encompass anemia (caused by nutrition, genetics, and comorbidities), transfusion reactions (ABO, Rh incompatibility), and cytopenias. All relevant diagnoses were considered in the analysis of RBC disorders in relation to other diagnoses.

[0176] The RBC disorder subnetwork indicates that the diagnosis of cellulitis leads to or follows the diagnosis of RBC disorders in a statistically significant proportion of patients. When cellulitis is untreated, bacteria from the infection enter the bloodstream and cause sepsis. Previous studies have shown that sepsis alters RBC morphology and rheology. It therefore seems reasonable that an increase in the diagnosis of cellulitis leads to an increase in the diagnosis of RBC disorders.

[0177] The RBC disorder subnetwork indicates that the diagnosis of RBC disorders leads to or follows the diagnosis of sepsis or severe sepsis in a statistically significant proportion of patients. Sepsis is likely in surgical patients, who can then undergo intravenous infusions. Multiple or prolonged intravenous infusions can cause pancytopenia, thereby linking the diagnosis of RBC disorders to sepsis.

[0178] The RBC disorder subnetwork also indicates that "other circulatory system diagnoses" lead to or follow the diagnosis of RBC disorders in a statistically significant proportion of patients. Circulatory system codes (i.e., other circulatory system diagnoses) include infections and abnormalities of heart tissue, as well as complications of cardiac surgery (placement of bypasses, shunts, implants, and valves). A possible link between circulatory disorders and RBC disorders involves hemolytic anemia from artificial valves. Specifically, iron-deficiency anemia (a type of RBC disorder) can cause rapid or irregular heartbeats, which can cause heart enlargement or heart failure.

[0179] The discovery of new interactions as identified in Examples 1 and 2 has important implications for both patients and providers. An important aspect of efficient health care delivery is predicting resource needs, determining strategies to maximize resource utilization, and availability of provider options. Currently, only some portions of medical insurance billing data, such as case mix indices, are used in planning. Leveraging advanced statistical analysis of billing data to augment such data can significantly improve the efficiency of health care management and result in significant savings for the health care system. The results demonstrate a new perspective on the use of advanced analysis of big data and, more importantly, demonstrate that the methods described herein can be effectively used to extract actionable information from big data. The inferred BNs, which are relational network models according to some embodiments, can also be analyzed to improve patient care by identifying adverse drug interactions, comorbidities, and disease causal relationships. The examples presented herein illustrate how embodiments can be used to understand and make sense of large, complex data sets to identify new interactions and produce actionable outputs for health care analytics and substantial health care economics.

[0180] In Examples 1 and 2, a meaningful relational network model representing interactions between diagnoses was constructed based on discharge volume according to embodiments. The relational network model was produced using a completely data-driven approach, completely free from bias by existing knowledge or assumptions. Results from the relational network model were validated using literature support. The examples illustrate that even a relatively small-scale analysis of low-resolution data leads to the identification of new interactions with clinical impact.

[0181] The example results illustrate how a purely data-driven approach in big data analysis according to the embodiments described herein can be used in health care to accelerate medical research and improve patient care by providing unique insights. The example results demonstrate a new perspective on the application of advanced analysis of big data and, more importantly, demonstrate that relational network models formed using Bayesian network algorithms can be effectively used to extract actionable information.

[0182] In this way, the health care analytics systems and methods disclosed herein can be used to analyze large data sets and gather new insights as relationships of data points in the data sets. Such analysis can be performed on data sets containing detailed DRG code information, patient-level claim information, patient clinical information such as diagnoses, medications, longitudinal data, clinical test results, and the like. Such analysis is beneficial to the pharmaceutical industry in terms of prescription recommendations, side effect and toxicity analysis, drug interactions, drug repositioning, patient groups for drug trials, and the like. Such analysis is beneficial to the hospital industry in terms of clinical decision support systems, outcome improvement for outcomes-based payments, improvement of standards of care, and the like.

[0183] Certain embodiments are described herein as including logic or a number of components, modules, or mechanisms. Modules can constitute either software modules (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware modules. A hardware module is a tangible unit capable of performing certain operations and can be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.

[0184] In various embodiments, a hardware module can be implemented mechanically or electronically. For example, a hardware module can include dedicated circuitry or logic that is permanently configured to perform certain operations, such as a special-purpose processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware module can also include programmable logic or circuitry that is temporarily configured by software to perform certain operations, such as a general- purpose processor configured by software. It will be appreciated that the decision to implement a hardware module mechanically or with temporary configuration software can be driven by cost and time considerations.

[0185] Accordingly, the term "hardware module" should be understood to encompass a tangible entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner and / or to perform certain operations described herein. Considering embodiments in which hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where a hardware module comprises a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor can be configured as

[0186] Hardware modules can provide information to, and receive information from, other hardware modules. Accordingly, the described hardware modules can be regarded as being communicatively coupled. Where multiple of such hardware modules exist contemporaneously, communications can be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communications between such hardware modules can be achieved, for example, through the storage and retrieval of information in memory structures to which such hardware modules have access. For example, one hardware module can perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware module can then access the memory device to retrieve and process the stored output. Hardware modules can also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information).

[0187] Various operations described herein can be performed, at least partially, by one or more processors executing relevant operations, either temporarily configured (e.g., by software) or permanently configured. Whether temporarily or permanently configured, such processors can constitute processor- implemented modules that operate to perform one or more operations or functions. In some example embodiments, modules referred to herein include processor- implemented modules.

[0188] Similarly, the methods described herein can be at least partially processor- implemented. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented modules. The performance of certain of the operations can be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processor or processors can be located in a single location (e.g., within a home environment, an office environment, or as a server farm), while in other embodiments the processors can be distributed across many

[0189] The one or more processors can also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS). For example, at least some of the operations can be performed by a group of computers (as examples of machines including processors), these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs).

[0190] Example embodiments can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Example embodiments can be implemented using a computer program product, e.g., a computer program tangibly embodied in information carrier, e.g., in a machine-readable medium for execution by, or to control the operation of, data processing apparatus, e.g., a programmable processor, a computer, or multiple computers.

[0191] A computer program can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed to be executed on one computer or on multiple computers (at one site or distributed across multiple sites and connected by a communication network).

[0192] In example embodiments, the operations can be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Method operations can also be performed by, and example embodiments can be implemented as, special purpose logic circuitry, e.g., an FPGA or an ASIC, to perform the functions.

[0193] A computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In embodiments deploying a programmable computing system, it will be understood that the hardware and software structures require consideration of the safety and security of the machine, as well as the accountability of the program and the people operating the program. Specifically, the selection of the hardware and software structures can involve a consideration of, among other things, the security of the data being stored and processed, the security of the machine, and the safety of the people operating the machine. The following notes describe hardware and software structures that can be employed in various example embodiments.

[0194] Figure 21900 is a block diagram of a machine in an example computer system within which instructions may be executed to cause the machine (e.g., client devices 110, 115, 120, 125; server 135; database server 140; database 130) to perform any one or more of the methodologies discussed herein. In alternative embodiments, the machine operates as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine may operate in the capacity of a server or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a PDA, a mobile phone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) specifying actions to be taken by the machine. Furthermore, while a single machine is shown, the term "machine" shall also be taken to include any collection of machines that, alone or in combination, execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0195] The example computer system 900 includes a processor 902 (e.g., a central processing unit (CPU), a multi-core processor, and / or a graphics processing unit (GPU)), a main memory 904, and a static memory 906, which are connected to each other via a bus 908. The computer system 900 may further include a video display unit 910 (e.g., a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)). The computer system 900 also includes a letter input device 912 (e.g., a physical or virtual keyboard), a user interface (UI) navigation device 914 (e.g., a mouse), a disk drive unit 916, a signal generating device 918 (e.g., a speaker), and a network interface device 920.

[0196] The disk drive unit 916 includes a machine-readable medium 922 on which is stored one or more sets of instructions and data structures (e.g., software) 924 embodying or used by any one or more of the methodologies or functions described herein. The instructions 924 may also reside, completely or at least partially, within the main memory 904, static memory 906, and / or processor 902, which also constitute machine-readable media, during execution thereof by the computer system 900, the main memory 904, and the processor 902.

[0197] Although the machine-readable medium 922 is shown in an example implementation to be a single medium, the term "machine-readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store the one or more instructions or data structures. The term "machine-readable medium" shall also be taken to include any tangible medium that is capable of storing, encoding or carrying instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present application, or that is capable of storing, encoding or carrying data structures used by or associated with such instructions. The term "machine-readable medium" shall accordingly be taken to include, but not be limited to, solid-state memories, and optical and magnetic media. Specific examples of machine-readable media include non-volatile memory, including by way of example semiconductor memory devices, e.g., erasable programmable read only memories (EPROM), electrically erasable programmable read only memories (EEPROM), and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0198] The instructions 924 can further be transmitted or received using a transmission medium via the network 926. The instructions 924 can be transmitted using the network interface device 920 and any one of a number of well-known transfer protocols (e.g., HTTP). Examples of communication networks include a LAN, a WAN, the Internet, mobile telephone networks, plain old telephone (POTS) networks, and wireless data networks (e.g., WiFi and WiMax networks). The term "transmission medium" shall be taken to include any intangible medium that is capable of storing, encoding or carrying the instructions for execution by the machine, and includes digital or analog communications signals or other intangible media to facilitate communication of such software.

[0199] While the present application has been described with reference to certain implementations, it will be apparent to those skilled in the art that various changes and modifications can be made and equivalents employed, without departing from the scope of the present application. Accordingly, the specification and drawings are to be regarded in an illustrative, rather than a restrictive sense.

[0200] It will be appreciated that, for clarity, purposes the above description has described some embodiments with reference to different functional units or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, processors or domains can be used without detracting from the application. For example, functionality illustrated to be performed by separate processors or controllers can be performed by the same processor or controller. Hence, references to specific functional units are only to be seen as references to suitable means for providing the described functionality rather than indicative of a strict logical or physical structure or organization.

[0201] While embodiments have been described with reference to particular embodiments, it will be apparent to those of ordinary skill in the art that various changes and modifications can be made to the embodiments without departing from the spirit and scope of the invention. Accordingly, the specification and drawings are to be regarded in an illustrative, rather than a restrictive, sense. The accompanying drawings, which form a part of the specification, are shown by way of illustration of certain embodiments in which the subject matter can be practiced. The implementations illustrated are set forth with specific details so as to provide a thorough understanding of the teachings. Other embodiments can be used and derived without departing from the scope of the disclosure. Accordingly, the detailed description is to be regarded as illustrative rather than restrictive. The scope of the embodiments is defined by the appended claims, and, accordingly, all equivalents and alternatives falling within the scope of these claims are included. The claims are not limited to the embodiments described herein but include any and all adaptations and modifications within the scope of the claims.

[0202] The embodiments of the inventive subject matter can be referred to herein, individually and / or collectively, by the term "application" merely for convenience and without intending to voluntarily limit the application to any single application or inventive concept if indeed more than one application or inventive concept is, in fact, disclosed. Thus, although specific embodiments have been illustrated and described herein, it should be appreciated that any arrangement can be substituted for the specific embodiments shown and described without departing from the spirit and scope of the application. The application disclosure is intended to embrace by inclusion all such substitutions and modifications.

[0203] In this document, the terms "a" or "an" are used, as is common in patent documents, to include one or more than one, independent of any other instances or usages of "at least one" or "one or more." In this document, the term "or" is used to mean, and is used in the same sense as "and / or" to include separate items listed with "or" without implying an exclusivity or necessity between alternatives. In the appended claims, the terms "including" and "in which" are used as the plain-English equivalents of the respective terms "comprising" and "wherein." Also, in the following claims, the terms "including" and "comprising" are open-ended; that is, a system, device, article, or process that includes elements in addition to those listed after such a term in a claim are still deemed to fall within the scope of that claim. Moreover, in the following claims, the terms "first," "second," and "third," etc. are used merely as labels, and are not intended to impose numerical requirements on their objects.

[0204] The abstract provided is to enable the reader to quickly ascertain the nature of the technical disclosure. It is not intended to be used to interpret or limit the scope or meaning of the claims. Furthermore, it is to be understood that the description of the state of the art in this section is to be interpreted as a representation of the close prior art. It is not intended to be used to limit or restrict the scope of the claims in any way. Moreover, it is to be understood that the description of the described embodiments is to be interpreted as an embodiment of the inventive concept. It is not intended to be used to limit or restrict the scope of the claims in any way.

Claims

1. A computer-implemented method for generating a causal network model based on patient data, the method comprising: receiving data corresponding to a plurality of patients comprising 50-1,000,000 patients, the data including diagnostic information and / or treatment information for each patient; parsing the data to generate normalized data for a plurality of variables including at least one variable related to the diagnosis or treatment of each patient, wherein for each patient, the normalized data is generated for more than one variable; Generating a causal network model involving the plurality of variables based on the generated standardized data using a programmable computing system, the programmable computing system comprising a memory containing network model building code and a plurality of processors configured to execute the network model building code, the generated causal network model including variables associated with a plurality of medical conditions, the generating the causal network model involving the plurality of variables comprising forming and evolving a full set of Bayesian networks based on the generated standardized data for 50 to 1,000,000 patients, at least some of the full set of Bayesian networks being evolved in parallel on the plurality of processors, wherein generating the causal network model involving the plurality of variables based on the generated standardized data comprises: forming a network segment library based on the variables through a Bayesian segment counting process; forming a set of test networks, each test network constructed from a different subset of network fragments in the library; and globally optimizing the entire set of trial networks by evolving each trial network through local transformations via simulated annealing to produce a consistent causal network model; determining a subnetwork from the generated causal network model, wherein one or more variables in the subnetwork are associated with the selected medical condition, or one or more variables in the subnetwork are associated with the selected drug; and Relationships in the subnetwork are explored to determine one or more predictors for the selected medical condition or the selected drug.

2. The method of claim 1, wherein the causal network model includes relationships indicative of one or more predictors for each of the plurality of medical conditions.

3. The method of claim 1, wherein the received data is not pre-selected as being relevant to one or more of the plurality of medical conditions.

4. The method of claim 1, wherein the plurality of patients comprises a first subset of patients each having data indicative of a diagnosis of a patient's medical condition and a second subset of patients each having data not indicative of a diagnosis of a patient's medical condition.

5. The method of claim 1, further comprising: receiving additional data corresponding to one or more other patients; and The causal network model is updated based on the additional data.

6. The method of claim 1, further comprising: receiving updated or additional data corresponding to one or more of the plurality of patients; and The generated causal network model is updated based on the updated or additional data.

7. The method of claim 1, wherein the causal network model is generated based solely on the generated standardized data.

8. The method of claim 1, wherein the one or more predictors for a selected medical condition are indicative of medical conditions that co-occur with the selected medical condition.

9. The method of claim 1, wherein the extent of the sub-network is determined based on the one or more variables associated with the selected medical condition and the strength of the relationship between the one or more variables and other variables in the generated causal network model.

10. The method of claim 1, wherein the subnetwork includes the one or more variables associated with the selected medical condition, a first set of other variables each having a first degree relationship with the one or more variables, and a second set of other variables each having a second degree relationship with the one or more variables.

11. The method of claim 1, wherein at least one of the one or more predictors is not previously known to be a predictor for the selected medical condition.

12. The method of claim 1, wherein at least one of the one or more predictors is newly identified as a predictor for the medical condition.

13. The method of claim 1, wherein the number of predictors is less than the number of variables.

14. The method of claim 1, further comprising: The one or more predictors are displayed in a user interface, the display including a graphical representation of the one or more variables, the one or more predictors, and a relationship among the one or more variables and the one or more predictors.

15. The method of claim 1, further comprising displaying a graphical representation of the subnetwork in a user interface.

16. The method of claim 1, further comprising ranking the one or more predictors based on the strength of the relationship between the one or more variables and the one or more predictors.

17. The method of claim 1, wherein the one or more predictors associated with the selected drug are indicative of drugs that are administered in combination with the selected drug.

18. The method of claim 1, wherein the one or more predictors are indicative of an adverse drug interaction between the selected drug and one or more other drugs.

19. The method of claim 1, wherein at least one of the one or more predictors is newly identified as a predictor for an adverse drug interaction between the selected drug and the one or more other drugs.

20. The method of claim 1, wherein the extent of the sub-network is determined based on the one or more variables associated with the selected drug and the strength of the relationship between the one or more variables and other variables in the generated causal network model.

21. The method of claim 1, wherein the subnetwork includes the one or more variables associated with the selected drug, a first set of other variables each having a first degree relationship with the one or more variables, and a second set of other variables each having a second degree relationship with the one or more variables.

22. The method of claim 1, wherein at least one of the one or more predictors is not previously known to be a predictor for the selected drug.

23. The method of any one of claims 1 to 22, wherein the causal network model is generated based on at least 50 variables.

24. The method of any one of claims 1 to 22, wherein the causal network model is generated based on at least 100 variables.

25. The method of any one of claims 1 to 22, wherein the causal network model is generated based on at least 1000 variables.

26. The method of any one of claims 1 to 22, wherein the causal network model is generated based on at least 100,000 variables.

27. The method of any one of claims 1 to 22, wherein the causal network model is generated based on 50 variables to 1,000,000 variables.

28. The method of any one of claims 1 to 22, wherein the received data comprises information from a patient's electronic health record.

29. The method of any one of claims 1 to 22, wherein the received data further comprises at least one of the following information for at least some of the plurality of patients: patient demographic data, medical history, patient family medical history, current medication information, non-current past medication information, allergy information, immune status information, laboratory test results, radiological images, vital signs information, patient weight, billing information, lifestyle information, habit information, insurance claim information, and pharmacy information.

30. The method of claim 29, wherein the patient demographic data comprises at least one of patient age, patient race, and patient ethnicity.

31. The method of any one of claims 1 to 22, wherein the received data comprises information from a medical record.

32. The method of claim 31, the information from the patient's medical record comprising at least one of a health care professional's notes, a health care professional's observations, medications and treatments administered, an order in which medications and treatments were administered, test results, and x-rays.

33. The method of any one of claims 1 to 22, wherein the received data comprises patient discharge information.

34. The method of claim 33, wherein the patient discharge information includes at least one of a diagnosis code, a treatment code, an insurance premium code, a diagnosis-related group code, and an International Classification of Diseases code.

35. The method of any one of claims 1 to 22, wherein the received data relates to a plurality of patients from a selected hospital.

36. The method of any one of claims 1 to 22, wherein the received data relates to a plurality of patients from a selected geographic region.

37. The method of claim 1, wherein generating the causal network model further comprises: The consistent causal network model is computer simulated based on the input data to provide a predicted confidence level for one or more causal relationships within the resulting generated causal network model.

38. The method of claim 1, further comprising: A selection of a medical condition is received from a plurality of medical conditions identifying the selected medical condition prior to determining the sub-network.

39. The method of claim 38, wherein the selection of the medical condition is received from a user via a user interface.

40. The method of claim 38, wherein at least one of the one or more predictors is not previously known to be a predictor for the selected medical condition.

41. The method of claim 38, wherein at least one of the one or more predictors is newly identified as a predictor for the selected medical condition.

42. The method of claim 38, wherein the number of predictors is less than the number of variables.

43. The method of claim 38, further comprising: The predictors are displayed in a user interface, the display including a graphical representation of the one or more selected variables, the one or more predictors, and a relationship between the one or more selected variables and the predictors.

44. The method of claim 38, further comprising displaying a graphical representation of the subnetwork in a user interface.

45. The method of claim 38, further comprising ranking the one or more predictors based on the strength of the relationship between the one or more selected variables and the one or more predictors.

46. ​​The method of claim 38, wherein the one or more predictors are associated with one or more medical drugs.

47. The method of claim 38, wherein the predictor is associated with one or more medical conditions.

48. The method of claim 1, further comprising: A query related to a medical condition is received from a plurality of medical conditions and the selected medical condition is identified based on the received query before determining the sub-network.

49. The method of claim 1, further comprising: receiving information related to a medical drug and identifying the selected drug based on the received information before determining the subnetwork; and The one or more predictors for the medical drug are stored.

50. A system for generating a causal network model based on patient data, the system comprising: a data receiving module configured to receive data related to more than 50-1,000,000 patients, wherein the data includes diagnosis information and / or treatment information for each patient; a parsing module configured to parse the data to generate normalized data for a plurality of variables including at least one variable related to the diagnosis or treatment of each patient, wherein, for each patient, the normalized data is generated for more than one variable; and A processor-implemented relationship network module configured to generate a causal relationship network model involving the plurality of variables based on generated standardized data, the module comprising a plurality of processors, the generated causal relationship network model including variables associated with a plurality of medical conditions, the generating the causal relationship network model involving the plurality of variables comprising forming and evolving a full set of Bayesian networks based on the generated standardized data for 50 to 1,000,000 patients, at least some of the full set of Bayesian networks being evolved in parallel on the plurality of processors, wherein generating the causal relationship network model involving the plurality of variables based on the generated standardized data comprises: forming a network segment library based on the variables through a Bayesian segment counting process; forming a set of test networks, each test network constructed from a different subset of network fragments in the library; and globally optimizing the entire set of trial networks by evolving each trial network through local transformations via simulated annealing to produce a consistent causal network model; a sub-network module configured to determine a sub-network from the generated causal network model, wherein one or more variables in the sub-network are associated with the selected medical condition, or one or more variables in the sub-network are associated with the selected drug; and A predictor module is configured to explore relationships in the sub-network to determine one or more predictors for the selected medical condition or the selected medication.

51. The system of claim 50, wherein the generated causal network model includes relationships indicative of one or more predictors for each of the plurality of medical conditions.

52. A non-transitory machine-readable storage medium storing at least one program, which, when executed by at least one processor, causes the at least one processor to perform the method of any one of claims 1 to 22.

Citation Information

Patent Citations

  • Healthcare information technology system for predicting development of cardiovascular conditions

    CN103493054A

  • Interrogatory cell-based assays and uses thereof

    CN103501859A