Scalable visual analytics pipeline for big data sets
By generating explicit time-series-aware graph data structures and graph-based visual analytics pipelines, the problem of difficulty in discovering patient care pathways and causal correlations in large-scale EHR data is solved, enabling efficient data analysis and visualization, supporting clinical experts to quickly identify patterns, and improving the efficiency of decision support systems.
Patent Information
- Application Number
- CN202211087196.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-13
- Filing Date
- 2022-09-07
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2042-09-07
AI Technical Summary
Existing decision support systems struggle to efficiently identify patient care pathways and potential causal or risk factor correlations when processing large-scale electronic health record (EHR) data, especially across large patient groups or different patient populations. Traditional visualization tools fail to effectively present the timeline of events and interactive queries, making it difficult for clinical experts to extract valuable patterns from large amounts of data.
By generating explicit time-aware graph data structures, and utilizing a time-aware graph query (CGQ) engine and a pattern discovery and visualization engine, a graph-based visual analytics pipeline is provided. This pipeline can efficiently process and visualize EHR data, generate Sankey graph visualizations, and support interactive pattern discovery by domain experts.
It enables efficient analysis of large-scale EHR data, quickly identifies patient care pathways and potential causal relationships, and improves the efficiency and accuracy of clinical decision support, especially significantly improving computational efficiency when processing historical data spanning tens of thousands of patients.
Smart Images

Figure CN115809239B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates generally to an improved data processing apparatus and method, and more specifically to a computer mechanism for providing a scalable visual analytic pipeline for large data sets. The scalable visual analytic pipeline can be part of or otherwise operate in conjunction with a decision support system or other cognitive computing system, for example. BACKGROUND
[0002] Decision support systems exist in many different industries where human experts need help retrieving and analyzing information. An example is diagnostic systems employed in the healthcare industry. Diagnostic systems can be classified into systems that use structured knowledge, systems that use unstructured knowledge, and systems that use clinical decision formulas, rules, trees, or algorithms. The earliest diagnostic systems used structured knowledge or classical hand-built knowledge bases. The Internist-I system, developed in the 1970s, used disease-finding relationships and disease-disease relationships. The MYCIN system, developed in the 1970s for diagnosing infectious diseases, used structured knowledge in the form of production rules that stated that if certain facts were true, then certain other facts could be derived with a given certainty factor. DXplain, developed starting in the 1980s, used structured knowledge similar to Internist-I but added a hierarchical dictionary of findings.
[0003] Iliad, developed starting in the 1990s, added more complex probabilistic reasoning in which each disease had an associated prior probability of the disease in the population for which Iliad was designed, and a list of findings along with the proportion of patients with the disease who had the finding (sensitivity), and the proportion of patients without the disease who had the finding (1 - specificity).
[0004] In 2000, diagnostic systems using unstructured knowledge began to appear. These systems used some structuring of knowledge, such as, for example, entities (such as findings and symptoms) tagged in documents to facilitate retrieval. ISABEL, for example, uses Autonomy information retrieval software and a database of medical textbooks to retrieve appropriate diagnoses under a given input finding. Autonomy Auminence uses Autonomy technology to retrieve diagnoses under a given finding and organizes the diagnoses by body system. The first CONSULT allowed people to search a large collection of medical books, journals, and guidelines by chief complaint and age group to get possible diagnoses. PEPID DDX is a diagnosis generator based on PEPID's independent clinical content.
[0005] Clinical decision rules have been developed for many medical conditions, and computer systems have been developed to help practitioners and patients apply these rules. The Acute Coronary Ischemia Time-Insensitive Predictive Instrument (ACI-TIPI) takes clinical and ECG features as input and produces a probability of acute coronary ischemia as output to aid in the differential classification of patients with chest pain or other symptoms suggestive of acute coronary ischemia. ACI-TIPI is incorporated into many commercial cardiac monitors / defibrillators. The CaseWalker system uses four questionnaires to diagnose major depressive disorder. The PKC Advisor provides guidance on 98 patient questions such as abdominal pain and vomiting. SUMMARY
[0006] This Summary is provided to introduce a selection of concepts that are further described in the detailed description. This Summary is not intended to identify key factors or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0007] In one illustrative embodiment, a method in a data processing system including at least one processor and at least one memory including instructions executed by the at least one processor to specifically configure the at least one processor to implement a visual analytics pipeline that performs the method is provided. The method includes generating a plurality of record's time-aware graph data structures from an input database of records based on features specified in an ontology data structure. The time-aware graph data structures include vertices representing events or records based on one or more of the features corresponding to the events, and edges representing temporal relationships between the events. The method further includes executing a time-aware graph query on the time-aware graph data structures to generate a set of filtered vertices and corresponding features corresponding to criteria of the time-aware graph query. The method further includes performing a pattern discovery operation on the set of filtered vertices and corresponding features to identify a subset of vertices and corresponding features corresponding to a set of relatively higher frequency event path patterns. Moreover, the method includes generating a visual analytics graphical representation for the subset of vertices and corresponding features in a visual analytics output.
[0008] In other illustrative embodiments, a computer program product including a computer usable or readable medium having computer readable program code is provided. The computer readable program code, when executed on a computer, causes the computer to perform various ones of and combinations of the operations outlined with regard to the method illustrative embodiments.
[0009] In yet another illustrative embodiment, a system / apparatus is provided. The system / apparatus can include one or more processors and memory coupled to the one or more processors. The memory can include instructions that when executed by the one or more processors cause the one or more processors to perform various ones of and combinations of the operations outlined above with regard to the method illustrative embodiment.
[0010] These and other features and advantages of the present application will be described in, or be apparent from, the following detailed description of the exemplary embodiments of the application. BRIEF DESCRIPTION OF DRAWINGS
[0011] The present application will be best understood by reference to the following detailed description of illustrative embodiments of the application in conjunction with the accompanying drawings, in which:
[0012] Figure 1 is an example block diagram illustrating the main operational elements of a scalable visual analytics pipeline according to one illustrative embodiment;
[0013] Figure 2A and 2B illustrates an example graphical representation of a graph pattern used to generate a time-aware data structure according to one illustrative embodiment;
[0014] Figure 3 is an example graph of a time-aware graph query (CGQ) algorithm implemented by CGQ logic of a CGQ engine according to one illustrative embodiment;
[0015] Figure 4 is an example graph of a pattern discovery algorithm implemented by pattern discovery logic of a pattern discovery and visualization engine according to one illustrative embodiment;
[0016] Figure 5 is an example graph of a visual analytics graphical representation of a coherent-aware graph query according to one illustrative embodiment;
[0017] Figures 6A-6C is an example of a visual analytics graphical representation for three group comparison according to one illustrative embodiment;
[0018] Figure 7 is a flow diagram outlining example operations of a scalable visual analytics pipeline according to one illustrative embodiment; and
[0019] Figure 8 is a block diagram of an example data processing system that can implement aspects of the illustrative embodiments. DETAILED DESCRIPTION
[0020] Decision support system operations depend on the digitization of health information and the implementation of computer tools to help understand patterns and correlations in the large and often complex information that exists in digitized health information. One key driver for digitizing health information from routine care delivery is to facilitate better understanding of the large variation in health care delivery practices, costs, and patient outcomes across health systems and patient populations. Traditionally, epidemiology has been a resource-intensive study of diseases, risk factors, and outcomes by comparing different patient population cohorts. Some of this knowledge can quickly become outdated as potential health risk exposures and medical practices change over time. An automated data processing mechanism that rapidly revalidates or discovers new relationships between emerging diseases, risk factors, management practices, and patient outcomes would add tremendous value to digitized health data and digital health data-based research and patient treatment computing systems, such as artificial intelligence and machine learning-based systems, decision support systems, and the like. Such an automated data processing mechanism would also require temporal awareness, as the evaluation of temporal relationships is critical to discovering potential causal or risk factor correlations between clinical data points when comparing patient cohorts.
[0021] By using large-scale digitized datasets, increasingly higher-capability machine learning tools can be used to automatically learn underlying data patterns to perform different disease prediction and / or clustering tasks, and for other domains of such prediction and / or clustering tasks, such as resource utilization prediction / clustering, and the like (for the purposes of the current description, the mechanisms of the illustrative embodiments will be assumed to be employed in the medical domain, but are not limited to these). However, it is often the case that such machine learning operations are performed without a sufficient understanding of the interaction between the data points and the process by which the data was collected. Thus, these machine learning mechanisms learn to make predictions using spurious correlations that do not have a directly actionable causal relationship when consulted by experts in the domain, such as clinical domain experts. On the other hand, clinicians also cannot easily examine these underlying data patterns from tabular data or even worse, large amounts of free-form text.
[0022] Visual analytics provides a way to include domain experts in the loop for the task of analyzing clinical feature and care path differences between patient cohorts that can explain different outcomes downstream. However, healthcare focused visual analytics tools only visualize patient data points at one cross-sectional time point for different forms of health data (e.g., age, gender, disease, labs, etc.) or present longitudinal time series line graphs for primarily tracking numerical data points (e.g., blood pressure, labs, etc.). Such visual analytics work well for dashboard understanding of individual patient health status and can be easily queried from tabular databases. However, these visualizations are limited for helping discover clinical use cases such as management differences and potential causal or risk factor level correlations between patient outcomes, especially across large patient groups or different patient cohorts. Such understanding requires visualizations of clinical or care paths that present the timing of events and have the flexibility and efficiency of creating clinical comparison cohorts from a range of different epidemiological style interactive queries.
[0023] A clinical or care path is a multidisciplinary management tool for evidence-based medical practice based on a group of patients with a predictable clinical course, where different interventions by medical professionals involved in patient care are defined, optimized, and ordered according to a specific desired timing, which relates outcomes to specific interventions. More simply, in relation to the timing-aware graph data structure mechanism of the illustrative embodiments, a care path or simply a "path" is a time series of patient visits and their outcomes (i.e., interactions between patients and practicing physicians and / or medical facilities for the purpose of medical care for the patients) and instances of features associated with those patient visits. It should be appreciated that a care path can be pre-defined by one or more medical experts, while actual path or care path instances can be extracted as observations from reality of patient data. Extracted actual paths or care path instances can deviate from the care path pre-defined by the medical experts.
[0024] One of the main technical challenges with existing mechanisms is the large number of medical features that exist for such machine learning, on the order of hundreds of thousands. This is combined with electronic health record (EHR) data samples that span hundreds of visits over years, resulting in a large and practically unmanageable number of possible patterns, e.g., over a billion, to explore through frequent sequence mining.
[0025] The illustrative embodiments address these issues by providing an improved computing tool solution based on an improved time-aware EHR database and an improved graph-based mechanism to perform data mining operations to mine EHR input data for possible patterns in patient healthcare trajectories. The illustrative embodiments provide a mechanism for generating explicit time-ordered graph representations of patient healthcare trajectories from EHR data. The illustrative embodiments further provide an improved computing tool mechanism that operates to express a range of common epidemiological style health pattern discovery problems, provide a mechanism to efficiently retrieve patient care paths, and be able to handle tens of thousands of patients with historical data spanning years. In addition, the illustrative embodiments provide a graph-based visual analytics pipeline that presents query results across time as a graph visualization diagram (e.g., a Sankey graph visualization (SVG) diagram) that allows for interactive pattern discovery for domain experts.
[0026] Before beginning a discussion of the various aspects of the illustrative embodiments and the improved computer operations performed by the illustrative embodiments, it should first be understood that throughout this description the term "mechanism" will be used to refer to elements of the present invention that perform various operations, functions, etc. The term "mechanism" as used herein can be an implementation of a function or aspect of the illustrative embodiments in the form of an apparatus, a process, or a computer program product. In the case of a process, the process is implemented by one or more devices, apparatus, computers, data processing systems, etc. In the case of a computer program product, the logic represented by the computer code or instructions embodied in or on the computer program product is executed by one or more hardware devices in order to implement the function or perform the operations associated with a particular "mechanism." Thus, the mechanisms described herein can be implemented as special purpose hardware, software executing on hardware, software instructions stored on media for execution by a hardware device to perform the functions, or processes or methods for performing the functions, or any combination of the above, where the software, when executed on the hardware, configures the hardware to perform the special purpose functions of the present invention that the hardware is not otherwise capable of performing, and the software instructions stored on the media cause the instructions to be readily executed by the hardware to thereby specially configure the hardware to perform the functions and particular computer operations described herein.
[0027] The specification and claims can utilize terminology "one," "at least one," and "one or more" with reference to certain features and elements of the illustrative embodiments. It will be understood that such terminology refers to the presence of at least one of the specific feature or element, but it does not preclude the presence of more than one of the feature or element. That is, such terminology does not preclude multiple features / elements, nor does it require a single feature;element, although both are possible under the teachings of the specification and claims. Rather, such terminology is simply used to avoid having to repeatedly recite "at least one" before each and every feature;element in the specification and claims.
[0028] Furthermore, it should be understood that if the term "engine" is used herein to describe an embodiment or feature of the present application, it is not intended to limit such embodiment or feature to any particular implementation, but rather it is intended to drive home the point that an engine can be implemented in any manner to achieve the actions, functions, etc. attributed to and / or performed by the engine. An engine can be, but is not limited to, software, hardware, and / or firmware implemented in, for example, a general purpose processor, a content addressable memory, a programmable logic array, and / or any other suitable device. Furthermore, unless otherwise specified, any name attributed to an engine is, for convenience, a name assigned to that engine by a user of the engine and, as such, is not intended to be a name or description of any particular implementation of the engine. Also, any functionality attributed to an engine can be performed, e.g., by a plurality of engines, integrated into a single engine, and / or distributed among a plurality of engines in any manner.
[0029] Furthermore, it should be understood that the following description uses a number of different examples of various elements to further assist in understanding the illustrative implementations of the illustrative embodiments and the mechanisms of the illustrative embodiments. These examples are intended to be non-limiting and are not exhaustive of the various possibilities for implementing the mechanisms of the illustrative embodiments. Given the nature of this specification, numerous alternatives will be apparent to those skilled in the art that will fall within the spirit and scope of the present application. Moreover, the above description is intended to be illustrative, and not restrictive. Many mechanical and structural changes can be made to the mechanisms disclosed herein, as well as many alternatives set forth nothwithstanding specific implementation described herein. Accordingly, the scope of the present application should be given by the appended claims along with their full scope of equivalents, wherein reference to an element in the singular is not intended to mean "one and only one" unless explicitly so stated, but rather one or more. The use of the term "adapted" herein is intended to mean "programmed" or "configured" unless otherwise indicated.
[0030] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present application.
[0031] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non- exhaustive list of more specific examples of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch cards or raised structures in grooves of a groove medium or other encoding methods, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0032] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0033] Computer readable program instructions for carrying out operations of the present application can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.
[0034] The computer readable program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0035] These computer readable program instructions can be provided to a processor of a computer, or other programmable data processing apparatus, to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include, without limitation, random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage technology. When the computer readable program instructions are executed by the computer or other programmable data processing apparatus, a computer-implemented process can be performed.
[0036] The computer readable program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include, without limitation, random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage technology. When the computer readable program instructions are executed by the computer or other programmable data processing apparatus, a computer-implemented process can be performed.
[0037] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, depending on the functions involved, two consecutively shown blocks may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0038] As described above, the present invention provides a scalable graph-based visual analytics (VA) pipeline operable on large electronic health record (EHR) databases with hundreds, thousands, or tens of thousands of features. Illustrative embodiments provide explicit time-series graph representations of patient healthcare trajectories from EHR databases. Illustrative embodiments can express health pattern discovery problems across a range of common epidemiological styles and can retrieve patient care pathways more efficiently than other query language-based systems (e.g., SQL systems). Illustrative embodiments can handle tens of thousands of patients with historical data spanning several years. Illustrative embodiments provide a graph-based visual analytics pipeline that presents query results across time as graph visualizations to allow interactive pattern discovery by domain experts.
[0039] Figure 1 This is an example block diagram illustrating the main operational elements of a scalable visual analytics pipeline according to an illustrative embodiment. To illustrate the operation of the main operational elements according to the illustrative embodiment, the following description will use an example medical condition known as transient ischemic attack (TIA) or “mini-stroke” as a case study. A TIA is caused by a temporary interruption of blood supply to a portion of the brain, in which neurological symptoms completely subside after the event. However, TIA is considered an early warning sign that a patient may be at risk of a fully eruptive life-changing event, such as stroke, death, or another cardiovascular-related event.
[0040] like Figure 1 As shown, the scalable visual analytics pipeline 100 includes a time-aware graph database generator 110, a time-aware graph database 120, a time-aware graph query (CGQ) engine 130, a pattern discovery and visualization engine 140, and a group query engine 150. It should be understood that, although in Figure 1These particular elements are shown, but the scalable visual analytics pipeline 100 can include additional logical elements (e.g., operating system logic, device drivers, communication interface logic, etc.) not explicitly shown and that provide underlying computer functionality and capability upon which the depicted elements are built.
[0041] Figure 1 The elements of FIG. 1 can be implemented as automated computer logic, including one or more of software logic and / or hardware logic specifically configured to perform the operations and functions attributed to these elements in this specification in accordance with one or more illustrative embodiments. With respect to software logic, this logic is executed on computer hardware to specifically configure the computer hardware to perform the operations and functions described herein. With respect to hardware logic implementation of one or more of the described operations or functions, if utilized, this hardware logic can be implemented as, for example, an application specific integrated circuit (ASIC) or other circuitry or physical hardware specifically configured to perform the described operations / functions. Regardless of whether hardware logic, software logic, or any combination of hardware logic and software logic is utilized, the resulting computing device is a specifically configured computing device to perform the particular operations / functions described herein.
[0042] The temporal-aware graph database generator 110 includes logic to generate one or more patient-centric graph data structures with temporal awareness for each patient in the EHR database source, which are stored in corresponding temporal-aware graph databases 120 for use by downstream logic of the pipeline 100 to generate analytical visualizations. The temporal-aware graph database generator 110 includes an ontology ingestion engine 112, an EHR data element extractor 114, and a temporal-aware graph data structure generator 116. The ontology ingestion engine 112 ingests one or more ontology data structures 111 representing domain knowledge for a particular domain of the input database source from one or more ontology source computing systems 103. The ontology source computing systems 103 can be any established source of ontology data structures that specify relationships for concepts for a particular domain of interest, for example, the Systematized Nomenclature of Medicine (SNOMED) Clinical Terms (CT)® computing system provides ontology data structures for medical applications corresponding to the SNOMED CT standard. Health Information Technology Standards Panel. SNOMED CT is owned and operated by SNOMED International, a not-for-profit organization.
[0043] A particular domain and corresponding ontology can have a level of specificity desired for a particular implementation. For example, in one implementation, the domain can be the medical domain, which includes a large number of different potential features that can be extracted from EHR data in one or more EHR database sources. In other illustrative embodiments, the domain can be more specific, such as a sub-group of one or more medical condition types patients, e.g., brain injury, cardiac medical conditions, cancer, etc. As noted above, while illustrative embodiments will be described with respect to patients and the medical domain, illustrative embodiments are not so limited and can be applied to other domains outside of the medical domain, such as resource utilization domains, etc. In general, illustrative embodiments can be applied to any data having a temporal dimension associated with data samples. As one further example, the mechanisms of illustrative embodiments can be implemented with respect to information technology (IT) support data samples having temporal characteristics associated therewith, e.g., servers or server components analogous to patients in this specification, and events recorded by the servers or server components analogous to patient visits discussed herein. Thus, for example, in an IT implementation of illustrative embodiments, the plurality of records can be a plurality of information technology records (e.g., log records, etc.) specifying events occurring in an information technology environment, and the events themselves can be information technology environment events occurring within one or more information technology software or hardware systems.
[0044] The ontology data structure 111 for a given domain specifies the types of features that can be extracted from input EHR data structures 104 of input database sources (e.g., EHR data source computing systems 102) and their relationships. In example illustrative embodiments, the domain is the medical domain, specifically the medical domain of a medical condition known as transient ischemic attack (TIA) or "mini-stroke," the ontology data structure 111 specifies features that can be extracted from patient EHR data structures 104, including entries for patient features and visit features, where visit features are features corresponding to visits between a patient and one or more medical personnel or medical facilities. The ontology data structure 111 can also include one or more ontology data structures specifying general features of a patient to be extracted from patient EHR data structures 104, such as patient name, address, age, gender, and various other types of demographics, including but not limited to height, weight, etc. The patient EHR data structures 104 can be obtained from any source computing system 102 of such a database storing patient EHR data, such as various healthcare databases maintained by healthcare provider organizations, medical facilities, doctor's offices, etc.
[0045] The ontology ingestion engine 112 ingests one or more ontology data structures 111 for a particular domain and configures the EHR data element extractor 114 to extract data elements from input EHR data structures 104 from one or more EHR database source computing systems 102 that correspond to the features specified in the one or more ontology data structures 111. The EHR data element extractor 114 can implement logic for extracting data elements corresponding to the features specified in the ontology data structures 111 from structured and / or unstructured EHR data structures. For example, the EHR data element extractor 114 can operate on table-based or otherwise structured EHR data structures 104 to identify fields having labels or annotations corresponding to the features specified in the ontology data structures 111 and can retrieve the corresponding data elements. In the case that the EHR data structures 104 include unstructured content, such as natural language content, the EHR data element extractor 114 can configure and implement computerized natural language processing logic and / or other artificial intelligence machine learning computing techniques to identify portions of the natural language content of the EHR data structures 104 that correspond to the features specified in the ontology data structures 111, for example, by machine learning or artificial intelligence based classification based on the syntax and semantics of the natural language content, annotating the natural language content with labels or annotations corresponding to the ontology features to which they correspond, and using the annotations to select data elements corresponding to the features specified in the ontology data structures 111. Of course, other types of EHR data structure 104 analysis can be used to extract data elements corresponding to features of interest as can be specified in the ontology data structures 111 without departing from the spirit and scope of the application, as will be apparent to one of ordinary skill in the art in view of this description.
[0046] The extracted data elements are related to the feature types and are used to generate a time-series awareness graph data structure corresponding to the patients. That is, the EHR data element extractor 114 extracts, for each patient in the input database, e.g., from the source computing systems 102, data element instances for the features specified in the ontology 111. These data element instances and their mapping to the features in the ontology 111 are then input to the time-series awareness graph data structure generator 116, which generates a time-series awareness data structure from these data element instances and feature mappings, which can be a graph-based data structure, which is then stored in the time-series awareness graph database 120.
[0047] The temporal-aware graph data structure generator 116 includes logic to identify instances of features and their relationships to a patient and to a patient encounter with a medical professional / medical facility (hereinafter simply “encounter”). Thus, for example, the ontology data structure 111 can indicate that a patient has a certain set of features defining the patient including, but not limited to, name, address, date of birth, gender, age, height, weight. Similarly, the ontology data structure 111 can indicate that a medical professional / medical facility has another set of features to a patient encounter, such as medical codes for diagnosis, lab results, vital signs, symptoms, dates, times, etc. These different features can be used as a basis for extracting instances of these features (i.e., data elements matching the feature types specified in the ontology data structure 111) from the EHR data structure 104. The temporal aspect of the extracted data elements corresponding to the features can be maintained so that the data elements can be related to a temporal ordering of the encounters and thereby provide temporal awareness.
[0048] The temporal-aware graph data structure generator 116 generates a temporal-aware data structure for a given patient in the input EHR data structure 104 based on these identified instances of data elements and their relationships according to a temporal-aware patient-centric graph schema that organizes the data elements and relationships in a temporal order with respect to the patient and to the patient’s encounters. Data elements representing features of the patient are connected to a node representing the patient, data elements representing features of the patient with respect to the patient’s encounters are connected to nodes representing various patient encounters. Each encounter is connected to the patient node by an edge. The temporal ordering of the encounters is modeled as directed edges between the encounter nodes (or vertices) from the oldest encounter node to the newest encounter node in the temporal ordering. This temporal-aware data structure for a given patient can be generated for each patient in the input EHR data structure 104 in the input EHR database (e.g., which can be maintained by the source computing system 102) so that multiple patients can have their EHR data structure 104 represented as a respective temporal-aware data structure, which can be stored in the temporal-aware graph database 120. These temporal-aware data structures for each patient can be stored individually in the database 120 or can be combined into an overall temporal-aware data structure having the extracted EHR data elements for multiple patients but maintaining the association of the data elements to their corresponding particular patient.
[0049] In some illustrative embodiments, with respect to the example temporal-aware data structure representing a graph structure, each feature (e.g., ontology concept) provides a node (feature node) in the graph, and each visit (per patient) containing that feature will be linked with the feature node. This type of temporal-aware graph data structure allows for searching a given feature to identify all relevant visits for all patients with one graph traversal. In other illustrative embodiments, multiple graphs are kept separate, with each graph containing a patient cohort. An advantage of these other illustrative embodiments is that searches can be initiated in parallel for all graphs at the unique cost of replicating the feature nodes in all graphs.
[0050] The temporal-aware data structure representation of data elements of a patient EHR data structure facilitates more efficient searching of features of patients and patient visits, enabling visual analytics and temporal-aware query processing. That is, by generating temporal-aware data structure representations such as the temporal-aware data structure representation of illustrative embodiments, querying feature instances of these temporal-aware data structure representations can be achieved through selective inference, i.e., searching for patterns of interest in the data without having to perform multiple join operations in temporal steps as required in common data models. That is, instead of navigating through a graph by blindly adding all discovered nodes after each traversal, illustrative embodiments provide a mechanism to perform selective inference for each path and at each traversal step, with selection conditions based on, for example, temporal distance and outcome node metadata.
[0051] That is, in cases where a temporal awareness data structure representation that can be represented as a temporal awareness graph has been generated, a temporal awareness graph query (CGQ) engine 130 can be employed to process temporal awareness graph queries for patient encounters having feature instances (data elements) corresponding to a specified range and combinations of such feature instances. For example, in some illustrative embodiments, a temporal awareness graph query (CGQ) can be defined as part of a visual analytics request 172 associated with a user (such as via a graphical user interface, etc.) in association with the scalable visual analytics pipeline 100 and / or the decision support computing system 160 with which the pipeline 100 is associated. That is, a user can access the decision support computing system 160 via a visual analytics client computing system 170 and one or more data networks (e.g., the Internet, one or more LANs, one or more WANs, etc.) and interact with the decision support computing system 160 and the scalable visual analytics pipeline 100 via one or more user interfaces to define queries of interest to the user. Alternatively, such queries can be automatically generated by the visual analytics client computing system 170 by an automated process. Whether user-generated or automatically process-generated, the generated visual analytics request 172 can specify certain input parameters defining the temporal awareness graph query (CGQ) to be processed, such as the start and / or end vertices of a patient care path (a sequence of patient encounters and feature instances associated with those encounters), the maximum duration of the path to be retrieved, the type of vertices (see examples in Table 1 below), the length of edges (such as in days, months, years, etc.) (relative times between encounter vertices) (this is also referred to herein as the maximum time t), whether the temporal awareness graph query is to be performed for past and / or future encounters, values of patient features of interest (where "values" are also referred to as feature instances or data elements corresponding to features) (such as, for example, age ranges and gender), or any other combination of patient features of interest, and the number of vertices and edges to be included in the visual analytics graphical representation (or "visualization"). It will be appreciated that the visual analytics request 172 can also specify parameters for cohort queries to compare visual analytics representations for multiple patient EHR cohorts, as described in more detail below with respect to the cohort query engine 150.
[0052] A temporal graph query (CGQ) operation to collect all feature instances associated with patient visits that occur between the presence of a feature instance "START_V" and the first occurrence of an ending feature instance "END_V," where "START_V" specifies a start vertex or node in the temporal graph data structure for one or more patients and "END_V" specifies an ending vertex or node for all patients within a maximum time period t, which can be specified in days, weeks, or any other time frame as desired by a particular implementation. In collecting all feature instances associated with such patients, information is maintained that specifies which patient is associated with which set of feature instances and records the time of the feature instance. For example, as shown in the example of Figure 5 START_V is "Cardiac examination and assessment (procedure)" and END_V is "Aortic valve stenosis, non-rheumatic (disorder)." It should be appreciated that the CGQ can generate a visualization for a single patient or a group of patients, and thus, the visual representation for visual analytics purposes can display information for more than one patient, such as for an epidemiological group, etc.
[0053] The CGQ is iteratively performed by the CGQ engine 130 for each vertex from the START_V vertex to the END_V vertex. The CGQ explores and looks at the vertices connected to the output of the selected vertex and determines: (1) whether the next vertex along the path is a visit vertex and is not the END_V, and (2) the difference in time of the visit corresponding to the visit vertex and the start vertex time is less than the maximum time t. If both conditions are met, the feature instance of that vertex is stored in the stored path. If the next vertex along the path is not a visit vertex, it is stored in the stored neighbor data structure. If the next vertex is the ending vertex END_V, or the time exceeds the maximum time t from the start vertex time, the stored path and the stored neighbors are returned. The "path" is a list of node identifiers. Each list of node identifiers in the "stored path" describes which visits were traversed for one path, all of whose visits belong to the same patient. The "stored neighbors" contain information about which features are contained in which traversed visits, where this feature information drives the visualizations generated by the mechanisms of the illustrative embodiments. Both the stored path and the stored neighbors are retrieved together during the same graph traversal for computational efficiency.
[0054] The CGQ allows both the features of the next visit and the current visit to be obtained during the same traversal of the time-aware graph data structure, e.g., the next visit is identified in the stored path and the features are stored in the stored neighbors, which improves the efficiency of performing the evaluation. For example, the CGQ is more efficient than database operations based on a general data format, e.g., SQL-based operations, which would require multiple join operations to identify the features of the next visit and the current visit. It has been determined that path retrieval using the CGQ mechanism of illustrative embodiments provides an improved improvement in efficiency, e.g., a reduction in computational time cost, as the number of features increases. That is, when it is necessary to extract temporal paths for large populations and long time windows (involving quantities on the order of 10,000 features), the CGQ shows up to a hundred-fold improvement in time cost over general data format-based mechanisms.
[0055] As previously described, illustrative embodiments provide a time-aware graph database mechanism that generates time-aware data structures based on processing of electronic health records (EHRs) of patients obtained from EHR data structure source computing systems and stores them in the database 120. The time-aware data structures can be represented as graph data structures and obtained by the time-aware graph database mechanism that performs data transformation operations that transform data from a general data model used in the EHR data structure source computing systems to a patient-centric graph data structure that explicitly models the temporal ordering of visit events as edges.
[0056] Figure 2A and 2B An example of a graphical representation of a graph schema used to generate time-aware data structures is shown in accordance with one illustrative embodiment. Figure 2A A first graphical representation highlighting a patient-centric graph schema with different vertex (node) types is shown. Figure 2B A second graphical representation highlighting the time-aware nature of the graph schema is shown, where visit nodes are connected by directed edges that specify an ordering of visit temporal ordering from oldest to newest.
[0057] As Figure 2A and 2B The graph schema uses a schema that models patient data and visits of a patient as nodes, where edges represent visit events that connect a patient to a visit. For example, in Figure 2Athe center vertex (node) types of the graph schema are the patient vertex (node) type 210 and the visit vertex (node) type 220. Patient-related data (e.g., patient birth year, gender, age, etc.) are connected as feature nodes directly to the patient vertex type 210. Visit-related data (e.g., medical codes associated with findings of a visit) are connected as feature nodes directly to the visit vertex type 220. A patient can have multiple visits occurring along a timeline, and this information is needed in order to retrieve the temporal progression between two vertices. In order to effectively utilize the temporal information at inference time, the timing of visits is modeled as Figure 2A and 2B edges in the graph data structure of Figure 2A , shown as edges 230 back to the visit nodes 220. However, it can be easier to see these edges in Figure 2B , Figure 2B shows a set of spiral directed visit edges spiraling away from the earliest visit to the latest visit. These visit edges allow navigation of searches and analysis along the timeline with simple selective inference, rather than the multiple joins per time step required in a general data model.
[0058] Referring again to Figure 1 As described above, based on this graph schema and the generation of this timing-aware graph data structure 120, the mechanisms of the illustrative embodiments further provide the CGQ engine 130 described above that provides the logic for performing CGQ on the timing-aware graph data structure 120 to collect features of patients and patient visits for the purpose of generating visual analytics output, such as in the form of Sankey diagrams and the like, as described below. Figure 3 An example of a CGQ algorithm that can be implemented using the CGQ engine 130 logic according to one illustrative embodiment is depicted in
[0059] As shown by the example algorithm of Figure 3 , the CGQ engine 130 logic operates to solve the following query abstraction: collect all features associated with visits that occurred between the presence of feature START_V and the first occurrence of feature END_V for all patients over a maximum time period of t days; keep information about which patient is associated with which set of features and keep information about the time at which the features were recorded. In this example, Figure 3 The premise of the operation of the algorithm of Figure 3As shown, the algorithm involves executing the function RETRIEVE_PATH (vertex_to_explore) on the given temporal-aware graph data structure, with the initial function input being (vertex_to_explore = START_V). This function retrieves the characteristics of the care path from the start vertex to the end vertex, where the care path is defined by the sequence or temporal ordering of patient visits. These care paths are stored in stored_paths and the characteristics of the different visits are maintained in association with the corresponding patient, while maintaining the temporal or chronological information. Meanwhile, if the next vertex is not a visit (i.e., vertex_type (next_vertx) is not equal to "visit"), then the next vertex is stored in stored_neighbors. Both the stored_paths and the stored_neighbors are returned.
[0060] Thus, the mechanisms of illustrative embodiments provide a temporal-aware data structure generation mechanism and a temporal-aware graph query mechanism. Moreover, illustrative embodiments provide a pattern discovery and visualization engine 140 based on the temporal-aware graph data structure generation and the temporal-aware query mechanism. The pattern discovery and visualization engine 140 operates on an initial set of features for which meaningful patterns are to be discovered in the collected instances of features, e.g., the data elements extracted from patient EHR data that match the features of the ontology data structure 111 and stored in the database 120 by the temporal-aware graph data structure generation engine 116 and within the time frame determined by the temporal-aware graph query engine 130. In some illustrative embodiments, this initial set of features can be all the features contained in the results generated by the CGQ engine 130, e.g., the features stored in stored_neighbors for visits in stored_paths, which represents the union of all the features (originally contained in stored_neighbors) for all the paths retrieved by the CGQ.
[0061] All possible patterns within the initial set of features are generated and structured as a graph, where the frequency of a pattern discovered in the collected feature data of the temporal-aware data structure (which can also be a graph data structure) in the database 120 is modeled as the weight of the edge between the vertices. In this graph, the nodes are the features, and the edges link the nodes that belong to the same pattern. The edge weights represent the frequency of occurrence of a given combination of features. For example, assume there are two feature patterns: F1-F2-F3 and F1-F2-F4. The graph data structure for this two-pattern example would contain the following edges: F1-F2, weight = 2; F2-F3, weight = 1; and F2-F4, weight = 1.
[0062] The pattern discovery and visualization engine 140 filters the pattern graph in order to select the most frequent patterns to provide visual analytics. Thus, the result generated by the CGQ engine 130 is a list of care pathways and corresponding features, and from these, patterns can be derived. However, the problem is that the number of possible patterns is huge. Therefore, a graph of patterns is generated for each query that fetches the output of the CGQ engine 130. The graph of patterns is then used to derive the most frequent patterns and at the same time eliminate patterns that are not included in the most frequent patterns, e.g. the top n patterns, where n is a tunable parameter.
[0063] For example, in one illustrative embodiment, all possible patterns are constructed as a graph, where the frequency of a pattern is modeled as the weight of the edge between vertices. The graph is filtered in order to display the most frequent patterns by traversing the vertices of each layer of the graph (where a layer is composed of all graph nodes resulting from one graph traversal step) to identify all neighbor vertices in the graph and rank the neighbors in descending order based on the edge weight of the neighbor vertices. Then, the top n vertices are selected for inclusion in the filtered set of vertices. It will be appreciated that there can be millions of nodes and tens of millions of edges in the graph of patterns. It is not feasible to visualize all possible patterns, and thus the present invention filters these patterns to those that exist between the top n vertices, or the top n patterns. In this way, a list of patterns is selected from the pattern graph.
[0064] In selecting the top n vertices, the value of n can be a monotonically decreasing value that decreases as the selection process continues to prevent oversampling in other layers. The remaining vertices in the filtered set of vertices are then graph visualized, such as by a Sankey diagram in which the size of the ribbons is proportional to the vertex degree. Modeling the patterns as one graph that contains the occurrence information as edge metadata, the solution provided by the illustrative embodiment is no longer affected by memory and time cost issues. In fact, it has been observed that a query involving up to 10^26 estimated number of patterns is represented in the graph and ranked in only a few seconds, with the pattern graph size being on the order of 10k vertices and 100 million edges.
[0065] Figure 4 is an example graph of a pattern discovery algorithm implemented by the pattern discovery logic of the pattern discovery and visualization engine according to one illustrative embodiment. As Figure 4 shown, the graph can be traversed to identify the most frequent patterns. Figure 1The algorithm implemented in the logic of the pattern discovery and visualization engine 140 in the begins with a filtered vertex set (filtered_vertices) that initially includes only the start vertex START_V. Then, for each layer in a range of layers, and for each vertex in the filtered vertex set, all neighbors of the vertex are identified and ranked according to the edge weights associated with those neighbors. Then, the top n neighbors are selected and stored in memory. Once all vertices in the filtered vertex set have been processed in this manner, for a layer, the filtered vertex set is updated with the stored neighbors (i.e., the top n neighbors for each vertex in the filtered vertex set for the current layer). This process is repeated for each layer. It will be appreciated that in each iteration for each layer, vertices that are not included in the filtered vertex set are removed.
[0066] The pattern discovery and visualization engine 140 then uses the resulting frequent patterns to generate a visual representation of the frequent patterns, such as a Sankey diagram or the like. This visual representation can be provided via one or more visual analytics outputs 174 that provide back to the visual analytics client computing system 170 or other source of the visual analytics request 172. The generated visual representation can take many different forms to represent the concepts, or features, and the strength of the relationships between the patterns of these concepts or features. Generally, the visual representation represents the strength of the relationships or associations as graphical representations having a size corresponding to the strength of the relationship or association. One example representation is a Sankey diagram, where graphical bands connect concepts / features, and the size of these bands corresponds to the strength of the relationships / associations between the concepts / features, e.g., the probability of one concept / feature connecting to another, the frequency of occurrence of the patterns or care paths involving these concepts / features, etc.
[0067] Figure 5 is an example diagram of a visual analytics graphical representation of a coherent perception graph query according to one illustrative embodiment. Figure 5 depicted in is a depiction of a Sankey diagram showing the strength of the relationships / associations between different concepts / features of a patient visit for a heart disease patient. For example, as Figure 5As shown, the concepts / features are represented as labeled rectangular blocks 510, 520 with labels such as cardiovascular exam and assessment, prosthetic heart valve history, preoperative assessment, acute pain, etc., with different categories of feature types, e.g., procedure, condition, disorder, finding, etc. These are the concepts / features specified in the ontology data structure, instances of which are identified in ingested patient EHR data structures, which are used to assess the probability or strength of associations / relationships between these concepts / features according to frequent patterns or care paths identified from the patient EHR data. The strips or channels 530, 540 connecting these concepts / features have widths representing the strength of the relationships / associations between these concepts / features, and thus the probability of the care path involving these concepts / features. Thus, for example, looking at the example in Figure 5 the care path involving cardiovascular exam and assessment, preoperative status, preoperative assessment, and aortic valve stenosis, non-rheumatic disorder, would be more likely to occur. The visualization allows the mechanisms of the illustrative embodiments and users viewing the visualization via the mechanisms of the illustrative embodiments to capture the dependencies of events at a glance, keeping in mind that in the example visualization, to the left of the nodes is the past and to the right is the future. Such a visualization is useful not only for projecting new patients on the resulting information from the historical cohort and thus comparing the likely next steps, but also for improving the efficiency of treatment at the hospital as a whole. For example, by observing the visualization, it can be determined how often and the most common reasons why a patient experiences a preoperative assessment multiple times (i.e., loops in the graph). Of course, many other assessments of care paths can be made by viewing the visualizations generated by the mechanisms of the illustrative embodiments.
[0068] Thus, in the visual representation of the visual analytics output 174, different features or concepts associated with features are represented, and the representations of these features / concepts within the visual representation have sizes proportional to the probabilities of occurrence of the features / concepts based on the extracted frequent patterns in which these features / concepts exist. For example, in a visual representation having a Sankey diagram type, such as Figure 5 As shown in and as described above, the widths of the channels or strips connecting the concepts / features are proportional to the number of occurrences of the concepts / features in the frequent patterns. Thus, the connections between the concepts / features represented in the visual representation have sizes proportional to the probabilities of occurrence of the concepts or features given the extracted care paths of interest and the frequent pattern finding mechanisms of the illustrative embodiments identified by the operation of the CGQ engine. In addition, the concepts / features are ordered in the visual representation with respect to temporal sequencing, as represented in the temporal aware data structure that allows the dependencies of events to be captured at a glance as described above.
[0069] Accordingly, illustrative embodiments provide mechanisms to convert patient longitudinal data from patient EHR data structures of one or more patient EHR data source computing systems into one or more time-aware data structures (which can be time-aware graph data structures) that link patient characteristics with patient visit characteristics while preserving the associations between patients and these characteristics and temporal characteristics that specify patient visit timing. Illustrative embodiments provide mechanisms for performing time-aware query processing to identify sets of characteristics of one or more patients that correspond to patient visits within a given timeframe at a starting point. Illustrative embodiments provide mechanisms to subsequently perform pattern discovery on these identified sets of characteristics to identify the most frequently occurring characteristics and characteristic patterns in the sets of characteristics generated by the time-aware query processing. Illustrative embodiments further provide mechanisms for visually representing the results of such pattern discovery as visual analytics output having graphical representations with sizes based on the frequency of occurrence of the patterns and characteristics / concepts.
[0070] Further, referring again to Figure 1 , illustrative embodiments further provide a cohort query engine 150. The cohort query engine 150 provides the logic for the operation of the other elements shown in Figure 1 . For purposes of illustration, it will be assumed that the process is performed with respect to two different cohorts, but it will be appreciated that the operation can be performed with respect to any number of patient EHR data cohorts without departing from the spirit and scope of the present invention.
[0071] The time-aware cohort query engine 150 performs the time-aware graph database generator 110, the time-aware graph query (CGQ) engine 130, and the pattern discovery and visualization engine 140 with respect to a defined cohort of interest. The cohort can be predefined with respect to different input patient EHR databases for different cohorts, or can be generated by classifying patient EHR data structures from one or more EHR data structure source computing systems according to predefined categories of patients, such as patients diagnosed with a transient ischemic attack (TIA) event within a timeframe of 2017-2019 as one cohort, and patients diagnosed with a TIA event within a timeframe of 2020-2021 as a second cohort. As patient EHR data is ingested, an initial analysis and classification of the patients can be performed, thereby associating patient EHR data structures with different cohorts of interest. For each cohort, the above-described processes are implemented to generate time-aware graph data structures 120 for the cohort and to perform time-aware graph queries (CGQs) on these data structures 120 and identify patterns and analytics visualizations for the cohort.
[0072] The cohort query engine 150 presents the resulting graph-based patterns via timestamped care path presentations with visualizations, such as the Sankey diagrams described previously, e.g., see the Sankey diagram described above inFigure 5 That is, similar to the CGQ described previously, a cohort query can be defined by a user or an automated process, such as by a subject matter expert via a user interface of the cohort query engine 150, with the following free parameters: the starting and / or ending vertex of the patient care pathway, the maximum duration of the pathway to be retrieved (cohort window), the type of vertex (see examples in Table 1 below), the length of the edges, whether the visit search is in the past and / or future, the patient age range, gender, etc., and the number of vertices and edges in the visualization. For a cohort query comparing two different cohorts, e.g., a first patient cohort from 2017 to 2019 and a second patient cohort from 2020 to 2021, separate graphical representations are generated for each cohort, with the results from the two cohorts presented via the same visualization, where the edge color of the same visualization encodes from which cohort each pathway comes, and the node color encodes whether there is any statistically significant difference in the care pathways between the two cohorts (proportion test based on normal (z) test for a significance level of 0.01 due to large dataset). For all graphical representations in the visualization, the size of the edges in the graphical representation (e.g., the width of the edges in a Sankey diagram) reflects the popularity of the care pathway normalized to the frequency of the target vertex (e.g., END V or START V), depending on the direction of the query (e.g., in the past or in the future). The frequency of the vertex is the width of the strip or channel, e.g., the number of times a feature appears in the pattern of the visualization.
[0073] With these parameters and visual analytic representations, the illustrative embodiments help domain experts explore a wide variety of epidemiological cohort exploratory analysis questions. For example, the illustrative embodiments provide a mechanism by which a domain expert can investigate a one-sided query for a target vertex, such as a query that answers what features (e.g., observations, medical history, and / or surgical history, etc.) do patients in a clinical cohort of interest have in a pre- or post-cohort window (number of days) relative to a target vertex (e.g., Transient Ischemic Attack (TIA) condition) selected by age range, gender, and target vertex. In addition, the illustrative embodiments provide a mechanism by which a domain expert can investigate a two-sided query for a care pathway between two target vertices, such as a query that answers what are the common features (clinical care pathways) between two different target vertices of a clinical cohort of interest within a given cohort window. Of course, other types of one-sided and two-sided queries can be answered using the mechanisms of the illustrative embodiments, these are just examples. For both one-sided and two-sided queries, the illustrative embodiments can also visualize and compare patient care pathways between different clinical cohorts, and ask which pathways are statistically more or less common between the two cohorts.
[0074] As an example, in an application of one illustrative embodiment, the first cohort includes all patients with a TIA event between September 2018 and December 2018 and includes their visits between June 2017 and December 2019. The second cohort includes all patients with a TIA event between February 2020 and March 2020 and includes their visits between January 2020 and February 2021. The resulting time-aware graph data structure resulting from querying using the above CGQ mechanism results in 21 different vertex types representing different medical features, with the graph for the first cohort containing 1,738,091 vertices and 41,538,032 edges, and the graph for the second cohort containing 359,504 vertices and 4,868,759 edges. Table 1 shows the distribution of the number of vertices for each vertex type as follows:
[0075] Vertex Type Group 2017-2019 Group 2020-2021 Diagnosis concept_name 8,656 5,185 Diagnosis snomed_id 8,656 5,185 Visit 1,119,601 184,572 Patient 12,029 2,833 Demographics birth_year 77 71 Demographics std_gender 2 2 Visit encounter_date 220,871 41,743 Visit std_encounter_type 54 43 Habits mapped_question_answer 10 10 History concept_name 5,282 3,060 History snomed_id 5,282 3,060 Observation loinc_id 5,182 2,464 Observation loinc-value-unit 232,424 67,888 Question List concept_name 4,703 440 Question List snomed_id 4,703 440 Procedure concept_name 6,817 3,636 Procedure cpt_concept_name 3,386 1,631 Procedure procedure_code 3,385 1,629 Procedure snomed_id 6,817 3,636 Surgical History concept_name 2,575 1,016 Surgical History snomed_id 2,575 1,016 Total Vertices 1,738,091 359,504 Total Edges 41,538,032 4,868,759
[0076] Table 1: Distribution of per-vertex types for graphs created from two TIA cohorts.
[0077] In the example shown in Table 1, it should be understood that“snomed” refers to the previously mentioned Systematized Nomenclature of Medicine (SNOMED) Clinical Terms (CT)® standard. For example, SNOMED CT can be used to define one or more input ontologies 111 for the mechanisms of the illustrative embodiments. The term“loinc” refers to Logical Observation Identifiers Names and Codes (LOINC)®, which is a registered trademark of the Regenstrief Institute, Inc.
[0078] To demonstrate the graph visualization analytics platform for TIA, different CGQ queries were executed on three example comparison cohorts, as described below. That is, the illustrative embodiments can help domain experts explore various epidemiological cohort exploratory analysis questions, which can be expressed as CGQ queries. Examples of these CGQ queries include:
[0079] 1. Unilateral queries for a single target vertex, e.g., what features (e.g., observations, medical history, and / or surgical history, etc.) do patients in a clinical cohort of interest have in a cohort pre-window or cohort post-window (number of days) relative to a target vertex (such as, for example, a TIA condition), selected (e.g., by age range, gender, and target vertex)?
[0080] 2. Bilateral queries for care paths between two target vertices, e.g., what are the common features (clinical care paths) between two different target vertices of a clinical cohort of interest within a given cohort window?
[0081] 3. For the above Type 1 and Type 2 queries, you can also request visualization and comparison of clinical groups and ask which pathway(s) are statistically more or less common between the two groups.
[0082] The clinical intent of these queries is to explore differences in patient characteristics (e.g., diagnoses) between groups and changes in management practices (e.g., procedures) over time. This example involves... Figures 6A-6C The first comparison shown visualizes common symptoms in patients diagnosed with a TIA within one year prior to the TIA and compares results from two predefined cohorts: age group 50–70 and age group 70–90. Furthermore, the example involves a second comparison visualizing the care pathways of patients diagnosed with a TIA within one year prior to the TIA and comparing results from two predefined cohorts: patients with recurrent TIAs and patients with a single TIA. Additionally, the example involves a third comparison visualizing the care pathways of patients diagnosed with a TIA within six months following the TIA and comparing results from two predefined cohorts: a cohort window from March 2019 to March 2020 (referred to as pre-pandemic) and a cohort window from March 2020 to February 2021 (referred to as pandemic).
[0083] Figures 6A-6C This is an example of a visual analysis graphical representation for comparing three groups according to an illustrative embodiment. Figures 6A-6C In each of the visual analytics graphical representations shown, rectangular elements corresponding to the concepts / features of a patient visit are shaded with respect to the number of patterns or care paths in which these concepts / features are identified. In the depicted example, blocks are shaded based on statistical significance, where statistical significance is correlated with the results of a normal test; for example, a p-value < 0.01 is considered significant. In each of these visual analytics graphical representations, the width of the bars or channels indicates the prevalence of the patterns involving the connected concepts / features, and thus the relative probability of these care paths occurring. Furthermore, the left-to-right ordering of concepts / features relative to each other indicates the temporal sequence of these concepts / features in the care paths.
[0084] Figure 6A The visual analysis was presented in the form of a Sankey diagram, in which the common primary symptoms of patients who had TIA in the year prior to their TIA event were shown in the first group or patient group aged 50–70 years (dark shaded bands / channels) and the second group or patient group aged 70–90 years (light shaded bands / channels). Figure 6BA visual analytics graphical representation is shown in the form of a Sankey diagram depicting common primary conditions for patients with a first occurrence of a TIA (first group; dark shaded bands / channels) and patients with recurrent TIA events (second group; light shaded bands / channels). Figure 6C A visual analytics graphical representation is shown in the form of a Sankey diagram depicting common primary processes occurring for patients after a first TIA event. Figure 6C A first patient group, Group 2019, prior to the pandemic is shown using dark shaded bands / channels, and a second patient group, Group 2020, during the pandemic is shown using light shaded bands / channels. It will be appreciated that, Figures 5-6C The concepts / features depicted in the middle are such as Figure 1 All of the concepts / features specified in the ontology data structure 111 of the ontology data structure in
[0085] Figure 7 is a flowchart outlining example operations of a scalable visual analytics pipeline according to one illustrative embodiment. Figure 7 The operations outlined in Figure 1 may be implemented by, for example, a scalable visual analytics pipeline 100 of Figure 7 The operations shown in Figure 1 The main operational elements of the pipeline 100 in (e.g., the temporal awareness graph database generator 110, the temporal awareness graph database 120, the temporal awareness graph query engine 130, the pattern discovery and visualization engine 140, and the cohort query engine 150) are intended to be implemented automatically by computer logic specifically configured to implement these different operational elements. This computer logic can be implemented as software executing on computer hardware and specifically configured to perform the operations and functions of these different operational elements, as special purpose computer hardware logic specifically configured to perform the operations and functions of these different operational elements, or as any combination of executing software and / or special purpose hardware.
[0086] As shown in Figure 7As shown, the operation begins with receiving an electronic health record (EHR) data structure from an EHR data source computing system and receiving one or more ontology data structures from one or more ontology data source computing systems (step 710). Feature instance or data element extraction is performed on the EHR data structure based on the concepts / features specified in the ontology data structures (step 720). A time-aware graph data structure for a patient in the EHR data structure is generated based on the extracted concept / feature instances (step 730). Frequent pattern discovery is performed on the time-aware graph data structure for one or more defined patient groups based on a recursive time-aware operation that collects both successive and current vertex features for each iteration (step 740). A visual analytic representation of the frequent patterns is then generated and output using the features collected for each pair of current and successive-in-time vertex, where the strength of association between features is indicated graphically with graphical elements of different sizes (e.g., differently width strips / tunnels connecting the features) (step 750). The operation then terminates.
[0087] From the above description, it can be appreciated that the illustrative embodiments provide a mechanism for creating visual analytics from large longitudinal datasets containing the following novel elements. The illustrative embodiments provide improved computational tools to generate explicit time-ordered graph representations of patient healthcare trajectories from structured and / or unstructured electronic health record data. The illustrative embodiments provide improved computational tools that can express a range of common epidemiological style health pattern discovery problems and retrieve patient care paths more efficiently than common data models using multiple join operations on patient electronic health record databases while handling thousands of patients with historical data spanning many years. Furthermore, the illustrative embodiments provide improved computational tools that implement a graph-based visual analytics pipeline that presents query results across time as visual representations that depict the frequency of patterns of extracted concepts / features from patient electronic health records, such as in one or more Sankey graph visualizations (SVG) diagrams that, for example, allow interactive pattern discovery by domain experts.
[0088] As noted above, the illustrative embodiments of the present application are directed specifically to an improved computing tool that automatically generates visual analytic representations of input data to aid in domain expert decision support operations and different interactive pattern discovery operations. While the improved computing tool can be utilized by and interact with humans, the actual operations of the scalable visual analytic pipeline mechanism are intended to be performed in an automated fashion using automated processes without human intervention, other than potentially receiving user input specifying a desired query to be resolved and providing output that can be used by a human. Further, while a human (e.g., a patient) can be the subject of a patient EHR data structure, the illustrative embodiments of the present application are not directed to actions performed by the patient or the user submitting the query, but rather to the logic and functionality specifically performed by the improved computing tool on the patient EHR data structure according to the parameters specified in the query. Further, while the present application can provide output to a visual analytic and / or decision support system that ultimately aids a human in evaluating a medical condition of a patient, the illustrative embodiments of the present application are not directed to actions performed by the human viewing the results of the processing performed by the visual analytic or decision support system, but rather to the specific operations performed by the particular improved computing tool of the present application that facilitate the generation of the visual analytic representations in an improved manner. Thus, the illustrative embodiments are not directed to organizing any human activity, but rather are directed to the automated logic and functionality of the improved computing tool.
[0089] Turning now to the drawings Figure 8 , a block diagram of a data processing system in accordance with one illustrative embodiment is depicted. The data processing system 800 can be used to implement a server computing device, such as one or more server computing devices that can be specifically configured to implement the scalable visual analytic pipeline 100 in Figure 1 . It should be appreciated that the pipeline 100 can be implemented on a single server computing device or can be distributed across multiple server computing devices, such as for example in a server farm or cloud computing implementation. Further, Figure 8 , the data processing system 800 in Figure 1 can also be used to implement a client computing device, such as the client computing device 170 in Figure 8 . It should be appreciated that the data processing system 800 in
[0090] In this illustrative example, the data processing system 800 includes a communication framework 802 that provides communications between a processor unit 804, a memory 806, a persistent storage 808, a communication unit 810, an input / output (I / O) unit 812, and a display 814. In this example, the communication framework 802 takes the form of a bus system.
[0091] The processor unit 804 is configured to execute the instructions of software that can be loaded into the memory 806. The processor unit 804 includes one or more processors. For example, the processor unit 804 can be selected from at least one of a multi-core processor, a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a digital signal processor (DSP), a network processor, or some other suitable type of processor. Moreover, in some illustrative embodiments, the processor unit 804 can be implemented using one or more heterogeneous processor systems in which a host processor is present on a single chip with a secondary processor. As another illustrative example, the processor unit 804 can be a symmetric multi-processor system that contains multiple processors of the same type on a single chip.
[0092] The memory 806 and the persistent storage 808 are examples of storage devices 816 that can include software logic executed by the processor unit 804 to specifically configure the processor unit 804 to implement one or more operational elements of one or more illustrative embodiments. A storage device is any hardware able to store information such as, for example and without limitation, both data and program code pieces in the form of functions, procedures, or other suitable information, either on a temporary basis, on a permanent basis, or on a temporary basis and a permanent basis. In these illustrative examples, the storage devices 816 can also be referred to as computer readable storage devices. In these examples, the memory 806 can be, for example, a random access memory, or any other suitable volatile or non-volatile storage device.
[0093] The persistent storage 808 can take a variety of forms, depending on the particular implementation. For example, the persistent storage 808 can include one or more components or devices. For example, the persistent storage 808 can be a hard disk drive, a solid-state drive (SSD), flash memory, a rewritable optical disk, a rewritable magnetic tape, or some combination of the above. The media used by the persistent storage 808 can also be removable. For example, a removable hard disk drive can be used for the persistent storage 808.
[0094] In these illustrative examples, the communication unit 810 provides communication with other data processing systems or devices. In these illustrative examples, the communication unit 810 is a network interface card that enables data communications with one or more data networks, such as the data networks described in the background section above. Figure 1
[0095] Input / output unit 812 allows for input and output of data with other devices that can be connected to data processing system 800. For example, input / output unit 812 can provide a connection for user input through a keyboard, mouse, or some other suitable input device. Further, input / output unit 812 can send output to a printer. Display 814 provides a mechanism to display information to a user.
[0096] Instructions for operating at least one of the system, application, or program can be in storage 816, which is in communication with processor unit 804 through communication framework 802. The processes of the different embodiments can be performed by the processor unit 804 using computer implemented instructions, which can be located in a memory, such as memory 806. These instructions are referred to as program code, computer-usable program code, or computer readable program code that can be read and executed by a processor in processor unit 804. The program code in the different embodiments can be embodied on different physical or computer readable storage media, such as memory 806 or persistent storage 808.
[0097] Program code 818 is located in functional form on selective removable computer readable media 820 and can be loaded onto or transferred to data processing system 800 for execution by processor unit 804. In these illustrative examples, program code 818 and computer readable media 820 form computer program product 822. In an illustrative example, computer readable media 820 is computer readable storage media 824.
[0098] In these illustrative examples, computer readable storage media 824 is a physical or tangible storage device used to store program code 818 rather than a medium that propagates or transmits program code 818. Alternatively, program code 818 can be transferred to data processing system 800 using computer readable signal media. Computer readable signal media can be, for example, a propagated data signal containing program code 818. For example, computer readable signal media can be an electromagnetic signal, an optical signal, or any other suitable type of signal. These signals can be transmitted over connections that can be wired, wireless, optical fiber cable, coaxial cable, electric cable, or any other suitable type of connection. Accordingly, computer readable storage media 824 is a non-transitory, tangible storage medium that is different from computer readable signal media, which is transitory.
[0099] The different components illustrated for data processing system 800 are not meant to provide architectural limitations to the manner in which different embodiments can be implemented. The components can be combined in a single package, or distributed in any manner as illustrated in the exemplary embodiments of the embodiments. For example, in some exemplary embodiments, the memory 806 or portions thereof can be combined with the processor unit 804. In different exemplary embodiments, different components of the data processing system can be combined, or distributed in a different manner. Figure 8 Other components illustrated for the data processing system 800 can be different in different illustrative embodiments. Different embodiments of the illustrative embodiments can be implemented using any hardware device or system that is capable of running the program code 818.
[0100] Accordingly, it should be appreciated that the illustrative embodiments can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment containing both hardware and software elements. In one exemplary embodiment, the mechanisms of the illustrative embodiments are implemented in software or program code, which includes but is not limited to firmware, resident software, microcode, etc.
[0101] A data processing system suitable for storing and / or executing program code will include at least one processor coupled directly or indirectly to memory elements through a communications bus or fabric. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memory, which provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during execution. The memory can be of various types, including, but not limited to, ROM, PROM, EPROM, EEPROM, DRAM, SRAM, flash memory, solid state memory, etc.
[0102] Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system either directly or through intervening I / O interfaces and / or controllers. The I / O devices can take many different forms, such as communication devices coupled by wire or wireless connections, including but not limited to smart phones, tablet computers, touch screen devices, voice recognition devices, etc. Any known or subsequently developed I / O devices are intended to be within the scope of the illustrative embodiments.
[0103] Network adapters can also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the currently available types of network adapters for wired communication. Network adapters based on wireless communication, including but not limited to 802.11 a / b / g / n wireless communication adapters, Bluetooth wireless adapters, and the like, can also be used. Any known or later developed network adapters are included within the spirit and scope of the present application.
[0104] The description of the application has been presented for purposes of illustration and description, and is not intended to be exhaustive or limited to the application in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The embodiment was chosen and described in order to best explain the principles of the application, the practical application, and to enable others skilled in the art to understand the application with various embodiments having suitable modifications in the various elements, and to best explain the principles of the application and its practical application to those familiar with the art. The terminology used herein was chosen to best explain the principles of the embodiment, the practical application, or a technical improvement over technology found in the marketplace, or to enable others skilled in the art to understand the embodiment disclosed herein.
Claims
1. A method in a data processing system comprising at least one processor and at least one memory, the at least one memory including instructions executed by the at least one processor to specifically configure the at least one processor to implement a visual analytics pipeline for performing the method, the method comprising: Generate multiple time-aware graph data structures based on record features specified in the ontology data structure from the input database of records, wherein the time-aware graph data structures include vertices representing one or more of the event-based or record-based features corresponding to the event, and edges representing the temporal relationships between events; Perform a time-aware graph query on the time-aware graph data structure to generate a standard, filtered set of vertices and corresponding features corresponding to the time-aware graph query; A pattern discovery operation is performed on the filtered set of vertices and their corresponding features to identify a subset of vertices and their corresponding features that correspond to a set of event path patterns with relatively high frequency. as well as Generate the vertex subset and corresponding feature visual analysis graphical representations in the visual analysis output. The input database for the records is an electronic health record database, and the plurality of records are multiple patient electronic health records, wherein the event is a patient visit between the patient and at least one of a healthcare professional or healthcare facility, and the plurality of record-based features include the patient's clinical characteristics.
2. The method of claim 1, wherein the record-based feature is a medical code associated with the patient's electronic healthcare record.
3. The method according to claim 1, wherein, The input database for the records is at least one information technology log database, and the plurality of records are a plurality of information technology records specifying events that occur in an information technology environment, wherein the events are information technology environment events that occur within one or more information technology software or hardware systems.
4. The method according to claim 1, wherein, The time-aware graph query is performed recursively, and during each iteration of the recursive execution, features for the current vertex of the iteration are collected and the next vertex along the care path specified in the time-aware graph data structure is identified, wherein the features collected for each iteration are combined into the filtered vertex set and corresponding features.
5. The method according to claim 1, wherein, Performing the pattern discovery operation on the filtered vertex set and corresponding features includes: Generate a pattern graph data structure, the pattern graph data structure including nodes corresponding to the filtered vertex set and edges connecting the vertices, wherein each edge has pattern frequency information specifying the frequency of the pattern of the vertices connected to the corresponding edge; For each path along the pattern graph data structure, determine the total pattern frequency along the path; and The generated pattern graph data structure is processed to select a predetermined number of vertex pattern paths that have the highest relative ranking total pattern frequency information in the pattern graph data structure.
6. The method according to claim 1, wherein, The data structure for generating the time-aware graph from the input database of the records includes: Associate the feature types specified in the ontology data structure with annotations or tags that are associated with the elements of the records in the input database of the records; Extract record-based features related to the specified feature types in the ontology data structure, and time information associated with the extracted record-based features; and The time-aware graph data structure is generated based on the extracted record-based features and corresponding time information.
7. The method according to claim 1, wherein, The visual analysis graphical representation includes graphical channels connecting graphical representations of concepts corresponding to vertices in the vertex subset, wherein each channel in the graphical channels has a width proportional to the probability of occurrence of the corresponding care path pattern of the connected concept corresponding to the channel.
8. The method according to claim 7, wherein, For each of the graphical channels, the graphical representation of the concept is organized in the visual analysis graphical representation according to the visit sequence of the corresponding nursing path pattern.
9. The method according to claim 1, wherein, The visual analytics pipeline is part of an artificial intelligence decision support computing system and provides the visual analytics graphical representation in a graphical user interface with user interface controls to evaluate a range of epidemiological health pattern discovery issues based on the visual analytics graphical representation.
10. A computer program product comprising: The instructions implemented therewith are executable by a processor to cause the processor to perform the method according to any one of claims 1-9.
11. A computer system, comprising: processor; and A memory that stores instructions that, when executed by the processor, perform the method according to any one of claims 1-9.
Citation Information
Patent Citations
Method and system for visual analysis of clinical episodes
CN104573306A
Method and device for managing graph data of temporary graph
CN105095371A