Dynamic clustering of a standard coding system

The method addresses interoperability issues by translating local coding systems to a standardized system for healthcare data, facilitating near-real-time patient and population health journey reconstruction and analysis, thus enhancing data management and reducing deployment time and cost.

WO2025147762A1PCT designated stage expired Publication Date: 2025-07-17VERTO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2025/050018
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2025-01-08
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Different organizations using varying coding systems in healthcare lead to interoperability issues, making integration, comparison, analysis, and presentation of medical data challenging across systems.

Method used

A method for dynamic clustering of a standard coding system involving clustering multiple codeable concepts, performing data store transformation, and journey reconstruction to translate local coding systems to a standardized system, enabling semantic understanding and clustering based on sameness and relatedness, and constructing journeys for patient and population health management.

Benefits of technology

Enables near-real-time reconstruction and query of patient and population health journeys, allowing for efficient deployment of population health initiatives and reducing the time and cost of data analysis from years to months.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025050018_17072025_PF_FP_ABST
    Figure CA2025050018_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A method implements dynamic clustering of a standard coding system. The method involves clustering multiple codeable concepts. The method further involves performing a data store transformation on the multiple codeable concepts. The method further involves performing journey reconstruction from the data store transformation. The method further involves performing journey discovery using the journey reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

DYNAMIC CLUSTERING OF A STANDARD CODING SYSTEMCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims benefit to U.S. Provisional Application No. 63 / 618,973, filed January 9, 2024, which is hereby incorporated by reference herein.BACKGROUND

[0002] Coding systems may be used in the healthcare industry to classify and document medical information, such as diagnoses, procedures, and treatments. The use of coding systems increases consistency and accuracy in the recording and sharing of health data across different systems and databases. Coding systems specify a common language that may be used to improve communication, streamline data entry and reporting, and enhance data analysis for patient care, research, and public health monitoring. The use of coding systems may streamline administrative processes, reduce errors, and improve the overall quality of healthcare services. An issue with using coding systems is that different organizations may use different coding systems, which can lead to interoperability issues. When different organizations use varying coding systems, the integration, comparison, analysis, presentation, and display of similar data in different systems becomes a challenge.SUMMARY

[0003] In general, in one or more aspects, the disclosure relates to a method for dynamic clustering of a standard coding system. The method involves clustering multiple codeable concepts. The method further involves performing a data store transformation on the multiple codeable concepts. The method further involves performing journey reconstruction from the data store transformation. The method further involves performing journey discovery using the journey reconstruction.

[0004] In general, in one or more aspects, the disclosure relates to a system that includes at least one processor and an application that executes on the at least one processor. Executing the application performs clustering multiple codeable concepts. Executing the application further performs performing a data store transformation on the multiple codeable concepts. Executing the application further performs performing journey reconstruction from the data store transformation. Executing the application further performs performing journey discovery using the journey reconstruction.

[0005] In general, in one or more aspects, the disclosure relates to a non- transitory computer readable medium including instructions executable by at least one processor. Executing the instructions performs clustering multiple codeable concepts. Executing the instructions further performs performing a data store transformation on the multiple codeable concepts. Executing the instructions further performs performing journey reconstruction from the data store transformation. Executing the instructions further performs performing journey discovery using the journey reconstruction.

[0006] Other aspects of one or more embodiments may be apparent from the following description and the appended claims.BRIEF DESCRIPTION OF DRAWINGS

[0007] FIG. 1 shows an example system in accordance with one or more embodiments.

[0008] FIG. 2 shows a method in accordance with the disclosure.

[0009] FIG. 3 shows a flow diagram of a concept cluster creation process in accordance with one or more embodiments.

[0010] FIG. 4 shows a flow diagram of a data transformation process in accordance with one or more embodiments.

[0011] FIG. 5 shows an example of a table of matched codeable concept pairs in accordance with one or more embodiments.

[0012] FIG. 6 shows an example of related codeable concepts in accordance with one or more embodiments.

[0013] FIG. 7 shows a flow diagram to reprocess unmatched codeable concepts in accordance with one or more embodiments.

[0014] FIG. 8 shows a flow diagram of a journey reconstruction process in accordance with one or more embodiments.

[0015] FIG. 9 shows an example of a reconstruction process in accordance with one or more embodiments.

[0016] FIG. 10 shows an example of a golden analytical model star schema diagram in accordance with one or more embodiments.

[0017] FIG. 11 shows an example flow diagram of a journey discovery process in accordance with one or more embodiments.

[0018] FIG. 12 shows an example of creating population sequences in accordance with one or more embodiments.

[0019] FIG. 13 is an example of a sequence graph in accordance with one or more embodiments.

[0020] FIG. 14 shows an example of a graph transformation in accordance with one or more embodiments.

[0021] FIG. 15 shows an example of graph consolidation in accordance with one or more embodiments.

[0022] FIG. 16 shows an example formula in accordance with one or more embodiments.

[0023] FIG. 17 shows an example of a journey diagram that may be used in accordance with one or more embodiments.

[0024] FIG. 18, FIG. 19, FIG. 20, and FIG. 21 show examples of user interfaces in accordance with one or more embodiments.

[0025] FIG. 22A and FIG. 22B show an example of a computing system that may execute one or more embodiments.

[0026] Similar elements in the various figures may be denoted by similar names and reference numerals. The features and elements described in one figure may extend to similarly named features and elements in different figures.DETAILED DESCRIPTION

[0027] One or more embodiments are directed to management of large volumes of patient data to track patient journeys. One or more embodiments translate local coding systems to a standardized coding system. To perform the translation, one or more embodiments build a semantic understanding of the local coding systems. Individual mappings are performed through analyzing custom terminology concurrently. Further, one or more embodiments perform clustering of codeable concepts based on local coding system distributions. Further, one or more embodiments translate a normalized data store to a standardized coding system into a format optimized for constructing journeys. Related parent codeable concepts may be enriched with children codeable concepts with a high degree of sameness. Further, one or more embodiments may filter events in the journeys based on contextual depth levels.

[0028] Journeys track the movement of an individual through a process over time. In healthcare, a patient care journey represents the care a patient receives over time or even between care settings. The patient care journeys may be aggregated into population journeys which may be used to better understand and manage populations of patients with common characteristics. Furthermore, the ability to understand population journeys provides planners and administrators with important information about how populations are interacting with the system in which their journey occurs.

[0029] In healthcare, an individual’s medical records are captured by different providers in different care settings at different times. In each of the caresettings, different systems (z.e., computing systems and applications) are used to capture information. Information may include terminology, such as diagnoses, medications, vaccinations, procedures, or other medical terms. The terms may be captured and stored in a number of different standards (according to a coding system) and non-standard (free text, custom codes) formats. Variations in data and the terminology differences within the standards and non-standard formats, make it difficult to interpret, aggregate, compare, and operate on the data in a consistent manner across a population.

[0030] Standardized coding systems provide industry specific codeable concepts, which are terms that have specific meaning with business values that are understood within that domain. Standardized coding systems also include hierarchies, definitions, attributes, and other relationships. An example of a standardized coding system in healthcare is the Unified Medical Language System® (UMLS®). UMLS includes codeable concepts from many vocabularies, including Current Procedural Terminology (CPT); International Classification of Diseases, version 10, Clinical Modification (ICD-10-CM); Logical Observation Identifiers Names and Codes (LOINC); Medical Subject Headings (MeSH); RxNorm; and Systematized Nomenclature of Medicine - Clinical Terms (SNOMED CT). Unified Medical Language System® and UMLS® are registered trademarks of the National Library of Medicine at 8600 Rockville Pike, Bethesda, Maryland, United States 20894.

[0031] In healthcare, the Health Level 7 (HL7) Fast Healthcare Interoperability Resource’s (FHIR) standard defines how healthcare information is exchanged between systems. HL7 FHIR defines a specific data type “CodeableConcept” to represent a value, which is defined by one or more references to terminologies, or ontologies, or which can be defined by text. The use of different terminologies, ontologies, or text to represent the same value is a common pattern in healthcare. The ability to map or translate these different terminologies, or ontologies, or text to a single value supports identification and normalization of medical codeable concepts across a heterogeneousecosystem of source systems. Thus “codeable concept” has a single meaning in the context of a standard coding system that may be defined by one or more codes formally defined in other healthcare standards such as LOINC or SNOMED CT or by text. The term “codeable concept” is recognized in industry to have this meaning. For example, Microsoft provides the following definition: “A Codeable Concept represents a value that is usually supplied by providing a reference to one or more terminologies, but may also be defined by the provision of text.” Further industry examples can be found, and as such, in this description, “codeable concept” is used to refer to values which may be defined by one or more terminologies, ontologies, or text, in accordance with industry standards and usage.

[0032] In UMLS, a codeable concept is defined as having a single meaning and contains all terminology from any source that represents that meaning in any way. Within a standard coding system, a codeable concept has a unique identifier and a natural language description. Use of UMLS supports identification and normalization of medical codeable concepts across a heterogeneous ecosystem of source systems to UMLS codeable concepts. Concept mapping is a means to normalize terminology against a standard coding system. Once local terminology is mapped to codeable concepts, clustering may be performed against the codeable concepts according to sameness and relatedness, resulting in concept clusters. In one or more embodiments, the concept clusters are then used to extract local terminology from data that is ingested from one or more source systems, and to translate the local terminology to codeable concepts using the concept clusters AThese codeable concepts are then organized as a time series timeline referred to as a golden analytical model (GAM).

[0033] The GAM stores medically relevant codeable concepts in a collectively exhaustive manner normalized against a standard coding system. The GAM offers value to health systems by storing the medically relevant information in a consistently queryable format, regardless of the system or setting that data iscollected from, to facilitate the reconstruction and analysis of population journeys. By adding visibility on timestamps for each of these codeable concepts in the GAM, combined with source (provider) attribution, the holy grail of queryable population health capabilities may become a near-real time capability using embodiments described herein.

[0034] Population health professionals like epidemiologists may use a four-step process for deploying population health initiatives: identification, stratification, interventions, and evaluation. This full population health cycle often takes health systems four or more years to complete the first intervention and corresponding evaluation results and often involves a heavily manual effort to collect information for each step of the journey. The GAM that one or more embodiments create in near-real time may allow population health teams to identify and stratify populations in hours, to deploy a mass-intervention within days, and to receive the evaluation data in near-real time from data feeds in the health system. One or more embodiments may allow health providers / payers to change a multi-year 10M+ population health initiative into a 3-6 month operational deployment at a fraction of the cost.

[0035] The ability to reconstruct and query journeys is used in supporting individual and population level surveillance, analysis, and management.

[0036] Turning to the Figures, FIG. 1 shows a diagram of a system in accordance with one or more embodiments.

[0037] In general, a data store or data repository is a type of storage unit or device (e.g., a file system, database, data structure, or any other storage mechanism) for storing data. The data store or data repository may include multiple different, potentially heterogeneous, storage units, and / or devices.

[0038] FIG. 1 shows a data store (102), which is a collection of entries which represent events of interest, where each entry is represented by information within the context and associated information captured in a source system during the execution of activities. An entry (106) in a data store (102) mayrepresent a corresponding event of interest. In healthcare, an event of interest could be an appointment for the delivery of care, a call into a virtual care line, an admission into a hospital, a diagnostic test conducted on a patient, or any number of clinical procedures that are undertaken by medical personnel. Source systems in healthcare include a Hospital Information System (HIS), an Electronic Medical Record (EMR), consumer wearables, telemetry generating medical devices, as well as any Information Communication Technology (ICT) system used to capture clinical, administrative, or patient related information. Information about those events and the care delivered is captured in the source system. Each event of interest may be related in the data store to an identifier that may be used to group journeys. In the healthcare industry, the identifier may be a patient identifier. The event of interest may also be directly associated with a timestamp. A record of the event of interest may also include other information, such as the location of the event, the modality the event was captured, timestamp information, who was responsible for logging the event, or any other metadata on the event as well.

[0039] In healthcare, other events of interest may refer to diagnoses, medications, vaccinations, procedures, or other medical terms where the event and associated information is captured in clinical or personal systems during the provision of care.

[0040] By way of an example, an entry in the data store represents Alice (the patient) emergency department admission (event of interest) on June 4th, 2023, 11 :53pm (timestamp), that was entered by Bob (the actor), at Acme Hospital (location).

[0041] Data in the data store (102) may be from one or more data sources (104). Examples of data sources include data lakes, clinical information systems, hospital information systems, etc. The data from the data sources may first be consolidated, such as by using the method described in PCT Publication PCT / US2021 / 027812 entitled “A Method for Consolidating Heterogeneous Healthcare Data.”

[0042] A data store may include local coding system specifications (108). A local coding system specification is a specification of a local coding system. A local coding system is a set of proprietary codes that may include a unique identifier, and / or a free text description, and that is not part of the standardized coding system. For example, a local coding system may define the list of medications in the given data store, and another local coding system that includes a list of medical conditions. A local coding system may be wholly or partially based on an existing standardized coding system, such as ICD10-CM, CPT, RxNORM, etc., but may also be based on or include terminology that is non-standard and specific to the local environment.

[0043] A standardized coding system specification (110) is a specification of a standardized coding system. Each industry may have the industry’s own standardized coding system that includes all relevant codeable concepts that are of business value. For example, in the healthcare industry, the Unified Medical Language System (UMLS) may be used. A standardized coding system specification may include terminologies and their associated definitions, defined ontological structures, or other information that the standardized coding system may expose. Any standardized coding system, appropriate to the industry, may be used.

[0044] A codeable concept is a terminology of the standardized coding system that is stored in the concept cluster. A codeable concept contains an identifier, a free text human-readable description, and additional metadata as needed, such as timestamps or longer text descriptions. For example, in UMLS the codeable concept with identifier C0011860 has a free text description of “Diabetes Mellitus, Non-Insulin-Dependent” while the UMLS codeable concept with identifier C0202054 has a free text description of “Glucohemoglobin measurement.”

[0045] The concept cluster is a representation of the local terminology from data sources standardized to codeable concepts in the standardized coding system, which have been mapped based on sameness and relatedness. Sameness meansthat the codeable concepts have the same or a highly similar meaning within the business context, while relatedness means that the codeable concepts in the pair have an already defined relationship within the business context. For example, the UMLS codeable concepts COO 11860 (Type II diabetes mellitus) and C0202054 (Eligibility Type 2 Diabetes) may be included in the concept cluster for diabetes based on relatedness. The concept cluster stores these mappings as codeable concept pairs in a translation table which also includes a confidence score (FIG. 5). One or more embodiments may include an application (112) that may be configured to operate on a computing system (described below). The application (112) may be configured to perform the following operations (114) (which may also be referred to as queries) on the concept cluster.

[0046] A first query, when given a codeable concept as input, returns the top X codeable concepts with the highest degree of sameness and / or relatedness to the input codeable concept. For example, given an input of “Type 2 DM”, one or more embodiments may output a list of related codeable concepts of a particular category such as Laboratory Procedure, in which the application may return a “Hemoglobin Ale” as a top result. In one embodiment, codeable concepts in the concept cluster may be retrieved using k-nearest neighbors or approximate nearest neighbor algorithms. In one embodiment, codeable concepts in the concept cluster may be retrieved using k-nearest neighbors or approximate nearest neighbor algorithms.

[0047] A second query, when given two codeable concepts as input, returns the sameness / relatedness factor between the two codeable concepts. For example, “Chronic Obstructive Airway Disease” and “Difficulty Breathing” may output a high degree of relatedness, where “Chronic Obstructive Airway Disease” and “Radionuclide Imaging” may return a low degree of relatedness.

[0048] A third query, when given three or more codeable concepts, will return a more general codeable concept that applies to all three concepts, also known as the closest ancestor codeable concept. For example, “Chronic ObstructiveAirway Disease” and “Type 2 Diabetes Mellitus” may return “Chronic disease” as the closest ancestor codeable concept. In one embodiment, the result of the query finds a candidate codeable concept that satisfies the following attributes with linear optimization below. In other words, the candidate codeable concept is identified by the minimization of the summation of the sameness between a first codeable concept with the candidate and the sameness of a second codeable concept with the candidate.

[0049] FIG. 2 shows a flowchart of a method (200) for dynamic clustering of a standard coding system. The method (200) of FIG. 2 may be implemented using the systems and components of FIG. 1, FIG. 22A, and FIG. 22B. One or more of the steps of the method (200) may be performed on, or received at, one or more computer processors. In an embodiment, a system may include at least one processor and an application that, when executing on the at least one processor, performs the method (200). In an embodiment, a non-transitory computer readable medium may include instructions that, when executed by one or more processors, perform the method (200). The outputs from various components (including models, functions, procedures, programs, processors, etc.) from performing the method (200) may be generated by applying a transformation to inputs using the components to create the outputs without using mental processes or human activities.

[0050] Block (202) includes clustering a plurality of codeable concepts. The codeable concepts, which may be the codeable concepts in a standardized coding system, may include several codeable concepts that may be similar to each other or that may be related to each other. A sameness factor may measure whether two codeable concepts are the same, and a relatedness factor may independently measure whether two codeable concepts are related. As an example, a disease may use different names in different codeable concepts that are identified as the same, whereas a disease and a treatment for the disease (intwo different codeable concepts) may be identified as related but are not the same.

[0051] Clustering the codeable concepts may include generating a plurality of concept clusters from the plurality of codeable concepts by calculating a sameness factor and a relatedness factor between pairs of codeable concepts from the plurality of codeable concepts. As further described with FIG. 3, the sameness factor and the relatedness factor may each be calculated as a combination of multiple heuristics.

[0052] Relationship heuristics may be related to existing ontological structures that are included within the standardized coding system. For example, a coding system may include a relationship two codeable concepts that identifies that one codeable concept is “also known as” a second codeable concept, which may be translated to a heuristic value of “1” for sameness between the two codeable concepts.

[0053] Shortest path heuristics may be related to the path length between codeable concepts in a graph generated from the standardized coding system. For example, each codeable concept may be a node in the graph, each relationship between a codeable concept may be an edge in the graph, and the shortest path between two codeable concepts may be the path with the fewest number of edges (and nodes) between the nodes representing the codeable concepts.

[0054] Distance heuristics may be related to the distances between vectors that represent the codeable concepts. For example, the vectors may be generated as embedding vectors using a language model that outputs an embedding vector for the natural language description of a codeable concept. The smaller the distance between the embedding vectors for two different codeable concepts, the more similar the two codeable concepts are.

[0055] Heuristics using lists of tuples may also be used. For example, a 4-tuple may include an identifier for a first codeable concept, an identifier for a secondcodeable concept, a sameness value, and a relatedness value. The sameness and relatedness values may be manually entered into the system.

[0056] The values from the different heuristics may be combined using a weighted combination. Different weights may be used to combine the heuristics to generate values for the sameness factor and for the relatedness factor. For example, a first set of weights may be used to generate the value for the sameness factor and a second set of weights may be used to generate the value for the relatedness factor.

[0057] Continuing with FIG. 2, block (205) includes performing a data store transformation on the plurality of codeable concepts. The datastore transformation may identify a mapping between the codeable concepts used by a local coding system to the codeable concepts used by the standardized coding system. Performing a data store transformation may include matching a plurality of codeable concepts using a local coding system to a plurality of concept clusters generated from a standardized coding system, which is further described in FIG. 4.

[0058] Matching codeable concepts using a local coding system to concept clusters generated from a standardized coding system may include extracting a plurality of identifiers using the local coding system from the plurality of codeable concepts using a local coding system. An identifier for each codeable concept for each entry (z.e., for each event) may be extracted.

[0059] Matching codeable concepts using a local coding system to concept clusters generated from a standardized coding system may further include extracting a plurality of descriptions using the local coding system from the plurality of codeable concepts using a local coding system. A description for each codeable concept for each entry (z.e., for each event) may be extracted from the local coding system.

[0060] Matching codeable concepts using a local coding system to concept clusters generated from a standardized coding system may further includematching the plurality of identifiers using the local coding system to a plurality of identifiers using a standardized coding system to generate a plurality of identifier match values. Matching between the identifiers of different coding systems (e.g., between the local and standardized coding systems) may be referred to as translation matching.

[0061] Matching codeable concepts using a local coding system to concept clusters generated from a standardized coding system may further include matching the plurality of descriptions using the local coding system to the plurality of concept clusters using the standardized coding system to generate a plurality of concept cluster match values. Heuristics may be used to match the codeable concepts using a local coding system to the concept clusters generated from a standardized coding system. The heuristics used here to match codeable concepts may be different from the previously described heuristics used to identify the codeable concepts. The previously described heuristics are used to identify codeable concepts and may include heuristics based on relationship, shortest path, vector, and manual scores for sameness and relatedness. The heuristics here may be used to match different codeable concepts (which is further described below) and may include the use of a vector database containing concept mapping vector embeddings; database full-text search, and n-gram matching algorithms.

[0062] Matching codeable concepts using a local coding system to concept clusters generated from a standardized coding system may further include generating a confidence value for each codeable concept from the plurality of codeable concepts using a local coding system based on the plurality of identifier match values and based on the plurality of concept cluster match values. The confidence value for each codeable concept may correspond to a label. For example, the confidence values may be a weighted combination of heuristics corresponding to the sameness or relatedness between different codeable concepts. The confidence values may be normalized to a range between zero and one and different thresholds may be used to identify differentlabels that correspond to the different confidence values. One threshold (e.g., 0.95) may be used to identify a match as a “full match” and another threshold value (e.g., 0.8) may be used to identify a match as a high confidence match. Codeable concepts that are not labeled as a “full match” or a “high confidence” match may be labeled as “unmatched”.

[0063] Performing a data store transformation may further include rematching the plurality of entries labeled as a first type of matches and a second type of matches to supplement the plurality of concept clusters, as further described in FIG. 7. The first type of matches may be full matches and the second type of matches may be high confidence matches. The rematching process may revise and supplement the previously generated concept clusters to include additional or different codeable concepts. The rematching process may use similar heuristics as those used to generate the concept clusters from the codeable concepts that use the standardized coding system.

[0064] Performing a data store transformation may further include reclassifying a plurality of unmatched codeable concepts subsequent to supplementing the plurality of concept clusters. As further described with FIG. 7, the unmatched codeable concepts may be labeled into different buckets based on different thresholds. The unmatched codeable concepts are codeable concepts that were labeled as unmatched with the previous matching process. The labels for the unmatched codeable concepts may include a “medium confidence” label for confidence values between or including 0.5 and 0.8, a low confidence label for confidence values between or including 0.3 and 0.5, and a “quarantined” label for confidence values that are below 0.3. The labels identifying the types of matches between the codeable concepts (e.g., full match, high confidence match, medium confidence match, low confidence match, etc.) may be stored in a graph with nodes representing the codeable concepts and edges representing the type of match between the codeable concepts.

[0065] Continuing with FIG. 2, block (208) includes performing journey reconstruction from the data store transformation, which is additionallydescribed with FIG. 8. After generating concept clusters and identifying the translations and mappings from the codeable concepts using the local coding system to the codeable concepts using the standardized coding system (as well as to the concept clusters), the data of the entries representing events are extracted and processed.

[0066] Performing the journey reconstruction may include extracting a plurality of entries comprising a plurality of local codeable concepts using a local coding system, along with the timestamps related to the local codeable concepts. The extraction of an entry may retrieve data for the entry from a database and store the data from an entry into a structured text file (e.g., a JavaScript object notation (JSON) file). Each entry may correspond to a single event and include multiple codeable concepts. Each event may be stored in one or more entries.

[0067] Performing the journey reconstruction may further include translating the plurality of local codeable concepts to a plurality of standardized codeable concepts. Translating a local codeable concept (i.e., a codeable concept using the local coding system) to a standardized codeable concept (i.e., a codeable concept using the standardized coding system) may include retrieving the identifier for the local codeable concept, identifying the identifier for the standardized codeable concept to which the identifier for the local codeable concept is mapped, and storing the identifier for the standardized codeable concept.

[0068] Performing the journey reconstruction may further include transforming a plurality of standardized codeable concepts using a plurality of dimensions. Each codeable concept is retrieved from a database and may be stored in the GAM as a record of a codeable concept using the standardized coding system (i.e., a record of a standardized codeable concept). The record for the standardized codeable concept that is stored in the GAM may include values for additional dimensions, as shown with FIG. 10. The dimensions of the record for the standardized codeable concept may include a subject dimension, a concept dimension, a time dimension, an actor dimension, and a locationdimension. The subject dimension may store values that identify a subject related to an event (e.g., a person (the subject) that received a treatment (the event)). The concept dimension may store identifiers of the codeable concepts using the standardized coding system. The time dimension may store values that identify when the event occurred and may include timestamps, which may identify the year, month, day, hour, minute, second, etc., that events occurred. The actor dimension may store values that identify an entity that performed an action (e.g., performed a treatment such as administering a drug), which may be a person (e.g., a nurse, a doctor, etc.). The location dimension may store values that identify where action is performed, which may be an organization (e.g., a hospital, a clinic, etc.).

[0069] Continuing with FIG. 2, block (210) includes performing journey discovery using the journey reconstruction, which is also discussed with FIG. 11. The process of journey discovery may create sequence graphs from the information pulled in from a database that used the local coding system. The sequence graphs may read the sequences of codeable concepts that correspond to and may be experienced by the patients (also referred to as subjects) identified in the data.

[0070] Performing the journey discovery may include generating population sequences.

[0071] The generation of population sequences may identify sequences of codeable concepts and statistics related to the sequences of codeable concepts, which is also described with FIG. 12. A population sequence may be a collection of sequences of codeable concepts gathered from a group (i.e., a population) of subjects.

[0072] Generating population sequences may include grouping a plurality of codeable concepts by subject to identify a plurality of subject sequences for a plurality of subjects (see the middle table of FIG. 12 for an example). Eachsubject may experience multiple sequences of codeable concepts and different subjects may experience the same or similar sequences of codeable concepts.

[0073] Generating population sequences may further include grouping the plurality of subject sequences by subject sequence to generate a plurality of sequence counts for the plurality of subject sequences (see the right table of FIG. 12 for an example). A sequence count may identify the number of times a sequence of codeable consequences occurs in the data and may correspond to the number of subjects that experience the sequence.

[0074] Performing the journey discovery may further include filtering the population sequences using one or more of a plurality of dimensions. Filter parameters may be specified to control the filtering of the sequences within a population sequence. Multiple types of filters may be used. Filters may screen the underlying data based on the dimensions of the records of the codeable concepts, which include dimensions for subject, concept, actor, and location. For example, a filter may screen for locations within a certain city, state, or groups thereof and another filter may filter for groups of subjects. As another example, a filter may screen for subjects that are older than a threshold age.

[0075] Performing the journey discovery may further include generating one or more sequence graphs from the plurality of population sequences. A sequence graph may be a graph of the sequences of codeable concepts that are experienced by subjects within a population. The nodes of a sequence graph may correspond to codeable concepts and the edges of a sequence graph may represent the number of subjects (z.e., patients) that transition between the codeable concepts represented by the nodes joined by the edge. The direction of an edge may represent that the codeable concept of the first node of the edge is experienced by a patient (z.e., in a first event) prior to the codeable concept of the second node of the edge being experienced by the subject (z.e., in a second or subsequent event).

[0076] Performing the journey discovery may further include consolidating a plurality of nodes to revise the one or more sequence graphs, examples of which are shown with FIG. 13, FIG. 14, and FIG. 15. Different nodes within a sequence graph may be duplicative representing the same or similar codeable concept. Nodes that are similar may be breadth wise consolidated (see FIG. 14), be depth wise consolidated (see FIG. 15), or a combination thereof.

[0077] The method (200) may further include presenting a codeable concept of the plurality of codeable concepts using a sequence graph in a view.

[0078] The codeable concept may be presented by one or more of transmitting and displaying information related to the codeable concept, including a name of the codeable concept, an identifier of the codeable concept, a description of the codeable concept, etc., examples of which are described with FIG. 17 through FIG. 21.

[0079] The method (200) may further include presenting a codeable concept in one of a timeline view and a patient summary view.

[0080] The timeline view may present a sequence of codeable concepts based on the temporal relationships between the events from which the codeable concepts were identified, an example of which is further described with FIG. 18. The patient summary view may present a summary of the multiple codeable concepts that a subject may have experienced, an example of which is further described with FIG. 19.

[0081] The method (200) may further include presenting a codeable concept in one of a population journey view and a dashboard view. The population journey view may portray a sequence of codeable concepts that a population of subjects experienced, an example of which is further described with FIG. 20. The dashboard view may portray multiple charts and graphs that may each depict information from a sequence of codeable concepts, an example of which is further described with FIG. 21.

[0082] FIG. 3 shows a diagram of the workflow (300) to calculate how the sameness factor and the relatedness factor given two codeable concepts from a set of heuristics. Turning to FIG. 3, a standardized coding system is selected in order to create a concept cluster. First, terminologies from a standardized coding system may be consumed through, but not limited to, tabular formats and graph formats. After the ingestion of a terminology, the concept cluster represents the terminology as a codeable concept.

[0083] The relatedness and sameness factor between two codeable concepts is computed as a linear combination of the results from all of the defined sameness and relatedness heuristics, namely ontology heuristic, shortest path heuristic, vector heuristic and manual heuristic. In one embodiment, the set of heuristics used to compute the sameness score and relatedness score includes ontology, shortest path, vector, and manual.

[0084] Ontology heuristic: ontological structures are defined by and exist within the standardized coding system. For example, the UMLS coding system defines relationships between its codeable concepts, such as “may_treat” or “also_known_as”. The “may_treat” relationship may increase the relatedness score, if this relationship exists between 2 codeable concepts, i.e., codeable concept 1 “may treat’ ’ codeable concept 2. The latter relationship may increase the sameness score if the relationship exists between 2 codeable concepts, for example, codeable concept 1 “also_known_as” codeable concept 2.

[0085] Shortest path heuristic: given the same existing ontological structures above, it is possible to create a graph with the codeable concepts as nodes, and relationships as edges. Then, a shortest path algorithm may be used to determine sameness and / or relatedness by traversing through the sameness / relatedness relationship types. Relationships may also be manually defined by a user as a sameness relationship type or a relatedness relationship type. If a simple path exists between 2 nodes, solely through connecting with the sameness relationship type edges, then it may be considered to have a high sameness factor. If a simple path exists between 2 nodes through the samenessrelationship type and relatedness relationship type, then it may be considered to have a high relatedness factor.

[0086] Vector heuristic: follows a 3-step process to generate a sameness or relatedness score. Firstly, a pretrained bidirectional encoder representations from transformers (BERT) or sentence-BERT (SBERT) model is selected, or is trained based on text related to the industry of interest, in this case healthcare.

[0087] Secondly, the vector embeddings for all codeable concepts are computed. Vector embeddings turn words and phrases into a vector of numbers, where vectors that are close to each other indicate that the underlying word or phrase represented by the vector embedding are similar, and vectors that are far from each other indicate that the underlying word or phrase represented by the vector embedding are not similar.

[0088] Thirdly, a textual similarity score between all pairs of vector embeddings is calculated. In one embodiment, the textual similarity score is calculated as the dot product of two vector embeddings. The resulting output is a textual similarity score that is a number from zero to one, where zero indicates that two concepts are an exact match, and a result of zero means that there is no relationship between the two concepts.

[0089] Manual heuristic: a person may manually specify a measurement of relatedness and sameness scores of two codeable concepts that may be manually measured by creating 4-tuples where (i) the first element is the first codeable concept; (2) the second element is the second codeable concept; (3) the third element is the relatedness score (from 0 to 1); and (4) the fourth element is the sameness score (from 0 to 1). The application may be configured to receive the relatedness and the sameness of the two codeable concepts from a user or other source.

[0090] In one embodiment, “Light Therapy” may not be initially related to “Diabetes” through the existing strategies. Manual creation of a 4-tuple may increase the relatedness or sameness between the two codeable concepts. Themanual heuristic is a means to fine tune sameness and relatedness scores based on the specific use case or implementation.

[0091] To facilitate queries, codeable concepts are turned into embeddings through precomputing pairwise sameness and relatedness scores and storing them in a 2D square matrix. The resulting scores contained in the 2D matrix may be used by the system to train or fine tune a sentence transformer model to more efficiently perform query retrieval. In one embodiment, the pair score method and the cosine similarity loss function may be used for fine tuning. After training, each codeable concept is now represented as a concept cluster vector embedding and may be inserted into a vector database for querying.

[0092] At the end of FIG. 3, terminologies are mapped to codeable concepts in a standard coding system. Codeable concepts are mapped into concept clusters based on sameness and relatedness scores.

[0093] Another process is a data store transformation as shown in FIG. 4. A data store (401) may be using local coding systems to represent its data rather than using the standardized coding system in the concept cluster. These local coding systems may be identified in the data store prior to concept mapping and data ingestion. The identification process scans through the data store and extracts relevant fields that may correspond to a local coding system of relevance for analysis. The scanning and extraction may be done through a combination of statistical analysis on the data store to find fields that are represented by categorical values, comparing them against other known local coding systems, and a manual review process by the implementer to determine if the attribute is relevant. The identification process may also be done automatically, through the use of the process set forth in PCT Patent Application No. PCT / US2021 / 027812, entitled “A Method for Consolidating Heterogeneous Healthcare Data.” The output of this process are the target local coding systems and the collected values to be mapped by the concept mapping process.

[0094] In data store transformation, local coding systems in the data store may be mapped to codeable concepts in the concept cluster one by one.

[0095] At least two phases for mapping codeable concepts in the concept mapping process may exist: the initial concept mapping phase and the subsequent cluster rematching phase. Both of these phases are described below.

[0096] In the initial concept mapping phase, a unique identifier matching process and a concept cluster matching process may be performed. The unique identifier matching process includes extracting unique identifiers (402) and matching them to a codeable concept in the concept cluster. This is useful when the unique identifier is coming from an existing local coding system that has known mappings to the standardized coding system used in the concept cluster.

[0097] The concept cluster matching process (405) includes extracting the description (403) and performing free text matching to match to an existing codeable concept in the concept cluster through use of a vector embedding model, a vector database containing concept cluster vector embeddings; a database full-text search; or an n-gram matching algorithm.

[0098] The sameness of the two codeable concepts may be resolved through the concept cluster matching process of blocks (406) and (408). The possible outcomes include a full match, a high confidence match, and unmatched.

[0099] A full match (407) means the following. A concept mapping is considered a “full match” if there is a translation match (404), and there are no descriptions to be matched; or if the translation match and the concept cluster match (405) has a high degree of sameness, as determined through the concept cluster. The confidence score will be between 0.95-1. In this case, a translation match may be a match between the identifier of the local coding system to an identifier of the standardized coding system.

[0100] A high confidence match (409) is defined as follows. A concept mapping is considered a “high confidence match” if the translation match does not have a high degree of sameness with the concept cluster match, but the confidencescore of one of the translation match and the concept cluster match passes a preconfigured threshold, and the other one does not. The confidence score will be above 0.8.

[0101] An unmatched (410) codeable concept is defined as follows. A codeable concept that does not meet the criteria for a full match or a high confidence match. There is no confidence score for unmatched codeable concepts.

[0102] If either of the unique identifier field, or the description field, does not exist then the possible outcome will either be a high confidence match or unmatched, but cannot be a full match.

[0103] For high confidence matches and full matches, the local coding system entry will be linked to a standardized coding system through a translation table (e.g., the table 500 of FIG. 5) alongside the confidence score.

[0104] In the cluster rematching phase shown in FIG. 7, codeable concepts that had a full match (704) or a high confidence match (703) are used to form one or more clusters (702) of codeable concepts, based on the codeable concepts relatedness and sameness in the concept cluster. One codeable concept may be part of multiple clusters. Clusters may be grouped together by the relation to codeable concepts of interest such as disorders, or by the center codeable concept in the cluster based on ontological or vector distances. In one embodiment, diabetes mellitus, non-insulin-dependent can be clustered as a diabetes related codeable concept, as well as a disorder (e.g., as shown in the clusters 600 of FIG. 6).

[0105] Turning to FIG. 7, after the codeable concept clusters are created, unmatched codeable concepts are reprocessed through the concept cluster based on context provided starting from the closest cluster to the unmatched codeable concept. The unmatched codeable concepts (701) may be reclassified into a medium confidence match bucket, a low confidence match bucket, and a quarantined unmatched bucket. Each of the buckets is described below.

[0106] The medium confidence match (705) is a codeable concept with a confidence score of over 0.5, but less than 0.8. The low confidence match (706) is a codeable concept with a confidence score of over 0.3, but less than 0.5.

[0107] The quarantined unmatched (707) is a codeable concept where no reasonable match is found. Human intervention may be performed to manually match the codeable concept, or to indicate that the codeable concept should be excluded as an event of interest.

[0108] Clustering codeable concepts, as performed above, may help narrow down the search. For example, if “DM” was an unmatched codeable concept and “Disorder” was a recognized cluster, then, “Diabetes Mellitus” could be considered a medium confidence match.

[0109] In one embodiment, “Mental Health” is a recognized cluster. Then, an unmatched codeable concept such as “ADM” could be considered as a “Mental Health Admission” as a low confidence match. Similar to full matches and high confidence matches, the local coding system entry with a medium confidence match will also be linked to a codeable concept in the concept cluster. At the end of the data store transformation process, each selected entry in the data store is mapped to one or more codeable concepts within the concept cluster through a translation table and assigned a confidence score.

[0110] Another process is the journey reconstruction process shown in FIG. 8. For the journey reconstruction process (800), entries in the data store, which now have local coding systems mapped to the concept cluster, are transformed into the golden analytical model database, which may be an embodiment of the golden analytical model (116) of FIG. 1. The journey reconstruction process (800) takes as input any collection of timestamped data store entries (z.e., through a real time HL7v2 stream, Kafka stream, logs, or tabular formats) to construct journeys. For meaningful analysis, the data must include timestamps of when the event of interest occurred (z.e., in healthcare, an event of interest may be an admission to the emergency department, or the diagnosis of aparticular condition) and are associated to an individual subject (z.e., a patient). Note that data may come from one or more sources of different types (such as data lake, clinical information systems, hospital information systems, etc.) which must be normalized and consolidated. The method for normalizing and consolidating data is not within the scope of this patent, but one such approach is described in PCT Patent Publication No. PCT / US2021 / 027812 directed to “A Method for Consolidating Heterogeneous Healthcare Data”.

[0111] The reconstruction process includes code extraction, concept translation and dimension transformation. Code extraction is responsible for extracting any local coding systems used in the data store entry alongside its timestamp. This may generate a JSON pay load (902) on the left side of the image shown in FIG. 9. In one embodiment, the extraction may refer to the translation table on all possible local coding system values. Concept translation takes the output of the code extraction step and translates one or more local coding system codes to standardized codeable concepts through the translation table. Once translated, codeable concepts along with their associated timestamps and subjects are passed to the dimension transformer. The confidence score is also captured and included in the final fact, which is described below. An example JSON output (905) is shown on the right side of the image shown in FIG. 9. The dimension transformer is responsible for creating models of the data in the golden analytical model (GAM) based on the codeable concepts obtained from concept translation. In one embodiment, one or more codeable concepts may be transformed into one or more facts which will be included in the golden analytical model. From the concept translation phase, the dimension transformer may categorize each codeable concept into multiple respective dimension.

[0112] In one embodiment, dimensions are subject (who), concept (what), location (where), and actor (who) which are associated with a fact which is augmented by a time dimension (when). “Emergency Department” may be categorized as a location, and “Admission” may be categorized as the codeableconcept. Then, the dimension transformer will create a fact with the location and concept mentioned above. In another embodiment, both “Triage” and “Acuity Level 3” are categorized as concepts. Then, the dimension transformer will create two facts instead of one. Collectively a fact and its associated dimension represent an event of interest. By way of an example, in healthcare, a fact may represent that Alice (who), was admitted to the emergency department (what), at Acme Hospital (where), by Bob (who), on June 4th, 2023, 11 :53pm, Eastern Standard Time (when).

[0113] At the end of the journey reconstruction process, data store entries of events of interest within the data store are fully transformed into the golden analytical model, ready for further processing and journey queries.

[0114] Returning to the data store (102) in FIG. 1, the golden analytical model (GAM) (116) is a model that is designed to support the queries used to reconstruct journeys and allows for flexibility in analysis by being able to define additional domains of interest through joining additional dimensions. In one embodiment, the model (116) may be a star schema. FIG. 10 shows an example schema (1000) of the GAM in a star schema. The fact table (001) holds the events of interest with timestamps, as well as foreign keys to the dimension tables. There are 4 dimension tables. 1) the subject dimension table, contains data related to the subject of the event of interest, such as subject gender, and the subject date of birth; 2) the concept dimension table, contains all concepts as defined in the concept cluster and refers to the codeable concepts that were created by the data store transformation process; 3) the actor dimension table, contains data related to the actor of the event of interest, such as the actor profession, actor team, and the actor identifier; 4) the location dimension table, contains data regarding the location in which the event of interest has happened. To enable historical consistency, the dimension tables are modeled with slowly changing dimension type 2 data. In another embodiment, these dimensions may be flattened and represented in a traditional big data columnar store.

[0115] At minimum, a codeable concept dimension is defined for journeys to be created. In some journey types, there would be some attribute to act as a grouping category for a subset of codeable concepts. In one embodiment, the attribute is the subject dimension. For example, in a healthcare context, this may be a patient, as usually the healthcare industry concerns with patientcentric journeys.

[0116] Additional dimensions may be identified as per the individual use case. For example, embodiments may support extracting information about the location of an individual at the time of the event, or additional actors that were involved, such as a clinician. Each dimension may also optionally refer to one or more concepts that were created in the concept cluster.

[0117] Using a time series backbone for the golden analytical model allows for the addition of other time-based dimensions that may be outside of the given input data. As the fact table has timestamps for when each codeable concept was applied, analysis may be performed against any other external set of time series data, such as weather data or data around other environmental scores. Similar to the additional dimensions, the time-based dimensions may also refer to codeable concepts that were created in the concept cluster. These additional dimensions can be added and removed as needed to support the population analysis needs of the user, however, the subject and concept dimensions will always be needed in the GAM model as a minimum schema.

[0118] FIG. 11 shows an example of the journey discovery process (1100). From the golden analytical model, the journey discovery process includes discovering concepts of value to one or more codeable concepts of interest. A codeable concept of interest is a codeable concept from the concept dimension that is important or that has value in the context of the domain. For example, a diabetes clinic may set “Type 2 Diabetes” and “Diabetic Retinopathy” as codeable concepts of interest, whereas an Emergency Department may set “Admission” and “Discharged” as codeable concepts of interest.

[0119] FIG. 12 shows an example of the workflow (1200) to create population sequences in accordance with one or more embodiments.

[0120] To perform journey discovery, population sequences are created from the golden analytical model using the following steps. First, the golden analytical model is grouped by the subject, and a sequence is created on a dimension and is ordered by the timestamps of the concepts in the sequence. Then, after the subject sequence is created, grouping populations by the unique sequences is performed to obtain, as output, a population with a count of how many subjects have gone through that sequence. This is referred to as the “population sequence”.

[0121] In the example of FIG. 12, aggregating by one dimension is shown. However, embodiments include aggregating in multiple dimensions at the same time. Similarly, it can be easily seen that the same thing can be done in timestamps, in order to find out the delta between each step in the sequence.

[0122] Returning to FIG. 11, one or more codeable concepts of interest are selected. Once the codeable concepts of interest are selected, the concepts that have a high degree of sameness, also known as child codeable concepts, according to the concept cluster are selected for filtering. Each codeable concept of interest, along with its associated child codeable concepts, may be referred to as a “codeable concept of interest family”.

[0123] The selected codeable concept of interest families may be chained together using the operators AND or OR, as a concept Boolean expression. If the AND operator is used, it means that the population sequence must contain both the first codeable concept of interest family, and the second codeable concept of interest family. If the OR operator is used, it means that it may contain either the first codeable concept of interest family, or the second codeable concept of interest family, or both. At the end of this process, zero or more population sequences that satisfy the concept Boolean expression will remain.

[0124] Continuing with FIG. 11, predefined filters, configured according to the codeable concepts of interest, are applied. The population sequences may be further filtered through applying filters. Filters may be applied on any dimension. An example of a filter in the subject dimension is age is greater than 60. An example of a filter in the GAM concept dimension is sameness with codeable concept of interest greater than 0.7. An example of a filter in the location dimension is any location that is not within an Emergency Department.

[0125] At the end of this process, zero or more population sequences that satisfy the filters remain. Next, sequence graphs are generated. From the remaining filtered population sequences, a series of sequence graphs, one per dimension, are generated.

[0126] Population sequences may be aggregated and represented as a directed acyclic graph (DAG), where a node represents a concept occurring in the nthstep of the sequence; an edge represents that there exists at least one transition between two concepts; and the edge weight represents the number of total transitions in the data set. The edge may also store additional information such as the time taken for each transition as a distribution. A constraint of the DAG is that the incoming edge weight is greater or equal to the outgoing edge weight.

[0127] Depending on the sparseness of the sequences, the number of nodes and the number of edges, a graph may be represented by a graph structure, an adjacency matrix, or an adjacency list to aid with transformation at a later stage.

[0128] An example of the Sequence graph is presented in FIG. 13. Turning to the above FIG. 13, node 1A represents the first steps of the population sequences, and 6A, 6B, and 6C represents the final steps in all population sequences. Consider a query where the concept selected was “Diabetes Mellitus”. Then, all types of diabetes mellitus, such as type 1 diabetes mellitus, type 2 diabetes mellitus, monogenic diabetes, diabetes mellitus ketosis-prone, etc., may be selected and consolidated into node 1A, as all population sequences contain node 1A. The edge weights represent the number of subjects that have made atransition from one concept to the next concept. For example, if the edge weight was equal to 50, then 50 subjects have gone from 1 A to 2B. At the end of this process, a graph representation of the population sequences may be generated for each dimension in the golden analytical model.

[0129] FIG. 14 shows an example of a graph transformation. One or more transformations may be performed on the graph representation (1300) for each dimension.

[0130] Firstly, two or more nodes in the same level may be consolidated based on its sameness score in the concept cluster. For example, if node 3 A and 3B represent “Hyperglycemia” and “Glucose in blood specimen above reference range” respectively, it may be consolidated into node 3AB as shown in the workflow (1400) of FIG. 14.

[0131] The edges may also be consolidated accordingly. In the example of FIG. 14, if node 2B has 50 subjects flowing to node 3A and 50 subjects flowing to node 3B, then after the transformation is applied, node 2B will have 100 subjects flowing into node 3 AB.

[0132] Secondly, nodes may be consolidated if the nodes have a high degree of relatedness, form a linear flow, and if time distribution is consistent. In the workflow (1500) of FIG. 15, if the conditions hold then, node 1A, node 2A, and node 3 A may be collapsed into node 3B.

[0133] Note that after nodes 1A, 2A, and 3A have been consolidated into node 3B, the edge which previously connected node 3 A to node 4A is now consistent with the edge from node 3B to node 4A and does not change its edge weight.

[0134] At the end of the transformation process for consolidating nodes, one or more transformations may have been performed on the graph representation for each dimension resulting in simplified graph representations ready for further processing.

[0135] The method of determining consolidation is as follows. First, one or more codeable concepts of interest is identified.

[0136] Next, a closeness score may be computed. For each dimension, a closeness score may be computed for each concept in the graph representation relative to the codeable concept of interest families.

[0137] The closeness score is computed by a linear combination applied against a set of heuristics. An example formula (1600) is described in FIG. 16. Each heuristic used contributes to the final value of the closeness score according to the weight assigned with the heuristic. The final value for the closeness score is the sum of the weighted results from all of the heuristics. In one embodiment, the set of heuristics used to compute the closeness score includes concept cluster sameness score, concept cluster relatedness score, distance, percentage of population, and time.

[0138] Sameness score and relatedness score: using the concept cluster, the system will determine the pairwise sameness scores and relatedness scores between a codeable concept of interest and every other node in the graph.

[0139] Distance: run a shortest path algorithm from the codeable concept of interest to a codeable concept in the directed graph representation, ignoring edge weights. The output of the distance heuristic is the number of nodes between the two nodes.

[0140] Percentage of population: the output of the percentage heuristic is the number of subjects that passed through the node divided by the total number of subjects represented in the graph. For example, if 1 subject had codeable concept X, and the graph represents 100 subjects in total, then the output of this heuristic will be 0.01.

[0141] Time: run a shortest path algorithm from the codeable concept of the interest to a codeable concept in the directed graph representation with the median time to transition as the edge weight. The output of the time heuristic is the median time taken to transition from a codeable concept to the codeableconcept of interest. After the closeness scores have been calculated in each dimension, the closeness scores are treated as input to a final heuristic function with adjustable weights. This final heuristic function returns the final closeness score for each dimension, relative to the codeable concept of interest families.

[0142] At the end of this last step, the final closeness score, the codeable concepts of interest and filters used are stored together in a data store. These three variables, taken together, are considered a journey which is centered around the codeable concepts of interest used. From the journey data store, sequence diagrams, such as the Sankey diagram displayed in the user interface (1700) shown in FIG. 17, may be generated to showcase only the codeable concepts that meet a closeness score threshold.

[0143] Turning to FIG. 18, the user interface (1800) includes the timeline view (1802) and may be displayed on a computing system. The horizontal axis of the timeline view identifies the dates that events occur, and the vertical axis identifies the type of information that is recorded for the events. The timeline view may be for a single subject identified in the window (1820).

[0144] The legend (1805) identifies the type of information represented by the points within the timeline view (1802). Different types of information may be coded with different colors. For example, “Activities” may be coded with the color yellow, “Disorders” may be coded with the color green, a “Finding” may be coded with the color purple, and an “Organization” may be coded with the color blue.

[0145] The point (1808) indicates that an event occurred for which a codeable concept relating to a disorder was recorded. Other types of information (e.g., for a finding) were also recorded on the same date as indicated by the additional points at the same time on the timeline (1802). The point (1808) may be selected to trigger the display of the window (1810).

[0146] The window (1810) displays additional information relating to the point (1808). The window (1810) indicates that the underlying concept was“respiratory distress” and was recorded at the date and time identified as “2023- 10-26 14:21 :00”.

[0147] Turning to FIG. 19, the user interface (1900) includes the summary view (1902) and may be displayed on a computing system. The summary view organizes the information by type into multiple sections. The summary view may be for a single subject identified in the window (1920).

[0148] The section (1902) is labeled “Health Concerns” and includes “35” items in a list. The list includes the item (1908).

[0149] The item (1908) is for a codeable concept identified by the text “Drusen of bilateral maculae (finding)”. The item (1908) is selected to trigger a display of the window (1910).

[0150] The window (1910) displays information related to the item (1908). The information in the window (1910) is displayed in a structured text format that includes several key value pairs of information. For example, the key “text” is paired with the value “Drusen of bilateral maculae (finding)”, which is displayed in the item (1908) on the user interface (1900). The window may display the local terminology “345261000119102” from the local coding system “http / / snomed.info / sct” which may be mapped to the codeable concept “C4749233” in the standard coding system.

[0151] Turning to FIG. 20, the user interface (2000) includes the population journey view (2002) shown as a Sankey diagram. The population journey view (2002) depicts a sequence of events that multiple subjects experienced.

[0152] The sequence of events includes five steps, labeled “Step 1” through “Step 5”. The date that each event occurred for each subject may be different, but the sequence of the types of events experienced by the different subjects may be similar.

[0153] At Step 1, the element (2022) indicates that 43,400 subjects in the population received urgent care service. Represented by other nodes in stepone, an additional 20 subjects received mental health service, 10 subjects received palliative care, and 1 subject received an urgent hospital admission.

[0154] At Step 2, the element (2025) indicates that 43,400 subjects identified with the element (2022) received additional urgent care service.

[0155] At Step 3, the element (2028) indicates that about 24,900 of the subjects identified with the element (2025) did not receive further care. The element (2030) indicates that about 18,500 of the subjects identified with the element (2025) received additional urgent care service.

[0156] At Step 4, the element (2032) indicates that about 5,500 of the subjects identified with the element (2030) did not receive further care. The element (2035) indicates that about 13,100 of the subjects identified with the element (2030) received additional urgent care service.

[0157] At Step 5, the element (2038) indicates that about 5,400 of the subjects identified with the element (2035) did not receive further care. The element (2040) indicates that about 7,700 of the subjects identified with the element (2035) received additional urgent care service.

[0158] Turning to FIG. 21, the user interface (2100) includes the dashboard view (2102). The dashboard view (2102) displays multiple tiles in multiple rows. The information within the individual tiles may be based on the filter parameters specified in the filter bar (2103) of the dashboard view (2102). A first row In the dashboard view (2102) includes the tiles (2105), (2108), and (2110). A second row in the dashboard view (2102) includes the tiles (2122) and (2125). A third row in the dashboard view (2102) includes the tile (2130). A fourth row in the dashboard view (2102) includes the tiles (2150) and (2152). The tile (2105) displays statistical information from the population. The tiles (2108), (2122), and (2152) display bar charts with information from the population. The tiles (2110), (2125), and (2150) display pie charts with information from the population. The tile (2130) displays a distribution chart.

[0159] The distribution chart includes the elements (2132), (2135), and (2138). The element (2132) indicates that 24,265 subjects of the population (amounting to about 99%) experienced an event in which a codeable concept that is labeled as a “disorder”. The element (2135) indicates that 11,445 of the subjects from the element (2132) (amounting to about 47%) experienced an event in which a codeable concept was recorded that is labeled as a “disease or syndrome”. The element (2138) indicates that 2,445 of the subjects from the element (2135) (amounting to about 10%) experienced an event in which a codeable concept was recorded that is labeled as a “circulatory disorder”.

[0160] One or more embodiments may be implemented on a computing system specifically designed to achieve an improved technological result. When implemented in a computing system, the features and elements of the disclosure provide a significant technological advancement over computing systems that do not implement the features and elements of the disclosure. Any combination of mobile, desktop, server, router, switch, embedded device, or other types of hardware may be improved by including the features and elements described in the disclosure.

[0161] For example, as shown in FIG. 22 A, the computing system (2200) may include one or more computer processor(s) (2202), non-persistent storage device(s) (2204), persistent storage device(s) (2206), a communication interface (2208) (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.), and numerous other elements and functionalities that implement the features and elements of the disclosure. The computer processor(s) (2202) may be an integrated circuit for processing instructions. The computer processor(s) (2202) may be one or more cores or micro-cores of a processor. The computer processor(s) (2202) includes one or more processors. The computer processor(s) (2202) may include a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), combinations thereof, etc.

[0162] The input device(s) (2210) may include a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device. The input device(s) (2210) may receive inputs from a user that are responsive to data and messages presented by the output device(s) (2212). The inputs may include text input, audio input, video input, etc., which may be processed and transmitted by the computing system (2200) in accordance with one or more embodiments. The communication interface (2208) may include an integrated circuit for connecting the computing system (2200) to a network (not shown) (e.g, a local area network (LAN), a wide area network (WAN) such as the Internet, mobile network, or any other type of network) or to another device, such as another computing device, and combinations thereof.

[0163] Further, the output device(s) (2212) may include a display device, a printer, external storage, or any other output device. One or more of the output devices may be the same or different from the input device(s) (2210). The input and output device(s) may be locally or remotely connected to the computer processor(s) (2202). Many different types of computing systems exist, and the aforementioned input and output device(s) may take other forms. The output device(s) (2212) may display data and messages that are transmitted and received by the computing system (2200). The data and messages may include text, audio, video, etc., and include the data and messages described above in the other figures of the disclosure.

[0164] Software instructions in the form of computer readable program code to perform embodiments may be stored, in whole or in part, temporarily or permanently, on a non-transitory computer readable medium such as a solid state drive (SSD), compact disk (CD), digital video disk (DVD), storage device, a diskette, a tape, flash memory, physical memory, or any other computer readable storage medium. Specifically, the software instructions may correspond to computer readable program code that, when executed by the computer processor(s) (2202), is configured to perform one or moreembodiments, which may include transmitting, receiving, presenting, and displaying data and messages described in the other figures of the disclosure.

[0165] The computing system (2200) in FIG. 22A may be connected to or be a part of a network. For example, as shown in FIG. 22B, the network (2220) may include multiple nodes (e.g., node X (2222), node Y (2224)). Each node may correspond to a computing system, such as the computing system shown in FIG. 22A, or a group of nodes combined may correspond to the computing system shown in FIG. 22A. By way of an example, embodiments may be implemented on a node of a distributed system that is connected to other nodes. By way of another example, embodiments may be implemented on a distributed computing system having multiple nodes, where each portion may be located on a different node within the distributed computing system. Further, one or more elements of the aforementioned computing system (2200) may be located at a remote location and connected to the other elements over a network.

[0166] The nodes (e.g., node X (2222), node Y (2224)) in the network (2220) may be configured to provide services for a client device (2226), including receiving requests and transmitting responses to the client device (2226). For example, the nodes may be part of a cloud computing system. The client device (2226) may be a computing system, such as the computing system shown in FIG. 22 A. Further, the client device (2226) may include or perform all or a portion of one or more embodiments.

[0167] The computing system of FIG. 22A may include functionality to present data (including raw data, processed data, and combinations thereof) such as results of comparisons and other processing. For example, presenting data may be accomplished through various methods. Specifically, data may be presented by being displayed in a user interface, transmitted to a different computing system, and stored. The user interface may include a graphical user interface (GUI) that displays information on a display device. The GUI may include various GUI widgets that organize what data is shown as well as how data is presented to a user. Furthermore, the GUI may present data directly to the user,e.g., data presented as actual data values through text, or rendered by the computing device into a visual representation of the data, such as through visualizing a data model.

[0168] As used herein, the term “connected to” contemplates multiple meanings. A connection may be direct or indirect (e.g., through another component or network). A connection may be wired or wireless. A connection may be a temporary, permanent, or semi-permanent communication channel between two entities.

[0169] The various descriptions of the figures may be combined and may include or be included within the features described in the other figures of the application. The various elements, systems, components, and steps shown in the figures may be omitted, repeated, combined, or altered as shown in the figures. Accordingly, the scope of the present disclosure should not be considered limited to the specific arrangements shown in the figures.

[0170] In the application, ordinal numbers (e.g., first, second, third, etc.) may be used as an adjective for an element (i.e., any noun in the application). The use of ordinal numbers is not to imply or create any particular ordering of the elements nor to limit any element to being only a single element unless expressly disclosed, such as by the use of the terms “before”, “after”, “single”, and other such terminology. Rather, ordinal numbers distinguish between the elements. By way of an example, a first element is distinct from a second element, and the first element may encompass more than one element and succeed (or precede) the second element in an ordering of elements.

[0171] Further, unless expressly stated otherwise, the conjunction “or” is an inclusive “or” and, as such, automatically includes the conjunction “and,” unless expressly stated otherwise. Further, items joined by the conjunction “or” may include any combination of the items with any number of each item, unless expressly stated otherwise.

[0172] In the above description, numerous specific details are set forth in order to provide a more thorough understanding of the disclosure. However, it will be apparent to one of ordinary skill in the art that the technology may be practiced without these specific details. In other instances, well-known features have not been described in detail to avoid unnecessarily complicating the description. Further, other embodiments not explicitly described above may be devised which do not depart from the scope of the claims as disclosed herein. Accordingly, the scope should be limited only by the attached claims.

Claims

CLAIMSWhat is claimed is:

1. A method comprising: clustering a plurality of codeable concepts; performing a data store transformation on the plurality of codeable concepts; performing journey reconstruction from the data store transformation; and performing journey discovery using the journey reconstruction.

2. The method of claim 1, wherein clustering the plurality of codeable concepts comprises: generating a plurality of concept clusters from the plurality of codeable concepts by calculating a sameness factor and a relatedness factor between pairs of codeable concepts from the plurality of codeable concepts.

3. The method of claim 1, wherein the plurality of codeable concepts use a standardized coding system and wherein performing a data store transformation comprises matching a plurality of codeable concepts using a local coding system to a plurality of concept clusters generated from a standardized coding system by: extracting a plurality of identifiers using the local coding system from the plurality of codeable concepts using a local coding system, extracting a plurality of descriptions using the local coding system from the plurality of codeable concepts using a local coding system, matching the plurality of identifiers using the local coding system to a plurality of identifiers using a standardized coding system to generate a plurality of identifier match values, matching the plurality of descriptions using the local coding system to the plurality of concept clusters using the standardized coding system to generate a plurality of concept cluster match values, and generating a confidence value for each codeable concept from the plurality of codeable concepts using a local coding system based on the plurality of identifier match values and based on theplurality of concept cluster match values, wherein the confidence value for each codeable concept corresponds to a label.

4. The method of claim 1, wherein performing a data store transformation comprises: rematching a plurality of entries labeled as a first type of matches and a second type of matches to supplement the plurality of concept clusters.

5. The method of claim 1, wherein performing a data store transformation comprises: reclassifying a plurality of unmatched codeable concepts subsequent to supplementing the plurality of concept clusters.

6. The method of claim 1, wherein performing journey reconstruction from the data store transformation comprises: extracting a plurality of entries comprising a plurality of local codeable concepts using a local coding system; translating the plurality of local codeable concepts to a plurality of standardized codeable concepts; and transforming a plurality of standardized codeable concepts using a plurality of dimensions.

7. The method of claim 1 , wherein performing the j ourney discovery using the j ourney reconstruction comprises: generating population sequences by: grouping a plurality of codeable concepts by subject to identify a plurality of subject sequences for a plurality of subjects, and grouping the plurality of subject sequences by subject sequence to generate a plurality of sequence counts for the plurality of subject sequences; filtering the population sequences using one or more of a plurality of dimensions; generating one or more sequence graphs from the plurality of population sequences; and consolidating a plurality of nodes to revise the one or more sequence graphs.

8. The method of claim 1, further comprising: presenting a codeable concept of the plurality of codeable concepts using a sequence graph in a view.

9. The method of claim 1, further comprising: presenting a codeable concept in a timeline view.

10. The method of claim 1, further comprising: presenting a codeable concept in one of a population journey view and a dashboard view.

11. A system comprising: at least one processor; and an application that, when executing on the at least one processor, performs operations comprising: clustering a plurality of codeable concepts, performing a data store transformation on the plurality of codeable concepts, performing journey reconstruction from the data store transformation, and performing journey discovery using the journey reconstruction.

12. The system of claim 11, wherein clustering the plurality of codeable concepts comprises: generating a plurality of concept clusters from the plurality of codeable concepts by calculating a sameness factor and a relatedness factor between pairs of codeable concepts from the plurality of codeable concepts.

13. The system of claim 11, wherein the plurality of codeable concepts use a standardized coding system and wherein performing a data store transformation comprises matching a plurality of codeable concepts using a local coding system to a plurality of concept clusters generated from a standardized coding system by: extracting a plurality of identifiers using the local coding system from the plurality of codeable concepts using a local coding system,extracting a plurality of descriptions using the local coding system from the plurality of codeable concepts using a local coding system, matching the plurality of identifiers using the local coding system to a plurality of identifiers using a standardized coding system to generate a plurality of identifier match values, matching the plurality of descriptions using the local coding system to the plurality of concept clusters using the standardized coding system to generate a plurality of concept cluster match values, and generating a confidence value for each codeable concept from the plurality of codeable concepts using a local coding system based on the plurality of identifier match values and based on the plurality of concept cluster match values, wherein the confidence value for each codeable concept corresponds to a label.

14. The system of claim 11, wherein performing a data store transformation comprises: rematching a plurality of entries labeled as a first type of matches and a second type of matches to supplement the plurality of concept clusters.

15. The system of claim 11, wherein performing a data store transformation comprises: reclassifying a plurality of unmatched codeable concepts subsequent to supplementing the plurality of concept clusters.

16. The system of claim 11, wherein performing journey reconstruction from the data store transformation comprises: extracting a plurality of entries comprising a plurality of local codeable concepts using a local coding system; translating the plurality of local codeable concepts to a plurality of standardized codeable concepts; and transforming a plurality of standardized codeable concepts using a plurality of dimensions.

17. The system of claim 11, wherein performing the journey discovery using the journey reconstruction comprises: generating population sequences by: grouping a plurality of codeable concepts by subject to identify a plurality of subject sequences for a plurality of subjects, and grouping the plurality of subject sequences by subject sequence to generate a plurality of sequence counts for the plurality of subject sequences; filtering the population sequences using one or more of a plurality of dimensions; generating one or more sequence graphs from the plurality of population sequences; and consolidating a plurality of nodes to revise the one or more sequence graphs.

18. The system of claim 11, further comprising: presenting a codeable concept of the plurality of codeable concepts using a sequence graph in a view.

19. The system of claim 11, further comprising: presenting a codeable concept in a timeline view.

20. The system of claim 11, further comprising: presenting a codeable concept in one of a population journey view and a dashboard view.

21. A non-transitory computer readable medium comprising stored instructions executable by at least one processor to perform: clustering a plurality of codeable concepts; performing a data store transformation on the plurality of codeable concepts; performing journey reconstruction from the data store transformation; and performing journey discovery using the journey reconstruction.

Citation Information

Patent Citations

  • Systems, methods, and apparatus for computer-assisted full medical code scheme to code scheme mapping

    US20120110016A1

  • Cognitive Mapping and Validation of Medical Codes Across Medical Systems

    US20170235887A1

  • Systems and Methods for Medical Concept Mapping

    US20220293283A1