DATA REPOSITORY, SYSTEM, AND METHOD FOR COHORT SELECTION - Patent application

JP2024532381A5Pending Publication Date: 2025-09-08GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024513211
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-29
Filing Date
2022-08-31
Publication Date
2025-09-08

AI Technical Summary

Technical Problem

Existing medical data systems struggle to efficiently identify and analyze patient cohorts based on complex and scattered medical records, making it difficult to derive insights for precision medicine and treatment strategies tailored to individual variability.

Method used

A computer system and method for cohort selection that integrates and analyzes health insurance claims data, medical data, and genomic data to identify patient cohorts based on diagnosis timing and codes, using engines and models to process and categorize patients with similar biological conditions.

Benefits of technology

Enables accurate identification of patient cohorts for precision medicine, allowing for tailored treatment strategies by efficiently processing and categorizing patients based on their medical history and genetic profiles, improving the effectiveness of treatment and prevention strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer may access one or more medical data tables that store medical insurance transaction data for a plurality of patients. The one or more medical data tables include a date column and a diagnosis column. The computer may identify a set of patients having a biological condition based on the diagnosis column. The set of patients can be from among the plurality of patients. The computer may determine, for each patient in the set of patients, a first date on which the patient was diagnosed with the biological condition. The computer may identify a cohort of patients from the set of patients based on the diagnosis column and the date column. The cohort of patients may lack a diagnosis from a set of biological conditions that is associated with a date that occurs during a predefined time window before the first date on which the patient was diagnosed with the biological condition.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (Claim of priority and incorporation by reference) This application claims priority to U.S. Provisional Patent Application No. 63 / 238,851, filed August 31, 2021, and entitled "Data Repository, System, and Method for Cohort Selection," U.S. Provisional Patent Application No. 63 / 250,912, filed September 30, 2021, and entitled "Computer Architecture for Generating a Reference Data Table," PCT Application No. PCT / US2022 / 032250, filed June 3, 2022, and entitled "Computer Architecture for Generating an Integrated Data Repository," and PCT Application No. PCT / US2022 / 038941, filed July 29, 2022, and entitled "Computer Architecture for Identifying Lines of Therapy," the entire contents of which are each incorporated herein by reference in their entirety.

[0002] Implementations relate to computer architectures. Some implementations relate to the use of computer systems for monitoring the treatment and progression of biological conditions, including medical conditions and diseases. Some implementations relate to data repositories, systems, and methods for cohort selection. [Background technology]

[0003] Precision medicine is an emerging approach to disease treatment and prevention that considers the individual variability in one or more of genes, environment, and lifestyle for each person. This approach can allow doctors and researchers to more accurately predict which treatment and prevention strategies for a particular disease will work in which groups of people. This is in contrast to a one-size-fits-all approach, where disease treatment and prevention strategies are developed for the average person, with little consideration given to differences between individuals. For some biological conditions, such as cancer, different people receive very different treatments. Identifying a cohort of patients with similar biological conditions (e.g., medical conditions, diseases, or genetic profiles) can be desirable for studying the treatment or progression of a biological condition.

[0004] When a patient undergoes a medical procedure, the medical service provider generates a medical record that indicates the treatment received by the patient. In addition, the medical record may indicate one or more diagnoses corresponding to the patient. The information contained within the medical record is typically complex and / or difficult to analyze so that insight can be determined regarding the medical treatment provided to the patient. Summary of the Invention [Means for solving the problem]

[0005] Detailed Description The following description and drawings sufficiently illustrate specific implementations to enable those skilled in the art to practice them. Other implementations may incorporate structural, logical, electrical, process, and other changes. Portions and features of some implementations may be included in or substituted for those of other implementations. Implementations set forth in the claims encompass all available equivalents of those claims. [Brief description of the drawings]

[0006] [Figure 1] FIG. 1 illustrates an exemplary system in which cohort selection can be implemented.

[0007] [Diagram 2] FIG. 2 illustrates an example of processing insurance claims data to extract information, according to one or more implementations.

[0008] [Diagram 3] FIG. 3 illustrates an example of patient information that may be stored, according to one or more implementations.

[0009] [Figure 4] FIG. 4 is a flowchart of an exemplary method for identifying primary lung cancer patients, according to one or more implementations.

[0010] [Diagram 5] FIG. 5 is a flowchart of an exemplary method for identifying primary lung cancer patients and rare cases, according to one or more implementations.

[0011] [Figure 6] FIG. 6 is a flowchart of an example method for identifying a last effective date, according to one or more implementations.

[0012] [Figure 7] FIG. 7 is a flowchart of a first exemplary process associated with assigning patients to cohorts, according to one or more implementations.

[0013] [Figure 8] FIG. 8 is a flowchart of a second exemplary process associated with assigning patients to cohorts, according to one or more implementations.

[0014] [Figure 9] FIG. 9 is a flowchart of an exemplary process associated with identifying a cohort of patients, according to one or more implementations.

[0015] [Figure 10] FIG. 10 illustrates an example medical data table, according to one or more implementations.

[0016] [Figure 11] FIG. 11 illustrates an example architecture for generating an integrated data repository containing multiple types of healthcare data, according to one or more implementations.

[0017] [Figure 12] FIG. 12 illustrates an exemplary framework that corresponds to an arrangement of data tables in a unified data repository, according to one or more implementations.

[0018] [Figure 13] FIG. 13 illustrates an architecture for generating one or more data sets from information retrieved from a data repository that integrates health-related data from a number of sources, in accordance with one or more implementations.

[0019] [Figure 14] FIG. 14 illustrates an architecture for generating an integrated data repository including de-identified health insurance claims data and de-identified genomics data, according to one or more implementations.

[0020] [Figure 15] FIG. 15 illustrates a framework for generating a dataset by a data pipeline system based on data stored by a unified data repository, according to some implementations.

[0021] [Figure 16] FIG. 16 illustrates a system for determining a cohort of patients having at least a primary diagnosis of a biological condition, according to one or more implementations.

[0022] [Figure 17] FIG. 17 is a block diagram of a computing machine according to one or more implementations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] As discussed above, identifying a cohort of patients having similar biological conditions (e.g., medical conditions or diseases) may be desirable to study treatment or progression of the biological conditions. Some implementations are directed to identifying the cohort of patients based on medical data, pharmacy data, and / or insurance transaction data. In some implementations, a computer accesses one or more medical data tables that store medical transaction data for a plurality of patients. The one or more medical data tables comprise a date column and a diagnosis column. The computer identifies a set of patients having a defined biological condition (e.g., lung cancer) based on the diagnosis column. The set of patients is from among the plurality of patients. For each patient in the set of patients, the computer determines the first date on which the patient received a diagnosis of the defined biological condition (e.g., the first date on which the patient was diagnosed with lung cancer). The computer also identifies a cohort of patients from the set of patients based on a diagnosis code included in the medical data table. In one or more embodiments, a cohort of patients may be determined based on the timing at which a diagnosis was recorded versus the timing at which one or more additional biological conditions were diagnosed for the patient. For example, in a situation where multiple biological conditions are diagnosed for a patient within a time window, the patient may be placed in a cohort that includes patients with multiple diagnoses. Additionally, in a situation where a single diagnosis is identified for a patient within a predefined time window, and the diagnosis is the first diagnosis for the patient, the computer may determine that the patient is included in a cohort that corresponds to the biological condition associated with the single diagnosis. The computer provides an output representative of the cohort.

[0024] Aspects of the present technology may be implemented as part of a computer system. The computer system may be one physical machine, or may be distributed among multiple physical machines, such as by role or function, or by process thread in the case of a cloud computing distributed model. In various implementations, aspects of the present technology may be configured to run within a virtual machine, which in turn runs on one or more physical machines. It will be understood by those skilled in the art that features of the present technology may be realized by a variety of different suitable machine implementations.

[0025] The system includes various engines, each constructed, programmed, configured, or otherwise adapted to perform a function or set of functions. The term "engine" as used herein means a tangible device, component, or arrangement of components implemented using hardware, such as by an application specific integrated circuit (ASIC) or field programmable gate array (FPGA), or as a combination of hardware and software, such as by a processor-based computing platform and a set of program instructions that transform the computing platform into a special purpose device and implement specific functionality. An engine may also be implemented as a combination of the two, with some functions facilitated solely by hardware and other functions facilitated by a combination of hardware and software.

[0026] In some embodiments, the software may reside on a tangible, machine-readable storage medium in executable or non-executable form. Software that resides in non-executable form may be compiled, translated, or otherwise converted into an executable form prior to or during run-time. In some embodiments, the software, when executed by the engine's underlying hardware, causes the hardware to perform specified operations. Thus, an engine is physically constructed or specifically configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in a specified manner or to perform some or all of any operations described herein in association with that engine.

[0027] Consider an example in which the engines are temporarily configured, where each of the engines may be instantiated at different moments in time. For example, if the engines comprise a general purpose hardware processor core that is configured using software, the general purpose hardware processor core may be configured as separate and distinct engines at different times. The software may thus configure the hardware processor core, e.g., configure a particular engine at one instance of time and configure a different engine at a different instance of time.

[0028] In some implementations, at least a portion, and in some cases all, of the engines may run on one or more computer processors that execute the operating system, system programs, and application programs, while also implementing the engines using multitasking, multithreading, distributed (e.g., cluster, peer-to-peer, cloud, etc.) processing or other such techniques, if necessary. Thus, each engine may be realized in a variety of suitable configurations and generally should not be limited to any particular implementation illustrated herein, unless such a limitation is expressly stated.

[0029] In addition, the engine itself may consist of one or more sub-engines, each of which may be considered an engine in itself. Furthermore, in the implementations described herein, the various engines each correspond to a defined functionality. However, it should be understood that in other contemplated implementations, each functionality may be distributed among more than one engine. Similarly, in other contemplated implementations, multiple defined functionalities may be implemented by a single engine that performs those multiple functions, possibly alongside other functions, or may be distributed among a set of engines in a manner different from that specifically illustrated in the examples herein.

[0030] As used herein, the term "model" encompasses its plain and ordinary meaning. A model may include, among other things, one or more engines that receive inputs and calculate outputs based on the inputs. The output may be a classification. For example, an image file may be classified as depicting a cat or not depicting a cat. Alternatively, an image file may be assigned a numerical score indicating the likelihood that the image file depicts a cat, and image files with scores above a threshold (e.g., 0.9 or 0.95) may be determined to depict a cat.

[0031] This document may refer to a specific number of things (e.g., "six mobile devices"). Unless expressly stated otherwise, the numbers provided are merely exemplary and may be substituted with any positive integer, whole number, or real number as would make sense for a given context. For example, "six mobile devices" may include any positive integer mobile devices in alternative implementations. Unless stated otherwise, an object referenced in the singular (e.g., "a computer" or "the computer") may include one or more objects (e.g., "a computer" may refer to one or more computers).

[0032] 1 illustrates an example system 100 in which cohort selection may be implemented. As shown, system 100 includes a data repository 110, a server 120, and a client computing device 140, which are connected to one another via a network 130. Network 130 may include one or more of a wired network, a wireless network, a local area network, a wide area network, a virtual private network, the Internet, an intranet, a Wi-Fi network, a cellular network, and the like. Data repository 110, server 120, and client computing device 140 may each include all or a portion of the components of a computing machine 1700 shown in FIG.

[0033] The data repository 110 may be a database or other data storage unit. The data repository 110 may store health insurance claims data, medical data, pharmacy data, genetic data, and the like. The data repository 110 may include a single data repository or multiple data repositories. The data repository 110 may store any of the data described herein, such as the data shown in FIG. 2, FIG. 3, and FIG. 10. The data repository 110 may include data stored in the additional data repository 1110, the health insurance claims data repository 1106, the reference information data repository 1112, the molecular data repository 1108, and the integrated data repository 1104 shown in FIG. 11.

[0034] The client computing devices 140 may include one or more of a laptop computer, a desktop computer, a mobile phone, a tablet computer, a smart watch, a smart TV including processing circuitry and memory, and the like. The server 120 may include one or more servers arranged in a server farm, for example. The server 120 may perform one or more of the processes described herein, for example, as shown in Figures 4-9.

[0035] As shown, data repositories 110 and servers 120 are all connected to network 130 and communicate with each other via network 130. In alternative implementations, one or more of data repositories 110 may be connected directly to server 120 (e.g., using a direct wired or wireless connection) without going through network 130. Data repositories that are directly connected to server 120 may or may not be connected to network 130.

[0036] Precision medicine is playing an increasingly prominent role in the treatment of certain biological conditions. Diverse patient subgroups with rare oncogenic driver mutations that are treatable with standard of care targeted therapies are now being identified.

[0037] To understand a patient's disease progression throughout treatment, it may be useful to consider the primary diagnosis, secondary / metastasis, molecular results, course of treatment, and procedures performed on the patient.

[0038] Currently, this information exists in patient claims records as codes in medical headers, medical summaries, and pharmacy records. The data is scattered across many columns, making it difficult to query and derive the types of information that could serve as real-world evidence. For example, codes must go through multiple translation steps before they can be meaningfully used in queries.

[0039] FIG. 2 illustrates an example of a process 200 of claims data to extract information. The process 200 begins at operation 202 with identifying National Drug Code (NDC) codes included within the claims data. The NDC codes may indicate drugs used to treat a biological condition. The NDC codes may have one or more defined formats, and the claims data may be analyzed against the one or more defined formats to identify the NDC codes within the claims data. Additionally, the NDC codes may be located within one or more defined columns of the claims data. In one or more examples, the one or more defined columns may be analyzed and rows within which values ​​exist may be identified. The NDC codes may then be extracted from the claims data. At operation 204, the NDC codes may be used to obtain drug name information, and at operation 206, the NDC codes may be used to obtain drug class information. The drug name information and drug class information can be stored by a data repository that is accessible using one or more application programming interfaces.

[0040] In operation 208, the claims data may be analyzed to determine start and stop dates for a given drug. In block 210, the claims data may be analyzed to determine drugs that are provided to treat a patient associated with a biological condition, such as cancer. In various embodiments, NDC codes may be analyzed to identify drugs that are provided to a patient associated with a biological condition. In block 212, a start date is determined for a drug that is provided to a patient associated with a given biological condition. In block 214, drug combinations are determined. For example, the claims data may be analyzed to determine multiple drugs that may be provided to a patient to treat a biological condition. The combinations (drugs and claims) are ranked against a primary diagnosis in block 216 and against tests (e.g., Guardant 360 (G360) tests by Guardant Health, Inc., of Redwood City, California) in block 218.

[0041] As the dataset is constructed (e.g., using process 200), some implementations formalize the logic used in translating (i) NDC codes into courses of care, (ii) International Classification of Diseases (ICD) Ninth Revision (ICD-9) and ICD-10 codes into primary diagnoses and metastases, and (iii) Healthcare Common Procedure Coding System (HCPCS), ICD-10, Personal Care Services (PCS), and Current Procedural Terminology (CPT) codes into procedures. The constructed dataset may include a subset of patients in a biological / medical data repository (e.g., database) who have a given biological condition (e.g., lung cancer) as a primary diagnosis in their insurance claims records. With respect to this subset, some implementations sort patient information, including diagnoses, treatments, procedures, and genomic tests.

[0042] 3 illustrates an example of patient information 300 that may be stored (e.g., in a data repository) according to some implementations. As shown, the patient information 300 includes a diagnosis 302, a treatment 304, a procedure 306, and a genomic test 308.

[0043] Some implementations attempt to build a framework for derived fields or derived content and evolve it over time. Healthcare data can be messy and incomplete. It can include diagnosis codes, treatment information, and dates. Nevertheless, the ability to derive valuable contextual information and present a patient's cancer history with a high degree of confidence is promising for real-world evidence. Some implementations describe what information to extract from insurance claims data and how to transform them into meaningful higher-level concepts that can be used to generate real-world evidence, measurements, and outcomes about patients.

[0044] A dataset for a given biological condition, e.g., lung cancer, is the subset of patients in the overall dataset who have lung cancer as a primary diagnosis. Some implementations consider medical headers. A patient may have multiple medical header records, with each row representing one claim. Each claim may have multiple diagnosis codes. Diagnoses can be ICD-9 or ICD-10 codes. A single claim may have either an ICD-9 or ICD-10 code, i.e., a mixture of both may be present in the same claim. In some implementations, the column icd_type or another code may be used to identify whether it is ICD-9 or ICD-10. The column claim_date will be used for the 6 month blackout period logic.

[0045] 4 is a flowchart of an exemplary method 400 for identifying a primary lung cancer patient (or a patient with another primary diagnosis) according to some implementations. Method 400 may be adjusted for other types of diagnoses that should not be combined with a set of other diagnoses. For example, method 400 may be used to identify a patient diagnosed with pneumonia who has not previously been diagnosed with influenza.

[0046] In block 402, a computing machine (e.g., computing machine 1700) sorts records associated with insurance procedures. In some implementations, the computing machine sorts the patient's medical headers in ascending order using a column for procedure date or billing date. One purpose of the sorting is to identify the first occurrence, if any, of a lung cancer diagnosis.

[0047] At block 404, the computing machine determines whether a header indicative of lung cancer (e.g., ICD code C34% or C33%) is present in the sorted insurance procedures for the given patient. In some implementations, values ​​in certain diagnosis code columns are analyzed to determine whether a C34 or C33 code or a 162 code is present in these columns. A header indicative of lung cancer may be identified based on an ICD-9 or ICD-10 code, for example, ICD code C34 "Malignant Neoplasm of the Bronchi and Lungs", ICD code C33 "Malignant Neoplasm of the Trachea", or ICD code 162 "Malignant Neoplasm of the Trachea, Bronchi, and Lungs". If no such code is present, at block 410, the computing machine determines that the given patient is not a lung cancer patient. If such a header is present, the method 400 continues to block 406.

[0048] In block 406, the computing machine determines whether a header indicating another cancer different from lung cancer (e.g., breast cancer, prostate cancer, skin cancer, and the like) occurred within six months from the procedure date associated with the patient's (temporally) earliest header indicating lung cancer. For example, the computing machine may look for any cancer-related ICD-10 or ICD-9 codes in the diagnosis column that occurred in the six months preceding the first lung cancer-related code identified in block 404. An ICD-10 cancer code may start with C00 and end with C76. An ICD-9 cancer code may have the number 140 and end with 195, i.e., 140-195. If such a header is identified, in block 412, the computing machine determines that this is not a primary lung cancer patient (e.g., another cancer that has metastasized to the patient's lungs). If such a code is not identified, in block 408, the computing machine determines that this is a primary lung cancer patient. After block 408, block 410, or block 412, the method 400 ends.

[0049] Some implementations are concerned with rare case handling. When a C34% code and another cancer code are present in the same first claim, some implementations may include the patient as a primary lung cancer patient. When a C34% code claim and another claim with the same claim date for a patient with a non-cancer code are present, some implementations may include the patient as a primary lung cancer patient. When a 34% code claim is present with another cancer claim within 6 months, then a claim for lung cancer 9 months later, some implementations may not include the patient as a primary lung cancer patient because some implementations look for the first occurrence of a lung cancer diagnosis and apply a 6 month washout period before that date.

[0050] 5 is a flowchart of an exemplary method 500 for identifying primary lung cancer patients and rare cases according to some implementations. Rare cases can be considered as either primary lung cancer patients or non-primary lung cancer patients according to the needs of the user.

[0051] In block 502, a computing machine (eg, computing machine 1700) sorts insurance procedures for a given patient.

[0052] At block 504, the computing machine determines whether a C34%, C33%, or 162 ICD code is present. If not, then at block 506, the computing machine determines that the given patient is not a lung cancer patient. After block 506, the method 500 ends. If so, the method 500 continues to block 508.

[0053] In block 508, the computing machine determines whether there is another C% on the same procedure or date. If so, the given patient is marked as a rare case in block 510. After block 510, method 500 continues at block 512. If not, method 500 continues at block 512.

[0054] In block 512, the computing machine determines whether there is a C% or 140-195 ICD code within 6 months of the first procedure with the C34%, C33%, or 162 ICD code identified in block 504. If not, then in block 514 the given patient is labeled as a primary lung cancer patient. After block 514, the method 500 ends. If so, then in block 516 the patient is not labeled as a primary lung cancer patient.

[0055] In block 518, the computing machine determines whether there is another C34% or C33% after six months or more. If not, the patient remains marked as not being a primary lung cancer patient according to block 516, and the method 500 ends. If so, the patient is marked as a rare case in block 520, while also remaining marked as not being a primary lung cancer patient according to block 516, and the method 500 ends. The rare case may be marked as either a lung cancer patient or a non-lung cancer patient, depending on whether false positives (not a lung cancer patient, incorrectly identifying a person as being such) or false negatives (lung cancer patient, incorrectly identifying a person as being such) are preferred.

[0056] With respect to a patient's mortality status, if a date of death exists for the patient, this value is set to "relevant". Otherwise, it is set to unknown. (In some cases, the mortality status for a patient may never be set to "not relevant". Alternatively, the mortality status may be set to "not relevant" if there is confirmation that the patient is alive within a threshold period (e.g., 1 day or 7 days) prior to the current date.) In both cases, some implementations set the quality metric associated with the mortality status to "high". In one or more embodiments, the source of the mortality data may include insurance claims information.

[0057] For patients who do not have a deceased status, the last effective date column stores the most recent date from the patient header and pharmacy claim in which the claim was generated. Some implementations include pharmacy claims, which are designated as paid, pending, or adjusted. Using the patient header data table, the receipt date column can be queried, if present, to determine the last effective date. In situations where a value does not exist in the receipt date column, the billing date column can be queried. Using the pharmacy billing data table, the usage date column can be queried to determine the last effective date.

[0058] The column containing the value for age represents the patient's age calculated from the year of birth. If the date of death is available, the value for age at death is the patient's age at the time of death.

[0059] The value for the patient's metastatic status is either "applicable", "not applicable", or "unknown", depending on the patient's metastatic state and whether it is known. In the first case, if the billing record has a secondary malignancy code reported, the patient is considered metastatic with high confidence. Secondary malignancies are identified by ICD-10 codes C77-C80 or ICD-9 196-198% recognized on or after the same date / same billing as the primary diagnosis. When some implementations set the patient metastatic status to "applicable" using the above logic, some implementations may set the metastatic quality measure as "high". In the second case,

[0060] If the claim record has any cancer ICD-10 / ICD-9 code (except skin and lung) that is different from the patient's primary diagnosis code within the past two years, some implementations may set the patient metastasis status as "applicable" and the metastasis quality measure as "low". Some ICD codes may be excluded. Lung cancer ICD-9 codes to be excluded include 162 malignant neoplasm of tracheobronchial and lung. Lung cancer ICD-10 codes to be excluded include C34 malignant neoplasm of bronchi and lung or C33 malignant neoplasm of trachea. Skin cancer ICD-9 codes to be excluded include 172 malignant melanoma of skin or 173 other and non-defined neoplasm of skin. Skin cancer ICD-10 codes to be excluded include C43 malignant melanoma of skin or C44 other and non-defined neoplasm of skin.

[0061] Some implementations may set the patient metastasis status to unknown and also set the metastasis quality measurements to "\n" since they are unknown.

[0062] The column,indicating whether the patient is enrolled in a clinical trial,is set to "true" if a Z006 ICD-10 or V707 ICD-9 code is present, otherwise,it is set to "false." The column,indicating the latest billing date associated with a clinical trial,corresponds to the latest procedure date when a Z006 ICD-10 or V707 ICD-9 code is present.

[0063] Data completeness may be based on the percentage of patients without a date of death and the percentage of patients with high confidence metastasis information. Data accuracy may be based on the count of patients suspected to be deceased and without a date of death. Demographic data quality may be based on the percentage of patients with high quality demographic data.

[0064] FIG. 6 is a flowchart of an example method 600 for identifying a last valid date to be included in a patient information data table, according to some implementations.

[0065] At block 602, a computing machine (eg, computing machine 1700) obtains a data set of medical header and pharmacy data.

[0066] At block 604, the computing machine filters out dispensed values ​​that are not paid, reserved, or adjusted. Dispensed values ​​that are paid, reserved, or adjusted remain in the data set.

[0067] At block 606, the computing machine determines whether a receipt date is in the header. If so, the method 600 continues at block 608. If not, the method 600 continues at block 612.

[0068] At block 608, the computing machine determines whether the header billing date is greater than the usage date in the prescription table. If so, the method 600 continues at block 610. If not, the method 600 continues at block 616.

[0069] In block 610, the computing machine determines that the billing date is the last valid date. After block 610, the method 600 ends.

[0070] At block 612, the computing machine determines whether the receipt date is after the usage date in the prescription table. If so, the method 600 continues at block 614. If not, the method 600 continues at block 616.

[0071] In block 614, the computing machine determines that the receipt date is the last valid date. After block 614, the method 600 ends.

[0072] In block 616, the computing machine determines that the usage date is the last valid date. After block 616, the method 600 ends.

[0073] In some implementations, a data table can be generated that includes patient information, and the data stored by the data table can be used to determine one or more cohorts for patients. One or more of the following patient information columns are included in the data table: year of birth, sex, date of death, death status, source of death data, last valid date, patient metastasis status, clinical research study enrollment status, date of most recent use, remission status, age, and / or age at death.

[0074] Note that some implementations determine a primary diagnosis based on treatment (e.g., if a patient takes a drug used for lung cancer but not for other conditions, the patient is likely to have lung cancer). Some implementations may include a single table that establishes the relationship between disease progression and treatment. The table may be generated based on common query patterns.

[0075] Some implementations relate to treatment data columns: A computing machine may create a lung cancer pharmacy table containing cancer-related treatments from claims pharmacy records for lung cancer patients.

[0076] One of the primary sources of information is the dispensing record. Oral drugs are typically present in the dispensing record. A few intravenous (IV) drugs may be present in the dispensing record. Some implementations include treatments from the dispensing record whose NDC codes have the class antineoplastic drug and the paid procedure type, or select the cancer drug names listed in Table 2. The list in Table 2 is not comprehensive. It may be missing some known previously existing cancer drugs and / or cancer drugs that will be developed or identified in the future. Table 3 illustrates an exemplary query. One purpose of the query in Table 3 is to extract all cancer-related drugs from the NDC lookup table (or some other data repository). The query may be based on drug name, drug class, and equivalents. The query in Table 3 may be based on the exemplary cancer drugs shown in Table 2. [Table 1-1] [Table 1-2]

[0077] Examples of dispensing data columns in the dispensing data table include patient identifier (ID), prescription fill date, prescription quantity, number of authorized refills, dispensing use date, NDC code, dispensing start date, dispensing end date, drug name, drug class, drug category, number of refills, days supply, quantity dispensed, and unit of measure. Some implementations remove duplicate dispensing claims. Duplicate dispensing transactions may be defined as transactions with the same drug name and the same days supply on the same use date.

[0078] Some implementations aggregate days of supply for dispensing claims for the same day. Some implementations aggregate all days of supply for a drug that appears multiple times on the same day with different days of supply. Table 4 illustrates an example of dispensing data that may be stored. Cancer-related treatments may be extracted from the set of items (e.g., for storage in a data structure such as Table 4). [Table 2-1] [Table 2-2]

[0079] 7 is a flow chart of an example process 700 associated with assigning patients to cohorts. In some implementations, one or more process blocks of FIG. 7 may be performed by a computing machine (e.g., computing machine 1700). In some implementations, one or more process blocks of FIG. 7 may be performed by another device or devices separate from or including the computing machine. Additionally or alternatively, one or more process blocks of FIG. 7 may be performed by one or more components of computing machine 1700 shown in FIG.

[0080] 7, process 700 may include accessing, by a computing machine, in a processing network, one or more medical data repositories that store data regarding a given patient from among a plurality of patients. The one or more medical data repositories store pharmacy data, clinic visit data, and medical insurance transaction data (block 710).

[0081] As further shown in FIG. 7, process 700 may include identifying, by a computing machine, one or more biological conditions and metastatic status for a given patient based on one or more disease codes in the clinic visit data or health insurance transaction data (block 720).

[0082] As further shown in FIG. 7, process 700 may include identifying, by the computing machine, one or more courses of treatment for a given patient based on one or more drug codes in the pharmacy data (block 730).

[0083] As further shown in FIG. 7, process 700 may include identifying, by the computing machine, one or more medical procedures received by a given patient based on one or more insurance codes in the medical insurance procedure data (block 740).

[0084] As further shown in FIG. 7, process 700 may include determining, by a computing machine, a primary diagnostic biological condition for a given patient based on a combination of one or more biological conditions, metastatic status, one or more courses of treatment, and one or more medical procedures (block 750).

[0085] As further shown in FIG. 7, process 700 may include assigning, by the computing machine, a given patient to a cohort of patients based on the primary diagnosed biological condition (block 760).

[0086] As further shown in FIG. 7, process 700 may include providing, by the computing machine, an output representative of the assigned cohort for a given patient (block 770).

[0087] Process 700 may include additional implementations, such as any single implementation or any combination of implementations described below and / or in conjunction with one or more other processes described anywhere herein.

[0088] In some implementations, determining the primary diagnosed biological condition includes creating a master gap table based on medical insurance procedure data, the master gap table indicating the time interval between two consecutive procedures for a given patient having the same procedure name and insurance code. The master gap table includes columns for procedure name, insurance code, unit, and gap length. Additionally, based on the master gap table, a median gap table can be generated indicating the median gap for each combination of procedure name and insurance code. The median gap table includes columns for procedure name, insurance code, unit, and gap length. The determination of the primary diagnosed biological condition can be based at least in part on the data in the median gap table. In some cases, the treatment that the patient receives can be an indicator of a biological condition present in the patient. For example, a patient who takes a known lung cancer drug and receives a known lung cancer therapy at a medical facility is likely to have lung cancer.

[0089] In some implementations, the processing circuitry comprises a plurality of multi-threaded graphic processing units (GPUs), and the method further includes determining, in parallel and using parallel threads of the plurality of multi-threaded GPUs, an assigned cohort for a plurality of patients, including the given patient, from the plurality of patients.

[0090] In some implementations, the disease codes comprise International Classification of Diseases (ICD) codes, the drug codes comprise National Drug Codes (NDC) codes, and the insurance codes comprise Healthcare Practice Coding System (HCPCS) codes.

[0091] In some implementations, process 700 includes identifying one or more biological conditions for a given patient based on one or more disease codes in clinic visit data or health insurance transaction data, including identifying that a given patient has lung cancer based on an ICD code that is associated with lung cancer.

[0092] In some implementations, process 700 includes identifying a metastatic status for a given patient based on one or more disease codes in clinic encounter data or medical insurance transaction data, which is based on a secondary malignancy ICD code or HCPCS code.

[0093] In some implementations, the primary diagnostic biological condition for a given patient is determined based on disease, drug, or insurance codes associated with dates within a predefined date range.

[0094] 7 shows example blocks of process 700, in some implementations process 700 may include additional, fewer, different, or differently arranged blocks than those depicted in FIG 7. Additionally or alternatively, two or more of the blocks of process 700 may be performed in parallel.

[0095] FIG. 8 is a flow chart of an example process 800 associated with assigning a given patient to a cohort. In some implementations, one or more block processes of FIG. 8 may be performed by a computing machine (e.g., computing machine 1700). In some implementations, one or more process blocks of FIG. 8 may be performed by another device or devices separate from or including the computing machine. Additionally or alternatively, one or more process blocks of FIG. 8 may be performed by one or more components of computing machine 1700 shown in FIG.

[0096] 8, process 800 may include analyzing, by a computing machine, one or more data tables containing medical transaction data for a given patient. The one or more data tables may be obtained from one or more medical data repositories (block 810).

[0097] As further shown in FIG. 8, process 800 may include determining, by the computing machine, one or more first insurance procedures including one or more first code identifiers contained within the one or more data tables, the one or more first code identifiers corresponding to the patient's diagnosis of one or more biological conditions (block 820).

[0098] As further shown in FIG. 8, process 800 may include determining, by the computing machine, one or more second insurance procedures including one or more second code identifiers contained within the one or more data tables, where the one or more second code identifiers correspond to medical procedures performed on the patient, the medical procedures being from the clinic visit data (block 830).

[0099] As further shown in FIG. 8, process 800 may include generating, by the computing machine, a medical header table including a first number of columns storing one or more first code identifiers, a second number of columns storing one or more second code identifiers, and a plurality of rows with each of the plurality of rows corresponding to a first medical insurance procedure of the one or more first insurance procedures or a second medical insurance procedure of the one or more second insurance procedures (block 840).

[0100] As further shown in FIG. 8, process 800 may include storing, by the computing machine, the medical header table in one or more medical data repositories (block 850).

[0101] As further shown in FIG. 8, process 800 may include determining, by the computing machine, a cohort for the patient based on the data in the medical header table (block 860).

[0102] Process 800 may include additional implementations, such as any single implementation or any combination of implementations described below and / or in conjunction with one or more other processes described anywhere herein.

[0103] In some implementations, the one or more first insurance procedures indicate a date of use of a respective first insurance procedure, and the one or more second insurance procedures indicate a date of use of a respective second insurance procedure.

[0104] In some implementations, the process 800 includes arranging the rows of the medical header table in ascending order based on the utilization date, such that the claim with the first utilization date is the first row in the medical header table and the procedure with the most recent utilization date is the last row in the medical header table.

[0105] In some implementations, process 800 may include analyzing the one or more first code identifiers and determining that the first code identifiers of the one or more first code identifiers are included within a group of insurance code identifiers that correspond to one or more biological conditions.

[0106] In some implementations, the first code identifier is arranged according to a first format that corresponds to a first classification of insurance code identifiers, and the first classification of insurance code identifiers corresponds to the International Classification of Diseases, Ninth Revision (ICD-9).

[0107] In some implementations, the first code identifiers are arranged according to a second format that corresponds to a second classification of insurance code identifiers, and the second classification of insurance code identifiers corresponds to the International Classification of Diseases, Tenth Revision (ICD-10).

[0108] In some implementations, the one or more biological conditions include a plurality of subtypes, each subtype of the plurality of subtypes corresponding to a subset of a group of insurance code identifiers corresponding to the one or more biological conditions, and the method further includes determining that the first code identifier is included within a first subset of the group of insurance codes corresponding to the first subtype of the biological condition.

[0109] In some implementations, the biological condition is cancer and the plurality of subtypes includes at least one of lung cancer, breast cancer, or colorectal cancer.

[0110] In some implementations, process 800 includes determining one or more third insurance procedures having use dates within a predefined period ending with a first use date, and analyzing one or more third insurance code identifiers of the third insurance procedures against the first code identifier and the second code identifier.

[0111] In some implementations, process 800 includes determining that one or more third insurance code identifiers are not included within the first code identifier and the second code identifier, and determining, based on the third insurance code identifier, that the patient is included within a cohort of patients in which one or more given subtypes of the biological condition are present.

[0112] In some implementations, the one or more third insurance code identifiers correspond to an additional biological condition.

[0113] In some implementations, process 800 includes determining that the one or more third insurance code identifiers are included in a portion of a group of insurance code identifiers that are not included in a subset of the group of insurance code identifiers, determining that a utilization date of at least one of the one or more third insurance procedures is the same as a utilization date of one of the one or more first insurance procedures, determining that there are no other additional insurance procedures having an insurance code identifier included in the group of code identifiers, and determining that the patient is included in a cohort of patients in which a subtype of the biological condition is present.

[0114] In some implementations, process 800 includes determining that the one or more third insurance code identifiers are included within a portion of a group of insurance code identifiers that are not included within a subset of the group of insurance code identifiers, determining that a utilization date of at least one of the one or more third insurance claims precedes a utilization date of one of the one or more first insurance procedures and is within a predefined time period, and determining that the patient is not included within a cohort of patients in which a subtype of the biological condition is present.

[0115] 8 shows example blocks of process 800, in some implementations process 800 may include additional, fewer, different, or differently arranged blocks than those depicted in FIG 8. Additionally or alternatively, two or more of the blocks of process 800 may be performed in parallel.

[0116] 9 is a flow chart of an example process 900 associated with identifying a cohort of patients. In some implementations, one or more process blocks of FIG. 9 may be performed by a computing machine (e.g., computing machine 1700). In some implementations, one or more process blocks of FIG. 9 may be performed by another device or devices separate from or including the computing machine. Additionally or alternatively, one or more process blocks of FIG. 9 may be performed by one or more components of computing machine 1700 shown in FIG.

[0117] 9, process 900 may include accessing, by a computing machine, in a processing network, one or more medical data tables that store medical insurance transaction data for a plurality of patients. The one or more medical data tables include a date column and a diagnosis column (block 910).

[0118] As further shown in FIG. 9, process 900 may include identifying, by the computing machine, using a processing circuitry and based on the diagnostic sequence, a set of patients having the defined biological condition, the set of patients being from among the plurality of patients (block 920).

[0119] As further shown in FIG. 9, process 900 may include determining, by the computing machine, for each patient in the set of patients, the first date that the patient was diagnosed with the defined biological condition (block 930).

[0120] As further shown in FIG. 9, process 900 may include identifying, by the computing machine, using processing circuitry and based on the diagnosis string and the date string, a cohort of patients from among the set of patients, the cohort of patients lacking a diagnosis from the set of biological conditions associated with a date occurring during a predefined time window prior to the first date that the patient was diagnosed with the specified biological condition (block 940).

[0121] As further shown in FIG. 9, process 900 may include providing, by the computing machine, an output representative of the cohort (block 950).

[0122] Process 900 may include additional implementations, such as any single implementation or any combination of implementations described below and / or in conjunction with one or more other processes described anywhere herein.

[0123] In some implementations, the diagnosis column stores International Classification of Diseases, Ninth Revision (ICD-9) or International Classification of Diseases, Tenth Revision (ICD-10) codes.

[0124] In some implementations, the defined biological condition is lung cancer, the set of biological conditions comprises a cancer different from lung cancer, and the predefined time window before the initial date is six months before the initial date.

[0125] In some implementations, the defined biological condition is a defined type of cancer, and the method further includes determining the metastatic status of at least one patient from the cohort.

[0126] In some implementations, metastatic status is determined based on secondary malignancies International Classification of Diseases (ICD) codes or Healthcare Common Criteria Coding System (HCPCS) codes.

[0127] In some implementations, identifying the cohort includes ordering rows associated with patients in the set by date, accessing rows associated with a predefined time window, and identifying patients in the set that lack a diagnosis from the set of biological conditions during the predefined time window.

[0128] 9 shows example blocks of process 900, in some implementations process 900 may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those depicted in FIG 9. Additionally or alternatively, two or more of the blocks of process 900 may be performed in parallel.

[0129] FIG. 10 illustrates an example medical data table 1000 according to some implementations. Table 1000 includes data (name, date, and diagnosis) for patients who are candidates for inclusion in a lung cancer primary diagnosis cohort: Albert, Betsy, Carlos, Debra, and Edward. Some implementations may store patient ID numbers instead of names to ensure patient privacy. Note that table 1000 is simplified to illustrate how some implementations operate. Other tables used with the techniques disclosed herein may include more rows, columns, patients, and data.

[0130] Albert's only diagnosis was lung cancer and therefore he is included within the lung cancer primary diagnosis cohort.

[0131] Betsy was diagnosed with influenza on December 17, 2015 and lung cancer on April 5, 2016. Although Betsy was diagnosed with influenza prior to being diagnosed with lung cancer, Betsy is still included within the lung cancer primary diagnosis cohort because influenza is not a type of cancer.

[0132] Carlos was diagnosed with liver cancer on June 8, 2017 and lung cancer on November 2, 2017. Carlos' first lung cancer diagnosis is liver cancer, not lung cancer, because he had another cancer diagnosis (liver cancer) within 6 months prior to his lung cancer diagnosis. Because Carlos' primary cancer is not lung cancer, Carlos is not included within the lung cancer primary diagnosis cohort.

[0133] Debra was diagnosed with lung cancer on July 1, 2017 and liver cancer on August 12, 2017. Because lung cancer was the first cancer Debra was diagnosed with, Debra is included in the lung cancer primary diagnosis cohort.

[0134] Edward was diagnosed with influenza on January 2, 2018, but was never diagnosed with lung cancer or any other cancer. Therefore, Edward is not included in the lung cancer primary diagnosis cohort. Based on Table 1000, the lung cancer primary diagnosis cohort includes Albert, Betsy, and Debra. The lung cancer primary diagnosis cohort does not include Carlos and Edward.

[0135] FIG. 11 illustrates an example architecture 1100 for generating a unified data repository including multiple types of healthcare data, according to one or more implementations. The architecture 1100 may include a data integration and analysis system 1102. The data integration and analysis system 1102 may obtain data from a number of data sources and integrate the data from the data sources into a unified data repository 1104. For example, the data integration and analysis system 1102 may obtain data from a health insurance claims data repository 1106. In various embodiments, the data integration and analysis system 1102 and the health insurance claims data repository 1106 may be created and maintained by different entities. In one or more additional embodiments, the data integration and analysis system 1102 and the health insurance claims data repository 1106 may be created and maintained by the same entity.

[0136] The data integration and analysis system 1102 may be implemented by one or more computing devices. The one or more computing devices may include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or a combination thereof. In some implementations, at least a portion of the one or more computing devices may be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices may be implemented in a cloud computing architecture. In a scenario where a computing system used to implement the data integration and analysis system 1102 is configured in a distributed computing architecture, processing operations may be performed in parallel by multiple virtual machines. In various embodiments, the data integration and analysis system 1102 may implement multi-threading techniques. The implementation of a distributed computing architecture and multi-threading techniques causes the data integration and analysis system 1102 to utilize fewer computing resources relative to a computing architecture that does not implement these techniques.

[0137] The health insurance claims data repository 1106 may store information obtained from one or more health insurance companies corresponding to claims made by subscribers of the one or more health insurance companies. The health insurance claims data repository 1106 may be arranged (e.g., sorted) by patient identifier. The patient identifier may be based on the patient's first name, last name, date of birth, social security number, address, employer, and the like. The data stored by the health insurance claims data repository 1106 may include structured data arranged in one or more data tables. The one or more data tables storing the structured data may include a number of rows and a number of columns indicating information about health insurance claims made by subscribers of the one or more health insurance companies related to procedures and / or treatments received by the subscribers from a health care provider. At least some of the rows and columns of the data tables stored by the health insurance claims data repository 1106 may include health insurance codes that may indicate diagnoses of biological conditions and treatments and / or procedures obtained by subscribers of the one or more health insurance companies. In various examples, the health insurance code may also indicate diagnostic procedures obtained by the individual related to one or more biological conditions that may be present in the individual. In one or more examples, the diagnostic procedures may provide information used in detecting the presence of a biological condition. The diagnostic procedures may also provide information used to determine the progression of the biological condition. In one or more illustrative examples, the diagnostic procedures may include one or more imaging procedures, one or more assays, one or more laboratory procedures, one or more combinations thereof, and the like.

[0138] The data integration and analysis system 1102 may also obtain information from a molecular data repository 1108. The molecular data repository 1108 may store a number of individual's data related to genomic, genetic, pathological (e.g., analysis of tissue slides), metabolomic, transcriptomic, fragmentomic, immune receptor, methylation, epigenomic, and / or proteomic information. In one or more embodiments, the data integration and analysis system 1102 and the molecular data repository 1108 may be created and maintained by different entities. In one or more additional embodiments, the data integration and analysis system 1102 and the molecular data repository 1108 may be created and maintained by the same entity.

[0139] The genomic information may indicate one or more mutations corresponding to the individual's genes. The individual's genetic mutations may correspond to differences between the individual's nucleic acid sequence and one or more reference genomes. The reference genome may include a known reference genome, such as hg119. In various embodiments, the individual's genetic mutations may correspond to differences in the individual's germline genes relative to the reference genome. In one or more additional embodiments, the reference genome may include the individual's germline genome. In one or more further embodiments, the individual's genetic mutations may include somatic mutations. The individual's genetic mutations may be associated with insertions, deletions, single base mutations, loss of heterozygosity, duplications, amplifications, translocations, fusion genes, or one or more combinations thereof.

[0140] In one or more illustrative examples, the genomic information stored by the molecular data repository 1108 may include a genomic profile of tumor cells present within the individual. In these circumstances, the genomic information may be derived from an analysis of genetic material, such as, but not limited to, deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA), from a tissue sample or a sample containing a tumor biopsy, circulating tumor cells (CTCs), exosomes or efferosomes, or from circulating nucleic acids (e.g., cell-free DNA) found in the individual's blood sample that are present due to the degradation of tumor cells present within the individual. In one or more examples, the genomic information of the individual's tumor cells may correspond to one or more target regions. One or more mutations present with respect to the one or more target regions may indicate the presence of tumor cells within the individual. The genomic information stored by the molecular data repository 1108 may be generated in association with an assay or other diagnostic test that may determine one or more mutations with respect to one or more target regions of a reference genome.

[0141] "Cell-free DNA", "cfDNA molecule", or simply "cfDNA" includes DNA molecules that occur in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or saliva), including DNA that is not contained within or otherwise bound to a cell at the time of isolation from the subject. The DNA was originally present within a cell or cells of a large complex biological organism (e.g., a mammal), or within other cells, such as bacteria that colonize the organism, but the DNA has been released from the cell into the fluid found within the organism. cfDNA includes, but is not limited to, the cell-free genomic DNA of a subject (e.g., the genomic DNA of a human subject) and the cell-free DNA of microorganisms, such as bacteria that inhabit the subject (whether pathogenic bacteria or bacteria that are normally found in commonly colonized locations such as the intestine or skin of healthy control groups), but does not include the cell-free DNA of microorganisms that simply contaminate a sample of bodily fluids. Typically, cfDNA can be obtained by obtaining a sample of a fluid without the need to perform an in vitro cell lysis step and including removal of cells present in the fluid (e.g., centrifugation of blood to remove cells).

[0142] In one or more additional examples, the data integration and analysis system 1102 may obtain information from one or more additional data repositories 1110. The one or more additional data repositories 1110 may store data related to an individual's electronic medical record for which data is present in at least one of the health insurance claims data repository 1106 or the molecular data repository 1108. Additionally, the one or more additional data repositories 1110 may store data related to an individual's pathology report for which data is present in at least one of the health insurance claims data repository 1106 or the molecular data repository 1108. In various examples, the one or more additional data repositories 1110 may store data related to a biological condition and / or a treatment related to a biological condition. In one or more examples, at least a portion of the data integration and analysis system 1102 and the one or more additional data repositories 1110 may be created and maintained by different entities. In one or more further embodiments, at least a portion of the data integration and analysis system 1102 and the one or more additional data repositories 1110 may be created and maintained by the same entity.

[0143] In one or more further implementations, the data integration and analysis system 1102 may obtain information from one or more reference information data repositories 1112. The one or more reference information data repositories 1112 may store information including definitions, standards, protocols, terminology tables, one or more combinations thereof, and the like. In various examples, the information stored by the one or more reference information data repositories may correspond to biological conditions and / or treatments for biological conditions. In one or more illustrative examples, the one or more reference information data repositories 1112 may include RxNorm. (RxNorm provides normalized names for clinical drugs and links the names to many of the drug terminology tables used in pharmacy management and drug interaction software.) In one or more examples, at least a portion of the data integration and analysis system 1102 and the one or more reference information data repositories 1112 may be created and maintained by different entities. In one or more further embodiments, at least a portion of the data integration and analysis system 1102 and the one or more reference information data repositories 1112 may be created and maintained by the same entity.

[0144] The data integration and analysis system 1102 may obtain data from at least one of the health insurance claims data repository 1106, the molecular data repository 1108, the one or more additional data repositories 1110, or the reference information data repository 1112 via one or more communication networks that are accessible to the data integration and analysis system 1102 and that are accessible to at least one of the health insurance claims data repository 1106, the molecular data repository 1108, the one or more additional data repositories 1110, or the reference information data repository 1112. The data integration and analysis system 1102 may also obtain data from at least one of the health insurance claims data repository 1106, the molecular data repository 1108, the one or more additional data repositories 1110, or the reference information data repository 1112 via one or more secure communication channels. Additionally, the data integration and analysis system 1102 may obtain data from at least one of a health insurance claims data repository 1106, a molecular data repository 1108, one or more additional data repositories 1110, or a reference information data repository 1112 via one or more application programming interface (API) calls.

[0145] The data integration and analysis system 1102 may include a data integration system 1114. The data integration system 1114 may obtain data from the health insurance claims data repository 1106 and the molecular data repository 1108 and generate the integrated data repository 1104. The data integration system 1114 may also obtain data from one or more additional data repositories 1110 and generate the integrated data repository 1104. In various embodiments, the data integration system 1114 may implement one or more natural language processing techniques to integrate data from the one or more additional data repositories 1110 into the integrated data repository 1104.

[0146] In one or more embodiments, the data integration system 1114 may generate one or more tokens to identify an individual having data stored in the health insurance claims data repository 1106 and having data stored in the molecular data repository 1108. In various embodiments, the data integration system 1114 may generate one or more tokens by implementing one or more hash functions. The data integration system 1114 may implement one or more hash functions to generate one or more tokens based on information stored by at least one of the health insurance claims data repository 1106 or the molecular data repository 1108. For example, the information used by the data integration system 1114 to generate the individual tokens by implementing a hash function may include at least one of an individual individual's identifier, the individual individual's date of birth, the individual individual's zip code, the individual individual's birth date, or the individual individual's gender. In one or more illustrative embodiments, the individual individual's identifier may include a combination of at least a portion of the individual individual's first name and at least a portion of the individual individual's last name. Tokens generated using data from different data repositories may correspond to the same or similar information or the same or similar types of information stored by the different data repositories. To illustrate, a token may be generated using a portion of an individual's name, date of birth, at least a portion of a zip code, and gender obtained from the health insurance claims data repository 1106 and the molecular data repository 1108.

[0147] The data integration system 1114 may integrate data from a number of different data sources by analyzing tokens generated by implementing one or more hash functions using data obtained from the number of different data sources. For example, the data integration system 1114 may obtain one or more first tokens generated from data stored by the health insurance claims data repository 1106 and one or more second tokens generated from data stored by the molecular data repository 1108. The data integration system 1114 may analyze the one or more first tokens with respect to the one or more second tokens to determine individual first tokens corresponding to the individual second tokens. In one or more illustrative examples, the data integration system 1114 may identify individual first tokens that match individual second tokens. A first token may match a second token when the data of the first token has at least a threshold amount of similarity with respect to the data of the second token. In one or more embodiments, a first token may match a second token when the data of the first token is identical to the data of the second token. To illustrate, a first token may match a second token when an alphanumeric string of the first token is identical to an alphanumeric string of the second token.

[0148] By determining a first token generated using data stored by the health insurance claims data repository 1106 that corresponds to a second token generated using data stored by the molecular data repository 1108, the data integration system 1114 may identify individuals who have data stored in both the health insurance claims data repository 1106 and the molecular data repository 1108. In this manner, the data integration system 1114 may obtain data from a number of individuals from the health insurance claims data repository 1106 and data from the molecular data repository 1108 from the same number of individuals and store the health insurance claims data and molecular data for that number of individuals in the integrated data repository 1104.

[0149] The data integration system 1114 may also integrate data stored by one or more additional data repositories 1110 with data from the health insurance claims data repository 1106 and the molecular data repository 1108 to generate the integrated data repository 1104. To illustrate, the data integration system 1114 may obtain one or more third tokens generated from data stored by the additional data repository 1110, such as a data repository that stores data corresponding to a pathology report. The data integration system 1114 may analyze the one or more third tokens with respect to the first token generated using information stored by the health insurance claims data repository 1106 and the second token generated using information stored by the molecular data repository 1108 to determine individual third tokens corresponding to each of the first tokens and each of the second tokens. In one or more illustrative embodiments, the data integration system 1114 may identify a third token generated using one or more hash functions and a common set of information obtained from the health insurance claims data repository 1106, the molecular data repository 1108, and the additional data repository 1110.

[0150] By determining a third token generated using data stored by the additional data repository 1110 that corresponds to the first token generated using data stored by the health insurance claims data repository 1106 and the second token generated using data stored by the molecular data repository 1108, the data integration system 1114 may identify individuals who have data stored in the health insurance claims data repository 1106, the molecular data repository 1108, and the additional data repository 1110. In this manner, the data integration system 1114 may obtain data from the health insurance claims data repository 1106 from a number of individuals and data from the molecular data repository 1108 and the additional data repository 1110 from a same number of individuals and store the health insurance claims data, molecular data, and additional data for that number of individuals in the integrated data repository 1104.

[0151] The data stored by the integrated data repository 1104 for the number of individuals may be accessible using the individual's individual identifier. The data integration system 1114 may implement a number of techniques as part of the de-identification process with respect to storing and retrieving the information of the individuals in the integrated data repository 1104. The individual's identifier may correspond to a key generated using at least one hash function. The individual's identifier may also be generated by implementing one or more salting processes with respect to the key generated using at least one hash function, the token generated using one or more hash functions, and a common set of information obtained from the health insurance claims data repository 1106, the molecular data repository 1108, and / or the additional data repository 1110. In one or more illustrative embodiments, the identifiers generated by the data integration system 1114 to access information about the individual individuals stored by the integrated data repository 1104 may be unique for each individual. In one or more embodiments, the individual's identifier may be generated using at least a portion of the information used to generate the token associated with the individual. In one or more additional embodiments, an individual's identifier may be generated using information that is different than the information used to generate a token associated with the individual.

[0152] The data integration system 1114 may also generate the integrated data repository 1104 from a number of different combinations of the data repositories in a similar manner. For example, the data integration system 1114 may obtain tokens generated from information stored by the health insurance claims data repository 1106 and additional tokens generated from information stored by one or more additional data stores 1110. The data integration system 1114 may determine individual tokens generated from information stored by the health insurance claims data repository 1106 that correspond to the individual additional tokens generated from information stored by the one or more additional data repositories 1110. By determining tokens generated using data stored by the health insurance claims data repository 1106 that correspond to additional tokens generated using data stored by the additional data repository 1110, the data integration system 1114 may identify individuals having data stored in both the health insurance claims data repository 1106 and the additional data repository 1110. In this manner, the data integration system 1114 may obtain data from the health insurance claims data repository 1106 from a number of individuals and data from the additional data repository 1110 from the same number of individuals and store the health insurance claims data and additional data for that number of individuals in the integrated data repository 1104. The health insurance claims data and additional data stored by the integrated data repository 1104 for that number of individuals may be accessible using the individuals' individual identifiers.

[0153] In one or more further embodiments, the data integration system 1114 may obtain tokens generated from information stored by the molecular data repository 1108 and tokens generated from information stored by the one or more additional data stores 1110. The data integration system 1114 may determine individual tokens generated from information stored by the molecular data repository 1108 that correspond to individual additional tokens generated from information stored by the one or more additional data repositories 1110. By determining tokens generated using data stored by the molecular data repository 1108 that correspond to additional tokens generated using data stored by the additional data repository 1110, the data integration system 1114 may identify individuals having data stored in both the molecular data repository 1108 and the additional data repository 1110. In this manner, the data integration system 1114 may obtain data from the molecular data repository 1108 from a number of individuals and data from the additional data repository 1110 from a same number of individuals and store molecular data and additional data for that number of individuals in the integrated data repository 1104. The molecular data and additional data stored by the integrated data repository 1104 for that number of individuals may be accessible using the individuals' individual identifiers.

[0154] The data stored by the integrated data repository 1104 may be stored in accordance with one or more regulatory frameworks that protect privacy and ensure security of individuals' medical records, health information, and insurance information. For example, the data may be stored by the integrated data repository 1104 in accordance with one or more government regulatory frameworks that are directed to protecting personal information, such as the Health Insurance Portability and Accountability Act (HIPAA) and / or the General Data Protection Regulation (GDPR). The integrated data repository 1104 also stores the data in an anonymous and de-identified manner to ensure protection of the privacy of individuals whose data is stored by the integrated data repository 1104. To further ensure privacy of individuals whose data is stored by the integrated data repository 1104, the data integration system 1114 may periodically regenerate the integrated data repository 1104. For example, the data integration system 1114 may create the integrated data repository 1104 once per quarter. In one or more additional embodiments, the data integration system 1114 may generate the integrated data repository 1104 monthly, weekly, or once every two weeks. By periodically regenerating the integrated data repository 1104, rather than simply refreshing it when new data is available, the integrated data repository 1104 enhances privacy protections for the data stored by the integrated data repository 1104. That is, in situations where the data repository is simply refreshed with new data, it may be possible to more easily track individuals associated with data newly added to the data repository, since the number of new individuals added at a given time is typically smaller than the existing number of individuals who already have data stored by the data repository.

[0155] In various embodiments, the data stored by the integrated data repository 1104 may be accessed via a database management system. Additionally, the integrated data repository 1104 may store data according to one or more database models. In one or more embodiments, the integrated data repository 1104 may store data according to one or more relational database technologies. For example, the integrated data repository 1104 may store data according to a relational database model. In one or more additional embodiments, the integrated data repository 1104 may store data according to an object-oriented database model. In one or more further embodiments, the integrated data repository 1104 may store data according to an extensible markup language (XML) database model. In yet additional embodiments, the integrated data repository 1104 may store data according to a structured query language (SQL) database model. In still further embodiments, the integrated data repository 1104 may store data according to an image database model.

[0156] The data integration system 1114 may generate the integrated data repository 1104 by generating a number of data tables and creating links between the data tables. The links may indicate logical connections between the data tables. The data integration system 1114 may generate the data tables by extracting a defined set of data from information obtained from the data repositories 1106, 1108, 1110, 1112 and storing the data in rows and columns of the respective data tables. In various embodiments, the logical connections between the data tables may include at least one of a one-to-one link, where a row of information in one data table corresponds to a row of information in another data table, a one-to-many link, where a row of information in one data table corresponds to multiple rows of information in another data table, or a many-to-many link, where multiple rows of information in one data table correspond to multiple rows of information in another data table.

[0157] A number of data tables may be arranged according to the data repository schema 1116. In the illustrative example of FIG. 1, the data repository schema 1114 includes a first data table 1118, a second data table 1120, a third data table 1122, a fourth data table 1124, and a fifth data table 1125. Although the illustrative example of FIG. 1 includes five data tables, in additional implementations, the data repository schema 1116 may include more or fewer data tables. The data repository schema 1116 may also include links between the data tables 1118, 1120, 1122, 1124, 1128. The links between the data tables 1118, 1120, 1122, 1124, 1126 may indicate that information retrieved from one of the data tables 1118, 1120, 1122, 1124, 1126 results in additional information being retrieved that is stored by one or more additional data tables 1118, 1120, 1122, 1124, 1126. Additionally, not all of the data tables 1118, 1120, 1122, 1124, 1126 may be linked to each of the other data tables 1118, 1120, 1120, 1122, 1124, 1126. 1 , the first data table 1118 is logically coupled to the second data table 1118 by a first link 1128, and the first data table 1118 is logically coupled to the fourth data table 1124 by a second link 1130. In addition, the second data table 1120 is logically coupled to the third data table 1122 via a third link 1132, and the fourth data table 1124 is logically coupled to the fifth data table 1126 via a fourth link 1134. Furthermore, the third data table 1122 is logically coupled to the fifth data table 1126 via a fifth link 1136.

[0158] In various examples, data tables are added and / or removed from the data repository schema 1116 such that additional links between data tables may be added or removed from the data repository schema 1116. In one or more illustrative examples, the integrated data repository 1104 may store data tables in accordance with the data repository schema 1116 for at least a portion of individuals for whom the data integration system 1114 obtained information from a combination of at least two of the health insurance claims data repository 1106, the molecular data repository 1108, the one or more additional data repositories 1110, and the one or more reference information data repositories 1112. As a result, the integrated data repository 1104 may store individual instances of the data tables 1118, 1120, 1122, 1124, 1126 in accordance with the data repository schema 1116 for thousands, tens of thousands, up to hundreds of thousands, or more individuals.

[0159] The data integration and analysis system 1102 may also include a data pipeline system 1138. The data pipeline system 1138 may include a number of algorithms, software code, scripts, macros, or other bundles of computer executable instructions that process information stored by the integrated data repository 1104 and generate additional data sets. The additional data sets may include information obtained from one or more of the data tables 1118, 1120, 1122, 1124, 1126. The additional data sets may also include information derived from data obtained from one or more of the data tables 1118, 1120, 1122, 1124, 1126. The components of the data pipeline system 1138 implemented to generate a first additional data set may be different from the components of the data pipeline system 1138 used to generate a second additional data set.

[0160] In one or more embodiments, the data pipeline system 1138 may generate a data set indicative of pharmaceutical treatments received by a number of individuals. In one or more illustrative embodiments, the data pipeline system 1138 may analyze information stored in at least one of the data tables 1118, 1120, 1122, 1124, 1126 to determine health insurance codes corresponding to pharmaceutical treatments received by a number of individuals. The data pipeline system 1138 may analyze health insurance codes corresponding to pharmaceutical treatments with respect to a library of data indicative of defined pharmaceutical treatments corresponding to the one or more health insurance codes to determine names of pharmaceutical treatments received by the individuals. In one or more additional embodiments, the data pipeline system 1138 may analyze information stored by the integrated data repository 1104 to determine medical procedures received by a number of individuals. To illustrate, the data pipeline system 1138 may analyze information stored by one of the data tables 1118, 1120, 1122, 1124, 1126 to determine treatments received by the individual via at least one of injections or intravenous. In one or more further examples, the data pipeline system 1138 may analyze information stored by the integrated data repository 1104 to determine episodes of treatment for the individual, lines of therapy received by the individual, progression of the biological condition, or time to next treatment. In various examples, the data sets generated by the data pipeline system 1138 may be different for different biological conditions. For example, the data pipeline system 1138 may generate a first number of data sets for a first type of cancer, such as lung cancer, and a second number of data sets for a second type of cancer, such as colorectal cancer.

[0161] The data pipeline system 1138 may also determine one or more confidence levels to assign to information associated with an individual having data stored by the integrated data repository 1104. The distinct confidence levels may correspond to different measures of accuracy for information associated with an individual having data stored by the integrated data repository 1104. The information associated with the distinct confidence levels may correspond to one or more characteristics of the individual derived from the data stored by the integrated data repository 1104. The confidence level values ​​for the one or more characteristics may be generated by the data pipeline system 1138 in conjunction with generating one or more data sets from the integrated data repository 1104. In one or more examples, the first confidence level may correspond to a first range of accuracy measures, the second confidence level may correspond to a second range of accuracy measures, and the third confidence level may correspond to a third range of accuracy measures. In one or more additional embodiments, the second range of the accuracy measure may include values ​​that are less than the values ​​of the first range of the accuracy measure, and the third range of the accuracy measure may include values ​​that are less than the values ​​of the second range of the accuracy measure. In one or more illustrative embodiments, the information corresponding to the first confidence level may be referred to as gold standard information, the information corresponding to the second confidence level may be referred to as silver standard information, and the information corresponding to the third confidence level may be referred to as bronze standard information.

[0162] The data pipeline system 1138 may determine a value for the confidence level of the individual's characteristic based on a number of factors. For example, a separate set of information may be used to determine the individual's characteristic. The data pipeline system 1138 may determine the confidence level of the individual's characteristic based on the amount of completeness of the separate set of information used to determine the characteristic for the individual. In a situation where one or more pieces of information are missing from a set of information associated with a first number of individuals, the confidence level for the characteristic may be lower than for a second number of individuals where no information is missing from the set of information. In one or more examples, the amount of missing information may be used by the data pipeline system 1138 to determine the confidence level of the individual's characteristic. To illustrate, a greater amount of missing information used to determine the individual's characteristic may cause a lower confidence level for the characteristic than in a situation where a smaller amount of missing information is used to determine the characteristic. Furthermore, different types of information may correspond to various confidence levels for the characteristic. In one or more examples, the presence of a first piece of information used to determine the individual's characteristic may result in a higher confidence level for the characteristic than the presence of a second piece of information used to determine the characteristic.

[0163] In one or more illustrative examples, the data pipeline system 1138 may determine a number of individuals included in the cohort with a primary diagnosis of lung cancer (or other biological condition). The data pipeline system 1138 may determine a confidence level for the individual with respect to being classified as having a primary diagnosis of lung cancer. The data pipeline system 1138 may use information from a number of columns included in the data tables 1118, 1120, 1122, 1124, 1126 to determine a confidence level for the inclusion of the individual in the lung cancer cohort. The number columns may include health insurance codes associated with the diagnosis of the biological condition and / or the treatment of the biological condition. Additionally, the number columns may correspond to the date of diagnosis and / or treatment for the biological condition. The data pipeline system 1138 may determine that the confidence level of the individual characterized as being part of the lung cancer cohort is higher in scenarios where information is available for the number of columns or at least for every threshold number of columns than in cases where information is available for less than the threshold number of columns. Further, the data pipeline system 1138 may determine a confidence level for an individual to be included in the lung cancer cohort based on the type of information associated with one or more columns and the availability of the information. To illustrate, in a situation where one or more diagnostic codes are present in association with one or more time periods for a group of individuals and one or more treatment codes are absent, the data pipeline system 1138 may determine that the confidence level of including the group of individuals in the lung cancer cohort is greater than in a situation where at least one of the diagnostic codes is absent and a treatment code used to determine whether the individual is included in the lung cancer cohort is present.

[0164] The data integration and analysis system 1102 may include a data analysis system 1140. The data analysis system 1140 may receive integrated data repository requests 1142 from one or more computing devices, such as an exemplary computing device 1144. The one or more integrated data repository requests 1142 may cause data to be read from the integrated data repository 1104. In various embodiments, the one or more integrated data repository requests 1142 may cause data to be read from one or more datasets generated by the data pipeline system 1138. The integrated data repository requests 1142 may specify data to be read from the integrated data repository 1104 and / or the one or more datasets generated by the data pipeline system 1138. In one or more additional embodiments, the integrated data repository requests 1142 may include one or more pre-built queries corresponding to computer-executable instructions to read a specified set of data from the one or more datasets generated by the integrated data repository 1104 and / or the data pipeline system 1138.

[0165] In response to the one or more integrated data repository requests 1142, the data analysis system 1140 may analyze data retrieved from at least one of the integrated data repository 1104 or the one or more data sets generated by the data pipeline system 1138 and generate data analysis results 1146. The data analysis results 1146 may be transmitted to one or more computing devices, such as the exemplary computing device 1148. Although the illustrative example of FIG. 1 shows the one or more integrated data repository requests 1142 and the data analysis results 1146 from one computing device 1144 being transmitted to another computing device 1148, in one or more additional implementations, the data analysis results 1146 may be received by the same computing device that sent the one or more integrated data repository requests 1142. The data analysis results 1146 may be displayed by one or more user interfaces rendered by the computing device 1144 or the computing device 1148.

[0166] In one or more examples, the data analysis system 1140 may implement at least one of one or more machine learning techniques or one or more statistical techniques to analyze the data retrieved in response to the one or more integrated data repository requests 1142. In one or more examples, the data analysis system 1140 may implement one or more artificial neural networks to analyze the data retrieved in response to the one or more integrated data repository requests 1142. To illustrate, the data analysis system 1140 may implement at least one of one or more convolutional neural networks or one or more residual neural networks to analyze the data retrieved from the integrated data repository 1104 in response to the one or more integrated data repository requests 1142. In at least some examples, the data analysis system 1140 may implement one or more random forest techniques, one or more support vector machines, or one or more hidden Markov models to analyze the data retrieved in response to the one or more integrated data repository requests 1142. One or more statistical models may also be implemented on the analyzed data retrieved in response to the one or more integrated data repository requests 1142 to identify at least one of a correlation or significance measure between the characteristics of the individuals. For example, a log-rank test may be applied to the data retrieved in response to the one or more integrated data repository requests 1142. In addition, a Cox proportional hazards model may be implemented on the data retrieved in response to the one or more integrated data repository requests 1142. Furthermore, a Wilcoxon signed rank test may be applied to the data retrieved in response to the one or more integrated data repository requests 1142. In yet other embodiments, z-score analysis may be performed on data retrieved in response to one or more integrated data repository requests 1142 .In still additional embodiments, a Kaplan Meier analysis may be performed on the data retrieved in response to the one or more integrated data repository requests 1142. In at least some embodiments, one or more machine learning techniques may be implemented in combination with one or more statistical techniques to analyze the data retrieved in response to the one or more integrated data repository requests 1142.

[0167] In one or more illustrative examples, the data analysis system 1140 may determine a survival rate of an individual with lung cancer in response to one or more treatments. In one or more additional illustrative examples, the data analysis system 1140 may determine a survival rate of an individual with one or more genomic region mutations with lung cancer in response to one or more treatments. In various examples, the data analysis system 1140 may generate data analysis results 1146 in situations where data retrieved from at least one of the one or more datasets generated by the integrated data repository 1104 or the data pipeline system 1138 meets one or more criteria. For example, the data analysis system 1140 may determine whether at least a portion of the data retrieved in response to one or more integrated data repository requests 1142 meets a threshold confidence level. In situations where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests 1142 is below a threshold confidence level, the data analysis system 1140 may refrain from generating at least a portion of the data analysis results 1146. In scenarios where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests 1142 is at least the threshold confidence level, the data analysis system 1140 may generate at least a portion of the data analysis results 1146. In various embodiments, the threshold confidence level may be related to the type of data analysis results 1146 being generated by the data analysis system 1140.

[0168] In one or more illustrative examples, the data analysis system 1140 may receive the integrated data repository request 1142 and generate a data analysis result 1146 indicating a survival rate of one or more individuals. In these cases, the data analysis system 1140 may determine whether the data stored by the integrated data repository 1104 and / or by one or more datasets generated by the data pipeline system 1138 meets a threshold confidence level, such as a gold standard confidence level. In one or more additional examples, the data analysis system 1140 may receive the integrated data repository request 1142 and generate a data analysis result 1146 indicating a treatment received by one or more individuals. In these implementations, the data analysis system 1140 may determine whether the data stored by the integrated data repository 1104 and / or by one or more datasets generated by the data pipeline system 1138 meets a lower threshold confidence level, such as a bronze standard confidence level.

[0169] In one or more additional illustrative examples, the data analysis system 1140 may receive the integrated data repository request 1142 to determine individuals who have one or more genomic mutations and have received one or more treatments for a biological condition. Continuing with this example, the data analysis system 1140 may determine a survival rate of an individual with one or more genomic mutations in association with one or more treatments that the individual has received. The data analysis system 1140 may then identify the effectiveness of a treatment for the individual in association with a genomic mutation that may be present in the individual based on the survival rate of the individual. In this manner, health outcomes for individuals may be improved by identifying potential treatments that may be more effective than current treatments provided to the individuals for a population of individuals with one or more genomic mutations.

[0170] 12 illustrates an example framework 1200 that corresponds to an arrangement of data tables in a unified data repository according to one or more implementations. In the illustrative example of FIG. 12, the framework 1200 includes a data repository schema 1202 that includes a first data table 204, a second data table 1206, a third data table 1208, a fourth data table 1210, a fifth data table 1212, a sixth data table 1214, and a seventh data table 1216. The illustrative example of FIG. 2 includes seven data tables, but in additional implementations, the data repository schema 1202 may include more or fewer data tables. The data repository schema 1202 may also include links between the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216. The links between the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 may indicate that information retrieved from one of the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 results in additional information being retrieved that is stored by one or more additional data tables 204, 1206, 1208, 1210, 1212, 1214, 1216. Additionally, not all of the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 may be linked to each of the other data tables 204, 1206, 1208, 1210, 1212, 1214, 1216. 2, the first data table 204 is logically coupled to the second data table 1206 by a first link 1218, and the third data table 1208 is logically coupled to the second data table 1206 by a second link 1220. The second data table 1206 is also logically coupled to the fourth data table 1210 by a third link 1222, the second data table 1206 is logically coupled to the fifth data table 1212 by a fourth link 1224, and the second data table 1206 is logically coupled to the sixth data table 1214 by a fifth link 1226.Additionally, the fifth data table 1212 is logically coupled to the sixth data table 1214 by a sixth link 1228, which is logically coupled to the seventh data table 1216 by a seventh link 1230. Additionally, the seventh data table 1216 is logically coupled to the fourth data table 1210 by an eighth link 1232. In various embodiments, data tables are added and / or removed from the data repository schema 1202, such that additional links between data tables may be added or removed from the data repository schema 1202. In one or more illustrative embodiments, the integrated data repository 1104 may store data tables according to the data repository schema 1202 for at least a portion of individuals for whom the data integration system 1114 obtained information from a combination of at least two of the health insurance claims data repository 1106, the molecular data repository 1108, and the one or more additional data repositories 1110. As a result, the integrated data repository 1104 may store individual instances of data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 according to the data repository schema 204 for thousands, tens of thousands, up to hundreds of thousands of individuals, or more.

[0171] In one or more embodiments, the first data table 204 may store data corresponding to genomics and genomics testing for an individual. For example, the first data table 204 may include columns containing information corresponding to the panel used to generate the genomics data, mutations in the genomic region, the type of mutation, copy number of the genomic region, coverage data indicating the number of nucleic acid molecules identified in the sample with one or more mutations, test date, and patient information. The first data table 204 may also include one or more columns containing health insurance data codes that may correspond to one or more diagnostic codes. In addition, the information in the first data table 204 may include at least one identifier for an individual associated with a case in the first data table 204.

[0172] The second data table 1206 may store data related to one or more patient visits by an individual to one or more health care providers. The third data table 1208 may store information corresponding to individual services provided to an individual in connection with one or more patient visits to one or more health care providers represented by the second data table 1206. To illustrate, an individual may visit a health care provider and multiple services may be performed on the individual at the visit. The second data table 1206 may include columns indicating information for each of multiple services performed during the patient visit. Multiple third data tables 1208 may be generated for a patient visit, including columns indicating a finer level of information regarding individual services provided during the patient visit than the information stored by the second data table 1206 associated with the patient visit. For example, the second data table 1206 may include multiple columns indicating health insurance codes for different services provided to an individual during a patient visit, and the third data table 1208 associated with one of the services may include multiple columns for additional health insurance codes corresponding to additional information related to the individual service. A second data table 1206 and a third data table 1208 relating to patient visits may indicate one or more dates of services corresponding to the patient visits.

[0173] The fourth data table 1210 may include columns indicating information about an individual whose information is stored by the integrated data repository 1104. For example, the fourth data table 1210 may include columns indicating information related to at least one of the individual's location, the individual's gender, the individual's date of birth, the individual's date of death (if applicable), or one or more keys associated with the individual. In one or more embodiments, the fourth data table 1210 may include one or more columns related to whether erroneous data has been identified for the individual. In various embodiments, a single fourth data table 1210 may be generated for an individual individual. Thus, the data repository schema 1202 may include multiple instances of the fourth data table 1210, such as thousands, tens of thousands, up to hundreds of thousands, or more.

[0174] The fifth data table 1212 may include columns that indicate information related to a health insurance company or government entity that made a payment for one or more services provided to an individual. For example, the fifth data table 1212 may include one or more payer identifiers. The sixth data table 1214 may include columns that include information corresponding to health insurance coverage information for an individual. In one or more embodiments, the sixth data table 1214 may include columns that indicate the presence of medical coverage for the individual, the presence of pharmacy coverage for the individual, and the type of health insurance plan associated with the individual, such as Health Maintenance Organization (HMO), Preferred Provider Organization (PPO), and the like.

[0175] The seventh data table 1216 may include columns indicating information related to drug treatments obtained by individual individuals. In one or more embodiments, the seventh data table 1216 may include one or more columns indicating health insurance codes corresponding to drug treatments available through dispensing. The health insurance codes may correspond to individual drug treatments. In addition, the health insurance codes may indicate a diagnosis of a biological condition for the individual. The seventh data table 1216 may also include additional information, such as at least one of dosage, days of supply, total amount prescribed, number of refills allowed, utilization date, or information related to the individual receiving the drug treatment.

[0176] In various embodiments, the data repository schema 1202 may provide the results of an analysis of the information stored by the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 in a more efficient manner than a typical data repository schema. For example, the logical connections between the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 may be arranged to efficiently retrieve related data across different data tables 204, 1206, 1208, 1210, 1212, 1214, 1216. In situations where the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 are arranged in a consecutive manner and / or where a greater number of the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 are logically connected, retrieving data from the integrated data repository 1104 from one or more of the data tables 204, 1206, 1208, 1210, 1212, 1214, 1216 to respond to a request for information from the integrated data repository 1104 may be less efficient than in situations where the data repository schema 1202 is implemented.

[0177] 13 illustrates an architecture 1300 for generating one or more data sets from information retrieved from a data repository that integrates health-related data from a number of sources, according to one or more implementations. The architecture 1300 may include a data integration and analysis system 1102 and an integrated data repository 1104. In addition, the data integration and analysis system 1102 may include at least a data pipeline system 1138 and a data analysis system 1140. The data pipeline system 1138 may include a number of sets of data processing instructions that are executable to generate individual data sets that can be analyzed by the data analysis system 1140 in response to an integrated data repository request 1142 to generate data analysis results 1146.

[0178] The data pipeline system 1138 may include a first data processing instruction 1302, a second data processing instruction 1304, and up to Nth data processing instruction 1306. The data processing instructions 1302, 1304, 1306 may be executable by one or more processing units to perform a number of operations to generate individual data sets using information retrieved from the unified data repository 1104. In one or more illustrative examples, the data processing instructions 1302, 1304, 1306 may include at least one of software code, scripts, API calls, macros, and the like. The first data processing instruction 1302 may be executable to generate a first data set 1308. In addition, the second data processing instruction 1304 may be executable to generate a second data set 1310. Furthermore, the Nth data processing instruction 1306 may be executable to generate an Nth data set 1312. In various embodiments, after the data integration and analysis system 1102 generates the integrated data repository 1104, the data pipeline system 1138 may execute the data processing instructions 1302, 1304, 1306 to generate the datasets 1308, 1310, 1312. In one or more embodiments, the datasets 1308, 1310, 1312 may be stored by the integrated data repository 1104 or by an additional data repository accessible to the data integration and analysis system 1102. At least some of the data processing instructions 1302, 1304, 1306 may analyze health insurance codes and generate at least some of the datasets 1308, 1310, 1312. Additionally, at least some of the data processing instructions 1302, 1304, 1306 may analyze genomics data and generate at least some of the datasets 1308, 1310, 1312.

[0179] In one or more examples, the first data processing instructions 1302 may be executable to read data from one or more first data tables stored by the integrated data repository 1104. The first data processing instructions 1302 may also be executable to read data from one or more defined columns of the one or more first data tables. In various examples, the first data processing instructions 1302 may be executable to identify individuals having health insurance codes stored in one or more column and row combinations that correspond to the one or more diagnostic codes. The first data processing instructions 1302 may then be executable to analyze the one or more diagnostic codes to determine the biological condition with which the individual has been diagnosed. In one or more illustrative examples, the first data processing instructions 1302 may be executable to analyze the one or more diagnostic codes with respect to a library of diagnostic codes that indicate one or more biological conditions that correspond to the individual diagnostic codes. The library of diagnostic codes may include hundreds up to thousands of diagnostic codes. The first data processing instructions 1302 may also be executable to determine individuals who have been diagnosed with a biological condition by analyzing individual timing information, such as date of treatment, date of diagnosis, date of death, one or more combinations thereof, and the like.

[0180] The second data processing instructions 1304 may be executable to read data from one or more second data tables stored by the integrated data repository 1104. The second data processing instructions 1304 may also be executable to read data from one or more defined columns of the one or more second data tables. In various embodiments, the second data processing instructions 1304 may be executable to identify individuals having health insurance codes stored in one or more column and row combinations that correspond to one or more treatment codes. The one or more treatment codes may correspond to treatments obtained from a pharmacy. In one or more additional embodiments, the one or more treatment codes may correspond to treatments received through a medical procedure, such as an injection or intravenous. The second data processing instructions 1304 may be executable to determine one or more treatments that correspond to individual health insurance codes contained in the one or more second data tables by analyzing the health insurance codes in association with a predetermined set of information. The predetermined set of information may include a data library indicating one or more treatments corresponding to one of hundreds up to thousands of health insurance codes. The second data processing instructions 1304 may generate a second data set 1310 to indicate individual treatments received by a group of individuals. In one or more illustrative examples, the group of individuals may correspond to individuals included in the first data set 1308. The second data set 1310 may be arranged in rows and columns, with one or more rows corresponding to a single individual and one or more columns indicating treatments received by the individual individuals.

[0181] The Nth processing instructions 1306 (N may be any positive integer) may be executable to generate the Nth dataset 1312 by combining information from a number of previously generated datasets, such as the first dataset 1308 and the second dataset 1310. Additionally, the Nth processing instructions 1306 may be executable to generate the Nth dataset 1312, retrieve additional information from one or more additional columns of the integrated data repository 1104, and combine the additional information from the integrated data repository 1104 with information obtained from the first dataset 1308 and the second dataset 1310. For example, the Nth processing instructions 1306 may be executable to identify individuals included in the first dataset 1308 who have been diagnosed with a biological condition, analyze defined columns of one or more additional data tables of the integrated data repository 1104, and determine treatment dates indicated in the second dataset 1210 that correspond to the individuals included in the first dataset 1308. In one or more further examples, the Nth processing instructions 1306 may be executable to analyze columns of one or more additional data tables in the integrated data repository 1104 to determine dosages of treatments indicated in the second dataset 1310 received by individuals included in the first dataset 1308. In this manner, the Nth processing instructions 1306 may be executable to generate an episode of treatment dataset based on information included in the cohort dataset and the treatment dataset.

[0182] In one or more illustrative examples, in response to receiving the integrated data repository request 1142, the data analysis system 1140 may determine one or more datasets corresponding to characteristics of a query associated with the integrated data repository request 1142. For example, the data analysis system 1140 may determine that information included in the first dataset 1308 and the second dataset 1310 is applicable to responding to the integrated data repository request 1142. In these scenarios, the data analysis system 1140 may analyze at least a portion of the data included in the first dataset 1308 and the second dataset 1310 to generate the data analysis results 1146. In one or more additional examples, the data analysis system 1140 may determine different datasets for responding to different queries included in the integrated data repository request 1142 to generate the data analysis results 1146.

[0183] The use of specific sets of data processing instructions to generate the individual data sets may reduce the number of inputs from users of the data integration and analysis system 1102 and may reduce computational burdens, such as the amount of processing resources and memory utilized to process the integrated data repository requests 1142. For example, without the specific architecture of the data pipeline system 1138, each time an integrated data repository request 1142 is received, the data utilized to respond to the integrated data repository request 1142 is assembled from the data repositories 1104. In contrast, by implementing the data pipeline system 1138 to execute the data processing instructions 1302, 1304, 1306 to generate the data sets 1308, 1310, 1312, the data needed to respond to the various integrated data repository requests 1142 may already be assembled and accessed by the data analysis system 1140 to respond to the integrated data repository requests 1142. Thus, fewer computing resources are used to respond to an integrated data repository request 1142 by implementing a data pipeline system 1138 to generate data sets 1308, 1310, 1312 than a typical system that performs an information analysis and collection process for each integrated data repository request 1142. Furthermore, in situations where the data pipeline system 1138 is not implemented, a user of the data integration and analysis system 1102 may need to submit multiple integrated data repository requests 1142 to analyze the information that the user intends to be analyzed, either because the ad-hoc collection of data to respond to an integrated data repository request 1142 in a typical system is inaccurate or because the data analysis system 1140 is invoked multiple times to perform an analysis of the information in a typical system that may be performed using a single integrated data repository request 1142 when the data pipeline system 1138 is implemented.

[0184] FIG. 14 illustrates an architecture 1400 for generating an integrated data repository including de-identified health insurance claims data and de-identified genomics data, according to one or more implementations. The architecture 1400 may include a data integration and analysis system 1102, a health insurance claims data repository 1106, and a molecular data repository 1108. The data integration and analysis system 1102 may obtain patient information 1402 from the molecular data repository 1108. The patient information 1402 may include genomics data 1404 regarding an individual having data stored by the molecular data repository 1108. The genomics data 1404 may represent the results of one or more nucleic acid sequencing operations that analyze the sequence of nucleic acid molecules contained in a sample obtained from the individual for one or more target genomic regions. In one or more embodiments, the sample may be obtained from tissue of one or more individuals. In one or more additional examples, the sample may be obtained from one or more individual fluids, such as blood or plasma. The one or more target genomic regions may correspond to genomic regions corresponding to the presence of one or more biological conditions. For example, the target regions may correspond to genomic regions of a reference genome having mutations present in individuals with a biological condition. In one or more illustrative examples, the target regions may correspond to genomic regions of a reference human genome having one or more mutations present in individuals with one or more forms of cancer. The patient information 1402 may also include information indicative of personal information about the individual with data stored by the molecular data repository 1108 and information corresponding to tests and analyses performed on the sample provided by the individual.

[0185] The data integration and analysis system 1102 may perform a de-identification process 1406 that anonymizes personal information obtained from the molecular data repository 1108. The data integration and analysis system 1102 may implement one or more computational techniques as part of the de-identification process to anonymize data related to individuals stored by the molecular data repository 1108 such that the de-identified data protects the privacy of the individuals and complies with one or more privacy regulatory frameworks. The de-identification process 1406 may include accessing 1408 a token. In various embodiments, the token may comprise an alphanumeric string. In one or more embodiments, the token may be generated by the data integration and analysis system 1102. In one or more additional embodiments, the token may be generated by a third party and obtained by the data integration and analysis system 1102.

[0186] The token may be generated using one or more hash functions in association with a subset 1410 of the patient information 1402. To illustrate, for an individual whose information is stored by the molecular data repository 1108, the token may be generated using a combination of at least a portion of the individual's first name, at least a portion of the individual's last name, at least a portion of the individual's date of birth, the individual's gender, and at least a portion of the individual's location identifier. The de-identification process 1406 may also include generating an identifier for the individual whose data is stored by the molecular data repository 1108, at 1412. The identifier may be generated by the data integration and analysis system 1102 using one or more hash functions that are different from the one or more hash functions used to generate the token. In one or more illustrative examples, the data integration and analysis system 1102 may use one or more hash functions to generate intermediate versions of the individual identifiers and then apply one or more salting techniques to the intermediate versions of the identifiers to generate a final version of the identifiers. The salt function comprises a function configured to add at least one random bit to each intermediate identifier to generate an individual final identifier. In various examples, the data integration and analysis system 1102 may generate 1412 the identifier using at least a portion of the information about the individual individual stored by the molecular data repository 1108. In one or more illustrative examples, the identifier may be generated based on a patient identifier included in the patient information 1402. The identifier generated by the data integration and analysis system 1102 may be unique with respect to the individual individual having data stored by the molecular data repository 1108.

[0187] In an operation 1414, the data integration and analysis system 1102 may generate corrected patient information 1416 based on the identifier. The corrected patient information 1416 may include genomics data 1404 associated with an individual associated with the molecular data repository 1108 and an identifier for the individual. The corrected patient information 1416 may have a data structure 1418. The data structure 1418 may include a column including the individual identifier for the individual associated with the molecular data repository 1108 and a number of columns including genomics data 1404 associated with the individual, such as identifiers of one or more genes, one or more genetic modifications, a type of genetic modification, etc.

[0188] The data integration and analysis system 1102 may generate a token file 1420. The token file 1420 may include a first token 1422, accessed in operation 1408, for an individual having data stored by the molecular data repository 1108. The token file 1420 may have a data structure 11424 including a number of columns that include information about the individual individual. The data structure 11424 may include a column indicating an individual identifier generated by the data integration and analysis system 1102 and a column indicating one or more first tokens 1422 associated with the individual identifier. The data integration and analysis system 1102 may transmit the token file 1420 to a health insurance claims data management system 1426 coupled to the health insurance claims data repository 1106. The health insurance claims data management system 1426 may analyze the first token 1422 for a corresponding second token 1428. The second token 1428 may be accessed by or generated by the health insurance claims data management system 1426. The second token 1428 may be generated using a subset of information about an individual having data stored in the health insurance claims data repository 1106 that is the same or similar to the subset 1410 of patient information 1402. For example, the second token 1428 may be generated using a combination of at least a portion of the individual's first name, at least a portion of the individual's last name, at least a portion of the individual's date of birth, the individual's gender, and at least a portion of the individual's location identifier.

[0189] In various embodiments, the health insurance claims data management system 1426 may retrieve health insurance claims data for an individual associated with an individual second token 1428 that matches a corresponding first token 1422 from the health insurance claims data repository 1106. A first token 1422 may match a second token 1428 when the data of the first token 1422 has at least a threshold amount of similarity with the data of the second token 1428. In one or more embodiments, a first token 1422 may match a second token 1428 when the data of the first token 1422 is identical to the data of the second token 1428.

[0190] In response to identifying health insurance claim data for an individual having a respective second token 1428 that corresponds to the respective first token 1422, the health insurance claim data management system 1426 may generate corrected health insurance claim data 1430. The health insurance claim data management system 1426 may transmit the corrected health insurance claim data 1430 to the data integration and analysis system 1102. In one or more embodiments, the corrected health insurance claim data 1430 may be formatted according to a data structure 1432. The data structure 1432 may include a column that includes a subset of the second tokens 1428 that correspond to the first tokens 1422 and a number of columns that include the health insurance claim data.

[0191] In operation 1434, the data integration and analysis system 1102 may integrate genomics data and health insurance claims data for individuals common to both the molecular data repository 1108 and the health insurance claims data repository 1106. The data integration and analysis system 1102 may determine individuals common to both the molecular data repository 1108 and the health insurance claims data repository 1106 by determining the genomics data and health insurance claims data that correspond to the common tokens. The data integration and analysis system 1102 may determine that a first token 1422 associated with a portion of the genomics data 1404 corresponds to a second token 1428 associated with a portion of the health insurance claims data by determining a measure of similarity between the first token 1422 and the second token 1428. In a scenario in which the first token 1422 has at least a threshold amount of similarity with respect to the second token 1428, the data integration and analysis system 1102 may store the corresponding portion of the genomics data 1404 and the corresponding portion of the health insurance claims data in association with the individual's identifier in an integrated data repository, such as the integrated data repository 1104 of Figures 1, 2, and 3.

[0192] An implementation of the architecture 1400 may implement a cryptographic protocol that allows de-identified information from disparate data repositories to be consolidated into a single data repository. In this manner, the security of the data stored by the consolidated data repository 1104 is increased. In addition, the cryptographic protocol implemented by the architecture 1400 may allow for more efficient reading and accurate analysis of the information stored by the consolidated data repository 1104 than in situations in which the cryptographic protocol of the architecture 1400 is not utilized. For example, by generating a token file 1420 including a first token 1422 using cryptographic techniques based on a defined set of information stored by the molecular data repository 1104, and utilizing a second token 1428 generated using the same or similar cryptographic techniques for a similar or identical set of information stored by the health insurance claims data repository 1106, the data integration and analysis system 1102 may match information stored by the disparate data repositories corresponding to the same individual. Without implementing the cryptographic protocols of architecture 1400, the probability of erroneously attributing information from a data repository to one or more individuals increases, which reduces the accuracy of results provided by the data integration and analysis system 1102 in response to an integrated data repository request 1142 sent to the data integration and analysis system 1102.

[0193] FIG. 15 illustrates a framework 1500 for generating a dataset by the data pipeline system 1138 based on data stored by the integrated data repository 1104, according to one or more implementations. The integrated data repository 1104 may store health insurance claims data and genomics data for a group of individuals 1502. For example, the integrated data repository 1104 may store information obtained from health insurance claims records 1504 for the group of individuals 1502. For each individual included in the group of individuals 1502, the integrated data repository 1104 may store information obtained from multiple health insurance claims records 1504. In various embodiments, the information stored by the integrated data repository 1104 may include and / or be derived from thousands, tens of thousands, hundreds of thousands, or up to millions of health insurance claims records 1504 for a number of individuals. In addition, each health insurance claim record may include multiple columns. As a result, the integrated data repository 1104 may be generated through the analysis of millions of columns of health insurance claims data.

[0194] Further, while the health insurance claims data may be organized according to a structured data format, the health insurance claims data is typically arranged to be viewed by health insurance providers, patients, and health care providers to show financial and insurance code information related to the services provided by the health care provider to the individual. Thus, the health insurance claims data is not easily analyzed to obtain insights that may be available in relation to the characteristics of an individual in whom a biological condition exists and that may aid in the treatment of the individual in relation to the biological condition. The integrated data repository 1104 may be generated and organized by analyzing and modifying the raw health insurance claims data in a manner that allows the data stored by the integrated data repository 1104 to be further analyzed to determine trends, characteristics, features, and / or insights regarding an individual in whom one or more biological conditions may exist. For example, health insurance codes may be stored within the integrated data repository 1104 in a manner such that at least one of a medical procedure, a biological condition, a treatment, a dosage, a drug manufacturer, a drug distributor, or a diagnosis may be determined for a given individual based on the health insurance claims data regarding the individual. In various embodiments, the data integration and analysis system 1102 may generate and implement one or more tables showing correlations between health insurance claims data and various treatments, symptoms, or biological conditions corresponding to the health insurance claims data. Additionally, the integrated data repository 1104 may be generated using the genomics data records 1506 of the group of individuals 1502. In various embodiments, large volumes of health insurance claims data may be matched with genomics data for the group of individuals 1502 to generate the integrated data repository 1104.

[0195] By integrating the genomics data records 1506 for a group of individuals 1502 with the health insurance claims records 1504, the data integration and analysis system 1102 may determine correlations between the presence of one or more biomarkers present in the genomics data records 1506 and other characteristics of the individuals indicated by the health insurance claims data records 1506 that existing systems typically cannot determine. For example, the data integration and analysis system 1102 may determine one or more genomic characteristics of the individuals corresponding to treatments received by the individuals, the timing of the treatments, the dosage of the treatments, the individual's diagnosis, smoking status, the presence of one or more biological conditions, the presence of one or more symptoms of the biological conditions, one or more combinations thereof, and the like. Based on the correlations determined by the data integration and analysis system 1102 using the integrated data repository 1104, cohorts of individuals that may benefit from one or more treatments that would not have been identified in existing systems may be identified. In one or more embodiments, the processes and techniques implemented to integrate the health insurance claim records 1504 and the genomics claim records 1506 to generate the integrated data repository 1104 may be complex, and efficiency-improving techniques, systems, and processes may be implemented to minimize the amount of computing resources used to generate the integrated data repository 1104.

[0196] In one or more illustrative examples, the data pipeline system 1138 may access information stored by the integrated data repository 1104 and generate a data set including a number of additional data records 1508 including information related to at least a portion of the group of individuals 1502. In the illustrative example of FIG. 5, the additional data records 1508 include information indicating whether the individual is included within a cohort of individuals in which lung cancer exists. The data pipeline system 1138 may execute a plurality of different sets of data processing instructions to determine the cohort of groups of individuals 1502 in which lung cancer exists. In various examples, the additional data records 1508 may indicate information used to determine the status of the individual 1502 with respect to lung cancer, such as one or more insurance procedure identifiers, one or more International Classification of Diseases (ICD) codes, and one or more health insurance procedure dates. In addition to including a column indicating whether the individual 1502 is included in the lung cancer cohort, the additional data record 1508 may include a column indicating a confidence level of the individual's 1502 status with respect to the presence of lung cancer.

[0197] FIG. 16 illustrates a system 1600 for determining a cohort of patients having at least a primary diagnosis of a biological condition, according to one or more implementations. The system 1600 can include a data integration and analysis system 1102 and an integrated data repository 1104. The data integration and analysis system 1102 can analyze information stored by the integrated data repository 1104 to determine a cohort of patients in which one or more biological conditions are present. For example, the data integration and analysis system 1102 can determine a first patient population having data stored by the integrated data repository 1104 in which a first biological condition is present, and a second patient population having data stored by the integrated data repository 1104 in which a second biological condition is present. The data analysis and integration system 1102 can include at least a data pipeline system 1138 and a data analysis system 1140. The data analysis system 1140 can generate a data analysis result 1146 based on the data obtained from the data pipeline system 1138. The data analysis system 1140 can also analyze the additional data obtained from the integrated data repository 1140 and generate data analysis results 1146 .

[0198] The data pipeline system 1138 can include a cohort selection system 1602 that analyzes data obtained from the integrated data repository 1104 to determine a cohort of patients in which one or more biological conditions are present. In various embodiments, the cohort selection system 1602 can analyze data from tens of thousands up to hundreds of thousands of patients or more to determine a cohort of patients. In one or more embodiments, hundreds of thousands up to millions of health insurance claim records or more are analyzed by the cohort selection system 1602 to determine one or more cohorts of patients. Additionally, the cohort selection system 1602 can analyze millions to tens of millions of health insurance claim codes or more to determine one or more cohorts of patients. In one or more illustrative embodiments, the cohort selection system 1602 can analyze information obtained from the integrated data repository 1104 according to a cohort identification framework 1604 to efficiently analyze such large amounts of information. The cohort identification framework 1604 includes at least one of a number of rules, a number of schemes, or logic by which data obtained from the integrated data repository 1104 can be analyzed. The cohort identification framework 1604 also provides a structure for identifying information stored in the integrated data repository 1104 that can be used to accurately identify patients in which a given biological condition is present. The cohort identification framework 1604 can be determined by implementing and / or training at least one of one or more machine learning techniques, one or more statistical techniques, or one or more additional computational techniques in association with a corpus of data related to the patient's diagnosis using the health insurance data. In various illustrative examples, the cohort identification framework 1604 can include at least a portion of the process described with respect to Figures 4-10.

[0199] In one or more examples, a medical header data table 1606 can be provided to the cohort selection system 1602 and analyzed according to the cohort identification framework 1604 to determine one or more cohorts of patients. The medical header data table 1606 can include multiple rows and multiple columns. The multiple rows can correspond to medical practices for a number of patients. A medical practice can correspond to visits to one or more healthcare providers, services provided by one or more healthcare providers, therapies provided to one or more patients, or one or more combinations thereof. In various examples, an individual medical practice can indicate charges for at least one of the services or products provided to a number of patients. For example, an individual medical practice can indicate one or more health insurance claims associated with the medical practice. In one or more additional examples, the data table 1606 can indicate that an individual patient is associated with one or more medical practices. An individual medical practice can correspond to a given medical practice legend and / or a given medical practice identifier. In at least some examples, an individual medical practice legend or an individual medical practice identifier can uniquely identify a given medical practice. The medical header data table 1606 can also include one or more columns indicating utilization dates for the individual medical practice. In addition, the medical header data table 1606 can include one or more columns indicating diagnosis codes for patients associated with the individual medical practice. To illustrate, the medical header data table 1606 can indicate one or more biological conditions for which the individual patient was treated.

[0200] The cohort identification framework 1604 can indicate one or more columns of a medical header data table 1606 to be analyzed by the cohort selection system 1602 to determine one or more cohorts of patients. To illustrate, the cohort identification framework 1604 can indicate one or more columns of a medical header data table 1606 that include health insurance codes corresponding to diagnoses of biological conditions. The cohort identification framework 1604 can also indicate the format of one or more types of health insurance diagnosis codes. For example, the cohort identification framework 1604 can indicate the format of an International Classification of Diseases (ICD) code. In one or more illustrative examples, the cohort identification framework 1604 can indicate the format of an ICD 9th edition code, the format of an ICD 10th edition code, the format of an ICD 11th edition code, or the format of another ICD version.

[0201] In addition, the cohort identification framework 1604 can indicate a diagnostic code corresponding to one or more biological conditions. In various examples, the cohort identification framework 1604 can indicate a diagnostic code corresponding to one or more forms of a given biological condition, such as one or more forms of cancer. To illustrate, the cohort identification framework 1604 can indicate one or more first diagnostic codes corresponding to a first form of cancer and one or more second diagnostic codes corresponding to a second form of cancer. In one or more additional illustrative examples, the cohort identification framework 1604 can indicate one or more first diagnostic codes in a first format corresponding to a first form of cancer, one or more second diagnostic codes in a second format corresponding to the first form of cancer, one or more third diagnostic codes in the first format corresponding to a second form of cancer, and one or more fourth diagnostic codes in the second format corresponding to a second form of cancer. In one or more examples, the diagnostic codes can correspond to a biological condition that is a primary diagnosis for the patient. In one or more examples, the cohort identification framework 1604 can include diagnostic codes that indicate one or more biological conditions that do not indicate the presence of a biological condition for one or more patients.

[0202] The cohort identification framework 1604 can indicate at least one of logic or rules for analyzing the information contained within the medical header data table 1606. For example, the cohort identification framework 1604 can indicate a threshold time period to be used to determine at least one of a primary diagnosis for the patient or a secondary diagnosis for the patient. In one or more illustrative examples, the cohort identification framework 1604 can indicate that a patient having a medical practice corresponding to a first diagnostic code and another medical practice corresponding to a second diagnostic code within the threshold time period may be excluded from a first cohort corresponding to a first biological condition associated with the first diagnostic code. Additionally, in these scenarios, the cohort identification framework 1604 can indicate that the patient may be excluded from a second cohort corresponding to a second biological condition associated with the second diagnostic code. Furthermore, the cohort identification framework 1604 can indicate that a patient having a medical practice indicating a new diagnostic code that exceeds an additional threshold time period after a previous medical practice indicating an initial diagnostic code may be excluded from a cohort corresponding to a biological condition associated with the initial diagnostic code. In various embodiments, the cohort identification framework 1604 can represent logic for determining how to categorize patients in situations where a patient receives treatment for multiple biological conditions on the same day.

[0203] Additionally, the cohort identification framework 1604 may include at least one of logic or one or more criteria for determining a quality metric for a diagnosis of a patient by the cohort selection system 1602 and including the patient in a cohort. The quality metric may correspond to a probability that a diagnosis of the patient by the cohort selection system 1602 corresponds to a biological condition present in the patient. In one or more examples, the quality metric can be a quantitative metric, such as a score or a range of probabilities. In one or more additional examples, the quality metric can be a qualitative metric, such as "low," "medium," or "high."

[0204] In one or more examples, the cohort selection system 1602 can implement the cohort identification framework 1604 and generate a first cohort data table 1608 corresponding to a first cohort of patients having health insurance records stored by the integrated data repository 1104. The first cohort data table 1608 can correspond to a first primary diagnosis 1610 of patients included in the first cohort. The cohort selection system 1602 can also implement the cohort identification framework 1604 and generate a second cohort data table 1612 corresponding to a second cohort of patients having health insurance records stored by the integrated data repository 1104. The second cohort data table 1612 can correspond to a second primary diagnosis 1614 of patients included in the second cohort. In addition, the cohort selection system 1602 can implement the cohort identification framework 1604 and generate a third cohort data table 1616 corresponding to a third cohort of patients having health insurance records stored by the integrated data repository 1104. The third cohort data table 1616 can correspond to patients having multiple diagnoses 1618, such as a primary diagnosis and a secondary diagnosis. In one or more illustrative examples, the first primary diagnosis 1610 can include a first biological condition, such as type 2 diabetes, and the second primary diagnosis can include a second biological condition, such as hypertension. In one or more additional illustrative examples, the multiple diagnoses 1618 can correspond to a primary diagnosis of type 2 diabetes and a secondary diagnosis of hypertension. In one or more further illustrative examples, the first primary diagnosis 1610 may correspond to a first form of cancer, the second primary diagnosis 1614 may correspond to a second form of cancer, and the multiple diagnoses 1618 may correspond to a patient having cancer that has metastasized, such that the patient has a primary diagnosis of a first form of cancer and a secondary diagnosis of a second form of cancer.

[0205] In various examples, at least one of the first cohort data table 1608, the second cohort data table 1612, or the third cohort data table 1616 can include information about patients included in an individual cohort. For example, the data tables 1608, 1612, 1616 can indicate identifiers for patients whose data is stored by the integrated data repository 1104. The data tables 1608, 1612, 1618 can also indicate patient personal information such as the patient's age, the patient's year of birth, the patient's date of birth, the patient's date of death, one or more dates of health insurance claim activity, the patient's primary diagnosis, the patient's secondary diagnosis, the patient's metastatic status, one or more combinations thereof, etc.

[0206] In one or more examples, the data tables 1608, 1612, 1616 can be provided to a data analysis system 1140, which can generate a data analysis result 1146 using at least a portion of the information stored by the data tables 1608, 1612, 1616. In various examples, the data analysis system 1140 can use at least the patient identifier to retrieve additional data corresponding to a patient contained in at least one of the data tables 1608, 1612, 1616 from the integrated data repository 1104. In one or more additional examples, the data analysis system 1140 can use the patient identifier in combination with additional information such as the patient's age, the patient's year of birth, the patient's date of birth, and the like to retrieve additional data corresponding to a patient contained in at least one of the data tables 1608, 1612, 1616 from the integrated data repository 1104. In one or more illustrative examples, the data analysis system 1140 can use information stored by at least one of the data tables 1608, 1612, 1616 to read at least one of the patient's genomic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, and / or proteomic information and generate data analysis results 1146.

[0207] In at least some examples, the cohort selection system 1602 can analyze the data contained in the medical header data table 1606 to determine patients having a set of health insurance billing codes in one or more diagnosis columns of the medical header data table 1604. The health insurance billing codes may correspond to a set of ICD 9th edition codes and / or a set of ICD 10th edition codes corresponding to a primary diagnosis of a biological condition. For example, the cohort selection system 1602 can analyze the medical header data table 1606 to identify patients having one or more ICD 9th edition diagnosis codes corresponding to non-small cell lung cancer and one or more ICD 10th edition diagnosis codes corresponding to non-small cell lung cancer.

[0208] The cohort selection system 1602 can generate an intermediate data table that stores identification information of a first patient number corresponding to a defined health insurance billing code. In one or more illustrative examples, the intermediate data table can be temporarily stored in a memory, such as in a cache memory, while additional analysis is performed by the cohort selection system 1602 on the data related to the first patient number. For example, the cohort selection system 1602 can analyze the additional health insurance claims data with respect to the first patient number according to the cohort identification framework 1604. In this manner, the cohort selection system 1602 can identify patients who may have a diagnosis of a biological condition corresponding to a defined health insurance billing code, but where the biological condition is not the patient's primary diagnosis. In this manner, the cohort selection system 1602 can implement a multi-step analysis that uses one or more intermediate data tables to accurately and efficiently determine patients for inclusion in the data tables 1608, 1612, 1616. In various examples, the cohort selection system 1602 can implement a cohort identification framework 1604 in conjunction with logic corresponding to the diagnosis date and / or in conjunction with additional health insurance billing codes to determine the number of second patients having a primary diagnosis of a biological condition.

[0209] In one or more further examples, the cohort selection system 1602 can analyze additional information to determine patients for inclusion in the data tables 1608, 1612, 1616. For example, the cohort selection system 1602 can analyze histology information stored by the integrated data repository 1104 in addition to health insurance claims data to identify patients for inclusion in the cohort who have a biological condition associated with the cohort. To illustrate, the histology records can also include diagnosis information. In these scenarios, the cohort selection system 1602 can analyze diagnosis information related to the biological condition stored by the integrated data repository 1104 in conjunction with health insurance claims data related to a diagnosis of the biological condition to determine patients for inclusion in the cohort. In one or more illustrative examples, the cohort selection system 1602 can analyze the histology information and health insurance claims data in accordance with the cohort identification framework 1604 to determine one or more patients in the cohort who have a primary diagnosis associated with the biological condition.

[0210] The cohort selection system 1602 can also generate one or more additional data tables. For example, the cohort selection system 1602 can generate a diagnosis data table indicating one or more diagnoses for an individual patient. To illustrate, the cohort selection system 1602 can determine that a patient is included in one or more cohorts. In these scenarios, the cohort selection system 1602 can indicate in the diagnosis data table that the patient is diagnosed with a biological condition corresponding to one or more cohorts that include the patient. The diagnosis data table can indicate a biological condition corresponding to the patient's primary diagnosis, an additional biological condition corresponding to the patient's secondary diagnosis, a metastatic condition of the patient, or one or more combinations thereof. The implementation for determining the diagnosis table for an individual patient also indicates the patient's different diagnoses over time. In addition to the patient's diagnosis, the diagnosis data table can also indicate a health insurance code corresponding to the diagnosis, a patient identifier, a date of treatment, a date associated with the patient's diagnosis, a most recent diagnosis, or one or more combinations thereof. In various examples, the data analysis system 1140 can analyze information contained in or derived from the diagnostic data tables to generate data analysis results 1146. In one or more illustrative examples, the diagnostic data tables can be used to generate real-world evidence measures, such as real-world overall survival (rwOS), to determine the data analysis results 1146.

[0211] In one or more illustrative examples, the data analysis system 1140 may analyze information stored by one or more of the data tables 1608, 1612, 1616 and / or information retrieved from the integrated data repository 1104 based on one or more of the data tables 1608, 1612, 1616 to determine a data analysis result 1146. In one or more examples, the data analysis system 1140 may receive a request to analyze information corresponding to a cohort of patients treated for a given biological condition. In response to the request, the data analysis system 1140 may analyze the information generated by the cohort selection system 1602 and generate a data analysis result 1146 including one or more quantitative measurements corresponding to the patients included in the one or more cohorts. To illustrate, the data analysis system 1140 may analyze the information generated by the cohort selection system 1602 and determine a real-world survival measurement for the patients included in the cohort. In various examples, the data analysis system 1140 may analyze information related to a cohort of patients and determine a survival probability over a period of time for patients included in the cohort. In one or more illustrative examples, the data analysis system 1140 may analyze information related to one or more cohorts of patients and determine a real-world overall survival measurement for patients included in the cohort. In one or more additional illustrative examples, the data analysis system 1140 may analyze information related to a cohort identified by the cohort selection system 1602 and determine a time to next treatment and / or a time to discontinuation measurement for patients included in one or more cohorts.

[0212] In various examples, the data analysis system 1140 may analyze information corresponding to patients included in a cohort identified by the cohort selection system 1602 to determine a degree of progression of a biological condition in at least a subset of patients included in the cohort. In one or more examples, the data analysis system 1140 may determine a degree of progression for a cohort of patients receiving one or more pharmaceutical substances as part of a course of therapy based on an analysis of the information generated by the cohort selection system 1602. Additionally, the data analysis system 1140 may determine a degree of progression for a cohort of patients having one or more genomic mutations based on an analysis of the information generated by the cohort selection system 1602. In one or more illustrative examples, the data analysis system 1140 may analyze at least one of a time to next treatment or a time to discontinuation measurement for the cohort of patients to determine a degree of progression of a biological condition for the patients in the cohort having a genomic mutation. In these cases, the data analysis system 1140 may query the integrated data repository 1104 to determine genomic data for patients included in the cohort and identify patients in the cohort who have one or more defined genomic mutations. The data analysis system 1140 may then analyze the time to next treatment, time to discontinuation, and / or real-world overall survival measurements of patients included in the cohort who have one or more genomic mutations to determine the progression of the biological condition for patients included in the cohort who have been treated for the biological condition.

[0213] In one or more further examples, the data analysis system 1140 may analyze information generated by the cohort selection system 1602 to determine a level of resistance expressed by one or more patients included in the cohort who have received one or more treatments for a biological condition associated with the cohort. For example, the data analysis system 1140 may analyze information of a cohort of patients identified by the cohort selection system 1602 to determine a level of resistance in one or more patients of the cohort who have received one or more pharmaceutical agents as part of a course of therapy to treat a biological condition of the patient included in the cohort. In various examples, the data analysis system 1140 may analyze at least one of a time to next treatment measurement, a time to discontinuation measurement, or a real-world survival measurement to determine a level of resistance expressed by the patients of the cohort who have received the treatment. In at least some examples, the data analysis system 1140 may also determine a level of resistance to one or more treatments for patients in the cohort who have one or more genomic mutations. In at least some embodiments, the level of resistance may be greater in situations where the time to next treatment or real-world survival rate has a lower value, and the level of resistance may be lower in situations where the time to next treatment or real-world survival rate has a relatively high value.

[0214] In at least some embodiments, the data analysis system 108 may analyze the set of therapy information stored by the one or more set of therapy data structures 836 corresponding to the biological condition and determine a recommendation for one or more therapies to administer to a patient diagnosed with the biological condition. In one or more embodiments, the data analysis system 1140 may analyze information about a cohort of patients identified by the cohort selection system 1602 and determine one or more characteristics of the patients in the cohort who have received one or more sets of therapies with a relatively low level of resistance and / or a relatively low degree of progression. The data analysis system 1140 may then analyze characteristics of one or more additional patients in the cohort who are diagnosed with the biological condition and determine whether the one or more sets of therapies should be recommended as treatment for the one or more additional patients. At least a portion of the one or more additional patients in the cohort may already be receiving treatment for the biological condition. In one or more additional examples, at least a portion of one or more additional patients of the cohort may not be receiving treatment for the biological condition associated with the cohort. In various examples, the data analysis system 1140 may also analyze information of patients included in a given cohort to determine the effectiveness of a course of therapy for patients included in the cohort. The effectiveness of a course of therapy may correspond to the probability of at least one of the course of therapy reducing or eliminating the effect of the biological condition on the patients of the cohort.

[0215] In various examples, the degree of progression of the biological condition, the effectiveness of a course of therapy to treat the biological condition, the probability of developing resistance to a course of therapy, or a combination thereof may be determined by the data analysis system 1140 using at least one of one or more statistical techniques or one or more machine learning techniques. To illustrate, the data analysis system 1140 may implement at least one of a Cox proportional hazards model, a chi-square test, a log-rank test, or a Kaplan-Meier method to determine at least one of the degree of progression of the biological condition, the effectiveness of a course of therapy to treat the biological condition, or the probability of developing resistance to a course of therapy. In one or more additional examples, the data analysis system 1140 may implement one or more neural networks, one or more convolutional neural networks, or one or more residual neural networks to determine at least one of the degree of progression of the biological condition, the effectiveness of a course of therapy to treat the biological condition, or the probability of developing resistance to a course of therapy.

[0216] In one or more illustrative examples, the data analysis system 1140 may determine one or more characteristics of patients having at least one of a below threshold probability of developing resistance to a course of therapy or at least an additional threshold amount of efficacy for the course of therapy. In one or more scenarios, the data analysis system 1140 may analyze information about a cohort of patients determined by the cohort selection system 1602 to determine one or more characteristics. In at least some examples, the data analysis system 1140 may implement at least one of one or more statistical techniques or one or more machine learning techniques to determine one or more characteristics of patients having at least one of a below threshold probability of developing resistance to a course of therapy or at least an additional threshold amount of efficacy for the course of therapy. In one or more examples, the data analysis system 1140 may implement at least one of one or more extraction algorithms or one or more classification algorithms to determine one or more characteristics. In various examples, the data analysis system 1140 may implement at least one of one or more neural networks, one or more feedforward neural networks, one or more recurrent neural networks, one or more residual networks, or one or more autoencoders to determine one or more characteristics having at least one of less than a threshold probability of developing resistance to a course of therapy or at least an additional threshold amount of efficacy for the course of therapy.

[0217] In one or more additional illustrative examples, the data analysis system 1140 may implement one or more log-rank tests to analyze the difference between the time to death and time to next treatment measurements determined based on information of one or more cohorts of patients determined by the cohort selection system 1602 having one or more genomic mutations and diagnosed with or suspected to have a given biological condition. In various examples, the patients included in the analysis may also have received one or more defined courses of therapy to treat the biological condition. In addition, the data analysis system 1140 may implement one or more chi-square tests to determine the proportion of patients included in the cohort that have one or more co-occurring genomic mutations among patients of the cohort that have one or more defined genomic mutations, and in at least some cases, one or more additional genomic characteristics, such as one or more clonal genomic mutations versus one or more subclonal genomic mutations. Additionally, one or more Cox proportional hazards models may be implemented by the data analysis system 1140 to determine a survival metric for the patient. In this manner, the effectiveness of one or more courses of therapies for treating a biological condition may be determined by the data analysis system 1140 based on the survival probabilities determined using the Cox proportional hazards models.

[0218] The cohort identification framework 1604, in addition to one or more computational techniques implemented by the cohort selection system 1602 and, in at least some cases, intermediate data tables generated by the cohort selection system 1602, may be used by the data analysis system 1140 to accurately generate the data analysis results 1146. That is, based on the cohorts identified by the cohort selection system 1602, real-world survival measures, disease progression measures, disease resistance measures, treatment efficacy levels, one or more combinations thereof, etc. may be accurately determined for the cohorts since patients included within the cohort have at least a threshold probability of a given biological condition being present. Accurate determination of these quantitative measures enables the data analysis system 1140 to provide treatment recommendations for patients that are accurate, effective, and result in improved outcomes for the patients. Without the procedures, rules, schemes, and protocols defined in the cohort identification framework 1604 and the computational techniques implemented by the cohort selection system 1602 and the data analysis system 1140, the treatment recommendations contained within the data analysis results 1146 are unlikely to improve outcomes for the patient. The cohort identification framework 1604 is generated over time using a number of computational techniques, training processes, and feedback loops to determine a defined set of criteria, rules, schemes, protocols, thresholds, and computational techniques that result in optimal treatment recommendations, provide accurate measurements indicative of the effectiveness of a course of therapy on outcomes for a cohort of patients, and provide accurate information regarding the impact of genomic mutations of a patient cohort on treatment outcomes.

[0219] FIG. 17 illustrates a circuit block diagram of a computing machine 1700 according to some implementations. In some implementations, components of the computing machine 1700 may house or be integrated within other components shown in the circuit block diagram of FIG. 17. For example, a portion of the computing machine 1700 may reside within the processor 1702 and may be referred to as "processing circuitry." The processing circuitry may include processing hardware, such as one or more central processing units (CPUs), one or more graphic processing units (GPUs), and the like. In alternative implementations, the computing machine 1700 may operate as a standalone device or may be connected (e.g., networked) to other computers. In a networked deployment, the computing machine 1700 may operate in the capacity of a server, a client, or both in a server-client network environment. In some implementations, the computing machine 1700 may act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. In this document, the terms "P2P", "Device to Device (D2D)", and "Sidelink" may be used interchangeably. The computing machine 1700 may be a special purpose computer, a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a mobile phone, a smart phone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) that define actions to be taken by the machine.

[0220] An embodiment as described herein may include or operate on logic or several components, modules, or mechanisms. Modules and components are tangible entities (e.g., hardware) capable of performing a specified operation and may be configured or arranged in a certain manner. In an embodiment, a circuit may be arranged in a specified manner (e.g., internally or with respect to an external entity such as other circuits) as a module. In an embodiment, all or part of one or more computer systems / devices (e.g., stand-alone, client, or server computer systems) or one or more hardware processors may be configured by firmware or software (e.g., instructions, application parts, or applications) as a module that operates to perform a specified operation. In an embodiment, the software may reside on a machine-readable medium. In an embodiment, the software, when executed by the underlying hardware of the module, causes the hardware to perform the specified operation.

[0221] Thus, the term "module" (and "component") is understood to encompass tangible entities that are entities that are physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transiently) configured (e.g., programmed) to operate in a specified manner or to perform some or all of any operations described herein. Considering examples in which modules are temporarily configured, the modules need not each be instantiated at any one moment in time. For example, if the modules comprise a general-purpose hardware processor core that is configured using software, the general-purpose hardware processor may be configured as separate and distinct modules at different times. The software may thus configure the hardware processor, e.g., configure a particular module at one instance of time and configure a different module at a different instance of time.

[0222] The computing machine 1700 may include a hardware processor 1702 (e.g., a central processing unit (CPU), a GPU, a hardware processor core, or any combination thereof), a main memory 1704, and a static memory 1706, some or all of which may communicate with each other via an interlink (e.g., a bus) 1708. Although not shown, the main memory 1704 may contain any removable and non-removable storage, volatile memory, or non-volatile memory. The computing machine 1700 may further include a video display unit 1710 (or other display unit), an alphanumeric input device 1712 (e.g., a keyboard), and a user interface (UI) navigation device 1714 (e.g., a mouse). In one embodiment, the display unit 1710, the input device 1712, and the UI navigation device 1714 may be touch screen displays. The computing machine 1700 may additionally include a storage device (e.g., a drive unit) 1716, a signal generating device 1718 (e.g., a speaker), a network interface device 1720, and one or more sensors 1621, such as a Global Positioning System (GPS) sensor, a compass, an accelerometer, or other sensor. The computing machine 1700 may include an output controller 1728, such as a serial (e.g., Universal Serial Bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, to communicate with or control one or more peripheral devices (e.g., a printer, a card reader, etc.).

[0223] The drive unit 1716 (e.g., a storage device) may include a machine-readable medium 1722 on which is stored one or more sets of data structures or instructions 1724 (e.g., software) that embody or are utilized by any one or more of the techniques or functions described herein. The instructions 1724 may also reside, completely or at least partially, within the main memory 1704, within the static memory 1706, or within the hardware processor 1702 during its execution by the computing machine 1700. In an embodiment, one or any combination of the hardware processor 1702, the main memory 1704, the static memory 1706, or the storage device 1716 may constitute a machine-readable medium.

[0224] Although the machine-readable medium 1722 is illustrated as a single medium, the term “machine-readable medium” may include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) configured to store one or more instructions 1724.

[0225] The term "machine-readable medium" may include any medium capable of storing, encoding, or carrying instructions for execution by the computing machine 1700 and capable of storing, encoding, or carrying data structures used by or associated with the computing machine 1700 to perform any one or more of the techniques of this disclosure. Non-limiting machine-readable medium examples may include solid-state memory and optical and magnetic media. Specific examples of machine-readable media may include non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) and flash memory devices), magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, random access memory (RAM), and CD-ROM and DVD-ROM disks. In some embodiments, the machine-readable medium may include non-transitory machine-readable media. In some embodiments, the machine-readable medium may include machine-readable media that are not transitory propagating signals.

[0226] The instructions 1724 may further be transmitted or received over a communications network 1726 using a transmission medium via a network interface device 1720 utilizing any one of a number of transport protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Exemplary communications networks may include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile telephone networks (e.g., cellular networks), plain old telephone (POTS) networks, and wireless data networks (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards known as Wi-Fi, the IEEE 802.16 family of standards known as WiMax), the IEEE 802.15.4 family of standards, the Long Term Evolution (LTE) family of standards, the Universal Mobile Telecommunications System (UMTS) family of standards, peer-to-peer (P2P) networks, among others. In one embodiment, the network interface device 1720 may include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas to connect to the communications network 1726.

[0227] Some implementations are described as numbered examples (Example 1, 2, 3, etc.), which are provided by way of example only and are not intended to limit the technology disclosed herein.

[0228] Example 1 is a method implemented in one or more computing machines having a processing circuitry and a memory, the method comprising: accessing, in the processing circuitry, one or more medical data repositories storing data for a given patient from a plurality of patients, the one or more medical data repositories storing pharmacy data, clinic visit data, and medical insurance transaction data; identifying one or more biological conditions and metastatic status for the given patient based on one or more disease codes in the clinic visit data or the medical insurance transaction data; identifying one or more courses of treatment for the given patient based on one or more insurance codes in the medical insurance transaction data; identifying one or more medical procedures received by the given patient based on one or more insurance codes in the medical insurance transaction data; determining a primary diagnostic biological condition for the given patient based on a combination of the one or more biological conditions, metastatic status, one or more courses of treatment, and one or more medical procedures; assigning the given patient to a patient cohort based on the primary diagnostic biological condition; and providing an output representative of the assigned cohort for the given patient.

[0229] In Example 2, which includes the subject matter of Example 1, determining the primary diagnostic biological condition includes creating a master gap table based on medical insurance procedure data, the master gap table indicating the time interval between two consecutive procedures for a given patient having the same procedure name and insurance code, the master gap table having columns for procedure name, insurance code, unit, and gap length; calculating a median gap table based on the master gap table, the median gap table indicating the median gap for each combination of procedure name and insurance code, the median gap table having columns for procedure name, insurance code, unit, and gap length; and determining the primary diagnostic biological condition based at least in part on the data in the median gap table.

[0230] In Example 3, which includes the subject matter of Examples 1-2, the processing circuitry includes a plurality of multi-threaded graphic processing units (GPUs), and the method further includes determining, in parallel and using parallel threads of the plurality of multi-threaded GPUs, an assigned cohort for a plurality of patients, including the given patient, from the plurality of patients.

[0231] In Example 4, which includes the subject matter of Examples 1-3, the disease codes include International Classification of Diseases (ICD) codes, the drug codes include National Drug Codes (NDC) codes, and the insurance codes include Healthcare Common Procedure Coding System (HCPCS) codes.

[0232] In Example 5, which includes the subject matter of Example 4, identifying one or more biological conditions for a given patient based on one or more disease codes in the clinic visit data or health insurance transaction data includes identifying that the given patient is suffering from lung cancer based on an ICD code that is associated with lung cancer.

[0233] In Example 6, which includes the subject matter of Example 5, identifying a metastatic status for a given patient based on one or more disease codes in clinic visit data or health insurance transaction data is based on a secondary malignancy ICD code or HCPCS code.

[0234] In Example 7, which includes the subject matter of Examples 1-6, a primary diagnostic biological condition for a given patient is determined based on disease codes, drug codes, or insurance codes associated with dates within a predefined date range.

[0235] In Example 8, which includes the subject matter of Examples 1-7, assigning the given patient to a cohort includes analyzing one or more data tables including medical insurance procedure data for the given patient, the one or more data tables being from one or more medical data repositories; determining one or more first insurance procedures including one or more first code identifiers included in the one or more data tables, the one or more first code identifiers corresponding to the patient's diagnosis of one or more biological conditions; and determining one or more second insurance procedures including one or more second code identifiers included in the one or more data tables. determining a second medical procedure to be performed on the patient, the one or more second code identifiers corresponding to medical procedures performed on the patient, the medical procedures being from the clinic visit data; generating a medical header table including a first number of columns storing the one or more first code identifiers, a second number of columns storing the one or more second code identifiers, and a plurality of rows with each row of the plurality of rows corresponding to a first medical procedure of the one or more first medical procedures or a second medical procedure of the one or more second medical procedures; storing the medical header table in one or more medical data repositories; and determining a cohort for the patient based on the data in the medical header table.

[0236] In Example 9, which includes the subject matter of Example 8, each first insurance procedure of the one or more first insurance procedures indicates a date of use of the each first insurance procedure, and each second insurance procedure of the one or more second insurance procedures indicates a date of use of the each second insurance procedure.

[0237] In Example 10, the subject matter of Example 9 includes arranging a plurality of rows in the medical header table in ascending order based on utilization date such that the claim with the first utilization date is the first row in the medical header table and the procedure with the most recent utilization date is the last row in the medical header table.

[0238] In Example 11, the subject matter of Examples 8-10 includes analyzing one or more first code identifiers and determining that the first code identifier of the one or more first code identifiers is included within a group of insurance code identifiers corresponding to one or more biological conditions.

[0239] In Example 12, which includes the subject matter of Example 11, the first code identifiers are arranged according to a first format that corresponds to a first classification of insurance code identifiers, and the first classification of insurance code identifiers corresponds to the International Classification of Diseases, Ninth Revision (ICD-9).

[0240] In Example 13, which includes the subject matter of Examples 11-12, the first code identifiers are arranged according to a second format that corresponds to a second classification of insurance code identifiers, and the second classification of insurance code identifiers corresponds to the International Classification of Diseases, Tenth Revision (ICD-10).

[0241] In Example 14, which includes the subject matter of Examples 11-13, the one or more biological conditions include a plurality of subtypes, each subtype of the plurality of subtypes corresponding to a subset of a group of insurance code identifiers corresponding to the one or more biological conditions, and the method further includes determining that a first code identifier is included within a first subset of the group of insurance codes corresponding to a first subtype of the biological condition.

[0242] In Example 15, which includes the subject matter of Examples 11-14, the biological condition is cancer and the multiple subtypes include at least one of lung cancer, breast cancer, or colorectal cancer.

[0243] In Example 16, the subject matter of Example 15 includes determining one or more third insurance procedures having use dates within a predefined period ending with a first use date, and analyzing one or more third insurance code identifiers of the third insurance procedures against the first code identifier and the second code identifier.

[0244] In Example 17, the subject matter of Example 16 includes determining that one or more third insurance code identifiers are not included within the first code identifier and the second code identifier, and determining, based on the third insurance code identifier, that the patient is included within a cohort of patients in which one or more given subtypes of the biological condition are present.

[0245] In Example 18, which includes the subject matter of Examples 16-17, the one or more third insurance code identifiers correspond to an additional biological condition.

[0246] In Example 19, the subject matter of Examples 16-18 includes determining that one or more third insurance code identifiers are included within a portion of a group of insurance code identifiers that are not included within a subset of the group of insurance code identifiers, determining that a date of use of at least one of the one or more third insurance procedures is the same as a date of use of one of the one or more first insurance procedures, determining that there are no other additional insurance procedures having an insurance code identifier included within the group of code identifiers, and determining that the patient is included within a cohort of patients in which a subtype of the biological condition is present.

[0247] In Example 20, the subject matter of Examples 16-19 includes determining that the one or more third insurance code identifiers are included within a portion of a group of insurance code identifiers that are not included within a subset of the group of insurance code identifiers, determining that a utilization date of at least one of the one or more third insurance procedures precedes a utilization date of one of the one or more first insurance procedures and is within a predefined time period, and determining that the patient is not included within a cohort of patients in which a subtype of the biological condition is present.

[0248] Example 21 is a method implemented in one or more computing machines having processing circuitry and memory, the method including: accessing, in the processing circuitry, one or more medical data tables storing medical insurance transaction data for a plurality of patients, the one or more medical data tables comprising a date column and a diagnosis column; identifying, using the processing circuitry and based on the diagnosis column, a set of patients suffering from a defined biological condition, the set of patients from among the plurality of patients; determining, for each patient in the set of patients, a first date on which the patient was diagnosed with the defined biological condition; identifying, using the processing circuitry and based on the diagnosis column and the date column, a cohort of patients from among the set of patients lacking a diagnosis from a set of biological conditions associated with a date occurring during a predefined time window before the first date on which the patient was diagnosed with the defined biological condition; and providing an output representative of the cohort.

[0249] In Example 22, which includes the subject matter of Example 21, the diagnosis column stores an International Classification of Diseases, Ninth Revision (ICD-9) or International Classification of Diseases, Tenth Revision (ICD-10) code.

[0250] In Example 23, which includes the subject matter of Examples 21-22, the defined biological condition is lung cancer, the set of biological conditions comprises a cancer different from lung cancer, and the predefined time window prior to the initial date is 6 months prior to the initial date.

[0251] In Example 24, which includes the subject matter of Examples 21-23, the defined biological condition is a defined type of cancer, and the method further includes determining the metastatic status of at least one patient from the cohort.

[0252] In Example 25, which includes the subject matter of Example 24, the metastatic status is determined based on secondary malignancies International Classification of Diseases (ICD) codes or Healthcare Common Criteria Coding System (HCPCS) codes.

[0253] In Example 26, which includes the subject matter of Examples 21-25, identifying the cohort includes ordering rows associated with patients in the set by date, and accessing rows associated with a predefined time window to identify patients in the set lacking a diagnosis from the set of biological conditions during the predefined time window.

[0254] Example 27 is at least one machine-readable medium comprising instructions that, when executed by a processing circuitry, cause the processing circuitry to perform operations to implement any of Examples 1-26.

[0255] Example 28 is an apparatus, comprising means for implementing any of the methods described in Examples 1-26.

[0256] Example 29 is a system for implementing any of the methods described in Examples 1-26.

[0257] Example 30 is a method for implementing any of the methods described in Examples 1-26.

[0258] Example 31. A method comprising: obtaining, by a computing system having one or more processors and a memory, health insurance claims data for a number of patients, the health insurance claims data indicating a number of health insurance codes for the number of patients; analyzing, by the computing system, the health insurance claims data to determine a patient cohort of the number of patients having a primary diagnosis corresponding to a biological condition; and determining, by the computing system, an identifier number for each patient included in the cohort, the identifier number uniquely identifying the individual patient in an integrated data repository, the integrated data repository storing genomics data for the number of patients. analyzing, by a computing system, the one or more real world evidence measures for each individual patient included in the cohort of patients, the one or more real world measures indicative of a degree of progression of a biological condition for each individual patient included in the cohort; and analyzing, by the computing system, the one or more real world measures in conjunction with the genomics data for the cohort of patients to determine one or more genomic mutations that correspond to a degree of progression of the biological condition for each individual patient included in the cohort.

[0259] Example 32. The method of Example 31, wherein the real world evidence measures include at least one of the time period between one or more treatments for a biological condition received by one or more first patients included in the patient cohort and the date of death of the one or more first patients, the time period between one or more first treatments received by one or more second patients included in the patient cohort and one or more second treatments received by the one or more second patients, or the time period between one or more treatments received by one or more third patients included in the patient cohort and the date of the last treatment received by the one or more third patients.

[0260] Example 33. The method of example 31 or 32, wherein the health insurance code includes a diagnosis code corresponding to a number of biological conditions.

[0261] Example 34. The method of any one of Examples 31-33, wherein the health insurance codes are stored in a data table including a number of rows corresponding to medical treatments of a number of patients, each medical treatment including a number of health insurance codes corresponding to at least one of medical services, medical procedures, or therapeutic agents provided to each of the number of patients in connection with treatment for one or more biological conditions.

[0262] Example 35. The method of any one of Examples 31-34, comprising determining, by a computing system, one or more candidate treatments for one or more patients included in the cohort based on at least one of the degree of progression of the biological condition for the one or more individual patients or the one or more genomic mutations present in the one or more patients.

[0263] Example 36. The method of any one of Examples 31-35, comprising determining, by a computing system, a cohort of patients by analyzing health insurance claims data according to a cohort identification framework, the cohort identification framework indicating at least one of one or more rules, one or more schemes, or logic to be applied to determine one or more patients to include in the patient cohort.

[0264] Example 37. The method of Example 36, wherein the cohort identification framework indicates one or more first health insurance diagnosis codes corresponding to a first biological condition and one or more second health insurance diagnosis codes corresponding to a second biological condition, the method including analyzing, by a computing system, health insurance claims data to determine patients having health insurance claims records including the one or more first health insurance diagnosis codes to be included in the patient cohort.

[0265] Example 38. The method of Example 37, comprising: determining, by a computing system, that one or more additional health insurance diagnosis codes corresponding to one or more additional biological conditions are not present in the health insurance claims data within a threshold period after a date of the initial health insurance claims data, wherein the initial health insurance claims data includes one or more first health insurance diagnosis codes; determining, by the computing system, that the one or more patients have a primary diagnosis corresponding to the biological condition; and determining, by the computing system, that the one or more patients are included in a cohort of patients.

[0266] Example 39. The method of Example 37, comprising: determining, by a computing system, that one or more additional health insurance diagnosis codes corresponding to the one or more additional biological conditions are present in the health insurance claims data within a threshold period after a date of the initial health insurance claims data, where the initial health insurance claims data includes one or more first health insurance diagnosis codes; determining, by the computing system, that one or more additional patients have a primary diagnosis corresponding to the additional biological condition that does not include the biological condition; and determining, by the computing system, that the one or more additional patients should be excluded from the cohort of patients.

[0267] Example 40. The method of Example 39, wherein the biological condition is a first form of cancer and the additional biological condition is a second form of cancer, and the method includes determining, by the computing system, that one or more additional patients are afflicted with cancer that has metastasized.

[0268] Example 41. The method of any one of Examples 31-40, comprising generating, by a computing system, a diagnostic data table including a plurality of rows, with each row corresponding to an individual patient of a number of patients, and with each row indicating one or more diagnoses of biological conditions for the individual patient associated with the each row.

[0269] Example 42. A system comprising one or more hardware processors, the system, when executed by the one or more hardware processors, comprising: obtaining health insurance claims data for a number of patients, the health insurance claims data indicating a number of health insurance codes for the number of patients; analyzing the health insurance claims data to determine a patient cohort of the number of patients having a primary diagnosis corresponding to a biological condition; and determining an identifier number for each patient included in the cohort, the identifier number uniquely identifying the individual patient in an integrated data repository; and the integrated data repository, when executed by the one or more hardware processors, comprising: 11. A system comprising: a memory storing computer readable instructions to perform operations including storing data; analyzing genomics data for a cohort of patients to determine one or more real world evidence measures for each individual patient included in the cohort of patients, the one or more real world measures indicative of a degree of progression of a biological condition for each individual patient included in the cohort; and analyzing the one or more real world measures in conjunction with the genomics data for the cohort of patients to determine one or more genomic mutations that correspond to a degree of progression of a biological condition for each individual patient included in the cohort.

[0270] Example 43. The system of Example 42, wherein the real world evidence measurements include at least one of the time period between one or more treatments for a biological condition received by one or more first patients included in the patient cohort and the date of death of the one or more first patients, the time period between one or more first treatments received by one or more second patients included in the patient cohort and one or more second treatments received by the one or more second patients, or the time period between one or more treatments received by one or more third patients included in the patient cohort and the date of the last treatment received by the one or more third patients.

[0271] Example 44. The system of example 42 or 43, wherein the health insurance code includes a diagnosis code corresponding to a number of biological conditions.

[0272] Example 45. The system of any one of Examples 42-44, wherein the health insurance codes are stored in a data table including a number of rows corresponding to medical treatments of a number of patients, each medical treatment including a number of health insurance codes corresponding to at least one of a medical service, a medical procedure, or a therapeutic agent provided to an individual of the number of patients in connection with treatment for one or more biological conditions.

[0273] Example 46. The system of any one of Examples 42-45, wherein the memory stores additional computer readable instructions that, when executed by the one or more hardware processors, perform additional operations including determining one or more candidate treatments for one or more patients included in the cohort based on at least one of the degree of progression of a biological condition for the one or more individual patients of the one or more patients or one or more genomic mutations present in the one or more patients.

[0274] Example 47. The system of any one of Examples 42-46, wherein the memory stores additional computer-readable instructions that, when executed by one or more hardware processors, perform additional operations including determining a cohort of patients by analyzing health insurance claims data according to a cohort identification framework, the cohort identification framework indicating at least one of one or more rules, one or more schemes, or logic to be applied to determine one or more patients for inclusion within the patient cohort.

[0275] Example 48. The system of Example 47, wherein the cohort identification framework indicates one or more first health insurance diagnosis codes corresponding to a first biological condition and one or more second health insurance diagnosis codes corresponding to a second biological condition, and the memory stores additional computer readable instructions that, when executed by the one or more hardware processors, perform additional operations including analyzing health insurance claims data and determining patients having health insurance claims records including the one or more first health insurance diagnosis codes for inclusion within the patient cohort.

[0276] Example 49. The system of Example 48, wherein the memory stores additional computer readable instructions that, when executed by the one or more hardware processors, perform additional operations including determining that one or more additional health insurance diagnosis codes corresponding to one or more additional biological conditions are not present in the health insurance claim data within a threshold period after a date of the initial health insurance claim data, where the initial health insurance claim data includes the one or more first health insurance diagnosis codes; determining that the one or more patients have a primary diagnosis corresponding to the biological condition; and determining that the one or more patients are included in a cohort of patients.

[0277] Example 50. The system of Example 48, wherein the memory stores additional computer readable instructions that, when executed by the one or more hardware processors, perform additional operations including determining that one or more additional health insurance diagnosis codes corresponding to the one or more additional biological conditions are present in the health insurance claims data within a threshold period after a date of the initial health insurance claims data, where the initial health insurance claims data includes the one or more first health insurance diagnosis codes; determining that one or more additional patients have a primary diagnosis corresponding to the additional biological condition that does not include the biological condition; and determining that the one or more additional patients should be excluded from the patient cohort.

[0278] Example 51. The system of Example 50, wherein the biological condition is a first form of cancer and the additional biological condition is a second form of cancer, and the memory stores additional computer readable instructions that, when executed by the one or more hardware processors, perform additional operations including determining that one or more additional patients have cancer that has metastasized.

[0279] Example 52. The system of any one of Examples 42-51, wherein the memory stores additional computer readable instructions that, when executed by the one or more hardware processors, perform additional operations including generating a diagnostic data table including a plurality of rows, with each row corresponding to an individual patient of a number of patients, and with each row indicating one or more diagnoses of biological conditions for the individual patient associated with the each row.

[0280] As used herein, a component may refer to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that provide partitioning or modularization of certain processing or control functions. Components may be combined with other components through their interfaces to perform machine processes. A component may be a packaged functional hardware unit designed for use with other components, and may be a part of a program that typically performs a specific function of the associated functionality. A component may constitute either a software component (e.g., code embodied on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit capable of performing an operation, and may be configured or arranged in a physical manner. In various exemplary implementations, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware component that operates to perform certain operations as described herein.

[0281] It should be understood that the individual steps used in the methods of the present teachings can be performed in any order and / or simultaneously so long as the present teachings remain operable. Further, it should be understood that the apparatus and methods of the present teachings can include any number or all of the described implementations so long as the present teachings remain operable.

[0282] The various steps of the methods disclosed herein or steps performed by the systems disclosed herein may be performed at the same or different times and / or in the same geographic location or different geographic locations, e.g., countries. The various steps of the methods disclosed herein can be performed by the same person or different people.

[0283] Various implementations of systems, devices, and methods are described herein. These implementations are provided as examples only and are not intended to limit the scope of the claimed invention. It should also be understood that various features of the described implementations can be combined in various ways to produce numerous additional implementations. Also, while various materials, dimensions, shapes, configurations, locations, etc. are described for use with the disclosed implementations, others than those disclosed may be utilized without departing from the scope of the claimed invention.

[0284] Those skilled in the art will recognize that implementations may comprise fewer features than those illustrated in any individual implementation described above. The implementations described herein are not meant to be an exhaustive listing of ways in which various features may be combined. Thus, implementations are not mutually exclusive combinations of features; rather, implementations may comprise combinations of different individual features selected from different individual implementations, as would be understood by one skilled in the art. Also, elements described with respect to one implementation may be implemented in other implementations, even when not described in such implementation, unless otherwise stated. Although a dependent claim may refer to a specific combination with one or more other claims in the claim, other implementations may also include a combination of the dependent claim with the subject matter of each other dependent claim or a combination of one or more features with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a specific combination is not intended. Moreover, it is also intended that features of a claim be included in any other independent claim, even if the claim is not made directly dependent on the independent claim.

[0285] Also, references herein to "one implementation," "an implementation," or "some implementations" mean that a particular feature, structure, or characteristic described in connection with an implementation is included in at least one implementation of the present teachings. The appearances of the phrase "in one implementation" in various places in the specification do not necessarily all refer to the same implementation.

[0286] The incorporation by reference of any of the above documents is limited such that it does not incorporate any subject matter contrary to the express disclosure of this specification. The incorporation by reference of any of the above documents is further limited such that it does not incorporate by reference any claims contained herein. The incorporation by reference of any of the above documents is further limited such that it does not incorporate by reference any definitions provided herein, unless expressly included herein.

[0287] Although the implementations have been described with reference to specific exemplary implementations, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the present disclosure. Thus, the specification and drawings are to be regarded in an illustrative and not a restrictive sense. The accompanying drawings, which form a part of this specification, show specific implementations in which the subject matter may be practiced, by way of example, and not by way of limitation. The illustrated implementations are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other implementations may be utilized and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. This detailed description is therefore not to be construed in a limiting sense, and the scope of various implementations is defined solely by the appended claims, along with the full scope of equivalents to which such claims are entitled.

[0288] Although specific implementations are illustrated and described herein, it should be understood that any arrangement calculated to achieve the same purpose may be substituted for the specific implementations shown. The present disclosure is intended to cover any and all adaptations or variations of the various implementations. Combinations of the above implementations and other implementations not specifically described herein will be apparent to those of skill in the art upon review of the above description.

[0289] The terms "a" or "an" are used herein to include "one or more than one," as is common in patent documents, independent of any other instance or usage of "at least one" or "one or more." The term "or" is used herein to refer to a non-exclusive or, such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise indicated. The terms "including" and "in which" are used herein as the plain English equivalents of the respective terms "comprising" and "wherein." Also, in the following claims, the terms "including" and "comprising" are open-ended, i.e., a system, user equipment (UE), article, composition, formulation, or process that includes elements in addition to those recited after such term in a claim, will still be considered to fall within the scope of that claim. Also, in the following claims, the terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements on their objects.

Claims

1. 1. A method implemented in one or more computing machines having processing circuitry and memory, the method comprising: accessing, in the processing circuitry, one or more medical data repositories storing data regarding a given patient from a plurality of patients, the one or more medical data repositories storing pharmacy data, clinic visit data, and health insurance transaction data; identifying one or more biological conditions and metastatic status for the given patient based on one or more disease codes in the clinic encounter data or the health insurance transaction data; identifying one or more courses of treatment for the given patient based on one or more drug codes in the pharmacy data; identifying one or more medical procedures received by the given patient based on one or more insurance codes in the medical insurance procedure data; determining a primary diagnostic biological condition for the given patient based on a combination of the one or more biological conditions, the metastatic status, the one or more courses of treatment, and the one or more medical procedures; assigning the given patient to a cohort of patients based on the primary diagnostic biological condition; providing an output representative of the assigned cohort for the given patient; A method comprising:

2. Determining the primary diagnostic biological condition comprises: creating a master gap table indicating a time interval between two consecutive procedures for the given patient having the same procedure name and insurance code based on the medical insurance procedure data, the master gap table comprising columns for procedure name, insurance code, unit, and gap length; Calculating a median gap table indicating a median gap for each combination of treatment name and insurance code based on the master gap table, the median gap table having columns for treatment name, insurance code, unit, and gap length; determining the primary diagnostic biological condition based, at least in part, on the data in the median gap table; and The method of claim 1 , comprising:

3. the processing circuitry comprises a plurality of multi-threaded graphics processing units (GPUs), and the method further comprises: determining, in parallel and using parallel threads of the plurality of multi-threaded GPUs, assigned cohorts for a plurality of patients from the plurality of patients, including the given patient; The method of claim 1.

4. the disease code comprises an International Classification of Diseases (ICD) code; the drug code comprises a National Drug Code (NDC) code; the insurance code comprises a Healthcare Common Procedure Coding System (HCPCS) code; The method of claim 1.

5. 5. The method of claim 4, wherein identifying the one or more biological conditions for the given patient based on the one or more disease codes in the clinic encounter data or the health insurance transaction data comprises identifying the given patient as having lung cancer based on an ICD code associated with lung cancer.

6. 6. The method of claim 5, wherein identifying the metastatic status for the given patient based on the one or more disease codes in the clinic encounter data or the health insurance transaction data is based on secondary malignancy ICD codes or HCPCS codes.

7. 10. The method of claim 1, wherein the primary diagnostic biological condition for the given patient is determined based on disease codes, drug codes, or insurance codes associated with dates within a predefined date range.

8. Allocating the given patient to the cohort comprises: analyzing one or more data tables containing medical insurance transaction data for the given patient, the one or more data tables being from the one or more medical data repositories; determining one or more first insurance procedures including one or more first code identifiers contained within the one or more data tables, the one or more first code identifiers corresponding to a diagnosis of the patient for the one or more biological conditions; determining one or more second insurance procedures including one or more second code identifiers contained within the one or more data tables, the one or more second code identifiers corresponding to medical procedures performed on the patient, the medical procedures being from the clinic encounter data; generating a medical header table, the medical header table including a first number of columns for storing the one or more first code identifiers, a second number of columns for storing the one or more second code identifiers, and a plurality of rows, each of which corresponds to a first medical insurance procedure of the one or more first insurance procedures or a second medical insurance procedure of the one or more second insurance procedures; storing the medical header table in the one or more medical data repositories; determining the cohort for the patient based on data in the medical header table; The method of claim 1 , comprising:

9. Each first insurance procedure of the one or more first insurance procedures indicates a date of use of the each first insurance procedure; Each of the one or more second insurance procedures indicates a date of use of the each second insurance procedure; The method of claim 8.

10. 10. The method of claim 9, further comprising arranging the rows of the medical header table in ascending order based on the utilization date, such that a claim with a first utilization date is the first row in the medical header table and a procedure with a most recent utilization date is the last row in the medical header table.

11. 10. The method of claim 8, further comprising analyzing the one or more first code identifiers and determining that a first code identifier of the one or more first code identifiers is included within a group of insurance code identifiers corresponding to the one or more biological conditions.

12. the first code identifiers are arranged according to a first format corresponding to a first classification of insurance code identifiers; the first classification of the insurance code identifier corresponds to the International Classification of Diseases, Ninth Revision (ICD-9); The method of claim 11.

13. the first code identifiers are arranged according to a second format corresponding to a second classification of insurance code identifiers; The second classification of the insurance code identifier corresponds to the International Classification of Diseases, Tenth Revision (ICD-10). The method of claim 11.

14. the one or more biological conditions include multiple subtypes; each subtype of the plurality of subtypes corresponds to a subset of the group of insurance code identifiers corresponding to the one or more biological conditions; The method further includes determining that the first code identifier is included in a first subset of the group of insurance codes corresponding to a first subtype of the biological condition. The method of claim 11.

15. 12. The method of claim 11, wherein the biological condition is cancer and the plurality of subtypes includes at least one of lung cancer, breast cancer, or colorectal cancer.

16. determining one or more third insurance procedures having use dates within a predefined period ending with the first use date; analyzing one or more third insurance code identifiers of the third insurance procedure against the first code identifier and the second code identifier; 16. The method of claim 15, further comprising:

17. determining that the one or more third insurance code identifiers are not included within the first code identifier and the second code identifier; determining, based on the third insurance code identifier, that the patient is included within a cohort of patients in which a given subtype of the one or more biological conditions is present; and 17. The method of claim 16, further comprising:

18. The method of claim 16 , wherein the one or more third insurance code identifiers correspond to an additional biological condition.

19. determining that the one or more third insurance code identifiers are included within a portion of the group of insurance code identifiers that is not included within the subset of the group of insurance code identifiers; Determining that a use date of at least one of the one or more third insurance procedures is the same as a use date of one of the one or more first insurance procedures; determining that there are no other additional insurance procedures having insurance code identifiers included within the group of code identifiers; determining that the patient is included within a cohort of patients in which a subtype of the biological condition is present; 17. The method of claim 16, further comprising:

20. determining that the one or more third insurance code identifiers are included within a portion of the group of insurance code identifiers that is not included within the subset of the group of insurance code identifiers; determining that a utilization date of at least one of the one or more third insurance claims precedes a utilization date of one of the one or more first insurance procedures and is within the predefined time period; determining that the patient is not included within a cohort of patients in which the subtype of the biological condition is present; 17. The method of claim 16, further comprising:

21. 1. A method implemented in one or more computing machines having processing circuitry and memory, the method comprising: accessing, in the processing circuitry, one or more medical data tables storing medical insurance transaction data for a plurality of patients, the one or more medical data tables comprising a date column and a diagnosis column; identifying a set of patients having a defined biological condition using the processing circuitry and based on the diagnostic sequence, the set of patients being from among the plurality of patients; and For each patient in the set of patients, determining the earliest date on which the patient was diagnosed with the defined biological condition; using the processing circuitry and based on the diagnosis sequence and the date sequence, identifying a cohort of patients from among the set of patients, the cohort of patients lacking a diagnosis from a set of biological conditions associated with a date occurring during a predefined time window before the first date on which the patient received the diagnosis of the specified biological condition; providing an output representative of said cohort; A method comprising:

22. 22. The method of claim 21, wherein the diagnosis column stores International Classification of Diseases, Ninth Revision (ICD-9) or International Classification of Diseases, Tenth Revision (ICD-10) codes.

23. 22. The method of claim 21 , wherein the defined biological condition is lung cancer, the set of biological conditions comprises cancers other than lung cancer, and the predefined time window before the initial date is six months before the initial date.

24. The defined biological condition is a defined type of cancer, and the method further comprises: determining the metastatic status of at least one patient from said cohort; 22. The method of claim 21.

25. 25. The method of claim 24, wherein the metastatic status is determined based on secondary malignancy International Classification of Diseases (ICD) codes or Healthcare Common Criteria Coding System (HCPCS) codes.

26. Identifying the cohort comprises: arranging rows associated with patients in said set by date; accessing rows associated with the predefined time window and identifying patients within the set that lack the diagnosis from the set of biological conditions during the predefined time window; 22. The method of claim 21, comprising:

27. 1. A system comprising: processing circuitry; a memory storing instructions that, when executed by the processing circuitry, cause the processing circuitry to: accessing, in the processing circuitry, one or more medical data repositories storing data regarding a given patient from a plurality of patients, the one or more medical data repositories storing pharmacy data, clinic visit data, and health insurance transaction data; identifying one or more biological conditions and metastatic status for the given patient based on one or more disease codes in the clinic encounter data or the health insurance transaction data; identifying one or more courses of treatment for the given patient based on one or more drug codes in the pharmacy data; identifying one or more medical procedures received by the given patient based on one or more insurance codes in the medical insurance procedure data; determining a primary diagnostic biological condition for the given patient based on a combination of the one or more biological conditions, the metastatic status, the one or more courses of treatment, and the one or more medical procedures; assigning the given patient to a cohort of patients based on the primary diagnostic biological condition; providing an output representative of the assigned cohort for the given patient; a memory and A system comprising: