Computer architecture for generating reference data tables

JP2024536149A5Pending Publication Date: 2025-10-02GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024519319
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-31
Filing Date
2022-09-30
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing healthcare systems face inefficiencies and inaccuracies in analyzing unstructured medical data, particularly health insurance claims data, which is crucial for determining treatments and diagnoses, leading to potential detrimental health effects due to imprecise analysis.

Method used

An integrated data repository system that combines structured health insurance claims data with molecular data, using algorithms to generate reference data tables that accurately analyze treatment effectiveness by correlating insurance code identifiers with treatment information, reducing computational resources and time required for analysis.

Benefits of technology

The system provides precise and efficient analysis of healthcare data, enabling confidential and anonymous insights into treatment effectiveness and patient outcomes, improving the accuracy of treatment recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An integrated data repository may be generated that includes genomic information and health insurance claims data information for a common group of individuals. The health insurance claims data may be analyzed to determine insurance code identifiers that correspond to treatments provided to individuals in whom the biological condition exists. The insurance code identifiers may be used to query one or more databases to obtain additional information regarding the treatments. A treatment reference table may be generated by using the insurance claims data and the additional information obtained from the one or more databases. The treatment reference table may also be used to identify populations of individuals who have received one or more treatments related to the biological condition and to determine one or more characteristics of the individuals included in the population.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims the benefit of PCT Application No. PCT / US2022 / 032250, entitled "Computer Architecture for Generating an Integrated Data Repository," filed June 3, 2022, PCT Application No. PCT / US2022 / 038941, entitled "Computer Architecture for Identifying Lines of Therapy," filed July 29, 2022, and U.S. Provisional Application No. 63 / 250,912, filed September 30, 2021, to PCT Application No. PCT / US2022 / 042262, entitled "Data Repository, System and Method for Cohort Selection," filed August 31, 2022, each of which is incorporated by reference in its entirety herein.

[0002] Implementations of the present disclosure relate generally to the field of computer architecture, and specifically to implementations of a computer architecture for generating a reference data table indicating an identifier for a treatment to be provided to a patient in which one or more biological conditions may be present. [Background technology]

[0003] When an individual visits a health care provider for treatment of one or more biological conditions, various types of documentation may be generated. For example, a medical record may be generated by the health care provider that includes clinical observations recorded by the health care provider, laboratory test results, diagnostic test information, imaging information, dental hygiene information, one or more combinations thereof, and the like. In addition, billing records may be generated that indicate payment information for at least one product or service provided by the health care provider to the individual. Furthermore, health insurance claims information may be generated that indicates information obtained by a health insurance company related to the treatment of the individual for one or more biological conditions. Summary of the Invention [Means for solving the problem]

[0004] The following description and the accompanying drawings sufficiently illustrate specific implementations to enable those skilled in the art to practice them. Other implementations may incorporate structural, logical, electrical process and other changes. Some portions and features of some implementations may be included in or substituted for some portions and features of other implementations. Implementations set forth in the claims encompass all available equivalents of those claims.

[0005] Analysis of healthcare data using existing systems and technologies is typically performed on medical records generated by healthcare providers. As used herein, a healthcare provider may refer to an entity, individual, or group of individuals involved in providing care to an individual related to at least one of the treatment or prevention of one or more biological conditions. In addition, as used herein, a biological condition may refer to an abnormality of function and / or structure in an individual to such an extent that it produces or threatens to produce a detectable feature of abnormality. A biological condition may be characterized by external and / or internal features, signs and / or symptoms that indicate a deviation from a biological norm in one or more populations. A biological condition may include at least one of one or more diseases, one or more disorders, one or more injuries, one or more syndromes, one or more disorders, one or more infections, one or more isolated symptoms, or other atypical deviations of an individual's biological structure and / or function. In addition, as used herein, a treatment may refer to a substance, procedure, routine, device, and / or other intervention that may be administered or performed with the intent of alleviating one or more effects of a biological condition in an individual. In one or more examples, the treatment may include a substance that is metabolized by the individual. The substance may include a composition, such as a pharmaceutical composition. The substance may be delivered to the individual through a number of methods, such as ingestion, injection, absorption, or inhalation. The treatment may also include a physical intervention, such as one or more surgeries.

[0006] Healthcare data typically analyzed by existing systems includes unstructured data. Unstructured data may include data that is not organized according to a predefined or standardized format. For example, unstructured data consisting of free text may include notes created by a healthcare provider. That is, the manner in which notes are taken does not include predefined inputs that are selectable by the healthcare provider (such as via a drop-down menu or via a list). Rather, the notes include text entered by the healthcare provider (which may include sentences, sentence fragments, words, characters, symbols, abbreviations, and combinations of one or more of these, etc.). In some cases, unstructured data may be partially structured. For example, a provider may select a billing code from a predefined list of billing codes and add an unstructured note to the data associated with that billing code.

[0007] Existing systems typically dedicate a large amount of computing resources to analyzing unstructured data to extract information that may be relevant to the analysis performed by the existing system. In some cases, the existing system may analyze and convert the unstructured data into a structured format to facilitate the analysis of the unstructured data. The analysis of unstructured data by the existing system may not only be inefficient but also inaccurate. In a scenario where the unstructured data is obtained from healthcare data, the importance of accurately analyzing the information is high because the analysis may relate to at least one of the treatment or diagnosis of many individuals with respect to one or more biological conditions. Thus, inaccurate analysis of healthcare data may have adverse effects on the health of individuals.

[0008] Implementations of the techniques, architectures, frameworks, systems, processes, and computer readable instructions described herein are directed to analyzing health insurance claims data to derive information regarding at least one of an individual's health or treatment. In contrast to existing systems, health insurance claims data is structured according to one or more formats and stored in many data tables. The data tables may include codes or other alphanumeric information indicating the treatment the individual received, dates of treatment, medication information, the individual's diagnosis of one or more biological conditions, information related to visits to a health care provider, dates of visits to a health care provider, billing information, and the like. Implementations described herein may be used to precisely analyze health insurance claims data for hundreds to thousands of individuals in which one or more biological conditions exist. In various examples, tens of thousands, hundreds of thousands, and even millions of rows and / or columns of health insurance claims data may be analyzed to determine health-related information for individuals in which one or more biological conditions exist.

[0009] In various examples, implementations described herein may integrate molecular data and health insurance claims data. The molecular data may include information derived from tissue samples extracted from many individuals. In addition, the molecular data may include information derived from blood samples extracted from many individuals. In one or more examples, the molecular data may include genomic data. Furthermore, in one or more examples, the health insurance claims data may be integrated with germline genetic information of many individuals.

[0010] An integrated data repository may be generated that combines the individual's health insurance claims data and the individual's molecular data. In one or more examples, an identifier may be generated for the individual that is associated with both the individual's health insurance claims data and the individual's molecular data. Both the molecular data and the health insurance claims data stored by the integrated data repository may be accessible by using a unique identifier for the individual. In one or more examples, the individual's identifier may include an encrypted security key. In various examples, the integrated data repository may include a number of data tables that correspond to various aspects of the data stored in the data repository. For example, a first data table may be generated that includes summary data (e.g., personal information) for the individual included in the integrated data repository, and a second data table may be generated that includes data corresponding to visits to a healthcare provider. In addition, a third data table may be generated that indicates medical procedures provided to the individual, and a fourth data table may be generated that indicates information related to prescriptions obtained by the individual. Furthermore, a fifth data table may be generated that includes molecular information for the individual.

[0011] The data tables included in the integrated data repository may be connected via logical links. In this way, a query to retrieve information from one data table may result in retrieval of information from one or more additional data tables. Information stored by the linked data tables may be accessed to generate many different data sets that may be used to analyze the information stored by the integrated data repository. In one or more examples, the information stored by the integrated data repository may be analyzed to extract biological meanings regarding a patient or group of patients. Additionally, the information stored by the integrated data repository may be analyzed to determine an individual's biological status. The biological status may correspond to determining whether a biological condition exists for a patient or group of patients.

[0012] In one or more alternative examples, the information stored by the integrated data repository may be analyzed by one or more algorithms to generate a dataset organized according to one or more schemes. The dataset may indicate the treatments received by an individual over a period of time for a biological condition. The dataset may also indicate a group of individuals included in the integrated data repository that have many common characteristics. In various examples, the dataset may consolidate and collate information from many different data sources, including the integrated data repository. The dataset may be analyzed for many queries to indicate information that may be of interest to at least one of a health care provider, a patient, or a provider of treatment for a biological condition. For example, one or more datasets may be integrated and analyzed to determine the survival rate of an individual (having a prescribed genomic profile in response to the presence of a biological condition and receiving a prescribed treatment).

[0013] The implementations described herein may provide a platform for integrating molecular data with an individual's health insurance claims data that is not found in existing systems that typically rely on electronic medical records that contain a certain amount of unstructured data. By generating and analyzing structured health insurance claims data integrated with molecular data, the implementations described herein may provide a more accurate characterization of the integrated data to existing systems that rely on relatively imprecise and unstructured electronic medical record data. In addition, the implementations described herein generate analytics ready data sets that enable analysis of health information about an individual in a private and anonymous manner.

[0014] In addition, the claims data may include information that may be used to determine at least one of one or more biological conditions present in an individual, treatments provided to an individual for which the one or more biological conditions exist, a timeline of treatments provided to the individual, or modifications to a biological condition present in an individual. However, claims data in its raw form may be difficult to interpret. Furthermore, a large amount of claims data may be generated for an individual over a period of time related to one or more treatments provided to the individual based on a biological condition present in the individual. In one or more examples, hundreds or even thousands of claims may be generated over a period of time based on one or more treatments provided to the individual for a biological condition. Also, it may be challenging to correlate one claim data of an individual with another claim data of the individual for a given biological condition, since the claims data is typically a series of codes that do not provide any indication as to how the codes relate to each other or what the codes themselves mean. Thus, using claims data to determine tangible insights, trends, correlations, etc. related to treatments received by an individual for one or more biological conditions is not straightforward and thus involves analysis of a large amount of disparate data. In situations where insurance claims data for multiple individuals is analyzed, the complexity, time and computing resources used to determine relationships, correlations, insights, etc. regarding an individual's treatment increases significantly, and even exponentially.

[0015] The techniques, systems, architectures, frameworks, processes, and methods described herein are also directed to generating a reference data table that can be used to analyze insurance claims data. In one or more examples, the reference data table can indicate information regarding one or more treatments provided to an individual in whom a biological condition exists. For example, the reference data table can include information regarding one or more medications provided to an individual in whom a biological condition exists. In various examples, the reference data table can indicate for a given insurance code identifier, a name of the treatment, a class of the treatment, one or more compositions contained in the treatment, a source of information regarding the treatment, one or more additional identifiers of the treatment, one or more combinations thereof, and the like.

[0016] In one or more implementations, the insurance code identifier may be extracted from one or more data tables that include a number of insurance code identifiers related to one or more treatments of the individual in which the biological condition exists. The insurance code identifier may be analyzed for one or more criteria related to at least one format of the insurance code identifier. For example, the insurance code identifier may be analyzed to determine whether the insurance code identifier corresponds to a National Drug Code (NDC) format. In one or more instances, the insurance code identifier may be analyzed to determine whether the insurance code identifier corresponds to an NDC9 format, an NDC10 format, an NDC11 format, or another NDC format. In one or more implementations, an application programming interface (API) request may be generated that queries the data repository to determine an updated insurance code identifier that corresponds to an initial version of the insurance code identifier. In at least some instances, a response to the API request may include an NDC code having a format that differs from the format of the NDC code used to generate the API request. For example, an API request may be generated and sent to the data repository management system that may include a modified version of the insurance code identifier that originally had an NDC9 format or an NDC10 format. In these scenarios, the response to the request may be a data file that includes a version of the insurance code identifier that corresponds to the NDC11 format. The insurance code identifier that corresponds to the NDC11 format may be used to generate an additional API request to retrieve information from one or more additional data sets stored by the data repository management system. The response to the additional API request may include at least one of an identifier of the procedure in the database management system, a name of the procedure or procedures, information related to the one or more classes of procedures, or information related to one or more sources of the one or more classes of procedures.

[0017] The reference data table may be generated by using at least a portion of information obtained from responses to API requests that retrieve information from one or more datasets. In one or more examples, rows may be generated in the reference data table for individual procedures having a valid insurance code identifier. Columns of the reference table may be populated by using information corresponding to the individual procedures (obtained from the API requests to retrieve data contained in one or more datasets). For example, the reference data table may include one or more columns indicating at least one of an original NDC code obtained from insurance claims data, an identifier of the procedure corresponding to the original NDC code, a name of the procedure corresponding to the original NDC code, or a class of procedure corresponding to the original NDC code.

[0018] In one or more instances, the analysis may be performed to determine the effectiveness of a treatment for an individual having one or more defined genetic mutations. To determine the effectiveness of the treatment, insurance claims data indicating the treatment provided to the individual may be analyzed. A lookup table may be used to determine the treatment provided to the individual by analyzing the insurance code identifier contained in the lookup table and the insurance code identifier contained in the insurance claim data related to the treatment. A population of individuals who will receive the treatment may then be identified by using the lookup table to determine the insurance code identifier corresponding to the treatment, and then identifying individuals who have insurance claim data that includes the insurance code identifier. Additional analyses may then be performed on the population of individuals to determine the effectiveness of the treatment provided to the individuals based on a number of additional criteria. For example, genetic material contained in at least one of the blood or tissue samples may be analyzed to determine whether individuals exhibiting a particular tumor who are treated are experiencing an increase or decrease in tumor size.

[0019] As a result, the reference data table may include a row for each NDC code included in the claims data and point to information related to the treatment that corresponds to each NDC code. In this manner, the claims data may be analyzed for treatments provided to individuals included in the claims data. Without the reference data table, each time the claims data is to be analyzed for treatments obtained by individuals, a data set containing the treatment information would be queried and analyzed for appropriateness. This is a time-consuming, computationally resource intensive, and inefficient endeavor. Thus, creating a reference data table that points to insurance code identifiers that correspond to each treatment and are easily accessible reduces the number of computational resources and time used to perform context-related analyses without the reference data table (correlation between each insurance code identifier and each treatment is determined for each analysis). [Brief description of the drawings]

[0020] [Figure 1] FIG. 1 illustrates an exemplary architecture for generating an integrated data repository containing multiple types of healthcare data, according to one or more implementations, and for generating a reference data table indicating identifiers of treatments to be provided to patients in whom one or more biological conditions may be present. [Diagram 2] 1 illustrates an exemplary framework that supports the arrangement of data tables within a unified data repository, according to one or more implementations. [Diagram 3] 1 illustrates an architecture for generating one or more data sets from information retrieved from a data repository that integrates health-related data from many sources, according to one or more implementations. [Figure 4] 1 illustrates an architecture for generating an integrated data repository containing de-identified health insurance claims data and de-identified genomic data, according to one or more implementations. [Diagram 5] 1 illustrates a framework for generating a data set by a data pipeline system based on data stored by an integrated data repository, in accordance with one or more implementations. [Figure 6] 1 illustrates an architecture for generating a look-up data table indicating an identifier for a treatment to be provided to a patient in which one or more biological conditions may be present, according to one or more implementations. [Figure 7] 1 is a flow diagram of an example process for generating a treatment reference table containing information regarding treatments provided to a patient in which one or more biological conditions may be present, according to one or more implementations. [Figure 8] 1 is a flow diagram of an example process for determining a drug identifier that corresponds to an insurance code identifier by using one or more application programming interface (API) requests, according to one or more implementations. [Figure 9] 1 is a flow diagram of an example process for determining a class corresponding to a treatment identifier and for including information related to the class in a reference data table that includes the treatment identifier, according to one or more implementations. [Figure 10] A diagrammatic representation of a machine in the form of a computer system in which a set of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein in accordance with one or more implementations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0021] 1 illustrates an example architecture 100 for generating a unified data repository including multiple types of healthcare data, according to one or more implementations. The architecture 100 may include a data integration & analysis system 102. The data integration & analysis system 102 may obtain data from many data sources and ingest the data from the data sources into a unified data repository 104. For example, the data integration & analysis system 102 may obtain data from a health insurance claims data repository 106. In various examples, the data integration & analysis system 102 and the health insurance claims data repository 106 may be generated and maintained by different entities. In one or more additional examples, the data integration & analysis system 102 and the health insurance claims data repository 106 may be generated and maintained by the same entity.

[0022] The data integration & analysis system 102 may be implemented by one or more computing devices. The one or more computing devices may include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or a combination thereof. In some implementations, at least a portion of the one or more computing devices may be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices may be implemented in a cloud computing architecture. In scenarios where the computing system used to implement the data integration & analysis system 102 is configured in a distributed computing architecture, processing operations may be performed by multiple virtual machines simultaneously. In various examples, the data integration & analysis system 102 may implement multithreading technology. The implementation of distributed computing architectures and multithreading technology allows the data integration & analysis system 102 to utilize fewer computing resources compared to computing architectures that do not implement these technologies.

[0023] The health insurance claims data repository 106 may store information obtained from one or more health insurance companies corresponding to claims made by subscribers of one or more health insurance companies. The health insurance claims data repository 106 may be arranged (e.g., sorted) by patient identifier. The patient identifier may be based on the patient's first name, last name, date of birth, social security number, address, employer, etc. The data stored by the health insurance claims data repository 106 may include structured data arranged in one or more data tables. The one or more data tables storing the structured data may include a number of rows and columns that indicate information regarding health insurance claims made by subscribers of one or more health insurance companies related to treatments and / or care received by the subscribers from a healthcare provider. At least some of the rows and columns of the data tables stored by the health insurance claims data repository 106 may include health insurance codes, which may indicate diagnoses of biological conditions and treatments and / or care obtained by the subscribers of the one or more health insurance companies. In various examples, the health insurance code may also indicate a diagnostic procedure obtained by the individual that is related to one or more biological conditions that may exist within the individual. In one or more examples, the diagnostic procedure may provide information used in detecting the presence of the biological condition. The diagnostic procedure may also provide information used to determine the progression of the biological condition. In one or more instances, the diagnostic procedure may include one or more imaging procedures, one or more analyses, one or more laboratory procedures, one or more combinations thereof, and so forth.

[0024] The data integration & analysis system 102 may also obtain information from a molecular data repository 108. The molecular data repository 108 may store many individuals' data related to genomic, genetic, metabolic, transcriptomic, fragmentomic, immune receptor, methylation, epigenomic, and / or proteomic information, immunohistochemistry (IHC) and immunofluorescence (IF). In one or more examples, the data integration & analysis system 102 and the molecular data repository 108 may be generated and maintained by different entities. In one or more additional examples, the data integration & analysis system 102 and the molecular data repository 108 may be generated and maintained by the same entity. As used herein, "fragmentomic information" may include, among other things, information related to analysis of DNA or RNA fragment lengths to determine the presence or absence of a tumor and to determine tumor characteristics. In one or more instances, the fragmentomics information may correspond to nucleosomal structures and transcription factor binding sites.

[0025] The genomic information may indicate one or more mutations corresponding to the individual's genes. The mutations to the individual's genes may correspond to differences between the individual's nucleic acid sequence and one or more reference genomes. The reference genome may include a known reference genome, such as hgl9. In various examples, the mutations to the individual's genes may correspond to differences in the individual's germline genes relative to the reference genome. In one or more additional examples, the reference genome may include the individual's germline genome. In one or more other examples, the mutations to the individual's genes may include somatic mutations. The mutations to the individual's genes may relate to insertions, deletions, single nucleotide changes, loss of heterozygosity, duplications, amplifications, translocations, fusion genes, or one or more combinations thereof. In at least some examples, the genomic information may correspond to non-coding regions of the genome. The non-coding regions may relate to the regulation of one or more genes. In one or more examples, analysis of the non-coding regions may detect one or more epigenetic signatures of one or more patients.

[0026] In one or more instances, the genomic information stored by the molecular data repository 108 may include a genomic profile of tumor cells present in the individual. In these circumstances, the genomic information may be derived from an analysis of genetic material, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA), found in the individual's blood sample that is present due to the degradation of tumor cells present in the individual. In one or more instances, the genomic information of the individual's tumor cells may correspond to one or more target sites. One or more mutations present with respect to the one or more target sites may indicate the presence of tumor cells in the individual. The genomic information stored by the molecular data repository 108 may be generated in association with an analysis or other diagnostic test that may determine one or more mutations with respect to one or more target sites of a reference genome.

[0027] In one or more instances, the genetic material may be derived from a specimen (including, but not limited to, a tissue specimen or tumor biopsy specimen, circulating tumor cells (CTCs), exosomes or efferosomes) or from circulating nucleic acids. In various instances, circulating nucleic acids may be referred to herein as "cell-free DNA." "Cell-free DNA," "cfDNA molecules," or simply "cfDNA" includes DNA molecules that occur within a subject in an extracellular form (e.g., within blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or saliva) and includes DNA that is not contained within or otherwise bound to a cell at the time it is isolated from the subject. DNA was originally present within a cell or group of cells of a large complex biological organism (e.g., a mammal) or other cells that colonize the organism (e.g., bacteria), but the DNA has been released from the cells into the fluids found within the biological organism. cfDNA includes, but is not limited to, cell-free genomic DNA of a subject (e.g., genomic DNA of a human subject) and cell-free DNA of microorganisms, such as bacteria (pathogenic bacteria or bacteria commonly found in colonized locations, such as the internal organs or skin of healthy controls) that infest the subject, but does not include cell-free DNA of microorganisms that have contaminated a sample of bodily fluid only. Typically, cfDNA can be obtained by obtaining a sample of a fluid without the need to perform an in vitro cell lysis step, and also includes removal of cells present within the fluid (e.g., centrifugation of blood to remove cells).

[0028] In one or more additional examples, the data integration & analysis system 102 may obtain information from one or more additional data repositories 110. The one or more additional data repositories 110 may store data related to an individual's electronic medical record (wherein the data is present in at least one of the health insurance claims data repository 106 or the molecular data repository 108). Additionally, the one or more additional data repositories 110 may store data related to an individual's pathology report (wherein the data is present in at least one of the health insurance claims data repository 106 or the molecular data repository 108). In various examples, the one or more additional data repositories 110 may store data related to a biological condition and / or a treatment of a biological condition. In one or more examples, the data integration & analysis system 102 and the one or more additional data repositories 110 may be at least partially generated and maintained by various entities. In one or more alternative examples, the data integration and analysis system 102 and the one or more additional data repositories 110 may be created and maintained, at least in part, by the same entity.

[0029] In one or more alternative implementations, the data integration & analysis system 102 may obtain information from one or more reference information data repositories 112. The one or more reference information data repositories 112 may store information including definitions, standards, protocols, glossaries, one or more combinations thereof, and the like. In various examples, the information stored by the one or more reference information data repositories may correspond to biological conditions and / or treatments of biological conditions. In one or more examples, the one or more reference information data repositories 112 may include RxNorm, which provides standardized names for clinical drugs and links the standardized names to many of the drug terms used within pharmacy management and drug interaction software. In one or more examples, the data integration & analysis system 102 and at least a portion of the one or more reference information data repositories 112 may be generated and maintained by various entities. In one or more examples, the data integration and analysis system 102 and at least a portion of the one or more reference information data repositories 112 may be created and maintained by the same entity.

[0030] The data integration & analysis system 102 may obtain data from at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112 via one or more communication networks accessible to the data integration & analysis system 102 and accessible to at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112. The data integration & analysis system 102 may also obtain data from at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112 via one or more secure communication channels. Additionally, the data integration and analysis system 102 may obtain data from at least one of the health insurance claims data repository 106, the molecular data repository 108, one or more additional data repositories 110, or the reference information data repository 112 via one or more application programming interface (API) calls.

[0031] The data integration & analysis system 102 may include a data integration system 114. The data integration system 114 may obtain data from the health insurance claims data repository 106 and the molecular data repository 108 to generate the integrated data repository 104. The data integration system 114 may also obtain data from one or more additional data repositories 110 to generate the integrated data repository 104. In various examples, the data integration system 114 may implement one or more natural language processing techniques to integrate data from the one or more additional data repositories 110 into the integrated data repository 104.

[0032] In one or more examples, the data integration system 114 may generate one or more tokens to identify an individual having data stored in the health insurance claims data repository 106 and data stored in the molecular data repository 108. In various examples, the data integration system 114 may generate one or more tokens by performing one or more hash functions. The data integration system 114 may perform one or more hash functions to generate one or more tokens based on information stored by at least one of the health insurance claims data repository 106 or the molecular data repository 108. For example, the information used by the data integration system 114 to generate the individual tokens by performing a hash function may include at least one of an identifier for each individual, a date of birth for each individual, a zip code for each individual, a date of birth for each individual, or a gender for each individual. In one or more examples, the identifier for each individual may include a combination of at least a portion of a first name for each individual and at least a portion of a last name for each individual. The tokens generated using data from the various data repositories may correspond to the same or similar information or the same or similar types of information stored by the various data repositories. For example, a token may be generated using a portion of an individual's name, date of birth, at least a portion of a zip code, and gender obtained from the health insurance claims data repository 106 and the molecular data repository 108.

[0033] The data integration system 114 may integrate data from many different data sources by analyzing tokens generated by performing one or more hash functions using data obtained from many different data sources. For example, the data integration system 114 may obtain one or more first tokens generated from data stored by the health insurance claims data repository 106 and one or more second tokens generated from data stored by the molecular data repository 108. The data integration system 114 may analyze the one or more first tokens with respect to the one or more second tokens to determine the individual first tokens that correspond to the individual second tokens. In one or more instances, the data integration system 114 may identify individual first tokens that match individual second tokens. A first token may match a second token if the data of the first token has at least a threshold amount of similarity to the data of the second token. In one or more instances, a first token may match a second token if the data of the first token is the same as the data of the second token. For example, a first token may match a second token if the alphanumeric string of the first token is the same as the alphanumeric string of the second token.

[0034] By determining which first token generated by using data stored by the health insurance claims data repository 106 corresponds to which second token generated by using data stored by the molecular data repository 108, the data integration system 114 may identify individuals who have data stored in both the health insurance claims data repository 106 and the molecular data repository 108. In this manner, the data integration system 114 may obtain data from the health insurance claims data repository 106 from many individuals and data from the molecular data repository 108 from many individuals, and store the health insurance claims data and molecular data of many individuals in the integrated data repository 104.

[0035] The data integration system 114 may also integrate data stored by one or more additional data repositories 110 with data from the health insurance claims data repository 106 and the molecular data repository 108 to generate the integrated data repository 104. For example, the data integration system 114 may obtain one or more third tokens generated from data stored by the additional data repository 110, such as a data repository that stores data corresponding to pathology reports. The data integration system 114 may analyze the one or more third tokens for the first token generated by using the information stored by the health insurance claims data repository 106 and the second token generated by using the information stored by the molecular data repository 108 to determine respective third tokens corresponding to respective first tokens and respective second tokens. In one or more instances, the data integration system 114 may identify a third token generated by using one or more hash functions and a common set of information obtained from the health insurance claims data repository 106, the molecular data repository 108, and the additional data repository 110.

[0036] By determining that the third token generated using data stored by the additional data repository 110 corresponds to the first token generated using data stored by the health insurance claims data repository 106 and the second token generated using data stored by the molecular data repository 108, the data integration system 114 may identify individuals having data stored in the health insurance claims data repository 106, the molecular data repository 108, and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the health insurance claims data repository 106 from many individuals, and data from the molecular data repository 108 and the additional data repository 110 from many individuals, and may store the health insurance claims data, molecular data, and additional data for many individuals in the integrated data repository 104.

[0037] Data stored by the integrated data repository 104 for many individuals may be accessible by using the individual's respective identifier. The data integration system 114 may implement many techniques as part of the de-identification process for storing and retrieving an individual's information in the integrated data repository 104. The individual's identifier may correspond to a key generated by using at least one hash function. The individual's identifier may also be generated by performing one or more salting processes on the key generated by using at least one hash function, a token generated by using one or more hash functions, and a common set of information obtained from the health insurance claims data repository 106, the molecular data repository 108, and / or the additional data repository 110. In one or more instances, the identifier generated by the data integration system 114 to access each individual's information stored by the integrated data repository 104 may be unique for each individual. In one or more instances, the individual's identifier may be generated by using at least a portion of the information used to generate a token related to the individual. In one or more additional instances, the individual's identifier may be generated by using various information from the information used to generate a token related to the individual.

[0038] The data integration system 114 may also generate the integrated data repository 104 from many different combinations of data repositories in a similar manner. For example, the data integration system 114 may obtain tokens generated from information stored by the health insurance claims data repository 106 and additional tokens generated from information stored by one or more additional data stores 110. The data integration system 114 may determine which individual tokens generated from the information stored by the health insurance claims data repository 106 correspond to which individual additional tokens generated from the information stored by the one or more additional data repositories 110. By determining which tokens generated by using data stored by the health insurance claims data repository 106 correspond to the additional tokens generated by using data stored by the additional data repository 110, the data integration system 114 may identify individuals who have data stored in both the health insurance claims data repository 106 and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the health insurance claims data repository 106 from many individuals and data from the additional data repository 110 from many individuals and store the health insurance claims data and additional data for many individuals in the integrated data repository 104. The health insurance claims data and additional data stored by the integrated data repository 104 for many individuals may be accessible using the individuals' respective identifiers.

[0039] In one or more alternative examples, the data integration system 114 may obtain tokens generated from information stored by the molecular data repository 108 and tokens generated from information stored by one or more additional data stores 110. The data integration system 114 may determine individual tokens generated from information stored by the molecular data repository 108 that correspond to individual additional tokens generated from information stored by the one or more additional data repositories 110. By determining tokens generated by using data stored by the molecular data repository 108 that correspond to additional tokens generated by using data stored by the additional data repository 110, the data integration system 114 may identify individuals having data stored in both the molecular data repository 108 and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the molecular data repository 108 from many individuals and data from the additional data repositories 110 from many individuals, and store the molecular data and additional data of many individuals in the integrated data repository 104. The molecular data and additional data stored by the integrated data repository 104 for many individuals may be accessible by using the individuals' respective identifiers.

[0040] The data stored by the integrated data repository 104 may be stored in compliance with one or more regulatory frameworks that protect the privacy and ensure the security of an individual's medical records, health information, and insurance information. For example, the data may be stored by the integrated data repository 104 in compliance with one or more government regulatory frameworks directed at protecting personal information, such as Health Insurance Portability and Account (HIPAA) and / or General Data Protection Regulation (GDPR). The integrated data repository 104 also stores the data in an anonymized manner to ensure protection of the privacy of individuals whose data is stored by the integrated data repository 104. To further ensure the privacy of individuals whose data is stored by the integrated data repository 104, the data integration system 114 may periodically regenerate the integrated data repository 104. For example, the data integration system 114 may generate the integrated data repository 104 once every quarter. In one or more additional examples, the data integration system 114 may generate the integrated data repository 104 on a monthly, weekly, or biweekly basis. By regenerating the integrated data repository 104 on a periodic basis, rather than simply refreshing it when new data is available, the integrated data repository 104 provides enhanced privacy protection for the data stored by the integrated data repository 104. That is, in situations where the data repository is simply refreshed with new data, it may be easier to track individuals associated with data newly added to the data repository, since the number of new individuals added at a given time is typically smaller than the existing number of individuals who already have data stored by the data repository.

[0041] In various examples, the data stored by the integrated data repository 104 may be accessed via a database management system. Additionally, the integrated data repository 104 may store data according to one or more database models. In one or more examples, the integrated data repository 104 may store data according to one or more relational database technologies. For example, the integrated data repository 104 may store data according to a relational database model. In one or more additional examples, the integrated data repository 104 may store data according to an object-oriented database model. In one or more other examples, the integrated data repository 104 may store data according to an extensible markup language (XML) database model. In yet another example, the integrated data repository 104 may store data according to a structured query language (SQL) database model. In yet another example, the integrated data repository may store data according to an image database model.

[0042] The data integration system 114 may generate the integrated data repository 104 by generating a number of data tables and generating links between the data tables. The links may indicate logical connections between the data tables. The data integration system 114 may generate the data tables by extracting a defined set of data from information obtained from the data repositories 106, 108, 110, 112 and storing the data in columns and rows of the respective data tables. In various examples, the logical connections between the data tables may include at least one of a one-to-one link, where a row of information in one data table corresponds to a row of information in another data table, a one-to-many link, where a row of information in one data table corresponds to multiple rows of information in another data table, or a many-to-many link, where multiple rows of information in one data table correspond to multiple rows of information in another data table.

[0043] A number of data tables may be arranged according to the data repository scheme 116. In the example of FIG. 1, the database scheme 114 includes a first data table 118, a second data table 120, a third data table 122, a fourth data table 124, and a fifth data table 124. Although the example of FIG. 1 includes five data tables, in additional implementations, the data repository scheme 116 may include more or fewer data tables. The data repository scheme 116 may also include links between the data tables 118, 120, 122, 124, 128. The links between the data tables 118, 120, 122, 124, 126 may indicate that information retrieved from one of the data tables 118, 120, 122, 124, 126 results in additional information stored by one or more additional data tables 118, 120, 122, 124, 126 to be retrieved. In addition, not all of the data tables 118, 120, 122, 124, 126 may be linked to each of the other data tables 118, 120, 120, 122, 124, 126. In the example of Figure 1, the first data table 118 is logically coupled to the second data table 118 by a first link 128, and the first data table 118 is logically coupled to the fourth data table 124 by a second link 130. In addition, the second data table 120 is logically coupled to the third data table 122 by a third link 132, and the fourth data table 124 is logically coupled to the fifth data table 126 by a fourth link 134. Furthermore, the third data table 122 is logically coupled to the fifth data table 126 by a fifth link 136.

[0044] In various examples, data tables are added and / or removed from the data repository scheme 116, so that additional links between data tables may be added or removed from the data repository scheme 116. In one or more instances, the integrated data repository 104 may store personal data tables in accordance with the data repository scheme 116 for at least some of the individuals for whom the data integration system 114 obtained information from a combination of at least two of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, and the one or more reference information data repositories 112. As a result, the integrated data repository 104 may store thousands, tens of thousands, and up to hundreds of thousands or more instances of each of the personal data tables 118, 120, 122, 124, 126 in accordance with the data repository scheme 116.

[0045] The data integration & analysis system 102 may also include a data pipeline system 138. The data pipeline system 138 may include any number of algorithms, software codes, scripts, macros, or any other set of computer executable instructions that process information stored by the integrated data repository 104 to generate additional data sets. The additional data sets may include information obtained from one or more of the data tables 118, 120, 122, 124, 126. The additional data sets may also include information derived from data obtained from one or more of the data tables 118, 120, 122, 124, 126. The parts of the data pipeline system 138 implemented to generate a first additional data set may be different from the parts of the data pipeline system 138 used to generate a second additional data set.

[0046] In one or more examples, the data pipeline system 138 may generate a data set indicating medical treatments received by a number of individuals. In one or more instances, the data pipeline system 138 may analyze information stored in at least one of the data tables 118, 120, 122, 124, 126 to determine health insurance codes corresponding to medical treatments received by a number of individuals. The data pipeline system 138 may analyze health insurance codes corresponding to medical treatments with respect to a library of data indicating prescribed medical treatments corresponding to the one or more health insurance codes to determine names of medical treatments received by the individuals. In one or more additional examples, the data pipeline system 138 may analyze information stored by the integrated data repository 104 to determine medical procedures received by a number of individuals. For example, the data pipeline system 138 may analyze information stored by one of the data tables 118, 120, 122, 124, 126 to determine treatments received by the individuals via at least one of an injection or an intravenous infusion. In one or more other examples, the data pipeline system 138 may analyze the information stored by the integrated data repository 104 to determine an individual's episodes of care, a course of therapy that the individual receives, the progression of a biological condition, or time to next treatment. In various examples, the datasets generated by the data pipeline system 138 may differ for different biological conditions. For example, the data pipeline system 138 may generate a first number of datasets for a first type of cancer, such as lung cancer, and a second number of datasets for a second type of cancer, such as colorectal cancer.

[0047] The data pipeline system 138 may also determine one or more confidence levels to assign to information associated with an individual having data stored by the integrated data repository 104. Each confidence level may correspond to a different measure of accuracy of information associated with an individual having data stored by the integrated data repository 104. The information associated with each confidence level may correspond to one or more features of the individual derived from the data stored by the integrated data repository 104. Confidence level values ​​for the one or more features may be generated by the data pipeline system 138 in conjunction with generating one or more data sets from the integrated data repository 104. In one or more examples, the first confidence level may correspond to a first range measure of accuracy, the second confidence level may correspond to a second range measure of accuracy, and the third confidence level may correspond to a third range measure of accuracy. In one or more additional examples, the second range measure of accuracy may include values ​​less than the values ​​of the first range measure of accuracy, and the third range measure of accuracy may include values ​​less than the values ​​of the second range measure of accuracy. In one or more examples, information corresponding to a first confidence level may be referred to as gold criteria information, information corresponding to a second confidence level may be referred to as silver criteria information, and information corresponding to a third confidence level may be referred to as bronze criteria information.

[0048] The data pipeline system 138 may determine the confidence level value of the individual's characteristic based on a number of factors. For example, each set of information may be used to determine the individual's characteristic. The data pipeline system 138 may determine the confidence level of the individual's characteristic based on a certain amount of completeness of each set of information used to determine the individual's characteristic. In a situation where one or more pieces of information are missing from a set of information associated with a first number of individuals, the confidence level of the characteristic may be lower than the confidence level of a second number of individuals where no information is missing from the set of information. In one or more examples, a certain amount of missing information may be used by the data pipeline system 138 to determine the confidence level of the individual's characteristic. For example, a large amount of missing information used to determine the individual's characteristic may result in a lower confidence level of the characteristic than a situation where a smaller amount of missing information is used to determine the characteristic. Furthermore, different types of information may correspond to different confidence levels of the characteristic. In one or more examples, the presence of a first piece of information used to determine the individual's characteristic may result in a higher confidence level of the characteristic than the presence of a second piece of information used to determine the characteristic.

[0049] In one or more instances, the data pipeline system 138 may determine a number of individuals included in the population with a primary diagnosis of lung cancer (or other biological condition). The data pipeline system 138 may determine a confidence level for each individual with respect to being classified as having a primary diagnosis of lung cancer. The data pipeline system 138 may use information from a number of columns included in the data tables 118, 120, 122, 124, and 126 to determine a confidence level of the individual's inclusion in the lung cancer population. The number of columns may include health insurance codes related to the diagnosis of the biological condition and / or the treatment of the biological condition. In addition, the number of columns may correspond to a date of diagnosis and / or treatment of the biological condition. The data pipeline system 138 may determine that the confidence level of the individual characterized as part of the lung cancer population in scenarios where information is available for each of the many columns, or at least a threshold number of columns, is higher than in instances where information is available for less than the threshold number of columns. Additionally, the data pipeline system 138 may determine a confidence level of the individual's inclusion in the lung cancer population based on the type of information and the availability of information associated with one or more columns. For example, in a situation where one or more diagnostic codes are present in association with one or more time periods for a group of individuals and one or more treatment codes are absent, the data pipeline system 138 may determine that "the confidence level of including the group of individuals in the lung cancer population is greater than in a situation where at least one of the diagnostic codes is absent and a treatment code used to determine whether an individual is included in the lung cancer population is present."

[0050] The data integration & analysis system 102 may include a data analysis system 140. The data analysis system 140 may receive an integrated data repository request 142 from one or more computing devices, such as an exemplary computing device 144. The one or more integrated data repository requests 142 may cause data to be retrieved from the integrated data repository 104. In various examples, the one or more integrated data repository requests 142 may cause data to be retrieved from one or more datasets generated by the data pipeline system 138. The integrated data repository request 142 may specify the data to be retrieved from the integrated data repository 104 and / or the one or more datasets generated by the data pipeline system 138. In one or more additional examples, the integrated data repository request 142 may include one or more pre-built queries corresponding to computer executable instructions to retrieve a specified set of data from the integrated data repository 104 and / or one or more datasets generated by the data pipeline system 138.

[0051] In response to the one or more integrated data repository requests 142, the data analysis system 140 may analyze data retrieved from at least one of the integrated data repositories 104 or one or more data sets generated by the data pipeline system 138 to generate data analysis results 146. The data analysis results 146 may be transmitted to one or more computing devices, such as the exemplary computing device 148. Although the example of FIG. 1 shows one or more integrated data repository requests 142 and the data analysis results 146 from one computing device 144 being transmitted to another computing device 148, in one or more additional implementations, the data analysis results 146 may be received by the same computing device that sent the one or more integrated data repository requests 142. The data analysis results 146 may be displayed by one or more user interfaces rendered by the computing device 144 or by the computing device 148.

[0052] In one or more examples, the data analysis system 140 may implement at least one of one or more machine learning techniques or one or more statistical techniques to analyze the data retrieved in response to one or more integrated data repository requests 142. In one or more examples, the data analysis system 140 may determine the survival rate of an individual having lung cancer in response to one or more treatments. In one or more additional examples, the data analysis system 140 may determine the survival rate of an individual having one or more genomic region mutations in response to one or more treatments. In various examples, the data analysis system 140 may generate data analysis results 146 in situations where the data retrieved from at least one of the integrated data repositories 104 or one or more data sets generated by the data pipeline system 138 meet one or more criteria. For example, the data analysis system 140 may determine whether at least a portion of the data retrieved in response to one or more integrated data repository requests 142 meets a threshold confidence level. In situations where the confidence level of at least some of the dates retrieved in response to one or more integrated data repository requests 142 is below a threshold confidence level, the data analysis system 140 may refrain from generating at least some of the data analysis results 146. In scenarios where the confidence level of at least some of the data retrieved in response to one or more integrated data repository requests 142 is at least the threshold confidence level, the data analysis system 140 may generate at least some of the data analysis results 146. In various examples, the threshold confidence level may be related to the type of data analysis results 146 generated by the data analysis system 140.

[0053] In one or more instances, the data analysis system 140 may receive an integrated data repository request 142 to generate a data analysis result 146 that indicates a survival rate of one or more individuals. In these instances, the data analysis system 140 may determine whether the data stored by the integrated data repository 104 and / or the data stored by one or more datasets generated by the data pipeline system 138 meets a threshold confidence level, such as a gold benchmark confidence level. In one or more additional examples, the data analysis system 140 may receive an integrated data repository request 142 to generate a data analysis result 146 that indicates a treatment to be received by one or more individuals. In these implementations, the data analysis system 140 may determine whether the data stored by the integrated data repository 104 and / or the data stored by one or more datasets generated by the data pipeline system 138 meets a lower threshold confidence level, such as a bronze benchmark confidence level.

[0054] The data integration & analysis system 102 may also include a treatment reference table system 150. Although the treatment reference table system 150 is shown in the illustrative example of FIG. 1 as separate from the data integration system 114 and the data pipeline system 138, at least some of the operations performed by the treatment reference table system 150 may be performed by at least one of the data integration system 114 or the data pipeline system 138. The treatment reference table system 150 may analyze information obtained from the health insurance claims data repository 106 and the at least one reference information data repository 112 to generate the treatment reference table 152. The data included in the treatment reference table 152 may be at least one of the data accessed or provided to the data analysis system 140 in response to one or more integrated data repository requests 142 to generate the data analysis results 146.

[0055] In one or more examples, the treatment reference table system 150 may analyze the health insurance claims data to determine an insurance code identifier corresponding to the treatment. The treatment reference table system 150 may then extract the insurance code identifier from the health insurance claims data corresponding to the treatment provided to the individual having data stored by the integrated data repository 104. The insurance code identifier corresponding to the treatment may be analyzed with respect to one or more criteria related to a format of the insurance code identifier. In addition, the insurance code identifier may be used in generating one or more API requests to retrieve information from the reference information data repository 112 related to the insurance code identifier and related to the treatment corresponding to the insurance code identifier. In one or more examples, the treatment reference table system 150 may determine that the insurance code identifier corresponding to the treatment is formatted according to the NDC9 format or the NDC10 format. The treatment reference table system 150 may generate an API request that is sent to the reference information data repository 112 and a response may be returned that includes information from the reference information data repository 112 having an additional insurance code identifier having an NDC11 format that corresponds to an initial insurance code identifier having an NDC9 format or an NDC10 format.

[0056] The treatment reference table system 150 may generate an additional API request by using a valid insurance code identifier that satisfies one or more formatting criteria. The additional API request may be sent to the reference information data repository 112 to obtain additional information related to the valid insurance code identifier. For example, the treatment reference table system 150 may obtain additional information related to the insurance code identifier including at least one of a name of the treatment corresponding to the insurance code identifier, one or more compositions of the treatment corresponding to the insurance code identifier, at least one class of the treatment corresponding to the insurance code identifier, at least one source related to the class information, an additional identifier for the insurance code identifier in the reference information data repository 112, or a term type related to the treatment.

[0057] The treatment reference table system 150 may generate the treatment reference table 152 by using information obtained in response to API requests sent to one or more reference information data repositories 112. In one or more additional examples, the treatment reference table system 150 may generate the treatment reference table 152 by using information obtained from the health insurance claims data repository 106. For example, the treatment reference table system 150 may generate rows of the treatment reference table 152 for unique insurance code identifiers included in data obtained from the health insurance claims data repository 106. The unique insurance code identifiers may include insurance code identifiers obtained from the health insurance claims data repository 106 that do not include the same set of alphanumeric characters or other symbols arranged in the same order as any other insurance code identifiers obtained from the health insurance claims data repository 106.

[0058] The treatment reference table 152 may also include a number of columns corresponding to individual rows of the treatment reference table 152. In one or more examples, the treatment reference table 152 may include a column indicating an insurance code identifier obtained from the health insurance claims data repository 106 and an additional identifier corresponding to the insurance code identifier obtained from the reference information data repository 112. In addition, the treatment reference table 152 may include a column indicating a name of the treatment corresponding to the insurance code identifier. Further, the treatment reference table 152 may include a column indicating a class of the treatment and a column indicating a source of the class. In various examples, the source of the class may correspond to a classification scheme used to organize and / or characterize the class of treatments. In one or more additional examples, the treatment reference table 152 may include a column including a comment generated by the treatment reference table system 150. The comment may relate to the treatment corresponding to a given row and / or information regarding the treatment. In one or more examples, one or more columns may include null values.

[0059] In one or more implementations, the reference information data repository 112 that stores the insurance code identifiers and requested information regarding the procedure may be external to an entity associated with the data integration & analysis system 102. In one or more additional examples, the reference information data repository 112 that stores the insurance code identifiers and requested information regarding the procedure may be an internal data repository controlled and maintained by an entity that controls, implements, and / or maintains the data integration & analysis system 102. For example, the internal reference information data repository 112 that stores the insurance code identifiers and information related to the procedure may store a copy of the information obtained from the external reference information data repository 112. In various examples, the internal reference information data repository 112 may be periodically updated using a number of API requests to the external reference information data repository 112 that stores the insurance code identifiers and procedure information.

[0060] In one or more instances, the data integration and analysis system 102 may obtain an integrated data repository request 142 that includes at least one of the following: a name of one or more respective therapies, a composition included in the one or more therapies, or a class of one or more therapies. The use of the names, compositions, or classes of the therapies may be more commonly known than insurance code identifiers that may be related to the therapies, thus allowing queries of the integrated data repository 104 to be generated more easily than in situations where insurance code identifiers are used that often change and / or are not readily available. After receiving the integrated data repository request 142 that includes at least one of the names, compositions, or classes of the one or more respective therapies, the data analysis system 140 may use the treatment reference table 152 to determine one or more insurance code identifiers that correspond to the information included in the integrated data repository request 142. The data analysis system 140 may then query the integrated data repository 104 to determine individuals that correspond to the therapies. In this manner, a population of individuals corresponding to a treatment may be determined based on a query to the data analysis system 150 including at least one of the names of the respective treatments, the compositions included in the treatments, or the classes of the treatments. Additional analysis of one or more of the information relating to the individuals included in the population may then be performed. For example, genetic information of the individuals included in the population may be analyzed. In addition, dosage information and / or frequency of the treatments taken for the individuals included in the population may be analyzed. Furthermore, diagnostic information of the individuals included in the population may be analyzed. Results of at least some of the analyses performed by the data analysis system 140 using the treatment reference table 152 may be included in the data analysis results 146.

[0061] In one or more additional instances, the data analysis system 140 may receive a request to analyze information corresponding to a population of patients. One or more genomic mutations may be present in the population of individuals. Additionally, the patients in the population may have received treatment for a given biological condition. In response to the request, the data analysis system 140 may analyze the information stored by the integrated data repository 104 to generate a data analysis result 146 that includes one or more quantitative measures corresponding to the patients in the population. For example, the data analysis system 140 may determine current world survival metrics for the patients in the population. In various instances, the data analysis system 140 may analyze information related to a population of patients to determine a survival probability over a period of time for the patients in the population. In one or more instances, the data analysis system 140 may analyze information related to one or more populations of patients to determine current world overall survival metrics for the patients in the population. In one or more other instances, the data analysis system 140 may analyze information related to a population to determine a "time to next treatment" metric and / or a "time to discontinuation" metric for the patients in the population.

[0062] In various examples, the data analysis system 140 may analyze information corresponding to the patients in the population to determine the evolution of the biological condition in at least a subset of the patients in the population. In one or more examples, the data analysis system 140 may determine the evolution of the population of patients who receive one or more pharmaceutical agents as part of a course of therapy. In one or more instances, the data analysis system 140 may analyze at least one of a "time to next treatment" metric or a "time to discontinuation" metric of the population of patients to determine the evolution of the biological condition of the patients in the population. In these instances, the data analysis system 140 may query the integrated data repository 1104 to determine the genomic data of the patients in the population and to identify the patients in the population who have one or more defined genomic mutations. The data analysis system 140 may then analyze the "time to next treatment" metric, the "time to discontinuation" metric, and / or the current global overall survival metric of the patients in the population who have one or more genomic mutations to determine the evolution of the biological condition of the patients in the population who have been treated for the biological condition.

[0063] In one or more further examples, the data analysis system 140 may analyze information about a population to determine a level of resistance developed by one or more patients in the population receiving one or more treatments for a biological condition. In various examples, the data analysis system 140 may analyze at least one of a "time to next treatment" metric, a "time to discontinuation" metric, or a current-world survival metric to determine a level of resistance developed by patients in the population receiving a treatment for a biological condition. In at least some examples, the data analysis system 140 may also determine a level of resistance for one or more treatments for patients in the population with one or more genomic mutations. In at least some examples, the level of resistance may be greater in situations where the "time to next treatment" or current-world survival rate has a lower value, and the level of resistance may be lower in situations where the value of the "time to next treatment" or current-world survival rate is relatively high.

[0064] In at least some examples, the data analysis system 140 may analyze the set of therapy information to determine a recommendation of one or more treatments to administer to a patient diagnosed with a biological condition. In one or more examples, the data analysis system 140 may analyze information about a population of patients to determine one or more characteristics of patients in the population who have received one or more sets of therapies with a relatively low level of resistance and / or a relatively low amount of progression. The data analysis system 140 may then analyze characteristics of one or more additional patients in the population analyzed by biological condition to determine whether to recommend one or more sets of therapies as treatment for the one or more additional patients. At least a portion of the one or more additional patients in the population may have already received treatment for the biological condition. In one or more additional examples, at least a portion of the one or more additional patients in the population may not have received treatment for the biological condition associated with the population. In various examples, the data analysis system 140 may also analyze information of patients in a given population to determine the effectiveness of a set of therapies for patients in the population. The efficacy of a course of therapy may correspond to the probability of the course of therapy (at least one of reducing the effect of the biological condition on the patients in the population or eliminating the biological condition on the patients in the population).

[0065] In various examples, the progression of the biological condition, the effectiveness of a course of therapy to treat the biological condition, the probability of developing resistance to a course of therapy, or a combination thereof may be determined by the data analysis system 140 by using at least one of one or more statistical techniques or one or more machine learning techniques. For example, the data analysis system 140 may implement at least one of a Cox proportional hazards model, a chi-square test, a log-rank test, or a Kaplan-Meier method to determine at least one of the progression of the biological condition, the effectiveness of a course of therapy to treat the biological condition, or the probability of developing resistance to a course of therapy. In one or more additional examples, the data analysis system 140 may implement one or more neural networks, one or more convolutional neural networks, or one or more residual neural networks to determine at least one of the progression of the biological condition, the effectiveness of a course of therapy to treat the biological condition, or the probability of developing resistance to a course of therapy.

[0066] One or more therapeutically effective amounts of one or more therapies may be administered to one or more patients in the population based on the evolution of the biological condition, the level of effectiveness of one or more courses of therapies, or the level of resistance determined for one or more patients. The therapeutically effective amount of one or more therapies may correspond to a new course of therapy or an additional course of therapy for one or more patients. In various examples, a therapeutically effective amount of one or more therapies may be provided to replace an ineffective therapy previously provided to one or more patients (e.g., due to one or more patients having developed a threshold level of resistance to one or more previous therapies provided to one or more patients).

[0067] In one or more examples, the data integration and analysis system 102 may identify treatments for administration to a patient having a given disease, disorder, or biological condition. Essentially any cancer treatment (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these treatments. In various examples, the treatment administered to the patient may include at least one chemotherapy drug. In at least some examples, the chemotherapy drugs may include alkylating agents (such as, but not limited to, chlorambucil, cyclophosphamide, cisplatin, and carboplatin), nitrosoureas (such as, but not limited to, carmustine and lomustine), antimetabolites (such as, but not limited to, fluorauracil, methotrexate, and fludarabine), plant alkaloids and natural products (such as, but not limited to, vincristine, paclitaxel, and topotecan), antitumor antibiotics (such as, but not limited to, bleomycin, adriamycin, and mitoxantrone), hormonal agents (such as, but not limited to, prednisone, dexamethasone, tamoxifen, and leuprolide), and biological response modifiers (such as, but not limited to, herceptin and avastin, erbitux, and rituxan). In one or more additional examples, the chemotherapy administered to the subject may include FOLFOX or FOLFIRI. The treatment may also include various polyadenosine diphosphate ribose polymerase (PARP) inhibitors, such as rucaparib and niraparib, in addition to kinase inhibitors, such as larotrectinib, binimetinib, encorafenib, and tofacitinib. In various examples, one or more treatments may be administered to treat one or more symptoms of cancer (e.g., entrecitinib, dacomitinib, and topotecan for treating lung cancer; triflurezin / tipiracil and iricatecan for treating colon cancer; apalutamide, degarelix, abiraterone, and enzalutamide for treating prostate cancer; and tamenotscatinib, talazoparib, and olaparib for treating breast cancer).

[0068] In one or more further examples, the one or more treatments may include at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to methods of enhancing immune responses against a given cancer type. In at least some implementations, immunotherapy refers to methods of enhancing T-cell responses against a tumor or cancer. In various examples, immunotherapy or immunotherapeutic agents target immune checkpoint molecules. Some tumors can evade the immune system by employing immune checkpoint pathways. Thus, targeting immune checkpoints has emerged as an effective approach to combat the ability of tumors to evade the immune system and to activate anti-turn or immunity against some cancers. Pardoll, Nature Reviews Cancer, 2012, 12:252-264.

[0069] In various examples, the immune checkpoint molecule is an inhibitory molecule that reduces signals involved in T cell responses to antigens. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 limits the activity of T cells in peripheral tissues during inflammatory responses. In addition, ligands for PD-1 (PD-L1 or PD-L2) are typically upregulated on the surface of many different tumors, resulting in downregulation of anti-tumor immune responses within the tumor microenvironment. In one or more examples, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In one or more additional examples, the inhibitory immune checkpoint molecule is a ligand for PD-1, such as PD-L1 or PD-L2. In at least some examples, the inhibitory immune checkpoint molecule is a ligand for CTLA4, such as CD80 or CD86. In yet other examples, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin like receptor (KIR), T cell membrane protein 3 (TIM3), galectin 9 (GAL9), or adenosine A2a receptor (A2aR).

[0070] Antagonists targeting these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against some cancers. Thus, in one or more implementations, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In various examples, the inhibitory immune checkpoint molecule is PD-1. In one or more examples, the inhibitory immune checkpoint molecule is PD-L1. In one or more additional examples, the antagonist of an inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In one or more other examples, the antibody or monoclonal antibody is an anti-CTLA4, anti-PD-1, anti-PD-Ll, or anti-PD-L2 antibody. In yet other examples, the antibody is a monoclonal anti-PD-1 antibody. In at least some examples, the antibody is a monoclonal anti-PD-Ll antibody. In various examples, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, an anti-CTLA4 antibody and an anti-PD-Ll antibody, or an anti-PD-Ll antibody and an anti-PD-1 antibody. In one or more instances, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In various scenarios, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In at least some implementations, the anti-PD-Ll antibody is one or more of atezomab (Tecentriq®), avelumab (Bavencio®), or durvalumab (Imfinzi®).

[0071] In one or more further examples, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., an antibody) to CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one or more additional examples, the antagonist is a soluble version of an inhibitory immune checkpoint molecule (such as a soluble fusion protein comprising an extracellular region of an inhibitory immune checkpoint molecule and an Fc region of an antibody). In one or more further examples, the soluble fusion protein comprises the extracellular region of CTLA4, PD-1, PD-L1, or PD-L2. In at least some examples, the soluble fusion protein comprises the extracellular region of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In various implementations, the soluble fusion protein comprises the extracellular region of PD-L2 or LAG3.

[0072] In one or more implementations, the immune checkpoint molecule is a costimulatory molecule that amplifies signals involved in a T cell response to an antigen. For example, CD28 is a co-stimulatory receptor expressed on T cells. When a T cell binds to an antigen through its T cell receptor, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on an antigen-presenting cell to amplify the T cell receptor signal and promote T cell activation. Because CD28 binds to the same ligands as CTLA4 (CD80 and CD86), CTLA4 can counteract or regulate the costimulatory signaling mediated by CD28. In at least some examples, the immune checkpoint molecule is a costimulatory molecule selected from CD28, inducible T cell costimulator (ICOS), CD137, OX40, or CD27. In various examples, the immune checkpoint molecule is a ligand of a costimulatory molecule, including, for example, CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.

[0073] Agents targeting these costimulatory checkpoint molecules can be used to enhance antigen-specific T cell responses to some cancers. Thus, in one or more examples, the immunotherapy or immunotherapeutic agent is an agonist of a costimulatory checkpoint molecule. In one or more additional examples, the agonist of a costimulatory checkpoint molecule is an agonist antibody, and preferably a monoclonal antibody. In one or more alternative implementations, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In at least some examples, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti-OX40, or anti-CD27 antibody. In various examples, the agonist antibody or monoclonal antibody is an anti-CD80, anti-CD86, anti-B7RP1, anti-B7-H3, anti-B7-H4, anti-CD137L, anti-OX40L, or anti-CD70 antibody.

[0074] Treatment options for treating particular genetic diseases, disorders or conditions other than cancer are generally well known to those of skill in the art and thus will be apparent given the particular disease, disorder or condition under consideration.

[0075] In one or more implementations, the customized therapy described herein is typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions including immunotherapeutic agents are typically administered intravenously. Some therapeutic agents are administered orally. The customized therapy (e.g., immunotherapeutic agents, etc.) may also be administered by any method known in the art, including, for example, buccal, sublingual, rectal, vaginal, urethral, ​​topical, intraocular, intranasal, and / or intraauricular, which may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, ointments, salves, etc.

[0076] In one or more further examples, the data integration and analysis system 102 may analyze data stored by one or more of the data repositories 106, 108, 110, 112 to generate additional data tables indicating information regarding one or more patients. In at least some examples, the one or more patients may have received one or more treatments for one or more biological conditions. In one or more examples, the additional data tables may include information regarding a population of patients in which the biological condition is present and who have received one or more prescribed treatments related to the biological condition. In one or more additional examples, the additional data tables may include information corresponding to an additional population of individuals in which the biological condition is not present. In various examples, the individuals in the additional population may be labeled as healthy individuals.

[0077] The additional data table may include many columns with individual columns corresponding to individual patients. The additional data table may also include many rows with individual rows corresponding to features of individual patients. The features may include numerical indicators corresponding to genomic mutations, biometric data, analytical test results, diagnostic imaging procedures, other diagnostic test results, patient physical characteristics, patient personal information, quantitative bioinformatics information, and one or more combinations thereof, and the like. In this manner, the additional data table may include a high-dimensional data matrix of column vectors representing many features of individual patients. In various examples, the information stored by the additional data table may be analyzed to determine the biological status of each individual. The biological status may be determined according to many different criteria. For example, the biological status of an individual patient may correspond to an overall level of health, where data relating to an individual in the absence of one or more defined biological conditions is used to determine a baseline level of health, and data of patients in the presence of one or more defined biological conditions is measured against the baseline level of health. The biological status of an individual patient may also be determined in relation to at least one of the presence of one or more biological conditions, the level at which one or more biological conditions are present, or the absence of one or more defined biological conditions. In at least some instances, the biological condition of a patient may vary according to at least one of age, diet, ethnic background, disease state, lifestyle choices, location, or environment.

[0078] Figure 2 illustrates an exemplary framework 200 that supports the arrangement of data tables within a unified data repository according to one or more implementations. In the example of Figure 2, the framework 200 includes a database scheme 202 that includes a first data table 204, a second data table 206, a third data table 208, a fourth data table 210, a fifth data table 212, a sixth data table 214, and a seventh data table 216. Although the example of Figure 2 includes seven data tables, in additional implementations the data repository scheme 202 may include more or fewer data tables. The data repository scheme 202 may also include links between the data tables 204, 206, 208, 210, 212, 214, 216. The links between the data tables 204, 206, 208, 210, 212, 214, 216 may indicate that information retrieved from one of the data tables 204, 206, 208, 210, 212, 214, 216 may result in additional information stored in one or more additional data tables 204, 206, 208, 210, 212, 214, 216 being retrieved. In addition, not all of the data tables 204, 206, 208, 210, 212, 214, 216 may be linked to each of the other data tables 204, 206, 208, 210, 212, 214, 216. 2, the first data table 204 is logically coupled to the second data table 206 by a first link 218, and the third data table 208 is logically coupled to the second data table 206 by a second link 220. The second data table 206 is also logically coupled to the fourth data table 210 by a third link 222, and the second data table 206 is logically coupled to the fifth data table 212 by a fourth link 224, and the second data table 206 is logically coupled to the sixth data table 214 by a fifth link 226. In addition, the fifth data table 212 is logically coupled to the sixth data table 214 by a sixth link 228, and the sixth data table 214 is logically coupled to the seventh data table 216 by a seventh link 230. Furthermore, the seventh data table 216 is logically coupled to the fourth data table 210 by an eighth link 232.In various examples, as data tables are added and / or removed from the data repository scheme 202, additional links between data tables may be added to or removed from the data repository scheme 202. In one or more instances, the integrated data repository 104 may store data tables according to the data repository scheme 202 for at least a portion of individuals for whom the data integration system 114 obtained information from a combination of at least two of the health insurance claims data repository 106, the molecular data repository 108, and the one or more additional data repositories 110. As a result, the integrated data repository 104 may store instances of each of the data tables 204, 206, 208, 210, 212, 214, 216 according to the data repository scheme 204 for thousands, tens of thousands, and up to hundreds of thousands or more individuals.

[0079] In one or more examples, the first data table 204 may store data corresponding to an individual's genome and genomic testing. For example, the first data table 204 may include columns containing information corresponding to the panel used to generate the genomic data, mutations in the genomic region, the type of mutation, copy number of the genomic region, coverage data indicating the number of nucleic acid molecules identified in the specimen having one or more mutations, test date, and patient information. The first data table 204 may also include one or more columns containing health insurance data codes that may correspond to one or more diagnostic codes. Additionally, the information in the first data table 204 may include at least one identifier of an individual associated with an instance of the first data table 204.

[0080] The second data table 206 may store data related to one or more patient visits by an individual to one or more healthcare providers. The third data table 208 may store information corresponding to each service provided to an individual in connection with one or more patient visits to one or more healthcare providers indicated by the second data table 206. For example, an individual may visit a healthcare provider and multiple services may be performed on the individual at the visit. The second data table 206 may include a column indicating information for each service of multiple services performed during the patient visit. Multiple third data tables 208 may be generated in connection with a patient visit and include columns indicating information at a higher level of granularity for each service provided during the patient visit than the information stored by the second data table 206 related to the patient visit. For example, the second data table 206 may include multiple columns indicating health insurance codes for various services provided to an individual during a patient visit, and the third data table 208 related to one of the services may include multiple columns of additional health insurance codes corresponding to additional information related to each service. The second data table 206 and the third data table 208 of the patient visit may indicate one or more dates of services corresponding to the patient visit.

[0081] The fourth data table 210 may include columns indicating information about the individual stored by the integrated data repository 104. For example, the fourth data table 210 may include columns indicating information related to at least one of the individual's location, the individual's gender, the individual's date of birth, the individual's date of death (if applicable), or one or more keys associated with the individual. In one or more examples, the fourth data table 210 may include one or more columns related to whether erroneous data was identified for the individual. In various examples, a single fourth data table 210 may be generated for each individual. Thus, the data repository scheme 202 may include multiple (e.g., thousands, tens of thousands, and up to hundreds of thousands or more) instances of the fourth data table 210.

[0082] The fifth data table 212 may include columns indicating information related to a health insurance company or government entity that paid for one or more services provided to each individual. For example, the fifth data table 212 may include one or more payer identifiers. The sixth data table 214 may include columns containing information corresponding to each individual's health insurance coverage information. In one or more examples, the sixth data table 214 may include columns indicating whether the individual has medical coverage, whether the individual has pharmaceutical coverage, and the type of health insurance plan associated with the individual (health maintenance organization (HMO), preferred provider organization (PPO), etc.).

[0083] The seventh data table 216 may include columns indicating information related to the medical treatment obtained by each individual. In one or more examples, the seventh data table 216 may include one or more columns indicating health insurance codes corresponding to the medical treatments available through the pharmacy. The health insurance codes may correspond to the individual medical treatments. In addition, the health insurance codes may indicate a diagnosis of a biological condition related to the individual. The seventh data table 216 may also include additional information, such as at least one of dosage, days of supply, amount dispensed, number of authorized refills, date of service, or information related to the individual receiving the medical treatment.

[0084] 3 illustrates an architecture 300 for generating one or more data sets from information retrieved from a data repository that integrates health-related data from multiple sources, according to one or more implementations. The architecture 300 may include a data integration & analysis system 102 and an integrated data repository 104. In addition, the data integration & analysis system 102 may include at least a data pipeline system 138 and a data analysis system 140. The data pipeline system 138 may include multiple sets of data processing instructions that are executable to generate respective data sets that may be analyzed by the data analysis system 140 in response to the integrated data repository requests 142 to generate data analysis results 146.

[0085] The data pipeline system 138 may include a first data processing instruction 302, a second data processing instruction 304, ... an Nth data processing instruction 306. The data processing instructions 302, 304, 306 may be executable by one or more processing units to perform a number of operations to generate respective data sets by using information retrieved from the integrated data repository 104. In one or more examples, the data processing instructions 302, 304, 306 may include at least one of software code, scripts, API calls, macros, etc. The first data processing instruction 302 may be executable to generate a first data set 308. In addition, the second data processing instruction 304 may be executable to generate a second data set 310. Furthermore, the Nth data processing instruction 306 may be executable to generate an Nth data set 312. In various examples, after the data integration & analysis system 102 generates the integrated data repository 104, the data pipeline system 138 may execute the data processing instructions 302, 304, 306 to generate the data sets 308, 310, 312. In one or more examples, the data sets 308, 310, 312 may be stored by the integrated data repository 104 or by an additional data repository accessible to the data integration & analysis system 102. At least some of the data processing instructions 302, 304, 306 may analyze health insurance codes to generate at least some of the data sets 308, 310, 312. Additionally, at least some of the data processing instructions 302, 304, 306 may analyze genomic data to generate at least some of the data sets 308, 310, 312.

[0086] In one or more examples, the first data processing instructions 302 may be executable to retrieve data from one or more first data tables stored by the integrated data repository 104. The first data processing instructions 302 may also be executable to retrieve data from one or more defined columns of the one or more first data tables. In various examples, the first data processing instructions 302 may be executable to identify individuals having health insurance codes stored in one or more column and row combinations corresponding to one or more diagnostic codes. The first data processing instructions 302 may then be executable to analyze the one or more diagnostic codes to determine the biological conditions for which the individual is analyzed. In one or more examples, the first data processing instructions 302 may be executable to analyze the one or more diagnostic codes for a library of diagnostic codes indicating one or more biological conditions corresponding to each diagnostic code. The library of diagnostic codes may include hundreds up to thousands of diagnostic codes. The first data processing instructions 302 may also be executable to determine diagnosed individuals having a biological condition by analyzing individual timing information (such as date of treatment, date of diagnosis, date of death, and any combination or combinations thereof).

[0087] The second data processing instructions 304 may be executable to retrieve data from one or more second data tables stored by the integrated data repository 104. The second data processing instructions 304 may also be executable to retrieve data from one or more defined columns of the one or more second data tables. In various examples, the second data processing instructions 304 may be executable to identify individuals having health insurance codes stored in one or more column and row combinations corresponding to one or more treatment codes. The one or more treatment codes may correspond to treatments obtained from a pharmacy. In one or more additional examples, the one or more treatment codes may correspond to treatments received by a medical procedure (such as an injection or intravenous injection). The second data processing instructions 304 may be executable to determine one or more treatments corresponding to each health insurance code included in the one or more second data tables by analyzing the health insurance codes related to the predetermined set of information. The predetermined set of information may include a data library indicating one or more treatments corresponding to one from hundreds up to thousands of health insurance codes. The second data processing instructions 304 may generate a second dataset 310 for prescribing respective treatments to be received by a group of individuals. In one or more instances, the group of individuals may correspond to individuals included in the first dataset 308. The second dataset 310 may be arranged in rows and columns, with one or more rows corresponding to a single individual and one or more columns prescribing the treatments each individual will receive.

[0088] The Nth processing instructions 306 (N may be any positive integer) may be executable to generate the Nth dataset 312 by combining information from a number of previously generated datasets (e.g., the first dataset 308 and the second dataset 310). In addition, the Nth processing instructions 306 may be executable to retrieve additional information from one or more additional columns of the integrated data repository 104 and generate the Nth dataset 312 to coordinate the additional information from the integrated data repository 104 with the information obtained from the first dataset 308 and the second dataset 310. For example, the Nth processing instructions 306 may be executable to identify individuals included in the first dataset 308 diagnosed by a biological condition and analyze defined columns of one or more additional data tables of the integrated data repository 104 to determine a date of treatment indicated in the second dataset 210 corresponding to the individuals included in the first dataset 308. In one or more other examples, the Nth processing instructions 306 may be executable to analyze columns of one or more additional data tables in the integrated data repository 104 to determine dosages of treatments prescribed in the second dataset 310 received by individuals included in the first dataset 308. In this manner, the Nth processing instructions 306 may be executable to generate an episode of care dataset based on information included in the population dataset and the treatment dataset.

[0089] In one or more instances, in response to receiving the integrated data repository request 142, the data analysis system 140 may determine one or more datasets corresponding to features of the query related to the integrated data repository request 142. For example, the data analysis system 140 may determine that information included in the first dataset 308 and the second dataset 310 is applicable to responding to the integrated data repository request 142. In these scenarios, the data analysis system 140 may analyze at least a portion of the data included in the first dataset 308 and the second dataset 310 to generate the data analysis result 146. In one or more additional examples, the data analysis system 140 may determine various datasets for responding to various queries included in the integrated data repository request 142 to generate the data analysis result 146.

[0090] FIG. 4 illustrates an architecture 400 for generating an integrated data repository including de-identified health insurance claims data and de-identified genomic data, according to one or more implementations. The architecture 400 may include a data integration & analysis system 102, a health insurance claims data repository 106, and a molecular data repository 108. The data integration & analysis system 102 may obtain patient information 402 from the molecular data repository 108. The patient information 402 may include genomic data 404 of an individual having data stored by the molecular data repository 108. The genomic data 404 may indicate the results of one or more nucleic acid sequence operations that analyze the sequence of nucleic acid molecules contained in a specimen obtained from the individual for one or more target genomic regions. In one or more examples, the specimen may be obtained from tissue of one or more individuals. In one or more additional examples, the specimen may be obtained from a fluid (such as blood or plasma) of one or more individuals. The one or more target genomic regions may correspond to genomic regions that correspond to the presence of one or more biological conditions. For example, the target site may correspond to a genomic region of a reference genome having a mutation present in an individual in whom a biological condition exists. In one or more instances, the target site may correspond to a genomic region of a reference human genome in which one or more mutations are present in an individual in whom one or more forms of cancer exist. The patient information 402 may also include information indicating personal information about an individual having data stored by the molecular data repository 108 and information corresponding to tests and analyses performed on a specimen provided by the individual.

[0091] The Data Integration & Analysis System 102 may perform an anonymization process 406 to anonymize personal information obtained from the molecular data repository 108. The Data Integration & Analysis System 102 may implement one or more computational techniques as part of the anonymization process to anonymize data related to individuals stored by the molecular data repository 108 such that the anonymized data protects the privacy of the individuals and complies with one or more privacy regulatory frameworks. The anonymization process 406 may include accessing a token at 408. In various examples, the access token may include an alphanumeric string. In one or more examples, the access token may be generated by the Data Integration & Analysis System 102. In one or more additional examples, the access token may be generated by a third party and obtained by the Data Integration & Analysis System 102.

[0092] The access token may be generated by using one or more hash functions related to a subset 410 of the patient information 402. For example, for individuals having information stored by the molecular data repository 108, the access token may be generated by using a combination of at least a portion of each individual's first name, at least a portion of each individual's last name, at least a portion of each individual's date of birth, the individual's gender, and at least a portion of each individual's location identifier. The de-identification process 406 may also include generating 412 an identifier for the individual having data stored by the molecular data repository 108. The identifier may be generated by the data integration & analysis system 102 by using one or more hash functions that are different from the one or more hash functions used to generate the access token. In one or more instances, the data integration & analysis system 102 may generate intermediate versions of each identifier by using one or more hash functions, and then apply one or more salting techniques to the intermediate versions of the identifier to generate a final version of the identifier. In various examples, the data integration & analysis system 102 may generate 412 an identifier by using at least a portion of the information for each individual stored by the molecular data repository 108. In one or more examples, the identifier may be generated based on a patient identifier included in the patient information 402. The identifier generated by the data integration & analysis system 102 may be unique for each individual having data stored by the molecular data repository 108.

[0093] In operation 414, the data integration and analysis system 102 may generate corrected patient information 416 based on the identifiers. The corrected patient information 416 may include genomic data 404 related to the molecular data repository 108 and the individuals associated with the respective individual identifiers. The corrected patient information 416 may have a data structure 418. The data structure 418 may include a column containing the identifiers of each of the individuals associated with the molecular data repository 108 and a number of columns containing the genomic data 404 related to the individuals (e.g., identifiers of one or more genes, modifications to one or more genes, types of modifications to genes, etc.).

[0094] The data integration & analysis system 102 may generate a token file 420. The token file 420 may include first tokens 422 accessed in the operation 408 for each individual having data stored by the molecular data repository 108. The token file 420 may include a data structure 424 including a number of columns containing information for each individual. The data structure 424 may include a column indicating each identifier generated by the data integration & analysis system 102 and a column indicating one or more first tokens 422 associated with each identifier. The data integration & analysis system 102 may transmit the token file 420 to a claims data management system 426 coupled to the claims data repository 106. The claims data management system 426 may analyze the first tokens 422 for corresponding second tokens 428. The second tokens 428 may be accessed or generated by the claims data management system 426. The second token 428 may be generated by using the same or a similar subset of the information of the individuals having data stored in the health insurance claims data repository 106 as the subset 410 of the patient information 402. For example, the second token 428 may be generated by using a combination of at least a portion of each individual's first name, at least a portion of each individual's last name, at least a portion of each individual's date of birth, the individual's gender, and at least a portion of each individual's location identifier.

[0095] In various examples, the health insurance claims data management system 426 may retrieve health insurance claims data from the personal health insurance claims data repository 106 associated with each second token 428 that matches a corresponding first token 422. A first token 422 may match a second token 428 if the data of the first token 422 has at least a threshold similarity with respect to the data of the second token 428. In one or more examples, a first token 422 may match a second token 428 if the data of the first token 422 is the same as the data of the second token 428.

[0096] In response to identifying the individual's health insurance claim data having respective second tokens 428 corresponding to respective first tokens 422, the health insurance claim data management system 426 may generate corrected health insurance claim data 430. The health insurance claim data management system 426 may transmit the corrected health insurance claim data 430 to the data integration and analysis system 102. In one or more examples, the corrected health insurance claim data 430 may be formatted according to a data structure 432. The data structure 432 may include a column that includes a subset of the second tokens 428 that correspond to the first tokens 422 and a number of columns that include the health insurance claim data.

[0097] In operation 434, the data integration & analysis system 102 may integrate genomic data and health insurance claim data for individuals that are common to both the molecular data repository 108 and the health insurance claim data repository 106. The data integration & analysis system 102 may determine individuals that are common to both the molecular data repository 108 and the health insurance claim data repository 106 by determining genomic data and health insurance claim data that correspond to common tokens. The data integration & analysis system 102 may determine that a first token 422 related to a portion of the genomic data 404 corresponds to a second token 428 related to a portion of the health insurance claim data by determining a measure of similarity between the first token 422 and the second token 428. In a scenario in which the first token 422 has at least a threshold similarity with respect to the second token 428, the data integration and analysis system 102 may store the corresponding portions of the genomic data 404 and the corresponding portions of the health insurance claims data related to the individual's identifier in an integrated data repository (such as the integrated data repository 104 of Figures 1, 2, and 3).

[0098] FIG. 5 illustrates a framework 500 for generating a dataset by the data pipeline system 138 based on data stored by the integrated data repository 104, according to one or more implementations. The integrated data repository 104 may store health insurance claims data and genomic data for a group of individuals 502. For example, the integrated data repository 104 may store information obtained from health insurance claims records 504 for the group of individuals 502. For each individual included in the group of individuals 502, the integrated data repository 104 may store information obtained from multiple health insurance claims records 504. In various examples, the information stored by the integrated data repository 104 may include and / or be derived from thousands, tens of thousands, hundreds of thousands, and up to millions of health insurance claims records 504. Additionally, each health insurance claim record may include multiple columns. As a result, the integrated data repository 104 may be generated through the analysis of millions of columns of health insurance claims data.

[0099] Furthermore, while the health insurance claims data may be organized according to a structured data format, the health insurance claims data is typically arranged for viewing by health insurance providers, patients, and health care providers to show financial and insurance code information related to services provided by the health care provider to the individual. Thus, the health insurance claims data is not easily analyzed to obtain insights that may be available in relation to the characteristics of the individual for whom a biological condition exists and that may assist in the treatment of the individual for the biological condition. The integrated data repository 104 may be generated and organized by analyzing and modifying the raw health insurance claims data in a manner that "enables the data stored by the integrated data repository 104 to be further analyzed to determine trends, characteristics, features, and / or insights regarding the individual for whom one or more biological conditions may exist." Furthermore, the integrated data repository 104 may be generated by using the genomic data records 506 of the group of individuals 502. In various examples, a large amount of health insurance claims data may be matched with the genomic data of the group of individuals 502 to generate the integrated data repository 104. In one or more examples, the processes and techniques implemented to integrate the health insurance claim records 504 and the genomic request records 506 to generate the integrated data repository 104 may be complex, and therefore efficiency enhancing techniques, systems and processes may be implemented to minimize the amount of computing resources used to generate the integrated data repository 104.

[0100] In one or more examples, the data pipeline system 138 may access information stored by the integrated data repository 104 to generate a data set including a number of additional data records 508 including information related to at least a portion of the group of individuals 502. In the example of FIG. 5, the additional data records 508 include information indicating whether the individual is included in a population of individuals having lung cancer. The data pipeline system 138 may execute a number of different sets of data processing instructions to determine the population of the group of individuals 502 having lung cancer. In various examples, the additional data records 508 may indicate information used to determine the status of the individual 502 with respect to lung cancer, such as one or more transaction insurance identifiers, one or more international classification of diseases (ICD) codes, and one or more health insurance transaction dates. In addition to including a column indicating whether the individual 502 is included in a lung cancer population, the additional data records 508 may include a column indicating a confidence level of the individual 502's status with respect to the presence of lung cancer.

[0101] 6 illustrates an architecture 600 for generating a reference data table 152 indicating an identifier of a treatment to be provided to a patient for which one or more biological conditions may be present, according to one or more implementations. The architecture 600 may include a data integration & analysis system 102. The data integration & analysis system 102 may include a treatment reference table system 150 that may generate the treatment reference table 152. In one or more examples, the data integration & analysis system 102 may at least one of obtain or generate a data table 602 including insurance code identifiers. The data table 602 may include and / or may be generated by using information obtained from the health insurance claims data repository 106 of FIG. 1. In various examples, the data table 602 may include one or more data tables stored by the integrated data repository 104 of FIG. 1. For example, the data table 602 may include a pharmacy records data table and / or a line of service data table stored by the integrated data repository 104.

[0102] The treatment referencing system 150 may analyze the data table 602 to generate a subset of insurance code identifiers included in the data table 602. For example, the treatment referencing system 150 may analyze the data table 602 for one or more criteria. In one or more examples, the treatment referencing system 150 may analyze the information included in the data table 602 to identify insurance code identifiers that correspond to treatments provided to individuals in which a biological condition exists. For example, the insurance code identifiers associated with the treatments may have one or more defined formats. In one or more examples, the insurance code identifiers corresponding to the treatments may have a defined number of alphanumeric or other symbols (eight characters or symbols, nine characters or symbols, ten characters or symbols, eleven characters or symbols, twelve characters or symbols, etc.). The insurance code identifiers corresponding to the treatments may also have one or more configurations of alphanumeric symbols and / or characters.

[0103] In one or more additional examples, the insurance code identifier corresponding to the medical treatment may have multiple segments, with multiple alphanumeric characters and / or symbols included in each segment. In various examples, the insurance code identifier corresponding to the medical treatment may have at least one segment, at least two segments, at least three segments, or at least four segments. Each segment may include at least one alphanumeric symbol, at least two alphanumeric symbols, at least three alphanumeric symbols, or at least four alphanumeric symbols. In one or more examples, the segments of the insurance code identifier corresponding to the medical treatment may be separated by a symbol. In various examples, the segments of the insurance code identifier corresponding to the medical treatment may be separated by at least one of a dash, a comma, or a period.

[0104] In one or more implementations, the treatment reference table system 150 may analyze the information contained in the data table 602 to determine column and row values ​​that correspond to one or more criteria. The treatment reference table system 150 may generate a set of treatment insurance code identifiers 604 that includes insurance code identifiers in the data table 602 that satisfy the one or more criteria. For example, the treatment reference table system 150 may analyze the information stored by the data table 602 to determine insurance code identifiers that correspond to one or more formats. In one or more instances, the treatment reference table system 150 may analyze the information stored by the data table 602 to determine whether the insurance code identifiers correspond to at least one of the following formats: NDC9, NDC10, or NDC11. In various instances, at least tens of thousands of insurance code identifiers, hundreds of thousands of insurance code identifiers, and up to millions or more insurance code identifiers are analyzed by the treatment reference table system 150 to generate the set of treatment insurance code identifiers 604. In at least some instances, each identifier included in the set of medical reimbursement code identifiers may uniquely identify a medical treatment provided to a patient in whom a biological condition exists.

[0105] In one or more additional examples, the treatment reference table system 150 may determine one or more columns of the data table 602 that include insurance code identifiers that satisfy one or more criteria. For example, the treatment reference table system 150 may determine that a number of columns 606 include insurance code identifiers that correspond to the insurance code identifier formatting criteria. In one or more examples, the number of columns 606 may be determined based on user input obtained by the data integration and analysis system 102. In one or more other examples, the treatment reference table system 150 may analyze information included in at least some of the columns of the data table 602 to determine a number of columns 606 that include insurance code identifiers that satisfy one or more formatting criteria. In one or more instances, each treatment insurance code identifier included in the set of treatment code identifiers 604 may have a first format 608, a second format 610, and a third format 612. Although the example of FIG. 6 shows that the set of insurance code identifiers 604 can have three formats, in additional implementations, the set of insurance code identifiers 604 can have fewer or more formats.

[0106] The care reference table system 150 may also analyze the set of insurance code identifiers 604 to determine a subset of the insurance code identifiers that includes one or more unique insurance code identifiers. In one or more examples, the care reference table system 150 may perform one or more de-duplication processes to generate a de-duplicated insurance code identifier 614 based on the set of insurance code identifiers 604. In various examples, the de-duplicated insurance code identifiers 614 may include insurance code identifiers corresponding to the treatment that are different from the insurance code identifiers of the de-duplicated insurance code identifiers 614. For example, each unique insurance code identifier included in the de-duplicated insurance code identifiers 614 may include at least one alphanumeric character and / or other symbol located in at least one position that is not present in the corresponding position of other insurance code identifiers included in the de-duplicated insurance code identifiers 614.

[0107] The care reference table system 150 may analyze the de-duplicated insurance code identifiers 614 for one or more additional formatting criteria. In various examples, the de-duplicated insurance code identifiers 614 may include unique insurance code identifiers formatted according to one or more NDC formats, and the care reference table system 150 may analyze the de-duplicated insurance code identifiers 614 to determine the NDC format of each insurance code identifier included in the de-duplicated insurance code identifiers 614. In one or more examples, each NDC format may have one or more formatting characteristics that are different from the formatting characteristics of the additional NDC formats. For example, insurance code identifiers formatted according to an NDC9 format may have one or more first characteristics, insurance code identifiers formatted according to an NDC10 format may have one or more second characteristics, and insurance code identifiers formatted according to an NDC11 format may have one or more third characteristics. The one or more first features may be different from the one or more second features and the one or more third features, and the one or more second features may be different from the one or more third features. In one or more examples, the care reference table system 150 may determine a portion of the de-duplicated insurance code identifier 614 that corresponds to an NDC9 format. Additionally, the care reference table system 150 may determine a portion of the de-duplicated insurance code identifier 614 that corresponds to an NDC10 format. Furthermore, the care reference table system 150 may determine a portion of the de-duplicated insurance code identifier 614 that corresponds to an NDC11 format. In one or more additional examples, the care reference table system 150 may determine an NDC format of the insurance code identifier based at least in part on determining a number of alphanumeric characters present in the insurance code identifier. In one or more other examples, the care reference table system 150 may determine an NDC format of the insurance code identifier based at least in part on a number of segments included in the insurance code identifier and / or a number of alphanumeric characters present in individual segments of the insurance code identifier.In one or more scenarios, the first format 608 may correspond to the NDC9 format, the second format 610 may correspond to the NDC10 format, and the third format 612 may correspond to the NDC11 format.

[0108] The data integration and analysis system 102 may retrieve information from one or more reference information data repositories to obtain information that the treatment reference table system 150 uses to generate the treatment reference table 152. In various examples, individual treatment insurance code identifiers included in the de-duplicated treatment insurance code identifiers 614 may be used to retrieve information from one or more reference information data repositories. In one or more examples, the data integration and analysis system 102 may be in communication with a treatment classification data management system 616 via one or more communication networks. The treatment classification data management system 616 may be coupled to a treatment classification data repository 618. The treatment classification data management system 616 may manage the storage and retrieval of information stored by the treatment classification data repository 618. The treatment classification data repository 618 may store information related to treatment insurance code identifiers. In one or more examples, the treatment classification data repository 618 may store information related to treatments corresponding to the treatment insurance code identifiers. For example, the treatment classification data management system 616 may store one or more treatment data sets 620. In various examples, Therapy Classification Data Management System 616 may be maintained, controlled, and implemented by an entity that is external to the entity that controls, maintains, and implements Data Integration & Analysis System 102. In one or more additional examples, Therapy Classification Data Management System 616 may be internal to Data Integration & Analysis System 102 and may be controlled, maintained, and implemented by the same entity as Data Integration & Analysis System 102. In these scenarios, Therapy Classification Data Repository 618 may store a copy of the information obtained from the external data repository using one or more API requests.

[0109] The individual treatment data sets 620 may store information corresponding to individual insurance code identifiers. The individual treatment data sets 620 may correspond to an arrangement of data that may be accessed as a group in response to a query from the treatment classification data management system 616. In various examples, the treatment data sets 620 may include additional identifiers, such as a data management system (DMS) identifier 622 used by the treatment classification data management system 616 to store and retrieve information related to the insurance code identifiers. In one or more examples, the individual DMS identifiers 622 used by the treatment classification data management system 616 may correspond to one or more insurance code identifiers. In at least some scenarios, the DMS identifiers 622 used by the treatment classification data management system 616 to store and retrieve information related to the insurance code identifiers may have a format that is different from the format of the insurance code identifiers. In one or more examples, the DMS identifiers 622 may include one or more RxNorm concept unique identifiers (RxCUIs).

[0110] In one or more implementations, the treatment dataset 620 may store information for the insurance code identifier corresponding to one or more names of one or more treatments associated with the insurance code identifier, one or more compositions of one or more treatments associated with the insurance code identifier, one or more classes of one or more treatments associated with the insurance code identifier, one or more sources of the class or classes, one or more term types of one or more treatments associated with the insurance code identifier, or one or more combinations thereof. In various examples, the treatment dataset 620 may store a status of the treatment corresponding to the insurance code identifier, at least one of a start date or an end date that the insurance code identifier was active in the treatment classification data repository 618, a history of the insurance code identifier associated with the treatment classification data management system 616, or one or more combinations thereof. The status of the insurance code identifier may indicate whether the insurance code identifier is currently being used to identify information related to the respective treatment.

[0111] The therapy classification data management system 616 may implement one or more application programming interfaces (APIs) 624. The one or more APIs 624 may include calls that may be used by the therapy classification data management system 616 to request information from the therapy classification data repository 618. The calls of the one or more APIs 624 may include one or more fields and one or more formats for retrieving information stored by the therapy classification data repository 618. In one or more examples, the data integration & analysis system 102 may send one or more API requests 626 to the therapy classification data management system 616. The therapy classification data management system 616 may then generate queries to the therapy classification data repository 618 to retrieve data from the therapy classification data repository 618. In one or more examples, the queries generated by the therapy classification data management system 616 may correspond to retrieving one or more therapy datasets 620 in response to the API requests 626. The therapy classification data management system 616 may send one or more API responses 628 to the data integration and analysis system 102 based on the one or more API requests 626. The therapy reference table system 150 may generate at least a portion of the therapy reference table 152 by using the information included in the API responses 628.

[0112] In one or more examples, the API requests 626 can be generated according to one or more schemes, with each scheme being used to retrieve a defined set of data. In various examples, an API request 626 corresponding to a first scheme 630 can be used to retrieve a first treatment data set 632 stored by the treatment classification data repository 618. In addition, an API request 626 corresponding to a second scheme 634 can be used to retrieve a second treatment data set 636 stored by the treatment classification data repository 618. Furthermore, an API request 626 corresponding to a third scheme 638 can be used to retrieve a third treatment data set 640 stored by the treatment classification data repository 618.

[0113] In various examples, the treatment reference table system 150 may generate the API request 626 corresponding to the first scheme 630 by using the first set of information. Additionally, the treatment reference table system 150 may generate the API request 626 corresponding to the second scheme 634 by using the second set of information. Furthermore, the treatment reference table system 150 may generate the API request 626 corresponding to the third scheme 638 by using the third set of information. For example, the treatment reference table system 150 may generate the API request 626 according to the first scheme 630 by using at least one of the de-duplicated insurance code identifier 614 having the first format 608 or the de-duplicated insurance code identifier 614 having the second format 610. In one or more examples, the care reference table system 150 may modify the de-duplicated insurance code identifier 614 having the first format 608 and / or the de-duplicated insurance code identifier 614 having the second format 610 to generate the API request 626 according to the first scheme 630. For example, the care reference table system 150 may at least one of add or remove one or more symbols and / or one or more alphanumeric characters from the de-duplicated insurance code identifier 614 having the first format 608 or from the de-duplicated insurance code identifier 614 having the second format 610 to generate the API request 626 according to the first scheme 630. In one or more examples, the care reference table system 150 may add a hyphen to separate the last two alphanumeric characters of the de-duplicated insurance code identifier 614 having the first format 608 and / or the second format 610 to generate the API request 626 according to the first scheme 630.

[0114] In one or more instances, the treatment reference table system 150 may generate the API request 626 according to the first scheme 630 by using the de-duplicated insurance code identifier 614 having an NDC9 format or an NDC10 format. For example, the treatment reference table system 150 may generate the API request 626 including the de-duplicated insurance code identifier 614 having an NDC9 format or an NDC10 format, or a modified version of the de-duplicated insurance code identifier 614. The API request 626 may also include at least one of an identifier of the respective treatment data set 620 or instructions for retrieving information from the treatment classification data repository 618. The additional information may also be used to generate the API request 626 according to the first scheme 630. The API request 626 may be formatted as a hypertext transfer protocol (HTTP) request.

[0115] The API request 626 generated by the treatment reference table system 150 may be sent to the treatment classification data management system 616. The treatment classification data management system 616 may then retrieve information from the treatment classification data repository 618 corresponding to the first treatment data set 632 and send the first treatment data set 632 to the data integration and analysis system 102 via an API response 628. In various examples, the first treatment data set 632 may include an additional treatment insurance code identifier having an NDC11 format corresponding to the de-duplicated insurance code identifier having the first format 608 or the second format 610. Thus, in these scenarios, the API request 626 generated according to the first scheme 630 may be used to obtain a treatment insurance code identifier having an NDC11 format corresponding to the treatment insurance code identifier having an NDC9 format or an NDC10 format. In various examples, the treatment reference table system 150 may determine whether the additional treatment insurance code identifier included in the first treatment data set 632 is included in the set of insurance code identifiers 604. In scenarios where the additional insurance code identifier included in the first treatment data set 632 is already included in the set of insurance code identifiers 604, the additional insurance code identifier may be ignored. In one or more additional examples, in situations where the additional insurance code identifier included in the first treatment data set 632 is not included in the set of insurance code identifiers 604, the additional insurance code identifier may be added to the de-duplicated insurance code identifiers 614.

[0116] The first treatment data set 632 may include additional information. For example, the first treatment data set 632 may also include a DMS identifier 622 corresponding to the de-duplicated medical insurance code identifier 614 used to generate the API request 626 according to the first scheme 630. For example, the first treatment data set 632 may include an RxCUI corresponding to an NDC9 identifier, an NDC10 identifier, and / or an NDC11 identifier. In addition, the first treatment data set 632 may include one or more characteristics of the treatment corresponding to the de-duplicated medical insurance code identifier 614 used to generate the API request 626 according to the first scheme 630. In various examples, the one or more characteristics may indicate packaging of the treatment, physical characteristics of the treatment (e.g., color, shape, etc.), dosing characteristics of the treatment, one or more additional characteristics of the treatment (e.g., generic, active, inactive, etc.), or one or more combinations thereof.

[0117] The treatment reference table system 150 may generate the API request 626 according to the second scheme 634 by using the de-duplicated treatment insurance code identifier 614 having the third format 612. For example, the treatment reference table system 150 may generate the API request 626 according to the second scheme 634 by using the de-duplicated treatment insurance code identifier 614 having the NDC11 format. The API request 626 generated according to the second scheme 634 may also include at least one of an identifier of the respective treatment data set 620 or instructions for retrieving information from the treatment classification data repository 618. The additional information may also be used to generate the API request 626 according to the second scheme 634. The API request 626 may be formatted as a Hypertext Transfer Protocol (HTTP) request.

[0118] In response to receiving the API request 626 generated according to the second scheme 634 from the data integration and analysis system 102, the treatment classification data management system 616 may retrieve a second treatment data set 636 from the treatment classification data repository 618 corresponding to the de-duplicated treatment insurance code identifier 614 used to generate the API request 626 having the third format 612. In various examples, the information included in the second treatment data set 636 may indicate a status of the de-duplicated treatment insurance code identifier 614 having the third format 612 in the treatment classification data management system 616. For example, the second treatment data set 636 may indicate whether the de-duplicated treatment insurance code identifier 614 having the NDC11 format may be actively used to retrieve information from the treatment classification data repository 618 by using the de-duplicated treatment insurance code identifier 614. In these scenarios, the API request 626 corresponding to the second scheme 634 may be used to determine whether the de-duplicated treatment insurance code identifier 614 having the third format 612 is valid.

[0119] In one or more examples, the treatment reference table system 150 may analyze the second treatment data set 636 to determine the validity of the de-duplicated insurance code identifier 614 having the third format 612. For example, the treatment reference table system 150 may determine that the second treatment data set 636 indicates that the de-duplicated insurance code identifier 614 having the third format 612 was not found in the treatment classification data repository 618. In these circumstances, the treatment reference table system 150 may determine that the de-duplicated insurance code identifier 614 having the third format 612 is not valid. In addition, the treatment reference table system 150 may determine that a similar but not identical insurance code identifier exists in the treatment classification data repository 618. In these scenarios, the treatment reference table system 150 may also determine that the de-duplicated insurance code identifier 614 having the third format 612 is not valid. Further, the care reference table system 150 may determine that the second care data set 636 indicates that information related to the de-duplicated insurance code identifier 614 having the third format 612 is proprietary. As a result, the care reference table system 150 may determine that the de-duplicated insurance code identifier 614 having the third format 612 is not valid. The care reference table system 150 may determine which de-duplicated insurance code identifiers 614 are valid according to one or more criteria and generate a set of valid de-duplicated insurance code identifiers. For each valid de-duplicated insurance code identifier, the care reference table system 150 may identify and extract a respective DMS identifier 622. Each DMS identifier 622 corresponding to a valid de-duplicated insurance code identifier 614 may be used to generate a respective row of the care reference table 152. In various examples, for an invalid de-duplicated medical insurance code identifier 614, the medical reference table system 150 may generate a row in the medical reference table 152 and generate a comment indicating that the de-duplicated medical insurance code identifier 614 is invalid.In at least some scenarios, the medical reference table system 150 may generate a comment indicating why the de-duplicated medical insurance code identifier 614 is not valid (a similar medical insurance code identifier with different packaging exists in the medical classification data repository 618) or the de-duplicated medical insurance code identifier has a proprietary classification.

[0120] For each de-duplicated insurance code identifier 614 that the care reference table system 150 determines to be valid, the care reference table system 150 may determine a DMS identifier 622 that corresponds to the valid de-duplicated insurance code identifier 614. The DMS identifier 622 may then be used to generate an API request 626 according to a third scheme 638. In one or more instances, the API request 626 generated according to the third scheme 638 may include an RxCUI extracted from at least one of the first treatment data set 632 or the second treatment data set 636. The API request 626 generated according to the third scheme 638 may also include at least one of an identifier of the respective treatment data set 620 or instructions for retrieving information from the treatment classification data repository 618. The additional information may also be used to generate the API request 626 according to the third scheme 638. The API request 626 may be formatted as a Hypertext Transfer Protocol (HTTP) request.

[0121] In response to receiving an API request 626 generated according to the third scheme 638 from the data integration & analysis system 102, the therapeutic classification data management system 616 may retrieve a third therapeutic data set 640 from the therapeutic classification data repository 618. For example, the third therapeutic data set 640 may be included in an API response 628 generated by the therapeutic classification data management system 616 based on the API request 626 corresponding to the third scheme 638. In various examples, the information included in the third therapeutic data set 640 may indicate one or more classes of therapy corresponding to the one or more DMS identifiers 622 included in the API request 626 generated according to the third scheme 638. For example, the information included in the third data set 640 may include a class identifier, a name of the class, a type of the class, a source of the class, a drug term type of the therapy, a name of the therapy, or one or more combinations thereof. In at least some examples, the therapeutic reference table system 150 may analyze the third therapeutic data set 640 for one or more criteria. For example, the therapy lookup table system 150 may analyze the third therapy data set 640 to determine a therapy associated with a DMS identifier 622 that is also associated with a therapy-related term type. The therapy-related term type may indicate at least one of a composition of the therapy, a class of composition of the therapy, a dosage of the therapy, a form of the therapy (e.g., oral, infusion), a brand name of the therapy, or a synonym of the therapy.

[0122] The class information included in the third treatment data set 640 may be used to determine one or more types of treatments having respective sources. Different sources of types of treatments may indicate different information related to the treatments. In one or more examples, each source of a type of treatment may include different information regarding the treatment. For example, a first source of treatment information may indicate a pharmacological action corresponding to the treatment. In addition, a second source of treatment information may include one or more mechanisms of action of the treatment, chemical structures and classification schemes of drugs or other compositions included in the treatment, and physiological effects of drugs or compositions included in the treatment. The second source of treatment information may also include a pharmacological class (such as a United States Federal Drug Administration established pharmacological class) associated with the treatment. A third source of treatment information may indicate a group of therapies determined based on the organ or organ system on which the treatment acts. The third source of treatment information may also indicate the chemical, pharmacological, and therapeutic properties of the treatment. Additionally, a fourth source of treatment information may indicate the treatment according to the biological condition being treated, the mechanism of action of the treatment, and the chemical structure of the treatment. The fourth source of therapeutic information may also indicate the classification of the therapy and the effect of the therapy on tissues, organs and organ systems. In addition, the fourth source of therapeutic information may indicate the pharmacokinetics of the therapy, such as the absorption, distribution and elimination of the active components of the therapy.

[0123] In one or more examples, the third treatment data set 640 may be analyzed with respect to a prioritized list of sources of treatment information. For example, the treatment reference table system 150 may analyze the third treatment data set 640 according to a set of rules or protocols that analyze one or more fields of data contained in the third treatment data set 640. For example, the rules implemented by the treatment reference table system 150 may traverse one or more prescribed fields of the third treatment data set 640 to determine "whether the one or more prescribed fields contain a first identifier of a first source of treatment information that has the highest priority in the prioritized list of sources of treatment information." In the situation where the treatment reference table system 150 determines that the first identifier is within one or more prescribed fields of the third treatment data set 640, the treatment reference table system 150 may determine one or more first values ​​of one or more columns of the treatment reference table 152 of the treatment. Additionally, in a scenario where the treatment reference table system 150 determines that the first identifier is not present within one or more defined fields of the third treatment data set 640, the treatment reference table system 150 may analyze the one or more defined fields for a second source of treatment information included in the prioritized list of sources of treatment information. The second source of treatment information may be associated with a second priority and a second identifier. In the instance where the treatment reference table system 150 determines that the second identifier is present within one or more defined fields of the third treatment data set 640, the treatment reference table system 150 may determine one or more second values ​​for one or more columns of the treatment reference table 152. In instances where the treatment reference table system 150 determines that the second identifier is not present in the one or more defined fields of the third treatment data set 640, the treatment reference table system 150 may continue to analyze the one or more defined fields for the prioritized list of sources of treatment information until a source of treatment information included in the prioritized list of sources of treatment information is identified or until the treatment reference table system 150 determines that no sources of treatment information included in the prioritized list exist for the given treatment.

[0124] Based on the source of the treatment information or the absence of a source of treatment information included in the prioritized list, the treatment reference table system 150 may determine a respective value of one or more columns of the treatment reference table 152. The values ​​determined by the treatment reference table system 150 for a given treatment may be different for various sources of treatment information. For example, the treatment reference table system 150 may determine one or more first values ​​of one or more columns of the treatment reference table 152 in response to determining that the given treatment corresponds to a first source in the prioritized list of sources. In addition, the treatment reference table system 150 may determine one or more second values ​​of one or more columns of the treatment reference table 152 in response to determining that the given treatment corresponds to a second source in the prioritized list of sources. In one or more instances, the treatment reference table system 150 may determine that the one or more first values ​​corresponding to the first source of treatment information and / or the one or more second values ​​corresponding to the second source of treatment information relate to at least one of the comments column, the treatment type column, the treatment class, the treatment category, or the treatment identifier of the treatment reference table 152.

[0125] The treatment reference table system 150 may generate the treatment reference table 150 based on information obtained from the data table 602, the first treatment data set 632, the second treatment data set 636, and the third treatment data set 640. For example, the treatment reference table system 150 may generate rows of the treatment reference table 152 corresponding to at least some of the individual treatments associated with the de-duplicated treatment insurance code identifiers 614. In one or more examples, the treatment reference table system 150 may analyze the data table 602 to determine insurance code identifiers of the individual treatments and enter them into columns of the treatment reference table 152. In addition, the treatment reference table system 150 may determine individual names of the treatments, such as trade names of the treatments, by analyzing the second treatment data set 636, and populate columns of the treatment reference table 152 by using the names of the treatments. Furthermore, the treatment reference table system 150 may determine individual classes and / or categories of the treatments based on the third treatment data set 640, and populate one or more columns of the treatment reference table 152 by using the classes and / or categories. In various examples, the treatment reference table system 150 may determine the composition of individual treatments by analyzing the second treatment data set 636 and use the composition of treatments to populate the columns of the treatment reference table 152.

[0126] In one or more additional examples, the treatment reference table 152 may include at least one column that includes comments or various other information regarding an individual treatment. The treatment reference table system 150 may determine a value for the comments column by analyzing data included in at least one of the second treatment data set 636 or the third treatment data set 640. For example, the treatment reference table system 150 may determine a value for the comments column based on a source of treatment information for an individual treatment included in the treatment reference table 152. For example, the treatment reference table system 150 may determine a value for the comments column of the treatment reference table 152 by indicating a source of treatment information or by indicating a type of information associated with the source of treatment information. In one or more instances, the treatment reference table system 150 may determine that the source of information for a given treatment is a first source that includes the mechanism of action of the treatment. In these scenarios, the treatment reference table system 150 may determine a value for the comments column of an individual treatment to indicate that mechanism of action information is available for the treatment. In one or more additional examples, the treatment reference table system 150 may determine that the source of information for a given treatment is a second source that includes pharmacological class information. In these cases, the treatment reference table system 150 may determine a value in the comments column for the given treatment to indicate that pharmacological class information is available for the treatment.

[0127] In one or more other examples, the treatment reference table system 150 may determine that one or more values ​​of the columns of the treatment reference table 152 should be set to null. For example, the treatment reference table system 150 may determine the source of information corresponding to a given treatment and determine that the value of the treatment source category and / or the value of the comment of the given treatment should be set to null. In these situations, the treatment reference table system 150 may implement one or more rules or protocols that dictate that the one or more sources of treatment information correspond to a null value of the comment column. In addition, the treatment reference table system 150 may analyze at least one of the de-duplicated treatment insurance code identifier 614, the first treatment data set 632, the second treatment data set 636, or the third treatment data set 640, and determine that at least one identifier corresponding to the treatment does not exist. In these cases, the treatment reference table system 150 may determine that the value of the treatment column (such as the insurance code identifier and / or the DMS identifier 622) corresponding to the treatment identifier should be set to null.

[0128] In this manner, the treatment reference table system 150 may analyze treatment data obtained from many disparate sources and may generate the treatment reference table 152 by using certain portions of the treatment data to generate values ​​for columns of the treatment reference table 152. As a result, the data stored by the treatment reference table 152 may be used by the data integration & analysis system 102 to determine information regarding treatments provided to one or more populations of individuals in which one or more defined biological conditions exist. For example, the data integration & analysis system 102 may analyze one or more values ​​in one or more rows of the treatment reference table 152 to determine one or more insurance code identifiers that correspond to a name of the treatment or a class of treatment. In one or more examples, the data integration & analysis system 102 may analyze the values ​​of many rows of one or more data tables to determine one or more rows that include one or more insurance code identifiers, and may determine one or more identifiers of the individuals included in the one or more rows to generate a population of individuals who have received the treatment related to the biological condition. In various examples, the data integration and analysis system 102 may determine genomic information of a population of treated individuals related to a biological condition and may apply at least one of one or more statistical techniques or one or more machine learning techniques to determine one or more features of the population of individuals. In one or more examples, the one or more features may include at least one of a genetic mutation in the genome of each of the individuals in the population of individuals, a genetic mutation in cell-free deoxyribonucleic acid (DNA) in one or more samples obtained from the individuals in the population of individuals, an amount of cell-free DNA carrying the genetic mutation for each of the individuals in the population of individuals, or a change in the amount of cell-free DNA carrying the genetic mutation for each of the individuals in the population of individuals over a period of time.

[0129] 7, 8 and 9 illustrate an exemplary process for generating an integrated data repository and generating a data set used to analyze information stored by the integrated data repository. The exemplary process is illustrated as a collection of blocks in a logical flow diagram that represents a sequence of operations that may be performed in hardware, software, or a combination thereof. The blocks are referenced by numerals. In the context of software, the blocks represent computer-executable instructions stored on one or more computer-readable media, which perform the recited operations when executed by one or more processing units (such as hardware microprocessors). Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and thus any number of the described blocks may be combined in any order and / or in parallel to perform a process.

[0130] FIG. 7 is a flow diagram of an example process 700 for generating a treatment reference table including information regarding treatments provided to a patient for which one or more biological conditions may exist, according to one or more implementations. In operation 702, the process 700 may include analyzing one or more data tables including insurance claims data corresponding to an individual's treatment of a biological condition for multiple formats of insurance code identifiers. The insurance claims data may include insurance code identifiers corresponding to many different interventions and / or services provided to an individual associated with a healthcare provider. In various examples, insurance code identifiers corresponding to different treatments may have different formats. For example, insurance code identifiers corresponding to medical treatments may be formatted according to one or more NDC formats. Additionally, insurance code identifiers corresponding to medical procedures may be formatted according to one or more current procedural terminology (CPT) codes and / or one or more healthcare common procedure coding system (HCPCS) codes. Additionally, insurance code identifiers related to an individual's diagnosis for a biological condition may correspond to one or more international classification of diseases (ICD) codes. In one or more examples, the insurance code identifier may have at least a first format, a second format, and a third format. In one or more examples, the insurance code identifier may have at least one of an NDC9 format, an NDC10 format, or an NDC11 format.

[0131] At operation 704, process 700 includes determining a plurality of insurance code identifiers included in one or more data tables corresponding to respective formats of the plurality of formats. In one or more examples, the respective formats may correspond to at least one arrangement of alphanumeric characters or symbols. The arrangement of alphanumeric characters and / or symbols of the respective insurance code identifiers may be analyzed with respect to the arrangement of alphanumeric characters and / or symbols of the respective formats. In response to determining at least a threshold similarity between the insurance code identifiers and the respective insurance code identifier formats, the insurance code identifiers may be designated as having the respective formats. In various examples, a first number of insurance code identifiers may be identified as having a first format, a second number of insurance code identifiers may be identified as having a second format, and a third number of insurance code identifiers may be identified as having a third format. In one or more instances, a first group of insurance code identifiers may have an NDC9 format, a second group of insurance code identifiers may have an NDC10 format, and a third group of insurance code identifiers may have an NDC11 format.

[0132] Additionally, the process 700 may include, at operation 706, generating one or more requests of an application programming interface (API) that includes the insurance code identifier included in the plurality of insurance code identifiers. The API request may include at least one sequence of alphanumeric characters or symbols that includes at least a portion of the insurance code identifier. In various examples, the insurance code identifier may be included in various API requests that may be used to retrieve various information from a data repository. For example, at least a portion of the insurance code identifier may be included in a first API request to obtain another version of the insurance code identifier having a different format. For example, an API request may be generated by using the insurance code identifier having an NDC9 format or an NDC10 format to retrieve a version of the insurance code identifier having an NDC11 format. In one or more additional examples, the insurance code identifier may be used to generate a second API request that may be used to retrieve one or more identifiers of a procedure associated with the insurance code identifier. In one or more implementations, the insurance code identifier may be used to generate an API request to look up a trade name for the treatment, a standardized identifier (such as an Rx Norm concept unique identifier (RxCUI)) may be used by a data repository to store information related to the treatment, or both. In one or more other examples, the insurance code identifier may be used to generate an API request to look up a source of information related to the treatment corresponding to the insurance code identifier and / or to look up a category related to the treatment corresponding to the insurance code identifier.

[0133] At operation 708, process 700 may include retrieving one or more data files including information corresponding to the insurance code identifier in response to the one or more API requests, and at operation 710, process 700 may include extracting a therapy identifier from the data files. In one or more examples, the therapy identifier may include a name of the therapy (e.g., a trade name of the therapy or a name of a composition of the therapy). In one or more additional examples, the therapy identifier may include an RxCUI of the therapy.

[0134] Additionally, the process 700 may include, at operation 712, generating an additional data table having a row indicating that the insurance code identifier corresponds to an identifier of the treatment. In one or more examples, the additional data table may have many rows with each row corresponding to a single treatment. In various examples, the treatments included in the additional data table may be used to treat a biological condition. The additional data table may also include many columns having values ​​corresponding to information related to each treatment. In one or more examples, the additional data table may include a column including an insurance code identifier of the treatment and another column including a treatment identifier obtained from the data repository. In one or more instances, the row corresponding to the treatment may have a first column corresponding to an identifier having an NDC9 format, an NDC10 format, or an NDC11 format, and a second column corresponding to an RxCUI of the treatment. The row corresponding to the treatment may also include a third column including a name of the treatment. The additional columns of the additional data table may include values ​​corresponding to a source of information about the treatment, a category of the treatment, a condition of the treatment, a date the treatment was actively used, or one or more combinations thereof.

[0135] FIG. 8 is a flow diagram of an example process 800 for determining a drug identifier corresponding to an insurance code identifier by using one or more application programming interface (API) requests, according to one or more implementations. The process 800 may include, at operation 802, generating one or more data tables including insurance claims information for a number of patients. The one or more data tables may include information obtained from a data repository that stores health insurance claims data (such as the health insurance claims data repository 106 of FIG. 1). At operation 804, the process 800 may include determining one or more columns of the one or more data tables including identifiers having an NDC format. In various examples, the one or more data tables may be arranged such that the NDC identifiers are present in one or more defined columns of the one or more data tables. In one or more additional examples, values ​​of columns of the one or more data tables may be analyzed to identify values ​​of columns having one or more formats that correspond to the NDC identifiers. The one or more formats may include at least one of an NDC9 format, an NDC10 format, or an NDC11 format. Each identifier having an NDC format can correspond to a treatment for at least one biological condition. In one or more instances, the treatment can include a pharmaceutical agent capable of treating one or more biological conditions.

[0136] Additionally, process 800 may include, at operation 806, removing duplicate identifiers having an NDC format to generate a dataset including deduplicated identifiers having an NDC format. Further, at operation 808, process 800 may include analyzing identifiers having an NDC format included in the dataset to determine a respective format of the identifiers. In various examples, the NDC identifiers included in the deduplicated NDC identifier dataset may be grouped according to the NDC format of the identifiers. For example, a first set of identifiers having a first NDC format may be included in a first group of identifiers, a second set of identifiers having a second NDC format may be included in a second group of identifiers, and a third set of identifiers having a third NDC format may be included in a third group of identifiers.

[0137] In the situation where the NDC identifier has a first format, process 800 may move to process 810 where one or more first API requests are generated using the NDC identifier. In at least some examples, the first format may include an NDC11 format. The one or more API requests may be used to retrieve information from one or more datasets. The information obtained using the one or more API requests may include at least one of a source of the NDC identifier, a state of the treatment corresponding to the NDC identifier, a start date when the NDC identifier was activated, an end date when the NDC identifier is no longer active, one or more names of the treatment corresponding to the NDC identifier, or additional information regarding the NDC identifier. In one or more examples, the information obtained using the one or more API requests may include an additional identifier of the treatment corresponding to the initial NDC identifier. In operation 812, process 800 may include extracting the additional identifier from a response to the one or more first API requests. The additional identifier may be assigned to the treatment and / or the NDC identifier by a third party. In one or more examples, the additional identifier may be unique with respect to other identifiers of the treatment assigned by a third party. In one or more instances, the additional identifier may include the RxCUI of the treatment.

[0138] Process 800 may include, at operation 814, generating a lookup table including rows with identifiers for the treatments. The lookup table may include a number of rows with individual rows of the rows corresponding to individual treatments. The lookup table may also include a number of columns with values ​​having information about the individual treatments. For example, the lookup table may include columns with values ​​indicating one or more categories of the treatment, one or more additional identifiers for the treatment, a condition of the treatment, a composition of the treatment, one or more combinations thereof, and so on.

[0139] In a scenario where the result of operation 808 indicates that the format of the NDC identifier corresponds to a second format, process 800 may proceed from operation 808 to operation 816. Operation 816 may include modifying the format of the NDC identifier to generate a modified NDC identifier. In one or more examples, the second format may correspond to an NDC9 format or an NDC10 format. In various examples, modifying the NDC identifier may include removing one or more alphanumeric characters or symbols from the NDC identifier. Additionally, modifying the NDC identifier may include adding one or more alphanumeric characters or symbols to the NDC identifier. In one or more examples, the NDC identifier may be modified by adding a dash symbol before the last two digits of the NDC identifier. In one or more additional examples, the NDC identifier may be modified by adding one or more zeros to one or more segments of the NDC identifier. In one or more other examples, the NDC identifier may be modified by removing one or more alphanumeric characters and / or symbols from the NDC identifier.

[0140] At operation 818, the process 800 may include generating one or more second API requests by using the modified NDC identifier. The one or more second API requests may be used to obtain information from an additional data set. The additional data set may include many different identifiers related to the initial NDC identifier having the second format. For example, the additional data set may include additional identifiers having the first format. For example, the NDC identifier may have an NDC9 format or an NDC10 format, and the additional identifiers may have an NDC11 format. At operation 824, the process 800 may include extracting the additional NDC identifiers having the first format from a response to the one or more second API requests. Based on the additional NDC identifier having the first format, the process may move to operation 810, where one or more first API requests may be generated using the additional NDC identifier having the first format, and may proceed to operation 812 and operation 814, where a treatment reference table is generated using information obtained using the one or more first API requests generated based on the additional identifier having the first format.

[0141] FIG. 9 is a flow diagram of an example process 900 for determining a class corresponding to a therapy identifier and for including information related to the class in a reference data table that includes the therapy identifier, according to one or more implementations. In operation 902, the process 900 may include generating one or more API requests that include the therapy identifier. In one or more examples, the therapy identifier may correspond to an identifier obtained from a data repository that stores information about many therapies. In one or more examples, the therapy identifier may include an RxCUI for the therapy. The one or more API requests may be used to obtain a specific set of information from the data repository. The one or more API requests may have a format and / or structure that differs from the format and / or structure of other API requests. The one or more API requests may be used to search for information stored in a respective storage location.

[0142] At 904, the process 900 may include analyzing one or more first fields of the output file 906 to determine a grouping of the therapy identifier. In one or more examples, in response to one or more API requests, an output file may be received that includes a number of fields and respective values ​​for the one or more fields. In various examples, the output file 906 may include a field that indicates at least one of an identifier of a class corresponding to the therapy identifier, an identifier of a type of class corresponding to the therapy identifier, or a source of class relationship or other information corresponding to the therapy identifier. In one or more additional examples, the output file 906 may include a field having a value that corresponds to a name of a therapy related to the therapy identifier and / or a therapy term type corresponding to the therapy identifier.

[0143] The process 900 may also include, in operation 908, analyzing one or more second fields of the output file 906 according to a set of rules and for a prioritized list 910 of sources of classes of information. The prioritized list 910 may indicate ordinal sources of the classes of information of the therapeutic identifier. In a scenario where the RxNORM database is queried, the prioritized list may include therapeutic class related sources (such as Anatomical Therapeutic Chemical (ATC), Food and Drug Administration Structured Product Labeling (FDASPL), Federal Medication Terminologies Subject Matter Expert (FMTSME), Medication Reference Terminology (MEDRT), Medical Subject Headings (MeSH), RxNorm (by the National Library of Medicine), or SNOMEDCT (by the International Health Terminology Standards Development Organization)). The values ​​of the fields of the output file 906 corresponding to the class of information may be analyzed with respect to the priority list 910 to determine the source of the class of information (e.g., therapeutic class related information) in operation 912. The values ​​of the fields of the output file 906 corresponding to the class of information of the therapeutic identifier may also be used to determine the values ​​of one or more columns of a row of the therapeutic identifier lookup table 916 based on the source of the class of information in operation 914. For example, the values ​​of the fields of the output file 906 corresponding to the source of the class information of the therapeutic identifier may be used to determine the source of the class information of the therapeutic identifier. In one or more instances, the class name of the therapeutic identifier may be "pain" and the class type may be "disease". The source of the class information may indicate the mechanism of action by which the therapeutic functions or the pharmacological class associated with the therapeutic. The source of the class information may also indicate the chemical structure of one or more compositions of the therapeutic and / or pharmacokinetic information associated with the therapeutic.

[0144] In the example of FIG. 9 , the source of the therapeutic class information may be a first source 918 corresponding to a value 920. In operation 922, the value 920 may be used to populate a row of the lookup table 916. That is, for each source of therapeutic class information, a value for each of the columns of the lookup table 916 corresponding to the therapeutic may be determined. For example, a first column may be filled with a name of the first source 918, and a second column may be filled with a type associated with the first source 918. Additionally, a value associated with a third column corresponding to a comment associated with the therapeutic may be filled in accordance with the value 920. In various examples, the comment may indicate whether the first source 918 includes therapeutic mechanism of action information or pharmacological class information.

[0145] Fig. 10 illustrates a diagrammatic representation of a machine 1000 in the form of a computer system on which a set of instructions may be executed to cause the machine 1000 to perform any one or more of the methodologies discussed herein, according to one example according to an exemplary implementation. Specifically, Fig. 10 illustrates a diagrammatic representation of a machine 1000 in the form of an exemplary computer system on which instructions 1002 (e.g., software, programs, applications, applets, apps, or other executable code) may be executed to cause the machine 1000 to perform any one or more of the methodologies discussed herein. For example, the instructions 1002 may cause the machine 1000 to implement the architectures and frameworks 100, 200, 300, 400, 500, 600 described with respect to Figs. 1, 2, 3, 4, 5, and 6, respectively, and to perform the methods 700, 800, 900 described with respect to Figs. 7, 8, and 9, respectively.

[0146] The instructions 1002 transform a generic unprogrammed machine 1000 into a specific machine 1000 programmed to perform the described and illustrated functions in the described manner. In alternative implementations, the machine 1000 may operate as a stand-alone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1000 may operate in the capacity of a server machine or a client machine in a server / client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1000 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cell phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially executing instructions 1002 or otherwise prescribing actions to be taken by the machine 1000. Additionally, although only a single machine 1000 is shown, the term "machine" shall also be taken to include a collection of machines 1000 that individually or collectively execute instructions 1002 to perform any one or more of the methodologies discussed herein.

[0147] An example of the machine 1000 may include logic, one or more components, circuits (e.g., modules), or mechanisms. A circuit is a tangible entity configured to perform some operations. In an example, a circuit may be arranged in a prescribed manner (e.g., internally or relative to an external entity, such as another circuit). In an example, one or more computer systems (e.g., stand-alone, client, or server computer systems) or one or more hardware processors (processors) may be configured with software (e.g., instructions, application sections, or applications) as a circuit that operates to perform some operations described herein. In an example, the software may reside (1) on a non-transitory machine-readable medium or (2) in a transmission signal. In an example, the software, when executed by the inherent hardware of the circuit, causes the circuit to perform some operations.

[0148] In one example, a circuit may be implemented mechanically or electronically. For example, a circuit may include dedicated circuitry or logic (including, for example, a dedicated processor, a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC)) that is specifically configured to perform one or more techniques as discussed above. In one example, a circuit may include programmable logic (such as circuitry contained within a general purpose processor or other programmable processor) that can be temporarily configured (e.g., by software) to perform certain operations. It will be recognized that the decision to implement a circuit mechanically (e.g., in dedicated and permanently configured circuitry) or with temporarily configured (e.g., software configured) circuitry may be driven by cost and time considerations.

[0149] Thus, the term "circuitry" is understood to encompass tangible entities (e.g., entities that are physically constructed and permanently configured (e.g., hardwired) or temporarily (e.g., transiently) configured (e.g., programmed) to operate in a specified manner or to perform specified operations). In one example, given multiple temporarily configured circuits, each of these circuits need not be configured or instantiated at any one time. For example, if a circuit includes a general-purpose processor configured via software, the general-purpose processor may be configured as each of the different circuits at various times. Thus, the software may, for example, configure the processor to configure a particular circuit at one time and to configure a different circuit at a different time.

[0150] In one example, a circuit may provide information to and receive information from other circuits. In this example, a circuit may be considered to be communicatively coupled to one or more other circuits. When multiple such circuits are present simultaneously, communication may be achieved through signal transmission (e.g., over appropriate circuits and buses) connecting the circuits. In implementations in which multiple circuits are configured or instantiated at different times, communication between such circuits may be achieved, for example, through storage and retrieval of information in a memory structure accessible to the multiple circuits. For example, one circuit may perform an operation and store the output of the operation in a communicatively coupled memory device. Another circuit may then subsequently access the memory device to retrieve and process the stored output. In one example, some circuits may be configured to initiate or receive communication with an input or output device and may operate on a resource (e.g., a collection of information).

[0151] Various operations of example methods described herein may be performed, at least in part, by one or more processors (e.g., by software) that are temporarily or permanently configured to perform the associated operations. Such processors, whether temporarily or permanently configured, may constitute processor-implemented circuitry that operates to perform one or more operations or functions. In one example, circuitry referenced herein may include processor-implemented circuitry.

[0152] Similarly, the methods described herein may be at least partially processor-implemented. For example, at least some of the operations of the methods may be performed by either a processor or processor-implemented circuitry. Performance of some of the operations may be distributed among one or more processors that are not only resident within a single machine, but also deployed across many machines. In one example, a processor or processors may be located in a single location (e.g., in a home environment, in an office environment, or as a server farm), while in other examples, processors may be distributed across many locations.

[0153] The one or more processors may also operate to facilitate performance of related operations within a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a collection of computers (such as an example of a machine that includes a processor), and these operations may be accessible over a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application program interfaces (APIs)).

[0154] Exemplary implementations (e.g., devices, systems, or methods) can be implemented in digital electronic circuitry, in computer hardware, in firmware, in software, or in any combination of these. Exemplary implementations can be implemented using a computer program product (e.g., a computer program, a programmable processor, a computer program tangibly embodied in an information carrier or machine-readable medium for execution by or to control the operation of a data processing apparatus, such as a computer or multiple computers).

[0155] A computer program may be written in any type of programming language (including compiled or interpreted languages) and may be deployed in any form (e.g., as a stand-alone program, or as a software module, subroutine, or other unit suitable for use within a computing environment). A computer program may be deployed to be executed on one computer or multiple computers at one site, or may be distributed across multiple sites and interconnected by a communication network.

[0156] In one example, some operations may be performed by one or more programmable processors executing computer programs to perform some functions by operating on input data and generating output data. Examples of method operations may also be performed by, and example apparatus may be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or application specific integrated circuit (ASIC).

[0157] A computing system may include clients and servers. Clients and servers are typically remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client / server relationship to each other. It will be understood that in an implementation deploying a programmable computing system, both hardware and software architectures require consideration. In particular, it will be understood that the choice of whether to implement some functions in permanently configured hardware (e.g., ASICs), temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of permanently and temporarily configured hardware may be a design choice. Set forth below are hardware (e.g., computing device 800) and software architectures that may be deployed in an exemplary implementation.

[0158] In one example, machine 1000 may operate as a stand-alone device or may be connected (eg, networked) to other machines.

[0159] In a networked deployment, the machine 1000 may operate in the capacity of either a server or a client machine in a server / client network environment. In one example, the machine 1000 may act as a peer machine in a peer-to-peer (or other distributed) network environment. The machine 1000 may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (either sequentially or non-sequentially) that specify actions to be taken (e.g., performed) by the computing device 800. Furthermore, although only a single machine 1000 is shown, the term "computing device" shall also be deemed to include any collection of machines that individually or collectively execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein.

[0160] The exemplary machine 1000 may include a processor 1004 (e.g., a central processing unit CPU), a graphics processing unit (GPU), or both), a main memory 1006, and a static memory 1008, some or all of which may communicate with each other via a bus 1010. The machine 1000 further includes a display unit 1012, an alphanumeric input device 1014 (e.g., a keyboard), and a user interface (UI) navigation device 1016 (e.g., a mouse). In one example, the display unit 1012, the input device 1014, and the UI navigation device 1016 may be touch screen displays. The machine 1000 may additionally include a storage device (e.g., a drive unit) 1018, a signal generating device 1020 (e.g., a speaker), a network interface device 1022, and one or more sensors 1024 (such as a global positioning system (GPS) sensor, a compass, an accelerometer, or another sensor).

[0161] The storage device 1018 may include a machine-readable medium 1026 having stored thereon one or more sets of data structures or instructions 1002 (e.g., software) that embody or are utilized by any one or more of the methodologies or functions described herein. The instructions 1002 may also reside, completely or at least partially, within the static memory 1008, within the main memory 806, or within the processor 1004 during execution by the computing device 800. In one example, one or any combination of the processor 1004, the main memory 1006, the static memory 1008, or the storage device 1018 may constitute a machine-readable medium.

[0162] Although the machine-readable medium 1026 is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) configured to store one or more instructions 1002. The term "machine-readable medium" may also be interpreted to include any tangible medium that can store, encode, or carry instructions for execution by a machine and cause the machine to perform any one or more of the methodologies of this disclosure. Thus, the term "machine-readable medium" may be interpreted to include, but is not limited to, solid-state memory, as well as optical and magnetic media. Particular examples of machine-readable media may include non-volatile memory, including, by way of example, semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable PROM (EEPROM)), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0163] The instructions 1002 may further be transmitted or received over a communications network 1028 using a transmission medium via a network interface device 1022 utilizing any one of a number of transport protocols (e.g., Frame Relay, IP, TCP, UDP, HTTP, etc.). Exemplary communications networks may include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile telephone networks (e.g., cellular networks), plain old telephone (POTS) networks, as well as wireless data communications networks (e.g., the IEEE 802.11 family of standards known as Wi-Fi®, the IEEE 802.16 family of standards known as WiMax®), peer-to-peer (P2P) networks, among others. The term "transmission media" is intended to include any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communications signals or other intangible media for facilitating communication of such software.

[0164] Some implementations are described as numbered examples (Example 1, 2, 3, etc.) These are provided as examples only and thus do not limit the technology disclosed herein.

[0165] Aspect 1. A method including: analyzing, with a computing system including processing circuitry and a memory, one or more data tables stored by a data repository related to one or more criteria, the one or more data tables including insurance claims data corresponding to one or more treatments provided to an individual in a biological condition, the one or more criteria indicating one or more formats of an insurance code identifier; determining, by the computing system, a plurality of insurance code identifiers included in the one or more data tables that correspond to the one or more criteria, analyzing, by the computing system, a subset of the plurality of insurance code identifiers to determine one or more formats of the subset of the plurality of insurance code identifiers; generating, by the computing system, one or more requests for an application programming interface (API) call including insurance code identifiers of the subset of the plurality of insurance code identifiers; obtaining, by the computing system and in response to the one or more requests of the API, a data file including information corresponding to the insurance code identifiers; extracting, by the computing system, a treatment identifier from the data file; and generating, by the computing system, an additional data table having a row indicating that the insurance code identifier corresponds to the treatment identifier.

[0166] Example 2. The method of example 1, comprising determining, by a computing system, one or more columns of one or more data tables that include insurance code identifiers corresponding to treatment of the individual.

[0167] Aspect 3. The method of aspect 1 or 2, including determining, by the computing system, that a first insurance code identifier of the plurality of insurance code identifiers is a duplicate of a second insurance code identifier of the plurality of insurance code identifiers; and removing, by the computing system, the second insurance code identifier from the plurality of insurance code identifiers to generate the plurality of insurance code identifiers.

[0168] Aspect 4. The method of any one of aspects 1-3, wherein the plurality of health code identifiers corresponds to a plurality of National Drug Code (NDC) identifiers.

[0169] Aspect 5. The method of any one of Aspects 1-4, including identifying, by a computing system, a first insurance code identifier of a subset of the plurality of insurance code identifiers; determining, by the computing system, that the first insurance code identifier corresponds to a first format of insurance code identifiers; and modifying, by the computing system, the first insurance code identifier to generate a second insurance code identifier that corresponds to a second format of insurance code identifiers.

[0170] Aspect 6. The method of aspect 5, including generating, by the computing system, one or more additional requests of the additional API including the second insurance code identifier; obtaining, by the computing system and in response to the one or more additional calls of the additional API, an additional data file including information corresponding to a third insurance code identifier; and extracting, by the computing system, the insurance code identifier from the additional data file.

[0171] Aspect 7. The method of any one of Aspects 1-4, comprising generating, by a computing system, one or more additional calls of the API including the additional insurance code identifiers; obtaining, by the computing system and in response to the one or more additional calls of the API, an additional data file including additional information corresponding to the additional insurance code identifiers; and determining, by the computing system, that the at least one valid insurance code identifier is not included in the additional information.

[0172]

[0023] Aspect 8. The method of any one of Aspects 1-7, comprising generating, by a computing system, one or more additional calls to the API including additional insurance code identifiers of a subset of the plurality of insurance code identifiers; obtaining, by the computing system and in response to the one or more additional calls to the API, additional data files including additional information corresponding to the additional insurance code identifiers; analyzing, by the computing system, the additional information with respect to one or more additional criteria; and determining, by the computing system, that the additional information does not include the at least one medication identifier.

[0173] Aspect 9. The method of any one of aspects 1-8, comprising generating, by the computing system, one or more additional calls to an additional API that includes the treatment identifier; obtaining, by the computing system and in response to the one or more additional calls, an additional data file that includes additional information corresponding to the medication identifier; and determining, by the computing system and based on the additional information, a class of medication corresponding to the medication identifier.

[0174] Embodiment 10. The method of any one of embodiments 1-9, comprising: determining, by the computing system, a source of the drug class by analyzing, by the computing system, a source of the drug class for the prioritized number of drug sources and based on the additional information; extracting, by the computing system, at least a portion of the additional information from one or more fields of the data file; and adding, by the computing system, at least a portion of the additional information to a row of the database table.

[0175] Embodiment 11. The method of embodiment 10, wherein the prioritized number of drug sources comprises a first source of a class of drug having a first priority and a second source of a class of drug having a second priority lower than the first priority.

[0176]

[0023] Example 12. The method of example 10 or 11, comprising: analyzing, by a computing system, the additional information regarding a first source of the class of drugs; determining, by the computing system, that the additional information does not include a source of the class of drugs corresponding to the first class; analyzing, by the computing system, the additional information regarding a second source of the class of drugs; and extracting, by the computing system, at least a portion of the additional information from an additional data file related to the second source of the class of drugs.

[0177]

[0023] Embodiment 13. The method of any one of embodiments 1-12, comprising receiving, by a computing system, a request to identify a group of treated individuals responsive to a biological condition existing for the group of individuals, the request including a name of the treatment or a class of treatment; analyzing, by the computing system, one or more values ​​of one or more rows of the additional data table to determine one or more insurance code identifiers corresponding to the name of the treatment or class of treatment; analyzing, by the computing system, values ​​of many rows of the one or more data tables to determine one or more rows including the one or more insurance code identifiers; and determining, by the computing system, one or more identifiers of individuals included in the one or more rows to generate a population of treated individuals related to the biological condition.

[0178]

[0023] Example 14. The method of example 13, comprising: determining, by a computing system, genomic information of a population of treated individuals related to a biological condition; and applying, by the computing system, at least one of one or more statistical techniques or one or more machine learning techniques to determine one or more features of the population of individuals.

[0179]

[0023] Example 15. The method of example 14, wherein the one or more features comprise at least one of: a genetic mutation in the genome of each of the individuals in the population of individuals; a genetic mutation in cell-free deoxyribonucleic acid (DNA) in one or more samples obtained from the individuals in the population of individuals; an amount of cell-free DNA with the genetic mutation for each individual in the population of individuals; or a change in the amount of cell-free DNA with the genetic mutation for each individual in the population of individuals over a period of time.

[0180] Aspect 16. A system including one or more hardware processing units; and one or more computer readable storage media storing computer executable instructions that, when executed by the one or more hardware processing units, cause the system to perform operations including: analyzing one or more data tables stored by a data repository for one or more criteria, the one or more data tables including insurance claims data corresponding to one or more treatments provided to an individual in a biological condition, the one or more criteria indicating one or more formats of an insurance code identifier; determining a plurality of insurance code identifiers included in the one or more data tables that correspond to the one or more criteria, analyzing a subset of the plurality of insurance code identifiers to determine one or more formats of the subset of the plurality of insurance code identifiers; generating one or more requests for an application programming interface (API) call including the insurance code identifiers of the subset of the plurality of insurance code identifiers; obtaining a data file including information corresponding to the insurance code identifiers in response to the one or more requests for the API; extracting the treatment identifier from the data file; and generating an additional data table having a row indicating that the insurance code identifier corresponds to the treatment identifier.

[0181] Aspect 17. The system of aspect 16, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining one or more columns of one or more data tables that include insurance code identifiers corresponding to the individual's treatment.

[0182] Aspect 18. The system of aspect 16 or 17, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including: determining that a first insurance code identifier of the plurality of insurance code identifiers is a duplicate of a second insurance code identifier of the plurality of insurance code identifiers; and removing the second insurance code identifier from the plurality of insurance code identifiers to generate the plurality of insurance code identifiers.

[0183]

[0023] Aspect 19. The system of any one of aspects 16-18, wherein the plurality of health code identifiers correspond to a plurality of National Drug Code (NDC) identifiers.

[0184] Aspect 20. The system of any one of Aspects 16-19, wherein the one or more computer-readable storage media store additional computer-executable instructions which, when executed by the one or more hardware processing units, cause the system to perform additional operations including: identifying a first insurance code identifier of a subset of the plurality of insurance code identifiers; determining that the first insurance code identifier corresponds to a first format of insurance code identifiers; and modifying the first insurance code identifier to generate a second insurance code identifier that corresponds to a second format of insurance code identifiers.

[0185] Aspect 21. The system of Aspect 20, wherein the one or more computer-readable storage media store additional computer-executable instructions which, when executed by one or more hardware processing units, cause the system to perform additional operations including: generating one or more additional requests of an additional API including a second insurance code identifier; obtaining an additional data file including information corresponding to a third insurance code identifier in response to one or more additional calls of the additional API; and extracting the insurance code identifier from the additional data file.

[0186] Aspect 22. A system as described in any one of aspects 16-19, wherein one or more computer-readable storage media store additional computer-executable instructions, which when executed by one or more hardware processing units, cause the system to perform additional operations including: generating one or more additional calls to an API including the additional insurance code identifier; obtaining additional data files including additional information corresponding to the additional insurance code identifier in response to the one or more additional calls to the API; and determining that at least one valid insurance code identifier is not included within the additional information.

[0187] Aspect 23. The system of any one of aspects 16-22, wherein the one or more computer-readable storage media store additional computer-executable instructions which, when executed by one or more hardware processing units, cause the system to perform additional operations including: generating, by the computing system, one or more additional calls to the API including additional insurance code identifiers of a subset of the plurality of insurance code identifiers; obtaining, in response to the one or more additional calls to the API, additional data files including additional information corresponding to the additional insurance code identifiers; analyzing the additional information with respect to one or more additional criteria; and determining that the additional information does not include at least one medication identifier.

[0188] Aspect 24. The system of any one of Aspects 16-23, wherein one or more computer-readable storage media store additional computer-executable instructions which, when executed by one or more hardware processing units, cause the system to perform additional operations including: generating one or more additional calls to an additional API including the treatment identifier; obtaining an additional data file including additional information corresponding to the drug identifier in response to the one or more additional calls; and determining a class of drug corresponding to the drug identifier based on the additional information.

[0189] Embodiment 25. The system of any one of embodiments 16-24, wherein the one or more computer readable storage media store additional computer executable instructions which, when executed by the one or more hardware processing units, cause the system to perform additional operations including: determining a source of the drug class based on the additional information by analyzing the source of the drug class for a prioritized number of drug sources; extracting at least a portion of the additional information from one or more fields of a data file; and adding at least a portion of the additional information to a row of a database table.

[0190] Embodiment 26. The system of embodiment 25, wherein the prioritized number of drug sources includes a first source of a class of drug having a first priority and a second source of a class of drug having a second priority lower than the first priority.

[0191] Aspect 27. The system of aspect 25 or 26, wherein the one or more computer readable storage media store additional computer executable instructions which, when executed by the one or more hardware processing units, cause the system to perform additional operations including: analyzing the additional information regarding a first source of the drug class; determining that the additional information does not include a source of the drug class that corresponds to the first class; analyzing the additional information regarding a second source of the drug class; and extracting at least a portion of the additional information from an additional data file related to the second source of the drug class.

[0192] Example 28. The system of any one of Examples 16-27, wherein the one or more computer readable storage media store additional computer executable instructions which, when executed by the one or more hardware processing units, cause the system to perform additional operations including: receiving a request to identify a group of treated individuals responsive to that a biological condition exists for the group of individuals, the request including a name of the treatment or a class of treatment, receiving: analyzing one or more values ​​in one or more rows of the additional data table to determine one or more insurance code identifiers corresponding to the name of the treatment or class of treatment; analyzing values ​​of many rows of the one or more data tables to determine one or more rows including the one or more insurance code identifiers; and determining one or more identifiers of individuals included in the one or more rows to generate a population of treated individuals related to the biological condition.

[0193] Aspect 29. The system of aspect 28, wherein the one or more computer-readable storage media store additional computer-executable instructions which, when executed by the one or more hardware processing units, cause the system to perform additional operations including: determining genomic information of a population of treated individuals related to a biological condition; and applying at least one of one or more statistical techniques or one or more machine learning techniques to determine one or more features of the population of individuals.

[0194]

[0036] Example 30. The system of example 29, wherein the one or more features include at least one of: a genetic mutation in the genome of each of the individuals in the population of individuals; a genetic mutation in cell-free deoxyribonucleic acid (DNA) in one or more samples obtained from the individuals in the population of individuals; an amount of cell-free DNA carrying the genetic mutation in each individual in the population of individuals; or a change in the amount of cell-free DNA carrying the genetic mutation in each individual in the population of individuals over a period of time.

[0195] Aspect 31. One or more non-transitory computer readable storage media storing computer executable instructions, the additional computer executable instructions when executed by one or more hardware processing units, causing the system to perform additional operations including: analyzing one or more data tables stored by the data repository for one or more criteria, the one or more data tables including insurance claims data corresponding to one or more treatments provided to an individual in a biological condition, the one or more criteria indicating one or more formats of an insurance code identifier; determining a plurality of insurance code identifiers included in the one or more data tables that correspond to the one or more criteria, analyzing a subset of the plurality of insurance code identifiers to determine one or more formats of the subset of the plurality of insurance code identifiers; generating one or more requests for an application programming interface (API) call including insurance code identifiers of the subset of the plurality of insurance code identifiers; obtaining a data file including information corresponding to the insurance code identifiers in response to the one or more requests for the API; extracting the treatment identifier from the data file; and generating an additional data table having a row indicating that the insurance code identifier corresponds to the treatment identifier.

[0196] Aspect 32. The one or more non-transitory computer readable media of aspect 31 comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining one or more columns of one or more data tables that include an insurance code identifier corresponding to a treatment for the individual.

[0197] Aspect 33. The one or more non-transitory computer-readable media of aspect 31 or 32 comprising additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including: determining that a first insurance code identifier of the plurality of insurance code identifiers is a duplicate of a second insurance code identifier of the plurality of insurance code identifiers; and removing the second insurance code identifier from the plurality of insurance code identifiers to generate the plurality of insurance code identifiers.

[0198]

[0023] Aspect 34. The one or more non-transitory computer readable media of any one of aspects 31-33, wherein the plurality of insurance code identifiers corresponds to a plurality of National Drug Code (NDC) identifiers.

[0199] Aspect 35. The one or more non-transitory computer readable media of any one of aspects 31-34 including additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: identifying a first insurance code identifier of a subset of the plurality of insurance code identifiers; determining that the first insurance code identifier corresponds to a first format of insurance code identifiers; and modifying the first insurance code identifier to generate a second insurance code identifier that corresponds to a second format of insurance code identifiers.

[0200] Aspect 36. The one or more non-transitory computer readable media of aspect 35 comprising additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: generating one or more additional requests of the additional API including the second insurance code identifier; obtaining an additional data file in response to the one or more additional calls of the additional API including information corresponding to the third insurance code identifier; and extracting the insurance code identifier from the additional data file.

[0201]

[0023] Aspect 37. The one or more non-transitory computer readable media of any one of Aspects 31-34, comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: generating one or more additional calls to the API including the additional insurance code identifiers; obtaining additional data files in response to the one or more additional calls to the API including additional information corresponding to the additional insurance code identifiers; and determining that the at least one valid insurance code identifier is not included within the additional information.

[0202]

[0023] Embodiment 38. The one or more non-transitory computer readable media of any one of embodiments 31-37, comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: generating one or more additional invocations of the API including additional insurance code identifiers of a subset of the plurality of insurance code identifiers; obtaining additional data files in response to the one or more additional invocations of the API including additional information corresponding to the additional insurance code identifiers; analyzing the additional information with respect to one or more additional criteria; and determining that the additional information does not include the at least one medication identifier.

[0203]

[0023] Embodiment 39. The one or more non-transitory computer readable media of any one of embodiments 31-38, comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: generating one or more additional calls to an additional API that includes the treatment identifier; obtaining an additional data file in response to the one or more additional calls, the additional data file including additional information corresponding to the medication identifier; and determining a class of medication corresponding to the medication identifier based on the additional information.

[0204]

[0023] Embodiment 40. The one or more non-transitory computer readable media of any one of embodiments 31-39, comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: determining a source of the drug class based on the additional information by analyzing the source of the drug class for a prioritized number of drug sources; extracting at least a portion of the additional information from one or more fields of a data file; and adding at least a portion of the additional information to a row of a database table.

[0205] Embodiment 41. The one or more non-transitory computer-readable media of embodiment 40, wherein the prioritized number of drug sources includes a first source of a class of drug having a first priority and a second source of a class of drug having a second priority lower than the first priority.

[0206]

[0036] Aspect 42. The one or more non-transitory computer readable media of aspect 40 or 41 comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: analyzing the additional information regarding a first source of the class of drugs; determining that the additional information does not include a source of the class of drugs that corresponds to the first class; analyzing the additional information regarding a second source of the class of drugs; and extracting at least a portion of the additional information from an additional data file related to the second source of the class of drugs.

[0207]

[0023] Embodiment 43. The non-transitory computer readable medium of any one of embodiments 31-42 comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: receiving a request to identify a group of treated individuals responsive to a biological condition existing for the group of individuals, the request including a name of the treatment or a class of treatment, receiving: analyzing one or more values ​​in one or more rows of the additional data table to determine one or more insurance code identifiers corresponding to the name of the treatment or the class of treatment; analyzing values ​​of many rows of the one or more data tables to determine one or more rows including the one or more insurance code identifiers; and determining one or more identifiers of individuals included in the one or more rows to generate a population of treated individuals related to the biological condition.

[0208] Aspect 44. The one or more non-transitory computer readable media of aspect 43, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including: determining genomic information of a population of treated individuals related to a biological condition; and applying at least one of one or more statistical techniques or one or more machine learning techniques to determine one or more features of the population of individuals.

[0209]

[0036] Aspect 45. The one or more non-transitory computer readable media of aspect 44, wherein the one or more features comprise at least one of: a genetic mutation in a genome of each of the individuals in the population of individuals; a genetic mutation in cell-free deoxyribonucleic acid (DNA) in one or more samples obtained from the individuals in the population of individuals; an amount of cell-free DNA with the genetic mutation for each individual in the population of individuals; or a change in the amount of cell-free DNA with the genetic mutation for each individual in the population of individuals over a period of time.

[0210] As used herein, a component may refer to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that provide partitioning or modularization of specific processing or control functionality. A component may be combined with other components through the component's interfaces to perform machine processes. A component may be a packaged functional hardware unit designed for use with other components, and a part of a program that typically performs a specific function of related functionality. A component may constitute either a software part (e.g., code embodied on a machine-readable medium) or a hardware part. A "hardware part" may consist of a tangible unit capable of performing some operations and may be arranged in some physical manner. In various exemplary implementations, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware parts of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware part that operates to perform some operations described herein.

[0211] It should be understood that the individual steps used in the methods of the present teachings can be performed in any order and / or simultaneously so long as the present teachings are still operable. It should be further understood that "the apparatus and methods of the present teachings can include any number or all of the described implementations so long as the present teachings are still operable."

[0212] The various steps of the methods disclosed herein or steps performed by the systems disclosed herein may be performed at the same time or at different times, and / or in the same geographic location or in different geographic locations (e.g., countries). The various steps of the methods disclosed herein may be performed by the same person or by different people.

[0213] Various implementations of systems, devices, and methods have been described herein. These implementations are provided by way of example only, and are therefore not intended to limit the scope of the claimed invention. Furthermore, it should be recognized that various features of the described implementations can be combined in various ways to create countless additional implementations. Furthermore, while various materials, dimensions, shapes, configurations, locations, etc. have been described for use with the disclosed implementations, others in addition to those disclosed may be utilized without departing from the scope of the claimed invention.

[0214] Those skilled in the art will recognize that implementations may include fewer features than those shown in any individual implementation described above. The implementations described herein are not meant to be an exhaustive listing of ways in which various features may be combined. Thus, implementations are not mutually exclusive combinations of features; rather, implementations may include combinations of various individual features selected from various individual implementations, as would be understood by one skilled in the art. Furthermore, elements described with respect to one implementation may be implemented in other implementations, even if not described in such implementations, unless otherwise specified. Although a dependent claim may refer to a specific combination with one or more other claims in the claims, other implementations may also include a combination of the dependent claim with the subject matter of each other dependent claim, or a combination of one or more features with other dependent or independent claims. Such combinations are suggested herein, unless it is stated that a specific combination is not intended. Furthermore, including features of any other independent claim within any other independent claim is also intended, even if this claim is not made directly dependent on the independent claim.

[0215] Moreover, references herein to "one implementation," "an implementation," or "some implementations" mean that a particular feature, structure, or characteristic described in connection with an implementation is included in at least one implementation of the present teachings. The appearances of the phrase "in one implementation" in various places in the specification do not necessarily refer to the same implementation.

[0216] Any incorporation by reference of the above documents is limited so that no subject matter contrary to the express disclosure herein is incorporated. Any incorporation by reference of the above documents is further limited so that no claims contained in the above documents are incorporated by reference herein. Any incorporation by reference of the above documents is further limited so that no definitions provided in the above documents are incorporated by reference herein unless expressly expressed herein.

[0217] Although an implementation has been described with reference to certain exemplary implementations, it will become apparent that various modifications and changes may be made thereto without departing from the broad spirit and scope of the present disclosure. Accordingly, this specification and the accompanying drawings are to be regarded as illustrative rather than restrictive. The accompanying drawings, which form a part hereof, illustrate, by way of illustration, and not by way of limitation, specific implementations in which the present subject matter may be practiced. The illustrated implementations have been described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other implementations may be utilized and derived from the scope of the present disclosure, such that structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. Accordingly, this Detailed Description is not to be taken in a limiting sense, but rather the scope of various implementations is defined solely by the appended claims, along with the full scope of equivalents to which such claims are entitled.

[0218] Although specific implementations have been shown and described herein, it should be understood that any arrangement calculated to achieve the same purpose may be substituted for the specific implementations shown. The present disclosure is intended to cover any adaptations or variations of the various implementations. Combinations of the above implementations and other implementations not specifically described herein will be apparent to those of skill in the art upon reviewing the above description.

[0219] In this specification, the singular articles are used to include one or more (as commonly used in patent documents) independent of any other instance or use of "at least one" or "one or more." In this specification, the term "or" is used to refer to a non-exclusive, or to include "A but not B," "B but not A," and "A and B," unless "A or B" is otherwise indicated. In this specification, the terms "comprises" and "in" are used as the plain English equivalents of the respective terms "including" and "wherein." Also, in the following claims, the term "comprises" is open-ended, i.e., a system, user equipment (UE), article, composition, method, or process includes several elements in addition to those recited after the claim term is still deemed to fall within the scope of the claim. Moreover, in the following claims, the terms "first," "second," "third," etc. are used merely as labels and are not intended to impose numerical requirements on their objects.

Claims

1. A method comprising: analyzing, with a computing system including processing circuitry and memory, one or more data tables stored by a data repository with respect to one or more criteria, the one or more data tables including insurance claims data corresponding to one or more treatments provided to individuals in a biological condition, the one or more criteria indicating one or more formats of insurance code identifiers; determining, by the computing system, a plurality of insurance code identifiers included in the one or more data tables that correspond to the one or more criteria; analyzing, by the computing system, the subset of the plurality of insurance code identifiers to determine one or more formats of the subset of the plurality of insurance code identifiers; generating, by the computing system, one or more requests for an application programming interface (API) call that includes insurance code identifiers of the subset of the plurality of insurance code identifiers; obtaining, by the computing system and in response to the one or more requests of the API, a data file containing information corresponding to the insurance code identifier; extracting, by the computing system, a treatment identifier from the data file; generating, by said computing system, an additional data table having a row indicating that said insurance code identifier corresponds to said treatment identifier; A method comprising:

2. The method of claim 1 , comprising determining, by the computing system, one or more columns of the one or more data tables that include insurance code identifiers corresponding to treatments for individuals.

3. determining, by the computing system, that a first insurance code identifier of the plurality of insurance code identifiers is a duplicate of a second insurance code identifier of the plurality of insurance code identifiers; removing, by the computing system, the second insurance code identifier from the plurality of insurance code identifiers to generate the plurality of insurance code identifiers; The method of claim 1 , comprising:

4. The method of claim 1 , wherein the plurality of insurance code identifiers corresponds to a plurality of National Drug Code (NDC) identifiers.

5. identifying, by the computing system, a first insurance code identifier of the subset of the plurality of insurance code identifiers; determining, by the computing system, that the first insurance code identifier corresponds to a first format of an insurance code identifier; modifying, by the computing system, the first insurance code identifier to generate a second insurance code identifier corresponding to a second format of insurance code identifier; The method of claim 1 , comprising:

6. generating, by the computing system, one or more additional requests of an additional API that include the second insurance code identifier; obtaining, by the computing system and in response to the one or more additional invocations of the additional API, an additional data file including information corresponding to a third insurance code identifier; extracting, by the computing system, the insurance code identifier from the additional data file; The method of claim 5 , comprising:

7. generating, by the computing system, one or more additional calls to the API that include an additional insurance code identifier; obtaining, by the computing system and in response to the one or more additional calls to the API, an additional data file including additional information corresponding to the additional insurance code identifier; determining, by the computing system, that the additional information does not include at least one valid insurance code identifier; The method of claim 1 , comprising:

8. generating, by the computing system, one or more additional calls to the API that include additional insurance code identifiers of the subset of the plurality of insurance code identifiers; obtaining, by the computing system and in response to the one or more additional calls to the API, an additional data file including additional information corresponding to the additional insurance code identifier; analyzing, by the computing system, the additional information with respect to one or more additional criteria; determining, by the computing system, that the additional information does not include at least one medication identifier; The method of claim 1 , comprising:

9. generating, by the computing system, one or more additional calls to additional APIs that include the treatment identifier; obtaining, by the computing system and in response to the one or more additional calls, an additional data file including additional information corresponding to the medication identifier; determining, by said computing system and based on said additional information, a class of medication corresponding to said medication identifier; The method of claim 1 , comprising:

10. determining, by the computing system and based on the additional information, the source of the class of drugs by analyzing, by the computing system, the source of the class of drugs for a prioritized number of drug sources; extracting at least a portion of the additional information by the computing system and from one or more fields of the data file; adding, by the computing system, said at least some of said additional information to a row in a database table; The method of claim 1 , comprising:

11. 11. The method of claim 10, wherein the prioritized number of drug sources includes a first source of the class of drug having a first priority and a second source of the class of drug having a second priority lower than the first priority.

12. analyzing, by the computing system, the additional information regarding the first source of the class of drug; determining, by the computing system, that the additional information does not include a source of the class of drugs corresponding to the first class; analyzing, by the computing system, the additional information regarding the second source of the class of drug; extracting by said computing system at least a portion of said additional information from said additional data file relating to said second source of said class of drug; The method of claim 10, comprising:

13. receiving, by the computing system, a request to identify a group of individuals who have received a treatment responsive to a biological condition existing for the group of individuals, the request including a name of the treatment or a class of the treatment; analyzing, by the computing system, one or more values ​​in one or more rows of the additional data table to determine one or more insurance code identifiers corresponding to the name of the treatment or the class of the treatment; analyzing, by the computing system, values ​​of the rows of the one or more data tables to determine one or more rows that include the one or more insurance code identifiers; determining, by the computing system, one or more identifiers of individuals included in the one or more rows to generate a population of individuals who have received the treatment related to the biological condition; The method of claim 1 , comprising:

14. determining, by the computing system, genomic information of the population of treated individuals that is related to the biological condition; applying, by the computing system, at least one of one or more statistical techniques or one or more machine learning techniques to determine one or more features of the population of individuals; 14. The method of claim 13, comprising:

15. 15. The method of claim 14, wherein the one or more features comprise at least one of: a genetic mutation in the genome of each of the individuals in the population of individuals; a genetic mutation in cell-free deoxyribonucleic acid (DNA) in one or more samples obtained from individuals in the population of individuals; an amount of cell-free DNA carrying the genetic mutation for each of the individuals in the population of individuals; or a change in the amount of cell-free DNA carrying the genetic mutation for each of the individuals in the population of individuals over a period of time.

16. A system comprising: one or more hardware processing units; one or more computer-readable storage media storing computer-executable instructions; the computer-executable instructions, when executed by the one or more hardware processing units, analyzing one or more data tables stored by a data repository with respect to one or more criteria, the one or more data tables including insurance claims data corresponding to one or more treatments provided to individuals in a biological condition, the one or more criteria indicating one or more formats of insurance code identifiers; determining a plurality of insurance code identifiers included in the one or more data tables that correspond to the one or more criteria; and analyzing the subset of the plurality of insurance code identifiers to determine one or more formats of the subset of the plurality of insurance code identifiers; generating one or more requests for an application programming interface (API) call that includes insurance code identifiers of the subset of the plurality of insurance code identifiers; Obtaining a data file containing information corresponding to the insurance code identifier in response to the one or more requests of the API; extracting a treatment identifier from the data file; generating an additional data table having a row indicating that the insurance code identifier corresponds to the treatment identifier; A system that causes the system to perform operations including:

17. One or more non-transitory computer-readable storage media storing computer-executable instructions that, when executed by one or more hardware processing units, analyzing one or more data tables stored by a data repository with respect to one or more criteria, the one or more data tables including insurance claims data corresponding to one or more treatments provided to individuals in a biological condition, the one or more criteria indicating one or more formats of insurance code identifiers; determining a plurality of insurance code identifiers included in the one or more data tables that correspond to the one or more criteria; and analyzing the subset of the plurality of insurance code identifiers to determine one or more formats of the subset of the plurality of insurance code identifiers; generating one or more requests for an application programming interface (API) call that includes insurance code identifiers of the subset of the plurality of insurance code identifiers; Obtaining a data file containing information corresponding to the insurance code identifier in response to the one or more requests of the API; extracting a treatment identifier from the data file; generating an additional data table having a row indicating that the insurance code identifier corresponds to the treatment identifier; One or more non-transitory computer-readable storage media that cause the system to perform operations including:

18. 20. The one or more non-transitory computer-readable media of claim 17, comprising additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining one or more columns of the one or more data tables that include insurance code identifiers corresponding to treatments for individuals.

19. additional computer-executable instructions that, when executed by the one or more hardware processing units, determining that a first insurance code identifier of the plurality of insurance code identifiers is a duplicate of a second insurance code identifier of the plurality of insurance code identifiers; removing the second insurance code identifier from the plurality of insurance code identifiers to generate the plurality of insurance code identifiers; 20. The one or more non-transitory computer-readable media of claim 17, which causes the system to perform additional operations including:

20. 20. The one or more non-transitory computer-readable media of claim 17, wherein the plurality of insurance code identifiers correspond to a plurality of National Drug Code (NDC) identifiers.

21. additional computer-executable instructions that, when executed by the one or more hardware processing units, identifying a first insurance code identifier of the subset of the plurality of insurance code identifiers; determining that the first insurance code identifier corresponds to a first format of an insurance code identifier; modifying the first insurance code identifier to generate a second insurance code identifier corresponding to a second format of insurance code identifier; 20. The one or more non-transitory computer-readable media of claim 17, which causes the system to perform additional operations including:

22. additional computer-executable instructions that, when executed by the one or more hardware processing units, generating one or more additional requests of an additional API that include the second insurance code identifier; Obtaining an additional data file in response to the one or more additional invocations of the additional API, the additional data file including information corresponding to the third insurance code identifier; extracting the insurance code identifier from the additional data file; 22. The one or more non-transitory computer-readable media of claim 21, which causes the system to perform additional operations including: