Computer architecture for generating integrated data repositories
Patent Information
- Application Number
- JP2023574149
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-30
- Filing Date
- 2022-06-03
- Publication Date
- 2025-06-11
AI Technical Summary
Existing systems face inefficiencies and inaccuracies in analyzing unstructured medical data, particularly health insurance claims and genomic data, which are crucial for understanding tumor behavior and treatment outcomes, leading to suboptimal treatment decisions.
An integrated data repository system that combines structured health insurance claims data with molecular data, using hash functions and cryptographic protocols to ensure privacy, enabling accurate and efficient analysis of health and treatment information.
The system provides a more accurate characterization of health information, identifies treatment effectiveness based on genomic profiles, and improves understanding of tumor progression and drug resistance, leading to better treatment strategies.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Priority Claims and Incorporation by Reference This application claims priority to U.S. Provisional Patent Application No. 63 / 196,609, entitled "Computer Architecture for Generating an Integrated Data Repository," filed on June 3, 2021, U.S. Provisional Patent Application No. 63 / 227,860, entitled "Computer Architecture for Identifying Lines of Therapy," filed on July 30, 2021, U.S. Provisional Patent Application No. 63 / 238,851, entitled "Data Repository System, and Method for Cohort Selection," filed on August 31, 2021, and U.S. Provisional Patent Application No. 63 / 250,912, entitled "Computer Architecture for Generating a Reference Data Table," filed on September 30, 2021, the entire contents of each of which are incorporated herein by reference.
[0002] Technical Field TECHNICAL FIELD Implementations of the present disclosure relate generally to the field of computer architectures, and more particularly to computer architectures for generating data repositories that integrate multiple sources of medical data, including medical insurance claims data and genomics data. [Background technology]
[0003] background Various types of documentation may be generated in association with an individual's visit to a health care provider for treatment of one or more biological conditions. For example, a medical record may be created by the health care provider that includes clinical findings, laboratory test results, diagnostic test information, imaging information, dental hygiene information, one or more combinations thereof, etc. recorded by the health care provider. In addition, a billing record may be generated that indicates payment information for at least one product or service provided by the health care provider to the individual. Furthermore, health insurance claims information may be generated that indicates information obtained by a health insurance company in connection with the individual's treatment for one or more biological conditions. [Brief description of the drawings]
[0004] [Figure 1] FIG. 1 illustrates an example architecture for generating a unified data repository containing multiple types of medical data, according to one or more implementations.
[0005] [Diagram 2] FIG. 2 illustrates an exemplary framework that supports the arrangement of data tables in a unified data repository, according to one or more implementations.
[0006] [Diagram 3] FIG. 3 illustrates an architecture for generating one or more data sets from information retrieved from a data repository that integrates health-related data from multiple sources, according to one or more implementations.
[0007] [Figure 4] FIG. 4 illustrates an architecture for generating an integrated data repository containing de-identified health insurance claims data and de-identified genomics data, according to one or more implementations.
[0008] [Diagram 5]FIG. 5 illustrates a framework for generating a data set by a data pipeline system based on data stored by a unified data repository, according to one or more implementations.
[0009] [Figure 6] FIG. 6 is a schematic diagram of an architecture for integrating medical record data into a unified data repository.
[0010] [Figure 7] FIG. 7 is a data flow diagram of an exemplary process for generating an integrated data repository that stores health insurance claims data and genomics data, according to one or more implementations.
[0011] [Figure 8] FIG. 8 is a data flow diagram of an exemplary process for generating multiple data sets used to analyze information stored by an integrated data repository that stores health insurance claims data and genomics data, according to one or more implementations.
[0012] [Figure 9] FIG. 9 shows a diagrammatic representation of a machine in the form of a computer system within which a set of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein, in accordance with one or more implementations.
[0013] [Figure 10] Figure 10 shows Kaplan-Meier curves showing real-world overall survival values for patients who received 1L therapy to treat non-small cell lung cancer prior to receiving treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0014] [Figure 11]Figure 11 shows Kaplan-Meier curves showing real-world overall survival values for patients receiving 1L therapy to treat non-small cell lung cancer during treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0015] [Figure 12] Figure 12 shows Kaplan-Meier curves showing real-world overall survival values for patients who received osimertinib to treat non-small cell lung cancer prior to treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0016] [Figure 13] FIG. 13 shows Kaplan-Meier curves illustrating real-world overall survival values for patients receiving osimertinib to treat non-small cell lung cancer during treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0017] [Figure 14] FIG. 14 shows Kaplan-Meier curves illustrating real-world overall survival values for patients receiving chemotherapy to treat non-small cell lung cancer during treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0018] [Figure 15] FIG. 15 shows Kaplan-Meier curves illustrating real-world overall survival values for patients receiving chemotherapy to treat non-small cell lung cancer after treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0019] [Figure 16] FIG. 16 shows Kaplan-Meier curves illustrating real-world overall survival values for patients who received chemotherapy to treat non-small cell lung cancer prior to receiving treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0020] [Figure 17] FIG. 17 illustrates the frequency of selected alterations in a cohort of patients (n=637) diagnosed with advanced non-small cell lung cancer (NSCLC) who underwent liquid biopsy after initiation of first-line osimertinib treatment.
[0021] [Figure 18] FIG. 18 shows the frequency of selected mutations in the ligand-binding domain of a cohort of patients (n=4448) diagnosed with breast cancer who underwent liquid biopsy after documented aromatase inhibitor (AI) treatment.
[0022] [Figure 19] FIG. 19 shows changes associated with osimertinib resistance detected by liquid biopsy following treatment provided to women diagnosed with NSCLC.
[0023] [Figure 20] FIG. 20 shows ESR1 resistance mutations detected after the second course of treatment for women diagnosed with metastatic breast cancer and treated with aromatase inhibitors. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0024] Detailed Description The following description and drawings sufficiently describe particular implementations to enable those skilled in the art to practice them. Other implementations may incorporate structural, logical, electrical process and other changes. Portions and features of some implementations may be included in or substituted for those of other implementations. Implementations set forth in the claims encompass all possible equivalents of those claims.
[0025] More data are needed to understand tumor behavior and treatment outcomes outside the highly selective constraints of randomized, controlled trials, often designed and conducted by entities with a commercial interest in their success. Real-world evidence (RWE), specifically the use of databases with integrated clinical and molecular data, is playing an increasingly important role in precision oncology research. However, the majority of such databases comprise genomic information from tumors restricted to only one time point, typically the time of diagnosis, in part due to the practical challenges of genomic profiling of consecutive tumor specimens in real-world clinical practice. Despite evidence that treatments can significantly alter tumor genomic profiles and lead to drug resistance, genomic data on tumors is often limited to those naïve to systematic treatment. Combining data obtained with liquid biopsy assays with rich clinical information can help overcome these issues and improve our understanding of tumor evolution and expression of biomarkers that confer resistance, guiding the development of new therapies addressing areas of unmet need.
[0026] Analysis of medical data using existing systems and technologies is usually performed on medical records generated by health care providers. As used herein, a health care provider may refer to an entity, individual, or group of individuals involved in providing care to an individual in relation to at least one of the treatment or prevention of one or more biological conditions. Also, as used herein, a biological condition may refer to a functional and / or structural abnormality in an individual to such an extent that it causes or is likely to cause a detectable characteristic of the abnormality. A biological condition may be characterized by external and / or internal characteristics, signs, and / or symptoms that indicate a deviation from a biological norm in one or more populations. A biological condition may be characterized by external and / or internal characteristics, signs, and / or symptoms that indicate a deviation from a biological norm in one or more populations. In various examples, a biological condition may include one or more molecular phenotypes. For example, a biological condition may correspond to a genetic or epigenetic lesion. In one or more additional examples, the biological condition may include at least one of one or more diseases, one or more disorders, one or more injuries, one or more syndromes, one or more disabilities, one or more infections, one or more isolated symptoms, or other atypical variations in the biological structure and / or function of the individual. In addition, treatment as used herein may refer to a substance, procedure, routine, device, and / or other intervention that may be administered or performed with the intent of treating one or more effects of the biological condition of the individual. In one or more examples, the treatment may include a substance that is metabolized by the individual. The substance may include a composition of matter, such as a pharmaceutical composition. The substance may be delivered to the individual via some method, such as ingestion, injection, absorption, or inhalation. The treatment may include a physical intervention, such as one or more surgeries. In at least some examples, the treatment may include a therapeutically useful intervention.
[0027] Medical data typically analyzed by existing systems includes unstructured data. Unstructured data may include data that is not organized according to a predefined or standardized format. For example, unstructured data may include annotations created by a healthcare provider that consist of free text. That is, the manner in which the annotations are captured does not include predefined inputs selectable by the healthcare provider, such as via a drop-down menu or list. Instead, the annotations include text entered by the healthcare provider, which may include sentences, sentence fragments, words, characters, symbols, abbreviations, one or more combinations thereof, and the like. In some cases, unstructured data may be partially structured. For example, a provider may select a billing code from a predefined list of billing codes and add unstructured annotations to data associated with that billing code.
[0028] Existing systems typically devote significant computing resources to analyzing unstructured data to extract information that may be relevant to the analysis being performed by the existing system. In some cases, existing systems may analyze the unstructured data and convert it into a structured form to facilitate analysis of the previously unstructured data. Analysis of unstructured data by existing systems may be inefficient and inaccurate. In scenarios where the unstructured data is derived from medical data, it is important to accurately analyze the information, as the analysis may be relevant to the treatment and / or diagnosis of multiple individuals with respect to one or more biological conditions. Thus, inaccurate analysis of medical data may result in suboptimal treatment of individuals.
[0029] Implementations of the techniques, architectures, frameworks, systems, processes, and computer-readable instructions described herein are directed to analyzing health insurance claims data to derive information about at least one of an individual's health or treatment. In contrast to existing systems, health insurance claims data is structured according to one or more formats and stored by a number of data tables. The data tables may include codes or other alphanumeric information indicating the treatment the individual received, the date of treatment, dosage information, the individual's diagnosis of one or more biological conditions, information related to a visit to a health care provider, the date of the visit to a health care provider, billing information, and the like. Implementations described herein may be used to accurately analyze health insurance claims data for hundreds, thousands, or even tens of thousands of individuals or more, in which one or more biological conditions exist. In various examples, tens, hundreds of thousands, or even millions of rows and / or columns of health insurance claims data may be analyzed to determine health-related information for individuals in which one or more biological conditions exist.
[0030] In various examples, implementations described herein can integrate molecular data with health insurance claims data. The molecular data can include information derived from tissue samples extracted from multiple individuals. The molecular data can also include information derived from blood samples extracted from multiple individuals. In one or more illustrative examples, the molecular data can include genomics data. Additionally, in one or more examples, the health insurance claims data can be integrated with germline genetic information of multiple individuals.
[0031] An integrated data repository may be created that combines an individual's health insurance claims data with the individual's molecular data. In one or more examples, an individual's identifier may be generated that is associated with both the individual's health insurance claims data and the individual's molecular data. Both the molecular data and the health insurance claims data stored by the integrated data repository may be accessible using a single identifier for the individual. In one or more illustrative examples, the individual's identifier may include an encrypted security key. In various examples, the integrated data repository may include multiple data tables corresponding to various aspects of the data stored in the data repository. For example, a first data table may be generated that includes summary data for the individual included in the integrated data repository, such as personal information, and a second data table may be generated that includes data corresponding to visits to a healthcare provider. In addition, a third data table may be generated that indicates medical treatments provided to the individual, and a fourth data table may be generated that indicates information related to prescriptions obtained by the individual. Furthermore, a fifth data table may be generated that includes a multiomics profiling of the individual. The multiomics profile may include at least one of a genomic profile, a transcriptomics profile, an epigenetics profile, or a proteomics profile.
[0032] The data tables included in the integrated data repository may be linked via logical links. In that way, a query that retrieves information from one data table may cause information from one or more additional data tables to be retrieved. Information stored by the linked data tables may be accessed to generate a number of different data sets that may be used to analyze the information stored by the integrated data repository. For example, the information stored by the integrated data repository may be analyzed by one or more algorithms to generate a data set organized according to one or more schemas. The data set may represent treatments received by an individual over a period of time for a biological condition. The data set may also represent a cohort of individuals included in the integrated data repository that have a number of common characteristics. In various examples, the data set may synthesize and compose information from several different data sources, including the integrated data repository. The data set may be analyzed with respect to several queries to represent information that may be of interest to at least one of a health care provider, a patient, or a provider of treatment for a biological condition. For example, one or more data sets may be analyzed to more accurately determine the survival rate of individuals with a particular genomic profile in response to the presence of a biological condition and receiving a specified treatment.
[0033] The implementations described herein may provide a platform for integrating molecular data with an individual's health insurance claims data that is not found in existing systems that typically rely on electronic medical records that contain a certain amount of unstructured data. By generating and analyzing structured health insurance claims data integrated with molecular data, the implementations described herein may provide a more accurate characterization of the integrated data than existing systems that rely on relatively inaccurate unstructured electronic medical record data. Additionally, the implementations described herein generate analyzable data sets that enable analysis of health information about an individual in a confidential and de-identified manner.
[0034] 1 illustrates an example architecture 100 for generating a unified data repository containing multiple types of medical data, according to one or more implementations. The architecture 100 may include a data integration and analysis system 102. The data integration and analysis system 102 may obtain data from multiple data sources and consolidate the data from the data sources into a unified data repository 104. For example, the data integration and analysis system 102 may obtain data from a health insurance claims data repository 106. In various examples, the data integration and analysis system 102 and the health insurance claims data repository 106 may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system 102 and the health insurance claims data repository 106 may be created and maintained by the same entity.
[0035] The data integration and analysis system 102 may be implemented by one or more computing devices. The one or more computing devices may include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or a combination thereof. In certain implementations, at least a portion of the one or more computing devices may be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices may be implemented in a cloud computing architecture. In scenarios where the computing system used to implement the data integration and analysis system 102 is configured as a distributed computing architecture, processing operations may be performed by multiple virtual machines simultaneously. In various examples, the data integration and analysis system 102 may implement multi-threading techniques. The implementation of distributed computing architectures and multi-threading techniques allows the data integration and analysis system 102 to utilize fewer computing resources compared to computing architectures that do not implement these techniques.
[0036] The health insurance claims data repository 106 may store information obtained from one or more health insurance companies corresponding to claims made by subscribers of one or more health insurance companies. The health insurance claims data repository 106 may be arranged (e.g., sorted) by patient identifier. The patient identifier may be based on the patient's first name, last name, birth date, social security number, address, employer, etc. The data stored by the health insurance claims data repository 106 may include structured data arranged in one or more data tables. The one or more data tables storing the structured data may include multiple rows and multiple columns indicating information regarding health insurance claims made by subscribers of one or more health insurance companies related to treatments and / or care received by the subscribers from a healthcare provider. At least a portion of the rows and columns of the data tables stored by the health insurance claims data repository 106 may include health insurance codes that may indicate diagnoses of biological conditions, treatments and / or procedures obtained by subscribers of one or more health insurance companies. In various examples, the health insurance codes may also indicate diagnostic procedures obtained by the individual related to one or more biological conditions that may exist for the individual. In one or more examples, the diagnostic procedure may provide information used in detecting the presence of a biological condition. The diagnostic procedure may also provide information used to determine the progression of a biological condition. In one or more illustrative examples, the diagnostic procedure may include one or more imaging procedures, one or more assays, one or more testing procedures, one or more combinations thereof, and the like.
[0037] The data integration and analysis system 102 may also obtain information from a molecular data repository 108. The molecular data repository 108 may store data for multiple individuals relating to genomic, genetic, metabolomic, transcriptomic, fragmentomic, immune receptor, methylation, epigenomic, and / or proteomic information. In one or more examples, the data integration and analysis system 102 and the molecular data repository 108 may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system 102 and the molecular data repository 108 may be created and maintained by the same entity.
[0038] The genomic information may indicate one or more mutations corresponding to the individual's genes. The mutations to the individual's genes may correspond to differences between the sequence of the individual's nucleic acid and one or more reference genomes. The reference genome may include a known reference genome, such as hg19. In various examples, the mutations of the individual's genes may correspond to differences in the individual's germline genes compared to the reference genome. In one or more additional examples, the reference genome may include the individual's germline genome. In one or more further examples, the mutations of the individual's genes may include somatic mutations. The mutations of the individual's genes may involve insertions, deletions, single nucleotide variants, loss of heterozygosity, duplications, amplifications, translocations, fusion genes, or one or more combinations thereof.
[0039] In one or more illustrative examples, the genomic information stored by the molecular data repository 108 may include a genomic profile of tumor cells present in the individual. In such a situation, the genomic information may be derived from an analysis of genetic material, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA), from a tissue sample or tumor biopsy, a sample containing circulating tumor cells (CTCs), exosomes or efferosomes, or from circulating nucleic acids (e.g., cell-free DNA) found in the individual's blood sample that are present due to the degradation of tumor cells present in the individual. In one or more examples, the genomic information of the individual's tumor cells may correspond to one or more target regions. One or more mutations present with respect to one or more target regions may indicate the presence of tumor cells in the individual. The genomic information stored by the molecular data repository 108 may be generated in conjunction with an assay or other diagnostic test that may determine one or more mutations with respect to one or more target regions of a reference genome.
[0040] "Cell-free DNA," "cfDNA molecules," or simply "cfDNA" includes DNA molecules that occur in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum) and includes DNA that is not contained within a cell or otherwise bound to a cell at the time of isolation from the subject. The DNA was originally present within one or more cells of a large complex organism (e.g., a mammal) or other cells, such as bacteria that colonize the organism, but the DNA has undergone release from the cell and entered the fluids found within the organism. cfDNA includes, but is not limited to, the cell-free genomic DNA of a subject (e.g., the genomic DNA of a human subject) and the cell-free DNA of microorganisms, such as bacteria that inhabit the subject (whether pathogenic or typically found in commonly colonizing locations such as the digestive tract or skin of healthy controls), but does not include the cell-free DNA of microorganisms that simply contaminate a sample of bodily fluid. Typically, cfDNA can be obtained by obtaining a sample of a fluid without the need for an in vitro cell lysis step and includes removal of cells present in the fluid (e.g., centrifugation of blood to remove cells).
[0041] In one or more additional examples, the data integration and analysis system 102 may obtain information from one or more additional data repositories 110. The one or more additional data repositories 110 may store data related to an electronic medical record of an individual whose data is present in at least one of the health insurance claims data repository 106 or the molecular data repository 108. Additionally, the one or more additional data repositories 110 may store data related to a pathology report of an individual whose data is present in at least one of the health insurance claims data repository 106 or the molecular data repository 108. In various examples, the one or more additional data repositories 110 may store data related to a biological condition and / or a treatment of a biological condition. In one or more examples, the data integration and analysis system 102 and at least a portion of the one or more additional data repositories 110 may be created and maintained by different entities. In one or more further examples, the data integration and analysis system 102 and at least a portion of the one or more additional data repositories 110 may be created and maintained by the same entity.
[0042] In one or more further implementations, the data integration and analysis system 102 may obtain information from one or more reference information data repositories 112. The one or more reference information data repositories 112 may store information including definitions, standards, protocols, glossaries, one or more combinations thereof, and the like. In various examples, the information stored by the one or more reference information data repositories may correspond to a biological condition and / or a treatment of a biological condition. In one or more illustrated examples, the one or more reference information data repositories 112 may include RxNorm, which provides normalized names for clinical drugs and links the names to many of the drug terms used in pharmacy management and drug interaction software. In one or more examples, the data integration and analysis system 102 and at least a portion of the one or more reference information data repositories 112 may be created and maintained by different entities. In one or more further examples, the data integration and analysis system 102 and at least a portion of the one or more reference information data repositories 112 may be created and maintained by the same entity.
[0043] The data integration and analysis system 102 may retrieve data from at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112 via one or more communication networks accessible by the data integration and analysis system 102 and accessible by at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112. The data integration and analysis system 102 may also retrieve data from at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112 via one or more secure communication channels. Additionally, the data integration and analysis system 102 may retrieve data from at least one of the health insurance claims data repository 106, the molecular data repository 108, one or more additional data repositories 110, or the reference information data repository 112 via one or more application programming interface (API) calls.
[0044] The data integration and analysis system 102 may include a data integration system 114. The data integration system 114 may retrieve data from the health insurance claims data repository 106 and the molecular data repository 108 to generate the integrated data repository 104. The data integration system 114 may also retrieve data from one or more additional data repositories 110 to generate the integrated data repository 104. In various examples, the data integration system 114 may implement one or more natural language processing techniques to integrate data from the one or more additional data repositories 110 into the integrated data repository 104.
[0045] In one or more examples, the data integration system 114 may generate one or more tokens for identifying an individual having data stored in the health insurance claims data repository 106 and having data stored in the molecular data repository 108. In various examples, the data integration system 114 may generate one or more tokens by implementing one or more hash functions. The data integration system 114 may implement one or more hash functions to generate one or more tokens based on information stored by at least one of the health insurance claims data repository 106 or the molecular data repository 108. For example, the information used by the data integration system 114 to generate the individual tokens by implementing a hash function may include at least one of an identifier for each individual, a date of birth for each individual, a zip code for each individual, a date of birth for each individual, or a gender for each individual. In one or more illustrative examples, the identifier for each individual may include a combination of at least a portion of a first name for each individual and at least a portion of a last name for each individual. The tokens generated using data from multiple different data repositories may correspond to the same or similar information or to the same or similar types of information stored by the different data repositories. By way of example, a token may be generated using a portion of an individual's name, date of birth, at least a portion of their zip code, and gender, which are retrieved from the health insurance claims data repository 106 and the molecular data repository 108.
[0046] The data integration system 114 may integrate data from multiple different data sources by analyzing tokens generated by implementing one or more hash functions using data obtained from the multiple different data sources. For example, the data integration system 114 may obtain one or more first tokens generated from data stored by the health insurance claims data repository 106 and one or more second tokens generated from data stored by the molecular data repository 108. The data integration system 114 may analyze the one or more first tokens in relation to the one or more second tokens to determine individual first tokens that correspond to individual second tokens. In one or more illustrative examples, the data integration system 114 may identify individual first tokens that match individual second tokens. A first token may match a second token when the data of the first token has at least a threshold amount of similarity to the data of the second token. In one or more examples, a first token may match a second token when the data of the first token is the same as the data of the second token. To illustrate, a first token may match a second token when an alphanumeric string of the first token is the same as an alphanumeric string of the second token.
[0047] By determining a first token generated using data stored by the health insurance claims data repository 106 that corresponds to a second token generated using data stored by the molecular data repository 108, the data integration system 114 may identify individuals who have data stored in both the health insurance claims data repository 106 and the molecular data repository 108. In this manner, the data integration system 114 may obtain data from multiple individuals from the health insurance claims data repository 106 and data from the molecular data repository 108 from the same multiple individuals and store the health insurance claims data and molecular data for the multiple individuals in the integrated data repository 104.
[0048] The data integration system 114 may also integrate data stored by one or more additional data repositories 110 with data from the health insurance claims data repository 106 and the molecular data repository 108 to generate the integrated data repository 104. By way of example, the data integration system 114 may obtain one or more third tokens generated from data stored by the additional data repository 110, such as a data repository storing data corresponding to a pathology report. The data integration system 114 may analyze the one or more third tokens against a first token generated using information stored by the health insurance claims data repository 106 and a second token generated using information stored by the molecular data repository 108 to determine respective third tokens corresponding to respective first tokens and respective second tokens. In one or more illustrative examples, the data integration system 114 may identify the third tokens generated using one or more hash functions and a common set of information obtained from the health insurance claims data repository 106, the molecular data repository 108, and the additional data repository 110.
[0049] By determining a third token generated using data stored by the additional data repository 110 that corresponds to the first token generated using data stored by the health insurance claims data repository 106 and the second token generated using data stored by the molecular data repository 108, the data integration system 114 may identify an individual having data stored in the health insurance claims data repository 106, the molecular data repository 108, and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the health insurance claims data repository 106 from multiple individuals and data from the molecular data repository 108 and the additional data repository 110 from the same multiple individuals and store the health insurance claims data, molecular data, and additional data for the multiple individuals in the integrated data repository 104.
[0050] Data stored by the integrated data repository 104 for multiple individuals may be accessible using an identifier for each of the individuals. The data integration system 114 may implement multiple techniques as part of the de-identification process for storing and retrieving an individual's information in the integrated data repository 104. An individual's identifier may correspond to a key generated using at least one hash function. An individual's identifier may also be generated by implementing one or more salting processes on a key generated using at least one hash function. A token generated using one or more hash functions and a common set of information retrieved from the health insurance claims data repository 106, the molecular data repository 108, and / or the additional data repository 110. In one or more illustrative examples, an identifier generated by the data integration system 114 to access each individual's information stored by the integrated data repository 104 may be unique to each individual. In one or more examples, an individual's identifier may be generated using at least a portion of the information used to generate a token related to the individual. In one or more additional examples, an individual's identifier may be generated using information different from the information used to generate a token related to the individual.
[0051] The data integration system 114 may also generate the integrated data repository 104 in a similar manner from different combinations of data repositories. For example, the data integration system 114 may obtain tokens generated from information stored by the health insurance claims data repository 106 and additional tokens generated from information stored by one or more additional data repositories 110. The data integration system 114 may determine individual tokens generated from information stored by the health insurance claims data repository 106 that correspond to the individual additional tokens generated from information stored by the one or more additional data repositories 110. By determining tokens generated using data stored by the health insurance claims data repository 106 that correspond to the additional tokens generated using data stored by the additional data repository 110, the data integration system 114 may identify individuals having data stored in both the health insurance claims data repository 106 and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the health insurance claims data repository 106 from multiple individuals and data from the additional data repository 110 from the same multiple individuals and store the health insurance claims data and additional data for the multiple individuals in the integrated data repository 104. The health insurance claims data and additional data stored in the integrated data repository 104 for the multiple individuals may be accessible using an identifier for each of the individuals.
[0052] In one or more further examples, the data integration system 114 may obtain tokens generated from information stored by the molecular data repository 108 and tokens generated from information stored by one or more additional data stores 110. The data integration system 114 may determine individual tokens generated from information stored by the molecular data repository 108 that correspond to individual additional tokens generated from information stored by the one or more additional data repositories 110. By determining tokens generated using data stored by the molecular data repository 108 that correspond to additional tokens generated using data stored by the additional data repository 110, the data integration system 114 may identify individuals having data stored in both the molecular data repository 108 and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the molecular data repository 108 from multiple individuals and data from the additional data repositories 110 from the same multiple individuals and store the molecular data and additional data of the multiple individuals in the integrated data repository 104. The molecular data and additional data stored in the integrated data repository 104 for multiple individuals may be accessible using the individuals' respective identifiers.
[0053] The data stored by the integrated data repository 104 may be stored in accordance with one or more regulatory frameworks that protect privacy and ensure security of individuals' medical records, health information, and insurance information. For example, the data may be stored by the integrated data repository 104 in accordance with one or more government regulatory frameworks that are directed to protecting personal information, such as the Health Insurance Portability and Accountability Act (HIPAA) and / or the General Data Protection Regulation (GDPR). The integrated data repository 104 also stores the data in an anonymized and de-identified form to ensure protection of the privacy of individuals whose data is stored by the integrated data repository 104. To further ensure privacy of individuals whose data is stored by the integrated data repository 104, the data integration system 114 may periodically regenerate the integrated data repository 104. For example, the data integration system 114 may create the integrated data repository 104 once a quarter. In one or more additional examples, the data integration system 114 may generate the integrated data repository 104 monthly, weekly, or biweekly. By periodically regenerating the integrated data repository 104 and not simply updating the integrated data repository 104 as new data becomes available, the integrated data repository 104 provides enhanced privacy protection with respect to the data stored by the integrated data repository 104. That is, in situations where the data repository is simply updated with new data, it may be easier to track individuals associated with data newly added to the data repository, since the number of new individuals added at any one time is typically less than the number of existing individuals who already have data stored by the data repository.
[0054] In various examples, the data stored by the integrated data repository 104 may be accessed via a database management system. Additionally, the integrated data repository 104 may store data according to one or more database models. In one or more examples, the integrated data repository 104 may store data according to one or more relational database technologies. For example, the integrated data repository 104 may store data according to a relational database model. In one or more additional examples, the integrated data repository 104 may store data according to an object-oriented database model. In one or more further examples, the integrated data repository 104 may store data according to an extensible markup language (XML) database model. In yet an additional example, the integrated data repository 104 may store data according to a structured query language (SQL) database model. In yet a further example, the integrated data repository 104 may store data according to an image database model.
[0055] The data integration system 114 may generate the integrated data repository 104 by generating multiple data tables and creating links between the data tables. The links may indicate logical connections between the data tables. The data integration system 114 may generate the data tables by extracting a specified set of data from information obtained from the data repositories 106, 108, 110, 112 and storing the data in rows and columns of the respective data tables. In various examples, the logical connections between the data tables may include at least one of a one-to-one link, where a row of information in one data table corresponds to a row of information in another data table, a one-to-many link, where a row of information in one data table corresponds to multiple rows of information in another data table, or a many-to-many link, where multiple rows of information in one data table correspond to multiple rows of information in another data table.
[0056] The multiple data tables may be configured according to the data repository schema 116. In the illustrative example of FIG. 1, the data repository schema 114 includes a first data table 118, a second data table 120, a third data table 122, a fourth data table 124, and a fifth data table 124. Although the illustrative example of FIG. 1 includes five data tables, in additional implementations, the data repository schema 116 may include more or fewer data tables. The data repository schema 116 may also include links between the data tables 118, 120, 122, 124, 128. The links between the data tables 118, 120, 122, 124, 126 may indicate that information retrieved from one of the data tables 118, 120, 122, 124, 126 may result in retrieval of additional information stored by one or more of the additional data tables 118, 120, 122, 124, 126. Also, all of the data tables 118, 120, 122, 124, 126 may not be linked to each of the other data tables 118, 120, 120, 122, 124, 126. In the illustrative example of FIG. 1, the first data table 118 is logically coupled to the second data table 118 by a first link 128, and the first data table 118 is logically coupled to the fourth data table 124 by a second link 130. In addition, the second data table 120 is logically coupled to the third data table 122 via a third link 132, and the fourth data table 124 is logically coupled to the fifth data table 126 via a fourth link 134. Furthermore, the third data table 122 is logically coupled to the fifth data table 126 via a fifth link 136.
[0057] In various examples, as data tables are added and / or removed from the data repository schema 116, additional links between data tables may be added or removed from the data repository schema 116. In one or more illustrated examples, the integrated data repository 104 may store data tables in accordance with the data repository schema 116 for at least a portion of individuals for whom the data integration system 114 obtained information from a combination of at least two of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, and the one or more reference information data repositories 112. As a result, the integrated data repository 104 may store instances of each of the data tables 118, 120, 122, 124, 126 in accordance with the data repository schema 116 for individuals ranging from thousands, tens of thousands, to hundreds of thousands or more.
[0058] The data integration and analysis system 102 may also include a data pipeline system 138. The data pipeline system 138 may include a group of algorithms, software codes, scripts, macros, or other computer executable instructions that process information stored by the integrated data repository 104 to generate additional data sets. The additional data sets may include information obtained from one or more of the data tables 118, 120, 122, 124, 126. The additional data sets may also include information derived from data obtained from one or more of the data tables 118, 120, 122, 124, 126. The components of the data pipeline system 138 implemented to generate the first additional data set may be different from the components of the data pipeline system 138 used to generate the second additional data set.
[0059] In one or more examples, the data pipeline system 138 may generate a data set indicative of pharmaceutical treatments received by a plurality of individuals. In one or more illustrative examples, the data pipeline system 138 may analyze information stored in at least one of the data tables 118, 120, 122, 124, 126 to determine health insurance codes corresponding to pharmaceutical treatments received by a plurality of individuals. The data pipeline system 138 may analyze the health insurance codes corresponding to the pharmaceutical treatments in relation to a library of data indicative of the designated pharmaceutical treatments corresponding to the one or more health insurance codes to determine the names of pharmaceutical treatments received by the individuals. In one or more additional examples, the data pipeline system 138 may analyze information stored by the integrated data repository 104 to determine medical procedures received by a plurality of individuals. Illustratively, the data pipeline system 138 may analyze information stored by one of the data tables 118, 120, 122, 124, 126 to determine treatments received by the individuals by at least one of injections or intravenous infusions. In one or more further examples, the data pipeline system 138 may analyze the information stored by the integrated data repository 104 to determine an individual's episodes of care, a course of therapy received by the individual, the progression of a biological condition, or time to next treatment. In various examples, the datasets generated by the data pipeline system 138 may be different for different biological conditions. For example, the data pipeline system 138 may generate a first plurality of datasets for a first type of cancer, such as lung cancer, and a second plurality of datasets for a second type of cancer, such as colon cancer.
[0060] The data pipeline system 138 may also determine one or more confidence levels to assign to information relating to an individual having data stored by the integrated data repository 104. Each confidence level may correspond to a different measure of accuracy of information relating to an individual having data stored by the integrated data repository 104. The information associated with each confidence level may correspond to one or more characteristics of the individual derived from data stored by the integrated data repository 104. Confidence level values for the one or more characteristics may be generated by the data pipeline system 138 in conjunction with generation of one or more data sets from the integrated data repository 104. In one or more examples, the first confidence level may correspond to a first range of accuracy measures, the second confidence level may correspond to a second range of accuracy measures, and the third confidence level may correspond to a third range of accuracy measures. In one or more additional examples, the second range of accuracy measures may include values less than the values of the first range of accuracy measures, and the third range of accuracy measures may include values less than the values of the second range of accuracy measures. In one or more illustrative examples, information corresponding to a first confidence level may be referred to as the highest (Gold standard) information, information corresponding to a second confidence level may be referred to as medium (Silver standard) information, and information corresponding to a third confidence level may be referred to as low (Bronze standard) information.
[0061] The data pipeline system 138 may determine a confidence level value for an individual's characteristic based on multiple factors. For example, each set of information may be used to determine the individual's characteristic. The data pipeline system 138 may determine a confidence level for an individual's characteristic based on the amount of completeness of each set of information used to determine the characteristic for an individual. In a situation where one or more pieces of information are missing from a set of information associated with a first plurality of individuals, the confidence level of the characteristic may be lower than for a second plurality of individuals where no information is missing from the set of information. In one or more examples, the amount of missing information may be used by the data pipeline system 138 to determine a confidence level for an individual's characteristic. Illustratively, a greater amount of missing information used to determine an individual's characteristic may result in a lower confidence level for the characteristic than a situation where a lesser amount of missing information used to determine the characteristic. Furthermore, different types of information may correspond to different confidence levels for a characteristic. In one or more examples, the presence of a first piece of information used to determine an individual's characteristic may result in a higher confidence level for the characteristic than the presence of a second piece of information used to determine the characteristic.
[0062] In one or more illustrative examples, the data pipeline system 138 may determine a number of individuals included in the cohort who have a primary diagnosis of lung cancer (or other biological condition). The data pipeline system 138 may determine a confidence level for each individual for being classified as having a primary diagnosis of lung cancer. The data pipeline system 138 may use information from a number of columns included in the data tables 118, 120, 122, 124, 126 to determine a confidence level for the individual's inclusion in the lung cancer cohort. The multiple columns may include health insurance codes related to the diagnosis of a biological condition and / or the treatment of a biological condition. Additionally, the multiple columns may correspond to a date of diagnosis and / or a date of treatment of a biological condition. The data pipeline system 138 may determine that the confidence level of the individual being characterized as being part of the lung cancer cohort is higher in a scenario in which information is available for each of the multiple columns, or at least a threshold number of columns, than when information is available for a number of columns less than the threshold number. Further, the data pipeline system 138 may determine a confidence level for an individual to be included in the lung cancer cohort based on the type of information and the availability of information for one or more columns. By way of example, in a situation in which one or more diagnosis codes are present and one or more treatment codes are absent in association with one or more time periods for a group of individuals, the data pipeline system 138 may determine that the confidence level for including the individuals in the lung cancer cohort is higher than in a situation in which at least one of the diagnosis codes is absent and a treatment code used to determine whether the individual is included in the lung cancer cohort is present.
[0063] The data integration and analysis system 102 may include a data analysis system 140. The data analysis system 148 may receive integrated data repository requests 142 from one or more computing devices, such as the exemplary computing device 144. The one or more integrated data repository requests 142 may cause data to be retrieved from the integrated data repository 104. In various examples, the one or more integrated data repository requests 142 may cause data to be retrieved from one or more datasets generated by the data pipeline system 138. The integrated data repository requests 142 may specify data to be retrieved from the integrated data repository 104 and / or the one or more datasets generated by the data pipeline system 138. In one or more additional examples, the integrated data repository requests 142 may include one or more pre-constructed queries corresponding to computer-executable instructions to retrieve a specified set of data from the one or more datasets generated by the integrated data repository 104 and / or the data pipeline system 138.
[0064] In response to the one or more integrated data repository requests 142, the data analysis system 140 may analyze data retrieved from at least one of the integrated data repository 104 or one or more data sets generated by the data pipeline system 138 to generate data analysis results 146. The data analysis results 146 may be sent to one or more computing devices, such as the exemplary computing device 148. While the illustrative example of FIG. 1 shows the one or more integrated data repository requests 142 and the data analysis results 146 from one computing device 144 being sent to another computing device 148, in one or more additional implementations, the data analysis results 146 may be received by the same computing device that sent the one or more integrated data repository requests 142. The data analysis results 146 may be displayed by one or more user interfaces rendered by the computing device 144 or the computing device 148.
[0065] In one or more examples, the data analysis system 140 may implement at least one of one or more machine learning techniques or one or more statistical techniques to analyze the data retrieved in response to the one or more integrated data repository requests 142. In one or more examples, the data analysis system 140 may implement one or more artificial neural networks to analyze the data retrieved in response to the one or more integrated data repository requests 142. By way of example, the data analysis system 140 may implement at least one of one or more convolutional neural networks or one or more residual neural networks to analyze the data retrieved from the integrated data repository 104 in response to the one or more integrated data repository requests 142. In at least some examples, the data analysis system 140 may implement one or more random forest techniques, one or more support vector machines, or one or more hidden Markov models to analyze the data retrieved in response to the one or more integrated data repository requests 142. Also, one or more statistical models may be implemented to analyze the data retrieved in response to the one or more integrated data repository requests 142 to identify at least one correlation or index of significance between individual characteristics. For example, a log-rank test may be applied to the data retrieved in response to the one or more integrated data repository requests 142. Additionally, a Cox proportional hazards model may be implemented with respect to the data retrieved in response to the one or more integrated data repository requests 142. Furthermore, a Wilcoxon signed-rank test may be applied to the data retrieved in response to the one or more integrated data repository requests 142. In yet another example, a z-score analysis may be performed with respect to the data retrieved in response to the one or more integrated data repository requests 142. In yet an additional example, a Kaplan-Meier analysis may be performed with respect to the data retrieved in response to the one or more integrated data repository requests 142.In at least some examples, one or more machine learning techniques may be implemented in combination with one or more statistical techniques to analyze data retrieved in response to one or more integrated data repository requests 142.
[0066] In one or more illustrative examples, the data analysis system 140 may determine a survival rate of an individual with lung cancer in response to one or more treatments. In one or more additional illustrative examples, the data analysis system 140 may determine a survival rate of an individual with one or more genomic region mutations with lung cancer in response to one or more treatments. In various examples, the data analysis system 140 may generate a data analysis result 146 in a situation where data retrieved from at least one of the integrated data repository 104 or one or more datasets generated by the data pipeline system 138 meets one or more criteria. For example, the data analysis system 140 may determine whether at least a portion of the data retrieved in response to one or more integrated data repository requests 142 meets a threshold confidence level. In a situation where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests 142 is below the threshold confidence level, the data analysis system 140 may refrain from generating at least a portion of the data analysis result 146. In scenarios where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests 142 is at least the threshold confidence level, the data analysis system 140 may generate at least a portion of the data analysis results 146. In various examples, the threshold confidence level may be related to the type of data analysis results 146 generated by the data analysis system 140.
[0067] In one or more illustrated examples, the data analysis system 140 may receive an integrated data repository request 142 that generates a data analysis result 146 indicative of a survival rate for one or more individuals. In these cases, the data analysis system 140 may determine whether the data stored by one or more datasets generated by the integrated data repository 104 and / or the data pipeline system 138 meets a threshold confidence level, such as a highest confidence level. In one or more additional examples, the data analysis system 140 may receive an integrated data repository request 142 that generates a data analysis result 146 indicative of a treatment received by one or more individuals. In such implementations, the data analysis system 140 may determine whether the data stored by one or more datasets generated by the integrated data repository 104 and / or the data pipeline system 138 meets a lower threshold confidence level, such as a lower confidence level.
[0068] In one or more additional illustrative examples, the data analysis system 140 may receive an integrated data repository request 142 to determine individuals who have one or more genomic mutations and have received one or more treatments for a biological condition. Continuing with this example, the data analysis system 140 may determine the survival rate of individuals with one or more genomic mutations in relation to the one or more treatments that those individuals have received. The data analysis system 140 may then identify the effectiveness of a treatment for the individuals based on the survival rate of the individuals in relation to the genomic mutations that may be present in the individuals. In this manner, health outcomes for individuals may be improved by identifying predictive treatments that may be more effective for a population of individuals with one or more genomic mutations than current treatments provided to the individuals.
[0069] FIG. 2 illustrates an exemplary framework 200 that supports the arrangement of data tables in a unified data repository, according to one or more implementations. In the illustrated example of FIG. 2, the framework 200 includes a data repository schema 202 that includes a first data table 204, a second data table 206, a third data table 208, a fourth data table 210, a fifth data table 212, a sixth data table 214, and a seventh data table 216. Although the illustrated example of FIG. 2 includes seven data tables, in additional implementations, the data repository schema 202 may include more or fewer data tables. The data repository schema 202 may also include links between the data tables 204, 206, 208, 210, 212, 214, 216. The links between the data tables 204, 206, 208, 210, 212, 214, 216 may indicate that information retrieved from one of the data tables 204, 206, 208, 210, 212, 214, 216 may result in retrieval of additional information stored by one or more additional data tables 204, 206, 208, 210, 212, 214, 216. In addition, not all of the data tables 204, 206, 208, 210, 212, 214, 216 may be linked to each of the other data tables 204, 206, 208, 210, 212, 214, 216. 2, the first data table 204 is logically coupled to the second data table 206 by a first link 218, and the third data table 208 is logically coupled to the second data table 206 by a second link 220. The second data table 206 is also logically coupled to the fourth data table 210 by a third link 222, the second data table 206 is logically coupled to the fifth data table 212 by a fourth link 224, and the second data table 206 is logically coupled to the sixth data table 214 by a fifth link 226.Additionally, the fifth data table 212 is logically coupled to the sixth data table 214 by a sixth link 228, which is logically coupled to the seventh data table 216 by a seventh link 230. Additionally, the seventh data table 216 is logically coupled to the fourth data table 210 by an eighth link 232. In various examples, as data tables are added and / or removed from the data repository schema 202, additional links between data tables may be added or removed from the data repository schema 202. In one or more illustrated examples, the integrated data repository 104 may store data tables according to the data repository schema 202 for at least a portion of individuals for which the data integration system 114 obtained information from a combination of at least two of the health insurance claims data repository 106, the molecular data repository 108, and the one or more additional data repositories 110. As a result, the unified data repository 104 may store instances of each of the data tables 204, 206, 208, 210, 212, 214, 216 in accordance with the data repository schema 204 for individuals ranging from thousands, tens of thousands, to hundreds of thousands or more of individuals.
[0070] In one or more examples, the first data table 204 may store data corresponding to an individual's genomics and genomics testing. For example, the first data table 204 may include columns containing information corresponding to the panel used to generate the genomics data, the mutations in the genomic region, the type of mutation, the copy number of the genomic region, coverage data indicating the number of nucleic acid molecules identified in the sample with one or more mutations, the test date, and patient information. The first data table 204 may also include one or more columns containing health insurance data codes that may correspond to one or more diagnostic codes. In addition, the information in the first data table 204 may include at least one identifier of an individual associated with an instance of the first data table 204.
[0071] The second data table 206 may store data related to one or more visits by an individual to one or more health care providers. The third data table 208 may store information corresponding to each service provided to the individual in relation to one or more visits to one or more health care providers represented by the second data table 206. By way of example, an individual may visit a health care provider and multiple services may be performed on the individual during the visit. The second data table 206 may include a column indicating information about each of multiple services performed during the visit. Multiple third data tables 208 may be generated for the visit, including columns indicating a higher level of information about each service provided during the visit than the information stored by the second data table 206 related to the visit. For example, the second data table 206 may include multiple columns indicating health insurance codes for various services provided to the individual during the visit, and a third data table 208 related to one of the services may include multiple columns for additional health insurance codes corresponding to additional information related to the respective service. The second data table 206 and the third data table 208 for the visit may indicate one or more service dates corresponding to the visit.
[0072] The fourth data table 210 may include columns that indicate information about individuals about whom information is stored by the integrated data repository 104. For example, the fourth data table 210 may include columns that indicate information related to at least one of the individual's location, the individual's gender, the individual's birth date, the individual's date of death (if applicable), or one or more keys associated with the individual. In one or more examples, the fourth data table 210 may include one or more columns related to whether erroneous data has ever been identified for an individual. In various examples, a single fourth data table 210 may be generated for each individual. Thus, the data repository schema 202 may include multiple instances of the fourth data table 210, ranging from thousands, tens of thousands, to hundreds of thousands or more individuals.
[0073] The fifth data table 212 may include columns that indicate information related to a health insurance company or government agency that paid for one or more services provided to each individual. For example, the fifth data table 212 may include one or more payer identifiers. The sixth data table 214 may include columns that contain information corresponding to health insurance coverage information for each individual. In one or more examples, the sixth data table 214 may include columns that indicate the existence of medical coverage for the individual, the existence of drug coverage for the individual, and the type of health insurance plan associated with the individual, such as a health maintenance organization (HMO) or preferred provider organization (PPO).
[0074] The seventh data table 216 may include columns indicating information related to the pharmaceutical treatments obtained by each individual. In one or more examples, the seventh data table 216 may include one or more columns indicating health insurance codes corresponding to the pharmaceutical treatments available through the pharmacy. The health insurance codes may correspond to the individual pharmaceutical treatments. In addition, the health insurance codes may indicate a diagnosis of a biological condition for the individual. The seventh data table 216 may also include additional information, such as at least one of dosage, prescription days, dosage, number of authorized refills, date of service, or information related to the individual receiving the pharmaceutical treatment.
[0075] In various examples, the data repository schema 202 may provide analysis of the information stored by the data tables 204, 206, 208, 210, 212, 214, 216 in a more efficient manner than a typical data repository schema. For example, the logical connections between the data tables 204, 206, 208, 210, 212, 214, 216 are configured to efficiently search for related data across different data tables 204, 206, 208, 210, 212, 214, 216. In situations where the data tables 204, 206, 208, 210, 212, 214, 216 are configured in series and / or a greater number of the data tables 204, 206, 208, 210, 212, 214, 216 are logically connected, retrieving data from one or more of the data tables 204, 206, 208, 210, 212, 214, 216 of the integrated data repository 104 to respond to a request for information from the integrated data repository 104 will be less efficient than in situations where the data repository schema 202 is implemented.
[0076] 3 illustrates an architecture 300 for generating one or more data sets from information retrieved from a data repository that integrates health-related data from multiple sources, according to one or more implementations. The architecture 300 may include a data integration and analysis system 102 and an integrated data repository 104. In addition, the data integration and analysis system 102 may include at least a data pipeline system 138 and a data analysis system 140. The data pipeline system 138 may include multiple sets of data processing instructions that are executable to generate respective data sets that may be analyzed by the data analysis system 140 in response to an integrated data repository request 142 that generates data analysis results 146.
[0077] The data pipeline system 138 may include a first data processing instruction 302, a second data processing instruction 304, up to an Nth data processing instruction 306. The data processing instructions 302, 304, 306 may be executable by one or more processing devices to perform a number of operations to generate respective data sets using information retrieved from the integrated data repository 104. In one or more illustrated examples, the data processing instructions 302, 304, 306 may include at least one of software code, scripts, API calls, macros, and the like. The first data processing instruction 302 may be executable to generate a first data set 308. In addition, the second data processing instruction 304 may be executable to generate a second data set 310. Furthermore, the Nth data processing instruction 306 may be executable to generate an Nth data set 312. In various examples, after the data integration and analysis system 102 generates the integrated data repository 104, the data pipeline system 138 may execute the data processing instructions 302, 304, 306 to generate the data sets 308, 310, 312. In one or more examples, the data sets 308, 310, 312 may be stored by the integrated data repository 104 or by an additional data repository accessible by the data integration and analysis system 102. At least a portion of the data processing instructions 302, 304, 306 may analyze health insurance codes to generate at least a portion of the data sets 308, 310, 312. Additionally, at least a portion of the data processing instructions 302, 304, 306 may analyze genomics data to generate at least a portion of the data sets 308, 310, 312.
[0078] In one or more examples, the first data processing instructions 302 may be executable to retrieve data from one or more first data tables stored by the integrated data repository 104. The first data processing instructions 302 may also be executable to retrieve data from one or more specified columns of the one or more first data tables. In various examples, the first data processing instructions 302 may be executable to identify individuals having health insurance codes stored in one or more combinations of columns and rows corresponding to one or more diagnostic codes. The first data processing instructions 302 may then be executable to analyze the one or more diagnostic codes to determine a biological condition for which the individual has been diagnosed. In one or more illustrative examples, the first data processing instructions 302 may be executable to analyze the one or more diagnostic codes with respect to a library of diagnostic codes indicative of one or more biological conditions corresponding to each diagnostic code. The library of diagnostic codes may include hundreds to thousands of diagnostic codes. The first data processing instructions 302 may also be executable to determine individuals who have been diagnosed with a biological condition by analyzing individual timing information, such as date of treatment, date of diagnosis, date of death, or one or more combinations thereof.
[0079] The second data processing instructions 304 may be executable to retrieve data from one or more second data tables stored by the integrated data repository 104. The second data processing instructions 304 may also be executable to retrieve data from one or more specified columns of the one or more second data tables. In various examples, the second data processing instructions 304 may be executable to identify individuals having health insurance codes stored in one or more combinations of columns and rows corresponding to one or more treatment codes. The one or more treatment codes may correspond to treatments obtained from a pharmacy. In one or more additional examples, the one or more treatment codes may correspond to treatments received by medical procedure, such as injections or intravenous infusions. The second data processing instructions 304 may be executable to determine one or more treatments corresponding to each health insurance code contained in the one or more second data tables by analyzing the health insurance codes in relation to a set of predetermined information. The set of predetermined information may include a data library indicating one or more treatments corresponding to one of hundreds to thousands of health insurance codes. The second data processing instructions 304 may generate a second dataset 310 indicative of respective treatments received by a group of individuals. In one or more illustrative examples, the group of individuals may correspond to individuals included in the first dataset 308. The second dataset 310 may be organized into rows and columns, with one or more rows corresponding to an individual and one or more columns indicative of the treatments received by each individual.
[0080] The Nth processing instructions 306 (N may be any positive integer) may be executable to generate the Nth dataset 312 by combining information from multiple previously generated datasets, such as the first dataset 308 and the second dataset 310. Additionally, the Nth processing instructions 306 may be executable to retrieve additional information from one or more additional columns of the integrated data repository 104 and generate the Nth dataset 312 for integrating the additional information from the integrated data repository 104 with information from the first dataset 308 and the second dataset 310. For example, the Nth processing instructions 306 may be executable to identify individuals included in the first dataset 308 who have been diagnosed with a biological condition and analyze specified columns of one or more additional data tables of the integrated data repository 104 to determine dates of treatment indicated in the second dataset 210 that correspond to those individuals included in the first dataset 308. In one or more further examples, the Nth processing instructions 306 may be executable to analyze columns of one or more additional data tables in the integrated data repository 104 to determine dosages of treatments indicated in the second dataset 310 received by individuals included in the first dataset 308. In this manner, the Nth processing instructions 306 may be executable to generate a dataset of episodes of care based on information included in the cohort dataset and the treatment dataset.
[0081] In one or more illustrative examples, in response to receiving an integrated data repository request 142, the data analysis system 140 may determine one or more data sets corresponding to characteristics of a query related to the integrated data repository request 142. For example, the data analysis system 140 may determine that information included in the first data set 308 and the second data set 310 is applicable to responding to the integrated data repository request 142. In such a scenario, the data analysis system 140 may analyze at least a portion of the data included in the first data set 308 and the second data set 310 to generate the data analysis result 146. In one or more additional examples, the data analysis system 140 may determine several different data sets for responding to different queries included in the integrated data repository request 142 to generate the data analysis result 146.
[0082] The use of a specific set of data processing instructions to generate each data set may reduce the number of inputs from a user of the data integration and analysis system 102 as well as reduce computational burdens, such as the amount of processing resources and memory utilized to process the integrated data repository requests 142. For example, without this specific architecture of the data pipeline system 138, each time an integrated data repository request 142 is received, the data utilized to respond to that integrated data repository request 142 is assembled from the data repositories 104. In contrast, by implementing the data pipeline system 138 to execute the data processing instructions 302, 304, 306 to generate the data sets 308, 310, 312, the data required to respond to the various integrated data repository requests 142 has already been assembled and may be accessed by the data analysis system 140 to respond to the integrated data repository requests 142. Thus, by implementing a data pipeline system 138 to generate data sets 308, 310, 312, fewer computing resources are used to respond to integrated data repository requests 142 than a typical system that performs an information analysis and collection process for each integrated data repository request 142. Furthermore, in situations where a data pipeline system 138 is not implemented, a user of the data integration and analysis system 102 may need to submit multiple integrated data repository requests 142 to analyze the information that the user intended to have analyzed. This is either because the ad-hoc collection of data to respond to an integrated data repository request 142 in a typical system is inaccurate, or because the data analysis system 140 is called multiple times in a typical system to perform an analysis of the information that may be performed using only one integrated data repository request 142 if a data pipeline system 138 is implemented.
[0083] FIG. 4 illustrates an architecture 400 for generating an integrated data repository including de-identified health insurance claims data and de-identified genomics data, according to one or more implementations. The architecture 400 may include a data integration and analysis system 102, a health insurance claims data repository 106, and a molecular data repository 108. The data integration and analysis system 102 may retrieve patient information 402 from the molecular data repository 108. The patient information 402 may include genomics data 404 of an individual having data stored by the molecular data repository 108. The genomics data 404 may represent results of one or more nucleic acid sequencing operations that analyze sequences of nucleic acid molecules contained in a sample obtained from the individual for one or more target genomic regions. In one or more examples, the sample may be obtained from tissue of one or more individuals. In one or more additional examples, the sample may be obtained from a bodily fluid of one or more individuals, such as blood or plasma. The one or more target genomic regions may correspond to genomic regions that correspond to the presence of one or more biological conditions. For example, the target region may correspond to a genomic region of a reference genome having a mutation present in an individual having a biological condition. In one or more illustrative examples, the target region may correspond to a genomic region of a reference human genome having one or more mutations present in an individual having one or more forms of cancer. The patient information 402 may also include information indicative of personal information about an individual for whom data is stored by the molecular data repository 108, as well as information corresponding to tests and analyses performed on a sample provided by the individual.
[0084] The data integration and analysis system 102 may perform a de-identification process 406 to anonymize personal information retrieved from the molecular data repository 108. The data integration and analysis system 102 may implement one or more computational techniques as part of the de-identification process to anonymize data related to individuals stored by the molecular data repository 108 such that the de-identified data protects the privacy of the individuals and complies with one or more privacy regulatory frameworks. The de-identification process 406 may include accessing 408 a token. In various examples, the token may consist of an alphanumeric string. In one or more examples, the token may be generated by the data integration and analysis system 102. In one or more additional examples, the token may be generated by a third party and retrieved by the data integration and analysis system 102.
[0085] The token may be generated using one or more hash functions in relation to the subset 410 of the patient information 402. By way of example, for individuals having information stored by the molecular data repository 108, the token may be generated using a combination of at least a portion of the respective individual's first name, at least a portion of the respective individual's last name, at least a portion of the respective individual's birth date, at least a portion of the respective individual's gender, and at least a portion of the respective individual's location identifier. The de-identification process 406 may also include generating 412 an identifier for the individual having data stored by the molecular data repository 108. The identifier may be generated by the data integration and analysis system 102 using one or more hash functions that are different from the one or more hash functions used to generate the token. In one or more illustrative examples, the data integration and analysis system 102 may generate each of the intermediate versions of the identifier using one or more hash functions and then apply one or more salting techniques to the intermediate versions of the identifier to generate the final version of the identifier. The salt function includes a function configured to add at least one random bit to each of the intermediate identifiers to generate each of the final identifiers. In various examples, the data integration and analysis system 102 may generate 412 an identifier using at least a portion of the information about each individual stored by the molecular data repository 108. In one or more illustrated examples, the identifier may be generated based on a patient identifier included in the patient information 402. The identifier generated by the data integration and analysis system 102 may be unique to each individual having data stored by the molecular data repository 108.
[0086] In operation 414, the data integration and analysis system 102 may generate modified patient information 416 based on the identifier. The modified patient information 416 may include genomics data 404 related to the individuals associated with the molecular data repository 108 and an identifier for each individual. The modified patient information 416 may have a data structure 418. The data structure 418 may include a column including an identifier for each of the individuals associated with the molecular data repository 108 and multiple columns including genomics data 404 related to those individuals, such as identifiers for one or more genes, one or more genetic alterations, types of genetic alterations, etc.
[0087] The data integration and analysis system 102 may generate a token file 420. The token file 420 may include a first token 422 accessed in operation 408 for each individual having data stored by the molecular data repository 108. The token file 420 may have a data structure 424 including a number of columns including information about each individual. The data structure 424 may include a column indicating each identifier generated by the data integration and analysis system 102 and a column indicating one or more first tokens 422 associated with each identifier. The data integration and analysis system 102 may send the token file 420 to a claims data management system 426 coupled to the claims data repository 106. The claims data management system 426 may analyze the first tokens 422 for corresponding second tokens 428. The second tokens 428 may be accessed or generated by the claims data management system 426. The second token 428 may be generated using the same or a similar subset of information about the individuals having information stored in the health insurance claims data repository 106 as the subset 410 of the patient information 402. For example, the second token 428 may be generated using a combination of at least a portion of each individual's first name, at least a portion of each individual's last name, at least a portion of each individual's date of birth, each individual's gender, and at least a portion of each individual's location identifier.
[0088] In various examples, the health insurance claims data management system 426 may retrieve health insurance claims data from the health insurance claims data repository 106 for an individual associated with each second token 428 that matches a corresponding first token 422. A first token 422 may match a second token 428 when the data of the first token 422 has at least a threshold amount of similarity to the data of the second token 428. In one or more examples, a first token 422 may match a second token 428 when the data of the first token 422 is the same as the data of the second token 428.
[0089] In response to identifying health insurance claim data for an individual having a respective second token 428 corresponding to a respective first token 422, the health insurance claim data management system 426 may generate modified health insurance claim data 430. The health insurance claim data management system 426 may send the modified health insurance claim data 430 to the data integration and analysis system 102. In one or more examples, the modified health insurance claim data 430 may be formatted according to a data structure 432. The data structure 432 may include a column that includes a subset of the second tokens 428 that correspond to the first tokens 422 and multiple columns that include the health insurance claim data.
[0090] At operation 434, the data integration and analysis system 102 may integrate the genomics data and health insurance claims data for individuals common to both the molecular data repository 108 and the health insurance claims data repository 106. The data integration and analysis system 102 may determine the individuals common to both the molecular data repository 108 and the health insurance claims data repository 106 by determining the genomics data and the health insurance claims data that correspond to the common tokens. The data integration and analysis system 102 may determine that a first token 422 related to a portion of the genomics data 404 corresponds to a second token 428 related to a portion of the health insurance claims data by determining a similarity metric between the first token 422 and the second token 428. In a scenario in which the first token 422 has at least a threshold amount of similarity to the second token 428, the data integration and analysis system 102 may associate the corresponding portion of the genomics data 404 and the corresponding portion of the health insurance claims data with the individual's identifier and store them in an integrated data repository, such as the integrated data repository 104 of Figures 1, 2, and 3.
[0091] Implementations of the architecture 400 may implement a cryptographic protocol that allows de-identified information from disparate data repositories to be integrated into a single data repository. In this manner, the security of the data stored by the integrated data repository 104 is increased. Additionally, the cryptographic protocol implemented by the architecture 400 may allow for more efficient searching and accurate analysis of the information stored by the integrated data repository 104 than in situations where the cryptographic protocol of the architecture 400 is not utilized. For example, by generating a token file 420 including a first token 422 using cryptographic techniques based on a specified set of information stored by the molecular data repository 104, and utilizing a second token 428 generated using the same or similar cryptographic techniques for a similar or same set of information stored by the health insurance claims data repository 106, the data integration and analysis system 102 may match information stored by the disparate data repositories that correspond to one and the same individual. Failure to implement the cryptographic protocols of architecture 400 may increase the probability of inaccurately determining that information from a data repository belongs to one or more individuals, thereby reducing the accuracy of the results provided by the data integration and analysis system 102 in response to an integrated data repository request 142 sent to the data integration and analysis system 102.
[0092] FIG. 5 illustrates a framework 500 for generating a dataset by the data pipeline system 138 based on data stored by the integrated data repository 104, according to one or more implementations. The integrated data repository 104 may store health insurance claims data and genomics data for a group of individuals 502. For example, the integrated data repository 104 may store information obtained from health insurance claims records 504 for the group of individuals 502. For each individual included in the group of individuals 502, the integrated data repository 104 may store information obtained from multiple health insurance claims records 504. In various examples, the information stored by the integrated data repository 104 may include and / or be derived from thousands, tens of thousands, hundreds of thousands, or even millions of health insurance claims records 504 for multiple individuals. In addition, each health insurance claim record may include multiple columns. As a result, the integrated data repository 104 may be generated through the analysis of millions of columns of health insurance claims data.
[0093] Additionally, while the health insurance claims data may be organized according to a structured data format, the health insurance claims data is typically configured to be viewed by health insurance providers, patients, and healthcare providers to show financial information and insurance code information related to services provided by the healthcare provider to the individual. Thus, the health insurance claims data may be available related to characteristics of an individual for whom a biological condition exists and may not be readily analyzed to obtain insights related to the biological condition that may aid in the treatment of the individual. The integrated data repository 104 may be generated and organized by analyzing and modifying the raw health insurance claims data such that the data stored by the integrated data repository 104 may be further analyzed to determine trends, characteristics, features, and / or insights related to individuals for whom one or more biological conditions may exist. For example, health insurance codes may be stored in the integrated data repository 104 in a manner such that at least one of a medical procedure, a biological condition, a treatment, a dosage, a pharmaceutical manufacturer, a pharmaceutical distributor, or a diagnosis may be determined for a given individual based on the individual's health insurance claims data. In various examples, the data integration and analysis system 102 may generate and implement one or more tables showing correlations between health insurance claims data and various treatments, symptoms, or biological conditions corresponding to the health insurance claims data. Additionally, the integrated data repository 104 may be generated using the genomics data records 506 of the group of individuals 502. In various examples, a large volume of health insurance claims data may be matched with the genomics data of the group of individuals 502 to generate the integrated data repository 104.
[0094] By integrating the genomics data records 506 of a group of individuals 502 with the health insurance claims records 504, the data integration and analysis system 102 may determine correlations between the presence of one or more biomarkers present in the genomics data records 506 and other characteristics of the individuals indicated by the health insurance claims data records 506 that existing systems may not typically be able to determine. For example, the data integration and analysis system 102 may determine one or more genomic characteristics of the individuals that correspond to treatments received by the individuals, the timing of the treatments, the dosage of the treatments, the individual's diagnosis, smoking status, the presence of one or more biological conditions, the presence of one or more symptoms of the biological conditions, one or more combinations thereof, etc. Based on the correlations determined by the data integration and analysis system 102 using the integrated data repository 104, cohorts of individuals that may benefit from one or more treatments that would not be identified by existing systems may be identified. In one or more examples, the processes and techniques implemented to integrate the health insurance claims records 504 and the genomics claims records 506 to generate the integrated data repository 104 may be complex, and efficiency enhancing techniques, systems, and processes may be implemented to minimize the amount of computing resources used to generate the integrated data repository 104.
[0095] In one or more illustrative examples, the data pipeline system 138 may access information stored by the integrated data repository 104 to generate a data set including a plurality of additional data records 508 including information related to at least a portion of the group of individuals 502. In the illustrative example of FIG. 5, the additional data records 508 include information indicating whether the individual is included in a cohort of individuals with lung cancer. The data pipeline system 138 may execute a plurality of different sets of data processing instructions to determine the cohort of the group of individuals 502 with lung cancer. In various examples, the additional data records 508 may indicate information used to determine the status of the individual 502 with respect to lung cancer, such as one or more transaction insurance identifiers, one or more International Classification of Diseases (ICD) codes, and one or more health insurance transaction dates. In addition to including a column indicating whether the individual 502 is included in a lung cancer cohort, the additional data records 508 may include a column indicating a confidence level of the status of the individual 502 with respect to the presence of lung cancer.
[0096] 6 is a schematic diagram of a computing architecture 600 for integrating medical record data into an integrated data repository 104. In various examples, at least a portion of the operations of the computing architecture 600 may be performed by the data integration and analysis system 102 of FIGS. 1, 3, and 4. In one or more examples, at least a portion of the operations of the computing architecture 600 may be performed by one or more additional computing systems controlled, maintained, and / or implemented by a service provider that also controls, maintains, and / or implements the data integration and analysis system 102. In one or more additional examples, at least a portion of the operations of the computing architecture 600 may be performed by multiple servers in a distributed computing environment.
[0097] The computing architecture 600 may include a medical record data repository 602. The medical record data repository 602 may store medical record data from multiple individuals. The medical record data may include imaging information, laboratory test results, diagnostic test information, clinical findings, dental hygiene information, medical practitioner annotations, medical history forms, diagnosis request forms, medical procedure order forms, medical information charts, one or more combinations thereof, and the like. In various examples, for a given individual, the medical record data repository 602 may store information obtained from one or more medical practitioners related to that individual.
[0098] The computing architecture 600 may perform an operation 604 that includes retrieving a data package from the medical record data repository 602. In one or more examples, the data package may be retrieved in response to one or more requests sent to the medical record data repository 602 for medical records corresponding to one or more individuals. In one or more additional examples, the data package may be retrieved by the computing architecture 600 using one or more application programming interface (API) calls. In one or more illustrative examples, a first data package 606, a second data package 608, through an Nth data package 610 may be retrieved using the computing architecture 600. Each data package 606, 608, 610 may correspond to a medical record for a respective individual. For example, the first data package 606 may include a medical record for a first individual, the second data package 608 may include a medical record for a second individual, and the Nth data package 610 may include a medical record for a third individual.
[0099] Each data package 606, 608, 610 may include multiple components. In one or more examples, each data package 606, 608, 610 may include individual components corresponding to medical records from different health care providers. In one or more additional examples, each data package 606, 608, 610 may include individual components corresponding to different portions of medical records corresponding to one or more health care providers. In the illustrated example of FIG. 6, the second data package 608 may include a first component 612, a second component 614, through an Nth component 616. In one or more illustrated examples, the first component 612 may include a first portion of an individual's medical record, the second component 614 may include a second portion of an individual's medical record, and the Nth component 616 may include a third portion of an individual's medical record. In various examples, the first component 612 may correspond to a first healthcare provider's medical record for the individual, the second component 614 may correspond to a second healthcare provider's medical record for the individual, and the third component may correspond to a third healthcare provider's medical record for the individual. In one or more additional illustrative examples, the first component 612 may include a first section of the individual's medical record, such as one or more forms related to a diagnostic test or procedure, and the second component 614 may include a second section of the individual's medical record, such as a pathology report for the individual.
[0100] In operation 618, the computing architecture 600 may pre-process the individual data packages to identify a corpus 620 of information to be analyzed. In one or more examples, pre-processing the data packages retrieved from the medical record data repository 602 may include transforming the data included in the data packages. For example, pre-processing the data packages may include converting at least a portion of the data retrieved from the medical record data repository 602 into machine-encoded information. By way of example, pre-processing the data packages may include performing one or more optical character recognition (OCR) operations on at least a portion of the data packages retrieved from the medical record data repository 602. By converting at least a portion of the data packages retrieved from the medical record data repository 602 into machine-encoded information, the data packages may be subjected to a number of operations, such as one or more parsing operations to identify one or more characters or strings of characters and one or more editing operations that cannot be performed on at least a portion of the data packages retrieved from the medical record data repository 602.
[0101] In one or more examples, pre-processing of the individual data packages may include determining information contained in the individual data packages that should be excluded from further analysis by the computing architecture 600. In various examples, one or more components of the individual data packages may be excluded from the corpus of information to be analyzed 620. For example, with respect to the second data package 608, the computing architecture 600 may determine that the first component 612 should be excluded from further analysis by the computing architecture 600. In one or more examples, the computing architecture 600 may analyze the components 612, 614, and / or 616 for one or more keywords to identify at least one of the components 612, 614, and / or 616 that should be excluded from further analysis by the computing architecture 600. In one or more illustrated examples, the computing architecture 600 may parse the components 612, 614, and / or 616 to identify one or more keywords, and in response to identifying the one or more keywords in the components 612, 614, and / or 616, the computing architecture 600 may determine to exclude the respective components 612, 614, and / or 616 from further analysis by the computing architecture 600. For example, the computing architecture 600 may determine that the first component 612 of the second data package 608 is a test order form for one or more diagnostic procedures or tests. In such a scenario, the computing architecture 600 may determine that the first component 612 should be excluded from further analysis by the computing architecture 600. Additionally, the computing architecture 600 may determine that at least one of the second components 614 and / or 616 corresponds to one or more pathology reports for an individual based on one or more keywords included in at least one of the second component 614 or the Nth component 616.In these cases, the computing architecture 600 may determine that at least a portion of the second component 614 and / or at least a portion of the Nth component 616 should be included in the corpus of information 620 for further analysis by the computing architecture 600.
[0102] Additionally, a subset of components of individual data packages retrieved from the medical record data repository 602 may be included in the corpus of information 620. In various examples, one or more additional operations may be performed to narrow the corpus of information 620. For example, one or more queries may be applied to the subset of information retrieved from the medical record data repository 602. The one or more queries may extract information from the one or more data packages that satisfies the one or more queries. In at least some examples, the one or more queries may be a group of queries that are applied to individual components of the data packages. In one or more illustrative examples, the group of queries may determine information to be included in the corpus of information 620 and additional information to be excluded from the corpus of information 620. In one or more additional examples, one or more sections of at least one component of a data package may be excluded from the corpus of information 620.
[0103] In one or more additional illustrative examples, after determining that the first component 612 should be excluded from further analysis by the computing architecture 600, the computing architecture 600 may then cause one or more queries to be performed on at least one of the second component 614 or the Nth component 616. In such a scenario, the one or more queries may determine that a section of the second component 614, such as a section indicating a family history of one or more biological conditions, should be excluded from the corpus of information 620. In various examples, the one or more queries may be to identify multiple keywords and / or combinations of keywords included in at least one of the second component 614 or the Nth component 616. In these cases, the computing architecture 600 may exclude one or more portions of individual components of the data package that include one or more keywords or combinations of keywords from the corpus of information 620. In one or more additional examples, computing architecture 600 may exclude from corpus of information 620 words, characters, and / or symbols following one or more keywords included in one or more portions of individual components of the data package.
[0104] Further, at operation 622, the computing architecture 600 may analyze the corpus of information to determine characteristics of the individuals. In one or more examples, the computing architecture 600 may analyze the corpus of information 620 to determine individuals having one or more phenotypes. In various examples, the computing architecture 600 may analyze the corpus of information 620 to determine one or more biomarkers indicative of a biological state. For example, the computing architecture 600 may analyze the corpus of information 620 to determine individuals having one or more genetic traits. The one or more genetic traits may include at least one of one or more mutations in a genomic region corresponding to a biological state. In one or more illustrative examples, the one or more genetic traits may correspond to one or more mutations in a genomic region corresponding to a type of cancer. In one or more additional illustrative examples, the one or more biomarkers may correspond to a level of an analyte that is outside of a specified range. By way of example, the computing architecture 600 may analyze the corpus of information 620 to determine individuals having a level of one or more proteins and / or a level of one or more small molecules present that are indicative of a biological state. In such a scenario, computing architecture 600 may analyze laboratory test results to determine the level of an analyte for an individual. In one or more additional examples, computing architecture 600 may analyze corpus of information 620 to determine individuals in which one or more symptoms indicative of a biological condition are present. In one or more further examples, computing architecture 600 may analyze imaging information included in corpus of information 620 to determine individuals in which one or more biomarkers are present.
[0105] In one or more examples, computing architecture 600 may implement one or more machine learning techniques to analyze corpus of information 620. For example, computing architecture 600 may implement one or more artificial neural networks, such as at least one of one or more convolutional neural networks or one or more residual neural networks, to analyze corpus of information 620. Computing architecture 600 may also implement at least one of one or more random forest techniques, one or more hidden Markov models, or one or more support vector machines to analyze corpus of information 620.
[0106] In at least some implementations, the computing architecture 600 may analyze the corpus of information 620 by performing one or more queries on the corpus of information 620. The one or more queries may correspond to one or more keywords and / or keyword combinations. The one or more keywords and / or keyword combinations may correspond to at least one of characters or symbols corresponding to one or more biological conditions. By way of example, the keywords may correspond to characters related to mutations in a genomic region, such as HER2. In one or more additional illustrative examples, one or more criteria may be associated with the keyword combination. By way of example, the criteria corresponding to the keyword combination may include multiple words that are present within a specified distance from each other in a portion of the corpus of information 620 for an individual, such as the words "fatigue," "blood pressure," and "bloating" occurring within 100 characters of each other. In these cases, the computing architecture 600 may parse the corpus of information 620 for its one or more keywords and / or keyword combinations. In various examples, in response to determining that the one or more keywords and / or combinations of keywords are present according to one or more criteria, the computing architecture 600 may determine that a biological condition is present for a given individual.
[0107] In one or more additional examples, the one or more queries may be image-based, and the computing architecture 600 may analyze the images in the corpus of information 620 against a template image. The template image may be generated based on analyzing a plurality of images in which a biological condition is present and aggregating the plurality of images into one template image. In such a scenario, the computing architecture 600 may analyze the images in the corpus of information 620 against the one or more template images to determine a similarity metric between the images in the corpus of information 620 and the template image. In situations where the similarity metric for an individual is at least a threshold, the computing architecture 600 may determine that a characteristic of a biological condition is present in the individual.
[0108] After determining the individuals having the one or more characteristics, the computing architecture 600 may generate a data structure in operation 624 to store data regarding the individuals having the one or more characteristics. In one or more examples, the computing architecture 600 may generate data tables that indicate individuals having individual characteristics and / or individuals having groups of characteristics. For example, the computing architecture 600 may generate a first data table 626 and a second data table 628. The first data table 626 may indicate individuals having one or more first characteristics, and the second data table 628 may indicate individuals having one or more second characteristics. In one or more illustrated examples, the first data table 626 may indicate individuals having one or more first biomarkers for a biological condition, and the second data table 628 may indicate individuals having one or more second biomarkers for the biological condition. The one or more first biomarkers may correspond to one or more first genomic variants associated with the biological condition, and the one or more second biomarkers may correspond to one or more second genomic variants associated with the biological condition. In various examples, the data tables 626, 628 may indicate whether one or more characteristics associated with the respective data tables 626, 628 are present for the respective individual. Illustratively, the first data table 626 may include a first indication for individuals in which the one or more first genomic variants are present, and a second indication for individuals in which the one or more first genomic variants are absent. In one or more additional examples, the first data table 626 may indicate the smoking status of the individual, and the second data table 628 may indicate whether the respective individual has undergone one or more treatments for a biological condition.
[0109] In one or more illustrative examples, the first data table 626 and the second data table 628 may have rows corresponding to individual individuals. In at least some examples, an individual identifier may be present in each row. The individual identifier may include alphanumeric characters or symbols corresponding to an individual. In various examples, the individual identifier may be present in a data package corresponding to an individual. The columns of the first data table 626 and the second data table 628 may indicate the status of the individual with respect to one or more characteristics. For example, the columns of the data tables 626, 628 may include identifiers including alphanumeric characters or symbols indicating the presence or absence of one or more characteristics for a given individual. Additionally, while the illustrative example of FIG. 6 includes the first data table 626 and the second data table 628, the computing architecture 600 may generate more or fewer data tables.
[0110] At operation 630, computing architecture 600 may store the data structures in an additional data repository. For example, computing architecture 600 may store at least first data table 626 and / or second data table 628 in intermediate data repository 632. In various examples, first data table 626 and second data table 628 may be temporarily stored in intermediate data repository 632. In one or more illustrated examples, first data table 626 and second data table 628 may be stored in intermediate data repository 632 before being added to unified data repository 104. In one or more examples, unified data repository 104 may be generated and / or updated periodically. In such a scenario, data structures generated by computing architecture 600 based on analyzing corpus of information 620 may be stored in intermediate data repository 632 until such time that unified data repository 104 is generated and / or updated.
[0111] Prior to adding the data structures stored by the intermediate data repository 632 to the integrated data repository 104, the computing architecture 600 may perform one or more de-identification processes in operation 634. The data structures stored by the intermediate data repository 632 may be de-identified to protect the privacy of individuals. The one or more de-identification processes may include applying one or more electronically implemented cryptographic techniques to the information of individuals contained in the data structures stored by the intermediate data repository 632. In one or more examples, the computing architecture 600 may generate tokens corresponding to each individual having information stored in the data structures of the intermediate data repository 632. The tokens may be generated by applying one or more hash functions to information relating to each individual. In one or more examples, the one or more de-identification processes may include applying a salt function to information corresponding to each individual to generate a token for the individual. In various examples, the one or more cryptographic techniques applied to de-identify the data structures stored by the intermediate data repository 632 may be the same or similar to those applied to information retrieved from the health insurance claims data repository 106 of FIG. 1 and FIG. 4.
[0112] At operation 636, the computing architecture 600 may store the de-identified data structure with the integrated data repository 104. For example, information stored in the intermediate data repository 632 for a given individual may be stored in the integrated data repository 104 along with additional information about the given individual. Illustratively, the integrated data repository 104 may store information about a given individual obtained from at least two of the molecular data repository 108, the health insurance claims data repository 106, and the intermediate data repository 632. In this manner, information about a given individual obtained from multiple separate data repositories may be stored in the integrated data repository 104. As a result, information about an individual obtained from different data repositories may be analyzed together, rather than separately as in many existing systems.
[0113] In various examples, the information stored by the intermediate data repository 632 may be used to validate one or more determinations made by the data integration and analysis system 102. For example, the data integration and analysis system 102 may analyze the information obtained from the health insurance claims data repository 106 and the molecular data repository 108 to determine characteristics of individuals. The data integration and analysis system 102 may then analyze the information obtained from the intermediate data repository 632 to determine whether expected characteristics identified from the information obtained from the health insurance claims data repository 106 and the molecular data repository 108 correspond to characteristics of the same individuals for the information stored by the intermediate data repository 632.
[0114] The one or more cryptographic techniques applied to de-identify the data structures stored by the intermediate data repository 632 may utilize the same or similar information as that used to generate at least one of the first token 422 or the second token 428 of FIG. 4. For example, operation 634 may implement one or more cryptographic techniques using a combination of at least a portion of each individual's first name, at least a portion of each individual's last name, at least a portion of each individual's birth date, the individual's gender, and at least a portion of each individual's location identifier to de-identify the data structures of the intermediate data repository. By utilizing the same or similar cryptographic techniques and the same or similar subset of information used to generate at least one of the first token 422 or the second token 428 to de-identify the data structures stored by the intermediate data repository 632, the information stored by the intermediate data repository 632 may be synchronized with information about the same individuals whose information is stored in the integrated data repository 104. Both the integrated data repository 104 and the intermediate data repository 632 may store information about thousands, tens of thousands, or even millions of individuals. Thus, without the ability to synchronize individuals whose records are stored by the integrated data repository 104 and the intermediate data repository 632 through the use of specified cryptographic protocols as described herein, the data structures in the integrated data repository 104 and the intermediate data repository 632 relating to the same individual may not be stored in a manner such that the information stored by the integrated data repository 104 and the information stored by the intermediate data repository 632 may be searched together for a given individual, which may lead to inaccurate information being provided by the data integration and analysis system 102.The absence of a specified cryptographic protocol as described herein may also lead to the use of more computing resources to determine the information stored in the integrated data repository 104 from other data sources and the information stored by the intermediate data repository 632 that corresponds to a given individual. Figures 7 and 8 show an exemplary process for generating an integrated data repository and generating a data set used in the analysis of the information stored by the integrated data repository. This exemplary process is described as a collection of blocks of a logical flow graph representing a sequence of operations that may be implemented as hardware, software, or a combination thereof. The blocks are referenced by numbers. In the context of software, the blocks represent computer-executable instructions stored on one or more computer-readable media that perform the recited operations when executed by one or more processing devices (such as hardware microprocessors). Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described blocks may be combined in any order and / or in parallel to implement a process.
[0115] FIG. 7 is a data flow diagram of an example process 700 for generating an integrated data repository storing health insurance claims data and genomics data, according to one or more implementations. At operation 702, the process 700 may include generating a data file including tokens generated using a first hash function. An individual token may correspond to each individual of a group of individuals having data stored by the molecular data repository. In one or more examples, an individual having data stored by the molecular data repository may be associated with one or more tokens. The tokens may be generated by applying one or more first hash functions to a subset of information corresponding to the group of individuals stored by the genomics data repository. In various examples, the individual tokens may be generated by applying one or more first hash functions to one or more combinations of at least a portion of a first name of each individual of the group of individuals, at least a portion of a last name of each individual of the group of individuals, a location identifier of each individual of the group of individuals, a gender of each individual of the group of individuals, and a date of birth of each individual of the group of individuals. In one or more illustrative examples, the tokens may be generated by a data integration and analysis system coupled to the genomics data repository. In one or more additional illustrative examples, the token may be generated by a third party system and accessed by a data integration and analysis system coupled to the molecular data repository. Process 700 may also include sending the data file to a health insurance claims data management system at operation 704. The health insurance claims data management system may match the token included in the data file with a second token accessed by the health insurance data management system and generated based on information stored by the health insurance claims data repository.
[0116] Additionally, at operation 706, the process 700 may include obtaining first data corresponding to the group of individuals from the health insurance claims data management system responsive to the data file, the first data including health insurance claims data. In some implementations, affirmative consent is obtained from members of the group of individuals to have their data transferred from the health insurance claims data management system. In one or more examples, the data is transferred in an anonymized form such that the data cannot be traced to individual members. The health insurance claims data management system may be coupled to a health insurance claims data repository that stores health insurance claims information for a plurality of individuals. In one or more examples, the health insurance claims data management system may analyze the tokens of the data file for additional tokens generated by the health insurance claims data management system. The additional tokens may be generated based on the same set of information used to generate the tokens included in the data file. However, the identity of the individuals may not be determined based on the tokens. In various examples, the health insurance claims data management system may match tokens included in the data file with additional tokens generated based on information stored by the health insurance claims data repository to determine individuals having information stored by the health insurance claims data repository who also have information stored by the genomics data repository. The technology disclosed herein complies with legislative and best practice privacy standards, such as HIPAA and GDPR.
[0117] At operation 708, the process 700 may include generating a plurality of identifiers using a second hash function different from the first hash function. In one or more examples, the individual identifiers may correspond to one or more tokens relating to each individual of the group of individuals. The identifiers may be unique to a given individual of the group of individuals and are de-identified. Additionally, the identifiers may be generated using information about the group of individuals stored by the genomics data repository that is different from the information stored by the genomics data repository used to generate the tokens. In various examples, intermediate identifiers may be generated by applying the second hash function to the information of each group of individuals, and a final version of the identifier may be generated by applying one or more salting techniques to the intermediate identifiers. The information stored by the genomics data repository for each individual may be stored in association with the identifier such that at least a portion of the information of a given individual stored by the genomics data repository may be accessed using the given individual's respective identifier.
[0118] Additionally, process 700 may include, at operation 710, obtaining second data from a molecular data repository for the group of individuals using the plurality of identifiers, and at operation 712, process 700 may include determining respective portions of the first data that correspond to respective portions of the second data for the group of individuals. For example, for a given individual, first data corresponding to health insurance claims data for the given individual may be identified in addition to second data corresponding to molecular data for the given individual, such as genomics data. In this manner, both health insurance claims data and molecular data may be identified for a given individual.
[0119] The process 700 may include, at operation 714, generating an integrated data repository that stores the respective portions of the first data and the respective portions of the second data in association with respective identifiers of the plurality of identifiers. For example, the integrated data repository may store health insurance claims data and genomics claims data for a given individual in association with an identifier that may be used to access the health insurance claims data and genomics claims data for the given individual. The information stored by the integrated data repository may be organized according to a data repository schema. For example, the integrated data repository may store health insurance claims data and genomics data for a group of individuals in multiple data tables. In one or more examples, the information stored by the multiple data tables may be linked. Illustratively, information related to a given individual stored by a first data table of the data repository schema may be linked to additional information related to the given individual stored by a second data table of the data repository schema. In this manner, information accessed in one data table of the data repository schema may result in access to additional information stored in another data table of the data repository schema.
[0120] In one or more illustrative examples, the data repository schema may include a first data table that stores genomics data for a group of individuals. For example, the first data table may store information corresponding to the panel used to generate the genomics data, mutations in the genomic region, the type of mutation, the copy number of the genomic region, coverage data indicating the number of nucleic acid molecules identified in the sample with one or more mutations, the test date, and patient information. The data repository schema may also include a second data table that stores data related to one or more visits by the individual to one or more health care providers, and a third data table that stores information corresponding to each service provided to the individual in relation to the one or more visits to the one or more health care providers indicated by the second data table. In addition, the data repository schema may include a fourth data table that stores personal information for the group of individuals, and a fifth data table that stores information related to a health insurance company or government agency that paid for the services provided to the group of individuals. Furthermore, the data repository schema may include a sixth data table that stores information corresponding to health insurance coverage information for the group of individuals, such as the type of health insurance plan associated with the group of individuals. The data repository schema may also include a seventh data table that stores information related to pharmaceutical treatments obtained by a group of individuals.
[0121] In one or more examples, the integrated data repository may also store medical records corresponding to at least a portion of the group of individuals. In these examples, the medical records may be retrieved from one or more data repositories that store the medical records. One or more optical character recognition (OCR) operations may be performed on the medical records. In addition, the medical records may be analyzed to determine one or more portions of additional information to be removed to generate the corpus of information. In various examples, the corpus of information may be analyzed to determine a portion of the subset of the group of additional individuals that corresponds to one or more biomarkers.
[0122] One or more data structures may be generated from the corpus of information that stores an identifier of the portion of the subset of the additional group of individuals and an indication that the portion of the subset of the additional group of individuals corresponds to one or more biomarkers. The one or more data structures may be stored by the intermediate data repository. One or more de-identification operations may be performed with respect to the identifier of the portion of the additional group of individuals before modifying the integrated data repository to store at least a portion of the additional information of the medical records of the portion of the subset of the additional group of individuals in association with the plurality of identifiers. After de-identifying the information stored by the one or more data structures, the information stored by the integrated data repository may be added to the integrated data repository. In at least some examples, the de-identified medical record information may be added to the integrated data repository in addition to or instead of the health insurance claims data. In various examples, the one or more data structures storing the de-identified medical record information in association with the biomarker data may have one or more logical connections with other data structures stored in the integrated data repository.By way of example, the one or more data structures storing de-identified medical record information in association with the biomarker data may have one or more logical connections to at least one of: a first data table that may store information corresponding to the panel used to generate the genomics data, mutations in the genomic region, the type of mutation, copy number of the genomic region, coverage data indicating the number of nucleic acid molecules identified in the sample with one or more mutations, test date, and patient information; a second data table that stores data related to one or more visits by an individual to one or more healthcare providers; a third data table that stores information corresponding to each service provided to the individual in connection with the one or more visits to the one or more healthcare providers represented by the second data table; a fourth data table that stores personal information of a group of individuals; a fifth data table that stores information related to health insurance companies or government agencies that paid for services provided to the group of individuals; a sixth data table that stores information corresponding to health insurance coverage information for a group of individuals, such as the type of health insurance plan associated with the group of individuals; or a seventh data table that stores information related to pharmaceutical treatments obtained by the group of individuals.
[0123] In various examples, the medical record data may be added to the integrated data repository by generating a data file including a first token generated using a first hash function. Each first token may correspond to a respective individual of the group of individuals having data stored by the molecular data repository. In addition, the data file may be sent to a medical record data management system, and medical record data corresponding to the group of individuals may be retrieved from the medical record data management system in response to the data file. Furthermore, a plurality of identifiers may be generated using a second hash function different from the first hash function. Each identifier may correspond to one or more tokens relating to each individual of the group of individuals. Using the plurality of identifiers, second data may be retrieved from the molecular data repository for the group of individuals. In various examples, for the group of individuals, respective portions of the first data corresponding to respective portions of the second data may be determined. In this manner, an integrated data repository may be generated that stores respective portions of the first data and respective portions of the second data in association with respective identifiers of the plurality of identifiers.
[0124] After an integrated data repository storing the medical record data is generated, a request may be received to determine data for a plurality of individuals having data stored in the integrated data repository. The request may include one or more search criteria. In one or more examples, a subset of the plurality of individuals having one or more characteristics corresponding to the one or more search criteria may be determined, and information of the subset of the plurality of individuals may be analyzed to determine an indication of the significance of the one or more characteristics with respect to a biological condition.
[0125] In one or more illustrative examples, one or more genomic mutations may be determined to be present in a subset of the plurality of individuals, and a plurality of treatments provided to the subset of the plurality of individuals may also be determined. In various examples, a survival rate of each of the subset of the plurality of individuals may be determined, such as a real-world survival rate. In at least some examples, the significance indicator may correspond to a survival rate for a treatment of the plurality of treatments and a genomic mutation of the one or more genomic mutations. Based on the significance indicator, the effectiveness of the treatment for the subset of the plurality of individuals may be determined. In one or more examples, an individual in the subset of the plurality of individuals who has not received the treatment may be determined. One or more therapeutically effective amounts of the treatment may be administered to an individual in the subset of the plurality of individuals who has not received the treatment.
[0126] FIG. 8 is a data flow diagram of an exemplary process 800 for generating a plurality of data sets used to analyze information stored by an integrated data repository that stores health insurance claims data and genomics data, according to one or more implementations. The process 800 may include, at operation 802, determining a first set of data processing instructions executable in relation to a first data stored by the integrated data repository. The integrated data repository may store health insurance claims data and molecular data for a common group of individuals. In one or more examples, the first set of data processing instructions may be included in a plurality of sets of data processing instructions that are part of a data processing pipeline. Each of the sets of data processing instructions of the data processing pipeline may be executed to generate a respective analyzable data set. For example, each of the sets of data processing instructions of the data processing pipeline may be executable to generate a data set that includes a specified portion and / or combination of information stored by the integrated data repository. In one or more additional examples, each of the sets of data processing instructions of the data processing pipeline may be executable to analyze and modify a portion of the information stored by the integrated data repository to generate a respective data set. In addition, different sets of data processing instructions may be executable on different subsets of the information stored by the unified data repository.
[0127] The process 800 may also include, at operation 804, executing a first set of data processing instructions to generate a first data set. The first data set may be indicative of a subset of a group of individuals in which a biological condition exists. The first set of data processing instructions may be executed to analyze the data stored by the integrated data repository to identify a cohort of individuals in which a biological condition exists. In one or more illustrative examples, the biological condition may include cancer. Illustratively, the first set of data processing instructions may be executed to analyze the data stored by the integrated data repository to identify a cohort of individuals in which lung cancer exists. In various examples, the data processing pipeline may include multiple sets of data processing instructions for identifying cohorts of individuals in which different biological conditions exist.
[0128] In one or more examples, the first set of data processing instructions may be executed to analyze at least one of the health insurance claims data or the molecular data to determine a cohort of individuals in which the biological condition exists. For example, the first set of data processing instructions may be executed to identify individuals having one or more health insurance codes present in the health insurance claims data to determine a group of individuals in which the biological condition exists. Additionally, the first set of data processing instructions may be executed to identify individuals in which one or more mutations exist in a genomic region of a nucleic acid molecule derived from a sample obtained from the individual to determine a group of individuals in which the biological condition exists.
[0129] Process 800 may also include, at operation 806, determining a second set of data processing instructions executable in relation to the second data stored by the integrated data repository. The second set of data stored by the integrated data repository may be different than the first set of data stored by the integrated data repository and may be analyzed in relation to the first set of data processing instructions. For example, the first data may correspond to a first column of one or more first data tables stored by the integrated data repository and the second data may correspond to a second column of one or more second data tables stored by the integrated data repository.
[0130] At operation 808, the process 800 may include executing a second set of data processing instructions to generate a second data set indicative of one or more treatments provided to a second subset of the group of individuals. The second data set may be indicative of a subset of the group of individuals who have received the one or more treatments. The one or more treatments may be provided to individuals with one or more biological conditions. In one or more examples, the second set of data processing instructions may be executed to analyze data stored by the integrated data repository to identify a cohort of individuals who have received the one or more treatments. Illustratively, the second set of data processing instructions may be executed to analyze at least one of health insurance claims data or genomics data to determine a cohort of individuals who have received the one or more treatments. In one or more illustrative examples, the second set of data processing instructions may be executed to identify individuals having one or more health insurance codes present in the health insurance claims data to determine a group of individuals who have received the one or more treatments.
[0131] Additionally, the process 800 may include, at operation 810, determining a third subset of the group of individuals that includes a portion of the first subset of the group of individuals that overlaps with a portion of the second subset of the group of individuals. As a result, the third subset of the group of individuals corresponds to individuals for whom the biological condition exists and for whom the one or more treatments are provided. At 812, the process 800 may include analyzing the first and second datasets with respect to the third subset of the group of individuals to determine an indication of significance of the characteristic of the third subset of the group of individuals. In one or more examples, one or more machine learning or statistical techniques may be applied to information contained in at least one of the first and second datasets with respect to the third subset of the group of individuals. The indication of significance may correspond to a statistical indication of significance for the characteristic. In one or more additional examples, the indication of significance may correspond to a probability that the characteristic is present in individuals for whom the biological condition exists.
[0132] In one or more illustrative examples, the characteristic may include one or more treatments provided to an individual in whom the biological condition exists. In one or more additional illustrative examples, the characteristic may include the presence of a mutation in a genomic region of a nucleic acid molecule derived from a sample obtained from an individual in whom the biological condition exists. In various examples, the information contained in at least one of the first dataset or the second dataset may be analyzed to determine the impact of the characteristic on one or more measures. In one or more examples, the information contained in at least one of the first dataset or the second dataset may be analyzed to determine the amount of impact of a treatment on the survival rate of an individual in whom the biological condition exists. In one or more further examples, the information contained in at least one of the first dataset or the second dataset may be analyzed to determine the amount of impact of a mutation in a genomic region on the survival rate of an individual in whom the biological condition exists. Additionally, the information contained in the first dataset and the second dataset may be analyzed to determine the amount of impact of one or more treatments on an individual in whom the biological condition exists and in whom one or more genomic mutations also exist.
[0133] 9 shows a schematic representation of a machine 9900 in the form of a computer system in which a set of instructions may be executed to cause the machine 900 to perform one or more of the methodologies discussed in the present invention, according to an example, according to an exemplary implementation. In particular, FIG. 8 shows a schematic representation of a machine 900 in the form of an exemplary computer system in which instructions 902 (e.g., software, programs, applications, applets, apps, or other executable code) may be executed to cause the machine 900 to perform any one or more of the methodologies discussed in the present invention. For example, the instructions 902 may cause the machine 900 to implement the architectures and frameworks 100, 200, 300, 400, 500, 600 described with respect to FIG. 1, FIG. 2, FIG. 3, FIG. 4, FIG. 5, and FIG. 6, respectively, and to perform the methods 700, 800 described with respect to FIG. 7 and FIG. 8, respectively.
[0134] The instructions 902 transform a general unprogrammed machine 900 into a specific machine 900 programmed to perform the described and illustrated functions as described. In alternative implementations, the machine 900 may operate as a stand-alone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 900 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 900 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a mobile phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing the instructions 902 that specify the operations to be performed by the machine 900. Further, although only one machine 900 is illustrated, the term "machine" is intended to include a collection of machines 900 that individually or cooperatively execute instructions 902 to perform any one or more of the methodologies discussed herein.
[0135] An example of a computing device 900 may include logic, one or more components, circuits (e.g., modules), or mechanisms. A circuit is a tangible entity configured to perform a certain operation. In one example, a circuit may be configured in a specified manner (e.g., internally or with respect to external entities such as other circuits). In one example, one or more computer systems (e.g., stand-alone, client or server computer systems) or one or more hardware processors (processors) may be configured as a circuit that operates by software (e.g., instructions, application portions, or applications) to perform certain operations described herein. In one example, the software may reside (1) in a non-transitory machine-readable medium or (2) in a transmission signal. In one example, the software, when executed by the underlying hardware of the circuit, causes the circuit to perform certain operations.
[0136] In one example, the circuitry may be implemented mechanically or electronically. For example, the circuitry may comprise dedicated circuitry or logic specifically configured to perform one or more of the techniques described above, including special purpose processors, field programmable gate arrays (FPGAs), or application specific integrated circuits (ASICs), etc. In one example, the circuitry may comprise programmable logic (e.g., circuitry contained within a general purpose processor or other programmable processor) that can be temporarily configured (e.g., by software) to perform certain operations. It will be appreciated that the decision to implement a circuitry mechanically (e.g., as a dedicated permanently configured circuitry) or as a temporarily configured circuitry (e.g., configured by software) is influenced by cost and time considerations.
[0137] Thus, the term "circuitry" is understood to encompass tangible entities, whether physically constructed and permanently configured (e.g., hardwired) or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specified manner or to perform specified operations. In one example, given multiple temporarily configured circuits, each of the circuits need not be configured or instantiated at any one time. For example, if a circuit includes a general-purpose processor that is configured via software, the general-purpose processor may be configured as different circuits at different times. Thus, the software may, for example, configure the processor to configure a particular circuit at a particular point in time and a different circuit at another time.
[0138] In one example, a circuit may provide information to and receive information from other circuits. In this example, a circuit may be considered to be communicatively coupled to one or more other circuits. When multiple such circuits are present simultaneously, communication may be achieved by signal transmission (e.g., across appropriate circuits and buses) connecting the circuits. In implementations in which multiple circuits are configured or instantiated at different times, communication between such circuits may be achieved, for example, through the storage and retrieval of information in a memory structure that can be accessed by multiple circuits. For example, one circuit may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. Yet another circuit may then access the memory device at a later time to retrieve and process the stored output. In one example, a circuit may be configured to initiate or receive communication with an input device or an output device and act on a resource (e.g., a collection of information).
[0139] Various operations of the example methods described herein may be performed, at least in part, by one or more processors that are temporarily or permanently configured (e.g., by software) to perform the associated operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented circuitry that operates to perform one or more operations or functions. In one example, circuitry referred to herein may include processor-implemented circuitry.
[0140] Similarly, the methods described herein may be at least partially implemented by a processor. For example, at least some or all of the operations of the method may be performed by one or more processors or circuitry implemented by a processor. The performance of some of the operations may be distributed to one or more processors and may be deployed across multiple machines rather than residing solely in one machine. In one example, one or more processors may be located in a single location (e.g., in a home environment, an office environment, or a server farm), while in other examples, the processors may be distributed across multiple locations.
[0141] The one or more processors may also operate to assist in performing operations related thereto within a "cloud computing" environment or as a "software as a service" (SaaS).
[0142] For example, at least some of the operations may be performed by a group of computers (as an example of a machine that includes a processor), which operations are accessible over a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application program interfaces (APIs)).
[0143] Exemplary implementations (e.g., devices, systems, or methods) may be implemented in digital electronic circuitry, computer hardware, firmware, software, or any combination thereof. Exemplary implementations may be implemented using a computer program product (e.g., a computer program tangibly embodied in an information carrier or machine-readable medium for execution by or to control the operation of a data processing apparatus, such as a programmable processor, a computer, or multiple computers).
[0144] The computer program may be written in any type of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a software module, subroutine, or other unit suitable for use in a computing environment. The computer program may be deployed to be executed on one computer or multiple computers at one location, or may be distributed across multiple locations and interconnected by a communication network.
[0145] In one example, the operations may be performed by one or more programmable processors that execute computer programs to perform functions by operating on input data and generating output. Example method operations may also be performed by, or example apparatus may be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)).
[0146] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. It will be appreciated that in an implementation that deploys a programmable computing system, both hardware and software architectures need to be considered. In particular, it will be appreciated that the choice of whether to implement a particular function as permanently configured hardware (e.g., an ASIC), temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of permanently and temporarily configured hardware may be a design choice. Below, hardware (e.g., computing device 900) and software architectures that may be deployed in an exemplary implementation are described.
[0147] In one example, computing device 900 may operate as a stand-alone device, or computing device 900 may be connected (eg, networked) to other machines.
[0148] In a networked deployment, the computing device 900 may operate in the capacity of either a server machine or a client machine in a server-client network environment. In one example, the computing device 900 may operate as a peer machine in a peer-to-peer (or other distributed) network environment. The computing device 900 may be a personal computer (PC), a tablet PC, a set-top box (STB), a mobile phone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify operations taken (e.g., performed) by the computing device 900. Additionally, although only one computing device 900 is illustrated, the term "computing device" shall be construed to include any collection of machines that individually or cooperatively execute a set (or sets) of instructions to perform one or more of the methodologies discussed herein.
[0149] The exemplary computing device 900 may include a processor 904 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory 906, and a static memory 908, some or all of which may communicate with each other via a bus 910. The computing device 900 may further include a display device 912, an alphanumeric input device 914 (e.g., a keyboard), and a user interface (UI) navigation device 916 (e.g., a mouse). In one example, the display device 912, the input device 914, and the UI navigation device 916 may be touch screen displays. The computing device 900 may additionally include a storage device (e.g., a drive device) 918, a signal generating device 920 (e.g., a speaker), a network interface device 922, and one or more sensors 924, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or another sensor.
[0150] The storage device 918 may include a machine-readable medium 926 having stored thereon one or more sets of data structures or instructions 902 (e.g., software) that implement or are utilized by any one or more of the methodologies or functions described herein. The instructions 902 may also reside, completely or at least partially, within the main memory 906, within the static memory 908, or within the processor 904 when executed by the computing device 900. In one example, one or any combination of the processor 904, the main memory 906, the static memory 908, or the storage device 918 may constitute a machine-readable medium.
[0151] While the machine-readable medium 926 is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a central or distributed database, and / or associated caches and servers) configured to store one or more instructions 902. The term "machine-readable medium" may also be interpreted to include any tangible medium capable of storing, encoding, or retaining instructions for execution by a machine, causing a machine to perform one or more of the methodologies of the present disclosure, or capable of storing, encoding, or retaining data structures utilized by or associated with such instructions. Accordingly, the term "machine-readable medium" may be interpreted to include, but is not limited to, solid-state memory, and optical and magnetic media. Specific examples of machine-readable media may include non-volatile memory, including, by way of example, semiconductor memory devices (e.g., electrically programmable read-only memory).
[0152] (EPROM, Electrically Erasable Programmable Read Only Memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0153] The instructions 902 may further be transmitted or received over a communications network 828 using a transmission medium via a network interface device 822 utilizing any one of a number of transport protocols (e.g., Frame Relay, IP, TCP, UDP, HTTP, etc.). Exemplary communications networks may include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile telephone networks (e.g., cellular networks), plain old voice telephone (POTS) networks, and wireless data networks (e.g., the IEEE 802.11 family of standards known as Wi-Fi®, the IEEE 802.16 family of standards known as WiMax®), peer-to-peer (P2P) networks, among others. The term "transmission medium" shall be construed to include any intangible medium capable of storing, encoding, or retaining instructions for execution by a machine, including digital or analog communications signals or other intangible media facilitating the communication of such software.
[0154] As used herein, a component may refer to a device, a physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that allow partitioning or modularization of certain processing or control functions. A component may be combined with other components through its interfaces to perform a machine process. A component may be a packaged functional hardware unit designed as part of a program to perform a specific function, usually among related functions, for use with other components. A component may constitute either a software component (e.g., code embodied on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit capable of performing a specific operation and may be configured or arranged in a specific physical manner. In various exemplary implementations, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) may be configured as a hardware component that operates by software (e.g., an application or application portion) to perform certain operations described herein.
[0155] A non-limiting numbered list of aspects of the present subject matter is presented below.
[0156] Aspect 1. A method comprising: generating, by a computing system including a processing circuit and a memory, a data file including first tokens generated using a first hash function, each first token corresponding to a respective individual of a group of individuals having data stored by a molecular data repository; sending, by the computing system, the data file to a health insurance claims data management system; retrieving, by the computing system, health data corresponding to the group of individuals from the health insurance claims data management system responsive to the data file; generating, by the computing system, a plurality of identifiers using a second hash function different from the first hash function, each identifier corresponding to one or more tokens relating to each individual of the group of individuals; retrieving, by the computing system, second data for the group of individuals from the molecular data repository using the plurality of identifiers; determining, by the computing system, respective portions of the first data corresponding to respective portions of the second data for the group of individuals; and generating, by the computing system, a unified data repository storing respective portions of the first data and respective portions of the second data in association with respective identifiers of the plurality of identifiers.
[0157] Aspect 2. The method of aspect 1, comprising: determining, by a computing system, a first set of data processing instructions executable in relation to first data stored by the integrated data repository, causing the computing system to execute the first set of data processing instructions to analyze first health insurance claim codes included in the first data to determine a first subset of a group of individuals for which a biological condition exists, and generating, by the computing system, a first data set indicative of the subset of the group of individuals for which the biological condition exists.
[0158] Aspect 3. The method of aspect 2, comprising: determining, by the computing system, a second set of data processing instructions executable in relation to second data stored by the integrated data repository, causing the computing system to execute the second set of data processing instructions to analyze second health insurance claim codes included in the second data to determine one or more treatments provided to a second subset of the group of individuals, and generating, by the computing system, a second data set indicative of the one or more treatments provided to the second subset of the group of individuals.
[0159] Example 4. The method of example 3, comprising: determining, by a computing system, a third subset of the group of individuals that includes a portion of the first subset of the group of individuals that overlaps with a portion of the second subset of the group of individuals; receiving, by the computing system, a request to perform an analysis of the first dataset and the second dataset in relation to the third subset of the group of individuals; and, in response to the request, analyzing, by the computing system, the first dataset and the second dataset with respect to the third subset of the group of individuals to determine an indication of significance of a characteristic of the third subset of the group of individuals with respect to a biological condition.
[0160]
[0023] Example 5. The method of example 4, comprising: determining, by a computing system, one or more genomic mutations present in a third subset of the group of individuals; determining, by the computing system, a plurality of treatments provided to the third subset of the group of individuals; and determining, by the computing system, a survival rate for each of the third subset of the group of individuals.
[0161] Embodiment 6. The method of embodiment 5, wherein the indication of significance corresponds to a survival rate for a treatment of the plurality of treatments and a genomic mutation of the one or more genomic mutations.
[0162]
[0023] Example 7. The method of example 6, comprising determining, by the computing system, efficacy of the treatment for a third subset of the group of individuals based on the indication of significance.
[0163]
[0023] Embodiment 8. The method of embodiment 7, comprising determining, by the computing system, individuals among a third subset of the group of individuals that have not received the treatment.
[0164] Embodiment 9. The method of embodiment 8, comprising administering a therapeutically effective amount of one or more treatments to individuals in the third subset who have not received the treatment.
[0165] Embodiment 10. The method of any one of embodiments 1-9, wherein the integrated data repository is configured according to a data repository schema including a plurality of data tables and a plurality of logical links between the plurality of data tables, each logical link of the plurality of logical links indicating one or more rows of one of the plurality of data tables that correspond to one or more further rows of the further data table of the plurality of data tables.
[0166] Embodiment 11. The method of embodiment 10, wherein the plurality of data tables includes a first data table storing genomics data for the group of individuals, a second data table storing data relating to one or more visits by the individuals to one or more health care providers, a third data table storing information corresponding to each service provided to the individuals in relation to the one or more visits to the one or more health care providers indicated by the second data table, a fourth data table storing personal information of the group of individuals, a fifth data table storing information relating to health insurance companies or government agencies that paid for the services provided to the group of individuals, a sixth data table storing information corresponding to health insurance coverage information for the group of individuals, and a seventh data table storing information relating to pharmaceutical treatments obtained by the group of individuals.
[0167] Aspect 12. The method of any one of aspects 1-11, wherein the plurality of identifiers generated using the second hash function include intermediate identifiers, and the method includes applying, by the computing system, a salt function to the intermediate identifiers to generate a final set of identifiers.
[0168] Example 13. The method of any one of Examples 1-12, comprising: obtaining, by a computing system, information from an additional data repository containing electronic medical records of an additional group of individuals; determining, by the computing system, a subset of the additional group of individuals that corresponds to the group of individuals having data stored by the genomics data repository; and modifying, by the computing system, the integrated data repository to store at least a portion of the information of the medical records of the subset of the additional group of individuals in association with the plurality of identifiers.
[0169] Aspect 14. The method of aspect 13, comprising: performing, by the computing system, one or more optical character recognition operations on the additional information; and analyzing, by the computing system, the additional information obtained from the additional data repository to determine one or more portions of the additional information to be removed to generate the corpus of information.
[0170] Example 15. The method of example 14, comprising: analyzing, by a computing system, the corpus of information to determine a portion of the subset of the group of additional individuals that corresponds to the one or more biomarkers; and generating, by the computing system, one or more data structures that store identifiers of the portion of the subset of the group of additional individuals and that store an indication that the portion of the subset of the group of additional individuals corresponds to the one or more biomarkers.
[0171] Example 16. The method of example 15, including storing, by a computing system, the one or more data structures in an intermediate data repository, and performing, by the computing system, one or more de-identification operations with respect to the identifiers of the portion of the subset of the group of additional individuals before modifying the integrated data repository to store at least a portion of the additional information of the medical records of the portion of the subset of the group of additional individuals in association with the plurality of identifiers.
[0172] Embodiment 17. The method of any one of embodiments 1-16, wherein the molecular data repository stores at least one or more of genomic information, genetic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, or proteomic information.
[0173] Aspect 18. A system, comprising: one or more hardware processing devices; and one or more computer-readable storage media having computer-executable instructions stored thereon, the computer-executable instructions, when executed by the one or more hardware processing devices, cause the system to perform operations including: generating a data file including first tokens generated using a first hash function, each first token corresponding to a respective individual of a group of individuals having data stored by a molecular data repository; sending the data file to a health insurance claims data management system; obtaining health insurance claims data corresponding to the group of individuals from the health insurance claims data management system responsive to the data file; generating a plurality of identifiers using a second hash function different from the first hash function, each identifier corresponding to one or more tokens relating to a respective individual of the group of individuals; obtaining second data from the molecular data repository for the group of individuals using the plurality of identifiers; determining respective portions of the first data corresponding to respective portions of the second data for the group of individuals; and generating an integrated data repository that stores the respective portions of the first data and the respective portions of the second data in association with respective identifiers of the plurality of identifiers.
[0174] Aspect 19. The system of aspect 18, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining a first set of data processing instructions executable in relation to the first data stored by the integrated data repository, executing the first set of data processing instructions to analyze a first health insurance claim code included in the first data to determine a first subset of a group of individuals for whom a biological condition exists, and generating a first data set indicative of the subset of the group of individuals for whom the biological condition exists.
[0175] Aspect 20. The system of aspect 19, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including: determining, by the computing system, a second set of data processing instructions executable in relation to second data stored by the integrated data repository, causing the computing system to execute the second set of data processing instructions to analyze second health insurance claim codes included in the second data to determine one or more treatments provided to a second subset of the group of individuals, and generating, by the computing system, a second data set indicative of the one or more treatments provided to the second subset of the group of individuals.
[0176]
[0023] Embodiment 21. The system of embodiment 20, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including: determining, by the computing system, a third subset of the group of individuals that includes a portion of the first subset of the group of individuals that overlaps with a portion of the second subset of the group of individuals; receiving, by the computing system, a request to perform an analysis of the first dataset and the second dataset in relation to the third subset of the group of individuals; and, in response to the request, analyzing, by the computing system, the first dataset and the second dataset with respect to the third subset of the group of individuals to determine an indication of significance of a characteristic of the third subset of the group of individuals with respect to a biological condition.
[0177]
[0023] Example 22. The system of example 21, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining one or more genomic mutations present in a third subset of the group of individuals, determining a plurality of treatments provided to the third subset of the group of individuals, and determining a survival rate for each of the third subset of the group of individuals.
[0178]
[0023] Embodiment 23. The system of embodiment 22, wherein the indication of significance corresponds to a survival rate for a treatment of the plurality of treatments and a genomic mutation of the one or more genomic mutations.
[0179]
[0023] Example 24. The system of example 23, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining effectiveness of treatment for a third subset of the group of individuals based on the indication of significance.
[0180]
[0023] Example 25. The system of example 24, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining individuals in a third subset of the group of individuals that have not received the treatment.
[0181]
[0023] Aspect 26. The system of any one of aspects 18-25, wherein the integrated data repository is configured according to a data repository schema including a plurality of data tables and a plurality of logical links between the plurality of data tables, each logical link of the plurality of logical links indicating one or more rows of one of the plurality of data tables that correspond to one or more additional rows of the further data table of the plurality of data tables.
[0182]
[0023] Embodiment 27. The system of embodiment 26, wherein the plurality of data tables includes a first data table storing genomics data for a group of individuals, a second data table storing data relating to one or more visits by the individuals to one or more health care providers, a third data table storing information corresponding to each service provided to the individuals for the one or more visits to the one or more health care providers indicated by the second data table, a fourth data table storing personal information of the group of individuals, a fifth data table storing information relating to health insurance companies or government agencies that paid for services provided to the group of individuals, a sixth data table storing information corresponding to health insurance coverage information for the group of individuals, and a seventh data table storing information relating to pharmaceutical treatments obtained by the group of individuals.
[0183] Example 28. The system of any one of Examples 18-27, wherein the plurality of identifiers generated using the second hash function includes intermediate identifiers, and wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including applying a salt function to the intermediate identifiers to generate a final set of identifiers.
[0184] Embodiment 29. The system of any one of embodiments 18-28, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including obtaining information from an additional data repository containing electronic medical records of an additional group of individuals, determining a subset of the additional group of individuals that corresponds to the group of individuals having data stored by the genomics data repository, and modifying the integrated data repository to store at least a portion of the information of the medical records of the subset of the additional group of individuals in association with the plurality of identifiers.
[0185]
[0023] Aspect 30. The system of aspect 29, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including performing one or more optical character recognition operations on the additional information and analyzing the additional information obtained from the additional data repository to determine one or more portions of the additional information to be removed to generate the corpus of information.
[0186] Embodiment 31. The system of embodiment 30, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including: analyzing the corpus of information to determine a portion of the subset of the additional group of individuals that corresponds to the one or more biomarkers; and generating one or more data structures that store an identifier of the portion of the subset of the additional group of individuals and that store an indication that the portion of the subset of the additional group of individuals corresponds to the one or more biomarkers.
[0187] Aspect 32. The system of claim 31, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including storing the one or more data structures in the intermediate data repository and performing one or more de-identification operations with respect to identifiers of the portion of the additional subset of the group of individuals before modifying the integrated data repository to store at least a portion of additional information of the medical records of the portion of the subset of the group of individuals in association with the plurality of identifiers.
[0188] Example 33. The system of any one of examples 18-32, wherein the molecular data repository stores at least one or more of genomic information, genetic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, or proteomic information.
[0189] Aspect 34. One or more non-transitory computer-readable storage media having computer-executable instructions stored thereon that, when executed by one or more hardware processing devices, cause a system to perform operations including: generating a data file including first tokens generated using a first hash function, each first token corresponding to a respective individual of a group of individuals having data stored by a molecular data repository; sending the data file to a health insurance claims data management system; obtaining health insurance claims data corresponding to the group of individuals from the health insurance claims data management system responsive to the data file; generating a plurality of identifiers using a second hash function different from the first hash function, each identifier corresponding to one or more tokens relating to a respective individual of the group of individuals; obtaining second data from the molecular data repository for the group of individuals using the plurality of identifiers; determining respective portions of the first data corresponding to respective portions of the second data for the group of individuals; and generating an integrated data repository that stores respective portions of the first data and respective portions of the second data in association with respective identifiers of the plurality of identifiers.
[0190] Aspect 35. The one or more non-transitory computer readable media of aspect 34 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining a first set of data processing instructions executable in relation to the first data stored by the integrated data repository, executing the first set of data processing instructions to analyze a first health insurance claim code included in the first data to determine a first subset of a group of individuals for which a biological condition exists, and generating a first data set indicative of the subset of the group of individuals for which the biological condition exists.
[0191] Aspect 36. The one or more non-transitory computer readable media of aspect 35 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including: determining, by the computing system, a second set of data processing instructions executable in relation to second data stored by the integrated data repository; causing the computing system to execute the second set of data processing instructions to analyze second health insurance claim codes included in the second data to determine one or more treatments provided to a second subset of the group of individuals; and generating, by the computing system, a second data set indicative of the one or more treatments provided to the second subset of the group of individuals.
[0192]
[0023] Aspect 37. The one or more non-transitory computer readable media of aspect 36 comprising additional computer-executable instructions that when executed by the one or more hardware processing devices cause the system to perform additional operations including: determining, by the computing system, a third subset of the group of individuals that includes a portion of the first subset of the group of individuals that overlaps with a portion of the second subset of the group of individuals; receiving, by the computing system, a request to perform an analysis of the first dataset and the second dataset in relation to the third subset of the group of individuals; and analyzing, by the computing system, in response to the request, the first dataset and the second dataset with respect to the third subset of the group of individuals to determine an indication of significance of a characteristic of the third subset of the group of individuals with respect to a biological condition.
[0193]
[0023] Embodiment 38. The one or more non-transitory computer readable media of embodiment 37 comprising additional computer executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining one or more genomic mutations present in a third subset of the group of individuals, determining a plurality of treatments provided to the third subset of the group of individuals, and determining a survival rate for each of the third subset of the group of individuals.
[0194]
[0023] Embodiment 39. The one or more non-transitory computer readable media of embodiment 38, wherein the indication of significance corresponds to a survival rate for a treatment of the plurality of treatments and a genomic mutation of the one or more genomic mutations.
[0195] Aspect 40. The one or more non-transitory computer-readable media of claim 39 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining effectiveness of a treatment for a third subset of the group of individuals based on the indication of significance.
[0196]
[0036] Aspect 41. The one or more non-transitory computer readable media of aspect 40 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining individuals in a third subset of the group of individuals that have not received the treatment.
[0197]
[0023] Aspect 42. The one or more non-transitory computer readable media of aspect 34, wherein the integrated data repository is configured according to a data repository schema including a plurality of data tables and a plurality of logical links between the plurality of data tables, each logical link of the plurality of logical links indicating one or more rows of a data table of the plurality of data tables that correspond to one or more further rows of the further data table of the plurality of data tables.
[0198]
[0023] Embodiment 43. The one or more non-transitory computer readable media of embodiment 42, wherein the plurality of data tables includes a first data table storing genomics data for a group of individuals, a second data table storing data relating to one or more visits by the individuals to one or more healthcare providers, a third data table storing information corresponding to respective services provided to the individuals for the one or more visits to the one or more healthcare providers indicated by the second data table, a fourth data table storing personal information of the group of individuals, a fifth data table storing information relating to health insurance companies or government agencies that paid for services provided to the group of individuals, a sixth data table storing information corresponding to health insurance coverage information for the group of individuals, and a seventh data table storing information relating to pharmaceutical treatments obtained by the group of individuals.
[0199] Aspect 44. The one or more non-transitory computer-readable media of any one of Aspects 34-43, wherein the plurality of identifiers generated using the second hash function includes intermediate identifiers, the one or more non-transitory computer-readable storage media storing additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including applying a salt function to the intermediate identifiers to generate a final set of identifiers.
[0200]
[0036] Embodiment 45. The one or more non-transitory computer readable media of embodiment 44 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including obtaining information from an additional data repository containing electronic medical records of an additional group of individuals, determining a subset of the additional group of individuals that corresponds to the group of individuals having data stored by the genomics data repository, and modifying the integrated data repository to store at least a portion of the information of the medical records of the subset of the additional group of individuals in association with the plurality of identifiers.
[0201] Aspect 46. The one or more non-transitory computer-readable media of aspect 45 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including performing one or more optical character recognition operations on the additional information and analyzing the additional information obtained from the additional data repository to determine one or more portions of the additional information to be removed to generate a corpus of information.
[0202] Embodiment 47. The one or more non-transitory computer readable media of embodiment 46 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including: analyzing the corpus of information to determine a portion of the subset of the additional group of individuals that corresponds to the one or more biomarkers; and generating one or more data structures that store an identifier of the portion of the subset of the additional group of individuals and that store an indication that the portion of the subset of the additional group of individuals corresponds to the one or more biomarkers.
[0203] Aspect 48. The one or more non-transitory computer readable media of aspect 47 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including storing the one or more data structures in the intermediate data repository and performing one or more de-identification operations with respect to the identifiers of the portion of the additional subset of the group of individuals before modifying the integrated data repository to store at least a portion of the additional information of the medical records of the portion of the subset of the group of individuals in association with the plurality of identifiers.
[0204]
[0023] Embodiment 49. The one or more non-transitory computer readable media of any one of embodiments 34-48, wherein the molecular data repository stores at least one or more of genomic information, genetic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, or proteomic information.
[0205] Aspect 50. Generating, by a computing system including a processing circuit and a memory, a data file including first tokens generated using a first hash function, each first token corresponding to a respective individual of a group of individuals having data stored by a molecular data repository, sending, by the computing system, the data file to a medical record data management system, retrieving, by the computing system, medical record data corresponding to the group of individuals from the medical record data management system responsive to the data file, generating, by the computing system, a plurality of identifiers using a second hash function different from the first hash function, each identifier corresponding to one or more tokens relating to a respective individual of the group of individuals, and retrieving, by the computing system, the second data for the group of individuals from the molecular data repository using the plurality of identifiers. determining, by the computing system, respective portions of the first data corresponding to respective portions of the second data for a group of individuals; generating, by the computing system, an integrated data repository storing respective portions of the first data and respective portions of the second data in relation to respective identifiers of a plurality of identifiers; receiving, by the computing system, a request to determine data for a plurality of individuals having data stored in the integrated data repository, the request including one or more search criteria; determining, by the computing system, a subset of the plurality of individuals having one or more characteristics corresponding to the one or more search criteria; and analyzing, by the computing system, information of the subset of the plurality of individuals to determine an indication of significance of the one or more characteristics with respect to a biological condition.
[0206]
[0023] Embodiment 51. The method of embodiment 50, comprising: determining, by a computing system, one or more genomic mutations present in a subset of the plurality of individuals;
[0207] A method comprising: determining, by a computing system, a plurality of treatments provided to a subset of a plurality of individuals; and determining, by the computing system, a survival rate for each of the subset of the plurality of individuals.
[0208] Embodiment 52. The method of embodiment 51, wherein the indication of significance corresponds to a survival rate for a treatment of the plurality of treatments and a genomic mutation of the one or more genomic mutations.
[0209]
[0023] Example 53. The method of example 52, comprising determining, by a computing system, efficacy of a treatment for a subset of the plurality of individuals based on the indication of significance.
[0210]
[0023] Example 54. The method of example 53, comprising determining, by a computing system, individuals among the subset of the plurality of individuals that have not received the treatment.
[0211] Embodiment 55. The method of embodiment 54, comprising administering a therapeutically effective amount of one or more treatments to individuals in the subset of the plurality of individuals who have not received the treatment.
[0212]
[0023] Embodiment 56. The method of any one of embodiments 50-55, wherein the integrated data repository is configured according to a data repository schema including a plurality of data tables and a plurality of logical links between the plurality of data tables, each logical link of the plurality of logical links indicating one or more rows of one of the plurality of data tables that correspond to one or more further rows of the further data table of the plurality of data tables.
[0213] Embodiment 57. The method of embodiment 56, wherein the plurality of data tables includes a first data table storing genomics data for the group of individuals, a second data table storing data relating to one or more visits by the individuals to one or more health care providers, a third data table storing information corresponding to each service provided to the individuals in relation to the one or more visits to the one or more health care providers indicated by the second data table, a fourth data table storing personal information of the group of individuals, a fifth data table storing information relating to health insurance companies or government agencies that paid for the services provided to the group of individuals, a sixth data table storing information corresponding to health insurance coverage information for the group of individuals, and a seventh data table storing information relating to pharmaceutical treatments obtained by the group of individuals.
[0214]
[0036] Aspect 58. The method of any one of aspects 50-57, wherein the plurality of identifiers generated using the second hash function include intermediate identifiers, and the method includes applying, by the computing system, a salt function to the intermediate identifiers to generate a final set of identifiers.
[0215]
[0023] Embodiment 59. The method of any one of embodiments 50-58, comprising: obtaining, by a computing system, additional information from an additional data repository containing health insurance claims data for an additional group of individuals; determining, by the computing system, at least a subset of the additional group of individuals that corresponds to the group of individuals having data stored by the genomics data repository; and modifying, by the computing system, the integrated data repository to store at least a portion of the additional information of the health insurance claims data for the at least the subset of the additional group of individuals in association with the plurality of identifiers.
[0216] Example 60. The method of any one of examples 50-59, including performing, by a computing system, one or more optical character recognition operations on the medical record data; and analyzing, by the computing system, the medical record data to determine one or more portions of the medical record data to be removed to generate the corpus of information.
[0217] Example 61. The method of example 60, comprising: analyzing, by a computing system, the corpus of information to determine a portion of the subset of the group of individuals that corresponds to the one or more biomarkers; and generating, by the computing system, one or more data structures that store identifiers of the portion of the subset of the group of individuals and that store an indication that the portion of the subset of the group of individuals corresponds to the one or more biomarkers.
[0218] Aspect 62. The method of aspect 61, comprising storing, by a computing system, the one or more data structures in an intermediate data repository, and performing, by the computing system, one or more de-identification operations with respect to the identifiers of the portion of the subset of the group of individuals before modifying the integrated data repository to store at least a portion of the medical record data of the portion of the subset of the group of individuals in association with a plurality of identifiers.
[0219] Embodiment 63. The method of any one of embodiments 50-62, wherein the molecular data repository stores at least one or more of genomic information, genetic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, or proteomic information.
[0220] Aspect 64. A system, comprising: one or more hardware processing devices; and one or more computer-readable storage media having computer-executable instructions stored thereon, the computer-executable instructions, when executed by the one or more hardware processing devices, cause the system to generate a data file including first tokens generated using a first hash function, each first token corresponding to a respective individual of a group of individuals having data stored by a molecular data repository; sending the data file to a medical record data management system; retrieving medical record data corresponding to the group of individuals from the medical record data management system responsive to the data file; and generating a plurality of identifiers using a second hash function different from the first hash function, each identifier corresponding to one or more tokens relating to a respective individual of the group of individuals. receiving a request to determine data for a plurality of individuals having data stored in the integrated data repository, the request including one or more search criteria; determining a subset of the plurality of individuals having one or more characteristics corresponding to the one or more search criteria; and analyzing information of the subset of the plurality of individuals to determine an indication of significance of the one or more characteristics with respect to a biological condition.
[0221]
[0023] Embodiment 65. The system of embodiment 64, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining one or more genomic mutations present in a subset of the plurality of individuals, determining a plurality of treatments provided to the subset of the plurality of individuals, and determining a survival rate for each of the subset of the plurality of individuals.
[0222]
[0023] Embodiment 66. The system of embodiment 65, wherein the indication of significance corresponds to a survival rate for a treatment of the plurality of treatments and a genomic mutation of the one or more genomic mutations.
[0223]
[0023] Embodiment 67. The system of embodiment 66, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining efficacy of treatment for a subset of the plurality of individuals based on the indication of significance.
[0224]
[0023] Embodiment 68. The system of embodiment 67, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining individuals among the subset of the plurality of individuals that have not received the treatment.
[0225] Aspect 69. The system of any one of aspects 64-68, wherein the integrated data repository is configured according to a data repository schema including a plurality of data tables and a plurality of logical links between the plurality of data tables, each logical link of the plurality of logical links indicating one or more rows of one data table of the plurality of data tables that correspond to one or more further rows of the further data table of the plurality of data tables.
[0226]
[0023] Embodiment 70. The system of embodiment 69, wherein the plurality of data tables includes a first data table storing genomics data for a group of individuals, a second data table storing data relating to one or more visits by the individuals to one or more health care providers, a third data table storing information corresponding to each service provided to the individuals for the one or more visits to the one or more health care providers indicated by the second data table, a fourth data table storing personal information of the group of individuals, a fifth data table storing information relating to health insurance companies or government agencies that paid for services provided to the group of individuals, a sixth data table storing information corresponding to health insurance coverage information for the group of individuals, and a seventh data table storing information relating to pharmaceutical treatments obtained by the group of individuals.
[0227]
[0023] Example 71. The system of any one of Examples 64-70, wherein the plurality of identifiers generated using the second hash function includes intermediate identifiers, and wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including applying a salt function to the intermediate identifiers to generate a final set of identifiers.
[0228]
[0023] Embodiment 72. The system of any one of embodiments 64-71, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including obtaining additional information from an additional data repository containing health insurance claims data for an additional group of individuals, determining at least a subset of the additional group of individuals that corresponds to the group of individuals having data stored by the genomics data repository, and modifying the integrated data repository to store at least a portion of the additional information of the health insurance claims data for the at least a subset of the additional group of individuals in association with the plurality of identifiers.
[0229] Embodiment 73. The system of any one of embodiments 64-72, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including performing one or more optical character recognition operations on the medical record data and analyzing the medical record data to determine one or more portions of the medical record data to be removed to generate the corpus of information.
[0230] Embodiment 74. The system of embodiment 73, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including: analyzing a corpus of information to determine a portion of the subset of the group of individuals that corresponds to the one or more biomarkers; and generating one or more data structures that store an identifier of the portion of the subset of the group of individuals and that store an indication that the portion of the subset of the group of individuals corresponds to the one or more biomarkers.
[0231] Aspect 75. The system of aspect 74, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including storing one or more data structures in the intermediate data repository, and performing, by the computing system, one or more de-identification operations with respect to identifiers of the portion of the subset of the group of individuals prior to modifying the integrated data repository to store at least a portion of the medical record data of the portion of the subset of the group of individuals in association with a plurality of identifiers.
[0232]
[0023] Embodiment 76. The system of any one of embodiments 64-75, wherein the molecular data repository stores at least one or more of genomic information, genetic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, or proteomic information.
[0233] Aspect 77. One or more non-transitory computer-readable storage media having stored thereon computer-executable instructions that, when executed by one or more hardware processing units, cause a system to generate a data file including first tokens generated using a first hash function, each first token corresponding to a respective individual of a group of individuals having data stored by a molecular data repository; sending the data file to a medical record data management system; retrieving medical record data corresponding to the group of individuals from the medical record data management system responsive to the data file; generating a plurality of identifiers using a second hash function different from the first hash function, each identifier corresponding to one or more tokens relating to a respective individual of the group of individuals; and using the plurality of identifiers to retrieve medical record data corresponding to the respective individual of the group of individuals. one or more non-transitory computer-readable storage media configured to cause a computer to perform operations including: obtaining second data from a molecular data repository for the group; determining respective portions of the first data for the group of individuals that correspond to respective portions of the second data; generating an integrated data repository that stores respective portions of the first data and respective portions of the second data in relation to respective identifiers of a plurality of identifiers; receiving a request to determine data for a plurality of individuals having data stored in the integrated data repository, the request including one or more search criteria; determining a subset of the plurality of individuals having one or more characteristics that correspond to the one or more search criteria; and analyzing information of the subset of the plurality of individuals to determine an indication of significance of the one or more characteristics with respect to a biological condition.
[0234]
[0023] Embodiment 78. The one or more non-transitory computer readable media of embodiment 77, wherein the one or more computer readable storage media store additional computer executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining one or more genomic mutations present in a subset of the plurality of individuals, determining a plurality of treatments provided to the subset of the plurality of individuals, and determining a survival rate for each of the subset of the plurality of individuals.
[0235]
[0023] Embodiment 79. The one or more non-transitory computer readable media of embodiment 78, wherein the indication of significance corresponds to a survival rate for a treatment of the plurality of treatments and a genomic mutation of the one or more genomic mutations.
[0236]
[0023] Aspect 80. The one or more non-transitory computer readable media of aspect 79 comprising additional computer executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including determining efficacy of treatment for a subset of the plurality of individuals based on the indication of significance.
[0237]
[0023] Aspect 81. The one or more non-transitory computer readable media of aspect 80 comprising additional computer executable instructions that, when executed by one or more hardware processing devices, cause the system to perform additional operations including determining individuals among the subset of the plurality of individuals that have not received the treatment.
[0238]
[0023] Embodiment 82. The one or more non-transitory computer readable media of any one of embodiments 77-81, wherein the integrated data repository is configured according to a data repository schema including a plurality of data tables and a plurality of logical links between the plurality of data tables, each logical link of the plurality of logical links indicating one or more rows of a data table of the plurality of data tables that correspond to one or more further rows of the further data table of the plurality of data tables.
[0239]
[0023] Embodiment 83. The one or more non-transitory computer readable media of embodiment 82, wherein the plurality of data tables includes a first data table storing genomics data for a group of individuals, a second data table storing data relating to one or more visits by the individuals to one or more healthcare providers, a third data table storing information corresponding to respective services provided to the individuals for the one or more visits to the one or more healthcare providers indicated by the second data table, a fourth data table storing personal information of the group of individuals, a fifth data table storing information relating to health insurance companies or government agencies that paid for services provided to the group of individuals, a sixth data table storing information corresponding to health insurance coverage information for the group of individuals, and a seventh data table storing information relating to pharmaceutical treatments obtained by the group of individuals.
[0240]
[0036] Example 84. The one or more non-transitory computer-readable media of any one of Examples 77-83, wherein the plurality of identifiers generated using the second hash function includes intermediate identifiers, the one or more non-transitory computer-readable media comprising additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including applying a salt function to the intermediate identifiers to generate a final set of identifiers.
[0241]
[0023] Embodiment 85. The one or more non-transitory computer readable media of any one of embodiments 77-84 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including obtaining additional information from an additional data repository containing health insurance claims data for an additional group of individuals, determining at least a subset of the additional group of individuals that corresponds to the group of individuals having data stored by the genomics data repository, and modifying the integrated data repository to store at least a portion of the additional information of the health insurance claims data for the at least a subset of the additional group of individuals in association with the plurality of identifiers.
[0242]
[0023] Embodiment 86. The one or more non-transitory computer-readable media of any one of embodiments 77-85 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including performing one or more optical character recognition operations on the medical record data and analyzing the medical record data to determine one or more portions of the medical record data to be removed to generate a corpus of information.
[0243] Embodiment 87. The one or more non-transitory computer readable media of embodiment 86 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including: analyzing the corpus of information to determine a portion of the subset of the group of individuals that corresponds to the one or more biomarkers; and generating one or more data structures that store an identifier of the portion of the subset of the group of individuals and that store an indication that the portion of the subset of the group of individuals corresponds to the one or more biomarkers.
[0244] Aspect 88. The one or more non-transitory computer readable media of aspect 87 comprising additional computer-executable instructions that, when executed by the one or more hardware processing devices, cause the system to perform additional operations including storing the one or more data structures in the intermediate data repository and performing, by the computing system, one or more de-identification operations with respect to the identifiers of the portion of the subset of the group of individuals before modifying the integrated data repository to store at least a portion of the medical record data of the portion of the subset of the group of individuals in association with a plurality of identifiers.
[0245]
[0023] Embodiment 89. The one or more non-transitory computer readable media of any one of embodiments 77-88, wherein the molecular data repository stores at least one or more of genomic information, genetic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, or proteomic information. EXAMPLES
[0246] Example 1
[0247] Liquid biopsies offer a minimally invasive alternative to tissue biopsies for comprehensive genomic profiling (CGP) and also contain additional information in the form of circulating tumor DNA (ctDNA) levels. Qualitative and quantitative ctDNA levels have been shown to indicate tumor volume. Little is known about how ctDNA levels estimated from a single blood draw correlate with outcome in patients with late-stage metastatic non-small cell lung cancer (NSCLC) undergoing different treatment regimens.
[0248] Patients with NSCLC were identified via an integrated database and grouped according to whether they underwent liquid biopsy testing within 190 days before initiation of metastatic first-line (1L) therapy ("pre-1L"), within 90 days after initiation of 1L ("early 1L"), or between 90 and 190 days after initiation of 1L ("late 1L"). Kaplan-Meier and Cox proportional hazards modeling (CPH) were used to assess differences in real-world overall survival (rwOS). Sex and age were included as covariates in CPH. ctDNA levels were defined as the highest mutant allele fraction when used as a quantitative indicator, and a threshold of 4% was used to define ctDNA high / low groups when used as a categorical variable in NSCLC.
[0249] Patients with higher levels of ctDNA had worse rwOS, regardless of therapy or timing of blood draw relative to start of 1L treatment, but comparison of the 90-day post-1L osimertinib group with the chemotherapy group did not cross the significance cutoff (<0.05), likely due to the small number of patients in these groups. Patients with no detectable tumor-derived alterations had the longest rwOS and the lowest hazard ratio for high ctDNA (range: 0.16-0.46).
[0250] Figure 10 shows Kaplan-Meier curves showing real-world overall survival values for patients who received 1L therapy to treat non-small cell lung cancer prior to receiving treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0251] Figure 11 shows Kaplan-Meier curves showing real-world overall survival values for patients receiving 1L therapy to treat non-small cell lung cancer during treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0252] Figure 12 shows Kaplan-Meier curves showing real-world overall survival values for patients who received osimertinib to treat non-small cell lung cancer prior to treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0253] FIG. 13 shows Kaplan-Meier curves illustrating real-world overall survival values for patients receiving osimertinib to treat non-small cell lung cancer during treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0254] FIG. 14 shows Kaplan-Meier curves illustrating real-world overall survival values for patients receiving chemotherapy to treat non-small cell lung cancer during treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0255] FIG. 15 shows Kaplan-Meier curves illustrating real-world overall survival values for patients receiving chemotherapy to treat non-small cell lung cancer after treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0256] FIG. 16 shows Kaplan-Meier curves illustrating real-world overall survival values for patients who received chemotherapy to treat non-small cell lung cancer prior to receiving treatment for high ctDNA counts, low ctDNA counts, and no detectable ctDNA.
[0257] In addition to providing a less invasive alternative to tissue biopsy for CPGs, the highest mutant allele fraction reported in liquid biopsy testing, especially the absence of detected ctDNA, may be useful for providing prognostic information about patients and identifying high-risk patients who would benefit from more aggressive treatment regimens. [Table 1] [Table 2-1] [Table 2-2] [Table 3]
[0258] Example 2
[0259] Results from CLIA-certified, CAP-accredited, NYSDOH-approved circulating tumor DNA (ctDNA) testing required for patients with advanced solid tumors, performed on approximately 103,000 patients, were anonymized and tokenized using irreversible one-way hashing. Using secure, HIPAA-compliant, and approved methods, these results were linked to a de-identified patient episode encounter database housing medical and pharmacy claims to obtain a longitudinal view of patient history, including diagnosis, treatment, and real-world time-to-event data points in the integrated database. These de-identified, integrated data can then be used to explore disease-, biomarker-, and therapy-specific models of tumor progression, as well as drug resistance.
[0260] Figure 17 shows the frequency of selected alterations in a cohort of patients (n=637) diagnosed with advanced non-small cell lung cancer (NSCLC) who underwent liquid biopsy testing after initiation of first-line osimertinib therapy. Approximately 12% showed secondary EGFR mutations. Approximately 12% showed gene amplifications, namely HER2 and MET. Approximately 10% showed mutations in MAPK / PIK3CA genes. And approximately 17% showed alterations in cell cycle genes. These results were directionally consistent with those published in Ramalingam et al. The mean duration of treatment for these patients was approximately 8 months, which is consistent with published osimertinib studies.
[0261] FIG. 18 shows the frequency of selected mutations in the ligand-binding domain of a cohort (n=4448) of patients diagnosed with breast cancer who underwent liquid biopsy testing after documented treatment with aromatase inhibitors (AIs). In the case of metastatic breast cancer, we investigated data from 4,448 patients with a diagnosis of metastatic breast cancer who were prescribed aromatase inhibitors and subsequently underwent liquid biopsy testing. Mutations occurring in the ligand-binding domain of ESR1 are commonly observed resistance mechanisms associated with progression to aromatase inhibitors, and data suggest that these mutations are highly heterogeneous. Thus, we observed such heterogeneity in the liquid biopsy test results, with the D538G and Y537S mutations being the most frequently observed, as one might expect based on Toy et al.
[0262] To further clarify the utility of the database in investigating genomic alterations in the context of treatment, two patient cases were investigated. In the first case, as shown in FIG. 19, a patient is seen to exhibit a T790M mutation detected by liquid biopsy testing, treated with osimertinib, and subsequently develop a secondary C797S mutation as well as MET amplification. In the second case, as shown in FIG. 20, a patient with a record of treatment with aromatase inhibitors letrozole and exemestane is seen to subsequently develop a D538G mutation in the ESR1 gene. FIG. 19 shows changes associated with osimertinib resistance detected by liquid biopsy testing after treatment provided to a woman diagnosed with NSCLC. FIG. 20 shows ESR1 resistance mutations detected after a second course of treatment for a woman diagnosed with metastatic breast cancer and treated with aromatase inhibitors.
[0263] The integrated database contains integrated, de-identified clinical and genomic information from over 103,000 patients with advanced cancer, making it one of the largest databases of its kind. It will continue to grow and mature with the continued use of liquid biopsy testing and due to its unique and comprehensive incorporation of integrated clinical data (avoiding loss of follow-up and patient mobility).
[0264] The integrated database can be used to identify and study clinical outcomes based on genomic tumor characteristics using liquid biopsy test data. Unique agents and classes of therapy (TKIs, CDK4 / 6is, etc.) can be reliably identified, placed into relevant cohorts, and studied. This unique resource can investigate biological mechanisms of drug response and resistance relevant to the treatment of advanced cancer in real-world settings. Researchers can, among other uses, identify and characterize unmet medical needs and accelerate the development of new therapies through trial design optimization, outcome surveillance in the post-marketing setting, and identify promising novel combinations and treatment strategies (sequencing). Further directions include further validation of the data and the addition of ancillary source data to support deeper analysis.
[0265] It should be understood that the individual steps used in the methods of the present teachings may be performed in any order and / or simultaneously so long as the teachings remain functional. Further, it should be understood that the apparatus and methods of the present teachings may include any number or all of the described implementations so long as the teachings remain functional.
[0266] Various implementations of systems, devices, and methods have been described herein. These implementations are provided merely as examples and do not limit the scope of the claimed invention. Furthermore, it should be recognized that various features of the described implementations may be combined in various ways to create numerous additional implementations. Furthermore, while various materials, dimensions, shapes, configurations, and positions, etc., for use in the disclosed implementations have been described, other than those disclosed may be utilized without departing from the scope of the claimed invention.
[0267] Those skilled in the art will recognize that implementations may comprise fewer features than shown in any individual implementation described above. The implementations described herein are not intended to be an exhaustive presentation of the ways in which various features may be combined. Thus, implementations are not mutually exclusive combinations of features; rather, as will be understood by those skilled in the art, implementations may comprise combinations of different individual features selected from different individual implementations. Furthermore, elements described with respect to one implementation may be implemented in other implementations even if not described in such implementation, unless otherwise specified. Although a dependent claim may refer to a specific combination with one or more other claims in the claims, other implementations may include combinations of that dependent claim with the subject matter of each other dependent claim, or combinations of one or more features with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a particular combination is not intended. Furthermore, it is also intended that a feature of a claim be included in an independent claim, even if that claim is not directly dependent on any other independent claim.
[0268] Moreover, references in the specification to "one implementation," "one implementation," or "some implementations" mean that a particular feature, structure, or characteristic described in connection with that implementation is included in at least one implementation of the present teachings. The appearances of the phrase "in one implementation" in various places in the specification do not necessarily all refer to the same implementation.
[0269] Any incorporation of documents by reference above is limited so that it does not incorporate subject matter that is inconsistent with the explicit disclosure herein. Any incorporation of documents by reference above is further limited so that any claims contained therein are not incorporated by reference herein. Any incorporation of documents by reference above is further limited so that any definitions provided therein are not incorporated by reference herein, unless expressly included.
[0270] Although the implementations have been described with reference to certain exemplary implementations, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the present disclosure. Thus, the specification and drawings should be regarded in an illustrative and not restrictive sense. The accompanying drawings, which form a part hereof, show, by way of example and not of limitation, specific implementations in which the subject matter may be practiced. The illustrated implementations are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other implementations may be utilized and derived therefrom, and structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. Thus, this detailed description should not be construed in a limiting sense, and the scope of the various implementations is defined solely by the appended claims, along with the full range of equivalents to which such claims are entitled.
[0271] Although specific implementations have been illustrated and described herein, it should be appreciated that any configuration contemplated to achieve the same purpose may be used in place of the specific implementations shown. The present disclosure is intended to encompass all modifications and variations of the various implementations. Combinations of the above implementations and other implementations not specifically described herein will be apparent to those of skill in the art upon reviewing the above description.
[0272] In this document, the terms "a" or "an" are used to include one or more, as is common in patent documents, regardless of any other instance or use of "at least one" or "one or more." In this document, the term "or" is used to refer to a non-exclusive "or," and "A or B" includes "including A but not B," "including B but not A," and "A and B," unless otherwise noted. In this document, the terms "including" and "in which" are used as the plain English equivalents of the terms "comprising" and "wherein," respectively. Also, in the following claims, the terms "including" and "comprising" are open-ended, i.e., a system, user equipment (UE), article, composition, formulation, or process that includes elements in addition to those recited after such terms in a claim are also deemed to be within the scope of that claim. Moreover, in the following claims, the terms "first," "second," and "third," etc. are used merely as labels and do not impose numerical requirements on their objects.
Claims
1. Generating, by a computing system including a processing circuit and a memory, a data file including a first token generated using a first hash function, wherein each individual first token corresponds to each individual in a group of individuals having data stored by a molecular data repository, and sending, by the computing system, the data file to a health insurance claim data management system; obtaining, by the computing system, health data corresponding to the group of individuals from the health insurance claim data management system in response to the data file; generating, by the computing system, a plurality of identifiers using a second hash function different from the first hash function, wherein each identifier corresponds to one or more tokens associated with each individual in the group of individuals; obtaining, by the computing system, second data from the molecular data repository for the group of individuals using the plurality of identifiers; determining, by the computing system, respective portions of first data corresponding to respective portions of the second data for the group of individuals; generating, by the computing system, an integrated data repository storing the respective portions of the first data and the respective portions of the second data in relation to each identifier of the plurality of identifiers. A method comprising.
2. determining, by the computing system, a first set of executable data processing instructions in relation to first data stored by the integrated data repository; executing, by the computing system, the first set of data processing instructions to analyze a first health insurance claim code included in the first data to determine a first subset of the group of individuals in which a certain biological state exists; generating, by the computing system, a first data set indicating the subset of the group of individuals in which the biological state exists The method according to claim 1, comprising.
3. Determining, by the computing system, a second set of data processing instructions executable in relation to second data stored by the integrated data repository; Causing, by the computing system, the second set of data processing instructions to be executed to analyze second health insurance claim codes included in the second data to determine one or more treatments provided to a second subset of the group of individuals; Generating, by the computing system, a second data set indicative of the one or more treatments provided to the second subset of the group of individuals The method according to claim 2, comprising: **Claim 4** Determining, by the computing system, a third subset of the group of individuals, the third subset including a portion of the first subset of the group of individuals that overlaps with a portion of the second subset of the group of individuals; Receiving, by the computing system, a request to analyze the first data set and the second data set in relation to the third subset of the group of individuals; In response to the request, analyzing, by the computing system, the first data set and the second data set with respect to the third subset of the group of individuals to determine an indicator of significance of characteristics of the third subset of the group of individuals regarding the biological state The method according to claim 3, comprising: **Claim 5** Determining, by the computing system, one or more genomic mutations present in the third subset of the group of individuals; Determining, by the computing system, a plurality of treatments provided to the third subset of the group of individuals; Determining, by the computing system, the respective survival rates of the third subset of the group of individuals The method according to claim 4, comprising: **Claim 6** The method according to claim 5, wherein the indicator of significance corresponds to a survival rate for one of the plurality of treatments and one of the one or more genomic mutations. **Claim 7** The method according to claim 6, comprising determining, by the computing system, the effectiveness of the treatment for the third subset of the group of individuals based on an index of significance.
8. The method according to claim 7, comprising determining, by the computing system, individuals within a third subset of the group of individuals who have not received the treatment.
9. The method according to claim 8, wherein the individuals within the third subset who have not received the treatment are indicated to receive the treatment.
10. The integrated data repository is configured according to a data repository schema including a plurality of data tables and a plurality of logical links between the plurality of data tables, wherein each logical link among the plurality of logical links indicates one or more rows of one of the plurality of data tables corresponding to one or more further rows of a further data table among the plurality of data tables. The method according to any one of claims 1 to 9.
11. The plurality of data tables include a first data table storing genomic data of the group of individuals, a second data storing data related to one or more medical provider visits by an individual one or more times, a third data table storing information corresponding to each service provided to an individual regarding one or more medical provider visits indicated by the second data table, a fourth data table storing personal information of the group of individuals, a fifth data table storing information related to a health insurance company or government agency that has paid for services provided to the group of individuals, a sixth data table storing information corresponding to health insurance coverage information of the group of individuals, a seventh data table storing information related to pharmaceutical treatments received by the group of individuals. The method according to claim 10.
12. The plurality of identifiers generated using the second hash function include intermediate identifiers, and the method includes applying, by the computing system, a sort function to the intermediate identifiers to generate a set of final identifiers. The method according to any one of claims 1 to 9.
13. obtaining information from an additional data repository that includes electronic medical records of an additional group of individuals by the computing system; determining, by the computing system, a subset of the additional group of individuals corresponding to the group of individuals having data stored by the genomics data repository; modifying, by the computing system, the integrated data repository to store at least a portion of the information of the medical records of the subset of the additional group of individuals in relation to the plurality of identifiers; The method according to any one of claims 1 to 9, comprising:
14. performing, by the computing system, one or more optical character recognition operations on the additional information; analyzing, by the computing system, the additional information obtained from the additional data repository to determine one or more portions of the additional information to be removed to generate a corpus of information; The method according to claim 13, comprising:
15. analyzing, by the computing system, the corpus of information to determine a portion of the subset of the additional group of individuals corresponding to one or more biomarkers; generating, by the computing system, one or more data structures that store identifiers of the portion of the subset of the additional group of individuals and store an indication that the portion of the subset of the additional group of individuals corresponds to the one or more biomarkers; The method according to claim 14, comprising:
16. storing, by the computing system, the one or more data structures in an intermediate data repository; performing, by the computing system, one or more anonymization operations on the identifiers of the portion of the subset of the additional group of individuals before modifying the integrated data repository to store at least a portion of the additional information of the medical records of the portion of the subset of the additional group of individuals in relation to the plurality of identifiers; The method according to claim 15, comprising:
17. The method according to any one of claims 1 to 9, wherein the molecular data repository stores at least one or a plurality of genomic information, genetic information, metabolomics information, transcriptomics information, fragmentomics information, immune receptor information, methylation information, epigenomics information, or proteomics information.
18. A system comprising: one or more hardware processing devices; and one or more computer-readable storage media storing computer-executable instructions, which, when executed by the one or more hardware processing devices, cause the system to: generate a data file containing first tokens generated using a first hash function, wherein each individual first token corresponds to each individual in a group of individuals having data stored by a molecular data repository; send the data file to a health insurance claim data management system; obtain health insurance claim data corresponding to the group of individuals from the health insurance claim data management system in response to the data file; generate a plurality of identifiers using a second hash function different from the first hash function, wherein each identifier corresponds to one or more tokens associated with each individual in the group of individuals; obtain second data from the molecular data repository for the group of individuals using the plurality of identifiers; determine respective portions of first data corresponding to respective portions of the second data for the group of individuals; and generate an integrated data repository that stores the respective portions of the first data and the respective portions of the second data in relation to each identifier of the plurality of identifiers. A system that performs operations including these.
19. A computing system including a processing circuit and a memory generates a data file containing first tokens generated using a first hash function, wherein each individual first token corresponds to each individual in a group of individuals having data stored by a molecular data repository; The computing system sends the data file to a medical record data management system. The computing system obtains medical record data corresponding to the group of individuals from the medical record data management system in response to the data file. The computing system generates a plurality of identifiers by using a second hash function different from the first hash function, wherein each identifier corresponds to one or more tokens related to each individual in the group of individuals. The computing system uses the plurality of identifiers to obtain second data from the molecular data repository for the group of individuals. The computing system determines respective portions of first data corresponding to respective portions of the second data for the group of individuals. The computing system generates an integrated data repository that stores the respective portions of the first data and the respective portions of the second data in relation to each of the plurality of identifiers. The computing system receives a request to determine data regarding a plurality of individuals having data stored in the integrated data repository, the request including one or more search criteria. The computing system determines a subset of the plurality of individuals having one or more characteristics corresponding to the one or more search criteria. The computing system analyzes information of the subset of the plurality of individuals to determine an indicator of the significance of a characteristic among the one or more characteristics regarding a biological state. A method including this.
20. The computing system determines one or more genomic mutations present in the subset of the plurality of individuals. The computing system determines a plurality of treatments provided to the subset of the plurality of individuals. The computing system determines the respective survival rates of the subset of the plurality of individuals The method according to claim 19, including this.
21. The method according to claim 20, wherein the significance index corresponds to the survival rate for one of the plurality of treatments and one of the one or more genomic mutations. **Claim 22** The method according to claim 21, comprising determining, by the computing system, the effectiveness of the treatment for the subset of the plurality of individuals based on a significance index. **Claim 23** The method according to claim 22, comprising determining, by the computing system, individuals among the subset of the plurality of individuals who have not received the treatment. **Claim 24** The method according to claim 23, wherein the individuals among the subset of the plurality of individuals who have not received the treatment are shown to receive the treatment.