COMPUTER ARCHITECTURE FOR IDENTIFYING A LINE OF THERAPY - Patent application

JP2024528088A5Pending Publication Date: 2025-08-05GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024505385
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-03
Filing Date
2022-07-29
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing systems inefficiently and inaccurately analyze unstructured healthcare data, such as pharmacy and medical procedural data, which is crucial for understanding patient therapies and health outcomes, leading to potential detrimental effects on individual health.

Method used

A computing architecture that integrates structured health insurance claims data with de-identified genomics data to accurately analyze healthcare data, identifying therapy lines by mapping pharmacy and medical procedures onto a timeline, using data integration and analysis systems with data pipelines and machine learning techniques to enhance data processing efficiency and accuracy.

Benefits of technology

Enhances the analysis of healthcare data to accurately determine therapy lines, improving the understanding of patient treatments and health outcomes, reducing computational resources, and ensuring privacy through de-identification and regulatory compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computing machine accesses a pharmacy procedure dataset and a medical procedure procedure dataset. The computing machine filters procedures associated with a biological condition from the pharmacy dataset and the medical procedure procedure dataset. The computing machine identifies one or more lines of therapy from the filtered pharmacy dataset and the filtered medical procedure procedure dataset. One aspect of the invention is a method implemented in one or more computing machines comprising processing circuitry and memory.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This patent application claims the benefit of priority to U.S. Provisional Application No. 63 / 227,860, filed July 30, 2021, U.S. Provisional Application No. 63 / 238,851, filed August 31, 2021, U.S. Provisional Application No. 63 / 250,912, filed September 30, 2021, and PCT Application No. PCT / US2022 / 032250, filed June 3, 2022, each of which is incorporated by reference in its entirety herein.

[0002] TECHNICAL FIELD Implementations of the present disclosure relate generally to the field of computer architecture, and more specifically, to computer architectures for identifying lines of therapy from pharmacy and medical procedure data. [Background technology]

[0003] When an individual visits a health care provider to treat one or more biological conditions, various types of documentation may be generated. For example, pharmacy transaction data may be generated when an individual visits a health care provider and bills their health insurance carrier. Medical procedure transaction data may be generated when an individual visits a health care provider and / or undergoes a medical procedure and bills their health insurance carrier. Pharmacy transaction data and medical procedure data may be useful in understanding the therapy a patient receives. Techniques for identifying the therapy from the pharmacy transaction data and medical procedure transaction data may be desirable. Summary of the Invention [Means for solving the problem]

[0004] Detailed Description The following description and drawings sufficiently illustrate specific implementations to enable those skilled in the art to practice them. Other implementations may incorporate structural, logical, electrical, process, and other changes. Portions and features of some implementations may be included in or substituted for those of other implementations. Implementations set forth in the claims encompass all available equivalents of those claims. [Brief description of the drawings]

[0005] [Figure 1] FIG. 1 illustrates an example architecture for generating an integrated data repository containing multiple types of healthcare data, according to one or more implementations.

[0006] [Diagram 2] FIG. 2 illustrates an exemplary framework that corresponds to an arrangement of data tables in a unified data repository, according to one or more implementations.

[0007] [Diagram 3] FIG. 3 illustrates an architecture for generating one or more data sets from information retrieved from a data repository that integrates health-related data from a number of sources, according to one or more implementations.

[0008] [Figure 4] FIG. 4 illustrates an architecture for generating an integrated data repository including de-identified health insurance claims data and de-identified genomics data, according to one or more implementations.

[0009] [Diagram 5] FIG. 5 illustrates a framework for generating datasets by a data pipeline system based on data stored by a unified data repository, according to one or more implementations.

[0010] [Figure 6A] 6A-6B illustrate a flowchart of an example process associated with determining a line of therapy, according to one or more implementations. [Figure 6B] 6A-6B illustrate a flowchart of an example process associated with determining a line of therapy, according to one or more implementations.

[0011] [Figure 7A] 7A-7B are data flow diagrams of an example process for determining a line of therapy from pharmacy transaction data and medical procedure transaction data, according to one or more implementations. [Figure 7B] 7A-7B are data flow diagrams of an example process for determining a line of therapy from pharmacy transaction data and medical procedure transaction data, according to one or more implementations.

[0012] [Figure 7C] FIG. 7C illustrates an exemplary system for determining a line of therapy, according to one or more implementations.

[0013] [Figure 8] FIG. 8 illustrates a computing architecture having one or more systems for generating lines of therapy that can be analyzed to determine an outcome for a patient, according to one or more implementations.

[0014] [Figure 9] FIG. 9 illustrates a diagrammatic representation of a machine in the form of a computer system in which a set of instructions, in accordance with one or more implementations, may be executed to cause the machine to perform any one or more of the methodologies discussed herein.

[0015] [Figure 10] FIG. 10 illustrates a table showing the prevalence of MSI-H in approximately 30,000 advanced gastrointestinal cancer patients.

[0016] [Figure 11] FIG. 11 illustrates graphs and tables showing MSI-H scores, maximum variant allele fractions (VAFs), and intervention treatments for consecutive tests for MSI-H.

[0017] [Figure 12] FIG. 12 illustrates a Kaplan-Meier plot showing real-world time to death across treatment groups.

[0018] [Figure 13] FIG. 13 illustrates a Kaplan-Meier plot showing real-world time to next treatment across treatment groups.

[0019] [Figure 14] FIG. 14 illustrates a Kaplan-Meier plot showing overall survival probability across treatment groups.

[0020] [Figure 15] FIG. 15 illustrates a table showing the percentage of patients with one or more PIK3CA-mt mutations and indicating whether the mutations were clonal, subclonal, or both.

[0021] [Figure 16] FIG. 16 illustrates a Kaplan-Meier plot showing real-world time to next treatment for groups with clonal or subclonal PIK3CA-mt mutations.

[0022] [Figure 17] FIG. 17 illustrates a Kaplan-Meier plot showing real-world time to death for groups with clonal or subclonal PIK3CA-mt mutations.

[0023] [Figure 18]FIG. 18 illustrates graphs showing Cox proportional hazard ratios for time to next treatment and time to death for patients with clonal or subclonal PIK3CA-mt mutations.

[0024] [Figure 19] FIG. 19 illustrates a graph showing one or more additional genomic mutations present in patients with clonal or subclonal PIK3CA-mt mutations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] Analysis of health care data using existing systems and techniques is typically performed on medical records generated by health care providers. As used herein, a health care provider may refer to an entity, individual, or group of individuals involved in providing treatment to an individual in relation to at least one of the treatment or prevention of one or more biological conditions. In addition, as used herein, a biological condition may refer to a functional and / or structural abnormality in an individual to such an extent that it produces or is likely to produce a detectable characteristic of the abnormality. A biological condition may be characterized by external and / or internal characteristics, signs, and / or symptoms that indicate a deviation from a biological norm in one or more populations. A biological condition may be characterized by external and / or internal characteristics, signs, and / or symptoms that indicate a deviation from a biological norm in one or more populations. In various examples, a biological condition may include one or more molecular phenotypes. For example, a biological condition may correspond to a genetic or epigenetic pathology. In one or more additional examples, the biological condition may include at least one of one or more diseases, one or more disorders, one or more injuries, one or more syndromes, one or more disorders, one or more infections, one or more sporadic symptoms, or other atypical variations of the biological structure and / or function of the individual. In addition, treatment as used herein may refer to a substance, procedure, routine, device, and / or other intervention that may be administered or performed with the intent of treating one or more effects of a biological condition in an individual. In one or more examples, the treatment may include a substance that is metabolized by the individual. The substance may include a composition, such as a pharmaceutical composition. The substance may be delivered to the individual via a number of methods, such as ingestion, injection, absorption, or inhalation. The treatment may also include a physical intervention, such as one or more surgeries. In at least some examples, the treatment may include a therapeutically significant intervention.

[0026] Typically, the healthcare data analyzed by existing systems includes unstructured data. Unstructured data may include data that is not organized according to a predefined or standardized format. For example, unstructured data may include notes created by a healthcare provider that consist of free text. That is, the format in which the notes are captured does not include predefined inputs that are selectable by the healthcare provider, such as via a drop-down menu or via a list. Rather, the notes include text typed by the healthcare provider, which may include sentences, sentence fragments, words, characters, symbols, abbreviations, one or more combinations thereof, and the like. In some cases, the unstructured data may be partially structured. For example, a provider may select a billing code from a predefined list of billing codes and add unstructured notes to the data associated with that billing code.

[0027] Existing systems typically devote a large amount of computing resources to analyzing unstructured data to extract information that may be relevant to the analysis being performed by the existing system. In some cases, existing systems may analyze unstructured data and convert the unstructured data into a structured format to facilitate the analysis of the previously unstructured data. The analysis of unstructured data by existing systems may be inefficient and inaccurate. In a scenario where the unstructured data is obtained from healthcare data, the importance of accurately analyzing the information is high, since the analysis may be relevant to at least one of a number of individual treatments or diagnoses regarding one or more biological conditions. Thus, inaccurate analysis of healthcare data may have a detrimental effect on the health of an individual.

[0028] Implementations of the techniques, architectures, frameworks, systems, processes, and computer-readable instructions described herein are directed to analyzing health insurance claims data to derive information about at least one of an individual's health or treatment. In contrast to existing systems, health insurance claims data is structured according to one or more formats and stored by a number of data tables. The data tables may include codes or other alphanumeric information indicating the treatment the individual received, dates of treatment, dosage information, the individual's diagnosis of one or more biological conditions, information related to visits to a health care provider, dates of visits to a health care provider, billing information, and the like. Implementations described herein may be used to accurately analyze health insurance claims data for hundreds up to thousands of individuals for whom one or more biological conditions exist, one or more biological conditions are suspected to exist, and / or one or more biological conditions that may place the individual at risk. In various embodiments, tens of thousands, hundreds of thousands, or even millions of rows and / or columns of health insurance claims data may be analyzed to determine health-related information regarding individuals for whom one or more biological conditions exist.

[0029] In some implementations, an individual may transact with a pharmacy and a healthcare provider to treat a biological condition. A record of the transaction may be provided to the individual's health insurer (or another entity involved in handling or processing all or a portion of the payment for the transaction, e.g., a government agency). Pharmacy transaction data may be generated when an individual purchases at a pharmacy and bills their health insurer. Medical procedure transaction data may be generated when an individual sees a healthcare provider and / or undergoes a medical procedure and bills their health insurer. Pharmacy transaction data and medical procedure data may be useful in understanding the therapy a patient receives.

[0030] In some cases, a line of therapy may be identified from an individual's pharmacy transaction data and / or medical procedure transaction data. As used herein, the phrase "line of therapy" encompasses its plain and ordinary meaning. A line of therapy may include therapies (e.g., drugs, health care provider visits, medical therapies) administered for the same stage of a biological condition within a common time window. When a gap in therapy occurs, the line of therapy may be discontinued before the gap and a new line of therapy may be initiated after the gap. If the biological condition progresses or regresses (e.g., a change in stage of cancer), a new line of therapy may be initiated in response to the progression or regression of the biological condition. As used herein, a gap in therapy may include, among other things, a period of at least a predetermined length when there is no therapy being administered.

[0031] In some implementations, the computing machine may access, in the processing circuitry, from the memory, a pharmacy procedure dataset for a given patient. Each pharmacy procedure in the pharmacy procedure dataset may comprise at least a procedure date, a therapy type, and a therapy delivery duration. The computing machine may identify, from the pharmacy procedure dataset, a pharmacy procedure subset associated with the biological condition based on the therapy type. The computing machine may calculate an end date for at least one pharmacy procedure in the pharmacy procedure subset. The procedure date may correspond to a date when payment begins for the procedure. The end date may be determined based on the procedure date and a therapy delivery duration associated with the at least one pharmacy procedure. The computing machine may access, in the processing circuitry, from the memory, for the patient. Each medical procedure in the medical procedure dataset may comprise at least a medical procedure date range and a medical procedure type. The computing machine may identify, from the medical procedure procedure dataset, a medical procedure procedure subset associated with the biological condition based on the medical procedure type. The computing machine may adjust the medical procedure date range based on a period during which a medical procedure type is valid or repeated for at least one medical procedure in the medical procedure procedure subset. The computing machine may map the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in the memory. The timeline data structure may store the pharmacy procedures and the medical procedure procedures arranged by date. The computing machine may determine one or more therapy gaps in the timeline data structure during which there are no pharmacy procedures and there are no medical procedure procedures. Each therapy gap may comprise a number of consecutive days. In one or more embodiments, the number of consecutive days may be greater than a threshold number of days. The computing machine may determine one or more lines of therapy based on the one or more therapy gaps.Each line of therapy may comprise pharmacy procedures and medical procedure procedures that occur either between two therapy gaps, before the earliest temporary therapy gap, or after the latest temporary therapy gap. Each line of therapy is associated with a line date range. The computing machine may transmit to a data repository for storage therein a data structure that identifies the patient, the one or more lines of therapy, and the line date range for each one or more lines of therapy.

[0032] The memory may include (or be connected to) a pharmacy data repository that stores the pharmacy procedure dataset and a medical procedure procedure data repository that stores the medical procedure dataset. Each data repository may include a database or other data storage unit. The timeline data structure may include, among other things, multiple data items, each data item associated with a single date or date range.

[0033] FIG. 1 illustrates an example architecture 100 for generating a unified data repository containing multiple types of healthcare data, according to one or more implementations. The architecture 100 may include a data integration and analysis system 102. The data integration and analysis system 102 may obtain data from a number of data sources and consolidate the data from the data sources into a unified data repository 104. For example, the data integration and analysis system 102 may obtain data from a health insurance claims data repository 106. In various embodiments, the data integration and analysis system 102 and the health insurance claims data repository 106 may be created and maintained by different entities. In one or more additional embodiments, the data integration and analysis system 102 and the health insurance claims data repository 106 may be created and maintained by the same entity.

[0034] The data integration and analysis system 102 may be implemented by one or more computing devices. The one or more computing devices may include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or a combination thereof. In some implementations, at least a portion of the one or more computing devices may be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices may be implemented in a cloud computing architecture. In scenarios where the computing system used to implement the data integration and analysis system 102 is configured in a distributed computing architecture, processing operations may be performed in parallel by multiple virtual machines. In various embodiments, the data integration and analysis system 102 may implement multi-threading techniques. The implementation of distributed computing architectures and multi-threading techniques causes the data integration and analysis system 102 to utilize fewer computing resources relative to computing architectures that do not implement these techniques.

[0035] The health insurance claims data repository 106 may store information obtained from one or more health insurance companies corresponding to claims made by subscribers of the one or more health insurance companies. The health insurance claims data repository 106 may be arranged (e.g., sorted) by patient identifier. The patient identifier may be based on the patient's first name, last name, date of birth, social security number, address, employer, and the like. The data stored by the health insurance claims data repository 106 may include structured data arranged in one or more data tables. The one or more data tables storing the structured data may include a number of rows and a number of columns indicating information about health insurance claims made by subscribers of the one or more health insurance companies related to procedures and / or treatments received by the subscribers from a health care provider. At least some of the rows and columns of the data tables stored by the health insurance claims data repository 106 may include health insurance codes that may indicate diagnoses of biological conditions and treatments and / or procedures obtained by subscribers of the one or more health insurance companies. In various examples, the health insurance code may also indicate diagnostic procedures obtained by the individual related to one or more biological conditions that may be present in the individual. In one or more examples, the diagnostic procedures may provide information used in detecting the presence of a biological condition. The diagnostic procedures may also provide information used to determine the progression of the biological condition. In one or more illustrative examples, the diagnostic procedures may include one or more imaging procedures, one or more assays, one or more laboratory procedures, one or more combinations thereof, and the like.

[0036] The data integration and analysis system 102 may also obtain information from a molecular data repository 108. The molecular data repository 108 may store data for a number of individuals related to genomic, genetic, metabolomic, transcriptomic, fragmentomic, immune receptor, methylation, epigenomic, and / or proteomic information. In one or more embodiments, the data integration and analysis system 102 and the molecular data repository 108 may be created and maintained by different entities. In one or more additional embodiments, the data integration and analysis system 102 and the molecular data repository 108 may be created and maintained by the same entity.

[0037] The genomic information may indicate one or more mutations corresponding to the individual's genes. The individual's genetic mutations may correspond to differences between the individual's nucleic acid sequence and one or more reference genomes. The reference genome may include a known reference genome, such as hg19. In various embodiments, the individual's genetic mutations may correspond to differences in the individual's germline genes relative to the reference genome. In one or more additional embodiments, the reference genome may include the individual's germline genome. In one or more further embodiments, the individual's genetic mutations may include somatic mutations. The individual's genetic mutations may be associated with insertions, deletions, single base mutations, loss of heterozygosity, duplications, amplifications, translocations, fusion genes, or one or more combinations thereof.

[0038] In one or more illustrative examples, the genomic information stored by the molecular data repository 108 may include a genomic profile of tumor cells present in the individual. In these circumstances, the genomic information may be derived from an analysis of genetic material, such as, but not limited to, deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA), from a tissue sample or a sample containing a tumor biopsy, circulating tumor cells (CTCs), exosomes or efferosomes, or from circulating nucleic acids (e.g., cell-free DNA) found in the individual's blood sample that are present due to the degradation of tumor cells present in the individual. In one or more examples, the genomic information of the individual's tumor cells may correspond to one or more target regions. One or more mutations present with respect to the one or more target regions may indicate the presence of tumor cells in the individual. The genomic information stored by the molecular data repository 108 may be generated in association with an assay or other diagnostic test that may determine one or more mutations with respect to one or more target regions of a reference genome.

[0039] "Cell-free DNA", "cfDNA molecule", or simply "cfDNA" includes DNA molecules that occur in a subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or saliva), including DNA that is not contained within or otherwise bound to a cell at the time of isolation from the subject. The DNA was originally present within a cell or cells of a large complex biological organism (e.g., a mammal), or within other cells, such as bacteria that colonize the organism, but the DNA has been released from the cell into the fluid found within the organism. cfDNA includes, but is not limited to, the cell-free genomic DNA of a subject (e.g., the genomic DNA of a human subject) and the cell-free DNA of microorganisms, such as bacteria that inhabit the subject (whether pathogenic bacteria or bacteria that are normally found in commonly colonized locations such as the intestine or skin of healthy control groups), but does not include the cell-free DNA of microorganisms that simply contaminate a sample of bodily fluids. Typically, cfDNA can be obtained by obtaining a sample of a fluid without the need to perform an in vitro cell lysis step and including removal of cells present in the fluid (e.g., centrifugation of blood to remove cells).

[0040] In one or more additional embodiments, the data integration and analysis system 102 may obtain information from one or more additional data repositories 110. The one or more additional data repositories 110 may store data related to an electronic medical record of an individual whose data is present in at least one of the health insurance claims data repository 106 or the molecular data repository 108. Additionally, the one or more additional data repositories 110 may store data related to a pathology report of an individual whose data is present in at least one of the health insurance claims data repository 106 or the molecular data repository 108. In various embodiments, the one or more additional data repositories 110 may store data related to a biological condition and / or a treatment related to a biological condition. In one or more embodiments, at least a portion of the data integration and analysis system 102 and the one or more additional data repositories 110 may be created and maintained by different entities. In one or more further embodiments, at least a portion of the data integration and analysis system 102 and the one or more additional data repositories 110 may be created and maintained by the same entity.

[0041] In one or more further implementations, the data integration and analysis system 102 may obtain information from one or more reference information data repositories 112. The one or more reference information data repositories 112 may store information including definitions, standards, protocols, terminology tables, one or more combinations thereof, and the like. In various examples, the information stored by the one or more reference information data repositories may correspond to biological conditions and / or treatments for biological conditions. In one or more illustrative examples, the one or more reference information data repositories 112 may include RxNorm. (RxNorm provides normalized names for clinical drugs and links the names to many of the drug terminology tables used in pharmacy management and drug interaction software.) In one or more examples, at least a portion of the data integration and analysis system 102 and the one or more reference information data repositories 112 may be created and maintained by different entities. In one or more further embodiments, at least a portion of the data integration and analysis system 102 and the one or more reference information data repositories 112 may be created and maintained by the same entity.

[0042] The data integration and analysis system 102 may obtain data from at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112 via one or more communication networks accessible to the data integration and analysis system 102 and accessible to at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112. The data integration and analysis system 102 may also obtain data from at least one of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, or the reference information data repository 112 via one or more secure communication channels. Additionally, the data integration and analysis system 102 may obtain data from at least one of the health insurance claims data repository 106, the molecular data repository 108, one or more additional data repositories 110, or the reference information data repository 112 via one or more application programming interface (API) calls.

[0043] The data integration and analysis system 102 may include a data integration system 114. The data integration system 114 may obtain data from the health insurance claims data repository 106 and the molecular data repository 108 to generate the integrated data repository 104. The data integration system 114 may also obtain data from one or more additional data repositories 110 to generate the integrated data repository 104. In various embodiments, the data integration system 114 may implement one or more natural language processing techniques to integrate data from the one or more additional data repositories 110 into the integrated data repository 104.

[0044] In one or more embodiments, the data integration system 114 may generate one or more tokens to identify an individual having data stored in the health insurance claims data repository 106 and having data stored in the molecular data repository 108. In various embodiments, the data integration system 114 may generate one or more tokens by implementing one or more hash functions. The data integration system 114 may implement one or more hash functions to generate one or more tokens based on information stored by at least one of the health insurance claims data repository 106 or the molecular data repository 108. For example, the information used by the data integration system 114 to generate the individual tokens by implementing a hash function may include at least one of an individual individual's identifier, the individual individual's date of birth, the individual individual's zip code, the individual individual's birth date, or the individual individual's gender. In one or more illustrative embodiments, the individual individual's identifier may include a combination of at least a portion of the individual individual's first name and at least a portion of the individual individual's last name. Tokens generated using data from different data repositories may correspond to the same or similar information or the same or similar types of information stored by the different data repositories. To illustrate, a token may be generated using a portion of an individual's name, date of birth, at least a portion of a zip code, and gender obtained from the health insurance claims data repository 106 and the molecular data repository 108.

[0045] The data integration system 114 may integrate data from a number of different data sources by analyzing the tokens generated by implementing one or more hash functions using data obtained from the number of different data sources. For example, the data integration system 114 may obtain one or more first tokens generated from data stored by the health insurance claims data repository 106 and one or more second tokens generated from data stored by the molecular data repository 108. The data integration system 114 may analyze the one or more first tokens with respect to the one or more second tokens to determine individual first tokens that correspond to the individual second tokens. In one or more illustrative examples, the data integration system 114 may identify individual first tokens that match individual second tokens. A first token may match a second token when the data of the first token has at least a threshold amount of similarity with respect to the data of the second token. In one or more embodiments, a first token may match a second token when the data of the first token is identical to the data of the second token. To illustrate, a first token may match a second token when an alphanumeric string of the first token is identical to an alphanumeric string of the second token.

[0046] By determining a first token generated using data stored by the health insurance claims data repository 106 that corresponds to a second token generated using data stored by the molecular data repository 108, the data integration system 114 may identify individuals who have data stored in both the health insurance claims data repository 106 and the molecular data repository 108. In this manner, the data integration system 114 may obtain data from a number of individuals from the health insurance claims data repository 106 and data from the molecular data repository 108 from the same number of individuals and store the health insurance claims data and molecular data for that number of individuals in the integrated data repository 104.

[0047] The data integration system 114 may also integrate data stored by one or more additional data repositories 110 with data from the health insurance claims data repository 106 and the molecular data repository 108 to generate the integrated data repository 104. To illustrate, the data integration system 114 may obtain one or more third tokens generated from data stored by the additional data repository 110, such as a data repository that stores data corresponding to pathology reports. The data integration system 114 may analyze the one or more third tokens with respect to the first token generated using information stored by the health insurance claims data repository 106 and the second token generated using information stored by the molecular data repository 108 to determine individual third tokens corresponding to each of the first tokens and each of the second tokens. In one or more illustrative embodiments, the data integration system 114 may identify a third token generated using one or more hash functions and a common set of information obtained from the health insurance claims data repository 106, the molecular data repository 108, and the additional data repository 110.

[0048] By determining a third token generated using data stored by the additional data repository 110 that corresponds to the first token generated using data stored by the health insurance claims data repository 106 and the second token generated using data stored by the molecular data repository 108, the data integration system 114 may identify individuals having data stored in the health insurance claims data repository 106, the molecular data repository 108, and the additional data repository 110. In this manner, the data integration system 114 may obtain data from a number of individuals from the health insurance claims data repository 106 and data from the molecular data repository 108 and the additional data repository 110 from a same number of individuals and store the health insurance claims data, molecular data, and additional data for that number of individuals in the integrated data repository 104.

[0049] The data stored by the integrated data repository 104 for the number of individuals may be accessible using the individual's individual identifier. The data integration system 114 may implement a number of techniques as part of the de-identification process with respect to storing and retrieving the information of the individuals in the integrated data repository 104. The individual's identifier may correspond to a key generated using at least one hash function. The individual's identifier may also be generated by implementing one or more salting processes with respect to the key generated using at least one hash function, the token generated using one or more hash functions, and a common set of information obtained from the health insurance claims data repository 106, the molecular data repository 108, and / or the additional data repository 110. In one or more illustrative embodiments, the identifiers generated by the data integration system 114 to access information about the individual individuals stored by the integrated data repository 104 may be unique for each individual. In one or more embodiments, the individual's identifier may be generated using at least a portion of the information used to generate the token associated with the individual. In one or more additional embodiments, the individual's identifier may be generated using information different from the information used to generate the token associated with the individual.

[0050] The data integration system 114 may also generate the integrated data repository 104 from a number of different combinations of data repositories in a similar manner. For example, the data integration system 114 may obtain tokens generated from information stored by the health insurance claims data repository 106 and additional tokens generated from information stored by one or more additional data repositories 110. The data integration system 114 may determine individual tokens generated from information stored by the health insurance claims data repository 106 that correspond to individual additional tokens generated from information stored by the one or more additional data repositories 110. By determining tokens generated using data stored by the health insurance claims data repository 106 that correspond to additional tokens generated using data stored by the additional data repository 110, the data integration system 114 may identify individuals having data stored in both the health insurance claims data repository 106 and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the health insurance claims data repository 106 from a number of individuals and data from the additional data repository 110 from the same number of individuals and store the health insurance claims data and the additional data for that number of individuals in the integrated data repository 104. The health insurance claims data and the additional data stored by the integrated data repository 104 for that number of individuals may be accessible using the individuals' individual identifiers.

[0051] In one or more further embodiments, the data integration system 114 may obtain tokens generated from information stored by the molecular data repository 108 and tokens generated from information stored by the one or more additional data stores 110. The data integration system 114 may determine individual tokens generated from information stored by the molecular data repository 108 that correspond to individual additional tokens generated from information stored by the one or more additional data repositories 110. By determining tokens generated using data stored by the molecular data repository 108 that correspond to additional tokens generated using data stored by the additional data repository 110, the data integration system 114 may identify individuals having data stored in both the molecular data repository 108 and the additional data repository 110. In this manner, the data integration system 114 may obtain data from the molecular data repository 108 from a number of individuals and data from the additional data repositories 110 from a same number of individuals and store molecular data and additional data for that number of individuals in the integrated data repository 104. The molecular data and additional data stored by the integrated data repository 104 for that number of individuals may be accessible using the individuals' individual identifiers.

[0052] The data stored by the integrated data repository 104 may be stored in accordance with one or more regulatory frameworks that protect privacy and ensure security of individuals' medical records, health information, and insurance information. For example, the data may be stored by the integrated data repository 104 in accordance with one or more government regulatory frameworks that are directed to protecting personal information, such as the Health Insurance Portability and Accountability Act (HIPAA) and / or the General Data Protection Regulation (GDPR). The integrated data repository 104 also stores the data in an anonymous and de-identified manner to ensure protection of the privacy of individuals whose data is stored by the integrated data repository 104. To further ensure the privacy of individuals whose data is stored by the integrated data repository 104, the data integration system 114 may periodically regenerate the integrated data repository 104. For example, the data integration system 114 may create the integrated data repository 104 once per quarter. In one or more additional embodiments, the data integration system 114 may generate the integrated data repository 104 monthly, weekly, or once every two weeks. By periodically regenerating the integrated data repository 104, rather than simply refreshing it when new data is available, the integrated data repository 104 enhances privacy protections with respect to the data stored by the integrated data repository 104. That is, in situations where the data repository is simply refreshed with new data, it may be possible to more easily track individuals associated with data newly added to the data repository, since the number of new individuals added at a given time is typically smaller than the existing number of individuals who already have data stored by the data repository.

[0053] In various embodiments, the data stored by the integrated data repository 104 may be accessed via a database management system. Additionally, the integrated data repository 104 may store data according to one or more database models. In one or more embodiments, the integrated data repository 104 may store data according to one or more relational database technologies. For example, the integrated data repository 104 may store data according to a relational database model. In one or more additional embodiments, the integrated data repository 104 may store data according to an object-oriented database model. In one or more further embodiments, the integrated data repository 104 may store data according to an extensible markup language (XML) database model. In yet additional embodiments, the integrated data repository 104 may store data according to a structured query language (SQL) database model. In still further embodiments, the integrated data repository 104 may store data according to an image database model.

[0054] The data integration system 114 may generate the integrated data repository 104 by generating a number of data tables and creating links between the data tables. The links may indicate logical connections between the data tables. The data integration system 114 may generate the data tables by extracting a defined set of data from information obtained from the data repositories 106, 108, 110, 112 and storing the data in rows and columns of the respective data tables. In various embodiments, the logical connections between the data tables may include at least one of a one-to-one link, where a row of information in one data table corresponds to a row of information in another data table, a one-to-many link, where a row of information in one data table corresponds to multiple rows of information in another data table, or a many-to-many link, where multiple rows of information in one data table correspond to multiple rows of information in another data table.

[0055] The number of data tables may be arranged according to the data repository schema 116. In the illustrative example of FIG. 1, the data repository schema 114 includes a first data table 118, a second data table 120, a third data table 122, a fourth data table 124, and a fifth data table 124. Although the illustrative example of FIG. 1 includes five data tables, in additional implementations, the data repository schema 116 may include more or fewer data tables. The data repository schema 116 may also include links between the data tables 118, 120, 122, 124, 128. The links between the data tables 118, 120, 122, 124, 126 may indicate that information retrieved from one of the data tables 118, 120, 122, 124, 126 results in additional information stored by one or more additional data tables 118, 120, 122, 124, 126 being retrieved. In addition, not all of the data tables 118, 120, 122, 124, 126 may be linked to each of the other data tables 118, 120, 120, 122, 124, 126. In the illustrative example of FIG. 1 , the first data table 118 is logically coupled to the second data table 118 by a first link 128, and the first data table 118 is logically coupled to the fourth data table 124 by a second link 130. In addition, the second data table 120 is logically coupled to the third data table 122 via a third link 132, and the fourth data table 124 is logically coupled to the fifth data table 126 via a fourth link 134. Furthermore, the third data table 122 is logically coupled to the fifth data table 126 via a fifth link 136.

[0056] In various embodiments, data tables are added and / or removed from the data repository schema 116 such that additional links between data tables may be added to or removed from the data repository schema 116. In one or more illustrative embodiments, the integrated data repository 104 may store data tables in accordance with the data repository schema 116 for at least a portion of the individuals for whom the data integration system 114 obtained information from a combination of at least two of the health insurance claims data repository 106, the molecular data repository 108, the one or more additional data repositories 110, and the one or more reference information data repositories 112. As a result, the integrated data repository 104 may store individual instances of the data tables 118, 120, 122, 124, 126 in accordance with the data repository schema 116 for thousands, tens of thousands, up to hundreds of thousands, or more individuals.

[0057] The data integration and analysis system 102 may also include a data pipeline system 138. The data pipeline system 138 may include a number of algorithms, software codes, scripts, macros, or other bundles of computer executable instructions that process the information stored by the integrated data repository 104 and generate additional data sets. The additional data sets may include information obtained from one or more of the data tables 118, 120, 122, 124, 126. The additional data sets may also include information derived from data obtained from one or more of the data tables 118, 120, 122, 124, 126. The components of the data pipeline system 138 implemented to generate the first additional data set may be different from the components of the data pipeline system 138 used to generate the second additional data set.

[0058] In one or more embodiments, the data pipeline system 138 may generate a data set indicative of pharmacy treatments received by a number of individuals. In one or more illustrative embodiments, the data pipeline system 138 may analyze information stored in at least one of the data tables 118, 120, 122, 124, 126 to determine health insurance codes corresponding to drug treatments received by a number of individuals. The data pipeline system 138 may analyze the health insurance codes corresponding to the drug treatments with respect to a library of data indicative of defined drug treatments corresponding to the one or more health insurance codes to determine the names of the drug treatments received by the individuals. In one or more additional embodiments, the data pipeline system 138 may analyze information stored by the integrated data repository 104 to determine medical procedures received by a number of individuals. To illustrate, the data pipeline system 138 may analyze information stored by one of the data tables 118, 120, 122, 124, 126 to determine the treatments received by the individuals via at least one of an injection or an intravenous. In one or more further embodiments, the data pipeline system 138 may analyze the information stored by the integrated data repository 104 to determine an episode of treatment for an individual, a line of therapy the individual has received, the progression of a biological condition, or time to next treatment. In various embodiments, the datasets generated by the data pipeline system 138 may be different for different biological conditions. For example, the data pipeline system 138 may generate a first number of datasets for a first type of cancer, such as lung cancer, and a second number of datasets for a second type of cancer, such as colon cancer.

[0059] The data pipeline system 138 may also determine one or more confidence levels to assign to information associated with an individual having data stored by the integrated data repository 104. The distinct confidence levels may correspond to different measures of accuracy for the information associated with an individual having data stored by the integrated data repository 104. The information associated with the distinct confidence levels may correspond to one or more characteristics of the individual derived from the data stored by the integrated data repository 104. The confidence level values ​​for the one or more characteristics may be generated by the data pipeline system 138 in conjunction with generating one or more data sets from the integrated data repository 104. In one or more embodiments, the first confidence level may correspond to a first range of the accuracy measure, the second confidence level may correspond to a second range of the accuracy measure, and the third confidence level may correspond to a third range of the accuracy measure. In one or more additional embodiments, the second range of the accuracy measure may include values ​​that are less than the values ​​of the first range of the accuracy measure, and the third range of the accuracy measure may include values ​​that are less than the values ​​of the second range of the accuracy measure. In one or more illustrative embodiments, the information corresponding to the first confidence level may be referred to as gold standard information, the information corresponding to the second confidence level may be referred to as silver standard information, and the information corresponding to the third confidence level may be referred to as bronze standard information.

[0060] The data pipeline system 138 may determine a value for the confidence level of the individual's characteristic based on a number of factors. For example, a separate set of information may be used to determine the individual's characteristic. The data pipeline system 138 may determine the confidence level of the individual's characteristic based on the amount of completeness of the separate set of information used to determine the characteristic for the individual. In a situation where one or more pieces of information are missing from a set of information associated with a first number of individuals, the confidence level for the characteristic may be lower than for a second number of individuals where no information is missing from the set of information. In one or more embodiments, the amount of missing information may be used by the data pipeline system 138 to determine the confidence level of the individual's characteristic. To illustrate, a greater amount of missing information used to determine the individual's characteristic may cause a lower confidence level for the characteristic than in a situation where a smaller amount of missing information is used to determine the characteristic. Furthermore, different types of information may correspond to various confidence levels for the characteristic. In one or more embodiments, the presence of a first piece of information used to determine the individual's characteristic may result in a higher confidence level for the characteristic than the presence of a second piece of information used to determine the characteristic.

[0061] In one or more illustrative examples, the data pipeline system 138 may determine a number of individuals included in the cohort with a primary diagnosis of lung cancer (or other biological condition). The data pipeline system 138 may determine a confidence level for the individual with respect to being classified as having a primary diagnosis of lung cancer. The data pipeline system 138 may use information from a number of columns included in the data tables 118, 120, 122, 124, 126 to determine a confidence level for the inclusion of the individual in the lung cancer cohort. The number of columns may include health insurance codes associated with the diagnosis of the biological condition and / or the treatment of the biological condition. Additionally, the number of columns may correspond to the date of diagnosis and / or treatment for the biological condition. The data pipeline system 138 may determine that the confidence level of the individual characterized as being part of the lung cancer cohort is higher in scenarios where information is available for the number of columns or at least for every threshold number of columns than in cases where information is available for less than the threshold number of columns. Further, the data pipeline system 138 may determine a confidence level for an individual to be included in the lung cancer cohort based on the type of information associated with one or more columns and the availability of the information. To illustrate, in a situation where one or more diagnostic codes are present in association with one or more time periods for a group of individuals and one or more treatment codes are absent, the data pipeline system 138 may determine that the confidence level of including the group of individuals in the lung cancer cohort is greater than in a situation where at least one of the diagnostic codes is absent and a treatment code used to determine whether the individual is included in the lung cancer cohort is present.

[0062] The data integration and analysis system 102 may include a data analysis system 140. The data analysis system 148 may receive integrated data repository requests 142 from one or more computing devices, such as the exemplary computing device 144. The one or more integrated data repository requests 142 may cause data to be read from the integrated data repository 104. In various embodiments, the one or more integrated data repository requests 142 may cause data to be read from one or more datasets generated by the data pipeline system 138. The integrated data repository requests 142 may specify data to be read from the integrated data repository 104 and / or one or more datasets generated by the data pipeline system 138. In one or more additional embodiments, the integrated data repository requests 142 may include one or more pre-built queries corresponding to computer-executable instructions to read a specified set of data from the integrated data repository 104 and / or one or more datasets generated by the data pipeline system 138.

[0063] In response to the one or more integrated data repository requests 142, the data analysis system 140 may analyze data retrieved from at least one of the integrated data repository 104 or the one or more data sets generated by the data pipeline system 138 and generate data analysis results 146. The data analysis results 146 may be transmitted to one or more computing devices, such as the exemplary computing device 148. Although the illustrative example of FIG. 1 shows the one or more integrated data repository requests 142 and the data analysis results 146 from one computing device 144 being transmitted to another computing device 148, in one or more additional implementations, the data analysis results 146 may be received by the same computing device that sent the one or more integrated data repository requests 142. The data analysis results 146 may be displayed by one or more user interfaces rendered by the computing device 144 or the computing device 148.

[0064] In one or more examples, the data analysis system 140 may implement at least one of one or more machine learning techniques or one or more statistical techniques to analyze the data retrieved in response to the one or more integrated data repository requests 142. In one or more examples, the data analysis system 140 may implement one or more artificial neural networks to analyze the data retrieved in response to the one or more integrated data repository requests 142. To illustrate, the data analysis system 140 may implement at least one of one or more convolutional neural networks or one or more residual neural networks to analyze the data retrieved from the integrated data repository 104 in response to the one or more integrated data repository requests 142. In at least some embodiments, the data analysis system 140 may implement one or more random forest techniques, one or more support vector machines, or one or more hidden Markov models to analyze data retrieved in response to one or more integrated data repository requests 142.

[0065] In one or more illustrative examples, the data analysis system 140 may determine a survival rate of an individual with lung cancer in response to one or more treatments. In one or more additional illustrative examples, the data analysis system 140 may determine a survival rate of an individual with one or more genomic region mutations in response to one or more treatments. In various examples, the data analysis system 140 may generate the data analysis result 146 in a situation where the data retrieved from at least one of the one or more datasets generated by the integrated data repository 104 or the data pipeline system 138 meets one or more criteria. For example, the data analysis system 140 may determine whether at least a portion of the data retrieved in response to one or more integrated data repository requests 142 meets a threshold confidence level. In a situation where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests 142 is below the threshold confidence level, the data analysis system 140 may refrain from generating at least a portion of the data analysis result 146. In situations where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests 142 is at least the threshold confidence level, data analysis system 140 may generate at least a portion of data analysis results 146. In various embodiments, the threshold confidence level may be related to the type of data analysis results 146 being generated by data analysis system 140.

[0066] In one or more illustrative examples, the data analysis system 140 may receive the integrated data repository request 142 and generate a data analysis result 146 indicative of a survival rate for one or more individuals. In these cases, the data analysis system 140 may determine whether the data stored by the integrated data repository 104 and / or by one or more datasets generated by the data pipeline system 138 meets a threshold confidence level, such as a gold standard confidence level. In one or more additional examples, the data analysis system 140 may receive the integrated data repository request 142 and generate a data analysis result 146 indicative of a treatment received by one or more individuals. In these implementations, the data analysis system 140 may determine whether the data stored by the integrated data repository 104 and / or by one or more datasets generated by the data pipeline system 138 meets a lower threshold confidence level, such as a bronze standard confidence level.

[0067] In one or more additional illustrative examples, the data analysis system 140 may receive the integrated data repository request 142 to determine individuals who have one or more genomic mutations and have received one or more treatments for a biological condition. Continuing with this example, the data analysis system 140 may determine a survival rate of individuals with one or more genomic mutations in association with one or more treatments that the individuals have received. The data analysis system 140 may then identify the effectiveness of a treatment for the individual in association with a genomic mutation that may be present in the individual based on the survival rate of the individual. In this manner, health outcomes for individuals may be improved by identifying potential treatments that may be more effective than current treatments provided to the individuals for a population of individuals with one or more genomic mutations.

[0068] 2 illustrates an example framework 200 that corresponds to an arrangement of data tables in a unified data repository, according to one or more implementations. In the illustrative example of FIG. 2, the framework 200 includes a data repository schema 202 that includes a first data table 204, a second data table 206, a third data table 208, a fourth data table 210, a fifth data table 212, a sixth data table 214, and a seventh data table 216. The illustrative example of FIG. 2 includes seven data tables, although in additional implementations, the data repository schema 202 may include more or fewer data tables. The data repository schema 202 may also include links between the data tables 204, 206, 208, 210, 212, 214, 216. The links between the data tables 204, 206, 208, 210, 212, 214, 216 may indicate that information retrieved from one of the data tables 204, 206, 208, 210, 212, 214, 216 results in additional information being retrieved that is stored by one or more additional data tables 204, 206, 208, 210, 212, 214, 216. Additionally, not all of the data tables 204, 206, 208, 210, 212, 214, 216 may be linked to each of the other data tables 204, 206, 208, 210, 212, 214, 216. 2, the first data table 204 is logically coupled to the second data table 206 by a first link 218, and the third data table 208 is logically coupled to the second data table 206 by a second link 220. The second data table 206 is also logically coupled to the fourth data table 210 by a third link 222, the second data table 206 is logically coupled to the fifth data table 212 by a fourth link 224, and the second data table 206 is logically coupled to the sixth data table 214 by a fifth link 226.Additionally, the fifth data table 212 is logically coupled to the sixth data table 214 by a sixth link 228, which is logically coupled to the seventh data table 216 by a seventh link 230. Additionally, the seventh data table 216 is logically coupled to the fourth data table 210 by an eighth link 232. In various embodiments, data tables are added and / or removed from the data repository schema 202, such that additional links between data tables may be added or removed from the data repository schema 202. In one or more illustrative embodiments, the integrated data repository 104 may store data tables according to the data repository schema 202 for at least a portion of individuals for whom the data integration system 114 obtained information from a combination of at least two of the health insurance claims data repository 106, the molecular data repository 108, and the one or more additional data repositories 110. As a result, the integrated data repository 104 may store individual instances of data tables 204, 206, 208, 210, 212, 214, 216 according to the data repository schema 204 for thousands, tens of thousands, up to hundreds of thousands, or more individuals.

[0069] In one or more embodiments, the first data table 204 may store data corresponding to genomics and genomics testing for an individual. For example, the first data table 204 may include columns containing information corresponding to the panel used to generate the genomics data, mutations in the genomic region, the type of mutation, copy number of the genomic region, coverage data indicating the number of nucleic acid molecules identified in the sample with one or more mutations, test date, and patient information. The first data table 204 may also include one or more columns containing health insurance data codes that may correspond to one or more diagnostic codes. In addition, the information in the first data table 204 may include at least one identifier for an individual associated with a case in the first data table 204.

[0070] The second data table 206 may store data related to one or more patient visits by an individual to one or more health care providers. The third data table 208 may store information corresponding to individual services provided to an individual in connection with one or more patient visits to one or more health care providers represented by the second data table 206. To illustrate, an individual may visit a health care provider and multiple services may be performed on the individual at the visit. The second data table 206 may include columns indicating information for each of multiple services performed during the patient visit. Multiple third data tables 208 may be generated for a patient visit, including columns indicating a finer level of information related to individual services provided during the patient visit than the information stored by the second data table 206 related to the patient visit. For example, the second data table 206 may include multiple columns indicating health insurance codes related to different services provided to an individual during a patient visit, and the third data table 208 related to one of the services may include multiple columns for additional health insurance codes corresponding to additional information related to the individual service. The second data table 206 and the third data table 208 relating to patient visits may indicate one or more dates of service corresponding to the patient visit.

[0071] The fourth data table 210 may include columns indicating information about an individual whose information is stored by the integrated data repository 104. For example, the fourth data table 210 may include columns indicating information related to at least one of the individual's location, the individual's gender, the individual's date of birth, the individual's date of death (if applicable), or one or more keys associated with the individual. In one or more embodiments, the fourth data table 210 may include one or more columns related to whether erroneous data has been identified for the individual. In various embodiments, a single fourth data table 210 may be generated for an individual individual. Thus, the data repository schema 202 may include multiple instances of the fourth data table 210, such as thousands, tens of thousands, up to hundreds of thousands, or more.

[0072] The fifth data table 212 may include columns that indicate information related to a health insurance company or government entity that made a payment for one or more services provided to an individual. For example, the fifth data table 212 may include one or more payer identifiers. The sixth data table 214 may include columns that include information corresponding to health insurance coverage information for an individual. In one or more embodiments, the sixth data table 214 may include columns that indicate the presence of medical coverage for the individual, the presence of pharmacy coverage for the individual, and the type of health insurance plan associated with the individual, such as Health Maintenance Organization (HMO), Preferred Provider Organization (PPO), and the like.

[0073] The seventh data table 216 may include columns indicating information related to drug treatments obtained by individual individuals. In one or more embodiments, the seventh data table 216 may include one or more columns indicating health insurance codes corresponding to drug treatments available through the pharmacy. The health insurance codes may correspond to the individual drug treatments. In addition, the health insurance codes may indicate a diagnosis of a biological condition for the individual. The seventh data table 216 may also include additional information, such as at least one of dosage, days of supply, total amount prescribed, number of refills authorized, date of service, or information related to the individual receiving the drug treatment.

[0074] In various embodiments, the data repository schema 202 may provide the results of an analysis of the information stored by the data tables 204, 206, 208, 210, 212, 214, 216 in a more efficient manner than a typical data repository schema. For example, the logical connections between the data tables 204, 206, 208, 210, 212, 214, 216 may be arranged to efficiently retrieve related data across different data tables 204, 206, 208, 210, 212, 214, 216. In situations where the data tables 204, 206, 208, 210, 212, 214, 216 are arranged in a contiguous manner and / or where a greater number of the data tables 204, 206, 208, 210, 212, 214, 216 are logically connected, retrieving data from the integrated data repository 104 from one or more of the data tables 204, 206, 208, 210, 212, 214, 216 to respond to a request for information from the integrated data repository 104 may be less efficient than in situations where the data repository schema 202 is implemented.

[0075] 3 illustrates an architecture 300 for generating one or more data sets from information retrieved from a data repository that integrates health-related data from a number of sources, according to one or more implementations. The architecture 300 may include a data integration and analysis system 102 and an integrated data repository 104. In addition, the data integration and analysis system 102 may include at least a data pipeline system 138 and a data analysis system 140. The data pipeline system 138 may include a number of sets of data processing instructions executable to generate individual data sets that can be analyzed by the data analysis system 140 in response to an integrated data repository request 142 to generate data analysis results 146.

[0076] The data pipeline system 138 may include a first data processing instruction 302, a second data processing instruction 304, and up to Nth data processing instruction 306. The data processing instructions 302, 304, 306 may be executable by one or more processing units to perform a number of operations to generate individual data sets using information retrieved from the unified data repository 104. In one or more illustrative embodiments, the data processing instructions 302, 304, 306 may include at least one of software code, scripts, API calls, macros, and the like. The first data processing instruction 302 may be executable to generate a first data set 308. In addition, the second data processing instruction 304 may be executable to generate a second data set 310. Furthermore, the Nth data processing instruction 306 may be executable to generate an Nth data set 312. In various embodiments, after the data integration and analysis system 102 generates the integrated data repository 104, the data pipeline system 138 may execute the data processing instructions 302, 304, 306 to generate the data sets 308, 310, 312. In one or more embodiments, the data sets 308, 310, 312 may be stored by the integrated data repository 104 or by an additional data repository accessible to the data integration and analysis system 102. At least some of the data processing instructions 302, 304, 306 may analyze health insurance codes and generate at least some of the data sets 308, 310, 312. Additionally, at least some of the data processing instructions 302, 304, 306 may analyze genomics data and generate at least some of the data sets 308, 310, 312.

[0077] In one or more examples, the first data processing instructions 302 may be executable to read data from one or more first data tables stored by the integrated data repository 104. The first data processing instructions 302 may also be executable to read data from one or more defined columns of the one or more first data tables. In various examples, the first data processing instructions 302 may be executable to identify individuals having health insurance codes stored in one or more column and row combinations that correspond to the one or more diagnostic codes. The first data processing instructions 302 may then be executable to analyze the one or more diagnostic codes to determine the biological condition with which the individual has been diagnosed. In one or more illustrative examples, the first data processing instructions 302 may be executable to analyze the one or more diagnostic codes with respect to a library of diagnostic codes that indicate one or more biological conditions that correspond to the individual diagnostic codes. The library of diagnostic codes may include hundreds up to thousands of diagnostic codes. The first data processing instructions 302 may also be executable to determine individuals who have been diagnosed with a biological condition by analyzing individual timing information, such as date of treatment, date of diagnosis, date of death, one or more combinations thereof, and the like.

[0078] The second data processing instructions 304 may be executable to read data from one or more second data tables stored by the integrated data repository 104. The second data processing instructions 304 may also be executable to read data from one or more defined columns of the one or more second data tables. In various embodiments, the second data processing instructions 304 may be executable to identify individuals having health insurance codes stored in one or more column and row combinations that correspond to one or more treatment codes. The one or more treatment codes may correspond to treatments obtained from a pharmacy. In one or more additional embodiments, the one or more treatment codes may correspond to treatments received through a medical procedure, such as an injection or intravenous. The second data processing instructions 304 may be executable to determine one or more treatments that correspond to individual health insurance codes contained within the one or more second data tables by analyzing the health insurance codes in association with a predetermined set of information. The predetermined set of information may include a data library indicating one or more treatments corresponding to one of hundreds up to thousands of health insurance codes. The second data processing instructions 304 may generate a second data set 310 to indicate individual treatments received by a group of individuals. In one or more illustrative examples, the group of individuals may correspond to individuals included in the first data set 308. The second data set 310 may be arranged in rows and columns, with one or more rows corresponding to a single individual and one or more columns indicating treatments received by the individual individuals.

[0079] The Nth processing instructions 306 (N may be any positive integer) may be executable to generate the Nth dataset 312 by combining information from a number of previously generated datasets, such as the first dataset 308 and the second dataset 310. Additionally, the Nth processing instructions 306 may be executable to generate the Nth dataset 312, retrieve additional information from one or more additional columns of the integrated data repository 104, and combine the additional information from the integrated data repository 104 with the information obtained from the first dataset 308 and the second dataset 310. For example, the Nth processing instructions 306 may be executable to identify individuals included in the first dataset 308 who have been diagnosed with a biological condition, analyze defined columns of one or more additional data tables of the integrated data repository 104, and determine dates of treatment indicated in the second dataset 210 that correspond to the individuals included in the first dataset 308. In one or more further embodiments, the Nth processing instructions 306 may be executable to analyze columns of one or more additional data tables in the integrated data repository 104 to determine dosages of treatments indicated in the second dataset 310 received by individuals included in the first dataset 308. In this manner, the Nth processing instructions 306 may be executable to generate episodes of a treatment dataset based on information included in the cohort dataset and the treatment dataset.

[0080] In one or more illustrative examples, in response to receiving the integrated data repository request 142, the data analysis system 140 may determine one or more datasets corresponding to characteristics of the queries associated with the integrated data repository request 142. For example, the data analysis system 140 may determine that information included in the first dataset 308 and the second dataset 310 is applicable to responding to the integrated data repository request 142. In these scenarios, the data analysis system 140 may analyze at least a portion of the data included in the first dataset 308 and the second dataset 310 to generate the data analysis result 146. In one or more additional examples, the data analysis system 140 may determine different datasets for responding to different queries included in the integrated data repository request 142 to generate the data analysis result 146.

[0081] The use of specific sets of data processing instructions to generate the individual data sets may reduce the number of inputs from users of the data integration and analysis system 102 and may reduce the amount of processing resources and memory utilized to process the integrated data repository requests 142. For example, without the specific architecture of the data pipeline system 138, the data utilized to respond to the integrated data repository requests 142 is assembled from the data repositories 104 each time an integrated data repository request 142 is received. In contrast, by implementing the data pipeline system 138 to execute the data processing instructions 302, 304, 306 to generate the data sets 308, 310, 312, the data required to respond to the various integrated data repository requests 142 is already assembled and may be accessed by the data analysis system 140 to respond to the integrated data repository requests 142. Thus, the computing resources used to respond to the integrated data repository requests 142 by implementing the data pipeline system 138 to generate the data sets 308, 310, 312 are less than a typical system that performs an information analysis and collection process for each integrated data repository request 142. Further, in situations where the data pipeline system 138 is not implemented, a user of the data integration and analysis system 102 may need to submit multiple integrated data repository requests 142 to analyze the information that the user intends to be analyzed, either because the ad-hoc collection of data to respond to an integrated data repository request 142 in a typical system is inaccurate or because the data analysis system 140 is called multiple times to perform an analysis of the information in a typical system that may be performed using a single integrated data repository request 142 when the data pipeline system 138 is implemented.

[0082] FIG. 4 illustrates an architecture 400 for generating an integrated data repository including de-identified health insurance claims data and de-identified genomics data, according to one or more implementations. The architecture 400 may include a data integration and analysis system 102, a health insurance claims data repository 106, and a molecular data repository 108. The data integration and analysis system 102 may obtain patient information 402 from the molecular data repository 108. The patient information 402 may include genomics data 404 regarding an individual having data stored by the molecular data repository 108. The genomics data 404 may represent the results of one or more nucleic acid sequencing operations that analyze the sequence of nucleic acid molecules contained in a sample obtained from the individual for one or more target genomic regions. In one or more embodiments, the sample may be obtained from tissue of one or more individuals. In one or more additional embodiments, the sample may be obtained from a fluid of one or more individuals, such as blood or plasma. The one or more target genomic regions may correspond to genomic regions corresponding to the presence of one or more biological conditions. For example, the target regions may correspond to genomic regions of a reference genome having a mutation present in an individual in which a biological condition exists. In one or more illustrative examples, the target regions may correspond to genomic regions of a reference human genome having one or more mutations present in an individual in which one or more forms of cancer exist. The patient information 402 may also include information indicative of personal information about the individual with data stored by the molecular data repository 108 and information corresponding to tests and analyses performed on samples provided by the individual.

[0083] The data integration and analysis system 102 may perform a de-identification process 406 that anonymizes personal information obtained from the molecular data repository 108. The data integration and analysis system 102 may implement one or more computational techniques as part of the de-identification process to anonymize data related to individuals stored by the molecular data repository 108 such that the de-identified data protects the privacy of the individuals and complies with one or more privacy regulatory frameworks. The de-identification process 406 may include accessing 408 a token. In various embodiments, the token may comprise an alphanumeric string. In one or more embodiments, the token may be generated by the data integration and analysis system 102. In one or more additional embodiments, the token may be generated by a third party and obtained by the data integration and analysis system 102.

[0084] The token may be generated using one or more hash functions in association with the subset 410 of the patient information 402. To illustrate, for an individual having information stored by the molecular data repository 108, the token may be generated using a combination of at least a portion of the individual's first name, at least a portion of the individual's last name, at least a portion of the individual's date of birth, the individual's gender, and at least a portion of the individual's location identifier. The de-identification process 406 may also include, at 412, generating an identifier for the individual having data stored by the molecular data repository 108. The identifier may be generated by the data integration and analysis system 102 using one or more hash functions that are different from the one or more hash functions used to generate the token. In one or more illustrative examples, the data integration and analysis system 102 may use one or more hash functions to generate intermediate versions of the individual identifiers and then apply one or more salting techniques to the intermediate versions of the identifiers to generate a final version of the identifiers. The salt function comprises a function configured to add at least one random bit to each intermediate identifier to generate an individual final identifier. In various embodiments, the data integration and analysis system 102 may generate 412 the identifier using at least a portion of the information about the individual individual stored by the molecular data repository 108. In one or more illustrative embodiments, the identifier may be generated based on a patient identifier included in the patient information 402. The identifier generated by the data integration and analysis system 102 may be unique with respect to the individual individual having data stored by the molecular data repository 108.

[0085] In an operation 414, the data integration and analysis system 102 may generate corrected patient information 416 based on the identifier. The corrected patient information 416 may include genomics data 404 associated with the individual associated with the molecular data repository 108 and an identifier for the individual. The corrected patient information 416 may have a data structure 418. The data structure 418 may include a column including the individual identifier for the individual associated with the molecular data repository 108 and a number of columns including genomics data 404 associated with the individual, such as identifiers of one or more genes, one or more genetic modifications, a type of genetic modification, etc.

[0086] The data integration and analysis system 102 may generate a token file 420. The token file 420 may include a first token 422 accessed in operation 408 for an individual having data stored by the molecular data repository 108. The token file 420 may have a data structure 424 including a number of columns that include information about the individual individual. The data structure 424 may include a column indicating an individual identifier generated by the data integration and analysis system 102 and a column indicating one or more first tokens 422 associated with the individual identifier. The data integration and analysis system 102 may transmit the token file 420 to a claims data management system 426 coupled to the claims data repository 106. The claims data management system 426 may analyze the first token 422 for a corresponding second token 428. The second token 428 may be accessed by or generated by the claims data management system 426. The second token 428 may be generated using a subset of information about an individual having data stored in the health insurance claims data repository 106 that is the same or similar to the subset 410 of patient information 402. For example, the second token 428 may be generated using a combination of at least a portion of the individual's first name, at least a portion of the individual's last name, at least a portion of the individual's date of birth, the individual's gender, and at least a portion of the individual's location identifier.

[0087] In various embodiments, the health insurance claims data management system 426 may retrieve health insurance claims data for an individual associated with an individual second token 428 that matches a corresponding first token 422 from the health insurance claims data repository 106. A first token 422 may match a second token 428 when the data of the first token 422 has at least a threshold amount of similarity with the data of the second token 428. In one or more embodiments, a first token 422 may match a second token 428 when the data of the first token 422 is identical to the data of the second token 428.

[0088] In response to identifying health insurance claim data for an individual having a respective second token 428 that corresponds to the respective first token 422, the health insurance claim data management system 426 may generate corrected health insurance claim data 430. The health insurance claim data management system 426 may transmit the corrected health insurance claim data 430 to the data integration and analysis system 102. In one or more embodiments, the corrected health insurance claim data 430 may be formatted according to a data structure 432. The data structure 432 may include a column that includes a subset of the second tokens 428 that correspond to the first tokens 422 and a number of columns that include the health insurance claim data.

[0089] At operation 434, the data integration and analysis system 102 may integrate the genomics data and health insurance claims data for individuals common to both the molecular data repository 108 and the health insurance claims data repository 106. The data integration and analysis system 102 may determine the individuals common to both the molecular data repository 108 and the health insurance claims data repository 106 by determining the genomics data and health insurance claims data that correspond to the common tokens. The data integration and analysis system 102 may determine that a first token 422 associated with a portion of the genomics data 404 corresponds to a second token 428 associated with a portion of the health insurance claims data by determining a measure of similarity between the first token 422 and the second token 428. In a scenario in which the first token 422 has at least a threshold amount of similarity with respect to the second token 428, the data integration and analysis system 102 may store the corresponding portion of the genomics data 404 and the corresponding portion of the health insurance claims data in association with the individual's identifier in an integrated data repository, such as the integrated data repository 104 of Figures 1, 2, and 3.

[0090] An implementation of the architecture 400 may implement a cryptographic protocol that allows de-identified information from disparate data repositories to be consolidated into a single data repository. In this manner, the security of the data stored by the consolidated data repository 104 is increased. Additionally, the cryptographic protocol implemented by the architecture 400 may allow for more efficient retrieval and accurate analysis of the information stored by the consolidated data repository 104 than in situations in which the cryptographic protocol of the architecture 400 is not utilized. For example, by generating a token file 420 including a first token 422 using cryptographic techniques based on a defined set of information stored by the molecular data repository 104, and utilizing a second token 428 generated using the same or similar cryptographic techniques for a similar or identical set of information stored by the health insurance claims data repository 106, the data integration and analysis system 102 may match information stored by the disparate data repositories corresponding to the same individual. Without implementing the cryptographic protocols of architecture 400, the probability of erroneously attributing information from a data repository to one or more individuals increases, which reduces the accuracy of results provided by the data integration and analysis system 102 in response to an integrated data repository request 142 sent to the data integration and analysis system 102.

[0091] FIG. 5 illustrates a framework 500 for generating a dataset by the data pipeline system 138 based on data stored by the integrated data repository 104, according to one or more implementations. The integrated data repository 104 may store health insurance claims data and genomics data for a group of individuals 502. For example, the integrated data repository 104 may store information obtained from health insurance claims records 504 for the group of individuals 502. For each individual included in the group of individuals 502, the integrated data repository 104 may store information obtained from multiple health insurance claims records 504. In various embodiments, the information stored by the integrated data repository 104 may include and / or be derived from thousands, tens of thousands, hundreds of thousands, or up to millions of health insurance claims records 504 for a number of individuals. In addition, each health insurance claim record may include multiple columns. As a result, the integrated data repository 104 may be generated through the analysis of millions of columns of health insurance claims data.

[0092] Further, while the health insurance claims data may be organized according to a structured data format, the health insurance claims data is typically arranged to be viewed by health insurance providers, patients, and health care providers to show financial and insurance code information related to the services provided by the health care provider to the individual. Thus, the health insurance claims data is not easily analyzed to obtain insights that may be available in relation to the characteristics of an individual for whom a biological condition exists and that may aid in the treatment of the individual for the biological condition. The integrated data repository 104 may be generated and organized by analyzing and modifying the raw health insurance claims data in a manner that allows the data stored by the integrated data repository 104 to be further analyzed to determine trends, characteristics, features, and / or insights regarding an individual for whom one or more biological conditions may exist. For example, health insurance codes may be stored within the integrated data repository 104 in such a way that at least one of a medical procedure, a biological condition, a treatment, a dosage, a drug manufacturer, a drug distributor, or a diagnosis may be determined for a given individual based on the health insurance claims data for the individual. In various embodiments, the data integration and analysis system 102 may generate and implement one or more tables showing correlations between health insurance claims data and various treatments, symptoms, or biological conditions corresponding to the health insurance claims data. Additionally, the integrated data repository 104 may be generated using the genomics data records 506 of a group of individuals 502. In various embodiments, large volumes of health insurance claims data may be matched with genomics data for a group of individuals 502 to generate the integrated data repository 104.

[0093] By integrating the genomics data records 506 for a group of individuals 502 with the health insurance claims records 504, the data integration and analysis system 102 may determine correlations between the presence of one or more biomarkers present in the genomics data records 506 and other characteristics of the individuals indicated by the health insurance claims data records 506 that existing systems typically cannot determine. For example, the data integration and analysis system 102 may determine one or more genomic characteristics of the individuals corresponding to treatments received by the individuals, the timing of the treatments, the dosage of the treatments, the individual's diagnosis, smoking status, the presence of one or more biological conditions, the presence of one or more symptoms of the biological conditions, one or more combinations thereof, and the like. Based on the correlations determined by the data integration and analysis system 102 using the integrated data repository 104, cohorts of individuals that may benefit from one or more treatments that would not have been identified in existing systems may be identified. In one or more embodiments, the processes and techniques implemented to integrate the health insurance claim records 504 and the genomics claim records 506 to generate the integrated data repository 104 may be complex, and efficiency-improving techniques, systems, and processes may be implemented to minimize the amount of computing resources used to generate the integrated data repository 104.

[0094] In one or more illustrative examples, the data pipeline system 138 may access information stored by the integrated data repository 104 and generate a data set including a number of additional data records 508 including information related to at least a portion of the group of individuals 502. In the illustrative example of FIG. 5, the additional data records 508 include information indicating whether the individual is included within a cohort of individuals in which lung cancer exists. The data pipeline system 138 may execute a number of different sets of data processing instructions to determine the cohort of groups of individuals 502 in which lung cancer exists. In various examples, the additional data records 508 may indicate information used to determine the status of the individual 502 with respect to lung cancer, such as one or more health insurance procedure identifiers, one or more International Classification of Diseases (ICD) codes, and one or more health insurance procedure dates. In addition to including a column indicating whether the individual 502 is included within a lung cancer cohort, the additional data records 508 may include a column indicating a confidence level of the individual's 502 status with respect to the presence of lung cancer.

[0095] 6A-6B illustrate a flowchart of an example process 600 associated with determining a line of therapy. In some implementations, one or more process blocks of FIGS. 6A-6B may be performed by a computing machine (e.g., computing device 900 of FIG. 8) including processing circuitry (e.g., processor 904) and memory (e.g., main memory 906, static memory 908, or storage device 918). In some implementations, one or more process blocks of FIGS. 6A-6B may be performed by another device or group of devices separate from or including the computing machine. Additionally or alternatively, one or more process blocks of Figures 6A-6B may be performed by one or more components of the computing device 900, such as the processor 904, the main memory 906, the static memory 908, the network interface device 922, the sensors 924, the display unit 912, the alphanumeric input device 914, the user interface (UI) navigation device 916, the storage device 918, the signal generating device 920, and the output controller 930.

[0096] As shown in FIG. 6A, the process 600 may include accessing, in the processing circuitry, a pharmacy procedure dataset from memory for a patient. Each pharmacy procedure in the pharmacy procedure dataset may comprise at least a procedure date, a therapy type, and a therapy delivery duration (block 605). For example, a computing machine may access, in the processing circuitry, a pharmacy procedure dataset from memory for a given patient. In one or more embodiments, the pharmacy procedure dataset may include one or more data tables including health records of a plurality of individuals. The health records may include medical information obtained in connection with visits by a plurality of individuals to one or more clinical settings. In at least some embodiments, the one or more clinical settings may include a facility used to conduct clinical trial studies. In one or more additional embodiments, the one or more clinical settings may include a facility used primarily to provide treatment to individuals diagnosed with, at risk for, or suspected of having a biological condition. In one or more illustrative examples, the pharmacy transaction dataset may include one or more data tables including health insurance claims data for a plurality of individuals. The health insurance claims data may include health insurance codes. The health insurance codes may indicate a treatment received by the patient. In one or more illustrative examples, the health insurance codes may include a drug treatment received by the patient. To illustrate, the health insurance codes may indicate a National Drug Code (NDC) that corresponds to the drug treatment received by the patient. In one or more additional illustrative examples, the pharmacy transaction dataset may include information obtained from one or more electronic medical records. The electronic medical records may include imaging information, laboratory test results, diagnostic test information, clinical observations, dental health information, health care provider notes, medical history forms, diagnostic request forms, medical procedure order forms, medical information charts, one or more combinations thereof, and the like.

[0097] In various examples, the drug treatment may include an ingestible form of drug substance provided to the patient in the treatment of a biological condition. For example, the drug treatment may include one or more tablets or other forms of drug substance that can be ingested by the patient by mouth. Additionally, the drug treatment may include an inhalable form of a therapeutic agent. In one or more examples, the pharmacy transaction data may indicate the number of times per day that the drug treatment should be provided to the patient. In one or more further examples, the pharmacy transaction data may indicate a dosage of the drug treatment, such as milligrams of the drug treatment. The therapy delivery duration may correspond to the amount of time that the patient should receive the drug treatment. To illustrate, the therapy delivery duration may indicate the number of days, weeks, or months that the drug treatment should be provided to the patient. In at least some examples, the therapy type may indicate that the drug treatment is an anti-neoplastic drug or an immunotherapy.

[0098] A procedure date included in a pharmacy procedure dataset may indicate the date a claim was paid for a given pharmacy procedure. For example, the procedure date may indicate the date a health insurance provider and / or a patient paid for a drug treatment for a biological condition. In one or more additional examples, the procedure date may correspond to the date a patient received the drug treatment. To illustrate, the procedure date may indicate the date a prescription for a drug treatment was filled or the date a quantity of the drug treatment was picked up by a patient from a drug treatment provider.

[0099] As further shown in FIG. 6A, the process 600 may include identifying a pharmacy procedure subset associated with the biological condition based on a therapy type from the pharmacy procedure dataset (block 610). For example, the computing machine may identify a pharmacy procedure subset associated with the biological condition based on a therapy type from the pharmacy procedure dataset as described above. In one or more illustrative examples, the pharmacy procedure subset may include pharmacy procedures corresponding to one or more categories of chemotherapy provided to the patient. In various examples, one or more therapies may be excluded from the pharmacy procedure dataset. In one or more examples, a patient being treated for a biological condition may receive a number of different types of treatments for the biological condition. To illustrate, a patient may receive one or more primary treatments meant to directly treat the biological condition and one or more additional treatments intended to indirectly treat the biological condition and / or treat side effects caused by the one or more primary treatments. In one or more illustrative examples, a patient being treated for cancer may receive one or more chemotherapy drug treatments and one or more steroids, such as one or more glucocorticoids. In these scenarios, the pharmacy procedure subset may include the one or more chemotherapy drug treatments and exclude the one or more steroids.

[0100] As further shown in FIG. 6A, the process 600 may include calculating an end date for at least one pharmacy procedure in the pharmacy procedure subset. The end date may be determined based on a procedure date and a therapy supply duration associated with the at least one pharmacy procedure (block 615). For example, the computing machine may calculate an end date for at least one pharmacy procedure in the pharmacy procedure subset. In one or more embodiments, the end date for the therapy treatment may be determined based on the procedure date indicated by the pharmacy procedure dataset plus the number of days of supply of the therapy treatment.

[0101] As further shown in FIG. 6A, the process 600 may include accessing, in the processing circuitry, from a memory, a medical procedure procedure dataset for the patient. Each medical procedure in the medical procedure dataset may include at least a medical procedure date range and a medical procedure type (block 620). For example, the computing machine may access, in the processing circuitry, from a memory, a medical procedure procedure dataset for the patient. The medical procedure procedure dataset may include one or more data tables indicating medical procedures obtained by the patient related to the treatment of the biological condition. In one or more embodiments, the medical procedure procedure dataset may include a number of health records of the patient. In one or more illustrative embodiments, the medical procedure procedure data may include health insurance claims data for the patient, including health insurance codes corresponding to the medical procedures. In various embodiments, the procedure procedure dataset may include one or more Health Care Common Procedure Coding System (HCPCS) codes corresponding to medical procedures obtained by the patient in the course of treatment for the biological condition. In one or more embodiments, the medical procedure obtained by the patient may include a procedure that provides one or more therapeutic agents to the patient. For example, the patient may receive one or more injections of one or more therapeutic agents to treat a biological condition. In addition, the patient may receive one or more intravenous infusions of one or more therapeutic agents.

[0102] A medical procedure date range included within the medical procedure procedure data may indicate a period over which the medical procedure was performed. In at least some instances, the medical procedure date range may indicate that the medical procedure was performed over a single day. In one or more additional examples, the medical procedure date range may indicate that the medical procedure was performed over a number of days. A medical procedure type included within the medical procedure procedure data may indicate a category associated with a given medical procedure. In one or more examples, a medical procedure type may correspond to a medical procedure performed by a physician, such as a surgical procedure, a medical device provided to a patient, a pathology service provided to a patient, a radiology service provided to a patient, administration of a therapeutic agent to a patient, etc.

[0103] As further shown in FIG. 6A, the process 600 may include identifying a medical procedure procedure subset associated with the biological condition based on the medical procedure type from the medical procedure procedure dataset (block 625). In one or more embodiments, the medical procedure procedure subset may include medical procedure procedures corresponding to one or more defined health insurance codes. In one or more illustrative embodiments, a patient being treated for cancer may have medical procedure procedures corresponding to radiological procedures, radiation procedures, surgical procedures, chemotherapy procedures, oral drug procedures, one or more combinations thereof, and the like. In various embodiments, to determine one or more lines of therapy to be provided to the patient, the medical procedure procedure dataset may be analyzed with respect to one or more criteria to determine medical procedure procedures corresponding to the one or more lines of therapy. In various embodiments, the medical procedure procedure subset may include medical procedure procedures corresponding to one or more health insurance codes corresponding to chemotherapy treatments and / or parenterally administered drugs, such as drugs delivered by injection or delivered intravenously. In addition, at least some health insurance codes may be excluded from the medical procedure procedure subset. To illustrate, when determining a line of therapy for a patient being treated for cancer, medical procedure procedures having health insurance codes corresponding to radiation therapy, surgery, and / or radiological services may be excluded.

[0104] As shown in FIG. 6B, the process 600 may include, for at least one medical procedure in the medical procedure procedure subset, adjusting the medical procedure date range based on a period during which the medical procedure type is effective or repeated (block 630). In one or more embodiments, the medical procedure may be associated with a period during which the medical procedure is effective in treating a biological condition. The period may be determined based on information obtained from an individual who has previously undergone the medical procedure. In addition, at least some medical procedures may be repeated over a period of time at a given frequency. The frequency at which the medical procedure is performed and the overall period during which the medical procedure is performed are based on one or more therapeutic agents delivered during the medical procedure. For example, in a scenario in which the medical procedure includes an injection of a therapeutic agent and / or an intravenous administration of a therapeutic agent, the medical procedure data range may correspond to a dosage of the therapeutic agent. In various embodiments, the medical procedure date range may correspond to a period between two instances in which the medical procedure is performed. In one or more additional embodiments, the medical procedure data range may correspond to a time period over which a series of instances of the medical procedure are administered.

[0105] As further shown in FIG. 6B, the process 600 may include mapping the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in memory, the timeline data structure storing the pharmacy procedure and the medical procedure procedures arranged by date (block 635). In one or more examples, the timeline data structure may include one or more data tables indicating one or more first time periods over which one or more pharmacy treatments are provided to the patient and one or more second time periods over which one or more medical treatments are provided to the patient. In one or more illustrative examples, the timeline data structure may indicate a first time period during which the patient received an initial supply of a first drug treatment and two refills of the first drug treatment. The timeline data structure may also indicate a second time period during which the patient received an initial supply of a second drug treatment and one refill. In one or more further embodiments, the timeline data structure may indicate a third time period over which the medical procedure was performed on the patient, hi at least some embodiments, at least one of the first time period, the second time period, or the third time period overlap.

[0106] As further shown in FIG. 6B, process 600 may include determining one or more therapy gaps within the timeline data structure during which there are no pharmacy procedures and there are no medical procedure procedures, each therapy gap comprising a number of consecutive days, the number of consecutive days being greater than a threshold number of days (block 640). In one or more examples, a therapy gap for a therapeutic agent may include the number of days between the end of a supply of the therapeutic agent and the time when a new supply of the therapeutic agent is obtained by or prescribed for the patient. In one or more additional examples, a therapy gap for a medical procedure may correspond to the number of days between a first instance in which the medical procedure is performed and a second instance in which the medical procedure is performed. In at least some examples, a therapy gap may be identified without being able to determine the number of days of the therapy gap. In these scenarios, an imputed therapy gap may be determined based on therapy gaps for previous patients who obtained the same therapy. In one or more illustrative examples, the imputed therapy gap corresponds to a median therapy gap for patients who previously obtained the same treatment. In one or more additional examples, the imputed therapy gap may be determined based on the therapy gap of patients who previously received a number of different therapies. To illustrate, health insurance data for at least a subset of individuals who have received a number of different therapies may be analyzed to determine a median therapy gap for patients who previously received a number of different therapies. In various examples, the number of different therapies may be related to the treatment of a given biological condition. To illustrate, the number of different therapies may be related to the treatment of one or more cancers.

[0107] As further shown in FIG. 6B, the process 600 may include determining one or more lines of therapy based on one or more therapy gaps, each line of therapy comprising a pharmacy procedure and a medical procedure procedure occurring either between two therapy gaps, before the earliest temporary therapy gap, or after the latest temporary therapy gap, each line of therapy being associated with a line date range (block 645). In one or more embodiments, the individual lines of therapy may indicate one or more time periods during which a drug treatment was obtained by the patient. In at least some embodiments, the drug treatment may be obtained over a first time period and a second time period with a gap between the first time period and the second time period being less than a threshold gap. The individual lines of therapy may also include medical procedures provided to the patient over one or more time periods. The one or more time periods during which the medical procedures are performed on the patient may overlap with the time periods during which the drug treatment is provided to the patient. In one or more additional embodiments, the one or more time periods during which the medical procedures are performed on the patient may be within a treatment interval of the drug treatment. In one or more illustrative examples, a patient being treated for a biological condition may be provided with a line of therapy that includes at least one of an ingestible drug therapy or an inhalable drug therapy in addition to at least one of an injection of a therapeutic agent or an intravenous administration of a therapeutic agent.

[0108] In one or more examples, the patient may receive multiple lines of therapy. The additional lines of therapy may be identified in response to determining a change in therapy provided to the individual after the first line of therapy is terminated. In various examples, the end of a line of therapy may be identified based on determining that a therapy different from the initial therapy has been provided to the patient more than a threshold therapy gap after the last therapy of the initial line of therapy. For example, the patient may receive a first drug therapy as part of the first line of therapy. The patient may receive one or more doses of the first drug therapy. In at least some examples, there may be a therapy gap between two doses of the first drug therapy that is less than a threshold therapy gap. After more than a threshold amount of time since the last dose of the first drug therapy, the patient may receive a second drug therapy different from the first drug therapy. The patient may receive one or more doses of the second drug therapy. In one or more illustrative examples, the patient may receive a medical procedure as part of the first line of therapy or as part of the second line of therapy.

[0109] As further shown in FIG. 6B, process 600 may include transmitting a data structure identifying the patient, one or more lines of therapy, and line date ranges for one or more lines of therapy to a data repository for storage therein (block 650). In one or more embodiments, the data repository may store lines of therapy for a number of patients. In one or more illustrative embodiments, the data repository may store lines of therapy for hundreds of patients, thousands of patients, up to tens of thousands of patients or more. In various embodiments, the lines of therapy for one or more cohorts of patients may be analyzed to determine the effectiveness of treatment for the cohort of individuals. In one or more additional embodiments, the lines of therapy for one or more cohorts may be analyzed in conjunction with genomic data for individuals to determine the effectiveness of treatment for patients with one or more genomic mutations. In one or more further embodiments, the effectiveness of treatments for patients who have previously undergone pharmacy treatments and / or medical procedures for treating a biological condition may be analyzed in conjunction with the patient's genomic profile to determine a treatment for a patient newly diagnosed with a biological condition based on the new patient's genomic profile.

[0110] Process 600 may include additional implementations, such as any single implementation or any combination of implementations described below and / or in conjunction with one or more other processes described elsewhere herein.

[0111] In some implementations, the process 600 includes determining, within a single line of therapy, a first pharmacy procedure or medical procedure associated with a first biological condition stage and a second pharmacy procedure or medical procedure associated with a second biological condition stage, where a start date associated with the second pharmacy procedure or medical procedure is later than a start date associated with the first pharmacy procedure or medical procedure, and splitting the single line of therapy into two lines of therapy using the start date associated with the second pharmacy procedure.

[0112] In some implementations, the processing circuitry comprises a plurality of multi-threaded graphic processing units (GPUs). The patient is one of a plurality of patients. One or more lines of therapy for the patient are determined in parallel with determining lines of therapy for other patients from the plurality of patients using parallel threads of the plurality of multi-threaded GPUs.

[0113] In some implementations, determining a line of therapy for a plurality of patients includes generating a plurality of intermediate tables, each of which is stored in a data repository for reviewing and adjusting the performance of one or more computing machines. In one or more implementations, the plurality of intermediate tables may increase the speed of calculations associated with determining a line of therapy for a plurality of patients by storing intermediate calculation results.

[0114] For at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range includes a single date. The at least one medical procedure procedure is mapped onto the timeline data structure based on the single date. In one or more additional embodiments, for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range includes a medical procedure start date and a medical procedure end date. The at least one medical procedure procedure is mapped onto the timeline data structure based on the medical procedure start date and the medical procedure end date. Additionally, a pharmacy procedure from the pharmacy procedure subset is mapped onto the timeline data structure based on a procedure date. The at least one pharmacy procedure is mapped onto the timeline data structure based on a procedure date and an end date. The threshold number of days may be determined based on a condition type of the biological condition.

[0115] In some implementations, the therapy type is a National Drug Code (NDC) classification. The pharmacy procedure dataset comprises one or more tables. Process 600 further includes determining a set of columns in the one or more tables associated with drugs, parsing the set of columns to identify NDC classifications, determining a set of NDC classifications corresponding to the drugs, and identifying a subset of the set of NDC classifications corresponding to the drugs. The subset is associated with drugs associated with the biological condition. Process 600 further includes identifying rows in the one or more tables for placement into the pharmacy procedure subset based on a group of NDC classifications associated with the biological condition.

[0116] 6A-6B show example blocks of process 600, in some implementations process 600 may include additional blocks, fewer blocks, different blocks, or blocks that are arranged differently than depicted in FIGs. Additionally or alternatively, two or more of the blocks of process 600 may be performed in parallel.

[0117] 7A-7B are data flow diagrams of an example process 700 for determining a line of therapy from pharmacy and medical procedure data, according to one or more implementations. Process 700 may be implemented in a computing machine (e.g., computing device 900) that includes processing circuitry and memory.

[0118] As shown in FIG. 7A, in process 700, a computing machine accesses a pharmacy procedure (txn) dataset 710 and a medical procedure procedure dataset 720 for a patient with a biological condition (e.g., based on a formal or informal diagnosis or machine prediction). The pharmacy procedure dataset 710 and / or the medical procedure procedure dataset 720 may reside in the health insurance claims data repository 106, shown in FIG. 1. Alternatively, the pharmacy procedure dataset 710 and the medical procedure procedure dataset 720 may reside in separate data repositories. In some cases, data from the health insurance claims data repository 106 is provided to the computing machine in an anonymized format. Active consent from the patient is always obtained before sharing any patient data that is not anonymized or that may be used to identify an individual patient with their own medical data.

[0119] In block 730, the computing machine filters data related to the biological condition from the pharmacy procedure dataset 710 and the medical procedure procedure dataset 720, respectively, resulting in a pharmacy procedure data subset 740 and a medical procedure procedure data subset 750. In the filtering of block 730, procedures not related to the biological condition are removed, resulting in the subsets of procedures related to the biological condition in blocks 740 and 750. The filtering of block 730 may be based on drug codes (e.g., NDC classifications) or procedure codes (e.g., Healthcare Common Procedure Coding System (HCPCS) classifications) that are known to be related or not related to the biological condition. The known codes may be stored in a table or list data structure or data repository referenced by the computing machine. For example, if the biological condition is lung cancer, the filtering of block 730 may cause pharmacy procedures from the pharmacy procedure dataset 710 that represent purchases of lung cancer drugs to be placed in the pharmacy procedure data subset 740. However, a pharmacy transaction from the pharmacy transaction dataset 710 that represents influenza vaccination may not be placed in the pharmacy transaction data subset 740 (assuming influenza vaccination is not associated with lung cancer).

[0120] According to some embodiments, a data repository accessible to a computing machine stores a list (or other data structure) of NDC or HCPCS codes associated with different biological conditions (e.g., lung cancer, breast cancer, liver cancer, and the like) or a set of biological conditions (e.g., all cancers). The stored list of codes may be determined based on one or more of RxNorm and National Comprehensive Cancer Network (NCCN) Clinical Practice Guidelines for Oncology. A person (or an artificial intelligence engine) may review the NCCN guidelines and identify NDC or HCPCS codes associated with different biological conditions or a set of biological conditions.

[0121] 7A, a timeline data structure 760 is generated by a computing machine from the pharmacy procedure data subset 740 and / or the medical procedure procedure data subset 750. (Note that the timeline depicted in the timeline data structure 760 is not drawn to scale.) As shown, the timeline data structure 760 indicates that the patient took Drug A from 09-06-97 (dates in month-day-year format) to 11-15-97. This may be determined, for example, because the patient purchased a sufficient supply of Drug A on 09-06-97 to last until 11-15-97. The timeline data structure 760 also indicates that the patient took Drug B from 11-01-97 to 12-12-97. There is a gap from 12-12-97 to 06-15-98 where no pharmacy or medical procedure procedures related to the biological condition were recorded. The patient underwent medical procedure C (Med. Proc. C) from 06-15-98 to 07-01-98, and the patient took medication D from 06-17-98 to 07-22-98.

[0122] Based on the timeline data structure 760, a line of therapy 770 is identified, as shown in Figure 7B. The line of therapy 770 and / or the timeline data structure 760 may be provided as a visual output via a graphical user interface (GUI) on a display coupled to a computing machine. Alternatively, the line of therapy 770 and / or the timeline data structure 760 may be transmitted to one or more data repositories (e.g., the integrated data repository 104) for storage therein.

[0123] As shown, each therapy line from among therapy lines 770 is associated with a start date, an end date, and a therapy (e.g., a medication, a medical procedure visit, and a medical procedure). As shown, therapy line 771 has a start date of 09-06-97 and an end date of 11-01-97 and includes medication A. Therapy line 772 has a start date of 11-01-97 and an end date of 12-12-97 and includes medication B. Therapy line 773 has a start date of 06-15-98 and an end date of 07-22-98 and includes medication C and medication D.

[0124] The computing machine has determined that therapy lines 771 / 772 and 773 are distinct because the gap between 12-12-97 and 06-15-98 exceeds a predetermined minimum gap size (e.g., 30 or 60 days). The predetermined minimum gap size may be set by the researcher depending on the amount of confidence the researcher requires to determine that different therapies are associated with different lines. For example, if a high amount of confidence is desired, a minimum gap size of 90 days may be set. If a low amount of confidence is acceptable, a minimum gap size of 7 days may be set.

[0125] However, note that there is no gap between the therapy line 771 representing drug A and the therapy line 772 representing drug B. In fact, as shown in the timeline data structure 760, drug A and drug B would overlap from 11-01-97 to 11-15-97. The distinction between the therapy line 771 for drug A and the therapy line 772 for drug B may be made because drug A and drug B may be associated with different stages of a biological condition. For example, drug A may be associated with stage II cancer and drug B may be associated with stage III cancer. Based on this intelligence, which may be stored in the memory of the computing machine or in a data repository coupled to the computing machine, the computing machine may determine that on 11-01-97, it was determined (e.g., by a medical professional) that the patient's cancer had progressed from stage II to stage III and therefore the patient was switched from drug A to drug B. It is likely that the patient did not take any more drug A after 11-01-97 because it was not useful in treating the current stage of the cancer. (The patient obtained a supply of drug A (e.g., from a pharmacy) from 11-01-97 to 11-15-97, but the patient did not use this supply.) Based on this information, the computing machine assigned dates of 09-06-97 to 11-01-97 to therapy line 771 associated with drug A. The computing machine assigned dates of 11-01-97 to 12-12-97 to therapy line 772 associated with drug B (not drug A).

[0126] In some embodiments, the computing machine first uses the gap between 12-12-97 and 06-15-98 in the timeline data structure 760 to separate therapy line 771 / 772 and therapy line 773. The therapy line 771 is then identified as distinct from therapy line 772 by the computing machine, which determines that drug A and drug B are associated with different cancer stages. The start date for drug A of 09-06-97 is earlier than the start date for drug B of 11-01-97. Thus, therapy line 771 for drug A is assigned dates from 09-06-97 to 11-01-97. The distinct therapy line 772 for only drug B is assigned dates from the start date for drug B (11-01-97) to the end date for drug B (12-12-97).

[0127] In some cases, intermediate therapy lines (e.g., combinations of therapy lines 771 / 772) are stored in the intermediate table. A user of the computing machine may review the generated intermediate table to verify that process 700 is functioning correctly and / or to modify thresholds or intelligence used in process 700. For example, if a user determines that drug A and drug B could have been taken together (based on medical knowledge that existed in 1997), therapy lines 771 / 772 may be combined into a single therapy line with both drug A and drug B. Alternatively, if the minimum gap size is set to 240 days, therapy lines 772 / 773 may be combined into a single therapy line from 11-01-97 to 07-22-98, including drug B, medical procedure C, and drug D. The gap shown in timeline data structure 760 between 12-12-97 and 06-15-98 is 185 days long.

[0128] In addition to the above, other data may also be useful in identifying a line of therapy. For example, if a patient's pharmacy procedure data subset 740 indicates that the patient purchased a 30-day supply of drug E on 02-01-01, drug E is for stage III cancer, and the patient was hospitalized (based on medical procedure procedure data subset 750) from 02-15-01 to 04-15-01, the hospitalization may be assigned a different line of therapy than the use of drug E because drug E was not provided to the patient at the hospital. In other words, there would be a first line of therapy for drug E from 02-01-01 to 02-15-01 and a second line of therapy for the hospitalization from 02-15-01 to 04-15-01.

[0129] In addition to providing useful information for a user of a computing machine who may wish to adjust thresholds or intelligence used in process 700, the intermediate tables may increase the calculation speed of determining lines of therapy for multiple patients in parallel by storing intermediate calculation results. In some embodiments, the intermediate tables may store intermediate results of a calculation. These intermediate results may be used for different types of calculations, such as statistical or machine learning based calculations. Storing data in the intermediate tables may increase the calculation speed because the stored data may accelerate statistical or machine learning based calculations.

[0130] Some embodiments are described herein in conjunction with a biological condition that is cancer or lung cancer, however, it should be noted that the techniques described herein may be used in conjunction with any other biological condition, such as Alzheimer's disease, arthritis, influenza, and the like.

[0131] 7C illustrates an example system 700C for determining a line of therapy according to one or more implementations. As shown in FIG. 7C, the health insurance claims data repository 106 includes a plurality of pharmacy procedure datasets 710 for a plurality of patients and a plurality of medical procedure procedure datasets 720. A computing machine 790 accesses the pharmacy procedure datasets 710 and the medical procedure procedure datasets 720 for a plurality of patients. The computing machine 790 may include all or a portion of the components of a computing device 900.

[0132] As shown, the computing machine 790 includes a multi-threaded processing circuitry 792 and a memory 794. The multi-threaded processing circuitry 792 may include one or more multi-threaded graphic processing units (GPUs) and / or one or more multi-threaded central processing units (CPUs). The multi-threaded processing circuitry 792 processes multiple pharmacy procedure data sets 710 and multiple medical procedure procedure data sets 720 for multiple patients in parallel (e.g., as described in conjunction with FIGS. 6A-6B and 7A-7B) by executing instructions from the memory 794. The computing machine 790 outputs a generated timeline 760 and a line of therapy 770 for multiple patients. The generated timeline 760 and / or line of therapy 770 may be stored in the integrated data repository 104 as shown in FIG. 1. Alternatively, these data structures may be stored in another data repository and / or provided as a visual output to a user of the computing machine 790.

[0133] Using the architecture of system 700C, health insurance claims data for multiple patients can be processed in parallel, greatly increasing the speed at which outputs (e.g., timeline 760 and / or therapy line 770) are generated.

[0134] FIG. 8 illustrates a computing architecture 800 having one or more systems for generating a line of therapy that can be analyzed to determine an outcome for a patient, according to one or more implementations. The architecture 800 may include the data integration and analysis system 102 and the integrated data repository 104 described with respect to FIGS. 1-5. In addition, the data integration and analysis system 102 may include a line of therapy system 802. The line of therapy system 802 may analyze data obtained from the integrated data repository 104 and determine a line of therapy for the one or more patients. The line of therapy may refer to one or more treatments provided to a patient to treat a biological condition. The one or more treatments may include at least one of one or more medical procedures or one or more therapeutic substances. In one or more illustrative examples, the one or more therapeutic substances may include one or more drug substances. In one or more additional illustrative embodiments, the one or more medical procedures may correspond to administration of one or more drug substances.

[0135] The therapy line system 802 may include a procedure procedure system 804 and a pharmacy procedure system 806. The procedure procedure system 804 may analyze health procedure information related to medical procedures and generate one or more procedure procedure data tables 808. In addition, the pharmacy procedure system 806 may analyze health procedure information related to therapeutic substances provided to one or more patients and generate one or more pharmacy procedure data tables 810. At least one of the one or more procedure procedure data tables 808 or the pharmacy procedure data tables 810 may be analyzed to determine one or more lines of therapy for the one or more patients. In one or more embodiments, the health procedure information analyzed by the procedure procedure system 804 and by the pharmacy procedure system 806 may include at least one of health insurance claims data or electronic medical records. In various embodiments, the electronic medical records may include at least one of structured data or unstructured data. In one or more illustrative embodiments, the electronic medical record may include imaging information, laboratory test results, diagnostic test information, clinical observations, dental health information, health care practitioner notes, medical history forms, diagnosis request forms, medical procedure order forms, medical information charts, one or more combinations thereof, etc. In one or more additional embodiments, the health procedure information analyzed by at least one of the procedure procedure system 804 or the pharmacy procedure system 806 may be obtained from the patient's participation in one or more clinical trials.

[0136] In various embodiments, the procedure system 804 may analyze medical procedure records 812 retrieved from the integrated data repository 104 and generate a procedure procedure data table 808. The medical procedure records may be stored by a medical procedure data table stored by the integrated data repository 104. The medical procedure records 812 may include information corresponding to a number of health insurance procedures for a number of patients associated with medical procedures obtained by the patients. In one or more embodiments, the medical procedure records 812 may indicate medical procedures provided to a patient in at least one of a clinical healthcare setting or as part of a clinical trial.

[0137] Additionally, the pharmacy transaction system 806 may analyze pharmacy records 814 retrieved from the integrated data repository 104 to generate a pharmacy transaction data table 810. The pharmacy records 814 may be stored by a pharmacy record data table stored by the integrated data repository 104. The pharmacy records 814 may include information corresponding to a number of health insurance transactions for a number of patients related to therapeutic substances obtained by the patients. In various embodiments, the pharmacy records 814 may indicate drug substances provided to the patient in at least one of a clinical healthcare setting or as part of a clinical trial.

[0138] In addition, the line of therapy system 802 may analyze line of therapy analysis information 816 and generate one or more lines of therapy for one or more patients. At least a portion of the line of therapy analysis information 816 may be obtained from one or more third party sources. At least a portion of the line of therapy analysis information 816 may also be generated using one or more computational techniques. In one or more embodiments, the line of therapy analysis information 816 may be generated using one or more machine learning techniques. Furthermore, at least a portion of the line of therapy analysis information 816 may be obtained from a service provider that controls, maintains, manages, or creates at least one of the data integration and analysis system 102 and the integrated data repository 104. For example, at least a portion of the line of therapy analysis information 816 may include information curated by the service provider.

[0139] The therapy line analysis information 816 may include health insurance claims data 818. The health insurance claims data 818 may correspond to one or more therapeutic substances that may be used to treat one or more biological conditions. In addition, the health insurance claims data 818 may correspond to one or more medical procedures that may be used to treat one or more biological conditions. In various examples, the health insurance claims data 818 may include a set of health insurance codes that correspond to at least one of the therapeutic substances or medical procedures used to treat the biological conditions. In one or more examples, the health insurance claims data 818 may indicate one or more health insurance codes that should be used to identify procedures to be included in at least one of the medical procedure record 812 or the pharmacy record 814. In one or more additional examples, the health insurance claims data 818 may indicate one or more additional health insurance codes that should be used to exclude procedures stored in the integrated data repository 104 from the medical procedure record 812 and / or the pharmacy record 814. In one or more illustrative examples, the health insurance claims data 818 may include one or more NDC identifiers. In one or more additional illustrative examples, the health insurance claims data 818 may include one or more HCPCS codes. In one or more further examples, the health insurance claims data 818 may indicate one or more columns of one or more database tables stored by the integrated data repository 104 to be analyzed to determine the presence or absence of one or more health insurance codes.

[0140] In at least some embodiments, the health insurance claims data 818 may include a number of groups of health insurance codes. Each group of health insurance codes may correspond to a different biological condition. In various illustrative embodiments, the health insurance claims data 818 may include one or more health insurance codes corresponding to treatments provided to the patient during treatment of a defined biological condition. To illustrate, the health insurance claims data 818 may show a first set of health insurance codes corresponding to treatments provided to the patient in association with Alzheimer's disease and a second set of health insurance codes corresponding to treatments provided to the patient in association with type II diabetes. In one or more additional embodiments, the health insurance claims data 818 may include health insurance codes corresponding to different forms of biological conditions, such as different forms of cancer. For example, the health insurance claims data 818 may include a first set of health insurance codes corresponding to treatments provided to the patient in association with colon cancer and a second set of health insurance codes corresponding to treatments provided to the patient in association with breast cancer.

[0141] The therapy line analysis information 816 may also include treatment data 820. The treatment data 820 may include a list of treatments to be included within the health procedure information, comprising at least one of the medical procedure record 812 or the pharmacy record 814. In one or more additional examples, the treatment data 820 may include an additional list of treatments used to exclude from the health procedure information from at least one of the medical procedure record 812 or the pharmacy record 814. Furthermore, the treatment data 820 may indicate at least one of a primary treatment for the one or more biological conditions, a secondary treatment for the one or more biological conditions, or a tertiary treatment for the one or more biological conditions. In one or more illustrative examples, the primary treatment for the biological condition may indicate one or more first treatments to be provided to a patient diagnosed with or suspected of having a biological condition. In addition, a second line treatment for a biological condition may refer to one or more second treatments to be provided to a patient diagnosed with or suspected of having a biological condition when one or more first treatments are ineffective. Further, a third line treatment for a biological condition may refer to one or more third treatments to be provided to a patient diagnosed with or suspected of having a biological condition when one or more first treatments and one or more second treatments are ineffective.

[0142] In one or more additional examples, the line of therapy analysis information 816 may include data analysis rules and schema 822. The data analysis rules and schema 822 may provide a framework by which the line of therapy system 802 may analyze the medical procedure records 812 and the pharmacy records 814. For example, the data analysis rules and schema 822 may include a framework for determining a period during which a patient has taken a drug substance when a first date of supply of the drug substance overlaps with a second date of supply of the drug substance. In one or more additional examples, the data analysis rules and schema 822 may include a framework for determining a period during which a patient has received a drug substance via a medical procedure, such as via one or more injections or one or more intravenous administrations of a drug substance, when health insurance claims for the medical procedures overlap. In one or more further examples, the data analysis rules and schema 822 includes a framework for determining a length of treatment for patients who die before a next treatment is obtained or who die before a threshold period after the last treatment is obtained. The data analysis rules and schema 822 may also include a framework for determining the date of the next treatment for a patient. To illustrate, the time to next treatment may be censored at the date of death of the patient even when the patient's last active date is after the date of death. That is, the health insurance claim date may be after the patient's date of death and may indicate a real-world time to next treatment that is after the date of death. In these scenarios, the real-world time to next treatment is determined to be the date of death, not the date of the last active health insurance claim.

[0143] In one or more illustrative examples, the framework for determining the real-world time to next treatment may depend on whether the date of the patient's last activity is less than a threshold amount of time. In one or more examples, the threshold amount of time is at least 30 days, at least 45 days, at least 60 days, at least 75 days, at least 90 days, at least 105 days, at least 120 days, at least 135 days, at least 150 days, at least 175 days, or at least 190 days. In at least some examples, in situations where a patient receives treatment after the threshold period, it may be determined that the next treatment is part of a new line of therapy, rather than a continuation of the current line of therapy. In various examples, the threshold period may be determined based on an analysis of health insurance data of a number of patients who have previously received one or more treatments for a biological condition.

[0144] In one or more additional illustrative examples, the line of therapy analysis information 816 may include one or more rules for determining a patient's date of death. For example, in the situation of a patient where the date of death occurs before the start of treatment, the patient is excluded from the procedure procedure data table 808 and the pharmacy procedure data table 810. Additionally, in various examples, the integrated data repository 104 may indicate the month and year of death rather than the date of death. In one or more examples, the date of death may be determined as one day after the sample collection date for a patient who died in the same month that the sample was collected. The sample may be collected in conjunction with a diagnostic procedure, such as one or more liquid biopsy procedures. In one or more additional examples, the date of death of a patient may be estimated as a defined date of the month that the patient died, such as the first day of the month, the fifth day of the month, the tenth day of the month, or the fifteenth day of the month.

[0145] In one or more further illustrative examples, the line of therapy analysis information 816 may indicate that the real-world time to next treatment is defined as the time from the start of the first line of therapy to the start of the next line of therapy. Different biological conditions may have different rules or frameworks for determining the line of therapy. In one or more examples, a new line of therapy for a first biological condition, such as a first form of cancer, may be defined as a change or addition within a non-biologic drug category that indicates a new line of therapy if it is outside of a first threshold time window since the start of the initial line of therapy, the first threshold time window being at least 5 days, at least 10 days, at least 15 days, at least 20 days, at least 30 days, at least 45 days, or at least 60 days. Additionally, a new line of therapy for the first biological condition may be initiated in response to a change or addition in a biologic or immune checkpoint inhibitor (ICI) drug if it is outside a second threshold time window since the start of the initial line of therapy, the second threshold time window being different from the first threshold time window and being at least 15 days, at least 30 days, at least 45 days, at least 60 days, at least 75 days, or at least 90 days. Additionally, another rule regarding the line of therapy for the first biological condition may indicate that adding or removing chemotherapy is not compatible with a new line of therapy. In one or more embodiments, the first biological condition may include colon cancer, the chemotherapy may include fluorouracil, leucovorin, or levoleucovorin, the non-biologic may include fluorouracil, capecitabine, irinotecan, oxaliplatin, leucovorin, levoleucovorin, and the biologic may include at least a portion of other colon cancer treatments under the National Comprehensive Cancer Network (NCCN) guidelines that are not included within the non-biologic list.In yet an additional embodiment relating to the first biological condition, a treatment episode occurring after a third threshold time window is considered a new line of therapy, and the third threshold time window differs from the first threshold time window, the second threshold time window, and is at least 120 days, at least 135 days, at least 150 days, at least 165 days, at least 180 days, at least 195 days, at least 210 days, at least 225 days, or at least 240 days.

[0146] In addition, the line of therapy analysis information 816 may include additional rules of the framework for determining a new line of therapy for a second biological condition, such as a second form of cancer. In one or more embodiments, the additional rules may indicate that termination of the first line of therapy occurs when a first period of time has passed and treatment with the same drug treatment is resumed or treatment with an additional drug occurs after a second period of time after the first line of therapy has begun. The first period of time may be at least 20 days, at least 25 days, at least 30 days, at least 35 days, at least 40 days, at least 45 days, or at least 60 days. The second period of time may be different from the first period of time and may be at least 30 days, at least 35 days, at least 40 days, at least 45 days, at least 50 days, at least 60 days, at least 70 days, at least 80 days, or at least 90 days. Continuation of the same drug substance within the first period of time may indicate that the first line of therapy is continuing and the second line of therapy has not yet begun. In various examples, when the second biological condition is non-small cell lung cancer, the line of therapy will not be changed from the first line of therapy to the second line of therapy if cisplatin and carboplatin are substituted for each other, paclitaxel is substituted for nab-paclitaxel, or bevacizumab is added to the chemotherapy.

[0147] The procedure procedure system 804 may generate a procedure data table 824 included in the procedure procedure data table 808. The procedure data table 824 may indicate medical procedures provided to a patient. In one or more embodiments, the procedure data table 824 may indicate a number of procedures provided to a patient for a given biological condition. For example, the procedure data table 824 may indicate medical procedures provided to a patient to treat diabetes or to treat a form of cancer. In one or more illustrative embodiments, the procedure data table may indicate one or more medical procedures that included administration of a drug substance to a patient. In various embodiments, the procedure procedure system 804 may analyze one or more columns of a data table stored by the integrated data repository 104 that includes health insurance claims data for one or more criteria. For example, the procedure procedure system 804 may analyze one or more columns of a data table stored by the integrated data repository 104 to determine one or more columns of the data table that correspond to the one or more criteria. The one or more criteria may correspond to one or more columns including at least one of one or more NDC identifiers or one or more HCPCS codes. The procedure procedure system 804 may determine that individual rows of one or more data tables stored by the integrated data repository 104 correspond to health insurance procedures related to the treatment of a biological condition. At least a portion of the information included in the one or more rows may be included in the medical procedure record 812 and stored by the procedure data table 824. The information included in the one or more rows may be stored in the procedure data table 824 in association with an identifier of the patient who underwent the medical procedure.

[0148] In one or more illustrative examples, the procedure procedure system 804 may analyze a set of columns of one or more data tables stored by the integrated data repository 104, where the set of columns is defined in the health insurance claims data 818, to determine one or more rows of the one or more data tables that include a number of HCPCS codes included in the health insurance claims data 818 and that correspond to a given biological condition. The procedure procedure system 804 may extract information from the one or more rows to generate a row of a procedure data table 824 included in the procedure procedure data table 808. In a scenario where an HCPCS code included in a row of a data table stored by the integrated data repository 104 is not included in the health insurance claims data 818 for the biological condition, the procedure procedure system 804 may analyze the row and determine if the row includes an NDC identifier included in the line of therapy analysis information 816. In situations where a row corresponding to a medical procedure includes an HCPCS code that is not included in the health insurance claims data 818, but does not include an NDC identifier that is included in the health insurance claims data 818, the procedure procedures system 804 may extract the information from the row and include the information in the procedure data table 824. Thus, the NDC identifier may refer to a drug substance provided to treat a biological condition, but because the drug substance was administered as part of the medical procedure, the information from the corresponding row in the data table may be stored in conjunction with the procedure data table 824.

[0149] An individual row of procedure data table 824 may include information related to a medical procedure obtained by a patient, such as at least one of a patient identifier, one or more sources of the medical procedure (e.g., an identifier of a database table stored by the integrated data repository 104 that provided information for the row), one or more identifiers of the medical procedure, one or more classes of the medical procedure, one or more categories of the medical procedure, one or more dates of the medical procedure, one or more HSPCS codes associated with the medical procedure, one or more NDC identifiers associated with the medical procedure, or one or more locations where the medical procedure was administered. In situations where a drug substance was provided as part of a medical procedure, a row of procedure data table 824 may include at least one of one or more names of the drug substance, an indicator of whether the drug substance is a generic version, or one or more dosages of the drug substance.

[0150] The procedure procedure system 804 may also analyze the medical procedure records 812 and determine one or more treatment gap tables included in the procedure procedure data table 808. The treatment gap tables may include a first treatment gap data table 826 indicating the time period between two consecutive instances in which a medical procedure is administered for a patient having the same HCPCS code, the same procedure name, and the same dosage. In one or more embodiments, the first treatment gap data table 826 may include multiple rows corresponding to a single treatment when multiple individuals underwent the same medical procedure. In this manner, each row associated with a treatment indicates the time period between consecutive instances in which the medical procedure is administered for multiple patients undergoing the medical procedure.

[0151] In one or more additional embodiments, the procedure procedure system 804 may generate a second treatment gap data table 828 based on the information included in the first treatment gap table. In various embodiments, the second treatment gap data table 828 may indicate at least one of a median or mean gap between consecutive instances in which a medical procedure is administered. In these scenarios, the procedure procedure system 804 may analyze multiple rows of the first treatment gap table 826 corresponding to individual medical procedures and determine at least one of a median or mean gap for the medical procedure. In this manner, the second treatment gap data table 828 may include individual rows for each individual treatment indicating at least one of a median or mean gap between consecutive instances in which the medical procedure is administered to a patient. In one or more illustrative embodiments, in scenarios in which the treatment gap for a patient is unknown, the median or mean treatment gap indicated by the second treatment gap data table 828 may be used to indicate the treatment gap for the patient. Additionally, in situations where the treatment gap for a patient exceeds a threshold amount of time or the treatment gap is outside a threshold number of standard deviations of the mean treatment gap, the median treatment gap or mean treatment gap stored by the second treatment gap data table 828 may be substituted for the actual treatment gap for the patient.

[0152] The procedure procedure system 804 may utilize information stored by the procedure data table 824 and at least one of the first treatment gap data table 826 or the second treatment gap data table 828 to generate a procedure episode data table 830. The procedure episode data table 830 may include a number of rows corresponding to individual patients who have undergone one or more medical procedures during the course of treatment for a biological condition. Each row of the procedure episode data table 830 may indicate one or more instances in which a medical procedure is obtained by a patient and a time period over which the one or more instances of the medical procedure were administered. In situations in which a medical procedure involved administration of a drug substance, each row of the procedure episode data table 830 may indicate at least one of a drug substance name, a drug substance category, a drug substance class, or a drug substance dosage. In various embodiments, the procedure episode data table 830 may indicate multiple lines of therapy obtained by a patient with a medical procedure.

[0153] The pharmacy transaction system 806 may generate a pharmacy data table 832 included in the pharmacy transaction data table 810. The pharmacy transaction system 806 may generate the pharmacy data table 832 by analyzing information stored by the integrated data repository 104 and included in the pharmacy record 804. In one or more embodiments, the pharmacy transaction system 806 may analyze one or more columns of one or more data tables stored by the integrated data repository 104. The one or more columns analyzed by the pharmacy transaction system 806 may be identified within the health insurance claims data 818. In various embodiments, the pharmacy transaction system 806 may analyze one or more columns of a data table stored by the integrated data repository 104 and identify rows of the data table that include one or more NDC identifiers included within the health insurance claims data 818. In one or more illustrative embodiments, the one or more NDC identifiers may correspond to a drug substance provided to a patient being treated for a given biological condition. In one or more additional embodiments, the pharmacy transaction system 806 may also analyze one or more rows of one or more data tables stored by the integrated data repository 104 to determine at least one row of the one or more data tables that includes a drug substance name that corresponds to a drug substance name included in the therapy line analysis information 816.

[0154] The pharmacy transaction system 806 may also analyze individual rows of one or more data tables stored by the integrated data repository 104 to determine at least one of a days supply of a drug substance provided to the patient, a category of the drug substance, a name of the drug substance provided to the patient, or a date of service indicating the date the drug substance was provided to the patient. Based on the analysis by the pharmacy transaction system 806 of the information stored by the one or more data tables of the integrated data repository 104 in conjunction with the health insurance claims data 818, the pharmacy transaction system 806 may generate a pharmacy data table 832. In addition to at least a portion of the information extracted by the pharmacy transaction system 806 from the one or more data tables of the integrated data repository 104, the pharmacy data table 832 may also include start and end dates for the treatment of the patient using the drug substance.

[0155] In various examples, the pharmacy transaction system 806 may analyze the information contained in the pharmacy data table 832 and generate a pharmacy episode data table 834. The pharmacy episode data table 834 may indicate a number of episodes of treatment for a plurality of individuals. Each episode of treatment stored by the pharmacy episode data table 834 may indicate a drug substance provided to a patient in the course of treatment for a biological condition and an amount of time the patient received the drug substance. In one or more examples, the amount of time the patient received the drug substance may be determined by the pharmacy transaction system 806 using the data analysis rules and schema 822. For example, in a situation where a supply date for a first instance of treatment for a patient using a drug substance overlaps with a supply date for a second instance of treatment for the patient using the drug substance, the pharmacy transaction system 806 may determine that the overall cycle of treatment for the patient using the drug substance is a combination of the supply dates for the first and second instances of treatment using the drug substance, rather than the overall supply date including the overlapping cycle. If the overlapping cycle was used to determine the episodes of treatment for the patient, there may be an inaccuracy in the actual timing of the episodes of treatment.

[0156] In one or more examples, the therapy line system 802 may analyze the procedure episode data table 830 and the pharmacy episode data table 834 to determine one or more therapy line data structures 836. The therapy line data structure 836 may indicate one or more lines of therapy provided to one or more patients for the treatment of one or more biological conditions. Each therapy line stored by the therapy line data structure 836 may indicate one or more therapeutic substances provided to the patient during the course of treatment for the biological condition and / or one or more medical procedures obtained by the patient during the course of treatment for the biological condition. Each therapy line stored by the therapy line data structure 836 may also indicate a time period over which at least one of the therapeutic substances or medical procedures was obtained by the patient during the course of treatment for the biological condition. To generate the therapy line data structure 836, the therapy line system 802 may analyze thousands, tens of thousands, hundreds of thousands, or up to millions of health insurance claim records. For a single line of therapy data structure 832 corresponding to a given biological condition, the line of therapy system 802 may analyze health insurance claims data for hundreds, thousands, tens of thousands, up to hundreds of thousands, or more patients who have been diagnosed with the biological condition.

[0157] In one or more additional embodiments, the line of therapy data structure 836 may store multiple lines of therapy for an individual patient. In at least some embodiments, the line of therapy data structure 836 may include a first line of therapy corresponding to a primary course of treatment taken by the patient for a biological condition and a second line of therapy corresponding to a secondary course of treatment taken by the patient for a biological condition. In various embodiments, the line of therapy system 802 may analyze the procedure episode data table 830 and the pharmacy episode data table 834 according to the line of therapy analysis information 816 to determine different lines of therapy for the individual. For example, the line of therapy system 802 may determine that the patient received one or more first therapeutic substances during a first time period and one or more second therapeutic substances during a second time period.

[0158] In one or more illustrative examples, the therapy line system 802 may analyze the therapy line analysis information 816 and determine that the one or more first therapeutic substances and the one or more second therapeutic substances are both included within a primary course of treatment for a biological condition. The therapy line system 802 may also determine that a gap between the first time period and the second time period is less than a first threshold gap. In these scenarios, the therapy line system 802 may determine that the one or more first therapeutic substances and the one or more second therapeutic substances are part of the same therapy line. Thus, a patient may receive different therapeutic substances during different time periods that are part of the same therapy line. To illustrate, one or more therapeutic substances included within a primary course of treatment may be replaced with one or more additional therapeutic substances included within the primary course of treatment and still be considered part of the same therapy line. In addition, the therapy line system 802 may determine that the therapy line for the patient is discontinued in response to determining that the gap between the first time period and the second time period is at least a first threshold gap. In one or more embodiments, the first threshold gap may be at least 30 days, at least 45 days, at least 60 days, at least 75 days, at least 90 days, at least 105 days, at least 120 days, at least 135 days, at least 150 days, at least 165 days, or at least 180 days.

[0159] Further, the therapy line system 802 may determine that a line of therapy for a patient is discontinued by determining that the patient has died. In the situation where a patient has died, the therapy line system 802 may analyze one or more periods during which the patient received treatment in relation to the patient's date of death for the therapy line analysis information 816 to determine the date the line of therapy was discontinued. For example, the therapy line system 802 may determine that the end date of the line of therapy is the patient's date of death when the date of death is within a period during which the patient received one or more treatments for a biological condition. Additionally, the therapy line system 802 may determine that the date the line of therapy was discontinued is the patient's date of death in response to determining that the date of death is within a third threshold gap after the patient received the last treatment. To illustrate, in the situation where a patient dies at least 30 days, at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 150 days, or at least 180 days after the last treatment, the therapy line system 802 may determine that the line of therapy was discontinued on the date of the patient's death. In a scenario in which a patient dies at or after the third threshold gap associated with the last treatment for the biological condition, the therapy line system 802 may determine that the therapy line was discontinued on a date that is a period corresponding to the date of the last treatment + the third threshold gap, such as the date of the last treatment + 60 days or the date of the last treatment + 90 days.

[0160] In one or more additional illustrative examples, the therapy line system 802 may analyze the therapy line analysis information 816 and determine that the one or more first therapeutic substances are part of a first course of treatment for a biological condition and the one or more second therapeutic substances are part of a second course of treatment for a biological condition. In these circumstances, the therapy line system 802 may determine that the one or more first therapeutic substances are part of a first line of therapy and the one or more second therapeutic substances are part of a first line of therapy. In one or more further illustrative examples, the therapy line system 802 may analyze the gap between the first time period and the second time period for one or more threshold gaps included within the therapy line analysis information 816 and determine whether the one or more first therapeutic substances and the one or more second therapeutic substances are part of a same course of treatment for a biological condition or different courses of treatment. For example, the therapy line system 802 may determine that one or more first therapeutic substances are part of a first course of treatment for a biological condition and one or more second therapeutic substances are part of a second course of treatment. The therapy line system 802 may also determine that the gap between the first time period and the second time period is at least a second threshold gap. In these cases, the therapy line system 802 may determine that the one or more first therapeutic substances are part of a first line of therapy and the one or more second therapeutic substances are part of a second line of therapy. In one or more additional scenarios, the therapy line system 802 may determine that the one or more first therapeutic substances and the one or more second therapeutic substances are part of the same line of therapy in response to determining that the gap between the first time period and the second time period is less than the second threshold gap. The second threshold gap may be different from the first threshold gap.The second threshold gap may include at least 10 days, at least 15 days, at least 20 days, at least 25 days, at least 30 days, at least 40 days, at least 50 days, at least 60 days, at least 70 days, at least 80 days, or at least 90 days.

[0161] In one or more embodiments, the line of therapy data structure 836 may include one or more data tables. The one or more data tables may store line of therapy information for individual patients treated for a biological condition. In at least some embodiments, the line of therapy data structure 836 may include multiple data tables, with each data table of the multiple data tables corresponding to an individual biological condition. For example, the line of therapy data structure 836 may include a first data table indicating a line of therapy for patients treated for influenza and a second data table indicating a line of therapy for patients treated for diabetes. In one or more additional embodiments, the line of therapy data structure 836 may include a number of data tables corresponding to patients treated for different forms of biological conditions. To illustrate, the line of therapy data structure 836 may include a first data table indicating a line of therapy for patients treated for type I diabetes and a second data table indicating a line of therapy for patients treated for type II diabetes. In one or more additional embodiments, the line of therapy data structure 836 may include a number of data tables corresponding to patients receiving treatment for different forms of cancer. In one or more illustrative embodiments, the line of therapy data structure 836 may include a first data table indicating a line of therapy corresponding to patients receiving treatment for non-small cell lung cancer, a second data table indicating a line of therapy corresponding to patients receiving treatment for breast cancer, and a third data table indicating a line of therapy corresponding to patients receiving treatment for colon cancer.

[0162] In various embodiments, each data table included within the line of therapy data structure 836 may include a column indicating an identifier of a patient who received treatment for one or more biological conditions. Each row of the data table may correspond to a line of therapy for a given patient. In situations where a patient has received treatment corresponding to multiple lines of therapy, a data table included within the line of therapy data structure 836 may include multiple rows, with each row indicating a different line of therapy for the patient. The data table may also include one or more columns indicating a start date and an end date for the line of therapy. In addition, the data table may include one or more columns indicating the name of the treatment received in association with the line of therapy and one or more columns indicating the category of the treatment received by the patient in association with the line of therapy. Furthermore, the data table may include one or more columns indicating whether a combination of treatments was provided to the patient in association with the line of therapy, whether the patient received adjuvant treatment, or whether maintenance treatment was obtained by the patient.

[0163] In one or more embodiments, the data tables included in the therapy line data structure 836 may also include one or more columns indicating a time to next treatment for the patient. In at least some embodiments, the time to next treatment may indicate a period from a date a first therapy line is started to a subsequent date a second therapy line is started. In at least some embodiments, the therapy line system 802 may use the therapy line analysis information 816 to determine the time to next treatment. For example, the therapy line system 802 may determine that the time to next treatment is the date of death of the patient in a situation where the date of death is before the end date or last active date of the therapy line for the patient. The last active date for the patient may correspond to the last insurance claim paid to the patient in connection with treatment for the biological condition. Additionally, the end date of the therapy line may correspond to the number of days of supply of one or more therapeutic substances provided to the patient or the amount of time a medical procedure is expected to be effective for the patient. In one or more additional examples, the therapy line system 802 may determine that in scenarios where the date of death is within a threshold period of the end date of the line of therapy, the time to next treatment is also the date of death of the patient. Additionally, the therapy line system 802 may determine that the time to next treatment is the last active date for the patient in cases where the last active date is before the end date of the line of therapy. In these examples, the patient may not be followed up on the treatment regimen for the entire period. In still other examples, the therapy line system 802 may determine that the time to next treatment for the patient is a threshold period after the end date of the line of therapy, including in situations where a new line of therapy is started but after the threshold period, or when no new line of therapy has been started for the patient after the end date of the initial line of therapy. In these implementations, the therapy line system 802 may determine that the time to next treatment is the end date of the line of therapy plus the threshold period.In one or more illustrative embodiments, the threshold period may be 15 days, 30 days, 45 days, 60 days, 75 days, 90 days, 105 days, 120 days, 135 days, 150 days, 165 days, or 180 days.

[0164] The line of therapy system 802 may provide one or more line of therapy data structures 836 to the data analysis system 140. The data analysis system 140 may analyze the information stored by the one or more line of therapy data structures 836 to determine the data analysis results 146. In one or more examples, the data analysis system 140 may receive a request to analyze information corresponding to a line of therapy received by a patient treated for a given biological condition. In response to the request, the data analysis system 140 may query the line of therapy system 802 and obtain one or more line of therapy data structures 836 corresponding to the biological condition. For example, the data analysis system 140 may receive a request to analyze information related to a line of therapy for a patient treated for non-small cell lung cancer. In these scenarios, the data analysis system 140 may query the line of therapy system 802 and obtain one or more line of therapy data structures 836 corresponding to non-small cell lung cancer. Data analysis system 140 may then analyze the information stored by one or more line of therapy data structures 836 and generate data analysis results 146 .

[0165] In at least some examples, data analysis system 140 may analyze information stored by one or more line of therapy data structures 836 to generate data analysis results 146 including one or more quantitative measures corresponding to a patient included in the one or more line of therapy data structures 836 being analyzed. To illustrate, data analysis system 140 may analyze line of therapy data structures 836 to determine a real-world survival metric for a patient treated for a biological condition. In various examples, data analysis system 140 may analyze one or more line of therapy data structures 836 corresponding to a biological condition to determine a survival probability over a period of time for a patient receiving one or more lines of therapy to treat the biological condition. In one or more illustrative examples, data analysis system 140 may analyze line of therapy information stored by one or more line of therapy data structures 836 to determine a real-world overall survival metric for a patient. In one or more additional illustrative examples, data analysis system 140 may analyze line of therapy information stored by one or more line of therapy data structures 836 to determine time to next treatment metrics and / or time to discontinuation metrics for patients diagnosed with a biological condition.

[0166] In various examples, the data analysis system 140 may analyze information stored by one or more line of therapy data structures 836 corresponding to the biological condition to determine an amount of progression of the biological condition in at least a subset of patients included in the line of therapy data structure 836 corresponding to the biological condition. In one or more examples, the data analysis system 140 may determine an amount of progression for patients receiving one or more drug substances as part of a line of therapy based on an analysis of the line of therapy information stored by the one or more line of therapy data structures 836 corresponding to the biological condition. In addition, the data analysis system 140 may determine an amount of progression for patients having one or more genomic mutations based on an analysis of the line of therapy information stored by the one or more line of therapy data structures 836 corresponding to the biological condition. In one or more illustrative examples, the data analysis system 140 may analyze at least one of a time to next treatment metric or a time to discontinuation metric generated based on the line of therapy information stored by the one or more line of therapy data structures 836 corresponding to the biological condition to determine an amount of progression of the biological condition for patients having a genomic mutation. In these cases, data analysis system 140 may query integrated data repository 104 to determine patients with one or more genomic mutations. Data analysis system 140 may then query one or more line of therapy data structures 836 corresponding to the given biological condition for line of therapy information associated with patients with one or more genomic mutations, analyze the line of therapy information, and determine data analysis results 146. In at least some examples, data analysis system 140 may analyze at least one of a time to next treatment metric or a time to discontinuation metric to determine the progression of the biological condition for patients with one or more genomic mutations who have received treatment for the biological condition.

[0167] In one or more further examples, data analysis system 140 may analyze line of therapy information stored by one or more line of therapy data structures 836 corresponding to a biological condition to determine a level of resistance developed by one or more patients receiving one or more treatments for the biological condition. For example, data analysis system 140 may analyze line of therapy information stored by one or more line of therapy data structures 836 corresponding to a biological condition to determine a level of resistance in one or more patients receiving one or more drug substances as part of a line of therapy to treat the biological condition. In various examples, data analysis system 140 may analyze at least one of a time to next treatment metric, a time to discontinuation metric, or a real-world survival metric to determine a level of resistance developed by a patient receiving a treatment for the biological condition. In at least some examples, data analysis system 140 may also determine a level of resistance for one or more treatments for an individual with one or more genomic mutations. In at least some examples, the level of resistance may be greater in situations where the time to next treatment or real-world survival rate has a lower value, and the level of resistance may be lower in situations where the time to next treatment or real-world survival rate has a relatively higher value.

[0168] In at least some embodiments, data analysis system 108 may analyze line of therapy information stored by one or more line of therapy data structures 836 corresponding to a biological condition to determine a recommendation for one or more treatments to administer to a patient diagnosed with a biological condition. In one or more embodiments, data analysis system 140 may analyze line of therapy information stored by one or more line of therapy data structures 836 associated with a biological condition to determine one or more characteristics of patients who have received one or more lines of therapy with a relatively low level of resistance and / or a relatively low amount of progression. Data analysis system 140 may then analyze characteristics of one or more additional patients diagnosed with a biological condition to determine whether to recommend one or more lines of therapy as treatment for the one or more additional patients. At least a portion of the one or more additional patients may already be receiving treatment for the biological condition. In one or more additional embodiments, at least a portion of the one or more additional patients may not be receiving treatment for the biological condition. In various examples, data analysis system 140 may also analyze the line of therapy information stored by line of therapy data structure 836 corresponding to a biological condition to determine the effectiveness of the line of therapy for a patient diagnosed with a biological condition. The effectiveness of the line of therapy may correspond to a probability of the line of therapy at least one of reducing the effect of the biological condition or eliminating the biological condition for the patient.

[0169] In various examples, the amount of progression of the biological condition, the effectiveness of the line of therapy to treat the biological condition, the probability of developing resistance to the line of therapy, or a combination thereof may be determined by the data analysis system 140 using at least one of one or more statistical techniques or one or more machine learning techniques. To illustrate, the data analysis system 140 may implement at least one of a Cox proportional hazards model, a chi-square test, a log-rank test, or a Kaplan-Meier method to determine at least one of the amount of progression of the biological condition, the effectiveness of the line of therapy to treat the biological condition, or the probability of developing resistance to the line of therapy. In one or more additional examples, the data analysis system 140 may implement one or more neural networks, one or more convolutional neural networks, or one or more residual neural networks to determine at least one of the amount of progression of the biological condition, the effectiveness of the line of therapy to treat the biological condition, or the probability of developing resistance to the line of therapy.

[0170] In one or more illustrative examples, data analysis system 140 may determine one or more characteristics of patients who have at least one of: a below threshold probability of developing resistance to a line of therapy or at least an additional threshold amount of efficacy for the line of therapy. In one or more scenarios, data analysis system 140 may analyze the line of therapy information stored by one or more line of therapy data structures 836 to determine one or more characteristics. In at least some examples, data analysis system 140 may implement at least one of one or more statistical techniques or one or more machine learning techniques to determine one or more characteristics of patients who have at least one of: a below threshold probability of developing resistance to a line of therapy or at least an additional threshold amount of efficacy for the line of therapy. In one or more examples, data analysis system 140 may implement at least one of one or more extraction algorithms or one or more classification algorithms to determine one or more characteristics. In various embodiments, the data analysis system 140 may implement at least one of one or more neural networks, one or more feedforward neural networks, one or more recurrent neural networks, one or more residual networks, or one or more autoencoders to determine one or more characteristics having at least one of less than a threshold probability of developing resistance to the line of therapy or at least an additional threshold amount of effectiveness for the line of therapy.

[0171] In one or more additional illustrative examples, the data analysis system 140 may implement one or more log-rank tests to analyze the difference between a time-to-death metric and a time-to-next-treatment metric determined based on the one or more lines-of-therapy data structures 836 for patients having one or more genomic mutations and diagnosed with or suspected to have a given biological condition. In various examples, the patients included in the analysis may also be receiving one or more defined lines of therapy to treat the biological condition. In addition, the data analysis system 140 may implement one or more chi-square tests to determine the proportion of patients having one or more co-occurring genomic mutations in patients having one or more defined genomic mutations and, in at least some cases, having one or more additional genomic characteristics, such as one or more clonal genomic mutations versus one or more subclonal genomic mutations. Furthermore, one or more Cox proportional hazards models may be implemented by the data analysis system 140 to determine survival metrics for the patients. In this manner, the effectiveness of one or more lines of therapy for treating a biological condition may be determined by the data analysis system 140 based on the survival probability determined using the Cox proportional hazards model.

[0172] The line of therapy analysis information 816 used by the line of therapy system 802 includes a number of criteria, thresholds, and other information that enable the line of therapy system 802 to generate a line of therapy data structure 836 that can be used by the data analysis system 140 to accurately generate the data analysis results 146. That is, based on the line of therapy analysis information 816 and the computational techniques implemented by the line of therapy system 802, real-world survival metrics, disease progression metrics, disease resistance metrics, treatment efficacy levels, one or more combinations thereof, etc. may be accurately determined. The accurate determination of these quantitative measures enables the data analysis system 140 to provide treatment recommendations to the patient that are accurate, effective, and result in improved outcomes for the patient. Without the framework and protocols defined in the line of therapy analysis information 816 and the computational techniques implemented by the line of therapy system 802 and the data analysis system 140, the treatment recommendations included within the data analysis results 146 are unlikely to improve outcomes for the patient. The line of therapy analysis information 816 is generated over time using a number of computational techniques, training processes, and feedback loops to determine a defined set of criteria, frameworks, protocols, thresholds, and computational techniques that generate optimal treatment recommendations, provide accurate metrics indicating the effectiveness of the line of therapy on outcomes, and provide accurate information regarding the impact of genomic mutations on treatment outcomes.

[0173] 9 illustrates a diagrammatic representation of a computing device 900 in the form of a computer system in which a set of instructions may be executed to cause the computing device 900 to perform any one or more of the methodologies discussed herein, according to an example implementation, according to an example embodiment. Specifically, FIG. 9 shows a diagrammatic representation of a computing device 900 in the example form of a computer system in which instructions 902 (e.g., software, programs, applications, applets, apps, or other executable code) may be executed to cause the computing device 900 to perform any one or more of the methodologies discussed herein. For example, the instructions 902 may cause the computing device 900 to implement the architecture and framework 100, 200, 300, 400, 500 described with respect to FIGS. 1, 2, 3, 4, and 5, respectively, to execute the methods 600, 700 described with respect to FIGS. 6 and 7, respectively, and to implement the architecture 800 described with respect to FIG. 8.

[0174] The instructions 902 transform a generic unprogrammed computing device 900 into a specific computing device 900 that is programmed to perform the functions described and illustrated in the described manner. In alternative implementations, the computing device 900 may operate as a standalone device or be coupled (e.g., networked) to other machines. In a networked deployment, the computing device 900 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The computing device 900 may comprise, without limitation, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing the instructions 902 that specify actions to be taken by the computing device 900. Further, although only a single computing device 900 is illustrated, the term "machine" is also intended to include a collection of machines 900 that individually or jointly execute instructions 902 to perform any one or more of the methodologies discussed herein.

[0175] An embodiment of a computing device 900 may include logic, one or more components, circuits (e.g., modules), or mechanisms. A circuit is a tangible entity configured to perform an operation. In an embodiment, a circuit may be arranged (e.g., internally or relative to an external entity, such as other circuits) in a defined manner. In an embodiment, one or more computer systems (e.g., stand-alone, client, or server computer systems) or one or more hardware processors (processors) may be configured by software (e.g., instructions, application portions, or applications) as a circuit that operates to perform an operation as described herein. In an embodiment, the software may reside (1) on a non-transitory machine-readable medium or (2) within a transmission signal. In an embodiment, the software, when executed by the hardware underlying the circuit, causes the circuit to perform an operation.

[0176] In some embodiments, the circuitry can be implemented mechanically or electronically. For example, the circuitry can comprise dedicated circuitry or logic specifically configured to perform one or more techniques such as those discussed above, including dedicated processors, field programmable gate arrays (FPGAs), or application specific integrated circuits (ASICs), etc. In some embodiments, the circuitry can comprise programmable logic (e.g., circuitry such as that contained within a general-purpose processor or other programmable processor) that can be temporarily configured (e.g., by software) to perform certain operations. It should be understood that the decision to implement a circuitry mechanically (e.g., in dedicated and permanently configured circuitry) or in temporarily configured circuitry (e.g., configured by software) can be driven by cost and time considerations.

[0177] Thus, the term "circuitry" is understood to encompass tangible entities that are physically constructed, permanently configured (e.g., hardwired), or temporarily (e.g., transiently) configured (e.g., programmed) to operate in a specified manner or to perform specified operations. In some embodiments, given multiple temporarily configured circuits, the circuits need not each be configured or instantiated at any one instance in time. For example, if a circuit comprises a general-purpose processor that is configured via software, the general-purpose processor can be configured as separate and different circuits at different times. The software can accordingly configure the processor, for example, to perform a particular circuit at one instance in time and a different circuit at a different instance in time.

[0178] In some embodiments, a circuit can provide information to and receive information from other circuits. In this embodiment, a circuit can be considered as communicatively coupled to one or more other circuits. When multiple such circuits are present simultaneously, communication can be achieved through signal transmission (e.g., via appropriate circuits and buses) connecting the circuits. In implementations in which multiple circuits are configured or instantiated at different times, communication between such circuits can be achieved, for example, through the storage and retrieval of information in a memory structure to which the multiple circuits have access. For example, one circuit can perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. An additional circuit can then access the memory device at a later time to read and process the stored output. In some embodiments, a circuit can be configured to initiate or receive communication with an input or output device and can operate on a resource (e.g., a collection of information, for example).

[0179] Various operations of the method embodiments described herein may be performed, at least in part, by one or more processors that are temporarily (e.g., by software) or permanently configured to perform the associated operations. Such processors, whether temporarily or permanently configured, may constitute processor-implemented circuitry that operates to perform one or more operations or functions. In some embodiments, the circuitry referred to herein may comprise processor-implemented circuitry.

[0180] Similarly, the methods described herein can be at least partially processor-implemented. For example, at least some of the operations of the methods can be performed by one or more processors or processor-implemented circuits. The performance of some of the operations can be distributed among one or more processors that are spread across a number of machines, rather than just residing within a single machine. In some embodiments, the processor or processors can be located within a single location (e.g., in a home environment, an office environment, or as a server farm), while in other embodiments, the processors can be distributed across a number of locations.

[0181] The one or more processors may also operate to support performance of related operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a collection of computers (as an example of a machine that includes a processor), which are accessible over a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application program interfaces (APIs)).

[0182] Exemplary implementations (e.g., devices, systems, or methods) can be implemented in digital electronic circuitry, in computer hardware, in firmware, in software, or any combination thereof. Exemplary implementations can be implemented using a computer program product (e.g., a computer program tangibly embodied in an information carrier or machine-readable medium for execution by, or to control the operation of, a data processing apparatus, such as a programmable processor, a computer, or multiple computers).

[0183] The computer program can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a software module, subroutine, or other unit suitable for use in a computing environment. The computer program can be deployed to be executed on one computer, on multiple computers at one site, or distributed across multiple sites and interconnected by a communication network.

[0184] In some embodiments, the operations may be performed by one or more programmable processors executing computer programs to perform functions by operating on input data and generating output. Embodiments of the method operations may also be performed by, and example apparatus may be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)).

[0185] A computing system may include clients and servers. Clients and servers are generally remote from each other and generally interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In an implementation that deploys a programmable computing system, it should be understood that both hardware and software architectures require consideration. In particular, it should be understood that the choice of whether to implement a certain functionality in permanently configured hardware (e.g., ASIC), in temporarily configured hardware (e.g., a combination of software and programmable processor), or in a combination of permanently and temporarily configured hardware may be a design choice. Below, hardware (e.g., computing device 900) and software architectures that may be deployed in an exemplary implementation are described.

[0186] In some embodiments, computing device 900 may operate as a stand-alone device or computing device 900 may be connected (eg, networked) to other machines.

[0187] In a networked deployment, the computing device 900 can operate in the capacity of either a server or a client machine in a server-client network environment. In some embodiments, the computing device 900 can act as a peer machine in a peer-to-peer (or other distributed) network environment. The computing device 900 can be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, switch, or bridge, or any machine capable of executing (sequentially or otherwise) instructions that define actions to be taken (e.g., performed) by the computing device 900. Furthermore, although only a single computing device 900 is illustrated, the term "computing device" shall also be construed to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to implement any one or more of the methodologies discussed herein.

[0188] The exemplary computing device 900 may include a processor 904 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory 906, and a static memory 908, some or all of which may communicate with each other via a bus 910. The computing device 900 may further include a display unit 912, an alphanumeric input device 914 (e.g., a keyboard), and a user interface (UI) navigation device 916 (e.g., a mouse). In an embodiment, the display unit 912, the input device 914, and the UI navigation device 916 may be touch screen displays. The computing device 900 may additionally include a storage device (e.g., a drive unit) 918, a signal generation device 920 (e.g., a speaker), a network interface device 922, and one or more sensors 924, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or another sensor. Computing device 900 may include an output controller 930 that controls output generated by computing device 900 .

[0189] The storage device 918 may include a machine-readable medium 926 on which is stored one or more sets of data structures or instructions 902 (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. The instructions 902 may also reside, completely or at least partially, within the primary memory 906, within the static memory 908, or within the processor 904 during its execution by the computing device 900. In an embodiment, one or any combination of the processor 904, the primary memory 906, the static memory 908, or the storage device 918 may constitute a machine-readable medium.

[0190] While the machine-readable medium 926 is illustrated as a single medium, the term "machine-readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store one or more instructions 902. The term "machine-readable medium" can also be interpreted to include any tangible medium capable of storing, encoding, or carrying instructions for execution by a machine, causing a machine to perform any one or more of the methodologies of the present disclosure, or capable of storing, encoding, or carrying data structures utilized by or associated with such instructions. The term "machine-readable medium" can be interpreted accordingly to include, but is not limited to, solid-state memories, and optical and magnetic media. Specific examples of machine-readable media include, by way of example, semiconductor memory devices (e.g., electrically programmable read-only memories,

[0191] The memory may include non-volatile memory, including electrically erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM) and flash memory devices, magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0192] The instructions 902 may further be transmitted or received over a communications network 928 using a transmission medium via a network interface device 922 utilizing any one of a number of transport protocols (e.g., Frame Relay, IP, TCP, UDP, HTTP, etc.). Exemplary communications networks may include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile telephone networks (e.g., cellular networks), plain old telephone service (POTS) networks, and wireless data networks (e.g., the IEEE 902.11 family of standards known as Wi-Fi, the IEEE 902.16 family of standards known as WiMax), peer-to-peer (P2P) networks, among others. The term "transmission medium" shall be construed to include any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine, including digital or analog communications signals or other intangible media for facilitating the communication of such software.

[0193] As used herein, a component may refer to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that provide partitioning or modularization of certain processing or control functions. Components may be combined with other components through their interfaces to perform machine processes. A component may be a packaged functional hardware unit designed for use with other components, and may be a part of a program that typically performs a specific function of the associated functionality. A component may constitute either a software component (e.g., code embodied on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit capable of performing an operation, and may be configured or arranged in a physical manner. In various exemplary implementations, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware component that operates to perform certain operations as described herein.

[0194] It should be understood that the individual steps used in the methods of the present teachings can be performed in any order and / or simultaneously so long as the present teachings remain operable. Further, it should be understood that the apparatus and methods of the present teachings can include any number or all of the described implementations so long as the present teachings remain operable.

[0195] The various steps of the methods disclosed herein or steps performed by the systems disclosed herein may be performed at the same or different times and / or in the same geographic location or different geographic locations, e.g., countries. The various steps of the methods disclosed herein can be performed by the same person or different people.

[0196] Various implementations of systems, devices, and methods are described herein. These implementations are provided as examples only and are not intended to limit the scope of the claimed invention. It should also be understood that various features of the described implementations can be combined in various ways to create numerous additional implementations. Also, while various materials, dimensions, shapes, configurations, locations, etc. are described for use with the disclosed implementations, others than those disclosed may be utilized without departing from the scope of the claimed invention.

[0197] Those skilled in the art will recognize that implementations may comprise fewer features than those illustrated in any individual implementation described above. The implementations described herein are not meant to be an exhaustive listing of ways in which various features may be combined. Thus, implementations are not mutually exclusive combinations of features; rather, implementations may comprise combinations of different individual features selected from different individual implementations, as would be understood by one skilled in the art. Also, elements described with respect to one implementation may be implemented in other implementations, even when not described in such implementation, unless otherwise stated. Although a dependent claim may refer to a specific combination with one or more other claims in the claim, other implementations may also include a combination of the dependent claim with the subject matter of each other dependent claim or a combination of one or more features with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a specific combination is not intended. Moreover, it is also intended that the features of a claim be included in any other independent claim, even if the claim is not made directly dependent on the independent claim.

[0198] Also, references herein to "one implementation," "an implementation," or "some implementations" mean that a particular feature, structure, or characteristic described in connection with an implementation is included in at least one implementation of the present teachings. The appearances of the phrase "in one implementation" in various places in the specification do not necessarily all refer to the same implementation.

[0199] The incorporation by reference of any of the above documents is limited such that it does not incorporate any subject matter contrary to the express disclosure of this specification. The incorporation by reference of any of the above documents is further limited such that it does not incorporate by reference any claims contained herein. The incorporation by reference of any of the above documents is further limited such that it does not incorporate by reference any definitions provided herein, unless expressly included herein.

[0200] Although the implementations have been described with reference to specific exemplary implementations, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the present disclosure. Thus, the specification and drawings are to be regarded in an illustrative and not a restrictive sense. The accompanying drawings, which form a part of this specification, show specific implementations in which the subject matter may be practiced, by way of example, and not by way of limitation. The illustrated implementations are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other implementations may be utilized and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. This detailed description is therefore not to be construed in a limiting sense, and the scope of various implementations is defined solely by the appended claims, along with the full scope of equivalents to which such claims are entitled.

[0201] Although specific implementations are illustrated and described herein, it should be understood that any arrangement calculated to achieve the same purpose may be substituted for the specific implementations shown. The present disclosure is intended to cover any and all adaptations or variations of the various implementations. Combinations of the above implementations and other implementations not specifically described herein will be apparent to those of skill in the art upon review of the above description.

[0202] The terms "a" or "an" are used herein to include "one or more than one," as is common in patent documents, independent of any other instance or usage of "at least one" or "one or more." The term "or" is used herein to refer to a non-exclusive or, such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise indicated. The terms "including" and "in which" are used herein as the plain English equivalents of the respective terms "comprising" and "wherein." Also, in the following claims, the terms "including" and "comprising" are open-ended, i.e., a system, user equipment (UE), article, composition, formulation, or process that includes elements in addition to those recited after such term in a claim, will still be considered to fall within the scope of that claim. Also, in the following claims, the terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements on their objects.

[0203] Presented below is a non-limiting numbered list of aspects of the present subject matter. These numbered aspects are provided for illustrative purposes only and do not limit the technology disclosed herein.

[0204] Aspect 1. A method implemented in one or more computing machines comprising processing circuitry and a memory, comprising the steps of: accessing, in the processing circuitry, from the memory, a pharmacy procedure dataset for a patient, wherein each pharmacy procedure in the pharmacy procedure dataset comprises at least a procedure date, a therapy type, and a therapy delivery duration; identifying, from the pharmacy procedure dataset, a pharmacy procedure subset associated with the biological condition based on the therapy type; and calculating, for at least one pharmacy procedure in the pharmacy procedure subset, an end date, wherein the procedure date corresponds to a date when the patient begins payment, and the end date is determined based on the procedure date and the therapy delivery duration associated with the at least one pharmacy procedure; accessing, in the processing circuitry, from the memory, a medical procedure procedure dataset for the patient, wherein each medical procedure in the medical procedure dataset comprises at least a medical procedure date range and a medical procedure type; and identifying, from the medical procedure procedure dataset, a medical procedure subset associated with the biological condition based on the medical procedure type. adjusting, for at least one medical procedure in the medical procedure procedure subset, a medical procedure date range based on a period during which the medical procedure type is valid or repeated; mapping, by the processing circuitry, the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in memory, the timeline data structure storing pharmacy procedures and medical procedure procedures arranged by date; determining, by the processing circuitry, one or more therapy gaps in the timeline data structure between which there are no pharmacy procedures and between which there are no medical procedure procedures, each therapy gap comprising a number of consecutive days, the number of consecutive days being greater than a threshold number of days; determining, by the processing circuitry, one or more lines of therapy based on the one or more therapy gaps, each line of therapy comprising a pharmacy procedure and a medical procedure procedure that occur either between two therapy gaps, before an earliest temporary therapy gap, or after a latest temporary therapy gap;each line of therapy is associated with a line date range; and transmitting to a data repository for storage therein a data structure identifying the patient, the one or more lines of therapy, and a line date range for each of the one or more lines of therapy.

[0205] Aspect 2. The method of aspect 1, further comprising the steps of determining, within a single line of therapy, a first pharmacy procedure or medical procedure associated with a first biological condition stage and a second pharmacy procedure or medical procedure associated with a second biological condition stage, wherein a start date associated with the second pharmacy procedure or medical procedure is later than a start date associated with the first pharmacy procedure or medical procedure, and splitting the single line of therapy into two lines of therapy using the start date associated with the second pharmacy procedure.

[0206] Aspect 3. The method of any of Aspects 1-2, wherein the processing circuitry comprises a plurality of multi-threaded graphic processing units (GPUs), the patient is one of a plurality of patients, and one or more lines of therapy for the patient are determined in parallel with determining lines of therapy for other patients from the plurality of patients using parallel threads of the plurality of multi-threaded GPUs.

[0207] Aspect 4. The method of aspect 3, wherein the step of determining a line of therapy for a plurality of patients includes the step of generating a plurality of intermediate tables, each of the plurality of intermediate tables being stored in a data repository for reviewing and adjusting performance of one or more computing machines.

[0208] Aspect 5. The method of aspect 4, wherein the multiple intermediate tables increase the calculation speed for determining a line of therapy for multiple patients by storing intermediate calculation results.

[0209] Aspect 6. The method of any of aspects 1-5, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a single date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the single date.

[0210] Aspect 7. The method of any of aspects 1-6, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a medical procedure start date and a medical procedure end date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the medical procedure start date and the medical procedure end date.

[0211] Aspect 8. The method of any of aspects 1-7, wherein the pharmacy procedures from the pharmacy procedure subset are mapped onto a timeline data structure based on procedure date.

[0212] Aspect 9. The method of aspect 8, wherein at least one pharmacy procedure is mapped onto a timeline data structure based on a procedure date and a completion date.

[0213] Aspect 10. The method of any of Aspects 1-9, wherein the therapy type comprises a National Drug Code (NDC) classification and the pharmacy procedure dataset comprises one or more tables, the method further comprising determining a set of columns in the one or more tables associated with the drugs, parsing the set of columns to identify NDC classifications, determining a set of NDC classifications corresponding to the drugs, identifying a subset of the set of NDC classifications corresponding to the drugs, the subset being associated with drugs associated with the biological condition, and identifying rows in the one or more tables for placement into the pharmacy procedure subset based on a group of NDC classifications associated with the biological condition.

[0214] Aspect 11. The method of any of Aspects 1-10, wherein the threshold number of days is determined based on a condition type of the biological condition.

[0215] Aspect 12. A system, comprising: a processing circuitry that, when executed by the processing circuitry, causes the processing circuitry to access, for a patient, a pharmacy procedure dataset, each pharmacy procedure in the pharmacy procedure dataset comprising at least a procedure date, a therapy type, and a therapy delivery duration; identify from the pharmacy procedure dataset a pharmacy procedure subset associated with a biological condition based on the therapy type; calculate, for at least one pharmacy procedure in the pharmacy procedure subset, an end date, wherein the procedure date corresponds to a date when the patient begins payment, the end date being determined based on the procedure date and the therapy delivery duration associated with the at least one pharmacy procedure; access, in the processing circuitry from the memory, for the patient, a medical procedure procedure dataset, each medical procedure in the medical procedure dataset comprising at least a medical procedure date range and a medical procedure type; identify from the medical procedure procedure dataset a medical procedure procedure subset associated with the biological condition based on the medical procedure type, wherein for at least one medical procedure in the medical procedure procedure subset, the medical procedure type is valid or repeated; and adjusting a medical procedure date range based on a period of time that the pharmacy procedure subset and the medical procedure procedure subset are associated with a line date range; mapping the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in the memory, the timeline data structure storing the pharmacy procedures and the medical procedure procedures arranged by date; determining in the timeline data structure one or more therapy gaps between which there are no pharmacy procedures and which are not associated with a medical procedure procedure, each therapy gap comprising a number of consecutive days, the number of consecutive days being greater than a threshold number of days; determining one or more lines of therapy based on the one or more therapy gaps, each line of therapy comprising a pharmacy procedure and a medical procedure procedure that occur either between two therapy gaps, before an earliest temporary therapy gap, or after a latest temporary therapy gap, each line of therapy being associated with a line date range; and transmitting to a data repository for storage therein a data structure identifying the patient, the one or more lines of therapy, and the line date range for each of the one or more lines of therapy.

[0216] Aspect 13. The system of Aspect 12, wherein the memory stores additional instructions that, when executed by the processing circuitry, cause the processing circuitry to determine, within a single line of therapy, a first pharmacy procedure or medical procedure associated with a first biological condition stage and a second pharmacy procedure or medical procedure associated with a second biological condition stage, a start date associated with the second pharmacy procedure or medical procedure that is later than a start date associated with the first pharmacy procedure or medical procedure, and cause the single line of therapy to be split into two lines of therapy using the start date associated with the second pharmacy procedure.

[0217] Aspect 14. The system of any of aspects 12-13, wherein the processing circuitry comprises a plurality of multi-threaded graphic processing units (GPUs), the patient is one of a plurality of patients, and one or more lines of therapy for the patient are determined in parallel with determining lines of therapy for other patients from the plurality of patients using parallel threads of the plurality of multi-threaded GPUs.

[0218] Aspect 15. The system of aspect 14, wherein the step of determining a line of therapy for a plurality of patients includes the step of generating a plurality of intermediate tables, each of the plurality of intermediate tables being stored in a data repository for reviewing and adjusting performance of one or more computing machines.

[0219] Aspect 16. The system of aspect 15, wherein multiple intermediate tables increase the calculation speed of determining a line of therapy for multiple patients by storing intermediate calculation results.

[0220] Aspect 17. The system of any of aspects 12-16, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a single date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the single date.

[0221] Aspect 18. The system of any of aspects 12-17, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a medical procedure start date and a medical procedure end date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the medical procedure start date and the medical procedure end date.

[0222] Aspect 19. The system of any of aspects 12-18, wherein the pharmacy procedures from the pharmacy procedure subset are mapped onto the timeline data structure based on procedure date.

[0223] Aspect 20. The system of aspect 19, wherein at least one pharmacy procedure is mapped onto the timeline data structure based on a procedure date and a completion date.

[0224] Aspect 21. The system of any of aspects 12-20, wherein the therapy type comprises a National Drug Code (NDC) classification, and the pharmacy procedure dataset comprises one or more tables, and the memory stores additional instructions that, when executed by the processing circuitry, cause the processing circuitry to determine a set of columns in the one or more tables associated with the drug, analyze the set of columns to identify an NDC classification, determine a set of NDC classifications corresponding to the drug, and identify a subset of the set of NDC classifications corresponding to the drug, the subset being associated with a drug associated with a biological condition, and identify a row in the one or more tables for placement in the pharmacy procedure subset based on a group of NDC classifications associated with the biological condition.

[0225] Aspect 22. The system of any of aspects 12-21, wherein the threshold number of days is determined based on a condition type of the biological condition.

[0226] Aspect 23. One or more non-transitory machine-readable medium that, when executed by processing circuitry of one or more computing machines, causes the processing circuitry to access, for a patient, a pharmacy procedure dataset, each pharmacy procedure in the pharmacy procedure dataset comprising at least a procedure date, a therapy type, and a therapy delivery duration, identify from the pharmacy procedure dataset a pharmacy procedure subset associated with the biological condition based on the therapy type, and calculate, for at least one pharmacy procedure in the pharmacy procedure subset, an end date, the procedure date corresponding to a date when the patient began payment, the end date being determined based on the procedure date and the therapy delivery duration associated with the at least one pharmacy procedure, and access in the processing circuitry from the memory, for the patient, a medical procedure procedure dataset, each medical procedure in the medical procedure dataset comprising at least a medical procedure date range and a medical procedure type, identify from the medical procedure procedure dataset a medical procedure procedure subset associated with the biological condition based on the medical procedure type, and calculate, for at least one medical procedure in the medical procedure procedure subset, adjusting the medical procedure date range based on a period between which the medical procedure type is valid or repeated; mapping the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in memory, the timeline data structure storing the pharmacy procedures and the medical procedure procedures arranged by date; determining in the timeline data structure one or more therapy gaps between which there are no pharmacy procedures and between which there are no medical procedure procedures, each therapy gap comprising a number of consecutive days, the number of consecutive days being greater than a threshold number of days; determining one or more lines of therapy based on the one or more therapy gaps, each line of therapy comprising a pharmacy procedure and a medical procedure procedure that occur either between two therapy gaps, before an earliest temporary therapy gap, or after a latest temporary therapy gap, each line of therapy being associated with a line date range; and transmitting to a data repository for storage therein a data structure identifying the patient, the one or more lines of therapy, and the line date range for each of the one or more lines of therapy.one or more non-transitory machine-readable media;

[0227] Aspect 24. One or more machine-readable media of Aspect 23 storing additional instructions that, when executed by the processing circuitry, cause the processing circuitry to determine, within a single line of therapy, a first pharmacy procedure or medical procedure associated with a first biological condition stage and a second pharmacy procedure or medical procedure associated with a second biological condition stage, a start date associated with the second pharmacy procedure or medical procedure being later than a start date associated with the first pharmacy procedure or medical procedure, and to split the single line of therapy into two lines of therapy using the start date associated with the second pharmacy procedure.

[0228] Aspect 25. One or more machine-readable media of any of Aspects 23-24, wherein the processing circuitry comprises a plurality of multi-threaded graphic processing units (GPUs), the patient is one of a plurality of patients, and one or more lines of therapy for the patient are determined in parallel with determining lines of therapy for other patients from the plurality of patients using parallel threads of the plurality of multi-threaded GPUs.

[0229] Aspect 26. One or more machine-readable media of aspect 25, wherein the step of determining a line of therapy for a plurality of patients includes the step of generating a plurality of intermediate tables, each of the plurality of intermediate tables being stored in a data repository for reviewing and adjusting performance of one or more computing machines.

[0230] Aspect 27. One or more machine-readable media of aspect 26, wherein a plurality of intermediate tables increases the calculation speed of determining a line of therapy for a plurality of patients by storing intermediate calculation results.

[0231] Aspect 28. One or more machine-readable media of any of aspects 23-27, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a single date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the single date.

[0232] Aspect 29. One or more machine-readable media of any of aspects 23-28, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a medical procedure start date and a medical procedure end date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the medical procedure start date and the medical procedure end date.

[0233] Aspect 30. One or more machine-readable media of any of aspects 23-29, wherein the pharmacy procedures from the pharmacy procedure subset are mapped onto a timeline data structure based on procedure date.

[0234] Aspect 31. One or more machine-readable media according to aspects 23-30, wherein at least one pharmacy procedure is mapped onto a timeline data structure based on a procedure date and a completion date.

[0235] Aspect 32. The one or more machine-readable media of any of aspects 23-31, wherein the therapy type comprises a National Drug Code (NDC) classification and the pharmacy procedure dataset comprises one or more tables, the one or more machine-readable media storing additional instructions that, when executed by the processing circuitry, cause the processing circuitry to determine a set of columns in the one or more tables associated with the drug, analyze the set of columns to identify an NDC classification, determine a set of NDC classifications corresponding to the drug, and identify a subset of the set of NDC classifications corresponding to the drug, the subset being associated with a drug associated with a biological condition, and identify a row in the one or more tables for placement in the pharmacy procedure subset based on a group of NDC classifications associated with the biological condition.

[0236] Aspect 33. One or more machine-readable media of any of aspects 23-32, wherein the threshold number of days is determined based on a condition type of the biological condition.

[0237] Aspect 34. A method comprising: analyzing, by a computing system having one or more processors and a memory, health insurance profile data stored by an integrated data repository to determine a number of patients who are at least one of diagnosed with a biological condition or have one or more genomic mutations detected, where the integrated data repository stores the health insurance profile data along with genomic data for the number of patients; determining, by the computing system, a subset of the health insurance profile data corresponding to the number of patients; analyzing, by the computing system, the subset of the health insurance profile data to determine a medical procedure record for a first patient of the number of patients, where the medical procedure record indicates one or more medical procedures administered to treat the biological condition in the first patient and indicates one or more first dates of the one or more medical procedures; and analyzing, by the computing system, the subset of the health insurance profile data. determining, by the computing system, a pharmacy record for a second patient of the number of patients, the pharmacy record indicating one or more drug treatments provided to the second patient for treating the biological condition and indicating one or more second dates associated with the one or more drug treatments; determining, by the computing system, a line of therapy corresponding to the patients of the number of patients based on at least one of the patient's medical procedure record or the patient's pharmacy record, the line of therapy indicating at least one of (i) a medical procedure administered to the patient for treating the biological condition and a first period during which the medical procedure was administered, or (ii) a drug treatment provided to the individual for treating the biological condition and a second period during which the drug treatment should be provided to the patient; generating, by the computing system, a line of therapy data structure including the patient's line of therapy and a number of additional lines of therapy for the first patient and the second patient; analyzing, by the computing system, the information stored by the line of therapy data structure;1. A method comprising: determining one or more quantitative measures for a first patient and a second patient in association with at least one of a medical procedure included within the one or more medical procedures or a drug treatment of the one or more drug treatments, the one or more quantitative measures corresponding to at least one of a first probability of progression of a biological condition or a second probability of resistance to at least one of the medical procedures or the drug treatments; and determining, by a computing system, at least one of an amount of progression of a biological condition or a level of resistance to the drug treatment for at least one of the first patient or the second patient based on the one or more quantitative measures.

[0238] Aspect 35. The method of aspect 34, comprising determining, by a computing system, a level of effectiveness of at least one of the medical procedure or pharmaceutical treatment for at least one of the first patient or the second patient based on one or more quantitative measures.

[0239] Aspect 36. The method of aspect 34 or 35, wherein the pharmacy record indicates a first instance in which the patient receives a first supply of the drug substance and a second instance in which the patient receives a second supply of the drug substance, the first supply of the drug substance and the second supply of the drug substance spanning a number of days, the method including determining, by the computing system, that the first instance in which the patient receives the first supply occurs on a first date and that the second instance in which the patient receives the second supply occurs on a second date.

[0240] Aspect 37. The method of aspect 36, comprising: determining, by the computing system, that the second date is less than a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply; and determining, by the computing system, that a second instance in which the patient receives a second supply of the drug substance is part of the line of therapy.

[0241] Aspect 38. The method of aspect 37, including determining, by the computing system, that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining, by the computing system, that a second instance in which the patient receives a second supply of the drug substance is part of an additional line of therapy that is different from the line of therapy.

[0242] Aspect 39. The method of any one of aspects 34-38, wherein the pharmacy record indicates a first instance in which the patient receives a first supply of a drug substance and a second instance in which the patient receives a second supply of a second drug substance, the first supply of the drug substance and the second supply of the drug substance spanning a number of days, the method including determining, by the computing system, that the first instance in which the patient receives the first supply occurs on a first date and that the second instance in which the patient receives the second supply occurs on a second date.

[0243] Aspect 40. The method of aspect 39, including determining, by the computing system, that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining, by the computing system, that a second instance in which the patient receives a second supply of the second drug substance is part of an additional line of therapy that is different from the line of therapy.

[0244] Aspect 41. The method of aspect 39, including the steps of: determining, by a computing system, that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply; determining, by the computing system, that the first drug substance and the second drug substance are part of a primary course of treatment for the biological condition; and determining, by the computing system, that a second instance in which the patient receives a second supply of the second drug substance is part of the line of therapy.

[0245] Aspect 42. The method of any one of aspects 34-41, wherein the one or more quantitative measures include at least one of time to next treatment, time to death, or real-world overall survival.

[0246] Aspect 43. The method of aspect 42, comprising determining, by a computing system, at least one of an amount of progression of a biological condition for the patient or a level of resistance to a drug substance based on a time to next treatment for the patient over a period of time for the line of therapy and one or more additional lines of therapy.

[0247] Aspect 44. The method of any one of aspects 34-43, comprising: generating, by a computing system, a procedure data table that stores medical procedure records for a first patient; generating, by the computing system, a pharmacy data table that stores pharmacy records for a second patient; and analyzing, by the computing system, the first information stored by the procedure data table and the second information stored by the pharmacy data table according to a framework indicative of at least one of one or more threshold time periods, one or more criteria, or one or more data analysis rules corresponding to a biological condition, and generating a line of therapy data structure including a line of therapy.

[0248] Aspect 45. A system having one or more hardware processing units, the system, when executed by the one or more hardware processing units, includes the steps of: analyzing health insurance lateral data stored by an integrated data repository to determine a number of patients who are at least one of diagnosed with a biological condition or have one or more genomic mutations detected, the integrated data repository storing the health insurance lateral data along with genomic data for the number of patients; determining a subset of the health insurance lateral data corresponding to the number of patients; analyzing the subset of the health insurance lateral data to determine a medical procedure record for a first patient of the number of patients, the medical procedure record indicating one or more medical procedures administered to treat the biological condition in the first patient and indicating one or more first dates of the one or more medical procedures; and analyzing the subset of the health insurance lateral data to determine a pharmacy record for a second patient of the number of patients, the pharmacy record indicating one or more medical procedures administered to treat the biological condition in the first patient and indicating one or more first dates of the one or more medical procedures. determining, by the computing system, a line of therapy corresponding to the number of patients based on at least one of the patient's medical procedure records or the patient's pharmacy records, the line of therapy indicating at least one of (i) a medical procedure administered to the patient to treat a biological condition and a first time period during which the medical procedure was administered or (ii) a drug treatment provided to the individual to treat a biological condition and a second time period during which the drug treatment should be provided to the patient; generating a line of therapy data structure including the patient's line of therapy and a number of additional lines of therapy for the first patient and the second patient; analyzing the information stored by the line of therapy data structure to determine one or more quantitative measures for the first patient and the second patient associated with at least one of the medical procedures or the one or more drug treatments included within the one or more medical procedures,and one or more computer-readable storage media storing computer-executable instructions that cause the system to perform operations including: determining, based on the one or more quantitative measures, at least one of an amount of progression of the biological condition or a level of resistance to the medical procedure or the drug treatment for at least one of the first patient or the second patient.

[0249] Aspect 46. The system of Aspect 45, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining a level of effectiveness of at least one of a medical procedure or pharmaceutical treatment for at least one of the first patient or the second patient based on one or more quantitative measures.

[0250] Aspect 47. The system of aspect 45 or 46, wherein the pharmacy record indicates a first instance in which the patient receives a first supply of the drug substance and a second instance in which the patient receives a second supply of the drug substance, the first supply of the drug substance and the second supply of the drug substance spanning a number of days, and the one or more computer-readable storage media stores additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining that the first instance in which the patient receives the first supply occurs on a first date and the second instance in which the patient receives the second supply occurs on a second date.

[0251] Aspect 48. The system of Aspect 47, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is less than a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining that a second instance in which the patient receives a second supply of the drug substance is part of a line of therapy.

[0252] Aspect 49. The system of Aspect 48, wherein the one or more computer-readable storage media stores additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining that a second instance in which the patient receives a second supply of the drug substance is part of an additional line of therapy that is different from the line of therapy.

[0253] Aspect 50. The method of any one of aspects 45-49, wherein the pharmacy record indicates a first instance in which the patient receives a first supply of a drug substance and a second instance in which the patient receives a second supply of a second drug substance, the first supply of the drug substance and the second supply of the drug substance spanning a number of days, and the one or more computer-readable storage media stores additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform additional operations including determining that the first instance in which the patient receives the first supply occurs on a first date and the second instance in which the patient receives the second supply occurs on a second date.

[0254] Aspect 51. The system of Aspect 50, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining that a second instance in which the patient receives a second supply of the second drug substance is part of an additional line of therapy that is different from the line of therapy.

[0255] Aspect 52. The system of Aspect 51, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, determining that the first drug substance and the second drug substance are part of a primary course of treatment for the biological condition, and determining that a second instance in which the patient receives a second supply of the second drug substance is part of a line of therapy.

[0256] Aspect 53. The system of any one of aspects 45-52, wherein the one or more quantitative measures include at least one of time to next treatment, time to death, or real-world overall survival.

[0257] Aspect 54. The system of Aspect 53, wherein the one or more computer-readable storage media store additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations, including determining at least one of an amount of progression of a biological condition for the patient or a level of resistance to a drug substance based on a time to next treatment for the patient over a period of time for the therapy line and one or more additional therapy lines.

[0258] Aspect 55. The system of any one of Aspects 45-54, wherein the one or more computer-readable storage media stores additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including generating a procedure data table that stores medical procedure records for a first patient, generating a pharmacy data table that stores pharmacy records for a second patient, and analyzing the first information stored by the procedure data table and the second information stored by the pharmacy data table according to a framework indicative of at least one of one or more threshold time periods, one or more criteria, or one or more data analysis rules corresponding to a biological condition, and generating a therapy line data structure including a therapy line.

[0259] Aspect 56. One or more non-transitory computer readable storage media that, when executed by one or more hardware processing units, causes the system to include the steps of: analyzing health insurance aspect data stored by an integrated data repository to determine a number of patients who are at least one of diagnosed with a biological condition or have one or more genomic mutations detected, where the integrated data repository stores the health insurance aspect data along with genomic data for the number of patients; determining a subset of the health insurance aspect data corresponding to the number of patients; analyzing the subset of the health insurance aspect data to determine a medical procedure record for a first patient of the number of patients, where the medical procedure record indicates one or more medical procedures administered to treat the biological condition in the first patient and indicates one or more first dates of the one or more medical procedures; and analyzing the subset of the health insurance aspect data to determine a pharmacy record for a second patient of the number of patients, where the pharmacy record indicates a pharmacy record provided to the second patient to treat the biological condition. determining, by the computing system, a line of therapy corresponding to the number of patients based on at least one of the patient's medical procedure records or the patient's pharmacy records, the line of therapy indicating at least one of (i) a medical procedure administered to the patient to treat a biological condition and a first time period during which the medical procedure was administered, or (ii) a drug treatment provided to the individual to treat a biological condition and a second time period during which the drug treatment should be provided to the patient; generating a line of therapy data structure including the patient's line of therapy and a number of additional lines of therapy for the first patient and the second patient; analyzing the information stored by the line of therapy data structure to determine one or more quantitative measures for the first patient and the second patient associated with at least one of the medical procedures or the one or more drug treatments included within the one or more medical procedures, the one or more quantitative measures being:one or more non-transitory computer readable storage media storing computer executable instructions to perform operations including: determining at least one of an amount of progression of a biological condition or a level of resistance to at least one of a medical procedure or a drug treatment for at least one of the first patient or the second patient based on the one or more quantitative measures corresponding to at least one of a first probability of progression of a biological condition or a second probability of resistance to at least one of a medical procedure or a drug treatment;

[0260] Aspect 57. The one or more non-transitory computer-readable media of Aspect 56 comprising additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining a level of effectiveness of at least one of a medical procedure or drug treatment for at least one of the first patient or the second patient based on one or more quantitative measures.

[0261] Aspect 58. The one or more non-transitory computer readable media of aspect 56 or 57, wherein the pharmacy record indicates a first instance in which the patient receives a first supply of the drug substance and a second instance in which the patient receives a second supply of the drug substance, the first supply of the drug substance and the second supply of the drug substance spanning a number of days, the one or more non-transitory computer readable storage media storing additional computer executable instructions that, when executed by the one or more hardware processing units, cause the one or more hardware processing units to perform additional operations including determining that the first instance in which the patient receives the first supply occurs on a first date and that the second instance in which the patient receives the second supply occurs on a second date.

[0262] Aspect 59. One or more non-transitory computer readable media of aspect 58 comprising additional computer executable instructions which, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is less than a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining that a second instance in which the patient receives a second supply of the drug substance is part of a line of therapy.

[0263] Aspect 60. One or more non-transitory computer readable media of aspect 59 comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining that a second instance in which the patient receives a second supply of the drug substance is part of an additional line of therapy that is different from the line of therapy.

[0264] Aspect 61. The one or more non-transitory computer readable media of any one of aspects 56-60, wherein the pharmacy record indicates a first instance in which the patient receives a first supply of a drug substance and a second instance in which the patient receives a second supply of a second drug substance, the first supply of the drug substance and the second supply of the drug substance spanning a number of days, the one or more non-transitory computer readable storage media storing additional computer-executable instructions that, when executed by the one or more hardware processing units, cause the one or more hardware processing units to perform additional operations including determining that the first instance in which the patient receives the first supply occurs on a first date and that the second instance in which the patient receives the second supply occurs on a second date.

[0265] Aspect 62. One or more non-transitory computer readable media of aspect 61 comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply, and determining that a second instance in which the patient receives a second supply of the second drug substance is part of an additional line of therapy that is different from the line of therapy.

[0266] Aspect 63. One or more non-transitory computer readable media of aspect 56 comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining that the second date is at least a threshold period of time after an intermediate date comprising the first date plus the number of days of the first supply; determining that the first drug substance and the second drug substance are part of a primary course of treatment for the biological condition; and determining that a second instance in which the patient receives a second supply of the second drug substance is part of a line of therapy.

[0267] Aspect 64. One or more non-transitory computer readable media of any one of aspects 56-63, wherein the one or more quantitative measures include at least one of time to next treatment, time to death, or real-world overall survival.

[0268] Aspect 65. One or more non-transitory computer readable media of aspect 64 comprising additional computer executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including determining at least one of an amount of progression of a biological condition for a patient or a level of resistance to a drug substance based on a time to next treatment for the patient over a period of time for the therapy line and one or more additional therapy lines.

[0269] Aspect 66. One or more non-transitory computer readable media of any one of Aspects 56-65 comprising additional computer-executable instructions that, when executed by one or more hardware processing units, cause the system to perform additional operations including generating a procedure data table that stores medical procedure records for a first patient, generating a pharmacy data table that stores pharmacy records for a second patient, and analyzing the first information stored by the procedure data table and the second information stored by the pharmacy data table according to a framework indicative of at least one of one or more threshold time periods corresponding to a biological condition, one or more criteria, or one or more data analysis rules, and generating a therapy line data structure including a therapy line.

[0270] Side 67 is a device that includes means for mounting any of sides 1-11.

[0271] Side 68 is a device that includes means for mounting any of sides 34-44.

[0272] Working Example Example 1 Identification of high-frequency microsatellite instability (MSI-H) is clinically meaningful in patients with advanced gastrointestinal (aGI) cancer, given the associated approval of multiple immune checkpoint inhibitors (ICIs). MSI-H has long been assessed via tissue analysis, with insights from plasma-based approaches limited to small validation studies. We sought to assess the prevalence of early and potential acquired MSI-H status across the aGI and report real-world outcomes of colorectal cancer (CRC) patients who underwent ICI after MSI-H identification with a commercially available liquid biopsy assay. Assessment of MSI prevalence genomic results from patients with advanced GI cancer who had testing using a commercially available liquid biopsy assay as part of routine clinical treatment from October 2019 to September 2021 were queried to assess MSI-H prevalence and identify cases of potential acquired MSI-H. The prevalence of MSI-H in approximately 30,000 advanced GI patients is shown in Figure 10. Figure 11 shows the MSI-H score, maximum variant allele fraction (VAF), and intervention treatment for consecutive tests for MSI-H.

[0273] Assessment of outcomes for patients with colorectal cancer with MSI-H was identified from liquid biopsy assays. Real-world evidence (RWE) was obtained from a pooled data repository consisting of aggregated payer claims and de-identified records from >150,000 individuals with comprehensive cell-free circulating tumor DNA (ctDNA) test results via liquid biopsy assays from November 2018 to May 2021. Patients with plasma-identified MSI-H who initiated new therapy ≤60 days after the assay report date were sorted into groups, group 1 treated with chemotherapy ± biologic therapy ("chemotherapy") and group 2 treated with immunotherapy ("ICI") via pembrolizumab or nivolumab, and real-world time to discontinuation (rwTTD) and real-world time to next treatment (rwTTNT) were assessed as measures of progression-free survival. Log-rank tests were used to assess differences in rwTTD, rwTTNT, and overall survival.

[0274] Patients who received ICI following MSI-H detection on liquid biopsy had significantly longer rwTTD, rwTTNT compared to those who received chemotherapy. Of 1,227 patients with MSI-H in the pooled data repository, 222 had CRC and were eligible for RWE analysis. 89 / 222 (45%) started a new therapy within 60 days of the result, 42 (47%) received ICI, 39 (44%) received chemotherapy, and 8 (9%) received a mixed regimen. Figure 12 illustrates rwTTD across treatment groups. Patients who received ICI had significantly longer rwTTD than those who received chemotherapy, median months to discontinuation = 7.5 (95% CI 3.4-12.3) vs. 2 (95% 1.4-3.3), p<0.001. Figure 13 illustrates rwTTNT across treatment groups. Patients who received ICIs had significantly longer rwTTNT than those who received chemotherapy, median months to next treatment = 23.8 (95% 10.6-NA) vs. 4.5 (95% CI 2.9-NA), p = 0.006. In addition, Figure 14 illustrates overall survival across treatment groups. No significant overall survival differences were observed between patients treated with ICIs and those treated with chemotherapy (p = 0.559).

[0275] The present commercial liquid biopsy assay detects MSI-H at a frequency similar to published tissue cohorts and, although a rare event, may identify acquired MSI-H following early lines of therapy. Well-validated liquid biopsy is a viable tool for identifying early and acquired MSI-H in advanced gastric cancer and may expand the number of patients who may benefit from ICI therapy, especially in cases where access to tissue samples is not feasible. Serial analysis and / or testing of patients with aGI cancers using liquid biopsy at the time of progression may maximize the opportunity to access ICI. Patients who received ICI following identification of MSI-H via liquid biopsy achieved clinical responses consistent with published data in pretreated advanced colorectal cancer, and MSI-H, when identified on liquid biopsy, may be treated similarly to when it would be identified on tissue.

[0276] Example 2 Alpelisib is an alpha-selective PI3K inhibitor approved in combination with fulvestrant for PIK3CA-mt HR+ / HER2 advanced breast cancer. These mutations can be either trancal (clonal) or acquired (subclonal) under treatment pressure, however, data on alpelisib efficacy in these two populations is currently limited. This study utilized RWE to assess how the PIK3CA genomic environment influences alpelisib response.

[0277] RWE was obtained from a pooled data repository consisting of aggregated commercial payer health claims and de-identified records from over 140,000 individuals with comprehensive cell-free circulating tumor DNA (ctDNA) testing results via commercially available liquid biopsies. All HR+ / HER2 advanced breast cancer patients with 1 or more of the 11 PIK3CA-mt markers cited in the therascreen PIK3CA RGQ PCR Kit alpelisib companion diagnostic identified via ctDNA since May 2019 (the month alpelisib was FDA approved) were included.

[0278] PIK3CA-mt is >50% (clonal) or <The clonal proportion (copy number adjusted PIK3CA mutant allele proportion / maximum somatic mutant allele proportion) of 50% (subclonal) was defined. Real-world time to treatment discontinuation (rwTTD) and real-world time to next treatment (rwTTNT) were assessed as indicators for progression-free survival. Log-rank test was used to assess the difference in rwTTD and rwTTNT, and chi-square test was used to compare the proportion of PIK3CA-mt and other co-occurring alterations between patients with only clonal PIK3CA-mt and patients with only subclonal PIK3CA-mt. Cox proportional hazards model was used to obtain hazard ratios. Of the 223 eligible patients, 216 (96%) did not have any prior alpelisib exposure and were included for further analysis. As shown in Figure 15, most patients had one PIK3CA-mt (73%), and 82% carried only clonal mutations.

[0279] PIK3CA E454K and E454G alterations were significantly more likely to be subclonal rather than clonal (p=0.033 and 0.017, respectively), while H1047R, E542K, H1047L, and G546R were observed as both clonal and subclonal alterations. As shown in Figures 16 and 17, there were no significant differences in rwTTD or rwTTNT for alpelisib in patients with clonal versus subclonal PIK3CA-mt. > Using a cutoff for clonality of 90% clonal percentage also did not result in any significant differences in rwTTD and rwTTNT, as shown in Figure 18. There were also no significant differences in the frequency of co-occurring alterations between samples with clonal versus subclonal PIK3A-mt, as shown in Figure 19. Many alterations known to be associated with resistance to alpelisib and / or CDK4 / 6 inhibitors were observed, including RB1 and PTEN loss-of-function mutations.

[0280] Review of RWE data from over 200 advanced breast cancer patients treated with alpelisib following identification of PIK3CA-mt via ctDNA did not show any significant differences in alpelisib outcomes based on the clonality of the identified PIK3CA-mt, suggesting that PIK3CA-mt does not need to be clonal for patients to benefit from alpelisib therapy, which is noteworthy since only 16% of patients were identified with subclonal PIK3CA-mt. Some PIK3CA-mt were significantly more likely to be subclonal, suggesting potential underlying biological differences (e.g., treatment pressures) driving the development of these mutations. Patients with both clonal and subclonal PIK3CA-mt did not have any significant differences in co-occurring alterations, although both cohorts had alterations suggestive of primary and / or acquired resistance to common therapies, including alpelisib and CDK4 / 6 inhibitors.

Claims

1. 1. A method implemented in one or more computing machines having processing circuitry and memory, the method comprising: accessing, in the processing circuitry, from the memory, a pharmacy procedure dataset for the patient, each pharmacy procedure in the pharmacy procedure dataset comprising at least a procedure date, a therapy type, and a therapy delivery duration; identifying a subset of pharmacy procedures from the pharmacy procedure dataset that are associated with a biological condition based on the therapy type; calculating an end date for at least one pharmacy procedure in the pharmacy procedure subset, the procedure date corresponding to a date when the patient began payment, the end date being determined based on the procedure date and the therapy supply duration associated with the at least one pharmacy procedure; accessing, in the processing circuitry, from the memory, a medical procedure procedure dataset for the patient, wherein each medical procedure in the medical procedure dataset comprises at least a medical procedure date range and a medical procedure type; identifying a subset of medical procedure procedures from the medical procedure procedure dataset that are associated with the biological condition based on the medical procedure type; adjusting, for at least one medical procedure in the medical procedure procedure subset, a medical procedure date range based on a period during which the medical procedure type is valid or repeated; mapping, by the processing circuitry, the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in the memory, the timeline data structure storing pharmacy procedures and medical procedure procedures arranged by date; determining, by the processing circuitry, one or more therapy gaps in the timeline data structure in which no pharmacy procedures are present and no medical procedure procedures are present, each therapy gap comprising a number of consecutive days, the number of consecutive days being greater than a threshold number of days; determining, by the processing circuitry, one or more lines of therapy based on the one or more therapy gaps, each line of therapy comprising pharmacy procedures and medical procedure procedures occurring between two therapy gaps, either before an earliest temporary therapy gap or after a latest temporary therapy gap, each line of therapy being associated with a line date range; transmitting a data structure identifying the patient, the one or more lines of therapy, and the line date range for each of the one or more lines of therapy to a data repository for storage therein; A method comprising:

2. determining, within a single line of therapy, a first pharmacy procedure or medical procedure associated with a first biological condition stage and a second pharmacy procedure or medical procedure associated with a second biological condition stage, wherein a start date associated with the second pharmacy procedure or medical procedure is later than a start date associated with the first pharmacy procedure or medical procedure; dividing the single line of therapy into two lines of therapy using the start date associated with the second pharmacy procedure; and The method of claim 1 further comprising:

3. the processing circuitry comprises a plurality of multi-threaded graphics processing units (GPUs); the patient is one of a plurality of patients; the one or more lines of therapy for the patient are determined in parallel with determining lines of therapy for other patients from the plurality of patients using parallel threads of the plurality of multi-threaded GPUs. The method of claim 1.

4. 4. The method of claim 3, wherein determining the line of therapy for the plurality of patients includes generating a plurality of intermediate tables, each of the plurality of intermediate tables being stored in the data repository for reviewing and adjusting performance of the one or more computing machines.

5. The method of claim 4 , wherein the plurality of intermediate tables stores intermediate calculation results to increase the calculation speed for determining the line of therapy for the plurality of patients.

6. 2. The method of claim 1, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a single date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the single date.

7. 2. The method of claim 1, wherein, for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a medical procedure start date and a medical procedure end date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the medical procedure start date and the medical procedure end date.

8. The method of claim 1 , wherein pharmacy procedures from the pharmacy procedure subset are mapped onto the timeline data structure based on the procedure date.

9. 9. The method of claim 8, wherein the at least one pharmacy procedure is mapped onto the timeline data structure based on the procedure date and the end date.

10. The therapy type comprises a National Drug Code (NDC) classification, the pharmacy transaction dataset comprises one or more tables, and the method further comprises: determining a set of columns in the one or more tables that are related to medications; analyzing the set of columns to identify an NDC classification; determining said set of NDC classifications corresponding to a drug; identifying a subset of the set of NDC classifications corresponding to drugs, the subset being associated with drugs associated with the biological condition; identifying a row in the one or more tables for placement in the pharmacy procedure subset based on the group of NDC classifications associated with the biological condition; The method of claim 1 , comprising:

11. The method of claim 1 , wherein the threshold number of days is determined based on a condition type of the biological condition.

12. The method of claim 1, further comprising, based on a determination that the patient possesses a breast cancer mutation having a first clonal proportion below a clonal proportion threshold, recommending one or more treatments targeting the breast cancer mutation in the patient according to the one or more lines of therapy.

13. The clonal fraction threshold comprises a maximum breast cancer somatic mutant allele fraction of 50%; 13. The method of claim 12, wherein the one or more treatments comprise exposure to a PI3K inhibitor.

14. Identifying, based on the data structure, that another patient possesses a second clonal proportion of breast cancer mutations that is higher than the first clonal proportion; recommending one or more treatments targeting said breast cancer mutation in said another patient according to said one or more lines of therapy; further comprising 10. The method of claim 1, wherein the one or more treatments targeting the breast cancer mutations in the patient with the first clonal fraction and the one or more treatments targeting the breast cancer mutations in the other patient with the second clonal fraction result in substantially similar progression-free survival metrics.

15. The method of claim 14, wherein the progression-free survival metric comprises real-world time to discontinuation (rwTTD), real-world time to next treatment (rwTTNT), or a combination thereof.

16. The method described in claim 14, wherein the second clone proportion comprises a maximum breast cancer somatic mutant allele proportion of at least 50% or at least 90%.

17. 1. A system comprising: processing circuitry; Memory and Equipped with The memory stores instructions that, when executed by the processing circuitry, cause the processing circuitry to: accessing a pharmacy procedure dataset for a patient, wherein each pharmacy procedure in the pharmacy procedure dataset comprises at least a procedure date, a therapy type, and a therapy delivery duration; identifying a subset of pharmacy procedures from the pharmacy procedure dataset that are associated with a biological condition based on the therapy type; calculating an end date for at least one pharmacy procedure in the pharmacy procedure subset, the procedure date corresponding to a date when the patient began payment, the end date being determined based on the procedure date and the therapy supply duration associated with the at least one pharmacy procedure; accessing, in the processing circuitry, from the memory, a medical procedure procedure dataset for the patient, wherein each medical procedure in the medical procedure dataset comprises at least a medical procedure date range and a medical procedure type; identifying a subset of medical procedure procedures from the medical procedure procedure dataset that are associated with the biological condition based on the medical procedure type; adjusting, for at least one medical procedure in the medical procedure procedure subset, a medical procedure date range based on a period during which the medical procedure type is valid or repeated; mapping the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in the memory, the timeline data structure storing pharmacy procedures and medical procedure procedures arranged by date; determining one or more therapy gaps within the timeline data structure in which no pharmacy procedures are present and no medical procedure procedures are present, each therapy gap comprising a number of consecutive days, the number of consecutive days being greater than a threshold number of days; determining one or more lines of therapy based on the one or more therapy gaps, each line of therapy comprising a pharmacy procedure and a medical procedure procedure occurring either between two therapy gaps, before an earliest temporary therapy gap, or after a latest temporary therapy gap, each line of therapy being associated with a line date range; transmitting a data structure identifying the patient, the one or more lines of therapy, and the line date range for each of the one or more lines of therapy to a data repository for storage therein; A system configured to:

18. The memory stores additional instructions that, when executed by the processing circuitry, cause the processing circuitry to: determining, within a single line of therapy, a first pharmacy procedure or medical procedure associated with a first biological condition stage and a second pharmacy procedure or medical procedure associated with a second biological condition stage, wherein a start date associated with the second pharmacy procedure or medical procedure is later than a start date associated with the first pharmacy procedure or medical procedure; dividing the single line of therapy into two lines of therapy using the start date associated with the second pharmacy procedure; and 20. The system of claim 17, configured to:

19. the processing circuitry comprises a plurality of multi-threaded graphics processing units (GPUs); the patient is one of a plurality of patients; the one or more lines of therapy for the patient are determined in parallel with determining lines of therapy for other patients from the plurality of patients using parallel threads of the plurality of multi-threaded GPUs.

20. The system of claim 17.

20. determining the line of therapy for the plurality of patients includes generating a plurality of intermediate tables, each of the plurality of intermediate tables being stored in the data repository for reviewing and adjusting performance of the one or more computing machines; 20. The system of claim 19, wherein the plurality of intermediate tables increases the calculation speed for determining the line of therapy for the plurality of patients by storing intermediate calculation results.

21. 20. The system of claim 17, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a single date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the single date.

22. 20. The system of claim 17, wherein for at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a medical procedure start date and a medical procedure end date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the medical procedure start date and the medical procedure end date.

23. pharmacy procedures from the pharmacy procedure subset are mapped onto the timeline data structure based on the procedure date; 20. The system of claim 17, wherein the at least one pharmacy procedure is mapped onto the timeline data structure based on the procedure date and the end date.

24. the therapy type comprises a National Drug Code (NDC) classification, the pharmacy procedure dataset comprises one or more tables, and the memory stores additional instructions that, when executed by the processing circuitry, cause the processing circuitry to: determining a set of columns in the one or more tables that are related to medications; analyzing the set of columns to identify an NDC classification; determining said set of NDC classifications corresponding to a drug; identifying a subset of the set of NDC classifications corresponding to drugs, the subset being associated with drugs associated with the biological condition; identifying a row in the one or more tables for placement in the pharmacy procedure subset based on the group of NDC classifications associated with the biological condition; The system of claim 17 ,

25. The system of claim 17 , wherein the threshold number of days is determined based on a condition type of the biological condition.

26. One or more non-transitory machine-readable media having instructions stored thereon that, when executed by processing circuitry of one or more computing machines, cause the processing circuitry to: accessing a pharmacy procedure dataset for a patient, wherein each pharmacy procedure in the pharmacy procedure dataset comprises at least a procedure date, a therapy type, and a therapy delivery duration; identifying a subset of pharmacy procedures from the pharmacy procedure dataset that are associated with a biological condition based on the therapy type; calculating an end date for at least one pharmacy procedure in the pharmacy procedure subset, the procedure date corresponding to a date when the patient began payment, the end date being determined based on the procedure date and the therapy supply duration associated with the at least one pharmacy procedure; accessing, in the processing circuitry, from the memory, a medical procedure procedure dataset for the patient, wherein each medical procedure in the medical procedure dataset comprises at least a medical procedure date range and a medical procedure type; identifying a subset of medical procedure procedures from the medical procedure procedure dataset that are associated with the biological condition based on the medical procedure type; adjusting, for at least one medical procedure in the medical procedure procedure subset, a medical procedure date range based on a period during which the medical procedure type is valid or repeated; mapping the pharmacy procedure subset and the medical procedure procedure subset onto a timeline data structure stored in the memory, the timeline data structure storing pharmacy procedures and medical procedure procedures arranged by date; determining one or more therapy gaps within the timeline data structure in which no pharmacy procedures are present and no medical procedure procedures are present, each therapy gap comprising a number of consecutive days, the number of consecutive days being greater than a threshold number of days; determining one or more lines of therapy based on the one or more therapy gaps, each line of therapy comprising pharmacy procedures and medical procedure procedures occurring either between two therapy gaps, before an earliest temporary therapy gap, or after a latest temporary therapy gap, each line of therapy being associated with a line date range; transmitting a data structure identifying the patient, the one or more lines of therapy, and the line date range for each of the one or more lines of therapy to a data repository for storage therein; One or more non-transitory machine-readable media configured to cause

27. storing additional instructions that, when executed by the processing circuitry, cause the processing circuitry to: determining, within a single line of therapy, a first pharmacy procedure or medical procedure associated with a first biological condition stage and a second pharmacy procedure or medical procedure associated with a second biological condition stage, wherein a start date associated with the second pharmacy procedure or medical procedure is later than a start date associated with the first pharmacy procedure or medical procedure; dividing the single line of therapy into two lines of therapy using the start date associated with the second pharmacy procedure; and 27. The one or more non-transitory machine-readable media of claim 26, configured to cause:

28. the processing circuitry comprises a plurality of multi-threaded graphics processing units (GPUs); the patient is one of a plurality of patients; 27. The one or more non-transitory machine-readable media of claim 26, wherein the one or more lines of therapy for the patient are determined in parallel with determining lines of therapy for other patients from the plurality of patients using parallel threads of the multiple multi-threaded GPUs.

29. For at least one medical procedure procedure in the medical procedure procedure subset, the medical procedure date range comprises a medical procedure start date and a medical procedure end date, and the at least one medical procedure procedure is mapped onto the timeline data structure based on the medical procedure start date and the medical procedure end date; pharmacy procedures from the pharmacy procedure subset are mapped onto the timeline data structure based on the procedure date; 27. The one or more non-transitory machine-readable media of claim 26, wherein the at least one pharmacy procedure is mapped onto the timeline data structure based on the procedure date and the end date.