Joint modeling of longitudinal and time-to-event data to predict patient survival

By employing joint modeling of longitudinal and time-to-event data, the method effectively addresses the challenges of characterizing early cancer stages using ctDNA, improving predictive power for patient survival and treatment efficacy.

WO2025106263A1PCT designated stage expired Publication Date: 2025-05-22GUARDANT HEALTH INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/053540
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-10-30
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Current techniques for characterizing early stages of cancer using circulating tumor DNA (ctDNA) face challenges such as smaller numbers of aberrations, confounding phenomena like clonal non-tumorous tissue expansion, and the lack of understanding regarding the significance of driver alterations.

Method used

The method involves joint modeling of longitudinal and time-to-event data to predict patient survival, utilizing nucleic acid sequence information and biomarkers like ctDNA to determine patient responses. This includes the application of hierarchical random effects models and the use of databases containing medical and insurance records.

Benefits of technology

This approach enables improved characterization of early disease stages by deciphering temporal changes in biomarkers related to time-to-event responses, thereby enhancing the predictive power for patient survival and treatment efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000075_0001
    Figure IMGF000075_0001
  • Figure IMGF000076_0001
    Figure IMGF000076_0001
  • Figure IMGF000077_0001
    Figure IMGF000077_0001
Patent Text Reader

Abstract

Changes in ctDNA levels can fluctuate significantly over time from patient to patient, and the results can be difficult to interpret. Described herein are methods and techniques capable of capturing these complexities while accounting for a diverse set of patient traits. Furthermore, analyzing an observational dataset consisting of patients with cancer who received therapy, analytic results are capable of being presented graphically. These results demonstrate the utility of the described methods and techniques in acquiring a comprehensive understanding of how response patterns evolve and how different patient characteristics influence these evolutions.
Need to check novelty before this filing date? Find Prior Art

Description

JOINT MODELING OF LONGITUDINAL AND TIME-TO-EVENT DATA TOPREDICT PATIENT SURVIVALCROSS-REFERENCE TO RELATED APPLICATION

[0001] This patent application claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 599,353, filed on November 15, 2023, which are each incorporated by reference herein in their entirety.BACKGROUND

[0002] Today, there is increasing knowledge of the molecular pathogenesis of cancer and with next generation sequencing techniques, increases in the potential to study early molecular alterations in cancer development. This includes liquid biopsy in body fluids. Genetic and epigenetic alterations associated with cancer development can be found in cell-free DNA (cfDNA), such as those in plasma, serum, urine, etc. with the potential for use as diagnostic biomarkers. Non-invasive sampling methods foster patient compliance, as easier, faster, and more economical to perform.

[0003] Such liquid biopsy techniques support characterization of the genomic makeup of different tissues in the subject. While generally released by all types of cells, cfDNA can originate from necrotic or apoptotic cells, for identification of specific tumor-related alterations, such as mutations, methylation, and copy number variations (CNVs). Improved characterization of this circulating tumor DNA (ctDNA) is challenging given the need to differentiate the signal originating from a disease tissue, such as cancer, from signals originating from germline cells releasing cfDNA the wider range of tissues, such as healthy tissue and white blood cells undergoing hematopoiesis. One can enrich signals by identifying variant alleles having allele fractions that do not adhere to exemplary 1: 1 ratios for heterozygous alleles in the germline.

[0004] Despite these advances, the majority of cfDNA use as diagnostics focus on advanced tumor stages, with much less known as to characteristics of early malignant disease stages. Yet several hurdles exist for early stage detection, including smaller numbers of aberrations, confounding phenomena such as clonal non-tumorous tissue expansion, adventitious cancer-associated mutations, and lack of understanding as to significance of driver alterations. Thus, there is a great need in the art for improvedtechniques for characterizing early disease stages to support development of cfDNA and ctDNA related diagnostics.

[0005] Described herein is use of detection measurement, which can include a variety of parameters including longitudinal and time-to-event data, thereby supporting understanding of how temporal changes in a biomarker relate to a time-to-event response and patient outcomes. For example, methods and techniques described herein incorporate longitudinal and time-to-event data supports to decipher temporal changes in a biomarker as related to a time-to-event response. Additionally, methods and techniques described herein allows evaluation of patient characteristics such as age, gender, etc. in analyses. Repeated measures via liquid biopsy provide an opportunity to assess patient outcomes.SUMMARY OF THE INVENTION

[0006] Described herein is a method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker; and determining a patient response for the at least one patient. In various embodiments, the biomarker comprises ctDNA. In various embodiments, the biomarker comprises allele frequency and tumor fraction. In various embodiments, the method includes determining a patient response for the at least one patient comprises use of a database. In various embodiments, the method includes database comprises medical records and / or insurance records. In various embodiments, the method includes use of the database comprises application of a model. In various embodiments, the model is a hierarchal model. In various embodiments, the model is an effects model. In various embodiments, the model is a regression model. In various embodiments, the model is a joint model. In various embodiments, the hierarchal model is a hierarchical random effects model. In various embodiments, the model comprises a cubic spline. In various embodiments, the model comprises a regression model. In various embodiments, the hierarchal random effects model comprises generation of data from nucleic acid sequence information comprising temporal changes in a biomarker comprising circulating tumor DNA (ctDNA) from at least one subject in a plurality of subjects. In various embodiments, the generation of data comprises generation of a cubic spline for at least one subject in a plurality of subjects. In various embodiments, the generation of data comprises generation ofresponse parameters comprising one or more covariates. In various embodiments, the generation of data comprises generation of response parameters without covariates. In various embodiments, the response parameters apply a multivariate normal distribution. In various embodiments, the method includes determining a patient response for the at least one patient comprises generation of a velocity plot. In various embodiments, the method includes determining a patient response for the at least one patient comprises comparison to the model. In various embodiments, the joint model comprises at least two models. In various embodiments, the joint model comprises association factors between the at least two models. In various embodiments, the joint model comprises a cubic spline and a proportional hazard model. In various embodiments, the biomarker is measured with next-generation DNA sequencing. In various embodiments, nextgeneration DNA sequencing comprising ligation of non-unique barcodes to the ctDNA. In various embodiments, next-generation DNA sequencing comprising ligation of unique barcodes to the ctDNA. In various embodiments, next-generation DNA sequencing comprising ligation of non-unique barcodes to ctDNA fragments, wherein the nonunique barcodes are present in at least 20x, at least 30x, at least 50x, or at least lOOx molar excess.

[0007] A system comprising a machine comprising at least one processor and storage comprising instructions capable of performing any of the preceding methods. A computer readable medium comprising instructions capable of performing any of the preceding methods.

[0008] Described herein is a method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a hierarchal random effects model. In various embodiments, the hierarchal random effects model comprises generation of data from nucleic acid sequence information comprising temporal changes in ctDNA from at least one subject in a plurality of subjects. In various embodiments, the hierarchal random effects model comprises generation of a cubic spline for at least one subject in the plurality of subjects. In various embodiments, the hierarchal random effects model comprises response parameters comprising one ormore covariates for at least one subject in the plurality of subjects. In various embodiments, the database comprises medical records and / or insurance records for the plurality of subjects. Described herein is a system comprising a machine comprising at least one processor and storage comprising instructions capable of performing a method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a hierarchal random effects model. In various embodiments, the hierarchal random effects model comprises generation of data from nucleic acid sequence information comprising temporal changes in ctDNA from at least one subject in a plurality of subjects. In various embodiments, the hierarchal random effects model comprises generation of a cubic spline for at least one subject in the plurality of subjects. In various embodiments, the hierarchal random effects model comprises response parameters comprising one or more covariates for at least one subject in the plurality of subjects. In various embodiments, the database comprises medical records and / or insurance records for the plurality of subjects. Described herein is a computer readable medium comprising instructions capable of performing a method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a hierarchal random effects model. In various embodiments, the hierarchal random effects model comprises generation of data from nucleic acid sequence information comprising temporal changes in ctDNA from at least one subject in a plurality of subjects. In various embodiments, the hierarchal random effects model comprises generation of a cubic spline for at least one subject in the plurality of subjects. In various embodiments, the hierarchal random effects model comprises response parameters comprising one or more covariates for at least one subject in the plurality of subjects. In various embodiments, the database comprises medical records and / or insurance records for the plurality of subjects.

[0009] Described herein is a method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a joint model comprising a cubic spline and proportional hazard model generated from data from nucleic acid sequence information for at least one subject in a plurality of subjects. In various embodiments, the database comprises medical records and / or insurance records for the plurality of subjects. Described herein is a system comprising a machine comprising at least one processor and storage comprising instructions capable of performing is a method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a joint model comprising a cubic spline and proportional hazard model generated from data from nucleic acid sequence information for at least one subject in a plurality of subjects. In various embodiments, the database comprises medical records and / or insurance records for the plurality of subjects. Described herein is a computer readable medium comprising instructions capable of performing is a method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a joint model comprising a cubic spline and proportional hazard model generated from data from nucleic acid sequence information for at least one subject in a plurality of subjects. In various embodiments, the database comprises medical records and / or insurance records for the plurality of subjects.BRIEF DESCRIPTION OF THE FIGURES

[0010] Figure 1. Spaghetti Plots of Raw and Transformed TMSs Grouped by Deceased (No / Yes) and Disease Progression (No / Yes).

[0011] Figure 2. Dynamic Predictions for Two Different Patients — Predicting Overall Survival Probabilities Based on the Estimate of the Current tTMS.

[0012] Figure 3. Dynamic Predictions for Two Different Patients — Predicting Progression Free Survival Probabilities Based on the Estimate of the Current tTMS.

[0013] Figure 4. Dynamic Predictions for a Patient with Identical tTMSs, but Different Patient Characteristics.

[0014] Figure 5. Spaghetti Plots of Raw and Transformed TMSs and Density Plots of the Study Duration Distributions for Deceased Patients (Yes / No) and for Patients who Experience Disease Progression (Yes / No).

[0015] Figure 6. Dynamic Predictions for Two Different Patients — Predicting Overall Survival Probabilities Based on the Estimate of the Current tTMS.

[0016] Figure 7. Dynamic Predictions for Two Different Patients — Predicting Progression Free Survival Probabilities Based on the Estimate of the Current tTMS.

[0017] Figure 8. Dynamic Predictions for a Patient with Identical tTMSs, but Different Patient Characteristics — Younger and Healthier at Baseline.DETAILED DESCRIPTIONAnalysis

[0018] The present methods can be used to diagnose presence of conditions, particularly cancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally,if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.

[0019] The types and number of cancers that may be detected may include blood cancers, brain cancers, lung cancers, skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, solid state tumors, heterogeneous tumors, homogenous tumors and the like. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.

[0020] Genetic and other analyte data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.

[0021] The present analyses are also useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.

[0022] The present methods can also be used for detecting genetic variations in conditions other than cancer. Immune cells, such as B cells, may undergo rapid clonalexpansion upon the presence of certain diseases. Clonal expansions may be monitored using copy number variation detection and certain immune states may be monitored. In this example, copy number variation analysis may be performed over time to produce a profile of how a particular disease may be progressing. Copy number variation or even rare mutation detection may be used to determine how a population of pathogens changes during the course of infection. This may be particularly important during chronic infections, such as HIV / AIDS or Hepatitis infections, whereby viruses may change life cycle state and / or mutate into more virulent forms during the course of infection. The present methods may be used to determine or profile rejection activities of the host body, as immune cells attempt to destroy transplanted tissue to monitor the status of transplanted tissue as well as altering the course of treatment or prevention of rejection.

[0023] For example, numerous types of malfunctions and abnormalities that commonly occur in the cardiovascular system, wherein failure to diagnose or treat, will progressively decrease the body's ability to supply sufficient oxygen to satisfy the coronary oxygen demand when the individual encounters stress. The progressive decline in the cardiovascular system's ability to supply oxygen under stress conditions will ultimately culminate in a heart attack, i.e., myocardial infarction event that is caused by the interruption of blood flow through the heart resulting in oxygen starvation of the heart muscle tissue (i.e., myocardium). In many cases, permanent damage will occur to the cells comprising the myocardium that will subsequently predispose the individual's susceptibility to additional myocardial infarction events.

[0024] Methods of the disclosure can characterize malfunctions and abnormalities associated with the heart muscle and valve tissues (e.g., hypertrophy), the decreased supply of blood flow and oxygen supply to the heart are often secondary symptoms of debilitation and / or deterioration of the blood now and supply system caused by physical and biochemical stresses. Examples of cardiovascular diseases that are directly affected by these types of stresses include atherosclerosis, coronary artery disease, peripheral vascular disease and peripheral artery disease, along with various cardias and arrythmias which may represent other forms of disease and dysfunction.

[0025] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of extracellular polynucleotides derived from the subject,wherein the genetic profile includes a plurality of data resulting from copy number variation and rare mutation analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.

[0026] The present methods can be used to generate or profile, fingerprint or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation and mutation analyses alone or in combination.

[0027] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non- invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.Methods of Modified Nucleic Acid Analysis

[0028] The disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylated, linked to histones and other modifications discussed above). In some such methods, a population of nucleic acids bearing the modification to different extents (e.g., 0, 1, 2, 3, 4, 5 or more methyl groups per nucleic acid molecule) is contacted with adapters before fractionation of the population depending on the extent of the modification. Adapters attach to either one end or both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Following attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites within the adapters. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites, but preferably adapters include the same primer binding site. Following amplification,the nucleic acids are contacted with an agent that preferably binds to nucleic acids bearing the modification (such as the previously described such agents). The nucleic acids are separated into at least two partitions differing in the extent to which the nucleic acids bear the modification from binding to the agents. For example, if the agent has affinity for nucleic acids bearing the modification, nucleic acids overrepresented in the modification (compared with median representation in the population) preferentially bind to the agent, whereas nucleic acids underrepresented for the modification do not bind or are more easily eluted from the agent. Following separation, the different partitions can then be subject to further processing steps, which typically include further amplification, and sequence analysis, in parallel but separately. Sequence data from the different partitions can then be compared.

[0029] Nucleic acids can be linked at both ends to Y-shaped adapters including primer binding sites and tags. The molecules are amplified. The amplified molecules are then fractionated by contact with an antibody preferentially binding to 5-methylcytosine to produce two partitions. One partition includes original molecules lacking methylation and amplification copies having lost methylation. The other partition includes original DNA molecules with methylation. The two partitions are then processed and sequenced separately with further amplification of the methylated partition. The sequence data of the two partitions can then be compared. In this example, tags are not used to distinguish between methylated and unmethylated DNA but rather to distinguish between different molecules within these partitions so that one can determine whether reads with the same start and stop points are based on the same or different molecules.

[0030] The disclosure provides further methods for analyzing a population of nucleic acid in which at least some of the nucleic acids include one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications described previously. In these methods, the population of nucleic acids is contacted with adapters including one or more cytosine residues modified at the 5C position, such as 5- methylcytosine. Preferably all cytosine residues in such adapters are also modified, or all such cytosines in a primer binding region of the adapters are modified. Adapters attach to both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. The primer binding sites in suchadapters can be the same or different, but are preferably the same. After attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites of the adapters. The amplified nucleic acids are split into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. The sequence data on molecules in the first aliquot is thus determined irrespective of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosines to uracils. The bisulfite treated nucleic acids are then subjected to amplification primed by primers to the original primer binding sites of the adapters linked to nucleic acid. Only the nucleic acid molecules originally linked to adapters (as distinct from amplification products thereof) are now amplifiable because these nucleic acids retain cytosines in the primer binding sites of the adapters, whereas amplification products have lost the methylation of these cytosine residues, which have undergone conversion to uracils in the bisulfite treatment. Thus, only original molecules in the populations, at least some of which are methylated, undergo amplification. After amplification, these nucleic acids are subject to sequence analysis. Comparison of sequences determined from the first and second aliquots can indicate among other things, which cytosines in the nucleic acid population were subject to methylation.Partitioning the Sample into a Plurality of Subsamples; Aspects of Samples; Analysis of Epigenetic Characteristics

[0031] In certain embodiments described herein, a population of different forms of nucleic acids (e.g., hypermethylated and hypomethylated DNA in a sample, such as a captured set of cfDNA as described herein) can be physically partitioned based on one or more characteristics of the nucleic acids prior to further analysis, e.g., differentially modifying or isolating a nucleobase, tagging, and / or sequencing. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated. In some embodiments, hypermethylation variable epigenetic target regions are analyzed to determine whether they show hypermethylation characteristic of tumor cells and / or hypomethylation variable epigenetic target regions are analyzed to determine whether they show hypomethylation characteristic of tumor cells. Additionally, by partitioning a heterogeneous nucleic acid population, one may increase rare signals, e.g., by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, a genetic variation present inhyper-methylated DNA but less (or not) in hypomethylated DNA can be more easily detected by partitioning a sample into hyper-methylated and hypo-methylated nucleic acid molecules. By analyzing multiple fractions of a sample, a multi-dimensional analysis of a single locus of a genome or species of nucleic acid can be performed and hence, greater sensitivity can be achieved.

[0032] In some instances, a heterogeneous nucleic acid sample is partitioned into two or more partitions (e.g., at least 3, 4, 5, 6 or 7 partitions). In some embodiments, each partition is differentially tagged. Tagged partitions can then be pooled together for collective sample prep and / or sequencing. The partitioning-tagging-pooling steps can occur more than once, with each round of partitioning occurring based on a different characteristics (examples provided herein) and tagged using differential tags that are distinguished from other partitions and partitioning means.

[0033] Examples of characteristics that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. Resulting partitions can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), doublestranded DNA (dsDNA), shorter DNA fragments and longer DNA fragments. In some embodiments, partitioning based on a cytosine modification (e.g., cytosine methylation) or methylation generally is performed and is optionally combined with at least one additional partitioning step, which may be based on any of the foregoing characteristics or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids with one or more epigenetic modifications and without the one or more epigenetic modifications. Examples of epigenetic modifications include presence or absence of methylation; level of methylation; type of methylation (e.g., 5- methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and association and level of association with one or more proteins, such as histones. Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules devoid of nucleosomes. Alternatively or additionally, a heterogeneous population of nucleic acids may be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or additionally, a heterogeneous population of nucleic acids may be partitioned based on nucleic acidlength (e.g., molecules of up to 160 bp and molecules having a length of greater than 160 bp).

[0034] In some instances, each partition (representative of a different nucleic acid form) is differentially labelled, and the partitions are pooled together prior to sequencing. In other instances, the different forms are separately sequenced. In some embodiments, a population of different nucleic acids is partitioned into two or more different partitions. Each partition is representative of a different nucleic acid form, and a first partition (also referred to as a subsample) includes DNA with a cytosine modification in a greater proportion than a second subsample. Each partition is distinctly tagged. The first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. The tagged nucleic acids are pooled together prior to sequencing. Sequence reads are obtained and analyzed, including to distinguish the first nucleobase from the second nucleobase in the DNA of the first subsample, in silico. Tags are used to sort reads from different partitions. Analysis to detect genetic variants can be performed on a partiti on-by-partition level, as well as whole nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variants, such as CNV, SNV, indel, fusion in nucleic acids in each partition. In some instances, in silico analysis can include determining chromatin structure. For example, coverage of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or nucleosome depleted region (NDR).

[0035] Samples can include nucleic acids varying in modifications including postreplication modifications to nucleotides and binding, usually noncovalently, to one or more proteins.

[0036] In an embodiment, the population of nucleic acids is one obtained from a serum, plasma or blood sample from a subject suspected of having neoplasia, a tumor, or cancer or previously diagnosed with neoplasia, a tumor, or cancer. The population of nucleic acids includes nucleic acids having varying levels of methylation. Methylation can occur from any one or more post-replication or transcriptional modifications. Post-replicationmodifications include modifications of the nucleotide cytosine, particularly at the 5- position of the nucleobase, e.g., 5-methylcytosine, 5-hydroxymethylcytosine, 5- formylcytosine and 5-carboxylcytosine. The affinity agents can be antibodies with the desired specificity, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected e.g., by phage display to have specificity to a given target.

[0037] Examples of capture moieties contemplated herein include methyl binding domain (MBDs) and methyl binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies preferentially binding to 5-methylcytosine. Likewise, partitioning of different forms of nucleic acids can be performed using histone binding proteins which can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48 and SANT domain peptides. Although for some affinity agents and modifications, binding to the agent may occur in an essentially all or none manner depending on whether a nucleic acid bears a modification, the separation may be one of degree. In such instances, nucleic acids overrepresented in a modification bind to the agent at a greater extent that nucleic acids underrepresented in the modification. Alternatively, nucleic acids having modifications may bind in an all or nothing manner. But then, various levels of modifications may be sequentially eluted from the binding agent.

[0038] For example, in some embodiments, partitioning can be binary or based on degree / level of modifications. For example, all methylated fragments can be partitioned from unmethylated fragments using methyl-binding domain proteins (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning may involve eluting fragments having different levels of methylation by adjusting the salt concentration in a solution with the methyl-binding domain and bound fragments. As salt concentration increases, fragments having greater methylation levels are eluted. In some instances, the final partitions are representative of nucleic acids having different extents of modifications (overrepresentative or underrepresentative of modifications). Overrepresentation and underrepresentation can be defined by the number of modifications born by a nucleic acid relative to the median number of modifications per strand in a population. For example, if the median number of 5- methylcytosine residues in nucleic acid in a sample is 2, a nucleic acid including morethan two 5-methylcytosine residues is overrepresented in this modification and a nucleic acid with 1 or zero 5-methylcytosine residues is underrepresented. The effect of the affinity separation is to enrich for nucleic acids overrepresented in a modification in a bound phase and for nucleic acids underrepresented in a modification in an unbound phase (i.e. in solution). The nucleic acids in the bound phase can be eluted before subsequent processing.

[0039] When using MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific) various levels of methylation can be partitioned using sequential elutions. For example, a hypomethylated partition (e.g., no methylation) can be separated from a methylated partition by contacting the nucleic acid population with the MBD from the kit, which is attached to magnetic beads. The beads are used to separate out the methylated nucleic acids from the non- methylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids having different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is once again used to separate higher levels of methylated nucleic acids from those with lower level of methylation. The elution and magnetic separation steps can repeat themselves to create various partitions such as a hypomethylated partition (representative of no methylation), a methylated partition (representative of low level of methylation), and a hyper methylated partition (representative of high level of methylation).

[0040] In some methods, nucleic acids bound to an agent used for affinity separation are subjected to a wash step. The wash step washes off nucleic acids weakly bound to the affinity agent. Such nucleic acids can be enriched in nucleic acids having the modification to an extent close to the mean or median (i.e., intermediate between nucleic acids remaining bound to the solid phase and nucleic acids not binding to the solid phase on initial contacting of the sample with the agent). The affinity separation results in at least two, and sometimes three or more partitions of nucleic acids with different extents of a modification. While the partitions are still separate, the nucleic acids of at least one partition, and usually two or three (or more) partitions are linked to nucleic acid tags, usually provided as components of adapters, with the nucleic acids in different partitionsreceiving different tags that distinguish members of one partition from another. The tags linked to nucleic acid molecules of the same partition can be the same or different from one another. But if different from one another, the tags may have part of their code in common so as to identify the molecules to which they are attached as being of a particular partition. For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some embodiments, the nucleic acid molecules can be fractionated into different partitions based on the nucleic acid molecules that are bound to a specific protein or a fragment thereof and those that are not bound to that specific protein or fragment thereof.

[0041] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on a specific property of a protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation) or enzymatic activity. Examples of proteins which may bind to DNA and serve as a basis for fractionation may include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate the nucleic acid molecules based on protein bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein bound regions include, but are not limited to, SDS-PAGE, chromatin-immuno-precipitation (ChIP), heparin chromatography, and asymmetrical field flow fractionation (AF4).

[0042] In some embodiments, partitioning of the nucleic acids is performed by contacting the nucleic acids with a methylation binding domain (“MBD”) of a methylation binding protein (“MBP”). MBD binds to 5-methylcytosine (5mC). MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different extents of methylation can be performed by eluting fractions by increasing the NaCl concentration.

[0043] An exemplary method for molecular tag identification of MBD-bead partitioned libraries through NGS is as follows:

[0044] Physical partitioning of an extracted DNA sample (e.g., extracted blood plasma DNA from a human sample) using a methyl -binding domain protein-bead purification kit, saving all elutions from process for downstream processing.

[0045] Parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, the hypermethylated, residual methylation('wash'), and hypomethylated partitions are ligated with NGS-adapters with molecular tags.

[0046] Re-combining all molecular tagged partitions, and subsequent amplification using adapter-specific DNA primer sequences.

[0047] Enrichment / hybridization of re-combined and amplified total library, targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).

[0048] Re-amplification of the enriched total DNA library, appending a sample tag. Different samples are pooled and assayed in multiplex on an NGS instrument.

[0049] Bioinformatics analysis of NGS data, with the molecular tags being used to identify unique molecules, as well deconvolution of the sample into molecules that were differentially MBD-partitioned. This analysis can yield information on relative 5- methylcytosine for genomic regions, concurrent with standard genetic sequencing / variant detection.

[0050] Examples of MBPs contemplated herein include, but are not limited to:

[0051] (a) MeCP2 is a protein preferentially binding to 5-methyl-cytosine over unmodified cytosine.

[0052] (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 preferentially bind to 5- hydroxymethyl -cytosine over unmodified cytosine.

[0053] (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind to 5-formyl- cytosine over unmodified cytosine (lurlaro et al., Genome Biol. 14: R119 (2013)).

[0054] (d) Antibodies specific to one or more methylated nucleotide bases.

[0055] In general, elution is a function of number of methylated sites per molecule, with molecules having more methylation eluting under increased salt concentrations. To elute the DNA into distinct populations based on the extent of methylation, one can use a series of elution buffers of increasing NaCl concentration. Salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three (3) partitions. Molecules are contacted with a solution at a first salt concentration and including a molecule including a methyl binding domain, which molecule can be attached to a capture moiety, such as streptavidin. At the first salt concentration a population of molecules will bind to the MBD and a population will remain unbound. The unbound population can be separated as a “hypomethylated” population. Forexample, a first partition representative of the hypomethylated form of DNA is that which remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. A second partition representative of intermediate methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM concentration. This is also separated from the sample. A third partition representative of hypermethylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.

[0056] The disclosure provides further methods for analyzing a population of nucleic acids in which at least some of the nucleic acids include one or more modified cytosine residues, such as 5 -methylcytosine and any of the other modifications described previously. In these methods, after partitioning, the subsamples of nucleic acids are contacted with adapters including one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably all cytosine residues in such adapters are also modified, or all such cytosines in a primer binding region of the adapters are modified. Adapters attach to both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. The primer binding sites in such adapters can be the same or different, but are preferably the same. After attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites of the adapters. The amplified nucleic acids are split into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. The sequence data on molecules in the first aliquot is thus determined irrespective of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase includes a cytosine modified at the 5 position, and the second nucleobase includes unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. The nucleic acids subjected to the procedure are then amplified with primers to the original primer binding sites of the adapters linked to nucleic acid. Only the nucleic acid molecules originally linked to adapters (as distinct from amplification products thereof) are now amplifiable because these nucleic acids retain cytosines in the primer binding sites of the adapters, whereas amplification products have lost the methylation of these cytosine residues,which have undergone conversion to uracils in the bisulfite treatment. Thus, only original molecules in the populations, at least some of which are methylated, undergo amplification. After amplification, these nucleic acids are subject to sequence analysis. Comparison of sequences determined from the first and second aliquots can indicate among other things, which cytosines in the nucleic acid population were subject to methylation.

[0057] Such an analysis can be performed using the following exemplary procedure. After partitioning, methylated DNA is linked to Y-shaped adapters at both ends including primer binding sites and tags. The cytosines in the adapters are modified at the 5 position (e.g., 5-methylated). The modification of the adapters serves to protect the primer binding sites in a subsequent conversion step (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but affects unmodified cytosine). After attachment of adapters, the DNA molecules are amplified. The amplification product is split into two aliquots for sequencing with and without conversion. The aliquot not subjected to conversion can be subjected to sequence analysis with or without further processing. The other aliquot is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase includes a cytosine modified at the 5 position, and the second nucleobase includes unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. Only primer binding sites protected by modification of cytosines can support amplification when contacted with primers specific for original primer binding sites. Thus, only original molecules and not copies from the first amplification are subjected to further amplification. The further amplified molecules are then subjected to sequence analysis. Sequences can then be compared from the two aliquots. As in the separation scheme discussed above, nucleic acid tags in adapters are not used to distinguish between methylated and unmethylated DNA but to distinguish nucleic acid molecules within the same partition.Subjecting the First Subsample to a Procedure that Affects a First Nucleobase in the DNA Differently from a Second Nucleobase in the DNA of the First Subsample

[0058] Methods disclosed herein comprise a step of subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modifiedor unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, if the first nucleobase is a modified or unmodified adenine, then the second nucleobase is a modified or unmodified adenine; if the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine; if the first nucleobase is a modified or unmodified guanine, then the second nucleobase is a modified or unmodified guanine; and if the first nucleobase is a modified or unmodified thymine, then the second nucleobase is a modified or unmodified thymine (where modified and unmodified uracil are encompassed within modified thymine for the purpose of this step).

[0059] In some embodiments, the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine. For example, first nucleobase may comprise unmodified cytosine (C) and the second nucleobase may comprise one or more of 5 -methyl cytosine (mC) and 5 -hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, as indicated, e.g., in the Summary above and the following discussion, such as where one of the first and second nucleobases includes mC and the other includes hmC.

[0060] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g. 5-formyl cytosine (fC) or 5-carboxylcytosine (caC)) to uracil whereas other modified cytosines (e.g., 5 -methylcytosine, 5-hydroxylmethylcystosine) are not converted. Thus, where bisulfite conversion is used, the first nucleobase includes one or more of unmodified cytosine, 5-formyl cytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of bisulfite-treated DNA identifies positions that are read as cytosine as being mC or hmC positions. Meanwhile, positions that are read as T are identified as being T or a bisulfite- susceptible form of C, such as unmodified cytosine, 5-formyl cytosine, or 5- carboxylcytosine. Performing bisulfite conversion on a first subsample as described herein thus facilitates identifying positions containing mC or hmC using the sequencereads obtained from the first subsample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068..

[0061] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes oxidative bisulfite (Ox-BS) conversion. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes Tet-assisted bisulfite (TAB) conversion. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes Tet-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes chemical-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes APOBEC-coupled epigenetic (ACE) conversion.

[0062] In some embodiments, procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692vl. For example, TET2 and T4- PGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3 A), and then a deaminase (e.g., APOBEC3 A) can be used to deaminate unmodified cytosines converting them to uracils.

[0063] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample includes separating DNA originally including the first nucleobase from DNA not originally including the first nucleobase.

[0064] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).

[0065] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases such as mA from other DNA. See, e.g., Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161 : 868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies for various modified nucleobases, such as forms of thymine / uracil including halogenated forms such as 5 -bromouracil, are commercially available. Various modified bases can also be detected based on alterations in their base-pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read in sequencing as a G. See, e.g., US Patent 8,486,630;Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, N.Y., 2002, chapter 14, “Mutation, Repair, and Recombination.”Enriching / Capturing Step, Amplification., Adaptors, Barcodes

[0066] In some embodiments, methods disclosed herein comprise a step of capturing one or more sets of target regions of DNA, such as cfDNA. Capture may be performed using any suitable approach known in the art. In some embodiments, capturing includes contacting the DNA to be captured with a set of target-specific probes. The set of targetspecific probes may have any of the features described herein for sets of target-specific probes, including but not limited to in the embodiments set forth above and the sections relating to probes below. Capturing may be performed on one or more subsamples prepared during methods disclosed herein. In some embodiments, DNA is captured from at least the first subsample or the second subsample, e.g., at least the first subsample and the second subsample. Where the first subsample undergoes a separation step (e.g., separating DNA originally including the first nucleobase (e.g., hmC) from DNA not originally including the first nucleobase, such as hmC-seal), capturing may be performed on any, any two, or all of the DNA originally including the first nucleobase (e.g., hmC), the DNA not originally including the first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.

[0067] The capturing step may be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on features of the probes such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given general knowledge in the art regarding nucleic acid hybridization. In some embodiments, complexes of target-specific probes and DNA are formed.

[0068] In some embodiments, a method described herein includes capturing cfDNA obtained from a test subject for a plurality of sets of target regions. The target regions comprise epigenetic target regions, which may show differences in methylation levels and / or fragmentation patterns depending on whether they originated from a tumor or from healthy cells. The target regions also comprise sequence-variable target regions, which may show differences in sequence depending on whether they originated from a tumor or from healthy cells. The capturing step produces a captured set of cfDNA molecules, and the cfDNA molecules corresponding to the sequence-variable target region set are captured at a greater capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the epigenetic target region set. For additional discussion of capturing steps, capture yields, and related aspects, see W02020 / 160414, which is incorporated herein by reference for all purposes.

[0069] In some embodiments, a method described herein includes contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of targetspecific probes is configured to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set.

[0070] It can be beneficial to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set because a greater depth of sequencing may be necessary to analyze the sequence-variable target regions with sufficient confidence or accuracy than may be necessary to analyze the epigenetic target regions. The volume of data needed to determine fragmentation patterns (e.g., to test fsor perturbation of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated partitions) is generally less than the volume of data needed to determine the presence or absence of cancer-related sequence mutations. Capturing the target region sets at different yields can facilitate sequencing the target regions to differentdepths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).

[0071] In various embodiments, the methods further comprise sequencing the captured cfDNA, e.g., to different degrees of sequencing depth for the epigenetic and sequencevariable target region sets, consistent with the discussion herein. In some embodiments, complexes of target-specific probes and DNA are separated from DNA not bound to target-specific probes. For example, where target-specific probes are bound covalently or noncovalently to a solid support, a washing or aspiration step can be used to separate unbound material. Alternatively, where the complexes have chromatographic properties distinct from unbound material (e.g., where the probes comprise a ligand that binds a chromatographic resin), chromatography can be used.

[0072] As discussed in detail elsewhere herein, the set of target-specific probes may comprise a plurality of sets such as probes for a sequence-variable target region set and probes for an epigenetic target region set. In some such embodiments, the capturing step is performed with the probes for the sequence-variable target region set and the probes for the epigenetic target region set in the same vessel at the same time, e.g., the probes for the sequence-variable and epigenetic target region sets are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the sequence-variable target region set is greater that the concentration of the probes for the epigenetic target region set.

[0073] Alternatively, the capturing step is performed with the sequence-variable target region probe set in a first vessel and with the epigenetic target region probe set in a second vessel, or the contacting step is performed with the sequence-variable target region probe set at a first time and a first vessel and the epigenetic target region probe set at a second time before or after the first time. This approach allows for preparation of separate first and second compositions including captured DNA corresponding to the sequence-variable target region set and captured DNA corresponding to the epigenetic target region set. The compositions can be processed separately as desired (e.g., to fractionate based on methylation as described elsewhere herein) and recombined in appropriate proportions to provide material for further processing and analysis such as sequencing.

[0074] In some embodiments, the DNA is amplified. In some embodiments, amplification is performed before the capturing step. In some embodiments, amplification is performed after the capturing step.

[0075] In some embodiments, adapters are included in the DNA. This may be done concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer, e.g., as described above. Alternatively, adapters can be added by other approaches, such as ligation.

[0076] In some embodiments, tags, which may be or include barcodes, are included in the DNA. Tags can facilitate identification of the origin of a nucleic acid. For example, barcodes can be used to allow the origin (e.g., subject) whence the DNA came to be identified following pooling of a plurality of samples for parallel sequencing. This may be done concurrently with an amplification procedure, e.g., by providing the barcodes in a 5’ portion of a primer, e.g., as described above. In some embodiments, adapters and tags / barcodes are provided by the same primer or primer set. For example, the barcode may be located 3’ of the adapter and 5’ of the target-hybridizing portion of the primer. Alternatively, barcodes can be added by other approaches, such as ligation, optionally together with adapters in the same ligation substrate.

[0077] Additional details regarding amplification, tags, and barcodes are discussed in the “General Features of the Methods” section below, which can be combined to the extent practicable with any of the foregoing embodiments and the embodiments set forth in the introduction and summary section.Computer Systems, Processing of Real World Evidence (RWE)

[0078] Methods of the present disclosure can be implemented using, or with the aid of, computer systems. For example, such methods may comprise: partitioning the sample into a plurality of subsamples, including a first subsample and a second subsample, wherein the first subsample includes DNA with a cytosine modification in a greater proportion than the second subsample; subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and sequencing DNA in the first subsample and DNA inthe second subsample in a manner that distinguishes the first nucleobase from the second nucleobase in the DNA of the first subsample.

[0079] In an aspect, the present disclosure provides a non-transitory computer-readable medium including computer-executable instructions which, when executed by at least one electronic processor, perform at least a portion of a method including: collecting cfDNA from a test subject; capturing a plurality of sets of target regions from the cfDNA, wherein the plurality of target region sets includes a sequence-variable target region set and an epigenetic target region set, whereby a captured set of cfDNA molecules is produced; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a greater depth of sequencing than the captured cfDNA molecules of the epigenetic target region set; obtaining a plurality of sequence reads generated by a nucleic acid sequencer from sequencing the captured cfDNA molecules; mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence-variable target region set and to the epigenetic target region set to determine the likelihood that the subject has cancer.

[0080] The code can be pre-compiled and configured for use with a machine with a processer adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.

[0081] Additional details relating to computer systems and networks, databases, and computer program products are also provided in, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety. Further information is found in PCT Pub. No. US2022032250 and U.S. App. No. 17832498.

[0082] Described herein is a method to generate an integrated data repository and / or analysis system that includes multiple types of healthcare data, according to one or more implementations. The architecture may include a data integration and / or analysis system.The data integration and analysis system may obtain data from a number of data sources and integrate the data from the data sources into an integrated data repository. For example, the data integration and analysis system may obtain data from a health insurance claims data repository. In various examples, the data integration and analysis system and the health insurance claims data repository may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system and the health insurance claims data repository may be created and maintained by the same entity.

[0083] The data integration and analysis system may be implemented by one or more computing devices. The one or more computing devices may include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or combinations thereof. In certain implementations, at least a portion of the one or more computing devices may be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices may be implemented in a cloud computing architecture. In scenarios where the computing systems used to implement the data integration and analysis system are configured in a distributed computing architecture, processing operations may be performed concurrently by multiple virtual machines. In various examples, the data integration and analysis system may implement multithreading techniques. The implementation of a distributed computing architecture and multithreading techniques cause the data integration and analysis system to utilize fewer computing resources in relation to computing architectures that do not implement these techniques.

[0084] The health insurance claims data repository may store information obtained from one or more health insurance companies that corresponds to insurance claims made by subscribers of the one or more health insurance companies. The health insurance claims data repository may be arranged (e.g., sorted) by patient identifier. The patient identifier may be based on the patient’s first name, last name, date of birth, social security number, address, employer, and the like. The data stored by the health insurance claims data repository may include structured data that is arranged in one or more data tables. The one or more data tables storing the structured data may include a number of rows and a number of columns that indicate information about health insurance claims made by subscribers of one or more health insurance companies in relation to proceduresand / or treatments received by the subscribers from healthcare providers. At least a portion of the rows and columns of the data tables stored by the health insurance claims data repository may include health insurance codes that may indicate diagnoses of biological conditions, and treatments and / or procedures obtained by subscribers of the one or more health insurance companies. In various examples, the health insurance codes may also indicate diagnostic procedures obtained by individuals that are related to one or more biological conditions that may be present in the individuals. In one or more examples, a diagnostic procedure may provide information used in the detection of the presence of a biological condition. A diagnostic procedure may also provide information used to determine a progression of a biological condition. In one or more illustrative examples, a diagnostic procedure may include one or more imaging procedures, one or more assays, one or more laboratory procedures, one or more combinations thereof, and the like.

[0085] The data integration and analysis system may also obtain information from a molecular data repository. The molecular data repository may store data of a number of individuals related to genomic information, genetic information, metabolomic information, transcriptomic information, fragmentomic information, immune receptor information, methylation information, epigenomic information, and / or proteomic information. In one or more examples, the data integration and analysis system and the molecular data repository may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system and the molecular data repository may be created and maintained by a same entity.

[0086] The genomic and / or epigenomic information may indicate one or more mutations corresponding to genes of the individuals. A mutation to a gene of individuals may correspond to differences between a sequence of nucleic acids of the individuals and one or more reference genomes. The reference genome may include a known reference genome, such as hgl9. In various examples, a mutation of a gene of an individual may correspond to a difference in a germline gene of an individual in relation to the reference genome. In one or more additional examples, the reference genome may include a germline genome of an individual. In one or more further examples, a mutation to a gene of an individual may include a somatic mutation. Mutations to genes of individuals may be related to insertions, deletions, single nucleotide variants, loss ofheterozygosity, duplication, amplification, translocation, fusion genes, or one or more combinations thereof.

[0087] In one or more illustrative examples, genomic and / or epigenomic information stored by the molecular data repository may include genomic and / or epigenomic profiles of tumor cells present within individuals. In these situations, the genomic and / or epigenomic information may be derived from an analysis of genetic material, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA) from a sample, including, but not limited to, a tissue sample or tumor biopsy, circulating tumor cells (CTCs), exosomes or efferosomes, or from circulating nucleic acids (e.g., cell-free DNA) found in blood samples of individuals that is present due to the degradation of tumor cells present in the individuals. . In one or more examples, the genomic and / or epigenomic information of tumor cells of individuals may correspond to one or more target regions. One or more mutations present with respect to the one or more target regions may indicate the presence of tumor cells in individuals. The genomic and / or epigenomic information stored by the molecular data repository may be generated in relation to an assay or other diagnostic test that may determine one or more mutations with respect to one or more target regions of the reference genome.

[0088] The number of data tables may be arranged according to a data repository schema. In the illustrative example of the data repository schema includes a first data table, a second data table, a third data table , a fourth data table , and a fifth data table . Although the illustrative example of ncludes five data tables, in additional implementations, the data repository schema may include more data tables or fewer data tables. The data repository schema may also include links between the data tables. The links between the data tables may indicate that information retrieved from one of the data tables results in additional information stored by one or more additional data tables to be retrieved. Additionally, not all the data tables may be linked to each of the other data tables. In the illustrative example of the first data table is logically coupled to the second data table by a first link and the first data table is logically coupled to the fourth data table by a second link . In addition, the second data table is logically coupled to the third data table via a third link and the fourth data table is logically coupled to the fifth data table via a fourth link . Further, the third data table is logically coupled to the fifth data table via a fifth link.

[0089] In various examples, as data tables are added to and / or removed from the data repository schema, additional links between data tables may be added to or removed from the data repository schema. In one or more illustrative examples, the integrated data repository may store data tables according to the data repository schema for at least a portion of the individuals for which the data integration system obtained information from a combination of at least two of the health insurance claims data repository , the molecular data repository , the one or more additional data repositories , and the one or more reference information data repositories . As a result, the integrated data repository may store respective instances of the data tables according to the data repository schema for thousands, tens of thousands, up to hundreds of thousands or more individuals.

[0090] The data integration and analysis system may also include a data pipeline system. The data pipeline system may include a number of algorithms, software code, scripts, macros, or other bundles of computer-executable instructions that process information stored by the integrated data repository to generate additional datasets. The additional datasets may include information obtained from one or more of the data tables. The additional datasets may also include information that is derived from data obtained from one or more of the data tables. The components of the data pipeline system implemented to generate a first additional dataset may be different from the components of the data pipeline system used to generate a second additional dataset.

[0091] In one or more examples, the data pipeline system may generate a dataset that indicates pharmacy treatments received by a number of individuals. In one or more illustrative examples, the data pipeline system may analyze information stored in at least one of the data tables to determine health insurance codes corresponding to pharmaceutical treatments received by a number of individuals. The data pipeline system may analyze the health insurance codes corresponding to pharmaceutical treatments with respect to a library of data that indicates specified pharmaceutical treatments that correspond to one or more health insurance codes to determine names of pharmaceutical treatments that have been received by the individuals. In one or more additional examples, the data pipeline system may analyze information stored by the integrated data repository to determine medical procedures received by a number of individuals. To illustrate, the data pipeline system may analyze information stored by one of the data tables to determine treatments received by individuals via at least one injection or intravenously. In one or more further examples, the data pipeline system may analyzeinformation stored by the integrated data repository to determine episodes of care for individuals, lines of therapy received by individuals, progression of a biological condition, or time to next treatment. In various examples, the datasets generated by the data pipeline system may be different for different biological conditions. For example, the data pipeline system may generate a first number of datasets with respect to a first type of cancer, such as lung cancer, and a second number of datasets with respect to a second type of cancer, such as colorectal cancer.

[0092] The data pipeline system may also determine one or more confidence levels to assign to information associated with individuals having data stored by the integrated data repository. The respective confidence levels may correspond to different measures of accuracy for information associated with individuals having data stored by the integrated data repository. The information associated with the respective confidence levels may correspond to one or more characteristics of individuals derived from data stored by the integrated data repository. Values of confidence levels for the one or more characteristics may be generated by the data pipeline system in conjunction with generating one or more datasets from the integrated data repository. In one or more examples, a first confidence level may correspond to a first range of measures of accuracy, a second confidence level may correspond to a second range of measures of accuracy, and a third confidence level may correspond to a third range of measures of accuracy. In one or more additional examples, the second range of measures of accuracy may include values that are less values of the first range of measures of accuracy and the third range of measures of accuracy may include values that are less than values of the second range of measures of accuracy. In one or more illustrative examples, information corresponding to the first confidence level may be referred to as Gold standard information, information corresponding to the second confidence level may be referred to as Silver standard information, and information corresponding to the third confidence level may be referred to as Bronze standard information.

[0093] The data pipeline system may determine values for the confidence levels of characteristics of individuals based on a number of factors. For example, a respective set of information may be used to determine characteristics of individuals. The data pipeline system may determine the confidence levels of characteristics of individuals based on an amount of completeness of the respective set of information used to determine a characteristic for an individual. In situations where one or more pieces of information aremissing from the set of information associated with a first number of individuals, the confidence levels for a characteristic may be lower than for a second number of individuals where information is not missing from the set of information. In one or more examples, an amount of missing information may be used by the data pipeline system to determine confidence levels of characteristics of individuals. To illustrate, a greater amount of missing information used to determine a characteristic of an individual may cause confidence levels for the characteristic to be lower than in situations where the amount of missing information used to determine the characteristic is lower. Further, different types of information may correspond to various confidence levels for a characteristic. In one or more examples, the presence of a first piece of information used to determine a characteristic of an individual may result in confidence levels for the characteristic being higher than the presence of a second piece of information used to determine the characteristic.

[0094] In one or more illustrative examples, the data pipeline system may determine a number of individuals included in a cohort with a primary diagnosis of lung cancer (or other biological condition). The data pipeline system may determine confidence levels for respective individuals with respect to being classified as having a primary diagnosis of lung cancer. The data pipeline system may use information from a number of columns included in the data tables to determine a confidence level for the inclusion of individuals within a lung cancer cohort. The number of columns may include health insurance codes related to diagnosis of biological conditions and / or treatments of biological conditions. Additionally, the number of columns may correspond to dates of diagnosis and / or treatment for biological conditions. The data pipeline system may determine that a confidence level of an individual being characterized as being part of the lung cancer cohort is higher in scenarios where information is available for each of the number of columns or at least a threshold number of columns than in instances where information is available for less than a threshold number of columns. Further, the data pipeline system may determine confidence levels for individuals included in a lung cancer cohort based on the type of information and availability of information associated with one or more columns. To illustrate, in situations where one or more diagnosis codes are present in relation to one or more periods of time for a group of individuals and one or more treatment codes are absent, the data pipeline system may determine that the confidence level of including the group of individuals in the lung cancer cohort is greaterthan in situations where at least one of the diagnosis codes is absent and the treatment codes used to determine whether individuals are included in the lung cancer cohort are present.

[0095] The data analysis system may receive integrated data repository requests from one or more computing devices, such as an example computing device. The one or more integrated data repository requests may cause data to be retrieved from the integrated data repository. In various examples, the one or more integrated data repository requests may cause data to be retrieved from one or more datasets generated by the data pipeline system. The integrated data repository requests may specify the data to be retrieved from the integrated data repository and / or the one or more datasets generated by the data pipeline system. In one or more additional examples, the integrated data repository requests may include one or more prebuilt queries that correspond to computerexecutable instructions that retrieve a specified set of data from the integrated data repository and / or one or more datasets generated by the data pipeline system.

[0096] In response to one or more integrated data repository requests, the data analysis system may analyze data retrieved from at least one of the integrated data repository or one or more datasets generated by the data pipeline system to generate data analysis results . The data analysis results may be sent to one or more computing devices, such as example computing devices. Although the illustrative example of hows that the one or more integrated data repository requests from one computing device and the data analysis results being sent to another computing device , in one or more additional implementations, the data analysis results may be received by a same computing device that sent the one or more integrated data repository requests . The data analysis results may be displayed by one or more user interfaces rendered by the computing device or the computing device.

[0097] Described herein is a method for analysis of nucleic acid sequence information. In various embodiments, the method of analysis comprises one or more models, each of one or more including one or more of survival, sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. In various embodiments, a model includes hierarchal models (e.g., nested models, multilevel models), mixed models (e.g., regression such as logistic regression and Poisson regression, pooled, random effect, fixed effect, mixed effect, linear mixed effect,generalized linear mixed effect), hazard model, odds ratio models and / or repeated sample (e.g., repeated measures such as ANOVA). In various embodiments, the model is a hierarchical random effects model. In various embodiments, the model is a hierarchical cubic spline random effects model. In various embodiments, the model is a cubic spline model. In various embodiments, the model is a generalized linear effects model. In various embodiments, the model is a linear effects model. In various embodiments, the model is a Cox proportional hazard model. In various embodiments, the method of analysis comprises assembly of models together. In various embodiments assembly includes generation of association parameters. In one or more embodiments, the method of analysis includes patient survival information and patient genetic information. As an example, assembly of models together could include different models for the different types of cancers, including subtypes, represented in the patient survival information. Each of different models can be configured to determine correlations between genetic factors and the survival times of patients diagnosed with the respective types of cancers they are configured to evaluate. For example, genetic factors determined to have strong correlations to cancer survival times (e.g., relatively short survival times and / or relatively long survival times) can be recommended as potential therapeutic targets.

[0098] In various embodiments, analysis can include one or more of survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. For example modeling can facilitate applying the aforementioned, such as patient survival information and the patient genetic information. In various embodiments, a sub-modeling component can determines subsets of the patient survival information and the patient genetic information for generations different patient cohorts associated with different types of cancer and cancer subtypes. In various embodiments, a sub-model includes hierarchal models (e.g., nested models, multi-level models), mixed models (e.g., regression such as logistic regression and Poisson regression, pooled, random effect, fixed effect, mixed effect, linear mixed effect, generalized linear mixed effect), hazard model, odds ratio models and / or repeated sample (e.g., repeated measures such as ANOVA). In various embodiments, the sub-model is a hierarchical random effects model. In various embodiments, the sub-model is a hierarchical cubic spline random effects model. In various embodiments, the sub-model is a cubic spline model.In various embodiments, the sub-model is a generalized linear effects model. In various embodiments, the sub-model is a linear effects model. In various embodiments, the submodel is a Cox proportional hazard model. Each subset of the patient survival information and the patient genetic information can comprise information for patients diagnosed with a different type of cancer and cancer subtypes. For example, the submodeling component can further apply the subsets of the patient survival information and the patient genetic information to corresponding individual survival models developed for the different cancer types, including subtypes. In various embodiments, information generated for the method of analysis can be stored in memory (e.g., as model data). In various embodiments, and information generated for the method of analysis generates one or more survival models for individual subjects.

[0099] In various embodiments, analysis of the patient survival information and the patient genetic information using the survival models, include the disease node determination and identification component can identify, for each type of cancer, disease nodes included in the patient genetic information that are involved in the genetic mechanisms employed by the respective cancer types to proliferate. In various embodiments,, the disease node component identifies disease node based on observed correlations between genetic factors and the cancer survival times provided in the patient survival information. For example, a genetic factor that is frequently observed in association with short survival times of a specific type of cancer and less frequently observed in association with long survival times of the specific type of cancer can be identified as an active genetic factor having an active role in the genetic mechanism of the specific type of cancer, including subtypes.

[0100] In various embodiments, disease node determination and identification includes disease association parameters regarding associations between different cancer types to facilitate identifying the active genetic factors associated with the different cancer types. For example, cancer types which are highly associated can share one or more common critical underlying genetic factors. As readily appreciated by one of ordinary skill, models (e.g., survival model) of associated cancer types dialectically exchange information to determine and / or identify active genetic factors across types of cancer, including subtypes. In various embodiments, the disease association parameters applied by the disease node determination and identification is facilicated by modeling. In various embodiments, generation of individual survival models can employ one or moremachine learning algorithms to facilitate the determination and / or identification of the survival, modeling, disease node associated with a particular type of cancer, including subtypes, based on the, and the patient genetic information and the disease association parameters.

[0101] In some embodiments, in association with node determination and identification for cancer type, including subtypes, includes determination of a score system for the disease node(s). For example, a score for a disease node with respect to a specific type of cancer, including subtypes, reflects association of the disease node to the survival time of the specific type of cancer, including subtypes. In various embodiments, scores can be based on a frequency with which a particular genetic factor is directly or indirectly identified for patients diagnosed with a specific cancer type. In various embodiments, analysis includes the aforementioned survival, sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. can be related to, less than a defined threshold, greater than a defined threshold. For example, greater scores associated with a disease node and cancer type, including, the greater the contribution of the disease node to survival time. In various embodiments, formation regarding disease nodes for respective types of cancer, including subtypes and scores determined for the active genetic factors can be collated in a data structure, such as a database.

[0102] Described herein is a method of analysis that includes effects modeling. In various embodiments, the effects modeling includes random effect, fixed effect, mixed effect, linear mixed effect, and generalized linear mixed effect. In various embodiments, the effects includes cubic spline. In various embodiments, the effects modeling includes regression. In various embodiments, the effects modeling includes logistic and Poisson regression. In various embodiments, the model does not include covariates. In In various embodiments, the model includes covariates. In various embodiments, the covariates are information from medical records (including laboratory testing records such as genomic, epigenomic, nucleic acid and other analyte results), insurance records or the like. Examples include age, line of therapies, smoking status (yes / no), gender, and various scoring and / or staging systems that have been utilized for specific cancer disease patients, with an illustrative example including age (in years), line of anti-EGFR therapy, smoking status (yes / no), gender (female / male), and the Van Walraven Elixhauser Comorbidity (ELIX) score specific to lung cancer patients (expressed as a weightedmeasure across multiple common comorbidities. One of skill readily appreciates covariates can include any number of data elements for individuals and individuals in a population, such as that from medical records (including laboratory testing records such as genomic, epigenomic, nucleic acid and other analyte results), insurance records or the like.

[0103] In various embodiments, the method of analysis includes generation of a hierarchy including at least one first level equation. In various embodiments, a first level equation includes a truncated cubic spline. In various embodiments, the truncated cubic spline includes longitudinal data. This includes, for example, direct or indirect measurements of ctDNA levels, allele fractions, tumor fractions. In various embodiments, an additional level equation includes covariate. In various embodiments, the covariates are information for an individual, or individuals in a population, drawn and / or stored from medical records (including laboratory testing records such as genomic, epigenomic, nucleic acid and other analyte results), insurance records, or the like. Examples include age, line of therapies, smoking status (yes / no), gender, and various scoring and / or staging systems that have been utilized for specific cancer disease patients. In various embodiments, a velocity plot is generated. In various embodiments, the velocity plot is a derivative or one or equations, such as at least one first level equation. In various embodiments, the method of analysis includes one or more of Equations (1), (2) and (3) described in the Examples.

[0104] Described herein is a method of analysis that includes jointly solving different analysis components, including one or more of survival, modeling and sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. In various embodiments, the method of analysis includes jointly solving one or more different models for the different cancer types under a joint model framework. For example, the method of analysis could include jointly solving one or more different survival models for the different cancer types under a joint model framework. In various embodiments, the method includes determination of association parameters. In various embodiments, association parameters include, for example, the relationship between patient survival and the patient’s estimated current value of the biomarker, the relationship between patient survival and patient’s estimated current change over time with respect to the biomarker. In various embodiments, this includesthe slope, and the relationship between overall survival and current estimated area under a subject’s longitudinal trajectory as a surrogate for a biomarker’s cumulative effect. It is readily appreciated by one of ordinary skill that association parameters can undertake multiple forms, and can also be combined. For instance, one could examine the relationship between overall survival and estimated current value plus the estimated current slope of the patient’s longitudinal trajectory.

[0105] In one or more examples, the data analysis system may implement at least one of one or more machine learning techniques or one or more statistical techniques to analyze data retrieved in response to one or more integrated data repository requests. In one or more examples, the data analysis system may implement one or more artificial neural networks to analyze data retrieved in response to one or more integrated data repository requests. To illustrate, the data analysis system may implement at least one of one or more convolutional neural networks or one or more residual neural networks to analyze data retrieved from the integrated data repository in response to one or more integrated data repository requests. In at least some examples, the data analysis system may implement one or more random forests techniques, one or more support vector machines, or one or more Hidden Markov models to analyze data retrieved in response to one or more integrated data repository requests. One or more statistical models may also be implemented to analyze data retrieved in response to one or more integrated data repository requests to identify at least one of correlations or measures of significance between characteristics of individuals. For example, log rank tests may be applied to data retrieved in response to one or more integrated data repository requests. In addition, Cox proportional hazards models may be implemented with respect to date retrieved in response to one or more integrated data repository requests. Further, Wilcoxon signed rank tests may be applied to data retrieved in response to one or more integrated data repository requests. In still other examples, a z-score analysis may be performed with respect to data retrieved in response to one or more integrated data repository requests. In still additional examples, a Kaplan Meier analysis may be performed with respect to data retrieved in response to one or more integrated data repository requests. In at least some examples, one or more machine learning techniques may be implemented in combination with one or more statistical techniques to analyze data retrieved in response to one or more integrated data repository requests.

[0106] In one or more illustrative examples, the data analysis system may determine a rate of survival of individuals in which lung cancer is present in response to one or more treatments. In one or more additional illustrative examples, the data analysis system may determine a rate of survival of individuals having one or more genomic and / or epigenomic region mutations in which lung cancer is present in response to one or more treatments. In various examples, the data analysis system may generate the data analysis results in situations where the data retrieved from at least one of the integrated data repository or the one or more datasets generated by the data pipeline system satisfies one or more criteria. For example, the data analysis system may determine whether at least a portion of the data retrieved in response to one or more integrated data repository requests satisfies a threshold confidence level. In situations where the confidence level for at least a portion of the date retrieved in response to one or more integrated data repository requests is less than a threshold confidence level, the data analysis system may refrain from generating at least a portion of data analysis results. In scenarios where the confidence level for at least a portion of the data retrieved in response to one or more integrated data repository requests is at least a threshold confidence level, the data analysis system may generate at least a portion of the data analysis results. In various examples, the threshold confidence level may be related to the type of data analysis results being generated by the data analysis system.

[0107] In one or more illustrative examples, the data analysis system may receive an integrated data repository request to generate data analysis results that indicate a rate of survival of one or more individuals. In these instances, the data analysis system may determine whether the data stored by the integrated data repository and / or by one or more datasets generated by the data pipeline system satisfies a threshold confidence level, such as a Gold standard confidence level. In one or more additional examples, the data analysis system may receive an integrated data repository request to generate data analysis results that indicate a treatment received by one or more individuals. In these implementations, the data analysis system may determine whether the data stored by the integrated data repository and / or by one or more datasets generated by the data pipeline system satisfies a lower threshold confidence level, such as a Bronze standard confidence level.

[0108] In one or more additional illustrative examples, the data analysis system may receive an integrated data repository request to determine individuals having one or moregenomic and / or epigenomic mutations and that have received one or more treatments for a biological condition. Continuing with this example, the data analysis system can determine a survival rate of individuals with the one or more genomic and / or epigenomic mutations in relation to the one or more treatments received by the individuals. The data analysis system can then identify based on the survival rate of individuals and effectiveness of treatments for the individuals in relation to genomic and / or epigenomic mutations that may be present in the individuals. In this way, health outcomes of individuals may be improved by identifying prospective treatments that may be more effective for populations of individuals having one or more genomic and / or epigenomic mutations than current treatments being provided to the individuals.

[0109] The data pipeline system may include first data processing instructions , second data processing instructions , up to Nth data processing instructions . The data processing instructions, , may be executable by one or more processing units to perform a number of operations to generate respective datasets using information obtained from the integrated data repository . In one or more illustrative examples, the data processing instructions, , may include at least one of software code, scripts, API calls, macros, and so forth. The first data processing instructions may be executable to generate a first dataset . In addition, the second data processing instructions may be executable to generate a second dataset . Further, the Nth data processing instructions may be executable to generate an Nth dataset . In various examples, after the data integration and analysis system generates the integrated data repository , the data pipeline system may cause the data processing instructions , , to be executed to generate the datasets , , . In one or more examples, the datasets, , may be stored by the integrated data repository or by an additional data repository that is accessible to the data integration and analysis system . At least a portion of the data processing instructions may analyze health insurance codes to generate at least a portion of the datasets. Additionally, at least a portion of the data processing instructions may analyze genomics data to generate at least a portion of the datasets .

[0110] In one or more examples, the first data processing instructions may be executable to retrieve data from one or more first data tables stored by the integrated data repository . The first data processing instructions may also be executable to retrieve data from one or more specified columns of the one or more first data tables. In various examples, the first data processing instructions may be executable to identify individualsthat have a health insurance code stored in one or more column and row combinations that correspond to one or more diagnosis codes. The first data processing instructions may then be executable to analyze the one or more diagnosis codes to determine a biological condition for which the individuals have been diagnosed. In one or more illustrative examples, the first data processing instructions may be executable to analyze the one or more diagnosis codes with respect to a library of diagnosis codes that indicates one or more biological conditions that correspond to respective diagnosis codes. The library of diagnosis codes may include hundreds up to thousands of diagnosis codes. The first data processing instructions may also be executable to determine individuals diagnosed with a biological condition by analyzing timing information of the individuals, such as dates of treatment, dates of diagnosis, dates of death, one or more combinations thereof, and the like.[OHl] The second data processing instructions may be executable to retrieve data from one or more second data tables stored by the integrated data repository . The second data processing instructions may also be executable to retrieve data from one or more specified columns of the one or more second data tables. In various examples, the second data processing instructions may be executable to identify individuals that have a health insurance code stored in one or more columns and row combinations that correspond to one or more treatment codes. The one or more treatment codes may correspond to treatments obtained from a pharmacy. In one or more additional examples, the one or more treatment codes may correspond to treatments received by a medical procedure, such as an injection or intravenously. The second data processing instructions may be executable to determine one or more treatments that correspond to the respective health insurance codes included in the one or more second data tables by analyzing the health insurance code in relation to a predetermined set of information. The predetermined set of information may include a data library that indicates one or more treatments that correspond to one out of hundreds up to thousands of health insurance codes. The second data processing instructions may generate the second dataset to indicate respective treatments received by a group of individuals. In one or more illustrative examples, the group of individuals may correspond to the individuals included in the first dataset. The second dataset may be arranged in rows and columns with one or more rows corresponding to a single individual and one or more columns indicating the treatments received by the respective individual.

[0112] The Nth processing instructions (where N may be any positive integer) may be executable to generate the Nth dataset by combining information from a number of previously generated datasets, such as the first dataset and the second dataset . In addition, the Nth processing instructions may be executable to generate the Nth dataset to retrieve additional information from one or more additional columns of the integrated data repository and incorporate the additional information from the integrated data repository with information obtained from the first dataset and the second dataset . For example, the Nth processing instructions may be executable to identify individuals included in the first dataset that are diagnosed with a biological condition and analyze specified columns of one or more additional data tables of the integrated data repository to determine dates of the treatments indicated in the second dataset that correspond to the individuals included in the first dataset . In one or more further examples, the Nth processing instructions may be executable to analyze columns of one or more additional data tables of the integrated data repository to determine dosages of treatments indicated in the second dataset received by the individuals included in the first dataset . In this way, the Nth processing instructions may be executable to generate an episodes of care dataset based on information included in a cohort dataset and a treatments dataset.

[0113] In one or more illustrative examples, in response to receiving an integrated data repository request, the data analysis system may determine one or more datasets that correspond to the features of the query related to the integrated data repository request . For example, the data analysis system may determine that information included in the first dataset and the second dataset is applicable to responding to the integrated data repository request . In these scenarios, the data analysis system may analyze at least a portion of the data included in the first dataset and the second dataset to generate the data analysis results . In one or more additional examples, the data analysis system may determine different datasets to respond to different queries included in the integrated data repository request in order to generate the data analysis results .

[0114] The use of specific sets of data processing instructions to generate respective data sets may reduce the number of inputs from users of the data integration and analysis system as well as reduce the computational load, such as the amount of processing resources and memory, utilized to process integrated data repository requests . For example, without the specific architecture of the data pipeline system, each time an integrated data repository request is received, the data utilized to respond to theintegrated data repository request is assembled from the data repository . In contrast, by implementing the data pipeline system to execute the data processing instruction to generate the datasets the data needed to respond to various integrated data repository requests has already been assembled and may be accessed by the data analysis system to respond to the integrated data repository request . Thus, the computing resources used to respond to the integrated data repository request by implementing the data pipeline system to generate the datasets are less than typical systems that perform an information parsing and collecting process for each integrated data repository request . Further, in situations where the data pipeline system has not been implemented, users of the data integration and analysis system may need to submit multiple integrated data repository request in order to analyze the information that the users are intending to have analyzed either because the ad hoc collection of data to respond to an integrated data repository request in typical systems is inaccurate or because the data analysis system is called upon multiple times to perform an analysis of information in typical systems that may be performed using a single integrated data repository request when the data pipeline system is implemented.

[0115] At operation, the data integration and analysis system may integrate genomics data and health insurance claims data of individuals that are common to both the molecular data repository and the health insurance claims data repository . The data integration and analysis system may determine individuals that are common to both the molecular data repository and the health insurance claims data repository by determining genomics data and health insurance claims data corresponding to common tokens. The data integration and analysis system may determine that a first token related to a portion of the genomics data corresponds to a second token related to a portion of the health insurance claims data by determining a measure of similarity between the first token and the second token . In scenarios where the first token has at least a threshold amount of similarity with respect to the second token , the data integration and analysis system may store the corresponding portion of the genomics data and the corresponding portion of the health insurance claims data in relation to the identifier of the individual in an integrated data repository, such as an integrated data repository .

[0116] The implementation of the architecture may implement a cryptographic protocol that enables de-identified information from disparate data repositories to be integrated into a single data repository. In this way, the security of the data stored by the integrateddata repository is increased. Additionally, the cryptographic protocol implemented by the architecture may enable more efficient retrieval and accurate analysis of information stored by the integrated data repository than in situations where the cryptographic protocol of the architecture is not utilized. For example, by generating a token file that includes first tokens using a cryptographic technique based on a specified set of information stored by the molecular data repository and utilizing second tokens generated using a same or similar cryptographic technique with respect to the similar or same set of information stored by the health insurance claims data repository , the data integration and analysis system may match information stored by disparate data repositories that correspond to a same individual. Without implementing the cryptographic protocol of the architecture, the probability of incorrectly attributing information from one data repository to one or more individuals increases, which decreases the accuracy of results provided by the data integration and analysis system in response to integrated data repository requests sent to the data integration and analysis system .

[0117] Described herein is a framework to generate a dataset, by a data pipeline system , based on data stored by an integrated data repository , according to one or more implementations. The integrated data repository may store health insurance claims data and genomics data for a group of individuals . For example, the integrated data repository may store information obtained from health insurance claims records of the group of individuals . For each individual included in the group of individuals, the integrated data repository may store information obtained from multiple health insurance claim records . In various examples, the information stored by the integrated data repository may include and / or be derived from thousands, tens of thousands, hundreds of thousands, up to millions of health insurance claims records for a number of individuals. Additionally, each health insurance claim record may include multiple columns. As a result, the integrated data repository may be generated through the analysis of millions of columns of health insurance claims data.

[0118] Further, although the health insurance claims data may be organized according to a structured data format, health insurance claims data is typically arranged to be viewed by health insurance providers, patients, and healthcare providers in order to show financial information and insurance code information related to services provided to individuals by healthcare providers. Thus, health insurance claims data is not easilyanalyzed to gain insights that may be available in relation to characteristics of individuals in which a biological condition is present and that may aid in the treatment of the individuals with respect to the biological condition. The integrated data repository may be generated and organized by analyzing and modifying raw health insurance claims data in a manner that enables the data stored by the integrated data repository to be further analyzed to determine trends, characteristics, features, and / or insights with respect to individuals in which one or more biological conditions may be present. For example, health insurance codes may be stored in the integrated data repository in such a way that at least one of medical procedures, biological conditions, treatments, dosages, manufacturers of medications, distributors of medications, or diagnoses may be determined for a given individual based on health insurance claims data for the individual. In various examples, the data integration and analysis system may generate and implement one or more tables that indicate correlations between health insurance claims data and various treatments, symptoms, or biological conditions that correspond to the health insurance claims data. Further, the integrated data repository may be generated using genomics data records of the group of individuals . In various examples, the large amounts of health insurance claims data may be matched with genomics data for the group of individuals to generate the integrated data repository.

[0119] By integrating the genomics data records for the group of individuals with the health insurance claims records , the data integration and analysis system may determine correlations between the presence of one or more biomarkers that are present in the genomics data records with other characteristics of individuals that are indicated by the health insurance claims data records that existing systems are typically unable to determine. For example, the data integration and analysis system may determine one or more genomic and / or epigenomic characteristics of individuals that correspond to treatments received by individuals, timing of treatments, dosages of treatments, diagnoses of individuals, smoking status, presence of one or more biological conditions, presence of one or more symptoms of a biological condition, one or more combinations thereof, and the like. Based on the correlations determined by the data integration and analysis system using the integrated data repository, cohorts of individuals that may benefit from one or more treatments may be identified that would not have been identified in existing systems. In one or more examples, the processes and techniques implemented to integrate the health insurance claims records and the genomics claimsrecords in order to generate the integrated data repository may be complex and implement efficiency-enhancing techniques, systems, and processes in order to minimize the amount of computing resources used to generate the integrated data repository .

[0120] In one or more illustrative examples, the data pipeline system may access information stored by the integrated data repository to generate datasets that include a number of additional data records that include information related to at least a portion of the group of individuals . In an illustrative example, the additional data record includes information indicating whether individuals are included in a cohort of individuals in which lung cancer is present. The data pipeline system may execute a plurality of different sets of data processing instructions to determine a cohort of the group of individuals in which lung cancer is present. In various examples, the additional data record may indicate information used to determine a status of an individual with respect to lung cancer, such as one or more transaction insurance identifier, one or more international classification of diseases (ICD) codes, and one or more health insurance transaction dates. In addition to including a column that indicates whether an individual is included in the lung cancer cohort, the additional data record may include a column indicating a confidence level of the status of the individual with respect to the presence of lung cancer.

[0121] Described herein is a computing architecture to incorporate medical records data into an integrated data repository. In various examples, at least a portion of the operations of the computing architecture may be performed by the data integration and analysis system. In one or more examples, at least a portion of the operations of the computing architecture may be performed by one or more additional computing systems that are at least one of controlled, maintained, or implemented by a service provider that also at least one of controls, maintains, or implements the data integration and analysis system . In one or more additional examples, at least a portion of the operations of the computing architecture may be performed by a number of servers in a distributed computing environment.

[0122] The computing architecture may include a medical records data repository. The medical records data repository may store medical records data from a number of individuals. The medical records data may include imaging information, laboratory test results, diagnostic test information, clinical observations, dental health information, notes of healthcare practitioners, medical history forms, diagnostic request forms,medical procedure order forms, medical information charts, one or more combinations thereof, and so forth. In various examples, for a given individual, the medical records data repository may store information obtained from one or more healthcare practitioners that is related to the individual.

[0123] The computing architecture may perform an operation that includes obtaining data packages from the medical records data repository. In one or more examples, the data packages may be obtained in response to one or more requests sent to the medical records data repository for medical records that correspond to one or more individuals. In one or more additional examples, the data packages may be obtained by the computing architecture using one or more application programming interface (API) calls. In one or more illustrative examples, a first data package, a second data package, up to an Nth data package may be obtained using the computing architecture. The individual data packages, may correspond medical records of a respective individual. For example, the first data package may include medical records of a first individual, the second data package may include medical records of a second individual, and the Nth data package may include medical records of a third individual.

[0124] Individual data packages, may include a number of components. In one or more examples, individual data packages, may include individual components that correspond to medical records from different healthcare providers. In one or more additional examples, the individual data packages, may include individual components that correspond to different parts of medical records that correspond to one or more healthcare providers. In an illustrative example the second data package may include a first component, a second component , up to an Nth component . In one or more illustrative examples, the first component may include a first portion of medical records of an individual, the second component may include a second portion of medical records of an individual, and the Nth component may include a third portion of medical records of an individual. In various examples, the first component may correspond to medical records of a first healthcare provider for the individual, the second component may correspond to medical records of a second healthcare provider for the individual, and the third component may correspond to medical records of a third healthcare provider for the individual. In one or more additional illustrative examples, the first component may include a first section of medical records of the individual, such as one or more forms related to a diagnostic test or procedure, and the second component may include asecond section of medical records of the individual, such as a pathology report of the individual.

[0125] At operation, the computing architecture may preprocess individual data packages to identify a corpus of information to be analyzed. In one or more examples, the preprocessing of data packages obtained from the medical records data repository, may include transforming the data included in the data packages. For example, preprocessing the data packages may include transforming at least a portion of the data obtained from the medical records data repository to machine encoded information. To illustrate, preprocessing the data packages may include performing one or more optical character recognition (OCR) operations with respect to at least a portion of the data packages obtained from the medical records data repository. By converting at least a portion of the data packages obtained from the medical records data repository to machine encoded information, the data packages may be subjected to a number of operations, such as one or more parsing operations to identify one or more characters or strings of characters or one or more editing operations that are unable to be performed with respect to at least a portion of the data packages obtained from the medical records data repository .

[0126] In one or more examples, the preprocessing of individual data packages may include determining information included in individual data packages that is to be excluded from further analysis by the computing architecture. In various examples, one or more components of individual data packages may be excluded from a corpus of information to be analyzed. For example, with respect to the second data package, the computing architecture may determine that the first component is to be excluded from further analysis by the computing architecture . In one or more examples, the computing architecture may analyze the components , , and / or with respect to one or more keywords to identify at least one of the components , , and / or to exclude from further analysis by the computing architecture . In one or more illustrative examples, the computing architecture may parse the components , , and / or to identify one or more keywords and in response to identifying the one or more keywords in a component , , and / or , the computing architecture may determine to exclude the respective component , , and / or from further analysis by the computing architecture . For example, the computing architecture may determine that the first component of the second data package is a test requisition form for one or more diagnostic procedures or tests. Inthese scenarios, the computing architecture may determine that the first component is to be excluded from further analysis by the computing architecture . Additionally, the computing architecture may determine that at least one of the second component and / or correspond to one or more pathology reports for an individual based on one or more keywords included in at least one of the second component or the Nth component . In these instances, the computing architecture may determine that at least a portion of the second component and / or at least a portion of the Nth component is to be included in the corpus of information to be further analyzed by the computing architecture .

[0127] In addition, a subset of the components of individual data packages obtained from the medical records data repository may be included in the corpus of information . In various examples, one or more additional operations may be performed to narrow the corpus of information. For example, one or more queries may be applied to a subset of information obtained from the medical records data repository. The one or more queries may extract information from the one or more data packages that satisfy the one or more queries. In at least some examples, the one or more queries may be a group of queries that are applied to individual components of a data package. In one or more illustrative examples, the group of queries may determine information to be included in the corpus of information and additional information that is to be excluded from the corpus of information . In one or more additional examples, one or more sections of at least one component of a data package may be excluded from the corpus of information.

[0128] In one or more additional illustrative examples, after determining that the first component is to be excluded from further analysis by the computing architecture , the computing architecture may then cause one or more queries to be implemented with respect to at least one the second component or the Nth component . In these scenarios, the one or more queries may determine that a section of the second component , such as a section that indicates family history for one or more biological conditions, is to be excluded from the corpus of information . In various examples, the one or more queries may be directed to identifying a number of keywords and / or combinations of keywords included in at least one of the second component or the Nth component . In these instances, the computing architecture may exclude from the corpus of information one or more portions of the individual components of the data packages that include one or more keywords or combinations of keywords. In one or more additional examples, the computing architecture may exclude from the corpus of information a number of words,a number of characters, and / or a number of symbols following one or more keywords that are included in one or more portions of the individual components of the data packages.

[0129] Further, at operation, the computing architecture may analyze the corpus of information to determine characteristics of individuals. In one or more examples, the computing architecture may analyze the corpus of information to determine individuals that have one or more phenotypes. In various examples, the computing architecture may analyze the corpus of information to determine one or more biomarkers that are indicative of a biological condition. For example, the computing architecture may analyze the corpus of information to determine individuals having one or more genetic characteristics. The one or more genetic characteristics may include at least one of one or more variants of a genomic and / or epigenomic region that correspond to a biological condition. In one or more illustrative examples, the one or more genetic characteristics may correspond to one or more variants of a genomic and / or epigenomic region that correspond to a type of cancer. In one or more additional illustrative examples, the one or more biomarkers may correspond to levels of an analyte being outside of a specified range. To illustrate, the computing architecture may analyze the corpus of information to determine individuals having levels of one or more proteins and / or levels of one or more small molecules present that are indicative of a biological condition. In these scenarios, the computing architecture may analyze results of laboratory tests to determine levels of analytes of individuals. In one or more additional examples, the computing architecture may analyze the corpus of information to determine individuals in which one or more symptoms are present that are indicative of a biological condition. In one or more further examples, the computing architecture may analyze imaging information included in the corpus of information to determine individuals in which one or more biomarkers are present.

[0130] In one or more examples, the computing architecture may implement one or more machine learning techniques to analyze the corpus of information . For example, the computing architecture may implement one or more artificial neural networks, such as at least one of one or more convolutional neural networks or one or more residual neural networks to analyze the corpus of information . The computing architecture may also implement at least one of one or more random forests techniques, one or morehidden Markov models, or one or more support vector machines to analyze the corpus of information .

[0131] In at least some implementations, the computing architecture may analyze the corpus of information by performing one or more queries with respect to the corpus of information . The one or more queries may correspond to one or more keywords and / or combinations of keywords. The one or more keywords and / or combinations of keywords may correspond to at least one of characters or symbols that correspond to one or more biological conditions. To illustrate, a keyword may correspond to characters related to a mutation of a genomic and / or epigenomic region, such as HER2. In one or more additional illustrative examples, one or more criteria may be associated with combinations of keyworks. To illustrate, a criterion that corresponds to a combination of keywords may include a number of words being present within a specified distance of one another in a portion of the corpus of information for an individual, such as the words fatigue, blood pressure, and swelling occurring within characters of one another. In these instances, the computing architecture may parse the corpus of information for the one or more keywords and / or combinations of keywords. In various examples, in response to determining that the one or more keywords and / or combinations of keywords are present in accordance with one or more criteria, the computing architecture may determine that a biological condition is present with respect to a given individual.

[0132] In one or more additional examples, the one or more queries may be imagebased and the computing architecture may analyze images included in the corpus of information with respect to template images. The template images may be generated based on analyzing a number of images in which a biological condition is present and aggregating the number of images into a template image. In these scenarios, the computing architecture may analyze images included in the corpus of information with respect to one or more template images to determine a measure of similarity between the images included in the corpus of information and the template images. In situations where the measure of similarity for an individual is at least a threshold value, the computing architecture may determine that a characteristic of a biological condition is present in the individual.

[0133] After determining individuals having one or more characteristics, the computing architecture may, at operation , generate data structures that store data for individuals having the one or more characteristics. In one or more examples, the computingarchitecture may generate data tables that indicate individuals having an individual characteristics and / or individuals having a group of characteristics. For example, the computing architecture may generate a first data table and a second data table . The first data table may indicate individuals having one or more first characteristics and the second data table may indicate individuals having one or more second characteristics. In one or more illustrative examples, the first data table may indicate individuals having one or more first biomarkers for a biological condition and the second data table may indicate individual having one or more second biomarkers for the biological condition. The one or more first biomarkers may correspond to one or more first genomic and / or epigenomic variants that are associated with the biological condition and the one or more second biomarkers may correspond to one or more second genomic and / or epigenomic variants that are associated with the biological condition.

[0134] One or more data structures may be generated from the corpus of information that store identifiers of the portion of the subset of the additional group of individuals and that store an indication that the portion of the subset of the additional group of individuals corresponds to the one or more biomarkers. The one or more data structures may be stored by an intermediate data repository. One or more de-identification operations may be performed with respect to the identifiers of the portion of the subset of the additional group of individuals before modifying the integrated data repository to store at least a portion of the additional information of the medical records of the portion of the subset of the additional group of individuals in relation to the number of identifiers. After de-identification of the information stored by the one or more data structures, the information stored by the integrated data repository may be added to the integrated data repository. In at least some examples, the de-identified medical records information may be added to the integrated data repository in addition to or in lieu of the health insurance claims data. In various examples, the one or more data structures storing the de-identified medical records information with respect to the biomarker data may have one or more logical connections with other data structures stored in the integrated data repository. To illustrate, the one or more data structures storing the de- identified medical records information with respect to the biomarker data may have one or more logical connections with at least one of the first data table may store information corresponding to a panel used to generate genomics data, mutations of genomic and / or epigenomic regions, types of mutations, copy numbers of genomic and / or epigenomicregions, coverage data indicating numbers of nucleic acid molecules identified in a sample having one or more mutations, testing dates, and patient information, the second data that stores data related to one or more patient visits by individuals to one or more healthcare providers, the a third data table that stores information corresponding to respective services provided to individuals with respect to one or more patient visits to one or more healthcare providers indicated by the second data table, the fourth data table that stores personal information of the group of individuals, the fifth data table that stores information related to a health insurance company or governmental entity that made payment for services provided to the group of individuals, the sixth data table storing information corresponding to health insurance coverage information for the group of individuals, such as a type of health insurance plan related to the group of individuals, or the seventh data table that stores information related to pharmaceutical treatments obtained by the group of individuals.

[0135] Described herein is a machine in the form of a computer system within which a set of instructions may be executed for causing the machine to perform any one or more of the methodologies discussed herein, according to an example, according to an example implementation. For example, a machine in the example form of a computer system, within which instructions (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine to perform any one or more of the methodologies discussed herein may be executed. For example, the instructions may cause the machine to implement the architectures and frameworks described previously, and to execute the methods described with respect to previously. For example, one or more machine-executable components embodied within one or more machines (e.g., embodied in one or more computer-readable storage media associated with one or more machines). Such components, when executed by the one or more machines (e.g., processors, computers, computing devices, virtual machines, etc.) can cause the one or more machines to perform the operations described through instructions. For example, a machine can include a computing device with an analysis component. Analysis can include survival, modeling, sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. In various embodiments analysis components are embodied in machine-executable components within a system including various electronic data sources and data structures comprising informationcapable of use with the analysis component. Non-limiting examples include data sources and structures such survival information, genetic information, model data, sub-model, disease node determination and identification, disease association information, disease subtyping, recurrence, metastasis, time to next treatment, etc.

[0136] A computing device can include or be operatively coupled to at least one memory and at least one processor. The at least one memory stores executable instructions for performance of analysis when executed by the at least one processor. In some embodiments, the memory can also store the various data sources and / or structures of system. In other embodiments, the various data sources and structures of system can be stored in other memory (e.g., at a remote device or system), that is accessible to the computing device.

[0137] The instructions transform the general, non-programmed machine, such as a computing device, into a particular machine programmed to carry out the described and illustrated functions in the manner described. In alternative implementations, the machine operates as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may comprise, but not be limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set -top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions , sequentially or otherwise, that specify actions to be taken by the machine . One of skill appreciate that a machine iinclude a collection of machines that individually or jointly execute the instructions to perform any one or more of the methodologies discussed herein.

[0138] Examples of computing devices may include logic, one or more components, circuits (e.g., modules), or mechanisms. Circuits are tangible entities configured to perform certain operations. In an example, circuits may be arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner. In an example, one or more computer systems (e.g., a standalone, client or server computersystem) or one or more hardware processors (processors) may be configured by software (e.g., instructions, an application portion, or an application) as a circuit that operates to perform certain operations as described herein. In an example, the software may reside (1) on a non-transitory machine readable medium or (2) in a transmission signal. In an example, the software, when executed by the underlying hardware of the circuit, causes the circuit to perform the certain operations.

[0139] The various operations of method examples described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor- implemented circuits that operate to perform one or more operations or functions. In an example, the circuits referred to herein may comprise processor-implemented circuits.

[0140] Similarly, the methods described herein may be at least partially processor implemented. For example, at least some or all of the operations of a method may be performed by one or processors or processor-implemented circuits. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In an example, the processor or processors may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other examples the processors may be distributed across a number of locations.

[0141] The one or more processors may also operate to support performance of the relevant operations in a "cloud computing" environment or as a "software as a service.”

[0142] (SaaS). For example, at least some of the operations may be performed by a group of computers (as examples of machines including processors), with these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., Application Program Interfaces (APIs).)

[0143] Example implementations (e.g., apparatus, systems, or methods) may be implemented in digital electronic circuitry, in computer hardware, in firmware, in software, or in any combination thereof. Example implementations may be implemented using a computer program product (e.g., a computer program, tangibly embodied in an information carrier or in a machine readable medium, for execution by, or to control the operation of, data processing apparatus such as a programmable processor, a computer, or multiple computers).

[0144] A computer program may be written in any form of programming language, including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a software module, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.

[0145] In an example, operations may be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Examples of method operations may also be performed by, and example apparatus may be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)).

[0146] The computing system may include clients and servers. A client and server are generally remote from each other and generally interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship with each other. In implementations deploying a programmable computing system, it will be appreciated that both hardware and software architectures require consideration. Specifically, it will be appreciated that the choice of whether to implement certain functionality in permanently configured hardware (e.g., an ASIC), in temporarily configured hardware (e.g., a combination of software and a programmable processor), or a combination of permanently and temporarily configured hardware may be a design choice. Below are set out hardware (e.g., computing device) and software architectures that may be deployed in example implementations.

[0147] In an example, the computing device may operate as a standalone device or the computing device may be connected (e.g., networked) to other machines.

[0148] In a networked deployment, the computing device may operate in the capacity of either a server or a client machine in server-client network environments. In an example, computing device may act as a peer machine in peer-to-peer (or other distributed) network environments. The computing device may be a personal computer (PC), a tablet PC, a set-top box (STB), a mobile telephone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) specifying actions to be taken (e.g., performed) by the computing device .Further, while only a single computing device is illustrated, the term “computing device” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or

[0149] The computing device may additionally include a storage device (e.g., drive unit) , a signal generation device (e.g., a speaker), a network interface device , and one or more sensors , such as a global positioning system (GPS) sensor, compass, accelerometer, or another sensor. The storage device may include a machine readable medium on which is stored one or more sets of data structures or instructions (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. The instructions may also reside, completely or at least partially, within the main memory , within static memory , or within the processor during execution thereof by the computing device . In an example, one or any combination of the processor, the main memory , the static memory , or the storage device may constitute machine readable media.

[0150] While the machine readable medium is illustrated as a single medium, the term "machine readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that configured to store the one or more instructions . The term “machine readable medium” may also be taken to include any tangible medium that is capable of storing, encoding, or carrying instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present disclosure or that is capable of storing, encoding or carrying data structures utilized by or associated with such instructions.

[0151] As used herein, a component, may refer to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other technologies that provide for the partitioning or modularization of particular processing or control functions. Components may be combined via their interfaces with other components to carry out a machine process. A component may be a packaged functional hardware unit designed for use with other components and a part of a program that usually performs a particular function of related functions. Components may constitute either software components (e.g., code embodied on a machine-readable medium) or hardware components. A "hardware component" is a tangible unit capable of performing certain operations and may be configured or arranged in a certain physical manner. In various example implementations, one or more computer systems (e.g., a standalonecomputer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware component that operates to perform certain operations as described herein.Diseases

[0152] The present methods can be used to diagnose presence of conditions, in a subject, to characterize conditions, monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of nucleic acids, such as cell free nucleic acids, detected in subject's blood if the treatment is successful as diseased and dysfunctional die and shed DNA or otherwise exhibit chronic and acute signs of inflammation. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of disease types and sub-types over time. This correlation may be useful in selecting a therapy.

[0153] In some embodiments, the methods and systems disclosed herein may be used to identify customized or targeted therapies to treat a given disease or condition in patients based on the classification of a nucleic acid variant as being of somatic or germline origin. Typically, the disease under consideration is a type of cancer.

[0154] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genomic and epigenomic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile includes a plurality of data that can characterize malfunctions and abnormalities associated with the heart muscle and valve tissues (e.g., hypertrophy), the decreased supply of blood flow and oxygen supply to the heart are often secondary symptoms of debilitation and / or deterioration of the blood now and supply system caused by physical and biochemical stresses. Examples of cardiovascular diseases that are directly affected by these types of stresses include atherosclerosis, coronary artery disease, peripheral vascular disease and peripheral artery disease, along with various cardias and arrythmias which may represent other forms of disease and dysfunction. The present methods can be used to generate our profile, fingerprint or set of data that is a summation of genetic information derived fromdifferent cells in a heterogeneous disease. This set of data may comprise copy number variation, epigenetic variation, and mutation analyses alone or in combination.

[0155] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non- invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co-circulate with maternal molecules.

[0156] Non-limiting examples of other genetic-based diseases, disorders, or conditions that are optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cri du chat, Crohn's disease, cystic fibrosis, Dercum disease, down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson disease, or the like.Therapies and Related Administration

[0157] In certain embodiments, the methods disclosed herein relate to identifying and administering customized therapies to patients given the status of a nucleic acid variant as being of somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, customized therapies include at least one immunotherapy (or an immunotherapeutic agent). Immunotherapy refers generally to methods of enhancing an immune response against a given cancer type. In certainembodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer.

[0158] In certain embodiments, the status of a nucleic acid variant from a sample from a subject as being of somatic or germline origin may be compared with a database of comparator results from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the test subject and / or patients who are receiving, or who have received, the same therapy as the test subject. A customized or targeted therapy (or therapies) may be identified when the nucleic variant and the comparator results satisfy certain classification criteria (e.g., are a substantial or an approximate match).

[0159] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing an immunotherapeutic agent are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by methods such as, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraauricular, which administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like.

[0160] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in practicing the invention. It is therefore contemplated that the disclosure shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of theinvention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

[0161] While the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be clear to one of ordinary skill in the art from a reading of this disclosure that various changes in form and detail can be made without departing from the true scope of the disclosure and may be practiced within the scope of the appended claims. For example, all the methods, systems, computer readable media, and / or component features, steps, elements, or other aspects thereof can be used in various combinations.

[0162] In some cases, the cancer treatment includes, without limitation, imatinib, gefatinib, afatinib, dacomitinib, sunitinib, sorafenib, vandetanib, brivanib, cabozantib, neratinib, tivantinib, bevacizumab, cixutumumab, dalotuzumab, figitumumab, rilotumumab, onartuzumab, ganitumab, ramucirumab, ridaforolimus, tensirolimus, everolimus, BMS-690514, BMS-754807, EMD 525797, GDC-0973, GDC-0941, MK- 2206, AZD6244, GSK1120212, PX-866, XL821, IMC-A12, MM-121, PF-02341066, RG7160, and Sym004. Antibodies suitable for use as anti-EGFR therapy include cetuximab (Trade Name: Erbitux) and panitumumab (Trade Name: Vectibex). In some cases. In some cases, the cancer treatment includes EGFR tyrosine kinase inhibitors such as gefitinib (Trade Name: Iressa), erlotinib (Trade Name: Tarceva), lapatinib, canertinib, and cetuximab.

[0163] In some instances, therapties may be used in combination, such as an anti- EGFR therapy and an anti-EGFR therapy. Anti-EGFR therapy may be used in combination with any combination of chemotherapeutic agents or chemotherapeutic regimens, for example, FOLFOX (fluorouracil [5-FU] / leucovorin / oxaliplatin), FOLFIRI (5-FU / leucovorin / irinotecan), and the like.

[0164] In some embodiments, the therapy includes an epigenetic regulator, including HAT, HDAC inhibitors as examples. Other examples including, EP015666, LLY-283, JNJ-64619178, BRD0639, AMG193, TNG908, SCR— 6920, PRT543, PRT811, MRTX1719, cycloleucine, aminobicycle-hexane-carboxcyclic acid, FIDAS agents, PF- 9366, AGR-25696, AG-270, Compound 28, IDE397. Other examples include agents described in Bray et al., Front. Onco. 2023, which is fully incorporated by reference herein.

[0165] In some embodiments, the therapy includes paclitaxel (chemotherapeutic drug), ipatasertib (AKT inhibitor), PI3K-Beta inhibitor, AZD8186, docetaxel, tyrosine kinase inhibitor, pazopanib, mTOR inhibitor, everolimus (NCT01430572), PI3K-Beta inhibitor GSK2636771, and immunotherapy, pembrolizumab (NCT03131908), tastuzumab. Other examples include agents described in Ertay et al., Genes and Diseases. 2023, and Dillon and Miller Curr Drug Targets 2015, each of which is fully incorporated by reference herein.

[0166] In some aspects, a cancer treatment is administered to a subject. In some cases, the cancer treatment is administered in combination another therapy, such as a non-anti- EGFR therapy with anti-EGFR therapy.Biomarkers

[0167] The disclosure provides methods of using biomarkers for the diagnosis, prognosis, and therapy selection of a subject suffering from diseases, e.g., heart failure, cardiovascular disease, cancer, etc.. A biomarker may be any gene or variant of a gene whose presence, mutation, deletion, substitution, copy number, or translation (i.e., to a protein) is an indicator of a disease state. Biomarkers of the present disclosure may include the presence, mutation, deletion, substitution, copy number, or translation in any one or more of EGFR, KRAS, MET, BRAF, MYC, NRAS, ERBB2, ALK, Notch, PIK3CA, APC, and SMO.

[0168] A biomarker is a genetic variant. Biomarkers may be determined using any of several resources or methods. A biomarker may have been previously discovered or may be discovered de novo using experimental or epidemiological techniques. Detection of a biomarker may be indicative of a disease when the biomarker is highly correlated to the disease. Detection of a biomarker may be indicative of cancer when a biomarker in a region or gene occur with a frequency that is greater than a frequency for a given background population or dataset.

[0169] Publicly available resources such as scientific literature and databases may describe in detail genetic variants. Scientific literature may describe experiments or genome-wide association studies (GWAS) associating one or more genetic variants. Databases may aggregate information gleaned from sources such as scientific literature to provide a more comprehensive resource for determining one or more biomarkers. Non-limiting examples of databases include FANTOM, GT ex, GEO, Body Atlas, INSiGHT, OMIM (Online Mendelian Inheritance in Man, omim.org), cBioPortal(cbioportal.org), CIViC (Clinical Interpretations of Variants in Cancer, civic.genome.wustl.edu), DOCM (Database of Curated Mutations, docm.genome.wustl.edu), and ICGC Data Portal (dcc.icgc.org). In a further example, the COSMIC (Catalogue of Somatic Mutations in Cancer) database allows for searching of biomarkers by cancer, gene, or mutation type. Biomarkers may also be determined de novo by conducting experiments such as case control or association (e.g, genome-wide association studies) studies.

[0170] One or more biomarkers may be detected in the sequencing panel. A biomarker may be one or more genetic variants. Biomarkers can be selected from single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (e.g., indels), gene fusions and inversions. Biomarkers may affect the level of a protein. Biomarkers may be in a promoter or enhancer, and may alter the transcription of a gene. The biomarkers may affect the transcription and / or translation efficacy of a gene. The biomarkers may affect the stability of a transcribed mRNA. The biomarker may result in a change to the amino acid sequence of a translated protein. The biomarker may affect splicing, may change the amino acid coded by a particular codon, may result in a frameshift, or may result in a premature stop codon. The biomarker may result in a conservative substitution of an amino acid. One or more biomarkers may result in a conservative substitution of an amino acid. One or more biomarkers may result in a nonconservative substitution of an amino acid.

[0171] The frequency of a biomarker may be as low as 0.001%. The frequency of a biomarker may be as low as 0.005%. The frequency of a biomarker may be as low as 0.01%. The frequency of a biomarker may be as low as 0.02%. The frequency of a biomarker may be as low as 0.03%. The frequency of a biomarker may be as low as 0.05%. The frequency of a biomarker may be as low as 0.1%. The frequency of a biomarker may be as low as 1%.

[0172] No single biomarker may be present in more than 50%, of subjects having the cancer. No single biomarker may be present in more than 40%, of subjects having the cancer. No single biomarker may be present in more than 30%, of subjects having the cancer. No single biomarker may be present in more than 20%, of subjects having the cancer. No single biomarker may be present in more than 10%, of subjects having the cancer. No single biomarker may be present in more than 5%, of subjects having the cancer. A single biomarker may be present in 0.001% to 50% of subjects having cancer.A single biomarker may be present in 0.01% to 50% of subjects having cancer. A single biomarker may be present in 0.01% to 30% of subjects having cancer. A single biomarker may be present in 0.01% to 20% of subjects having cancer. A single biomarker may be present in 0.01% to 10% of subjects having cancer. A single biomarker may be present in 0.1% to 10% of subjects having cancer. A single biomarker may be present in 0.1% to 5% of subjects having cancer.Genetic Analysis

[0173] Genetic analysis includes detection of nucleotide sequence variants and copy number variations. Genetic variants can be determined by sequencing. The sequencing method can be massively parallel sequencing, that is, simultaneously (or in rapid succession) sequencing any of at least 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules. Sequencing methods may include, but are not limited to: high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, singlemolecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by- ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), Next-generation sequencing, Single Molecule Sequencing by Synthesis (SMSS)(Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms and any other sequencing methods known in the art.

[0174] Sequencing can be made more efficient by performing sequence capture, that is, the enrichment of a sample for target sequences of interest, e.g., sequences including the KRAS and / or EGFR genes or portions of them containing sequence variant biomarkers. Sequence capture can be performed using immobilized probes that hybridize to the targets of interest.

[0175] Cell free DNA can include small amounts of tumor DNA mixed with germline DNA. Sequencing methods that increase sensitivity and specificity of detecting tumor DNA, and, in particular, genetic sequence variants and copy number variation, can be useful in the methods of this invention. Such methods are described in, for example, in WO 2014 / 039556. These methods not only can detect molecules with a sensitivity of up to or greater than 0.1%, but also can distinguish these signals from noise typical in current sequencing methods. Increases in sensitivity and specificity from blood-based samples of cfDNA can be achieved using various methods. One method includes highefficiency tagging of DNA molecules in the sample, e.g., tagging at least any of 50%, 75% or 90% of the polynucleotides in a sample. This increases the likelihood that a low- abundance target molecule in a sample will be tagged and subsequently sequenced, and significantly increases sensitivity of detection of target molecules.

[0176] Another method involves molecular tracking, which identifies sequence reads that have been redundantly generated from an original parent molecule, and assigns the most likely identity of a base at each locus or position in the parent molecule. This significantly increases specificity of detection by reducing noise generated by amplification and sequencing errors, which reduces frequency of false positives.

[0177] Methods of the present disclosure can be used to detect genetic variation in nonuni quely tagged initial starting genetic material (e.g., rare DNA) at a concentration that is less than 5%, 1%, 0.5%, 0.1%, 0.05%, or 0.01%, at a specificity of at least 99%, 99.9%, 99.99%, 99.999%, 99.9999%, or 99.99999%. Sequence reads of tagged polynucleotides can be subsequently tracked to generate consensus sequences for polynucleotides with an error rate of no more than 2%, 1%, 0.1%, or 0.01%.

[0178] In other examples, a gene of interest may be amplified using primers that recognize the gene of interest. The primers may hybridize to a gene upstream and / or downstream of a particular region of interest (e.g., upstream of a mutation site). A detection probe may be hybridized to the amplification product. Detection probes may specifically hybridize to a wild-type sequence or to a mutated / variant sequence. Detection probes may be labeled with a detectable label (e.g., with a fluorophore). Detection of a wild-type or mutant sequence may be performed by detecting the detectable label (e.g., fluorescence imaging). In examples of copy number variation, a gene of interest may be compared with a reference gene. Differences in copy number between the gene of interest and the reference gene may indicate amplification or deletion / truncation of a gene. Examples of platforms suitable to perform the methods described herein include digital PCR platforms such as e.g., Fluidigm Digital Array.

[0179] Described herein is a method for analysis of nucleic acid sequence information. In various embodiments, the method of analysis comprises one or more models, each of one or more including one or more of survival, sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. In various embodiments, a model includes hierarchal models (e.g., nested models, multi-level models), mixed models (e.g., regression such as logistic regression and Poisson regression, pooled, random effect, fixed effect, mixed effect, linear mixed effect, generalized linear mixed effect), hazard model, odds ratio models and / or repeated sample (e.g., repeated measures such as ANOVA). In various embodiments, the model is a hierarchical random effects model. In various embodiments, the model is a hierarchical cubic spline random effects model. In various embodiments, the model is a cubic spline model. In various embodiments, the model is a generalized linear effects model. In various embodiments, the model is a linear effects model. In various embodiments, the model is a Cox proportional hazard model. In various embodiments, the method of analysis comprises assembly of models together. In various embodiments assembly includes generation of association parameters. In one or more embodiments, the method of analysis includes patient survival information and patient genetic information. As an example, assembly of models together could include different models for the different types of cancers, including subtypes, represented in the patient survival information. Each of different models can be configured to determine correlations between genetic factors and the survival times of patients diagnosed with the respective types of cancers they are configured to evaluate. For example, genetic factors determined to have strong correlations to cancer survival times (e.g., relatively short survival times and / or relatively long survival times) can be recommended as potential therapeutic targets.

[0180] In various embodiments, analysis can include one or more of survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. For example modeling can facilitate applying the aforementioned, such as patient survival information and the patient genetic information. In various embodiments, a sub-modeling component can determine subsets of the patient survival information and the patient genetic information for generations different patient cohorts associated with different types of cancer and cancer subtypes. In various embodiments, a sub-model includes hierarchal models (e.g., nested models, multi-level models), mixed models (e.g., regression such as logistic regression and Poisson regression, pooled, random effect, fixed effect, mixed effect, linear mixed effect, generalized linear mixed effect), hazard model, odds ratio models and / or repeated sample (e.g., repeated measures such as ANOVA). In various embodiments, the sub-model is a hierarchical randomeffects model. In various embodiments, the sub-model is a hierarchical cubic spline random effects model. In various embodiments, the sub-model is a cubic spline model. In various embodiments, the sub-model is a generalized linear effects model. In various embodiments, the sub-model is a linear effects model. In various embodiments, the submodel is a Cox proportional hazard model. Each subset of the patient survival information and the patient genetic information can comprise information for patients diagnosed with a different type of cancer and cancer subtypes. For example, the submodeling component can further apply the subsets of the patient survival information and the patient genetic information to corresponding individual survival models developed for the different cancer types, including subtypes. In various embodiments, information generated for the method of analysis can be stored in memory (e.g., as model data). In various embodiments, and information generated for the method of analysis generates one or more survival models for individual subjects.

[0181] In various embodiments, analysis of the patient survival information and the patient genetic information using the survival models, include the disease node determination and identification component can identify, for each type of cancer, disease nodes included in the patient genetic information that are involved in the genetic mechanisms employed by the respective cancer types to proliferate. In various embodiments,, the disease node component identifies disease node based on observed correlations between genetic factors and the cancer survival times provided in the patient survival information. For example, a genetic factor that is frequently observed in association with short survival times of a specific type of cancer and less frequently observed in association with long survival times of the specific type of cancer can be identified as an active genetic factor having an active role in the genetic mechanism of the specific type of cancer, including subtypes.

[0182] In various embodiments, disease node determination and identification includes disease association parameters regarding associations between different cancer types to facilitate identifying the active genetic factors associated with the different cancer types. For example, cancer types which are highly associated can share one or more common critical underlying genetic factors. As readily appreciated by one of ordinary skill, models (e.g., survival model) of associated cancer types dialectically exchange information to determine and / or identify active genetic factors across types of cancer, including subtypes. In various embodiments, the disease association parameters appliedby the disease node determination and identification is facilicated by modeling. In various embodiments, generation of individual survival models can employ one or more machine learning algorithms to facilitate the determination and / or identification of the survival, modeling, disease node associated with a particular type of cancer, including subtypes, based on the, and the patient genetic information and the disease association parameters.

[0183] In some embodiments, in association with node determination and identification for cancer type, including subtypes, includes determination of a score system for the disease node(s). For example, a score for a disease node with respect to a specific type of cancer, including subtypes, reflects association of the disease node to the survival time of the specific type of cancer, including subtypes. In various embodiments, scores can be based on a frequency with which a particular genetic factor is directly or indirectly identified for patients diagnosed with a specific cancer type. In various embodiments, analysis includes the aforementioned survival, sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. can be related to, less than a defined threshold, greater than a defined threshold. For example, greater scores associated with a disease node and cancer type, including, the greater the contribution of the disease node to survival time. In various embodiments, formation regarding disease nodes for respective types of cancer, including subtypes and scores determined for the active genetic factors can be collated in a data structure, such as a database.

[0184] Described herein is a method of analysis that includes effects modeling. In various embodiments, the effects modeling includes random effect, fixed effect, mixed effect, linear mixed effect, and generalized linear mixed effect. In various embodiments, the effects includes cubic spline. In various embodiments, the effects modeling includes regression. In various embodiments, the effects modeling includes logistic and Poisson regression. In various embodiments, the model does not include covariates. In In various embodiments, the model includes covariates. In various embodiments, the covariates are information from medical records (including laboratory testing records such as genomic, epigenomic, nucleic acid and other analyte results), insurance records or the like. Examples include age, line of therapies, smoking status (yes / no), gender, and various scoring and / or staging systems that have been utilized for specific cancer disease patients, with an illustrative example including age (in years), line of anti-EGFR therapy,smoking status (yes / no), gender (female / male), and the Van Walraven Elixhauser Comorbidity (ELIX) score specific to lung cancer patients (expressed as a weighted measure across multiple common comorbidities. One of skill readily appreciates covariates can include any number of data elements for individuals and individuals in a population, such as that from medical records (including laboratory testing records such as genomic, epigenomic, nucleic acid and other analyte results), insurance records or the like.

[0185] In various embodiments, the method of analysis includes generation of a hierarchy including at least one first level equation. In various embodiments, a first level equation includes a truncated cubic spline. In various embodiments, the truncated cubic spline includes longitudinal data. This includes, for example, direct or indirect measurements of ctDNA levels, allele fractions, tumor fractions. In various embodiments, an additional level equation includes covariate. In various embodiments, the covariates are information for an individual, or individuals in a population, drawn and / or stored from medical records (including laboratory testing records such as genomic, epigenomic, nucleic acid and other analyte results), insurance records, or the like. Examples include age, line of therapies, smoking status (yes / no), gender, and various scoring and / or staging systems that have been utilized for specific cancer disease patients. In various embodiments, a velocity plot is generated. In various embodiments, the velocity plot is a derivative or one or equations, such as at least one first level equation. In various embodiments, the method of analysis includes one or more of Equations (1), (2) and (3) described in the Examples.

[0186] Described herein is a method of analysis that includes jointly solving different analysis components, including one or more of survival, modeling and sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. as separate components. In various embodiments, the method of analysis includes jointly solving one or more different models for the different cancer types under a joint model framework. For example, the method of analysis could include jointly solving one or more different survival models for the different cancer types under a joint model framework. In various embodiments, the method includes determination of association parameters. In various embodiments, association parameters include, for example, the relationship between patient survival and the patient’s estimated current value of thebiomarker, the relationship between patient survival and patient’s estimated current change over time with respect to the biomarker. In various embodiments, this includes the slope, and the relationship between overall survival and current estimated area under a subject’s longitudinal trajectory as a surrogate for a biomarker’s cumulative effect. It is readily appreciated by one of ordinary skill that association parameters can undertake multiple forms, and can also be combined. For instance, one could examine the relationship between overall survival and estimated current value plus the estimated current slope of the patient’s longitudinal trajectory.Example 1 - Joint Modeling

[0187] Described herein is a hierarchal random effects cubic spline model with advantages in comparison to traditional longitudinal modeling approaches such as the ability to incorporate patient characteristics and create patient-specific results, which can be applied directly in precision oncology settings.

[0188] The aforementioned methodology is well adapted to handling complex longitudinal genomic data. One can apply the methodology in a real-world data setting for interrogation purposes as well as other data settings such as hypothesis generation, making statistical inference, and patient monitoring, each of which can generate patientlevel results, where such results can add to our understanding of ctDNA dynamics and enhance our ability to integrate ctDNA into clinical decision-making.

[0189] In an initial study, 167 patients with colorectal cancer were selected from a real- world database linking genomic and insurance claims data. All patients received chemotherapy and had at least three serial liquid biopsy tests completed via a genomic test. To meet model assumptions, ctDNA levels, measured by maximum variant allele frequency on each test, were transformed into logits. Due to the model’s hierarchal structure, an unconditional cubic spline model was fit first, producing an estimated response pattern for the cohort. First, an equation can assumes the form of a truncated cubic spline and captures how a particular patient’s ctDNA levels change over time. Next, as patient-level results are of interest, the unconditional model was built upon by fitting a conditional model that incorporated covariates consisting of demographic and health status information, which provided numerous patient-level response patterns. For example, a second equation can be a unique linear regression equation relating the distribution of each response parameter to the covariates. Subsequently, eachconstellation of covariate values governs the shape of the trajectory estimated by the model, meaning, that for each patient, a “custom” trajectory, based on his or her own characteristics, is rendered. Other examples for non-small cell lung cancer (NSCLC) models, cancer patients receiving anti-EGFR therapy, among many others.

[0190] A high volume including numerous patient-level projections are generated (each covariate value combination produces a unique projection), an R-Shiny application was developed to visually present and compare results. Additionally, to enhance the understanding of patient response patterns, velocity plots, which provides the instantaneous rate of change in ctDNA levels at different time points, are also provided.

[0191] The methods and systems described herein demonstrate that the proposed method can successfully be applied to genomic data to describe and explore complex patient-level temporal ctDNA patterns while accounting for the impact of covariate values have on these patterns. The proposed methodology can be applied in a wide variety of settings, ranging from hypothesis testing in clinical trials to patient monitoring. Regardless of the setting, results from the model can further our conceptualization of ctDNA dynamics and enhance our ability to integrate these results into targeted, patient centric, clinical decision-making.

[0192] The aforementioned approach can equally be applied to randomized studies where patients sampled are representative of a target population. Under such conditions, since the potential exists to generate and compare many different response patterns, the number of comparisons should be kept small, based on pre-determined hypotheses, and typical considerations in randomized designs such as controlling for type-I error should be made. Another application of this approach is patient monitoring. Here, each response pattern is a reasonable portrayal of a response pattern for a patient with the same set of characteristics — and — in this way — each response pattern serves as a baseline, or reference response pattern. Additionally, if survival information (deceased yes / no) is incorporated into the modeling, a reference response pattern for survivors and non-survivors can be created while keeping the remaining covariate values constant.Thus, if the response pattern of a new patient behaves like the reference response pattern of a survivor, this indicates that intervention is unnecessary. Alternatively, if his or her response pattern mirrors that of a non-survivor, this may serve as a warning that intervention is needed. Additionally, model adjustments can be made regardless of the application. First, an association between the response parameter and covariate may benon-linear, and therefore, imposing a linear restriction will fail to correctly specify the model. A convenient way to assess this relationship is to create a scatterplot, where the association between the response parameter and a covariate can be visually assessed, and based on this assessment, correct model adjustment made. Second, collinearity may be present in the second-level equations leading to inflated standard errors. Thus, removing the responsible covariates may improve the accuracy of results.Example 2 - Data Source, Patient Cohort, Response Variables and Study Covariates

[0193] In a study related to non-small cell lung cancer (NSCLC), data used for analysis included genomic, epigenomic, real-world outcomes, and exemplary patient specific information. Patients in a cohort diagnosed with stage III or stage IV NSCLC were treated with immune checkpoint inhibitors (ICI) . Including the baseline measure prior to initiating ICI treatment, patients had at least three serial Guardant Reveal genomic and epigenomic detection platform liquid biopsy tests. Due to the clinical relevancy of the window, only DNA methylation measures between 0 to 420 days post treatment initiation were considered as only patients who were still alive had TMSs captured beyond 420 days.

[0194] One can apply joint modeling, which is composed of sub-models, one for the longitudinal component and the other for the time-to-event component, three response variables were considered. The TMS was used in the analysis of the longitudinal component, and OS and PFS were used in the analysis of the time-to-event component, resulting in constructing a JM for OS and another for PFS. In some instances, the reported TMSs fell below the limit of detection (LoD) of the assay. When this occurred, scores were replaced with 0.05%, a value consistent with the LoD of the test. However, a sensitivity analysis was performed to assess the impact of using values other than 0.05%, where values ranging between 0.01% and 0.1% were investigated to ensure using different values near the LoD had minimal impact on results. Regarding the time-to- event outcomes, both OS and PFS were censored if the patient did not experience the event, or the patient dropped out of the study or was lost to follow-up. As the data originates from a prospective study, only right-censoring is applicable. Baseline covariates were collected pre-treatment and included age (in years), smoking status (yes / no), gender (female / male), cancer stage (III / IV) and a comorbidity index (CI), expressed as a sum score across multiple common comorbidities.Example 3 - Modeling Approach

[0195] While both frequentist and Bayesian JM approaches have been deployed, a Bayesian approach was applied due to ability to handle more complicated models. As described, the JM framework involves evaluating two sub-models, one for the longitudinal data and the other for the time-to-event data. Thes sub-models are analyzed simultaneously, and their information is integrated to assess the relationship between the two. Given the complex nature of the biomarker progression, both within and between patients, a hierarchical cubic spline random effects model (HCSREM) was adopted to assess the temporal progression of TMS. Concurrently, OS and PFS were analyzed with a Cox-proportional hazard (CPH) model. As is standard practice in hierarchical modeling, continuous variables were centered about their respective means and the TMSs underwent a logit transformation (tTMS) to better conform to the normality assumption of the HSCREM. Given that baseline covariates such as age, gender, smoking status, cancer stage, and patient health are associated with the longitudinal patterns of the biomarker and are also confounded with OS and PFS, these covariates were intentionally incorporated into both sub-models.

[0196] A primary objective when constructing a JM is to determine the nature of the association between the longitudinal and time-to-event data. This association is captured by what is called a functional form. Functional forms come in many flavors and include the relationship between patient outcome and at least one of the following: (1) the estimated current value of the biomarker, (2) the estimated current change over time with respect to the biomarker i.e., the slope, and (3) the current estimated area under the patient’s longitudinal trajectory, often used as surrogate for the biomarker’s cumulative effect. These functional forms can also be combined to explore more complex associations. For instance, one might examine the relationship between overall survival and estimated current value plus the estimated current slope of the patient’s longitudinal trajectory. In our analysis the Inventors explored the current value, slope, area, and their linear combinations and considered associated p-values less than an alpha-level of 0.05 to be actionable, where actionable p-values combined with information criteria informed the final model selection. Once a functional form is identified it is leveraged to produce dynamic predictions. That is, patient-level OS and PFS probabilities are predicted (in large part) based on the functional form that links the sub-models.Example 4 - Study Results

[0197] All statistical analysis was performed using R version 4.1.3 with the JMBayes2 package utilized for executing the JM. A complete case analysis was performed as patients with missing covariate values were excluded, though the missing data impact was minimal (see patient funnel in the supplementary information). The final cohort included 254 patients with 1000 longitudinal measures. The baseline characteristics of the cohort were as follows: average age of 69 years, 54% of are male, 42% had stage 3 cancer at baseline, 6% never smoked, and the average number of comorbidities was 5.2. Regarding the longitudinal response variable, the mean TMS was 2.25%, with a minimum value set to the LoD and a maximum value of 72.2%. Additionally, for the time-to-event variables, 28% of patients were deceased and 36% experienced disease progression at study completion. A more involved description of the baseline covariates can be found in Table 1 and spaghetti plots of the raw data, logit transformed data, and density plots capturing the distributions of the study duration times based on survival status and disease progression are displayed in Figure 1.Table 1. Cohort Characteristics Summarized by Survival Status (Alive vs Deceased) and Disease Progression (Not Progressed vs Progressed)Examples 5 - Analysis

[0198] When projected onto the ordinate, the spaghetti plot of the raw data (Figure la) reveals a pronounced skew in the TMSs where the spaghetti plot of the logit transformedTMSs (tTMSs) shows the skew is alleviated (Figure lb). The spaghetti plot in panel lb provides insight into the complexity of the data and highlights the large amount of variability in the longitudinal tTMSs both within and between patients. The density plots display the study duration distributions grouped by survival status (Figure 1c) and disease progression (Figure Id), where patients that survive, or do not experience disease progression, are ultimately censored in the time-to-event analyses. In general, density plots reveal what is expected: censored patients experience longer study durations compared to non-censored patients and disease progression times generally lag behind survival times.

[0199] As a cubic spline model is at the core of the HSCREM, Akaike information criteria (AIC) was used to identify the model that provides the best representation of biomarker evolution — found to be a basis cubic spline model with 5 degrees of freedom. This model was selected by comparing a number of natural and basis spline models, each with different degrees of freedom. Once an optimal spline model was found, study covariates were then incorporated. Since HCSREM parameter estimates are uninterpretable alone, results are displayed graphically (see Figures 2 & 3). Of the functional forms investigated, the current value and area under the longitudinal trajectory displayed highly significant associations (p-values < 0.0001) for the JMs based on OS and PFS respectively. However, for both OS and PFS, the JMs using the current value had the lowest Wantanabe-AIC (WAIC) and deviance information criteria (DIC). Consequently, the current value of the tTMSs, as estimated by the HCSREM, is used to predict both OS and PFS. The JM results for OS and PFS are provided in Table 2.Table 2 Joint Model Results for Overall Survival and Progression Free Survival Respectively1. SD = Standard Deviation2. CI = Credible Interval3. LB = Lower Bound of the CI4. UB = Upper Bound of the CI

[0200] Care was taken to ensure model parameters were estimated accurately where default non-informative priors contained within the JMBayes2 package were used. For each JM, 3 Markov chain Monte Carlo simulations were used to generate posterior distributions of the parameter estimates. Each chain consisted of 9000 burn-in iterations followed by 36,000 iterations. A thinning factor of 3 was used to mitigate autocorrelation. As evidenced by the R-hat values, model convergence for all parameter estimates was achieved.

[0201] For those unfamiliar with the presentation of JM results, the summarized output in Table 2 aligns closely with that of a traditional CPH model. For instance, if hazard ratios (HR) are desired, parameter estimates can be exponentiated. To demonstrate, since our focus is on the current value of the tTMSs, a one-unit-increase on the logit scale translates into a 1.40 (exp(0.336)) fold increase in the patient’s risk of experiencing death at a given time point, after controlling for study covariates. As results are Bayesian, credible intervals instead of frequenti st-based confidence intervals are reported. While the HR is useful for reflecting overall trends, from a precision oncologyperspective, the strength in the JM lies in its ability to produce patient-level dynamic predictions. As the concept of dynamic prediction is best understood visually, graphical depictions of dynamic predictions for the OS JM (Figure 2), and PFS JM (Figure 3) for two different patients is provided below.Example 6 - Further Analysis

[0202] “Longitudinal Trajectory” panels are not the actual biomarker values themselves. These functions are flexible and do well to map out the change in tTMSs over time — a direct consequence of using the HCSREM. Estimated longitudinal trajectories and corresponding survival curves are both accompanied by 95% credible bands. For the longitudinal outcome, credible bands remain relatively stable, but, for the survival curves, credible bands widen the further one migrates from the latest longitudinal estimate. Since results in Table 2.0 suggest OS and PFS predicted probabilities decrease when estimated tTMS values increase (and vice versa), patientlevel dynamic predictions should bear this out.

[0203] Transitioning from a broad overview to concentrate on specifics, the Inventors consider patient 1 in Figure 2. For this patient, the most recent tTMS estimate at 5 days (2a) is -7.75, closely aligning to the estimate at the landmark timepoint of 50 days (2c) of -7.55. To present results on a familiar scale, back-transformed estimated TMS values are also provided, where -7.75 back-transforms to a TMS of 0.043% and -7.55 back- transforms to a TMS of 0.053%. By examining the corresponding survival curves at 100 days post biomarker update i.e. at 105 (2b) and 150 (2d) days respectively, where the blue dashed lines intersect (used to compare within patient results), the Inventors observe estimated survival probabilities to be nearly identical (approximately 0.99). However, at 200 days (2e), the most recent estimated tTMS increases to -6.7 (TMS = 0.12%), resulting in a drop in the predicted survival probability to 0.92 (2f). Finally, at 400 days (2g), the estimate increases markedly to -2.3 (TMS=9.11%), and, consequently, the associated predicted survival probability reduces to 0.13 (2h). Results for patient 1 are consistent with what is expected — as estimated tTMSs increase predicted survival probability decreases. One should recognize that extending the timeframe beyond 100 days from the most recent biomarker update to assess survival probability is somewhat arbitrary, as a similar pattern is observed if the timeframe was set to 150 day (although uncertainty grows as the timeframe increases). In selecting a timeframe, the non-linearnature of the survival curve should be considered, as, in some cases, it can provide contradictory results occur. This is apparent by scrutinizing Figures 2j and 21. Using a timeframe of 500 days, the predicted survival probability in Figure 2j is 0.62 while in Figure 21 it is 0.65, despite the fact that the estimated tTMS in Figure 2i is lower than in 2k. In selecting a timeframe, one must also be mindful of the study end time. As the end time in our example is 550 days, selecting a 250 day timeframe would be nonsensical, as, at 400 days, using this timeframe would exceed the study end time. In the above example the Inventors examined OS for patient 1, however, inspection of the dynamic predictions based on PFS for the same patient also produce expected results (see Figure 3) as does scrutinizing the dynamic predictions for patient 2 for both OS (Figure 2) and PFS (Figure 3).

[0204] Of interesting is comparison between patients and evaluate OS between patients 1 and 2 (see Figure 2). When comparing patients, it is crucial to recognize that patient characteristics can impact results — a topic addressed shortly. It is often informative to use the green dashed line (although the blue dashed line can also be used), which corresponds to the “end of study” predicted survival probability, when between patient contrasts are desired. Furthermore, when comparing survival probabilities between patients, equivalent biomarker measurement times should be used. For example, at 5 days, the estimated tTMS for patients 1 (2a) and 2 (2i) are -7.75 (TMS=0.043%) and -4.9 (TMS=0.73%) respectively. Since patient 2 has a higher tTMS, as anticipated, the predicted survival probability for this patient (2j) is lower (0.57) as compared to patient 1 (0.62) (2b). This pattern continues as the biomarker evolves but is most prevalent at 400 days when the estimated tTMS for patient 1 (2g) is -2.6 (TMS=6.9%) and -7.5 (TMS=0.055%) for patient 2 (2o), resulting in corresponding predicted survival probabilities of 0.06 (2h) and 0.99 (2p). In Figure 3.0 the same patients are used to estimate PFS dynamic predictions. While the OS and PFS dynamic prediction patterns are similar (see Figure 2 and 3), PFS probabilities are generally lower than their OS counterparts, which aligns with expectations since patients typically undergo disease progression before death. To illustrate, examine patient 1 at 200 days; the predicted survival probability at 550 days (2f) is 0.65, but the predicted probability they remain disease free is 0.61 (Figure 3f).

[0205] The model yields consistent results when comparing within and between patients and between OS to PFS, but an important detail has been neglected — theinfluence of patient characteristics. As emphasized throughout the manuscript, an attractive JM feature is covariates are readily incorporated into both sub-models. The implication is, longitudinal trajectories are fit only fit to the data but are also informed by patient characteristics, and survival curves are statistically controlled / adjusted by the same characteristics. Thus, for patients who share identical longitudinal tTMSs, but differ in characteristics, the Inventors would not expect to observe the same estimated longitudinal trajectory or corresponding survival curve. To demonstrate, consider the dynamic predictions presented in Figure 4. The longitudinal measurements for patient 3 are identical to those of patient 1 (see Figure 2), however, patient 1 is 71 years old with a CI of 6, where patient 3 is 44 years old with a CI of 1. By examining results across like panels, though slight differences in the longitudinal trajectories are observed, patient 3 consistently maintains a better OS outlook — reflecting what would be expected in a younger, healthier patient. As a final note, it is worth mentioning that the JM can produce TMS estimates slightly below the LoD. This occurs because constraints are applied only to the lower limit of the TMS rather than to the logit scale itself, meaning it is possible for the model to produce TMS estimates below the LoD, but not at or below zero. In conclusion, while JM interpretation requires some guidance, and, at times disparate results do arise, more often than not, dynamic predictions align with expectations if a strong association between sub-models exists.Example 7 - Data Source and Patient Cohort, Response Variables and Study Covariates

[0206] In an additional study, the data originates from consists of a prospective single institutional cohort of ER+ / HER2- mBC patients who received endocrine therapy (ET) and CDK4 / 6-inhibitors (CDK4 / 6i). Contained within the data is information regarding genomic, epigenomic, real-world outcomes, and exemplary patient-specific details. Of the 57 patients in the study, 49 met the criteria of having at least three temporal measures (279 in total) and no missing covariate values, resulting in a complete case dataset. Note that some patients received subsequent therapy post-baseline, and for this analysis, their corresponding temporal measures were removed. The longitudinal response variable, the TMS, originally reported as a precent, was transformed onto the logit scale (tTMS) to better adhere to model assumptions — though back-transformed values are calculated to enhance result interpretation. Additionally, only longitudinal measures between 0 and1300 days were considered due to data sparsity issues beyond this timeframe. When reported TMSs fell below the assay’s limit of detection (LoD), scores were replaced with 0.05%, consistent with the LoD of the test. A sensitivity analysis was performed using values ranging between 0.01% and 0.1% to assess the impact of using other values near the LoD. Additional TMS information is presented in Table 1.0 accompanied by a spaghetti plot of the raw and transformed data (see Figure 5). The two time-to-event response variables are OS and PFS respectively, where patients that did not experience an event, dropped out of the study, or were lost to follow-up, are right censored. A summary of the time-to-event response variables can be found in Table 1.0. All covariates included in the JM were captured at baseline and include patient age (median=62 years), CDK4 / 6i drug used [Palbociclib (71.4%) vs Riboci clib or Abemaciclib], histology [ductal (81%) vs lobular or mixed histology], current line of therapy [1 (75.5%) vs 2 or more], and prior adjuvant treatment [no (57.1%) vs yes], A detailed accounting of the cohort characteristics can be found in Table 3.Example 8 - Modeling Approach

[0207] As described, JM is divided into two sub-models, one interrogates the longitudinal data while the other interrogates the time-to-event data. Since the temporal progression of the tTMS varies substantially between and within patients, a flexible hierarchical cubic spline random effects model (HCSREM) was employed to account for this variability. For the time-to-event sub-model, a traditional Cox-proportional hazard (CPH) model was utilized to analyze OS and PFS — with separate JMs constructed for each. A notable aspect of the sub-models is their ability to readily incorporate covariates, where study baseline covariates were added to each with the purpose of expanding the explanatory capabilities of the models and / or to serve as statistical controls. Ultimately, information from the sub-models is synthesized by relating the evolution of tTMS to the time-to-event outcome through an association structure, also known as a functional form. Although numerous association structures are available, the most common assess the relationship between patient outcomes and: (1) the estimated current value of the biomarker, (2) the estimated current instantaneous rate of change (slope) of the biomarker, and (3) the current estimated area under longitudinal trajectory, a measure of the biomarker’s cumulative effect. Linear combinations of these association structures are also possible. As an example, the association between overall survival andestimated current value plus the estimated current slope of the longitudinal trajectory could be explored. In our analysis the Inventors investigated the current value, slope, area, and their linear combinations and considered association structure p-values less than an alpha-level of 0.05 to be significant. Final model selection was based on a combination of significant p-values and information criteria.Example 9 - Study Results

[0208] A complete case analysis was performed where R version 4.1.3 was used to conduct all analyses, and more specifically, the JMBayes2 package executed the JM. Additional aspects of the cohort are described in Table 3, grouped by alive and deceased patients, and by patients whose breast cancer progressed vs. those that did not. To complement the information reported in Table 3, Figure 5 contains spaghetti plots that display the raw and logit transformed longitudinal data, grouped by survival and diseased progression status.Table 3. Cohort Characteristics Summarized by Survival Status (Alive vs Deceased) and Disease Progression (Not Progressed vs Progressed)Example 10 - Analysis

[0209] Referencing Figure 5a and 5c in Figure 5, the Inventors observe that by collapsing the temporal measures on the ordinate, TMSs are highly skewed. In contrast, by repeating this process for the logit transformed TMSs (tTMSs), the skew becomes attenuated (Figures 5b and 5d). The spaghetti plots in panels lb and Id provide a visual account of the complex biomarker evolution both within and across patients. Additionally, Figure 5b illustrates how temporal trajectories differ between survivors and non-survivors, and Figure 5d provides a similar insight into trajectories for those who experienced diseased progression vs those who do not.

[0210] A hierarchical cubic spline random effects model (HCSREM) was used for the longitudinal sub-model, where with its ability to incorporate study covariates, the malleable properties of the HCSREM motivated its selection. Since a cubic spline model underlies the process of capturing tTMS progression, Akaike information criteria (AIC) was employed to identify the best fit model without overcomplicating the model. By investigating different natural and cubic spline models, a basis cubic spline model with 6 degrees of freedom was found to be the most optimal. In conjunction with the longitudinal sub-model, a CPH model was used to assess OS and PFS, where study covariates were included into the model and serve as statistical controls.

[0211] In executing the JM, a Bayesian approach was adopted to determine the relationship between sub-models, as this approach offers more model specification options compared to frequentists approaches. During this process, various associationstructures and their linear combinations were explored. Among these structures, the current biomarker value and area under the longitudinal trajectory of the biomarker displayed highly significant relationships (p-values < 0.0001 for each JM i.e. one OS and other on PFS). However, for both OS and PFS, the JMs based on the current biomarker value demonstrated the lowest deviance and Watanabe-Akaike information criteria. As a result, the current value of the tTMSs was used as the basis for conducting dynamic predictions. The JM results for OS and PFS are provided in Table 4.Table 4. Joint Model Results for Overall Survival and Progression Free Survival Respectively1. SD = Standard Deviation5. CI = Credible Interval6. LB = Lower Bound of the CI7. UB = Upper Bound of the CI8. Note that, as is common with spline models, since results are uninterpretable, they are displayed graphically. Hence the results of the HCSREM portion of the JM are displayed graphically in Figures 2.0 & 3.0.

[0212] The model parameters in Table 4 were estimated using default non-informative priors within the JMBayes2 package. To ensure parameter estimate accuracy, three chains were used, each based on 15000 burn-in iterations followed by 45,000 iterations, where a thinning factor of 3 was used to alleviate autocorrelation issues. Convergence of all parameter estimates is supported by the R-hat values, and since a Bayesian analysis was performed, credible intervals instead of frequentist-based confidence intervals are reported.

[0213] The output in Table 4 can be likened to that of a traditional CPH model. For instance, hazard ratios (HR) can be estimated by exponentiating parameter estimates. To demonstrate, since our focus is on the current tTMS, a one-unit-increase on the logit scale translates into a 1.82 (exp(0.60)) fold increase in the patient’s risk of experiencing death at a given time point, holding the remaining covariates constant. Calculating the HR is informative if the desire is to capture a cohort-level trend, but, since the primary objective is to demonstrate the usefulness of patient-level dynamic predictions, the remainder of our discussion focuses on this aspect. As dynamic predictions are best understood visually, Figures 6 and 7 contains graphical depictions of dynamic predictions for two different patients based on the OS and PFS JMs, respectively.Example 11 -Representation of the biomarker evolution, event probability

[0214] Regarding reported probabilities, it is important to be mindful that probabilities are conditional in the sense that the survival curve is “re-set” each time the estimated biomarker value is updated. Note that estimated tTMSs are not biomarker values themselves (indicated by the dots) but instead are given by the blue functions (see the “Longitudinal Trajectory” panels). As a direct consequence of employing the HCSREM, these functions are agile and do well to capture the intricacies of the temporal progression of tTMSs. Longitudinal trajectories and their related survival curves are both accompanied by 95% credible bands. Credible bands remain relatively constant for the longitudinal outcomes but widen for the survival curves as projections extend further into the future. Results in Table 4 suggest that, in general, as tTMSs increase, OS and PFS outlooks declines. Conversely as tTMSs decrease, OS and PFS outlook improves, where both trends are reflected in the dynamic predictions. For instance, the estimated back-transformed TMS for patient 2 is nearly 0.25% at 5 days (Figure 6i), and their 350- day (approximately 1 year) OS outlook, marked by the intersection of the blue dashedlines, is 99% (Figure 6j). However, as additional longitudinal information is populated, a reduction in the patients updated 350-day OS outlook is observed, where, at 540 days, the estimated TMS is around 10% (Figure 6o), and their 350-day survival probability reduces to 78% (Figure 6p). In Figure 8, a timeframe of 350 days beyond the patient’s most recent longitudinal estimate is used to estimate the survival probability, though this timeframe is somewhat arbitrary and alternative timeframes could be also selected. This is not to say any timeframe will do. Timeframes should not exceed the study end time (given by the green dashed line), should be clinically relevant, and in line with the nature / intent of the outcome measure. For this reason, a clinically relevant approximate 1-year timeframe was chosen for OS, while 100 days was used for PFS, driven by the fact that disease progression precedes mortality. Finally, when comparing wi thin-patient results, timeframes should be equivalent, allowing survival probabilities to be directly compared after biomarker information is revised. For this reason, using the study end time when conducting within-patient dynamic predictions is not ideal because the interval between the most recent longitudinal estimate and the end of the study steadily decreases with each new longitudinal estimate generated.

[0215] Beyond within-patient comparisons, between-patient comparisons can also provide valuable insights. When different patients are contrasted it is often informative to evaluate survival curve probabilities estimated using the study end timeframe as a rough estimate (alternatively the blue dashed line can be used), provided equivalent biomarker measurement times are used in the comparison. The term “rough guideline” is used because patient characteristics can influence results, a topic to be addressed shortly. For example, the estimated TMS for patients 3 and 4 at day 5 are 4.6% (Figure 7a) and 0.2% (Figure 7b) respectively. Consequently, as expected, because patient 3 has a higher TMS, their end of study predicted PFS probability is higher (Figure 7b) as compared to that of patient 4 (Figure 7j). Similar conclusion can be drawn by comparing patients 1 and 2 at day 5.

[0216] Thus far, the results demonstrate that dynamic predictions yield anticipated results. An appealing feature of the JM is the ability to readily incorporate covariates into each sub-model, thereby improving prediction accuracy. To illustrate, analysis was performed on patients who share the same longitudinal biomarker profile but differ in their characteristics. In such a scenario, even though temporal measures are identical, the Inventors would not necessarily expect that patients would experience the sameoutcome, in fact, the Inventors would expect them not to. To demonstrate, one can draw distinctions between the PFS estimates for patient 3 who was 59 years old at baseline, was prescribed Palbociclib, had prior adjuvant therapy, is currently on their first line of therapy, and falls in the ILC / MDLC histology category with patient 3 “adjusted” (see Figure 8) who is 65 year old, was prescribed Ribociclib, did not receive prior adjuvant therapy, is in the IDC histology category, and is on their second line of therapy.Example 11 - Examining across equivalent biomarker measurement

[0217] Timepoints for patient 3 (Figure 7) and patient 3 “adjusted” (Figure 8), it is clear that PFS survival curves differ, meaning, although the current biomarker estimate modifies PFS, the biomarker does not act in isolation but instead acts in concert with patient factors.

[0218] These results further emphasize the robustness of the JM and reinforces its application in precision oncology, as not only are dynamic predictions driven by updated biomarker information, predictions are also adjusted based on a patient’s unique traits.

[0219] One of skill in the art will appreciate that there are numerous biomarkers, cancer types, and mutations available for investigation as the analysis performed here can be applied to other cancer types and mutations, and, in the process, additional relevant biomarkers may be identified. This approach support creation of patient-specific monitoring systems that are both custom-tailored to a specific cancer type and mutation combination.

Claims

THE CLAIMS1. A method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker; and determining a patient response for the at least one patient.

2. The method of claim of any preceding claim, wherein the biomarker comprises circulating tumor DNA (ctDNA).

3. The method of claim of any preceding claim, wherein the biomarker comprises allele frequency and tumor fraction.

4. The method of claim of any preceding claim, wherein determining a patient response for the at least one patient comprises use of a database.

5. The method of any preceding claim, wherein the database comprises medical records and / or insurance records.

6. The method of any preceding claims, wherein use of the database comprises application of a model.

7. The method of claim of any preceding claim, wherein the model is a hierarchal model.

8. The method of claim of any preceding claim, wherein the model is an effects model.

9. The method of claim of any preceding claim, wherein the model is a regression model.

10. The method of claim of any preceding claim, wherein the model is a joint model.

11. The method of claim of any preceding claim x, wherein the hierarchal model is a hierarchical random effects model.

12. The method of claim of any preceding claim, wherein the model comprises a cubic spline.

13. The method of claim of any preceding claim, wherein the model comprises a regression model.

14. The method of claim of any preceding claim, wherein the hierarchal random effects model comprises generation of data from nucleic acid sequence information comprising temporal changes in a biomarker comprising circulating tumor DNA (ctDNA) from at least one subject in a plurality of subjects.

15. The method of claim of any preceding claim, wherein the generation of data comprises generation of a cubic spline for at least one subject in a plurality of subjects.

16. The method of claim of any preceding claim, wherein the generation of data comprises generation of response parameters comprising one or more covariates.

17. The method of claim of any preceding claim, wherein the generation of data comprises generation of response parameters without covariates.

18. The method of claim of any preceding claim, wherein the response parameters apply a multivariate normal distribution.

19. The method of claim of any preceding claim, wherein the determining a patient response for the at least one patient comprises generation of a velocity plot.

20. The method of claim of any preceding claim, wherein the determining a patient response for the at least one patient comprises comparison to the model.

21. The method of claim of any preceding claim, wherein the joint model comprises at least two models.

22. The method of claim of any preceding claim, wherein the joint model comprises association factors between the at least two models.

23. The method of claim of any preceding claim, wherein the joint model comprises a cubic spline and a proportional hazard model.

24. The method of claim of any preceding claim, wherein the biomarker is measured with next-generation DNA sequencing.

25. The method of claim of any preceding claim, wherein next-generation DNA sequencing comprising ligation of non-unique barcodes to the ctDNA.

26. The method of claim of any preceding claim, wherein next-generation DNA sequencing comprising ligation of unique barcodes to the ctDNA.

27. The method of claim of any preceding claim, wherein next-generation DNA sequencing comprising ligation of non-unique barcodes to ctDNA fragments, wherein the non-unique barcodes are present in at least 20x, at least 30x, at least 50x, or at least lOOx molar excess.

28. A method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA (ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a hierarchal random effects model.

29. The method of claim of any preceding claim, wherein the hierarchal random effects model comprises generation of data from nucleic acid sequence information comprising temporal changes in ctDNA from at least one subject in a plurality of subjects.

30. The method of claim of any preceding claim, wherein the hierarchal random effects model comprises generation of a cubic spline for at least one subject in the plurality of subjects.

31. The method of claim x, wherein the hierarchal random effects model comprises response parameters comprising one or more covariates for at least one subject in the plurality of subjects.

32. The method of any preceding claim, wherein the database comprises medical records and / or insurance records for the plurality of subjects.

33. A method of determining a patient response in at least one patient, comprising, obtaining nucleic acid sequence information from at least one patient, comprising measurements of temporal changes in a biomarker comprising circulating tumor DNA(ctDNA); and determining a patient response for the at least one patient comprising use of a database comprising medical records and / or insurance record from a plurality of subjects wherein use of the database comprises application of a joint model comprising a cubic spline and proportional hazard model generated from data from nucleic acid sequence information for at least one subject in a plurality of subjects.

34. The method of any preceding claim, wherein the database comprises medical records and / or insurance records for the plurality of subjects.

35. A system comprising a machine comprising at least one processor and storage comprising instructions capable of performing any of the preceding methods.

36. A computer readable medium comprising instructions capable of performing any of the preceding methods.

Citation Information

Patent Citations

  • Vehicle Remote Function System and Method for Effectuating Vehicle Operations Based on Vehicle FOB Movement

    US20140253287A1

  • Methods for accurate sequence data and modified base position determination

    US8486630B2

  • Systems and methods to detect rare mutations and copy number variation

    WO2014039556A1

  • Methods and systems for analyzing nucleic acid molecules

    WO2018119452A2

  • Compositions and methods for isolating cell-free DNA

    WO2020160414A1