Joint modeling of longitudinal and time-to-event data to predict patient survival

By employing a hierarchical random-effects model with cubic splines to analyze ctDNA using longitudinal and time-to-event data, the method addresses the challenge of distinguishing cancerous signals from healthy tissues, enabling early cancer detection and effective treatment monitoring.

JP2026505886APending Publication Date: 2026-02-19GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025540190
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-01-11
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Current methods for characterizing circulating tumor DNA (ctDNA) in liquid biopsies struggle to distinguish signals from diseased tissues from those of healthy tissues, particularly in early stages of cancer, leading to challenges in early detection and characterization.

Method used

The use of longitudinal data and time-to-event data in conjunction with a hierarchical random-effects model and cubic splines to analyze ctDNA, incorporating medical and insurance records, allows for the characterization of temporal changes in biomarkers and patient response.

Benefits of technology

This approach enhances the ability to diagnose cancer at early stages, monitor treatment response, and predict prognosis by accurately tracking temporal changes in ctDNA, thereby improving patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505886000001_ABST
    Figure 2026505886000001_ABST
Patent Text Reader

Abstract

The changes in ctDNA levels can vary significantly over time for each patient, and the results can be difficult to interpret.Methods and techniques are described herein that are capable of capturing these complexities, while taking into account a diverse set of patient characteristics.Furthermore, the observational data set of cancer patients who have received treatment can be analyzed, and the analysis results can be displayed graphically.These results demonstrate the utility of the described methods and techniques in obtaining a comprehensive understanding of how response patterns evolve and how different patient characteristics affect these developments.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 479,470, filed January 11, 2023, U.S. Provisional Patent Application No. 63 / 496,765, filed April 18, 2023, and U.S. Provisional Patent Application No. 63 / 612,218, filed December 19, 2023, each of which is incorporated by reference herein in its entirety. [Background technology]

[0002] background Today, our understanding of the molecular pathogenesis of cancer is increasing, and with the advancement of next-generation sequencing techniques, the potential for studying early molecular alterations in cancer development is increasing. This includes liquid biopsies in body fluids. Genetic and epigenetic alterations associated with cancer development can be found in cell-free DNA (cfDNA), such as cell-free DNA in plasma, serum, and urine, which has the potential for use as a diagnostic biomarker. Non-invasive sample collection methods are easier, faster, and more economical to perform, thus promoting patient compliance.

[0003] Such liquid biopsy techniques support characterization of the genomic makeup of different tissues in a subject. While cfDNA is generally released by all cell types, it can originate from necrotic or apoptotic cells to identify specific tumor-associated alterations, such as mutations, methylation, and copy number variations (CNVs). Improving the characterization of this circulating tumor DNA (ctDNA) is challenging because it requires distinguishing signals originating from diseased tissues, such as cancer, from signals originating from healthy tissues and germ cells, which are widespread tissues that release cfDNA, such as hematopoietic leukocytes. Signal enrichment can be achieved by identifying variant alleles whose allele fractions do not follow the typical 1:1 ratio for heterozygous alleles in the germline.

[0004] Despite these advances, the majority of the use of cfDNA as a diagnostic method is focused on advanced tumor stage, and little is known about the characterization of early malignant disease stage.In addition, there are several obstacles to early stage detection, including fewer abnormalities, confounding phenomena such as clonal non-tumor tissue expansion, occasional cancer-related mutations, and lack of understanding of the importance of driver changes.Therefore, there is a great demand in the art for improved techniques for characterizing early disease stage, which supports the development of cfDNA and ctDNA-related diagnostic methods.

[0005] The use of detection measurements described herein can include various parameters, including longitudinal data and time-to-event data, thereby supporting understanding of how the temporal changes of biomarkers relate to time-to-event response and patient outcome.For example, the methods and techniques described herein incorporate longitudinal data and time-to-event data to support interpreting the temporal changes of biomarkers in relation to time-to-event response.In addition, the methods and techniques described herein allow for the evaluation of patient characteristics such as age and gender in analysis.Repeated measurements by liquid biopsy provide the opportunity to assess patient outcome. [Brief explanation of the drawings]

[0006] [Figure 1-1] Distribution of allele frequencies and tumor fractions. Figure 1A. Depiction of allele frequencies and tumor fractions for EGFR L858R. Figure 1B. Depiction of allele frequencies and log-transformation for EGFR L858R, KRAS G12D, KRAS G12V. [Figure 1-2] Distribution of allele frequencies and tumor fractions. Figure 1A. Depiction of allele frequencies and tumor fractions for EGFR L858R. Figure 1B. Depiction of allele frequencies and log-transformation for EGFR L858R, KRAS G12D, KRAS G12V.

[0007] [Figure 2-1] Spaghetti plots of allele frequencies and percent tumor. Figure 2A. Spaghetti plots of allele frequencies and percent tumor for EGFR L858R. Figure 2B. Spaghetti plots of allele frequencies and log-transformation for EGFR L858R, KRAS G12D, and KRAS G12V. [Figure 2-2] Spaghetti plots of allele frequencies and percent tumor. Figure 2A. Spaghetti plots of allele frequencies and percent tumor for EGFR L858R. Figure 2B. Spaghetti plots of allele frequencies and log-transformation for EGFR L858R, KRAS G12D, and KRAS G12V.

[0008] [Figure 3] Cubic spline-based GLMM results fitted for log-transformed biomarkers for EGFR L858R.

[0009] [Figure 4] Biomarker evolution and corresponding survival curves. Biomarker EGFR L858R evolution shown for one patient at 300, 600 and 900 days.

[0010] [Figure 5] Biomarker evolution and corresponding survival curves. Biomarker EGFR L858R evolution shown for one patient at 300, 600 and 900 days.

[0011] [Figure 6] Biomarker evolution and corresponding survival curves. Biomarker KRAS G12V evolution shown for one patient at 300, 600 and 900 days.

[0012] [Figure 7]Time-to-event submodel: overall survival. Delineation for EGFR L858R, KRAS G12D, and KRAS G12V.

[0013] [Figure 8] Random-effects modeling. Delineation for EGFR L858R, KRAS G12D, and KRAS G12V.

[0014] [Figure 9] Distribution of ctDNA levels and logit-transformed ctDNA levels and corresponding spaghetti plots.

[0015] [Figure 10] Unconditional model fits with and without data points for the non-small cell lung cancer (NSCLC) cohort. Unconditional model fits with and without data points for the NSCLC cohort. The black curve shows the response pattern for the cohort, while each black point represents a ctDNA level value. The purple area represents the 95% confidence band of the estimated trajectory.

[0016] [Figure 11-1] Response patterns for different values ​​of baseline age and ELIX score for female non-smokers receiving first-line anti-EGFR treatment: FIG. 11A: surviving patients, and FIG. 11B: deceased patients. [Figure 11-2] Response patterns for different values ​​of baseline age and ELIX score for female non-smokers receiving first-line anti-EGFR treatment: FIG. 11A: surviving patients, and FIG. 11B: deceased patients.

[0017] [Figure 12-1] Velocity (IRC) plots for different values ​​of baseline age and ELIX score for female non-smokers receiving first-line anti-EGFR treatment: FIG. 12A: surviving patients, and FIG. 12B: deceased patients. [Figure 12-2]Velocity (IRC) plots for different values ​​of baseline age and ELIX score for female non-smokers receiving first-line anti-EGFR treatment: FIG. 12A: surviving patients, and FIG. 12B: deceased patients. Summary of the Invention [Means for solving the problem]

[0018] Summary of the Invention Described herein are methods for determining patient response in at least one patient, the methods comprising obtaining nucleic acid sequence information from the at least one patient, the nucleic acid sequence information comprising measurements of temporal changes in biomarkers, and determining the patient response for the at least one patient. In various embodiments, the biomarkers comprise ctDNA. In various embodiments, the biomarkers comprise allele frequencies and tumor fractions. In various embodiments, the methods comprise determining the patient response for the at least one patient comprising use of a database. In various embodiments, the methods comprise a database comprising medical and / or insurance records. In various embodiments, the methods comprise use of a database comprising application of a model. In various embodiments, the model is a hierarchical model. In various embodiments, the model is an effect model. In various embodiments, the model is a regression model. In various embodiments, the model is a joint model. In various embodiments, the hierarchical model is a hierarchical random-effects model. In various embodiments, the model comprises cubic splines. In various embodiments, the model comprises a regression model. In various embodiments, the hierarchical random-effects model comprises generating data from nucleic acid sequence information comprising temporal changes in biomarkers, including circulating tumor DNA (ctDNA), from at least one subject among a plurality of subjects. In various embodiments, generating the data includes generating a cubic spline for at least one subject among the plurality of subjects. In various embodiments, generating the data includes generating response parameters including one or more covariates. In various embodiments, generating the data includes generating response parameters without covariates. In various embodiments, the response parameters are fitted to a multivariate normal distribution. In various embodiments, the method includes determining a patient response for at least one patient, including generating a rate plot. In various embodiments, the method includes determining a patient response for at least one patient, including comparing with a model. In various embodiments, the joint model includes at least two models.In various embodiments, the joint model includes a correlation factor between at least two models. In various embodiments, the joint model includes a cubic spline and a proportional hazards model. In various embodiments, the biomarkers are measured by next-generation DNA sequencing. In various embodiments, the next-generation DNA sequencing includes ligating non-unique barcodes to ctDNA. In various embodiments, the next-generation DNA sequencing includes ligating unique barcodes to ctDNA. In various embodiments, the next-generation DNA sequencing includes ligating non-unique barcodes to ctDNA fragments, wherein the non-unique barcodes are present in at least 20-fold, at least 30-fold, at least 50-fold, or at least 100-fold molar excess.

[0019] A system including a machine including at least one processor and storage including instructions capable of performing any of the above methods. A computer readable medium including instructions capable of performing any of the above methods.

[0020] Described herein are methods for determining patient response in at least one patient, the methods comprising: obtaining nucleic acid sequence information from the at least one patient, the nucleic acid sequence information comprising measurements of temporal changes in a biomarker, the biomarker comprising circulating tumor DNA (ctDNA); and determining the patient response for the at least one patient using a database comprising medical and / or insurance records from a plurality of subjects, wherein using the database comprises applying a hierarchical random-effects model. In various embodiments, the hierarchical random-effects model comprises generating data from nucleic acid sequence information comprising temporal changes in ctDNA from at least one subject among the plurality of subjects. In various embodiments, the hierarchical random-effects model comprises generating a cubic spline for at least one subject among the plurality of subjects. In various embodiments, the hierarchical random-effects model comprises a response parameter comprising one or more covariates for at least one subject among the plurality of subjects. In various embodiments, the database comprises medical and / or insurance records for the plurality of subjects. Described herein is a system comprising a machine including at least one processor; and a storage device including instructions capable of performing a method for determining patient response in at least one patient, the method comprising: obtaining nucleic acid sequence information from at least one patient, the nucleic acid sequence information including measurements of temporal changes in biomarkers, including circulating tumor DNA (ctDNA), and determining the patient response for at least one patient by using a database including medical and / or insurance records from a plurality of subjects, wherein the use of the database includes applying a hierarchical random effects model. In various embodiments, the hierarchical random effects model comprises generating data from nucleic acid sequence information including temporal changes in ctDNA from at least one subject among the plurality of subjects. In various embodiments, the hierarchical random effects model comprises generating a cubic spline for at least one subject among the plurality of subjects. In various embodiments, the hierarchical random effects model comprises a response parameter including one or more covariates for at least one subject among the plurality of subjects.In various embodiments, the database includes medical and / or insurance records for a plurality of subjects. Described herein is a computer-readable medium comprising instructions capable of performing a method for determining patient response in at least one patient, the method comprising obtaining nucleic acid sequence information from at least one patient, the nucleic acid sequence information including a measure of temporal change in a biomarker, the biomarker including circulating tumor DNA (ctDNA), and determining patient response for the at least one patient using a database including medical and / or insurance records from a plurality of subjects, wherein using the database includes applying a hierarchical random-effects model. In various embodiments, the hierarchical random-effects model includes generating data from nucleic acid sequence information including temporal change in ctDNA from at least one subject among the plurality of subjects. In various embodiments, the hierarchical random-effects model includes generating a cubic spline for at least one subject among the plurality of subjects. In various embodiments, the hierarchical random-effects model includes a response parameter including one or more covariates for at least one subject among the plurality of subjects. In various embodiments, the database includes medical and / or insurance records for a plurality of subjects.

[0021] Described herein is a method for determining the patient response of at least one patient, the method comprising: obtaining nucleic acid sequence information from at least one patient, comprising measurements of the temporal changes of biomarkers, including circulating tumor DNA (ctDNA); and determining the patient response for at least one patient, comprising using a database comprising medical records and / or insurance records from multiple subjects, wherein using the database comprises applying a joint model comprising a cubic spline and a proportional hazards model generated from data from the nucleic acid sequence information for at least one subject among the multiple subjects.In various embodiments, the database comprises the medical records and / or insurance records of multiple subjects. The present specification describes a system comprising a machine that includes at least one processor; and a storage device that includes instructions capable of performing a method for determining the patient response in at least one patient, the method comprising: obtaining nucleic acid sequence information from at least one patient, comprising measurements of the time change of biomarkers, including circulating tumor DNA (ctDNA), and determining the patient response for at least one patient, comprising using a database that includes medical records and / or insurance records from multiple subjects, wherein using the database comprises applying a joint model that includes a cubic spline and a proportional hazards model generated from data from the nucleic acid sequence information for at least one subject among the multiple subjects.In various embodiments, the database includes the medical records and / or insurance records of multiple subjects.Described herein is a computer-readable medium that includes instructions capable of performing a method for determining patient response in at least one patient, the method comprising: obtaining nucleic acid sequence information from at least one patient, the nucleic acid sequence information comprising measurements of temporal changes in biomarkers, including circulating tumor DNA (ctDNA), and determining the patient response for at least one patient by using a database that includes medical records and / or insurance records from multiple subjects, wherein the database use comprises applying a joint model that includes a cubic spline and a proportional hazards model generated from data from the nucleic acid sequence information for at least one subject among the multiple subjects.In various embodiments, the database includes medical records and / or insurance records for multiple subjects. DETAILED DESCRIPTION OF THE INVENTION

[0022] Detailed Description analysis The present method can be used to diagnose the presence of a condition, particularly cancer, in a subject, characterize the condition (e.g., determine the stage of the cancer or determine the heterogeneity of the cancer), monitor the response to treatment of the condition, and indicate the prognostic risk of developing the condition or the subsequent course of the condition. The present disclosure can also be useful in determining the effectiveness of a particular treatment option. If treatment is successful, the success of the treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood, as more cancer cells may die and shed DNA. In other instances, this may not occur. In another example, perhaps a particular treatment option can be correlated with the genetic profile of the cancer over time. This correlation can be useful in selecting a therapy. Additionally, if the cancer is observed to be in remission after treatment, the present method can be used to monitor for residual disease or disease recurrence.

[0023] The types and number of cancers that can be detected may include blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc. The type and / or stage of cancer may be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.

[0024] Genetic and other analyte data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage determination. Genetic profile data can enable characterization of specific subtypes of cancer, which can be important in diagnosing or treating that subtype. This information can also provide the subject or practitioner with clues about the prognosis of a particular type of cancer, allowing either the subject or practitioner to tailor treatment options according to the progression of the disease. Some cancers can progress and become more aggressive and genetically unstable. Other cancers can remain benign, inactive, or dormant. The systems and methods of the present disclosure can be useful in determining disease progression.

[0025] This analysis is also useful in determining the effectiveness of certain treatment options.If treatment is successful, more cancer cells will die and DNA may be lost, so the success of treatment options may increase the amount of copy number variations or rare mutations detected in the subject's blood.In other cases, this may not happen.In another example, perhaps a certain treatment option may be correlated with the genetic profile of cancer over time.This correlation may be useful in selecting treatment.In addition, if cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.

[0026] The method can also be used to detect genetic mutations in conditions other than cancer. Immune cells, such as B cells, can undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number variation detection, and certain immune conditions can be monitored. In this example, copy number variation analysis can be performed over time to produce a profile of how a particular disease may be progressing. Detection of copy number variations, or even rare mutations, can be used to determine how pathogen populations change during the course of infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, whereby the virus can change life cycle states and / or mutate to more virulent forms during the course of infection. As immune cells attempt to destroy transplanted tissue, the method can be used to determine or profile the host's body's rejection activity to monitor the status of the transplanted tissue or to modify the course of rejection treatment or prevention.

[0027] For example, many types of dysfunctions and abnormalities that commonly occur in the cardiovascular system, if not diagnosed or treated correctly, gradually reduce the body's ability to supply enough oxygen to meet the heart's oxygen demands when an individual experiences stress. This progressive decline in the cardiovascular system's ability to supply oxygen under stressful conditions ultimately leads to a heart attack, i.e., a myocardial infarction event caused by an interruption in blood flow through the heart, resulting in oxygen depletion of cardiac muscle tissue (i.e., the myocardium). In many cases, permanent damage occurs to the cells that comprise the myocardium, which in turn predisposes the individual to further myocardial infarction events.

[0028] The disclosed methods can characterize dysfunctions and abnormalities (e.g., hypertrophy) associated with cardiac muscle and valvular tissue, and reduced blood flow and oxygen delivery to the heart are often secondary to weakening and / or deterioration of the current blood and supply systems caused by physical and biochemical stress. Examples of cardiovascular diseases directly affected by these types of stress include atherosclerosis, coronary artery disease, peripheral vascular disease, and peripheral arterial disease, along with various cardiac and arrhythmias that may represent other forms of disease and dysfunction.

[0029] Furthermore, the methods of the present disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such a method may include, for example, generating a genetic profile of extracellular polynucleotides from a subject, where the genetic profile includes multiple data resulting from copy number variation and rare mutation analysis. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition may be a condition that results in a heterogeneous genomic population. In the example of cancer, some tumors are known to contain tumor cells at different stages of cancer. In other examples, the heterogeneity may include multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, with one or more foci likely being the result of metastasis spreading from the primary site.

[0030] The method can be used to generate or profile a fingerprint or dataset that is a summary of genetic information from different cells in a heterogeneous disease, which dataset can include copy number variation and mutation analysis, either alone or in combination.

[0031] This method can be used to diagnose, predict prognosis, monitor or observe cancer or other diseases.In some embodiments, the method herein does not involve diagnosing, predicting prognosis or monitoring fetus, and therefore is not intended for non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, predict prognosis, monitor or observe cancer or other diseases in the subject in utero, where DNA and other polynucleotides can co-circulate with maternal molecules. Methods for analyzing modified nucleic acids

[0032] The present disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylated, histone-linked, and other modifications discussed above). In some such methods, a population of nucleic acids with different degrees of modification (e.g., 0, 1, 2, 3, 4, 5, or more methyl groups per nucleic acid molecule) is contacted with adapters, and the population is then fractionated by degree of modification. The adapters are attached to either or both ends of the nucleic acid molecules in the population. Preferably, the adapters contain a sufficient number of different tags such that the number of tag combinations occurs with low probability; for example, 95, 99, or 99.9% of two nucleic acids with the same start and stop points receive the same combination of tags. After attachment of the adapters, the nucleic acids are amplified from primers that bind to the primer binding sites in the adapters. The adapters, whether with the same or different tags, may contain the same or different primer binding sites, but preferably the adapters contain the same primer binding sites. After amplification, the nucleic acids are preferably contacted with an agent that binds to nucleic acids with modifications (e.g., such agents as previously described). After binding to the drug, nucleic acid is separated into at least two partitions, which have different degrees of modification of nucleic acid.For example, if the drug has affinity for the nucleic acid with modification, the nucleic acid that is over-represented in modification (compared to the median representation in the population) will preferentially bind to the drug, while the nucleic acid that is under-represented in modification will not bind or will be more easily eluted from the drug.After separation, different partitions can then be subjected to further processing steps, typically including further amplification and sequence analysis in parallel but separately.The sequence data from different partitions can then be compared.

[0033] The nucleic acid can be ligated at both ends to a Y-shaped adapter containing a primer binding site and a tag. The molecule is amplified. The amplified molecule is then fractionated by contact with an antibody that preferentially binds to 5-methylcytosine to produce two partitions. One partition contains the original molecule lacking methylation and the amplified copy that has lost methylation. The other partition contains the original DNA molecule with methylation. The two partitions are then processed and sequenced separately, with further amplification of the methylated partition. The sequence data for the two partitions can then be compared. In this example, the tag is used not to distinguish between methylated and unmethylated DNA, but to distinguish between different molecules within these partitions, allowing for the determination of whether reads with the same start and stop points are based on the same or different molecules.

[0034] The present disclosure provides further methods for analyzing a population of nucleic acids, at least some of which contain one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications previously described. In these methods, the population of nucleic acids is contacted with adapters containing one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably, all cytosine residues in such adapters are also modified, or all such cytosines in the primer binding region of the adapter are modified. The adapters are attached to both ends of the nucleic acid molecules in the population. Preferably, the adapters contain a sufficient number of different tags such that the number of tag combinations occurs with a low probability, e.g., 95, 99, or 99.9% of two nucleic acids with the same start and stop points receive the same combination of tags. The primer binding sites in such adapters can be the same or different, but are preferably the same. After attachment of the adapters, the nucleic acids are amplified from primers that bind to the primer binding sites of the adapters. The amplified nucleic acids are divided into first and second aliquots. The first aliquot is assayed for sequence data, with or without further processing. Thus, sequence data for the molecules in the first aliquot is determined regardless of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosines to uracil. The bisulfite-treated nucleic acids are then subjected to amplification primed by a primer to the original primer binding site of the adapter linked to the nucleic acid. Because these nucleic acids retain the cytosines in the primer binding site of the adapter, only the nucleic acid molecules originally linked to the adapter (separate from the amplification product) can now be amplified, while the amplification product has lost the methylation of these cytosine residues and been converted to uracil during the bisulfite treatment. Thus, only the original molecules in the population, at least some of which are methylated, undergo amplification. After amplification, these nucleic acids are subjected to sequence analysis.Comparison of the sequences determined from the first and second aliquots can indicate, among other things, which cytosines in the nucleic acid population have been subjected to methylation. Partitioning of a sample into multiple sub-samples: Sample characteristics: Epigenetic signature analysis

[0035] In certain embodiments described herein, a population of nucleic acids with different forms (e.g., hypermethylated DNA and hypomethylated DNA in a sample, e.g., a set of captured cfDNA described herein) can be physically partitioned based on one or more characteristics of the nucleic acids, such as differential modification or isolation, tagging, and / or sequencing of nucleobases, prior to further analysis. This approach can be used, for example, to determine whether a particular sequence is hypermethylated or hypomethylated. In some embodiments, hypermethylated variable epigenetic target regions are analyzed to determine whether they are indicative of the hypermethylation characteristic of tumor cells, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they are indicative of the hypomethylation characteristic of tumor cells. In addition, partitioning a heterogeneous nucleic acid population can increase rare signals, for example, by enriching rare nucleic acid molecules that are more abundant in one fraction (or partition) of the population. For example, genetic variations present in hypermethylated DNA but less (or absent) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of a single locus or nucleic acid species of the genome can be performed, and thus greater sensitivity can be achieved.

[0036] In some cases, heterogeneous nucleic acid samples are partitioned into two or more partitions (for example, at least 3, 4, 5, 6 or 7 partitions).In some embodiments, each partition is differentially tagged.Tagged partitions can then be pooled together for collective sample preparation and / or sequencing.Partitioning-tagging-pooling process can be carried out two or more times, and each round of partitioning is carried out based on different characteristics (examples provided herein) and is tagged with a differential tag that distinguishes it from other partitions and partitioning means.

[0037] Examples of characteristics that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins binding to DNA. The resulting partitions may include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, partitioning based on cytosine modification (e.g., cytosine methylation) or methylation is performed generally, optionally combined with at least one additional partitioning step that may be based on any of the above characteristics or DNA forms. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and the level of association with one or more proteins, such as histones. Alternatively or in addition, heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules that lack nucleosomes.Alternatively or in addition, heterogeneous population of nucleic acids can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA).Alternatively or in addition, heterogeneous population of nucleic acids can be partitioned based on nucleic acid length (for example, molecules that are up to 160bp and molecules that have a length greater than 160bp).

[0038] In some cases, each partition (representing different nucleic acid forms) is differentially labeled, and the partitions are pooled together and then sequenced.In other cases, the different forms are sequenced separately.In some embodiments, a population of different nucleic acids is partitioned into two or more different partitions.Each partition represents a different nucleic acid form, and the first partition (also referred to as sub-sample) contains DNA with a greater proportion of cytosine modification than the second sub-sample.Each partition is separately tagged.The first sub-sample is subjected to a procedure that affects the first nucleic acid base in DNA differently from the second nucleic acid base in the DNA of the first sub-sample, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, and the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first nucleic acid base and the second nucleic acid base have the same base pairing specificity.The tagged nucleic acids are pooled together and then sequenced. Sequence read data are acquired and analyzed, which includes distinguishing a first nucleic acid base from a second nucleic acid base in the DNA of the first subsample in silico. The tags are used to classify the read data from different partitions. Analysis to detect genetic variants can be performed at the partition level and at the overall nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variants such as CNVs, SNVs, indels, and fusions in the nucleic acids in each partition. In some cases, in silico analysis can include determining chromatin structure. For example, coverage of sequence read data can be used to determine nucleosome positioning in chromatin. Higher coverage can correlate with higher nucleosome occupancy in a genomic region, while lower coverage can correlate with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).

[0039] The sample may contain nucleic acids with a variety of modifications, including post-replication modifications to nucleotides and attachment, usually non-covalently, to one or more proteins.

[0040] In one embodiment, the population of nucleic acids is obtained from a serum, plasma, or blood sample from a subject suspected of having or previously diagnosed with a neoplasia, tumor, or cancer. The population of nucleic acids includes nucleic acids with various levels of methylation. Methylation can result from any one or more post-replicative or post-transcriptional modifications. Post-replicative modifications include modifications of the nucleotide cytosine, particularly at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine. The affinity agent can be an antibody with the desired specificity, its natural binding partner, or a variant thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected, for example, by phage display, to have specificity for a given target.

[0041] Examples of capture moieties contemplated herein include the methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including proteins such as MeCP2 and antibodies that preferentially bind to 5-methylcytosine. Similarly, partitioning of various forms of nucleic acids can be performed using histone-binding proteins, which can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides. For some affinity drugs and modifications, binding to the drug can occur in an essentially all-or-none manner, depending on whether the nucleic acid has the modification, but separation can be a degree of separation. In such cases, nucleic acids that are overrepresented in the modification will bind to the drug to a greater extent than nucleic acids that are underrepresented in the modification. Alternatively, nucleic acids with the modification may bind in an all-or-none manner. However, various levels of modifications can then be sequentially eluted from the binding agent.

[0042] For example, in some embodiments, partitioning can be binary or based on the degree / level of modification. For example, all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning can involve eluting fragments with different levels of methylation by adjusting the salt concentration in a solution containing the methyl-binding domain and the bound fragments. As the salt concentration increases, fragments with higher methylation levels are eluted. In some cases, the final partitions represent nucleic acids with different degrees of modification (over- or under-represented in the modification). Over- and under-representation can be defined by the number of modifications made by a nucleic acid compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is 2, nucleic acids containing more than two 5-methylcytosine residues will be over-represented in this modification, and nucleic acids with one or zero 5-methylcytosine residues will be under-represented. The effect of affinity separation is to enrich for nucleic acids over-represented in the modification in the binding phase and enrich for nucleic acids under-represented in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted and then further processed.

[0043] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be partitioned using sequential elution. For example, the hypomethylated partition (e.g., no methylation) can be separated from the methylated partition by contacting the nucleic acid population with MBD from the kit attached to magnetic beads. The beads are used to separate the methylated nucleic acids from the unmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids with different levels of methylation. For example, the first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, for example, at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids have been eluted, magnetic separation is again used to separate the more highly methylated nucleic acids from nucleic acids with lower levels of methylation. The elution and magnetic separation steps may be repeated to create various partitions, e.g., hypomethylated partitions (representing no methylation), methylated partitions (representing low levels of methylation), and hypermethylated partitions (representing high levels of methylation).

[0044] In some methods, nucleic acids bound to the agent used for affinity separation are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched for nucleic acids with a degree of modification close to the average or median (i.e., intermediate between the nucleic acids that remain bound to the solid phase and the nucleic acids that are not bound to the solid phase upon initial contact of the sample with the agent). Affinity separation results in at least two, and sometimes three or more, partitions of nucleic acids with different degrees of modification. While the partitions are still separate, nucleic acids in at least one, and usually two or three (or more) partitions, are linked to nucleic acid tags, usually provided as components of adapters, and nucleic acids in different partitions receive different tags that distinguish members of one partition from those of another partition. Tags linked to nucleic acid molecules in the same partition can be the same or different from each other. However, if different from each other, the tags can share a portion of their code in common to identify the molecules to which they are attached as molecules of a specific partition. For further details about partitioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some embodiments, nucleic acid molecules may be fractionated into different partitions based on nucleic acid molecules that are bound to a particular protein or fragment thereof and nucleic acid molecules that are not bound to that particular protein or fragment thereof.

[0045] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on the specific properties of the proteins. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that bind to DNA and can serve as the basis for fractionation include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate nucleic acid molecules based on protein-binding regions. Examples of methods used to fractionate nucleic acid molecules based on protein-binding regions include, but are not limited to, SDS-PAGE, chromatin-immunoprecipitation (ChIP), heparin chromatography, and asymmetric flow field separation (AF4).

[0046] In some embodiments, nucleic acid partitioning is performed by contacting the nucleic acid with the methylation binding domain ("MBD") of methylation binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 streptavidin, via a biotin linker. Partitioning into fractions with different degrees of methylation can be performed by eluting the fractions with increasing NaCl concentrations.

[0047] An exemplary method for molecular tag identification of a library partitioned by MBD beads by NGS is as follows.

[0048] A methyl-binding domain protein-bead purification kit is used to physically partition extracted DNA samples (e.g., plasma DNA extracted from human samples) and preserve all eluates from the process for downstream processing.

[0049] To each partition, differential molecular tags and adapter sequences that enable NGS are applied in parallel. For example, hypermethylated, residually methylated ("washed"), and hypomethylated partitions are ligated with NGS-adapters bearing molecular tags.

[0050] All molecularly tagged partitions are remixed and then amplified using adaptor-specific DNA primer sequences.

[0051] The remixed and amplified total library is enriched / hybridized targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).

[0052] The enriched total DNA library is re-amplified and sample tags are added, and the different samples are pooled and multiplexed on an NGS machine.

[0053] Molecular tags are used to identify unique molecules, and bioinformatics analysis of NGS data is performed to further deconvolute samples into differentially MBD-partitioned molecules. This analysis can yield information about relative 5-methylcytosine content for genomic regions, in parallel with standard gene sequencing / variant detection.

[0054] Examples of MBPs contemplated herein include, but are not limited to:

[0055] (a) MeCP2 is a protein that preferentially binds 5-methyl-cytosine over unmodified cytosine.

[0056] (b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 bind preferentially to 5-hydroxymethyl-cytosine over unmodified cytosine.

[0057] (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind 5-formyl-cytosine over unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).

[0058] (d) an antibody specific for one or more methylated nucleotide bases;

[0059] Generally, elution is a function of the number of methylation sites per molecule, with molecules with more methylation eluting at higher salt concentrations. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into distinct populations based on the degree of methylation. Salt concentrations can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three partitions. Molecules are contacted with a solution at a first salt concentration and containing molecules containing a methyl-binding domain, which can be attached to a capture moiety, such as streptavidin. At the first salt concentration, some population of molecules bind to the MBD, while others remain unbound. The unbound population can be separated as a "hypomethylated" population. For example, the first partition, representing a hypomethylated form of DNA, is the partition that remains unbound at low salt concentrations, e.g., 100 mM or 160 mM. The second partition, representing intermediately methylated DNA, is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM, and is separated from the sample. The third partition, representing highly methylated forms of DNA, is eluted using a high salt concentration, e.g., at least about 2000 mM.

[0060] The present disclosure provides further methods for analyzing a population of nucleic acids, at least some of which contain one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications previously described. In these methods, after partitioning, a subsample of nucleic acids is contacted with an adapter containing one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably, all cytosine residues in such adapters are also modified, or all such cytosines in the primer binding region of the adapter are modified. The adapters are attached to both ends of the nucleic acid molecules in the population. Preferably, the adapters contain a sufficient number of different tags such that the number of tag combinations occurs with low probability, e.g., 95, 99, or 99.9% of two nucleic acids with the same start and stop points receive the same combination of tags. The primer binding sites in such adapters can be the same or different, but are preferably the same. After attachment of the adapters, the nucleic acids are amplified from primers that bind to the primer binding sites of the adapters. The amplified nucleic acid is divided into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. Thus, the sequence data for the molecules in the first aliquot is determined regardless of the initial methylation state of the nucleic acid molecule. The nucleic acid molecules in the second aliquot are subjected to a procedure that affects the first nucleobase in DNA differently from the second nucleobase in DNA, where the first nucleobase comprises a modified cytosine at position 5, and the second nucleobase comprises an unmodified cytosine. This procedure can be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. Then, the nucleic acid that has been subjected to the procedure is amplified by a primer that is attached to the original primer binding site of the adapter that is linked to the nucleic acid. These nucleic acids retain the cytosines in the primer binding sites of the adapters, so that only the nucleic acid molecules originally linked to the adapters (separate from their amplification products) can now be amplified, while the amplification products have lost methylation of these cytosine residues and undergone conversion to uracil during bisulfite treatment.Therefore, only the original molecules in the population are methylated at least in part and undergo amplification. After amplification, these nucleic acids are subjected to sequence analysis. The comparison of the sequences determined from the first and second aliquots can indicate, among other things, which cytosines in the nucleic acid population have been methylated.

[0061] Such analysis can be performed using the following exemplary procedure: After partitioning, the methylated DNA is ligated at both ends to Y-shaped adapters containing primer binding sites and tags. The cytosines in the adapters are modified at position 5 (e.g., 5-methylation). The adapter modification serves to protect the primer binding sites in the next conversion step (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosines but affects unmodified cytosines). After the adapters are attached, the DNA molecules are amplified. The amplification products are divided into two aliquots for sequencing with and without conversion. The aliquot that is not subjected to conversion can be subjected to sequence analysis with or without further processing. The other aliquot is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, where the first nucleobase contains a cytosine modified at position 5 and the second nucleobase contains an unmodified cytosine. This procedure can be bisulfite treatment or another procedure that converts unmodified cytosines to uracil. Only the primer binding sites protected by the cytosine modification can support amplification when contacted with a primer specific to the original primer binding site. Thus, only the original molecules are subjected to further amplification, and not the copies from the first amplification. The further amplified molecules are then subjected to sequence analysis. The sequences from the two aliquots can then be compared. Similar to the separation scheme discussed above, the nucleic acid tag in the adapter is not used to distinguish between methylated and unmethylated DNA, but is used to distinguish between nucleic acid molecules within the same partition. subjecting the first sub-sample to a procedure that affects a first nucleobase in the DNA differently than a second nucleobase in the DNA of the first sub-sample;

[0062] The methods disclosed herein include subjecting a first sub-sample to a procedure that affects a first nucleobase in the DNA differently than a second nucleobase in the DNA of the first sub-sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase that is different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base-pairing specificity. In some embodiments, when the first nucleobase is a modified or unmodified adenine, the second nucleobase is a modified or unmodified adenine; when the first nucleobase is a modified or unmodified cytosine, the second nucleobase is a modified or unmodified cytosine; when the first nucleobase is a modified or unmodified guanine, the second nucleobase is a modified or unmodified guanine; and when the first nucleobase is a modified or unmodified thymine, the second nucleobase is a modified or unmodified thymine (wherein modified and unmodified uracil are included in modified thymine for the purposes of this step).

[0063] In some embodiments, the first nucleobase is modified or unmodified cytosine, and the second nucleobase is then modified or unmodified cytosine.For example, the first nucleobase can comprise unmodified cytosine (C), and the second nucleobase can comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC).Alternatively, the second nucleobase can comprise C, and the first nucleobase can comprise one or more of mC and hmC.For example, as shown in the summary above and the following discussion, other combinations are also possible, such as when one of the first and second nucleobases comprises mC, and the other comprises hmC.

[0064] In some embodiments, the technique that affects the first nucleic acid base in DNA differently from the second nucleic acid base in the DNA of the first subsample comprises bisulfite conversion.Bisulfite treatment converts unmodified cytosine and certain modified cytosine nucleotides (for example, 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) into uracil, while other modified cytosines (for example, 5-methylcytosine, 5-hydroxymethylcytosine) are not converted.Therefore, when bisulfite conversion is used, the first nucleic acid base comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine or other cytosine forms that are affected by bisulfite, and the second nucleic acid base can comprise one or more of mC and hmC, for example, mC and optionally hmC.Sequencing of the DNA that has been treated with bisulfite identifies the position where cytosine is read as mC or hmC position. On the other hand, positions read as T are identified as T, or forms of C that are susceptible to bisulfite treatment, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Thus, performing bisulfite conversion on the first subsample as described herein facilitates identifying positions containing mC or hmC using sequence read data obtained from the first subsample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018;9:5068.

[0065] In some embodiments, the procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA of the first subsample comprises oxidative bisulfite (Ox-BS) conversion. In some embodiments, the procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA of the first subsample comprises Tet-assisted bisulfite (TAB) conversion. In some embodiments, the procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the DNA of the first subsample comprises Tet-assisted conversion with a displaced borane reducing agent, where optionally the displaced borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure affecting the first nucleobase in the DNA differently from the second nucleobase in the DNA of the first subsample comprises chemically assisted conversion with a displaced borane reducing agent, where optionally the displaced borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure affecting the first nucleobase in the DNA differently from the second nucleobase in the DNA of the first subsample comprises APOBEC-linked epigenetic (ACE) conversion.

[0066] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample comprises enzymatic conversion of the first nucleobase, such as in EM-Seq. See, for example, Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), which can then be used to deaminate unmodified cytosines, converting them to uracil.

[0067] In some embodiments, the procedure affecting a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first sub-sample comprises separating DNA that originally includes the first nucleobase from DNA that originally does not include the first nucleobase.

[0068] In some embodiments, the first nucleobase is modified or unmodified adenine, and the second nucleobase is modified or unmodified adenine.In some embodiments, the modified adenine is N6-methyladenine (mA).In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA) or N6-formyladenine (fA).

[0069] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases, such as mA, from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific to mA are described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies for various modified nucleobases, such as thymine / uracil forms, including halogenated forms such as 5-bromouracil, are commercially available. Various modified bases can also be detected based on changes in their base-pairing specificity. For example, hypoxanthine is a modified form of adenine, which can result from deamination and is read as G in sequencing. See, e.g., U.S. Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, NY, 2002, chapter 14, "Mutation, Repair, and Recombination." Enrichment / capture steps, amplification, adapters, barcodes

[0070] In some embodiments, the method disclosed herein comprises capturing one or more sets of target regions of DNA, such as cfDNA.Capturing can be performed using any suitable method known in the art.In some embodiments, capturing comprises contacting the DNA to be captured with a set of target-specific probes.The set of target-specific probes can have any of the characteristics described herein for a set of target-specific probes, including but not limited to the characteristics in the embodiments shown above and the probe section below.Capturing can be performed on one or more subsamples prepared during the method disclosed herein.In some embodiments, DNA is captured from at least a first subsample or a second subsample, for example, at least a first subsample and a second subsample. If the first subsample is subjected to a separation step (e.g., separating DNA originally comprising the first nucleobase (e.g., hmC) from DNA originally not comprising the first nucleobase, e.g., hmC-seal), capturing can be performed on any, any two, or all of the DNA originally comprising the first nucleobase (e.g., hmC), the DNA originally not comprising the first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.

[0071] The capturing step can be carried out under conditions suitable for specific nucleic acid hybridization, which generally depends to some extent on the characteristics of the probe, such as length, base composition, etc. Those skilled in the art will be familiar with suitable conditions based on general knowledge in the art of nucleic acid hybridization. In some embodiments, a complex is formed between the target-specific probe and DNA.

[0072] In some embodiments, the methods described herein include capturing cfDNA obtained from a test subject for a plurality of sets of target regions. The target regions include epigenetic target regions, which may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from a tumor or a healthy cell. The target regions also include sequence-variable target regions, which may exhibit differences in sequence depending on whether they originate from a tumor or a healthy cell. The capturing step produces a captured set of cfDNA molecules, and cfDNA molecules corresponding to the sequence-variable target region set are captured with a greater capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the epigenetic target region set. For further discussion of capturing steps, capture yields, and related aspects, see WO2020 / 160414, incorporated herein by reference for all purposes.

[0073] In some embodiments, the methods described herein include contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to a set of sequence-variable target regions with a greater capture yield than cfDNA corresponding to a set of epigenetic target regions.

[0074] Because analyzing sequence-variable target regions with sufficient reliability or accuracy may require a greater sequencing depth than that required for analyzing epigenetic target regions, it may be beneficial to capture the cfDNA corresponding to a set of sequence-variable target regions with a higher capture yield than the cfDNA corresponding to a set of epigenetic target regions.The amount of data required to determine fragmentation patterns (for example, to test for disruption of transcription start sites or CTCF binding sites) or fragment abundance (for example, in hypermethylated and hypomethylated partitions) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence mutations.Capturing target region sets with different yields can facilitate sequencing target regions to different sequencing depths in the same sequencing run (for example, using pooled mixtures and / or in the same sequencing cell).

[0075] In various embodiments, the method further comprises sequencing the captured cfDNA to different sequencing depths for epigenetic target region set and sequence variable target region set, for example, according to the discussion herein.In some embodiments, the complex of target-specific probe and DNA is separated from the DNA that is not bound to the target-specific probe.For example, when the target-specific probe is covalently or non-covalently bound to solid support, washing or suction step can be used to separate unbound material.Alternatively, when the complex has different chromatographic properties from unbound material (for example, when the probe comprises a ligand that binds to chromatographic resin), chromatography can be used.

[0076] As discussed in detail elsewhere herein, the set of target-specific probes may include multiple sets, for example, probes for a set of sequence-variable target regions and probes for a set of epigenetic target regions. In some such embodiments, the capturing step is performed simultaneously in the same container with probes for a set of sequence-variable target regions and probes for a set of epigenetic target regions, for example, the probes for a set of sequence-variable target regions and the set of epigenetic target regions are present in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of probes for the set of sequence-variable target regions is greater than the concentration of probes for the set of epigenetic target regions.

[0077] Alternatively, the capturing step is performed with a sequence variable target region probe set in a first container and an epigenetic target region probe set in a second container, or the contacting step is performed with a sequence variable target region probe set in a first container at a first time and an epigenetic target region probe set at a second time before or after the first time. This approach allows for the preparation of separate first and second compositions containing captured DNA corresponding to the sequence variable target region set and the captured DNA corresponding to the epigenetic target region set. The compositions can be processed separately if desired (e.g., to fractionate based on methylation, as described elsewhere herein) and remixed in appropriate proportions to provide material for further processing and analysis, such as sequencing.

[0078] In some embodiments, the DNA is amplified. In some embodiments, the amplification occurs before the capturing step. In some embodiments, the amplification occurs after the capturing step.

[0079] In some embodiments, the adapter is included in the DNA. This can be done simultaneously with the amplification procedure, for example, by providing the adapter in the 5' portion of the primer, as described above. Alternatively, the adapter can be added by other techniques, for example, ligation.

[0080] In some embodiments, a tag that can be or include a barcode is included in the DNA. The tag can facilitate identification of the origin of the nucleic acid. For example, barcodes can be used to allow the source (e.g., subject) of the DNA to be identified after pooling multiple samples for parallel sequencing. This can be performed simultaneously with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer, as described above. In some embodiments, the adapter and tag / barcode are provided by the same primer or primer set. For example, the barcode can be located 3' of the adapter and 5' of the target-hybridizing portion of the primer. Alternatively, the barcode can be added by other techniques, such as ligation, optionally with the adapter in the same ligation substrate.

[0081] Additional details regarding amplification, tags, and barcodes are discussed below in the "General Aspects of the Method" section, which, to the extent practicable, can be combined with any of the embodiments above and those set forth in the Introduction and Summary sections. Computer systems for processing real-world evidence (RWE)

[0082] The method of the present disclosure can be implemented by using computer system or with the assistance of computer system.For example, this method can include: partitioning sample into a plurality of sub-samples, including a first sub-sample and a second sub-sample, wherein the first sub-sample comprises the DNA with cytosine modification at a greater rate than the second sub-sample; subjecting the first sub-sample to a procedure that affects the first nucleobase in DNA differently from the second nucleobase in the DNA of the first sub-sample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity; and sequencing the DNA in the first sub-sample and the DNA in the second sub-sample, in a manner that distinguishes the first nucleobase from the second nucleobase in the DNA of the first sub-sample.

[0083] In one aspect, the present disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: collecting cfDNA from a test subject; capturing a plurality of sets of target regions from the cfDNA, wherein the plurality of sets of target regions includes a set of sequence variable target regions and a set of epigenetic target regions, thereby producing a captured set of cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the set of sequence variable target regions are sequenced to a greater sequencing depth than the captured cfDNA molecules of the set of epigenetic target regions; obtaining a plurality of sequence read data generated by a nucleic acid sequencer from sequencing the captured cfDNA molecules; mapping the plurality of sequence read data to one or more reference sequences to generate mapped sequence read data; and processing the mapped sequence read data corresponding to the set of sequence variable target regions and to the set of epigenetic target regions to determine the likelihood that the subject has cancer.

[0084] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be supplied in a programming language that may be selected to allow the code to be executed in a pre-compiled or as-compiled manner.

[0085] Additional details regarding computer systems and networks, databases, and computer program products are also provided in, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety. Further information can be found in PCT Publication No. US2022032250 and U.S. Application No. 17832498.

[0086] According to one or more implementations, a method for generating an integrated data repository and / or analysis system containing multiple types of medical data is described herein. The architecture may include a data integration and / or analysis system. The data integration and analysis system may obtain data from several data sources and integrate the data from the data sources into the integrated data repository. For example, the data integration and analysis system may obtain data from a medical claims data repository. In various examples, the data integration and analysis system and the medical claims data repository may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system and the medical claims data repository may be created and maintained by the same entity.

[0087] The data integration and analysis system may be implemented by one or more computing devices. The one or more computing devices may include one or more server computing devices, one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, or a combination thereof. In certain implementations, at least a portion of the one or more computing devices may be implemented in a distributed computing environment. For example, at least a portion of the one or more computing devices may be implemented in a cloud computing architecture. In scenarios where the computing system used to implement the data integration and analysis system is configured in a distributed computing architecture, processing operations may be performed simultaneously by multiple virtual machines. In various examples, the data integration and analysis system may implement multithreading techniques. The implementation of a distributed computing architecture and multithreading techniques allows the data integration and analysis system to utilize fewer computing resources than those associated with a computing architecture that does not implement these techniques.

[0088] The claims data repository may store information obtained from one or more health insurance companies corresponding to claims made by subscribers of one or more health insurance companies. The claims data repository may be organized (e.g., categorized) by patient identifier. The patient identifier may be based on the patient's first name, last name, date of birth, social security number, address, employer, etc. The data stored by the claims data repository may include structured data arranged in one or more data tables. The one or more data tables storing the structured data may include several columns and several rows indicating information about claims made by subscribers of one or more health insurance companies for procedures and / or treatments received by the subscribers from healthcare providers. At least some of the columns and rows of the data tables stored by the claims data repository may include health insurance codes that may indicate diagnoses of biological conditions and treatments and / or procedures obtained by subscribers of one or more health insurance companies. In various examples, the health insurance codes may also indicate diagnostic procedures obtained by individuals related to one or more biological conditions that may exist in the individual. In one or more examples, a diagnostic procedure may provide information used in detecting the presence of a biological condition. A diagnostic procedure may also provide information used to determine the progression of a biological condition. In one or more illustrative examples, a diagnostic procedure may include one or more imaging procedures, one or more assays, one or more laboratory procedures, one or more combinations thereof, etc.

[0089] The data integration and analysis system may also obtain information from a molecular data repository. The molecular data repository may store data for several individuals related to genomic, genetic, metabolomic, transcriptomic, fragmentomic, immune receptor, methylation, epigenomic, and / or proteomic information. In one or more examples, the data integration and analysis system and the molecular data repository may be created and maintained by different entities. In one or more additional examples, the data integration and analysis system and the molecular data repository may be created and maintained by the same entity.

[0090] The genomic information and / or epigenomic information may indicate one or more mutations corresponding to the individual's genes. The individual's genetic mutations may correspond to differences between the individual's nucleic acid sequence and one or more reference genomes. The reference genome may include a known reference genome, such as hg19. In various examples, the individual's genetic mutations may correspond to differences in the individual's germline genes relative to the reference genome. In one or more additional examples, the reference genome may include the individual's germline genome. In one or more further examples, the individual's genetic mutations may include somatic mutations. The individual's genetic mutations may be associated with insertions, deletions, single-base variants, loss of heterozygosity, duplications, amplifications, translocations, fusion genes, or one or more combinations thereof.

[0091] In one or more illustrative examples, the genomic and / or epigenomic information stored by the molecular data repository may include a genomic and / or epigenomic profile of tumor cells present in an individual. In these situations, the genomic and / or epigenomic information may be derived from an analysis of genetic material, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA), from samples including, but not limited to, tissue samples or tumor biopsies, circulating tumor cells (CTCs), exosomes, or efferosomes, or from circulating nucleic acids (e.g., cell-free DNA) found in an individual's blood sample, present due to the degradation of tumor cells present in the individual. In one or more examples, the genomic and / or epigenomic information of an individual's tumor cells may correspond to one or more target regions. One or more mutations present in one or more target regions may indicate the presence of tumor cells in the individual. The genomic and / or epigenomic information stored by the molecular data repository may be generated in connection with an assay or other diagnostic test that can determine one or more mutations for one or more target regions of a reference genome.

[0092] The number of data tables may be arranged according to the data repository schema. In the illustrative example, the data repository schema includes a first data table, a second data table, a third data table, a fourth data table, and a fifth data table. While the illustrative example includes five data tables, in additional implementations, the data repository schema may include more or fewer data tables. The data repository schema may also include linkages between the data tables. Linkages between the data tables may indicate that information retrieved from one of the data tables results in additional information being stored by one or more additional data tables being retrieved. In addition, not all data tables are linked with each of the other data tables. In the illustrative example, the first data table is logically coupled to the second data table by a first linkage, and the first data table is logically coupled to the fourth data table by a second linkage. Additionally, the second data table is logically coupled to the third data table via a third link, the fourth data table is logically coupled to the fifth data table via a fourth link, and the third data table is logically coupled to the fifth data table via a fifth link.

[0093] In various examples, data tables may be added to and / or removed from the data repository schema, such that additional linkages between data tables may be added to or removed from the data repository schema. In one or more illustrative examples, the integrated data repository may store data tables according to the data repository schema for at least some of the individuals for whom the data integration system retrieves information from a combination of at least two of the prescription data repository, the molecular data repository, one or more additional data repositories, and one or more reference information data repositories. As a result, the integrated data repository may store data tables according to the data repository schema for thousands, tens of thousands, up to hundreds of thousands, or more individuals.

[0094] The data integration and analysis system may also include a data pipeline system. The data pipeline system may include algorithms, software code, scripts, macros, or any number of other computer-executable instructions that process information stored by the integrated data repository to generate additional datasets. The additional datasets may include information obtained from one or more of the data tables. The additional datasets may also include information derived from data obtained from one or more of the data tables. The components of the data pipeline system implemented to generate a first additional dataset may be different from the components of the data pipeline system used to generate a second additional dataset.

[0095] In one or more examples, the data pipeline system may generate a dataset indicating pharmaceutical treatments received by several individuals. In one or more illustrative examples, the data pipeline system may analyze information stored in at least one of the data tables to determine health insurance codes corresponding to pharmaceutical treatments received by several individuals. The data pipeline system may analyze health insurance codes corresponding to pharmaceutical treatments for a library of data indicating identified pharmaceutical treatments corresponding to one or more health insurance codes to determine the names of pharmaceutical treatments received by the individuals. In one or more additional examples, the data pipeline system may analyze information stored by an integrated data repository to determine medical procedures received by several individuals. Illustratively, the data pipeline system may analyze information stored by one of the data tables to determine treatments received by the individuals via at least one injection or intravenous route. In one or more further examples, the data pipeline system may analyze information stored by an integrated data repository to determine episodes of care for the individuals, lines of therapy received by the individuals, progression of a biological condition, or time to next treatment. In various examples, the datasets generated by the data pipeline system may be different for different biological states. For example, the data pipeline system may generate a first number of datasets for a first type of cancer, such as lung cancer, and a second number of datasets for a second type of cancer, such as colon cancer.

[0096] The data pipeline system may also determine and assign one or more confidence levels to information associated with individuals whose data is stored by the integrated data repository. Each confidence level may correspond to a different measure of accuracy for the information associated with individuals whose data is stored by the integrated data repository. The information associated with each confidence level may correspond to one or more features of the individual derived from the data stored by the integrated data repository. Confidence level values ​​for the one or more features may be generated by the data pipeline system in conjunction with generating one or more datasets from the integrated data repository. In one or more examples, the first confidence level may correspond to a first range of the accuracy measure, the second confidence level may correspond to a second range of the accuracy measure, and the third confidence level may correspond to a third range of the accuracy measure. In one or more additional examples, the second range of the accuracy measure may include values ​​that are less than the values ​​in the first range of the accuracy measure, and the third range of the accuracy measure may include values ​​that are less than the values ​​in the second range of the accuracy measure. In one or more illustrative examples, information corresponding to a first confidence level may be referred to as gold standard information, information corresponding to a second confidence level may be referred to as silver standard information, and information corresponding to a third confidence level may be referred to as bronze standard information.

[0097] The data pipeline system may determine a value for the confidence level of an individual's feature based on several factors. For example, each set of information may be used to determine the individual's feature. The data pipeline system may determine the confidence level of the individual's feature based on the amount of completeness of each set of information used to determine the feature for the individual. In a situation where one or more pieces of information are missing from a set of information associated with a first number of individuals, the confidence level for the feature may be lower than for a second number of individuals for which no information is missing from the set of information. In one or more examples, the amount of missing information may be used by the data pipeline system to determine the confidence level of the individual's feature. Illustratively, due to a greater amount of missing information used to determine the individual's feature, the confidence level for the feature may be lower than in a situation where the amount of missing information used to determine the feature is lower. Furthermore, different types of information may correspond to different confidence levels for the feature. In one or more examples, the presence of a first piece of information used to determine the individual's feature may result in a higher confidence level for the feature than the presence of a second piece of information used to determine the feature.

[0098] In one or more illustrative examples, the data pipeline system may determine the number of individuals included in a cohort who have a primary diagnosis of lung cancer (or other biological condition). The data pipeline system may determine a confidence level for each individual about being classified as having a primary diagnosis of lung cancer. The data pipeline system may use information from several rows included in a data table to determine an individual's confidence level for inclusion in the lung cancer cohort. Some rows may include health insurance codes related to a diagnosis of a biological condition and / or a treatment for a biological condition. Additionally, some rows may correspond to a diagnosis date and / or a treatment date for a biological condition. The data pipeline system may determine that the confidence level that an individual has been characterized as part of the lung cancer cohort is higher in scenarios where information is available for each of several rows or for at least a threshold number of rows than when information is available for fewer than a threshold number of rows. Furthermore, the data pipeline system may determine a confidence level for an individual included in the lung cancer cohort based on the type of information associated with one or more rows and the availability of the information. By way of example, in a situation where one or more diagnostic codes are present and one or more treatment codes are absent for a group of individuals in association with one or more time periods, the data pipeline system may determine that the confidence level for including the group of individuals in the lung cancer cohort is greater than in a situation where at least one of the diagnostic codes is absent and the treatment code used to determine whether an individual is included in the lung cancer cohort is present.

[0099] The data analysis system may accept integrated data repository requests from one or more computing devices, such as the example computing device. One or more integrated data repository requests may retrieve data from the integrated data repository. In various examples, one or more integrated data repository requests may retrieve data from one or more datasets generated by the data pipeline system. The integrated data repository requests may specify data to be retrieved from the integrated data repository and / or one or more datasets generated by the data pipeline system. In one or more additional examples, the integrated data repository requests may include one or more pre-built queries corresponding to computer-executable instructions for retrieving a specified set of data from the integrated data repository and / or one or more datasets generated by the data pipeline system.

[0100] In response to one or more integrated data repository requests, the data analysis system may analyze data retrieved from the integrated data repository or at least one of the one or more datasets generated by the data pipeline system to generate data analysis results. The data analysis results may be transmitted to one or more computing devices, such as the example computing device. While the illustrative example shows one or more integrated data repository requests and data analysis results from one computing device being transmitted to another computing device, in one or more additional implementations, the data analysis results may be received by the same computing device that sent the one or more integrated data repository requests. The data analysis results may be displayed in one or more user interfaces provided by the computing device or by the computing device.

[0101] Methods for analyzing nucleic acid sequence information are described herein. In various embodiments, the analysis method includes one or more models, each of which includes one or more of the following as separate components: survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. In various embodiments, the model includes a hierarchical model (e.g., nested model, multilevel model), a mixed model (e.g., regression such as logistic regression and Poisson regression, pooled, random effects, fixed effects, mixed effects, linear mixed effects, generalized linear mixed effects), a hazard model, an odds ratio model, and / or a replicated sample (e.g., repeated measures such as analysis of variance). In various embodiments, the model is a hierarchical random effects model. In various embodiments, the model is a hierarchical cubic spline random effects model. In various embodiments, the model is a cubic spline model. In various embodiments, the model is a generalized linear effects model. In various embodiments, the model is a linear effects model. In various embodiments, the model is a Cox proportional hazards model. In various embodiments, the analytical method includes assembly with a model. In various embodiments, the assembly includes generation of associated parameters. In one or more embodiments, the analytical method includes patient survival information and patient genetic information. As an example, the assembly with a model may include different models for different types of cancer, including subtypes, represented in the patient survival information. Each of the different models may be configured to determine a correlation between genetic factors and survival time of patients diagnosed with each type of cancer that the method is configured to evaluate. For example, genetic factors determined to have a strong correlation to cancer survival time (e.g., relatively short survival time and / or relatively long survival time) may be recommended as potential therapeutic targets.

[0102] In various embodiments, the analysis may include, as separate components, one or more of survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. For example, modeling may facilitate applying the above, such as patient survival information and patient genetic information. In various embodiments, the submodeling component may determine subsets of patient survival information and patient genetic information to generate different patient cohorts associated with different types of cancer and cancer subtypes. In various embodiments, the submodel includes a hierarchical model (e.g., nested model, multilevel model), a mixed model (e.g., regression such as logistic regression and Poisson regression, pooled, random effects, fixed effects, mixed effects, linear mixed effects, generalized linear mixed effects), a hazard model, an odds ratio model, and / or a replicated sample (e.g., repeated measures such as analysis of variance). In various embodiments, the submodel is a hierarchical random effects model. In various embodiments, the submodel is a hierarchical cubic spline random effects model. In various embodiments, the submodel is a cubic spline model. In various embodiments, the sub-model is a generalized linear effects model. In various embodiments, the sub-model is a linear effects model. In various embodiments, the sub-model is a Cox proportional hazards model. Each subset of patient survival information and patient genetic information may include information about patients diagnosed with different types of cancer and cancer subtypes. For example, the sub-modeling component can further apply the subset of patient survival information and patient genetic information to corresponding individual survival models developed for different cancer types, including subtypes. In various embodiments, information generated about the analytical method may be stored in memory (e.g., as model data). In various embodiments, information generated about the analytical method generates one or more survival models for individual subjects.

[0103] In various embodiments, the analysis of patient survival information and patient genetic information using a survival model includes a disease node determination and identification component that can identify, for each type of cancer, disease nodes included in the patient genetic information that are involved in the genetic mechanisms used to grow by each cancer type. In various embodiments, the disease node component identifies disease nodes based on observed correlations between genetic factors and cancer survival times provided in the patient survival information. For example, genetic factors that are frequently observed in association with short survival times for a particular type of cancer and less frequently observed in association with long survival times for a particular type of cancer can be identified as active genetic factors that play an active role in the genetic mechanisms of a particular type of cancer, including subtypes.

[0104] In various embodiments, disease node determination and identification includes disease-related parameters for the association between different cancer types to facilitate identifying active genetic factors associated with different cancer types. For example, highly associated cancer types may share one or more common critical underlying genetic factors. As will be readily understood by those skilled in the art, models of associated cancer types (e.g., survival models) dialectically exchange information to determine and / or identify active genetic factors across cancer types, including subtypes. In various embodiments, the disease-related parameters applied by disease node determination and identification are facilitated by modeling. In various embodiments, the generation of individual survival models may utilize one or more machine learning algorithms to facilitate the determination and / or identification of disease nodes associated with specific types of cancer, including subtypes, based on patient genetic information and disease-related parameters.

[0105] In some embodiments, the determination of a score system for a disease node is included in association with the node determination and identification of a cancer type, including subtypes. For example, the score of a disease node for a particular type of cancer, including subtypes, reflects the association of the disease node with the survival time of the particular type of cancer, including subtypes. In various embodiments, the score may be based on the frequency with which a particular genetic factor is directly or indirectly identified for patients diagnosed with a particular cancer type. In various embodiments, the analysis may include the above-mentioned survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc., and may be related to values ​​below or above a specified threshold. For example, the higher the score associated with the included disease node and cancer type, the greater the contribution of the disease node to survival time. In various embodiments, the formation of disease nodes for each type of cancer, including subtypes, and the scores determined for active genetic factors may be collated in a data structure, such as a database.

[0106] Analytical methods including effect modeling are described herein. In various embodiments, effect modeling includes random effects, fixed effects, mixed effects, linear mixed effects, and generalized linear mixed effects. In various embodiments, effects include cubic splines. In various embodiments, effect modeling includes regression. In various embodiments, effect modeling includes logistic regression and Poisson regression. In various embodiments, the model does not include covariates. In various embodiments, the model includes covariates. In various embodiments, the covariates are information from medical records (including clinical laboratory records, e.g., genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc. Examples include age, line of treatment, smoking status (yes / no), gender, and various scoring and / or staging systems utilized for specific cancer disease patients, with illustrative examples including age (years), line of anti-EGFR treatment, smoking status (yes / no), gender (female / male), and the lung cancer patient-specific Van Walraven Elixhauser Comorbidity (ELIX) score (expressed as a weighted measure across multiple common comorbidities). Those skilled in the art will readily appreciate that covariates can include any number of data elements for individuals and individuals in a population, such as data elements from medical records (laboratory records, including, e.g., genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc.

[0107] In various embodiments, the analytical method includes generating a hierarchy including at least one first-level equation. In various embodiments, the first-level equation includes a truncated cubic spline. In various embodiments, the truncated cubic spline includes longitudinal data. This includes, for example, direct or indirect measurements of ctDNA levels, allele fractions, and tumor fractions. In various embodiments, additional level equations include covariates. In various embodiments, the covariates are information about an individual or individuals in a population derived from and / or stored in medical records (laboratory records, e.g., including genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc. Examples include age, line of treatment, smoking status (yes / no), gender, and various scoring and / or staging systems utilized for patients with a particular cancer disease. In various embodiments, a velocity plot is generated. In various embodiments, the velocity plot is a derivative or equation, e.g., at least one first-level equation. In various embodiments, the analytical method includes one or more of equations (1), (2) and (3) as described in the Examples.

[0108] Described herein are analytical methods that include jointly solving different analytical components, including, as separate components, one or more of survival, modeling and submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. In various embodiments, the analytical method includes jointly solving one or more different models for different cancer types under a joint model framework. For example, the analytical method may include jointly solving one or more different survival models for different cancer types under a joint model framework. In various embodiments, the method includes determining relevant parameters. In various embodiments, the relevant parameters include, for example, the relationship between patient survival and an estimate of the current value of the patient's biomarker, or the relationship between patient survival and an estimate of the current change over time for the patient's biomarker. In various embodiments, this includes the slope, and the relationship between overall survival and a current estimate of the patient's longitudinal area under the trajectory as a proxy for the cumulative effect of the biomarker. It will be readily apparent to those skilled in the art that the relevant parameters can take many forms and can be combined. For example, one can examine the relationship between overall survival and an estimate of the current value and an estimate of the current slope of a patient's longitudinal trajectory.

[0109] In one or more examples, the data analysis system may implement at least one of one or more machine learning techniques or one or more statistical techniques to analyze data retrieved in response to one or more integrated data repository requests. In one or more examples, the data analysis system may implement one or more artificial neural networks to analyze data retrieved in response to one or more integrated data repository requests. For example, the data analysis system may implement at least one of one or more convolutional neural networks or one or more residual neural networks to analyze data retrieved from the integrated data repository in response to one or more integrated data repository requests. In at least some examples, the data analysis system may implement one or more random forest techniques, one or more support vector machines, or one or more hidden Markov models to analyze data retrieved in response to one or more integrated data repository requests. Also, one or more statistical models may be implemented to analyze data retrieved in response to one or more integrated data repository requests to identify at least one of correlations or measures of significance between individual characteristics. For example, a log-rank test may be applied to data retrieved in response to one or more integrated data repository requests. Additionally, a Cox proportional hazards model may be implemented on data retrieved in response to one or more integrated data repository requests. Furthermore, a Wilcoxon signed-rank test may be applied to data retrieved in response to one or more integrated data repository requests. In yet another example, a z-score analysis may be performed on data retrieved in response to one or more integrated data repository requests. In yet an additional example, a Kaplan-Meier analysis may be performed on data retrieved in response to one or more integrated data repository requests.In at least some examples, one or more machine learning techniques may be implemented in combination with one or more statistical techniques to analyze data retrieved in response to one or more integrated data repository requests.

[0110] In one or more illustrative examples, the data analysis system may determine the survival rate of individuals with lung cancer in response to one or more treatments. In one or more additional illustrative examples, the data analysis system may determine the survival rate of individuals with lung cancer and one or more genomic and / or epigenomic region mutations in response to one or more treatments. In various examples, the data analysis system may generate data analysis results in situations where data retrieved from at least one of the one or more datasets generated by the integrated data repository or data pipeline system meets one or more criteria. For example, the data analysis system may determine whether at least some of the data retrieved in response to one or more integrated data repository requests meets a threshold confidence level. In situations where the confidence level for at least some of the data retrieved in response to one or more integrated data repository requests is lower than the threshold confidence level, the data analysis system may refrain from generating at least some of the data analysis results. In a scenario where the confidence level for at least some of the data retrieved in response to one or more integrated data repository requests is at least a threshold confidence level, the data analysis system may generate at least some of the data analysis results. In various examples, the threshold confidence level may be related to the type of data analysis results generated by the data analysis system.

[0111] In one or more illustrative examples, the data analysis system may accept an integrated data repository request to generate data analysis results indicative of survival rates for one or more individuals. In these cases, the data analysis system may determine whether the data stored by the integrated data repository and / or by one or more datasets generated by the data pipeline system meets a threshold confidence level, e.g., a gold standard confidence level. In one or more additional examples, the data analysis system may accept an integrated data repository request to generate data analysis results indicative of treatments received by one or more individuals. In these implementations, the data analysis system may determine whether the data stored by the integrated data repository and / or by one or more datasets generated by the data pipeline system meets a lower threshold confidence level, e.g., a bronze standard confidence level.

[0112] In one or more additional illustrative examples, the data analysis system may receive an integrated data repository request to determine individuals who have one or more genomic and / or epigenomic variations and who have received one or more treatments for a biological condition. Continuing with this example, the data analysis system may determine the survival rate of individuals with one or more genomic and / or epigenomic variations with respect to one or more treatments received by the individuals. The data analysis system may then identify, based on the survival rate of the individuals and the effectiveness of treatments for the individuals, with respect to genomic and / or epigenomic variations that may be present in the individuals. In this manner, individual health outcomes may be improved by identifying potential treatments that may be more effective for a population of individuals with one or more genomic and / or epigenomic variations than current treatments provided to individuals.

[0113] The data pipeline system may include first data processing instructions, second data processing instructions, and up to Nth data processing instructions. The data processing instructions may be executable by one or more processing units to perform several operations to generate each data set using information retrieved from the integrated data repository. In one or more illustrative examples, the data processing instructions may include at least one of software code, scripts, API calls, macros, etc. The first data processing instructions may be executable to generate the first data set. Additionally, the second data processing instructions may be executable to generate the second data set. Furthermore, the Nth data processing instructions may be executable to generate the Nth data set. In various examples, after the data integration and analysis system generates the integrated data repository, the data pipeline system may execute the data processing instructions to generate the data sets. In one or more examples, the data sets may be stored by the integrated data repository or by an additional data repository accessible to the data integration and analysis system. At least some of the data processing instructions may parse health insurance codes to generate at least some of the data sets. Additionally, at least some of the data processing instructions may analyze the genomics data to generate at least some of the data sets.

[0114] In one or more examples, the first data processing instructions may be executable to retrieve data from one or more first data tables stored by the integrated data repository. The first data processing instructions may also be executable to retrieve data from one or more specific rows of the one or more first data tables. In various examples, the first data processing instructions may be executable to identify individuals having health insurance codes stored in one or more row and column combinations corresponding to one or more diagnostic codes. The first data processing instructions may then be executable to analyze the one or more diagnostic codes to determine the biological conditions with which the individuals have been diagnosed. In one or more illustrative examples, the first data processing instructions may be executable to analyze one or more diagnostic codes against a library of diagnostic codes indicating one or more biological conditions corresponding to each diagnostic code. The library of diagnostic codes may include hundreds, or up to thousands, of diagnostic codes. The first data processing instructions may also be executable to determine individuals diagnosed with a certain biological condition by analyzing timing information for the individuals, such as date of treatment, date of diagnosis, date of death, or one or more combinations thereof.

[0115] The second data processing instructions may be executable to retrieve data from one or more second data tables stored by the integrated data repository. The second data processing instructions may also be executable to retrieve data from one or more specific rows of the one or more second data tables. In various examples, the second data processing instructions may be executable to identify individuals having health insurance codes stored in one or more row and column combinations corresponding to one or more procedure codes. The one or more procedure codes may correspond to treatments obtained from a pharmacy. In one or more additional examples, the one or more procedure codes may correspond to medical procedures, such as treatments received by injection or intravenously. The second data processing instructions may be executable to determine one or more treatments corresponding to each health insurance code included in the one or more second data tables by analyzing the health insurance codes with respect to a predetermined set of information. The predetermined set of information may include a data library indicating one or more treatments corresponding to one of hundreds, or up to thousands, of health insurance codes. The second data processing instructions may generate a second data set indicating each treatment received by a group of individuals. In one or more illustrative examples, the group of individuals may correspond to the individuals included in the first dataset. The second dataset may be arranged in columns and rows, with one or more columns corresponding to a single individual and one or more rows indicating the treatment received by each individual.

[0116] The Nth processing instructions (where N may be any positive integer) may be executable to generate the Nth dataset by combining information from several previously generated datasets, e.g., the first dataset and the second dataset. Additionally, the Nth processing instructions may be executable to generate the Nth dataset to retrieve additional information from one or more additional rows of the integrated data repository and combine the additional information from the integrated data repository with information obtained from the first dataset and the second dataset. For example, the Nth processing instructions may be executable to identify individuals included in the first dataset who are diagnosed with a certain biological condition and analyze specific rows of one or more additional data tables of the integrated data repository to determine dates of treatments indicated in the second dataset that correspond to individuals included in the first dataset. In one or more further examples, the Nth processing instructions may be executable to analyze rows of one or more additional data tables of the integrated data repository to determine dosages of treatments indicated in the second dataset that will be received by individuals included in the first dataset. In this manner, the Nth processing instructions may be executable to generate an episode of care dataset based on information contained in the cohort dataset and the treatment dataset.

[0117] In one or more illustrative examples, in response to receiving the integrated data repository request, the data analysis system may determine one or more datasets that correspond to the nature of the query related to the integrated data repository request. For example, the data analysis system may determine that information included in a first dataset and a second dataset is applicable to responding to the integrated data repository request. In these scenarios, the data analysis system may analyze at least a portion of the data included in the first dataset and the second dataset to generate data analysis results. In one or more additional examples, the data analysis system may determine different datasets to respond to different queries included in the integrated data repository request to generate data analysis results.

[0118] The use of a specific set of data processing instructions to generate each dataset may reduce the number of inputs from users of the data integration and analysis system and may also reduce the computational burden, such as the amount of processing resources and memory, required to process integrated data repository requests. For example, without the specific architecture of a data pipeline system, the data required to respond to an integrated data repository request is assembled from the data repository each time an integrated data repository request is received. In contrast, by implementing a data pipeline system to execute data processing instructions and generate datasets, the data required to respond to various integrated data repository requests is already assembled and can be accessed by the data analysis system to respond to the integrated data repository request. Therefore, the computational resources required to respond to integrated data repository requests by implementing a data pipeline system to generate datasets are fewer than those required by a typical system that performs information parsing and collection processes for each integrated data repository request. Furthermore, in situations where a data pipeline system is not implemented, due to the imprecision of ad hoc collection of data to respond to an integrated data repository request in a typical system, or because the data analysis system is called multiple times in a typical system to perform analysis of information that could be performed using a single integrated data repository request if a data pipeline system were implemented, users of the data integration and analysis system may be required to submit multiple integrated data repository requests to analyze the information that the user intended to analyze.

[0119] In operation, the data integration and analysis system may integrate genomics data and medical claims data for individuals common to both the molecular data repository and the medical claims data repository. The data integration and analysis system may determine individuals common to both the molecular data repository and the medical claims data repository by determining the genomics data and medical claims data corresponding to common tokens. The data integration and analysis system may determine that a first token associated with a portion of the genomics data corresponds to a second token associated with a portion of the medical claims data by determining a measure of similarity between the first token and the second token. In a scenario in which the first token has at least a threshold amount of similarity with the second token, the data integration and analysis system may store the corresponding portion of the genomics data and the corresponding portion of the medical claims data with respect to the individual's identifier in an integrated data repository, such as an integrated data repository.

[0120] An implementation of the architecture may implement a cryptographic protocol that allows de-identified information from disparate data repositories to be integrated into a single data repository. In this way, the security of the data stored by the integrated data repository is increased. Additionally, the cryptographic protocol implemented by the architecture may enable more efficient retrieval and accurate analysis of information stored by the integrated data repository than in situations where the architecture's cryptographic protocol is not utilized. For example, by generating a token file containing a first token using cryptographic techniques based on a particular set of information stored by a molecular data repository, and utilizing a second token generated using the same or similar cryptographic techniques for a similar or identical set of information stored by a medical receipt data repository, the data integration and analysis system may match information stored by disparate data repositories that correspond to the same individual. Failure to implement the architecture's cryptographic protocols increases the probability of inaccurately attributing information from one data repository to one or more individuals, which reduces the accuracy of the results provided by the data integration and analysis system in response to an integrated data repository request sent to the data integration and analysis system.

[0121] According to one or more implementations, a framework for generating a dataset by a data pipeline system based on data stored by an integrated data repository is described herein. The integrated data repository may store medical claims data and genomics data for a group of individuals. For example, the integrated data repository may store information obtained from medical claims records for the group of individuals. For each individual included in the group of individuals, the integrated data repository may store information obtained from multiple medical claims records. In various examples, the information stored by the integrated data repository may include and / or be derived from thousands, tens of thousands, hundreds of thousands, or even millions of medical claims records for several individuals. Additionally, each medical claims record may include multiple rows. As a result, the integrated data repository may be generated through analysis of millions of rows of medical claims data.

[0122] Furthermore, while claims data may be organized according to a structured data format, it is typically arranged for viewing by health insurance organizations, patients, and healthcare providers to display accounting and insurance code information related to services provided to individuals by healthcare providers. Therefore, it is not easy to analyze claims data to obtain insights that may be available related to characteristics of individuals in whom a biological condition exists and that may assist in treating the individual for the biological condition. An integrated data repository may be generated and organized by analyzing and modifying raw claims data to enable the data stored by the integrated data repository to be further analyzed to determine trends, characteristics, traits, and / or insights about individuals in whom one or more biological conditions may exist. For example, health insurance codes may be stored in the integrated data repository in a manner such that at least one of a medical procedure, a biological condition, a treatment, a dosage, a drug manufacturer, a drug distributor, or a diagnosis can be determined for a given individual based on the claims data for the individual. In various examples, the data integration and analysis system may generate and implement one or more tables showing correlations between medical claims data and various treatments, symptoms, or biological conditions corresponding to the medical claims data. Additionally, the integrated data repository may be generated using genomics data records for a group of individuals. In various examples, large amounts of medical claims data may be matched with genomics data for a group of individuals to generate the integrated data repository.

[0123] By integrating genomics data records for a group of individuals with medical insurance claims records, a data integration and analysis system can determine correlations between the presence of one or more biomarkers present in the genomics data records and other characteristics of the individuals indicated by the medical insurance claims data records that existing systems typically cannot determine. For example, the data integration and analysis system can determine one or more genomic and / or epigenomic characteristics of the individuals that correspond to the treatment received by the individual, the timing of the treatment, the dosage of the treatment, the individual's diagnosis, smoking status, the presence of one or more biological conditions, the presence of one or more symptoms of the biological conditions, one or more combinations thereof, etc. Based on the correlations determined by the data integration and analysis system using the integrated data repository, cohorts of individuals that may benefit from one or more treatments that were not identified in existing systems can be identified. In one or more examples, the processes and techniques implemented to integrate prescription records and genomics claims records to generate an integrated data repository are complex, and efficiency-enhancing techniques, systems, and processes may be implemented to minimize the amount of computational resources used to generate the integrated data repository.

[0124] In one or more illustrative examples, the data pipeline system may access information stored by the integrated data repository to generate a dataset including several additional data records containing information related to at least a portion of a group of individuals. In an illustrative example, the additional data records include information indicating whether the individual is included in a cohort of individuals in which lung cancer is present. The data pipeline system may execute multiple different sets of data processing instructions to determine the cohort of individuals in which lung cancer is present. In various examples, the additional data records may indicate information used to determine the individual's status with respect to lung cancer, such as one or more transaction insurance identifiers, one or more International Classification of Diseases (ICD) codes, and one or more health insurance transaction dates. In addition to including a row indicating whether the individual is included in a lung cancer cohort, the additional data record may include a row indicating a confidence level of the individual's status with respect to the presence of lung cancer.

[0125] A computing architecture for aggregating medical record data into an integrated data repository is described herein. In various examples, at least a portion of the operations of the computing architecture may be performed by a data integration and analysis system. In one or more examples, at least a portion of the operations of the computing architecture may be performed by one or more additional computing systems controlled, maintained, or implemented by a service provider that also controls, maintains, or implements the data integration and analysis system. In one or more additional examples, at least a portion of the operations of the computing architecture may be performed by several servers in a distributed computing environment.

[0126] The computing architecture may include a medical record data repository. The medical record data repository may store medical record data from several individuals. The medical record data may include imaging information, laboratory test results, diagnostic test information, clinical findings, dental hygiene information, healthcare practitioner notes, medical questionnaires, diagnostic requests, medical procedure orders, medical charts, one or more combinations thereof, etc. In various examples, for a given individual, the medical record data repository may store information related to the individual obtained from one or more healthcare practitioners.

[0127] The computing architecture may perform operations including retrieving a data package from a medical record data repository. In one or more examples, the data package may be retrieved in response to one or more requests sent to the medical record data repository for medical records corresponding to one or more individuals. In one or more additional examples, the data package may be retrieved by the computing architecture using one or more application programming interface (API) calls. In one or more illustrative examples, a first data package, a second data package, and up to an Nth data package may be retrieved using the computing architecture. Each data package may correspond to a medical record for each individual. For example, the first data package may include a medical record for a first individual, the second data package may include a medical record for a second individual, and the Nth data package may include a medical record for a third individual.

[0128] An individual data package may include several components. In one or more examples, an individual data package may include individual components corresponding to medical records from different healthcare providers. In one or more additional examples, an individual data package may include individual components corresponding to different portions of medical records corresponding to one or more healthcare providers. In an illustrative example, a second data package may include a first component, a second component, up to an Nth component. In one or more illustrative examples, the first component may include a first portion of an individual's medical record, the second component may include a second portion of the individual's medical record, and the Nth component may include a third portion of the individual's medical record. In various examples, the first component may correspond to a medical record of a first healthcare provider for the individual, the second component may correspond to a medical record of a second healthcare provider for the individual, and the third component may correspond to a medical record of a third healthcare provider for the individual. In one or more additional illustrative examples, the first component may include a first section of the individual's medical record, such as one or more forms related to a diagnostic test or procedure, and the second component may include a second section of the individual's medical record, such as a pathology report for the individual.

[0129] In operation, the computing architecture may preprocess individual data packages to identify a corpus of information to be analyzed. In one or more examples, preprocessing the data packages retrieved from the medical record data repository may include transforming the data included in the data packages. For example, preprocessing the data packages may include converting at least a portion of the data retrieved from the medical record data repository into machine-coded information. Illustratively, preprocessing the data packages may include performing one or more optical character recognition (OCR) operations on at least a portion of the data packages retrieved from the medical record data repository. By converting at least a portion of the data packages retrieved from the medical record data repository into machine-coded information, the data packages may be subjected to several operations, such as one or more parsing operations to identify one or more characters or strings of characters, or one or more editing operations, that cannot be performed on at least a portion of the data packages retrieved from the medical record data repository.

[0130] In one or more examples, preprocessing of the individual data packages may include determining information contained in the individual data packages that should be excluded from further analysis by the computing architecture. In various examples, one or more components of the individual data packages may be excluded from the corpus of information to be analyzed. For example, for a second data package, the computing architecture may determine that a first component should be excluded from further analysis by the computing architecture. In one or more examples, the computing architecture may analyze the components to identify at least one of the components for one or more keywords to remove from further analysis by the computing architecture. In one or more illustrative examples, the computing architecture may parse the components to identify one or more keywords, and in response to identifying the one or more keywords in the component, the computing architecture may determine to exclude each component from further analysis by the computing architecture. For example, the computing architecture may determine that a first component of the second data package is a test bill for one or more diagnostic procedures or tests. In these scenarios, the computing architecture may determine that the first component should be excluded from further analysis by the computing architecture. Additionally, the computing architecture may determine that at least one of the second components corresponds to one or more pathology reports for the individual based on one or more keywords included in the second component or at least one of the Nth component. In these cases, the computing architecture may determine that at least a portion of the second component and / or at least a portion of the Nth component should be included in the corpus of information for further analysis by the computing architecture.

[0131] Additionally, a subset of components of individual data packages retrieved from the medical record data repository may be included in the corpus of information. In various examples, one or more additional operations may be performed to narrow the corpus of information. For example, one or more queries may be applied to the subset of information retrieved from the medical record data repository. The one or more queries may extract information from the one or more data packages that satisfies the one or more queries. In at least some examples, the one or more queries may be a group of queries applied to individual components of the data packages. In one or more illustrative examples, the group of queries may determine information to be included in the corpus of information and additional information to be excluded from the corpus of information. In one or more additional examples, one or more sections of at least one component of the data package may be excluded from the corpus of information.

[0132] In one or more additional illustrative examples, after determining that a first component should be excluded from further analysis by the computing architecture, the computing architecture may then implement one or more queries on at least one of the second component or the Nth component. In these scenarios, the one or more queries may determine that a section of the second component, such as a section indicating family history of one or more biological conditions, should be excluded from the corpus of information. In various examples, the one or more queries may be directed to identifying keywords and / or combinations of keywords included in at least one of the second component or the Nth component. In these cases, the computing architecture may exclude from the corpus of information one or more portions of the individual components of the data package that include one or more keywords or combinations of keywords. In one or more additional examples, the computing architecture may exclude from the corpus of information words, characters, and / or symbols following one or more keywords included in one or more portions of the individual components of the data package.

[0133] Further, in operation, the computing architecture may analyze the corpus of information to determine characteristics of individuals. In one or more examples, the computing architecture may analyze the corpus of information to determine individuals having one or more phenotypes. In various examples, the computing architecture may analyze the corpus of information to determine one or more biomarkers indicative of a biological state. For example, the computing architecture may analyze the corpus of information to determine individuals having one or more genetic characteristics. The one or more genetic characteristics may include at least one of one or more variants in a genomic region and / or an epigenomic region corresponding to a biological state. In one or more illustrative examples, the one or more genetic characteristics may correspond to one or more variants in a genomic region and / or an epigenomic region corresponding to a type of cancer. In one or more additional illustrative examples, the one or more biomarkers may correspond to a level of an analyte outside a particular range. Illustratively, the computing architecture may analyze the corpus of information to determine individuals having levels of one or more proteins and / or levels of one or more small molecules indicative of a biological state present. In these scenarios, the computing architecture may analyze laboratory test results to determine the level of an analyte in an individual. In one or more additional examples, the computing architecture may analyze a corpus of information to determine the presence of one or more symptoms indicative of a biological condition in an individual. In one or more further examples, the computing architecture may analyze imaging information included in the corpus of information to determine the presence of one or more biomarkers in an individual.

[0134] In one or more examples, the computing architecture may implement one or more machine learning techniques to analyze the corpus of information. For example, the computing architecture may implement one or more artificial neural networks, such as at least one of one or more convolutional neural networks or one or more residual neural networks, to analyze the corpus of information. The computing architecture may also implement at least one of one or more random forest techniques, one or more hidden Markov models, or one or more support vector machines to analyze the corpus of information.

[0135] In at least some implementations, the computing architecture may analyze the corpus of information by performing one or more queries on the corpus of information. The one or more queries may correspond to one or more keywords and / or keyword combinations. The one or more keywords and / or keyword combinations may correspond to at least one of characters or symbols corresponding to one or more biological conditions. For example, the keywords may correspond to characters associated with mutations in genomic and / or epigenomic regions, such as HER2. In one or more additional illustrative examples, one or more criteria may be associated with a keyword combination. For example, criteria corresponding to a keyword combination may include several words that are within a certain distance of each other in a portion of the corpus of information about an individual, such as the words fatigue, blood pressure, and swelling, that are within a certain character of each other. In these cases, the computing architecture may parse the corpus of information for one or more keywords and / or keyword combinations. In various examples, in response to determining that one or more keywords and / or combinations of keywords are present consistent with one or more criteria, the computing architecture may determine that a biological condition is present for a given individual.

[0136] In one or more additional examples, the one or more queries may be image-based, and the computing architecture may analyze images included in the corpus of information for a template image. The template image may be generated based on analyzing several images in which a biological state is present and aggregating the several images into a template image. In these scenarios, the computing architecture may analyze the images included in the corpus of information for one or more template images to determine a similarity measure between the images included in the corpus of information and the template image. In situations where the similarity measure for the individual is at least a threshold, the computing architecture may determine that a characteristic of the biological state is present in the individual.

[0137] After determining the individuals having the one or more characteristics, the computing architecture may, in operation, generate a data structure that stores data about the individuals having the one or more characteristics. In one or more examples, the computing architecture may generate a data table that indicates individuals having individual characteristics and / or individuals having groups of characteristics. For example, the computing architecture may generate a first data table and a second data table. The first data table may indicate individuals having one or more first characteristics, and the second data table may indicate individuals having one or more second characteristics. In one or more illustrative examples, the first data table may indicate individuals having one or more first biomarkers for a biological state, and the second data table may indicate individuals having one or more second biomarkers for the biological state. The one or more first biomarkers may correspond to one or more first genomic and / or epigenomic variants associated with the biological state, and the one or more second biomarkers may correspond to one or more second genomic and / or epigenomic variants associated with the biological state.

[0138] One or more data structures may be generated from the corpus of information that stores identifiers for portions of the subset of additional groups of individuals and that stores an indication that the portions of the subset of additional groups of individuals correspond to one or more biomarkers. The one or more data structures may be stored by an intermediate data repository. After performing one or more de-identification operations on the identifiers for the portions of the subset of additional groups of individuals, the integrated data repository may be modified to store at least some of the additional information for the medical records of the portions of the subset of additional groups of individuals in association with some identifiers. After de-identification of the information stored by the one or more data structures, the information stored by the integrated data repository may be added to the integrated data repository. In at least some examples, de-identified medical record information may be added to the integrated data repository in addition to or instead of medical claims data. In various examples, the one or more data structures that store de-identified medical record information for biomarker data may have one or more logical connections with other data structures stored in the integrated data repository.By way of example, the one or more data structures storing de-identified medical record information about biomarker data may have one or more logical connections to at least one of: a first data table that may store information corresponding to genomics data, mutations in genomic and / or epigenomic regions, the type of mutation, copy numbers of genomic and / or epigenomic regions, coverage data indicating the number of nucleic acid molecules identified in the sample with one or more mutations, test dates, and the panel used to generate the patient information; a second data table that stores data related to one or more visits by an individual to one or more healthcare provider institutions; a third data table that stores information corresponding to each service provided to an individual for one or more visits to one or more healthcare provider institutions indicated by the second data table; a fourth data table that stores personal information of a group of individuals; a fifth data table that stores information related to health insurance companies or government agencies that paid for services provided to the group of individuals; a sixth data table that stores information corresponding to health insurance coverage information for a group of individuals, such as the type of health insurance plan associated with the group of individuals; or a seventh data table that stores information related to medical treatments obtained by a group of individuals.

[0139] Described herein is a machine in the form of a computer system on which a set of instructions can be executed to cause the machine to perform any one or more of the methodologies discussed herein, according to an example implementation. For example, a machine in the form of an example computer system on which instructions (e.g., software, programs, applications, applets, apps, or other executable code) can be executed to cause the machine to perform any one or more of the methodologies discussed herein. For example, the instructions can cause the machine to implement the architecture and framework previously described and perform the methods previously described. For example, one or more machine-executable components embodied within one or more machines (e.g., embodied in one or more computer-readable storage media associated with one or more machines). Such components, when executed by one or more machines (e.g., processors, computers, computing devices, virtual machines, etc.), can cause the one or more machines to perform the operations described through the instructions. For example, a machine can include a computing device with an analysis component. The analysis may include survival, modeling, sub-modeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. In various embodiments, the analysis component is embodied in a machine-executable component within a system that includes various electronic data sources and data structures containing information usable by the analysis component. Non-limiting examples include data sources and structures, such as survival information, genetic information, model data, sub-models, disease node determination and identification, disease association information, disease subtyping, recurrence, metastasis, time to next treatment, etc.

[0140] The computing device may include or be operably coupled to at least one memory and at least one processor. The at least one memory stores executable instructions for performance of the analysis when executed by the at least one processor. In some embodiments, the memory may also store various data sources and / or structures of the system. In other embodiments, various data sources and structures of the system may be stored in other memory accessible to the computing device (e.g., in a remote device or system).

[0141] The instructions transform a general, unprogrammed machine, e.g., a computing device, into a specific machine programmed to perform the functions described and illustrated in the described manner. In alternative implementations, the machine may operate as a standalone device or may be coupled (e.g., networked) to other machines. In a network deployment, the machine may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a mobile phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing instructions, sequentially or otherwise, that specify actions to be taken by the machine. Those skilled in the art will appreciate that a machine includes a collection of machines that individually or collectively execute instructions to perform any one or more of the methodologies discussed herein.

[0142] Examples of computing devices may include logic, one or more components, circuits (e.g., modules), or mechanisms. A circuit is a tangible entity configured to perform a certain operation. In one example, a circuit may be arranged in a particular manner (e.g., internally or with respect to an external entity such as another circuit). In one example, one or more computer systems (e.g., standalone, client, or server computer systems) or one or more hardware processors (processors) may be configured by software (e.g., instructions, application portions, or applications) as circuits that operate to perform certain operations described herein. In one example, the software may reside (1) in a non-transitory machine-readable medium or (2) in a transmission signal. In one example, the software, when executed by the hardware underlying the circuit, causes the circuit to perform a certain operation.

[0143] Various operations of example methods described herein may be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented circuitry that operates to perform one or more operations or functions. In one example, circuitry referred to herein may include processor-implemented circuitry.

[0144] Similarly, the methods described herein may be at least partially implemented by a processor. For example, at least some or all of the operations of a method may be performed by one or more processors or processor-implemented circuitry. Certain performance of operations may reside not only within a single machine, but may also be deployed across several machines and distributed among one or more processors. In one example, the processor(s) may be located at a single location (e.g., in a home environment, an office environment, or as a server farm), while in another example, the processor(s) may be distributed across several locations.

[0145] One or more processors may also operate to support the performance of related operations in a "cloud computing" environment or as "software as a service."

[0146] For example, at least some of the operations may be performed by a group of computers (as examples of machines that include processors), and these operations are accessible over a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application program interfaces (APIs)).

[0147] Example implementations (e.g., devices, systems, or methods) can be implemented in digital electronic circuitry, in computer hardware, in firmware, in software, or in any combination of these. Example implementations can also be implemented using a computer program product (e.g., a computer program tangibly embodied in an information carrier or machine-readable medium for execution by, or to control the operation of, a data processing apparatus such as a programmable processor, a computer, or multiple computers).

[0148] A computer program may be written in any form of programming language, including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a software module, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communications network.

[0149] In one example, operations may be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Example method operations may also be performed by, and example apparatus may be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)).

[0150] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. It is understood that in implementations deploying a programmable computing system, both the hardware and software architectures need consideration. In particular, it is understood that the choice of whether to implement certain functionality in permanently configured hardware (e.g., ASICs), temporarily configured hardware (e.g., a combination of software and programmable processors), or a combination of permanently and temporarily configured hardware may be a design choice. Hardware architectures (e.g., computing devices) and software architectures that may be deployed in example implementations are shown below.

[0151] In one example, the computing device may operate as a stand-alone device, or the computing device may be connected (eg, networked) to other machines.

[0152] In a network deployment, a computing device can operate as either a server machine or a client machine in a server-client network environment. In one example, a computing device can act as a peer machine in a peer-to-peer (or other distributed) network environment. A computing device can be a personal computer (PC), a tablet PC, a set-top box (STB), a mobile phone, a web appliance, a network router, switch, or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken (e.g., performed) by the computing device. Furthermore, while only a single computing device is illustrated, the term "computing device" should also be interpreted to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one of the tasks.

[0153] The computing device may further include a storage device (e.g., a drive unit), a signal generating device (e.g., a speaker), a network interface device, and one or more sensors, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or another sensor. The storage device may include a machine-readable medium on which is stored one or more data structures or sets of instructions (e.g., software) that embody or are utilized by any one or more of the methodologies or functions described herein. Also, the instructions may reside, completely or at least partially, within main memory, static memory, or within the processor during execution thereof by the computing device. In one example, one or any combination of the processor, main memory, static memory, or storage device may constitute a machine-readable medium.

[0154] Although the machine-readable medium is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store one or more instructions. The term "machine-readable medium" may also be interpreted to include any tangible medium capable of storing, encoding, or transmitting instructions for execution by a machine and causing a machine to perform any one or more of the methodologies of the present disclosure or storing, encoding, or transmitting data structures used by or associated with such instructions.

[0155] As used herein, a component may refer to a device, physical entity, or logic with boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that provide partitioning or modularization of specific processing or control functions. A component may be combined with other components through their interfaces to perform machine processes. A component may be a packaged functional hardware unit designed for use with other components, and may be part of a program that performs a specific function, usually associated with that function. A component may constitute either a software component (e.g., code embodied in a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit capable of performing a specific operation and may be configured or arranged in a specific physical manner. In various example implementations, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems), or one or more hardware components of a computer system (e.g., a processor or group of processors), may be configured by software (e.g., an application or application portion) as hardware components that operate to perform certain operations described herein. disease

[0156] This method can be used to diagnose the presence of a condition in a subject, characterize the condition, monitor the response to treatment for the condition, and indicate the prognostic risk of developing the condition or the next stage of the condition.The present disclosure can also be useful in determining the effectiveness of a particular treatment option.If treatment is successful, the amount of nucleic acid, such as cell-free nucleic acid, detected in the subject's blood may increase, as diseased and dysfunctional cells die and release DNA, or otherwise show chronic and acute signs of inflammation.In other examples, this may not occur.In another example, perhaps a particular treatment option can be correlated with the genetic profile of disease type and subtype over time.This correlation can be useful in selecting a treatment.

[0157] In some embodiments, the methods and systems disclosed herein can be used to identify customized or targeted therapies for treating a given disease or condition in a patient based on the classification of nucleic acid variants as being of somatic or germline origin. Typically, the disease under consideration is a type of cancer.

[0158] Furthermore, the disclosed methods can be used to characterize heterogeneity in abnormal conditions in a subject. Such methods may include, for example, generating a genomic and epigenomic profile of extracellular polynucleotides derived from a subject, where the genetic profile includes multiple data that can characterize dysfunctions and abnormalities (e.g., hypertrophy) associated with cardiac muscle and valve tissue. Reduced blood flow and oxygen supply to the heart are often secondary symptoms of weakening and / or deterioration of the current blood and supply system caused by physical and biochemical stress. Examples of cardiovascular diseases directly affected by these types of stress include atherosclerosis, coronary artery disease, peripheral vascular disease, and peripheral arterial disease, along with various cardiac and cardiac arrhythmias that may represent other forms of disease and dysfunction. The disclosed methods can be used to generate a profile, fingerprint, or data set that summarizes genetic information derived from different cells in a heterogeneous disease. This data set can include copy number variation, epigenetic mutations, and mutation analysis, alone or in combination.

[0159] This method can be used to diagnose, predict prognosis, monitor or observe cancer or other diseases.In some embodiments, the method herein does not involve diagnosing, predicting prognosis or monitoring fetus, and therefore is not intended for non-invasive prenatal testing.In other embodiments, these methodologies can be used in pregnant subjects to diagnose, predict prognosis, monitor or observe cancer or other diseases in the subject in utero, where DNA and other polynucleotides may co-circulate with maternal molecules.

[0160] Non-limiting examples of other gene-based diseases, disorders, or conditions that may be optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth disease (CMT), cri-a-cat syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down's syndrome, Duane's syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, and fragile X. Syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland syndrome, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatocardiofacial syndrome, WAGR syndrome, Wilson's disease, etc. Treatment and Related Administration

[0161] In certain embodiments, the methods disclosed herein relate to identifying and administering customized treatments to patients whose status has been confirmed as having a nucleic acid variant of somatic or germline origin. In some embodiments, essentially any cancer treatment (e.g., surgical treatment, radiation therapy, chemotherapy, etc.) can be included as part of these methods. Typically, customized treatments include at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of improving the immune response to a given cancer type. In certain embodiments, immunotherapy refers to a method of improving T-cell responses against tumors or cancer.

[0162] In certain embodiments, the status of the nucleic acid variants from the sample from the subject that are of somatic origin or germline origin can be compared with the database of comparator results from a reference population to identify the customized or targeted treatment for the subject.Typically, the reference population comprises patients with the same cancer or disease type as the test subject, and / or patients who are receiving or have received the same treatment as the test subject.Customized or targeted treatment(s) can be identified if the nucleic acid variants and comparator results meet certain classification criteria (for example, are substantially or approximately compatible).

[0163] In certain embodiments, the customized therapy described herein is typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapy (e.g., immunotherapeutic agents, etc.) may also be administered by methods such as oral, sublingual, rectal, vaginal, intraurethral, ​​topical, intraocular, intranasal, and / or intraauricular administration, and may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, ointments, etc.

[0164] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided herein. While the present invention has been described with reference to the above-referenced specifications, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions shown herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be used in practicing the present invention. Therefore, it is intended that the present disclosure also cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are intended to be covered thereby.

[0165] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent to those skilled in the art upon reading this disclosure that various changes in form and detail can be made without departing from the true scope of the disclosure and can be practiced within the purview of the appended claims. For example, all methods, systems, computer-readable media and / or component features, steps, elements or other aspects thereof can be used in various combinations. Biomarkers

[0166] The present disclosure provides methods of using biomarkers for the diagnosis, prognosis, and treatment selection of subjects suffering from diseases such as heart failure, cardiovascular disease, cancer, etc. A biomarker can be any gene or gene variant whose presence, mutation, deletion, substitution, copy number, or translation (i.e., into protein) indicates a disease state. Biomarkers of the present disclosure can include the presence, mutation, deletion, substitution, copy number, or translation in any one or more of EGFR, KRAS, MET, BRAF, MYC, NRAS, ERBB2, ALK, Notch, PIK3CA, APC, and SMO.

[0167] Biomarkers are genetic variants. Biomarkers can be determined using any of several resources or methods. Biomarkers can be previously discovered or newly discovered using experimental or epidemiological techniques. If a biomarker is highly correlated with a disease, the detection of the biomarker can indicate the disease. If a biomarker in a region or gene is present at a frequency greater than the frequency for a given background population or data set, the detection of the biomarker can indicate cancer.

[0168] Publicly available resources, such as scientific literature and databases, can provide detailed descriptions of genetic variants. Scientific literature can describe experiments or genome-wide association studies (GWAS) that link one or more genetic variants. Databases can aggregate information collected from sources such as scientific literature to provide a more comprehensive resource for determining one or more biomarkers. Non-limiting examples of databases include FANTOM, GTex, GEO, Body Atlas, INSiGHT, OMIM (Online Mendelian Inheritance in Man, omim.org), cBioPortal (cbioportal.org), CIViC (Clinical Interpretations of Variants in Cancer, civic.genome.wustl.edu), DOCM (Database of Curated Mutations, docm.genome.wustl.edu) and ICGC Data Portal (dcc.icgc.org). In a further example, the COSMIC (Catalogue of Somatic Mutations in Cancer) database allows for searching biomarkers by cancer, gene, or mutation type. Biomarkers may also be determined de novo by conducting experiments such as case-control or association (e.g., genome-wide association studies) studies.

[0169] One or more biomarkers may be detected in a sequencing panel. The biomarkers may be one or more genetic variants. The biomarkers may be selected from single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (e.g., indels), gene fusions, and inversions. The biomarkers may affect protein levels. The biomarkers may be present in promoters or enhancers and may alter gene transcription. The biomarkers may affect gene transcription and / or translation efficiency. The biomarkers may affect the stability of transcribed mRNA. The biomarkers may result in changes to the amino acid sequence of a translated protein. The biomarkers may affect splicing, change the amino acids encoded by specific codons, result in frameshifts, or result in premature stop codons. The biomarkers may result in conservative amino acid substitutions. One or more biomarkers may result in conservative amino acid substitutions. One or more biomarkers may result in non-conservative amino acid substitutions.

[0170] The frequency of a biomarker can be as low as 0.001%. The frequency of a biomarker can be as low as 0.005%. The frequency of a biomarker can be as low as 0.01%. The frequency of a biomarker can be as low as 0.02%. The frequency of a biomarker can be as low as 0.03%. The frequency of a biomarker can be as low as 0.05%. The frequency of a biomarker can be as low as 0.1%. The frequency of a biomarker can be as low as 1%.

[0171] A single biomarker may be absent in more than 50% of subjects with cancer. A single biomarker may be absent in more than 40% of subjects with cancer. A single biomarker may be absent in more than 30% of subjects with cancer. A single biomarker may be absent in more than 20% of subjects with cancer. A single biomarker may be absent in more than 10% of subjects with cancer. A single biomarker may be absent in more than 5% of subjects with cancer. A single biomarker may be present in 0.001% to 50% of subjects with cancer. A single biomarker may be present in 0.01% to 50% of subjects with cancer. A single biomarker may be present in 0.01% to 30% of subjects with cancer. A single biomarker may be present in 0.01% to 20% of subjects with cancer. A single biomarker may be present in 0.01% to 10% of subjects with cancer. A single biomarker may be present in 0.1% to 10% of subjects with cancer. A single biomarker may be present in 0.1% to 5% of subjects with cancer. Genetic analysis

[0172] Genetic analysis includes the detection of nucleotide sequence variants and copy number variations.Genetic variants can be determined by sequencing.Sequencing method can be massively parallel sequencing, that is, simultaneous (or continuous) sequencing of at least 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules. Sequencing methods may include, but are not limited to, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis (SBS), single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing, Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent or Nanopore platforms, and any other sequencing method known in the art.

[0173] Sequencing can be made more efficient by sequence capture, i.e., enriching the sample for sequences containing target sequences of interest, such as KRAS and / or EGFR genes or portions thereof containing sequence variant biomarkers. Sequence capture can be performed using immobilized probes that hybridize to the target of interest.

[0174] Cell-free DNA may contain a small amount of tumor DNA mixed with germline DNA.Sequencing methods that increase the sensitivity and specificity of detecting tumor DNA, and in particular gene sequence variants and copy number variations, can be useful in the methods of the present invention.Such methods are described, for example, in WO2014 / 039556.These methods can not only detect molecules with a sensitivity of up to 0.1% or more than 0.1%, but also distinguish these signals from the typical noise in current sequencing methods.Increasing the sensitivity and specificity of cfDNA blood-based samples can be achieved using various methods.Some methods involve highly efficient tagging of DNA molecules in samples, for example, tagging at least 50%, 75% or 90% of the polynucleotides in samples.This increases the possibility that low-abundance target molecules in samples can be tagged and then sequenced, significantly increasing the sensitivity of detecting target molecules.

[0175] Another method involves molecular tracking, which identifies sequence reads generated in duplicate from the original parent molecule and assigns the most likely identity of the base at each locus or position in the parent molecule. This significantly increases the specificity of detection by reducing noise generated by amplification and sequencing errors, which reduces the frequency of false positives.

[0176] The disclosed methods can be used to detect genetic variations in non-uniquely tagged initial starting genetic material (e.g., rare DNA) at concentrations that are less than 5%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% with specificity of at least 99%, 99.9%, 99.99%, 99.999%, 99.9999%, or 99.99999%. Sequence read data of tagged polynucleotides can subsequently be tracked to generate a consensus sequence for the polynucleotide with an error rate of 2%, 1%, 0.1%, or 0.01% or less.

[0177] In another example, a gene of interest can be amplified using primers that recognize the gene of interest. The primers can hybridize to the gene upstream and / or downstream of a specific region of interest (e.g., upstream of a mutation site). A detection probe can hybridize to the amplification product. The detection probe can specifically hybridize to the wild-type sequence or the mutated / variant sequence. The detection probe can be labeled with a detectable label (e.g., a fluorophore). Detection of the wild-type or mutant sequence can be performed by detecting the detectable label (e.g., fluorescent imaging). In the example of copy number variation, the gene of interest can be compared to a reference gene. The difference in copy number between the gene of interest and the reference gene can indicate gene amplification or deletion / truncation. Examples of platforms suitable for performing the methods described herein include digital PCR platforms, such as Fluidigm Digital Array.

[0178] Methods for analyzing nucleic acid sequence information are described herein. In various embodiments, the analysis method includes one or more models, each of which includes one or more of the following as separate components: survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. In various embodiments, the model includes a hierarchical model (e.g., nested model, multilevel model), a mixed model (e.g., regression such as logistic regression and Poisson regression, pooled, random effects, fixed effects, mixed effects, linear mixed effects, generalized linear mixed effects), a hazard model, an odds ratio model, and / or a replicated sample (e.g., repeated measures such as analysis of variance). In various embodiments, the model is a hierarchical random effects model. In various embodiments, the model is a hierarchical cubic spline random effects model. In various embodiments, the model is a cubic spline model. In various embodiments, the model is a generalized linear effects model. In various embodiments, the model is a linear effects model. In various embodiments, the model is a Cox proportional hazards model. In various embodiments, the analytical method includes assembly with a model. In various embodiments, the assembly includes generation of associated parameters. In one or more embodiments, the analytical method includes patient survival information and patient genetic information. As an example, the assembly with a model may include different models for different types of cancer, including subtypes, represented in the patient survival information. Each of the different models may be configured to determine a correlation between genetic factors and survival time of patients diagnosed with each type of cancer that the method is configured to evaluate. For example, genetic factors determined to have a strong correlation to cancer survival time (e.g., relatively short survival time and / or relatively long survival time) may be recommended as potential therapeutic targets.

[0179] In various embodiments, the analysis may include, as separate components, one or more of survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. For example, modeling may facilitate applying the above, such as patient survival information and patient genetic information. In various embodiments, the submodeling component may determine subsets of patient survival information and patient genetic information to generate different patient cohorts associated with different types of cancer and cancer subtypes. In various embodiments, the submodel includes a hierarchical model (e.g., nested model, multilevel model), a mixed model (e.g., regression such as logistic regression and Poisson regression, pooled, random effects, fixed effects, mixed effects, linear mixed effects, generalized linear mixed effects), a hazard model, an odds ratio model, and / or a replicated sample (e.g., repeated measures such as analysis of variance). In various embodiments, the submodel is a hierarchical random effects model. In various embodiments, the submodel is a hierarchical cubic spline random effects model. In various embodiments, the submodel is a cubic spline model. In various embodiments, the sub-model is a generalized linear effects model. In various embodiments, the sub-model is a linear effects model. In various embodiments, the sub-model is a Cox proportional hazards model. Each subset of patient survival information and patient genetic information may include information about patients diagnosed with different types of cancer and cancer subtypes. For example, the sub-modeling component can further apply the subset of patient survival information and patient genetic information to corresponding individual survival models developed for different cancer types, including subtypes. In various embodiments, information generated about the analytical method may be stored in memory (e.g., as model data). In various embodiments, information generated about the analytical method generates one or more survival models for individual subjects.

[0180] In various embodiments, the analysis of patient survival information and patient genetic information using a survival model includes a disease node determination and identification component that can identify, for each type of cancer, disease nodes included in the patient genetic information that are involved in the genetic mechanisms used to grow by each cancer type. In various embodiments, the disease node component identifies disease nodes based on observed correlations between genetic factors and cancer survival times provided in the patient survival information. For example, genetic factors that are frequently observed in association with short survival times for a particular type of cancer and less frequently observed in association with long survival times for a particular type of cancer can be identified as active genetic factors that play an active role in the genetic mechanisms of a particular type of cancer, including subtypes.

[0181] In various embodiments, disease node determination and identification includes disease-related parameters for the association between different cancer types to facilitate identifying active genetic factors associated with different cancer types. For example, highly associated cancer types may share one or more common critical underlying genetic factors. As will be readily understood by those skilled in the art, models of associated cancer types (e.g., survival models) dialectically exchange information to determine and / or identify active genetic factors across cancer types, including subtypes. In various embodiments, the disease-related parameters applied by disease node determination and identification are facilitated by modeling. In various embodiments, the generation of individual survival models may utilize one or more machine learning algorithms to facilitate the determination and / or identification of disease nodes associated with specific types of cancer, including subtypes, based on patient genetic information and disease-related parameters.

[0182] In some embodiments, the determination of a score system for a disease node is included in association with the node determination and identification of a cancer type, including subtypes. For example, the score of a disease node for a particular type of cancer, including subtypes, reflects the association of the disease node with the survival time of the particular type of cancer, including subtypes. In various embodiments, the score may be based on the frequency with which a particular genetic factor is directly or indirectly identified for patients diagnosed with a particular cancer type. In various embodiments, the analysis may include the above-mentioned survival, submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc., and may be related to values ​​below or above a specified threshold. For example, the higher the score associated with the included disease node and cancer type, the greater the contribution of the disease node to survival time. In various embodiments, the formation of disease nodes for each type of cancer, including subtypes, and the scores determined for active genetic factors may be collated in a data structure, such as a database.

[0183] Analytical methods including effect modeling are described herein. In various embodiments, effect modeling includes random effects, fixed effects, mixed effects, linear mixed effects, and generalized linear mixed effects. In various embodiments, effects include cubic splines. In various embodiments, effect modeling includes regression. In various embodiments, effect modeling includes logistic regression and Poisson regression. In various embodiments, the model does not include covariates. In various embodiments, the model includes covariates. In various embodiments, the covariates are information from medical records (including clinical laboratory records, e.g., genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc. Examples include age, line of treatment, smoking status (yes / no), gender, and various scoring and / or staging systems utilized for specific cancer disease patients, with illustrative examples including age (years), line of anti-EGFR treatment, smoking status (yes / no), gender (female / male), and the lung cancer patient-specific Van Walraven Elixhauser Comorbidity (ELIX) score (expressed as a weighted measure across multiple common comorbidities). Those skilled in the art will readily appreciate that covariates can include any number of data elements for individuals and individuals in a population, such as data elements from medical records (laboratory records, including, e.g., genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc.

[0184] In various embodiments, the analytical method includes generating a hierarchy including at least one first-level equation. In various embodiments, the first-level equation includes a truncated cubic spline. In various embodiments, the truncated cubic spline includes longitudinal data. This includes, for example, direct or indirect measurements of ctDNA levels, allele fractions, and tumor fractions. In various embodiments, additional level equations include covariates. In various embodiments, the covariates are information about an individual or individuals in a population derived from and / or stored in medical records (laboratory records, e.g., including genomic, epigenomic, nucleic acid, and other analyte results), insurance records, etc. Examples include age, line of treatment, smoking status (yes / no), gender, and various scoring and / or staging systems utilized for patients with a particular cancer disease. In various embodiments, a velocity plot is generated. In various embodiments, the velocity plot is a derivative of one or more equations, e.g., at least one first-level equation. In various embodiments, the analytical method includes one or more of equations (1), (2) and (3) as described in the Examples.

[0185] Described herein are analytical methods that include jointly solving different analytical components, including, as separate components, one or more of survival, modeling and submodeling, disease node determination and identification (e.g., driver mutations), disease association, disease subtyping, recurrence, metastasis, time to next treatment, etc. In various embodiments, the analytical method includes jointly solving one or more different models for different cancer types under a joint model framework. For example, the analytical method may include jointly solving one or more different survival models for different cancer types under a joint model framework. In various embodiments, the method includes determining relevant parameters. In various embodiments, the relevant parameters include, for example, the relationship between patient survival and an estimate of the current value of the patient's biomarker, or the relationship between patient survival and an estimate of the current change over time for the patient's biomarker. In various embodiments, this includes the slope, and the relationship between overall survival and a current estimate of the patient's longitudinal area under the trajectory as a proxy for the cumulative effect of the biomarker. It will be readily apparent to those skilled in the art that the relevant parameters can take many forms and can be combined. For example, one can examine the relationship between overall survival and an estimate of the current value and an estimate of the current slope of a patient's longitudinal trajectory. [Example]

[0186] Example 1 Joint Modeling The inventors have applied joint modeling (JM) of longitudinal and time-to-event data combined with next-generation sequencing (NGS) genetic testing to demonstrate the ability to detect changes in a biomarker (or several biomarkers) over time in relation to the probability of survival for a particular patient. The detection and characterization of genomic biomarkers using methods and techniques demonstrates how the evolution of such biomarkers can be associated with and predict patient survival. As an example, a real-world application of joint modeling generates a patient-level monitoring system designed to improve clinician decision-making capabilities.

[0187] In particular, JM includes the ability to appropriately accommodate endogenous time-varying covariates. Since most biomarkers fall into this category, this leads to reduced bias in parameter estimates, improved statistical inference, and the ability to perform dynamic patient-level predictions, where predictions are based on partial or complete biomarker history. Joint modeling is flexible in that both frequentist and Bayesian approaches have been deployed. Here, we adopted a Bayesian approach based on the Markov Chain Monte Carlo sampling algorithm for computational efficiency.

[0188] Example 2 Genetic testing using next-generation sequencing We selected a patient cohort from the Real-World Evidence Database, which contains real-world outcomes, de-identified genomic data, and structured payer claims data for over 240,000 patients. For demonstration purposes, distinct target populations within this dataset included patients diagnosed with non-small cell lung cancer (NSCLC) harboring the EGFR L858R mutation and colorectal cancer (CRC) harboring KRAS G12D and KRAS12V mutations, respectively. For the longitudinal component of this study, only patients with a minimum of three temporal measurements were included. Upon fulfilling these criteria, the resulting cohort consisted of 252 patients. The biomarkers of interest, i.e., longitudinal outcomes, were patient variant allele frequency (AF) and tumor fraction (TF), and it was the progression of these biomarkers over time that we intended to correlate with patient survival.

[0189] Example 3 method The joint modeling framework is divided into two submodels, where, once analyzed, information from these submodels is combined to determine whether an association exists between the two. More specifically, the first submodel focuses on providing a sufficient representation of longitudinal data (at the patient level), while the second submodel assesses patient survival. In this study, a general linear mixed model (GLMM) was used to assess the temporal progression of each biomarker, while a Cox proportional hazards (CPH) model examined patient survival. It is important to note that because the distribution of each biomarker was highly skewed, the analysis was based on a log transformation of both AF and TP to conform to the GLMM normality assumption. Furthermore, for some patients with complex biomarker progression, a cubic spline model was used to describe the patient-level response. Furthermore, because factors such as age and gender are often confounding with survival, these factors were included in the CPH model to act as statistical controls.

[0190] As will be readily understood by those skilled in the art, the methods and techniques described herein support the determination of associations between longitudinal data and time-to-event data, referred to as association structures. Examples of association structures include, but are not limited to, the relationship between patient survival and an estimate of the current value of a patient's biomarker, the relationship between patient survival and an estimate of the current change over time for a patient's biomarker, e.g., slope, and the relationship between overall survival and an estimate of the current area under the patient's longitudinal trajectory, which is often used as a proxy for the cumulative effect of a biomarker. Associate structures can take many forms and can be combined. For example, one can examine the relationship between overall survival and an estimate of the current value and current slope of a patient's longitudinal trajectory. Here, we describe association structures for current values, slopes, and combinations thereof; however, those skilled in the art will understand that numerous association structures are available for exploration, many of which are not explicitly mentioned above but will be readily apparent to those skilled in the art. After establishing suitable JMs for each biomarker, these JMs are then used to inform dynamic predictions. That is, overall survival is predicted for each patient depending on the nature of the association structure between the longitudinal and time-to-event data, or more specifically, survival is predicted for a given patient using measurements captured up to a given time point, and as additional measurements are collected, the patient survival prediction is adjusted accordingly - hence the term "dynamic prediction."

[0191] Example 4 statistical analysis All statistical analyses were performed using R version 4.1.3, where the JMBayes2 package performed joint modeling. As previously mentioned, for the longitudinal component of the study, each patient had a minimum of three temporal measurements, where the first measurement coincided with the patient's initial Guardant360 test, while the remaining measurements followed accordingly. In all, 252 patients met these criteria, resulting in a total of 909 measurements collected for AF and TP, respectively, ranging from November 19, 2014, to September 30, 2022. The distribution of each biomarker is shown in Figure 1, followed by relevant summary statistics (see Table 1). JM results indicate that the most recent longitudinal change in each biomarker is associated with patient survival (AF: p-value = 0.0139; TF: p-value = 0.0332). Through these associations, graphical depictions of patient-level survival curves can be displayed, allowing assessment of clinical outcomes based on a patient's specific biomarker evolution. [Table 1] Example 5 result

[0192] The distributions of allele frequencies and tumor fractions are shown in Figure 1. To complement the descriptive statistics, patient-level longitudinal data for each biomarker are displayed as spaghetti plots in Figure 2, which illustrate the complexity of patient-level longitudinal progression for each biomarker and reinforce the skewed nature of the data. To comply with the normality assumption required by GLMM, a log transformation was applied to each biomarker, and to correct for the complexity observed in patient-level evolution, natural cubic splines were utilized to model each patient's longitudinal characteristics within the GLMM structure. Plots of the fitted GLMM results for each patient are found in Figure 3, and the GLMM fixed and random effects for each biomarker are depicted in Figure 8. Because biomarkers were collected for the same set of patients, only a single CPH model needed to be fitted. Of the 252 patients, 99 experienced an event (death), while the remaining observations were censored. Because both age and gender are often confounded with survival, the initial CPH model included these covariates as statistical controls. However, analysis of the initial model revealed that neither age (p-value = 0.519) nor gender (p-value = 0.310) was statistically significant at the 0.05 level. Similarly, models that included age and gender separately produced similar results (age, p-value = 0.56) and (gender, p-value = 0.33). Subsequently, a null CPH model (a model in which no covariates were present) was used for joint modeling.

[0193] Example 6 fitting The fitted cubic spline-based GLMM results for the log-transformed biomarkers are shown in Figure 3. For each biomarker, three JMs were analyzed, each fitted to the relevant structures (and combinations thereof) cited above. Because the analysis was performed under a Bayesian paradigm, care was taken to ensure that model parameters were accurately estimated. In doing so, each model consisted of two strands, each with 9,000 burn-in iterations followed by 90,000 iterations, and a rarefaction factor of 3 was implemented to account for potential autocorrelation issues. Similarly, inspection of the trace plots visually confirmed that the model parameters converged satisfactorily. The joint modeling results for each respective biomarker (numbered 1–3) are summarized in Tables 2 and 3. Because a Bayesian approach was employed, 95% credible intervals are reported rather than frequentist-based confidence intervals.

[0194] The results in Tables 2 and 3 reveal that the second JM for each biomarker shows promise, as evidenced by the respective p-values ​​(0.0139 and 0.0332), suggesting the existence of an association between the current slope and patient survival. Further information can be extracted from these tables as well. That is, it is possible to calculate the hazard ratio corresponding to each of these association structures. For example, referring to the mean values ​​in Table 2, if the current rate of change in allele frequency increases by 10% over 100 days, the resulting hazard ratio is 1.19, meaning that the hazard of death associated with such an increase is 19% higher. Similar calculations can be performed for maximum tumor percentage. [Table 2] [Table 3]

[0195] Example 7 Dynamic Forecasting While HRs closely reflect overall trends, from a precision medicine perspective, the real strength of the JM methodology lies in producing dynamic predictions. Because the concept of dynamic prediction is best understood through a visual display, a graphical depiction of this process is provided in Figures 4 and 5.

[0196] The top panel of Figure 4 depicts the longitudinal trajectory (as seen by the blue line) associated with a patient's biomarker evolution; as additional measurements are captured, the trajectory adjusts accordingly. It is important to note that the focus is on the current slope of the trajectory, since the JM used in generating dynamic predictions is built on this relational structure. In this example, we examine time frames spanning 0-300, 600, and 900 days, respectively. Directly beneath each trajectory, i.e., the bottom panel, is the fitted survival curve. Note that each curve is updated as new biomarker information becomes available. For example, as shown by Figure 4, from 0-300 days, the trajectory for patient 106 is declining. Through inspection of the corresponding survival curve, we make predictions from approximately 1000 days; i.e., if survival is assessed at 1300 days, the patient's probability of survival is approximately 0.71, or 71%. Similarly, at 600 days, additional biomarker values ​​are captured, which changes the trajectory; now, even though the trend remains downward, the slope is increasing. Evaluating at 1000 days (1600 days) reveals a 6% drop in the patient's survival estimate, from 71% to 65%. Generally, as the slope increases, survival decreases, so such an outcome is expected. Finally, the final set of measurements, collected up to 800 days, results in a slight increase in the slope, so the patient's survival estimate decreases only slightly, from 65% to 64%. Here, a 1000-day projection was used, but survival trends remain relatively comparable regardless of the projected time frame.

[0197] In contrast to patient 106, the slope of the trajectory for patient 94 (see Figure 5) remains fairly consistent over the period considered, although a slight increase in slope is observed. Therefore, we should expect little change in survival probability. If we were to forecast from 1000 days onwards as before, the predicted survival probabilities would be 71%, 70%, and 69%, respectively - which fits the prediction. As with the HR calculation, a similar dynamic prediction can be made based on maximum tumor percent.

[0198] Using the methods and techniques described herein, JM results indicate that the most recent change in each biomarker over time is associated with patient survival (AF: p-value = 0.0139; TF: p-value = 0.0332). Through these associations, a graphical depiction of patient-level survival curves can be displayed, allowing for assessment of clinical outcomes based on a patient's unique biomarker evolution.

[0199] Example 8 Consideration In addition to the many available JM options, dynamic predictive capabilities are particularly beneficial because they are well suited to improving clinicians' decision-making capabilities. This is because, in a real-world medical environment, a patient's condition is constantly changing, and as a result, it is often in the patient's best interest to make informed decisions using the most recent data available. As demonstrated, JM inherently captures the ever-changing patient landscape, and as changes occur, JM adapts accordingly. Therefore, by utilizing JM's ability to link up-to-date information to patient survival, clinicians may be well-positioned to modify and / or adjust treatment plans with the ultimate goal of improving patient survival. Furthermore, the application of techniques such as JM supports the generation of vast amounts of genetic data. Those skilled in the art will understand that the analysis performed herein can be applied to other cancer types and mutations, and therefore, there are numerous biomarkers, cancer types, and mutations available for investigation, and additional relevant biomarkers may be identified in the process. This approach supports the creation of patient-specific monitoring systems tailored to specific cancer types and mutation combinations.

[0200] Example 9 Hierarchical cubic spline random effects model The use of a hierarchical cubic spline random-effects model (HCSREM) applied to a retrospective, real-world cohort of patients diagnosed with advanced non-small cell lung cancer (NSCLC) is described herein. While those skilled in the art will appreciate that the proposed framework can be applied to longitudinal biomarkers, or any combination of biomarkers, we are interested in ctDNA levels measured by the maximum variant allele fraction of all somatic variants detected through liquid biopsies. A key advantage of the method is its ability to incorporate patient information, taking into account several relevant covariates. Finally, to improve interpretation, model results are graphically displayed in the form of longitudinal prediction estimates, each based on a distinct set of patient characteristics. Within this process, patient-level predictions are directly compared, where the comparison is enhanced by subsequently defined velocity plots.

[0201] Example 10 Data sources and patient cohorts The cohort used to illustrate the utility of the methodology is based on observational data and sourced from the Real World Evidence De-identified Clinical Genomics Database, which includes structured commercial payer claims collected from inpatient and outpatient facilities in both academic and community settings.

[0202] Patients selected for the cohort were diagnosed with advanced non-small cell lung cancer (NSCLC) and underwent at least three genomic liquid biopsy tests in the United States between June 1, 2014, and June 30, 2023. Only patients receiving targeted therapy for EGFR mutations were included, with the following treatments considered: osimertinib, afatinib, dacomitinib, erlotinib, gefitinib, and amivantamab. All patients were required to have at least three blood samples during a specific anti-EGFR treatment line, or within 30 days before and after the start of that line. Patients whose first genomic test for that line of treatment occurred more than 120 days after the start of that line of treatment were excluded. For patients with multiple lines of treatment who met these criteria, the earliest line of treatment was selected for inclusion in the study. Finally, patients with suspected germline mutations were removed from the cohort.

[0203] Example 11 Response variables and study covariates The response variable, ctDNA measurements captured over time, is reported as a percentage. If a sample contained a ctDNA level below the assay's detection limit, the value was replaced with 0.04%, the lowest value in the cohort and consistent with the test's detection limit. All covariates except mortality were captured at baseline, where the baseline period was defined as 6 months before the index date, i.e., the date of the patient's first genomic test. Baseline covariates included age (years), anti-EGFR treatment line, smoking status (yes / no), gender (female / male), and the Van Walraven Elixhauser Comorbidity (ELIX) score (expressed as a weighted measure across multiple common comorbidities), which is specific to lung cancer patients. Because the cohort is based on real-world data, it is not possible to directly align the treatment start date with the patient's first genomic date, as is achievable in prospective studies. Therefore, the number of days between the first genomic test and the start of treatment was added as a covariate and used as a statistical control, set to day 0 in the analysis to mimic the post-treatment scenario. Patient mortality, captured as alive versus dead within the study time frame, was also included.

[0204] Example 12 Example Statistical Model The mathematical details of HCSREM are described herein, which are flexible enough to capture non-linear trends in variability, allowing for the direct incorporation of patient characteristics in the form of covariates. In addition to these properties, this model can provide a unique corresponding temporal ctDNA pattern for each combination of covariate values. It is this type of ability to provide patient-specific information that makes this methodology attractive in targeted oncology efforts.

[0205] The model is partitioned into first and second level equations, which create a hierarchical structure. The first level equation assumes the form of a truncated cubic spline and captures how a particular patient's ctDNA levels change over time (see equation (1)). At a high level, this is achieved by creating a function that is divided into various compartments across the abscissa. Within each compartment, the data is fit using a cubic polynomial, with the ends of the successive cubic polynomials connected by knots. While "automated" methods exist for determining the amount and placement of knots, the location and number of knots can be strategically devised based on data exploration. Ultimately, the cubic spline model combines the separate compartments to form a single uniform function that represents the data.

number

number

[0206] In equation (1), the ctDNA measurements (or variations thereof) captured over time are expressed as Y ij where i is used to index the patient and j indexes the measurement occasion. The time points captured within a patient are denoted by t ij is shown by

number

[0207] The importance of the second level equations is that they contain information about individual patient characteristics and relate these characteristics to the response parameters themselves. The second level equations are shown below.

number

number

[0208] If a model contains covariates, it is called a conditional model; otherwise, it is an unconditional model. Unconditional models provide outcomes at the cohort level, while conditional models are responsible for producing patient-level outcomes.

[0209] Furthermore, the described velocity plots are useful for examining the direction and speed at which ctDNA levels change at a given time point, i.e., for the purpose of instantaneous rate of change (IRC). Each model that generated patient trajectories has a cubic spline at its root. An advantageous property of cubic splines is that they are twice differentiable, thus allowing the IRC at a given time point to be calculated.20 For the spline model employed, this quantity, up to taking the first derivative of equation (1) with respect to time, becomes:

number

[0210] The value of IRC is given by the slope of a line tangent to the patient trajectory, where positive values ​​correspond to an increasing IRC, negative values ​​correspond to a decrease, and an IRC value of 0 indicates either a peak or trough has been reached or the trajectory is flat. The further the IRC value is from 0, the more extreme the rate of change.

[0211] Example 13 Statistical analysis and results Data were extracted using SAS software package 9.4 (SAS Institute, Cary, NC, USA), and all statistical analyses for HCSREM were performed using R version 4.1.3. A total of 400 patients with advanced NSCLC who had at least three G360 tests were identified from the GuardantINFORM database. 73 patients were excluded because their first test was more than 120 days after treatment initiation, and 5 patients were excluded due to germline mutations. Of the remaining patients, 163 received anti-EGFR treatment and had a total of 561 longitudinal ctDNA measurements; these 163 patients defined the cohort used in the analysis. The mean age of these patients was 62 years, 66% were female, the mean anti-EGFR treatment line was 1, and the mean time between G360 testing and treatment initiation was 0 days (range, -115 to 30 days) (Table 1). [Table 4]

[0212] We developed an unconditional model that was fitted to the transformed data using knots set at 50, 125, 250, 500, 750, 1000, and 1250 days, respectively, as shown in Figure 9. Other knot orientations were explored to ensure consistency, but different orientations did not significantly alter the results. Because spline model parameter estimates are difficult to interpret, results are presented graphically; however, parameter estimates and associated outputs are provided in the Supplementary Information for reference. A graphical representation of the unconditional model, referred to as the response pattern, is shown in Figure 10. Here, the black curve represents the response pattern for the cohort, while each black dot represents a ctDNA level value. The purple area represents the 95% confidence band for the estimated trajectory.

[0213] The response pattern suggests a substantial decline in ctDNA levels between the initial G360 test and 30 days, followed by a rapid rise until 150 days, at which point ctDNA levels decline slightly and then rise again, albeit at a less extreme rate, around 300 days. Additionally, ctDNA levels decline between 550 and 1000 days, then rise again between 1000 and 1600 days. Over time, as the number of data points decreases, the corresponding 95% confidence bands expand. The flexibility built into the unconditional model revealed details hidden in the data that simpler models would not detect. Despite this, the unconditional model only estimates the response pattern for the cohort and does not account for the possibility that patients with different characteristics may exhibit different response patterns. To assess the impact of incorporating patient characteristics, a conditional model incorporating all baseline covariates was fitted to the data. As is typical in hierarchical models, all numerical covariates were centered about their respective means.

[0214] Example 14 Age and health status, response patterns Here, Figure 11 shows how baseline age and health status, as measured by ELIX score, affect response patterns in female non-smokers receiving their first line of EGFR-TKI treatment. Results are divided into surviving patients versus deceased patients. Because data become sparse after 400 days, we examine only the first 400 days. The example shown above reveals that patients with different characteristics have different response patterns. In the upper left panel, we contrast response curves for 30- and 80-year-olds with average ELIX scores.

[0215] These results suggest that 80-year-old patients did not show an initial post-treatment drop in ctDNA levels compared to 30-year-old patients who demonstrated a rapid decrease followed by a rapid increase. The upper center panel shows the response pattern for patients with the average age; a maximum ELIX score of 13 appears significantly different compared to the same patient with a minimum ELIX score of 0, suggesting that patients with many comorbidities showed a delayed treatment response. In the upper right panel, response patterns are displayed for older patients with a high comorbidity burden and younger, otherwise healthy patients, illustrating how age / health status combinations amplify differences in response patterns. Although not shown, a decreasing trend in ctDNA values ​​is observed for patients surviving at the end of the study beyond 400 days, while the trend increases for patients who died before the end of the study.

[0216] Example 15 Velocity Plot To focus on the behavior of response patterns, we generated rate plots displaying the IRCs for the corresponding response patterns (Figure 12). While the information shown in the rate plots can be gleaned from the response patterns themselves, differences in response patterns are highlighted when they are examined through an IRC lens. Therefore, comparing rate plots can provide additional clues about where response patterns are similar and where they diverge, based on IRC values. Another advantage of utilizing rate plots exists when the baseline values ​​between response patterns are dissimilar, and therefore, differences between response patterns may be due to the fact that biomarker values ​​were different at the start. In these cases, using rate plots to make comparisons may be more appropriate because the IRCs are invariant with respect to the baseline values ​​of the biomarkers.

[0217] Interpreting kinetic plots is understood by those skilled in the art. We can now focus on the leftmost panel. Over the first 100 days, the kinetic plots (red curves) for 80-year-old surviving and deceased patients showed different patterns. For survivors, the IRC was initially positive but slowed to 0 around day 20 (see dashed line, which indicates the peak in the corresponding response curve), and then declined, with the most rapid rate of decline (-0.026 logits per day) occurring around day 43. After day 43, the IRC continued to decline and remained relatively flat even after day 100. In contrast, the kinetic plot for the 80-year-old deceased patient displayed an almost opposite pattern.

[0218] Example 16 Consideration Methods and techniques applicable to the analysis of complex longitudinal genomic data are described herein.As shown, the inventors analyzed observational data and demonstrated their use in different data settings, including hypothesis generation, statistical inference and patient monitoring.Here, the 95% confidence intervals that the inventors utilized could not retain their traditional inferential meaning, but instead were used as "guidelines" in identifying differences in response patterns.This supported the generation of thousands of response patterns.

[0219] Those skilled in the art will easily understand that the framework described can also be applied to each cohort.When statistical inference is the goal, there is the potential to generate and compare numerous response patterns, so based on a priori hypotheses, the number of comparisons should be minimized, and common considerations such as controlling type I error should be taken into account.Hypothesis can include comparing response patterns between patients with a given set of covariate values ​​(covariates from other studies can be used as statistical controls), but can also include making hypotheses about the nature of the relationship between response pattern behavior and covariate values ​​themselves.

[0220] Another example involves patient monitoring. The general idea is that each response pattern is a reasonable representation of a patient, explained by its own unique set of characteristics—thus, the same response pattern can serve as a reference for new patients who share these characteristics. In addition, if survival status (death or not) is incorporated into the model, reference response patterns for survivors and non-survivors can be generated. Thus, if a new patient's response pattern matches that of survivors, no intervention is required, but if the response pattern closely resembles that of non-survivors, intervention may be required. Using rate plots to compare response patterns—especially when the baseline values ​​between response patterns are dissimilar—can further enhance this process. To ensure reliable classification, such monitoring systems should undergo internal and external validation. Internal validation can be achieved by generating training and test datasets and then applying approximately k-fold cross-validation to assess classification accuracy. If an acceptable level of accuracy is achieved, external validation can be achieved to determine whether new patients, i.e., patients not involved in cross-validation, are also classified with high accuracy.

[0221] As explained, changes in ctDNA levels can vary significantly from patient to patient over time. Here, the methods and techniques mentioned above generate patient-level results, where such results reveal ctDNA dynamics for clinical decision-making.

Claims

1. 1. A method for determining patient response in at least one patient, comprising: obtaining nucleic acid sequence information from at least one patient, the nucleic acid sequence information comprising measurements of temporal changes in biomarkers; and determining a patient response for said at least one patient; A method comprising:

2. 10. The method of any preceding claim, wherein the biomarker comprises circulating tumor DNA (ctDNA).

3. 10. The method of any preceding claim, wherein the biomarkers comprise allele frequencies and tumor fractions.

4. 10. The method of any preceding claim, wherein determining the patient response for the at least one patient comprises use of a database.

5. 10. The method of any preceding claim, wherein the database includes medical records and / or insurance records.

6. 10. A method according to any preceding claim, wherein using the database comprises applying a model.

7. 10. The method of any preceding claim, wherein the model is a hierarchical model.

8. 10. The method of any preceding claim, wherein the model is an effects model.

9. 10. The method of any preceding claim, wherein the model is a regression model.

10. 10. The method of any preceding claim, wherein the model is a joint model.

11. 10. The method of any preceding claim, wherein the hierarchical model is a hierarchical random effects model.

12. 10. The method of any preceding claim, wherein the model comprises cubic splines.

13. 10. The method of any preceding claim, wherein the model comprises a regression model.

14. 10. The method of any preceding claim, wherein the hierarchical random effects model comprises generating data from nucleic acid sequence information comprising temporal changes in a biomarker comprising circulating tumor DNA (ctDNA) from at least one subject among a plurality of subjects.

15. 10. A method according to any preceding claim, wherein said generating data comprises generating a cubic spline for at least one subject of a plurality of subjects.

16. 10. The method of any preceding claim, wherein said generating data comprises generating response parameters that include one or more covariates.

17. 10. The method of any preceding claim, wherein said generating data comprises generating response parameters without covariates.

18. 10. The method of any preceding claim, wherein the response parameters are fitted to a multivariate normal distribution.

19. 10. The method of any preceding claim, wherein said determining a patient response for said at least one patient comprises generating a rate plot.

20. 10. The method of any preceding claim, wherein said determining a patient response for said at least one patient comprises a comparison with said model.

21. 10. The method of any preceding claim, wherein the joint model comprises at least two models.

22. 10. The method of any preceding claim, wherein the joint model comprises a correlation factor between the at least two models.

23. 10. The method of any preceding claim, wherein the joint model comprises a cubic spline and a proportional hazards model.

24. 10. The method of any preceding claim, wherein the biomarkers are measured by next generation DNA sequencing.

25. 10. The method of any preceding claim, wherein next generation DNA sequencing comprises ligating a non-unique barcode to the ctDNA.

26. 10. The method of any preceding claim, wherein next generation DNA sequencing comprises ligating a unique barcode to the ctDNA.

27. 10. The method of any preceding claim, wherein next generation DNA sequencing comprises ligating non-unique barcodes to ctDNA fragments, wherein the non-unique barcodes are present in at least a 20-fold, at least a 30-fold, at least a 50-fold, or at least a 100-fold molar excess.

28. 1. A method for determining patient response in at least one patient, comprising: obtaining nucleic acid sequence information from at least one patient, the nucleic acid sequence information including measurements of temporal changes in biomarkers including circulating tumor DNA (ctDNA); and determining a patient response for said at least one patient comprising using a database comprising medical and / or insurance records from a plurality of subjects, wherein said using the database comprises applying a hierarchical random-effects model. A method comprising:

29. 10. The method of any preceding claim, wherein the hierarchical random effects model comprises generating data from nucleic acid sequence information comprising temporal changes in ctDNA from at least one subject among a plurality of subjects.

30. 10. The method of any preceding claim, wherein the hierarchical random effects model comprises generating a cubic spline for at least one subject in the plurality of subjects.

31. 10. The method of claim x, wherein the hierarchical random-effects model comprises a response parameter that includes one or more covariates for at least one subject among the plurality of subjects.

32. 10. The method of any preceding claim, wherein the database includes medical and / or insurance records for the plurality of subjects.

33. 1. A method for determining patient response in at least one patient, comprising: obtaining nucleic acid sequence information from at least one patient, the nucleic acid sequence information including measurements of temporal changes in biomarkers including circulating tumor DNA (ctDNA); and determining a patient response for said at least one patient comprising using a database comprising medical and / or insurance records from a plurality of subjects, wherein said using the database comprises applying a joint model comprising a cubic spline and a proportional hazards model generated from data from nucleic acid sequence information for at least one subject among the plurality of subjects. A method comprising:

34. 10. The method of any preceding claim, wherein the database includes medical and / or insurance records for the plurality of subjects.

35. A system including a machine including at least one processor and storage containing instructions capable of performing any of the preceding methods.

36. A computer-readable medium containing instructions capable of performing any of the preceding methods.