Multi-tiered testing for tracking cancer heterogeneity
The multiple-tiered analysis method efficiently identifies and tracks tumor heterogeneity in large populations by using a first screen and intra-individual analyses to reduce resource consumption and enhance diagnostic accuracy for rare cancers.
Patent Information
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- FLAGSHIP PIONEERING INNOVATIONS VI LLC
- Filing Date
- 2025-01-03
- Publication Date
- 2026-07-16
AI Technical Summary
Current diagnostic technologies, both point-of-care (POC) tests and complex centralized tests, fail to accurately identify individuals with rare cancers in large populations due to low accuracy and high resource consumption, respectively.
A multiple-tiered analysis method involving a first screen to eliminate non-cancerous individuals, followed by intra-individual analyses using target and reference nucleic acids to generate background-corrected methylation information, and a second analysis to track tumor heterogeneity over time.
This approach significantly reduces resource consumption by rapidly identifying a large proportion of non-cancerous individuals and accurately diagnosing cancerous individuals, achieving improved sensitivity, specificity, and positive predictive value compared to conventional methods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 636,405 filed April 19, 2024, and U.S. Provisional Patent Application No. 63 / 617,989 filed January 5, 2024 the entire disclosure of each of which is hereby incorporated by reference in its entirety for all purposes. BACKGROUND
[0002] Diagnostic technologies include simple, point of care (POC) tests applied to large populations to identify relatively common diseases as well as complex, centralized tests applied to select populations. However, although POC tests can be applied to large populations, they are incapable of identifying individuals for cancer at a high enough accuracy to be feasible for implementation. Similarly, although complex, centralized testing can be deployed for rare population testing, such testing is often invasive, expensive, and fails when applied for detecting rare cancers in large patient populations. For example, complex, centralized testing suffers from poor performance (e.g., high number of false positives and / or low positive predictive value) when attempting to diagnose rare cancers in large patient populations. Thus, current POC tests are not suitable for identifying individuals with cancer and for tracking such individuals over time. SUMMARY
[0003] Disclosed herein are methods involving a multiple tiered analysis for tracking tumor heterogeneity in subjects. In particular, the methods disclosed herein involving a multiple tiered analysis are useful for tracking tumor heterogeneity in individuals from a large population (e.g., millions of individuals) who have a rare cancer. The multiple tiered analysis involves a first screen, which eliminates a large proportion of individuals who are identified as negative for cancer. For subjects that are identified as not negative for cancer, they can be provided an intervention (e.g., a tumor therapeutic). These subjects undergo additional analyses (e.g., one or more intra-individual analysis and / or a second analysis) which can be performed using samples obtained from the subjects across different timepoints. For example, intra-individual analyses can be conducted for each sample obtained from the subject. By doing so, a change in tumor heterogeneity can be determined which is informative for determining the efficacy of the provided intervention. Altogether, the multiple tiered analysis can be useful e.g., for guided therapy. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] These and other features, aspects, and advantages of the present invention will become better understood with regard to the following description and accompanying drawings. It is noted that wherever practicable, similar or like reference numbers may be used in the figures and may indicate similar or like functionality. For example, a letter after a reference numeral, such as “third party entity 155A,” indicates that the text refers specifically to the element having that particular reference numeral. A reference numeral in the text without a following letter, such as “third party entity 155,” refers to any or all of the elements in the figures bearing that reference numeral (e.g. “third party entity 155” in the text refers to reference numerals “third party entity 155A” and / or “third party entity 155B” in the figures).
[0005] Figure (FIG.) 1A depicts an overall flow process of the multiple-tiered process for tracking tumor heterogeneity, in accordance with an embodiment.
[0006] FIG. IB depicts an overall flow process of the multiple-tiered process for tracking tumor heterogeneity, in accordance with a second embodiment.
[0007] FIG. IC depicts an overall system environment including a tumor heterogeneity system, in accordance with an embodiment.
[0008] FIG. 2A depicts a block diagram of the tumor heterogeneity system, in accordance with an embodiment.
[0009] FIG. 2B depicts an example conversion of nucleic acids, in accordance with an embodiment.
[0010] FIG. 2C shows the results of nitrite conversion on select nucleotides, in accordance with a second embodiment. Figure adapted from Li et al. (2022) Genome Biology 23:122.
[0011] FIG. 3 A depicts example methylation information useful for determining whether an individual is at risk for cancer, in accordance with an embodiment.
[0012] FIG. 3B shows an example flow process for determining whether an individual is at risk for cancer, in accordance with an embodiment.
[0013] FIG. 3C depicts an example process of combining sequence information of target nucleic acids and reference nucleic acids to generate a signal informative for determining presence or absence of cancer, in accordance with an embodiment.
[0014] FIG. 3D is an illustrative example of a signal informative for cancer, in accordance with an embodiment.
[0015] FIG. 3E shows aligned sequence reads of an analyte and a corresponding window of a kmer size, in accordance with an embodiment.
[0016] FIG. 3F shows the generation of metrics from sequence reads across 2k possible patterns, in accordance with an embodiment.
[0017] FIG. 3G shows an example data structure including information useful for training machine learning models, in accordance with an embodiment.
[0018] FIG. 4A shows an example flow process involving a first and second intra-individual analyses, in accordance with a first embodiment.
[0019] FIG. 4B shows an example flow process involving a first and second intra-individual analyses, in accordance with a second embodiment.
[0020] FIG. 5 illustrates an example computer for implementing the entities shown in FIGs. 1A-1C, 2A, 3A-3G, and 4A-4B.
[0021] FIG. 6 shows example performance of different tiers of the multiple tier analysis for diagnosing individuals with cancer (e.g., prostate cancer).
[0022] FIG. 7 depicts performance of a single tier analysis and a two-tier analysis of a population involving 1046 samples.
[0023] FIG. 8 shows an example sample from which target nucleic acids and reference nucleic acids are obtained. DETAILED DESCRIPTION Definitions
[0024] Terms used in the claims and specification are defined as set forth below unless otherwise specified.
[0025] The terms “subject,” “patient,” and “individual” are used interchangeably and encompass a cell, tissue, or organism, human or non-human, male or female.
[0026] The term “sample” can include a single cell or multiple cells or fragments of cells or an aliquot of body fluid, such as a blood sample, taken from a subject, by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, or intervention or other means known in the art. Examples of an aliquot of body fluid include amniotic fluid, aqueous humor, bile, lymph, breast milk, interstitial fluid, blood, blood plasma, cerumen (earwax), Cowper’s fluid (pre-ejaculatory fluid), chyle, chyme, female ejaculate, menses, mucus, saliva, urine, vomit, tears, vaginal lubrication, sweat, serum, semen, sebum, pus, pleural fluid, cerebrospinal fluid, synovial fluid, intracellular fluid, and vitreous humour.
[0027] The term “obtaining information,” “obtaining marker information,” and “obtaining sequence information” encompasses obtaining information that is determined from at least one sample. Obtaining information (e.g., marker information or sequence information) encompasses obtaining a sample and processing the sample to experimentally determine the information (e.g., marker information or sequence information). The phrase also encompasses receiving the information, e.g., from a third party that has processed the sample to experimentally determine the information.
[0028] The terms “marker,” “markers,” “biomarker,” and “biomarkers” encompass, without limitation, lipids, lipoproteins, proteins, cytokines, chemokines, growth factors, peptides, nucleic acids (e.g., DNA or RNA), genes, and oligonucleotides, together with their related complexes, metabolites, mutations, variants, polymorphisms, modifications, fragments, subunits, degradation products, elements, and other analytes or sample-derived measures. A marker can also include mutated proteins, mutated nucleic acids, variations in copy numbers, and / or transcript variants, in circumstances in which such mutations, variations in copy number and / or transcript variants are useful for generating a prediction model, or are useful in prediction models developed using related markers (e.g., non-mutated versions of the proteins or nucleic acids, alternative transcripts, etc.).
[0029] The term “screen” or a “first analysis” refers to a step in the first tier of a multiple tiered analysis. The screen achieves a high specificity and removes a large majority of true negatives (e.g., individuals not at risk of a cancer). In various embodiments, the “screen” refers to an in silico screen that involves application of a machine learning model. For example, such a machine learning model may analyze sequence information (e.g., methylation information) and predicts whether individuals are likely to be at risk of the cancer.
[0030] The phrase “second analysis” refers to a step in the second tier of a multiple tiered analysis. The second analysis is performed on individuals who were identified, using the screen, as not negative for cancer. Thus, the second analysis achieves a higher positive predictive value than the screen, given that the screen removes a large proportion of the true negatives. In various embodiments, the “second analysis” refers to an in silico analysis that involves application of a machine learning model that analyzes sequence information (e.g., methylation information). The second analysis can predict whether individuals have cancer. In various embodiments, the second analysis is implemented to predict a change in tumor heterogeneity for purposes of tracking tumor heterogeneity in a subject.
[0031] The phrase “intra-individual analysis” refers to an analysis performed for an individual that removes baseline biological signatures that are less informative for determining whether the individual is at risk for cancer. In various embodiments, the intraindividual analysis involves combining information from target nucleic acids and reference nucleic acids of an individual to generate a signal informative for determining presence or absence of cancer within the individual. By combining the information from the target nucleic acids and the reference nucleic acids, the generated signal can be more informative of presence or absence of cancer in comparison to a signal derived from the target nucleic acids alone.
[0032] The phrase “target nucleic acids” refers to nucleic acids of an individual that contain at least signatures that may be informative for determining presence or absence of cancer. The target nucleic acids may further include baseline biological signatures of the individual that are not informative or less informative. In various embodiments, target nucleic acids may be nucleic acids derived from a diseased cell that is associated with cancer. For example, target nucleic acids may be cell-free nucleic acids originating from cancer cells. Target nucleic acids can be any of DNA, cDNA, or RNA. In particular embodiments, target nucleic acids include DNA.
[0033] The phrase “reference nucleic acids” refers to nucleic acids of an individual that contain baseline biological signatures of the individual. Here, the baseline biological signatures of the individual may be present when the individual is healthy, and therefore, the baseline biological signatures are less informative for determining presence or absence of cancer in comparison to sequence information of the target nucleic acids. Reference nucleic acids can be any of DNA, cDNA, or RNA. In particular embodiments, reference nucleic acids include DNA.
[0034] It must be noted that, as used in the specification, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Overview of Multiple Tier Analysis
[0035] Disclosed herein is a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints. For example, methods disclosed herein are useful for detecting circulating tumor DNA from samples obtained from a subject across two or more timepoints. Determining the change in circulating tumor DNA from samples obtained from the subject across two or more timepoints enables tracking of the tumor heterogeneity. In various embodiments, tracking tumor heterogeneity is informative for determining whether an intervention (e.g., a tumor therapeutic) is efficacious. Therefore, tracking tumor heterogeneity can be useful for e.g., guided therapy.
[0036] In various embodiments, the tiered, multipart method involves performing a first analysis of nucleic acid sequence information that was derived from a first assay performed on a biological sample obtained from the subject. This first analysis identifies whether the biological sample is at risk or not at risk of containing circulating tumor DNA. In various embodiments, for a biological sample that is determined as not negative for containing circulating tumor DNA, the multipart method further includes performing an intra-individual analysis and a second analysis. In various embodiments, the intra-individual analysis includes obtaining target nucleic acids and reference nucleic acids from the biological sample or an additional biological sample obtained from the individual; processing the target nucleic acids and reference nucleic acids to generate a dataset comprising methylation information from the target nucleic acids and methylation information from the reference nucleic acids; and using a computer processor, combining the methylation information from the target nucleic acids and the methylation information from the reference nucleic acids to generate background-corrected methylation information for the target nucleic acids. Here, the background-corrected methylation information is more informative for determining presence or absence of cancer within the individual. In various embodiments, performing the second analysis comprises analyzing the background-corrected methylation information to detect the presence of the circulating tumor DNA in the biological sample. By detecting presence of circulating tumor DNA in the biological sample, the individual can be identified as having cancer.
[0037] Generally, multi-tier testing methodologies described herein achieve significant improvements in comparison to conventional testing methodologies (e.g., single tier testing methodologies). For example, the multi-tier testing methodologies described herein achieve improved performance metrics (e.g., sensitivity, specificity, positive predictive value (PPV), and / or negative predictive value (NPV)) in comparison to conventional methodologies. In particular embodiments, the combination of a first tier and a second tier testing achieves improved specificity (e.g., true negative rate reported as a proportion of correctly identified negatives) in comparison to conventional methodologies.
[0038] In some scenarios, the multi-tier testing methodologies described herein rapidly and accurately screen out a large proportion of individuals in a first tier through a more efficient, lower cost tier 1 test, followed by a more rigorous tier 2 test on the remaining subpopulation of patients. Here, the multi-tier testing methodology can achieve overall performance metrics that are comparable to or not substantially less than the overall performance metrics of conventional methodologies. Altogether, by rapidly and accurately screening out a large proportion of individuals in a first tier, only a small number of individuals undergo the more rigorous tier 2 testing. This represents an improvement in comparison to conventional methodologies that attempt to apply rigorous tests across the entire population, which requires substantial resources. Thus, even in scenarios where the multi-tier testing methodologies achieve performance metrics comparable to those of conventional methodologies, the multi-tier testing methodologies deliver improved performance as a function of resource consumption. Examples of resource consumption include time resources, monetary resources, resources of consumable goods (e.g., consumable assay reagents). In various embodiments, the multi-tier testing methodologies disclosed herein achieve at least a 10% reduction in resource consumption in comparison to a corresponding single-tier test. In various embodiments, the multi-tier testing methodologies disclosed herein achieve at least a 20% reduction, at least a 30% reduction, at least a 40% reduction, at least a 50% reduction, at least a 60% reduction, at least a 70% reduction, at least a 80% reduction, or at least a 90% reduction in resource consumption in comparison to a corresponding singletier test. In various embodiments, the multi-tier testing methodologies disclosed herein achieve at least a 60% reduction in resource consumption in comparison to a corresponding single-tier test. In particular embodiments, the multiple-tiered process disclosed herein is useful for detecting rare or low incidence cancers. For example, the rare or low incidence cancers may have an incidence rate of 1 in 100, 1 in 1,000, 1 in 10,000 individuals, 1 in 100,000 individuals, 1 in 1,000,000 individuals, 1 in 10,000,000 individuals, 1 in 100,000,000 individuals or 1 in 1,000,000,000 individuals. Therefore, the disclosed multipletiered process represents a significant improvement over current methodologies that suffer from poor specificity or sensitivity which contributes to their inability to detect rare or low incidence conditions with sufficient positive predictive value.
[0039] In various embodiments, subjects that were not screened out in the first tier further undergo subsequent analysis to track tumor heterogeneity. For example, the intra-individual analysis may be performed again to analyze a second sample obtained from the same subject at a second timepoint. Here, the second timepoint is subsequent to a first timepoint when the first sample was obtained. Performing the intra-individual analysis using the second sample generates background-corrected methylation information for the second sample. Therefore, by comparing the background-corrected methylation information of the first sample to the background-corrected methylation information of the second sample, a change in the background-corrected methylation information across the two samples is generated. Here, the change in the background-corrected methylation information across the two samples is informative for the change in tumor heterogeneity across the two timepoints from when the two samples were respectively obtained.
[0040] Figure (FIG.) 1A depicts an overall flow process 100 of the multiple-tiered process for tracking tumor heterogeneity, in accordance with an embodiment. Although FIG. 1A shows the flow process in relation to a single subject 110, in various embodiments, the flow process can be performed for more than a single subject 110 (e.g., for thousands, millions, tens of millions, or hundreds of millions of individuals).
[0041] FIG. 1A introduces a first sample 115A, an assay 120A, a first tier (e.g., screen 125), an intra-individual analysis 128A, a second sample 115B, an assay 120B, and a second tier (e.g., second analysis 130) of the multiple-tiered analysis. Generally, the second tier involves a more complex molecular test and analysis in comparison to the first tier. In various embodiments, the more complex molecular test of the second tier is more expensive to perform than the simpler molecular test of the first tier. By employing a cheaper and less complex test, the first tier can identify and remove of individuals that are not at risk of cancer. The more complex molecular test and analysis of the second tier enables more accurate identification of the remaining individuals for purposes of tracking tumor heterogeneity. As shown in FIG. 1 A, the method may involve two or more intra-individual analyses performed on different samples. Here, an intra-individual analysis removes baseline biological signatures. For example, the intra-individual analysis can be performed to remove baseline biological signatures in sequencing information (hereafter referred to as “background-corrected information”) prior to the performance of the second tier. Thus, the more complex molecular test of the second tier can be applied to analyze the background-corrected information of two or more intra-individual analyses to more accurately track tumor heterogeneity in a subject.
[0042] Although FIG. 1A shows a first tier and a second tier of a multiple-tiered analysis, in various embodiments, there may be additional tiers for further classifying individuals. In various embodiments, the multiple-tiered analysis includes three or more tiers, includes four or more tiers, includes five or more tiers, includes six or more tiers, includes seven or more tiers, includes eight or more tiers, includes nine or more tiers, or includes ten or more tiers.
[0043] In various embodiments, the combination of the first tier and the second tier enables the ultimate high performance (e.g., high positive predictive value) of the multiple-tier analysis. In various embodiments, the first tier and the second tier interrogate different markers from samples obtained from subjects. This can be beneficial because different markers can provide different information. In some cases, different markers can be informative for different predictions. As an example, the first tier may analyze protein markers from samples obtained from subjects whereas the second tier may analyze sequencing data derived from nucleic acids in the samples obtained from subjects.
[0044] In various embodiments, the first tier and second tier interrogate the same type of markers from samples obtained from subjects, but at different levels of detail. For example, the first tier may involve the analysis of methylation statuses for a limited, pre-selected set of genomic sites. The differential methylation of the limited, pre-selected set of genomic sites is sufficient to enable identification of subjects not at risk of cancer. Additionally, the second tier may involve the analysis of methylation statuses for a larger set of genomic sites. In one scenario, the second tier involves analysis of methylation statuses for the whole genome (e.g., through whole genome bisulfite sequencing). The differential methylation of the larger set of genomic sites enables more accurate tracking of tumor heterogeneity in the remaining subjects. As another example, the first tier may involve the analysis of shallow sequencing data. Here, shallow sequencing data is sufficient to identify and remove subjects who are not at risk or who do not have cancer. The second tier may involve analysis of sequencing data derived from deeper sequencing, which is sufficient to track tumor heterogeneity for subjects who have cancer.
[0045] FIG. 1A introduces a subject 110. One or more samples (e.g., sample 115A and / or sample 115B) are obtained from the subject 110. In various embodiments, a sample is any of a blood sample, a stool sample, a urine sample, a mucous sample, or a saliva sample. In particular embodiments, each sample obtained from the subject 110 is a blood sample. The sample can be obtained by the individual or by a third party, e.g., a medical professional. Examples of medical professionals include physicians, emergency medical technicians, nurses, first responders, psychologists, phlebotomist, medical physics personnel, nurse practitioners, surgeons, dentists, and any other obvious medical professional as would be known to one skilled in the art. In various embodiments, the one or more samples can be obtained from the subject 110 by a reference lab.
[0046] In various embodiments, the sample obtained from the subject is a liquid biopsy sample obtained at a first point in time. In various embodiments, the liquid biopsy sample may include various biomarkers, examples of which include proteins, metabolites, and / or nucleic acids. In particular embodiments, the liquid biopsy sample includes cell-free DNA (cfDNA) fragments. In particular embodiments, the cfDNA fragments include genomic sequences corresponding to CpG islands for which methylation states are informative of the cancer.
[0047] In various embodiments, a plurality of samples are obtained from the subject 110 at a plurality of different points in time. For example, a sample (e.g., sample 115A) can be obtained at a first timepoint and at least a second sample (e.g., sample 115B) can be obtained from the subject 110 at a second timepoint. In such embodiments, the first sample can be used for performing the assay 120A, the screen 125, and the intra-individual analysis 128A. Additionally, the second sample 115B can be used to perform an assay 120B, and a second intra-individual analysis 128B. The second analysis 130 can then be performed using the results from each of the two or more intra-individual analyses (e.g., intra-individual analysis 128A and intra-individual analysis 128B). Obtaining a plurality of liquid biopsy samples from the individual at a plurality of different points in time includes obtaining a number M of liquid biopsy samples, wherein M is one of: 2, 3, 4, ... , N-l, N, wherein N is a positive integer.
[0048] In various embodiments, sample 115A and / or sample 115B may be processed to extract target nucleic acids and reference nucleic acids. In various embodiments, samples can undergo cellular disruption methods (e.g., to obtain genomic DNA) involving chemical methods or mechanical methods. Example chemical methods include osmotic shock, enzymatic digestion, detergents, or alkali treatment. Example mechanical methods include homogenization, ultrasonication or cavitation, pressure cell, or ball mill. In various embodiments, samples can undergo removal of membrane lipids or proteins or nucleic acid purification. Example chemical methods for removing membrane lipids or proteins and methods for nucleic acid purification include guanidine thiocyanate (GuSCN)-phenol-chloroform extraction, alkaline extraction, cesium chloride gradient centrifugation with ethidium bromide, Chelex® extraction, or cetyltrimethylammonium bromide extraction. Example physical methods for removing membrane lipids or proteins and methods for nucleic acid purification include solid-phase extraction methods using any of silica matrices, glass particles, diatomaceous earth, magnetic beads, anion exchange material, or cellulose matrix. Further details of nucleic acid extraction methods are described in Ali et al, Current Nucleic Acid Extraction Methods and Their Implications to Point-of-Care Diagnostics, Biomed Res. Int. 2017; 2017:9306564, which is hereby incorporated by reference in its entirety.
[0049] Assay 120A and / or assay 120B are performed on the obtained sample 115A and 115B, respectively, to generate marker information. An example of marker information can include quantitative levels of a biomarker, such as a protein biomarker, nucleic acid biomarker, metabolite biomarker, that is present in the sample. Another examples of marker information is sequence information for a plurality of genomic sites. In various embodiments, given that the assay 120 may be performed on a large number of samples (e.g., millions of samples) obtained from a large patient population, the assay 120 be a simplified molecular test that generates marker information that can rapidly distinguish between individuals at risk and individuals not at risk for cancer. For example, the marker information can include quantitative levels of a biomarker, such as a protein biomarker, nucleic acid biomarker, metabolite biomarker, that can rapidly guide the identification and removal of individuals not at risk for the cancer. As another example, the marker information can be sequence information for a limited number of genomic sites that are sufficient for identifying individuals who are not at risk for the cancer (e.g., true negatives). In particular embodiments, the sequence information for a plurality of genomic sites includes methylation information, such as methylation statuses for the plurality of genomic sites. In various embodiments, the plurality of genomic sites include a plurality of CpG islands (CGIs) whose differential methylation status may be indicative of risk for the cancer.
[0050] In particular embodiments, assay 120A and / or assay 120B are performed to generate sequence information for target nucleic acids and to generate sequence information for reference nucleic acids. Thus, sequence information of target and reference nucleic acids can be used to perform the intra-individual analysis 128A and / or intra-individual analysis 128B. In particular embodiments, sequence information includes statuses for a plurality of genomic sites, such as epigenetic statuses for a plurality of CpG sites. In various embodiments, epigenetic statuses refer to methylation statuses. In particular embodiments, sequence information of the target nucleic acids and sequence information of the reference nucleic includes statuses for two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more common genomic sites. In particular embodiments, sequence information of the target nucleic acids and sequence information of the reference nucleic each includes statuses for 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 750 or more, 1000 or more, 2000 or more, 3000 or more, 4000 or more, 5000 or more, 6000 or more, 7000 or more, 8000 or more, 9000 or more, 10000 or more, 11000 or more, 12000 or more, 13000 or more, 14000 or more, 15000 or more, 16000 or more, 17000 or more, 18000 or more, 19000 or more, or 20000 or more genomic sites. In particular embodiments, sequence information of the target nucleic acids and sequence information of the reference nucleic each includes statuses for 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 750 or more, 1000 or more, 2000 or more, 3000 or more, 4000 or more, 5000 or more, 6000 or more, 7000 or more, 8000 or more, 9000 or more, 10000 or more, 11000 or more, 12000 or more, 13000 or more, 14000 or more, 15000 or more, 16000 or more, 17000 or more, 18000 or more, 19000 or more, or 20000 or more of the same genomic sites or overlapping genomic sites. In various embodiments, the plurality of genomic sites include a plurality of CpG islands (CGIs) whose differential methylation status may be indicative of a cancer.
[0051] A screen 125 is performed to analyze the marker information generated by the assay 120A. For example, the screen 125 can involve an in silico analysis of the marker information. In various embodiments, the marker information includes quantitative values of biomarkers. Therefore, the screen 125 can identify and remove individuals whose quantitative values of biomarkers indicate that the individuals are not at risk of the cancer. In various embodiments, the marker information is sequence information for a plurality of genomic sites. Therefore, the screen 125 involves deploying a trained machine learning model that analyzes the sequence information for the plurality of genomic sites and predicts whether an individual is at risk for a cancer. If the screen 125 identifies the individual as not at risk for cancer (as indicated in FIG. 1A as “If negative”), then the subject 110 can be reported as not at risk for the cancer. The process can terminate for this subject and therefore, additional resources need not be further devoted to this subject.
[0052] Alternatively, if the screen identifies the subject as at risk for cancer (as indicated in FIG. 1A as “If not negative” following screen 125), then the subject 110 undergoes at least another tier of testing. As shown in FIG. 1 A, an intra-individual analysis 128A and a second analysis 130 can be performed for subjects identified as at risk for cancer. In particular embodiments, a second sample 115B, assay 120B and second intra-individual analysis 128B are performed for the subject after having determined that the subject is not negative based on the results of the screen 125.
[0053] In various embodiments, as shown in FIG. 1 A, the subject 110 receives an intervention 112. In various embodiments, the subject 110 receives the intervention 112 after the screen determines that the subject 110 is not negative for cancer. Thus, the subject 110 may have been selected and provided the intervention to treat for the cancer and / or to reduce the risk for cancer . An example of an intervention 112 is a tumor therapeutic (e.g., a cancer therapeutic, a chemotherapy, and / or a gene therapy).
[0054] Referring to the intra-individual analysis 128A and intra-individual analysis 128B, the analysis is conducted for a specific subject, such as a subject identified via the screen 125 as at risk for the cancer. Therefore, for a particular subject, the intra-individual analysis is performed to remove baseline biological signatures that are present in the subject. Here, the baseline biological signatures are present irrespective of whether the subject has or does not have cancer. These baseline biological signatures would be confounding signals if analyzed to generate predictions for the patient. Thus, performing the intra-individual analysis 128 for individual samples (e.g., sample 115A or sample 115B) eliminates these confounding baseline biological signatures while keeping signatures that are more informative for determining presence or absence of cancer. For example, in processing nucleic acid sequencing information to generate a signal that may be detected, the resulting signal may comprise a mixture of baseline biological signatures (e.g., germline methylation in a patient) that represent a form of background noise and signatures informative of a cancer(e.g., cancer). Such background noise can obscure a signal informative of a cancer. Advantageously, in certain embodiments, methods described herein contemplate subtracting such background noise from a patient’s nucleic acid sequencing information, thereby improving the signal-to-noise ratio of the signal informative of a cancer.
[0055] In contrast to an inter-individual analysis, where, for example, to determine a presence or absence of cancer within a patient, an average of baseline signatures from a group of normal subjects are removed from the nucleic acid sequencing information of the patient, it has been discovered that performing an intra-individual analysis can significantly improve the sensitivity or specificity of detecting a signal informative for determining presence or absence of cancer.
[0056] Generally, the intra-individual analysis 128A or intra-individual analysis 128B involves generating information from at least target nucleic acids and reference nucleic acids from a corresponding sample (e.g., sample 115 A and sample 115B) obtained from the patient. In various embodiments, the intra-individual analysis 128A and intra-individual analysis 128B is performed on sequence information. Such sequence information may be generated by assay 120A and assay 120B, as shown in FIG. 1 A.
[0057] In various embodiments, the intra-individual analysis 128A and intra-individual analysis 128B involve combining information from target nucleic acids and the reference nucleic acids to generate a signal informative for determining presence or absence of cancer within the patient. By combining the information from the target nucleic acids and the reference nucleic acids, the generated signal can be more informative of presence or absence of a cancer in comparison to a signal derived from the target nucleic acids alone. For example, the information from the reference nucleic acids can represent baseline biology of the patient. By combining the information from the target nucleic acids and the reference nucleic acids, the baseline biology of the patient, which may not be informative for the presence or absence of a cancer, is removed from the generated signal. Thus, information of the target nucleic acids that are not attributable to the patient’s baseline biology remains and is included in the generated signal for determining presence or absence of cancer in the patient.
[0058] Referring next to the second analysis 130, the second analysis 130 is implemented to determine a change in tumor heterogeneity 135 in the subject 110. In various embodiments, the second analysis 130 determines a change in signal between a first set of background-corrected methylation information generated from the first intra-individual analysis 128A and a second set of background-corrected methylation information generated from the second intra-individual analysis 128B. For example, as shown in FIG. 1A, the output of each of the intra-individual analysis 128A and intra-individual analysis 128B can be combined to determine the change in signal. The change in signal can be provided for the second analysis 130 and can be indicative of whether the tumor heterogeneity in the subject is increasing, decreasing, or remaining stable.
[0059] Referring next to FIG. IB, it depicts an overall flow process of the multiple-tiered process for tracking tumor heterogeneity, in accordance with a second embodiment. Here, FIG. IB differs from FIG. 1A in that the second analysis 130 is individually performed to analyze the results of each respective intra-individual analysis e.g., intra-individual analysis 128A and intra-individual analysis 128B. Therefore, as shown in FIG. IB, the output of the second analysis 130A can be combined with the output of second analysis 130B to determine a change in tumor heterogeneity 135 for the subject 110.
[0060] Altogether, the multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128, and second analysis 130) enables the rapid identification of a large proportion of individuals (e.g., greater than 80% of the patient population) representing true negatives, and further enables the accurate identification and diagnosis of a subset of the population representing true positives. The overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multipletiered analysis involving each of the screen 125, intra-individual analysis 128A, intraindividual analysis 128B, and second analysis 130) achieves one or more performance metrics, such as metrics of sensitivity, specificity, positive predictive value (PPV), and / or negative predictive value (NPV). Sensitivity is the true positive rate, reported as a proportion of correctly identified positives. Specificity is the true negative rate reported as a proportion of correctly identified negatives. Positive predictive value refers to the number of true positives divided by the sum of true positives and false positives. Negative predictive value refers to the true negative rate divided by the sum of true negatives and false negatives.
[0061] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% sensitivity in detecting presence of a cancer. In various embodiments, the overall multiple-tiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 70% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 71% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 72% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 73% sensitivity. In particular embodiments, the overall multipletiered analysis achieves at least 74% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 75% sensitivity.
[0062] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% specificity in excluding individuals without the cancer. In various embodiments, the overall multiple-tiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% specificity. In particular embodiments, the overall multiple-tiered analysis achieves at least 99% specificity. In particular embodiments, the overall multiple-tiered analysis achieves at least 99.5% specificity. In particular embodiments, the overall multipletiered analysis achieves at least 99.9% specificity.
[0063] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves a particular sensitivity and a particular specificity. The combination of the sensitivity and specificity limits both the number of false positives and the number of false negatives. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 75% to 89% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 80% to 88% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 83% to 87% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 84% to 86% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves about 85% sensitivity and between 90% to 100% specificity.
[0064] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves between 70% to 90% sensitivity and between 91% to 99% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 92% to 98% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 93% to 97% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 97% to 96% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and about 95% specificity.
[0065] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves between 75% to 89% sensitivity and between 91% to 99% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 80% to 88% sensitivity and between 92% to 98% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 83% to 87% sensitivity and between 93% to 97% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 84% to 86% sensitivity and between 94% to 96% specificity. In various embodiments, the overall multiple-tiered analysis achieves about 85% sensitivity and about 95% specificity.
[0066] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% positive predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 20% positive predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, or at least 40% positive predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 40% positive predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, or at least 60% positive predictive value. In various embodiments, the overall multipletiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 80% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 81% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 82% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 83% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 84% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 85% positive predictive value.
[0067] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% negative predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% negative predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 98% negative predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 99% negative predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 99.4% negative predictive value. System Environment Overview
[0068] FIG. IC depicts an overall system environment 150 including a tumor heterogeneity system 170, in accordance with an embodiment. The overall system environment 150 includes a tumor heterogeneity system 170 for at least performing one or more steps shown in FIG. 1A, and one or more third party entities 155A and 155B in communication with one another through a network 160. FIG. IB depicts one embodiment of the overall system environment 150 in which two third party entities 155A and 155B are involved. In other embodiments, additional or fewer third party entities 155 in communication with the tumor heterogeneity system 170 can be included. The third party entities 155 may communicate with the tumor heterogeneity system 170 to enable the tumor heterogeneity system 170 to perform a screen, one or more intra-individual analyses, and / or second analysis. Third Party Entity
[0069] A third party entity 155 represents a partner entity of the tumor heterogeneity system 170 that can operate upstream, downstream, or both upstream and downstream of the operations of the tumor heterogeneity system 170. As one example, the third party entity 155 operates upstream of the tumor heterogeneity system 170 and provides samples obtained from patients to the tumor heterogeneity system 170. Thus, the tumor heterogeneity system 170 can perform assays, a screen, one or more intra-individual analyses, and / or a second analysis to track tumor heterogeneity of subjects. As another example, the third party entity 155 may process samples obtained from subjects by performing one or more assays on the samples to generate data. Thus, the third party entity 155 can provide the data derived from the assays to the tumor heterogeneity system 170 such that the tumor heterogeneity system 170 can perform a screen, one or more intra-individual analyses, and / or second analysis.
[0070] As another example, the third party entity 155 operates downstream of the tumor heterogeneity system 170. In this scenario, the tumor heterogeneity system 170 may perform a screen and determine whether a subject is at risk for cancer. The tumor heterogeneity system 170 can provide an indication to the third party entity 155 that identifies the subject at risk for the cancer. The third party entity 155 may notify the subject regarding a follow-up appointment such that an additional sample (e.g., sample 115B shown in FIG. 1 A) can be obtained from the subject at the follow-up appointment for subsequent analysis. Network
[0071] This disclosure contemplates any suitable network 160 that enables connection between the tumor heterogeneity system 170 and third party entities 155. The network 160 may comprise any combination of local area and / or wide area networks, using both wired and / or wireless communication systems. In one embodiment, the network 160 uses standard communications technologies and / or protocols. For example, the network 160 includes communication links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, code division multiple access (CDMA), digital subscriber line (DSL), etc. Examples of networking protocols used for communicating via the network 160 include multiprotocol label switching (MPLS), transmission control protocol / Internet protocol (TCP / IP), hypertext transport protocol (HTTP), simple mail transfer protocol (SMTP), and file transfer protocol (FTP). Data exchanged over the network 160 may be represented using any suitable format, such as hypertext markup language (HTML) or extensible markup language (XML). In some embodiments, all or some of the communication links of the network 160 may be encrypted using any suitable technique or techniques. Tumor Heterogeneity System
[0072] FIG. 2A depicts a block diagram of the tumor heterogeneity system 170, in accordance with an embodiment. The block diagram of the tumor heterogeneity system 170 is introduced to show an embodiment in which the tumor heterogeneity system 170 includes one or more assay apparatuses 205 communicatively coupled to a computational system 202. The computational system 202 can further include computational modules, such as a screen module 210, intra-individual analysis module 215, second analysis module 220, and a tumor tracking module 230. The computational system 202 can further include data stores such as a machine learning model store 240 for storing one or more trained machine learning models. FIG. 2A depicts an embodiment in which the tumor heterogeneity system 170 performs one or more assays (e.g., assay 120A or 120B described in FIG. 1 A), performs the screen (e.g., screen 125 described in FIG. 1 A), performs the one or more intra-individual analyses (e.g., intra-individual analysis 128A and / or intra-individual analysis 128B described in FIG. 1A), and performs the second analysis (e.g., second analysis 130 described in FIG. 1 A).
[0073] In various embodiments, the tumor heterogeneity system 170 may be differently configured than shown in FIG. 2A. For example, although the tumor heterogeneity system 170 shown in FIG. 2A includes three different assay apparatuses 205, in various embodiments, the tumor heterogeneity system 170 includes fewer or additional assay apparatuses. In various embodiments, the tumor heterogeneity system 170 does not include an assay apparatus. In such embodiments, the tumor heterogeneity system 170 includes only the computational system 202. In these embodiments in which the tumor heterogeneity system 170 does not include an assay apparatus, the tumor heterogeneity system 170 may perform the screen (e.g., screen 125 described in FIG. 1 A), one or more intra-individual analyses (e.g., intra-individual analysis 128A and / or intra-individual analysis 128B described in FIG. 1 A), and the second analysis (e.g., second analysis 130 described in FIG. 1 A). However, the tumor heterogeneity system 170 does not perform an assay. The assay apparatus 205 may be operated and used by a different entity, such as a third party entity (e.g., third party entity 155 described in FIG. IC). Thus, the third party entity can perform assays using one or more assay apparatus 205 and then transmits the data generated from the assays to the tumor heterogeneity system 170 for performing the screen and / or second analysis. Assays
[0074] Methods disclosed herein involve performing an assay to generate marker information. Assays described in this section can refer to either assay 120A, assay 120B, or both assay 120A and assay 120B shown in FIGs. 1A and IB. Referring to FIG. 2A, performing an assay can involve employing one or more assay apparatuses 205 to perform the assay. In various embodiments, marker information refers to quantitative values of biomarkers, such as protein biomarkers, nucleic acid biomarkers, or metabolite biomarkers. Thus, the quantitative values of biomarkers in a sample can be used to determine whether the individual is at risk for a cancer. In various embodiments, to determine quantitative values of protein biomarkers, performing an assay can include performing one or more of an immunoassay, a protein-binding assay, an antibody-based assay, an antigen-binding proteinbased assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), or a Western blot. To determine quantitative values of nucleic acid biomarkers, performing an assay can include performing one or more of quantitative PCR (qPCR) or digital PCR (dPCR). To determine quantitative values of metabolites, performing an assay can include performing NMR, mass spectrometry, LC-MS, or UPLC-MS / MS.
[0075] In various embodiments, marker information refers to sequence information for a plurality of genomic sites. The sequence information can then be analyzed to generate a prediction for an individual (e.g., whether an individual is negative for cancer or whether the individual is not negative for cancer). In particular embodiments, performing the assay results in generation of methylation sequence information. Methylation sequence information includes methylation statuses for a plurality of genomic sites. In various embodiments, the plurality of genomic sites are previously identified and selected. For example, the plurality of genomic sites may be one or more CpG sites whose differential methylation are informative for determining whether an individual is at risk for a cancer. A CpG site is portion of a genome that has cytosine and guanine separated by only one phosphate group and is often denoted as “5'—C—phosphate—G—3'”, or “CpG” for short. Regions with a high frequency of CpG sites are commonly referred to as “CG islands” or “CGIs”. It has been found that certain CGIs and certain features of certain CGIs in tumor cells tend to be different from the same CGIs or features of the CGIs in healthy cells. Herein, such CGIs and features of the genome are referred to herein as “cancer informative CGIs.”
[0076] Reference is made to FIG. 3 A, which depicts example methylation information useful for determining whether an individual is at risk for a cancer, in accordance with an embodiment. Specifically, FIG. 3A shows that across various types of cancers (e.g., bladder, cervical, colorectal, endometrial, gastric, lung, ovarian, and prostate cancers), sub-regions within a particular CGI can exhibit differential methylation in comparison to normal plasma. Thus, FIG. 3 A depicts an example cancer informative CGI such that performing the assay results in the generation of methylation sequence information corresponding to the cancer informative CGI.
[0077] In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites includes the steps of processing nucleic acids of a sample, enriching the processed nucleic acids for pre-selected genomic sequences (e.g., pre-selected informative CGIs), amplifying the genomic sequences to generate amplicons, and quantifying the amplicons including the genomic sequences (e.g., via sequencing or via quantitative methods such as an ELISA, quantitative PCR, or DNA or RNA-based assay). In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites involves a subset of the previously mentioned steps. For example, enriching the processed nucleic acids can be omitted. Therefore, performing an assay may include processing nucleic acids of a sample, amplifying the pre-selected genomic sequences, and quantifying the amplicons including the genomic sequences.
[0078] Referring again to FIG. 1A or IB, in various embodiments, assay 120A and assay 120B may both involve performing steps of processing nucleic acids of a sample, enriching the processed nucleic acids for pre-selected genomic sequences (e.g., pre-selected informative CGIs), amplifying the genomic sequences to generate amplicons, and quantifying the amplicons including the genomic sequences. In various embodiments, assay 120A and assay 120B involve quantifying the amplicons by performing an ELISA assay, by performing quantitative PCR, or by performing next generation sequencing.
[0079] A methylated nucleic acid is a nucleic acid having a modification in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5-methylcytosine. Methylation can occur at dinucleotides of cytosine and guanine referred to herein as “CpG sites”, which can be a target for enrichment. Methylation of cytosine can occur in cytosines in other sequence contexts, for example, 5'-CHG-3' and 5'-CHH-3', where H is adenine, cytosine or thymine. Cytosine methylation can also be in the form of 5-hydroxymethylcytosine. Methylation of DNA can include methylation of non-cytosine nucleotides, such as A6-methyladenine (6mA). Anomalous cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which may be indicative of cancer status. As is well known in the art, DNA methylation anomalies (compared to healthy controls) can cause different effects, which may contribute to cancer.
[0080] In certain embodiments, the nucleic acid comprises a CpG site (ie., cytosine and guanine separated by only one phosphate group). In certain embodiments, the nucleic acid comprises a CpG island (also referred to as a “CG islands” or “CGI”) or a portion thereof, which is the target for enrichment. Because certain CGIs and certain features of certain CGIs in tumor cells tend to be different from the same CGIs or features of the CGIs in healthy cells, detection of such CGIs can be informative of a cancer. In certain embodiments, the CGI is a “cancer informative CGIs”, which is defined and described in more detail below. In certain embodiments, the CpG is an “informative CpG”, e.g., a “cancer informative CGI”. Such CGIs may have methylation patterns in tumor cells that are different from the methylation patterns in healthy cells. Accordingly, detection of a cancer informative CGI can be informative regarding a subject’s risk of developing cancer or can be indicative that the subject has cancer. Exemplary cancer informative CGIs, which can be target sequences as described herein, are identified in, e.g., Table 1 of U.S. Patent Publication 2020 / 0109456A1, Tables 2 and 3 of WO2022 / 133315, and Tables 1-4 provided herein.
[0081] In certain aspects, the nucleic acids have been treated to convert one or more unmethylated nucleotides (e.g., cytosines) to another nucleotide (a “converted nucleotide”, as used herein, such as a uracil), for example, prior to amplification. Example conversions include bisulfite conversion, enzymatic conversion, or nitrite conversion, further details of which are described herein. In certain embodiments, one or more unmethylated cytosines are converted to a nucleotide that pairs with adenine (e.g., the unmethylated cytosine may be converted to uracil). In certain embodiments, one or more unmethylated adenines are converted to a base that pairs with cytosine (e.g., the unmethylated adenine may be converted to inosine (I)). In certain embodiments, one or more methylated cytosines (e.g., a 5-methylcytosine (5mC)) is converted to a thymine, which pairs with adenine. In certain embodiments, methylated cytosines are protected from conversion (e.g., deamination) during the conversion step.
[0082] In various embodiments, nucleic acids undergo a bisulfite conversion. Bisulfite conversion is performed on DNA by denaturation using high heat, preferential deamination (at an acidic pH) of unmethylated cytosines, which are then converted to uracil by desulfonation (at an alkaline pH). Methylated cytosines remain unchanged on the singlestranded DNA (ssDNA) product.
[0083] In some embodiments the methods include treatment of the sample with bisulfite (e.g., sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium metabisulfite, potassium metabisulfite, ammonium metabisulfite, magnesium metabisulfite and the like). Unmethylated cytosine is converted to uracil through a three-step process during sodium bisulfite modification. As shown in FIG. 2B, the steps are sulphonation to convert cytosine to cytosine sulphonate, deamination to convert cytosine sulphonate to uracil sulphonate and alkali desulphonation to convert uracil sulphonate to uracil. Conversion on methylated cytosine is much slower and is not observed at significant levels in a 4-16 hour reaction. (See Clark et al.. Nucleic Acids Res., 22(15):2990-7 (1994).) If the cytosine is methylated it will remain a methylated cytosine. If the cytosine is unmethylated it will be converted to uracil. When the modified strand is copied, for example, through extension of a locus specific primer, a random or degenerate primer or a primer to an adaptor, a G will be incorporated in the interrogation position (opposite the C being interrogated) if the C was methylated and an A will be incorporated in the interrogation position if the C was unmethylated and converted to U. When the double stranded extension product is amplified those Cs that were converted to Us and resulted in incorporation of A in the extended primer will be replaced by Ts during amplification. Those Cs that were not converted (i.e., the methylated Cs) and resulted in the incorporation of G will be replaced by unmethylated Cs during amplification.
[0084] In various embodiments, nucleic acids undergo an enzymatic conversion. In certain embodiments, the enzymatic treatment with a cytidine deaminase enzyme is used to convert cytosine to uracil. Enzymatic conversion can include an oxidation step, in which Tet methylcytosine dioxygenase 2 (TET2) catalyzes the oxidation of 5mC to 5hmC to protect methylated cytosines from conversion by subsequent exposure to a cytidine deaminase. Other protection steps known in the art can be used in addition to or in place of oxidation by TET2. After the oxidation step, the nucleic acid is treated with the cytidine deaminase to convert one or more unmethylated cytosines to uracils. As with bisulfite conversion, when the modified strand is copied, a G will be incorporated in the interrogation position (opposite the C being interrogated) if the C was methylated and an A will be incorporated in the interrogation position if the C was unmethylated. When the double stranded extension product is amplified those Cs that were converted to Us and resulted in incorporation of A in the extended primer will be replaced by Ts during amplification. Those Cs that were not modified and resulted in the incorporation of G will remain as C.
[0085] In certain embodiments the cytidine deaminase may be APOBEC. In certain embodiments the cytidine deaminase includes activation induced cytidine deaminase (AID) and apolipoprotein B mRNA editing enzymes, catalytic polypeptide-like (APOBEC). In certain embodiments, the APOBEC enzyme is selected from the human APOBEC family consisting of APOBEC-1 (Apol), APOBEC-2 (Apo2), AID, APOBEC-3 A, -3B, -3C, -3DE, -3F, -3G, -3H and APOBEC-4 (Apo4). In certain embodiments, the APOBEC enzyme is APOBEC-seq.
[0086] In certain embodiments, nitrite treatment is used to deaminate adenine and cytosine. As shown in FIG. 2C, deamination of an A results in conversion to an inosine (I), which is read by a polymerase as a G, whereas deamination of a methylated A (A6-methyladenine (6mA)) results in a nitrosylated 6mA (6mA-N0), which causes the base to be read by a polymerase as an A. Deamination of a C results in conversion to a uracil, which is read by a polymerase as a T, whereas deamination of a A^-methylcytosine (4mC) to 4mC-N0 or a 5-methylcytosine (5mC) to a T causes the base to be read by a polymerase as a C or a T, respectively. For 5mC bases, the C to T ratio at the 5mC position is about 40% higher than other cytosine positions, allowing 5mC to be differentiated from C. (See, Li et al. (2022) Genome Biology 23:122.)
[0087] In various embodiments, performing the assay includes enriching for specific genomic sequences, such as genomic sequences of pre-selected CGIs. In various embodiments, enrichment of pre-selected CGIs can be accomplished via hybrid capture. Examples of such hybrid capture probe sets include the KAPA HyperPrep Kit and SeqCAP Epi Enrichment System from Roche Diagnostics (Pleasanton, CA). For example, hybrid capture probe sets can be designed to target (e.g., hybridize with) selected genomic sequences, thereby capturing and enriching the selected genomic sequences.
[0088] In various embodiments, performing the assay includes a step of nucleic acid amplification. During amplification, the converted nucleotide pairs with its complementary nucleotide, and in the next round of amplification, the complementary nucleotide pairs with a replacement nucleotide. For example, following the conversion of an unmethylated cytosine to a uracil, the nucleic acid may be amplified such that an adenine pairs with the uracil in the first round of replication, and in the second round of replication, the adenine pairs with a thymine. Accordingly, the thymine replaces the uracil in the original nucleic acid sequence, and is referred to herein as a “replacement nucleotide”.
[0089] Examples of such assays include, but are not limited to performing PCR assays, Realtime PCR assays, Quantitative real-time PCR (qPCR) assays, digital PCR (dPCR), Allelespecific PCR assays, Reverse-transcription PCR assays and reporter assays. For example, given the processed nucleic acids (e.g., bisulfite converted nucleic acids) that are enriched for pre-selected genomic sequences, a PCR assay is performed to amplify the pre-selected genomic sequences to generate amplicons. Here, PCR primers are added to initiate the amplification. In various embodiments, the PCR primers are whole genome primers that enable whole genome amplification. In various embodiments, the PCR primers are genespecific primers that result in amplification of sequences of specific genes. In various embodiments, the PCR primers are allele-specific primers. For example, allele specific primers can target a genomic sequence corresponding to a pre-selected CGI, such that performing nucleic acid amplification results in amplification of the genomic sequence of the pre-selected CGI.
[0090] In various embodiments, performing the assay includes quantifying the nucleic acids including the pre-selected genomic sequences (e.g., informative CGIs). In some embodiments, quantifying the nucleic acids to generate sequence information comprises performing an enzyme-linked immunosorbent assay (ELISA). In some embodiments, quantifying the nucleic acids to generate sequence information comprises performing quantitative PCR (qPCR) or digital PCR (dPCR). Therefore, the number of methylated, unmethylated, or partially methylated pre-selected genomic sequences can be quantified.
[0091] In various embodiments, quantifying the nucleic acids comprises sequencing the nucleic acids including the pre-selected genomic sequences. Thus, the sequenced reads can be aligned to a reference library and methylation sequence information including methylation statuses of the informative CGIs can be determined. Therefore, the number of methylated, unmethylated, or partially methylated pre-selected genomic sequences can be quantified via the sequenced reads.
[0092] FIG. 3B shows an example flow process for determining whether an individual is at risk for a cancer, in accordance with an embodiment. Here, specific genomic regions of an indexed library of nucleic acids (e.g., DNA) are targeted. For example, locus 1 can refer to a reference genomic location. Here, a reference genomic location serves as a control. For example, the reference genomic location is not differentially methylated in healthy individuals in comparison to individuals with the cancer. Locus 2 can refer to a pre-selected genomic location, such as a pre-selected informative CGI.
[0093] Performing the assay further includes performing nucleic acid amplification (e.g., PCR) to generate marker information. In various embodiments, nucleic acid amplification includes either qPCR or dPCR. This quantifies the number of methylated, unmethylated, or partially methylated sequences at locus 1 (reference) and at locus 2. In various embodiments, performing the assay includes performing an ELISA to quantify the number of methylated, unmethylated, or partially methylated sequences at locus 1 (reference) and at locus 2. Assays for Generating Sequencing Information for Performing Intra-Individual Analysis
[0094] In particular embodiments, assays disclosed herein (e.g., assay 120A or 120B shown in FIGs. 1A-1B) are useful for generating sequencing information for performing an intraindividual analysis (e.g., one or both of intra-individual analysis 128A and intra-individual analysis 128B shown in FIGs. 1A-1B). For example, an assay is performed to generate sequence information for target nucleic acids and / or reference nucleic acids.
[0095] In various embodiments, sequence information of target nucleic acids and / or sequence information of reference nucleic acids refer to statuses for a plurality of genomic sites. Sequence information of target nucleic acids refers to epigenetic statuses (e.g., methylation statuses) across a plurality of genomic sites in the target nucleic acids. Sequence information of reference nucleic acids refers to epigenetic statuses (e.g., methylation statuses) across a plurality of genomic sites in the reference nucleic acids. In various embodiments, the plurality of genomic sites are previously identified and selected. For example, the plurality of genomic sites may be one or more CpG sites whose differential methylation are informative for determining whether an individual has a cancer. A CpG site is portion of a genome that has cytosine and guanine separated by only one phosphate group and is often denoted as “5'—C—phosphate—G—3'”, or “CpG” for short. Regions with a high frequency of CpG sites are commonly referred to as “CG islands” or “CGIs”. It has been found that certain CGIs and certain features of certain CGIs in tumor cells tend to be different from the same CGIs or features of the CGIs in healthy cells. Herein, such CGIs and features of the genome are referred to herein as “cancer informative CGIs.” Cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. Example CGIs include, but are not limited to, the CGIs shown in the accompanying tables (referred to herein as Tables 1-4) which lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.
[0096] In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites includes the steps of processing nucleic acids of a sample, enriching the processed nucleic acids for pre-selected genomic sequences (e.g., pre-selected informative CGIs), amplifying the genomic sequences to generate amplicons, and quantifying the amplicons including the genomic sequences (e.g., via sequencing such as next generation sequencing or via quantitative methods such as an ELISA, quantitative PCR, allele-specific PCR, or DNA or RNA-based assay). In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites involves a subset of the previously mentioned steps. For example, enriching the processed nucleic acids can be omitted. Therefore, performing an assay may include processing nucleic acids of a sample, amplifying the pre-selected genomic sequences, and quantifying the amplicons including the genomic sequences.
[0097] In various embodiments, performing an assay (e.g., assay 120A or assay 120B) involves processing target nucleic acids and / or reference nucleic acids. In various embodiments, processing target nucleic acids and / or reference nucleic acids to capture methylation modifications includes performing a nucleic acid conversion (e.g., any of bisulfite conversion, enzymatic conversion, or nitrite conversion). In various embodiments, processing target nucleic acids and / or reference nucleic acids to capture methylation modifications includes performing any of nucleic acid amplification, polymerase chain reaction (PCR), methylation specific PCR, bisulfite pyrosequencing, single-strand conformation polymorphism (SSCP) analysis, methylation-sensitive single-strand conformation analysis restriction analysis, high resolution melting analysis, methylationsensitive single-nucleotide primer extension, restriction analysis, microarray technology, next generation methylation sequencing, nanopore sequencing, and combinations thereof.
[0098] In various embodiments, performing the assay includes enriching for specific sequences in the target nucleic acids and / or reference nucleic acids. In various embodiments, the specific sequences refer to sequences of pre-selected CGIs. In various embodiments, enrichment of pre-selected CGIs can be accomplished via hybrid capture. Examples of such hybrid capture probe sets include the KAPA HyperPrep Kit and SeqCAP Epi Enrichment System from Roche Diagnostics (Pleasanton, CA). For example, hybrid capture probe sets can be designed to hybridize with particular sequences of the target nucleic acids and / or reference nucleic acids, thereby capturing and enriching the particular sequences.
[0099] In various embodiments, performing the assay includes performing nucleic acid amplification to amplify the particular sequences of the target nucleic acids and / or reference nucleic acids. Examples of such assays include, but are not limited to performing PCR assays, Real-time PCR assays, Quantitative real-time PCR (qPCR) assays, digital PCR (dPCR), Allele-specific PCR assays, Reverse-transcription PCR assays and reporter assays. For example, given the processed nucleic acids (e.g., bisulfite converted nucleic acids) that are enriched for pre-selected sequences, a PCR assay is performed to amplify the pre-selected sequences to generate amplicons. Here, PCR primers are added to initiate the amplification. In various embodiments, the PCR primers are whole genome primers that enable whole genome amplification. In various embodiments, the PCR primers are gene-specific primers that result in amplification of sequences of specific genes. In various embodiments, the PCR primers are allele-specific primers. For example, allele specific primers can target a genomic sequence corresponding to a pre-selected CGI, such that performing nucleic acid amplification results in amplification of the sequence of the pre-selected CGI.
[00100] In various embodiments, performing the assay includes quantifying the nucleic acids including the pre-selected sequences (e.g., informative CGIs). In some embodiments, quantifying the nucleic acids to generate sequence information comprises performing any of real-time PCR assay, quantitative real-time PCR (qPCR) assay, digital PCR (dPCR) assay, allele-specific PCR assay, or reverse-transcription PCR assay. Therefore, the number of methylated, hypermethylated, unmethylated, or partially methylated pre-selected sequences are quantified.
[00101] In various embodiments, quantifying the nucleic acids comprises sequencing the nucleic acids including the pre-selected sequences. Thus, the sequenced reads are aligned to a reference library and sequence information including methylation statuses of the informative CGIs of amplicons derived from the target nucleic acids and / or reference nucleic acids can be determined. Therefore, the number of methylated, hypermethylated, unmethylated, or partially methylated pre-selected sequences of the target nucleic acids and the reference nucleic acids can be quantified via the sequenced reads. Assays for Generating Sequencing Information for Phased Sequencing
[00102] In various embodiments, performing the assay comprises sequencing the target nucleic acids and / or reference nucleic acids. In various embodiments, sequencing comprises performing next generation sequencing methods to generate sequence reads from the target nucleic acids and / or reference nucleic acids. As described herein, sequence reads from reference nucleic acids may be long sequence reads (e.g., greater than 500 bases in length). Generally, long sequence reads include an average read length that is longer than sequence reads obtained through standard sequencing methods. In various embodiments, the long sequence reads of reference nucleic acids refer to sequence reads of at least 500 bases, at least 1 kilobase, at least 2 kilobases (kb), at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, at least 12 kb, at least 15 kb, at least 20 kb, at least 25 kb, at least 30 kb, at least 40 kb, at least 50 kb, at least 60 kb, at least 70 kb, at least 80 kb, at least 90 kb, at least 100 kb, at least 200 kb, at least 300 kb, at least 400 kb, at least 500 kb, at least 600 kb, at least 700 kb, at least 800 kb, at least 900 kb, at least 1000 kb, at least 1500 kb, or at least 2000 kb. In particular embodiments, the long sequence reads of reference nucleic acids refer to sequence reads of between 5 kb and 100 kb, between 10 kb and 80 kb, between 20 kb and 70 kb, between 30 kb and 60 kb, or between 40 kb and 50 kb. In particular embodiments, long sequence reads of reference nucleic acids refer to sequence reads of greater than about 8 kb, greater than about 9 kb or greater than about 10 kb. In particular embodiments, long sequence reads of reference nucleic acids refer to sequence reads between about 10 kb and about 100 kb, or between about 10 kb and about 2 MB. In various embodiments, generating long sequence reads of reference nucleic acids involves performing nanopore sequencing. Methods for long-read sequencing are known in the art and such methods can be performed using, for example, an Oxford Nanopore instrument (e.g., PromethlON™) or Pacific Biosciences Single-Molecule Real-Time (SMRT) sequencing technology.
[00103] In various embodiments, performing the assay includes generating phased sequencing information for target nucleic acids and / or reference nucleic acids. As used herein, “phased sequencing information,” also referred to herein as “haplotype sequencing information,” refers to sequencing information derived specifically from a particular source. For example, phased sequencing information or haplotype sequencing information can refer to sequencing information derived from either the maternal or paternal chromosome. Generally, phased sequencing information of target nucleic acids may be useful for determining presence or absence of a cancer because signals originating from the same source (e.g., maternal or paternal chromosome) may provide additional information in comparison to other approaches that merely analyze signals irrespective of the source.
[00104] In various embodiments, the phased sequencing information comprises mutation sequence information of the cell-free DNA. For example, mutation sequence information can include one or more mutations present across a plurality of genomic sites. In particular embodiments, the mutation sequence information includes one or more mutations that originate from a common source (e.g., a maternal chromosome or a paternal chromosome). Here, two or more genomic sites derived from a common source that have a particular pattern of mutations (e.g., each having a mutation, some pattern of mutated / non-mutated, or all nonmutated) can be referred to as coupled genomic sites. In various embodiments, a mutation can be any of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.
[00105] In various embodiments, the phased sequencing information comprises methylation sequence information of the cell-free DNA. Methylation sequence information can include methylation statuses across a plurality of genomic sites. In particular embodiments, the methylation sequence information includes methylation statuses of genomic sites from a common source (e.g., a maternal chromosome or a paternal chromosome). As a specific example, methylation at a first genomic site may be coupled with methylation at a second genomic site on the same maternal or paternal chromosome. Two or more genomic sites with a particular methylation pattern (e.g., all methylated, partially methylated, or non-methylated) that originate from the same maternal or paternal chromosome is referred to herein as coupled methylation sites. Example coupled methylation sites may be two or more CGIs disclosed herein (e.g., two or more CGIs disclosed in any of Tables 1-4). In various embodiments, two or more genomic sites of coupled methylation sites may be separated by tens, hundreds, or even thousands of bases. Thus, coupled methylation sites include two or more genomic sites from a common source and need not be limited to genomic sites that are close in proximity (e.g., adjacent CpG sites). In various embodiments, coupled methylation sites include 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more methylation sites from a common source. Thus, detecting these coupled methylation sites may provide disease diagnostic utility.
[00106] In various embodiments, generating phased sequencing information for target nucleic acids comprises aligning sequence reads of target nucleic acids to long sequence reads of reference nucleic acids derived from different sources (e.g., either the maternal or paternal chromosome). Long sequence reads of reference nucleic acids originating from different sources can be distinguished due to sequence differences present in the long sequence reads. For example, given a particular chromosome, long sequence reads derived from a maternal chromosome would have sequence differences in comparison to long sequence reads derived from a paternal chromosome. Here, sequence differences can refer to mutations that are present in long sequence reads from one source, but not present in long sequence reads from the second source, and vice versa. Thus, the presence or absence of certain mutations can be useful for distinguishing whether a long sequence read originated from a first source or a second source. Altogether, by comparing sequences of long sequence reads, a first set of long sequence reads with a set of common sequences can be attributed to a first source (e.g., a maternal chromosome) whereas a second set of long sequence reads with a different set of common sequences can be attributed to a second source (e.g., a paternal chromosome). In various embodiments, the different sets of long sequence reads need not specifically be attributed to a maternal chromosome and a paternal chromosome; rather, it is sufficient to distinguish different sets of long sequence reads from a first source and a second source. These long sequence reads from a first source or a second source have sufficiently different sequences to enable phasing of the target nucleic acids (e.g., to determine sources from which target nucleic acids were derived from).
[00107] By aligning sequence reads of target nucleic acids to long sequence reads of reference nucleic acids, the long sequence reads of reference nucleic acids serve as digital guides to phase e.g., determine the source of target nucleic acids. For example, target nucleic acids from a first common source (e.g., from a maternal chromosome) can be categorized together based on sequence similarities between the target nucleic acids and the long sequence reads of reference nucleic acids from the first source. Additionally, target nucleic acids from a second common source (e.g., from a paternal chromosome) can be categorized together based on sequence similarities between the target nucleic acids and the long sequence reads of reference nucleic acids from the second source. In contrast to using the standard human genome to align sequence reads of target nucleic acids, using long reads of reference nucleic acids would enable alignment of reference nucleic acids to sequences of the maternal or paternal chromosome Individual-specific differences between target nucleic acids deriving from the maternal and paternal chromosomes could be used as markers to create haplotype-specific sequence information that is informative for determining presence or absence of a cancer.
[00108] In various embodiments, phased sequencing information includes phased methylation sequencing information of cfDNA, where at least a first set of the phased methylation sequencing information of cfDNA originates from a first source and at least a second set of the phased methylation sequencing information of cfDNA originates from a second source. In various embodiments, methods for generating phased sequencing information can further include comparing the first set of the phased methylation sequencing information of cfDNA from the first source to the second set of the phased methylation sequencing information of cfDNA from the second source. In particular embodiments, generating phased sequencing information further includes comparing methylation statuses of two or more genomic sites from a first source to methylation statuses of the same two or more genomic sites from a second source. Differences in methylation statuses of genomic sites from the first source and the second source can be valuable for inclusion in the signal informative for determining presence or absence of a cancer. For example, if multiple genomic sites from a first source are methylated but the same genomic sites from a second source are unmethylated, this may be an informative signal for presence or absence of a cancer. Screen
[00109] The description in this section pertains to the performance of a screen, such as screen 125 described in FIG. 1 A, which can be performed by the screen module 210 described in FIG. 2A. Generally, a screen is performed on marker information generated by the assay (e.g., assay 120A). In various embodiments, the screen is performed to determine whether a biological sample is at risk or not at risk of containing a signal indicative of a cancer. For example, the screen is performed to determine whether a biological sample is at risk or not at risk of containing circulating tumor DNA. Circulating DNA within the biological sample may indicate that the individual (e.g., individual from whom the biological sample is obtained) may be at risk of a cancer. In various embodiments, the screen is performed to classify the subject as negative for cancer or not negative for cancer.
[00110] In various embodiments, the marker information represents quantified values of biomarkers. For example, depending on the type of biomarker, the quantified values may be generated via one or more of: an immunoassay, a protein-binding assay, an antibody-based assay, an antigen-binding protein-based assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), a Western blot, quantitative PCR (qPCR) or digital PCR (dPCR), NMR, mass spectrometry, LC-MS, or UPLC-MS / MS.
[00111] In various embodiments, performing the screen involves comparing the quantified values of biomarkers to one or more reference values or to threshold values. For example, a reference value can be a statistical measure of quantified biomarker values corresponding to individuals known to be at risk for cancer. Therefore, if the comparison identifies that the quantified values of biomarkers for an individual is statistically significantly different from the reference value corresponding to individuals known to be at risk for cancer, then the screen can identify the cancer as negative for cancer.
[00112] In various embodiments, the marker information represents sequencing information for one or more genomic locations, such as one or more CpG islands. In various embodiments, performing the screen involves comparing methylation information at one or more pre-selected genomic locations to quantified values of reference genomic locations. For example, referring again to FIG. 3B, an assay may have been performed that generates methylation information for locus 1 corresponding to a reference genomic location and for locus 2 corresponding to a pre-selected genomic location (e.g., a pre-selected informative CGI). Thus, the methylation information at locus 1 is compared to methylation information at locus 2. Based on the comparison, the screen can identify the subject as not negative for cancer.
[00113] In various embodiments, the screen can be a cheaper and less complex test in comparison to the second tier analysis (e.g., the second analysis). The screen can analyze marker information at a low resolution for purposes of identifying and removing large proportions of individuals that are not at risk of cancer. In various embodiments, the screen analyzes methylation information across a plurality of genomic locations and determines a measure of overall methylation across the plurality of genomic locations. Here, the measure of overall methylation across the plurality of genomic sites can represent methylation information of low resolution. Specifically, the measure of overall methylation provides a metric for methylation across the plurality of genomic sites, but may not provide information as to methylation status at each individual genomic site. The measure of overall methylation can be sufficient for identifying and removing large proportions of individuals not at risk for cancer. In various embodiments, the overall methylation across the plurality of genomic sites can be a total number of methylated CpG sites. In various embodiments, the overall methylation across the plurality of genomic sites can be a total number of methylated CpG sites across the plurality of genomic sites located in a subset of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the overall methylation across the plurality of genomic sites can be a total number of methylated CpG sites across the plurality of genomic sites located in all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the overall methylation across the plurality of genomic sites can be an average number of methylated CpG sites (e.g., an average number of methylated CpG sites within a target region or a CGI). In various embodiments, the overall methylation across the plurality of genomic sites can be an average number of methylated CpG sites across the plurality of genomic sites located in a subset of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the overall methylation across the plurality of genomic sites can be an average number of methylated CpG sites across the plurality of genomic sites located in all of the CGIs in any one of Tables 1, 2, 3, or 4.
[00114] In various embodiments, performing the screen involves performing whole genome sequencing or whole genome bisulfite sequencing and determining the overall methylation across the whole genome. Thus, in such embodiments, performing the screen is not limited to only analyzing CGIs or portions thereof; rather, performing the screen involves analyzing methylation statuses across the whole genome. In various embodiments, analyzing the methylation statuses across the whole genome can involve determining a quantifiable measure of the overall methylation across the whole genome. In various embodiments, the quantifiable measure of overall methylation is a score, such as a whole genome methylation burden score. In various embodiments, the higher the whole genome methylation burden score, the more likely the biological sample is at risk for containing circulating tumor DNA. In various embodiments, the lower the whole genome methylation burden score, the less likely the biological sample is at risk for containing circulating tumor DNA. In various embodiments, the biological sample is classified as negative (e.g., not at risk for containing circulating tumor DNA) or not negative (e.g., at risk for containing circulating tumor DNA) based on the determined whole genome methylation burden score. For example, if the whole genome methylation burden score for the biological sample is above a threshold score, the biological sample can be classified as not negative. As another example, if the whole genome methylation burden score for the biological sample is below a threshold score, the biological sample can be classified as negative.
[00115] In various embodiments, the measure of overall methylation across one or more preselected genomic locations and methylation information for reference genomic locations can be a cycle threshold (Ct) value. Cycle threshold refers to the number of PCR cycles needed for a sample to amplify and cross a threshold. In various embodiments, if a difference between the Ct value of the methylation sequences of the pre-selected genomic locations and the Ct value of the reference genomic locations is greater than a threshold, then the screen identifies the subject as not negative for cancer. If a difference between the Ct value of the methylation sequences of the pre-selected genomic locations and the Ct value of the reference genomic locations is less than a threshold, then the screen identifies the subject as negative for cancer.
[00116] In various embodiments, a screen is performed on sequence information generated via sequencing (e.g., next generation sequencing) of sequences at the one or more genomic locations, such as one or more CpG islands. In various embodiments, such a screen is performed using a system comprising a computer storage and a processing system. The screen can further involve the implementation of a machine learning model. For example, the computer storage can store sequence information corresponding to a processed sample, the processed sample including cell-free DNA fragments originating from a liquid biopsy of an individual and having been processed to enrich for cancer informative CGIs, the sequencer information comprising, for each sequenced cell-free DNA fragment corresponding to the cancer informative CGIs, a respective position on the genome for the cell-free DNA fragment and methylation information for the cell-free DNA fragment. The processing system can compute values of the cancer informative CGIs for the individual and applies the values as input to a trained machine learning model. The machine learning model provides a predicted output as to whether the individual is at risk for cancer based on the values of the cancer informative CGIs.
[00117] In various embodiments, performing the screen involves analyzing a plurality of CGIs. For example, performing the screen involves analyzing methylation statuses of a plurality of CGIs. Cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. The accompanying tables (e.g., Tables 1-4) lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.
[00118] In some embodiments, performing the screen involves analyzing a plurality of CGIs including one or more CGIs that are methylated in the genome of extraembryonic ectoderm (ExE). Here, such example CGIs may be differentially methylated in the genome of ExE and not methylated in corresponding epiblast or adult tissue. Example CGIs that are methylated in the genome of ExE are further disclosed in Table 3 of WO2022133315, which is hereby incorporated by reference in its entirety.
[00119] In various embodiments, performing the screen involves analyzing all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 1. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 1. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 2. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 2. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 3. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 3. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 4. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 4. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Tables 2 and 3. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Tables 2 and 3.
[00120] In various embodiments, performing the screen involves analyzing 1 CGI, 2 CGIs, 3 CGIs, 4 CGIs, 5 CGIs, 6 CGIs, 7 CGIs, 8 CGIs, 9 CGIs, 10 CGIs, 11 CGIs, 12 CGIs, 13 CGIs, 14 CGIs, 15 CGIs, 16 CGIs, 17 CGIs, 18 CGIs, 19 CGIs, 20 CGIs, 21 CGIs, 22 CGIs, 23 CGIs, 24 CGIs, 25 CGIs, 26 CGIs, 27 CGIs, 28 CGIs, 29 CGIs, 30 CGIs, 31 CGIs, 32 CGIs, 33 CGIs, 34 CGIs, 35 CGIs, 36 CGIs, 37 CGIs, 38 CGIs, 39 CGIs, 40 CGIs, 41 CGIs, 42 CGIs, 43 CGIs, 44 CGIs, 45 CGIs, 46 CGIs, 47 CGIs, 48 CGIs, 49 CGIs, or 50 CGIs (e.g., CGIs as shown in any of Tables 1-4 or portions of CGIs shown in any of Tables 1-4). In various embodiments, performing the screen involves analyzing at most 2 CGIs, at most 5 CGIs, at most 10 CGIs, at most 15 CGIs, at most 20 CGIs, at most 25 CGIs, at most 30 CGIs, at most 35 CGIs, at most 40 CGIs, at most 45 CGIs, or at most 50 CGIs (e.g., CGIs as shown in any of Tables 1-4 or portions of CGIs shown in any of Tables 1-4). In various embodiments, performing the screen involves analyzing at most 50 CGIs, at most 100 CGIs, at most 150 CGIs, at most 200 CGIs, at most 300 CGIs, at most 400 CGIs, at most 500 CGIs, at most 600 CGIs, at most 700 CGIs, at most 800 CGIs, at most 900 CGIs, at most 1000 CGIs, at most 1500 CGIs, at most 2000 CGIs, at most 2500 CGIs, at most 3000 CGIs, at most 3500 CGIs, at most 4000 CGIs, at most 4500 CGIs, at most 5000 CGIs, at most 5500 CGIs, or at most 6000 CGIs (e.g., CGIs as shown in any of Tables 1-4 or portions of CGIs shown in any of Tables 1-4). In particular embodiments, performing the screen involves analyzing at most 500 CGIs.
[00121] In various embodiments, the screen achieves at least 60% sensitivity in detecting presence of a cancer. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sensitivity. In particular embodiments, the screen achieves at least 75% sensitivity. In particular embodiments, the screen achieves at least 76% sensitivity. In particular embodiments, the screen achieves at least 77% sensitivity. In particular embodiments, the screen achieves at least 78% sensitivity. In particular embodiments, the screen achieves at least 79% sensitivity. In particular embodiments, the screen achieves at least 80% sensitivity.
[00122] In various embodiments, the screen achieves at least 60% specificity in excluding individuals without cancer. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% specificity. In particular embodiments, the screen achieves at least 90% specificity. In particular embodiments, the screen achieves at least 91% specificity. In particular embodiments, the screen achieves at least 92% specificity. In particular embodiments, the screen achieves at least 93% specificity. In particular embodiments, the screen achieves at least 94% specificity. In particular embodiments, the screen achieves at least 95% specificity.
[00123] In various embodiments, the screen achieves at least 15% positive predictive value. In various embodiments, the screen achieves at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, or at least 40% positive predictive value. In particular embodiments, the screen achieves at least 20% positive predictive value. In particular embodiments, the screen achieves at least 21% positive predictive value. In particular embodiments, the screen achieves at least 22% positive predictive value. In particular embodiments, the screen achieves at least 23% positive predictive value. In particular embodiments, the screen achieves at least 24% positive predictive value. In particular embodiments, the screen achieves at least 25% positive predictive value. In particular embodiments, the screen achieves at least 26% positive predictive value. In particular embodiments, the screen achieves at least 27% positive predictive value. In particular embodiments, the screen achieves at least 28% positive predictive value. In particular embodiments, the screen achieves at least 29% positive predictive value. In particular embodiments, the screen achieves at least 30% positive predictive value. In particular embodiments, the screen achieves at least 31% positive predictive value. In particular embodiments, the screen achieves at least 32% positive predictive value. In particular embodiments, the screen achieves at least 33% positive predictive value. In particular embodiments, the screen achieves at least 34% positive predictive value. In particular embodiments, the screen achieves at least 35% positive predictive value. In particular embodiments, the screen achieves at least 36% positive predictive value. In particular embodiments, the screen achieves at least 37% positive predictive value. In particular embodiments, the screen achieves at least 38% positive predictive value. In particular embodiments, the screen achieves at least 39% positive predictive value. In particular embodiments, the screen achieves at least 40% positive predictive value.
[00124] In various embodiments, the screen achieves at least 60% negative predictive value. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% negative predictive value. In particular embodiments, the screen achieves at least 95% negative predictive value. In particular embodiments, the screen achieves at least 96% negative predictive value. In particular embodiments, the screen achieves at least 97% negative predictive value. In particular embodiments, the screen achieves at least 98% negative predictive value. In particular embodiments, the screen achieves at least 99% negative predictive value. Intra-Individual Analysis
[00125] The description in this section pertains to the performance of an intra-individual analysis, such as an intra-individual analysis 128A and / or intra-individual analysis 128B described in FIGs. 1A-1B. In general, the intra-individual analyses are conducted for subjects that were previously determined (e.g., via screen 125 as shown in FIG. 1 A) as not negative for cancer. The intra-individual analysis removes baseline biological signatures that are specific for a subject to generate a background-corrected signal. Thus, the second analysis involves analyzing the background-corrected signal to determine whether the individual has cancer.
[00126] In various embodiments, an intra-individual analysis is conducted using a single sample, such as a blood sample. The sample may contain target nucleic acids and reference nucleic acids. Target nucleic acids may include signatures that are informative of determining presence or absence of a cancer, and can further include baseline biological signatures. Here, target nucleic acids in the blood sample may be derived from a diseased cell which is associated with the cancer. For example, target nucleic acids can include cell-free DNA in the blood that originates from a diseased cell. In particular embodiments, target nucleic acids are cell-free DNA in the blood that originates from a cancer cell. Reference nucleic acids in the sample refer to nucleic acids that contain baseline biological signatures of the individual. For example, baseline biological signatures of the individual may be present in nucleic acids irrespective of whether the nucleic acids originate from a diseased source, or a non-diseased source. The baseline biological signatures of the reference nucleic acids are generally less informative for determining presence or absence of a cancer in comparison to the informative signatures present in the target nucleic acids. In various embodiments, reference nucleic acids refer to cellular genomic DNA derived from a healthy cell from the individual. In various embodiments, reference nucleic acids found in the sample derive from a cell in a healthy organ of the individual. Example organs include the brain, heart, thorax, lung, abdomen, colon, cervix, pancreas, kidney, liver, muscle, lymph nodes, esophagus, intestine, spleen, stomach, and gall bladder. In particular embodiments, reference nucleic acids are found in the sample and refer to cellular genomic DNA derived from peripheral blood mononuclear cells (PBMCs) (e.g., lymphocytes or monocytes) or polymorphonuclear cells (e.g., eosinophils or neutrophils).
[00127] In various embodiments, target nucleic acids and reference nucleic acids are separately obtained from the single sample. In various embodiments, the sample is processed to separate the target nucleic acids and reference nucleic acids. For example, the sample may be processed through any one of centrifugation, filtration, gel electrophoresis, bead capture, or matrix extraction. In particular embodiments, target nucleic acids are cell-free nucleic acids and therefore, can be obtained from the supernatant of the separated sample. In particular embodiments, reference nucleic acids are cellular genomic nucleic acids and therefore, can be obtained from a different portion of the separated sample that contains cells.
[00128] Generally, an intra-individual analysis is performed on sequence information of target nucleic acids and sequence information of reference nucleic acids. In particular embodiments, the sequence information of target nucleic acids comprise sequence WO 2025 / 147572 PCT / US2025 / 010181 information of cell free DNA. In particular embodiments, the sequence information of reference nucleic acids comprise sequence information of cells, such as peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells.
[00129] The intra-individual analysis involves combining the sequence information of target nucleic acids and sequence information of reference nucleic acids to generate a background-corrected signal informative for determining presence or absence of a cancer. In various embodiments, combining the sequence information of target nucleic acids and sequence information of reference nucleic acids involves differentiating between signatures present or absent in the sequence information of target nucleic acids and signatures present or absent in the sequence information of the reference nucleic acids. For example, if particular signatures are present in the sequence information of target nucleic acids, and the signatures are also present in the sequence information of reference nucleic acids, the signatures in both the target nucleic acids and reference nucleic acids may represent baseline biological signatures. Thus, these signatures may be excluded from the resulting signal informative of determining presence or absence of the cancer. As another example, if particular signatures are present in the sequence information of target nucleic acids, but those signatures are absent in the sequence information of reference nucleic acids, the signatures may not be baseline biological signatures. Thus, these signatures may be included in the resulting signal informative of determining presence or absence of the cancer.
[00130] In various embodiments, combining the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids includes aligning the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids. For example, aligning the sequence information involves aligning sequences of a plurality of pre-selected genomic sites for the target nucleic acids and sequences of the same or overlapping plurality of pre-selected genomic sites for the reference nucleic acids.
[00131] In various embodiments, both the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are aligned to a reference genome library (e.g., a reference assembly) with known sequences. Therefore, sequence information of the target nucleic acids are aligned to the sequence information of the reference nucleic acids via the reference genome library. In various embodiments, the sequence information of the target nucleic acids is aligned directly with the sequence information of the reference nucleic acids. In such embodiments, a reference genome library need not be used.
[00132] In various embodiments, combining the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids includes determining a difference between the sequence information of the target nucleic acids to the sequence information of the reference nucleic acids.
[00133] In various embodiments, differences between the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are performed on a per-position basis. For example, at a first position of a genomic site, the difference between the sequence information of the target nucleic acids at the first position and the sequence information of the reference nucleic acid at the same first position is determined. The process can then be further repeated for additional positions (e.g., for additional positions across the plurality of genomic sites). In various embodiments, the differences are determined on a perposition basis if the sequence information of the target nucleic acids and reference nucleic acids were generated using a sequencing assay (e.g., next generation sequencing) which provides base-level resolution of the sequences.
[00134] In various embodiments, differences between the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are performed on a per-CGI basis. For example, at a first CGI of a genomic site, the difference between the sequence information of the target nucleic acids at the first CGI and the sequence information of the reference nucleic acid at the same CGI or overlapping portion of the first CGI is determined. The process can then be further repeated for additional CGIs (e.g., for additional CGIs across the plurality of genomic sites). In various embodiments, the differences are determined on a per-CGI basis if the sequence information of the target nucleic acids and reference nucleic acids were generated using a quantitative assay (e.g., qPCR assay).
[00135] In various embodiments, differences between the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are performed on a per-allele basis. For example, at a first allele of a genomic site, the difference between the sequence information of the target nucleic acids at the first allele and the sequence information of the reference nucleic acid at the same allele or overlapping portion of the first allele is determined. The process can then be further repeated for additional alleles (e.g., for additional alleles across the plurality of genomic sites). In various embodiments, the differences are determined on a per-allele basis if the sequence information of the target nucleic acids and reference nucleic acids were generated using a quantitative assay (e.g., qPCR assay or allele-specific PCR assay).
[00136] In various embodiments, the intra-individual analysis generates a background-corrected signal that comprises phased sequencing information. As described herein, phased sequence information is derived specifically from a particular source and therefore, may be useful for determining presence or absence of a cancer because signals originating from the same source (e.g., maternal or paternal chromosome) may provide additional information in comparison to other approaches that merely analyze signals irrespective of the source. In various embodiments, performing the intra-individual analysis includes removing baseline biological signatures that would otherwise have been interpreted as being derived from a particular source. As described herein, phased sequencing information can include coupled genomic sites and / or coupled methylation sites from common sources. Therefore, by performing the intra-individual analysis, the coupled genomic sites and / or coupled methylation sites can be informative signatures deriving from common sources as opposed to baseline biological signatures.
[00137] Reference is now made to FIG. 3C, which depicts an example combining of sequence information of target nucleic acids and reference nucleic acids to generate a signal informative for a cancer, in accordance with an embodiment. The sequence information of the target nucleic acids and the sequence information of the reference nucleic acids include methylation statuses across a plurality of genomic sites. FIG. 3C shows an example genomic site in which nucleotide bases may be differentially methylated in the target nucleic acid and the reference nucleic acid. For example, as shown in FIG. 3C, the nucleotide base at the second position is methylated (as represented by the presence of a cytosine base which arises following bisulfite conversion) in both the target nucleic acid and the reference nucleic acid. Given that the methylation at the second position occurs in both the target nucleic acid and the reference nucleic acid, this may be a baseline biological signature. Conversely, the target nucleic acid may additionally be methylated at the sixth position and the ninth position, whereas the reference nucleic acid is unmethylated at the sixth position and the ninth position. Here, given that the reference nucleic acid is not methylated at the sixth and ninth position, the presence of the methylated nucleotide bases in the target nucleic acid may represent signatures that are informative of presence or absence of the cancer. Additionally, at the eleventh nucleotide position, the target nucleic acid is unmethylated whereas the reference nucleic acid is methylated. Here, the methylation of the reference nucleic acid can be interpreted as a baseline biological signature.
[00138] The differences between the methylation status at each position of the target nucleic acid and the reference nucleic acid can represent the cancer signal. As shown in FIG. 3C, the cancer signal includes methylation statuses at the genomic site, wherein the sixth and ninth position are methylated. Thus, the cancer signal includes signatures from the target nucleic acids that are likely informative of the cancer (e.g., methylated statuses of the sixth and ninth nucleotide bases), and further excludes baseline biological signatures (e.g., baseline biological signatures present in reference nucleic acids such as methylated statuses of the second and eleventh nucleotide bases). Second Analysis
[00139] The description in this section pertains to the performance of a second analysis, such as second analysis 130 described in FIG. 1 A, which can be performed by the second analysis module 220 described in FIG. 2A. Generally, a second analysis is performed on sequence information generated by the assay (e.g., assay 120A or assay 120B). In various embodiments, the second analysis is performed to determine whether a biological sample obtained from an individual contains a signal indicative of a cancer. For example, the screen is performed to determine whether a biological sample contains circulating tumor DNA. Circulating DNA within the biological sample may indicate that the individual (e.g., individual from whom the biological sample is obtained) has cancer. In various embodiments, the second analysis is performed on background-corrected methylation information from an intra-individual analysis to classify the subject as having cancer or not having cancer. In various embodiments, the second analysis is performed to analyze a change in background-corrected methylation information from two or more intra-individual analyses. By analyzing a change in background-corrected methylation information, the second analysis can predict a change in tumor heterogeneity e.g., for tracking tumor heterogeneity in the subject for guided therapy.
[00140] In various embodiments, a second analysis is performed on background-corrected sequence information generated via sequencing (e.g., next generation sequencing) of sequences at the one or more genomic locations, such as one or more CpG islands. In various embodiments, the background-corrected sequence information is generated as a result of whole genome sequencing and therefore, a second analysis is performed on sequences of one or more genomic locations across the whole genome.
[00141] Generally, the second analysis is a more expensive and / or a more complex test in comparison to the first tier (e.g., screen). By implementing a more complex second analysis, the second analysis can achieve a higher positive predictive value than the first tier. In various embodiments, performing the second analysis involves analyzing methylation information across a plurality of genomic locations that represents a higher resolution in comparison to the lower resolution information analyzed in the first tier. For example, the second analysis may determine a high resolution measure of methylation across the plurality of genomic sites that distinguishes individuals having cancer from other individuals not having cancer in accordance with a high performance metric (e.g., high PPV or high sensitivity). Here, the high resolution measure of methylation can provide information as to methylation status at each individual genomic site and / or methylation statuses across a group of genomic sites.
[00142] In various embodiments, the high resolution measure of methylation can be a total quantity of consecutively methylated CpG sites within target regions. In some embodiments, the high resolution measure of methylation can be a total quantity of 3 consecutively methylated CpG sites (referred to as “K3N3”) within target regions. In some embodiments, the high resolution measure of methylation can be a total quantity of 4 consecutively methylated CpG sites (referred to as “K4N4”) within target regions. In some embodiments, the high resolution measure of methylation can be a total quantity of 5 consecutively methylated CpG sites (referred to as “K5N5”) within target regions. For example, the high resolution measure of methylation can be a total quantity of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within a subset of the CGIs in any one of Tables 1, 2, 3, or 4. As another example, the high resolution measure of methylation can be a total quantity of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within all of the CGIs in any one of Tables 1, 2, 3, or 4. In some embodiments, the high resolution measure of methylation can be a proportion of 3 consecutively methylated CpG sites (referred to as “K3N3”) within target regions. In some embodiments, the high resolution measure of methylation can be a proportion of 4 consecutively methylated CpG sites (referred to as “K4N4”) within target regions. In some embodiments, the high resolution measure of methylation can be a proportion of 5 consecutively methylated CpG sites (referred to as “K5N5”) within target regions. For example, the high resolution measure of methylation can be a proportion of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within a subset of the CGIs in any one of Tables 1, 2, 3, or 4. As another example, the high resolution measure of methylation can be a proportion of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within all of the CGIs in any one of Tables 1, 2, 3, or 4.
[00143] In some embodiments, the high resolution measure of methylation can be a total quantity of consecutively methylated CpG sites within one or more CGIs that are methylated in the genome of extraembryonic ectoderm (ExE). Here, such example CGIs may be differentially methylated in the genome of ExE and not methylated in corresponding epiblast or adult tissue. Example CGIs that are methylated in the genome of ExE are further disclosed in Table 3 of WO2022133315, which is hereby incorporated by reference in its entirety.
[00144] In various embodiments, the high resolution measure of methylation can include methylation statuses of a plurality of CpG sites from a haplotype (e.g., inherited from either a maternal or paternal source). In various embodiments, the high resolution measure of methylation refers to methylation statuses of at least a portion of the CpGs within a CGI within at least a portion of one or more regions in Tables 1-4 from a common haplotype. In various embodiments, the high resolution measure of methylation refers to methylation statuses of all CpGs within a CGI within at least a portion of one or more regions in Tables 14 from a common haplotype. In various embodiments, the high resolution measure of methylation refers to methylation statuses of all CpGs within a CGI within one or more regions in Tables 1-4 from a common haplotype.
[00145] In various embodiments, the second analysis is performed using a system comprising a computer storage and a processing system. The second analysis can involve the implementation of trained machine learning models, details of which are described in further detail herein. For example, the computer storage can store sequence information corresponding to a processed sample, the processed sample including cell-free DNA fragments originating from a liquid biopsy of an individual and having been processed to enrich for cancer informative CGIs, the sequencer information comprising, for each sequenced cell-free DNA fragment corresponding to the cancer informative CGIs, a respective position on the genome for the cell-free DNA fragment and methylation information for the cell-free DNA fragment.
[00146] In particular embodiments, the second analysis further reveals, for individuals who are determined to have the cancer, a tissue of origin of the cancer. The second analysis may identify a tissue of origin of the cancer according to the methylation statuses of the cancer informative CGIs. For example, particular methylation patterns across the cancer informative CGIs are attributable to certain tissues, examples of which include the nervous tissue (e.g., brain, spinal cord, nerves), muscle tissue (cardiac muscle, smooth muscle, skeletal muscle), epithelial tissue (e.g., GI tract lining, skin), and connective tissue (e.g., fat, bone, tendon, and ligaments). As a particular example, in patients with brain cancer, a first set of CGIs may be frequently methylated. Therefore, if a similar methylation pattern is observed across the first set of CGIs for an individual who is under analysis, the second analysis can identify that the individual has cancer, and furthermore, that the cancer is localized to the brain.
[00147] In various embodiments, the second analysis involves analyzing a plurality of CGIs. For example, the second analysis involves analyzing methylation statuses of a plurality of CGIs. Cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. The accompanying tables (e.g., Tables 1-4) lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In various embodiments, the second analysis involves analyzing all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 1. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 1. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 2. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 2. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 3. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 3. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 4. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 4. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Tables 2 and 3. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Tables 2 and 3.
[00148] In various embodiments, the second analysis involves analyzing at least 100 CGIs (e.g., CGIs as shown in any of Tables 1-4). In various embodiments, the second analysis involves analyzing at least 100 CGIs, at least 150 CGIs, at least 200 CGIs, at least 300 CGIs, at least 400 CGIs, at least 500 CGIs, at least 600 CGIs, at least 700 CGIs, at least 800 CGIs, at least 900 CGIs, at least 1000 CGIs, at least 1500 CGIs, at least 2000 CGIs, at least 2500 CGIs, at least 3000 CGIs, at least 3500 CGIs, at least 4000 CGIs, at least 4500 CGIs, at least 5000 CGIs, at least 5500 CGIs, or at least 6000 CGIs (e.g., CGIs as shown in any of Tables 1-4). In particular embodiments, performing the screen involves analyzing at least 500 CGIs. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.
[00149] In various embodiments, the second analysis involves analyzing more CGIs in comparison to the quantity of CGIs analyzed during the screen. For example, the CGIs analyzed during the screen can represent a subset of the CGIs analyzed during the second analysis. In some scenarios, every CpG island analyzed during the screen is further analyzed when performing the second analysis. Therefore, the second analysis represents a more robust and rigorous analysis in comparison to the more rapid and cost-effective screen. In various embodiments, the second analysis involves analyzing at least 2 times the number of CGIs analyzed during the screen. In various embodiments, the second analysis involves analyzing at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least 14 times at least 15 times, at least 16 times, at least 17 times, at least 18 times, at least 19 times, at least 20 times, at least 21 times, at least 22 times, at least 23 times, at least 24 times, at least 25 times, at least 26 times, at least 27 times, at least 28 times, at least 29 times, at least 30 times, at least 31 times, at least 32 times, at least 33 times, at least 34 times, at least 35 times, at least 36 times, at least 37 times, at least 38 times, at least 39 times, or at least 40 times the number of CGIs analyzed during the screen. In particular embodiments, the second analysis involves analyzing at least 5 times the number of CGIs analyzed during the screen. For example, the screen may involve analyzing at least 100 CGIs and the second analysis may involve analyzing at least 500 CGIs.
[00150] In various embodiments, the second analysis achieves at least 60% sensitivity in detecting presence of a cancer. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sensitivity. In particular embodiments, the second analysis achieves at least 85% sensitivity. In particular embodiments, the second analysis achieves at least 86% sensitivity. In particular embodiments, the second analysis achieves at least 87% sensitivity. In particular embodiments, the second analysis achieves at least 88% sensitivity. In particular embodiments, the second analysis achieves at least 89% sensitivity. In particular embodiments, the second analysis achieves at least 90% sensitivity.
[00151] In various embodiments, the second analysis achieves at least 60% specificity in excluding individuals without the cancer. In various embodiments, the second analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% specificity. In particular embodiments, the second analysis achieves at least 90% specificity. In particular embodiments, the second analysis achieves at least 91% specificity. In particular embodiments, the second analysis achieves at least 92% specificity. In particular embodiments, the second analysis achieves at least 93% specificity. In particular embodiments, the second analysis achieves at least 94% specificity. In particular embodiments, the second analysis achieves at least 95% specificity.
[00152] In various embodiments, the second analysis achieves at least 60% positive predictive value. In various embodiments, the second analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% positive predictive value. In particular embodiments, the second analysis achieves at least 80% positive predictive value. In particular embodiments, the second analysis achieves at least 81% positive predictive value. In particular embodiments, the second analysis achieves at least 82% positive predictive value. In particular embodiments, the second analysis achieves at least 83% positive predictive value. In particular embodiments, the second analysis achieves at least 84% positive predictive value. In particular embodiments, the second analysis achieves at least 85% positive predictive value.
[00153] In various embodiments, the second analysis achieves at least 60% negative predictive value. In various embodiments, the second analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% negative predictive value. In particular embodiments, the second analysis achieves at least 90% negative predictive value. In particular embodiments, the second analysis achieves at least 91% negative predictive value. In particular embodiments, the second analysis achieves at least 92% negative predictive value. In particular embodiments, the second analysis achieves at least 93% negative predictive value. In particular embodiments, the second analysis achieves at least 94% negative predictive value. In particular embodiments, the second analysis achieves at least 95% negative predictive value. In particular embodiments, the second analysis achieves at least 96% negative predictive value. In particular embodiments, the second analysis achieves at least 97% negative predictive value. In particular embodiments, the second analysis achieves at least 98% negative predictive value. In particular embodiments, the second analysis achieves at least 99% negative predictive value. Longitudinal Analysis
[00154] In various embodiments, methods disclosed herein are valuable for performing longitudinal analysis for a subject. For example, a subject who was determined to have a presence of cancer (e.g., through the screen or through the second analysis) can be further tracked through a longitudinal analysis. In various embodiments, an additional sample is obtained from the subject at a subsequent timepoint, and the second analysis can be further performed for the subject using the additional sample. Thus, the second analysis performed for the additional sample can determine a change in the cancer for the subject over the intervening timeframe.
[00155] In various embodiments, a longitudinal analysis can be performed for subjects who may have been identified as not having cancer. In various embodiments, a longitudinal analysis is performed for subjects who were identified as negative through the screen (e.g., first analysis). In various embodiments, a longitudinal analysis is performed for subjects who were identified as negative through the second analysis. In various embodiments, a longitudinal analysis is performed for subjects who were identified as not negative through the screen and then further identified as negative through the second analysis. By longitudinally tracking subjects who may have been identified as not having cancer, any false negative subjects can potentially be identified through subsequent testing of one or more additional samples obtained at one or more subsequent timepoints. For example, a subject can be identified as not negative through the screen, and through the longitudinal analysis (e.g., at a subsequent timepoint), an additional sample of the subject can be analyzed using either the methodology described in reference to the screen or the second analysis to identify the subject as a false negative. As another example, a subject can be identified as not negative through the second analysis, and through the longitudinal analysis (e.g., at a subsequent timepoint), an additional sample of the subject can be analyzed using either the methodology described in reference to the screen or the second analysis to identify the subject as a false negative.
[00156] Reference is now made to the tumor tracking module 230, which represents a module of the tumor heterogeneity system 170 as shown in FIG. 2A. In various embodiments, tracking tumor heterogeneity over two or more timepoints enables the determination of whether an intervention is efficacious. Given a subject who has previously received the intervention (e.g., a tumor therapeutic) for treating cancer, tracking tumor heterogeneity over two or more timepoints using the methods disclosed herein is informative for determining whether the intervention is efficacious for treating the cancer. Generally, a subject exhibiting a reduction in tumor heterogeneity over two or more timepoints is indicative that the tumor subclones are decreasing and that the intervention is effective. Alternatively, a subject who does not exhibit a reduction in tumor heterogeneity (e.g., stable or increase tumor heterogeneity) is indicative that the tumor subclones is unchanging or is increasing. In this scenario, the intervention lacks efficacy. Thus, methods for tracking tumor heterogeneity can be useful for e.g., guided therapy.
[00157] In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining samples from the subject across two or more timepoints, performing intraindividual analysis for one or more of the obtained samples, and generating predictions across at least the two or more timepoint. The predictions can be informative for the subject’s tumor heterogeneity. In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining three or more samples from a subject across at least three timepoints, performing intra-individual analysis for the three or more samples, and generating predictions across the at least three timepoints. In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining four or more samples from a subject across at least four timepoints, performing intra-individual analysis for the four or more samples, and generating predictions across the at least four timepoints. In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining samples from a subject, performing intra-individual analysis for each of the obtained samples, and generating predictions across at least five timepoints, at least six timepoints, at least seven timepoints, at least eight timepoints, at least nine timepoints, at least ten timepoints, at least eleven timepoints, at least twelve timepoints, at least thirteen timepoints, at least fourteen timepoints, at least fifteen timepoints, at least sixteen timepoints, at least seventeen timepoints, at least eighteen timepoints, at least nineteen timepoints, or at least twenty timepoints.
[00158] In various embodiments, the time between any two timepoints can be between 1 day and 12 months, between 5 days and 8 months, between 10 days and 6 months, between 15 days and 4 months, between 20 days and 3 months, between 30 days and 2 months. In various embodiments, the time between any two timepoints can be between 1 days and 10 days, between 10 days and 20 days, between 20 days and 30 days, between 30 days and 40 days, between 40 days and 50 days, or between 50 days and 60 days. In various embodiments, the time between any two timepoints can be between 1 day and 100 days, between 5 day and 80 days, between 10 days and 70 days, between 15 days and 60 days, between 20 days and 50 days, between 25 days and 40 days, or between 30 days and 35 days. In various embodiments, the time between any two timepoints can be between 1 days and 10 days, between 10 days and 20 days, between 20 days and 30 days, between 30 days and 40 days, between 40 days and 50 days, or between 50 days and 60 days. In various embodiments, the time between any two timepoints can be between 1 month and 2 months.
[00159] In various embodiments, methods for tracking tumor heterogeneity involve obtaining a sample from the subject at a first timepoint (e.g., an initial timepoint), performing an intraindividual analysis using the obtained sample, and generating a cancer prediction for the sample obtained at the first timepoint. In various embodiments, the first timepoint may refer to a timepoint prior to which the subject receives an intervention, such as a tumor therapeutic. Thus, the generated for the sample obtained at the first timepoint may represent a baseline prediction prior to any therapeutic treatment. In various embodiments, the first timepoint may refer to a timepoint immediately after the subject receives an intervention, such as a tumor therapeutic. In this context, “immediately after” the subject receives an intervention can refer to a timeframe within 1 day after the subject receives the intervention. In various embodiments, “immediately after” refers to a timeframe within 12 hours, within 8 hours, within 6 hours, within 4 hours, within 3 hours, within 2 hours, within 1 hour, within 30 minutes, within 15 minutes, within 10 minutes, within 5 minutes, or within 1 minute of the subject receiving the therapeutic.
[00160] In particular embodiments, methods for tracking tumor heterogeneity further involve obtaining one or more subsequent samples from the subject after the first timepoint (e.g., at a second timepoint, at a third timepoint, at a fourth timepoint, etc.), performing intra-individual analyses for a subsequent sample, and generating predictions for the one or more subsequent samples. In this scenario, the change in the predictions for the one or more subsequent samples in comparison to the prediction of the first sample can be indicative of the change in tumor heterogeneity. In various embodiments, the one or more subsequent samples are obtained from the subject after the subject has received an intervention, such as a tumor therapeutic. Thus, the change in tumor heterogeneity can be reflective of the efficacy, or lack thereof, of the intervention provided to the subject. Machine Learning Models for Analyzing Sequence Information
[00161] In various embodiments, trained machine learning models can be deployed to analyze sequence information for tracking tumor heterogeneity for a subject across two or more timepoints. In various embodiments, the sequence information includes methylation statuses of plurality of genomic sites. Therefore, trained machine learning models analyze differential methylation of the plurality of genomic sites to output predictions.
[00162] In various embodiments, a trained machine learning model is deployed as part of a screen (e.g., screen 125 as shown in FIG. 1 A). Thus, the trained machine learning model can analyze sequence information generated via an assay (e.g., assay 120A shown in FIG. 1 A) to determine whether a subject is negative or not negative for a cancer. In various embodiments, a trained machine learning model is deployed as part of a second analysis (e.g., second analysis 130 shown in FIG. 1 A). Therefore, the trained machine learning model can analyze sequence information including methylation statuses for a plurality of genomic sites, such as a plurality of CpG sites disclosed herein. In various embodiments, the sequence information includes background-corrected sequence information generated via an intraindividual analysis (e.g., intra-individual analysis 128A and / or intra-individual analysis 128B shown in FIG. 1 A). In some embodiments, the trained machine learning model analyzes a difference between background-corrected sequence information determined from two intraindividual analyses (as shown in FIG. 1 A). In some embodiments, the trained machine learning model analyzes background-corrected sequence information from a single intraindividual analysis (as shown in FIG. IB).
[00163] In various embodiments, a machine learning model is any one of a regression model (e.g., linear regression, logistic regression, or polynomial regression), decision tree, random forest, support vector machine, Naive Bayes model, k-means cluster, or neural network (e.g., feed-forward networks, convolutional neural networks (CNN), deep neural networks (DNN), autoencoder neural networks, generative adversarial networks, or recurrent networks (e.g., long short-term memory networks (LSTM), bi-directional recurrent networks, deep bidirectional recurrent networks).
[00164] The machine learning model can be trained using a machine learning implemented method, such as any one of a linear regression algorithm, logistic regression algorithm, decision tree algorithm, support vector machine classification, Naive Bayes classification, K-Nearest Neighbor classification, random forest algorithm, deep learning algorithm, gradient boosting algorithm, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or combinations thereof. In various embodiments, the machine learning model is trained using supervised learning algorithms, unsupervised learning algorithms, semi-supervised learning algorithms (e.g., partial supervision), weak supervision, transfer, multi-task learning, or any combination thereof.
[00165] In various embodiments, the machine learning model has one or more parameters, such as hyperparameters or model parameters. Hyperparameters are generally established prior to training. Examples of hyperparameters include the learning rate, depth or leaves of a decision tree, number of hidden layers in a deep neural network, number of clusters in a k-means cluster, penalty in a regression model, and a regularization parameter associated with a cost function. Model parameters are generally adjusted during training. Examples of model parameters include weights associated with nodes in layers of neural network, support vectors in a support vector machine, and coefficients in a regression model. The model parameters of the machine learning model are trained (e.g., adjusted) using the training data to improve the predictive power of the machine learning model.
[00166] In particular embodiments, trained machine learning models analyze methylation statuses of a plurality of genomic sites to generate predictions. The methylation statuses can correspond to a set of cancer informative CpG islands (CGIs), wherein the cancer informative CGIs are selected from a group consisting of a ranked set of candidate CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 50 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 100 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 150 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 200 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 250 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 300 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 400 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 500 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 600 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 700 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 800 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 900 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 1000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 2500 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 5000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 7500 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 10000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 15000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 20000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 25000 CGIs.
[00167] In various embodiments, a machine learning model analyzes methylation statuses for CGIs across the whole genome. For example, a machine learning model may be implemented to analyze sequencing data generated from whole genome sequencing (e.g., whole genome bisulfite sequencing).
[00168] Additionally disclosed herein are particular genomic sites, such as CpG islands (CGIs) whose methylation statuses can be informative for determining whether a subject is at risk of a cancer or whether the individual has a cancer. In some embodiments, methylation statuses of the informative CGIs representing a signal in a sample can be indicative of a presence of the cancer. In some embodiments, methylation statuses of the informative CGIs representing a signal in a sample can be indicative of an absence of the cancer. In various embodiments, methods disclosed herein, such as methods involving the multiple-tiered analysis, are useful for detecting or identifying the signal (e.g., methylation statuses of the informative CGIs) in a sample. In various embodiments, methods disclosed herein, such as methods involving the multiple-tiered analysis, are useful for increasing the probability that the detected signal (e.g., methylation statuses of the informative CGIs) in the sample is authentic. A signal (e.g., methylation statuses of the informative CGIs) detected by the multiple-tiered analysis can be confidently trusted as present in the sample. Thus, by tracking the change in methylation statuses for the subject across multiple timepoints, a change in the subject’s risk for cancer or a change in the subject’s cancer can be more accurately determined.
[00169] Methylation statuses of cancer informative CGIs can be useful for predicting whether an individual has a cancer or is at risk for a cancer. In various embodiments, the methylation statuses of cancer informative CGIs are background-corrected methylation statuses of cancer informative CGIs. For example, background-corrected methylation statuses of cancer informative CGIs can be determined via an intra-individual analysis. For example, background-corrected methylation statuses of cancer informative CGIs can be determined by combining methylation information of cancer informative CGIs of target nucleic acids and methylation information of cancer informative CGIs of reference nucleic acids.
[00170] In various embodiments, each cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. The accompanying tables (e.g., Tables 1-4) lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.
[00171] Reference is now made to FIG. 3D, which is an illustrative example of a signal informative for a cancer. In various embodiments, the signal informative for a cancer shown in FIG. 3D can be generated as a result of the intra-individual analysis. Thus, the signal informative for a cancer represents background-corrected sequence information e.g., corrected via an intra-individual analysis that combines sequence information from target nucleic acids and reference nucleic acids. In various embodiments, the signal informative for a cancer shown in FIG. 3D can represent sequence information of target nucleic acids. In such embodiments, the signal is not derived from an intra-individual analysis.
[00172] As shown in FIG. 3D, for each instance of an analyte, e.g., a cell-free DNA fragment, there is data indicating, for each of a plurality of positions along the instance of the analyte, e.g., distinct CpG sites along a DNA fragment, information about a marker at that position, e.g., whether that CpG is methylated or unmethylated. An instance of an analyte can be a single sequenced DNA fragment or a portion of a single sequenced DNA fragment. In various embodiments, the DNA fragment may be a bisulfite converted DNA fragment. Therefore, an instance of an analyte may refer to a sequenced bisulfite converted DNA fragment or a portion thereof.
[00173] Conceptually, using methylation of CpGs in cell-free DNA as an illustrative example, the signal illustrated in FIG. 3D includes a row, e.g., row 240, for each instance of an analyte, such as a single sequenced DNA fragment. Thus, in FIG. 3D, data for sixteen instances of an analyte are shown, e.g., sixteen DNA fragments. Each circle corresponds to a position along the analyte, such as a CpG site. In this example, whether the circle is illustrated as black or white in FIG. 3D, is indicative of whether the CpG site is methylated (black) or unmethylated (white). In some instances, information about a marker at a position in a nucleic acid may not be binary.
[00174] The information about the markers for each instance of an analyte in a sample can result in a large amount of data. As an example, in practice, in the case of obtaining methylation state of CpGs in cell-free DNA from a blood sample using deep sequencing, using a DNA sequencer that outputs such data into a FASTQ format data file, the signal generated by processing a single blood sample can be many gigabytes, e.g., 20 to 30 gigabytes, of data.
[00175] FIG. 3D also illustrates a relative alignment among the distinct instances of the analyte. In the example of DNA, for example, the position of a DNA fragment within a genome for the individual from which a sample originated can be determined, and each position within the genome can have a respective set of coordinates identifying it. Thus, DNA fragments can be assigned coordinates based on their respective positions within the genome, and then aligned or grouped by those coordinates. Thus, in FIG. 3D, column 242 indicates a position on an analyte, such as a single CpG site in a genome, and the distinct instances of the analyte are illustrated as aligned by position on the analyte.
[00176] By using the position information for each instance of an analyte, distinct instances of the analyte can be grouped into regions within the analyte. Typically, markers related to cancer are localized within identifiable regions of analytes, such as specific genes or regions within the genome. Thus, the signals generated for each instance of an analyte can be grouped and processed by cancer-informative regions. In particular embodiments, an informative region is a CGI (or at least a portion thereof) as disclosed in any of Tables 1-4. The example in FIG. 3D can be considered to illustrate data about methylation at CpG sites within one informative region of the genome, for multiple DNA fragments obtained from a biological sample. There can be multiple cancer-informative regions.
[00177] As disclosed herein, trained machine learning models are deployed to generate informative predictions regarding presence or absence of cancer. To use a trained machine learning model in this context, there are several technical problems that arise relating to encoding the signal resulting from processing a biological sample into features. Some problems arise because the signal includes a large amount of information. One of the challenges involves reducing the volume of data into a set of informative features. However, as the number of features increases, the complexity of the computational model increases. However, as the number of features decreases, information relevant to detection of a cancer may be lost. Some problems arise because of uncertainty around which metrics and which regions of an analyte are truly informative of a cancer. Omission of some metrics or some regions from the set of features may adversely impact the performance of a trained computational model.
[00178] To address such problems, in various embodiments, very particularly engineered features are generated from a biological sample. Such engineered features may be dependent on one or more health-condition-informative regions (e.g., CGIs) and / or one or more distinct windows within the health-condition informative regions (e.g., CGIs). Each window may have a specified range of positions within a health-condition informative region, and a specified size. The size is specified in terms of a number of consecutive sites of interest within the analyte. A metric is thus computed for a plurality of windows within the healthcondition informative region. Thus, in particular embodiments, the engineered features, representing metrics within a particular window within a health-condition informative region (e.g., CGIs), are informative for a cancer.
[00179] To train a machine learning model, in some embodiments, a first set of features is computed for a training set, which can include several candidate features. The candidate features can include one or more candidate metrics, or one or more candidate healthcondition-informative regions, or combinations of both. A computational model can be trained using candidate features, and then analyzed to determine which candidate features were more influential in the output of the trained computational model. Such analysis can be used to identify features which are more influential to the model, whether due to the metric or due to the health-condition-informative region. A second set of features can be defined by reducing the first set of features based on those identified features which are more influential, and the trained machine learning model can be built using the second set of features.
[00180] In various embodiments, to generate data for a machine learning model (e.g., for training or for deployment), the methodology includes computing, for one or more instances of an analyte in a window of a plurality of windows on a target region of the analyte, a metric specific for the window and the target region. The specific metrics used, and healthcondition-informative regions selected can depend on a variety of factors and may be experimentally determined. The machine learning model can be implemented to analyze at least the metric specific for the window and the target region. In various embodiments, the metric specific for the window and the target region includes a proportion of a count of DNA fragments having a specific count of methylated CpGs to a count of DNA fragments for the window of the target region. In various embodiments, the metric specific for the window and the target region comprises a proportion of a count of DNA fragments having a specific pattern of methylation to a count of DNA fragments for the window of the target region. As described in further detail below, computing the metric can involve applying two or more functions. For example, computing the metric specific for the window and the target region can involve performing a first function to quantify a count of occurrences of methylated CpGs within the window of the target region. As another example, computing the metric specific for the window and the target region can involve performing a second function to normalize the count of occurrences of methylated CpGs relative to a count of DNA fragments for the window of the target region.
[00181] In various embodiments, to generate features, each instance of the analyte (e.g., cell-free DNA) is processed. For each instance of an analyte in the biological sample, and for each window of a plurality of windows on health-condition-informative regions of the analyte, a respective value is generated. After processing instances of the analyte, the feature computation module then computes, for each window of the plurality of windows on the health-condition-informative region, one or more respective metrics for the window based on a first function and / or a second function for instances of the analyte for the window. In various embodiments, a first function quantifies markers within a window. As a specific example, a first function refers to a quantification of a number of methylated CpG sites within a window. In various embodiments, a second function computes a proportion of the quantified markers within the window in relation to other quantified markers. As a specific example, a second function computes the proportion of the number of methylated CpG sites within a window relative to other numbers of methylated CpG sites within a window.
[00182] Example implementations will now be described in reference to FIGs. 3E and 3F. Here, in FIG. 3E, illustrative marker information for instances of an analyte are shown schematically for the purposes of simplifying this explanation. In this example, there are ten (10) instances of an analyte, each having a length of six (6) sites of interest, at which marker information is a binary value, indicated by a black or white circle. FIG. 3E shows aligned instances of an analyte, along with the designation of a window with a particular kmer size (e.g., K=3). Each window has a size of three (3) consecutive sites of interest within the analyte. In other embodiments, smaller or larger window sizes may be implemented for the analysis. There are four (4) windows of size three (3) (i.e., a first window that includes the first, second, and third sites of interest from the left, a second window that includes the second, third, and fourth sites of interest from the left, a third window that includes the third, fourth, and fifth sites of interest from the left, and a fourth window that includes the fourth, fifth, and sixth sites of interest from the left), but computations for three (3) windows are shown.
[00183] In FIG. 3E an example of a first function applied to an instance of an analyte is a count of occurrences of marker information within the instance of the analyte within the window. For example, where the marker information is methylation of a CpG site, this function can be a count of methylated CpGs in the window. That is, if the window has a size of three sites of interest, then there are four possible counts: 0, 1,2, and 3. Note that inverse results would be obtained if the count was of unmethylated CpGs in the window, but such results when used in training would have the same effect.
[00184] In FIG. 3E, the second function computes counts of the number of instances having each possible count resulting from the first function. That is, if the window has a size of three sites of interest, for which there are four possible counts (0, 1, 2, and 3), for that window the second function computes a count of the number of instances with a count of zero, a count of the number of instances with a count of one, a count of the number of instances with a count of two, and a count of the number of instances with a count of three. The second function divides the respective number of instances computed for possible counts by the total number of instances, thus providing a fractional value for each of the possible counts for this window.
[00185] In this example in FIG. 3E, for this health-condition-informative region (referred to as “HC1”), there are windows “Wl”, “W2”, and “W3”, each of which has four (4) values, representing the respective count for each possible count of methylated CpGs among the instances that overlap that window. Because there are ten (10) instances, each of these values is divided by 10 in the second function, to provide the respective final four output values for each window. As shown in FIG. 3E, referring to the example of Window 1 (Wl), the final four output values are 0.3 (0 methylated CpG sites in the window), 0.1 (1 methylated CpG sites in the window), 0.1 (2 methylated CpG sites in the window), and 0.5 (3 fully methylated CpG sites in the window). Here, the proportion of fully methylated CpG sites, proportion of fully non-methylated CpG sites, and proportion of partially methylated CpG sites (e.g., either 1 or 2 methylated CpG sites in the window) can be metrics informative for a cancer.
[00186] Reference is now made to FIG. 3F, which shows an example application of a first function and second function to instances of an analyte. Here, the bottom of FIG. 3F shows patterns of the marker information in the instance, from among a set of possible patterns. A pattern is a unique sequence of marker information along the sites of interest in a window. For example, as shown in FIG. 3F, if the window has a size of three sites of interest, and if the marker information for the sites of information is binary, then there are eight possible patterns. For example, where the marker information is methylation of a CpG site, each possible pattern of methylation in a window is a distinct sequence of the methylation state (e.g., methylated or unmethylated) of the CpG sites along the sequence of consecutive CpG sites in the window. When the marker information is methylation of CpG sites, the first function, applied to an instance of a DNA fragment in a window, outputs an indication of which of the possible patterns of methylation of CpGs is present in the window in that DNA fragment.
[00187] The second function computes a count of the number of instances having each possible pattern in a window. That is, for that window, the second function produces a count of the number of instances with the first pattern, a count of the number of instances with the second pattern, and so on. The second function then divides the respective number of instances identified for each possible pattern by the total number of instances, thus providing a fractional value for each of the possible patterns for this window, as shown in the bottom panel of FIG. 3F.
[00188] In this example in FIG. 3F, for this health-condition-informative region (say, “HC1”), there are windows “Wl”, “W2”, and “W3”, each of which has eight values, representing the respective number of occurrences each possible pattern among the instances that overlap that window divided by the number of instances, in this case ten (10).
[00189] In any of the foregoing example implementations, and in other implementations, a size of a health-condition-informative region, in terms of a number of sites of interest within an instance of an analyte, can vary. For example, cancer-informative regions of DNA may be as small as a single CpG site, and may include several 10’s, 100’s, or 1000’s of CpG sites. Within a set of features, there may be a plurality of health-condition-informative regions, each having its own respective size.
[00190] In any of the foregoing example implementations, and in other implementations, a size of a window in a health-condition-informative region, in terms of a number of sites of interest within an instance of an analyte, can vary. Generally, the number of sites of interest is a positive integer number that ranges between 1 and N. In some example implementations, N is less than or equal to 10, or 9, or 8, or 7, or 6, or 5, or 4, or 3. In various embodiments, a window within a health-condition informative region includes a specific numbers of CpG sites. In various embodiments, TV is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites. In various embodiments, TV is between 1 and 100, between 2 and 80, between 3 and 60, between 4 and 40, between 5 and 20, or between 6 and 10 CpG sites. In various embodiments, N is between 1 and 10, between 2 and 9, between 3 and 8, between 4 and 7, or between 4 and 6 CpG sites. Within a set of features, there may be a plurality of health-condition-informative regions, each having its own respective window size or set of window sizes. Different window sizes may be used in different regions. The same window size may be used in different regions. A region may have metrics computed for it for multiple different window sizes. Windows may be over-lapping or non-overlapping.
[00191] In various embodiments, a metric represents an input vector that can be provided as input to a machine learning model (e.g., either during training or deployment of the machine learning model). Here, the metric may be specific for a window and a target region of interest (e.g., a target region comprising one or more CpG sites). For example, the input vector of the metric may include a set of values representing the proportion of counts of methylated CpGs in the window relative to a total count (e.g., total count of DNA fragments for the window of the target region). In various embodiments, the input vector of the metric may include a set of values representing proportions of DNA fragments having specific counts of methylated CpGs out of all possible CpG methylation patterns in the window. The all possible CpG methylation patterns are 2k possible patterns, where k refers to a number of CpG sites in the window. Referring against to the bottom panel of FIG. 3F, an input vector of a metric can be generated for a particular window. Taking the first window (e.g., left-most window shown in FIG. 3F) as an example, the input vector of the metric may include the proportion vales shown in the left most column in the bottom panel of FIG. 3F. Thus, the input vector of the metric may be represented as [0.3, 0.1, 0, 0, 0, 0.1, 0, 0.5], Similar input vectors for other metrics can be generated using the values of other windows.
[00192] The computed sets of values for the set of features for samples can be stored in a data structure, which can be stored in a database, memory, or other computer storage for use in connection with the computational model, or for other purposes.
[00193] In some implementations, the sets of values for the set of features for a sample can be stored in association with an identifier of the subject, or an identifier of the sample, or both, so that the identifier of the subject or the identifier of the sample, or both, can be used to access the set of values from the computer storage. In some implementations, each computed value can be associated with an identifier of the cancer-informative region, and an identifier of the window within that region, to which the value corresponds.
[00194] Accordingly, an example implementation of such a data structure is shown in FIG. 3G. A set of values for a set of features is stored for a biological sample originating from a subject. The data structure can include an optional identifier for the subject, and an optional identifier for the biological sample. The latter identifier is useful when there are multiple samples for a single subject. For a sample, as indicated at 250, the set of features includes one or more metrics, for each of one or more windows 254, e.g., window “W-l-1”, within each of one or more health-condition-informative regions 252A, e.g., region “RI” or 252B e.g., region “R2”. For each feature, e.g., R-l, W-l-1, Metric, the computed value, e.g., Value 256, is stored. The number of windows in each region can be different for each region. The size of the window can be different for each window. The metric(s) computed for the window can be different for each window. Example Methods for Conducting Two or More Intra-Individual Analyses
[00195] As disclosed herein, methods involve tracking tumor heterogeneity in a subject by conducting intra-individual analyses for two or more samples obtained from the subject across two or more timepoints. For example, a first intra-individual analysis can be performed for a first sample obtained from the subject at a first timepoint and a second intraindividual analysis can be performed for a second sample obtained from the subject at a second timepoint. Thus, the change in results from each intra-individual analysis can be informative for tracking tumor heterogeneity in the subject.
[00196] FIG. 4A shows an example flow process involving a first and second intra-individual analyses, in accordance with a first embodiment. In this first embodiment, the flow process involves performing separate intra-individual analyses for first and second samples obtained from the subject at two different timepoints and performing a second analysis on the difference between the results of the separate intra-individual analyses.
[00197] Step 410 involves performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA.
[00198] Next, at step 415, if the first biological sample is not identified as not at risk, perform a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information.
[00199] Step 420 involves performing a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information, the second biological sample obtained from the subject at a second timepoint subsequent to the first timepoint.
[00200] Step 425 involves determining a change in signal between the first set of background-corrected methylation information and the second set of background-corrected methylation information.
[00201] Step 430 involves performing a second analysis comprising analyzing the determined change in signal to track tumor heterogeneity.
[00202] Reference is now made to FIG. 4B, which shows an example flow process involving a first and second intra-individual analyses, in accordance with a second embodiment. In this second embodiment, the flow process involves performing separate intra-individual analyses for first and second samples obtained from the subject at two different timepoints and performing a second analysis on each of the results of the separate intra-individual analyses.
[00203] Step 450 involves performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA.
[00204] Step 455 involves performing a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information.
[00205] Step 460 involves performing a second analysis to predict a tumor heterogeneity state.
[00206] Step 465 involves performing a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information, the second biological sample obtained from the subject at a second timepoint subsequent to the first timepoint.
[00207] Step 470 involves performing a second analysis to predict an updated tumor heterogeneity state.
[00208] Step 475 involves determining a change in signal between the first set of background-corrected methylation information and the second set of background-corrected methylation information. Guided Therapy
[00209] In various embodiments, the methods disclosed herein for performing a multipletiered analysis (e.g., screening and / or intra-individual analysis) to track tumor heterogeneity of one or more cancers in one or more subjects are informative for identifying an intervention for the subject. In various embodiments, an intervention may be any intervention known to those of ordinary skill in the art. Non-limiting examples of interventions include surgery (e.g., excising diseased or pre-disease tissue from an individual), a tumor therapeutic (e.g., chemotherapy, gene therapy, or gene editing), radiation therapy, or a lifestyle intervention (e.g., change in behavior or habits). In particular embodiments, the intervention comprises a tumor therapeutic.
[00210] In various embodiments, the methods disclosed herein are performed for a subject who previously received a tumor therapeutic. Thus, tracking the tumor heterogeneity of one or more cancers for the subject can be informative for determining whether the previously provided tumor therapeutic is efficacious. For example, if the tumor heterogeneity of a cancer is not decreasing (e.g., is increasing or is remaining stable) over the two or more timepoints, the tumor therapeutic is deemed non-efficacious. In this example, methods can involve selecting a new intervention, such as a new or different tumor therapeutic, for treatment of the subject’s cancer. As another example, if the tumor heterogeneity is decreasing over the two or more timepoints, the tumor therapeutic can be deemed efficacious. In this example, methods can involve selecting the tumor therapeutic that was previously provided to subject. Thus, the tumor therapeutic can continue to be provided to the subject to treat the cancer. In some embodiments, methods can involve selecting a new or different tumor therapeutic for treatment of the subject’s cancer. In some embodiments, methods can involve selecting a new or different intervention in addition to the previously provided tumor therapeutic. Thus, the new or different intervention and the previously provided tumor therapeutic can be provided to the subject to treat the cancer. Cancers
[00211] The disclosure provides methods for performing a multiple-tiered analysis (e.g., screening and / or intra-individual analysis) to track tumor heterogeneity of one or more cancers in one or more subjects. In various embodiments, the subject may have been previously diagnosed with a cancer and receives an intervention for treating the cancer. For example, the subject may have previously received a tumor therapeutic for treating the cancer. In various embodiments, the subject may be suspected of having a cancer, but may not have been previously diagnosed with a cancer. In various embodiments, the subject is healthy and is not yet suspected of having a cancer. In certain embodiments, a cancer is an early-stage health cancer, e.g., prior to development of symptoms.
[00212] In various embodiments, the cancer is an early stage cancer. In various embodiments, the cancer is a preclinical phase cancer. In various embodiments, the cancer is a stage I cancer. In various embodiments, the cancer is a stage II cancer. Thus, the methods disclosed herein enable the screening and tracking of tumor heterogeneity of a subject for an early stage or preclinical stage cancer.
[00213] In various embodiments, the cancer is any of an acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, cardiac cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumor, gestational trophoblastic cancer, hairy cell leukemia, liver cancer, Hodgkin lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, mouth cancer, multiple endocrine neoplasia syndromes, multiple myeloma neoplasms, myelodysplastic neoplasms, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer. Computer Implementation
[00214] The methods of the invention, including the methods of performing a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints, are, in some embodiments, performed on one or more computers. In particular embodiments, the steps of performing a screen (e.g., screen 125 shown in FIG. 1 A), performing an intra-individual analysis (e.g., intra-individual analysis 128A or intraindividual analysis 128B shown in FIG. 1 A), and performing a second analysis (e.g., second analysis 130 shown in FIG. 1 A) are performed on one or more computers. The steps of performing an assay (e.g., assay 120A and / or assay 120B shown in FIG. 1 A) are not performed on one or more computers.
[00215] In various embodiments, the performance of the screen, the intra-individual analysis, and / or the second analysis can be implemented in hardware or software, or a combination of both. In one embodiment, a machine-readable storage medium is provided, the medium comprising a data storage material encoded with machine readable data which, when using a machine programmed with instructions for using said data, is capable of displaying data (e.g., methylation data) and results of the screen, intra-individual analysis, and / or second analysis (e.g., tracked tumor heterogeneity). Such data can be used for a variety of purposes, such as determining an efficacy of a tumor therapeutic, or selecting a new intervention for the subject. The invention can be implemented in computer programs executing on programmable computers, comprising a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), a graphics adapter, a pointing device, a network adapter, at least one input device, and at least one output device. A display is coupled to the graphics adapter. Program code is applied to input data to perform the functions described above and generate output information. The output information is applied to one or more output devices, in known fashion. The computer can be, for example, a personal computer, microcomputer, or workstation of conventional design.
[00216] Each program can be implemented in a high level procedural or object oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or device (e.g., ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The system can also be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
[00217] The signature patterns and databases thereof can be provided in a variety of media to facilitate their use. “Media” refers to a manufacture that contains the signature pattern information of the present invention. The databases of the present invention can be recorded on computer readable media, e.g., any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; and hybrids of these categories such as magnetic / optical storage media. One of skill in the art can readily appreciate how any of the presently known computer readable mediums can be used to create a manufacture comprising a recording of the present database information. “Recorded” refers to a process for storing information on computer readable medium, using any such methods as known in the art. Any convenient data storage structure can be chosen, based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g. word processing text file, database format, etc.
[00218] In some embodiments, the methods disclosed herein, are performed on one or more computers in a distributed computing system environment (e.g., in a cloud computing environment). In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared set of configurable computing resources. Cloud computing can be employed to offer on-demand access to the shared set of configurable computing resources. The shared set of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly. A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“laaS”). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed. Example Computer
[00219] FIG. 5 illustrates an example computer for implementing the entities shown in FIGs. 1A-1C, 2A, 3A-3G, and 4A-4B. In particular embodiments, the example computer 500 can represent computational system 202 described in FIG. 2A. The computer 500 includes at least one processor 502 coupled to a chipset 504. The chipset 504 includes a memory controller hub 520 and an input / output (I / O) controller hub 422. A memory 506 and a graphics adapter 512 are coupled to the memory controller hub 520, and a display 518 is coupled to the graphics adapter 512. A storage device 508, an input device 514, and network adapter 516 are coupled to the I / O controller hub 522. Other embodiments of the computer 500 have different architectures.
[00220] The storage device 508 is a non-transitory computer-readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or a solid-state memory device. The memory 506 holds instructions and data used by the processor 502. The input device 514 is a touch-screen interface, a mouse, track ball, or some combination thereof, and is used to input data into the computer 500. The keyboard 510 may be another device for inputting data into the computer 500. In some embodiments, the computer 500 may be configured to receive input (e.g., commands) from the input device 514 via gestures from the user. The graphics adapter 512 displays images and other information on the display 518. The network adapter 516 couples the computer 500 to one or more computer networks.
[00221] The computer 500 is adapted to execute computer program modules for providing functionality described herein. As used herein, the term “module” refers to computer program logic used to provide the specified functionality. Thus, a module can be implemented in hardware, firmware, and / or software. In one embodiment, program modules are stored on the storage device 508, loaded into the memory 506, and executed by the processor 502. A module can be implemented as computer program code processed by the processing system(s) of one or more computers. Computer program code includes computerexecutable instructions and / or computer-interpreted instructions, such as program modules, which instructions are processed by a processing system of a computer. Generally, such instructions define routines, programs, objects, components, data structures, and so on, that, when processed by a processing system, instruct the processing system to perform operations on data or configure the processor or computer to implement various components or data structures in computer storage. A data structure is defined in a computer program and specifies how data is organized in computer storage, such as in a memory device or a storage device, so that the data can accessed, manipulated, and stored by a processing system of a computer.
[00222] The types of computers 500 used by the entities of FIG. IC can vary depending upon the embodiment and the processing power required by the entity. For example, the tumor heterogeneity system 170 can run in a single computer 500 or multiple computers 500 communicating with each other through a network such as in a server farm. The computers 500 can lack some of the components described above, such as graphics adapters 512, and displays 518. Kit Implementation
[00223] Also disclosed herein are kits for performing a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints. Such kits can include equipment to draw a sample from a patient. For example, kits can include syringes and / or needles for obtaining a sample from a patient. Kits can include detection reagents for determining marker information using the sample obtained from the patient.
[00224] For example, detection reagents can include antibody reagents for performing a protein immunoassay. As another example, detection reagents can be a set of primers that, when combined with the sample, allows detection of a plurality of sites in cell-free DNA in the sample. In particular embodiments, the detection reagents enable detection of methylated or unmethylated target sites (e.g., methylated or unmethylated informative CpGs including one or more CGIs selected from Tables 1-4, or one or more CpGs within at least a portion of a region in Tables 1-4). Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. For example, the detection reagents may be primers that target specific known sequences of target sites, thereby enabling nucleic acid amplification of the target sites. Thus, the use of the detection reagents results in generation of methylation information of the patient corresponding to the target sites.
[00225] A kit can include instructions for use of one or more sets of detection reagents. For example, a kit can include instructions for performing at least one detection assay such as a nucleic acid amplification assay (e.g., polymerase chain reaction assay including any of realtime PCR assays, quantitative real-time PCR (qPCR) assays, allele-specific PCR assays, and reverse-transcription PCR assays), nucleic acid sequencing (e.g., targeted gene sequencing, targeted amplicon sequencing, whole genome sequencing, or whole genome bisulfite sequencing), hybrid capture, an immunoassay, a protein-binding assay, an antibody-based assay, an antigen-binding protein-based assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), reporter assays, flow cytometry, a protein array, a blot, a Western blot, nephelometry, turbidimetry, chromatography, NMR, mass spectrometry, LC-MS, UPLC-MS / MS, enzymatic activity, proximity extension assay, and an immunoassay selected from RIA, immunofluorescence, immunochemiluminescence, immunoelectrochemiluminescence, immunoelectrophoretic, a competitive immunoassay, and immunoprecipitation.
[00226] Kits can further include instructions for accessing computer program instructions stored on a computer storage medium. In various embodiments, the computer program instructions, when executed by a processor of a computer system, cause the processor to perform one or more intra-individual analyses, generate background corrected methylation information, and / or track tumor heterogeneity across two or more timepoints.
[00227] In various embodiments, the kits include instructions for practicing the methods disclosed herein (e.g., performing an assay, screen, or diagnostic assay). These instructions can be present in the kits in a variety of forms, one or more of which can be present in the kit. One form in which these instructions can be present is as printed information on a suitable medium or substrate, e.g., a piece or pieces of paper on which the information is printed, in the packaging of the kit, in a package insert, etc. Yet another means would be a computer readable medium, e.g., diskette, CD, hard-drive, network data storage, etc., on which the information has been recorded. Yet another means that can be present is a website address which can be used via the internet to access the information at a removed site. Any convenient means can be present in the kits. Systems
[00228] Further disclosed herein are systems for performing a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints. In various embodiments, such a system can include one or more sets of detection reagents for determining genomic information using a sample obtained from the patient, an apparatus configured to receive a mixture of the one or more sets of detection reagents and the sample WO 2025 / 147572 PCT / US2025 / 010181 obtained from a subject to generate methylation information of the subject, and a computer system communicatively coupled to the apparatus to generate background-corrected methylation information and / or to track the change in tumor heterogeneity.
[00229] The one or more sets of detection reagents enable the determination of marker information using the sample obtained from the patient. For example, detection reagents can include antibody reagents for performing a protein immunoassay. For example, detection reagents can be a set of primers that, when combined with the sample, allows detection of a plurality of sites in cell-free DNA in the sample. In particular embodiments, the detection reagents enable detection of methylated or methylated target sites (e.g., methylated or unmethylated informative CpGs including one or more CGI’s selected from Tables 1-4 or one or more CpGs within at least a portion of a region in Tables 1-4). Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety.
[00230] The apparatus is configured to determine the methylation information from a mixture of the detection reagents and sample. For example, the apparatus can be configured to perform one or more of a nucleic acid amplification assay (e.g., polymerase chain reaction assay), nucleic acid sequencing (e.g., targeted gene sequencing, whole genome sequencing, or whole genome bisulfite sequencing), and hybrid capture to determine methylation information.
[00231] The mixture of the detection reagents and sample may be presented to the apparatus through various conduits, examples of which include wells of a well plate (e.g., 96 well plate), a vial, a tube, and integrated fluidic circuits. As such, the apparatus may have an opening (e.g., a slot, a cavity, an opening, a sliding tray) that can receive the container including the reagent test sample mixture and perform a reading. Examples of an apparatus include one or more of a sequencer, an incubator, plate reader (e.g., a luminescent plate reader, absorbance plate reader, fluorescence plate reader), a spectrometer, or a spectrophotometer.
[00232] The computer system, such as example computer 500 described in FIG. 5, communicates with the apparatus to receive the methylation information. The computer system generates background-corrected methylation information and can further track the change in tumor heterogeneity (e.g., based on the change of the background-corrected methylation information across two or more timepoints). EXAMPLES
[00233] Below are examples of specific embodiments for carrying out the present invention. The examples are offered for illustrative purposes only and are not intended to limit the scope of the present invention in any way. Efforts have been made to ensure accuracy with respect to numbers used (e.g., percentages, etc.), but some experimental error and deviation should be allowed for. Example 1: Overall performance of two-tier screening and diagnosis of patients with protstate cancer
[00234] FIG. 6 shows example performance of different tiers of the multiple tier analysis for diagnosing individuals with cancer (e.g., prostate cancer). Here, the process begins with 19 million individuals who underwent testing. At a 2% incidence rate, of the 19 million individuals, 380,000 are true positives, and 18.6 million are true negatives.
[00235] The multi-tiered analysis involves performing a screen by analyzing methylation data (generated via an assay) of the patients. Here, the screen is designed to achieve 80% sensitivity and 95% specificity, thereby identifying 1.2 million out of the original 19 million individuals as at risk for prostate cancer. Additionally, the screen identifies 17.8 million out of the original 19 million individuals as not at risk for prostate cancer. Thus, these 17.8 million individuals need not undergo further analysis. Altogether, the screen achieves a 25% positive predictive rate and a 99% negative predictive rate.
[00236] The 1.2 million individuals identifies as at risk for prostate cancer further undergo a second test in the form of the second analysis. The second analysis achieves a 90% sensitivity and a 95% specificity. Of the 1.2 million individuals, -320,000 individuals are identified as having prostate cancer. This represents a 85% positive predictive rate as 273,600 individuals were true positives and 47,000 were false positives. Additionally, the second analysis identifies 945,000 negatives, of which 884,450 were true negatives, and 30,400 were false negatives, thereby representing a 97% negative predictive value.
[00237] Altogether, the overall performance of the multi-tier screen and second analysis includes 72% sensitivity, 99.9% specificity, 85% positive predictive value, and 99.4% negative predictive value.
[00238] Example steps for performing the multiple-tier analysis shown in FIG. 6 are detailed below. Prepare tarset specimen
[00239] The target specimen type (e.g. DNA, RNA, protein, exosomes, metabolites, etc.) is isolated from a patient’s biological source (e.g. tissue, blood, plasma, serum, saliva, feces, etc.). That specimen can be isolated by a CRO or private or service laboratory or hospital or isolated internally using an internal procedure. Target specimens are assayed for quality and quantity measurements. Phase 1 testins
[00240] Phase 1 testing is a relatively quick, non-invasive assay with simple technology, using small amounts of the target specimen. The result of this assay can be both qualitative and quantitative. Phase 1 testing is typically lower specificity (e.g. 95% specificity, 5% false positives) but higher sensitivity (e.g. 80% sensitivity, 20% false negatives) in order to screen a large proportion of the testing population rapidly and inexpensively. The phase 1 assay will overall increase the incidence of the target population (e.g. diseased) for the phase 2 assay, which will then increase the positive predictive value (PPV). Examples of the Phase 1 assay include but are not limited to ELISA assays, PCR assays, Real-time PCR assays, Quantitative real-time PCR (qPCR) assays, Allele-specific PCR assays, Reverse-transcription PCR assays and reporter assays. Phase 2 testing
[00241] Phase 2 testing is a more complex, potentially invasive assay with complex technology, potentially using larger amounts of the target specimen. The result of this assay is both qualitative and quantitative. Phase 2 testing is typically higher specificity (e.g. 95% specificity, 10% false positives) but lower sensitivity (e.g. 90% sensitivity, 10% false negatives) in order to limit false positives. By screening out a large volume of the testing population, the target population has higher target incidence than the general population, which increases positive predictive value (PPV). Phase 2 Protocol
[00242] Examples of the phase 2 assay include but are not limited to Next Generation Sequencing assays utilizing target enrichment technologies, targeted amplicon sequencing technologies, whole genome sequencing, and whole genome bisulfite sequencing.
[00243] The target specimen for library construction is dsDNA isolated from formalin-fixed paraffin-embedded (FFPE) tissue. Alternatively, cfDNA is isolated from blood. For FFPE, the dsDNA is first mechanically sheared by the Covaris instrument utilizing adaptive focused acoustics to a target insert size of 200 base pairs. Post-shearing, a solid-phase reversible immobilization (SPRI) selection is done to remove smaller DNA fragments remaining in solution. For blood DNA, cfDNA is isolated. The fragmented DNA is then end-repaired and A-tailed (ERAT) to produce 5’-phosphorylated, 3’-dA-tailed dsDNA fragments. After ERAT, dsDNA unique dual index adapters with 3’-dTMP overhangs are then ligated to 3’-dA-tailed dsDNA fragments. Indices allow for sample multiplex for the downstream assay. Postligation, a solid-phase reversible immobilization (SPRI) selection is done to remove unwanted DNA fragments, excess adapters and molecules. PCR amplification is performed with a high-fidelity, low-bias polymerase at 10 cycles. Post-PCR, a SPRI selection is done to remove unwanted DNA fragments, excess primers, excess adapters and excess molecules. After library construction, the library quality and quantity are evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively.
[00244] Libraries that pass quality control checks move forward to target enrichment through hybridization capture. Target enrichment by hybridization capture is defined as a positive selection strategy to enrich low abundance regions of interest from NGS libraries, allowing for more accurate sequencing analysis of these target regions. Indexed libraries are multiplexed and hybridized to a custom, sequence specific, biotinylated probeset. The vast excess of probes drives their hybridization to complementary library fragments. The library fragment-biotinylated probe hybrid is pulled down by streptavidin beads, thereby capturing the target regions of interest. The streptavidin bead-bound library is sequentially washed with buffers to remove non-specifically associated library fragments. Following washes and recovery of captured libraries, samples are enriched for on target fragments and depleted for off-target fragments. Depletion of off-target fragments reduces overall library yield, requiring post-capture library amplification by PCR. The final amplified library is enriched for regions of interest. The hybrid captured library quality and quantity is evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively. Additionally, the enrichment efficiency is evaluated using an iSeq Sequencing run and calculation of percent of reads within target enrichment panel. Measuring percent on-target is a good first approximation of target enrichment efficiency because the reads aligning to the target enrichment (bait) region indicate efficient hybridization and subsequent capture.
[00245] Target enriched libraries that pass quality control checks move forward to NovaSeq sequencing. Captured libraries with non-overlapping indices from library construction are pooled to multiplex for sequencing. Sequencing is completed on the NovaSeq 6000 instrument using paired end 150x150 base sequencing with a 10% PhiX spike-in. Sequencing data generated is then demultiplexed utilizing the assigned index, aligned to the human genome and trimmed to enrich for insert sample data only. This cleaned-up data is then processed through a quality pipeline to collapse duplicate reads and evaluate the sequencing data generated. Once the data is collapsed, the data is processed through a proprietary biomarker analysis pipeline to identify differences from the reference alignment (e.g. mutations, chemical modifications, etc). A report is then generated with the specific biomarker analysis per sample that confirms the results of the phase 1 assay or identifies true false positives from the phase 1 assay. Phase 1 Protocol:
[00246] An example protocol of an Allele-specific Real-Time PCR assay is as follows: 1. This assay runs DNA samples in triplicate with 2ng input in 5uL for the reference and mutation assays. 2. Combine 900nmol / L unspecific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the reference control assay. 3. Combine 450nmol / L allele-specific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the mutation assay. 4. Mix each reaction 10X and centrifuge to collect volume at the bottom of the well or tube. 5. Run the real-time PCR on a calibrated Real-Time PCR system under the following conditions: (1) 95°C for 10 minutes followed by (2) 50 cycles of90°C for 15 seconds and 60°C for 1 minute with fluorescence detection using FAM / VIC fluorophores. 6. Cycle threshold (Ct) values are recorded by the system and exported into an analysis program (e.g. Excel). 7. Average the Ct values between sample replicates for the reference and mutation assays. 8. Calculate the ACt between the sample average allele-specific Ct minus the sample average unspecific (reference) Ct. 9. Positive mutation results are identified by the ACt cut off > 3 cycles and will move forward to phase 2 testing.
[00247] Allele-specific real-time PCR can be performed by combining library DNA with PCR reagents and primers specific for target sequences. The primers are designed to have single-base discrimination between tumor and non-tumor sequences. Perform real-time PCR (or digital PCR) for 30-50 cycles and monitor the output for signal via fluorescence from amplified target DNA or probe sequence. Cycle threshold values (Ct) are recorded and exported for analysis. The delta-Ct between negative control, positive control, and sample are calculated to determine presence or absence of target tumor sequences. Slight modifications of this protocol will allow for end-point PCR detection of RNA or DNA of tumor sequences. Phase 1 detection will be designed to remove 90-95% of non-cancer patient samples from moving forward for further testing.
[00248] ELISA assay detection of target molecules can be performed by coating an immunoassay well with monoclonal antibody designed to specifically detect target molecules, followed by blocking against non-specific binding. Next, target sample is introduced to the well, incubated and washed away. Any bound target can then be bound by a polyclonal antibody specific for the target. Additional secondary antibodies with color or fluorescent tags can be used to detect the presence of target molecules. Interpreting results for phase 1 and phase 2 assays
[00249] Two positive signals from the phase 1 assay and phase 2 assay can be determined as a true positive sample with an 85% probability of being accurate.
[00250] One negative signal from the phase 1 assay can be determined as a true negative sample with a 99% probability of being accurate.
[00251] One positive signal from the phase 1 assay and one negative signal from the phase 2 assay can be determined as an indeterminate sample with a 97% probability of a false positive in phase 1 assay. Example 2: Two-tier Analysis Achieves Improved Performance in Comparison to Single Tier Analysis
[00252] Samples were obtained from patients of a patient population with an assumed 1.3% cancer prevalence. In total, 1046 samples obtained from the patients underwent either a single tier analysis or a two-tier analysis. The performance metrics (as measured by specificity, positive predictive value (PPV), and negative predictive value (NPV)) of each of the methodologies were determined.
[00253] Reference is now made to FIG. 7, which depicts performance of a single tier and two-tier analysis of a population involving 1046 samples. The Tier 1 analysis focused on analyzing signal from a subset of the 4059 CGIs shown in Tables 2 and 3. In particular, 130 regions were analyzed to estimate tumor content according to methylation statuses of the regions, and estimated tumor content was used to distinguish patients that were negative or not negative for cancer. Logistic regression was performed to assess performance at 90% specificity (e.g., true negative rate reported as a proportion of correctly identified negatives). Performance was estimated to be about 63% sensitivity. For the single tier analysis (including only the Tier 1 analysis), it achieved a PPV (defined as number of true positives divided by the sum of true positives and false positives) of 0.0761 and a NPV (defined as true negative rate divided by the sum of true negatives and false negatives) of 0.9946. Thus, the single tier analysis was capable of successfully screening out a large proportion of samples that were negative for cancer. However, based on the low PPV, it had room for improvement in identifying samples that were true positives. The single tier analysis (including only a Tier 2 analysis) was additionally performed. Specifically, for each sample, signal of the 4059 CGIs was analyzed using a machine learning algorithm to distinguish samples having a cancer signal from samples not having a cancer signal. The single tier (Tier 2 analysis) achieved a PPV of 0.1858 and a NPV of 0.9969. Thus, the more costly Tier 2 analysis achieved a higher PPV in comparison to the less costly Tier 1 analysis without sacrificing the NPV metric.
[00254] Referring to the two-tier analysis, it involved performing the Tier 1 analysis (analyzing subset of top features) and samples deemed to be negative for cancer were screened out. An additional Tier 2 analysis was then performed. Specifically, for each sample, signal of the 4059 CGIs were analyzed using a machine learning algorithm to distinguish samples having a cancer signal from samples not having a cancer signal. Here, the Tier 2 analysis achieved a high specificity of 96%. For the two-tier analysis (including both the Tier 1 and Tier 2 analyses), the methodology achieved a PPV (defined as number of true positives divided by the sum of true positives and false positives) of 0.2421 and a NPV (defined as true negative rate divided by the sum of true negatives and false negatives) of 0.9942. Here, the two-tier analysis exhibited a significant improvement in comparison to the single-tier analysis. Specifically, the two-tier analysis achieved a higher specificity (e.g., 96% versus 90%). Furthermore, the two-tier analysis exhibited an improved PPV (0.2421 versus 0.0761) without adversely impacting the NPV (0.9942 versus 0.9946). Example 3: Example Samples and Assays for Conducting an Intra-Individual Analysis
[00255] Blood samples are obtained from individuals. FIG. 8 shows an example sample from which target nucleic acids and reference nucleic acids are obtained. Shown on the left in FIG. 8 is a tube of blood obtained from an individual, the tube including diluted peripheral blood of the individual and separation medium. The tube undergoes centrifugation to separate different components of the diluted peripheral blood. For example, at a speed of 2200 rpm, the diluted peripheral blood is fractionated into plasma (including platelets, cytokines, hormones, and electrolytes), peripheral blood mononuclear cells (PBMCs), the separation medium, and polymorphonuclear cells. Here, target nucleic acids in the form of cell free DNA is found in the plasma whereas reference nucleic acids in the form of cellular genomic DNA is found in PBMCs.
[00256] Examples of an assay for generating sequence information from the target nucleic acids and the reference nucleic acids include but are not limited to Allele-specific PCR assays, Next Generation Sequencing assays, such as target enrichment technologies, targeted amplicon sequencing technologies, and whole genome sequencing.
[00257] An example protocol of an Allele-specific Real-Time PCR assay is as follows: 1. This assay runs all cfDNA samples in triplicate with 2ng input in 5uL for the reference and hypermethylation assays. 2. Combine 900nmol / L unspecific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the reference control assay. 3. Combine 450nmol / L allele-specific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the mutation assay. 4. Mix each reaction 10X and centrifuge to collect volume at the bottom of the well or tube. 5. Run the real-time PCR on a calibrated Real-Time PCR system under the following conditions: (1) 95°C for 10 minutes followed by (2) 50 cycles of90°C for 15 seconds and 60°C for 1 minute with fluorescence detection using FAM / VIC fluorophores. 6. Cycle threshold (Ct) values are recorded by the system and exported into an analysis program (e.g. Excel). 7. Average the Ct values between sample replicates for the reference and mutation assays. 8. Calculate the DCt between the sample average allele-specific Ct minus the sample average unspecific (reference) Ct. 9. Positive hypermethylation results are identified by the DCt cut off > 3 cycles and will be compared to the patients individual PBMC natural signal.
[00258] An example protocol of an Allele-specific Real-Time PCR assay is as follows: Allele-specific real-time PCR can be performed by combining library from cfDNA with PCR reagents and primers specific for target sequences. The primers are designed to have singlebase discrimination between tumor and non-tumor sequences. Perform real-time PCR (or digital PCR) for 30-50 cycles and monitor the output for signal via fluorescence from amplified target DNA or probe sequence. Cycle threshold values (Ct) are recorded and exported for analysis. The delta-Ct between negative control, positive control, and sample are calculated to determine presence or absence or absence of target tumor sequences. Slight modifications of this protocol will allow for end-point PCR detection of RNA or DNA of tumor sequences.
[00259] An example protocol of a next generation sequencing (NGS) Target Enrichment assay is as follows: The target specimen for library construction is dsDNA isolated from PBMCs. The dsDNA is first mechanically sheared by the Covaris instrument utilizing adaptive focused acoustics to a target insert size of 200 base pairs. Post-shearing, a solidphase reversible immobilization (SPRI) selection is done to remove smaller DNA fragments remaining in solution. The fragmented DNA is then end-repaired and A-tailed (ERAT) to produce 5’-phosphorylated, 3’-dA-tailed dsDNA fragments. After ERAT, dsDNA unique dual index adapters with 3’-dTMP overhangs are then ligated to 3 ’-dA-tailed dsDNA fragments. Indices allow for sample multiplex for the downstream assay. Post-ligation, a solid-phase reversible immobilization (SPRI) selection is done to remove unwanted DNA fragments, excess adapters and molecules. PCR amplification is performed with a high-fidelity, low-bias polymerase at 10 cycles. Post-PCR, a SPRI selection is done to remove unwanted DNA fragments, excess primers, excess adapters and excess molecules. After library construction, the library quality and quantity are evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively.
[00260] Libraries that pass quality control checks move forward to target enrichment through hybridization capture. Target enrichment by hybridization capture is defined as a positive selection strategy to enrich low abundance regions of interest from NGS libraries, allowing for more accurate sequencing analysis of these target regions. Indexed libraries are multiplexed and hybridized to a custom, sequence specific, biotinylated probeset. The vast excess of probes drives their hybridization to complementary library fragments. The library fragment-biotinylated probe hybrid is pulled down by streptavidin beads, thereby capturing the target regions of interest. The streptavidin bead-bound library is sequentially washed with buffers to remove non-specifically associated library fragments. Following washes and recovery of captured libraries, samples are enriched for on target fragments and depleted for off-target fragments. Depletion of off-target fragments reduces overall library yield, requiring post-capture library amplification by PCR. The final amplified library is enriched for regions of interest. The hybrid captured library quality and quantity is evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively. Additionally, the enrichment efficiency is evaluated using an iSeq Sequencing run and calculation of percent of reads within target enrichment panel. Measuring percent on-target is a good first approximation of target enrichment efficiency because the reads aligning to the target enrichment (bait) region indicate efficient hybridization and subsequent capture.
[00261] Target enriched libraries that pass quality control checks move forward to NovaSeq sequencing. Captured libraries with non-overlapping indices from library construction are pooled to multiplex for sequencing. Sequencing is completed on the NovaSeq 6000 instrument using paired end 150x150 base sequencing with a 10% PhiX spike-in. Sequencing data generated is then demultiplexed utilizing the assigned index, aligned to the human genome and trimmed to enrich for insert sample data only. This cleaned-up data is then processed through a quality pipeline to collapse duplicate reads and evaluate the sequencing data generated. Once the data is collapsed, the data is processed through a proprietary analysis pipeline to identify differences from the reference alignment (e.g. mutations, chemical modifications, etc.). A report is then generated with the specific signal informative for determining presence or absence of cancer. Attorney Docket No. ELCr-022WU TABLE 1 - List of CGIs Pos (hgl9 coordinates) chrl3:108518334-108518633 chr6:137242315-137245442 chr2:177016416-177016632 chr5:2738953-2741237 chr4:111553079-111554210 chrl5:96909815-96910030 chr6:42072032-42072701 chrl0:123922850-123923542 chrl6:86612188-86613821 chrl9:47151768-47153125 chrl:110610265-110613303 chr5:3594467-3603054 chr9:126773246-126780953 chr3:138656627-138659107 chr4:4859632-4860191 chrl0:118895963-118898037 chr7:103086344-103086840 chrl9:407011-409511 chrl0:22764708-22767050 chrl6:86549069-86550512 chr9:96713326-96718186 chr8:139508795-139509774 chr2:73143055-73148260 chr8:26721642-26724566 chr9:129386112-129389231 chrl2:49483601-49484255 chrl6:54325040-54325703 chr8:72468560-72469561 chrl8:70533965-70536871 chr9:98111364-98112362 chrl:50882997-50883426 chrl0:88122924-88127364 chrll:31839363-31839813 chrl0:101290025-101290338 chr6:41528266-41528900 chrl6:51183699-51188763 chr5:140346105-140346931 chr9:23820691-23822135 chr20:690575-691099 chrl:177133392-177133846 chr5:45695394-45696510 chr2:45395869-45398186 chr20:48184193-48184833 chr6:6002471-6005125 chrl4:101192851-101193499 chr8:4848968-4852635 chr8:53851701-53854426 chrl2:186863-187610 chr5:54519054-54519628 chr6:108485671-108490539 chr3:157815581-157816095 chrll:626728-628037 chr2:177012371-177012675 chrl7:59531723-59535254 chrl6:55364823-55365483 chr8:99960497-99961438 chr7:42267546-42267823 chrl7:14202632-14203258 chrl0:102891010-102891794 chr5:174158680-174159729 chrl4:33402094-33404079 chr2:177036254-177037213 chrl0:106399567-106402812 chr6:166579973-166583423 chrll:123066517-123066986 chrll:44327240-44327932 chrl4:95237622-95238211 chr9:102590742-102591303 chrl5:76630029-76630970 chr4:24801109-24801902 chr8:97169731-97170432 chr3:6902823-6903516 chr22:48884884-48887043 chrl5:45408573-45409528 chr9:100610696-100611517 chr4:174448333-174448845 chrl6:20084707-20085305 chr4:174439812-174440249 chr6:10381558-10382354 chrl5:35046443-35047480 chrl0:119494493-119494991 chr5:72676120-72678421 chrll:44325657-44326517 chrl7:46670522-46671458 chrl4:92789494-92790712 chr4:174459200-174460054 chr2:80549578-80549798 chr7:153748407-153750444 chr6:1389139-1391393 chrl6:49314037-49316543 chr2:105459127-105461770 chr21:38079941-38081833 chr4:174427891-174428192 chrl4:60973772-60974123 chr8:99985733-99986983 chr2:63281034-63281347 chrl2:101109863-101111622 chrl:119549144-119551320 chr5:38257825-38259136 chr5:54522302-54523533 chrl:165324191-165326328 chrl5:33602816-33604003 chrl0:118030732-118034230 chr2:45240372-45241579 chr4:174430386-174430861 chr6:50810642-50810994 chr5:122430676-122431443 chrl0:109674196-109674964 chr8:97172634-97173880 chr8:11536767-11538961 chr5:180486154-180486892 chr2:38301276-38304518 chrl0:1778784-1780018 chrl2:54424610-54425173 chrl7:46669434-46669811 chrll:8190226-8190671 chr8:25900562-25905842 chrl2:81102034-81102716 chr7:27199661-27200960 chrl0:119311204-119312104 chrl2:130387609-130389139 chr7:155258827-155261403 chr6:117591533-117592279 chrl0:111216604-111217083 chrl:29585897-29586598 chr2:144694666-144695180 chrl2:48397889-48398731 chr5:2748368-2757024 chrl2:114845861-114847650 chr2:80529677-80530846 chr5:1874907-1879032 chr6:100905952-100906686 chrl5:96904722-96905050 chr5:134374385-134376751 chr2:66652691-66654218 chrl2:54440642-54441543 chr6:108495654-108495986 chrl7:70112824-70114271 chr3:87841796-87842563 chr7:96650221-96651551 chr4:110222970-110224257 chr6:78172231-78174088 chr7:155164557-155167854 chrl2:113900750-113906442 chr9:112081402-112082905 chrl2:114886354-114886579 chr5:3590644-3592000 chr2:119592602-119593845 chr20:21485932-21496714 chrl8:11148307-11149936 chrl7:46824785-46825372 chrl0:100992156-100992687 chrl4:36986362-36990576 chrl8:55094825-55096310 chrl5:96895306-96895729 chrl7:36717727-36718593 chr2:223183013-223185468 chr7:30721372-30722445 chrl:53527572-53528974 chrl8:56939624-56941540 chr5:175085004-175085756 chrl0:50817601-50820356 chrl4:60975732-60978180 chrl5:89920793-89922768 chr9:122131086-122132214 chrl:217311467-217311773 chrl4:38724254-38725537 chrl4:61103978-61104663 chrl8:73167402-73167920 chrl:50880916-50881516 chr2:241758141-241760783 chrll:31825743-31826967 chr7:27260101-27260467 chr20:41817475-41819212 chr3:238391-240140 chr7:121950249-121950927 chr5:72526203-72526497 chrl5:96903311-96903711 chrl0:26504383-26507434 chr6:100915602-100915883 chrl:18962842-18963481 chr3:127794369-127796136 chr7:27203915-27206462 chr8:25899335-25899692 chrl2:114838312-114838889 chr6:38682949-38683265 chrll:31841315-31842003 chr4:174451828-174452962 chr9:129372737-129378106 chr2:176964062-176965509 chr2:176931575-176932663 chrl2:114833911-114834210 chrll:79148358-79152200 chr2:177024501-177025692 chr5:172672311-172672971 chr7:27291119-27292197 chrl:180198119-180204975 chrl4:37126786-37128274 chr2:200333687-200334172 chrl4:58331676-58333121 chr3:147131066-147131333 chrl3:109147798-109149019 chrl4:48143433-48145589 chr6:100905444-100905697 chrl7:14200579-14200996 chr6:1379693-1380014 chrl:34642382-34643024 chr2:119599059-119599299 chr2:119613031-119615565 chr4:85413997-85414874 chr9:17906419-17907488 chrl2:29302034-29302954 chr20:10200088-10200384 chr8:57358126-57359415 chrl0:63212495-63213009 chr2:176936246-176936809 chrll:20618197-20619920 chrl8:19744936-19752363 chrl4:29234889-29235908 chrl7:46673532-46674181 chr4:144620822-144622218 chrl6:82660651-82661813 chr3:192125821-192127994 chr2:119599458-119600966 chr22:44257942-44258612 chrl9:13616752-13617267 chr3:147138916-147139564 chr9:969529-973276 chrl8:55103154-55108853 chr4:174422024-174422443 chr4:57521621-57522703 chrl5:79724099-79725643 chrl4:37135513-37136348 chrl0:23480697-23482455 chr2:45169505-45171884 chrl8:30349690-30352302 chr6:99291327-99291737 chr9:21970913-21971190 chr4:107146-107898 chrl2:117798076-117799448 chr2:219736132-219736592 chrl0:118892161-118892639 chrll:27743472-27744564 chrl2:65218245-65219143 chrl2:75601081-75601752 chr7:54612324-54612558 chr6:100912071-100913337 chrl0:102905714-102906693 chr8:87081653-87082046 chr6:50818180-50818431 chrl:91189139-91189400 chr2:118981769-118982466 chrl0:50602989-50606783 chrl7:59528979-59530266 chr4:147559205-147561901 chrl:4713989-4716555 chrl3:102568425-102569495 chrl6:6068914-6070401 chr22:29709281-29712013 chrl0:100993820-100994188 chr6:391188-393790 chr2:176977284-176977540 chr4:4868440-4869173 chr6:137809342-137810204 chrl2:54321301-54321721 chr2:105468851-105473488 chr8:55366180-55367628 chrl2:72665683-72667551 chr4:54966163-54968063 chr5:134366913-134367438 chrl:226075150-226075680 chr20:17206528-17206952 chr4:172733734-172735118 chrl8:55019707-55021605 chr2:162279835-162280709 chr6:1381743-1385211 chr7:103968783-103969959 chr6:150358872-150359394 chr2:119914126-119916663 chr7:27278945-27279469 chrl2:114851957-114852360 chrl6:24267040-24267527 chr6:7229877-7230865 chr2:45227644-45228783 chr4:174450046-174451469 chr4:154712073-154712706 chr3:22413492-22414365 chr20:21694472-21695344 chr6:1378445-1379318 chr8:70981873-70984888 chrl2:53107912-53108471 chrl0:102996034-102996646 chr3:157821232-157821604 chr4:111554965-111555504 chrl3:58206526-58208930 chrl0:22634000-22634862 chr9:22005887-22006229 chr5:159399004-159399928 chr2:31805293-31806403 chr6:100903491-100903713 chr5:77268350-77268787 chrl4:85997468-85998637 chr5:92923487-92924497 chrll:64480199-64481344 chrl3:28366549-28368505 chr5:77805753-77806313 chr9:79633326-79636030 chr4:93226348-93227007 chr2:223170486-223171140 chrl:91172102-91172771 chrl:1181756-1182470 chr8:65281903-65283043 chrl0:948 25546-94826320 chr6:108491033-108491410 chr21:38076762-38077685 chrl:91183240-91184540 chr3:147136903-147137328 chrl5:96911511-96911808 chrl4:57274607-57276840 chrl3:112726281-112728419 chr2:171672310-171675447 chr8:11559596-11562956 chrl0:48438411-48439320 chrl8:59000683-59001692 chrl5:91642908-91643702 chr5:3592391-3592644 chrl9:56988313-56989741 chr6:26614013-26614851 chrll:27742059-27742273 chr3:147113608-147114479 chrl4:57264638-57265561 chr7:155302253-155303158 chrll:31848487-31848776 chrl6:54970301-54972846 chrl9:30715549-30715753 chr9:96710811-96711717 chrl8:77557780-77558948 chr20:21686199-21687689 chrll:31847132-31847958 chrl6:86530747-86532994 chrl:203044722-203045390 chrl5:53096014-53096482 chr7:97361132-97363018 chrl4:29236835-29237832 chrl3:79182859-79183880 chrll:69517840-69519929 chrl:231296559-231297345 chrl9:8675333-8675699 chrl:63795363-63796140 chr4:90228714-90229010 chr3:62362610-62363082 chrl9:5827754-5828405 chrl0:125732220-125732843 chr9:136293566-136294160 chrl:63782394-63790471 chr4:4867386-4867673 chr9:133534534-133542394 chrl5:100913438-100914022 chrl0:101279941-101280382 chrl3:53419897-53422872 chrl:77747314-77748224 chrl4:36974548-36975425 chrl2:57618769-57619402 chr7:49813008-49815752 chr4:188916605-188916876 chrll:31831620-31839038 chr8:132052203-132054749 chr2:237071794-237078762 chr20:39994545-39995810 chrll:132812662-132813075 chr5:170735169-170739863 chrl:221051966-221053673 chr5:72529099-72529976 chrl4:36973169-36973740 chr4:158141404-158141836 chrl4:103655241-103655928 chrl:65731411-65731849 chrl:38218190-38218977 chr3:128719865-128721245 chrl5:33009530-33011696 chr2:162275161-162275596 chr7:155241323-155243757 chrl9:46001830-46002686 chr6:137814355-137815202 chr7:70596228-70598382 chrl5:96959341-96960531 chrl6:66612749-66613412 chr6:110299365-110301267 chrl5:27215951-27216856 chrll:88241710-88242562 chr2:124782252-124783255 chrl7:70111979-70112308 chr2:63283936-63284147 chrl7:46800945-46801288 chr6:1393049-1394170 chr3:137489594-137491004 chrl5:60296135-60298520 chrl2:106979429-106981086 chrl2:54360374-54360660 chrl4:36991594-36992488 chr4:156129168-156130209 chr4:54975387-54976202 chr3:137482964-137484454 chrl0:118893527-118894432 chrl8:76737005-76741244 chrl0:110671724-110672326 chr5:71014917-71015715 chr6:50787286-50788091 chrl9:3868586-3869217 chr4:5894071-5895116 chrll:131780328-131781532 chr6:101846766-101847135 chrll:71952112-71952528 chr5:172663616-172664584 chr9:23822412-23822667 chr4:5891981-5892365 chrl:217310749-217311178 chrl0:108923780-108924805 chr6:100038655-100039477 chr7:121945345-121946235 chr3:147126988-147128999 chr7:121956543-121957341 chr4:156680095-156681386 chr4:85404986-85405252 chrl:221064889-221065600 chrl7:73749618-73750178 chr8:55370170-55372525 chr6:70992040-70992912 chrl6:55513220-55513526 chr6:106433984-106434459 chrl4:29254365-29255069 chr6:33655966-33656238 chr9:19788215-19789288 chrll:115630398-115631117 chrl:34628783-34630976 chrl4:101923575-101925995 chrl7:72855621-72858012 chr2:223162946-223163912 chr4:85417659-85420799 chrl:156390403-156391581 chr3:147130342-147130577 chr2:119602616-119604486 chr9:120175253-120177496 chr4:174443365-174443948 chr5:145724294-145724551 chrll:32454874-32457311 chr2:176949511-176949795 chrl:18436551-18437673 chr3:26665950-26666164 chr3:170303044-170303249 chr2:223176493-223177515 chr2:182321761-182323029 chrl8:44789742-44790678 chrl7:46796234-46797292 chrl8:44772992-44775577 chr8:101117922-101118693 chr7:27134097-27134303 chrl0:102507482-102509646 chrl9:39754973-39756540 chr7:26415746-26416891 chrl4:37116188-37117628 chr4:174421347-174421559 chr6:85472702-85474132 chr20:22557517-22559240 chr6:117198089-117198705 chrl0:71331926-71333392 chrl9:36334994-36335321 chr4:46995128-46995872 chr9:135455164-135458586 chr8:65290108-65290946 chrl0:94828102-94829040 chrl:116380359-116382364 chrl5:47476369-47477499 chr3:147115764-147116421 chrl7:59485573-59485780 chrl0:23983366-23984978 chr2:176949993-176950336 chr9:137967110-137967727 chr2:176957054-176958279 chrll:119293320-119293943 chrll:132813562-132814395 chr2:237068071-237068834 chrl0:27547668-27548402 chr4:4866438-4866813 chr21:19617098-19617874 chrl:91185156-91185577 chrl9:15292399-15292632 chrl:145075483-145075845 chr2:19560963-19561650 chrl4:57260878-57262123 chr8:55378928-55380186 chr6:99290279-99290771 chrl9:13124959-13125259 chrl5:27112030-27113479 chr8:145925410-145926101 chrll:124629723-124629926 chr4:109093038-109094546 chr3:62356773-62357315 chrl4:37131181-37132785 chrl0:124905634-124906161 chr7:35296921-35298218 chrl9:36248979-36249307 chrl2:15475318-15475901 chr5:87985470-87985810 chrl2:54423427-54423712 chr7:96653467-96654199 chr2:45155195-45157049 chrl5:96896928-96897301 chrl2:58004982-58005351 chr2:176933131-176933449 chr2:176962179-176962487 chr20:25063838-25065525 chrl2:5153012-5154346 chr3:154146347-154146965 chrl:165323486-165323811 chr21:38065179-38066185 chrl0:119000435-119001530 chrl2:45444202-45445386 chr4:158143296-158144053 chr5:76932317-76933523 chr5:172659049-172660277 chr2:223168653-223169008 chrl:248020330-248021252 chrl8:904578-909574 chrl2:127940451-127940907 chr9:135461934-135462909 chrl7:48041282-48043064 chr4:94755786-94756310 chrl0:130338695-130338994 chr2:119616133-119616826 chr2:177042751-177043444 ch r2:105478600-105479188 chr5:172670829-172671824 chr2:176952695-176953297 chrl3:28549839-28550246 chrl3:112720564-112723582 chr6:100895773-100896062 chr7:136553854-136556194 chr6:127441553-127441760 chrl:119526782-119527192 chrl2:49484920-49485178 chr9:23850910-23851522 chr2:220299483-220300243 chr5:1881924-1887743 chr8:57360585-57360815 chrl8:74961556-74963822 chr5:172660720-172661133 chrl7:75277317-75278172 chrl0:99789614-99791320 chr2:176944087-176948446 chr4:154709512-154710827 chr5:140798757-140799359 chr3:44063314-44063837 chrl5:79574830-79575211 chr2:223161531-223161919 chr6:134210639-134211218 chrl0:102899177-102899489 chrl3:79181944-79182222 chr7:71800757-71802768 chr3:186078710-186080111 chrl:24229115-24229537 chrl6:48844551-48845264 chr7:113724924-113727795 chr22:44726724-44727590 chr4:15779998-15780729 chr4:41869174-41869459 chrl:38941919-38942404 chr2:176971706-176972305 chr2:119607378-119607910 chr5:76934581-76935296 chrl2:103696090-103696418 chr5:63255044-63255407 chrl:221067447-221068185 chr2:119611296-119611881 chrl0:124907283-124911035 chrl2:114878143-114879155 chrl2:49371690-49375550 chrl7:36719544-36719938 chrl7:46696553-46696926 chr3:147142181-147142391 chr8:9762661-9764748 chrl4:74706188-74708192 chr3:12837992-12838359 chr20:37352130-37357372 chrl0:8077829-8078378 chr4:4864456-4864834 chr4:13524062-13526083 chrl:66258440-66258918 chrll:17740789-17743779 chrl2:106975195-106975714 chr9:91792662-91793611 chrl:149333785-149334111 chr3:170303532-170303768 chr5:72594147-72595808 chr5:145725286-145725852 chrl0:23462224-23463889 chr20:21689758-21690048 chrl5:53080458-53083699 chr2:154727906-154728271 chr5:170743178-170744107 chrl0:102899822-102900263 chr5:134368578-134370466 chr2:66808568-66809404 chr7:96651963-96652246 chrl:91190489-91192804 chrl7:75368688-75370506 chr4:185939222-185942747 chr7:43152020-43153340 chrl3:84453664-84453897 chr2:176956504-176956707 chr7:87563342-87564571 chr20:17208550-17208756 chr22:19746924-19747141 chr2:223159725-223160487 chrl2:131200509-131200726 chrl8:44336183-44337110 chr2:63285949-63287097 chr4:13526553-13526770 chrl5:89949373-89951130 chrl9:55815940-55816277 chrl7:50235175-50236466 chrl9:58545115-58545897 chrl2:113592203-113592620 chrl2:115109503-115110061 chr4:164264821-164265772 chrl:2772126-2772665 chr3:71834068-71834653 chrl2:5018585-5021171 chrl5:74419870-74423044 chr3:147108511-147111703 chr5:88185224-88185589 chrl2:54354529-54355491 chrl0:101290625-101291178 chr8:11557852-11558252 chr8:105478672-105479340 chrll:20181200-20182325 chrl9:54483021-54483572 chrl3:112707804-112708696 chrl6:22824616-22826459 chr4:66536065-66536674 chr4:154713537-154714240 chr7:12151220-12151559 chrl2:119212110-119212393 chrl7:14201726-14202052 chr20:21376358-21378245 chrl3:36045931-36046143 chrl5:60287107-60287663 chr9:100613938-100614622 chrl0:102475276-102475579 chr7:121940006-121940648 chr5:37834671-37835128 chrl:197887088-197887791 chrl2:99139386-99139769 chr6:1619093-1621094 chrl2:113917394-113918107 chrl4:24044886-24046760 chr5:77253832-77254049 chr4:85403830-85404524 chr6:166666837-166667541 chrl8:77547965-77549038 chr2:219848919-219850541 chrl7:7832532-7833164 chr5:134363092-134365146 chrl0:103043990-103044480 chr8:97171805-97172022 chr20:57089460-57090237 chrl2:114840853-114841063 chr4:66535193-66535620 chr8:85096759-85097247 chr6:10881846-10882051 chrl3:28498226-28499046 chrl:161695637-161697298 chrll:2890388-2891337 chrl7:5000369-5001205 chrl3:27334226-27335205 chrl0:22623350-22625875 chr2:157185557-157186355 chr7:20370003-20371504 chr4:961347-962155 chrl2:49485766-49485977 chr3:62356119-62356378 chrll:14995128-14995908 chrl2:53359192-53359507 chrl6:51168266-51169110 chrl4:57278709-57279116 chr6:37616722-37617179 chrl8:11750953-11752756 chrl9:45260352-45261809 chrl:119531991-119532196 chrl9:36523391-36523887 chrl2:52652018-52652743 chr8:49468683-49468959 chr8:9760750-9761643 chr7:19146923-19147308 chrl3:32889533-32889900 chr5:140797162-140797701 chr21:42218489-42219222 chrl9:54411376-54411968 chr3:62354291-62355012 chrl2:113590806-113591304 chrl:225865068-225865328 chr7:130790358-130792773 chrl5:53076187-53077926 chrl:214158726-214159080 chrl2:3308812-3310270 chrl :39044059-39044561 chrl0:119312766-119313563 chrl2:65514878-65515863 chrl2:54366815-54369103 chrl2:114885105-114885418 chrl6:2228190-2230946 chrll:68622722-68623252 chr2:25499763-25500429 chr5:172661486-172662228 chrl7:46691520-46692097 chrl2:75602991-75603344 chr2:80531367-80531719 chr5:158478378-158478630 chr2:177017266-177017489 chr2:63282514-63283122 chr7:155595692-155599414 chr5:172665306-172666072 chrl2:114843022-114843610 chrl3:112758598-112760491 chr4:4858389-4858893 chrl6:55365814-55366022 chr9:96108466-96108992 chrl2:3475010-3475654 chr9:86152353-86153777 chr6:10384965-10385492 chr22:31500396-31501239 chr5:179228283-179229003 chr6:137816474-137817223 chr2:106681982-106682403 chrl4:95239375-95239679 chr7:154001964-154002281 chrl:1476093-1476669 chrl5:89904822-89906050 chrll:89224416-89224718 chr9:100615234-100617510 chr3:172165372-172166738 chrl:202678881-202679769 chrl4:37053134-37053690 chr4:41875445-41875794 chr2:162273294-162273725 chrl:181287300-181287873 chrl3:79181327-79181614 chr8:145103285-145108027 chr22:42305617-42307254 chr8:102505512-102506430 chrl7:74533281-74534566 chrl:214156000-214156851 chr20:2780978-2781497 chr4:4861227-4862241 chrl9:13215244-13215543 chr7:121943867-121944538 chrl7:71948478-71949255 chr2:127413696-127414171 chrl:113286332-113287172 chrl:47009575-47010132 chrl6:62069121-62070634 chrl6:3013651-3015131 chrl8:76732970-76734765 chr4:155664819-155665833 chr6:72298274-72298528 chrl5:89147660-89149198 chrl7:33775294-33775794 chrl8:44337510-44338100 chrl0:8076002-8077261 chrl3:112717125-112717421 chrl5:89914363-89915061 chrl:228785986-228786204 chrl:156358050-156358252 chr7:751712-752150 chr3:137489051-137489409 chrl7:7905927-7907445 chrl8:35144907-35147628 chr3:9177691-9178189 chr6:10390888-10391098 chrl4:37052537-37052838 chrl:47909712-47911020 chrl3:93879245-93880877 chrl:50893468-50893745 chr7:27282086-27283136 chr4:147558231-147558583 chrl9:13124569-13124788 chrl7:46619087-46619314 chr3:44596535-44597018 chrl4:24803678-24804353 chr2:3286324-3286530 chrl2:14134626-14135242 chrl2:114881649-114881937 chr20:22548967-22549720 chr8:37822486-37824008 chrl3:100641334-100642188 chr4:206377-206892 chr3:11034446-11035384 chr7:152622343-152623305 chrl0:22629360-22630328 chr4:140201064-140201449 chrl9:46318490-46319266 chr3:121902742-121903645 chr9:77112712-77113583 chr2:114256775-114258043 chrl0:15761423-15762101 chrl:115880167-115881332 chr6:50791110-50791573 chr6:55039170-55039392 chr2:176980765-176981423 chr8:86350765-86351196 chr8:24812946-24814299 chr7:19184818-19185033 chr5:76936126-76936984 chr5:87980878-87981272 chr9:77111778-77112042 chrll:20622720-20623399 chrl:50882433-50882660 chrl7:35291899-35300875 chrl7:46675044-46675589 chr20:5296266-5297798 chr7:156871054-156871297 chr4:681313-681514 chr2:177039551-177039951 chrl7:46695325-46695553 chrl:41283840-41284591 chr9:16726859-16727273 chrl:65991001-65991811 chrl:181452706-181453073 chr8:120428398-120429178 chr3:32863174-32863415 chr4:134069162-134070442 chrl2:123754049-123754373 chr5:63256548-63257886 chr5:1879689-1879928 chrl0:118899247-118900329 chr20:2731063-2731395 chr5:134385967-134386370 chr2:177014948-177015214 chrl:67218079-67218293 chrll:65408344-65408631 chr7:156801418-156801632 chrl8:54788959-54789194 chr2:220173870-220174283 chr2:220173021-220173271 chrl2:113908887-113910681 chr6:100897080-100897621 chrl:155290606-155291001 chr2:130763483-130763764 chrl2:129337870-129338653 chr21:34395128-34400245 chrl2:52115410-52115679 chr3:126113547-126113967 chrl6:3220438-3221356 chrl:119543056-119543454 chrl4:62279476-62280019 chrll:636906-640628 chrl0:102893660-102895059 chr3:3840513-3842772 chrl:119529819-119530712 chr9:32782936-32783625 chrl9:1064897-1065191 chr5:54527319-54527760 chr7:156795355-156799394 chrl:155147185-155147444 chr9:37002489-37002957 chrll:69831571-69832484 chr2:128421719-128422182 chr22:38476836-38478839 chrl9:54412710-54413087 chr9:123656750-123656972 chr7:129422997-129423355 chrl9:36336275-36337138 chr2:50574045-50574817 chrl0:102975969-102978096 chr6:5996185-5996486 chr3:26664104-26664796 chr7:155170623-155170939 chr8:65286067-65286659 chrl4:37125219-37125661 chrll:65816404-65816665 chr6:41908745-41909711 chrl7:46620367-46621373 chr2:142887724-142888553 chrl:221050448-221050864 chrl2:106974412-106974951 chrl4:57278068-57278287 chrl:67773329-67773767 chrl7:40936445-40936668 chr20:2729997-2730797 chrl2:113013099-113013529 chr7:155244046-155244357 chrl:214153214-214153668 chrl:156863415-156863711 chrl:114695136-114696672 chrl4:85996494-85996958 chr7:100823307-100823701 chr20:52789252-52790986 chr5:178421225-178422337 chrll:36397926-36399398 chrl3:36052553-36053119 chrl4:57283967-57284558 chr4:25090106-25090510 chr2:5831187-5831413 chr6:117869097-117869530 chrl9:58094739-58095764 chr4:85422929-85423190 chrl3:100547172-100547431 chr8:68864584-68864946 chrl6:49311413-49312308 chr7:19184221-19184686 chr2:19562749-19562965 chrl9:54481412-54481955 chrl0:124901907-124902617 chr3:62357639-62359774 chrll:31827696-31827921 chrl7:43037166-43037740 chr7:37955622-37956555 chr6:106429111-106429772 chr6:50682334-50683214 chr5:76923887-76924502 chr6:168841818-168843100 chr7:19145872-19146256 chr20:32856659-32857248 chrl7:79859808-79860963 chr7:95225503-95226194 chrl4:105167663-105168129 chrl7:14248391-14248721 chrl6:84002269-84002860 chr9:104499849-104501076 chrl7:46604362-46604881 chr2:87015974-87018182 chrl4:36990873-36991209 chr5:52777788-52777996 chrl9:35633847-35634629 chrl:221055492-221055800 chrl:146551476-146551764 chrl3:100642774-100643094 chrl4:85999532-86000478 chrl3:36049570-36050159 chr2:119606038-119606313 chrll:123065426-123066184 chr3:172167526-172167866 chr4:41882450-41882964 chr8:142528185-142529029 chr9:79637814-79638169 chr3:19189688-19190100 chr4:122301567-122302290 chrl0:130339526-130339777 chr9:35846310-35846638 chrl5:53097561-53098476 chr2:157184389-157184632 chr5:145718289-145720095 chrll:105481126-105481422 chr5:170741603-170742751 chr3:62355315-62355534 chrl:38219702-38220012 chr4:41881177-41881418 chrl3:112715359-112716234 chrl7:1880789-1881116 chrl8:56887091-56887665 chr6:10390038-10390565 chrll:69516931-69517218 chrl9:39737689-39739288 chr3:157812053-157812764 chrl4:37049333-37051726 chr7:156409023-156409294 chrll:46366876-46367101 chr5:50685453-50686148 chr4:41883492-41884570 chrl3:112709884-112712665 chr22:44287497-44288061 chr22:46440393-46441019 chr8:23562475-23565175 chr2:207506774-207507422 chr4:169799086-169799625 chr3:133393118-133393657 chr8:41424341-41425300 chr4:100870377-100871994 chr4:107956555-107957453 chrl7:79314962-79320653 chr2:30453566-30455655 chrl:18956895-18959829 chrl2:41086522-41087102 chr22:42685894-42686095 chr6:100914946-100915245 chrl:46951168-46951792 chr4:41749184-41749811 chrll:128419198-128419513 chr2:171671598-171671804 chrl:170630456-170630851 chr20:44657463-44659243 chr9:139096665-139096993 chr7:155174128-155175248 chrl4:36993488-36994488 chr3:138654837-138655363 chr4:5709985-5710495 chrl5:23157794-23158624 chr20:9496471-9496893 chr4:174437914-174438346 ch r5:140305712-140307193 chrl5:79576059-79576270 chrl4:38678245-38680937 chrl0:102473206-102474026 chrl7:59486727-59487132 chr3:64253533-64253819 chrl0:102484200-102484476 chr7:27198182-27198514 chr2:97192977-97193383 chr9:77113709-77113927 chr6:154360586-154361008 chrll:44324875-44325087 chr2:182521221-182521927 chr7:124404700-124406189 chr2:132182327-132183101 chr7:101005899-101007443 ch r7:149744402-149746469 chr8:50822270-50822860 chr7:27227520-27229043 chr6:134212690-134213098 chrl3:36044844-36045481 chrll:132934059-132934291 chrl6:51189800-51190260 chrl:155145342-155145938 chr4:682724-683079 chr5:92939795-92940216 chrl0:134597357-134602649 chrl:200009807-200010036 chrl9:12666243-12666682 chr9:97401286-97402067 chr2:107103833-107104053 chrl5:89910521-89912177 chr5:140789094-140789762 chr2:114033359-114033617 chrl7:12568667-12569335 chrll:68622108-68622339 chrl: 160340604-160340843 chr7:103085710-103086132 chrl5:76628998-76629207 chr20:10198135-10198984 chr20:44660342-44660948 chrl7:35290403-35290663 chrl7:933026-933236 chr4:128544031-128544903 chrl:50881884-50882103 chrl0:125425495-125426642 chrl7:46801784-46802071 chrl:25255527-25259005 chr3:32861141-32861429 chrl7:70116274-70119998 chrl0:75407413-75407706 chr2:467849-468659 chrll:132952538-132953307 chr3:6904133-6904641 chrl0:120353692-120355821 chr7:20830567-20830817 chrll:71950815-71951408 chrl4:95240083-95240341 chrl9:5829048-5829474 chr20:9495253-9495597 chr9:112083333-112083549 chrl5:96873408-96877721 chrl6:67208067-67208678 chrl:175568376-175568808 chr6:5999149-5999787 chr3:129693127-129694841 chr6:10383525-10384114 chrll:636435-636668 chrl:181451311-181452049 chr9:135464586-135466240 chrl5:60289325-60289533 chrl6:49309123-49309353 chrl:243646394-243646888 chrl2:54071053-54071265 chrl:91176404-91176701 chr5:140864527-140864748 chr4:47034427-47034940 chrl0:102489343-102491011 chrl0:102419147-102419668 chrl2:81471569-81472119 chr6:50813314-50813699 chr5:158526133-158526431 chrl:119543821-119544339 chr5:77140542-77140914 chr8:23567180-23567678 chrl:41831976-41832542 chr2:139537692-139538650 chr7:100075303-100075551 chr2:176969217-176969895 chr7:27284639-27286237 chr5:31193952-31194419 chr6:37616393-37616621 chrl9:1748167-1750243 chrl0:101281181-101282116 chr21:31311386-31312106 chr2:176973427-176973718 chrl5:96900142-96900644 chr7:158936507-158938492 chr3:63263989-63264205 chrl6:71459781-71460338 chr7:155601175-155603235 chrl2:54447744-54448091 chrl2:53491572-53491955 chrl0:16561604-16563822 chrll:133994709-133995090 chr2:137522460-137523696 chrl7:12877270-12877773 chr8:98289604-98290404 chr4:185937242-185937750 chr3:185911344-185912228 chrl2:54378696-54380102 chrl:221060850-221061071 chrl2:63543636-63544967 chr6:6006689-6007043 chrl9:51169659-51172023 chrl:1474962-1475220 chrl4:54418677-54418881 chr6:108497595-108497996 chrl7:37764092-37764304 chr4:109092578-109092839 chrl:91182097-91182364 chrl3:112760865-112761113 chrl2:122018170-122018457 chr7:142494563-142495248 chrl3:58203586-58204322 chrl:92945907-92952609 chrl2:106977388-106977713 chr5:76925445-76926875 chrl6:3190765-3191389 chrl:12123488-12124148 chrl7:48545570-48546900 chrl2:113916433-113916717 chr4:41747508-41747944 chrl9:46916587-46916862 chrl5:49254984-49255564 chrl9:8674332-8674764 chr2:223167205-223167560 chrl7:1173535-1174733 chr3:75955759-75956308 chr5:115697134-115697589 chr8:21644908-21647845 chr5:59189046-59189894 chrl2:54338761-54339168 chrl6:31053479-31053800 chrl:50892437-50893243 chrl7:40935964-40936180 chrl9:44203558-44203987 chr4:81109887-81110460 chrl:2979275-2980758 chrl6:49872449-49872926 chrl:200008392-200009047 chrl6:49316997-49317263 ch r2:114034594-114036041 chr2:105480197-105480760 chrl8:44777632-44778084 chrl9:13213450-13213821 chrl7:6616422-6617471 chrl4:36977518-36977996 chrl:214160798-214161034 chrl:91182509-91182857 chrl0:130508443-130508658 chr2:154728944-154729328 chrl5:89952271-89953061 chrl8:55102427-55102708 chr22:31198491-31199033 chrl0:50821487-50821688 chr7:100076454-100076785 chrl8:13641584-13642415 chrl8:13868532-13869026 chr6:168841438-168841699 chrl:61515875-61516831 chr7:32110063-32110910 chr7:56355508-56355798 chrl9:12767749-12767980 chrl9:19371675-19372393 chrl4:69256676-69257036 chrl7:75447477-75447821 chrl4:24801680-24802153 chr5:148033472-148034080 chrl0:125650820-125651373 chrll:43568921-43569854 chr22:37212769-37213467 chr2:162283581-162284677 chr8:130995921-130996149 chrll:70508328-70508617 chrl6:88943427-88943669 chrl9:42891311-42891646 chrl5:53079220-53079579 chrl7:46690390-46691055 chr4:41880224-41880500 chrl:156105707-156106171 chr6:5997027-5997414 chrl:18964180-18964401 chrl4:36983440-36983738 chrl2:54445876-54446113 chr5:87968635-87968907 chrl:29587087-29587412 chrll:60718428-60718888 chr2:66672431-66673636 chr4:81119095-81119391 chrl0:76573195-76573507 chr22:42322043-42322909 chrl9:45898879-45900315 chrl4:95826675-95826941 chrl7:48194634-48195085 chrl9:49669275-49669552 chrl5:96897596-96898046 chrl9:40314926-40315144 chr9:120507227-120507642 chr5:145722467-145722925 chr3:19188246-19188772 chr5:140787447-140788044 chrl9:50881418-50881664 chrl0:102896342-102896665 chr7:53286851-53287192 chrl5:89903446-89903720 chrl0:23461300-23461610 chr2:127783081-127783311 chrll:72532612-72533774 chr2:119605200-119605620 chrl8:12254147-12255089 chr7:100817759-100817975 chrl4:77736733-77737772 chrl2:127212279-127212529 chr2:119606569-119606826 chrl:155264318-155265536 chrl2:131199824-131200157 chrl:91300979-91301891 chr6:100909210-100909444 chr6:4079052-4079443 chr2:233251361-233253414 chr4:960505-960836 chrl9:21769189-21769786 chrl0:102279162-102279730 chrl2:127210778-127211651 chrl2:54069625-54070177 chrl5:53087211-53087488 chrl3:28365545-28365785 chrl2:113913615-113914322 chrl4:51338712-51339146 chr7:155604725-155605095 chr3:62364017-62364316 chr6:6008857-6009299 chr3:46618307-46618669 chrl7:33776553-33776888 chrl2:58158855-58160000 chr2:219857682-219858917 chrl9:44278273-44278777 chrl0:101282725-101282934 chr20:2539133-2539877 chrl2:58003880-58004249 chrl6:51147490-51147944 chrl:179544720-179545307 chr2:71787430-71787897 chrl0:129534410-129537366 chr6:42145847-42146053 chrl4:24802927-24803159 chr22:29707479-29707797 chr9:132459587-132460017 chrl7:40937258-40937480 chr4:151504011-151505085 chrl:18967251-18968119 chrl9:56598038-56600296 chrl9:35633409-35633697 chr2:171678546-171680358 chr6:134638797-134639021 chrl:36549554-36549965 chrl9:12833104-12833574 chr3:137487429-137488021 chr9:139715663-139716441 chr6:37617863-37618147 chrl7:32484007-32484280 chr7:156409577-156409865 chr5:11384681-11385521 chr8:102504478-102504841 chr20:33296514-33298242 chr20:57415135-57417153 chrl0:71331449-71331691 chr3:75667777-75669067 chrl6:67571252-67572728 chrl9:36500169-36500530 chr2:154729613-154729918 chrl2:48399168-48399372 chr4:41867385-41867586 chrl7:46800533-46800746 chr20:44685771-44687610 chrl9:10406934-10407342 chr6:108496715-108497320 chr5:158523906-158524598 chr9:124413512-124414193 chr20:57427691-57427995 chrl6:10912159-10912719 chr7:149389654-149389976 chrl:173638662-173639045 chrl9:55597977-55598887 chrl4:62279037-62279339 chr3:13114627-13115245 chr2:3750828-3751927 chr4:85402764-85403175 chrl7:74017769-74018658 chr5:54523676-54523901 chr7:89747892-89749036 chrl8:72916107-72917233 chr9:136294738-136295236 chrl:201252452-201253648 chr5:146888750-146889840 chrl4:52734207-52735486 chrl3:20875518-20876214 chrl8:77560088-77560292 chr2:102803672-102804556 chr2:176982107-176982402 chrl7:6679205-6679710 chrl9:10463626-10464378 chr5:140810494-140812617 chrll:46299544-46300216 chrll:64136814-64138187 chr6:6007387-6007797 chrl7:37321482-37322099 chrl0:94455524-94455896 chrl3:51417371-51418149 chr8:11565217-11567212 chrl:226127112-226127695 chr2:3287874-3288228 chr6:10882926-10883149 chr22:19746155-19746369 chr3:12838471-12838782 chr9:36739534-36739782 chr9:134429866-134430491 chrll:70672834-70673055 chrl4:24641053-24642220 chr7:27283408-27283614 chrl2:49182421-49182658 chrl:44031286-44031853 chrl:114696886-114697185 chrl5:89901914-89902785 chrll:65352231-65353134 chr7:72838383-72838815 chr22:38379093-38379964 chr4:155663809-155664315 chr9:100619984-100620192 chr7:143582125-143582610 chr7:23287221-23287508 chrll:64815040-64815722 chr2:87088816-87089037 chr20:57426729-57427047 chrl0:43428167-43429460 chrl0:121577529-121578385 chr4:190939801-190940591 chr6:100037323-100037544 chrl9:12880574-12880888 chr2:171670110-171670549 chr7:124404174-124404432 chr7:97840559-97840845 chrl9:50879606-50880094 chrl:113265573-113265787 chrl9:2424005-2427983 chr3:127633993-127634588 chrl0:50817095-50817309 chr2:171676552-171676980 chrl:86621278-86622871 chrl:164545540-164545917 chr22:19967279-19967808 chrll:67350928-67351953 chr20:36226617-36226841 chrl9:14089570-14089796 chrl9:38700333-38700577 chrl:18435566-18435904 chr8:21905461-21905757 chr2:176950595-176950846 chrl7:75251958-75252180 chrl5:37390175-37390380 chr9:98113447-98113662 chrl:40235767-40237190 chr8:144811237-144811446 chr8:99984584-99985072 chr7:152621916-152622149 chrl:40769186-40769871 chrl9:2428349-2428731 chrl7:15820620-15821325 chr22:25081850-25082112 chrl:19203874-19204234 chr20:61703526-61704022 chr2:237080188-237080432 chrl:156338758-156339251 chrl:149332993-149333389 chr22:50496441-50497393 chr7:27146069-27146600 chrl3:100547633-100548911 chr4:190939007-190939274 chr7:73894815-73895110 chrl9:35632356-35632572 chrl6:67918679-67918909 chr2:108602824-108603467 chr2:238864315-238865170 chr8:144808221-144810978 chr8:145101631-145101834 chrl2:132905449-132906206 chr6:99275763-99276038 chr5:140800760-140801072 chrl7:75242871-75243613 chrl7:41278134-41278460 chrl2:122016170-122017693 chrl0:131264948-131265710 chrl7:46631800-46632212 chrl4:105167277-105167501 chrl0:23982382-23982589 chrl9:50931270-50931638 chr3:27771638-27771942 chrl8:74799144-74800038 chrl:21616380-21617101 chrl:147782066-147782473 chr7:6590563-6590957 chr7:97839862-97840222 chrl2:113914440-113914657 chrl9:7933263-7934898 chr20:22559553-22560001 chrl5:53086629-53086858 chrl0:94180315-94180754 chr5:140052059-140053381 chrl0:101287162-101287920 chrl4:38677154-38677787 chr22:39262338-39263211 chrl8:74153239-74155073 chrl5:59157045-59157594 chr4:963804-964115 chrll:624780-625053 chr7:1362811-1363643 chrl9:36246328-36247982 chr5:54528095-54528404 chrl2:54359658-54359906 chr2:127782613-127782829 chrl9:406131-406611 chrl7:46697413-46697701 chrl8:43608140-43608510 chrl6:23724270-23724775 chrl8:55922987-55924068 chrl5:60291879-60292167 chrl4:92788913-92789204 chrl9:1108394-1109610 chrll:124628367-124629590 chrl:32052471-32052771 chrl9:11594372-11594987 chrl9:870774-871318 chr2:54086776-54087266 chr2:241459632-241460047 chr7:127990926-127992616 chrl:208132327-208133117 chr7:90893567-90896683 chrl:41284847-41285149 chrll:32452144-32452708 chr5:77146998-77147785 chrl9:45901452-45901688 chr7:6661875-6662695 chr6:161188084-161188639 chrl7:934417-935088 chrll:65409636-65410127 chrl7:19883325-19883610 chrl8:77549524-77550299 chrl:38461584-38461988 chrl9:10464666-10464927 chrl7:70120139-70120442 chr7:27147589-27148389 chr2:31806545-31806782 chrll:119292689-119292891 chrl9:18979351-18981200 chr6:42879279-42879623 chrl2:130908777-130909191 chrl7:46629553-46629816 chrl:202162958-202163390 chrl7:21367114-21367592 chrl6:84001805-84002011 chrl:221057463-221057757 chrl7:27899511-27900067 chrl5:40268581-40269061 chr22:37465056-37465331 chrl7:77805866-77809046 chrl9:13198699-13198999 chr3:184056419-184056671 chr22:37911979-37912258 chrl9:19368708-19369681 chrll:64135815-64136381 chrl8:77552401-77552603 chrl9:58554354-58554587 chr20:57414595-57414896 chr4:190938106-190938848 chr5:172110282-172111166 chrl6:68480864-68482822 chr9:139395020-139395287 chrl2:113515164-113515970 chrl:221054554-221054888 chr8:144990270-145002135 chr9:131154346-131155923 chr6:150335525-150336278 chr9:115824684-115825033 chrl2:54519768-54520457 chr6:35479872-35480154 chrl9:3870788-3871043 chrl9:48965002-48965792 chr6:35479388-35479678 chrl2:52408381-52408675 chrl:221068782-221069159 chr6:46655262-46656738 chr3:55508336-55508708 chrl:39980365-39981768 chrl6:3067521-3068358 chrl:1473107-1473342 chrl0:105362549-105362827 chrl7:46698880-46699083 chr2:198029068-198029438 chr20:17209418-17209622 chrl2:49183049-49183282 chrl6:58030214-58031633 chrl0:94820026-94823252 chrll:725596-726870 chr6:170732119-170732442 chrl2:120835586-120835927 chr20:36012595-36013439 ch r8:143545445-143546178 chr6:27228100-27228364 chr21:32624144-32624382 chr9:95477296-95477708 chrl0:105420685-105421076 chrl:1470604-1471450 chrl:146552328-146552577 chrl9:33625467-33625805 chrll:64478843-64479598 chr20:57428308-57428516 chr7:27182613-27185562 chrl9:51815157-51815458 chrl7:46607804-46608390 chrl2:52408860-52409121 chrl9:10405924-10406398 chrll:14993452-14993661 chrl9:13135317-13136169 chr7:750788-751237 chrl:53742297-53742845 chrl:200010625-200010832 chr5:139138875-139139242 chrl7:45949676-45949885 chr3:128722283-128723036 chrl5:89312719-89313183 chr9:135039673-135039978 chrl9:12831793-12832225 chr20:51589707-51590020 chr20:3145121-3145746 chr8:65710990-65711722 chrll:128694084-128694688 chr2:20870006-20871280 chrl9:18977466-18977833 chr3:49947621-49948430 chr6:30139718-30140263 chrl2:104697348-104697984 chrl0:105361784-105362188 chr6:29894140-29895117 chr4:187219320-187219745 chrl5:67073306-67073943 chr2:220412341-220412678 chr6:170730395-170730887 chr9:115822071-115823416 chrl:10764449-10764925 chrl7:46627787-46628444 chrl9:51601822-51602260 chrl9:55814067-55814278 chr6:138745348-138745593 chr9:124987743-124991086 chr22:46318693-46319087 chrl6:3013016-3013228 chr4:114900355-114900810 chrl9:1063544-1064265 chrl9:1110399-1110701 chr7:97841636-97842005 chr8:57359899-57360114 chrl7:72915568-72916510 chrl:16860873-16862296 chrl7:75398284-75398527 chr9:139397412-139397710 chr6:33393592-33393908 chr6:29595298-29595795 chrl2:6438272-6438931 chr3:113160299-113160641 chrl:55505060-55506015 chrll:132951692-132952260 chr4:81118137-81118603 chrl9:38876070-38876332 chrl9:58549305-58549712 chrl7:43472527-43474343 chr9:139396205-139397040 chrl6:3192181-3192669 chr6:33048416-33048814 chr7:128555329-128556650 chrl9:46915311-46915802 chr6:30095173-30095610 Table 2: Example CGIs Human CGI (hgl9) chrl 1181756-1182470 chrl 2 103696090-103696418 chrl 1470604-1471450 chrl 2 104697348-104697984 chrl 2772126-2772665 chrl 2 106974412-106974951 chrl 4713989-4716555 chrl 2 113013099-113013529 chrl 18436551-18437673 chrl 2 113515164-113515970 chrl 18956895-18959829 chrl 2 113916433-113916717 chrl 18962842-18963481 chrl 2 114833911-114834210 chrl 18967251-18968119 chrl 2 114838312-114838889 chrl 19203874-19204234 chrl 2 114843022-114843610 chrl 21616380-21617101 chrl 2 114845861-114847650 chrl 25255527-25259005 chrl 2 114851957-114852360 chrl 29585897-29586598 chrl 2 114881649-114881937 chrl 34628783-34630976 chrl 2 114885105-114885418 chrl 39980365-39981768 chrl 2 119212110-119212393 chrl 40235767-40237190 chrl 2 123754049-123754373 chrl 41831976-41832542 chrl 2 127210778-127211651 chrl 46951168-46951792 chrl 2 127940451-127940907 chrl 47909712-47911020 chrl 2 129337870-129338653 chrl 53742297-53742845 chrl 2 131199824-131200157 chrl 55505060-55506015 chrl 2 132905449-132906206 chrl 61515875-61516831 chrl 3 20875518-20876214 chrl 63782394-63790471 chrl 3 28366549-28368505 chrl 65731411-65731849 chrl 3 28549839-28550246 chrl 66258440-66258918 chrl 3 36044844-36045481 chrl 77747314-77748224 chrl 3 51417371-51418149 chrl 91172102-91172771 chrl 3 53419897-53422872 chrl 91176404-91176701 chrl 3 58203586-58204322 chrl 92945907-92952609 chrl 3 58206526-58208930 chrl 115880167-115881332 chrl 3 79181944-79182222 chrl 116380359-116382364 chrl 3 93879245-93880877 chrl 156105707-156106171 chrl 3 100547633-100548911 chrl 156338758-156339251 chrl 3 100641334-100642188 chrl 156358050-156358252 chrl 3 102568425-102569495 chrl 156390403-156391581 chrl 3 112707804-112708696 chrl 160340604-160340843 chrl 3 112709884-112712665 chrl 161695637-161697298 chrl 3 112715359-112716234 chrl 177133392-177133846 chrl 3 112717125-112717421 chrl 180198119-180204975 chrl 3 112720564-112723582 chrl 197887088-197887791 chrl 3 112726281-112728419 chrl 201252452-201253648 chrl 3 112758598-112760491 chrl 202678881-202679769 chrl 3 112760865-112761113 chrl 214156000-214156851 chrl 4 24044886-24046760 chrl: 214158726-214159080 chrl 4 24641053-24642220 chrl: 221057463-221057757 chrl 4 24803678-24804353 chrl: 221067447-221068185 chrl 4 29236835-29237832 chrl: 226075150-226075680 chrl 4 29254365-29255069 chrl: 248020330-248021252 chrl 4 33402094-33404079 chrlO 50602989-50606783 chrl 4 36973169-36973740 chrlO 50817601-50820356 chrl 4 36983440-36983738 chrlO 71331926-71333392 chrl 4 36990873-36991209 chrlO 88122924-88127364 chrl 4 36993488-36994488 chrlO 94820026-94823252 chrl 4 37053134-37053690 chrlO 101279941-101280382 chrl 4 37126786-37128274 chrlO 101281181-101282116 chrl 4 37135513-37136348 chrlO 102419147-102419668 chrl 4 38724254-38725537 chrlO 102473206-102474026 chrl 4 48143433-48145589 chrlO 102484200-102484476 chrl 4 51338712-51339146 chrlO 102489343-102491011 chrl 4 52734207-52735486 chrlO 102507482-102509646 chrl 4 57260878-57262123 chrlO 102893660-102895059 chrl 4 57264638-57265561 chrlO 102896342-102896665 chrl 4 57278709-57279116 chrlO 102899822-102900263 chrl 4 58331676-58333121 chrlO 102975969-102978096 chrl 4 60973772-60974123 chrlO 105361784-105362188 chrl4 60975732-60978180 chrlO 105420685-105421076 chrl 4 61103978-61104663 chrlO 106399567-106402812 chrl 4 62279476-62280019 chrlO 118899247-118900329 chrl 4 77736733-77737772 chrlO 119000435-119001530 chrl 4 85997468-85998637 chrlO 119311204-119312104 chrl 4 85999532-86000478 chrlO 119312766-119313563 chrl 4 92789494-92790712 chrlO 124905634-124906161 chrl 4 95239375-95239679 chrlO 124907283-124911035 chrl 4 95826675-95826941 chrlO 129534410-129537366 chrl 4 101192851-101193499 chrl 1 725596-726870 chrl 4 101923575-101925995 chrll 8190226-8190671 chrl 4 103655241-103655928 chrll 17740789-17743779 chrl 5 23157794-23158624 chrl 1 20181200-20182325 chrl 5 27112030-27113479 chrll 20622720-20623399 chrl 5 27215951-27216856 chrl 1 31825743-31826967 chrl 5 33602816-33604003 chrl 1 31839363-31839813 chrl 5 35046443-35047480 chrll 31848487-31848776 chrl 5 37390175-37390380 chrl 1 32452144-32452708 chrl 5 53076187-53077926 chrll 32454874-32457311 chrl 5 53079220-53079579 chrll 36397926-36399398 chrl 5 53080458-53083699 chrl 1 44327240-44327932 chrl 5 53087211-53087488 chrll 46299544-46300216 chrl 5 53097561-53098476 chrl 1 46366876-46367101 chrl 5 59157045-59157594 chrll 64136814-64138187 chrl 5 76630029-76630970 chrll 65352231-65353134 chrl 5 79574830-79575211 chrll 69517840-69519929 chrl 5 89147660-89149198 chrl 1 69831571-69832484 chrl 5 89312719-89313183 chrl 1 70672834-70673055 chrl 5 89903446-89903720 chrll 72532612-72533774 chrl 5 89910521-89912177 chrl 1 79148358-79152200 chrl 5 89952271-89953061 chrll 124629723-124629926 chrl 5 96895306-96895729 chrl 2 3475010-3475654 chrl 5 96903311-96903711 chrl 2 5018585-5021171 chrl 5 96904722-96905050 chrl 2 6438272-6438931 chrl 5 96909815-96910030 chrl 2 15475318-15475901 chrl 5 96959341-96960531 chrl 2 29302034-29302954 chrl 5 100913438-100914022 chrl 2 45444202-45445386 chrl 6 3067521-3068358 chrl 2 49183049-49183282 chrl 6 3220438-3221356 chrl 2 49371690-49375550 chrl 6 6068914-6070401 chrl 2 49484920-49485178 chrl 6 10912159-10912719 chrl 2 53491572-53491955 chrl 6 20084707-20085305 chrl 2 54338761-54339168 chrl 6 23724270-23724775 chrl 2 54366815-54369103 chrl 6 24267040-24267527 chrl 2 54378696-54380102 chrl 6 31053479-31053800 chrl 2 54423427-54423712 chrl 6 49309123-49309353 chrl 2 54440642-54441543 chrl 6 49316997-49317263 chrl 2 54447744-54448091 chrl 6 51183699-51188763 chrl 2 54519768-54520457 chrl 6 54325040-54325703 chrl 2 57618769-57619402 chrl 6 55364823-55365483 chrl 2 58003880-58004249 chrl 6 66612749-66613412 chrl 2 58158855-58160000 chrl 6 67918679-67918909 chrl 2 63543636-63544967 chrl 6 71459781-71460338 chrl 2 75602991-75603344 chrl 6 82660651-82661813 chrl 2 99139386-99139769 chrl 6 84002269-84002860 chrl 2 101109863-101111622 chrl 7 934417-935088 chrl 2 106979429-106981086 chrl 7 1173535-1174733 chrl 2 113590806-113591304 chrl 7 1880789-1881116 chrl 2 113900750-113906442 chrl 7 5000369-5001205 chrl 2 113908887-113910681 chrl 7 6616422-6617471 chrl 2 113913615-113914322 chrl 7 6679205-6679710 chrl 2 114878143-114879155 chrl 7 7832532-7833164 chrl 2 114886354-114886579 chrl 7 7905927-7907445 chrl 2 115109503-115110061 chrl 7 12877270-12877773 chrl 2 117798076-117799448 chrl 7 14201726-14202052 chrl 2 120835586-120835927 chrl 7 15820620-15821325 chrl 2 122016170-122017693 chrl 7 19883325-19883610 chrl 2 130387609-130389139 chrl 7 21367114-21367592 chrl 2 130908777-130909191 chrl 7 27899511-27900067 chrl3 27334226-27335205 chrl7 33776553-33776888 chrl3 28498226-28499046 chrl7 36717727-36718593 chrl3 36049570-36050159 chrl7 37321482-37322099 chrl3 36052553-36053119 chrl7 43037166-43037740 chrl3 79182859-79183880 chrl7 46604362-46604881 chrl3 84453664-84453897 chrl7 46627787-46628444 chrl3 108518334-108518633 chrl7 46673532-46674181 chrl3 109147798-109149019 chrl7 46697413-46697701 chrl4 36974548-36975425 chrl7 46796234-46797292 chrl4 36986362-36990576 chrl7 46800533-46800746 chrl4 37049333-37051726 chrl7 46824785-46825372 chrl4 37116188-37117628 chrl7 48041282-48043064 chrl4 38678245-38680937 chrl7 48545570-48546900 chrl4 54418677-54418881 chrl7 59531723-59535254 chrl4 57274607-57276840 chrl7 70111979-70112308 chrl4 57283967-57284558 chrl7 70112824-70114271 chrl4 69256676-69257036 chrl7 71948478-71949255 chrl4 74706188-74708192 chrl7 73749618-73750178 chrl4 95237622-95238211 chrl7 74533281-74534566 chrl4 105167663-105168129 chrl8 904578-909574 chrl5 33009530-33011696 chrl8 11148307-11149936 chrl5 40268581-40269061 chrl8 11750953-11752756 chrl5 45408573-45409528 chrl8 12254147-12255089 chrl5 47476369-47477499 chrl8 13641584-13642415 chrl5 49254984-49255564 chrl8 13868532-13869026 chrl5 60287107-60287663 chrl8 43608140-43608510 chrl5 60296135-60298520 chrl8 44336183-44337110 chrl5 67073306-67073943 chrl8 44337510-44338100 chrl5 74419870-74423044 chrl8 44772992-44775577 chrl5 79724099-79725643 chrl8 44777632-44778084 chrl5 89914363-89915061 chrl8 44789742-44790678 chrl5 89920793-89922768 chrl8 54788959-54789194 chrl5 89949373-89951130 chrl8 55019707-55021605 chrl5 91642908-91643702 chrl8 55094825-55096310 chrl5 96873408-96877721 chrl8 56887091-56887665 chrl6 2228190-2230946 chrl8 56939624-56941540 chrl6 3013016-3013228 chrl8 70533965-70536871 chrl6 3190765-3191389 chrl8 72916107-72917233 chrl6 22824616-22826459 chrl8 73167402-73167920 chrl6 48844551-48845264 chrl8 74799144-74800038 chrl6 49311413-49312308 chrl8 76732970-76734765 chrl6 49314037-49316543 chrl8 76737005-76741244 chrl6 49872449-49872926 chrl8 77547965-77549038 chrl6 51147490-51147944 chrl8 77557780-77558948 chrl6 51168266-51169110 chrl9 870774-871318 chrl6 54970301-54972846 chrl9 3868586-3869217 chrl6 55513220-55513526 chrl9 5829048-5829474 chrl6 58030214-58031633 chrl9 8674332-8674764 chrl6 62069121-62070634 chrl9 10406934-10407342 chrl6 67208067-67208678 chrl9 10463626-10464378 chrl6 67571252-67572728 chrl9 12666243-12666682 chrl6 68480864-68482822 chrl9 12767749-12767980 chrl6 86530747-86532994 chrl9 12831793-12832225 chrl6 86549069-86550512 chrl9 12880574-12880888 chrl6 86612188-86613821 chrl9 13124959-13125259 chrl6 88943427-88943669 chrl9 13616752-13617267 chrl7 12568667-12569335 chrl9 14089570-14089796 chrl7 14248391-14248721 chrl9 19371675-19372393 chrl7 32484007-32484280 chrl9 21769189-21769786 chrl7 35291899-35300875 chrl9 33625467-33625805 chrl7 37764092-37764304 chrl9 36246328-36247982 chrl7 40937258-40937480 chrl9 36523391-36523887 chrl7 43472527-43474343 chrl9 38700333-38700577 chrl7 45949676-45949885 chrl9 39737689-39739288 chrl7 46607804-46608390 chrl9 39754973-39756540 chrl7 46620367-46621373 chrl9 40314926-40315144 chrl7 46631800-46632212 chrl9 44203558-44203987 chrl7 46669434-46669811 chrl9 44278273 -44278777 chrl7 46691520-46692097 chrl9 45260352-45261809 chrl7 48194634-48195085 chrl9 46001830-46002686 chrl7 50235175-50236466 chrl9 46318490-46319266 chrl7 59485573-59485780 chrl9 46915311-46915802 chrl7 59528979-59530266 chrl9 47151768-47153125 chrl7 70116274-70119998 chrl9 49669275-49669552 chrl7 70120139-70120442 chrl9 51601822-51602260 chrl7 72855621-72858012 chrl9 51815157-51815458 chrl7 72915568-72916510 chrl9 54412710-54413087 chrl7 74017769-74018658 chrl9 54481412-54481955 chrl7 77805866-77809046 chrl9 54483021-54483572 chrl7 79314962-79320653 chrl9 55597977-55598887 chrl7 79859808-79860963 chrl9 56988313-56989741 chrl8 19744936-19752363 chrl9 58094739-58095764 chrl8 30349690-30352302 chrl9 58545115-58545897 chrl8 35144907-35147628 chrl9 58554354-58554587 chrl8 55103154-55108853 chr2: 467849-468659 chrl8 55922987-55924068 chr2: 3286324-3286530 chrl8 59000683-59001692 chr2: 5831187-5831413 chrl8 74153239-74155073 chr2: 19560963-19561650 chrl8 74961556-74963822 chr2: 20870006-20871280 chrl9 407011-409511 chr2: 25499763-25500429 chrl9: 1063544-1064265 chr2 31805293-31806403 chrl9: 1 108394-1109610 chr2 45169505-45171884 chrl9: 1748167-1750243 chr2 45227644-45228783 chrl9: 2424005-2427983 chr2 45240372-45241579 chrl9: 7933263-7934898 chr2 54086776-54087266 chrl9: 1 1594372-11594987 chr2 63282514-63283122 chrl9: 13135317-13136169 chr2 63283936-63284147 chrl9: 13198699-13198999 chr2 63285949-63287097 chrl9: 13213450-13213821 chr2 66652691-66654218 chrl9: 18979351-18981200 chr2 66672431-66673636 chrl9: 19368708-19369681 chr2 80549578-80549798 chrl9: 30715549-30715753 chr2 87015974-87018182 chrl9: 35633409-35633697 chr2 87088816-87089037 chrl9: 36336275-36337138 chr2 97192977-97193383 chrl9: 36500169-36500530 chr2 105480197-105480760 chrl9: 38876070-38876332 chr2 106681982-106682403 chrl9: 42891311-42891646 chr2 107103833-107104053 chrl9: 45898879-45900315 chr2 114033359-114033617 chrl9: 48965002-48965792 chr2 114034594-114036041 chrl9: 50881418-50881664 chr2 114256775-114258043 chrl9: 50931270-50931638 chr2 118981769-118982466 chrl9: 51169659-51172023 chr2 119592602-119593845 chrl9: 55815940-55816277 chr2 119599059-119599299 chrl9: 56598038-56600296 chr2 119602616-119604486 chr2: 3750828-3751927 chr2 119606569-119606826 chr2: 30453566-30455655 chr2 119611296-119611881 chr2: 38301276-38304518 chr2 119616133-119616826 chr2: 45155195-45157049 chr2 119914126-119916663 chr2: 45395869-45398186 chr2 124782252-124783255 chr2: 50574045-50574817 chr2 127413696-127414171 chr2: 66808568-66809404 chr2 127782613-127782829 chr2: 71787430-71787897 chr2 128421719-128422182 chr2: 73143055-73148260 chr2 130763483-130763764 chr2: 80529677-80530846 chr2 132182327-132183101 chr2: 102803672-102804556 chr2 139537692-139538650 chr2: 105459127-105461770 chr2 154727906-154728271 chr2: 105468851-105473488 chr2 154728944-154729328 chr2: 108602824-108603467 chr2 162279835-162280709 chr2: 119599458-119600966 chr2 162283581-162284677 chr2: 137522460-137523696 chr2 171671598-171671804 chr2: 142887724-142888553 chr2 171678546-171680358 chr2: 144694666-144695180 chr2 176931575-176932663 chr2: 157185557-157186355 chr2 176936246-176936809 chr2: 162273294-162273725 chr2 176944087-176948446 chr2: 176949511-176949795 chr2 176949993-176950336 chr2: 176964062-176965509 chr2: 176956504-176956707 chr2: 176969217-176969895 chr2: 177012371-177012675 chr2: 176977284-176977540 chr2: 177016416-177016632 chr2: 176982107-176982402 chr2: 177024501-177025692 chr2: 177036254-177037213 chr2: 198029068-198029438 chr2: 177042751-177043444 chr2: 200333687-200334172 chr2: 182321761-182323029 chr2: 207506774-207507422 chr2: 182521221-182521927 chr2: 220173 870-220174283 chr2: 219736132-219736592 chr2: 223159725-223160487 chr2: 219848919-219850541 chr2: 223162946-223163912 chr2: 219857682-219858917 chr2: 223167205-223167560 chr2: 220299483-2203 00243 chr2: 223168653-223169008 chr2: 220412341-220412678 chr2: 223176493-223177515 chr2: 223183013-223185468 chr2: 233251361-233253414 chr2: 237071794-237078762 chr2: 237068071-237068834 chr2: 241758141-241760783 chr2: 238864315-238865170 chr20 3145121-3145746 chr2: 241459632-241460047 chr20 21485932-21496714 chr20 690575-691099 chr20 21686199-21687689 chr20 2539133-2539877 chr20 22557517-22559240 chr20 2729997-2730797 chr20 33296514-33298242 chr20 2780978-2781497 chr20 37352130-37357372 chr20 5296266-5297798 chr20 39994545-39995810 chr20 9496471-9496893 chr20 44657463-44659243 chr20 10198135-10198984 chr20 44685771-44687610 chr20 17206528-17206952 chr20 51589707-51590020 chr20 17208550-17208756 chr20 52789252-52790986 chr20 21376358-21378245 chr20 57415135-57417153 chr20 21694472-21695344 chr21 31311386-31312106 chr20 22548967-22549720 chr21 32624144-32624382 chr20 25063838-25065525 chr21 3 8065179-38066185 chr20 32856659-32857248 chr22 19967279-19967808 chr20 36012595-36013439 chr22 29709281-29712013 chr20 36226617-36226841 chr22 31198491-31199033 chr20 41817475-41819212 chr22 31500396-31501239 chr20 48184193-48184833 chr22 37212769-37213467 chr20 57089460-57090237 chr22 37911979-37912258 chr20 57426729-57427047 chr22 38476836-38478839 chr20 61703526-61704022 chr22 42305617-42307254 chr21 19617098-19617874 chr22 42322043-42322909 chr21 34395128-34400245 chr22 44726724-44727590 chr21 3 8076762-38077685 chr22 46318693-46319087 chr21 38079941-38081833 chr22 46440393-46441019 chr21 42218489-42219222 chr3: 3840513-3842772 chr22 19746155-19746369 chr3: 6902823-6903516 chr22 25081850-25082112 chr3 13114627-13115245 chr22: 37465056-37465331 chr3 19189688-19190100 chr22: 38379093-38379964 chr3 49947621-49948430 chr22: 39262338-39263211 chr3 55508336-55508708 chr22: 42685894-42686095 chr3 62354291-62355012 chr22: 44257942-44258612 chr3 62357639-62359774 chr22: 44287497-44288061 chr3 71834068-71834653 chr22: 48884884-48887043 chr3 87841796-87842563 chr22: 50496441-50497393 chr3 137482964-137484454 chr3 238391-240140 chr3 137489594-137491004 chr3 6904133-6904641 chr3 147108511-147111703 chr3 9177691-9178189 chr3 147113608-147114479 chr3 11034446-11035384 chr3 147130342-147130577 chr3 12838471-12838782 chr3 147131066-147131333 chr3 22413492-22414365 chr3 154146347-154146965 chr3 26664104-26664796 chr3 157821232-157821604 chr3 27771638-27771942 chr3 170303044-170303249 chr3 32861141-32861429 chr3 172165372-172166738 chr3 44063314-44063837 chr4 4868440-4869173 chr3 44596535-44597018 chr4 25090106-25090510 chr3 46618307-46618669 chr4 41749184-41749811 chr3 62356119-62356378 chr4 47034427-47034940 chr3 62356773-62357315 chr4 54966163-54968063 chr3 62362610-62363082 chr4 81119095-81119391 chr3 63263989-63264205 chr4 90228714-90229010 chr3 64253533-64253819 chr4 94755786-94756310 chr3 75667777-75669067 chr4 100870377-100871994 chr3 75955759-75956308 chr4 107956555-107957453 chr3 113160299-113160641 chr4 109093038-109094546 chr3 121902742-121903645 chr4 114900355-114900810 chr3 126113547-126113967 chr4 1223 01567-122302290 chr3 127633993-127634588 chr4 128544031-128544903 chr3 127794369-127796136 chr4 144620822-144622218 chr3 128719865-128721245 chr4 147559205-147561901 chr3 129693127-129694841 chr4 156680095-156681386 chr3 133393118-133393657 chr4 164264821-164265772 chr3 138656627-138659107 chr4 172733734-172735118 chr3 147126988-147128999 chr4 174430386-174430861 chr3 147138916-147139564 chr4 185939222-185942747 chr3 147142181-147142391 chr5 1879689-1879928 chr3 157812053-157812764 chr5 1881924-1887743 chr3 170303532-170303768 chr5 2748368-2757024 chr3 184056419-184056671 chr5 37834671-37835128 chr3 185911344-185912228 chr5 38257825-38259136 chr3 186078710-186080111 chr5 52777788-52777996 chr3 192125821-192127994 chr5 54527319-54527760 chr4 107146-107898 chr5 59189046-59189894 chr4 206377-206892 chr5 63256548-63257886 chr4 682724-683079 chr5 71014917-71015715 chr4 961347-962155 chr5 72529099-72529976 chr4 4859632-4860191 chr5 76932317-76933523 chr4 5709985-5710495 chr5 76934581-76935296 chr4 5891981-5892365 chr5 77805753-77806313 chr4 5894071-5895116 chr5 92923487-92924497 chr4 13524062-13526083 chr5 92939795-92940216 chr4 15779998-15780729 chr5 134363092-134365146 chr4 24801109-24801902 chr5 134366913-134367438 chr4 41869174-41869459 chr5 134374385-134376751 chr4 41875445-41875794 chr5 139138875-139139242 chr4 41880224-41880500 chr5 140052059-140053381 chr4 41882450-41882964 chr5 140305712-140307193 chr4 46995128-46995872 chr5 140798757-140799359 chr4 54975387-54976202 chr5 140810494-140812617 chr4 57521621-57522703 chr5 145718289-145720095 chr4 66535193-66535620 chr5 145725286-145725852 chr4 81109887-81110460 chr5 158523906-158524598 chr4 85403830-85404524 chr5 172665306-172666072 chr4 85413997-85414874 chr5 179228283-179229003 chr4 85422929-85423190 chr6 391188-393790 chr4 93226348-93227007 chr6 1381743-1385211 chr4 110222970-110224257 chr6 5997027-5997414 chr4 111554965-111555504 chr6 6007387-6007797 chr4 134069162-134070442 chr6 7229877-7230865 chr4 140201064-140201449 chr6 10390038-10390565 chr4 151504011-151505085 chr6 29894140-29895117 chr4 154709512-154710827 chr6 33393592-33393908 chr4 154712073-154712706 chr6 33655966-33656238 chr4 154713537-154714240 chr6 41908745-41909711 chr4 155663809-155664315 chr6 42072032-42072701 chr4 156129168-156130209 chr6 46655262-46656738 chr4 158143296-158144053 chr6 50682334-50683214 chr4 169799086-169799625 chr6 50791110-50791573 chr4 174422024-174422443 chr6 55039170-55039392 chr4 174427891-174428192 chr6 99275763-99276038 chr4 174437914-174438346 chr6 101846766-101847135 chr4 174439812-174440249 chr6 108485671-108490539 chr4 174448333-174448845 chr6 108491033-108491410 chr4 174450046-174451469 chr6 108497595-108497996 chr4 174451828-174452962 chr6 117198089-117198705 chr4 174459200-174460054 chr6 117591533-117592279 chr4 185937242-185937750 chr6 134210639-134211218 chr4 187219320-187219745 chr6 134638797-134639021 chr4 188916605-188916876 chr6 13 7242315-137245442 chr4 190938106-190938848 chr6 137814355-137815202 chr4 190939801-190940591 chr6 138745348-138745593 chr5 1874907-1879032 chr7 1362811-1363643 chr5 2738953-2741237 chr7 6590563-6590957 chr5 3590644-3592000 chr7 6661875-6662695 chr5 3594467-3603054 chr7 19145872-19146256 chr5 11384681-11385521 chr7 20370003-20371504 chr5 31193952-31194419 chr7 20830567-20830817 chr5 45695394-45696510 chr7 26415746-26416891 chr5 50685453-50686148 chr7 27146069-27146600 chr5 54519054-54519628 chr7 27182613-27185562 chr5 63255044-63255407 chr7 27227520-27229043 chr5 72526203-72526497 chr7 27278945-27279469 chr5 72594147-72595808 chr7 27282086-27283136 chr5 72676120-72678421 chr7 30721372-30722445 chr5 76923 887-76924502 chr7 37955622-37956555 chr5 76936126-76936984 chr7 49813008-49815752 chr5 77140542-77140914 chr7 56355508-56355798 chr5 77146998-77147785 chr7 87563342-87564571 chr5 77253832-77254049 chr7 90893567-90896683 chr5 77268350-77268787 chr7 95225503-95226194 chr5 87968635-87968907 chr7 96650221-96651551 chr5 87980878-87981272 chr7 96651963-96652246 chr5 87985470-87985810 chr7 97841636-97842005 chr5 88185224-88185589 chr7 113724924-113727795 chr5 115697134-115697589 chr7 130790358-130792773 chr5 12243 0676-122431443 chr7 136553854-136556194 chr5 134385967-134386370 chr7 155595692-155599414 chr5 140346105-140346931 chr7 155604725-155605095 chr5 140787447-140788044 chr7 156795355-156799394 chr5 140864527-140864748 chr8 21905461-21905757 chr5 146888750-146889840 chr8 25900562-25905842 chr5 148033472-148034080 chr8 55366180-55367628 chr5 158478378-158478630 chr8 65710990-65711722 chr5 159399004-159399928 chr8 70981873-70984888 chr5 170735169-170739863 chr8 105478672-105479340 chr5 170741603-170742751 chr8 120428398-120429178 chr5 170743178-170744107 chr8 143545445-143546178 chr5 172110282-172111166 chr8 144808221-144810978 chr5 172659049-172660277 chr8 144990270-145002135 chr5 172660720-172661133 chr9 17906419-17907488 chr5 172661486-172662228 chr9 21970913-21971190 chr5 172672311 -172672971 chr9: 22005887-22006229 chr5 174158680-174159729 chr9: 86152353-86153777 chr5 175085004-175085756 chr9: 95477296-95477708 chr5 178421225-178422337 chr9: 96713326-96718186 chr5 180486154-180486892 chr9: 97401286-97402067 chr6 1378445-1379318 chr9: 102590742-102591303 chr6 1393049-1394170 chr9: 112081402-112082905 chr6 1619093-1621094 chr9: 120175253-120177496 chr6 4079052-4079443 chr9: 122131086-122132214 chr6 5999149-5999787 chr9: 124413512-124414193 chr6 10381558-10382354 chr9: 124987743-124991086 chr6 10881846-10882051 chr9: 126773246-126780953 chr6 26614013-26614851 chr9: 129372737-129378106 chr6 27228100-27228364 chr9: 129386112-129389231 chr6 29595298-29595795 chr9: 131154346-131155923 chr6 30095173-30095610 chr9: 132459587-132460017 chr6 30139718-30140263 chr9: 133534534-133542394 chr6 33048416-33048814 chr9: 135039673-135039978 chr6 35479388-35479678 chr9: 135455164-135458586 chr6 37616722-37617179 chr9: 135461934-135462909 chr6 38682949-38683265 chr9: 135464586-135466240 chr6 41528266-41528900 chr9: 139096665-139096993 chr6 42145847-42146053 chr9: 139396205-139397040 chr6 42879279-42879623 chrX: 673 52650-67352923 chr6 50787286-50788091 chrX: 99891299-99891794 chr6 50810642-50810994 chrX: 152612775-152613464 chr6 50813314-50813699 chrl: 1474962-1475220 chr6 50818180-50818431 chrl: 2979275-2980758 chr6 70992040-70992912 chrl: 10764449-10764925 chr6 72298274-72298528 chrl: 12123488-12124148 chr6 78172231-78174088 chrl: 16860873-16862296 chr6 85472702-85474132 chrl: 18964180-18964401 chr6 99290279-99290771 chrl: 24229115-24229537 chr6 100038655-100039477 chrl: 32052471-32052771 chr6 100897080-100897621 chrl: 34642382-34643024 chr6 100903491-100903713 chrl: 36549554-36549965 chr6 100905444-100905697 chrl: 38219702-38220012 chr6 100905952-100906686 chrl: 38461584-38461988 chr6 100914946-100915245 chrl: 38941919-38942404 chr6 106429111-106429772 chrl: 39044059-39044561 chr6 106433984-106434459 chrl: 40769186-40769871 chr6 108495654-108495986 chrl: 41284847-41285149 chr6 110299365-110301267 chrl: 44031286-44031853 chr6 117869097-117869530 chrl: 47009575-47010132 chr6 127441553-127441760 chrl: 50880916-50881516 chr6 137809342-137810204 chrl 50881884-50882103 chr6 137816474-137817223 chrl 50892437-50893243 chr6 150335525-150336278 chrl 53527572-53528974 chr6 150358872-150359394 chrl 63795363-63796140 chr6 154360586-154361008 chrl 65991001-65991811 chr6 161188084-161188639 chrl 67218079-67218293 chr6 166579973-166583423 chrl 67773329-67773767 chr6 166666837-166667541 chrl 86621278-86622871 chr6 168841438-168841699 chrl 91183240-91184540 chr6 170732119-170732442 chrl 91185156-91185577 chr7 751712-752150 chrl 91190489-91192804 chr7 12151220-12151559 chrl 91300979-91301891 chr7 19184818-19185033 chrl 110610265-110613303 chr7 23287221-23287508 chrl 113265573-113265787 chr7 27134097-27134303 chrl 113286332-113287172 chr7 27147589-27148389 chrl 114695136-114696672 chr7 27198182-27198514 chrl 119526782-119527192 chr7 27203915-27206462 chrl 119529819-119530712 chr7 27260101-27260467 chrl 119543056-119543454 chr7 27291119-27292197 chrl 119549144-119551320 chr7 32110063-32110910 chrl 145075483-145075845 chr7 35296921-35298218 chrl 146552328-146552577 chr7 42267546-42267823 chrl 147782066-147782473 chr7 43152020-43153340 chrl 149332993-149333389 chr7 53286851-53287192 chrl 155147185-155147444 chr7 54612324-54612558 chrl 155264318-155265536 chr7 70596228-70598382 chrl 155290606-155291001 chr7 71800757-71802768 chrl 156863415-156863711 chr7 72838383-72838815 chrl 164545540-164545917 chr7 73894815-73895110 chrl 165324191-165326328 chr7 89747892-89749036 chrl 170630456-170630851 chr7 97361132-97363018 chrl 173638662-173639045 chr7 100075303-100075551 chrl 175568376-175568808 chr7 100817759-100817975 chrl 179544720-179545307 chr7 100823307-100823701 chrl 181287300-181287873 chr7 101005899-101007443 chrl 181452706-181453073 chr7 103085710-103086132 chrl 200009807-200010036 chr7 103968783-103969959 chrl 202162958-202163390 chr7 121940006-121940648 chrl 203044722-203045390 chr7 121950249-121950927 chrl 208132327-208133117 chr7 121956543-121957341 chrl 214153214-214153668 chr7 124404174-124404432 chrl 217310749-217311178 chr7 127990926-127992616 chrl 221050448-221050864 chr7 128555329-128556650 chrl 221060850-221061071 chr7 129422997-129423355 chrl 225865068-225865328 chr7 142494563-142495248 chrl: 226127112-226127695 chr7 143582125-143582610 chrl: 228785986-228786204 chr7 149389654-149389976 chrl: 231296559-231297345 chr7 149744402-149746469 chrl: 243646394-243646888 chr7 152621916-152622149 chrlO 1778784-1780018 chr7 153748407-153750444 chrlO 8076002-8077261 chr7 154001964-154002281 chrlO 8077829-8078378 chr7 155164557-155167854 chrlO 15761423-15762101 chr7 155174128-155175248 chrlO 16561604-16563822 chr7 155241323-155243757 chrlO 22623350-22625875 chr7 155258827-155261403 chrlO 22634000-22634862 chr7 155302253-155303158 chrlO 22764708-22767050 chr7 156409023-156409294 chrlO 23461300-23461610 chr7 156409577-156409865 chrlO 23462224-23463889 chr7 156801418-156801632 chrlO 23480697-23482455 chr7 156871054-156871297 chrlO 23983366-23984978 chr7 158936507-158938492 chrlO 26504383-26507434 chr8 4848968-4852635 chrlO 27547668-27548402 chr8 9760750-9761643 chrlO 43428167-43429460 chr8 9762661-9764748 chrlO 48438411-48439320 chr8 11536767-11538961 chrlO 63212495-63213009 chr8 11557852-11558252 chrlO 71331449-71331691 chr8 11565217-11567212 chrlO 75407413-75407706 chr8 21644908-21647845 chrlO 76573195-76573507 chr8 23562475-23565175 chrlO 94180315-94180754 chr8 23567180-23567678 chrlO 94455524-94455896 chr8 24812946-24814299 chrlO 94828102-94829040 chr8 26721642-26724566 chrlO 99789614-99791320 chr8 37822486-37824008 chrlO 100992156-100992687 chr8 41424341-41425300 chrlO 101282725-101282934 chr8 49468683-49468959 chrlO 101290025-101290338 chr8 50822270-50822860 chrlO 102279162-102279730 chr8 53851701-53854426 chrlO 102475276-102475579 chr8 55370170-55372525 chrlO 102891010-102891794 chr8 55378928-55380186 chrlO 102905714-102906693 chr8 57358126-57359415 chrlO 102996034-102996646 chr8 65281903-65283043 chrlO 103043990-103044480 chr8 65286067-65286659 chrlO 108923780-108924805 chr8 65290108-65290946 chrlO 109674196-109674964 chr8 68864584-68864946 chrlO 110671724-110672326 chr8 72468560-72469561 chrlO 111216604-111217083 chr8 85096759-85097247 chrlO 118030732-118034230 chr8 86350765-86351196 chrlO 118892161-118892639 chr8 87081653-87082046 chrlO 118893527-118894432 chr8 97169731-97170432 chrlO 119494493-119494991 chr8 97171805-97172022 chrlO 120353692-120355821 chr8: 98289604-98290404 chrlO 121577529-121578385 chr8: 99960497-99961438 chrlO 123922850-123923542 chr8: 99984584-99985072 chrlO 124901907-124902617 chr8: 99985733-99986983 chrlO 125425495-125426642 chr8: 101117922-101118693 chrlO 125650820-125651373 chr8: 130995921-130996149 chrlO 125732220-125732843 chr8: 132052203-132054749 chrlO 130338695-130338994 chr8: 139508795-139509774 chrlO 130508443-130508658 chr8: 142528185-142529029 chrlO 134597357-134602649 chr8: 145103285-145108027 chrll 626728-628037 chr8: 145925410-145926101 chrll 636435-636668 chr9: 969529-973276 chrl 1 636906-640628 chr9: 16726859-16727273 chrll 2890388-2891337 chr9: 19788215-19789288 chrl 1 14995128-14995908 chr9: 23820691-23822135 chrl 1 20618197-20619920 chr9: 23850910-23851522 chrll 27743472-27744564 chr9: 32782936-32783625 chrl 1 31827696-31827921 chr9: 36739534-36739782 chrll 31841315-31842003 chr9: 37002489-37002957 chrll 31847132-31847958 chr9: 77112712-77113583 chrl 1 43568921-43569854 chr9: 77113709-77113927 chrll 44325657-44326517 chr9: 79633326-79636030 chrll 60718428-60718888 chr9: 79637814-79638169 chrl 1 64478843-64479598 chr9: 91792662-91793611 chrll 64815040-64815722 chr9: 96108466-96108992 chrl 1 65409636-65410127 chr9: 96710811-96711717 chrl 1 65816404-65816665 chr9: 98111364-98112362 chrll 68622108-68622339 chr9: 100610696-100611517 chrl 1 70508328-70508617 chr9: 100619984-100620192 chrll 71952112-71952528 chr9: 104499849-104501076 chrll 88241710-88242562 chr9: 115822071-115823416 chrl 1 89224416-89224718 chr9: 120507227-120507642 chrll 105481126-105481422 chr9: 123656750-123656972 chrll 115630398-115631117 chr9: 134429866-134430491 chrl 1 119293320-119293943 chr9: 136294738-136295236 chrll 123066517-123066986 chr9: 137967110-137967727 chrl 1 128419198-128419513 chr9: 139715663-139716441 chrl 1 128694084-128694688 chrll 131780328-131781532 chrl 1 132813562-132814395 chrll 132934059-132934291 chrll 132952538-132953307 chrl 1 133994709-133995090 chrl 2 186863-187610 chrl 2 3308812-3310270 chrl2: 5153012-5154346 chrl2: 14134626-14135242 chrl2: 41086522-41087102 chrl2: 483991...
Claims
1. A tiered, multipart method for tracking tumor heterogeneity across at least first and second biological samples obtained from a subject, the method comprising:(a) performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained from the subject at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA;(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing, a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein the second biological sample was obtained from the subject at a second timepoint subsequent to the first timepoint;(iii) determining a change in signal between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background-corrected methylation information from the second intra-individual analysis; and(iv) performing a second analysis comprising analyzing the determined change in signal to track tumor heterogeneity across the first biological sample and the second biological sample.
2. A method for tracking tumor heterogeneity in a patient during or subsequent to administration of a tumor therapeutic, or in a patient being considered for administration of a tumor therapeutic, comprising:(a) confirming that a first biological sample of the patient is not identified as not at risk of containing circulating tumor DNA;WO 2025 / 147572 PCT / US2025 / 010181(b) responsive to determining that the first biological sample is not identified as not atrisk:(i) performing, at a baseline timepoint, a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing, at a second timepoint, a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein between the baseline timepoint and second timepoint the patient may be administered, or continues to be administered, one or more tumor therapeutics;(iii) determining a change in signal between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background-corrected methylation information from the second intra-individual analysis; and(iv) performing a second analysis comprising analyzing the determined change in signal to assess tumor heterogeneity across the first biological sample and the second biological sample and therefore track the patient’s therapeutic progress and / or assess the tumor therapeutic.
3. The method of claim 1 or 2, wherein determining the change in signal comprises determining a difference between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background-corrected methylation information from the second intra-individual analysis.
4. The method of any one of claims 1-3, wherein the first set of background-corrected methylation information or the second set of background-corrected methylation information comprises methylation statuses for a plurality of genomic sites.
5. The method of claim 4, wherein the plurality of genomic sites comprise a plurality of CpG sites.WO 2025 / 147572 PCT / US2025 / 0101816. The method of claim 5, wherein the plurality of CpG sites are located in one or more CpG islands or portions of one or more CpG islands shown in Tables 1-4.
7. The method of claim 4, wherein the first set of background-corrected methylation information and the second set of background-corrected methylation information comprises methylation statuses for a plurality of CpG sites.
8. The method of claim 7, wherein the plurality of CpG sites of the first set of background-corrected methylation information are the same plurality of CpG sites of the second set of background-corrected methylation information.
9. The method of any one of claims 1-8, wherein performing the first intra-individual analysis comprises:obtaining target nucleic acids and reference nucleic acids from the first biological sample obtained from the subject;performing bisulfite conversion of the target nucleic acids and the reference nucleic acids;selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids;generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; andusing a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the first set of background-corrected methylation information.
10. The method of claim 9, wherein the reference nucleic acids from the first biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.
11. The method of any one of claims 1-10, wherein the first set of background-corrected methylation information comprises phased sequencing information.
12. The method of claim 11, wherein the phased sequencing information of the first set of background-corrected methylation information is generated by:obtaining or having obtained sequence reads of cell-free DNA from the first sample;WO 2025 / 147572 PCT / US2025 / 010181obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length;attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; andaligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.
13. The method of claim 12, wherein the phased sequencing information of cell-free DNA comprises methylation statuses for a plurality of genomic sites of the cell-free DNA.
14. The method of claim 13, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.
15. The method of claim 12, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.
16. The method of claim 15, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.
17. The method of claim 16, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.
18. The method of claim 16 or 17, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.
19. The method of any one of claims 12-18, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.
20. The method of any one of claims 12-19, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, atWO 2025 / 147572 PCT / US2025 / 010181least 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.
21. The method of any one of claims 1-8, wherein performing the second intra-individual analysis comprises:obtaining target nucleic acids and reference nucleic acids from the second biological sample obtained from the subject;performing bisulfite conversion of the target nucleic acids and the reference nucleic acids;selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids;generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; andusing a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the second set of background-corrected methylation information.
22. The method of claim 21, wherein the reference nucleic acids from the second biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.
23. The method of any one of claims 1-22, wherein the first set of background-corrected methylation information and / or the second set of background-corrected methylation information comprise a high resolution measure of methylation.
24. The method of claim 23, wherein the high resolution measure of methylation comprises a total quantity of consecutively methylated CpG sites within target regions.
25. The method of claim 24, wherein the total quantity of consecutively methylated CpG sites within target regions comprises the total quantity of 3, 4, or 5 consecutively methylated CpG sites within target regions.
26. The method of claim 23, wherein the high resolution measure of methylation comprises methylation statuses of a plurality of CpG sites from a haplotype.WO 2025 / 147572 PCT / US2025 / 01018127. The method of any one of claims 1-26, wherein the second set of background-corrected methylation information comprises phased sequencing information.
28. The method of claim 27, wherein the phased sequencing information of the second setof background-corrected methylation information is generated by:obtaining or having obtained sequence reads of cell-free DNA from the second sample;obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length;attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; andaligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.
29. The method of claim 28, wherein the phased sequencing information of cell-free DNA comprises methylation statuses for a plurality of genomic sites of the cell-free DNA.
30. The method of claim 29, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.
31. The method of claim 28, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.
32. The method of claim 31, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.
33. The method of claim 32, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.
34. The method of claim 32 or 33, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.WO 2025 / 147572 PCT / US2025 / 01018135. The method of any one of claims 28-34, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.
36. The method of any one of claims 28-35, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, at least 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.
37. The method of any one of claims 1-36, wherein the nucleic acid sequence information of the first analysis comprises methylation sequence information.
38. The method of claim 37, wherein the methylation sequence information of the first analysis comprises methylation statuses for a plurality of genomic sites.
39. The method of claim 38, wherein the plurality of genomic sites comprise a plurality of CpG sites.
40. The method of claim 38, wherein the nucleic acid sequence information of the first analysis comprises a measure of overall methylation across the plurality of genomic sites.
41. The method of claim 40, wherein the measure of overall methylation comprises a total number of methylated genomic sites or an average number of methylated genomic sites.
42. The method of any one of claims 1-41, wherein performing the first analysis of nucleic acid sequence information comprises applying a trained machine learning model.
43. The method of any one of claims 1-42, wherein the method delivers improved performance as a function of resource consumption in comparison to the single tier method.
44. The method of any one of claims 1-42, wherein the method achieves an improved performance metric in comparison to a single tier method.
45. The method of any one of claims 1-42, wherein the method tracks tumor heterogeneity of one or more of the early stage cancers.
46. The method of claim 45, wherein the one or more of the early stage cancers is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, cardiac cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumor, gestational trophoblastic cancer, hairy cell leukemia, liver cancer, Hodgkin lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, mouth cancer, multiple endocrine neoplasia syndromes, multiple myeloma neoplasms, myelodysplastic neoplasms, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.
47. The method of claim 45, wherein the one or more early stage cancers is a preclinical phase cancer.
48. The method of claim 47, wherein the preclinical phase cancer is stage I or stage II cancer.
49. The method of any one of claims 1-48, wherein the nucleic acid sequence information, the background-corrected methylation information of the first intra-individual analysis, and / or the background-corrected methylation information of the second intraindividual analysis is obtained from an assay, wherein the assay comprises performing one or more of:a. sequencing of nucleic acids;b. hybrid capture;c. methylation-specific PCR;d. an assay that generates methylation information; ande. sequencing a clone library generated from a template immortalized library.WO 2025 / 147572 PCT / US2025 / 01018150. The method of any one of claims 1-49, wherein each of the first biological sample and the second biological sample independently comprises any one of a blood sample, a stool sample, a urine sample, a mucous sample, or a saliva sample.
51. The method of any one of claims 1-50, wherein each of the first biological sample and the second biological sample is a blood sample.
52. The method of claim 51, wherein each of the first biological sample and the second biological sample does not comprise an invasive biopsy sample.
53. The method of any one of claims 1-52, wherein the second analysis comprises whole genome sequencing, optionally whole genome bisulfite sequencing.
54. The method of any one of claim 1-53, wherein the subject received a tumor therapeutic prior to the first timepoint.
55. The method of any one of claim 1-53, wherein subsequent to the first timepoint and prior to the second timepoint, the subject received a tumor therapeutic.
56. The method of claim 54 or 55, further comprising determining an efficacy of the tumor therapeutic based on the tracked tumor heterogeneity.
57. The method of claim 56, wherein if the tracked tumor heterogeneity indicates a stable or increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the tumor therapeutic lacks efficacy.
58. The method of claim 57, further comprising selecting a new tumor therapeutic for the subject responsive to determining that the tumor therapeutic lacks efficacy.
59. The method of claim 56, wherein if the tracked tumor heterogeneity indicates a reducing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the tumor therapeutic achieves therapeutic efficacy.
60. The method of any one of claims 1-59, wherein prior to (a), a prior sample obtained from the subject was previously determined to be not at risk for containing circulating tumor DNA.WO 2025 / 147572 PCT / US2025 / 01018161. The method of claim 60, wherein further responsive to determining that the first biological sample is not identified as not at risk, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.
62. The method of claim 60, wherein if the tracked tumor heterogeneity indicates an increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.
63. A tiered, multipart method for assessing tumor heterogeneity across at least first and second biological samples obtained from a subject, the method comprising:(a) performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA,(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing a second analysis of the first biological sample comprising analyzing the background-corrected methylation information to predict a tumor heterogeneity state;(c) determining an updated tumor heterogeneity state by:(i) performing a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein the second biological sample was obtained from the subject at a second timepoint subsequent to the first timepoint; and(ii) performing a second analysis of the second biological sample comprising analyzing the background-corrected methylation information to predict the updated tumor heterogeneity state; and(d) comparing the tumor heterogeneity state from the first biological sample to the updated tumor heterogeneity state from the second biological sample to track tumor heterogeneity across the first biological sample and the second biological sample.
64. A method for assessing tumor heterogeneity in a patient during or subsequent to administration of a tumor therapeutic, or in a patient being considered for administration of a tumor therapeutic, comprising:(a) confirming that a first biological sample of the patient is not identified as not at risk of containing circulating tumor DNA;(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing, at a baseline timepoint, a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing, at a second timepoint, a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein between the baseline timepoint and second timepoint the patient may be administered, or continues to be administered, one or more tumor therapeutics;(iii) determining a change in signal between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background-corrected methylation information from the second intra-individual analysis; and(iv) performing a second analysis comprising analyzing the determined change in signal to assess tumor heterogeneity across the first biological sample and thesecond biological sample and therefore track the patient’s therapeutic progress and / or assess the tumor therapeutic.
65. The method of claim 63 or 64, wherein the first set of background-corrected methylation information or the second set of background-corrected methylation information comprises methylation statuses for a plurality of genomic sites.
66. The method of claim 65, wherein the plurality of genomic sites comprise a plurality of CpG sites.
67. The method of claim 66, wherein the plurality of CpG sites are located in one or more CpG islands or portions of one or more CpG islands shown in Tables 1-4.
68. The method of claim 65, wherein the first set of background-corrected methylation information and the second set of background-corrected methylation information comprises methylation statuses for the plurality of CpG sites.
69. The method of claim 68, wherein the plurality of CpG sites of the first set of background-corrected methylation information are the same plurality of CpG sites of the second set of background-corrected methylation information.
70. The method of any one of claims 63-69, wherein performing the first intra-individual analysis comprises:obtaining target nucleic acids and reference nucleic acids from the first biological sample obtained from the subject;performing bisulfite conversion of the target nucleic acids and the reference nucleic acids;selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids;generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; andusing a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the first set of background-corrected methylation information.WO 2025 / 147572 PCT / US2025 / 01018171. The method of claim 70, wherein the reference nucleic acids from the first biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.
72. The method of any one of claims 63-71, wherein the first set of background-corrected methylation information comprise a high resolution measure of methylation.
73. The method of claim 72, wherein the high resolution measure of methylation comprises a total quantity of consecutively methylated CpG sites within target regions.
74. The method of claim 73, wherein the total quantity of consecutively methylated CpG sites within target regions comprises the total quantity of 3, 4, or 5 consecutively methylated CpG sites within target regions.
75. The method of claim 72, wherein the high resolution measure of methylation comprises methylation statuses of a plurality of CpG sites from a haplotype.
76. The method of any one of claims 63-75, wherein the first set of background-corrected methylation information comprises phased sequencing information.
77. The method of claim 76, wherein the phased sequencing information of the first set of background-corrected methylation information is generated by:obtaining or having obtained sequence reads of cell-free DNA from the first sample;obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length;attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; andaligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.
78. The method of claim 77, wherein the phased sequencing information comprises methylation statuses for a plurality of genomic sites.
79. The method of claim 78, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.WO 2025 / 147572 PCT / US2025 / 01018180. The method of claim 77, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.
81. The method of claim 80, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.
82. The method of claim 81, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.
83. The method of claim 81 or 82, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.
84. The method of any one of claims 77-83, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.
85. The method of any one of claims 77-84, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, at least 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.
86. The method of any one of claims 63-69, wherein performing the second intraindividual analysis comprises:obtaining target nucleic acids and reference nucleic acids from the second biological sample obtained from the subject;performing bisulfite conversion of the target nucleic acids and the reference nucleic acids;selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids;generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; andWO 2025 / 147572 PCT / US2025 / 010181using a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the first set of background-corrected methylation information.
87. The method of claim 86, wherein the reference nucleic acids from the second biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.
88. The method of any one of claims 86-87, wherein the second set of background-corrected methylation information comprises a high resolution measure of methylation.
89. The method of claim 88, wherein the high resolution measure of methylation comprises a total quantity of consecutively methylated CpG sites within target regions.
90. The method of claim 89, wherein the total quantity of consecutively methylated CpG sites within target regions comprises the total quantity of 3, 4, or 5 consecutively methylated CpG sites within target regions.
91. The method of claim 88, wherein the high resolution measure of methylation comprises methylation statuses of a plurality of CpG sites from a haplotype.
92. The method of any one of claims 63-91, wherein the second set of background-corrected methylation information comprises phased sequencing information.
93. The method of claim 92, wherein the phased sequencing information of the second set of background-corrected methylation information is generated by:obtaining or having obtained sequence reads of cell-free DNA from the second sample;obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length;attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; andaligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.WO 2025 / 147572 PCT / US2025 / 01018194. The method of claim 93, wherein the phased sequencing information of cell-free DNA comprises methylation statuses for a plurality of genomic sites of the cell-free DNA.
95. The method of claim 94, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.
96. The method of claim 93, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.
97. The method of claim 96, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.
98. The method of claim 97, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.
99. The method of claim 97 or 98, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.
100. The method of any one of claims 93-99, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.
101. The method of any one of claims 93-100, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, at least 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.
102. The method of any one of claims 63-101, wherein the nucleic acid sequence information of the first analysis comprises methylation sequence information.
103. The method of claim 102, wherein the methylation sequence information of the first analysis comprises methylation statuses for a plurality of genomic sites.
104. The method of claim 103, wherein the plurality of genomic sites comprise a plurality of CpG sites.
105. The method of claim 103, wherein the nucleic acid sequence information of the first analysis comprises a measure of overall methylation across the plurality of genomic sites.
106. The method of claim 105, wherein the measure of overall methylation comprises a total number of methylated genomic sites or an average number of methylated genomic sites.
107. The method of any one of claims 63-106, wherein performing the first analysis of nucleic acid sequence information comprises applying a trained machine learning model.
108. The method of any one of claims 63-107, wherein the tiered, multipart method delivers improved performance as a function of resource consumption in comparison to the single tier method.
109. The method of any one of claims 63-107, wherein the tiered, multipart method achieves an improved performance metric in comparison to a single tier method.
110. The method of any one of claims 63-107, wherein the tiered, multipart method tracks tumor heterogeneity of one or more of the early stage cancers.
111. The method of claim 110, wherein the one or more of the early stage cancers is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, cardiac cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumor, gestational trophoblastic cancer, hairy cell leukemia, liver cancer, Hodgkin lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, mouth cancer, multiple endocrine neoplasiasyndromes, multiple myeloma neoplasms, myelodysplastic neoplasms, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.
112. The method of claim 110, wherein the one or more early stage cancers is a preclinical phase cancer.
113. The method of claim 112, wherein the preclinical phase cancer is stage I or stage II cancer.
114. The method of any one of claims 63-113, wherein the nucleic acid sequence information, the background-corrected methylation information of the first intra-individual analysis, and / or the background-corrected methylation information of the second intraindividual analysis is obtained from an assay, wherein the assay comprises performing one or more of:a. sequencing of nucleic acids;b. hybrid capture;c. methylation-specific PCR;d. an assay that generates methylation information; ande. sequencing a clone library generated from a template immortalized library.
115. The method of any one of claims 63-114, wherein each of the first biological sample and the second biological sample independently comprises any one of a blood sample, a stool sample, a urine sample, a mucous sample, or a saliva sample.
116. The method of any one of claims 63-115, wherein each of the first biological sample and the second biological sample is a blood sample.
117. The method of claim 116, wherein each of the first biological sample and the second biological sample does not comprise an invasive biopsy sample.
118. The method of any one of claims 63-117, wherein the second analysis of the first biological sample or the second analysis of the second biological sample comprises whole genome sequencing, optionally whole genome bisulfite sequencing.
119. The method of any one of claim 63-118, wherein the subject received a tumor therapeutic prior to the first timepoint.
120. The method of any one of claim 63-118, wherein subsequent to the first timepoint and prior to the second timepoint, the subject received a tumor therapeutic.
121. The method of claim 119 or 120, further comprising determining an efficacy of the tumor therapeutic based on the assessed tumor heterogeneity.
122. The method of claim 121, wherein if the tracked tumor heterogeneity indicates a stable or increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the assessed tumor therapeutic lacks efficacy.
123. The method of claim 122, further comprising selecting a new intervention for the subject responsive to determining that the assessed tumor therapeutic lacks efficacy.
124. The method of claim 121, wherein if the tracked tumor heterogeneity indicates a reducing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the assessed tumor therapeutic achieves therapeutic efficacy.
125. The method of any one of claims 63-124, wherein prior to (a), a prior sample obtained from the subject was previously determined to be not at risk for containing circulating tumor DNA.
126. The method of claim 125, wherein further responsive to determining that the first biological sample is not identified as not at risk, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.
127. The method of claim 125, wherein if the tracked tumor heterogeneity indicates an increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.
128. A tiered, multipart method for determining whether a prior sample obtained from a subject was a false negative sample, the method comprising:(a) performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained from the subject at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA, wherein the prior sample was obtained from the subject prior to the first biological sample,(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing one or more intra-individual analyses using the first biological sample or one or more additional biological samples, wherein performing each intraindividual analysis involves generating background-corrected methylation information representing a difference between methylation information from target nucleic acids and methylation information from reference nucleic acids from one of the first biological sample or one or more additional biological samples;(ii) performing a longitudinal analysis comprising analyzing generated background-corrected methylation information from each of the one or more intra-individual analyses; and(iii) determining that the prior sample obtained from a subject was a false negative sample.
129. The method of claim 128, wherein determining that the prior sample obtained from a subject was a false negative sample is responsive to determining that the first biological sample is not identified as not at risk.
130. The method of claim 128, wherein determining that the prior sample obtained from a subject was a false negative sample is responsive to performing the longitudinal analysis.
131. The method of any one of claims 128-130, wherein performing the longitudinal analysis comprising tracking tumor heterogeneity across the first biological sample and one or more additional biological samples.
132. A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-127.13 3. A sy stem compri sing:a processor; anda non-transitory computer readable medium comprising instructions that, when executedby a processor, cause the processor to perform the method of any one of claims 1-127.