Multi-tiered testing for tracking cancer heterogeneity

A multi-tiered analysis using nucleic acid screening and intra-individual methods effectively identifies and tracks tumor heterogeneity in large populations, enhancing cancer diagnosis accuracy and reducing resource use for rare cancer detection.

WO2025147572A1PCT designated stage expired Publication Date: 2025-07-10FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/010181
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2025-01-03
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Current diagnostic technologies struggle to accurately identify individuals with rare cancers in large populations due to poor performance and high resource consumption, with POC tests lacking accuracy and centralized tests being invasive and expensive.

Method used

A multi-tiered analysis approach involving a first screen to eliminate negative individuals, followed by intra-individual analyses using target and reference nucleic acids to generate background-corrected methylation information, and a second analysis to track tumor heterogeneity over time.

Benefits of technology

This method significantly improves sensitivity, specificity, and resource efficiency by rapidly identifying a large proportion of non-cancer cases while accurately diagnosing cancer, reducing resource consumption by up to 90% compared to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025010181_10072025_PF_FP_ABST
    Figure US2025010181_10072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints. Each sample undergoes at least an intra-individual analysis to generate background-corrected methylation information. The change in the background-corrected methylation information across the different samples is informative for tracking a change in the tumor heterogeneity. The change in tumor heterogeneity is useful e.g., for providing a guided therapy.
Need to check novelty before this filing date? Find Prior Art

Description

MULTI-TIERED TESTING FOR TRACKING CANCER HETEROGENEITYCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 636,405 filed April 19, 2024, and U.S. Provisional Patent Application No. 63 / 617,989 filed January 5, 2024 the entire disclosure of each of which is hereby incorporated by reference in its entirety for all purposes.BACKGROUND

[0002] Diagnostic technologies include simple, point of care (POC) tests applied to large populations to identify relatively common diseases as well as complex, centralized tests applied to select populations. However, although POC tests can be applied to large populations, they are incapable of identifying individuals for cancer at a high enough accuracy to be feasible for implementation. Similarly, although complex, centralized testing can be deployed for rare population testing, such testing is often invasive, expensive, and fails when applied for detecting rare cancers in large patient populations. For example, complex, centralized testing suffers from poor performance (e.g., high number of false positives and / or low positive predictive value) when attempting to diagnose rare cancers in large patient populations. Thus, current POC tests are not suitable for identifying individuals with cancer and for tracking such individuals over time.SUMMARY

[0003] Disclosed herein are methods involving a multiple tiered analysis for tracking tumor heterogeneity in subjects. In particular, the methods disclosed herein involving a multiple tiered analysis are useful for tracking tumor heterogeneity in individuals from a large population (e.g., millions of individuals) who have a rare cancer. The multiple tiered analysis involves a first screen, which eliminates a large proportion of individuals who are identified as negative for cancer. For subjects that are identified as not negative for cancer, they can be provided an intervention (e.g., a tumor therapeutic). These subjects undergo additional analyses (e.g., one or more intra-individual analysis and / or a second analysis) which can be performed using samples obtained from the subjects across different timepoints. For example, intra-individual analyses can be conducted for each sample obtained from the subject. By doing so, a change in tumor heterogeneity can be determined which isinformative for determining the efficacy of the provided intervention. Altogether, the multiple tiered analysis can be useful e.g., for guided therapy.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] These and other features, aspects, and advantages of the present invention will become better understood with regard to the following description and accompanying drawings. It is noted that wherever practicable, similar or like reference numbers may be used in the figures and may indicate similar or like functionality. For example, a letter after a reference numeral, such as “third party entity 155 A,” indicates that the text refers specifically to the element having that particular reference numeral. A reference numeral in the text without a following letter, such as “third party entity 155,” refers to any or all of the elements in the figures bearing that reference numeral (e.g. “third party entity 155” in the text refers to reference numerals “third party entity 155 A” and / or “third party entity 155B” in the figures).

[0005] Figure (FIG.) 1 A depicts an overall flow process of the multiple-tiered process for tracking tumor heterogeneity, in accordance with an embodiment.

[0006] FIG. IB depicts an overall flow process of the multiple-tiered process for tracking tumor heterogeneity, in accordance with a second embodiment.

[0007] FIG. 1C depicts an overall system environment including a tumor heterogeneity system, in accordance with an embodiment.

[0008] FIG. 2A depicts a block diagram of the tumor heterogeneity system, in accordance with an embodiment.

[0009] FIG. 2B depicts an example conversion of nucleic acids, in accordance with an embodiment.

[0010] FIG. 2C shows the results of nitrite conversion on select nucleotides, in accordance with a second embodiment. Figure adapted from Li el al. (2022) Genome Biology 23 : 122.

[0011] FIG. 3 A depicts example methylation information useful for determining whether an individual is at risk for cancer, in accordance with an embodiment.

[0012] FIG. 3B shows an example flow process for determining whether an individual is at risk for cancer, in accordance with an embodiment.

[0013] FIG. 3C depicts an example process of combining sequence information of target nucleic acids and reference nucleic acids to generate a signal informative for determining presence or absence of cancer, in accordance with an embodiment.

[0014] FIG. 3D is an illustrative example of a signal informative for cancer, in accordance with an embodiment.

[0015] FIG. 3E shows aligned sequence reads of an analyte and a corresponding window of a kmer size, in accordance with an embodiment.

[0016] FIG. 3F shows the generation of metrics from sequence reads across 2kpossible patterns, in accordance with an embodiment.

[0017] FIG. 3G shows an example data structure including information useful for training machine learning models, in accordance with an embodiment.

[0018] FIG. 4A shows an example flow process involving a first and second intra-individual analyses, in accordance with a first embodiment.

[0019] FIG. 4B shows an example flow process involving a first and second intra-individual analyses, in accordance with a second embodiment.

[0020] FIG. 5 illustrates an example computer for implementing the entities shown in FIGs. 1A-1C, 2A, 3A-3G, and 4A-4B.

[0021] FIG. 6 shows example performance of different tiers of the multiple tier analysis for diagnosing individuals with cancer (e.g., prostate cancer).

[0022] FIG. 7 depicts performance of a single tier analysis and a two-tier analysis of a population involving 1046 samples.

[0023] FIG. 8 shows an example sample from which target nucleic acids and reference nucleic acids are obtained.DETAILED DESCRIPTIONDefinitions

[0024] Terms used in the claims and specification are defined as set forth below unless otherwise specified.

[0025] The terms “subject,” “patient,” and “individual” are used interchangeably and encompass a cell, tissue, or organism, human or non-human, male or female.

[0026] The term “sample” can include a single cell or multiple cells or fragments of cells or an aliquot of body fluid, such as a blood sample, taken from a subject, by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, or intervention or other means known in the art. Examples of an aliquot of body fluid include amniotic fluid, aqueous humor, bile, lymph, breast milk, interstitial fluid, blood, blood plasma, cerumen (earwax), Cowper’s fluid (pre-ejaculatoryfluid), chyle, chyme, female ejaculate, menses, mucus, saliva, urine, vomit, tears, vaginal lubrication, sweat, serum, semen, sebum, pus, pleural fluid, cerebrospinal fluid, synovial fluid, intracellular fluid, and vitreous humour.

[0027] The term “obtaining information,” “obtaining marker information,” and “obtaining sequence information” encompasses obtaining information that is determined from at least one sample. Obtaining information (e.g., marker information or sequence information) encompasses obtaining a sample and processing the sample to experimentally determine the information (e.g., marker information or sequence information). The phrase also encompasses receiving the information, e.g., from a third party that has processed the sample to experimentally determine the information.

[0028] The terms “marker,” “markers,” “biomarker,” and “biomarkers” encompass, without limitation, lipids, lipoproteins, proteins, cytokines, chemokines, growth factors, peptides, nucleic acids (e.g., DNA or RNA), genes, and oligonucleotides, together with their related complexes, metabolites, mutations, variants, polymorphisms, modifications, fragments, subunits, degradation products, elements, and other analytes or sample-derived measures. A marker can also include mutated proteins, mutated nucleic acids, variations in copy numbers, and / or transcript variants, in circumstances in which such mutations, variations in copy number and / or transcript variants are useful for generating a prediction model, or are useful in prediction models developed using related markers (e.g., non-mutated versions of the proteins or nucleic acids, alternative transcripts, etc.).

[0029] The term “screen” or a “first analysis” refers to a step in the first tier of a multiple tiered analysis. The screen achieves a high specificity and removes a large majority of true negatives (e.g., individuals not at risk of a cancer). In various embodiments, the “screen” refers to an in silico screen that involves application of a machine learning model. For example, such a machine learning model may analyze sequence information (e.g., methylation information) and predicts whether individuals are likely to be at risk of the cancer.

[0030] The phrase “second analysis” refers to a step in the second tier of a multiple tiered analysis. The second analysis is performed on individuals who were identified, using the screen, as not negative for cancer. Thus, the second analysis achieves a higher positive predictive value than the screen, given that the screen removes a large proportion of the true negatives. In various embodiments, the “second analysis” refers to an in silico analysis that involves application of a machine learning model that analyzes sequence information (e.g.,methylation information). The second analysis can predict whether individuals have cancer. In various embodiments, the second analysis is implemented to predict a change in tumor heterogeneity for purposes of tracking tumor heterogeneity in a subject.

[0031] The phrase “intra-individual analysis” refers to an analysis performed for an individual that removes baseline biological signatures that are less informative for determining whether the individual is at risk for cancer. In various embodiments, the intra- individual analysis involves combining information from target nucleic acids and reference nucleic acids of an individual to generate a signal informative for determining presence or absence of cancer within the individual. By combining the information from the target nucleic acids and the reference nucleic acids, the generated signal can be more informative of presence or absence of cancer in comparison to a signal derived from the target nucleic acids alone.

[0032] The phrase “target nucleic acids” refers to nucleic acids of an individual that contain at least signatures that may be informative for determining presence or absence of cancer. The target nucleic acids may further include baseline biological signatures of the individual that are not informative or less informative. In various embodiments, target nucleic acids may be nucleic acids derived from a diseased cell that is associated with cancer. For example, target nucleic acids may be cell-free nucleic acids originating from cancer cells. Target nucleic acids can be any of DNA, cDNA, or RNA. In particular embodiments, target nucleic acids include DNA.

[0033] The phrase “reference nucleic acids” refers to nucleic acids of an individual that contain baseline biological signatures of the individual. Here, the baseline biological signatures of the individual may be present when the individual is healthy, and therefore, the baseline biological signatures are less informative for determining presence or absence of cancer in comparison to sequence information of the target nucleic acids. Reference nucleic acids can be any of DNA, cDNA, or RNA. In particular embodiments, reference nucleic acids include DNA.

[0034] It must be noted that, as used in the specification, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise.Overview of Multiple Tier Analysis

[0035] Disclosed herein is a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints. For example, methods disclosed herein are useful for detecting circulating tumor DNA from samples obtained from a subjectacross two or more timepoints. Determining the change in circulating tumor DNA from samples obtained from the subject across two or more timepoints enables tracking of the tumor heterogeneity. In various embodiments, tracking tumor heterogeneity is informative for determining whether an intervention (e.g., a tumor therapeutic) is efficacious. Therefore, tracking tumor heterogeneity can be useful for e.g., guided therapy.

[0036] In various embodiments, the tiered, multipart method involves performing a first analysis of nucleic acid sequence information that was derived from a first assay performed on a biological sample obtained from the subject. This first analysis identifies whether the biological sample is at risk or not at risk of containing circulating tumor DNA. In various embodiments, for a biological sample that is determined as not negative for containing circulating tumor DNA, the multipart method further includes performing an intra-individual analysis and a second analysis. In various embodiments, the intra-individual analysis includes obtaining target nucleic acids and reference nucleic acids from the biological sample or an additional biological sample obtained from the individual; processing the target nucleic acids and reference nucleic acids to generate a dataset comprising methylation information from the target nucleic acids and methylation information from the reference nucleic acids; and using a computer processor, combining the methylation information from the target nucleic acids and the methylation information from the reference nucleic acids to generate background-corrected methylation information for the target nucleic acids. Here, the background-corrected methylation information is more informative for determining presence or absence of cancer within the individual. In various embodiments, performing the second analysis comprises analyzing the background-corrected methylation information to detect the presence of the circulating tumor DNA in the biological sample. By detecting presence of circulating tumor DNA in the biological sample, the individual can be identified as having cancer.

[0037] Generally, multi-tier testing methodologies described herein achieve significant improvements in comparison to conventional testing methodologies (e.g., single tier testing methodologies). For example, the multi-tier testing methodologies described herein achieve improved performance metrics (e.g., sensitivity, specificity, positive predictive value (PPV), and / or negative predictive value (NPV)) in comparison to conventional methodologies. In particular embodiments, the combination of a first tier and a second tier testing achieves improved specificity (e.g., true negative rate reported as a proportion of correctly identified negatives) in comparison to conventional methodologies.

[0038] In some scenarios, the multi-tier testing methodologies described herein rapidly and accurately screen out a large proportion of individuals in a first tier through a more efficient, lower cost tier 1 test, followed by a more rigorous tier 2 test on the remaining subpopulation of patients. Here, the multi-tier testing methodology can achieve overall performance metrics that are comparable to or not substantially less than the overall performance metrics of conventional methodologies. Altogether, by rapidly and accurately screening out a large proportion of individuals in a first tier, only a small number of individuals undergo the more rigorous tier 2 testing. This represents an improvement in comparison to conventional methodologies that attempt to apply rigorous tests across the entire population, which requires substantial resources. Thus, even in scenarios where the multi-tier testing methodologies achieve performance metrics comparable to those of conventional methodologies, the multi-tier testing methodologies deliver improved performance as a function of resource consumption. Examples of resource consumption include time resources, monetary resources, resources of consumable goods (e.g., consumable assay reagents). In various embodiments, the multi-tier testing methodologies disclosed herein achieve at least a 10% reduction in resource consumption in comparison to a corresponding single-tier test. In various embodiments, the multi-tier testing methodologies disclosed herein achieve at least a 20% reduction, at least a 30% reduction, at least a 40% reduction, at least a 50% reduction, at least a 60% reduction, at least a 70% reduction, at least a 80% reduction, or at least a 90% reduction in resource consumption in comparison to a corresponding singletier test. In various embodiments, the multi-tier testing methodologies disclosed herein achieve at least a 60% reduction in resource consumption in comparison to a corresponding single-tier test. In particular embodiments, the multiple-tiered process disclosed herein is useful for detecting rare or low incidence cancers. For example, the rare or low incidence cancers may have an incidence rate of 1 in 100, 1 in 1,000, 1 in 10,000 individuals, 1 in 100,000 individuals, 1 in 1,000,000 individuals, 1 in 10,000,000 individuals, 1 in 100,000,000 individuals or 1 in 1,000,000,000 individuals. Therefore, the disclosed multipletiered process represents a significant improvement over current methodologies that suffer from poor specificity or sensitivity which contributes to their inability to detect rare or low incidence conditions with sufficient positive predictive value.

[0039] In various embodiments, subjects that were not screened out in the first tier further undergo subsequent analysis to track tumor heterogeneity. For example, the intra-individual analysis may be performed again to analyze a second sample obtained from the same subjectat a second timepoint. Here, the second timepoint is subsequent to a first timepoint when the first sample was obtained. Performing the intra-individual analysis using the second sample generates background-corrected methylation information for the second sample. Therefore, by comparing the background-corrected methylation information of the first sample to the background-corrected methylation information of the second sample, a change in the background-corrected methylation information across the two samples is generated. Here, the change in the background-corrected methylation information across the two samples is informative for the change in tumor heterogeneity across the two timepoints from when the two samples were respectively obtained.

[0040] Figure (FIG.) 1A depicts an overall flow process 100 of the multiple-tiered process for tracking tumor heterogeneity, in accordance with an embodiment. Although FIG. 1 A shows the flow process in relation to a single subject 110, in various embodiments, the flow process can be performed for more than a single subject 110 (e.g., for thousands, millions, tens of millions, or hundreds of millions of individuals).

[0041] FIG. 1 A introduces a first sample 115A, an assay 120 A, a first tier (e.g., screen 125), an intra-individual analysis 128A, a second sample 115B, an assay 120B, and a second tier (e.g., second analysis 130) of the multiple-tiered analysis. Generally, the second tier involves a more complex molecular test and analysis in comparison to the first tier. In various embodiments, the more complex molecular test of the second tier is more expensive to perform than the simpler molecular test of the first tier. By employing a cheaper and less complex test, the first tier can identify and remove of individuals that are not at risk of cancer. The more complex molecular test and analysis of the second tier enables more accurate identification of the remaining individuals for purposes of tracking tumor heterogeneity. As shown in FIG. 1 A, the method may involve two or more intra-individual analyses performed on different samples. Here, an intra-individual analysis removes baseline biological signatures. For example, the intra-individual analysis can be performed to remove baseline biological signatures in sequencing information (hereafter referred to as “background-corrected information”) prior to the performance of the second tier. Thus, the more complex molecular test of the second tier can be applied to analyze the background- corrected information of two or more intra-individual analyses to more accurately track tumor heterogeneity in a subject.

[0042] Although FIG. 1 A shows a first tier and a second tier of a multiple-tiered analysis, in various embodiments, there may be additional tiers for further classifying individuals. Invarious embodiments, the multiple-tiered analysis includes three or more tiers, includes four or more tiers, includes five or more tiers, includes six or more tiers, includes seven or more tiers, includes eight or more tiers, includes nine or more tiers, or includes ten or more tiers.

[0043] In various embodiments, the combination of the first tier and the second tier enables the ultimate high performance (e.g., high positive predictive value) of the multiple-tier analysis. In various embodiments, the first tier and the second tier interrogate different markers from samples obtained from subjects. This can be beneficial because different markers can provide different information. In some cases, different markers can be informative for different predictions. As an example, the first tier may analyze protein markers from samples obtained from subjects whereas the second tier may analyze sequencing data derived from nucleic acids in the samples obtained from subjects.

[0044] In various embodiments, the first tier and second tier interrogate the same type of markers from samples obtained from subjects, but at different levels of detail. For example, the first tier may involve the analysis of methylation statuses for a limited, pre-selected set of genomic sites. The differential methylation of the limited, pre-selected set of genomic sites is sufficient to enable identification of subjects not at risk of cancer. Additionally, the second tier may involve the analysis of methylation statuses for a larger set of genomic sites. In one scenario, the second tier involves analysis of methylation statuses for the whole genome (e.g., through whole genome bisulfite sequencing). The differential methylation of the larger set of genomic sites enables more accurate tracking of tumor heterogeneity in the remaining subjects. As another example, the first tier may involve the analysis of shallow sequencing data. Here, shallow sequencing data is sufficient to identify and remove subjects who are not at risk or who do not have cancer. The second tier may involve analysis of sequencing data derived from deeper sequencing, which is sufficient to track tumor heterogeneity for subjects who have cancer.

[0045] FIG. 1 A introduces a subject 110. One or more samples (e.g., sample 115A and / or sample 115B) are obtained from the subject 110. In various embodiments, a sample is any of a blood sample, a stool sample, a urine sample, a mucous sample, or a saliva sample. In particular embodiments, each sample obtained from the subject 110 is a blood sample. The sample can be obtained by the individual or by a third party, e.g., a medical professional. Examples of medical professionals include physicians, emergency medical technicians, nurses, first responders, psychologists, phlebotomist, medical physics personnel, nurse practitioners, surgeons, dentists, and any other obvious medical professional as would beknown to one skilled in the art. In various embodiments, the one or more samples can be obtained from the subject 110 by a reference lab.

[0046] In various embodiments, the sample obtained from the subject is a liquid biopsy sample obtained at a first point in time. In various embodiments, the liquid biopsy sample may include various biomarkers, examples of which include proteins, metabolites, and / or nucleic acids. In particular embodiments, the liquid biopsy sample includes cell-free DNA (cfDNA) fragments. In particular embodiments, the cfDNA fragments include genomic sequences corresponding to CpG islands for which methylation states are informative of the cancer.

[0047] In various embodiments, a plurality of samples are obtained from the subject 110 at a plurality of different points in time. For example, a sample (e.g., sample 115A) can be obtained at a first timepoint and at least a second sample (e.g., sample 115B) can be obtained from the subject 110 at a second timepoint. In such embodiments, the first sample can be used for performing the assay 120A, the screen 125, and the intra-individual analysis 128A. Additionally, the second sample 115B can be used to perform an assay 120B, and a second intra-individual analysis 128B. The second analysis 130 can then be performed using the results from each of the two or more intra-individual analyses (e.g., intra-individual analysis 128A and intra-individual analysis 128B). Obtaining a plurality of liquid biopsy samples from the individual at a plurality of different points in time includes obtaining a number M of liquid biopsy samples, wherein M is one of: 2, 3, 4, ... , N-l, N, wherein N is a positive integer.

[0048] In various embodiments, sample 115A and / or sample 115B may be processed to extract target nucleic acids and reference nucleic acids. In various embodiments, samples can undergo cellular disruption methods (e.g., to obtain genomic DNA) involving chemical methods or mechanical methods. Example chemical methods include osmotic shock, enzymatic digestion, detergents, or alkali treatment. Example mechanical methods include homogenization, ultrasonication or cavitation, pressure cell, or ball mill. In various embodiments, samples can undergo removal of membrane lipids or proteins or nucleic acid purification. Example chemical methods for removing membrane lipids or proteins and methods for nucleic acid purification include guanidine thiocyanate (GuSCN)-phenol- chloroform extraction, alkaline extraction, cesium chloride gradient centrifugation with ethidium bromide, Chelex® extraction, or cetyltrimethylammonium bromide extraction. Example physical methods for removing membrane lipids or proteins and methods fornucleic acid purification include solid-phase extraction methods using any of silica matrices, glass particles, diatomaceous earth, magnetic beads, anion exchange material, or cellulose matrix. Further details of nucleic acid extraction methods are described in Ali et al, Current Nucleic Acid Extraction Methods and Their Implications to Point-of-Care Diagnostics, Biomed Res. Int. 2017; 2017:9306564, which is hereby incorporated by reference in its entirety.

[0049] Assay 120A and / or assay 120B are performed on the obtained sample 115A and 115B, respectively, to generate marker information. An example of marker information can include quantitative levels of a biomarker, such as a protein biomarker, nucleic acid biomarker, metabolite biomarker, that is present in the sample. Another examples of marker information is sequence information for a plurality of genomic sites. In various embodiments, given that the assay 120 may be performed on a large number of samples (e.g., millions of samples) obtained from a large patient population, the assay 120 be a simplified molecular test that generates marker information that can rapidly distinguish between individuals at risk and individuals not at risk for cancer. For example, the marker information can include quantitative levels of a biomarker, such as a protein biomarker, nucleic acid biomarker, metabolite biomarker, that can rapidly guide the identification and removal of individuals not at risk for the cancer. As another example, the marker information can be sequence information for a limited number of genomic sites that are sufficient for identifying individuals who are not at risk for the cancer (e.g., true negatives). In particular embodiments, the sequence information for a plurality of genomic sites includes methylation information, such as methylation statuses for the plurality of genomic sites. In various embodiments, the plurality of genomic sites include a plurality of CpG islands (CGIs) whose differential methylation status may be indicative of risk for the cancer.

[0050] In particular embodiments, assay 120A and / or assay 120B are performed to generate sequence information for target nucleic acids and to generate sequence information for reference nucleic acids. Thus, sequence information of target and reference nucleic acids can be used to perform the intra-individual analysis 128A and / or intra-individual analysis 128B. In particular embodiments, sequence information includes statuses for a plurality of genomic sites, such as epigenetic statuses for a plurality of CpG sites. In various embodiments, epigenetic statuses refer to methylation statuses. In particular embodiments, sequence information of the target nucleic acids and sequence information of the reference nucleic includes statuses for two or more, three or more, four or more, five or more, six or more,seven or more, eight or more, nine or more, or ten or more common genomic sites. In particular embodiments, sequence information of the target nucleic acids and sequence information of the reference nucleic each includes statuses for 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 750 or more, 1000 or more, 2000 or more, 3000 or more, 4000 or more, 5000 or more, 6000 or more, 7000 or more, 8000 or more, 9000 or more, 10000 or more, 11000 or more, 12000 or more, 13000 or more, 14000 or more, 15000 or more, 16000 or more, 17000 or more, 18000 or more, 19000 or more, or 20000 or more genomic sites. In particular embodiments, sequence information of the target nucleic acids and sequence information of the reference nucleic each includes statuses for 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 750 or more, 1000 or more, 2000 or more, 3000 or more, 4000 or more, 5000 or more, 6000 or more, 7000 or more, 8000 or more, 9000 or more, 10000 or more, 11000 or more, 12000 or more, 13000 or more, 14000 or more, 15000 or more, 16000 or more, 17000 or more, 18000 or more, 19000 or more, or 20000 or more of the same genomic sites or overlapping genomic sites. In various embodiments, the plurality of genomic sites include a plurality of CpG islands (CGIs) whose differential methylation status may be indicative of a cancer.

[0051] A screen 125 is performed to analyze the marker information generated by the assay 120 A. For example, the screen 125 can involve an in silico analysis of the marker information. In various embodiments, the marker information includes quantitative values of biomarkers. Therefore, the screen 125 can identify and remove individuals whose quantitative values of biomarkers indicate that the individuals are not at risk of the cancer. In various embodiments, the marker information is sequence information for a plurality of genomic sites. Therefore, the screen 125 involves deploying a trained machine learning model that analyzes the sequence information for the plurality of genomic sites and predicts whether an individual is at risk for a cancer. If the screen 125 identifies the individual as not at risk for cancer (as indicated in FIG. 1 A as “If negative”), then the subject 110 can be reported as not at risk for the cancer. The process can terminate for this subject and therefore, additional resources need not be further devoted to this subject.

[0052] Alternatively, if the screen identifies the subject as at risk for cancer (as indicated in FIG. 1 A as “If not negative” following screen 125), then the subject 110 undergoes at least another tier of testing. As shown in FIG. 1 A, an intra-individual analysis 128A and a secondanalysis 130 can be performed for subjects identified as at risk for cancer. In particular embodiments, a second sample 115B, assay 120B and second intra-individual analysis 128B are performed for the subject after having determined that the subject is not negative based on the results of the screen 125.

[0053] In various embodiments, as shown in FIG. 1 A, the subject 110 receives an intervention 112. In various embodiments, the subject 110 receives the intervention 112 after the screen determines that the subject 110 is not negative for cancer. Thus, the subject 110 may have been selected and provided the intervention to treat for the cancer and / or to reduce the risk for cancer . An example of an intervention 112 is a tumor therapeutic (e.g., a cancer therapeutic, a chemotherapy, and / or a gene therapy).

[0054] Referring to the intra-individual analysis 128A and intra-individual analysis 128B, the analysis is conducted for a specific subject, such as a subject identified via the screen 125 as at risk for the cancer. Therefore, for a particular subject, the intra-individual analysis is performed to remove baseline biological signatures that are present in the subject. Here, the baseline biological signatures are present irrespective of whether the subject has or does not have cancer. These baseline biological signatures would be confounding signals if analyzed to generate predictions for the patient. Thus, performing the intra-individual analysis 128 for individual samples (e.g., sample 115A or sample 115B) eliminates these confounding baseline biological signatures while keeping signatures that are more informative for determining presence or absence of cancer. For example, in processing nucleic acid sequencing information to generate a signal that may be detected, the resulting signal may comprise a mixture of baseline biological signatures (e.g., germline methylation in a patient) that represent a form of background noise and signatures informative of a cancer(e.g., cancer). Such background noise can obscure a signal informative of a cancer.Advantageously, in certain embodiments, methods described herein contemplate subtracting such background noise from a patient’s nucleic acid sequencing information, thereby improving the signal-to-noise ratio of the signal informative of a cancer.

[0055] In contrast to an inter-individual analysis, where, for example, to determine a presence or absence of cancer within a patient, an average of baseline signatures from a group of normal subjects are removed from the nucleic acid sequencing information of the patient, it has been discovered that performing an intra-individual analysis can significantly improve the sensitivity or specificity of detecting a signal informative for determining presence or absence of cancer.

[0056] Generally, the intra-individual analysis 128A or intra-individual analysis 128B involves generating information from at least target nucleic acids and reference nucleic acids from a corresponding sample (e.g., sample 115A and sample 115B) obtained from the patient. In various embodiments, the intra-individual analysis 128A and intra-individual analysis 128B is performed on sequence information. Such sequence information may be generated by assay 120A and assay 120B, as shown in FIG. 1 A.

[0057] In various embodiments, the intra-individual analysis 128A and intra-individual analysis 128B involve combining information from target nucleic acids and the reference nucleic acids to generate a signal informative for determining presence or absence of cancer within the patient. By combining the information from the target nucleic acids and the reference nucleic acids, the generated signal can be more informative of presence or absence of a cancer in comparison to a signal derived from the target nucleic acids alone. For example, the information from the reference nucleic acids can represent baseline biology of the patient. By combining the information from the target nucleic acids and the reference nucleic acids, the baseline biology of the patient, which may not be informative for the presence or absence of a cancer, is removed from the generated signal. Thus, information of the target nucleic acids that are not attributable to the patient’s baseline biology remains and is included in the generated signal for determining presence or absence of cancer in the patient.

[0058] Referring next to the second analysis 130, the second analysis 130 is implemented to determine a change in tumor heterogeneity 135 in the subject 110. In various embodiments, the second analysis 130 determines a change in signal between a first set of background- corrected methylation information generated from the first intra-individual analysis 128A and a second set of background-corrected methylation information generated from the second intra-individual analysis 128B. For example, as shown in FIG. 1A, the output of each of the intra-individual analysis 128A and intra-individual analysis 128B can be combined to determine the change in signal. The change in signal can be provided for the second analysis 130 and can be indicative of whether the tumor heterogeneity in the subject is increasing, decreasing, or remaining stable.

[0059] Referring next to FIG. IB, it depicts an overall flow process of the multiple-tiered process for tracking tumor heterogeneity, in accordance with a second embodiment. Here, FIG. IB differs from FIG. 1 A in that the second analysis 130 is individually performed to analyze the results of each respective intra-individual analysis e.g., intra-individual analysis128A and intra-individual analysis 128B. Therefore, as shown in FIG. IB, the output of the second analysis 130A can be combined with the output of second analysis 130B to determine a change in tumor heterogeneity 135 for the subject 110.

[0060] Altogether, the multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128, and second analysis 130) enables the rapid identification of a large proportion of individuals (e.g., greater than 80% of the patient population) representing true negatives, and further enables the accurate identification and diagnosis of a subset of the population representing true positives. The overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multipletiered analysis involving each of the screen 125, intra-individual analysis 128A, intra- individual analysis 128B, and second analysis 130) achieves one or more performance metrics, such as metrics of sensitivity, specificity, positive predictive value (PPV), and / or negative predictive value (NPV). Sensitivity is the true positive rate, reported as a proportion of correctly identified positives. Specificity is the true negative rate reported as a proportion of correctly identified negatives. Positive predictive value refers to the number of true positives divided by the sum of true positives and false positives. Negative predictive value refers to the true negative rate divided by the sum of true negatives and false negatives.

[0061] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% sensitivity in detecting presence of a cancer. In various embodiments, the overall multiple-tiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 70% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 71% sensitivity. In particular embodiments, the overall multiple-tiered analysisachieves at least 72% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 73% sensitivity. In particular embodiments, the overall multipletiered analysis achieves at least 74% sensitivity. In particular embodiments, the overall multiple-tiered analysis achieves at least 75% sensitivity.

[0062] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% specificity in excluding individuals without the cancer. In various embodiments, the overall multiple-tiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% specificity. In particular embodiments, the overall multiple-tiered analysis achieves at least 99% specificity. In particular embodiments, the overall multiple-tiered analysis achieves at least 99.5% specificity. In particular embodiments, the overall multipletiered analysis achieves at least 99.9% specificity.

[0063] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves a particular sensitivity and a particular specificity. The combination of the sensitivity and specificity limits both the number of false positives and the number of false negatives. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 75% to 89% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 80% to 88% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 83% to 87% sensitivity and between 90% to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 84% to 86% sensitivity and between 90%to 100% specificity. In various embodiments, the overall multiple-tiered analysis achieves about 85% sensitivity and between 90% to 100% specificity.

[0064] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves between 70% to 90% sensitivity and between 91% to 99% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 92% to 98% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 93% to 97% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and between 97% to 96% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 70% to 90% sensitivity and about 95% specificity.

[0065] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves between 75% to 89% sensitivity and between 91% to 99% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 80% to 88% sensitivity and between 92% to 98% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 83% to 87% sensitivity and between 93% to 97% specificity. In various embodiments, the overall multiple-tiered analysis achieves between 84% to 86% sensitivity and between 94% to 96% specificity. In various embodiments, the overall multiple-tiered analysis achieves about 85% sensitivity and about 95% specificity.

[0066] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% positive predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 20% positive predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, or at least 40% positivepredictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 40% positive predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, or at least 60% positive predictive value. In various embodiments, the overall multipletiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 80% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 81% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 82% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 83% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 84% positive predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 85% positive predictive value.

[0067] In various embodiments, the overall multiple-tiered analysis (e.g., multiple-tiered analysis involving the screen 125 and second analysis 130 or multiple-tiered analysis involving each of the screen 125, intra-individual analysis 128A, intra-individual analysis 128B, and second analysis 130) achieves at least 60% negative predictive value. In various embodiments, the overall multiple-tiered analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% negative predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 98% negative predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 99% negative predictive value. In particular embodiments, the overall multiple-tiered analysis achieves at least 99.4% negative predictive value.System Environment Overview

[0068] FIG. 1C depicts an overall system environment 150 including a tumor heterogeneity system 170, in accordance with an embodiment. The overall system environment 150 includes a tumor heterogeneity system 170 for at least performing one or more steps shown in FIG. 1A, and one or more third party entities 155 A and 155B in communication with one another through a network 160. FIG. IB depicts one embodiment of the overall system environment 150 in which two third party entities 155 A and 155B are involved. In other embodiments, additional or fewer third party entities 155 in communication with the tumor heterogeneity system 170 can be included. The third party entities 155 may communicate with the tumor heterogeneity system 170 to enable the tumor heterogeneity system 170 to perform a screen, one or more intra-individual analyses, and / or second analysis.Third Party Entity

[0069] A third party entity 155 represents a partner entity of the tumor heterogeneity system 170 that can operate upstream, downstream, or both upstream and downstream of the operations of the tumor heterogeneity system 170. As one example, the third party entity 155 operates upstream of the tumor heterogeneity system 170 and provides samples obtained from patients to the tumor heterogeneity system 170. Thus, the tumor heterogeneity system 170 can perform assays, a screen, one or more intra-individual analyses, and / or a second analysis to track tumor heterogeneity of subjects. As another example, the third party entity 155 may process samples obtained from subjects by performing one or more assays on the samples to generate data. Thus, the third party entity 155 can provide the data derived from the assays to the tumor heterogeneity system 170 such that the tumor heterogeneity system 170 can perform a screen, one or more intra-individual analyses, and / or second analysis.

[0070] As another example, the third party entity 155 operates downstream of the tumor heterogeneity system 170. In this scenario, the tumor heterogeneity system 170 may perform a screen and determine whether a subject is at risk for cancer. The tumor heterogeneity system 170 can provide an indication to the third party entity 155 that identifies the subject atrisk for the cancer. The third party entity 155 may notify the subject regarding a follow-up appointment such that an additional sample (e.g., sample 115B shown in FIG. 1 A) can be obtained from the subject at the follow-up appointment for subsequent analysis.Network

[0071] This disclosure contemplates any suitable network 160 that enables connection between the tumor heterogeneity system 170 and third party entities 155. The network 160 may comprise any combination of local area and / or wide area networks, using both wired and / or wireless communication systems. In one embodiment, the network 160 uses standard communications technologies and / or protocols. For example, the network 160 includes communication links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, code division multiple access (CDMA), digital subscriber line (DSL), etc. Examples of networking protocols used for communicating via the network 160 include multiprotocol label switching (MPLS), transmission control protocol / Internet protocol (TCP / IP), hypertext transport protocol (HTTP), simple mail transfer protocol (SMTP), and file transfer protocol (FTP). Data exchanged over the network 160 may be represented using any suitable format, such as hypertext markup language (HTML) or extensible markup language (XML). In some embodiments, all or some of the communication links of the network 160 may be encrypted using any suitable technique or techniques.Tumor Heterogeneity System

[0072] FIG. 2A depicts a block diagram of the tumor heterogeneity system 170, in accordance with an embodiment. The block diagram of the tumor heterogeneity system 170 is introduced to show an embodiment in which the tumor heterogeneity system 170 includes one or more assay apparatuses 205 communicatively coupled to a computational system 202. The computational system 202 can further include computational modules, such as a screen module 210, intra-individual analysis module 215, second analysis module 220, and a tumor tracking module 230. The computational system 202 can further include data stores such as a machine learning model store 240 for storing one or more trained machine learning models. FIG. 2A depicts an embodiment in which the tumor heterogeneity system 170 performs one or more assays (e.g., assay 120A or 120B described in FIG. 1 A), performs the screen (e.g., screen 125 described in FIG. 1 A), performs the one or more intra-individual analyses (e.g.,intra-individual analysis 128A and / or intra-individual analysis 128B described in FIG. 1A), and performs the second analysis (e.g., second analysis 130 described in FIG. 1 A).

[0073] In various embodiments, the tumor heterogeneity system 170 may be differently configured than shown in FIG. 2A. For example, although the tumor heterogeneity system 170 shown in FIG. 2A includes three different assay apparatuses 205, in various embodiments, the tumor heterogeneity system 170 includes fewer or additional assay apparatuses. In various embodiments, the tumor heterogeneity system 170 does not include an assay apparatus. In such embodiments, the tumor heterogeneity system 170 includes only the computational system 202. In these embodiments in which the tumor heterogeneity system 170 does not include an assay apparatus, the tumor heterogeneity system 170 may perform the screen (e.g., screen 125 described in FIG. 1 A), one or more intra-individual analyses (e.g., intra-individual analysis 128A and / or intra-individual analysis 128B described in FIG. 1 A), and the second analysis (e.g., second analysis 130 described in FIG. 1 A). However, the tumor heterogeneity system 170 does not perform an assay. The assay apparatus 205 may be operated and used by a different entity, such as a third party entity (e.g., third party entity 155 described in FIG. 1C). Thus, the third party entity can perform assays using one or more assay apparatus 205 and then transmits the data generated from the assays to the tumor heterogeneity system 170 for performing the screen and / or second analysis.Assays

[0074] Methods disclosed herein involve performing an assay to generate marker information. Assays described in this section can refer to either assay 120 A, assay 120B, or both assay 120A and assay 120B shown in FIGs. 1 A and IB. Referring to FIG. 2A, performing an assay can involve employing one or more assay apparatuses 205 to perform the assay. In various embodiments, marker information refers to quantitative values of biomarkers, such as protein biomarkers, nucleic acid biomarkers, or metabolite biomarkers. Thus, the quantitative values of biomarkers in a sample can be used to determine whether the individual is at risk for a cancer. In various embodiments, to determine quantitative values of protein biomarkers, performing an assay can include performing one or more of an immunoassay, a protein-binding assay, an antibody -based assay, an antigen-binding proteinbased assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), or a Western blot. To determine quantitative values of nucleic acid biomarkers, performing an assay can include performing one or more of quantitative PCR (qPCR) or digital PCR(dPCR). To determine quantitative values of metabolites, performing an assay can include performing NMR, mass spectrometry, LC-MS, or UPLC-MS / MS.

[0075] In various embodiments, marker information refers to sequence information for a plurality of genomic sites. The sequence information can then be analyzed to generate a prediction for an individual (e.g., whether an individual is negative for cancer or whether the individual is not negative for cancer). In particular embodiments, performing the assay results in generation of methylation sequence information. Methylation sequence information includes methylation statuses for a plurality of genomic sites. In various embodiments, the plurality of genomic sites are previously identified and selected. For example, the plurality of genomic sites may be one or more CpG sites whose differential methylation are informative for determining whether an individual is at risk for a cancer. A CpG site is portion of a genome that has cytosine and guanine separated by only one phosphate group and is often denoted as “5' — C — phosphate — G — 3'”, or “CpG” for short. Regions with a high frequency of CpG sites are commonly referred to as “CG islands” or “CGIs”. It has been found that certain CGIs and certain features of certain CGIs in tumor cells tend to be different from the same CGIs or features of the CGIs in healthy cells. Herein, such CGIs and features of the genome are referred to herein as “cancer informative CGIs.”

[0076] Reference is made to FIG. 3 A, which depicts example methylation information useful for determining whether an individual is at risk for a cancer, in accordance with an embodiment. Specifically, FIG. 3A shows that across various types of cancers (e.g., bladder, cervical, colorectal, endometrial, gastric, lung, ovarian, and prostate cancers), sub-regions within a particular CGI can exhibit differential methylation in comparison to normal plasma. Thus, FIG. 3 A depicts an example cancer informative CGI such that performing the assay results in the generation of methylation sequence information corresponding to the cancer informative CGI.

[0077] In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites includes the steps of processing nucleic acids of a sample, enriching the processed nucleic acids for pre-selected genomic sequences (e.g., pre-selected informative CGIs), amplifying the genomic sequences to generate amplicons, and quantifying the amplicons including the genomic sequences (e.g., via sequencing or via quantitative methods such as an ELISA, quantitative PCR, or DNA or RNA-based assay). In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites involves a subset of the previously mentioned steps. For example, enrichingthe processed nucleic acids can be omitted. Therefore, performing an assay may include processing nucleic acids of a sample, amplifying the pre-selected genomic sequences, and quantifying the amplicons including the genomic sequences.

[0078] Referring again to FIG. 1 A or IB, in various embodiments, assay 120A and assay 120B may both involve performing steps of processing nucleic acids of a sample, enriching the processed nucleic acids for pre-selected genomic sequences (e.g., pre-selected informative CGIs), amplifying the genomic sequences to generate amplicons, and quantifying the amplicons including the genomic sequences. In various embodiments, assay 120 A and assay 120B involve quantifying the amplicons by performing an ELISA assay, by performing quantitative PCR, or by performing next generation sequencing.

[0079] A methylated nucleic acid is a nucleic acid having a modification in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5- methylcytosine. Methylation can occur at dinucleotides of cytosine and guanine referred to herein as “CpG sites”, which can be a target for enrichment. Methylation of cytosine can occur in cytosines in other sequence contexts, for example, 5'-CHG-3' and 5'-CHH-3', where H is adenine, cytosine or thymine. Cytosine methylation can also be in the form of 5- hydroxymethylcytosine. Methylation of DNA can include methylation of non-cytosine nucleotides, such as A6-methyladenine (6mA). Anomalous cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which may be indicative of cancer status. As is well known in the art, DNA methylation anomalies (compared to healthy controls) can cause different effects, which may contribute to cancer.

[0080] In certain embodiments, the nucleic acid comprises a CpG site (ie., cytosine and guanine separated by only one phosphate group). In certain embodiments, the nucleic acid comprises a CpG island (also referred to as a “CG islands” or “CGI”) or a portion thereof, which is the target for enrichment. Because certain CGIs and certain features of certain CGIs in tumor cells tend to be different from the same CGIs or features of the CGIs in healthy cells, detection of such CGIs can be informative of a cancer. In certain embodiments, the CGI is a “cancer informative CGIs”, which is defined and described in more detail below. In certain embodiments, the CpG is an “informative CpG”, e.g., a “cancer informative CGI”. Such CGIs may have methylation patterns in tumor cells that are different from the methylation patterns in healthy cells. Accordingly, detection of a cancer informative CGI can be informative regarding a subject’s risk of developing cancer or can be indicative that the subject has cancer. Exemplary cancer informative CGIs, which can be target sequences asdescribed herein, are identified in, e.g., Table 1 of U.S. Patent Publication 2020 / 0109456A1, Tables 2 and 3 of WO2022 / 133315, and Tables 1-4 provided herein.

[0081] In certain aspects, the nucleic acids have been treated to convert one or more unmethylated nucleotides (e.g., cytosines) to another nucleotide (a “converted nucleotide”, as used herein, such as a uracil), for example, prior to amplification. Example conversions include bisulfite conversion, enzymatic conversion, or nitrite conversion, further details of which are described herein. In certain embodiments, one or more unmethylated cytosines are converted to a nucleotide that pairs with adenine (e.g., the unmethylated cytosine may be converted to uracil). In certain embodiments, one or more unmethylated adenines are converted to a base that pairs with cytosine (e.g., the unmethylated adenine may be converted to inosine (I)). In certain embodiments, one or more methylated cytosines (e.g., a 5- methylcytosine (5mC)) is converted to a thymine, which pairs with adenine. In certain embodiments, methylated cytosines are protected from conversion (e.g., deamination) during the conversion step.

[0082] In various embodiments, nucleic acids undergo a bisulfite conversion. Bisulfite conversion is performed on DNA by denaturation using high heat, preferential deamination (at an acidic pH) of unmethylated cytosines, which are then converted to uracil by desulfonation (at an alkaline pH). Methylated cytosines remain unchanged on the singlestranded DNA (ssDNA) product.

[0083] In some embodiments the methods include treatment of the sample with bisulfite (e.g., sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium metabisulfite, potassium metabisulfite, ammonium metabisulfite, magnesium metabisulfite and the like). Unmethylated cytosine is converted to uracil through a three-step process during sodium bisulfite modification. As shown in FIG. 2B, the steps are sulphonation to convert cytosine to cytosine sulphonate, deamination to convert cytosine sulphonate to uracil sulphonate and alkali desulphonation to convert uracil sulphonate to uracil. Conversion on methylated cytosine is much slower and is not observed at significant levels in a 4-16 hour reaction. (See Clark et al., Nucleic Acids Res., 22(15):2990-7 (1994).) If the cytosine is methylated it will remain a methylated cytosine. If the cytosine is unmethylated it will be converted to uracil. When the modified strand is copied, for example, through extension of a locus specific primer, a random or degenerate primer or a primer to an adaptor, a G will be incorporated in the interrogation position (opposite the C being interrogated) if the C was methylated and an A will be incorporated in the interrogation position if the C wasunmethylated and converted to U. When the double stranded extension product is amplified those Cs that were converted to Us and resulted in incorporation of A in the extended primer will be replaced by Ts during amplification. Those Cs that were not converted (z.e., the methylated Cs) and resulted in the incorporation of G will be replaced by unmethylated Cs during amplification.

[0084] In various embodiments, nucleic acids undergo an enzymatic conversion. In certain embodiments, the enzymatic treatment with a cytidine deaminase enzyme is used to convert cytosine to uracil. Enzymatic conversion can include an oxidation step, in which Tet methylcytosine dioxygenase 2 (TET2) catalyzes the oxidation of 5mC to 5hmC to protect methylated cytosines from conversion by subsequent exposure to a cytidine deaminase.Other protection steps known in the art can be used in addition to or in place of oxidation by TET2. After the oxidation step, the nucleic acid is treated with the cytidine deaminase to convert one or more unmethylated cytosines to uracils. As with bisulfite conversion, when the modified strand is copied, a G will be incorporated in the interrogation position (opposite the C being interrogated) if the C was methylated and an A will be incorporated in the interrogation position if the C was unmethylated. When the double stranded extension product is amplified those Cs that were converted to Us and resulted in incorporation of A in the extended primer will be replaced by Ts during amplification. Those Cs that were not modified and resulted in the incorporation of G will remain as C.

[0085] In certain embodiments the cytidine deaminase may be APOBEC. In certain embodiments the cytidine deaminase includes activation induced cytidine deaminase (AID) and apolipoprotein B mRNA editing enzymes, catalytic polypeptide-like (APOBEC). In certain embodiments, the APOBEC enzyme is selected from the human APOBEC family consisting of APOBEC-1 (Apol), APOBEC-2 (Apo2), AID, APOBEC-3 A, -3B, -3C, -3DE, -3F, -3G, -3H and APOBEC-4 (Apo4). In certain embodiments, the APOBEC enzyme is APOBEC-seq.

[0086] In certain embodiments, nitrite treatment is used to deaminate adenine and cytosine. As shown in FIG. 2C, deamination of an A results in conversion to an inosine (I), which is read by a polymerase as a G, whereas deamination of a methylated A (A6-methyladenine (6mA)) results in a nitrosylated 6mA (6mA-N0), which causes the base to be read by a polymerase as an A. Deamination of a C results in conversion to a uracil, which is read by a polymerase as a T, whereas deamination of a ^-methylcytosine (4mC) to 4mC-N0 or a 5- methylcytosine (5mC) to a T causes the base to be read by a polymerase as a C or a T,respectively. For 5mC bases, the C to T ratio at the 5mC position is about 40% higher than other cytosine positions, allowing 5mC to be differentiated from C. (See, Li et al. (2022) Genome Biology 23 : 122.)

[0087] In various embodiments, performing the assay includes enriching for specific genomic sequences, such as genomic sequences of pre-selected CGIs. In various embodiments, enrichment of pre-selected CGIs can be accomplished via hybrid capture. Examples of such hybrid capture probe sets include the KAPA HyperPrep Kit and SeqCAP Epi Enrichment System from Roche Diagnostics (Pleasanton, CA). For example, hybrid capture probe sets can be designed to target (e.g., hybridize with) selected genomic sequences, thereby capturing and enriching the selected genomic sequences.

[0088] In various embodiments, performing the assay includes a step of nucleic acid amplification. During amplification, the converted nucleotide pairs with its complementary nucleotide, and in the next round of amplification, the complementary nucleotide pairs with a replacement nucleotide. For example, following the conversion of an unmethylated cytosine to a uracil, the nucleic acid may be amplified such that an adenine pairs with the uracil in the first round of replication, and in the second round of replication, the adenine pairs with a thymine. Accordingly, the thymine replaces the uracil in the original nucleic acid sequence, and is referred to herein as a “replacement nucleotide”.

[0089] Examples of such assays include, but are not limited to performing PCR assays, Realtime PCR assays, Quantitative real-time PCR (qPCR) assays, digital PCR (dPCR), Allelespecific PCR assays, Reverse-transcription PCR assays and reporter assays. For example, given the processed nucleic acids (e.g., bisulfite converted nucleic acids) that are enriched for pre-selected genomic sequences, a PCR assay is performed to amplify the pre-selected genomic sequences to generate amplicons. Here, PCR primers are added to initiate the amplification. In various embodiments, the PCR primers are whole genome primers that enable whole genome amplification. In various embodiments, the PCR primers are genespecific primers that result in amplification of sequences of specific genes. In various embodiments, the PCR primers are allele-specific primers. For example, allele specific primers can target a genomic sequence corresponding to a pre-selected CGI, such that performing nucleic acid amplification results in amplification of the genomic sequence of the pre-selected CGI.

[0090] In various embodiments, performing the assay includes quantifying the nucleic acids including the pre-selected genomic sequences (e.g., informative CGIs). In someembodiments, quantifying the nucleic acids to generate sequence information comprises performing an enzyme-linked immunosorbent assay (ELISA). In some embodiments, quantifying the nucleic acids to generate sequence information comprises performing quantitative PCR (qPCR) or digital PCR (dPCR). Therefore, the number of methylated, unmethylated, or partially methylated pre-selected genomic sequences can be quantified.

[0091] In various embodiments, quantifying the nucleic acids comprises sequencing the nucleic acids including the pre-selected genomic sequences. Thus, the sequenced reads can be aligned to a reference library and methylation sequence information including methylation statuses of the informative CGIs can be determined. Therefore, the number of methylated, unmethylated, or partially methylated pre-selected genomic sequences can be quantified via the sequenced reads.

[0092] FIG. 3B shows an example flow process for determining whether an individual is at risk for a cancer, in accordance with an embodiment. Here, specific genomic regions of an indexed library of nucleic acids (e.g., DNA) are targeted. For example, locus 1 can refer to a reference genomic location. Here, a reference genomic location serves as a control. For example, the reference genomic location is not differentially methylated in healthy individuals in comparison to individuals with the cancer. Locus 2 can refer to a pre-selected genomic location, such as a pre-selected informative CGI.

[0093] Performing the assay further includes performing nucleic acid amplification (e.g., PCR) to generate marker information. In various embodiments, nucleic acid amplification includes either qPCR or dPCR. This quantifies the number of methylated, unmethylated, or partially methylated sequences at locus 1 (reference) and at locus 2. In various embodiments, performing the assay includes performing an ELISA to quantify the number of methylated, unmethylated, or partially methylated sequences at locus 1 (reference) and at locus 2.Assays for Generating Sequencing Information for Performing Intra-Individual Analysis

[0094] In particular embodiments, assays disclosed herein (e.g., assay 120A or 120B shown in FIGs. 1 A-1B) are useful for generating sequencing information for performing an intraindividual analysis (e.g., one or both of intra-individual analysis 128A and intra-individual analysis 128B shown in FIGs. 1 A-1B). For example, an assay is performed to generate sequence information for target nucleic acids and / or reference nucleic acids.

[0095] In various embodiments, sequence information of target nucleic acids and / or sequence information of reference nucleic acids refer to statuses for a plurality of genomic sites. Sequence information of target nucleic acids refers to epigenetic statuses (e.g., methylation statuses) across a plurality of genomic sites in the target nucleic acids. Sequence information of reference nucleic acids refers to epigenetic statuses (e.g., methylation statuses) across a plurality of genomic sites in the reference nucleic acids. In various embodiments, the plurality of genomic sites are previously identified and selected. For example, the plurality of genomic sites may be one or more CpG sites whose differential methylation are informative for determining whether an individual has a cancer. A CpG site is portion of a genome that has cytosine and guanine separated by only one phosphate group and is often denoted as “5' — C — phosphate — G — 3'”, or “CpG” for short. Regions with a high frequency of CpG sites are commonly referred to as “CG islands” or “CGIs”. It has been found that certain CGIs and certain features of certain CGIs in tumor cells tend to be different from the same CGIs or features of the CGIs in healthy cells. Herein, such CGIs and features of the genome are referred to herein as “cancer informative CGIs.” Cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. Example CGIs include, but are not limited to, the CGIs shown in the accompanying tables (referred to herein as Tables 1-4) which lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0096] In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites includes the steps of processing nucleic acids of a sample, enriching the processed nucleic acids for pre-selected genomic sequences (e.g., pre-selected informative CGIs), amplifying the genomic sequences to generate amplicons, and quantifying the amplicons including the genomic sequences (e.g., via sequencing such as next generation sequencing or via quantitative methods such as an ELISA, quantitative PCR, allele-specificPCR, or DNA or RNA-based assay). In various embodiments, performing an assay to generate sequence information for a plurality of genomic sites involves a subset of the previously mentioned steps. For example, enriching the processed nucleic acids can be omitted. Therefore, performing an assay may include processing nucleic acids of a sample, amplifying the pre-selected genomic sequences, and quantifying the amplicons including the genomic sequences.

[0097] In various embodiments, performing an assay (e.g., assay 120A or assay 120B) involves processing target nucleic acids and / or reference nucleic acids. In various embodiments, processing target nucleic acids and / or reference nucleic acids to capture methylation modifications includes performing a nucleic acid conversion (e.g., any of bisulfite conversion, enzymatic conversion, or nitrite conversion). In various embodiments, processing target nucleic acids and / or reference nucleic acids to capture methylation modifications includes performing any of nucleic acid amplification, polymerase chain reaction (PCR), methylation specific PCR, bisulfite pyrosequencing, single-strand conformation polymorphism (SSCP) analysis, methylation-sensitive single-strand conformation analysis restriction analysis, high resolution melting analysis, methylationsensitive single-nucleotide primer extension, restriction analysis, microarray technology, next generation methylation sequencing, nanopore sequencing, and combinations thereof.

[0098] In various embodiments, performing the assay includes enriching for specific sequences in the target nucleic acids and / or reference nucleic acids. In various embodiments, the specific sequences refer to sequences of pre-selected CGIs. In various embodiments, enrichment of pre-selected CGIs can be accomplished via hybrid capture. Examples of such hybrid capture probe sets include the KAPA HyperPrep Kit and SeqCAP Epi Enrichment System from Roche Diagnostics (Pleasanton, CA). For example, hybrid capture probe sets can be designed to hybridize with particular sequences of the target nucleic acids and / or reference nucleic acids, thereby capturing and enriching the particular sequences.

[0099] In various embodiments, performing the assay includes performing nucleic acid amplification to amplify the particular sequences of the target nucleic acids and / or reference nucleic acids. Examples of such assays include, but are not limited to performing PCR assays, Real-time PCR assays, Quantitative real-time PCR (qPCR) assays, digital PCR (dPCR), Allele-specific PCR assays, Reverse-transcription PCR assays and reporter assays. For example, given the processed nucleic acids (e.g., bisulfite converted nucleic acids) that are enriched for pre-selected sequences, a PCR assay is performed to amplify the pre-selectedsequences to generate amplicons. Here, PCR primers are added to initiate the amplification. In various embodiments, the PCR primers are whole genome primers that enable whole genome amplification. In various embodiments, the PCR primers are gene-specific primers that result in amplification of sequences of specific genes. In various embodiments, the PCR primers are allele-specific primers. For example, allele specific primers can target a genomic sequence corresponding to a pre-selected CGI, such that performing nucleic acid amplification results in amplification of the sequence of the pre-selected CGI.

[0100] In various embodiments, performing the assay includes quantifying the nucleic acids including the pre-selected sequences (e.g., informative CGIs). In some embodiments, quantifying the nucleic acids to generate sequence information comprises performing any of real-time PCR assay, quantitative real-time PCR (qPCR) assay, digital PCR (dPCR) assay, allele-specific PCR assay, or reverse-transcription PCR assay. Therefore, the number of methylated, hypermethylated, unmethylated, or partially methylated pre-selected sequences are quantified.

[0101] In various embodiments, quantifying the nucleic acids comprises sequencing the nucleic acids including the pre-selected sequences. Thus, the sequenced reads are aligned to a reference library and sequence information including methylation statuses of the informative CGIs of amplicons derived from the target nucleic acids and / or reference nucleic acids can be determined. Therefore, the number of methylated, hypermethylated, unmethylated, or partially methylated pre-selected sequences of the target nucleic acids and the reference nucleic acids can be quantified via the sequenced reads.Assays for Generating Sequencing Information for Phased Sequencing

[0102] In various embodiments, performing the assay comprises sequencing the target nucleic acids and / or reference nucleic acids. In various embodiments, sequencing comprises performing next generation sequencing methods to generate sequence reads from the target nucleic acids and / or reference nucleic acids. As described herein, sequence reads from reference nucleic acids may be long sequence reads (e.g., greater than 500 bases in length). Generally, long sequence reads include an average read length that is longer than sequence reads obtained through standard sequencing methods. In various embodiments, the long sequence reads of reference nucleic acids refer to sequence reads of at least 500 bases, at least 1 kilobase, at least 2 kilobases (kb), at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, at least 12 kb, at least 15 kb, at least 20kb, at least 25 kb, at least 30 kb, at least 40 kb, at least 50 kb, at least 60 kb, at least 70 kb, at least 80 kb, at least 90 kb, at least 100 kb, at least 200 kb, at least 300 kb, at least 400 kb, at least 500 kb, at least 600 kb, at least 700 kb, at least 800 kb, at least 900 kb, at least 1000 kb, at least 1500 kb, or at least 2000 kb. In particular embodiments, the long sequence reads of reference nucleic acids refer to sequence reads of between 5 kb and 100 kb, between 10 kb and 80 kb, between 20 kb and 70 kb, between 30 kb and 60 kb, or between 40 kb and 50 kb. In particular embodiments, long sequence reads of reference nucleic acids refer to sequence reads of greater than about 8 kb, greater than about 9 kb or greater than about 10 kb. In particular embodiments, long sequence reads of reference nucleic acids refer to sequence reads between about 10 kb and about 100 kb, or between about 10 kb and about 2 MB. In various embodiments, generating long sequence reads of reference nucleic acids involves performing nanopore sequencing. Methods for long-read sequencing are known in the art and such methods can be performed using, for example, an Oxford Nanopore instrument (e.g., PromethlON™) or Pacific Biosciences Single-Molecule Real-Time (SMRT) sequencing technology.

[0103] In various embodiments, performing the assay includes generating phased sequencing information for target nucleic acids and / or reference nucleic acids. As used herein, “phased sequencing information,” also referred to herein as “haplotype sequencing information,” refers to sequencing information derived specifically from a particular source. For example, phased sequencing information or haplotype sequencing information can refer to sequencing information derived from either the maternal or paternal chromosome. Generally, phased sequencing information of target nucleic acids may be useful for determining presence or absence of a cancer because signals originating from the same source (e.g., maternal or paternal chromosome) may provide additional information in comparison to other approaches that merely analyze signals irrespective of the source.

[0104] In various embodiments, the phased sequencing information comprises mutation sequence information of the cell-free DNA. For example, mutation sequence information can include one or more mutations present across a plurality of genomic sites. In particular embodiments, the mutation sequence information includes one or more mutations that originate from a common source (e.g., a maternal chromosome or a paternal chromosome). Here, two or more genomic sites derived from a common source that have a particular pattern of mutations (e.g., each having a mutation, some pattern of mutated / non-mutated, or all nonmutated) can be referred to as coupled genomic sites. In various embodiments, a mutationcan be any of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.

[0105] In various embodiments, the phased sequencing information comprises methylation sequence information of the cell-free DNA. Methylation sequence information can include methylation statuses across a plurality of genomic sites. In particular embodiments, the methylation sequence information includes methylation statuses of genomic sites from a common source (e.g., a maternal chromosome or a paternal chromosome). As a specific example, methylation at a first genomic site may be coupled with methylation at a second genomic site on the same maternal or paternal chromosome. Two or more genomic sites with a particular methylation pattern (e.g., all methylated, partially methylated, or non-m ethylated) that originate from the same maternal or paternal chromosome is referred to herein as coupled methylation sites. Example coupled methylation sites may be two or more CGIs disclosed herein (e.g., two or more CGIs disclosed in any of Tables 1-4). In various embodiments, two or more genomic sites of coupled methylation sites may be separated by tens, hundreds, or even thousands of bases. Thus, coupled methylation sites include two or more genomic sites from a common source and need not be limited to genomic sites that are close in proximity (e.g., adjacent CpG sites). In various embodiments, coupled methylation sites include 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more methylation sites from a common source. Thus, detecting these coupled methylation sites may provide disease diagnostic utility.

[0106] In various embodiments, generating phased sequencing information for target nucleic acids comprises aligning sequence reads of target nucleic acids to long sequence reads of reference nucleic acids derived from different sources (e.g., either the maternal or paternal chromosome). Long sequence reads of reference nucleic acids originating from different sources can be distinguished due to sequence differences present in the long sequence reads. For example, given a particular chromosome, long sequence reads derived from a maternal chromosome would have sequence differences in comparison to long sequence reads derived from a paternal chromosome. Here, sequence differences can refer to mutations that are present in long sequence reads from one source, but not present in long sequence reads from the second source, and vice versa. Thus, the presence or absence of certain mutations can beuseful for distinguishing whether a long sequence read originated from a first source or a second source. Altogether, by comparing sequences of long sequence reads, a first set of long sequence reads with a set of common sequences can be attributed to a first source (e.g., a maternal chromosome) whereas a second set of long sequence reads with a different set of common sequences can be attributed to a second source (e.g., a paternal chromosome). In various embodiments, the different sets of long sequence reads need not specifically be attributed to a maternal chromosome and a paternal chromosome; rather, it is sufficient to distinguish different sets of long sequence reads from a first source and a second source. These long sequence reads from a first source or a second source have sufficiently different sequences to enable phasing of the target nucleic acids (e.g., to determine sources from which target nucleic acids were derived from).

[0107] By aligning sequence reads of target nucleic acids to long sequence reads of reference nucleic acids, the long sequence reads of reference nucleic acids serve as digital guides to phase e.g., determine the source of target nucleic acids. For example, target nucleic acids from a first common source (e.g., from a maternal chromosome) can be categorized together based on sequence similarities between the target nucleic acids and the long sequence reads of reference nucleic acids from the first source. Additionally, target nucleic acids from a second common source (e.g., from a paternal chromosome) can be categorized together based on sequence similarities between the target nucleic acids and the long sequence reads of reference nucleic acids from the second source. In contrast to using the standard human genome to align sequence reads of target nucleic acids, using long reads of reference nucleic acids would enable alignment of reference nucleic acids to sequences of the maternal or paternal chromosome Individual-specific differences between target nucleic acids deriving from the maternal and paternal chromosomes could be used as markers to create haplotype-specific sequence information that is informative for determining presence or absence of a cancer.

[0108] In various embodiments, phased sequencing information includes phased methylation sequencing information of cfDNA, where at least a first set of the phased methylation sequencing information of cfDNA originates from a first source and at least a second set of the phased methylation sequencing information of cfDNA originates from a second source. In various embodiments, methods for generating phased sequencing information can further include comparing the first set of the phased methylation sequencing information of cfDNA from the first source to the second set of the phased methylationsequencing information of cfDNA from the second source. In particular embodiments, generating phased sequencing information further includes comparing methylation statuses of two or more genomic sites from a first source to methylation statuses of the same two or more genomic sites from a second source. Differences in methylation statuses of genomic sites from the first source and the second source can be valuable for inclusion in the signal informative for determining presence or absence of a cancer. For example, if multiple genomic sites from a first source are methylated but the same genomic sites from a second source are unmethylated, this may be an informative signal for presence or absence of a cancer.Screen

[0109] The description in this section pertains to the performance of a screen, such as screen 125 described in FIG. 1 A, which can be performed by the screen module 210 described in FIG. 2A. Generally, a screen is performed on marker information generated by the assay (e.g., assay 120A). In various embodiments, the screen is performed to determine whether a biological sample is at risk or not at risk of containing a signal indicative of a cancer. For example, the screen is performed to determine whether a biological sample is at risk or not at risk of containing circulating tumor DNA. Circulating DNA within the biological sample may indicate that the individual (e.g., individual from whom the biological sample is obtained) may be at risk of a cancer. In various embodiments, the screen is performed to classify the subject as negative for cancer or not negative for cancer.

[0110] In various embodiments, the marker information represents quantified values of biomarkers. For example, depending on the type of biomarker, the quantified values may be generated via one or more of: an immunoassay, a protein-binding assay, an antibody -based assay, an antigen-binding protein-based assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), a Western blot, quantitative PCR (qPCR) or digital PCR (dPCR), NMR, mass spectrometry, LC-MS, or UPLC-MS / MS.

[0111] In various embodiments, performing the screen involves comparing the quantified values of biomarkers to one or more reference values or to threshold values. For example, a reference value can be a statistical measure of quantified biomarker values corresponding to individuals known to be at risk for cancer. Therefore, if the comparison identifies that the quantified values of biomarkers for an individual is statistically significantly different fromthe reference value corresponding to individuals known to be at risk for cancer, then the screen can identify the cancer as negative for cancer.

[0112] In various embodiments, the marker information represents sequencing information for one or more genomic locations, such as one or more CpG islands. In various embodiments, performing the screen involves comparing methylation information at one or more pre-selected genomic locations to quantified values of reference genomic locations. For example, referring again to FIG. 3B, an assay may have been performed that generates methylation information for locus 1 corresponding to a reference genomic location and for locus 2 corresponding to a pre-selected genomic location (e.g., a pre-selected informative CGI). Thus, the methylation information at locus 1 is compared to methylation information at locus 2. Based on the comparison, the screen can identify the subject as not negative for cancer.

[0113] In various embodiments, the screen can be a cheaper and less complex test in comparison to the second tier analysis (e.g., the second analysis). The screen can analyze marker information at a low resolution for purposes of identifying and removing large proportions of individuals that are not at risk of cancer. In various embodiments, the screen analyzes methylation information across a plurality of genomic locations and determines a measure of overall methylation across the plurality of genomic locations. Here, the measure of overall methylation across the plurality of genomic sites can represent methylation information of low resolution. Specifically, the measure of overall methylation provides a metric for methylation across the plurality of genomic sites, but may not provide information as to methylation status at each individual genomic site. The measure of overall methylation can be sufficient for identifying and removing large proportions of individuals not at risk for cancer. In various embodiments, the overall methylation across the plurality of genomic sites can be a total number of methylated CpG sites. In various embodiments, the overall methylation across the plurality of genomic sites can be a total number of methylated CpG sites across the plurality of genomic sites located in a subset of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the overall methylation across the plurality of genomic sites can be a total number of methylated CpG sites across the plurality of genomic sites located in all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the overall methylation across the plurality of genomic sites can be an average number of methylated CpG sites (e.g., an average number of methylated CpG sites within a target region or a CGI). In various embodiments, the overall methylation across the plurality of genomicsites can be an average number of methylated CpG sites across the plurality of genomic sites located in a subset of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the overall methylation across the plurality of genomic sites can be an average number of methylated CpG sites across the plurality of genomic sites located in all of the CGIs in any one of Tables 1, 2, 3, or 4.

[0114] In various embodiments, performing the screen involves performing whole genome sequencing or whole genome bisulfite sequencing and determining the overall methylation across the whole genome. Thus, in such embodiments, performing the screen is not limited to only analyzing CGIs or portions thereof; rather, performing the screen involves analyzing methylation statuses across the whole genome. In various embodiments, analyzing the methylation statuses across the whole genome can involve determining a quantifiable measure of the overall methylation across the whole genome. In various embodiments, the quantifiable measure of overall methylation is a score, such as a whole genome methylation burden score. In various embodiments, the higher the whole genome methylation burden score, the more likely the biological sample is at risk for containing circulating tumor DNA. In various embodiments, the lower the whole genome methylation burden score, the less likely the biological sample is at risk for containing circulating tumor DNA. In various embodiments, the biological sample is classified as negative (e.g., not at risk for containing circulating tumor DNA) or not negative (e.g., at risk for containing circulating tumor DNA) based on the determined whole genome methylation burden score. For example, if the whole genome methylation burden score for the biological sample is above a threshold score, the biological sample can be classified as not negative. As another example, if the whole genome methylation burden score for the biological sample is below a threshold score, the biological sample can be classified as negative.

[0115] In various embodiments, the measure of overall methylation across one or more preselected genomic locations and methylation information for reference genomic locations can be a cycle threshold (Ct) value. Cycle threshold refers to the number of PCR cycles needed for a sample to amplify and cross a threshold. In various embodiments, if a difference between the Ct value of the methylation sequences of the pre-selected genomic locations and the Ct value of the reference genomic locations is greater than a threshold, then the screen identifies the subject as not negative for cancer. If a difference between the Ct value of the methylation sequences of the pre-selected genomic locations and the Ct value of the referencegenomic locations is less than a threshold, then the screen identifies the subject as negative for cancer.

[0116] In various embodiments, a screen is performed on sequence information generated via sequencing (e.g., next generation sequencing) of sequences at the one or more genomic locations, such as one or more CpG islands. In various embodiments, such a screen is performed using a system comprising a computer storage and a processing system. The screen can further involve the implementation of a machine learning model. For example, the computer storage can store sequence information corresponding to a processed sample, the processed sample including cell-free DNA fragments originating from a liquid biopsy of an individual and having been processed to enrich for cancer informative CGIs, the sequencer information comprising, for each sequenced cell-free DNA fragment corresponding to the cancer informative CGIs, a respective position on the genome for the cell-free DNA fragment and methylation information for the cell-free DNA fragment. The processing system can compute values of the cancer informative CGIs for the individual and applies the values as input to a trained machine learning model. The machine learning model provides a predicted output as to whether the individual is at risk for cancer based on the values of the cancer informative CGIs.

[0117] In various embodiments, performing the screen involves analyzing a plurality of CGIs. For example, performing the screen involves analyzing methylation statuses of a plurality of CGIs. Cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. The accompanying tables (e.g., Tables 1-4) lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0118] In some embodiments, performing the screen involves analyzing a plurality of CGIs including one or more CGIs that are methylated in the genome of extraembryonic ectoderm (ExE). Here, such example CGIs may be differentially methylated in the genome of ExE andnot methylated in corresponding epiblast or adult tissue. Example CGIs that are methylated in the genome of ExE are further disclosed in Table 3 of WO2022133315, which is hereby incorporated by reference in its entirety.

[0119] In various embodiments, performing the screen involves analyzing all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 1. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 1. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 2. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 2. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 3. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 3. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Table 4. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Table 4. In various embodiments, performing the screen involves analyzing at most 10% of the CGIs in Tables 2 and 3. In various embodiments, performing the screen involves analyzing at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, or at most 99% of the CGIs in Tables 2 and 3.

[0120] In various embodiments, performing the screen involves analyzing 1 CGI, 2 CGIs, 3 CGIs, 4 CGIs, 5 CGIs, 6 CGIs, 7 CGIs, 8 CGIs, 9 CGIs, 10 CGIs, 11 CGIs, 12 CGIs, 13 CGIs, 14 CGIs, 15 CGIs, 16 CGIs, 17 CGIs, 18 CGIs, 19 CGIs, 20 CGIs, 21 CGIs, 22 CGIs,23 CGIs, 24 CGIs, 25 CGIs, 26 CGIs, 27 CGIs, 28 CGIs, 29 CGIs, 30 CGIs, 31 CGIs, 32CGIs, 33 CGIs, 34 CGIs, 35 CGIs, 36 CGIs, 37 CGIs, 38 CGIs, 39 CGIs, 40 CGIs, 41 CGIs,42 CGIs, 43 CGIs, 44 CGIs, 45 CGIs, 46 CGIs, 47 CGIs, 48 CGIs, 49 CGIs, or 50 CGIs(e.g., CGIs as shown in any of Tables 1-4 or portions of CGIs shown in any of Tables 1-4). In various embodiments, performing the screen involves analyzing at most 2 CGIs, at most 5 CGIs, at most 10 CGIs, at most 15 CGIs, at most 20 CGIs, at most 25 CGIs, at most 30 CGIs, at most 35 CGIs, at most 40 CGIs, at most 45 CGIs, or at most 50 CGIs (e.g., CGIs as shown in any of Tables 1-4 or portions of CGIs shown in any of Tables 1-4). In various embodiments, performing the screen involves analyzing at most 50 CGIs, at most 100 CGIs, at most 150 CGIs, at most 200 CGIs, at most 300 CGIs, at most 400 CGIs, at most 500 CGIs, at most 600 CGIs, at most 700 CGIs, at most 800 CGIs, at most 900 CGIs, at most 1000 CGIs, at most 1500 CGIs, at most 2000 CGIs, at most 2500 CGIs, at most 3000 CGIs, at most 3500 CGIs, at most 4000 CGIs, at most 4500 CGIs, at most 5000 CGIs, at most 5500 CGIs, or at most 6000 CGIs (e.g., CGIs as shown in any of Tables 1-4 or portions of CGIs shown in any of Tables 1-4). In particular embodiments, performing the screen involves analyzing at most 500 CGIs.

[0121] In various embodiments, the screen achieves at least 60% sensitivity in detecting presence of a cancer. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sensitivity. In particular embodiments, the screen achieves at least 75% sensitivity. In particular embodiments, the screen achieves at least 76% sensitivity. In particular embodiments, the screen achieves at least 77% sensitivity. In particular embodiments, the screen achieves at least 78% sensitivity. In particular embodiments, the screen achieves at least 79% sensitivity. In particular embodiments, the screen achieves at least 80% sensitivity.

[0122] In various embodiments, the screen achieves at least 60% specificity in excluding individuals without cancer. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% specificity. In particular embodiments, the screen achieves at least 90% specificity. In particular embodiments, the screen achieves at least 91% specificity. In particular embodiments, the screen achieves at least 92% specificity. In particular embodiments, the screen achieves at least 93% specificity. In particular embodiments, the screen achieves at least 94% specificity. In particular embodiments, the screen achieves at least 95% specificity.

[0123] In various embodiments, the screen achieves at least 15% positive predictive value. In various embodiments, the screen achieves at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least39%, or at least 40% positive predictive value. In particular embodiments, the screen achieves at least 20% positive predictive value. In particular embodiments, the screen achieves at least 21% positive predictive value. In particular embodiments, the screen achieves at least 22% positive predictive value. In particular embodiments, the screen achieves at least 23% positive predictive value. In particular embodiments, the screen achieves at least 24% positive predictive value. In particular embodiments, the screen achieves at least 25% positive predictive value. In particular embodiments, the screen achieves at least 26% positive predictive value. In particular embodiments, the screen achieves at least 27% positive predictive value. In particular embodiments, the screen achieves at least 28% positive predictive value. In particular embodiments, the screen achieves at least 29% positive predictive value. In particular embodiments, the screen achieves at least 30% positive predictive value. In particular embodiments, the screen achieves at least 31% positive predictive value. In particular embodiments, the screen achieves at least 32% positive predictive value. In particular embodiments, the screenachieves at least 33% positive predictive value. In particular embodiments, the screen achieves at least 34% positive predictive value. In particular embodiments, the screen achieves at least 35% positive predictive value. In particular embodiments, the screen achieves at least 36% positive predictive value. In particular embodiments, the screen achieves at least 37% positive predictive value. In particular embodiments, the screen achieves at least 38% positive predictive value. In particular embodiments, the screen achieves at least 39% positive predictive value. In particular embodiments, the screen achieves at least 40% positive predictive value.

[0124] In various embodiments, the screen achieves at least 60% negative predictive value. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% negative predictive value. In particular embodiments, the screen achieves at least 95% negative predictive value. In particular embodiments, the screen achieves at least 96% negative predictive value. In particular embodiments, the screen achieves at least 97% negative predictive value. In particular embodiments, the screen achieves at least 98% negative predictive value. In particular embodiments, the screen achieves at least 99% negative predictive value.Intra-Individual Analysis

[0125] The description in this section pertains to the performance of an intra-individual analysis, such as an intra-individual analysis 128A and / or intra-individual analysis 128B described in FIGs. 1A-1B. In general, the intra-individual analyses are conducted for subjects that were previously determined (e.g., via screen 125 as shown in FIG. 1 A) as not negative for cancer. The intra-individual analysis removes baseline biological signatures that are specific for a subject to generate a background-corrected signal. Thus, the second analysis involves analyzing the background-corrected signal to determine whether the individual has cancer.

[0126] In various embodiments, an intra-individual analysis is conducted using a single sample, such as a blood sample. The sample may contain target nucleic acids and reference nucleic acids. Target nucleic acids may include signatures that are informative of determining presence or absence of a cancer, and can further include baseline biological signatures. Here, target nucleic acids in the blood sample may be derived from a diseased cell which is associated with the cancer. For example, target nucleic acids can include cell-free DNA in the blood that originates from a diseased cell. In particular embodiments, target nucleic acids are cell-free DNA in the blood that originates from a cancer cell. Reference nucleic acids in the sample refer to nucleic acids that contain baseline biological signatures of the individual. For example, baseline biological signatures of the individual may be present in nucleic acids irrespective of whether the nucleic acids originate from a diseased source, or a non-diseased source. The baseline biological signatures of the reference nucleic acids are generally less informative for determining presence or absence of a cancer in comparison to the informative signatures present in the target nucleic acids. In various embodiments, reference nucleic acids refer to cellular genomic DNA derived from a healthy cell from the individual. In various embodiments, reference nucleic acids found in the sample derive from a cell in a healthy organ of the individual. Example organs include the brain, heart, thorax, lung, abdomen, colon, cervix, pancreas, kidney, liver, muscle, lymph nodes, esophagus, intestine, spleen, stomach, and gall bladder. In particular embodiments, reference nucleic acids are found in the sample and refer to cellular genomic DNA derived from peripheral blood mononuclear cells (PBMCs) (e.g., lymphocytes or monocytes) or polymorphonuclear cells (e.g., eosinophils or neutrophils).

[0127] In various embodiments, target nucleic acids and reference nucleic acids are separately obtained from the single sample. In various embodiments, the sample is processed to separate the target nucleic acids and reference nucleic acids. For example, the sample may be processed through any one of centrifugation, filtration, gel electrophoresis, bead capture, or matrix extraction. In particular embodiments, target nucleic acids are cell-free nucleic acids and therefore, can be obtained from the supernatant of the separated sample. In particular embodiments, reference nucleic acids are cellular genomic nucleic acids and therefore, can be obtained from a different portion of the separated sample that contains cells.

[0128] Generally, an intra-individual analysis is performed on sequence information of target nucleic acids and sequence information of reference nucleic acids. In particular embodiments, the sequence information of target nucleic acids comprise sequenceinformation of cell free DNA. In particular embodiments, the sequence information of reference nucleic acids comprise sequence information of cells, such as peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells.

[0129] The intra-individual analysis involves combining the sequence information of target nucleic acids and sequence information of reference nucleic acids to generate a background- corrected signal informative for determining presence or absence of a cancer. In various embodiments, combining the sequence information of target nucleic acids and sequence information of reference nucleic acids involves differentiating between signatures present or absent in the sequence information of target nucleic acids and signatures present or absent in the sequence information of the reference nucleic acids. For example, if particular signatures are present in the sequence information of target nucleic acids, and the signatures are also present in the sequence information of reference nucleic acids, the signatures in both the target nucleic acids and reference nucleic acids may represent baseline biological signatures. Thus, these signatures may be excluded from the resulting signal informative of determining presence or absence of the cancer. As another example, if particular signatures are present in the sequence information of target nucleic acids, but those signatures are absent in the sequence information of reference nucleic acids, the signatures may not be baseline biological signatures. Thus, these signatures may be included in the resulting signal informative of determining presence or absence of the cancer.

[0130] In various embodiments, combining the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids includes aligning the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids. For example, aligning the sequence information involves aligning sequences of a plurality of pre-selected genomic sites for the target nucleic acids and sequences of the same or overlapping plurality of pre-selected genomic sites for the reference nucleic acids.

[0131] In various embodiments, both the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are aligned to a reference genome library (e.g., a reference assembly) with known sequences. Therefore, sequence information of the target nucleic acids are aligned to the sequence information of the reference nucleic acids via the reference genome library. In various embodiments, the sequence information of the target nucleic acids is aligned directly with the sequence information of the reference nucleic acids. In such embodiments, a reference genome library need not be used.

[0132] In various embodiments, combining the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids includes determining a difference between the sequence information of the target nucleic acids to the sequence information of the reference nucleic acids.

[0133] In various embodiments, differences between the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are performed on a per-position basis. For example, at a first position of a genomic site, the difference between the sequence information of the target nucleic acids at the first position and the sequence information of the reference nucleic acid at the same first position is determined. The process can then be further repeated for additional positions (e.g., for additional positions across the plurality of genomic sites). In various embodiments, the differences are determined on a per- position basis if the sequence information of the target nucleic acids and reference nucleic acids were generated using a sequencing assay (e.g., next generation sequencing) which provides base-level resolution of the sequences.

[0134] In various embodiments, differences between the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are performed on a per-CGI basis. For example, at a first CGI of a genomic site, the difference between the sequence information of the target nucleic acids at the first CGI and the sequence information of the reference nucleic acid at the same CGI or overlapping portion of the first CGI is determined. The process can then be further repeated for additional CGIs (e.g., for additional CGIs across the plurality of genomic sites). In various embodiments, the differences are determined on a per-CGI basis if the sequence information of the target nucleic acids and reference nucleic acids were generated using a quantitative assay (e.g., qPCR assay).

[0135] In various embodiments, differences between the sequence information of the target nucleic acids and the sequence information of the reference nucleic acids are performed on a per-allele basis. For example, at a first allele of a genomic site, the difference between the sequence information of the target nucleic acids at the first allele and the sequence information of the reference nucleic acid at the same allele or overlapping portion of the first allele is determined. The process can then be further repeated for additional alleles (e.g., for additional alleles across the plurality of genomic sites). In various embodiments, the differences are determined on a per-allele basis if the sequence information of the target nucleic acids and reference nucleic acids were generated using a quantitative assay (e.g., qPCR assay or allele-specific PCR assay).

[0136] In various embodiments, the intra-individual analysis generates a background- corrected signal that comprises phased sequencing information. As described herein, phased sequence information is derived specifically from a particular source and therefore, may be useful for determining presence or absence of a cancer because signals originating from the same source (e.g., maternal or paternal chromosome) may provide additional information in comparison to other approaches that merely analyze signals irrespective of the source. In various embodiments, performing the intra-individual analysis includes removing baseline biological signatures that would otherwise have been interpreted as being derived from a particular source. As described herein, phased sequencing information can include coupled genomic sites and / or coupled methylation sites from common sources. Therefore, by performing the intra-individual analysis, the coupled genomic sites and / or coupled methylation sites can be informative signatures deriving from common sources as opposed to baseline biological signatures.

[0137] Reference is now made to FIG. 3C, which depicts an example combining of sequence information of target nucleic acids and reference nucleic acids to generate a signal informative for a cancer, in accordance with an embodiment. The sequence information of the target nucleic acids and the sequence information of the reference nucleic acids include methylation statuses across a plurality of genomic sites. FIG. 3C shows an example genomic site in which nucleotide bases may be differentially methylated in the target nucleic acid and the reference nucleic acid. For example, as shown in FIG. 3C, the nucleotide base at the second position is methylated (as represented by the presence of a cytosine base which arises following bisulfite conversion) in both the target nucleic acid and the reference nucleic acid. Given that the methylation at the second position occurs in both the target nucleic acid and the reference nucleic acid, this may be a baseline biological signature. Conversely, the target nucleic acid may additionally be methylated at the sixth position and the ninth position, whereas the reference nucleic acid is unmethylated at the sixth position and the ninth position. Here, given that the reference nucleic acid is not methylated at the sixth and ninth position, the presence of the methylated nucleotide bases in the target nucleic acid may represent signatures that are informative of presence or absence of the cancer. Additionally, at the eleventh nucleotide position, the target nucleic acid is unmethylated whereas the reference nucleic acid is methylated. Here, the methylation of the reference nucleic acid can be interpreted as a baseline biological signature.

[0138] The differences between the methylation status at each position of the target nucleic acid and the reference nucleic acid can represent the cancer signal. As shown in FIG. 3C, the cancer signal includes methylation statuses at the genomic site, wherein the sixth and ninth position are methylated. Thus, the cancer signal includes signatures from the target nucleic acids that are likely informative of the cancer (e.g., methylated statuses of the sixth and ninth nucleotide bases), and further excludes baseline biological signatures (e.g., baseline biological signatures present in reference nucleic acids such as methylated statuses of the second and eleventh nucleotide bases).Second Analysis

[0139] The description in this section pertains to the performance of a second analysis, such as second analysis 130 described in FIG. 1 A, which can be performed by the second analysis module 220 described in FIG. 2A. Generally, a second analysis is performed on sequence information generated by the assay (e.g., assay 120A or assay 120B). In various embodiments, the second analysis is performed to determine whether a biological sample obtained from an individual contains a signal indicative of a cancer. For example, the screen is performed to determine whether a biological sample contains circulating tumor DNA. Circulating DNA within the biological sample may indicate that the individual (e.g., individual from whom the biological sample is obtained) has cancer. In various embodiments, the second analysis is performed on background-corrected methylation information from an intra-individual analysis to classify the subject as having cancer or not having cancer. In various embodiments, the second analysis is performed to analyze a change in background-corrected methylation information from two or more intra-individual analyses. By analyzing a change in background-corrected methylation information, the second analysis can predict a change in tumor heterogeneity e.g., for tracking tumor heterogeneity in the subject for guided therapy.

[0140] In various embodiments, a second analysis is performed on background-corrected sequence information generated via sequencing (e.g., next generation sequencing) of sequences at the one or more genomic locations, such as one or more CpG islands. In various embodiments, the background-corrected sequence information is generated as a result of whole genome sequencing and therefore, a second analysis is performed on sequences of one or more genomic locations across the whole genome.

[0141] Generally, the second analysis is a more expensive and / or a more complex test in comparison to the first tier (e.g., screen). By implementing a more complex second analysis, the second analysis can achieve a higher positive predictive value than the first tier. In various embodiments, performing the second analysis involves analyzing methylation information across a plurality of genomic locations that represents a higher resolution in comparison to the lower resolution information analyzed in the first tier. For example, the second analysis may determine a high resolution measure of methylation across the plurality of genomic sites that distinguishes individuals having cancer from other individuals not having cancer in accordance with a high performance metric (e.g., high PPV or high sensitivity). Here, the high resolution measure of methylation can provide information as to methylation status at each individual genomic site and / or methylation statuses across a group of genomic sites.

[0142] In various embodiments, the high resolution measure of methylation can be a total quantity of consecutively methylated CpG sites within target regions. In some embodiments, the high resolution measure of methylation can be a total quantity of 3 consecutively methylated CpG sites (referred to as “K3N3”) within target regions. In some embodiments, the high resolution measure of methylation can be a total quantity of 4 consecutively methylated CpG sites (referred to as “K4N4”) within target regions. In some embodiments, the high resolution measure of methylation can be a total quantity of 5 consecutively methylated CpG sites (referred to as “K5N5”) within target regions. For example, the high resolution measure of methylation can be a total quantity of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within a subset of the CGIs in any one of Tables 1, 2, 3, or 4. As another example, the high resolution measure of methylation can be a total quantity of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within all of the CGIs in any one of Tables 1, 2, 3, or 4. In some embodiments, the high resolution measure of methylation can be a proportion of 3 consecutively methylated CpG sites (referred to as “K3N3”) within target regions. In some embodiments, the high resolution measure of methylation can be a proportion of 4 consecutively methylated CpG sites (referred to as “K4N4”) within target regions. In some embodiments, the high resolution measure of methylation can be a proportion of 5 consecutively methylated CpG sites (referred to as “K5N5”) within target regions. For example, the high resolution measure of methylation can be a proportion of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within a subset of the CGIs in any one of Tables 1, 2, 3, or 4. As another example, the high resolution measure of methylation can bea proportion of 3, 4, 5, 6, 7, 8, 9, or 10 consecutively methylated CpG sites within all of the CGIs in any one of Tables 1, 2, 3, or 4.

[0143] In some embodiments, the high resolution measure of methylation can be a total quantity of consecutively methylated CpG sites within one or more CGIs that are methylated in the genome of extraembryonic ectoderm (ExE). Here, such example CGIs may be differentially methylated in the genome of ExE and not methylated in corresponding epiblast or adult tissue. Example CGIs that are methylated in the genome of ExE are further disclosed in Table 3 of WO2022133315, which is hereby incorporated by reference in its entirety.

[0144] In various embodiments, the high resolution measure of methylation can include methylation statuses of a plurality of CpG sites from a haplotype (e.g., inherited from either a maternal or paternal source). In various embodiments, the high resolution measure of methylation refers to methylation statuses of at least a portion of the CpGs within a CGI within at least a portion of one or more regions in Tables 1-4 from a common haplotype. In various embodiments, the high resolution measure of methylation refers to methylation statuses of all CpGs within a CGI within at least a portion of one or more regions in Tables 1- 4 from a common haplotype. In various embodiments, the high resolution measure of methylation refers to methylation statuses of all CpGs within a CGI within one or more regions in Tables 1-4 from a common haplotype.

[0145] In various embodiments, the second analysis is performed using a system comprising a computer storage and a processing system. The second analysis can involve the implementation of trained machine learning models, details of which are described in further detail herein. For example, the computer storage can store sequence information corresponding to a processed sample, the processed sample including cell-free DNA fragments originating from a liquid biopsy of an individual and having been processed to enrich for cancer informative CGIs, the sequencer information comprising, for each sequenced cell-free DNA fragment corresponding to the cancer informative CGIs, a respective position on the genome for the cell-free DNA fragment and methylation information for the cell-free DNA fragment.

[0146] In particular embodiments, the second analysis further reveals, for individuals who are determined to have the cancer, a tissue of origin of the cancer. The second analysis may identify a tissue of origin of the cancer according to the methylation statuses of the cancer informative CGIs. For example, particular methylation patterns across the cancer informative CGIs are attributable to certain tissues, examples of which include the nervous tissue (e.g.,brain, spinal cord, nerves), muscle tissue (cardiac muscle, smooth muscle, skeletal muscle), epithelial tissue (e.g., GI tract lining, skin), and connective tissue (e.g., fat, bone, tendon, and ligaments). As a particular example, in patients with brain cancer, a first set of CGIs may be frequently methylated. Therefore, if a similar methylation pattern is observed across the first set of CGIs for an individual who is under analysis, the second analysis can identify that the individual has cancer, and furthermore, that the cancer is localized to the brain.

[0147] In various embodiments, the second analysis involves analyzing a plurality of CGIs. For example, the second analysis involves analyzing methylation statuses of a plurality of CGIs. Cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. The accompanying tables (e.g., Tables 1-4) lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In various embodiments, the second analysis involves analyzing all of the CGIs in any one of Tables 1, 2, 3, or 4. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 1. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 1. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 2. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 2. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 3. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 3.In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Table 4. In various embodiments, the second analysis involves analyzing at least 10%, atleast 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Table 4. In various embodiments, the second analysis involves analyzing at least 10% of the CGIs in Tables 2 and 3. In various embodiments, the second analysis involves analyzing at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the CGIs in Tables 2 and 3.

[0148] In various embodiments, the second analysis involves analyzing at least 100 CGIs (e.g., CGIs as shown in any of Tables 1-4). In various embodiments, the second analysis involves analyzing at least 100 CGIs, at least 150 CGIs, at least 200 CGIs, at least 300 CGIs, at least 400 CGIs, at least 500 CGIs, at least 600 CGIs, at least 700 CGIs, at least 800 CGIs, at least 900 CGIs, at least 1000 CGIs, at least 1500 CGIs, at least 2000 CGIs, at least 2500 CGIs, at least 3000 CGIs, at least 3500 CGIs, at least 4000 CGIs, at least 4500 CGIs, at least 5000 CGIs, at least 5500 CGIs, or at least 6000 CGIs (e.g., CGIs as shown in any of Tables 1-4). In particular embodiments, performing the screen involves analyzing at least 500 CGIs. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0149] In various embodiments, the second analysis involves analyzing more CGIs in comparison to the quantity of CGIs analyzed during the screen. For example, the CGIs analyzed during the screen can represent a subset of the CGIs analyzed during the second analysis. In some scenarios, every CpG island analyzed during the screen is further analyzed when performing the second analysis. Therefore, the second analysis represents a more robust and rigorous analysis in comparison to the more rapid and cost-effective screen. In various embodiments, the second analysis involves analyzing at least 2 times the number of CGIs analyzed during the screen. In various embodiments, the second analysis involves analyzing at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least 14 times at least 15 times, at least 16 times, at least 17 times, at least 18 times,at least 19 times, at least 20 times, at least 21 times, at least 22 times, at least 23 times, at least 24 times, at least 25 times, at least 26 times, at least 27 times, at least 28 times, at least 29 times, at least 30 times, at least 31 times, at least 32 times, at least 33 times, at least 34 times, at least 35 times, at least 36 times, at least 37 times, at least 38 times, at least 39 times, or at least 40 times the number of CGIs analyzed during the screen. In particular embodiments, the second analysis involves analyzing at least 5 times the number of CGIs analyzed during the screen. For example, the screen may involve analyzing at least 100 CGIs and the second analysis may involve analyzing at least 500 CGIs.

[0150] In various embodiments, the second analysis achieves at least 60% sensitivity in detecting presence of a cancer. In various embodiments, the screen achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sensitivity. In particular embodiments, the second analysis achieves at least 85% sensitivity. In particular embodiments, the second analysis achieves at least 86% sensitivity. In particular embodiments, the second analysis achieves at least 87% sensitivity. In particular embodiments, the second analysis achieves at least 88% sensitivity. In particular embodiments, the second analysis achieves at least 89% sensitivity. In particular embodiments, the second analysis achieves at least 90% sensitivity.

[0151] In various embodiments, the second analysis achieves at least 60% specificity in excluding individuals without the cancer. In various embodiments, the second analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% specificity. In particular embodiments, the second analysis achievesat least 90% specificity. In particular embodiments, the second analysis achieves at least 91% specificity. In particular embodiments, the second analysis achieves at least 92% specificity. In particular embodiments, the second analysis achieves at least 93% specificity. In particular embodiments, the second analysis achieves at least 94% specificity. In particular embodiments, the second analysis achieves at least 95% specificity.

[0152] In various embodiments, the second analysis achieves at least 60% positive predictive value. In various embodiments, the second analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% positive predictive value. In particular embodiments, the second analysis achieves at least 80% positive predictive value. In particular embodiments, the second analysis achieves at least 81% positive predictive value. In particular embodiments, the second analysis achieves at least 82% positive predictive value. In particular embodiments, the second analysis achieves at least 83% positive predictive value. In particular embodiments, the second analysis achieves at least 84% positive predictive value. In particular embodiments, the second analysis achieves at least 85% positive predictive value.

[0153] In various embodiments, the second analysis achieves at least 60% negative predictive value. In various embodiments, the second analysis achieves at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% negative predictive value. In particular embodiments, the second analysis achieves at least 90% negative predictive value. In particular embodiments, the second analysis achieves at least 91% negative predictive value. In particular embodiments, the second analysis achieves atleast 92% negative predictive value. In particular embodiments, the second analysis achieves at least 93% negative predictive value. In particular embodiments, the second analysis achieves at least 94% negative predictive value. In particular embodiments, the second analysis achieves at least 95% negative predictive value. In particular embodiments, the second analysis achieves at least 96% negative predictive value. In particular embodiments, the second analysis achieves at least 97% negative predictive value. In particular embodiments, the second analysis achieves at least 98% negative predictive value. In particular embodiments, the second analysis achieves at least 99% negative predictive value.Longitudinal Analysis

[0154] In various embodiments, methods disclosed herein are valuable for performing longitudinal analysis for a subject. For example, a subject who was determined to have a presence of cancer (e.g., through the screen or through the second analysis) can be further tracked through a longitudinal analysis. In various embodiments, an additional sample is obtained from the subject at a subsequent timepoint, and the second analysis can be further performed for the subject using the additional sample. Thus, the second analysis performed for the additional sample can determine a change in the cancer for the subject over the intervening timeframe.

[0155] In various embodiments, a longitudinal analysis can be performed for subjects who may have been identified as not having cancer. In various embodiments, a longitudinal analysis is performed for subjects who were identified as negative through the screen (e.g., first analysis). In various embodiments, a longitudinal analysis is performed for subjects who were identified as negative through the second analysis. In various embodiments, a longitudinal analysis is performed for subjects who were identified as not negative through the screen and then further identified as negative through the second analysis. By longitudinally tracking subjects who may have been identified as not having cancer, any false negative subjects can potentially be identified through subsequent testing of one or more additional samples obtained at one or more subsequent timepoints. For example, a subject can be identified as not negative through the screen, and through the longitudinal analysis (e.g., at a subsequent timepoint), an additional sample of the subject can be analyzed using either the methodology described in reference to the screen or the second analysis to identify the subject as a false negative. As another example, a subject can be identified as not negative through the second analysis, and through the longitudinal analysis (e.g., at asubsequent timepoint), an additional sample of the subject can be analyzed using either the methodology described in reference to the screen or the second analysis to identify the subject as a false negative.

[0156] Reference is now made to the tumor tracking module 230, which represents a module of the tumor heterogeneity system 170 as shown in FIG. 2A. In various embodiments, tracking tumor heterogeneity over two or more timepoints enables the determination of whether an intervention is efficacious. Given a subject who has previously received the intervention (e.g., a tumor therapeutic) for treating cancer, tracking tumor heterogeneity over two or more timepoints using the methods disclosed herein is informative for determining whether the intervention is efficacious for treating the cancer. Generally, a subject exhibiting a reduction in tumor heterogeneity over two or more timepoints is indicative that the tumor subclones are decreasing and that the intervention is effective. Alternatively, a subject who does not exhibit a reduction in tumor heterogeneity (e.g., stable or increase tumor heterogeneity) is indicative that the tumor subclones is unchanging or is increasing. In this scenario, the intervention lacks efficacy. Thus, methods for tracking tumor heterogeneity can be useful for e.g., guided therapy.

[0157] In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining samples from the subject across two or more timepoints, performing intraindividual analysis for one or more of the obtained samples, and generating predictions across at least the two or more timepoint. The predictions can be informative for the subject’s tumor heterogeneity. In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining three or more samples from a subject across at least three timepoints, performing intra-individual analysis for the three or more samples, and generating predictions across the at least three timepoints. In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining four or more samples from a subject across at least four timepoints, performing intra-individual analysis for the four or more samples, and generating predictions across the at least four timepoints. In various embodiments, tracking tumor heterogeneity for a subject comprises obtaining samples from a subject, performing intra-individual analysis for each of the obtained samples, and generating predictions across at least five timepoints, at least six timepoints, at least seven timepoints, at least eight timepoints, at least nine timepoints, at least ten timepoints, at least eleven timepoints, at least twelve timepoints, at least thirteen timepoints, at least fourteen timepoints, at least fifteen timepoints, at leastsixteen timepoints, at least seventeen timepoints, at least eighteen timepoints, at least nineteen timepoints, or at least twenty timepoints.

[0158] In various embodiments, the time between any two timepoints can be between 1 day and 12 months, between 5 days and 8 months, between 10 days and 6 months, between 15 days and 4 months, between 20 days and 3 months, between 30 days and 2 months. In various embodiments, the time between any two timepoints can be between 1 days and 10 days, between 10 days and 20 days, between 20 days and 30 days, between 30 days and 40 days, between 40 days and 50 days, or between 50 days and 60 days. In various embodiments, the time between any two timepoints can be between 1 day and 100 days, between 5 day and 80 days, between 10 days and 70 days, between 15 days and 60 days, between 20 days and 50 days, between 25 days and 40 days, or between 30 days and 35 days. In various embodiments, the time between any two timepoints can be between 1 days and 10 days, between 10 days and 20 days, between 20 days and 30 days, between 30 days and 40 days, between 40 days and 50 days, or between 50 days and 60 days. In various embodiments, the time between any two timepoints can be between 1 month and 2 months.

[0159] In various embodiments, methods for tracking tumor heterogeneity involve obtaining a sample from the subject at a first timepoint (e.g., an initial timepoint), performing an intraindividual analysis using the obtained sample, and generating a cancer prediction for the sample obtained at the first timepoint. In various embodiments, the first timepoint may refer to a timepoint prior to which the subject receives an intervention, such as a tumor therapeutic. Thus, the generated for the sample obtained at the first timepoint may represent a baseline prediction prior to any therapeutic treatment. In various embodiments, the first timepoint may refer to a timepoint immediately after the subject receives an intervention, such as a tumor therapeutic. In this context, “immediately after” the subject receives an intervention can refer to a timeframe within 1 day after the subject receives the intervention. In various embodiments, “immediately after” refers to a timeframe within 12 hours, within 8 hours, within 6 hours, within 4 hours, within 3 hours, within 2 hours, within 1 hour, within 30 minutes, within 15 minutes, within 10 minutes, within 5 minutes, or within 1 minute of the subject receiving the therapeutic.

[0160] In particular embodiments, methods for tracking tumor heterogeneity further involve obtaining one or more subsequent samples from the subject after the first timepoint (e.g., at a second timepoint, at a third timepoint, at a fourth timepoint, etc.), performing intra-individual analyses for a subsequent sample, and generating predictions for the one or more subsequentsamples. In this scenario, the change in the predictions for the one or more subsequent samples in comparison to the prediction of the first sample can be indicative of the change in tumor heterogeneity. In various embodiments, the one or more subsequent samples are obtained from the subject after the subject has received an intervention, such as a tumor therapeutic. Thus, the change in tumor heterogeneity can be reflective of the efficacy, or lack thereof, of the intervention provided to the subject.Machine Learning Models for Analyzing Sequence Information

[0161] In various embodiments, trained machine learning models can be deployed to analyze sequence information for tracking tumor heterogeneity for a subject across two or more timepoints. In various embodiments, the sequence information includes methylation statuses of plurality of genomic sites. Therefore, trained machine learning models analyze differential methylation of the plurality of genomic sites to output predictions.

[0162] In various embodiments, a trained machine learning model is deployed as part of a screen (e.g., screen 125 as shown in FIG. 1 A). Thus, the trained machine learning model can analyze sequence information generated via an assay (e.g., assay 120 A shown in FIG. 1 A) to determine whether a subject is negative or not negative for a cancer. In various embodiments, a trained machine learning model is deployed as part of a second analysis (e.g., second analysis 130 shown in FIG. 1 A). Therefore, the trained machine learning model can analyze sequence information including methylation statuses for a plurality of genomic sites, such as a plurality of CpG sites disclosed herein. In various embodiments, the sequence information includes background-corrected sequence information generated via an intraindividual analysis (e.g., intra-individual analysis 128A and / or intra-individual analysis 128B shown in FIG. 1 A). In some embodiments, the trained machine learning model analyzes a difference between background-corrected sequence information determined from two intra- individual analyses (as shown in FIG. 1 A). In some embodiments, the trained machine learning model analyzes background-corrected sequence information from a single intra- individual analysis (as shown in FIG. IB).

[0163] In various embodiments, a machine learning model is any one of a regression model (e.g., linear regression, logistic regression, or polynomial regression), decision tree, random forest, support vector machine, Naive Bayes model, k-means cluster, or neural network (e.g., feed-forward networks, convolutional neural networks (CNN), deep neural networks (DNN), autoencoder neural networks, generative adversarial networks, or recurrent networks (e.g.,long short-term memory networks (LSTM), bi-directional recurrent networks, deep bidirectional recurrent networks).

[0164] The machine learning model can be trained using a machine learning implemented method, such as any one of a linear regression algorithm, logistic regression algorithm, decision tree algorithm, support vector machine classification, Naive Bayes classification, K- Nearest Neighbor classification, random forest algorithm, deep learning algorithm, gradient boosting algorithm, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or combinations thereof. In various embodiments, the machine learning model is trained using supervised learning algorithms, unsupervised learning algorithms, semi-supervised learning algorithms (e.g., partial supervision), weak supervision, transfer, multi-task learning, or any combination thereof.

[0165] In various embodiments, the machine learning model has one or more parameters, such as hyperparameters or model parameters. Hyperparameters are generally established prior to training. Examples of hyperparameters include the learning rate, depth or leaves of a decision tree, number of hidden layers in a deep neural network, number of clusters in a k- means cluster, penalty in a regression model, and a regularization parameter associated with a cost function. Model parameters are generally adjusted during training. Examples of model parameters include weights associated with nodes in layers of neural network, support vectors in a support vector machine, and coefficients in a regression model. The model parameters of the machine learning model are trained (e.g., adjusted) using the training data to improve the predictive power of the machine learning model.

[0166] In particular embodiments, trained machine learning models analyze methylation statuses of a plurality of genomic sites to generate predictions. The methylation statuses can correspond to a set of cancer informative CpG islands (CGIs), wherein the cancer informative CGIs are selected from a group consisting of a ranked set of candidate CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 50 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 100 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 150 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 200 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 250 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 300 CGIs. In variousembodiments, a machine learning model analyzes methylation statuses for at least 400 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 500 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 600 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 700 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 800 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 900 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 1000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 2500 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 5000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 7500 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 10000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 15000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 20000 CGIs. In various embodiments, a machine learning model analyzes methylation statuses for at least 25000 CGIs.

[0167] In various embodiments, a machine learning model analyzes methylation statuses for CGIs across the whole genome. For example, a machine learning model may be implemented to analyze sequencing data generated from whole genome sequencing (e.g., whole genome bisulfite sequencing).

[0168] Additionally disclosed herein are particular genomic sites, such as CpG islands (CGIs) whose methylation statuses can be informative for determining whether a subject is at risk of a cancer or whether the individual has a cancer. In some embodiments, methylation statuses of the informative CGIs representing a signal in a sample can be indicative of a presence of the cancer. In some embodiments, methylation statuses of the informative CGIs representing a signal in a sample can be indicative of an absence of the cancer. In various embodiments, methods disclosed herein, such as methods involving the multiple-tiered analysis, are useful for detecting or identifying the signal (e.g., methylation statuses of the informative CGIs) in a sample. In various embodiments, methods disclosed herein, such as methods involving the multiple-tiered analysis, are useful for increasing the probability that the detected signal (e.g., methylation statuses of the informative CGIs) in the sample is authentic. A signal (e.g., methylation statuses of the informative CGIs) detected by themultiple-tiered analysis can be confidently trusted as present in the sample. Thus, by tracking the change in methylation statuses for the subject across multiple timepoints, a change in the subject’s risk for cancer or a change in the subject’s cancer can be more accurately determined.

[0169] Methylation statuses of cancer informative CGIs can be useful for predicting whether an individual has a cancer or is at risk for a cancer. In various embodiments, the methylation statuses of cancer informative CGIs are background-corrected methylation statuses of cancer informative CGIs. For example, background-corrected methylation statuses of cancer informative CGIs can be determined via an intra-individual analysis. For example, background-corrected methylation statuses of cancer informative CGIs can be determined by combining methylation information of cancer informative CGIs of target nucleic acids and methylation information of cancer informative CGIs of reference nucleic acids.

[0170] In various embodiments, each cancer informative CGI can be a “CGI identifier” or reference number to allow referencing CGIs during data processing by their respective unique CGI identifiers. The accompanying tables (e.g., Tables 1-4) lists, for each CGI, its respective location in the human genome. Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. In some embodiments, methylation statuses of a plurality of CpGs within a CGI may be analyzed. In some embodiments, at least a portion of the CpGs within a CGI may be analyzed. In other embodiments, all of the CpGs within a CGI may be analyzed. In some embodiments, an analysis of a CGI as contemplated herein may comprise analyzing CpGs within at least a portion of one or more regions in Tables 1-4.

[0171] Reference is now made to FIG. 3D, which is an illustrative example of a signal informative for a cancer. In various embodiments, the signal informative for a cancer shown in FIG. 3D can be generated as a result of the intra-individual analysis. Thus, the signal informative for a cancer represents background-corrected sequence information e.g., corrected via an intra-individual analysis that combines sequence information from target nucleic acids and reference nucleic acids. In various embodiments, the signal informative for a cancer shown in FIG. 3D can represent sequence information of target nucleic acids. In such embodiments, the signal is not derived from an intra-individual analysis.

[0172] As shown in FIG. 3D, for each instance of an analyte, e.g., a cell-free DNA fragment, there is data indicating, for each of a plurality of positions along the instance of theanalyte, e.g., distinct CpG sites along a DNA fragment, information about a marker at that position, e.g., whether that CpG is methylated or unmethylated. An instance of an analyte can be a single sequenced DNA fragment or a portion of a single sequenced DNA fragment. In various embodiments, the DNA fragment may be a bisulfite converted DNA fragment. Therefore, an instance of an analyte may refer to a sequenced bisulfite converted DNA fragment or a portion thereof.

[0173] Conceptually, using methylation of CpGs in cell-free DNA as an illustrative example, the signal illustrated in FIG. 3D includes a row, e.g., row 240, for each instance of an analyte, such as a single sequenced DNA fragment. Thus, in FIG. 3D, data for sixteen instances of an analyte are shown, e.g., sixteen DNA fragments. Each circle corresponds to a position along the analyte, such as a CpG site. In this example, whether the circle is illustrated as black or white in FIG. 3D, is indicative of whether the CpG site is methylated (black) or unmethylated (white). In some instances, information about a marker at a position in a nucleic acid may not be binary.

[0174] The information about the markers for each instance of an analyte in a sample can result in a large amount of data. As an example, in practice, in the case of obtaining methylation state of CpGs in cell-free DNA from a blood sample using deep sequencing, using a DNA sequencer that outputs such data into a FASTQ format data file, the signal generated by processing a single blood sample can be many gigabytes, e.g., 20 to 30 gigabytes, of data.

[0175] FIG. 3D also illustrates a relative alignment among the distinct instances of the analyte. In the example of DNA, for example, the position of a DNA fragment within a genome for the individual from which a sample originated can be determined, and each position within the genome can have a respective set of coordinates identifying it. Thus, DNA fragments can be assigned coordinates based on their respective positions within the genome, and then aligned or grouped by those coordinates. Thus, in FIG. 3D, column 242 indicates a position on an analyte, such as a single CpG site in a genome, and the distinct instances of the analyte are illustrated as aligned by position on the analyte.

[0176] By using the position information for each instance of an analyte, distinct instances of the analyte can be grouped into regions within the analyte. Typically, markers related to cancer are localized within identifiable regions of analytes, such as specific genes or regions within the genome. Thus, the signals generated for each instance of an analyte can be grouped and processed by cancer-informative regions. In particular embodiments, aninformative region is a CGI (or at least a portion thereof) as disclosed in any of Tables 1-4. The example in FIG. 3D can be considered to illustrate data about methylation at CpG sites within one informative region of the genome, for multiple DNA fragments obtained from a biological sample. There can be multiple cancer-informative regions.

[0177] As disclosed herein, trained machine learning models are deployed to generate informative predictions regarding presence or absence of cancer. To use a trained machine learning model in this context, there are several technical problems that arise relating to encoding the signal resulting from processing a biological sample into features. Some problems arise because the signal includes a large amount of information. One of the challenges involves reducing the volume of data into a set of informative features. However, as the number of features increases, the complexity of the computational model increases. However, as the number of features decreases, information relevant to detection of a cancer may be lost. Some problems arise because of uncertainty around which metrics and which regions of an analyte are truly informative of a cancer. Omission of some metrics or some regions from the set of features may adversely impact the performance of a trained computational model.

[0178] To address such problems, in various embodiments, very particularly engineered features are generated from a biological sample. Such engineered features may be dependent on one or more health-condition-informative regions (e.g., CGIs) and / or one or more distinct windows within the health-condition informative regions (e.g., CGIs). Each window may have a specified range of positions within a health-condition informative region, and a specified size. The size is specified in terms of a number of consecutive sites of interest within the analyte. A metric is thus computed for a plurality of windows within the healthcondition informative region. Thus, in particular embodiments, the engineered features, representing metrics within a particular window within a health-condition informative region (e.g., CGIs), are informative for a cancer.

[0179] To train a machine learning model, in some embodiments, a first set of features is computed for a training set, which can include several candidate features. The candidate features can include one or more candidate metrics, or one or more candidate health- condition-informative regions, or combinations of both. A computational model can be trained using candidate features, and then analyzed to determine which candidate features were more influential in the output of the trained computational model. Such analysis can be used to identify features which are more influential to the model, whether due to the metric ordue to the health-condition-informative region. A second set of features can be defined by reducing the first set of features based on those identified features which are more influential, and the trained machine learning model can be built using the second set of features.

[0180] In various embodiments, to generate data for a machine learning model (e.g., for training or for deployment), the methodology includes computing, for one or more instances of an analyte in a window of a plurality of windows on a target region of the analyte, a metric specific for the window and the target region. The specific metrics used, and health- condition-informative regions selected can depend on a variety of factors and may be experimentally determined. The machine learning model can be implemented to analyze at least the metric specific for the window and the target region. In various embodiments, the metric specific for the window and the target region includes a proportion of a count of DNA fragments having a specific count of methylated CpGs to a count of DNA fragments for the window of the target region. In various embodiments, the metric specific for the window and the target region comprises a proportion of a count of DNA fragments having a specific pattern of methylation to a count of DNA fragments for the window of the target region. As described in further detail below, computing the metric can involve applying two or more functions. For example, computing the metric specific for the window and the target region can involve performing a first function to quantify a count of occurrences of methylated CpGs within the window of the target region. As another example, computing the metric specific for the window and the target region can involve performing a second function to normalize the count of occurrences of methylated CpGs relative to a count of DNA fragments for the window of the target region.

[0181] In various embodiments, to generate features, each instance of the analyte (e.g., cell- free DNA) is processed. For each instance of an analyte in the biological sample, and for each window of a plurality of windows on health-condition-informative regions of the analyte, a respective value is generated. After processing instances of the analyte, the feature computation module then computes, for each window of the plurality of windows on the health-condition-informative region, one or more respective metrics for the window based on a first function and / or a second function for instances of the analyte for the window. In various embodiments, a first function quantifies markers within a window. As a specific example, a first function refers to a quantification of a number of methylated CpG sites within a window. In various embodiments, a second function computes a proportion of the quantified markers within the window in relation to other quantified markers. As a specificexample, a second function computes the proportion of the number of methylated CpG sites within a window relative to other numbers of methylated CpG sites within a window.

[0182] Example implementations will now be described in reference to FIGs. 3E and 3F. Here, in FIG. 3E, illustrative marker information for instances of an analyte are shown schematically for the purposes of simplifying this explanation. In this example, there are ten (10) instances of an analyte, each having a length of six (6) sites of interest, at which marker information is a binary value, indicated by a black or white circle. FIG. 3E shows aligned instances of an analyte, along with the designation of a window with a particular kmer size (e.g., K=3). Each window has a size of three (3) consecutive sites of interest within the analyte. In other embodiments, smaller or larger window sizes may be implemented for the analysis. There are four (4) windows of size three (3) (i.e., a first window that includes the first, second, and third sites of interest from the left, a second window that includes the second, third, and fourth sites of interest from the left, a third window that includes the third, fourth, and fifth sites of interest from the left, and a fourth window that includes the fourth, fifth, and sixth sites of interest from the left), but computations for three (3) windows are shown.

[0183] In FIG. 3E an example of a first function applied to an instance of an analyte is a count of occurrences of marker information within the instance of the analyte within the window. For example, where the marker information is methylation of a CpG site, this function can be a count of methylated CpGs in the window. That is, if the window has a size of three sites of interest, then there are four possible counts: 0, 1, 2, and 3. Note that inverse results would be obtained if the count was of unmethylated CpGs in the window, but such results when used in training would have the same effect.

[0184] In FIG. 3E, the second function computes counts of the number of instances having each possible count resulting from the first function. That is, if the window has a size of three sites of interest, for which there are four possible counts (0, 1, 2, and 3), for that window the second function computes a count of the number of instances with a count of zero, a count of the number of instances with a count of one, a count of the number of instances with a count of two, and a count of the number of instances with a count of three. The second function divides the respective number of instances computed for possible counts by the total number of instances, thus providing a fractional value for each of the possible counts for this window.

[0185] In this example in FIG. 3E, for this health-condition-informative region (referred to as “HC1”), there are windows “Wl”, “W2”, and “W3”, each of which has four (4) values,representing the respective count for each possible count of methylated CpGs among the instances that overlap that window. Because there are ten (10) instances, each of these values is divided by 10 in the second function, to provide the respective final four output values for each window. As shown in FIG. 3E, referring to the example of Window 1 (Wl), the final four output values are 0.3 (0 methylated CpG sites in the window), 0.1 (1 methylated CpG sites in the window), 0.1 (2 methylated CpG sites in the window), and 0.5 (3 fully methylated CpG sites in the window). Here, the proportion of fully methylated CpG sites, proportion of fully non-methylated CpG sites, and proportion of partially methylated CpG sites (e.g., either 1 or 2 methylated CpG sites in the window) can be metrics informative for a cancer.

[0186] Reference is now made to FIG. 3F, which shows an example application of a first function and second function to instances of an analyte. Here, the bottom of FIG. 3F shows patterns of the marker information in the instance, from among a set of possible patterns. A pattern is a unique sequence of marker information along the sites of interest in a window. For example, as shown in FIG. 3F, if the window has a size of three sites of interest, and if the marker information for the sites of information is binary, then there are eight possible patterns. For example, where the marker information is methylation of a CpG site, each possible pattern of methylation in a window is a distinct sequence of the methylation state (e.g., methylated or unmethylated) of the CpG sites along the sequence of consecutive CpG sites in the window. When the marker information is methylation of CpG sites, the first function, applied to an instance of a DNA fragment in a window, outputs an indication of which of the possible patterns of methylation of CpGs is present in the window in that DNA fragment.

[0187] The second function computes a count of the number of instances having each possible pattern in a window. That is, for that window, the second function produces a count of the number of instances with the first pattern, a count of the number of instances with the second pattern, and so on. The second function then divides the respective number of instances identified for each possible pattern by the total number of instances, thus providing a fractional value for each of the possible patterns for this window, as shown in the bottom panel of FIG. 3F.

[0188] In this example in FIG. 3F, for this health-condition-informative region (say, “HC1”), there are windows “Wl”, “W2”, and “W3”, each of which has eight values, representing the respective number of occurrences each possible pattern among the instances that overlap that window divided by the number of instances, in this case ten (10).

[0189] In any of the foregoing example implementations, and in other implementations, a size of a health-condition-informative region, in terms of a number of sites of interest within an instance of an analyte, can vary. For example, cancer-informative regions of DNA may be as small as a single CpG site, and may include several 10’s, 100’s, or 1000’s of CpG sites. Within a set of features, there may be a plurality of health-condition-informative regions, each having its own respective size.

[0190] In any of the foregoing example implementations, and in other implementations, a size of a window in a health-condition-informative region, in terms of a number of sites of interest within an instance of an analyte, can vary. Generally, the number of sites of interest is a positive integer number that ranges between 1 and N. In some example implementations, N is less than or equal to 10, or 9, or 8, or 7, or 6, or 5, or 4, or 3. In various embodiments, a window within a health-condition informative region includes a specific numbers of CpG sites. In various embodiments, TV is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites. In various embodiments, TV is between 1 and 100, between 2 and 80, between 3 and 60, between 4 and 40, between 5 and 20, or between 6 and 10 CpG sites. In various embodiments, N is between 1 and 10, between 2 and 9, between 3 and 8, between 4 and 7, or between 4 and 6 CpG sites. Within a set of features, there may be a plurality of health-condition-informative regions, each having its own respective window size or set of window sizes. Different window sizes may be used in different regions. The same window size may be used in different regions. A region may have metrics computed for it for multiple different window sizes. Windows may be over-lapping or non-overlapping.

[0191] In various embodiments, a metric represents an input vector that can be provided as input to a machine learning model (e.g., either during training or deployment of the machine learning model). Here, the metric may be specific for a window and a target region of interest (e.g., a target region comprising one or more CpG sites). For example, the input vector of the metric may include a set of values representing the proportion of counts of methylated CpGs in the window relative to a total count (e.g., total count of DNA fragments for the window of the target region). In various embodiments, the input vector of the metric may include a set of values representing proportions of DNA fragments having specific counts of methylated CpGs out of all possible CpG methylation patterns in the window. The all possible CpG methylation patterns are 2kpossible patterns, where k refers to a number of CpG sites in the window. Referring against to the bottom panel of FIG. 3F, an input vector of a metric can be generated for a particular window. Taking the first window (e.g., left-mostwindow shown in FIG. 3F) as an example, the input vector of the metric may include the proportion vales shown in the left most column in the bottom panel of FIG. 3F. Thus, the input vector of the metric may be represented as [0.3, 0.1, 0, 0, 0, 0.1, 0, 0.5], Similar input vectors for other metrics can be generated using the values of other windows.

[0192] The computed sets of values for the set of features for samples can be stored in a data structure, which can be stored in a database, memory, or other computer storage for use in connection with the computational model, or for other purposes.

[0193] In some implementations, the sets of values for the set of features for a sample can be stored in association with an identifier of the subject, or an identifier of the sample, or both, so that the identifier of the subject or the identifier of the sample, or both, can be used to access the set of values from the computer storage. In some implementations, each computed value can be associated with an identifier of the cancer-informative region, and an identifier of the window within that region, to which the value corresponds.

[0194] Accordingly, an example implementation of such a data structure is shown in FIG. 3G. A set of values for a set of features is stored for a biological sample originating from a subject. The data structure can include an optional identifier for the subject, and an optional identifier for the biological sample. The latter identifier is useful when there are multiple samples for a single subject. For a sample, as indicated at 250, the set of features includes one or more metrics, for each of one or more windows 254, e.g., window “W-l-1”, within each of one or more health-condition-informative regions 252A, e.g., region “Rl” or 252B e.g., region “R2”. For each feature, e.g., R-l, W-l-1, Metric, the computed value, e.g., Value 256, is stored. The number of windows in each region can be different for each region. The size of the window can be different for each window. The metric(s) computed for the window can be different for each window.Example Methods for Conducting Two or More Intra-Individual Analyses

[0195] As disclosed herein, methods involve tracking tumor heterogeneity in a subject by conducting intra-individual analyses for two or more samples obtained from the subject across two or more timepoints. For example, a first intra-individual analysis can be performed for a first sample obtained from the subject at a first timepoint and a second intra- individual analysis can be performed for a second sample obtained from the subject at a second timepoint. Thus, the change in results from each intra-individual analysis can be informative for tracking tumor heterogeneity in the subject.

[0196] FIG. 4A shows an example flow process involving a first and second intra-individual analyses, in accordance with a first embodiment. In this first embodiment, the flow process involves performing separate intra-individual analyses for first and second samples obtained from the subject at two different timepoints and performing a second analysis on the difference between the results of the separate intra-individual analyses.

[0197] Step 410 involves performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA.

[0198] Next, at step 415, if the first biological sample is not identified as not at risk, perform a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information.

[0199] Step 420 involves performing a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information, the second biological sample obtained from the subject at a second timepoint subsequent to the first timepoint.

[0200] Step 425 involves determining a change in signal between the first set of background-corrected methylation information and the second set of background-corrected methylation information.

[0201] Step 430 involves performing a second analysis comprising analyzing the determined change in signal to track tumor heterogeneity.

[0202] Reference is now made to FIG. 4B, which shows an example flow process involving a first and second intra-individual analyses, in accordance with a second embodiment. In this second embodiment, the flow process involves performing separate intra-individual analyses for first and second samples obtained from the subject at two different timepoints and performing a second analysis on each of the results of the separate intra-individual analyses.

[0203] Step 450 involves performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA.

[0204] Step 455 involves performing a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information.

[0205] Step 460 involves performing a second analysis to predict a tumor heterogeneity state.

[0206] Step 465 involves performing a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information, the second biological sample obtained from the subject at a second timepoint subsequent to the first timepoint.

[0207] Step 470 involves performing a second analysis to predict an updated tumor heterogeneity state.

[0208] Step 475 involves determining a change in signal between the first set of background-corrected methylation information and the second set of background-corrected methylation information.Guided Therapy

[0209] In various embodiments, the methods disclosed herein for performing a multipletiered analysis (e.g., screening and / or intra-individual analysis) to track tumor heterogeneity of one or more cancers in one or more subjects are informative for identifying an intervention for the subject. In various embodiments, an intervention may be any intervention known to those of ordinary skill in the art. Non-limiting examples of interventions include surgery (e.g., excising diseased or pre-disease tissue from an individual), a tumor therapeutic (e.g., chemotherapy, gene therapy, or gene editing), radiation therapy, or a lifestyle intervention (e.g., change in behavior or habits). In particular embodiments, the intervention comprises a tumor therapeutic.

[0210] In various embodiments, the methods disclosed herein are performed for a subject who previously received a tumor therapeutic. Thus, tracking the tumor heterogeneity of one or more cancers for the subject can be informative for determining whether the previously provided tumor therapeutic is efficacious. For example, if the tumor heterogeneity of a cancer is not decreasing (e.g., is increasing or is remaining stable) over the two or more timepoints, the tumor therapeutic is deemed non-efficacious. In this example, methods can involve selecting a new intervention, such as a new or different tumor therapeutic, for treatment of the subject’s cancer. As another example, if the tumor heterogeneity is decreasing over the two or more timepoints, the tumor therapeutic can be deemed efficacious. In this example, methods can involve selecting the tumor therapeutic that was previously provided to subject. Thus, the tumor therapeutic can continue to be provided to the subject totreat the cancer. In some embodiments, methods can involve selecting a new or different tumor therapeutic for treatment of the subject’s cancer. In some embodiments, methods can involve selecting a new or different intervention in addition to the previously provided tumor therapeutic. Thus, the new or different intervention and the previously provided tumor therapeutic can be provided to the subject to treat the cancer.Cancers

[0211] The disclosure provides methods for performing a multiple-tiered analysis (e.g., screening and / or intra-individual analysis) to track tumor heterogeneity of one or more cancers in one or more subjects. In various embodiments, the subject may have been previously diagnosed with a cancer and receives an intervention for treating the cancer. For example, the subject may have previously received a tumor therapeutic for treating the cancer. In various embodiments, the subject may be suspected of having a cancer, but may not have been previously diagnosed with a cancer. In various embodiments, the subject is healthy and is not yet suspected of having a cancer. In certain embodiments, a cancer is an early-stage health cancer, e.g., prior to development of symptoms.

[0212] In various embodiments, the cancer is an early stage cancer. In various embodiments, the cancer is a preclinical phase cancer. In various embodiments, the cancer is a stage I cancer. In various embodiments, the cancer is a stage II cancer. Thus, the methods disclosed herein enable the screening and tracking of tumor heterogeneity of a subject for an early stage or preclinical stage cancer.

[0213] In various embodiments, the cancer is any of an acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, cardiac cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumor, gestational trophoblastic cancer, hairy cell leukemia, liver cancer, Hodgkin lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, mouth cancer, multiple endocrine neoplasia syndromes, multiple myeloma neoplasms, myelodysplastic neoplasms, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer,prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.Computer Implementation

[0214] The methods of the invention, including the methods of performing a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints, are, in some embodiments, performed on one or more computers. In particular embodiments, the steps of performing a screen (e.g., screen 125 shown in FIG.1 A), performing an intra-individual analysis (e.g., intra-individual analysis 128A or intraindividual analysis 128B shown in FIG. 1 A), and performing a second analysis (e.g., second analysis 130 shown in FIG. 1 A) are performed on one or more computers. The steps of performing an assay (e.g., assay 120A and / or assay 120B shown in FIG. 1 A) are not performed on one or more computers.

[0215] In various embodiments, the performance of the screen, the intra-individual analysis, and / or the second analysis can be implemented in hardware or software, or a combination of both. In one embodiment, a machine-readable storage medium is provided, the medium comprising a data storage material encoded with machine readable data which, when using a machine programmed with instructions for using said data, is capable of displaying data (e.g., methylation data) and results of the screen, intra-individual analysis, and / or second analysis (e.g., tracked tumor heterogeneity). Such data can be used for a variety of purposes, such as determining an efficacy of a tumor therapeutic, or selecting a new intervention for the subject. The invention can be implemented in computer programs executing on programmable computers, comprising a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), a graphics adapter, a pointing device, a network adapter, at least one input device, and at least one output device. A display is coupled to the graphics adapter. Program code is applied to input data to perform the functions described above and generate output information. The output information is applied to one or more output devices, in known fashion. The computer can be, for example, a personal computer, microcomputer, or workstation of conventional design.

[0216] Each program can be implemented in a high level procedural or object oriented programming language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language can be a compiled or interpreted language. Each such computer program is preferably stored on astorage media or device (e.g., ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The system can also be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.

[0217] The signature patterns and databases thereof can be provided in a variety of media to facilitate their use. “Media” refers to a manufacture that contains the signature pattern information of the present invention. The databases of the present invention can be recorded on computer readable media, e.g., any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; and hybrids of these categories such as magnetic / optical storage media. One of skill in the art can readily appreciate how any of the presently known computer readable mediums can be used to create a manufacture comprising a recording of the present database information. “Recorded” refers to a process for storing information on computer readable medium, using any such methods as known in the art. Any convenient data storage structure can be chosen, based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g. word processing text file, database format, etc.

[0218] In some embodiments, the methods disclosed herein, are performed on one or more computers in a distributed computing system environment (e.g., in a cloud computing environment). In this description, “cloud computing” is defined as a model for enabling on- demand network access to a shared set of configurable computing resources. Cloud computing can be employed to offer on-demand access to the shared set of configurable computing resources. The shared set of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly. A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“laaS”). A cloud-computingmodel can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.Example Computer

[0219] FIG. 5 illustrates an example computer for implementing the entities shown in FIGs. 1A-1C, 2A, 3A-3G, and 4A-4B. In particular embodiments, the example computer 500 can represent computational system 202 described in FIG. 2A. The computer 500 includes at least one processor 502 coupled to a chipset 504. The chipset 504 includes a memory controller hub 520 and an input / output (I / O) controller hub 422. A memory 506 and a graphics adapter 512 are coupled to the memory controller hub 520, and a display 518 is coupled to the graphics adapter 512. A storage device 508, an input device 514, and network adapter 516 are coupled to the I / O controller hub 522. Other embodiments of the computer 500 have different architectures.

[0220] The storage device 508 is a non-transitory computer-readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or a solid-state memory device. The memory 506 holds instructions and data used by the processor 502. The input device 514 is a touch-screen interface, a mouse, track ball, or some combination thereof, and is used to input data into the computer 500. The keyboard 510 may be another device for inputting data into the computer 500. In some embodiments, the computer 500 may be configured to receive input (e.g., commands) from the input device 514 via gestures from the user. The graphics adapter 512 displays images and other information on the display 518. The network adapter 516 couples the computer 500 to one or more computer networks.

[0221] The computer 500 is adapted to execute computer program modules for providing functionality described herein. As used herein, the term “module” refers to computer program logic used to provide the specified functionality. Thus, a module can be implemented in hardware, firmware, and / or software. In one embodiment, program modules are stored on the storage device 508, loaded into the memory 506, and executed by the processor 502. A module can be implemented as computer program code processed by the processing system(s) of one or more computers. Computer program code includes computerexecutable instructions and / or computer-interpreted instructions, such as program modules, which instructions are processed by a processing system of a computer. Generally, such instructions define routines, programs, objects, components, data structures, and so on, that, when processed by a processing system, instruct the processing system to perform operationson data or configure the processor or computer to implement various components or data structures in computer storage. A data structure is defined in a computer program and specifies how data is organized in computer storage, such as in a memory device or a storage device, so that the data can accessed, manipulated, and stored by a processing system of a computer.

[0222] The types of computers 500 used by the entities of FIG. 1C can vary depending upon the embodiment and the processing power required by the entity. For example, the tumor heterogeneity system 170 can run in a single computer 500 or multiple computers 500 communicating with each other through a network such as in a server farm. The computers 500 can lack some of the components described above, such as graphics adapters 512, and displays 518.Kit Implementation

[0223] Also disclosed herein are kits for performing a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints. Such kits can include equipment to draw a sample from a patient. For example, kits can include syringes and / or needles for obtaining a sample from a patient. Kits can include detection reagents for determining marker information using the sample obtained from the patient.

[0224] For example, detection reagents can include antibody reagents for performing a protein immunoassay. As another example, detection reagents can be a set of primers that, when combined with the sample, allows detection of a plurality of sites in cell-free DNA in the sample. In particular embodiments, the detection reagents enable detection of methylated or unmethylated target sites (e.g., methylated or unmethylated informative CpGs including one or more CGIs selected from Tables 1-4, or one or more CpGs within at least a portion of a region in Tables 1-4). Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety. For example, the detection reagents may be primers that target specific known sequences of target sites, thereby enabling nucleic acid amplification of the target sites.Thus, the use of the detection reagents results in generation of methylation information of the patient corresponding to the target sites.

[0225] A kit can include instructions for use of one or more sets of detection reagents. For example, a kit can include instructions for performing at least one detection assay such as anucleic acid amplification assay (e.g., polymerase chain reaction assay including any of realtime PCR assays, quantitative real-time PCR (qPCR) assays, allele-specific PCR assays, and reverse-transcription PCR assays), nucleic acid sequencing (e.g., targeted gene sequencing, targeted amplicon sequencing, whole genome sequencing, or whole genome bisulfite sequencing), hybrid capture, an immunoassay, a protein-binding assay, an antibody-based assay, an antigen-binding protein-based assay, a protein-based array, an enzyme-linked immunosorbent assay (ELISA), reporter assays, flow cytometry, a protein array, a blot, a Western blot, nephelometry, turbidimetry, chromatography, NMR, mass spectrometry, LC- MS, UPLC-MS / MS, enzymatic activity, proximity extension assay, and an immunoassay selected from RIA, immunofluorescence, immunochemiluminescence, immunoelectrochemiluminescence, immunoelectrophoretic, a competitive immunoassay, and immunoprecipitation.

[0226] Kits can further include instructions for accessing computer program instructions stored on a computer storage medium. In various embodiments, the computer program instructions, when executed by a processor of a computer system, cause the processor to perform one or more intra-individual analyses, generate background corrected methylation information, and / or track tumor heterogeneity across two or more timepoints.

[0227] In various embodiments, the kits include instructions for practicing the methods disclosed herein (e.g., performing an assay, screen, or diagnostic assay). These instructions can be present in the kits in a variety of forms, one or more of which can be present in the kit. One form in which these instructions can be present is as printed information on a suitable medium or substrate, e.g., a piece or pieces of paper on which the information is printed, in the packaging of the kit, in a package insert, etc. Yet another means would be a computer readable medium, e.g., diskette, CD, hard-drive, network data storage, etc., on which the information has been recorded. Yet another means that can be present is a website address which can be used via the internet to access the information at a removed site. Any convenient means can be present in the kits.Systems

[0228] Further disclosed herein are systems for performing a tiered, multipart method for tracking tumor heterogeneity across samples obtained from a subject at different timepoints. In various embodiments, such a system can include one or more sets of detection reagents for determining genomic information using a sample obtained from the patient, an apparatus configured to receive a mixture of the one or more sets of detection reagents and the sampleobtained from a subject to generate methylation information of the subject, and a computer system communicatively coupled to the apparatus to generate background-corrected methylation information and / or to track the change in tumor heterogeneity.

[0229] The one or more sets of detection reagents enable the determination of marker information using the sample obtained from the patient. For example, detection reagents can include antibody reagents for performing a protein immunoassay. For example, detection reagents can be a set of primers that, when combined with the sample, allows detection of a plurality of sites in cell-free DNA in the sample. In particular embodiments, the detection reagents enable detection of methylated or methylated target sites (e.g., methylated or unmethylated informative CpGs including one or more CGI’s selected from Tables 1-4 or one or more CpGs within at least a portion of a region in Tables 1-4). Additional example CGIs are disclosed in WO2018209361 (see Table 1) and WO2022133315 (see Table 2 entitled “TOO Methylation Sites” and Table 3 entitled “Pan Cancer Methylation Sites”), each of which is hereby incorporated by reference in its entirety.

[0230] The apparatus is configured to determine the methylation information from a mixture of the detection reagents and sample. For example, the apparatus can be configured to perform one or more of a nucleic acid amplification assay (e.g., polymerase chain reaction assay), nucleic acid sequencing (e.g., targeted gene sequencing, whole genome sequencing, or whole genome bisulfite sequencing), and hybrid capture to determine methylation information.

[0231] The mixture of the detection reagents and sample may be presented to the apparatus through various conduits, examples of which include wells of a well plate (e.g., 96 well plate), a vial, a tube, and integrated fluidic circuits. As such, the apparatus may have an opening (e.g., a slot, a cavity, an opening, a sliding tray) that can receive the container including the reagent test sample mixture and perform a reading. Examples of an apparatus include one or more of a sequencer, an incubator, plate reader (e.g., a luminescent plate reader, absorbance plate reader, fluorescence plate reader), a spectrometer, or a spectrophotometer.

[0232] The computer system, such as example computer 500 described in FIG. 5, communicates with the apparatus to receive the methylation information. The computer system generates background-corrected methylation information and can further track the change in tumor heterogeneity (e.g., based on the change of the background-corrected methylation information across two or more timepoints).EXAMPLES

[0233] Below are examples of specific embodiments for carrying out the present invention. The examples are offered for illustrative purposes only and are not intended to limit the scope of the present invention in any way. Efforts have been made to ensure accuracy with respect to numbers used (e.g., percentages, etc.), but some experimental error and deviation should be allowed for.Example 1: Overall performance of two-tier screening and diagnosis of patients with protstate cancer

[0234] FIG. 6 shows example performance of different tiers of the multiple tier analysis for diagnosing individuals with cancer (e.g., prostate cancer). Here, the process begins with 19 million individuals who underwent testing. At a 2% incidence rate, of the 19 million individuals, 380,000 are true positives, and 18.6 million are true negatives.

[0235] The multi-tiered analysis involves performing a screen by analyzing methylation data (generated via an assay) of the patients. Here, the screen is designed to achieve 80% sensitivity and 95% specificity, thereby identifying 1.2 million out of the original 19 million individuals as at risk for prostate cancer. Additionally, the screen identifies 17.8 million out of the original 19 million individuals as not at risk for prostate cancer. Thus, these 17.8 million individuals need not undergo further analysis. Altogether, the screen achieves a 25% positive predictive rate and a 99% negative predictive rate.

[0236] The 1.2 million individuals identifies as at risk for prostate cancer further undergo a second test in the form of the second analysis. The second analysis achieves a 90% sensitivity and a 95% specificity. Of the 1.2 million individuals, -320,000 individuals are identified as having prostate cancer. This represents a 85% positive predictive rate as 273,600 individuals were true positives and 47,000 were false positives. Additionally, the second analysis identifies 945,000 negatives, of which 884,450 were true negatives, and 30,400 were false negatives, thereby representing a 97% negative predictive value.

[0237] Altogether, the overall performance of the multi-tier screen and second analysis includes 72% sensitivity, 99.9% specificity, 85% positive predictive value, and 99.4% negative predictive value.

[0238] Example steps for performing the multiple-tier analysis shown in FIG. 6 are detailed below.Prepare target specimen

[0239] The target specimen type (e.g. DNA, RNA, protein, exosomes, metabolites, etc.) is isolated from a patient’s biological source (e.g. tissue, blood, plasma, serum, saliva, feces, etc.). That specimen can be isolated by a CRO or private or service laboratory or hospital or isolated internally using an internal procedure. Target specimens are assayed for quality and quantity measurements.Phase 1 testing

[0240] Phase 1 testing is a relatively quick, non-invasive assay with simple technology, using small amounts of the target specimen. The result of this assay can be both qualitative and quantitative. Phase 1 testing is typically lower specificity (e.g. 95% specificity, 5% false positives) but higher sensitivity (e.g. 80% sensitivity, 20% false negatives) in order to screen a large proportion of the testing population rapidly and inexpensively. The phase 1 assay will overall increase the incidence of the target population (e.g. diseased) for the phase 2 assay, which will then increase the positive predictive value (PPV). Examples of the Phase 1 assay include but are not limited to ELISA assays, PCR assays, Real-time PCR assays, Quantitative real-time PCR (qPCR) assays, Allele-specific PCR assays, Reverse-transcription PCR assays and reporter assays.Phase 2 testing

[0241] Phase 2 testing is a more complex, potentially invasive assay with complex technology, potentially using larger amounts of the target specimen. The result of this assay is both qualitative and quantitative. Phase 2 testing is typically higher specificity (e.g. 95% specificity, 10% false positives) but lower sensitivity (e.g. 90% sensitivity, 10% false negatives) in order to limit false positives. By screening out a large volume of the testing population, the target population has higher target incidence than the general population, which increases positive predictive value (PPV).Phase 2 Protocol

[0242] Examples of the phase 2 assay include but are not limited to Next Generation Sequencing assays utilizing target enrichment technologies, targeted amplicon sequencing technologies, whole genome sequencing, and whole genome bisulfite sequencing.

[0243] The target specimen for library construction is dsDNA isolated from formalin-fixed paraffin-embedded (FFPE) tissue. Alternatively, cfDNA is isolated from blood. For FFPE, the dsDNA is first mechanically sheared by the Covaris instrument utilizing adaptive focusedacoustics to a target insert size of 200 base pairs. Post-shearing, a solid-phase reversible immobilization (SPRI) selection is done to remove smaller DNA fragments remaining in solution. For blood DNA, cfDNA is isolated. The fragmented DNA is then end-repaired and A-tailed (ERAT) to produce 5 ’-phosphorylated, 3’-dA-tailed dsDNA fragments. After ERAT, dsDNA unique dual index adapters with 3’-dTMP overhangs are then ligated to 3’-dA-tailed dsDNA fragments. Indices allow for sample multiplex for the downstream assay. Postligation, a solid-phase reversible immobilization (SPRI) selection is done to remove unwanted DNA fragments, excess adapters and molecules. PCR amplification is performed with a high-fidelity, low-bias polymerase at 10 cycles. Post-PCR, a SPRI selection is done to remove unwanted DNA fragments, excess primers, excess adapters and excess molecules. After library construction, the library quality and quantity are evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively.

[0244] Libraries that pass quality control checks move forward to target enrichment through hybridization capture. Target enrichment by hybridization capture is defined as a positive selection strategy to enrich low abundance regions of interest from NGS libraries, allowing for more accurate sequencing analysis of these target regions. Indexed libraries are multiplexed and hybridized to a custom, sequence specific, biotinylated probeset. The vast excess of probes drives their hybridization to complementary library fragments. The library fragment-biotinylated probe hybrid is pulled down by streptavidin beads, thereby capturing the target regions of interest. The streptavidin bead-bound library is sequentially washed with buffers to remove non-specifically associated library fragments. Following washes and recovery of captured libraries, samples are enriched for on target fragments and depleted for off-target fragments. Depletion of off-target fragments reduces overall library yield, requiring post-capture library amplification by PCR. The final amplified library is enriched for regions of interest. The hybrid captured library quality and quantity is evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively. Additionally, the enrichment efficiency is evaluated using an iSeq Sequencing run and calculation of percent of reads within target enrichment panel. Measuring percent on-target is a good first approximation of target enrichment efficiency because the reads aligning to the target enrichment (bait) region indicate efficient hybridization and subsequent capture.

[0245] Target enriched libraries that pass quality control checks move forward to NovaSeq sequencing. Captured libraries with non-overlapping indices from library construction are pooled to multiplex for sequencing. Sequencing is completed on the NovaSeq 6000instrument using paired end 150x150 base sequencing with a 10% PhiX spike-in. Sequencing data generated is then demultiplexed utilizing the assigned index, aligned to the human genome and trimmed to enrich for insert sample data only. This cleaned-up data is then processed through a quality pipeline to collapse duplicate reads and evaluate the sequencing data generated. Once the data is collapsed, the data is processed through a proprietary biomarker analysis pipeline to identify differences from the reference alignment (e.g. mutations, chemical modifications, etc). A report is then generated with the specific biomarker analysis per sample that confirms the results of the phase 1 assay or identifies true false positives from the phase 1 assay.Phase 1 Protocol:

[0246] An example protocol of an Allele-specific Real-Time PCR assay is as follows:1. This assay runs DNA samples in triplicate with 2ng input in 5uL for the reference and mutation assays.2. Combine 900nmol / L unspecific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the reference control assay.3. Combine 450nmol / L allele-specific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the mutation assay.4. Mix each reaction 10X and centrifuge to collect volume at the bottom of the well or tube.5. Run the real-time PCR on a calibrated Real-Time PCR system under the following conditions: (1) 95°C for 10 minutes followed by (2) 50 cycles of 90°C for 15 seconds and 60°C for 1 minute with fluorescence detection using FAM / VIC fluorophores.6. Cycle threshold (Ct) values are recorded by the system and exported into an analysis program (e.g. Excel).7. Average the Ct values between sample replicates for the reference and mutation assays.8. Calculate the ACt between the sample average allele-specific Ct minus the sample average unspecific (reference) Ct.9. Positive mutation results are identified by the ACt cut off > 3 cycles and will move forward to phase 2 testing.

[0247] Allele-specific real-time PCR can be performed by combining library DNA with PCR reagents and primers specific for target sequences. The primers are designed to have single-base discrimination between tumor and non-tumor sequences. Perform real-time PCR (or digital PCR) for 30-50 cycles and monitor the output for signal via fluorescence from amplified target DNA or probe sequence. Cycle threshold values (Ct) are recorded and exported for analysis. The delta-Ct between negative control, positive control, and sample are calculated to determine presence or absence of target tumor sequences. Slight modifications of this protocol will allow for end-point PCR detection of RNA or DNA of tumor sequences.Phase 1 detection will be designed to remove 90-95% of non-cancer patient samples from moving forward for further testing.

[0248] ELISA assay detection of target molecules can be performed by coating an immunoassay well with monoclonal antibody designed to specifically detect target molecules, followed by blocking against non-specific binding. Next, target sample is introduced to the well, incubated and washed away. Any bound target can then be bound by a polyclonal antibody specific for the target. Additional secondary antibodies with color or fluorescent tags can be used to detect the presence of target molecules.Interpreting results for phase 1 and phase 2 assays

[0249] Two positive signals from the phase 1 assay and phase 2 assay can be determined as a true positive sample with an 85% probability of being accurate.

[0250] One negative signal from the phase 1 assay can be determined as a true negative sample with a 99% probability of being accurate.

[0251] One positive signal from the phase 1 assay and one negative signal from the phase 2 assay can be determined as an indeterminate sample with a 97% probability of a false positive in phase 1 assay.Example 2: Two-tier Analysis Achieves Improved Performance in Comparison to Single Tier Analysis

[0252] Samples were obtained from patients of a patient population with an assumed 1.3% cancer prevalence. In total, 1046 samples obtained from the patients underwent either a single tier analysis or a two-tier analysis. The performance metrics (as measured by specificity, positive predictive value (PPV), and negative predictive value (NPV)) of each of the methodologies were determined.

[0253] Reference is now made to FIG. 7, which depicts performance of a single tier and two-tier analysis of a population involving 1046 samples. The Tier 1 analysis focused on analyzing signal from a subset of the 4059 CGIs shown in Tables 2 and 3. In particular, 130 regions were analyzed to estimate tumor content according to methylation statuses of the regions, and estimated tumor content was used to distinguish patients that were negative or not negative for cancer. Logistic regression was performed to assess performance at 90% specificity (e.g., true negative rate reported as a proportion of correctly identified negatives). Performance was estimated to be about 63% sensitivity. For the single tier analysis (including only the Tier 1 analysis), it achieved a PPV (defined as number of true positives divided by the sum of true positives and false positives) of 0.0761 and a NPV (defined as true negative rate divided by the sum of true negatives and false negatives) of 0.9946. Thus, thesingle tier analysis was capable of successfully screening out a large proportion of samples that were negative for cancer. However, based on the low PPV, it had room for improvement in identifying samples that were true positives. The single tier analysis (including only a Tier 2 analysis) was additionally performed. Specifically, for each sample, signal of the 4059 CGIs was analyzed using a machine learning algorithm to distinguish samples having a cancer signal from samples not having a cancer signal. The single tier (Tier 2 analysis) achieved a PPV of 0.1858 and a NPV of 0.9969. Thus, the more costly Tier 2 analysis achieved a higher PPV in comparison to the less costly Tier 1 analysis without sacrificing the NPV metric.

[0254] Referring to the two-tier analysis, it involved performing the Tier 1 analysis (analyzing subset of top features) and samples deemed to be negative for cancer were screened out. An additional Tier 2 analysis was then performed. Specifically, for each sample, signal of the 4059 CGIs were analyzed using a machine learning algorithm to distinguish samples having a cancer signal from samples not having a cancer signal. Here, the Tier 2 analysis achieved a high specificity of 96%. For the two-tier analysis (including both the Tier 1 and Tier 2 analyses), the methodology achieved a PPV (defined as number of true positives divided by the sum of true positives and false positives) of 0.2421 and a NPV (defined as true negative rate divided by the sum of true negatives and false negatives) of 0.9942. Here, the two-tier analysis exhibited a significant improvement in comparison to the single-tier analysis. Specifically, the two-tier analysis achieved a higher specificity (e.g., 96% versus 90%). Furthermore, the two-tier analysis exhibited an improved PPV (0.2421 versus 0.0761) without adversely impacting the NPV (0.9942 versus 0.9946).Example 3: Example Samples and Assays for Conducting an Intra-Individual Analysis

[0255] Blood samples are obtained from individuals. FIG. 8 shows an example sample from which target nucleic acids and reference nucleic acids are obtained. Shown on the left in FIG. 8 is a tube of blood obtained from an individual, the tube including diluted peripheral blood of the individual and separation medium. The tube undergoes centrifugation to separate different components of the diluted peripheral blood. For example, at a speed of 2200 rpm, the diluted peripheral blood is fractionated into plasma (including platelets, cytokines, hormones, and electrolytes), peripheral blood mononuclear cells (PBMCs), the separation medium, and polymorphonuclear cells. Here, target nucleic acids in the form ofcell free DNA is found in the plasma whereas reference nucleic acids in the form of cellular genomic DNA is found in PBMCs.

[0256] Examples of an assay for generating sequence information from the target nucleic acids and the reference nucleic acids include but are not limited to Allele-specific PCR assays, Next Generation Sequencing assays, such as target enrichment technologies, targeted amplicon sequencing technologies, and whole genome sequencing.

[0257] An example protocol of an Allele-specific Real-Time PCR assay is as follows:1. This assay runs all cfDNA samples in triplicate with 2ng input in 5uL for the reference and hypermethylation assays.2. Combine 900nmol / L unspecific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the reference control assay.3. Combine 450nmol / L allele-specific primer(s), lOOnmol / L target probe(s), 2X polymerase enzyme(s), 2X dNTPs, 2X passive reference dyes, lOuL water and 2ng sample DNA at a prespecified reaction volume as the mutation assay.4. Mix each reaction 10X and centrifuge to collect volume at the bottom of the well or tube.5. Run the real-time PCR on a calibrated Real-Time PCR system under the following conditions: (1) 95°C for 10 minutes followed by (2) 50 cycles of 90°C for 15 seconds and 60°C for 1 minute with fluorescence detection using FAM / VIC fluorophores.6. Cycle threshold (Ct) values are recorded by the system and exported into an analysis program (e.g. Excel).7. Average the Ct values between sample replicates for the reference and mutation assays.8. Calculate the DCt between the sample average allele-specific Ct minus the sample average unspecific (reference) Ct.9. Positive hypermethylation results are identified by the DCt cut off > 3 cycles and will be compared to the patients individual PBMC natural signal.

[0258] An example protocol of an Allele-specific Real-Time PCR assay is as follows: Allele-specific real-time PCR can be performed by combining library from cfDNA with PCR reagents and primers specific for target sequences. The primers are designed to have singlebase discrimination between tumor and non-tumor sequences. Perform real-time PCR (or digital PCR) for 30-50 cycles and monitor the output for signal via fluorescence from amplified target DNA or probe sequence. Cycle threshold values (Ct) are recorded and exported for analysis. The delta-Ct between negative control, positive control, and sample are calculated to determine presence or absence or absence of target tumor sequences. Slightmodifications of this protocol will allow for end-point PCR detection of RNA or DNA of tumor sequences.

[0259] An example protocol of a next generation sequencing (NGS) Target Enrichment assay is as follows: The target specimen for library construction is dsDNA isolated from PBMCs. The dsDNA is first mechanically sheared by the Covaris instrument utilizing adaptive focused acoustics to a target insert size of 200 base pairs. Post-shearing, a solidphase reversible immobilization (SPRI) selection is done to remove smaller DNA fragments remaining in solution. The fragmented DNA is then end-repaired and A-tailed (ERAT) to produce 5 ’-phosphorylated, 3’-dA-tailed dsDNA fragments. After ERAT, dsDNA unique dual index adapters with 3’-dTMP overhangs are then ligated to 3 ’-d A-tailed dsDNA fragments. Indices allow for sample multiplex for the downstream assay. Post-ligation, a solid-phase reversible immobilization (SPRI) selection is done to remove unwanted DNA fragments, excess adapters and molecules. PCR amplification is performed with a high- fidelity, low-bias polymerase at 10 cycles. Post-PCR, a SPRI selection is done to remove unwanted DNA fragments, excess primers, excess adapters and excess molecules. After library construction, the library quality and quantity are evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively.

[0260] Libraries that pass quality control checks move forward to target enrichment through hybridization capture. Target enrichment by hybridization capture is defined as a positive selection strategy to enrich low abundance regions of interest from NGS libraries, allowing for more accurate sequencing analysis of these target regions. Indexed libraries are multiplexed and hybridized to a custom, sequence specific, biotinylated probeset. The vast excess of probes drives their hybridization to complementary library fragments. The library fragment-biotinylated probe hybrid is pulled down by streptavidin beads, thereby capturing the target regions of interest. The streptavidin bead-bound library is sequentially washed with buffers to remove non-specifically associated library fragments. Following washes and recovery of captured libraries, samples are enriched for on target fragments and depleted for off-target fragments. Depletion of off-target fragments reduces overall library yield, requiring post-capture library amplification by PCR. The final amplified library is enriched for regions of interest. The hybrid captured library quality and quantity is evaluated using the Agilent TapeStation and Qubit Fluorometer, respectively. Additionally, the enrichment efficiency is evaluated using an iSeq Sequencing run and calculation of percent of reads within target enrichment panel. Measuring percent on-target is a good first approximation of targetenrichment efficiency because the reads aligning to the target enrichment (bait) region indicate efficient hybridization and subsequent capture.

[0261] Target enriched libraries that pass quality control checks move forward to NovaSeq sequencing. Captured libraries with non-overlapping indices from library construction are pooled to multiplex for sequencing. Sequencing is completed on the NovaSeq 6000 instrument using paired end 150x150 base sequencing with a 10% PhiX spike-in. Sequencing data generated is then demultiplexed utilizing the assigned index, aligned to the human genome and trimmed to enrich for insert sample data only. This cleaned-up data is then processed through a quality pipeline to collapse duplicate reads and evaluate the sequencing data generated. Once the data is collapsed, the data is processed through a proprietary analysis pipeline to identify differences from the reference alignment (e.g. mutations, chemical modifications, etc.). A report is then generated with the specific signal informative for determining presence or absence of cancer.Attorney Docket No. ELG-022WOTABLE 1 - List of CGIsReference Pos (hgl9 coordinates)1 chrl3:108518334-1085186332 chr6:137242315-1372454423 chr2:177016416-1770166324 chr5:2738953-27412375 chr4:111553079-1115542106 chrl5:96909815-969100307 chr6:42072032-420727018 chrl0:123922850-1239235429 chrl6:86612188-8661382110 chrl9:47151768-4715312511 chrl:110610265-11061330312 chr5:3594467-360305413 chr9:126773246-12678095314 chr3:138656627-13865910715 chr4:4859632-486019116 chrl0:118895963-11889803717 chr7:103086344-10308684018 chrl9:407011-40951119 chrl0:22764708-2276705020 chrl6:86549069-8655051221 chr9:96713326-9671818622 chr8:139508795-13950977423 chr2:73143055-7314826024 chr8:26721642-2672456625 chr9:129386112-12938923126 chrl2:49483601-49484255 1 chrl6:54325040-5432570328 chr8:72468560-7246956129 chrl8:70533965-7053687130 chr9:98111364-9811236231 chrl:50882997-5088342632 chrl0:88122924-8812736433 chrll:31839363-3183981334 chrl0:101290025-10129033835 chr6:41528266-4152890036 chrl6:51183699-5118876337 chr5:140346105-14034693138 chr9:2382O691-2382213539 chr20:690575-69109940 chrl:177133392-17713384641 chr5:45695394-4569651042 chr2:45395869-4539818643 chr20:48184193-4818483344 chr6:6002471-600512545 chrl4:101192851-101193499chr8:4848968-4852635 chr8:53851701-53854426 chrl2:186863-187610 chr5:54519054-54519628 chr6:108485671-108490539 chr3:157815581-157816095 chrll:626728-628037 chr2:177012371-177012675 chrl7:59531723-59535254 chrl6:55364823-55365483 chr8:99960497-99961438 chr7:42267546-42267823 chrl7:14202632-14203258 chrl0:102891010-102891794 chr5:174158680-174159729 chrl4:33402094-33404079 chr2:177036254-177037213 chrl0:106399567-106402812 chr6:166579973-166583423 chrll:123066517-123066986 chrll:44327240-44327932 chrl4:95237622-95238211 chr9:102590742-102591303 chrl5:76630029-76630970 chr4:24801109-24801902 chr8:97169731-97170432 chr3:6902823-6903516 chr22:48884884-48887043 chrl5:45408573-45409528 chr9:100610696-100611517 chr4: 174448333-174448845 chrl6:20084707-20085305 chr4: 174439812-174440249 chr6:10381558-10382354 chrl5:35046443-35047480 chrl0:119494493-119494991 chr5:72676120-72678421 chrll:44325657-44326517 chrl7:46670522-46671458 chrl4:92789494-92790712 chr4: 174459200-174460054 chr2:80549578-80549798 chr7: 153748407-153750444 chr6:1389139-1391393 chrl6:49314037-49316543 chr2:105459127-105461770 chr21:38079941-38081833chr4:174427891-174428192 chrl4:60973772-60974123 chr8:99985733-99986983 chr2:63281034-63281347 chrl2:101109863-101111622 chrl:119549144-119551320 chr5:38257825-38259136 chr5:54522302-54523533 chrl:165324191-165326328 chrl5:33602816-33604003 chrl0:118030732-118034230 chr2:45240372-45241579 chr4:174430386-174430861 chr6:50810642-50810994 chr5:122430676-122431443 chrl0:109674196-109674964 chr8:97172634-97173880 chr8:11536767-11538961 chr5:180486154-180486892 chr2:38301276-38304518 chrl0:1778784-1780018 chrl2:54424610-54425173 chrl7:46669434-46669811 chrll:8190226-8190671 chr8:25900562-25905842 chrl2:81102034-81102716 chr7:27199661-27200960 chrl0:119311204-119312104 chrl2:130387609-130389139 chr7:155258827-155261403 chr6:117591533-117592279 chrl0:111216604-111217083 chrl:29585897-29586598 chr2:144694666-144695180 chrl2:48397889-48398731 chr5:2748368-2757024 chrl2:114845861-114847650 chr2:80529677-80530846 chr5:1874907-1879032 chr6:100905952-100906686 chrl5:96904722-96905050 chr5:134374385-134376751 chr2:66652691-66654218 chrl2:54440642-54441543 chr6:108495654-108495986 chrl7:70112824-70114271 chr3:87841796-87842563chr7:96650221-96651551 chr4:110222970-110224257 chr6:78172231-78174088 chr7:155164557-155167854 chrl2:113900750-113906442 chr9:112081402-112082905 chrl2:114886354-114886579 chr5:3590644-3592000 chr2:119592602-119593845 chr20:21485932-21496714 chrl8:11148307-11149936 chrl7:46824785-46825372 chrl0:100992156-100992687 chrl4:36986362-36990576 chrl8:55094825-55096310 chrl5:96895306-96895729 chrl7:36717727-36718593 chr2:223183013-223185468 chr7:30721372-30722445 chrl:53527572-53528974 chrl8:56939624-56941540 chr5:175085004-175085756 chrl0:50817601-50820356 chrl4:60975732-60978180 chrl5:89920793-89922768 chr9:122131086-122132214 chrl:217311467-217311773 chrl4:38724254-38725537 chrl4:61103978-61104663 chrl8:73167402-73167920 chrl:50880916-50881516 chr2:241758141-241760783 chrll:31825743-31826967 chr7:27260101-27260467 chr20:41817475-41819212 chr3:238391-240140 chr7:121950249-121950927 chr5:72526203-72526497 chrl5:96903311-96903711 chrl0:26504383-26507434 chr6:100915602-100915883 chrl:18962842-18963481 chr3:127794369-127796136 chr7:27203915-27206462 chr8:25899335-25899692 chrl2:114838312-114838889 chr6:38682949-38683265chrll:31841315-31842003 chr4:174451828-174452962 chr9:129372737-129378106 chr2:176964062-176965509 chr2:176931575-176932663 chrl2:114833911-114834210 chrll:79148358-79152200 chr2:177024501-177025692 chr5:172672311-172672971 chr7:27291119-27292197 chrl:180198119-180204975 chrl4:37126786-37128274 chr2:200333687-200334172 chrl4:58331676-58333121 chr3:147131066-147131333 chrl3:109147798-109149019 chrl4:48143433-48145589 chr6 : 100905444-100905697 chrl7:14200579-14200996 chr6:1379693-1380014 chrl:34642382-34643024 chr2:119599059-119599299 chr2:119613031-119615565 chr4:85413997-85414874 chr9:17906419-17907488 chrl2:29302034-29302954 chr20:10200088-10200384 chr8:57358126-57359415 chrl0:63212495-63213009 chr2:176936246-176936809 chrll:20618197-20619920 chrl8:19744936-19752363 chrl4:29234889-29235908 chrl7:46673532-46674181 chr4:144620822-144622218 chrl6:82660651-82661813 chr3:192125821-192127994 chr2:119599458-119600966 chr22:44257942-44258612 chrl9:13616752-13617267 chr3:147138916-147139564 chr9:969529-973276 chrl8:55103154-55108853 chr4:174422024-174422443 chr4:57521621-57522703 chrl5:79724099-79725643 chrl4:37135513-37136348chrl0:23480697-23482455 chr2:45169505-45171884 chrl8:30349690-30352302 chr6:99291327-99291737 chr9:21970913-21971190 chr4:107146-107898 chrl2:117798076-117799448 chr2:219736132-219736592 chrl0:118892161-118892639 chrll:27743472-27744564 chrl2:65218245-65219143 chrl2:75601081-75601752 chr7:54612324-54612558 chr6:100912071-100913337 chrl0:102905714-102906693 chr8:87081653-87082046 chr6:50818180-50818431 chrl:91189139-91189400 chr2:118981769-118982466 chrl0:50602989-50606783 chrl7:59528979-59530266 chr4:147559205-147561901 chrl:4713989-4716555 chrl3:102568425-102569495 chrl6:6068914-6070401 chr22:29709281-29712013 chrl0:100993820-100994188 chr6:391188-393790 chr2:176977284-176977540 chr4:4868440-4869173 chr6:137809342-137810204 chrl2:54321301-54321721 chr2:105468851-105473488 chr8:55366180-55367628 chrl2:72665683-72667551 chr4:54966163-54968063 chr5:134366913-134367438 chrl:226075150-226075680 chr20:17206528-17206952 chr4:172733734-172735118 chrl8:55019707-55021605 chr2:162279835-162280709 chr6:1381743-1385211 chr7:103968783-103969959 chr6:150358872-150359394 chr2:119914126-119916663 chr7:27278945-27279469chrl2:114851957-114852360 chrl6:24267040-24267527 chr6:7229877-7230865 chr2:45227644-45228783 chr4:174450046-174451469 chr4:154712073-154712706 chr3:22413492-22414365 chr20:21694472-21695344 chr6:1378445-1379318 chr8:70981873-70984888 chrl2:53107912-53108471 chrl0:102996034-102996646 chr3:157821232-157821604 chr4:111554965-111555504 chrl3:58206526-58208930 chrl0:22634000-22634862 chr9:22005887-22006229 chr5:159399004-159399928 chr2:31805293-31806403 chr6:100903491-100903713 chr5:77268350-77268787 chrl4:85997468-85998637 chr5:92923487-92924497 chrll:64480199-64481344 chrl3:28366549-28368505 chr5:77805753-77806313 chr9:79633326-79636030 chr4:93226348-93227007 chr2:223170486-223171140 chrl:91172102-91172771 chrl:1181756-1182470 chr8:65281903-65283043 chrl0:94825546-94826320 chr6:108491033-108491410 chr21:38076762-38077685 chrl:91183240-91184540 chr3:147136903-147137328 chrl5:96911511-96911808 chrl4:57274607-57276840 chrl3:112726281-112728419 chr2:171672310-171675447 chr8:11559596-11562956 chrl0:48438411-48439320 chrl8:59000683-59001692 chrl5:91642908-91643702 chr5:3592391-3592644 chrl9:56988313-56989741chr6:26614013-26614851 chrll:27742059-27742273 chr3:147113608-147114479 chrl4:57264638-57265561 chr7:155302253-155303158 chrll:31848487-31848776 chrl6:54970301-54972846 chrl9:30715549-30715753 chr9:96710811-96711717 chrl8:77557780-77558948 chr20:21686199-21687689 chrll:31847132-31847958 chrl6:86530747-86532994 chrl:203044722-203045390 chrl5:53096014-53096482 chr7:97361132-97363018 chrl4:29236835-29237832 chrl3:79182859-79183880 chrll:69517840-69519929 chrl:231296559-231297345 chrl9:8675333-8675699 chrl:63795363-63796140 chr4:90228714-90229010 chr3:62362610-62363082 chrl9:5827754-5828405 chrl0:125732220-125732843 chr9:136293566-136294160 chrl:63782394-63790471 chr4:4867386-4867673 chr9:133534534-133542394 chrl5:100913438-100914022 chrl0:101279941-101280382 chrl3:53419897-53422872 chrl:77747314-77748224 chrl4:36974548-36975425 chrl2:57618769-57619402 chr7:49813008-49815752 chr4:188916605-188916876 chrll:31831620-31839038 chr8:132052203-132054749 chr2:237071794-237078762 chr20:39994545-39995810 chrll:132812662-132813075 chr5:170735169-170739863 chrl:221051966-221053673 chr5:72529099-72529976 chrl4:36973169-36973740chr4:158141404-158141836 chrl4:103655241-103655928 chrl:65731411-65731849 chrl:38218190-38218977 chr3:128719865-128721245 chrl5:33009530-33011696 chr2:162275161-162275596 chr7:155241323-155243757 chrl9:46001830-46002686 chr6:137814355-137815202 chr7:70596228-70598382 chrl5:96959341-96960531 chrl6:66612749-66613412 chr6:110299365-110301267 chrl5:27215951-27216856 chrll:88241710-88242562 chr2:124782252-124783255 chrl7:70111979-70112308 chr2:63283936-63284147 chrl7:46800945-46801288 chr6:1393049-1394170 chr3:137489594-137491004 chrl5:60296135-60298520 chrl2:106979429-106981086 chrl2:54360374-54360660 chrl4:36991594-36992488 chr4:156129168-156130209 chr4:54975387-54976202 chr3:137482964-137484454 chrl0:118893527-118894432 chrl8:76737005-76741244 chrl0:110671724-110672326 chr5:71014917-71015715 chr6:50787286-50788091 chrl9:3868586-3869217 chr4:5894071-5895116 chrll:131780328-131781532 chr6:101846766-101847135 chrll:71952112-71952528 chr5:172663616-172664584 chr9:23822412-23822667 chr4:5891981-5892365 chrl:217310749-217311178 chrl0:108923780-108924805 chr6:100038655-100039477 chr7:121945345-121946235 chr3:147126988-147128999chr7:121956543-121957341 chr4:156680095-156681386 chr4:85404986-85405252 chrl:221064889-221065600 chrl7:73749618-73750178 chr8:55370170-55372525 chr6:70992040-70992912 chrl6:55513220-55513526 chr6 : 106433984-106434459 chrl4:29254365-29255069 chr6:33655966-33656238 chr9:19788215-19789288 chrll:115630398-115631117 chrl:34628783-34630976 chrl4:101923575-101925995 chrl7:72855621-72858012 chr2:223162946-223163912 chr4:85417659-85420799 chrl:156390403-156391581 chr3:147130342-147130577 chr2:119602616-119604486 chr9:120175253-120177496 chr4:174443365-174443948 chr5:145724294-145724551 chrll:32454874-32457311 chr2:176949511-176949795 chrl:18436551-18437673 chr3:26665950-26666164 chr3:170303044-170303249 chr2:223176493-223177515 chr2:182321761-182323029 chrl8:44789742-44790678 chrl7:46796234-46797292 chrl8:44772992-44775577 chr8:101117922-101118693 chr7:27134097-27134303 chrl0:102507482-102509646 chrl9:39754973-39756540 chr7:26415746-26416891 chrl4:37116188-37117628 chr4:174421347-174421559 chr6:85472702-85474132 chr20:22557517-22559240 chr6:117198089-117198705 chrl0:71331926-71333392 chrl9:36334994-36335321 chr4:46995128-46995872chr9:135455164-135458586 chr8:65290108-65290946 chrl0:94828102-94829040 chrl:116380359-116382364 chrl5:47476369-47477499 chr3:147115764-147116421 chrl7:59485573-59485780 chrl0:23983366-23984978 chr2:176949993-176950336 chr9:137967110-137967727 chr2:176957054-176958279 chrll:119293320-119293943 chrll:132813562-132814395 chr2:237068071-237068834 chrl0:27547668-27548402 chr4:4866438-4866813 chr21:19617098-19617874 chrl:91185156-91185577 chrl9:15292399-15292632 chrl:145075483-145075845 chr2:19560963-19561650 chrl4:57260878-57262123 chr8:55378928-55380186 chr6:99290279-99290771 chrl9:13124959-13125259 chrl5:27112030-27113479 chr8:145925410-145926101 chrll:124629723-124629926 chr4:109093038-109094546 chr3:62356773-62357315 chrl4:37131181-37132785 chrl0:124905634-124906161 chr7:35296921-35298218 chrl9:36248979-36249307 chrl2:15475318-15475901 chr5:87985470-87985810 chrl2:54423427-54423712 chr7:96653467-96654199 chr2:45155195-45157049 chrl5:96896928-96897301 chrl2:58004982-58005351 chr2:176933131-176933449 chr2:176962179-176962487 chr20:25063838-25065525 chrl2:5153012-5154346 chr3:154146347-154146965 chrl:165323486-165323811chr21:38065179-38066185 chrl0:119000435-119001530 chrl2:45444202-45445386 chr4:158143296-158144053 chr5:76932317-76933523 chr5:172659049-172660277 chr2:223168653-223169008 chrl:248020330-248021252 chrl8:904578-909574 chrl2:127940451-127940907 chr9:135461934-135462909 chrl7:48041282-48043064 chr4:94755786-94756310 chrl0:130338695-130338994 chr2:119616133-119616826 chr2:177042751-177043444 ch r2 : 105478600-105479188 chr5:172670829-172671824 chr2:176952695-176953297 chrl3:28549839-28550246 chrl3:112720564-112723582 chr6:100895773-100896062 chr7:136553854-136556194 chr6:127441553-127441760 chrl:119526782-119527192 chrl2:49484920-49485178 chr9:23850910-23851522 chr2:220299483-220300243 chr5:1881924-1887743 chr8:57360585-57360815 chrl8:74961556-74963822 chr5:172660720-172661133 chrl7:75277317-75278172 chrl0:99789614-99791320 chr2:176944087-176948446 chr4:154709512-154710827 chr5:140798757-140799359 chr3:44063314-44063837 chrl5:79574830-79575211 chr2:223161531-223161919 chr6:134210639-134211218 chrl0:102899177-102899489 chrl3:79181944-79182222 chr7:71800757-71802768 chr3:186078710-186080111 chrl:24229115-24229537 chrl6:48844551-48845264chr7:113724924-113727795 chr22:44726724-44727590 chr4:15779998-15780729 chr4:41869174-41869459 chrl:38941919-38942404 chr2:176971706-176972305 chr2:119607378-119607910 chr5:76934581-76935296 chrl2:103696090-103696418 chr5:63255044-63255407 chrl:221067447-221068185 chr2:119611296-119611881 chrl0:124907283-124911035 chrl2:114878143-114879155 chrl2:49371690-49375550 chrl7:36719544-36719938 chrl7:46696553-46696926 chr3:147142181-147142391 chr8:9762661-9764748 chrl4:74706188-74708192 chr3:12837992-12838359 chr20:37352130-37357372 chrl0:8077829-8078378 chr4:4864456-4864834 chr4:13524062-13526083 chrl:66258440-66258918 chrll:17740789-17743779 chrl2:106975195-106975714 chr9:91792662-91793611 chrl:149333785-149334111 chr3:170303532-170303768 chr5:72594147-72595808 chr5:145725286-145725852 chrl0:23462224-23463889 chr20:21689758-21690048 chrl5:53080458-53083699 chr2:154727906-154728271 chr5: 170743178-170744107 chrl0:102899822-102900263 chr5:134368578-134370466 chr2:66808568-66809404 chr7:96651963-96652246 chrl:91190489-91192804 chrl7:75368688-75370506 chr4:185939222-185942747 chr7:43152020-43153340 chrl3:84453664-84453897chr2:176956504-176956707 chr7:87563342-87564571 chr20:17208550-17208756 chr22:19746924-19747141 chr2:223159725-223160487 chrl2:131200509-131200726 chrl8:44336183-44337110 chr2:63285949-63287097 chr4:13526553-13526770 chrl5:89949373-89951130 chrl9:55815940-55816277 chrl7:50235175-50236466 chrl9:58545115-58545897 chrl2:113592203-113592620 chrl2:115109503-115110061 chr4:164264821-164265772 chrl:2772126-2772665 chr3:71834068-71834653 chrl2:5018585-5021171 chrl5:74419870-74423044 chr3:147108511-147111703 chr5:88185224-88185589 chrl2:54354529-54355491 chrl0:101290625-101291178 chr8:11557852-11558252 chr8:105478672-105479340 chrll:20181200-20182325 chrl9:54483021-54483572 chrl3:112707804-112708696 chrl6:22824616-22826459 chr4:66536065-66536674 chr4:154713537-154714240 chr7:12151220-12151559 chrl2:119212110-119212393 chrl7:14201726-14202052 chr20:21376358-21378245 chrl3:36045931-36046143 chrl5:60287107-60287663 chr9:100613938-100614622 chrl0:102475276-102475579 chr7:121940006-121940648 chr5:37834671-37835128 chrl:197887088-197887791 chrl2:99139386-99139769 chr6:1619093-1621094 chrl2:113917394-113918107 chrl4:24044886-24046760chr5:77253832-77254049 chr4:85403830-85404524 chr6:166666837-166667541 chrl8:77547965-77549038 chr2:219848919-219850541 chrl7:7832532-7833164 chr5:134363092-134365146 chrl0:103043990-103044480 chr8:97171805-97172022 chr20:57089460-57090237 chrl2:114840853-114841063 chr4:66535193-66535620 chr8:85096759-85097247 chr6:10881846-10882051 chrl3:28498226-28499046 chrl:161695637-161697298 chrll:2890388-2891337 chrl7:5000369-5001205 chrl3:27334226-27335205 chrl0:22623350-22625875 chr2:157185557-157186355 chr7:20370003-20371504 chr4:961347-962155 chrl2:49485766-49485977 chr3:62356119-62356378 chrll:14995128-14995908 chrl2:53359192-53359507 chrl6:51168266-51169110 chrl4:57278709-57279116 chr6:37616722-37617179 chrl8:11750953-11752756 chrl9:45260352-45261809 chrl:119531991-119532196 chrl9:36523391-36523887 chrl2:52652018-52652743 chr8:49468683-49468959 chr8:9760750-9761643 chr7:19146923-19147308 chrl3:32889533-32889900 chr5:140797162-140797701 chr21:42218489-42219222 chrl9:54411376-54411968 chr3:62354291-62355012 chrl2:113590806-113591304 chrl:225865068-225865328 chr7:130790358-130792773 chr!5:53076187-53077926chrl:214158726-214159080 chrl2:3308812-3310270 chrl :39044059-39044561 chrl0:119312766-119313563 chrl2:65514878-65515863 chrl2:54366815-54369103 chrl2:114885105-114885418 chrl6:2228190-2230946 chrll:68622722-68623252 chr2:25499763-25500429 chr5:172661486-172662228 chrl7:46691520-46692097 chrl2:75602991-75603344 chr2:80531367-80531719 chr5:158478378-158478630 chr2:177017266-177017489 chr2:63282514-63283122 chr7:155595692-155599414 chr5:172665306-172666072 chrl2:114843022-114843610 chrl3:112758598-112760491 chr4:4858389-4858893 chrl6:55365814-55366022 chr9:96108466-96108992 chrl2:3475010-3475654 chr9:86152353-86153777 chr6:10384965-10385492 chr22:31500396-31501239 chr5:179228283-179229003 chr6:137816474-137817223 chr2:106681982-106682403 chrl4:95239375-95239679 chr7:154001964-154002281 chrl:1476093-1476669 chrl5:89904822-89906050 chrll:89224416-89224718 chr9:100615234-100617510 chr3:172165372-172166738 chrl:202678881-202679769 chrl4:37053134-37053690 chr4:41875445-41875794 chr2:162273294-162273725 chrl:181287300-181287873 chrl3:79181327-79181614 chr8:145103285-145108027 chr22:42305617-42307254 chr8:102505512-102506430chrl7:74533281-74534566 chrl:214156000-214156851 chr20:2780978-2781497 chr4:4861227-4862241 chrl9:13215244-13215543 chr7:121943867-121944538 chrl7:71948478-71949255 chr2:127413696-127414171 chrl:113286332-113287172 chrl:47009575-47010132 chrl6:62069121-62070634 chrl6:3013651-3015131 chrl8:76732970-76734765 chr4:155664819-155665833 chr6:72298274-72298528 chrl5:89147660-89149198 chrl7:33775294-33775794 chrl8:44337510-44338100 chrl0:8076002-8077261 chrl3:112717125-112717421 chrl5:89914363-89915061 chrl:228785986-228786204 chrl:156358050-156358252 chr7:751712-752150 chr3:137489051-137489409 chrl7:7905927-7907445 chrl8:35144907-35147628 chr3:9177691-9178189 chr6:10390888-10391098 chrl4:37052537-37052838 chrl:47909712-47911020 chrl3:93879245-93880877 chrl:50893468-50893745 chr7:27282086-27283136 chr4:147558231-147558583 chrl9:13124569-13124788 chrl7:46619087-46619314 chr3:44596535-44597018 chrl4:24803678-24804353 chr2:3286324-3286530 chrl2:14134626-14135242 chrl2:114881649-114881937 chr20:22548967-22549720 chr8:37822486-37824008 chrl3:100641334-100642188 chr4:206377-206892 chr3:11034446-11035384chr7:152622343-152623305 chrl0:22629360-22630328 chr4:140201064-140201449 chrl9:46318490-46319266 chr3:121902742-121903645 chr9:77112712-77113583 chr2:114256775-114258043 chrl0:15761423-15762101 chrl:115880167-115881332 chr6:50791110-50791573 chr6:55039170-55039392 chr2:176980765-176981423 chr8:86350765-86351196 chr8:24812946-24814299 chr7:19184818-19185033 chr5:76936126-76936984 chr5:87980878-87981272 chr9:77111778-77112042 chrll:20622720-20623399 chrl:50882433-50882660 chrl7:35291899-35300875 chrl7:46675044-46675589 chr20:5296266-5297798 chr7:156871054-156871297 chr4:681313-681514 chr2:177039551-177039951 chrl7:46695325-46695553 chrl:41283840-41284591 chr9:16726859-16727273 chrl:65991001-65991811 chrl:181452706-181453073 chr8:120428398-120429178 chr3:32863174-32863415 chr4:134069162-134070442 chrl2:123754049-123754373 chr5:63256548-63257886 chr5:1879689-1879928 chrl0:118899247-118900329 chr20:2731063-2731395 chr5:134385967-134386370 chr2:177014948-177015214 chrl:67218079-67218293 chrll:65408344-65408631 chr7:156801418-156801632 chrl8:54788959-54789194 chr2:220173870-220174283 chr2:220173021-220173271chrl2:113908887-113910681 chr6:100897080-100897621 chrl:155290606-155291001 chr2:130763483-130763764 chrl2:129337870-129338653 chr21:34395128-34400245 chrl2:52115410-52115679 chr3:126113547-126113967 chrl6:3220438-3221356 chrl:119543056-119543454 chrl4:62279476-62280019 chrll:636906-640628 chrl0:102893660-102895059 chr3:3840513-3842772 chrl:119529819-119530712 chr9:32782936-32783625 chrl9:1064897-1065191 chr5:54527319-54527760 chr7:156795355-156799394 chrl:155147185-155147444 chr9:37002489-37002957 chrll:69831571-69832484 chr2:128421719-128422182 chr22:38476836-38478839 chrl9:54412710-54413087 chr9:123656750-123656972 chr7:129422997-129423355 chrl9:36336275-36337138 chr2:50574045-50574817 chrl0:102975969-102978096 chr6:5996185-5996486 chr3:26664104-26664796 chr7:155170623-155170939 chr8:65286067-65286659 chrl4:37125219-37125661 chrll:65816404-65816665 chr6:41908745-41909711 chrl7:46620367-46621373 chr2:142887724-142888553 chrl:221050448-221050864 chrl2:106974412-106974951 chrl4:57278068-57278287 chrl:67773329-67773767 chrl7:40936445-40936668 chr20:2729997-2730797 chrl2:113013099-113013529 chr7:155244046-155244357chrl:214153214-214153668 chrl:156863415-156863711 chrl:114695136-114696672 chrl4:85996494-85996958 chr7:100823307-100823701 chr20:52789252-52790986 chr5:178421225-178422337 chrll:36397926-36399398 chrl3:36052553-36053119 chrl4:57283967-57284558 chr4:25090106-25090510 chr2:5831187-5831413 chr6:117869097-117869530 chrl9:58094739-58095764 chr4:85422929-85423190 chrl3:100547172-100547431 chr8:68864584-68864946 chrl6:49311413-49312308 chr7:19184221-19184686 chr2:19562749-19562965 chrl9:54481412-54481955 chrl0:124901907-124902617 chr3:62357639-62359774 chrll:31827696-31827921 chrl7:43037166-43037740 chr7:37955622-37956555 chr6:106429111-106429772 chr6:50682334-50683214 chr5:76923887-76924502 chr6:168841818-168843100 chr7:19145872-19146256 chr20:32856659-32857248 chrl7:79859808-79860963 chr7:95225503-95226194 chrl4:105167663-105168129 chrl7:14248391-14248721 chrl6:84002269-84002860 chr9 : 104499849-104501076 chrl7:46604362-46604881 chr2:87015974-87018182 chrl4:36990873-36991209 chr5:52777788-52777996 chrl9:35633847-35634629 chrl:221055492-221055800 chrl:146551476-146551764 chrl3:100642774-100643094 chrl4:85999532-86000478chrl3:36049570-36050159 chr2:119606038-119606313 chrll:123065426-123066184 chr3:172167526-172167866 chr4:41882450-41882964 chr8:142528185-142529029 chr9:79637814-79638169 chr3:19189688-19190100 chr4:122301567-122302290 chrl0:130339526-130339777 chr9:35846310-35846638 chrl5:53097561-53098476 chr2:157184389-157184632 chr5:145718289-145720095 chrll:105481126-105481422 chr5:170741603-170742751 chr3:62355315-62355534 chrl:38219702-38220012 chr4:41881177-41881418 chrl3:112715359-112716234 chrl7:1880789-1881116 chrl8:56887091-56887665 chr6:10390038-10390565 chrll:69516931-69517218 chrl9:39737689-39739288 chr3:157812053-157812764 chrl4:37049333-37051726 chr7:156409023-156409294 chrll:46366876-46367101 chr5:50685453-50686148 chr4:41883492-41884570 chrl3:112709884-112712665 chr22:44287497-44288061 chr22:46440393-46441019 chr8:23562475-23565175 chr2:207506774-207507422 chr4:169799086-169799625 chr3:133393118-133393657 chr8:41424341-41425300 chr4:100870377-100871994 chr4:107956555-107957453 chrl7:79314962-79320653 chr2:30453566-30455655 chrl:18956895-18959829 chrl2:41086522-41087102 chr22:42685894-42686095 chr6:100914946-100915245986 chrl:46951168-46951792987 chr4:41749184-41749811988 chrll:128419198-128419513989 chr2:171671598-171671804990 chrl:170630456-170630851991 chr20:44657463-44659243992 chr9:139096665-139096993993 chr7:155174128-155175248994 chrl4:36993488-36994488995 chr3:138654837-138655363996 chr4:5709985-5710495997 chrl5:23157794-23158624998 chr20:9496471-9496893999 chr4:174437914-1744383461000 ch r5 : 140305712-1403071931001 chrl5:79576059-795762701002 chrl4:38678245-386809371003 chrl0:102473206-1024740261004 chrl7:59486727-594871321005 chr3:64253533-642538191006 chrl0:102484200-1024844761007 chr7:27198182-271985141008 chr2:97192977-971933831009 chr9:77113709-771139271010 chr6:154360586-1543610081011 chrll:44324875-443250871012 chr2:182521221-1825219271013 chr7 : 124404700-1244061891014 chr2:132182327-1321831011015 chr7:101005899-1010074431016 ch r7 : 149744402-1497464691017 chr8:50822270-508228601018 chr7:27227520-272290431019 chr6:134212690-1342130981020 chrl3:36044844-360454811021 chrll:132934059-1329342911022 chrl6:51189800-511902601023 chrl:155145342-1551459381024 chr4:682724-6830791025 chr5:92939795-929402161026 chrl0:134597357-1346026491027 chrl:200009807-2000100361028 chrl9:12666243-126666821029 chr9:97401286-974020671030 chr2:107103833-1071040531031 chrl5:89910521-899121771032 chr5:140789094-1407897621033 chr2:114033359-1140336171034 chrl7:12568667-125693351035 chrll:68622108-686223391036 chrl : 160340604-1603408431037 chr7:103085710-1030861321038 chrl5:76628998-766292071039 chr20:10198135-101989841040 chr20:44660342-446609481041 chrl7:35290403-352906631042 chrl7:933026-9332361043 chr4:128544031-1285449031044 chrl:50881884-508821031045 chrl0:125425495-1254266421046 chrl7:46801784-468020711047 chrl:25255527-252590051048 chr3:32861141-328614291049 chrl7:70116274-701199981050 chrl0:75407413-754077061051 chr2:467849-4686591052 chrll:132952538-1329533071053 chr3:6904133-69046411054 chrl0:120353692-1203558211055 chr7:20830567-208308171056 chrll:71950815-719514081057 chrl4:95240083-952403411058 chrl9:5829048-58294741059 chr20:9495253-94955971060 chr9:112083333-1120835491061 chrl5:96873408-968777211062 chrl6:67208067-672086781063 chrl:175568376-1755688081064 chr6:5999149-59997871065 chr3:129693127-1296948411066 chr6:10383525-103841141067 chrll:636435-6366681068 chrl:181451311-1814520491069 chr9:135464586-1354662401070 chrl5:60289325-602895331071 chrl6:49309123-493093531072 chrl:243646394-2436468881073 chrl2:54071053-540712651074 chrl:91176404-911767011075 chrS : 140864527-1408647481076 chr4:47034427-470349401077 chrl0:102489343-1024910111078 chrl0:102419147-1024196681079 chrl2:81471569-814721191080 chr6:50813314-508136991081 chr5:158526133-1585264311082 chrl:119543821-1195443391083 chr5:77140542-771409141084 chr8:23567180-235676781085 chrl:41831976-418325421086 chr2:139537692-1395386501087 chr7:100075303-1000755511088 chr2:176969217-1769698951089 chr7:27284639-272862371090 chr5:31193952-311944191091 chr6:37616393-376166211092 chrl9:1748167-17502431093 chrl0:101281181-1012821161094 chr21:31311386-313121061095 chr2:176973427-1769737181096 chrl5:96900142-969006441097 chr7:158936507-1589384921098 chr3:63263989-632642051099 chrl6:71459781-714603381100 chr7:155601175-1556032351101 chrl2:54447744-544480911102 chrl2:53491572-534919551103 chrl0:16561604-165638221104 chrll:133994709-1339950901105 chr2:137522460-1375236961106 chrl7:12877270-128777731107 chr8:98289604-982904041108 chr4:185937242-1859377501109 chr3:185911344-1859122281110 chrl2:54378696-543801021111 chrl:221060850-2210610711112 chrl2:63543636-635449671113 chr6:6006689-60070431114 chrl9:51169659-511720231115 chrl:1474962-14752201116 chrl4:54418677-544188811117 chr6:108497595-1084979961118 chrl7:37764092-377643041119 chr4:109092578-1090928391120 chrl:91182097-911823641121 chrl3:112760865-1127611131122 chrl2:122018170-1220184571123 chr7:142494563-1424952481124 chrl3:58203586-582043221125 chrl:92945907-929526091126 chrl2:106977388-1069777131127 chr5:76925445-769268751128 chrl6:3190765-31913891129 chrl:12123488-121241481130 chrl7:48545570-485469001131 chrl2:113916433-1139167171132 chr4:41747508-417479441133 chrl9:46916587-469168621134 chrl5:49254984-492555641135 chrl9:8674332-86747641136 chr2:223167205-2231675601137 chrl7:1173535-11747331138 chr3:75955759-759563081139 chr5:115697134-1156975891140 chr8:21644908-216478451141 chr5:59189046-591898941142 chrl2:54338761-543391681143 chrl6:31053479-310538001144 chrl:50892437-508932431145 chrl7:40935964-409361801146 chrl9:44203558-442039871147 chr4:81109887-811104601148 chrl:2979275-29807581149 chrl6:49872449-498729261150 chrl:200008392-2000090471151 chrl6:49316997-493172631152 ch r2 : 114034594-1140360411153 chr2:105480197-1054807601154 chrl8:44777632-447780841155 chrl9:13213450-132138211156 chrl7:6616422-66174711157 chrl4:36977518-369779961158 chrl:214160798-2141610341159 chrl:91182509-911828571160 chrl0:130508443-1305086581161 chr2:154728944-1547293281162 chrl5:89952271-899530611163 chrl8:55102427-551027081164 chr22:31198491-311990331165 chrl0:50821487-508216881166 chr7 : 100076454-1000767851167 chrl8:13641584-136424151168 chrl8:13868532-138690261169 chr6:168841438-1688416991170 chrl:61515875-615168311171 chr7:32110063-321109101172 chr7:56355508-563557981173 chrl9:12767749-127679801174 chrl9:19371675-193723931175 chrl4:69256676-692570361176 chrl7:75447477-754478211177 chrl4:24801680-248021531178 chr5 : 148033472-1480340801179 chrl0:125650820-1256513731180 chrll:43568921-435698541181 chr22:37212769-372134671182 chr2:162283581-1622846771183 chr8:130995921-1309961491184 chrll:70508328-705086171185 chrl6:88943427-889436691186 chrl9:42891311-428916461187 chrl5:53079220-530795791188 chrl7:46690390-466910551189 chr4:41880224-418805001190 chrl:156105707-1561061711191 chr6:5997027-59974141192 chrl:18964180-189644011193 chrl4:36983440-369837381194 chrl2:54445876-544461131195 chr5:87968635-879689071196 chrl:29587087-295874121197 chrll:60718428-607188881198 chr2:66672431-666736361199 chr4:81119095-811193911200 chrl0:76573195-765735071201 chr22:42322043-423229091202 chrl9:45898879-459003151203 chrl4:95826675-958269411204 chrl7:48194634-481950851205 chrl9:49669275-496695521206 chrl5:96897596-968980461207 chrl9:40314926-403151441208 chr9:120507227-1205076421209 chr5:145722467-1457229251210 chr3:19188246-191887721211 chr5 : 140787447-1407880441212 chrl9:50881418-508816641213 chrl0:102896342-1028966651214 chr7:53286851-532871921215 chrl5:89903446-899037201216 chrl0:23461300-234616101217 chr2:127783081-1277833111218 chrll:72532612-725337741219 chr2:119605200-1196056201220 chrl8:12254147-122550891221 chr7:100817759-1008179751222 chrl4:77736733-777377721223 chrl2:127212279-1272125291224 chr2:119606569-1196068261225 chrl:155264318-1552655361226 chrl2:131199824-1312001571227 chrl:91300979-913018911228 chr6:100909210-1009094441229 chr6:4079052-40794431230 chr2:233251361-2332534141231 chr4:960505-9608361232 chrl9:21769189-217697861233 chrl0:102279162-1022797301234 chrl2:127210778-1272116511235 chrl2:54069625-540701771236 chrl5:53087211-530874881237 chrl3:28365545-283657851238 chrl2:113913615-1139143221239 chrl4:51338712-513391461240 chr7:155604725-1556050951241 chr3:62364017-623643161242 chr6:6008857-60092991243 chr3:46618307-466186691244 chrl7:33776553-337768881245 chrl2:58158855-581600001246 chr2:219857682-2198589171247 chrl9:44278273-442787771248 chrl0:101282725-1012829341249 chr20:2539133-25398771250 chrl2:58003880-580042491251 chrl6:51147490-511479441252 chrl:179544720-1795453071253 chr2:71787430-717878971254 chrl0:129534410-1295373661255 chr6:42145847-421460531256 chrl4:24802927-248031591257 chr22:29707479-297077971258 chr9:132459587-1324600171259 chrl7:40937258-409374801260 chr4:151504011-1515050851261 chrl:18967251-189681191262 chrl9:56598038-566002961263 chrl9:35633409-356336971264 chr2:171678546-1716803581265 chr6:134638797-1346390211266 chrl:36549554-365499651267 chrl9:12833104-128335741268 chr3:137487429-1374880211269 chr9:139715663-1397164411270 chr6:37617863-376181471271 chrl7:32484007-324842801272 chr7:156409577-1564098651273 chr5:11384681-113855211274 chr8:102504478-1025048411275 chr20:33296514-332982421276 chr20:57415135-574171531277 chrl0:71331449-713316911278 chr3:75667777-756690671279 chrl6:67571252-675727281280 chrl9:36500169-365005301281 chr2:154729613-1547299181282 chrl2:48399168-483993721283 chr4:41867385-418675861284 chrl7:46800533-468007461285 chr20:44685771-446876101286 chrl9:10406934-104073421287 chr6:108496715-1084973201288 chr5:158523906-1585245981289 chr9:124413512-1244141931290 chr20:57427691-574279951291 chrl6:10912159-109127191292 chr7:149389654-1493899761293 chrl:173638662-1736390451294 chrl9:55597977-555988871295 chrl4:62279037-622793391296 chr3:13114627-131152451297 chr2:3750828-37519271298 chr4:85402764-854031751299 chrl7:74017769-740186581300 chr5:54523676-545239011301 chr7:89747892-897490361302 chrl8:72916107-729172331303 chr9:136294738-1362952361304 chrl:201252452-2012536481305 chr5: 146888750-1468898401306 chrl4:52734207-527354861307 chrl3:20875518-208762141308 chrl8:77560088-775602921309 chr2:102803672-1028045561310 chr2:176982107-1769824021311 chrl7:6679205-66797101312 chrl9:10463626-104643781313 chr5: 140810494-1408126171314 chrll:46299544-463002161315 chrll:64136814-641381871316 chr6:6007387-60077971317 chrl7:37321482-373220991318 chrl0:94455524-944558961319 chrl3:51417371-514181491320 chr8:11565217-115672121321 chrl:226127112-2261276951322 chr2:3287874-32882281323 chr6:10882926-108831491324 chr22:19746155-197463691325 chr3:12838471-128387821326 chr9:36739534-367397821327 chr9:134429866-1344304911328 chrll:70672834-706730551329 chrl4:24641053-246422201330 chr7:27283408-272836141331 chrl2:49182421-491826581332 chrl:44031286-440318531333 chrl:114696886-1146971851334 chrl5:89901914-899027851335 chrll:65352231-653531341336 chr7:72838383-728388151337 chr22:38379093-383799641338 chr4:155663809-1556643151339 chr9:100619984-1006201921340 chr7:143582125-1435826101341 chr7:23287221-232875081342 chrll:64815040-648157221343 chr2:87088816-870890371344 chr20:57426729-574270471345 chrl0:43428167-434294601346 chrl0:121577529-1215783851347 chr4:190939801-1909405911348 chr6:100037323-1000375441349 chrl9:12880574-128808881350 chr2:171670110-1716705491351 chr7:124404174-1244044321352 chr7:97840559-978408451353 chrl9:50879606-508800941354 chrl:113265573-1132657871355 chrl9:2424005-24279831356 chr3:127633993-1276345881357 chrl0:50817095-508173091358 chr2:171676552-1716769801359 chrl:86621278-866228711360 chrl:164545540-1645459171361 chr22:19967279-199678081362 chrll:67350928-673519531363 chr20:36226617-362268411364 chrl9:14089570-140897961365 chrl9:38700333-387005771366 chrl:18435566-184359041367 chr8:21905461-219057571368 chr2:176950595-1769508461369 chrl7:75251958-752521801370 chrl5:37390175-373903801371 chr9:98113447-981136621372 chrl:40235767-402371901373 chr8:144811237-1448114461374 chr8:99984584-999850721375 chr7:152621916-1526221491376 chrl:40769186-407698711377 chrl9:2428349-24287311378 chrl7:15820620-158213251379 chr22:25081850-250821121380 chrl:19203874-192042341381 chr20:61703526-617040221382 chr2:237080188-2370804321383 chrl:156338758-1563392511384 chrl:149332993-1493333891385 chr22:50496441-504973931386 chr7:27146069-271466001387 chrl3:100547633-1005489111388 chr4:190939007-1909392741389 chr7:73894815-738951101390 chrl9:35632356-356325721391 chrl6:67918679-679189091392 chr2:108602824-1086034671393 chr2:238864315-2388651701394 chr8:144808221-1448109781395 chr8:145101631-1451018341396 chrl2:132905449-1329062061397 chr6:99275763-992760381398 chr5:140800760-1408010721399 chrl7:75242871-752436131400 chrl7:41278134-412784601401 chrl2:122016170-1220176931402 chrl0:131264948-1312657101403 chrl7:46631800-466322121404 chrl4:105167277-1051675011405 chrl0:23982382-239825891406 chrl9:50931270-509316381407 chr3:27771638-277719421408 chrl8:74799144-748000381409 chrl:21616380-216171011410 chrl:147782066-1477824731411 chr7:6590563-65909571412 chr7:97839862-978402221413 chrl2:113914440-1139146571414 chrl9:7933263-79348981415 chr20:22559553-225600011416 chrl5:53086629-530868581417 chrl0:94180315-941807541418 chr5:140052059-1400533811419 chrl0:101287162-1012879201420 chrl4:38677154-386777871421 chr22:39262338-392632111422 chrl8:74153239-741550731423 chrl5:59157045-591575941424 chr4:963804-9641151425 chrll:624780-6250531426 chr7:1362811-13636431427 chrl9:36246328-362479821428 chr5:54528095-545284041429 chrl2:54359658-543599061430 chr2:127782613-1277828291431 chrl9:406131-4066111432 chrl7:46697413-466977011433 chrl8:43608140-436085101434 chrl6:23724270-237247751435 chrl8:55922987-559240681436 chrl5:60291879-602921671437 chrl4:92788913-927892041438 chrl9:1108394-11096101439 chrll:124628367-1246295901440 chrl:32052471-320527711441 chrl9:11594372-115949871442 chrl9:870774-8713181443 chr2:54086776-540872661444 chr2:241459632-2414600471445 chr7:127990926-1279926161446 chrl:208132327-2081331171447 chr7:90893567-908966831448 chrl:41284847-412851491449 chrll:32452144-324527081450 chr5:77146998-771477851451 chrl9:45901452-459016881452 chr7:6661875-66626951453 chr6:161188084-1611886391454 chrl7:934417-9350881455 chrll:65409636-654101271456 chrl7:19883325-198836101457 chrl8:77549524-775502991458 chrl:38461584-384619881459 chrl9:10464666-104649271460 chrl7:70120139-701204421461 chr7:27147589-271483891462 chr2:31806545-318067821463 chrll:119292689-1192928911464 chrl9:18979351-189812001465 chr6:42879279-428796231466 chrl2:130908777-1309091911467 chrl7:46629553-466298161468 chrl:202162958-2021633901469 chrl7:21367114-213675921470 chrl6:84001805-840020111471 chrl:221057463-2210577571472 chrl7:27899511-279000671473 chrl5:40268581-402690611474 chr22:37465056-374653311475 chrl7:77805866-778090461476 chrl9:13198699-131989991477 chr3:184056419-1840566711478 chr22:37911979-379122581479 chrl9:19368708-193696811480 chrll:64135815-641363811481 chrl8:77552401-775526031482 chrl9:58554354-585545871483 chr20:57414595-574148961484 chr4:190938106-1909388481485 chr5:172110282-1721111661486 chrl6:68480864-684828221487 chr9:139395020-1393952871488 chrl2:113515164-1135159701489 chrl:221054554-2210548881490 chr8:144990270-1450021351491 chr9:131154346-1311559231492 chr6:150335525-1503362781493 chr9:115824684-1158250331494 chrl2:54519768-545204571495 chr6:35479872-354801541496 chrl9:3870788-38710431497 chrl9:48965002-489657921498 chr6:35479388-354796781499 chrl2:52408381-524086751500 chrl:221068782-2210691591501 chr6:46655262-466567381502 chr3:55508336-555087081503 chrl:39980365-399817681504 chrl6:3067521-30683581505 chrl:1473107-14733421506 chrl0:105362549-1053628271507 chrl7:46698880-466990831508 chr2:198029068-1980294381509 chr20:17209418-172096221510 chrl2:49183049-491832821511 chrl6:58030214-580316331512 chrl0:94820026-948232521513 chrll:725596-7268701514 chr6:170732119-1707324421515 chrl2:120835586-1208359271516 chr20:36012595-360134391517 ch r8 : 143545445-1435461781518 chr6:27228100-272283641519 chr21:32624144-326243821520 chr9:95477296-954777081521 chrl0:105420685-1054210761522 chrl:1470604- 14714501523 chrl:146552328-1465525771524 chrl9:33625467-336258051525 chrll:64478843-644795981526 chr20:57428308-574285161527 chr7:27182613-271855621528 chrl9:51815157-518154581529 chrl7:46607804-466083901530 chrl2:52408860-524091211531 chrl9:10405924-104063981532 chrll:14993452-149936611533 chrl9:13135317-131361691534 chr7:750788-7512371535 chrl:53742297-537428451536 chrl:200010625-2000108321537 chr5:139138875-1391392421538 chrl7:45949676-459498851539 chr3:128722283-1287230361540 chrl5:89312719-893131831541 chr9:135039673-1350399781542 chrl9:12831793-128322251543 chr20:51589707-515900201544 chr20:3145121-31457461545 chr8:65710990-657117221546 chrll:128694084-1286946881547 chr2:20870006-208712801548 chrl9:18977466-189778331549 chr3:49947621-499484301550 chr6:30139718-301402631551 chrl2:104697348-1046979841552 chrl0:105361784-1053621881553 chr6:29894140-298951171554 chr4:187219320-1872197451555 chrl5:67073306-670739431556 chr2:220412341-2204126781557 chr6:170730395-1707308871558 chr9:115822071-1158234161559 chrl:10764449-107649251560 chrl7:46627787-466284441561 chrl9:51601822-516022601562 chrl9:55814067-558142781563 chr6:138745348-1387455931564 chr9:124987743-1249910861565 chr22:46318693-463190871566 chrl6:3013016-30132281567 chr4:114900355-1149008101568 chrl9:1063544-10642651569 chrl9:1110399-11107011570 chr7:97841636-978420051571 chr8:57359899-573601141572 chrl7:72915568-729165101573 chrl:16860873-168622961574 chrl7:75398284-753985271575 chr9:139397412-1393977101576 chr6:33393592-333939081577 chr6:29595298-295957951578 chrl2:6438272-64389311579 chr3:113160299-1131606411580 chrl:55505060-555060151581 chrll:132951692-1329522601582 chr4:81118137-811186031583 chrl9:38876070-388763321584 chrl9:58549305-585497121585 chrl7:43472527-434743431586 chr9:139396205-1393970401587 chrl6:3192181-31926691588 chr6:33048416-330488141589 chr7:128555329-1285566501590 chrl9:46915311-469158021591 chr6:30095173-30095610Table 2: Example CGIsTable 3: Additional Example CGTsTable 4: Additional Example CGIs

Claims

CLAIMS1. A tiered, multipart method for tracking tumor heterogeneity across at least first and second biological samples obtained from a subject, the method comprising:(a) performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained from the subject at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA;(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing, a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein the second biological sample was obtained from the subject at a second timepoint subsequent to the first timepoint;(iii) determining a change in signal between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background-corrected methylation information from the second intra-individual analysis; and(iv) performing a second analysis comprising analyzing the determined change in signal to track tumor heterogeneity across the first biological sample and the second biological sample.

2. A method for tracking tumor heterogeneity in a patient during or subsequent to administration of a tumor therapeutic, or in a patient being considered for administration of a tumor therapeutic, comprising:(a) confirming that a first biological sample of the patient is not identified as not at risk of containing circulating tumor DNA;(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing, at a baseline timepoint, a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing, at a second timepoint, a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein between the baseline timepoint and second timepoint the patient may be administered, or continues to be administered, one or more tumor therapeutics;(iii) determining a change in signal between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background-corrected methylation information from the second intra-individual analysis; and(iv) performing a second analysis comprising analyzing the determined change in signal to assess tumor heterogeneity across the first biological sample and the second biological sample and therefore track the patient’s therapeutic progress and / or assess the tumor therapeutic.

3. The method of claim 1 or 2, wherein determining the change in signal comprises determining a difference between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background- corrected methylation information from the second intra-individual analysis.

4. The method of any one of claims 1-3, wherein the first set of background-corrected methylation information or the second set of background-corrected methylation information comprises methylation statuses for a plurality of genomic sites.

5. The method of claim 4, wherein the plurality of genomic sites comprise a plurality of CpG sites.

6. The method of claim 5, wherein the plurality of CpG sites are located in one or more CpG islands or portions of one or more CpG islands shown in Tables 1-4.

7. The method of claim 4, wherein the first set of background-corrected methylation information and the second set of background-corrected methylation information comprises methylation statuses for a plurality of CpG sites.

8. The method of claim 7, wherein the plurality of CpG sites of the first set of background-corrected methylation information are the same plurality of CpG sites of the second set of background-corrected methylation information.

9. The method of any one of claims 1-8, wherein performing the first intra-individual analysis comprises: obtaining target nucleic acids and reference nucleic acids from the first biological sample obtained from the subject; performing bisulfite conversion of the target nucleic acids and the reference nucleic acids; selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids; generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; and using a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the first set of background-corrected methylation information.

10. The method of claim 9, wherein the reference nucleic acids from the first biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.

11. The method of any one of claims 1-10, wherein the first set of background-corrected methylation information comprises phased sequencing information.

12. The method of claim 11, wherein the phased sequencing information of the first set of background-corrected methylation information is generated by: obtaining or having obtained sequence reads of cell-free DNA from the first sample;obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length; attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; and aligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.

13. The method of claim 12, wherein the phased sequencing information of cell-free DNA comprises methylation statuses for a plurality of genomic sites of the cell-free DNA.

14. The method of claim 13, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.

15. The method of claim 12, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.

16. The method of claim 15, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.

17. The method of claim 16, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.

18. The method of claim 16 or 17, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.

19. The method of any one of claims 12-18, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.

20. The method of any one of claims 12-19, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, atleast 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.

21. The method of any one of claims 1-8, wherein performing the second intra-individual analysis comprises: obtaining target nucleic acids and reference nucleic acids from the second biological sample obtained from the subject; performing bisulfite conversion of the target nucleic acids and the reference nucleic acids; selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids; generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; and using a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the second set of background-corrected methylation information.

22. The method of claim 21, wherein the reference nucleic acids from the second biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.

23. The method of any one of claims 1-22, wherein the first set of background-corrected methylation information and / or the second set of background-corrected methylation information comprise a high resolution measure of methylation.

24. The method of claim 23, wherein the high resolution measure of methylation comprises a total quantity of consecutively methylated CpG sites within target regions.

25. The method of claim 24, wherein the total quantity of consecutively methylated CpG sites within target regions comprises the total quantity of 3, 4, or 5 consecutively methylated CpG sites within target regions.

26. The method of claim 23, wherein the high resolution measure of methylation comprises methylation statuses of a plurality of CpG sites from a haplotype.

27. The method of any one of claims 1-26, wherein the second set of background- corrected methylation information comprises phased sequencing information.

28. The method of claim 27, wherein the phased sequencing information of the second set of background-corrected methylation information is generated by: obtaining or having obtained sequence reads of cell-free DNA from the second sample; obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length; attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; and aligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.

29. The method of claim 28, wherein the phased sequencing information of cell-free DNA comprises methylation statuses for a plurality of genomic sites of the cell-free DNA.

30. The method of claim 29, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.

31. The method of claim 28, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.

32. The method of claim 31, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.

33. The method of claim 32, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.

34. The method of claim 32 or 33, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.

35. The method of any one of claims 28-34, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.

36. The method of any one of claims 28-35, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, at least 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.

37. The method of any one of claims 1-36, wherein the nucleic acid sequence information of the first analysis comprises methylation sequence information.

38. The method of claim 37, wherein the methylation sequence information of the first analysis comprises methylation statuses for a plurality of genomic sites.

39. The method of claim 38, wherein the plurality of genomic sites comprise a plurality of CpG sites.

40. The method of claim 38, wherein the nucleic acid sequence information of the first analysis comprises a measure of overall methylation across the plurality of genomic sites.

41. The method of claim 40, wherein the measure of overall methylation comprises a total number of methylated genomic sites or an average number of methylated genomic sites.

42. The method of any one of claims 1-41, wherein performing the first analysis of nucleic acid sequence information comprises applying a trained machine learning model.

43. The method of any one of claims 1-42, wherein the method delivers improved performance as a function of resource consumption in comparison to the single tier method.

44. The method of any one of claims 1-42, wherein the method achieves an improved performance metric in comparison to a single tier method.

45. The method of any one of claims 1-42, wherein the method tracks tumor heterogeneity of one or more of the early stage cancers.

46. The method of claim 45, wherein the one or more of the early stage cancers is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, cardiac cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumor, gestational trophoblastic cancer, hairy cell leukemia, liver cancer, Hodgkin lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, mouth cancer, multiple endocrine neoplasia syndromes, multiple myeloma neoplasms, myelodysplastic neoplasms, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

47. The method of claim 45, wherein the one or more early stage cancers is a preclinical phase cancer.

48. The method of claim 47, wherein the preclinical phase cancer is stage I or stage II cancer.

49. The method of any one of claims 1-48, wherein the nucleic acid sequence information, the background-corrected methylation information of the first intra-individual analysis, and / or the background-corrected methylation information of the second intra- individual analysis is obtained from an assay, wherein the assay comprises performing one or more of: a. sequencing of nucleic acids; b. hybrid capture; c. methylation-specific PCR; d. an assay that generates methylation information; and e. sequencing a clone library generated from a template immortalized library.

50. The method of any one of claims 1-49, wherein each of the first biological sample and the second biological sample independently comprises any one of a blood sample, a stool sample, a urine sample, a mucous sample, or a saliva sample.

51. The method of any one of claims 1-50, wherein each of the first biological sample and the second biological sample is a blood sample.

52. The method of claim 51, wherein each of the first biological sample and the second biological sample does not comprise an invasive biopsy sample.

53. The method of any one of claims 1-52, wherein the second analysis comprises whole genome sequencing, optionally whole genome bisulfite sequencing.

54. The method of any one of claim 1-53, wherein the subject received a tumor therapeutic prior to the first timepoint.

55. The method of any one of claim 1-53, wherein subsequent to the first timepoint and prior to the second timepoint, the subject received a tumor therapeutic.

56. The method of claim 54 or 55, further comprising determining an efficacy of the tumor therapeutic based on the tracked tumor heterogeneity.

57. The method of claim 56, wherein if the tracked tumor heterogeneity indicates a stable or increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the tumor therapeutic lacks efficacy.

58. The method of claim 57, further comprising selecting a new tumor therapeutic for the subject responsive to determining that the tumor therapeutic lacks efficacy.

59. The method of claim 56, wherein if the tracked tumor heterogeneity indicates a reducing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the tumor therapeutic achieves therapeutic efficacy.

60. The method of any one of claims 1-59, wherein prior to (a), a prior sample obtained from the subject was previously determined to be not at risk for containing circulating tumor DNA.

61. The method of claim 60, wherein further responsive to determining that the first biological sample is not identified as not at risk, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.

62. The method of claim 60, wherein if the tracked tumor heterogeneity indicates an increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.

63. A tiered, multipart method for assessing tumor heterogeneity across at least first and second biological samples obtained from a subject, the method comprising:(a) performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA,(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing a second analysis of the first biological sample comprising analyzing the background-corrected methylation information to predict a tumor heterogeneity state;(c) determining an updated tumor heterogeneity state by:(i) performing a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein the second biological sample was obtained from the subject at a second timepoint subsequent to the first timepoint; and(ii) performing a second analysis of the second biological sample comprising analyzing the background-corrected methylation information to predict the updated tumor heterogeneity state; and(d) comparing the tumor heterogeneity state from the first biological sample to the updated tumor heterogeneity state from the second biological sample to track tumor heterogeneity across the first biological sample and the second biological sample.

64. A method for assessing tumor heterogeneity in a patient during or subsequent to administration of a tumor therapeutic, or in a patient being considered for administration of a tumor therapeutic, comprising:(a) confirming that a first biological sample of the patient is not identified as not at risk of containing circulating tumor DNA;(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing, at a baseline timepoint, a first intra-individual analysis using the first biological sample to generate a first set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the first biological sample and methylation information from reference nucleic acids from the first biological sample;(ii) performing, at a second timepoint, a second intra-individual analysis using a second biological sample to generate a second set of background-corrected methylation information representing a difference between methylation information from target nucleic acids from the second biological sample and methylation information from reference nucleic acids from the second biological sample, wherein between the baseline timepoint and second timepoint the patient may be administered, or continues to be administered, one or more tumor therapeutics;(iii) determining a change in signal between the first set of background-corrected methylation information from the first intra-individual analysis and the second set of background-corrected methylation information from the second intra-individual analysis; and(iv) performing a second analysis comprising analyzing the determined change in signal to assess tumor heterogeneity across the first biological sample and thesecond biological sample and therefore track the patient’s therapeutic progress and / or assess the tumor therapeutic.

65. The method of claim 63 or 64, wherein the first set of background-corrected methylation information or the second set of background-corrected methylation information comprises methylation statuses for a plurality of genomic sites.

66. The method of claim 65, wherein the plurality of genomic sites comprise a plurality of CpG sites.

67. The method of claim 66, wherein the plurality of CpG sites are located in one or more CpG islands or portions of one or more CpG islands shown in Tables 1-4.

68. The method of claim 65, wherein the first set of background-corrected methylation information and the second set of background-corrected methylation information comprises methylation statuses for the plurality of CpG sites.

69. The method of claim 68, wherein the plurality of CpG sites of the first set of background-corrected methylation information are the same plurality of CpG sites of the second set of background-corrected methylation information.

70. The method of any one of claims 63-69, wherein performing the first intra-individual analysis comprises: obtaining target nucleic acids and reference nucleic acids from the first biological sample obtained from the subject; performing bisulfite conversion of the target nucleic acids and the reference nucleic acids; selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids; generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; and using a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the first set of background-corrected methylation information.

71. The method of claim 70, wherein the reference nucleic acids from the first biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.

72. The method of any one of claims 63-71, wherein the first set of background-corrected methylation information comprise a high resolution measure of methylation.

73. The method of claim 72, wherein the high resolution measure of methylation comprises a total quantity of consecutively methylated CpG sites within target regions.

74. The method of claim 73, wherein the total quantity of consecutively methylated CpG sites within target regions comprises the total quantity of 3, 4, or 5 consecutively methylated CpG sites within target regions.

75. The method of claim 72, wherein the high resolution measure of methylation comprises methylation statuses of a plurality of CpG sites from a haplotype.

76. The method of any one of claims 63-75, wherein the first set of background-corrected methylation information comprises phased sequencing information.

77. The method of claim 76, wherein the phased sequencing information of the first set of background-corrected methylation information is generated by: obtaining or having obtained sequence reads of cell-free DNA from the first sample; obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length; attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; and aligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.

78. The method of claim 77, wherein the phased sequencing information comprises methylation statuses for a plurality of genomic sites.

79. The method of claim 78, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.

80. The method of claim 77, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.

81. The method of claim 80, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.

82. The method of claim 81, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.

83. The method of claim 81 or 82, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.

84. The method of any one of claims 77-83, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.

85. The method of any one of claims 77-84, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, at least 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.

86. The method of any one of claims 63-69, wherein performing the second intraindividual analysis comprises: obtaining target nucleic acids and reference nucleic acids from the second biological sample obtained from the subject; performing bisulfite conversion of the target nucleic acids and the reference nucleic acids; selectively amplifying target regions comprising a plurality of CpG sites of the bisulfite converted target nucleic acids and reference nucleic acids; generating a dataset comprising methylation information of the plurality of CpG sites from the target nucleic acids and methylation information of the plurality of CpG sites from the reference nucleic acids; andusing a computer processor, combining the methylation information of the plurality of CpG sites from the target nucleic acids and the methylation information of the plurality of CpG sites from the reference nucleic acids to generate the first set of background-corrected methylation information.

87. The method of claim 86, wherein the reference nucleic acids from the second biological sample comprise genomic DNA from peripheral blood mononuclear cells (PBMCs) or polymorphonuclear cells of the subject.

88. The method of any one of claims 86-87, wherein the second set of background- corrected methylation information comprises a high resolution measure of methylation.

89. The method of claim 88, wherein the high resolution measure of methylation comprises a total quantity of consecutively methylated CpG sites within target regions.

90. The method of claim 89, wherein the total quantity of consecutively methylated CpG sites within target regions comprises the total quantity of 3, 4, or 5 consecutively methylated CpG sites within target regions.

91. The method of claim 88, wherein the high resolution measure of methylation comprises methylation statuses of a plurality of CpG sites from a haplotype.

92. The method of any one of claims 63-91, wherein the second set of background- corrected methylation information comprises phased sequencing information.

93. The method of claim 92, wherein the phased sequencing information of the second set of background-corrected methylation information is generated by: obtaining or having obtained sequence reads of cell-free DNA from the second sample; obtaining or having obtained long sequence reads of reference nucleic acids from the second sample, wherein the long sequence reads of reference nucleic acids are at least 500 bases in length; attributing long sequence reads of reference nucleic acids to one of two or more different sources of the subject; and aligning the obtained sequence reads of cell-free DNA to the long sequence reads of reference nucleic acids.

94. The method of claim 93, wherein the phased sequencing information of cell-free DNA comprises methylation statuses for a plurality of genomic sites of the cell-free DNA.

95. The method of claim 94, wherein the methylation statuses for the plurality of genomic sites comprise at least one coupled genomic site representing two or more methylated genomic sites originating from a common source.

96. The method of claim 93, wherein the phased sequencing information comprises mutation sequence information of the cell-free DNA.

97. The method of claim 96, wherein the mutation sequence information comprises a plurality of mutations present across the plurality of genomic sites.

98. The method of claim 97, wherein the plurality of mutations present across the plurality of genomic sites comprise coupled genomic sites representing two or more mutated genomic sites originating from a common source.

99. The method of claim 97 or 98, wherein the plurality of mutations comprise one or more of a single nucleotide polymorphism (SNP), single nucleotide variant (SNV), insertion, deletion, copy number variation (CNV), duplication, or translocation.

100. The method of any one of claims 93-99, wherein the two or more different sources of the subject comprise a maternal chromosome source or a paternal chromosome source.

101. The method of any one of claims 93-100, wherein the long sequence reads of reference nucleic acids comprise at least 500 bases, at least 1000 bases, at least 2000 bases, at least 3000 bases, at least 4000 bases, at least 5000 bases, at least 6000 bases, at least 7000 bases, at least 8000 bases, at least 9000, at least 10,000 bases, at least 12,000 bases, at least 15,000 bases, at least 20,000 bases, at least 25,000 bases, at least 30,000 bases, at least 40,000 bases, at least 50,000 bases, at least 60,000 bases, at least 70,000 bases, at least 80,000 bases, at least 90,000 bases, or at least 100,000 bases.

102. The method of any one of claims 63-101, wherein the nucleic acid sequence information of the first analysis comprises methylation sequence information.

103. The method of claim 102, wherein the methylation sequence information of the first analysis comprises methylation statuses for a plurality of genomic sites.

104. The method of claim 103, wherein the plurality of genomic sites comprise a plurality of CpG sites.

105. The method of claim 103, wherein the nucleic acid sequence information of the first analysis comprises a measure of overall methylation across the plurality of genomic sites.

106. The method of claim 105, wherein the measure of overall methylation comprises a total number of methylated genomic sites or an average number of methylated genomic sites.

107. The method of any one of claims 63-106, wherein performing the first analysis of nucleic acid sequence information comprises applying a trained machine learning model.

108. The method of any one of claims 63-107, wherein the tiered, multipart method delivers improved performance as a function of resource consumption in comparison to the single tier method.

109. The method of any one of claims 63-107, wherein the tiered, multipart method achieves an improved performance metric in comparison to a single tier method.

110. The method of any one of claims 63-107, wherein the tiered, multipart method tracks tumor heterogeneity of one or more of the early stage cancers.

111. The method of claim 110, wherein the one or more of the early stage cancers is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, soft tissue sarcoma, lymphoma, anal cancer, gastrointestinal cancer, brain cancer, skin cancer, bile duct cancer, bladder cancer, bone cancer, breast cancer, lung cancer, cardiac cancer, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative neoplasms, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, eye cancer, fallopian tube cancer, gallbladder cancer, gastric cancer, germ cell tumor, gestational trophoblastic cancer, hairy cell leukemia, liver cancer, Hodgkin lymphoma, intraocular melanoma, pancreatic cancer, kidney cancer, leukemia, mesothelioma, metastatic cancer, mouth cancer, multiple endocrine neoplasiasyndromes, multiple myeloma neoplasms, myelodysplastic neoplasms, ovarian cancer, parathyroid cancer, penile cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasm, primary peritoneal cancer, prostate cancer, rectal cancer, retinoblastoma, sarcoma, small intestine cancer, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, urethral cancer, uterine cancer, vaginal cancer, and vulvar cancer.

112. The method of claim 110, wherein the one or more early stage cancers is a preclinical phase cancer.

113. The method of claim 112, wherein the preclinical phase cancer is stage I or stage II cancer.

114. The method of any one of claims 63-113, wherein the nucleic acid sequence information, the background-corrected methylation information of the first intra-individual analysis, and / or the background-corrected methylation information of the second intra- individual analysis is obtained from an assay, wherein the assay comprises performing one or more of: a. sequencing of nucleic acids; b. hybrid capture; c. methylation-specific PCR; d. an assay that generates methylation information; and e. sequencing a clone library generated from a template immortalized library.

115. The method of any one of claims 63-114, wherein each of the first biological sample and the second biological sample independently comprises any one of a blood sample, a stool sample, a urine sample, a mucous sample, or a saliva sample.

116. The method of any one of claims 63-115, wherein each of the first biological sample and the second biological sample is a blood sample.

117. The method of claim 116, wherein each of the first biological sample and the second biological sample does not comprise an invasive biopsy sample.

118. The method of any one of claims 63-117, wherein the second analysis of the first biological sample or the second analysis of the second biological sample comprises whole genome sequencing, optionally whole genome bisulfite sequencing.

119. The method of any one of claim 63-118, wherein the subject received a tumor therapeutic prior to the first timepoint.

120. The method of any one of claim 63-118, wherein subsequent to the first timepoint and prior to the second timepoint, the subject received a tumor therapeutic.

121. The method of claim 119 or 120, further comprising determining an efficacy of the tumor therapeutic based on the assessed tumor heterogeneity.

122. The method of claim 121, wherein if the tracked tumor heterogeneity indicates a stable or increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the assessed tumor therapeutic lacks efficacy.

123. The method of claim 122, further comprising selecting a new intervention for the subject responsive to determining that the assessed tumor therapeutic lacks efficacy.

124. The method of claim 121, wherein if the tracked tumor heterogeneity indicates a reducing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the assessed tumor therapeutic achieves therapeutic efficacy.

125. The method of any one of claims 63-124, wherein prior to (a), a prior sample obtained from the subject was previously determined to be not at risk for containing circulating tumor DNA.

126. The method of claim 125, wherein further responsive to determining that the first biological sample is not identified as not at risk, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.

127. The method of claim 125, wherein if the tracked tumor heterogeneity indicates an increasing tumor heterogeneity in the subject across the first biological sample and the second biological sample, determining that the prior sample previously determined to be not at risk for containing circulating tumor DNA was a false negative.

128. A tiered, multipart method for determining whether a prior sample obtained from a subject was a false negative sample, the method comprising:(a) performing a first analysis of nucleic acid sequence information that was derived from an assay performed on a first biological sample obtained from the subject at a first timepoint to identify whether the biological sample is not at risk of containing circulating tumor DNA, wherein the prior sample was obtained from the subject prior to the first biological sample,(b) responsive to determining that the first biological sample is not identified as not at risk:(i) performing one or more intra-individual analyses using the first biological sample or one or more additional biological samples, wherein performing each intra- individual analysis involves generating background-corrected methylation information representing a difference between methylation information from target nucleic acids and methylation information from reference nucleic acids from one of the first biological sample or one or more additional biological samples;(ii) performing a longitudinal analysis comprising analyzing generated background- corrected methylation information from each of the one or more intra-individual analyses; and(iii) determining that the prior sample obtained from a subject was a false negative sample.

129. The method of claim 128, wherein determining that the prior sample obtained from a subject was a false negative sample is responsive to determining that the first biological sample is not identified as not at risk.

130. The method of claim 128, wherein determining that the prior sample obtained from a subject was a false negative sample is responsive to performing the longitudinal analysis.

131. The method of any one of claims 128-130, wherein performing the longitudinal analysis comprising tracking tumor heterogeneity across the first biological sample and one or more additional biological samples.

132. A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-127.

133. A sy stem compri sing : a processor; and a non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-127.

Citation Information

Patent Citations

  • Universal early cancer diagnostics

    US20200109456A1

  • Methods of cancer detection using extraembryonically methylated CPG islands

    WO2022133315A1

  • Methods and systems for tumor detection

    US20180237863A1

  • Universal early cancer diagnostics

    WO2018209361A2

  • Cancer classification using patch convolutional neural networks

    WO2021119471A1