Methods and systems for analyzing nucleic acid molecules

The method improves the detection of minimal residual disease in cancer patients by analyzing phased or graded variants in cell-free nucleic acids, addressing limitations in current detection methods and enhancing sensitivity and specificity for early disease detection.

JP7825272B2Active Publication Date: 2026-03-06THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-11-06
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Current methods for detecting minimal residual disease (MRD) in cancer patients using cell-free nucleic acids are limited by low input DNA quantities and high background error rates, leading to insufficient detection of residual disease recurrence or progression, exemplified by false-negative rates in diffuse large B-cell lymphoma, colon, and breast cancer.

Method used

A method and system for analyzing cell-free nucleic acids (cfDNA, cfRNA) using sequencing data to identify phased or graded variants relative to a reference genome, allowing for improved sensitivity and specificity in detecting cancer-derived nucleic acids, with detection limits as low as 1 in 50,000 observations.

Benefits of technology

Enhances the detection of minimal residual disease by improving sensitivity and specificity, enabling early detection of disease recurrence or progression, particularly in conditions like diffuse large B-cell lymphoma, through the analysis of phased or graded variants in cell-free nucleic acids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007825272000228
    Figure 0007825272000228
  • Figure 0007825272000229
    Figure 0007825272000229
  • Figure 0007825272000230
    Figure 0007825272000230
Patent Text Reader

Abstract

Processes and materials for detecting cancer from biopsies are described. In some cases, cell-free nucleic acids can be sequenced, and sequencing results can be used to detect sequences derived from neoplasms. Detection of in-phase somatic variants can indicate the presence of cancer in diagnostic scans, allowing clinical intervention. The present disclosure provides methods and systems for analyzing cell-free nucleic acids (e.g., cfDNA, cfRNA) from a subject. The disclosed methods and systems can utilize sequencing results from a subject to detect cancer-derived nucleic acids (e.g., ctDNA, ctRNA), for example, for disease diagnosis, disease monitoring, or treatment decisions for the subject. The disclosed methods and systems can exhibit improved sensitivity, specificity, and / or reliability in detecting cancer-derived nucleic acids.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 931,688, filed November 6, 2019, which is incorporated herein by reference in its entirety. Sequence Listing

[0002] This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy, created on November 3, 2020, is named 58626-702_601_SL.txt and is 307,199 bytes in size. Government Rights

[0003] This invention was made with government support under Grants CA233975, CA241076, and CA188298 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]

[0004] background Noninvasive blood tests capable of detecting somatic alterations (e.g., mutant nucleic acids) based on the analysis of cell-free nucleic acids (e.g., cell-free deoxyribonucleic acid (cfDNA) and cell-free ribonucleic acid (cfRNA)) are attractive candidates for cancer screening applications due to the relative ease of obtaining biological specimens (e.g., biological fluids). Circulating tumor nucleic acids (e.g., ctDNA or ctRNA, i.e., nucleic acids derived from cancerous cells) can be sensitive and specific biomarkers in numerous cancer subtypes. However, current methods for minimal residual disease (MRD) detection from ctDNA can be limited by one or more factors, such as low input DNA quantities and high background error rates.

[0005] Recent approaches have improved ctDNA MRD performance by tracking multiple somatic mutations using error-suppression sequencing, resulting in detection limits as low as 4 in 100,000 from limited cfDNA input. Detection of residual disease during or after treatment is a powerful tool, and detectable MRD represents an adverse prognostic sign, even during radiological remission. However, current detection limits may be insufficient to universally detect residual disease in patients whose disease is set to recur or progress. This "loss of detection" is exemplified in diffuse large B-cell lymphoma (DLBCL), where ctDNA detection after two cycles of curative-intent therapy is a strong prognostic marker. Despite this, nearly one-third of patients experiencing disease progression do not have detectable ctDNA at this landmark, representing a "false-negative" test. Similar false-negative rates have been observed in colon and breast cancer. Summary of the Invention [Means for solving the problem]

[0006] overview The present disclosure provides a method and system for analyzing cell-free nucleic acid (for example, cfDNA, cfRNA) from subject.The method and system of the present disclosure can utilize the sequencing results from subject to detect cancer-derived nucleic acid (for example, ctDNA, ctRNA), for example, for subject disease diagnosis, disease monitoring or treatment determination.The method and system of the present disclosure can show improved sensitivity, specificity and / or reliability of detecting cancer-derived nucleic acid.

[0007] In one aspect, the disclosure provides a method comprising: (a) obtaining, by a computer system, sequencing data from a plurality of cell-free nucleic acid molecules obtained or derived from a subject; (b) processing, by the computer system, the sequencing data to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of phased variants relative to a reference genome sequence, and wherein at least about 10% of the one or more cell-free nucleic acid molecules comprise a first phased variant of the plurality of phased variants and a second phased variant of the plurality of phased variants separated by at least one nucleotide; and (c) analyzing, by the computer system, the identified one or more cell-free nucleic acid molecules to determine a status of the subject.

[0008] In some embodiments of any one of the methods disclosed herein, at least about 10% of the cell-free nucleic acid molecules comprise at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of the one or more cell-free nucleic acid molecules.

[0009] In one aspect, the disclosure provides a method including: (a) obtaining, by a computer system, sequencing data from a plurality of cell-free nucleic acid molecules obtained or derived from a subject; (b) processing, by the computer system, the sequencing data to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of phased variants relative to a reference genome sequence that are separated by at least one nucleotide; and (c) analyzing, by the computer system, the identified one or more cell-free nucleic acid molecules to determine a status of the subject.

[0010] In one aspect, the disclosure provides a method including: (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained or derived from a subject; (b) processing the sequencing data to identify one or more cell-free nucleic acid molecules from the plurality of cell-free nucleic acid molecules at a limit of detection of less than about 1 in 50,000 observations from the sequencing data; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine a status of the subject.

[0011] In some embodiments of any one of the methods disclosed herein, the detection limit of the identification step is less than about 1 in 100,000, less than about 1 in 500,000, less than about 1 in 1 million, less than about 1 in 1.5 million, or less than about 1 in 2 million observations from the sequencing data.

[0012] In some embodiments of any one of the methods disclosed herein, each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence. In some embodiments of any one of the methods disclosed herein, a first graded variant of the plurality of graded variants and a second graded variant of the plurality of graded variants are separated by at least one nucleotide.

[0013] In some embodiments of any one of the methods disclosed herein, processes (a)-(c) are performed by a computer system.

[0014] In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on nucleic acid amplification.In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on polymerase chain reaction.In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on amplicon sequencing.

[0015] In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on next-generation sequencing (NGS). Alternatively, in some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on non-hybridization-based NGS.

[0016] In some embodiments of any one of the methods disclosed herein, the sequencing data is generated without molecular barcoding of at least some of the plurality of cell-free nucleic acid molecules. In some embodiments of any one of the methods disclosed herein, the sequencing data is obtained without sample barcoding of at least some of the plurality of cell-free nucleic acid molecules.

[0017] In some embodiments of any one of the methods disclosed herein, the sequencing data is obtained without in silico removal or suppression of (i) background errors or (ii) sequencing errors.

[0018] In one aspect, the disclosure provides a method of treating a condition in a subject, the method comprising: (a) identifying a subject for treatment for the condition, wherein the subject is determined to have the condition based on identification of one or more cell-free nucleic acid molecules from a plurality of cell-free nucleic acid molecules obtained or derived from the subject, wherein each of the identified one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence that are separated by at least one nucleotide, and the presence of the plurality of graded variants is indicative of the condition in the subject; and (b) subjecting the subject to treatment based on the identification of (a).

[0019] In one aspect, the disclosure provides a method of monitoring the progression of a condition of a subject, the method comprising: (a) determining a first status of the subject's condition based on identification of a first set of one or more cell-free nucleic acid molecules from a first plurality of cell-free nucleic acid molecules obtained or derived from the subject; (b) determining a second status of the subject's condition based on identification of a second set of one or more cell-free nucleic acid molecules from a second plurality of cell-free nucleic acid molecules obtained or derived from the subject, wherein the second plurality of cell-free nucleic acid molecules are obtained from the subject after obtaining the first plurality of cell-free nucleic acid molecules from the subject; and (c) determining the progression of the condition based on the first status of the condition and the second status of the condition, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence separated by at least one nucleotide.

[0020] In some embodiments of any one of the methods disclosed herein, the progression of the condition is a worsening of the condition.

[0021] In some embodiments of any one of the methods disclosed herein, the progression of the condition is at least partial remission of the condition.

[0022] In some embodiments of any one of the methods disclosed herein, the presence of multiple graded variants is indicative of the first status or the second status of the subject's condition.

[0023] In some embodiments of any one of the methods disclosed herein, the second plurality of cell-free nucleic acid molecules is obtained from the subject at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 4 weeks, at least about 2 months, or at least about 3 months after obtaining the first plurality of cell-free nucleic acid molecules.

[0024] In some embodiments of any one of the methods disclosed herein, the subject is subjected to treatment for the condition (i) before obtaining the second plurality of cell-free nucleic acid molecules from the subject, and (ii) after obtaining the first plurality of cell-free nucleic acid molecules from the subject.

[0025] In some embodiments of any one of the methods disclosed herein, the progression of the condition indicates minimal residual disease of the subject's condition. In some embodiments of any one of the methods disclosed herein, the progression of the condition indicates tumor or cancer burden in the subject.

[0026] In some embodiments of any one of the methods disclosed herein, the one or more cell-free nucleic acid molecules are captured from among the plurality of cell-free nucleic acid molecules with a set of nucleic acid probes, wherein the set of nucleic acid probes are configured to hybridize to at least a portion of the cell-free nucleic acid molecule that includes one or more genomic regions associated with the condition.

[0027] In one aspect, the disclosure provides a method comprising: (a) providing a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained or derived from a subject, wherein each individual nucleic acid probe of the set of nucleic acid probes is designed to hybridize to at least a portion of a target cell-free nucleic acid molecule comprising a plurality of graded variants relative to a reference genome sequence separated by at least one nucleotide, and wherein each individual nucleic acid probe comprises an activatable reporter agent, and activation of the activatable reporter agent is selected from the group consisting of: (i) hybridization of each individual nucleic acid probe to the plurality of graded variants; and (ii) dehybridization of at least a portion of each individual nucleic acid probe hybridized to the plurality of graded variants; (b) detecting the activated activatable reporter agent to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine the status of the subject.

[0028] In one aspect, the disclosure provides (a) a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained or derived from a subject, wherein each nucleic acid probe of the set of nucleic acid probes is designed to hybridize to at least a portion of a target cell-free nucleic acid molecule that comprises a plurality of graded variants relative to a reference genome sequence, and each nucleic acid probe comprises an activatable reporter agent, and activation of the activatable reporter agent is determined by (i) hybridization of the individual nucleic acid probe to the plurality of graded variants, and (ii) hybridization of the individual nucleic acid probe to the plurality of graded variants. (b) detecting the activated activatable reporter agent to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants, and the detection limit of the identification step is less than about 1 in 50,000 cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine the status of the subject.

[0029] In some embodiments of any one of the methods disclosed herein, the detection limit of the identifying step is less than about 1 in 100,000, less than about 1 in 500,000, less than about 1 in 1 million, less than about 1 in 1.5 million, or less than about 1 in 2 million cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules.

[0030] In some embodiments of any one of the methods disclosed herein, the first stepped variant of the plurality of stepped variants and the second stepped variant of the plurality of stepped variants are separated by at least one nucleotide.

[0031] In some embodiments of any one of the methods disclosed herein, the activatable reporter agent is activated upon hybridization of an individual nucleic acid probe to the plurality of graded variants.

[0032] In some embodiments of any one of the methods disclosed herein, the activatable reporter agent is activated upon dehybridization of at least some of the individual nucleic acid probes hybridized to the plurality of graded variants.

[0033] In some embodiments of any one of the methods disclosed herein, the method further comprises mixing (1) the set of nucleic acid probes and (2) the plurality of cell-free nucleic acid molecules.

[0034] In some embodiments of any one of the methods disclosed herein, the activatable reporter agent is a fluorophore.

[0035] In some embodiments of any one of the methods disclosed herein, analyzing the identified one or more cell-free nucleic acid molecules comprises analyzing (i) the identified one or more cell-free nucleic acid molecules and (ii) other cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules that do not comprise the plurality of graded variants, as different variables.

[0036] In some embodiments of any one of the methods disclosed herein, the analysis of the identified one or more cell-free nucleic acid molecules is not based on other cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules that do not comprise the plurality of graded variants.

[0037] In some embodiments of any one of the methods disclosed herein, the number of identified graded variants from the one or more cell-free nucleic acid molecules indicates the status of the subject. In some embodiments, the ratio of (i) the number of graded variants from the one or more cell-free nucleic acid molecules to (ii) the number of single base variants (SNVs) from the one or more cell-free nucleic acid molecules indicates the status of the subject.

[0038] In some embodiments of any one of the methods disclosed herein, the frequency of the plurality of graded variants in the identified one or more cell-free nucleic acid molecules indicates the condition of the subject. In some embodiments, the frequency indicates disease cells associated with the condition. In some embodiments, the condition is diffuse large B-cell lymphoma, and the frequency indicates whether the one or more cell-free nucleic acid molecules are derived from germinal center B cells (GCB) or activated B cells (ABC).

[0039] In some embodiments of any one of the methods disclosed herein, the genomic origin of the identified one or more cell-free nucleic acid molecules is indicative of the subject's condition.

[0040] In some embodiments of any one of the methods disclosed herein, the first and second stepwise variants are separated by at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 nucleotides. In some embodiments of any one of the methods disclosed herein, the first and second stepwise variants are separated by at most about 180 nucleotides, at most about 170 nucleotides, at most about 160 nucleotides, at most about 150 nucleotides, or at most about 140 nucleotides.

[0041] In some embodiments of any one of the methods disclosed herein, at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants comprise a single-base variant (SNV) that is at least two nucleotides away from an adjacent SNV.

[0042] In some embodiments of any one of the methods disclosed herein, the plurality of graded variants comprises at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, or at least 25 graded variants within the same cell-free nucleic acid molecule.

[0043] In some embodiments of any one of the methods disclosed herein, the one or more identified cell-free nucleic acid molecules include at least 2, at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 500, or at least 1,000 cell-free nucleic acid molecules.

[0044] In some embodiments of any one of the methods disclosed herein, the reference genome sequence is derived from a reference cohort. In some embodiments, the reference genome sequence comprises a consensus sequence from the reference cohort. In some embodiments, the reference genome sequence comprises at least a portion of the hg19 human genome, the hg18 genome, the hg17 genome, the hg16 genome or the hg38 genome.

[0045] In some embodiments of any one of the methods disclosed herein, the reference genome sequence is derived from a sample of the subject.

[0046] In some embodiments of any one of the methods disclosed herein, the sample is a healthy sample. In some embodiments, the sample comprises healthy cells. In some embodiments, the healthy cells comprise healthy white blood cells.

[0047] In some embodiments of any one of the methods disclosed herein, the sample is a disease sample. In some embodiments, the disease sample comprises disease cells. In some embodiments, the disease cells comprise tumor cells. In some embodiments, the disease sample comprises a solid tumor.

[0048] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes is designed based on a plurality of staged variants identified by comparing (i) sequencing data from a solid tumor, lymphoma, or hematologic tumor of the subject with (ii) sequencing data from healthy cells of the subject or a healthy cohort. In some embodiments, the healthy cells are from the subject. In some embodiments, the healthy cells are from a healthy cohort.

[0049] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes is designed to hybridize to at least a portion of a sequence of a genomic locus associated with a condition, in some embodiments, the genomic locus associated with the condition is known to exhibit aberrant somatic hypermutation when a subject has the condition.

[0050] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes is designed to hybridize to at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants.

[0051] In some embodiments of any one of the methods disclosed herein, each nucleic acid probe of the set of nucleic acid probes has at least about 70%, at least about 80%, at least about 90% sequence identity, at least about 95% sequence identity, or about 100% sequence identity to a probe sequence selected from Table 6.

[0052] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes comprises at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the probe sequences in Table 6.

[0053] In some embodiments of any one of the methods disclosed herein, the method further comprises determining that the subject has a condition, or determining the degree or status of the subject's condition, based on the identified one or more cell-free nucleic acid molecules containing a plurality of graded variants. In some embodiments, the method further comprises determining that the one or more cell-free nucleic acid molecules are derived from a sample associated with the condition based on performing a statistical model analysis of the identified one or more cell-free nucleic acid molecules. In some embodiments, the statistical model analysis comprises Monte Carlo statistical analysis.

[0054] In some embodiments of any one of the methods disclosed herein, the method further includes monitoring the progression of the subject's condition based on the identified one or more cell-free nucleic acid molecules.

[0055] In some embodiments of any one of the methods disclosed herein, the method further comprises performing a different procedure to confirm the subject's condition. In some embodiments, the different procedure comprises a blood test, a genetic test, medical imaging, a physical exam, or a tissue biopsy.

[0056] In some embodiments of any one of the methods disclosed herein, the method further includes determining a treatment for the subject's condition based on the identified one or more cell-free nucleic acid molecules.

[0057] In some embodiments of any one of the methods disclosed herein, the subject has been subjected to treatment for the condition prior to (a).

[0058] In some embodiments of any one of the methods disclosed herein, the treatment comprises chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, adoptive cellular therapy, hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusion, or medical surveillance.

[0059] In some embodiments of any one of the methods disclosed herein, the plurality of cell-free nucleic acid molecules comprises a plurality of cell-free deoxyribonucleic acid (DNA) molecules.

[0060] In some embodiments of any one of the methods disclosed herein, the condition comprises a disease.

[0061] In some embodiments of any one of the methods disclosed herein, the plurality of cell-free nucleic acid molecules are derived from a bodily sample of the subject. In some embodiments, the bodily sample comprises plasma, serum, blood, cerebrospinal fluid, lymph, saliva, urine, or feces.

[0062] In some embodiments of any one of the methods disclosed herein, the subject is a mammal. In some embodiments of any one of the methods disclosed herein, the subject is a human.

[0063] In some embodiments of any one of the methods disclosed herein, the condition comprises a neoplasm, cancer, or tumor. In some embodiments, the condition comprises a solid tumor. In some embodiments, the condition comprises a lymphoma. In some embodiments, the condition comprises a B-cell lymphoma. In some embodiments, the condition comprises a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt's lymphoma, and B-cell chronic lymphocytic leukemia.

[0064] In some embodiments of any one of the methods disclosed herein, the plurality of staged variants have been previously identified in the tumor from sequencing of a previous tumor sample or a cell-free nucleic acid sample.

[0065] In one aspect, the present disclosure provides a composition comprising a bait set comprising a set of nucleic acid probes designed to capture cell-free DNA molecules from at least about 5% of the genomic regions specified in (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants.

[0066] In some embodiments of any of the compositions disclosed herein, the set of nucleic acid probes is designed to pull down cell-free DNA molecules from at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of the genomic regions specified in (i) the genomic regions identified in Table 1, (ii) the genomic regions identified in Table 3, or (iii) the genomic regions identified in Table 3 as having multiple graded variants.

[0067] In some embodiments of any of the compositions disclosed herein, the set of nucleic acid probes is designed to capture one or more cell-free DNA molecules from up to about 10%, up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 80%, up to about 90%, or about 100% of a genomic region specified in (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants.

[0068] In some embodiments of any of the compositions disclosed herein, the bait set comprises at most 5, at most 10, at most 50, at most 100, at most 500, at most 1000, or at most 2000 nucleic acid probes.

[0069] In some embodiments of any of the compositions disclosed herein, an individual nucleic acid probe of the set of nucleic acid probes comprises a pull-down tag.

[0070] In some embodiments of any of the compositions disclosed herein, the pull-down tag comprises a nucleic acid barcode.

[0071] In some embodiments of any of the compositions disclosed herein, the pull-down tag comprises biotin.

[0072] In some embodiments of any of the compositions disclosed herein, each of the cell-free DNA molecules is between about 100 nucleotides and about 180 nucleotides in length.

[0073] In some embodiments of any of the compositions disclosed herein, the genomic region is associated with a condition.

[0074] In some embodiments of any of the compositions disclosed herein, when the subject has the condition, the genomic region exhibits aberrant somatic hypermutation.

[0075] In some embodiments of any of the compositions disclosed herein, the condition comprises a B-cell lymphoma, hi some embodiments, the condition comprises a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt's lymphoma, and B-cell chronic lymphocytic leukemia.

[0076] In some embodiments of any of the compositions disclosed herein, the composition further comprises a plurality of cell-free DNA molecules obtained or derived from the subject.

[0077] In one aspect, the disclosure provides a method of performing a clinical procedure on an individual, the method comprising: (a) obtaining, or having obtained, targeted sequencing results of a collection of cell-free nucleic acid molecules, wherein the collection of cell-free nucleic acid molecules is sourced from a fluid or waste biopsy of the individual, and wherein the targeted sequencing is performed utilizing nucleic acid probes to pull down sequences of genomic loci known to undergo aberrant somatic hypermutation in B-cell cancer; (b) identifying, or having identified, multiple variants in phase within the cell-free nucleic acid sequencing results; (c) determining, or having determined, utilizing a statistical model and the identified phased variants, that the cell-free nucleic acid sequencing results comprise nucleotides derived from a neoplasm; and (d) performing the clinical procedure on the individual to confirm the presence of B-cell cancer based on determining that the cell-free nucleic acid sequencing results comprise nucleic acid sequences likely derived from B-cell cancer.

[0078] In some embodiments of any of the compositions disclosed herein, the biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces.

[0079] In some embodiments of any of the compositions disclosed herein, the genomic locus is selected from (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants.

[0080] In some embodiments of any of the compositions disclosed herein, the sequence of the nucleic acid probe is selected from Table 6.

[0081] In some embodiments of any of the compositions disclosed herein, the clinical is procedure is a blood test, medical imaging, or a physical exam.

[0082] In one aspect, the disclosure provides a method of treating an individual for B-cell cancer, the method comprising: (a) obtaining or having obtained targeted sequencing results of a collection of cell-free nucleic acid molecules, wherein the collection of cell-free nucleic acid molecules is sourced from a fluid or waste biopsy of the individual, and the targeted sequencing is performed utilizing nucleic acid probes to pull down sequences of genomic loci known to undergo aberrant somatic hypermutation in B-cell cancer; (b) identifying or having identified multiple variants in phase within the cell-free nucleic acid sequencing results; (c) determining or having determined, utilizing a statistical model and the identified phased variants, that the cell-free nucleic acid sequencing results comprise nucleotides derived from the neoplasm; and (d) treating the individual to reduce the B-cell cancer based on determining that the cell-free nucleic acid sequencing results comprise nucleic acid sequences derived from the B-cell cancer.

[0083] In some embodiments of any of the compositions disclosed herein, the biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces.

[0084] In some embodiments of any of the compositions disclosed herein, the genomic locus is selected from (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants.

[0085] In some embodiments of any of the compositions disclosed herein, the sequence of the nucleic acid probe is selected from Table 6.

[0086] In some embodiments of any of the compositions disclosed herein, the treatment is chemotherapy, radiation therapy, immunotherapy, hormone therapy, targeted drug therapy, or medical surveillance.

[0087] In one aspect, the disclosure provides a method of detecting cancerous minimal residual disease in an individual and treating the individual for cancer, the method comprising: (a) obtaining, or having obtained, targeted sequencing results of a collection of cell-free nucleic acid molecules, wherein the collection of cell-free nucleic acid molecules is sourced from a fluid or waste biopsy of the individual, wherein the fluid or waste biopsy is sourced after a series of treatments to detect minimal residual disease, and wherein the targeted sequencing is performed utilizing nucleic acid probes to pull down sequences of genomic loci determined to contain multiple variants in phase, as determined by previous sequencing results on the previous biopsy derived from the cancer; (b) identifying, or having identified, at least one set of multiple variants in phase within the cell-free nucleic acid sequencing results; and (c) treating the individual to reduce the cancer based on determining that the cell-free nucleic acid sequencing results contain nucleic acid sequences derived from the cancer.

[0088] In some embodiments of any of the compositions disclosed herein, the fluid or waste biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces.

[0089] In some embodiments of any of the compositions disclosed herein, the treatment is chemotherapy, radiation therapy, immunotherapy, hormone therapy, targeted drug therapy, or medical surveillance.

[0090] In one aspect, the present disclosure provides a computer program product, the computer program product including a non-transitory computer-readable medium having computer-executable code encoded therein, the computer-executable code adapted and executed to implement any one of the methods disclosed herein.

[0091] In one aspect, the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, the computer memory comprising machine-executable code that, when executed by the one or more computer processors, implements any one of the methods disclosed herein.

[0092]

[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. Incorporation by Reference

[0093] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, and patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or supersede such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The various features of the present invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG."). In certain embodiments, for example, the following items are provided: (Item 1) (a) obtaining, with a computer system, sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from a subject; (b) processing, with the computer system, the sequencing data to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, each of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants relative to a reference genome sequence, and at least about 10% of the one or more cell-free nucleic acid molecules comprising a first graded variant of the plurality of graded variants and a second graded variant of the plurality of graded variants separated by at least one nucleotide; (c) analyzing, by the computer system, the identified one or more cell-free nucleic acid molecules to determine a status of the subject. (Item 2) 2. The method of claim 1, wherein said at least about 10% of the cell-free nucleic acid molecules comprises at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of said one or more cell-free nucleic acid molecules. (Item 3) (a) obtaining, with a computer system, sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from a subject; (b) processing, with the computer system, the sequencing data to identify one or more cell-free nucleic acid molecules among the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence that are separated by at least one nucleotide; (c) analyzing, by the computer system, the identified one or more cell-free nucleic acid molecules to determine a status of the subject. (Item 4) (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from a subject; (b) processing the sequencing data to identify one or more cell-free nucleic acid molecules among the plurality of cell-free nucleic acid molecules at a limit of detection of less than about 1 in 50,000 observations from the sequencing data; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine a status of the subject. (Item 5) 5. The method of claim 4, wherein the detection limit of the identifying step is less than about 1 in 100,000, less than about 1 in 500,000, less than about 1 in 1 million, less than about 1 in 1.5 million, or less than about 1 in 2 million observations from the sequencing data. (Item 6) 6. The method of any one of items 4 to 5, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence. (Item 7) 7. The method of claim 6, wherein a first stepped variant of the plurality of stepped variants and a second stepped variant of the plurality of stepped variants are separated by at least one nucleotide. (Item 8) The method according to any one of items 4 to 7, wherein (a) to (c) are performed by a computer system. (Item 9) 10. The method of any one of the preceding items, wherein the sequencing data is generated based on nucleic acid amplification. (Item 10) 10. The method of any one of the preceding items, wherein the sequencing data is generated based on polymerase chain reaction. (Item 11) 10. The method of any one of the preceding items, wherein the sequencing data is generated based on amplicon sequencing. (Item 12) 10. The method of any one of the preceding items, wherein the sequencing data is generated based on next generation sequencing (NGS). (Item 13) 10. The method of any one of the preceding items, wherein the sequencing data is generated based on non-hybridization-based NGS. (Item 14) 10. The method of any one of the preceding items, wherein the sequencing data is generated without the use of molecular barcoding of at least some of the plurality of cell-free nucleic acid molecules. (Item 15) 10. The method of any one of the preceding items, wherein the sequencing data is obtained without sample barcoding of at least some of the plurality of cell-free nucleic acid molecules. (Item 16) 10. The method of any one of the preceding items, wherein the sequencing data is obtained without in silico removal or suppression of (i) background errors or (ii) sequencing errors. (Item 17) 1. A method of treating a condition in a subject, comprising: (a) identifying the subject for treatment of the condition, wherein the subject is determined to have the condition based on identification of one or more cell-free nucleic acid molecules from a plurality of cell-free nucleic acid molecules obtained or derived from the subject; each of the identified one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence that are separated by at least one nucleotide; identifying, wherein the presence of said plurality of graded variants is indicative of said condition in said subject; (b) subjecting said subject to said treatment based on said identification of (a). (Item 18) 1. A method for monitoring the progression of a condition in a subject, comprising: (a) determining a first status of the condition in the subject based on identification of a first set of one or more cell-free nucleic acid molecules from a first plurality of cell-free nucleic acid molecules obtained or derived from the subject; (b) determining a second state of the condition in the subject based on the identification of a second set of one or more cell-free nucleic acid molecules from a second plurality of cell-free nucleic acid molecules obtained or derived from the subject, determining that the second plurality of cell-free nucleic acid molecules is obtained from the subject after obtaining the first plurality of cell-free nucleic acid molecules from the subject; (c) determining the progression of the condition based on the first status of the condition and the second status of the condition; determining that each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence that are separated by at least one nucleotide. (Item 19) 19. The method of claim 18, wherein the progression of the condition is a worsening of the condition. (Item 20) 19. The method of claim 18, wherein the progression of the condition is at least partial remission of the condition. (Item 21) 21. The method according to any one of items 18 to 20, wherein the presence of the plurality of graded variants indicates the first status or the second status of the condition of the subject. (Item 22) 22. The method of any one of Items 18 to 21, wherein the second plurality of cell-free nucleic acid molecules is obtained from the subject at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 4 weeks, at least about 2 months, or at least about 3 months after obtaining the first plurality of cell-free nucleic acid molecules from the subject. (Item 23) 23. The method of any one of items 18 to 22, wherein the subject is subjected to treatment for the condition (i) before obtaining the second plurality of cell-free nucleic acid molecules from the subject, and (ii) after obtaining the first plurality of cell-free nucleic acid molecules from the subject. (Item 24) 24. The method of any one of items 18 to 23, wherein said progression of said condition indicates minimal residual disease of said condition in said subject. (Item 25) 25. The method of any one of items 18 to 24, wherein the progression of the condition indicates tumor burden or cancer burden in the subject. (Item 26) 10. The method of any one of the preceding items, wherein the one or more cell-free nucleic acid molecules are captured from among the plurality of cell-free nucleic acid molecules with a set of nucleic acid probes, the set of nucleic acid probes configured to hybridize to at least a portion of the cell-free nucleic acid molecule comprising one or more genomic regions associated with the condition. (Item 27) (a) providing a mixture comprising: (1) a set of nucleic acid probes; and (2) a plurality of cell-free nucleic acid molecules obtained or derived from a subject; each nucleic acid probe of the set of nucleic acid probes is designed to hybridize to at least a portion of a target cell-free nucleic acid molecule that comprises a plurality of graded variants relative to a reference genome sequence that are separated by at least one nucleotide; the individual nucleic acid probes comprise an activatable reporter agent, and activation of the activatable reporter agent is selected from the group consisting of: (i) hybridization of the individual nucleic acid probes to the plurality of graded variants, and (ii) dehybridization of at least a portion of the individual nucleic acid probes hybridized to the plurality of graded variants; (b) detecting the activated activatable reporter agent to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, each of the one or more cell-free nucleic acid molecules comprising the plurality of graded variants; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine the status of the subject. (Item 28) (a) providing a mixture comprising: (1) a set of nucleic acid probes; and (2) a plurality of cell-free nucleic acid molecules obtained or derived from a subject; each nucleic acid probe of the set of nucleic acid probes is designed to hybridize to at least a portion of a target cell-free nucleic acid molecule that comprises a plurality of graded variants relative to a reference genome sequence; the individual nucleic acid probes comprise an activatable reporter agent, and activation of the activatable reporter agent is selected from the group consisting of: (i) hybridization of the individual nucleic acid probes to the plurality of graded variants, and (ii) dehybridization of at least a portion of the individual nucleic acid probes hybridized to the plurality of graded variants; (b) detecting the activated activatable reporter agent to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, each of the one or more cell-free nucleic acid molecules comprising the plurality of graded variants, wherein the detection limit of the identifying step is less than about 1 in 50,000 cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules; (c) analyzing the identified one or more cell-free nucleic acid molecules to determine the status of the subject. (Item 29) 29. The method of claim 28, wherein the detection limit of the identifying step is less than about 1 in 100,000, less than about 1 in 500,000, less than about 1 in 1 million, less than about 1 in 1.5 million, or less than about 1 in 2 million cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules. (Item 30) 30. The method of claim 28 or 29, wherein a first stepped variant of the plurality of stepped variants and a second stepped variant of the plurality of stepped variants are separated by at least one nucleotide. (Item 31) 31. The method of any one of items 27 to 30, wherein the activatable reporter agent is activated upon hybridization of the individual nucleic acid probes to the plurality of graded variants. (Item 32) 31. The method of any one of items 27 to 30, wherein the activatable reporter agent is activated upon dehybridization of at least a portion of the individual nucleic acid probes hybridized to the plurality of stepwise variants. (Item 33) 33. The method according to any one of Items 27 to 32, further comprising mixing (1) the set of nucleic acid probes and (2) the plurality of cell-free nucleic acid molecules. (Item 34) 34. The method of any one of items 27 to 33, wherein the activatable reporter agent is a fluorophore. (Item 35) The method of any one of the preceding items, wherein analyzing the identified one or more cell-free nucleic acid molecules comprises analyzing, as different variables, (i) the identified one or more cell-free nucleic acid molecules and (ii) other cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules that do not contain the plurality of graded variants. (Item 36) 10. The method of any one of the preceding items, wherein the analysis of the identified one or more cell-free nucleic acid molecules is not based on other cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules that do not contain the plurality of graded variants. (Item 37) 10. The method of any one of the preceding items, wherein the number of graded variants from the identified one or more cell-free nucleic acid molecules is indicative of the condition of the subject. (Item 38) 38. The method of claim 37, wherein a ratio of (i) the number of the plurality of graded variants from the one or more cell-free nucleic acid molecules to (ii) the number of single base variants (SNVs) from the one or more cell-free nucleic acid molecules is indicative of the condition of the subject. (Item 39) 2. The method of any one of the preceding items, wherein the frequency of the plurality of graded variants in the identified one or more cell-free nucleic acid molecules is indicative of the condition in the subject. (Item 40) 40. The method of claim 39, wherein the frequency indicates disease cells associated with the condition. (Item 41) 41. The method of item 40, wherein the condition is diffuse large B-cell lymphoma and the frequency indicates whether the one or more cell-free nucleic acid molecules are derived from a germinal center B cell (GCB) or an activated B cell (ABC). (Item 42) 2. The method of any one of the preceding items, wherein the genomic origin of the identified one or more cell-free nucleic acid molecules is indicative of the condition in the subject. (Item 43) 10. The method of any one of the preceding items, wherein the first graded variant and the second graded variant are separated by at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 nucleotides. (Item 44) 10. The method of any one of the preceding items, wherein the first and second graded variants are separated by at most about 180 nucleotides, at most about 170 nucleotides, at most about 160 nucleotides, at most about 150 nucleotides, or at most about 140 nucleotides. (Item 45) 2. The method of any one of the preceding items, wherein at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants comprise a single nucleotide variant (SNV) that is at least two nucleotides away from an adjacent SNV. (Item 46) 10. The method of any one of the preceding items, wherein the plurality of graded variants comprises at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, or at least 25 graded variants within the same cell-free nucleic acid molecule. (Item 47) 2. The method of any one of the preceding items, wherein the one or more identified cell-free nucleic acid molecules comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 500, or at least 1,000 cell-free nucleic acid molecules. (Item 48) 10. The method of any one of the preceding items, wherein the reference genome sequence is derived from a reference cohort. (Item 49) 49. The method of claim 48, wherein the reference genome sequence comprises a consensus sequence from the reference cohort. (Item 50) 49. The method of claim 48, wherein the reference genome sequence comprises at least a portion of the hg19 human genome, the hg18 genome, the hg17 genome, the hg16 genome, or the hg38 genome. (Item 51) 10. The method of any one of the preceding items, wherein the reference genome sequence is derived from a sample of the subject. (Item 52) 52. The method of claim 51, wherein the sample is a healthy sample. (Item 53) 53. The method of claim 52, wherein the sample comprises healthy cells. (Item 54) 54. The method of claim 53, wherein the healthy cells comprise healthy white blood cells. (Item 55) 52. The method of claim 51, wherein the sample is a disease sample. (Item 56) 56. The method of claim 55, wherein the disease sample comprises disease cells. (Item 57) 57. The method of claim 56, wherein the diseased cells comprise tumor cells. (Item 58) 56. The method of claim 55, wherein the disease sample comprises a solid tumor. (Item 59) 10. The method of any one of the preceding items, wherein the set of nucleic acid probes is designed based on the plurality of graded variants identified by comparing (i) sequencing data from a solid tumor, lymphoma, or hematological tumor of the subject and (ii) sequencing data from healthy cells of the subject or a healthy cohort. (Item 60) 60. The method of claim 59, wherein the healthy cells are from the subject. (Item 61) 60. The method of claim 59, wherein the healthy cells are from the healthy cohort. (Item 62) 10. The method of any one of the preceding items, wherein the set of nucleic acid probes is designed to hybridize to at least a portion of the sequence of a genomic locus associated with the condition. (Item 63) 63. The method of item 62, wherein the genomic locus associated with the condition is known to exhibit aberrant somatic hypermutation when the subject has the condition. (Item 64) The method of any one of the preceding items, wherein the set of nucleic acid probes is designed to hybridize to at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants. (Item 65) 10. The method of any one of the preceding items, wherein each nucleic acid probe of the set of nucleic acid probes has at least about 70%, at least about 80%, at least about 90% sequence identity, at least about 95% sequence identity, or about 100% sequence identity to a probe sequence selected from Table 6. (Item 66) 7. The method of any one of the preceding items, wherein the set of nucleic acid probes comprises at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the probe sequences in Table 6. (Item 67) The method of any one of the preceding items, further comprising determining that the subject has the condition, or determining the degree or status of the condition in the subject, based on the identified one or more cell-free nucleic acid molecules comprising the plurality of graded variants. (Item 68) 68. The method of claim 67, further comprising determining that the one or more cell-free nucleic acid molecules are derived from a sample associated with the condition based on performing a statistical model analysis of the identified one or more cell-free nucleic acid molecules. (Item 69) Item 69. The method of item 68, wherein the statistical model analysis comprises Monte Carlo statistical analysis. (Item 70) The method of any one of the preceding items, further comprising monitoring the progression of the condition in the subject based on the identified one or more cell-free nucleic acid molecules. (Item 71) 10. The method of any one of the preceding items, further comprising performing a different procedure to confirm the condition in the subject. (Item 72) 72. The method of claim 71, wherein the different procedure comprises a blood test, a genetic test, medical imaging, a physical examination, or a tissue biopsy. (Item 73) 10. The method of any one of the preceding items, further comprising determining a treatment for the condition in the subject based on the identified one or more cell-free nucleic acid molecules. (Item 74) The method of any one of the preceding items, wherein the subject is subjected to treatment for the condition prior to (a). (Item 75) The method of any one of the preceding items, wherein the treatment comprises chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, adoptive cell therapy, hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusion, or medical surveillance. (Item 76) 10. The method of any one of the preceding items, wherein the plurality of cell-free nucleic acid molecules comprises a plurality of cell-free deoxyribonucleic acid (DNA) molecules. (Item 77) The method of any one of the preceding items, wherein the condition comprises a disease. (Item 78) 10. The method of any one of the preceding items, wherein the plurality of cell-free nucleic acid molecules is derived from a bodily sample of the subject. (Item 79) 80. The method of claim 78, wherein the body sample comprises plasma, serum, blood, cerebrospinal fluid, lymph, saliva, urine, or feces. (Item 80) 10. The method of any one of the preceding items, wherein the subject is a mammal. (Item 81) The method of any one of the preceding items, wherein the subject is a human. (Item 82) The method of any one of the preceding items, wherein the condition comprises a neoplasm, cancer, or tumor. (Item 83) 83. The method of claim 82, wherein the condition comprises a solid tumor. (Item 84) 83. The method of claim 82, wherein the condition comprises lymphoma. (Item 85) 85. The method of item 84, wherein the condition comprises B-cell lymphoma. (Item 86) 86. The method of item 85, wherein the condition comprises a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt's lymphoma, and B-cell chronic lymphocytic leukemia. (Item 87) 10. The method of any one of the preceding items, wherein the plurality of staged variants have been previously identified as tumors from sequencing of a previous tumor sample or a cell-free nucleic acid sample. (Item 88) A composition comprising a bait set comprising a set of nucleic acid probes designed to capture cell-free DNA molecules from at least about 5% of the genomic regions specified in (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple phased variants. (Item 89) 89. The composition of claim 88, wherein the set of nucleic acid probes is designed to pull down cell-free DNA molecules from at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of the genomic regions specified in (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants. (Item 90) 90. The composition of any one of paragraphs 88-89, wherein the set of nucleic acid probes is designed to capture one or more cell-free DNA molecules from up to about 10%, up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 80%, up to about 90%, or up to about 100% of the genomic regions specified in (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants. (Item 91) 91. The composition of any one of items 88 to 90, wherein the bait set comprises at most 5, at most 10, at most 50, at most 100, at most 500, at most 1000, or at most 2000 nucleic acid probes. (Item 92) 92. The composition of any one of items 88 to 91, wherein each nucleic acid probe of the set of nucleic acid probes comprises a pull-down tag. (Item 93) 93. The composition of any one of items 88 to 92, wherein the pull-down tag comprises a nucleic acid barcode. (Item 94) 94. The composition of any one of items 88 to 93, wherein the pull-down tag comprises biotin. (Item 95) 95. The composition of any one of items 88 to 94, wherein each of the cell-free DNA molecules is from about 100 nucleotides to about 180 nucleotides in length. (Item 96) 96. The composition of any one of items 88 to 95, wherein the genomic region is associated with a condition. (Item 97) 97. The composition of any one of items 88 to 96, wherein the genomic region exhibits aberrant somatic hypermutation when the subject has the condition. (Item 98) 98. The composition of any one of items 88 to 97, wherein the condition comprises B-cell lymphoma. (Item 99) 99. The composition of item 98, wherein the condition comprises a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt's lymphoma, and B-cell chronic lymphocytic leukemia. (Item 100) 99. The composition of any one of items 88-99, further comprising a plurality of cell-free DNA molecules obtained or derived from a subject. (Item 101) 1. A method of performing a clinical procedure on an individual, comprising: Obtaining or having obtained targeted sequencing results of a collection of cell-free nucleic acid molecules, the collection of cell-free nucleic acid molecules is sourced from a fluid or waste biopsy of an individual; and the targeted sequencing is performed, obtained, or obtained utilizing nucleic acid probes to pull down sequences of genomic loci known to undergo aberrant somatic hypermutation in B-cell cancers; Identifying or having identified multiple variants in phase within the cell-free nucleic acid sequencing results; Utilizing a statistical model and the identified stepwise variants, determining, or having determined, that the cell-free nucleic acid sequencing results comprise nucleotides from a neoplasm; and and performing a clinical procedure on the individual to confirm the presence of the B-cell cancer based on determining that the cell-free nucleic acid sequencing results contain nucleic acid sequences that may be derived from the B-cell cancer. (Item 102) 102. The method of claim 101, wherein the biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces. (Item 103) 102. The method of claim 101, wherein the genomic locus is selected from (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants. (Item 104) 102. The method of claim 101, wherein the sequence of the nucleic acid probe is selected from Table 6. (Item 105) 102. The method of claim 101, wherein the clinical procedure is a blood test, medical imaging, or physical examination. (Item 106) 1. A method of treating an individual for B-cell cancer, comprising: Obtaining or having obtained targeted sequencing results of a collection of cell-free nucleic acid molecules, the collection of cell-free nucleic acid molecules is sourced from a fluid or waste biopsy of an individual; and the targeted sequencing is performed, obtained, or obtained utilizing nucleic acid probes to pull down sequences of genomic loci known to undergo aberrant somatic hypermutation in B-cell cancers; Identifying or having identified multiple variants in phase within the cell-free nucleic acid sequencing results; Utilizing a statistical model and the identified stepwise variants, determining, or having determined, that the cell-free nucleic acid sequencing results comprise nucleotides from a neoplasm; and and treating the individual to reduce the B-cell cancer based on determining that the cell-free nucleic acid sequencing results contain nucleic acid sequences derived from the B-cell cancer. (Item 107) 107. The method of claim 106, wherein the biopsy is one of blood, serum, cerebrospinal fluid, lymphatic fluid, urine, or feces. (Item 108) 107. The method of claim 106, wherein the genomic locus is selected from (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple graded variants. (Item 109) 107. The method of claim 106, wherein the sequence of the nucleic acid probe is selected from Table 6. (Item 110) 107. The method of claim 106, wherein the treatment is chemotherapy, radiation therapy, immunotherapy, hormone therapy, targeted drug therapy, or medical surveillance. (Item 111) 1. A method of detecting cancerous minimal residual disease in an individual and treating said individual for cancer, comprising: Obtaining or having obtained targeted sequencing results of a collection of cell-free nucleic acid molecules, said collection of cell-free nucleic acid molecules is sourced from a fluid or waste biopsy of an individual; The fluid or waste biopsy is provided after a course of treatment to detect minimal residual disease; and the targeted sequencing is performed, obtained, or obtained utilizing nucleic acid probes to pull down sequences of genomic loci determined to contain multiple variants in phase, as determined by previous sequencing results on a previous biopsy from the cancer; Identifying or having identified at least one set of said plurality of variants in phase within said cell-free nucleic acid sequencing results; and treating the individual to reduce the cancer based on determining that the cell-free nucleic acid sequencing results contain nucleic acid sequences derived from the cancer. (Item 112) 112. The method of claim 111, wherein the fluid or waste biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces. (Item 113) 112. The method of claim 111, wherein the treatment is chemotherapy, radiation therapy, immunotherapy, hormone therapy, targeted drug therapy, or medical surveillance. (Item 114) 1. A computer program product comprising: a non-transitory computer-readable medium having computer-executable code encoded therein, the computer-executable code adapted and executed to implement a method according to any one of the preceding items. (Item 115) 1. A system comprising one or more computer processors and a computer memory coupled thereto, the computer memory comprising machine-executable code that, when executed by the one or more computer processors, implements a method according to any one of the preceding items. [Brief explanation of the drawings]

[0095] [Figure 1-1]Figure 1 illustrates the discovery of phased variants and their mutational signatures through analysis of whole-genome sequencing data. Figure 1A is a schematic diagram showing the difference in detection between single-nucleotide variants (SNVs) (top) and multiple "in-phase" variants (phased variants, PVs; bottom) on individual cell-free DNA molecules. Theoretically, the detection of a PV is a more specific event than the detection of a single SNV. Figure 1B is a scatterplot showing the distribution of PV counts from WGS data for 24 different cancer histologies, normalized by the total number of SNVs. Bars indicate the median and interquartile range. (FL-NHL, follicular lymphoma; DLBCL-NHL, diffuse large B-cell lymphoma; Burkitt-NHL, Burkitt lymphoma; Lung-SCC, squamous cell lung cancer; Lung-Adeno, lung adenocarcinoma; Kidney-RCC, renal cell carcinoma; Bone-Os teosarc, osteosarcoma; Liver-HCC, hepatocellular carcinoma; Breast-Adeno, pancreatic adenocarcinoma; Head-SCC, head and neck squamous cell carcinoma; Ovary-Adeno, ovarian adenocarcinoma; Eso-Adeno, esophageal adenocarcinoma; Uterus-Ad (Eneno, uterine adenocarcinoma; Stomach-Adeno, gastric adenocarcinoma; CLL, chronic lymphocytic leukemia; ColoRect-Adeno, colorectal adenocarcinoma; Prost-Adeno, prostate adenocarcinoma; CNS-GBM, glioblastoma multiforme; Panc-Endocrine, pancreatic neuroendocrine tumor; Thy-Adeno, thyroid adenocarcinoma; CNS-PiloAstro, piloastrocytoma; CNS-Medullo, medulloblastoma.) Figure 1C is a heatmap showing the enrichment of single-nucleotide substitution (SBS) mutation signatures of PVs for single SNVs across multiple cancer types. Blue represents signatures enriched for PVs in a particular histology. Darker gray represents unphased, single-SNV enriched signatures. Red represents singly occurring SNVs. Only signatures with significant differences between PVs and unphased SNVs after correcting for multiple hypotheses are shown. Other signatures are in grey. Signatures associated with smoking, AID / AICDA, and APOBEC are shown.Figure 1D demonstrates a bar plot showing the distribution of PVs occurring in typological regions across the genome in B-lymphoid malignancies and lung adenocarcinoma. In this plot, the genome was divided into 1000-bp bins, and the proportion of samples of a given histology that contained PVs in each 1000-bp bin was calculated. Only bins with a recurrence frequency of at least 2% in any cancer subtype are shown. Significant genomic loci are also labeled. Figure 1E compares double-stranded sequencing with stepwise variant sequencing. A schematic diagram comparing error-suppression sequencing versus stepwise variant recovery by double-stranded sequencing. Double-stranded sequencing requires the recovery of a single SNV observed on both strands of the original DNA duplex (i.e., in trans). This requires the independent recovery of two molecules by sequencing, as the plus and minus strands of the original DNA molecule are independently passed through library preparation and PCR. In contrast, PV recovery requires multiple SNVs observed on the same single strand of DNA (i.e., in cis). Therefore, recovery of only the plus or minus strand (but not both) is sufficient for identification of PV. [Figure 1-2] Same as above. [Figure 1-3] Same as above.

[0096] [Figure 2-1]Figure 2 illustrates the design, validation, and application of stepwise variant enrichment sequencing. Figure 2A is a schematic of the design for PhaseED-Seq. WGS data from DLBCL tumor samples were aggregated (left), and regions of putative recurrent PVs were identified (center). An assay was then designed to capture the genomic regions most recurrently containing PVs (right), resulting in approximately 7,500-fold enrichment in PVs compared to WGS. The top right panel shows the in silico expected number of PVs per case per kilobase of panel size (y-axis) for increasing panel size (x-axis). Dashed lines indicate selected regions in the PhaseED-Seq panel. The bottom right panel shows the expected total number of PVs per case (y-axis, assessed in silico from WGS data) for increasing panel size (y-axis). The darkened areas indicate selected regions in the PhaseED-Seq panel. Figure 2B illustrates two panels showing the SNV (left) and PV (right) yields for sequencing tumor DNA and matched germline DNA using a previously established lymphoma CAPP-Seq panel or a PhaseED-Seq panel. Values ​​are assessed in silico by confining WGS to the target space of interest. PVs reported in the right panel include doublet, triplet, and quadruplet phased events. Figure 2C, similar to Figure 2B, shows the SNV (left) and PV (right) yields from experimental sequencing of tumor and / or cell-free DNA from CAPP-Seq versus PhaseED-Seq. Figure 2D is a scatterplot showing the frequency of PVs by genomic location (in 1000-bp bins) for patients with DLBCL identified by WGS or PhaseED-Seq. PVs from IGH, BCL2, MYC, and BCL6 are highlighted. Figure 2E illustrates a scatterplot comparing the frequency of PVs by genomic location (in 50-bp bins) for patients with different types of lymphoma. Colored circles indicate the relative frequency of PVs in 50-bp bins from specific genes of interest. Other (gray) circles indicate the relative frequency of PVs in 50-bp bins from the remainder of the PhasED-Seq sequencing panel.Figure 2F illustrates a volcano plot summarizing the differences in relative frequency of PVs at specific loci between lymphoma types, including ABC-DLBCL vs. GCB-DLBCL (dark gray, left), PMBCL vs. DLBCL (dark gray, center), and HL vs. DLBCL (dark gray, right). The x-axis indicates the relative enrichment of PVs at specific loci, and the y-axis indicates the statistical significance of this association. (Example 10). [Figure 2-2] Same as above. [Figure 2-3] Same as above. [Figure 2-4] Same as above.

[0097] [Figure 3-1]Figure 3 illustrates the technical performance of PhaseED-Seq for disease detection. Figure 3A illustrates a bar plot showing the performance of hybrid capture sequencing for the recovery of synthetic 150-bp oligonucleotides from two loci (MYC and BCL6) with increasing levels of mutations / non-reference bases. Error bars represent 95% confidence intervals (n = 3 replicates of each condition in separate samples). Figure 3B illustrates a plot revealing the background error rates for different types of error suppression (Example 10) from 12 healthy control cell-free DNA samples sequenced with the PhaseED-Seq panel. "PhasED-Seq 2x" or "doublet" represents the detection of two mutations in phase on the same DNA molecule. "PhasED-Seq 3x" or "triplet" represents the detection of three mutations in phase on the same DNA molecule. Figure 3C illustrates a bar plot showing the depth of unique molecule recovery (e.g., depth after barcode-mediated PCR duplicate removal) from sequencing data from 12 cell-free DNA samples for different types of error suppression, including barcode deduplication, double-stranded sequencing, and recovery of PVs with increasing maximum distances between in-phase SNVs. Figure 3D illustrates a bar plot showing the cumulative percentage of PVs with a maximum distance between SNVs less than the number of base pairs indicated on the x-axis. Figure 3E illustrates a plot showing the results of a limiting dilution series simulating cell-free DNA samples containing patient-specific tumor fractions ranging from 1 x 10 to 0.5 x 10. cfDNA from three independent patient samples was used at each dilution. The same sequencing data was analyzed using various error suppression methods for recovery of expected tumor fractions, including iDES, double-stranded sequencing, and PhasED-Seq (both for recovery of doublet and triplet molecules). Points and error bars represent the mean, minimum, and maximum across the three patient-specific tumor mutations considered. Differences between observed and expected tumor fractions for samples <1:10,000 were compared by paired t test. *, P<0.05; **, P<0.005; ***, P<0.0005.Figure 3F illustrates a plot showing the background signal for the detection of tumor-specific alleles in 12 unrelated healthy cell-free DNA samples and a healthy cfDNA sample used in a limiting dilution series (n = 13 total samples). For each sample, tumor-specific SNVs or PVs from the three patient samples utilized in the limiting dilution experiment shown in Figure 3E were evaluated for a total of 39 assessments. Bars represent the arithmetic mean across all 39 assessments. Statistical comparisons were performed using the Wilcoxon rank-sum test. *, P < 0.05; **, P < 0.005; ***, P < 0.0005. Figure 3G illustrates a plot showing the theoretical detection rate for samples with a given number of PV-containing regions by simple binomial sampling. This plot is generated by assuming a unique sequencing depth of 5000x (line) along with a varying number of independent 150-bp PV-containing regions, ranging from 3 regions (blue) to 67 regions (purple). The confidence envelope considers a depth of 4000-6000x. A 5% false-positive rate is also assumed. Figure 3H illustrates a plot showing the observed detection rate (y-axis) for samples of a given true tumor fraction (x-axis) with various numbers of PV-containing regions. For each number of tumor reporter regions, ranging from 3 to 67, 150-bp windows of this number were randomly sampled 25 times from each of the three patient-specific PV reporter lists and used to evaluate tumor detection at each dilution. Solid points represent "wet" dilution series experiments, and open points represent in silico dilution experiments. Points and error bars represent the mean, minimum, and maximum of the three patient-specific PV reporter lists used in the original sampling. Figure 3I illustrates a scatterplot comparing the predicted and observed detection rates for samples from the dilution series shown in panels 3G and 3H. Further details of this experiment are provided in Example 10. [Figure 3-2] Same as above. [Figure 3-3] Same as above.

[0098] [Figure 4-1]Figure 4 illustrates the clinical application of PhaseED-Seq for ultrasensitive disease detection and response monitoring in DLBCL. Figure 4A illustrates a plot showing ctDNA levels in DLBCL patients who responded to first-line immunochemotherapy and subsequently relapsed. Levels measured by CAPP-Seq are indicated by darker gray circles, while levels measured by PhaseED-Seq are indicated by lighter gray circles. Open circles represent undetectable levels by CAPP-Seq. Figure 4B illustrates a univariate scatter plot showing the mean tumor allele fraction measured by PhaseED-Seq for clinical samples at the time of minimal disease (i.e., after one or two cycles of treatment). The plot is divided into samples detected and undetected by standard CAPP-Seq. P values ​​from Wilcoxon rank-sum tests. Figure 4C illustrates bar plots showing the percentage of DLBCL patients with detectable ctDNA by CAPP-Seq after one or two cycles of treatment (dark gray bars), as well as the additional percentage of patients with detectable disease when PhaseED-Seq was added to standard CAPP-Seq (medium gray bars). P values ​​represent Fisher's exact test for detection by CAPP-Seq alone versus the combination of PhaseED-Seq and CAPP-Seq in 171 samples after one or two cycles of treatment. Figure 4D illustrates a waterfall plot showing the change in ctDNA levels measured by CAPP-Seq after two cycles of first-line treatment in DLBCL patients. Patients with undetectable ctDNA by CAPP-Seq are indicated in darker colors as "ND" ("not detected"). The bar color also indicates the final clinical outcome of these patients. Figure 4E illustrates a Kaplan-Meier plot showing relapse-free survival for 52 DLBCL patients with undetectable ctDNA measured by CAPP-Seq after two cycles.Figure 4F illustrates a Kaplan-Meier plot showing recurrence-free survival for the 52 patients shown in Figure 4E (ctDNA undetectable by CAPP-Seq) stratified by ctDNA detection by PhaseED-Seq at this same time point (cycle 3, day 1). Figure 4G illustrates a Kaplan-Meier plot showing recurrence-free survival for 89 DLBCL patients stratified by ctDNA at cycle 3 and day 1 into three strata—patients failing to achieve a major molecular response (dark gray), patients with a major molecular response and still detectable ctDNA by PhaseED-Seq and / or CAPP-Seq (light gray), and patients with stringent molecular remission (ctDNA undetectable by PhaseED-Seq and CAPP-Seq; medium gray). [Figure 4-2] Same as above.

[0099] [Figure 5]Figure 5 illustrates the enumeration of SNVs and PVs in diverse cancers from WGS. Figures 5A-5C show univariate scatter plots showing the number of SNVs (Figure 5A), PVs (Figure 5B), and PVs controlling the total number of SNVs (Figure 5C) from WGS data for cancers of 24 different histologies. Bars indicate the median and interquartile range. (FL-NHL, follicular lymphoma; DLBCL-NHL, diffuse large B-cell lymphoma; Burkitt-NHL, Burkitt lymphoma; Lung-SCC, squamous cell lung carcinoma; Lung-Adenocarcinoma, lung adenocarcinoma; Kidney-RCC, renal cell carcinoma; Bone-Osteosarcoma, osteosarcoma; Liver-HCC, hepatocellular carcinoma; Breast-Adenocarcinoma, breast adenocarcinoma; Panc-Adenocarcinoma, pancreatic adenocarcinoma; Head-SCC, head and neck squamous cell carcinoma; Ovary-A deno, ovarian adenocarcinoma; Eso-Adeno, esophageal adenocarcinoma; Uterus-Adeno, Stomach-Adeno, gastric adenocarcinoma; CLL, chronic lymphocytic leukemia; ColoRect-Adeno, colorectal adenocarcinoma; Prost-Ade no, prostate adenocarcinoma; CNS-GBM, glioblastoma multiforme; Panc-Endocrine, pancreatic neuroendocrine tumor; Thy-Adeno, thyroid adenocarcinoma; CNS-PiloAstro, pilocytic astrocytoma; CNS-Medullo, medulloblastoma).

[0100] [Figure 6-1] Figure 6 illustrates the contribution of mutational signatures to staged and unstaged SNVs in WGS (Figures 6A-6WW). Scatter plots show the contribution of established single-nucleotide substitution (SBS) mutational signatures to SNVs found within PVs, shown in darker colors, and SNVs found outside possible staged relationships, shown in lighter colors, from WGS. This is presented for 49 SBS mutational signatures across 24 cancer subtypes. Mutational signatures showing significant differences in contribution between staged and unstaged SNVs after multiple hypothesis testing correction are indicated with a *. These figures represent the raw data summarized in Figure 1C. [Figure 6-2] Same as above. [Figure 6-3] Same as above. [Figure 6-4] Same as above. [Figure 6-5] Same as above. [Figure 6-6] Same as above.

[0101] [Figure 7] Figure 7 illustrates the distribution of PVs in typological regions across the genome. Bar plots show the distribution of PVs occurring in typological regions across the genome for multiple cancer types. For this plot, the genome was divided into 1000-bp bins, and the proportion of samples of a given histology that had PVs in each 1000-bp bin was calculated. Only bins with a recurrence frequency of at least 2% in any cancer subtype are shown. The histologies shown are the same as in Figure 1E. The activated B-cell (ABC) and germinal center B-cell (GCB) subtypes of DLBCL are also shown.

[0102] [Figure 8] Figure 8 illustrates the abundance and genomic location of PVs from WGS in lymphoid malignancies. Figure 8A illustrates a bar plot showing the number of independent 1000-bp regions across the genome that recurrently contain PVs for DLBCL, FL, BL, and CLL (n = 68, 74, 36, and 151, respectively). Figures 8B-8D illustrate plots showing the frequency of PVs for multiple lymphoid malignancies with specific loci, including Figure 8B: BCL2, Figure 8C: MYC, and Figure 8D: ID3. The transcript location of a given gene is shown below the gray plot. Exons are shown in darker gray. *Indicates regions with significantly more PVs in a given cancer histology compared to all other histologies by Fisher's exact test (P < 0.05). Similar to Figure 8E and Figures 8B-8D, these plots show the frequency of PVs across lymphoma subtypes. Shown here are the IGH loci, consisting of the IGHV, IGHD, and IGHJ segments, for ABC and GCB subtype DLBCL (n = 25 and 25, respectively). The coding regions of the Ig segments, including the Ig constant region and V genes, are shown. (DLBCL, diffuse large B-cell lymphoma; FL, follicular lymphoma; BL, Burkitt lymphoma; CLL, chronic lymphocytic leukemia).

[0103] [Figure 9-1] Figure 9 illustrates the performance of PhaseED-Seq for the recovery of PVs across lymphoma. Figure 9A illustrates a univariate scatterplot showing the percentage of total PVs across the genome (n = 79) identified by WGS using a previously reported lymphoma CAPP-Seq panel (left) compared to PhaseED-Seq (right). Figure 9B illustrates the expected yield per case of SNVs identified from WGS using a previously established lymphoma CAPP-Seq panel or a PhaseED-Seq panel. Figure 9C illustrates the expected yield per case of PVs identified from WGS using a previously established lymphoma CAPP-Seq panel or a PhaseED-Seq panel. Data from three independent publicly available cohorts are shown in Figures 9A-9C. Figures 9D-9F illustrate plots showing the improvement in PV recovery by PhaseED-Seq compared to CAPP-Seq in 16 patients sequenced by both assays. This includes improvements for d) two SNVs in phase (e.g., 2x or "doublet PVs"), e) three SNVs in phase (3x or "triplet PVs"), and f) four SNVs in phase (e.g., 4x or "quadruplet PVs"). Figures 9G-9K illustrate panels showing the number of SNVs and PVs identified for patients with different types of lymphoma. These panels show the number of g) SNVs, h) doublet PVs, i) triplet PVs, j) quadruplet PVs, and k) all PVs. *, P<0.05; **, P<0.01; ***, P<0.001. (DLBCL, diffuse large B-cell lymphoma; GCB, germinal center B-cell-like DLBCL; ABC, activated B-cell-like DLBCL; PMBCL, primary mediastinal B-cell lymphoma; HL, Hodgkin's lymphoma). [Figure 9-2] Same as above.

[0104] [Figure 10-1]Figure 10 illustrates the location-specific differences in PVs between ABC-DLBCL and GCB-DLBCL (Figures 10A-10Y). Similar to Figure 2D, these scatterplots compare the frequency of PVs by genomic location (in 50-bp bins) for patients with different types of lymphoma. In this figure, the differences between ABC-DLBCL and GCB-DLBCL are shown. Red circles indicate the relative frequency of PVs in 50-bp bins from specific genes of interest. Other (gray) circles indicate the relative frequency of PVs in 50-bp bins from the remainder of the PhasED-Seq sequencing panel. Only genes with statistically significant differences in PVs between ABC-DLBCL and GCB-DLBCL are shown. P values ​​represent a Wilcoxon rank-sum test of the 50-bp bin from a given gene against all other 50-bp bins. See Example 10. [Figure 10-2] Same as above. [Figure 10-3] Same as above.

[0105] [Figure 11-1] Figure 11 illustrates the location-specific differences in PVs between DLBCL and PMBCL (Figures 11A-11X). Similar to Figure 2D, these scatter plots compare the frequency of PVs by genomic location (in 50-bp bins) for patients with different types of lymphoma. In this figure, the differences between DLBCL and PMBCL are shown. Blue circles indicate the relative frequency of PVs in 50-bp bins from specific genes of interest. Other (gray) circles indicate the relative frequency of PVs in 50-bp bins from the remainder of the PhasED-Seq sequencing panel. Only genes with statistically significant differences in PVs between DLBCL and PMBCL are shown. P values ​​represent Wilcoxon rank-sum tests of 50-bp bins from a given gene against all other 50-bp bins. See Example 10. [Figure 11-2] Same as above. [Figure 11-3] Same as above.

[0106] [Figure 12-1]Figure 12 illustrates the location-specific differences in PVs between DLBCL and HL. Similar to Figure 2D, the scatterplots in Figures 12A-12NN compare the frequency of PVs by genomic location (in 50-bp bins) for patients with different types of lymphoma. In this figure, the differences between DLBCL and HL are shown. Green circles indicate the relative frequency of PVs in 50-bp bins from specific genes of interest. Other (gray) circles indicate the relative frequency of PVs in 50-bp bins from the remainder of the PhasED-Seq sequencing panel. Only genes with statistically significant differences in PVs between DLBCL and HL are shown. P values ​​represent Wilcoxon rank-sum tests of the 50-bp bin from a given gene against all other 50-bp bins. See Example 10. [Figure 12-2] Same as above. [Figure 12-3] Same as above. [Figure 12-4] Same as above. [Figure 12-5] Same as above.

[0107] [Figure 13] Figure 13 illustrates the differences in PVs across lymphoma types in IGH locus mutations. This figure shows the frequency of PVs from PhaseED-Seq across the @IGH locus for different types of B-cell lymphoma. The bottom track shows the structure of the @IGH locus and the genetic portion including Ig constant and V genes. The next (outlined) track shows the frequency of PVs in this genomic region from WGS data (ICGC cohort). The remaining tracks show the frequency of PVs from PhaseED-Seq targeted sequencing data, including 1) DLBCL, GCB-DLBCL, ABC-DLBCL, PMBCL, and HL. The regions targeted by the PhasED-Seq panel are indicated at the top. Specific histologies are labeled with selected immunoglobulin moieties with enriched PVs (i.e., IGHV4-34, Sε, Sγ3, and Sγ1).

[0108] [Figure 14-1]Figure 14 illustrates the technical aspects of PhaseED-Seq by hybrid capture sequencing. Figure 14A shows a plot of the theoretical energy of binding for a typical 150-mer across the genome, with increasing percentages of mutated bases from the reference genome. Mutations were spread across the 150-mer, either clustered at one end of the sequence, clustered in the middle of the sequence, or randomly clustered across the entire sequence. The points and error bars represent the median and interquartile range from 10,000 in silico simulations. Figure 14B illustrates a plot showing two histograms of summary metrics of mutation rates in 151-bp windows across the PhasED-Seq panel across all patients in this study. The light gray histogram shows the maximum percent mutated in any 151-bp window for all patients in this study. The dark gray histogram shows the 95th percentile mutation rate across all mutated 151-bp windows. Figure 14C is a plot showing percentiles of mutation rates across all mutated 151-bp windows across all patients in this study. Figure 14D illustrates a heatmap showing the relative error rates (as log10(error rate)) for single SNVs (left, "RED"), doublet PVs (center, "YELLOW"), and triplet PVs (right, "BLUE"). Figure 14D demonstrates that analyses based on multiple stepwise variants (e.g., doublet or triplet PVs) yield lower error rates than analyses based on single SNVs. Furthermore, Figure 14D demonstrates that analyses using a larger number of stepwise variant sets (e.g., triplet PVs labeled "BLUE") yield lower error rates than analyses based on a smaller number of stepwise variant sets (e.g., doublet PVs labeled "YELLOW"). Error rates for single SNVs from sequencing with multiple error suppression methods, including barcode deduplication, iDES, and double-stranded sequencing, are shown. Error rates are aggregated by mutation type. For triplet PVs, the x and y axes of the heatmap represent the first and second types of base changes in the PV.The third change is averaged over all 12 possible base changes. Figure 14E illustrates a plot showing the error rate for doublets / 2x PVs as a function of genomic distance between the component SNVs. [Figure 14-2] Same as above.

[0109] [Figure 15] Figure 15 illustrates a comparison of ctDNA quantification by PhaseED-Seq with CAPP-Seq and clinical applications. Figure 15 illustrates the detection rate of ctDNA from pretreatment samples across 107 patients with large B-cell lymphoma by standard CAPP-Seq (green) and PhaseED-Seq using doublet (light blue), triplet (medium blue), and quadruplet (dark blue) sequences. The specificity of ctDNA detection is also shown. The bottom two plots show the false detection rate in 40 retained healthy control cfDNA samples. The size of each bar in these two plots indicates the detection rate of patient-specific cfDNA mutations in these 40 retained controls across all 107 cases. [Figure 16]Figure 16 illustrates a comparison of ctDNA quantification by PhaseED-Seq with CAPP-Seq and clinical applications. Figure 16A illustrates a table summarizing the sensitivity and specificity of ctDNA detection in pretreatment samples by CAPP-Seq and PhaseED-Seq using doublets, triplets, and quadruplets, as shown in panel A. Sensitivity is calculated across all 107 cases, while specificity is calculated across 40 retained control samples, evaluating a total of 4,280 independent tests for each of the 107 independent patient-specific mutation lists. Figure 16B illustrates a scatterplot showing the amount of ctDNA (measured as log10 (haploid genome equivalents / mL)) measured by CAPP-Seq versus PhaseED-Seq in individual samples. Samples taken before cycle 1 (i.e., pretreatment), cycle 2, and cycle 3 of RCHOP treatment are shown in separate colors (blue, green, and red, respectively; a total of 278 samples). Undetectable levels are on the axis. Spearman correlations and P values ​​are shown.

[0110] [Figure 17]Figure 17 illustrates the detection of ctDNA after two cycles of systemic therapy. Figure 17A illustrates a scatterplot showing the log fold change in ctDNA after two cycles of treatment (i.e., major molecular response or MMR) as measured by CAPP-Seq or PhaseED-Seq for patients receiving RCHOP therapy. The dotted line indicates the previously established threshold of a 2.5-log reduction in ctDNA for MMR. Undetectable samples are on the axis, and the correlation coefficient represents the Spearman rho for 33 samples detected by both CAPP-Seq and PhaseED-Seq. Figure 17B illustrates a 2x2 table summarizing the detection rate of ctDNA samples after two cycles of treatment by PhaseED-Seq versus CAPP-Seq. Patients who eventually progressed are shown in the lower panel, and patients who did not eventually progress are shown in the upper panel. Figure 17C illustrates a bar plot showing the area under the receiver operator curve (AUC) for classification of patients for recurrence-free survival at 24 months based on CAPP-Seq (light colors) or PhasED-Seq (dark colors) after two cycles of treatment. Classifications of both all patients (n=89, left) and only patients who achieved MMR (n=69, right) are shown. Figure 17D illustrates a Kaplan-Meier plot showing recurrence-free survival for 69 patients who achieved MMR stratified by ctDNA detection by CAPP-Seq (top) or PhasED-Seq (bottom).

[0111] [Figure 18-1]Figure 18 illustrates the detection of ctDNA after one cycle of systemic therapy. Figure 18A illustrates a scatterplot showing the log fold change in ctDNA after one cycle of treatment (i.e., early molecular response or EMR) as measured by CAPP-Seq or PhaseED-Seq for patients undergoing RCHOP treatment. The dotted line indicates the previously established threshold of a 2-log reduction in ctDNA relative to EMR. Undetectable samples are on the axis, and the correlation coefficient represents the Spearman rho for 45 samples detected by both CAPP-Seq and PhaseED-Seq. Figure 18B illustrates a 2x2 table summarizing the detection rate of ctDNA samples after one cycle of treatment by PhaseED-Seq versus CAPP-Ceq. Patients with eventual disease progression are shown in red, and patients without eventual disease progression are shown in blue. Figure 18C illustrates a bar plot showing the area under the receiver operating curve (AUC) for patient classification for recurrence-free survival at 24 months based on CAPP-Seq (light color) or PhaseED-Seq (dark color) after one cycle of treatment. Classifications for all patients (n = 82, left) and only patients who achieved EMR (n = 63, right) are shown. Figure 18D illustrates a Kaplan-Meier plot showing recurrence-free survival for 63 patients who achieved EMR stratified by ctDNA detection by CAPP-Seq (top) or ctDNA detection by PhaseED-Seq (bottom). Figure 18E illustrates a waterfall plot showing the change in ctDNA levels measured by CAPP-Seq after one cycle of first-line treatment in DLBCL patients. Patients with undetectable ctDNA by CAPP-Seq are indicated in darker colors as "ND" ("not detected"). The bar color also indicates the final clinical outcome for these patients. Figure 18F illustrates a Kaplan-Meier plot showing relapse-free survival for 33 DLBCL patients with undetectable ctDNA measured by CAPP-Seq after one cycle of treatment.Figure 18G illustrates a Kaplan-Meier plot showing recurrence-free survival for the 33 patients shown in Figure 18F (ctDNA undetectable by CAPP-Seq) stratified by ctDNA detection by PhaseED-Seq at this same time point (Cycle 2, Day 1). Figure 18H illustrates a Kaplan-Meier plot showing recurrence-free survival for 82 DLBCL patients stratified by ctDNA at Cycle 2, Day 1, divided into three strata—patients who fail to achieve an early molecular response, patients who have an early molecular response and still have detectable ctDNA by PhaseED-Seq and / or CAPP-Seq, and patients who have stringent molecular remission (ctDNA undetectable by PhaseED-Seq and CAPP-Seq). [Figure 18-2] Same as above.

[0112] [Figure 19] Figure 19 illustrates the percentage of patients for which PhaseED-Seq would achieve a lower LOD than double-stranded sequencing tracking SNVs based on PCAWG data (whole genome sequencing) that quantified the number of SNVs and phased variants (PVs) in different tumor types.

[0113] [Figure 20] FIG. 20 illustrates the improved LOD achieved in lung cancer (adenocarcinoma, abbreviated "A" and squamous cell carcinoma, abbreviated "S") compared to double-stranded sequencing of whole genome sequencing data.

[0114] [Figure 21] Figure 21 illustrates empirical data from an experiment in which WGS was performed on tumor tissue and a custom panel was designed for 5 patients with solid tumors (5 lung cancers) to examine and compare the LOD of custom CAPP-Seq vs. PhaseED-Seq, showing approximately 10-fold lower LOD using PhaseED-Seq in 5 / 5 patients.

[0115] [Figure 22]Figure 22A illustrates patient photographs from a proof-of-principle example comparing the use of custom CAPP-Seq and PhaseED-Seq for disease surveillance in lung cancer, showing earlier detection of recurrence using PhaseED-Seq. Figure 22B illustrates patient photographs from a proof-of-principle example comparing the use of custom CAPP-Seq and PhaseED-Seq for early detection of disease in breast cancer, showing earlier detection of disease with PhaseED-Seq.

[0116] [Figure 23] 23A-23B illustrate that the methods described herein (eg, the methods shown to generate FIGS. 3E and 3F) do not require barcode-mediated error suppression.

[0117] [Figure 24] FIG. 24 illustrates a flow diagram of a process for performing clinical intervention and / or treatment on an individual based on the detection of circulating tumor nucleic acid sequences in sequencing results, according to one embodiment.

[0118] [Figure 25A] 25A-25C show an exemplary flowchart of a method for determining the status of a subject based on one or more cell-free nucleic acid molecules that contain multiple variants.

[0119] [Figure 25B] 25A-25C show an exemplary flowchart of a method for determining the status of a subject based on one or more cell-free nucleic acid molecules that contain multiple variants. [Figure 25C] 25A-25C show an exemplary flowchart of a method for determining the status of a subject based on one or more cell-free nucleic acid molecules that contain multiple variants.

[0120] [Figure 25D]FIG. 25D shows an exemplary flowchart of a method of treating a condition in a subject based on one or more cell-free nucleic acid molecules comprising multiple variants.

[0121] [Figure 25E] FIG. 25E shows an exemplary flowchart of a method for determining the progression (eg, progression or regression) of a subject's condition based on one or more cell-free nucleic acid molecules comprising multiple variants.

[0122] [Figure 25F] 25F and 25G show an exemplary flowchart of a method for determining the status of a subject based on one or more cell-free nucleic acid molecules comprising multiple variants. [Figure 25G] 25F and 25G show an exemplary flowchart of a method for determining the status of a subject based on one or more cell-free nucleic acid molecules comprising multiple variants.

[0123] [Figure 26] 26A and 26B schematically illustrate different fluorescent probes for identifying one or more cell-free nucleic acid molecules containing multiple graded variants.

[0124] [Figure 27] FIG. 27 illustrates a computer system programmed or otherwise configured to implement the methods provided herein. DETAILED DESCRIPTION OF THE INVENTION

[0125] Detailed Description While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may be made by those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be used.

[0126] The terms "about" or "approximately" generally mean within an acceptable error range for a particular value, which may depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" can mean within one standard deviation or more than one standard deviation, in accordance with practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2-fold of a value. When particular values ​​are described in this application and claims, unless otherwise specified, the term "about" can be assumed to mean within an acceptable error range for the particular value.

[0127] The terms "phased variants," "variants in phase," "PV," or "somatic variants in phase," as used interchangeably herein, generally refer to two or more mutations (e.g., SNVs or indels) that occur in cis (i.e., on the same strand of the nucleic acid molecule) within a single cell-free nucleic acid molecule. In some cases, the cell-free nucleic acid molecule may be a cell-free deoxyribonucleic acid (cfDNA) molecule. In some cases, the cfDNA molecule may be derived from diseased tissue, such as a tumor (e.g., a circulating tumor DNA (ctDNA) molecule).

[0128] The terms "biological sample" or "body sample," used interchangeably herein, generally refer to a tissue or fluid sample derived from a subject. A biological sample can be obtained directly from a subject. Alternatively, a biological sample can be derived from a subject (e.g., by processing an initial biological sample obtained from the subject). A biological sample can be or contain one or more nucleic acid molecules, such as DNA or ribonucleic acid (RNA) molecules. A biological sample can be derived from any organ, tissue, or biological fluid. A biological sample can include, for example, a body fluid or a solid tissue sample. An example of a solid tissue sample is a tumor sample, for example, from a solid tumor biopsy. Non-limiting examples of body fluids include blood, serum, plasma, tumor cells, saliva, urine, cerebrospinal fluid, lymph, prostatic fluid, semen, milk, sputum, feces, tears, and derivatives thereof. In some cases, one or more cell-free nucleic acid molecules, such as those disclosed herein, can be derived from a biological sample.

[0129] The term "subject" as used herein generally refers to any animal, mammal, or human. A subject may have, potentially have, or be suspected of having one or more conditions, such as a disease. In some cases, the subject's condition may be cancer, a symptom(s) associated with cancer, or be asymptomatic or undiagnosed with cancer (e.g., not diagnosed with cancer). In some cases, the subject may have cancer, the subject may exhibit a symptom(s) associated with cancer, the subject may not have a symptom associated with cancer, or the subject may not have been diagnosed with cancer. In some examples, the subject is a human.

[0130] The term "cell-free DNA" or "cfDNA" used interchangeably herein generally refers to DNA fragments that circulate freely in the bloodstream of a subject.Cell-free DNA fragments can have dinucleosome protection (for example, fragment size of at least 240 base pairs ("bp")).These cfDNA fragments with dinucleosome protection may not be cut between nucleosomes, resulting in longer fragment lengths (for example, have a typical size distribution centered around 334 bp).Cell-free DNA fragments can have mononucleosome protection (for example, fragment size of less than 240 base pairs ("bp")).These cfDNA fragments with mononucleosome protection may be cut between nucleosomes, resulting in shorter fragment lengths (for example, have a typical size distribution centered around 167 bp).

[0131] As used herein, the term "sequencing data" generally refers to the "raw sequence reads" and / or "consensus sequences" of nucleic acids, such as cell-free nucleic acids or their derivatives. Raw sequence reads are the output of a DNA sequencer and typically contain redundant sequences of the same parent molecule, for example, after amplification. A "consensus sequence" is a sequence derived from redundant sequences of parent molecules intended to represent the sequence of the original parent molecule. Consensus sequences can be generated by voting (where each majority nucleotide in the sequence, e.g., the most commonly observed nucleotide at a given base position, is the consensus nucleotide) or other approaches, such as comparison with a reference genome. In some cases, consensus sequences can be generated by tagging the original parent molecule with unique or non-unique molecular tags, allowing tracking of the tags and / or tracking of progeny sequences (e.g., after amplification) using internal information from the sequence reads.

[0132] As used herein, the term "reference genome sequence" generally refers to a nucleotide sequence to which a nucleotide sequence of interest is compared.

[0133] As used herein, the term "genomic region" generally refers to any region of a genome (e.g., a range of base pair positions), such as an entire genome, a chromosome, a gene, or an exon. A genomic region can be a contiguous or discontinuous region. A "genetic locus" (or "locus") can be part or an entire genomic region (e.g., a gene, a portion of a gene, or a single base of a gene).

[0134] As used herein, the term "likelihood" generally refers to probability, relative probability, presence or absence, or degree.

[0135] As used herein, the term "liquid biopsy" generally refers to a non-invasive or minimally invasive laboratory test or assay (e.g., of a biological sample or cell-free nucleic acid). A "liquid biopsy" assay can report the detection or measurement (e.g., minor allele frequency, gene expression, or protein expression) of one or more marker genes associated with a condition of interest (e.g., cancer or tumor-associated marker genes). A. Introduction

[0136] Alterations (e.g., mutations) in genomic DNA can be manifested in the formation and / or progression of one or more conditions (e.g., diseases such as cancer or tumors) in a subject. The present disclosure provides methods and systems for analyzing cell-free nucleic acid molecules, such as cfDNA, from a subject to determine the presence or absence of a condition in the subject, a prognosis for a diagnosed condition in the subject, the progression of a condition in the subject over time, therapeutic treatment for a diagnosed condition in the subject, or a predicted treatment outcome for a condition in the subject.

[0137] Analysis of cell-free nucleic acids, such as cfDNA, has been developed for a wide range of applications, for example, in prenatal testing, organ transplantation, infectious diseases, and oncology. In terms of detecting or monitoring target diseases, such as cancer, circulating tumor DNA (ctDNA) can be a highly sensitive and specific biomarker in many cancer types. In some cases, ctDNA can be used to detect the presence of minimal residual disease (MRD) or tumor burden after treatment, such as chemotherapy or surgical resection of solid tumors. However, the limit of detection (LOD) of ctDNA analysis can be limited by several factors, including (i) the low amount of input DNA from typical blood draws and (ii) the background error rate from sequencing.

[0138] In some cases, ctDNA-based cancer detection can be improved by tracking multiple somatic mutations using commercially available panels or personalized assays, for example, with error-suppression sequencing at an LOD of approximately 2 in 100,000 from cfDNA input. However, in some cases, the current LOD of the ctDNA of interest may be insufficient to broadly detect MRD in patients whose disease is expected to relapse or progress. For example, such "loss of detection" can be exemplified in diffuse large B-cell lymphoma (DLBCL). In DLBCL, interim ctDNA detection after only two cycles of curative-intent therapy can represent a major molecular response (MMR) and be a strong prognostic marker for ultimate clinical outcome. Despite this, nearly one-third of patients who ultimately experience disease progression do not have detectable ctDNA at this interim landmark using available technologies (e.g., Cancer Personalized Profiling by Deep Sequencing (CAPP-Seq)), thus representing a "false-negative" measurement. Such high false-negative rates have also been observed in DLBCL patients with alternative methods, such as monitoring ctDNA for immunoglobulin gene rearrangements. Therefore, improved methods for ctDNA-based cancer detection with higher sensitivity are needed.

[0139] Somatic variants detected on both complementary strands of the parent DNA duplex can be used to lower the LOD of ctDNA detection, thereby advantageously increasing the sensitivity of ctDNA detection. Such "double-stranded sequencing" can reduce the background error profile due to the requirement for two matching events for single-nucleotide variant (SNV) detection. However, because recovery of both original strands may occur in a small number of all recovered molecules, double-stranded sequencing approaches alone can be limited by inefficient recovery of DNA duplexes. Therefore, double-stranded sequencing may be suboptimal and inefficient for real-world ctDNA detection with limited amounts of starting sample, where input DNA from actual blood volumes (e.g., approximately 4,000 to approximately 8,000 genomes per standard 10 milliliter (mL) blood collection tube) is limited and maximum genome recovery is essential.

[0140] Thus, there remains a significant unmet need for detection and analysis of ctDNA with a low LOD (e.g., thereby providing high sensitivity) to, for example, determine the presence or absence of disease in a subject, the prognosis of the disease, the treatment of the disease, and / or the predicted outcome of the treatment. B. Methods and Systems for Determining or Monitoring Conditions

[0141] The present disclosure describes a method and system for detecting and analyzing cell-free nucleic acid with multiple stepwise variants as a characteristic of a subject's condition.In some embodiments, cell-free nucleic acid molecules can include cfDNA molecules, such as ctDNA molecules.The method and system disclosed herein can utilize the sequencing data from a plurality of cell-free nucleic acid molecules of a subject to identify a subset of a plurality of cell-free nucleic acid molecules with multiple stepwise variants, thereby determining the subject's condition.The method and system disclosed herein can directly detect, and in some cases pull down (or capture) such subset of a plurality of cell-free nucleic acid molecules that show multiple stepwise variants, thereby determining the subject's condition with or without sequencing.The method and system disclosed herein can reduce the background error rate that is often involved in the detection and analysis of cell-free nucleic acid molecules, such as cfDNA.

[0142] In some aspects, methods and systems for cell-free nucleic acid sequencing and cancer detection are provided. In some embodiments, cell-free nucleic acid (e.g., cfDNA or cfRNA) can be extracted from an individual's liquid biopsy and prepared for sequencing. The results of cell-free nucleic acid sequencing can be analyzed to detect in-phase somatic variants (i.e., phase variants as disclosed herein) as indicators of circulating tumor nucleic acid (ctDNA or ctRNA) sequences (i.e., sequences derived from or originated from nucleic acids in cancer cells). Thus, in some cases, cancer can be detected in an individual by extracting a liquid biopsy from the individual and sequencing the cell-free nucleic acid from the liquid biopsy to detect circulating tumor nucleic acid sequences, and the presence of circulating tumor nucleic acid sequences can indicate that the individual has cancer (e.g., a particular type of cancer). In some cases, clinical intervention and / or treatment can be determined and / or implemented for the individual based on the detection of cancer.

[0143] As disclosed herein, the presence of in-phase somatic variants can be a strong indicator that a nucleic acid containing such a stepwise variant originates from a bodily sample having a condition, such as a cancerous cell (or that the nucleic acid is derived from a bodily sample obtained or derived from a subject having a condition, such as cancer). Because stepwise mutations are unlikely to occur within small genetic windows, the approximate size of a typical cell-free nucleic acid molecule (e.g., about 170 bp or less), detection of stepwise somatic variants can increase the signal-to-noise ratio of cell-free nucleic acid detection methods (e.g., by reducing or eliminating spurious "noise" signals).

[0144] In some embodiments, some genomic regions can be used as hotspots for detecting stage variants, particularly in various cancers, such as lymphoma.In some cases, enzymes (such as AID, Apobec3a) can mutate the DNA of specific genes and positions in a typical manner, leading to the occurrence of specific cancers.Therefore, the cell-free nucleic acid derived from such hotspot genomic regions can be captured or targeted (for example, with or without deep sequencing) for cancer detection and / or monitoring.Alternatively, capture or targeted sequencing can be performed on the region where stage variants have been previously detected from a specific individual's cancer source (for example, tumor) to detect the individual's cancer.

[0145] In some embodiments, capture sequencing of cell-free nucleic acid can be carried out as screening diagnostics.In some cases, screening diagnostics can be developed and used to detect circulating tumor nucleic acid for cancer with typical region of stage variants.In some cases, capture sequencing of cell-free nucleic acid can be carried out as diagnostics to detect MRD or tumor burden and determine whether certain disease exists during or after treatment.In some cases, capture sequencing of cell-free nucleic acid can be carried out as diagnostics to determine the progress (for example, progress or relapse) of treatment.

[0146] In some embodiments, cell-free nucleic acid sequencing results can be analyzed to detect whether stepwise somatic single-nucleotide variants (SNVs) or other mutations or variants (e.g., indels) are present in a cell-free nucleic acid sample. In some cases, the presence of specific somatic SNVs or other variants can indicate a circulating tumor nucleic acid sequence and thus a tumor present in a subject. In some cases, at least two variants can be detected in phase on a cell-free nucleic acid molecule. In some cases, at least three variants can be detected in phase on a cell-free nucleic acid molecule. In some cases, at least four variants can be detected in phase on a cell-free nucleic acid molecule. In some cases, at least five or more variants can be detected in phase on a cell-free nucleic acid molecule. In some cases, the greater the number of stepwise variants detected on a cell-free nucleic acid molecule, the greater the likelihood that the cell-free nucleic acid molecule is derived from cancer, as opposed to detecting harmless sequences of somatic variants resulting from molecular preparation of a sequence library or random biological errors. Thus, the likelihood of a false positive detection can decrease with the detection of more in-phase variants within a molecule (e.g., thereby increasing the specificity of detection).

[0147] In some embodiments, cell-free nucleic acid sequencing results can be analyzed to detect whether one or more nucleic acid bases (i.e., indels) are inserted or deleted in a cell-free nucleic acid sample, for example, relative to a reference genome sequence. Without wishing to be bound by theory, in some cases, the presence of indels in a cell-free nucleic acid molecule (e.g., cfDNA) can indicate a subject's condition, for example, a disease such as cancer. In some cases, the genetic variation resulting from indels can be treated as a variant or mutation, and thus, two indels can be treated as two phase variants as disclosed herein. In some examples, within a cell-free nucleic acid molecule, the first genetic variation (first phase variant) from the first indel and the second genetic variation (second phase variant) from the second indel can be separated from each other by at least one nucleotide.

[0148] As disclosed herein, within a single cell-free nucleic acid molecule (e.g., a single cfDNA molecule), the first step variant can be an SNV, and the second step variant can be part of a different small nucleotide polymorphism, such as another SNV or a multinucleotide variant (MNV). A multinucleotide variant can be a cluster of two or more (e.g., at least 2, 3, 4, 5, or more) adjacent variants present within the same strand of a nucleic acid molecule. In some cases, the first step variant and the second step variant can be part of the same MNV within a single cell-free nucleic acid molecule. In some cases, the first step variant and the second step variant can be derived from two different MNVs within a single cell-free nucleic acid molecule.

[0149] In some embodiments, statistical methods can be used to calculate the likelihood that a detected stepwise variant is cancer-derived and not random or artificial (e.g., from sample preparation or sequencing errors). In some cases, Monte Carlo sampling methods can be used to determine the likelihood that a detected stepwise variant is cancer-derived and not random or artificial.

[0150] The present disclosure provides for the identification or detection of cell-free nucleic acid (for example, cfDNA molecule) with multiple step variants, for example, from a liquid biopsy of a subject.In some cases, the first step variant of the multiple step variants and the second step variant of the multiple step variants can be directly adjacent to each other (for example, adjacent SNV).In some cases, the first step variant of the multiple step variants and the second step variant of the multiple step variants can be separated by at least one nucleotide.The interval between the first step variant and the second step variant can be limited by the length of the cell-free nucleic acid molecule.

[0151] As disclosed herein, within a single cell-free nucleic acid molecule (e.g., a single cfDNA molecule), the first graded variant and the second graded variant can be at least or up to about 1 nucleotide, at least or up to about 2 nucleotides, at least or up to about 3 nucleotides, at least or up to about 4 nucleotides, at least or up to about 5 nucleotides, at least or up to about 6 nucleotides, at least or up to about 7 nucleotides, at least or up to about 8 nucleotides, at least or up to about 9 nucleotides, at least or up to about 10 nucleotides, at least or up to about 11 nucleotides, at least or up to about 12 nucleotides, at least or up to about 13 nucleotides, at least or up to about 14 nucleotides, at least or up to about 15 nucleotides, at least or up to about 20 nucleotides, at least or up to about 25 nucleotides, nucleotides, at least or up to about 30 nucleotides, at least or up to about 35 nucleotides, at least or up to about 40 nucleotides, at least or up to about 45 nucleotides, at least or up to about 50 nucleotides, at least or up to about 60 nucleotides, at least or up to about 70 nucleotides, at least or up to about 80 nucleotides, at least or up to about 90 nucleotides, at least or up to about 100 nucleotides, at least or up to about 110 nucleotides, at least or up to about 120 nucleotides, at least or up to about 130 nucleotides, at least or up to about 140 nucleotides, at least or up to about 150 nucleotides, at least or up to about 160 nucleotides, at least or up to about 170 nucleotides, or at least or up to about 180 nucleotides. Alternatively, or in addition, within a single cell-free nucleic acid molecule, the first and second stepwise variants may not or need not be separated by one or more nucleotides and may therefore be directly adjacent to each other.

[0152] A single cell-free nucleic acid molecule (e.g., a single cfDNA molecule) disclosed herein can comprise at least or up to about 2 graded variants, at least or up to about 3 graded variants, at least or up to about 4 graded variants, at least or up to about 5 graded variants, at least or up to about 6 graded variants, at least or up to about 7 graded variants, at least or up to about 8 graded variants, at least or up to about 9 graded variants, at least or up to about 10 graded variants, at least or up to about 12 graded variants, at least or up to about 12 graded variants, at least or up to about 13 graded variants, at least or up to about 14 graded variants, at least or up to about 15 graded variants, at least or up to about 20 graded variants, or at least or up to about 25 graded variants within the same molecule.

[0153] For each cell-free nucleic acid molecule identified as containing a plurality of stepwise variants from the plurality of cell-free nucleic acid molecules obtained (e.g., from a liquid biopsy of a subject), on average, at least or up to about 2 stepwise variants, at least or up to about 3 stepwise variants, at least or up to about 4 stepwise variants, at least or up to about 5 stepwise variants, at least or up to about 6 stepwise variants, at least or up to about 7 stepwise variants, at least or up to about 8 stepwise variants, at least or up to about 9 stepwise variants, at least or up to about 10 stepwise variants, at least or up to about 11 stepwise variants, at least or up to about 12 stepwise variants, at least or up to about 13 stepwise variants, at least or up to about 14 stepwise variants, at least or up to about 15 stepwise variants, at least or up to about 16 stepwise variants, at least or up to about 17 stepwise variants, at least or up to about 18 stepwise variants, at least or up to about 19 stepwise variants, at least or up to about 20 stepwise variants, at least or up to about 21 stepwise variants, at least or up to about 22 stepwise variants, at least or up to about 23 stepwise variants, at least or up to about 24 stepwise variants, at least or up to about 25 stepwise variants, at least or up to about 26 stepwise variants, at least or up to about 27 stepwise variants, at least or up to about 28 stepwise variants, at least or up to about 29 stepwise variants, at least or up to about 30 stepwise variants, at least or up to about 31 stepwise variants, at least or up to about 32 stepwise variants, at least or up to about 33 stepwise variants, at least or up to about 34 stepwise variants, at least or up to about 35 Two or more (e.g., 10 or more, 1,000 or more, 10,000 or more) cell-free nucleic acid molecules having at least or up to about 10 stepwise variants, at least or up to about 12 stepwise variants, at least or up to about 12 stepwise variants, at least or up to about 13 stepwise variants, at least or up to about 14 stepwise variants, at least or up to about 15 stepwise variants, at least or up to about 20, or at least or up to about 25 stepwise variants can be identified.

[0154] In some cases, a plurality of cell-free nucleic acid molecules (e.g., cfDNA molecules) can be obtained from a biological sample of a subject (e.g., a solid tumor or liquid biopsy). Of the plurality of cell-free nucleic acid molecules, at least or at most 1, at least or at most 2, at least or at most 3, at least or at most 4, at least or at most 5, at least or at most 6, at least or at most 7, at least or at most 8, at least or at most 9, at least or at most 10, at least or at most 15, at least or at most 20, at least or at most 25, at least or at most 30, at least or at most 35, at least or at most 40, at least or at most 45, at least or at most 50, at least or at most 60, at least or at most 70, at least or at most 80, At least or up to 90, at least or up to 100, at least or up to 150, at least or up to 200, at least or up to 300, at least or up to 400, at least or up to 500, at least or up to 600, at least or up to 700, at least or up to 800, at least or up to 900, at least or up to 1,000, at least or up to 5,000, at least or up to 10,000, at least or up to 50,000, or at least or up to 100,000 cell-free nucleic acid molecules can be identified, each identified cell-free nucleic acid molecule comprising a plurality of graded variants as disclosed herein.

[0155] In some cases, a plurality of cell-free nucleic acid molecules (e.g., cfDNA molecules) can be obtained from a biological sample of a subject (e.g., a solid tumor or liquid biopsy). Of the plurality of cell-free nucleic acid molecules, at least or at most 1, at least or at most 2, at least or at most 3, at least or at most 4, at least or at most 5, at least or at most 6, at least or at most 7, at least or at most 8, at least or at most 9, at least or at most 10, at least or at most 15, at least or at most 20, at least or at most 25, at least or at most 30, at least or at most 35, at least or at most 40, at least or at most 45, at least or at most 50, at least or at most 60, at least or at most 70, at least or up to 80, at least or up to 90, at least or up to 100, at least or up to 150, at least or up to 200, at least or up to 300, at least or up to 400, at least or up to 500, at least or up to 600, at least or up to 700, at least or up to 800, at least or up to 900, or at least or up to 1,000 cell-free nucleic acid molecules can be identified from a target genomic region (e.g., a target genomic locus), each identified cell-free nucleic acid molecule comprising a plurality of graded variants as disclosed herein.

[0156] Figure 1A and 1E show the example of (i) a cfDNA molecule that contains SNV and (ii) another cfDNA molecule that contains multiple stepwise variants.Each variant identified in cfDNA can indicate the existence of another genetic mutation in the cell from which cfDNA is derived.In alternative embodiments, one or more stepwise variants can be insertion or deletion (indel) instead of SNV.

[0157] In one embodiment, the present disclosure provides a method for determining a subject's status, as shown by flowchart 2510 in FIG. 25A. The method may include (a) obtaining, by a computer system, sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject (process 2512). The method may further include (b) processing, by a computer system, the sequencing data to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, each of the identified one or more cell-free nucleic acid molecules comprising a plurality of staged variants relative to a reference genome sequence (process 2514). In some cases, as disclosed herein, at least a portion of the one or more cell-free nucleic acid molecules may comprise a first staged variant of the plurality of staged variants and a second staged variant of the plurality of staged variants separated by at least one nucleotide. The method may optionally include (c) analyzing, by a computer system, at least a portion of the identified one or more cell-free nucleic acid molecules to determine the subject's status (process 2516).

[0158] In some cases, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, at least or up to about 50%, at least or up to about 60%, at least or up to about 70%, at least or up to about 80%, at least or up to about 90%, at least or up to about 95%, at least or up to about 99%, or about 100% of one or more cell-free nucleic acid molecules can comprise a first graded variant of the plurality of graded variants and a second graded variant of the plurality of graded variants separated by at least one nucleotide, as disclosed herein. In some examples, the multiple step variants in a single cfDNA molecule can include (i) a first multiple step variants separated from each other by at least one nucleotide, and (ii) a second multiple step variants adjacent to each other (e.g., two step variants in MNV). In some examples, the multiple step variants in a single cfDNA molecule can consist of step variants separated from each other by at least one nucleotide.

[0159] In one embodiment, the present disclosure provides a method for determining a subject's status, as shown by flowchart 2520 in Figure 25B. The method may include (a) obtaining, by a computer system, sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject (process 2522). The method may further include (b) processing, by a computer system, the sequencing data to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of staged variants relative to a reference genome sequence (process 2524). In some cases, as disclosed herein, a first staged variant of the plurality of staged variants and a second staged variant of the plurality of staged variants may be separated by at least one nucleotide. The method may optionally include (c) analyzing, by a computer system, at least a portion of the identified one or more cell-free nucleic acid molecules to determine the subject's status (process 2526).

[0160] In one aspect, the present disclosure provides a method for determining a subject's status, as shown by flowchart 2530 in Figure 25C. The method may include (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject (process 2532). The method may further include (b) processing the sequencing data to identify one or more cell-free nucleic acid molecules from the plurality of cell-free nucleic acid molecules having an LOD that is less than about 1 in 50,000 observations (or cell-free nucleic acid molecules) from the sequencing data (process 2534). In some cases, each of the one or more cell-free nucleic acid molecules comprises multiple graded variants relative to a reference genome sequence. The method may optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules to determine the subject's status (process 2536).

[0161] In some cases, as disclosed herein, the LOD of an operation identifying one or more cell-free nucleic acid molecules can be less than about 1 in 60,000, less than 1 in 70,000, less than 10 in 80,000, less than 1 in 90,000, less than 1 in 100,000, less than 1 in 150,000, less than 1 in 200,000, less than 1 in 300,000, less than 1 in 400,000, less than 1 in 500,000, less than 1 in 600,000, less than 1 in 700,000, less than 1 in 800,000, less than 1 in 900,000, less than 1 in 1 million, less than 1 in 1 million, less than 1 in 1.1 million, less than 1.2 million, less than 1.3 million, less than 1.4 million, less than 1.5 million, or less than 1 in 2 million observations from the sequencing data.

[0162] In some cases, as disclosed herein, at least one cell-free nucleic acid molecule of the identified one or more cell-free nucleic acid molecules may comprise a first stepwise variant of the plurality of stepwise variants and a second stepwise variant of the plurality of stepwise variants separated by at least one nucleotide.

[0163] In some cases, one or more of the subject method operations (a)-(c) can be performed by a computer system. In one example, all of the subject method operations (a)-(c) can be performed by a computer system.

[0164] The sequencing data disclosed herein can be obtained from one or more sequencing methods. The sequencing method can be a first-generation sequencing method (e.g., Maxam-Gilbert sequencing, Sanger sequencing). The sequencing method can be a high-throughput sequencing method, such as next-generation sequencing (NGS) (e.g., sequencing by synthesis). A high-throughput sequencing method can simultaneously (or substantially simultaneously) sequence at least about 10,000, at least about 100,000, at least about 1 million, at least about 10 million, at least about 100 million, at least about 1 billion, or more polynucleotide molecules (e.g., cell-free nucleic acid molecules or derivatives thereof). NGS can be any generation of sequencing technology (e.g., second-generation sequencing technology, third-generation sequencing technology, fourth-generation sequencing technology, etc.). Non-limiting examples of high-throughput sequencing methods include massively parallel signature sequencing, polony sequencing, pyrosequencing, sequencing-by-synthesis, combinatorial probe anchor synthesis (cPAS), sequencing-by-ligation (e.g., sequencing by oligonucleotide ligation and detection (SOLiD) sequencing), semiconductor sequencing (e.g., Ion Torrent semiconductor sequencing), DNA nanoball sequencing, and single molecule sequencing, sequencing-by-hybridization.

[0165] In some embodiments of any one of the methods disclosed herein, sequencing data can be obtained based on any of the disclosed sequencing methods that utilize nucleic acid amplification (e.g., polymerase chain reaction (PCR)). Non-limiting examples of such sequencing methods can include 454 pyrosequencing, polony sequencing, and SoLiD sequencing. In some cases, sequencing data can be generated by generating amplicons (e.g., derivatives of multiple cell-free nucleic acid molecules obtained from or derived from a subject, as disclosed herein) corresponding to a genomic region of interest (e.g., a genomic region associated with a disease) by PCR, optionally pooling, and then sequencing. In some examples, because the region of interest is amplified into amplicons by PCR before sequencing, the nucleic acid sample is already enriched for the region of interest, and therefore, additional pooling (e.g., hybridization) prior to sequencing may not, and need not, be required (e.g., non-hybridization-based NGS). Alternatively, hybridization pooling can be further performed for further enrichment prior to sequencing. Alternatively, sequencing data can be obtained without generating PCR copies, for example, by cPAS sequencing.

[0166] Some embodiments utilize capture hybridization technology to perform targeted sequencing. When sequencing cell-free nucleic acids, library products can be captured by hybridization before sequencing to enhance the resolution of specific genomic loci. Capture hybridization can be particularly useful when attempting to detect rare and / or somatic stage variants at specific genomic loci from a sample. In some situations, detecting rare and / or somatic stage variants indicates the source of nucleic acid, including nucleic acid derived from a cancer source. Therefore, capture hybridization is a tool that can enhance the detection of circulating tumor nucleic acid in cell-free nucleic acid.

[0167] Various types of cancer repeatedly undergo aberrant somatic hypermutation, particularly at genomic loci. For example, enzyme activation-induced deaminase induces aberrant somatic hypermutation in B cells, resulting in various B cell lymphomas, including but not limited to diffuse large B cell lymphoma (DLBCL), follicular lymphoma (FL), Burkitt's lymphoma (BL), and B cell chronic lymphocytic leukemia (CLL). Therefore, in many embodiments, probes are designed to pull down (or capture) genomic loci known to undergo aberrant somatic hypermutation in lymphomas. Figure 1D and Table 1 list several regions that undergo aberrant somatic hypermutation in DLBCL, FL, BL, and CLL. Table 6 provides a list of nucleic acid probes that can be used to pull down (or capture) genomic loci to detect aberrant somatic hypermutation in B cell cancers.

[0168] Capture sequencing can also be performed using personalized nucleic acid probes designed to detect the presence of cancer in individuals. Individuals with cancer can have their cancer biopsied and sequenced to detect somatic stage variants accumulated in the cancer. Based on the sequencing results, in some embodiments, nucleic acid probes are designed and synthesized that can pull down genomic loci containing the positions where stage variants exist. These personalized designed and synthesized nucleic acid probes can be used to detect circulating tumor nucleic acids from the individual's liquid biopsy. Thus, personalized nucleic acid probes can be useful for determining treatment response and / or detecting MRD after treatment.

[0169] In some embodiments of any one of the methods disclosed herein, sequencing data can be obtained based on any sequencing method that utilizes an adapter. A nucleic acid sample (e.g., a plurality of cell-free nucleic acid molecules from a subject as disclosed herein) can be conjugated with one or more adapters (or adapter sequences) for recognizing (e.g., by hybridization) the sample or any derivatives thereof (e.g., amplicons). In some examples, for example, the nucleic acid sample can be tagged with a molecular barcode, so that each cell-free nucleic acid molecule of the plurality of cell-free nucleic acid molecules can have a unique barcode. Alternatively, or in addition, for example, the nucleic acid sample can be tagged with a sample barcode, so that a plurality of cell-free nucleic acid molecules from a subject (e.g., a plurality of cell-free nucleic acid molecules obtained from a specific body tissue of a subject) can have the same barcode.

[0170] In alternative embodiments, the methods for identifying one or more cell-free nucleic acid molecules comprising multiple graded variants disclosed herein can be performed without molecular barcoding, without sample barcoding, or without molecular barcoding and sample barcoding, due at least in part to the high specificity and low LOD achieved by relying on the identification of graded variants as opposed to, e.g., single SNVs.

[0171] In some embodiments of any one of the methods disclosed herein, sequencing data can be obtained and analyzed without in silico removal or suppression of (i) background errors and / or (ii) sequencing errors, due at least in part to the high specificity and low LOD achieved by relying on identifying stepwise variants as opposed to, e.g., single SNVs or indels.

[0172] In some embodiments of any one of the methods disclosed herein, using multiple variants as a condition for identifying target cell-free nucleic acid molecules with specific mutations of interest without in silico methods of error suppression can result in a background error rate that is at least about 5-fold, at least about 10-fold, at least about 20-fold, at least about 30-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, at least about 100-fold, at least about 200-fold, at least about 400-fold, at least about 600-fold, at least about 800-fold, or at least about 1,000-fold lower than the background error rate of (i) barcode deduplication, (ii) integrated digital error suppression, or (iii) double-stranded sequencing. This approach can advantageously increase the signal-to-noise ratio (thereby increasing sensitivity and / or specificity) for identifying target cell-free nucleic acid molecules with specific mutations of interest.

[0173] In some embodiments of any one of the methods disclosed herein, increasing the minimum number of stepwise variants per cell-free nucleic acid molecule required to identify a target cell-free nucleic acid molecule with a specific mutation of interest (e.g., from at least two stepwise variants to at least three stepwise variants) can reduce the background error rate by at least about 5-fold, at least about 10-fold, at least about 20-fold, at least about 30-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, or at least about 100-fold. This approach can advantageously increase the signal-to-noise ratio for identifying a target cell-free nucleic acid molecule with a specific mutation of interest (thereby increasing sensitivity and / or specificity).

[0174] In one embodiment, the present disclosure provides a method of treating a condition in a subject, as shown in flowchart 2540 of Figure 25D. The method can include (a) identifying a subject for treatment of a condition, the subject having been determined to have the condition based on the identification of one or more cell-free nucleic acid molecules from a plurality of cell-free nucleic acid molecules obtained or derived from the subject (process 2542). Each of the identified one or more cell-free nucleic acid molecules can comprise multiple stage variants relative to a reference genome sequence. As disclosed herein, at least some (e.g., partial or all) of the multiple stage variants can be separated by at least one nucleotide, such that a first stage variant of the multiple stage variants and a second stage variant of the multiple stage variants are separated by at least one nucleotide. In some cases, the presence of multiple stage variants indicates a condition (e.g., a disease such as cancer) in the subject. The method can further include (b) subjecting the subject to treatment based on step (a) (process 2544). Examples of such treatments for subject conditions are disclosed elsewhere in this disclosure.

[0175] In one aspect, the present disclosure provides a method for monitoring the progression (e.g., progression or regression) of a condition in a subject, as shown in flowchart 2550 of FIG. 25E. The method may include (a) determining a first status of the subject's condition based on identification of a first set of one or more cell-free nucleic acid molecules from a first plurality of cell-free nucleic acid molecules obtained or derived from the subject (process 2552). The method may further include (b) determining a second status of the subject's condition based on identification of a second set of one or more cell-free nucleic acid molecules from a second plurality of cell-free nucleic acid molecules obtained or derived from the subject (process 2554). The second plurality of cell-free nucleic acid molecules can be obtained from the subject after obtaining the first plurality of cell-free nucleic acid molecules from the subject. The method optionally may include (c) determining the progression (e.g., progression or regression) of the condition based at least in part on the first status of the condition and the second status of the condition (process 2556). In some cases, each of the identified one or more cell-free nucleic acid molecules (e.g., each of the first set of identified one or more cell-free nucleic acid molecules, each of the second set of identified one or more cell-free nucleic acid molecules) may contain multiple graded variants relative to the reference genome sequence. As disclosed herein, at least a portion (e.g., partial or all) of the identified one or more cell-free nucleic acid molecules may be separated by at least one nucleotide. In some cases, the presence of multiple graded variants may indicate the status of the subject's condition.

[0176] In some cases, a first plurality of cell-free nucleic acid molecules can be obtained from a subject (e.g., by blood biopsy) and analyzed to determine (e.g., diagnose) a first status of the subject's condition (e.g., a disease such as cancer). The first plurality of cell-free nucleic acid molecules can be analyzed using any of the methods disclosed herein (e.g., with or without sequencing) to identify a first set of one or more cell-free nucleic acid molecules comprising a plurality of staged variants, and the presence or characteristics of the first set of one or more cell-free nucleic acid molecules can be used to determine a first status (e.g., initial diagnosis) of the subject's condition. Based on the determined first status of the condition, the subject can be subjected to one or more treatments (e.g., chemotherapy) disclosed herein. Following the one or more treatments, a second plurality of cell-free nucleic acid molecules can be obtained from the subject.

[0177] In some cases, the subject may be subjected to at least or up to about 1 treatment, at least or up to about 2 treatments, at least or up to about 3 treatments, at least or up to about 4 treatments, at least or up to about 5 treatments, at least or up to about 6 treatments, at least or up to about 7 treatments, at least or up to about 8 treatments, at least or up to about 9 treatments, or at least or up to about 10 treatments based on the first status of the determined condition. In some cases, a subject may receive multiple treatments based on the first status of the determined condition, and the first treatment of the multiple treatments and the second treatment of the multiple treatments may be separated by at least or up to about 1 day, at least or up to about 7 days, at least or up to about 2 weeks, at least or up to about 3 weeks, at least or up to about 4 weeks, at least or up to about 2 months, at least or up to about 3 months, at least or up to about 4 months, at least or up to about 5 months, at least or up to about 6 months, at least or up to about 12 months, at least or up to about 2 years, at least or up to about 3 years, at least or up to about 4 years, at least or up to about 5 years, or at least or up to about 10 years. The multiple treatments for a subject may be the same. Alternatively, the multiple treatments may differ by drug type (e.g., different chemotherapy drugs), drug dosage (e.g., increased dosage, decreased dosage), the presence or absence of co-therapeutic agents (e.g., chemotherapy and immunotherapy), mode of administration (e.g., intravenous administration vs. oral administration), frequency of administration (e.g., daily, weekly, monthly), etc.

[0178] In some cases, the subject may not, and need not, be treated for the condition between the determination of the first status of the condition and the determination of the second status of the condition. For example, without intervening treatment, a second plurality of cell-free nucleic acid molecules from the subject (e.g., by liquid biopsy) can be included to determine whether the subject still exhibits symptoms of the first status of the condition.

[0179] In some cases, the second plurality of cell-free nucleic acid molecules from the subject can be obtained (e.g., by blood biopsy) at least or up to about 1 day, at least or up to about 7 days, at least or up to about 2 weeks, at least or up to about 3 weeks, at least or up to about 4 weeks, at least or up to about 2 months, at least or up to about 3 months, at least or up to about 4 months, at least or up to about 5 months, at least or up to about 6 months, at least or up to about 12 months, at least or up to about 2 years, at least or up to about 3 years, at least or up to about 4 years, at least or up to about 5 years, or at least or up to about 10 years after obtaining the first plurality of cell-free nucleic acid molecules from the subject.

[0180] In some cases, as disclosed herein, different samples comprising at least or up to about 2, at least or up to about 3, at least or up to about 4, at least or up to about 5, at least or up to about 6, at least or up to about 7, at least or up to about 8, at least or up to about 9, or at least or up to about 10 plurality of nucleic acid molecules (e.g., at least a first plurality of cell-free nucleic acid molecules and a second plurality of cell-free nucleic acid molecules) can be obtained over time (e.g., once every month for six months, once every two months for one year, once every three months for one year, once every six months for one year or more) to monitor the progression of a subject's condition.

[0181] In some cases, determining the progression of a condition based on the first condition and the second condition can include comparing one or more characteristics of the first condition and the second condition, such as (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants in each condition (e.g., per equal weight or volume of the original biological sample, per equal number of analyzed initial cell-free nucleic acid molecules, etc.), (ii) the average number of multiple stepwise variants per each cell-free nucleic acid molecule identified as containing multiple stepwise variants (i.e., two or more stepwise variants), or (iii) the number of cell-free nucleic acid molecules identified as containing multiple stepwise variants divided by the total number of cell-free nucleic acid molecules containing mutations that overlap with some of the multiple stepwise variants (i.e., stepwise variant allele frequency). Based on this comparison, the MRD of the subject's condition (e.g., cancer or tumor) can be determined. For example, the subject's tumor burden or cancer burden can be determined based on this comparison.

[0182] In some cases, the progression of the condition can be progression or worsening of the condition. In one example, the worsening of the condition can include the progression of cancer from an early stage to a later stage, such as from stage I cancer to stage III cancer. In another example, the worsening of the condition can include an increase in size (e.g., volume) of a solid tumor. In yet another example, the worsening of the condition can include cancer metastasis from one location to another within the subject's body.

[0183] In some examples, (i) the total number of cell-free nucleic acid molecules identified as comprising a plurality of graded variants from the second state of the subject's condition is at least or at most about 0.1-fold, at least or at most about 0.2-fold, at least or at most about 0.3-fold, at least or at most about 0.4-fold, at least or at most about 0.5-fold, at least or at most about 0.6-fold, at least or at most about 0.7-fold, at least or at most about 0.8-fold, at least or at most about 0.9-fold, at least or at most about 1-fold, at least or at most about 2-fold, at least or at most about 3-fold, at least or at most about 4-fold less than (ii) the total number of cell-free nucleic acid molecules identified as comprising a plurality of graded variants from the first state of the subject's condition. The antibody may be at least or up to about 5 times greater, at least or up to about 6 times greater, at least or up to about 7 times greater, at least or up to about 8 times greater, at least or up to about 9 times greater, at least or up to about 10 times greater, at least or up to about 15 times greater, at least or up to about 20 times greater, at least or up to about 30 times greater, at least or up to about 40 times greater, at least or up to about 50 times greater, at least or up to about 60 times greater, at least or up to about 70 times greater, at least or up to about 80 times greater, at least or up to about 90 times greater, at least or up to about 100 times greater, at least or up to about 200 times greater, at least or up to about 300 times greater, at least or up to about 400 times greater, or at least or up to about 500 times greater.

[0184] In some examples, (i) the average number of stepwise variants per each cell-free nucleic acid molecule identified as comprising a plurality of stepwise variants from the second state of the subject's condition is at least or at most about 0.1-fold, at least or at most about 0.2-fold, at least or at most about 0.3-fold, at least or at most about 0.4-fold, at least or at most about 0.5-fold, at least or at most about 0.6-fold, at least or at most about 0.7-fold, at least or at most about 0.8-fold, at least or at most about 0.9-fold, at least or at most about 1-fold, at least or at most about 2-fold, at least or at most about 3 ... It may be at least or up to about 4 times, at least or up to about 5 times, at least or up to about 6 times, at least or up to about 7 times, at least or up to about 8 times, at least or up to about 9 times, at least or up to about 10 times, at least or up to about 15 times, at least or up to about 20 times, at least or up to about 30 times, at least or up to about 40 times, at least or up to about 50 times, at least or up to about 60 times, at least or up to about 70 times, at least or up to about 80 times, at least or up to about 90 times, at least or up to about 100 times, at least or up to about 200 times, at least or up to about 300 times, at least or up to about 400 times, or at least or up to about 500 times greater.

[0185] In some cases, progression of the condition can be regression or at least partial remission of the condition. In one example, at least partial remission of the condition can include downstaging of the cancer from a later stage to an earlier stage, such as from stage IV cancer to stage II cancer. Alternatively, at least partial remission of the condition can be complete remission from the cancer. In another example, at least partial remission of the condition can include a decrease in the size (e.g., volume) of a solid tumor.

[0186] In some examples, (i) the total number of cell-free nucleic acid molecules identified as comprising a plurality of graded variants from the second state of the subject's condition is at least or at most about 0.1-fold, at least or at most about 0.2-fold, at least or at most about 0.3-fold, at least or at most about 0.4-fold, at least or at most about 0.5-fold, at least or at most about 0.6-fold, at least or at most about 0.7-fold, at least or at most about 0.8-fold, at least or at most about 0.9-fold, at least or at most about 1-fold, at least or at most about 2-fold, at least or at most about 3-fold, at least or at most about 4-fold less than (ii) the total number of cell-free nucleic acid molecules identified as comprising a plurality of graded variants from the first state of the subject's condition. The size may be at least or up to about 5 times smaller, at least or up to about 6 times smaller, at least or up to about 7 times smaller, at least or up to about 8 times smaller, at least or up to about 9 times smaller, at least or up to about 10 times smaller, at least or up to about 15 times smaller, at least or up to about 20 times smaller, at least or up to about 30 times smaller, at least or up to about 40 times smaller, at least or up to about 50 times smaller, at least or up to about 60 times smaller, at least or up to about 70 times smaller, at least or up to about 80 times smaller, at least or up to about 90 times smaller, at least or up to about 100 times smaller, at least or up to about 200 times smaller, at least or up to about 300 times smaller, at least or up to about 400 times smaller, or at least or up to about 500 times smaller.

[0187] In some examples, (i) the average number of stepwise variants per each cell-free nucleic acid molecule identified as comprising a plurality of stepwise variants from the second state of the subject's condition is at least or at most about 0.1-fold, at least or at most about 0.2-fold, at least or at most about 0.3-fold, at least or at most about 0.4-fold, at least or at most about 0.5-fold, at least or at most about 0.6-fold, at least or at most about 0.7-fold, at least or at most about 0.8-fold, at least or at most about 0.9-fold, at least or at most about 1-fold, at least or at most about 2-fold, at least or at most about 3 ... It may be at least or up to about 4 times, at least or up to about 5 times, at least or up to about 6 times, at least or up to about 7 times, at least or up to about 8 times, at least or up to about 9 times, at least or up to about 10 times, at least or up to about 15 times, at least or up to about 20 times, at least or up to about 30 times, at least or up to about 40 times, at least or up to about 50 times, at least or up to about 60 times, at least or up to about 70 times, at least or up to about 80 times, at least or up to about 90 times, at least or up to about 100 times, at least or up to about 200 times, at least or up to about 300 times, at least or up to about 400 times, or at least or up to about 500 times smaller.

[0188] In some cases, the progression of the condition can remain substantially the same between the two states of the subject's condition. In some examples, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from the second state of the subject's condition can be approximately the same as (ii) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from the first state of the subject's condition. In some examples, (i) the average number of multiple stepwise variants per each cell-free nucleic acid molecule identified as containing multiple stepwise variants from the second state of the subject's condition can be approximately the same as (ii) the average number of multiple stepwise variants per each cell-free nucleic acid molecule identified as containing multiple stepwise variants from the first state of the subject's condition.

[0189] In some embodiments of any one of the methods disclosed herein, one or more cell-free nucleic acid molecules containing multiple stepwise variants can be identified from the multiple cell-free nucleic acid molecules by one or more sequencing methods. Alternatively, or in addition, one or more cell-free nucleic acid molecules containing multiple stepwise variants can be identified by being pulled down from (or captured among) the multiple cell-free nucleic acid molecules using a set of nucleic acid probes. The pull-down (or capture) method via a set of nucleic acid probes may be sufficient to identify one or more cell-free nucleic acid molecules of interest without sequencing. In some cases, the set of nucleic acid probes may be configured to hybridize to at least a portion of cell-free nucleic acid (e.g., cfDNA) molecules from one or more genomic regions associated with the subject's condition. Thus, the presence of one or more cell-free nucleic acid molecules pulled down by the set of nucleic acid probes may indicate that the one or more cell-free nucleic acid molecules are derived from the condition (e.g., ctDNA or ctRNA). Further details of nucleic acid probe sets are disclosed elsewhere in this disclosure.

[0190] In some embodiments of any one of the methods disclosed herein, based on sequencing data derived from a plurality of cell-free nucleic acid molecules (e.g., cfDNA) obtained from or derived from a subject, (i) one or more cell-free nucleic acid molecules identified as containing a plurality of stepwise variants can be separated in silico from (ii) one or more other cell-free nucleic acid molecules not identified as containing a plurality of stepwise variants (or one or more other cell-free nucleic acid molecules that do not contain a plurality of stepwise variants). In some cases, the method may further include (i) generating additional data that includes sequencing information of only one or more other cell-free nucleic acid molecules not identified as containing a plurality of stepwise variants (or one or more other cell-free nucleic acid molecules that do not contain a plurality of stepwise variants). In some cases, the method may further include (ii) generating different data that includes sequencing information of only one or more other cell-free nucleic acid molecules not identified as containing a plurality of stepwise variants (or one or more other cell-free nucleic acid molecules that do not contain a plurality of stepwise variants).

[0191] In one embodiment, the present disclosure provides a method for determining a subject's status, as shown by flowchart 2560 in Figure 25F. The method may include (a) providing a mixture containing (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained or derived from the subject (process 2562). In some cases, individual nucleic acid probes of the set of nucleic acid probes can be designed to hybridize to a target cell-free nucleic acid molecule containing multiple stepwise variants relative to a reference genome sequence, separated by at least one nucleotide. Thus, as disclosed herein, a first stepwise variant of the multiple stepwise variants and a second stepwise variant of the multiple stepwise variants can be separated by at least one nucleotide. In some cases, the individual nucleic acid probes can include an activatable reporter agent. The activatable reporter agent can be activated by either (i) hybridization of the individual nucleic acid probes to the multiple stepwise variants and (ii) dehybridization of at least some of the individual nucleic acid probes hybridized to the multiple stepwise variants. The method may further include (b) detecting the activated reporter agent to identify one or more cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules (process 2564). Each of the one or more cell-free nucleic acid molecules may comprise a plurality of graded variants. The method may optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules to determine the subject's status (process 2566).

[0192] In one aspect, the present disclosure provides a method for determining a subject's status, as shown by flowchart 2570 in Figure 25G. The method may include (a) providing a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained or derived from the subject (process 2572). In some cases, individual nucleic acid probes of the set of nucleic acid probes can be designed to hybridize to target cell-free nucleic acid molecules comprising a plurality of graded variants relative to a reference genome sequence. In some cases, the individual nucleic acid probes can comprise an activatable reporter agent. The activatable reporter agent can be activated by either (i) hybridization of the individual nucleic acid probes to the plurality of graded variants and (ii) dehybridization of at least some of the individual nucleic acid probes hybridized to the plurality of graded variants. The method may further include (b) detecting the activated reporter agent to identify one or more cell-free nucleic acid molecules among the plurality of cell-free nucleic acid molecules (process 2574). Each of the one or more cell-free nucleic acid molecules can comprise a plurality of graded variants, and as disclosed herein, the LOD of the identifying step can be less than about 1 in 50,000 cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules. The method can optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules to determine the status of the subject (process 2576).

[0193] In some cases, as disclosed herein, a first graded variant of the plurality of graded variants and a second graded variant of the plurality of graded variants are separated by at least one nucleotide.

[0194] In some cases, the LOD for identifying one or more cell-free nucleic acid molecules as disclosed herein is less than about 1 in 60,000, less than 1 in 70,000, less than 10 in 80,000, less than 1 in 90,000, less than 1 in 100,000, less than 1 in 150,000, less than 1 in 200,000, less than 1 in 300,000, less than 1 in 400,000, less than 1 in 500,000, less than 1 in 600,000, less than 1 in 70,000, less than 1 ...100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in 100,000, less than 1 in The LOD may be less than 1 in 0 million, less than 1 in 800,000, less than 1 in 900,000, less than 1 in 1 million, less than 1 in 1 million, less than 1 in 1.1 million, less than 1.2 million, less than 1.3 million, less than 1.4 million, less than 1.5 million, less than 2 million, less than 1 in 2.5 million, less than 1 in 3 million, less than 1 in 4 million, or less than 1 in 5 million cell-free nucleic acid molecules. Generally, the lower the LOD, the higher the sensitivity of such detection.

[0195] In some embodiments of any one of the methods disclosed herein, the method can further include mixing (1) the set of nucleic acid probes and (2) the plurality of cell-free nucleic acid molecules.

[0196] In some embodiments of any one of the methods disclosed herein, the activatable reporter agent of the nucleic acid probe can be activated upon hybridization to multiple graded variants of an individual nucleic acid probe. Non-limiting examples of such nucleic acid probes can include molecular beacons, eclipse probes, amplifluor probes, scorpion PCR primers, and light-upon-extension fluorogenic PCR primers (LUX primers).

[0197] For example, the nucleic acid probe can be a molecular beacon, as shown in Figure 26A. The molecular beacon can be a fluorescently labeled (e.g., dye-labeled) oligonucleotide probe that contains complementarity to a target cell-free nucleic acid molecule 2603 within a region containing multiple stepwise variants. The molecular beacon can have a length of about 25 nucleotides to about 50 nucleotides. The molecular beacon can also be designed to be partially self-complementary, forming a hairpin structure with stem 2601a and loop 2601b. The 5' and 3' ends of the molecular beacon probe can have complementary sequences (e.g., about 5-6 nucleotides) that form stem structure 2601a. ​​The loop portion 2601b of the hairpin can be designed to specifically hybridize to a portion (e.g., about 15-30 nucleotides) of the target sequence containing two or more stepwise variants. The hairpin can be designed to hybridize to a portion containing at least two, three, four, five, or more stepwise variants. A fluorescent reporter molecule can be attached to the 5' end of the molecular beacon probe, and a quencher that quenches the fluorescence of the fluorescent reporter can be attached to the 3' end of the molecular beacon probe. Thus, the formation of a hairpin can bring the fluorescent reporter and quencher together, resulting in no fluorescence emission. However, during the annealing operation of an amplification reaction of multiple cell-free nucleic acid molecules obtained from or derived from a subject, the loop portion of the molecular beacon can bind to its target sequence and denature the stem. Thus, the reporter and quencher can be separated, quenching can be disabled, and the fluorescent reporter is activated and detectable. Since the fluorescence of the fluorescent reporter is only emitted from the molecular beacon probe when the probe binds to the target sequence, the amount or level of detected fluorescence can be proportional to the amount of target in the reaction (e.g., as disclosed herein, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants in each situation, or (ii) the average number of multiple stepwise variants per each cell-free nucleic acid molecule identified as containing multiple stepwise variants).

[0198] In some embodiments of any one of the methods disclosed herein, the activatable reporter agent can be activated upon dehybridization of at least a portion of the individual nucleic acid probes hybridized to the multiple stepwise variants. In other words, when the individual nucleic acid probes are hybridized to a portion of a target cell-free nucleic acid molecule containing multiple stepwise variants, dehybridization of at least a portion of the individual nucleic acid probes and the target cell-free nucleic acid can activate the activatable reporter agent. Non-limiting examples of such nucleic acid probes include hydrolysis probes (e.g., TaqMan probes), dual hybridization probes, and QZyme PCR primers.

[0199] For example, the nucleic acid probe can be a hydrolysis probe, as shown in Figure 26B. Hydrolysis probe 2611 can be a fluorescently labeled oligonucleotide probe capable of specifically hybridizing to a portion (e.g., about 10 to about 25 nucleotides) of a target cell-free nucleic acid molecule 2613, where the hybridized portion contains two or more graded variants. Hydrolysis probe 2611 can be labeled with a fluorescent reporter at the 5' end and a quencher at the 3' end. When the hydrolysis probe is intact (e.g., not cleaved), the reporter fluorescence is quenched due to its proximity to the quencher (Figure 26B). During the annealing step of an amplification reaction of multiple cell-free nucleic acid molecules obtained from or derived from a subject, the 5' to 3' exonuclease activity of a specific thermostable polymerase (e.g., Taq or Tth) is activated. Amplification reactions of multiple cell-free nucleic acid molecules obtained from or derived from a subject can include a combined annealing / extension procedure in which a hydrolysis probe hybridizes to the target cell-free nucleic acid molecule and the dsDNA-specific 5' to 3' exonuclease activity of a thermostable polymerase (e.g., Taq or Tth) cleaves the fluorescent reporter from the hydrolysis probe, resulting in separation of the fluorescent reporter from the quencher and the generation of a fluorescent signal proportional to the amount of target in the sample (e.g., as disclosed herein, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants in each instance, or (ii) the average number of multiple stepwise variants per each cell-free nucleic acid molecule identified as containing multiple stepwise variants).

[0200] In some embodiments of any one of the methods disclosed herein, the reporter agent can comprise a fluorescent reporter. Non-limiting examples of fluorescent reporters include fluorescein amidite (FAM, 2-[3-(dimethylamino)-6-dimethyliminio-xanthen-9-yl]benzoate TAMRA, (2E)-2-[(2E,4E)-5-(2-tert-butyl-9-ethyl-6,8,8-trimethyl-pyrano[3,2-g]quinolin-1-ium-4-yl)penta-2,4-dienylidene]-1-(6-hydroxy-6-oxo-hexyl)-3,3-dimethyl-indoline-5-sulfonate Dy 750, 6-carboxy-2',4,4',5',7,7'-hexachlorofluorescein, 4,5,6,7-tetrachlorofluorescein TET™, sulforhodamine 101 acid chloride succinimidyl ester Texas Red-X, and ALEXA. Dyes, Bodipy Dyes, Cyanine Dyes, Rhodamine 123 (hydrochloride), Well RED Dyes, MAX, and TEX 613. In some cases, the reporter agent further comprises a quencher as disclosed herein. Non-limiting examples of quenchers include Black Hole Quencher, Iowa Black Quencher, and 4-dimethylaminoazobenzene-4'-sulfonyl chloride (DABCYL).

[0201] In some embodiments of any one of the methods disclosed herein, any PCR reaction utilizing the set of nucleic acid probes can be performed using real-time PCR (qPCR). Alternatively, the PCR reaction utilizing the set of nucleic acid probes can be performed using digital PCR (dPCR).

[0202] Figure 24 provides an example flowchart of a process for performing clinical intervention and / or treatment based on the detection of circulating tumor nucleic acid in an individual's biological sample. In some embodiments, the detection of circulating tumor nucleic acid is determined by the detection of an in-phase somatic variant in a cell-free nucleic acid sample. In many embodiments, the detection of circulating tumor nucleic acid indicates the presence of cancer, and therefore appropriate clinical intervention and / or treatment can be performed.

[0203] Referring to Figure 24, process 2400 can begin with obtaining, preparing, and sequencing (2401) cell-free nucleic acids obtained from a non-invasive biopsy (e.g., a liquid or waste biopsy) using a capture sequencing approach across a region shown to have multiple genetic mutations or variants occurring in phase. In some embodiments, cfDNA and / or cfRNA are extracted from plasma, blood, lymph, saliva, urine, feces, and / or other suitable bodily fluids. The cell-free nucleic acids can be isolated and purified by any suitable means. In some embodiments, column purification is utilized (e.g., QIAamp Circulating Nucleic Acid Kit, Qiagen, Hilden, Germany). In some embodiments, the isolated RNA fragments can be converted to complementary DNA for further downstream analysis.

[0204] In some embodiments, a biopsy is taken before any signs of cancer. In some embodiments, a biopsy is taken to provide an early screening for detecting cancer. In some embodiments, a biopsy is taken to detect whether residual cancer is present after treatment. In some embodiments, a biopsy is taken during treatment to determine whether the treatment is providing the desired response. Screening for any particular cancer can be performed. In some embodiments, screening is performed to detect cancers that express somatic stage variants in typical regions of the genome, such as (for example) lymphoma. In some embodiments, screening is performed to detect cancers in which somatic stage variants are found, utilizing a previously taken cancer biopsy.

[0205] In some embodiments, the biopsy is taken from an individual determined to be at risk for developing cancer, such as an individual with a family history of the disorder or a determined risk factor (e.g., exposure to a carcinogen). In many embodiments, the biopsy is taken from any individual within the general population. In some embodiments, the biopsy is taken from an individual within a specific age group at higher risk for cancer, such as an elderly individual over the age of 50. In some embodiments, the biopsy is taken from an individual who has been diagnosed with and treated for cancer.

[0206] In some embodiments, the extracted cell-free nucleic acid is prepared for sequencing. Thus, the cell-free nucleic acid is converted into a molecular library for sequencing. In some embodiments, adapters and / or primers are attached to the cell-free nucleic acid to facilitate sequencing. In some embodiments, targeted sequencing of specific genomic loci is to be performed, and therefore, specific sequences corresponding to specific loci are captured by hybridization prior to sequencing (e.g., capture sequencing). In some embodiments, capture sequencing is performed using a probe set that pulls down (or captures) regions discovered to share stage variants for a particular cancer (e.g., lymphoma). In some embodiments, capture sequencing is performed using a probe set that pulls down (or captures) regions discovered to share stage variants previously determined by sequencing cancer biopsies. A more detailed discussion of capture sequencing and probes is provided in the section entitled "Capture Sequencing."

[0207] In some embodiments, any suitable sequencing technology capable of detecting stage variants indicative of circulating tumor nucleic acids can be utilized, including, but not limited to, 454 sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent sequencing, single-read sequencing, paired-end sequencing, etc.

[0208] Process 2400 analyzes (2403) the cell-free nucleic acid sequencing results to detect circulating tumor nucleic acid sequences, as determined by detection of in-phase somatic variants. As cancers actively grow and spread, neoplastic cells often release biomolecules (especially nucleic acids) into the vascular, lymphatic, and / or waste systems. Furthermore, due to biophysical constraints in their local environment, neoplastic cells often rupture, releasing their internal cellular contents into the vascular, lymphatic, and / or waste systems. Thus, it is possible to detect distant primary tumors and / or metastases from liquid or waste biopsies.

[0209] Detection of circulating tumor nucleic acid sequences indicates the presence of cancer in the individual being tested. Thus, based on the detection of circulating tumor nucleic acids, clinical intervention and / or treatment can be performed (2405). In some embodiments, clinical procedures such as (e.g.,) blood tests, genetic testing, medical imaging, physical examination, tumor biopsy, or any combination thereof are performed. In some embodiments, a diagnosis is performed to determine the specific stage of the cancer. In some embodiments, a treatment is performed such as (e.g.,) chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, hormone therapy, targeted drug therapy, surgery, transplant, blood transfusion, medical monitoring, or any combination thereof. In some embodiments, the individual is evaluated and / or treated by a medical professional, such as a doctor, physician, physician's assistant, nurse practitioner, nurse, caregiver, nutritionist, etc.

[0210] Various embodiments of the present disclosure are directed to utilizing cancer detection to perform clinical intervention. In some embodiments, an individual has a fluid or waste biopsy screened and processed by the methods described herein to indicate that the individual has cancer and therefore should undergo intervention. Clinical intervention includes clinical procedures and treatments. Clinical procedures include, but are not limited to, blood tests, genetic testing, medical imaging, physical examinations, and tumor biopsies. Treatments include, but are not limited to, chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, hormone therapy, targeted drug therapy, surgery, transplants, blood transfusions, and medical surveillance. In some embodiments, a diagnosis is performed to determine the specific stage of cancer. In some embodiments, the individual is evaluated and / or treated by a medical professional, such as a physician, doctor, physician's assistant, nurse practitioner, nurse, caregiver, or nutritionist.

[0211] In some embodiments described herein, cancer can be detected using sequencing results of cell-free nucleic acid derived from blood, serum, cerebrospinal fluid, lymph, urine, or feces. In many embodiments, cancer is detected when the sequencing results contain one or more somatic variants present in phase within a short genetic window, such as the length of the cell-free molecule (e.g., approximately 170 bp). In many embodiments, statistical methods are used to determine whether the presence of phased variants is due to a cancerous source (as opposed to a molecular artifact or other biological source). Various embodiments utilize Monte Carlo sampling as a statistical method to determine whether the cell-free nucleic acid sequencing results contain a sequence of circulating tumor nucleic acid based on a score determined by the presence of phased variants. Thus, in some embodiments, cell-free nucleic acid is extracted, processed, and sequenced, and the sequencing results are analyzed to detect cancer. This process is particularly useful in clinical practice to provide diagnostic scans.

[0212] An exemplary procedure for diagnostic scanning of an individual for B-cell cancer is as follows. (a) extracting a fluid or waste biopsy from an individual; (b) preparing and performing targeted sequencing of cell-free nucleic acid from the biopsy using a nucleic acid probe specific for the B-cell cancer; (c) detecting staged variants in sequencing results representing circulating tumor nucleic acid sequences; (d) implementing clinical interventions based on the detection of circulating tumor nucleic acid sequences;

[0213] An exemplary procedure for personalized diagnostic scanning of an individual for cancer that has been previously sequenced to detect staged variants at specific genomic loci is as follows. Extracting cancer biopsies from individual sequences to detect staged variants accumulated in cancer (a) designing and synthesizing a nucleic acid probe for a genomic locus that includes the location of the detected phase variant; (b) extracting a fluid or waste biopsy from the individual; (c) preparing and performing targeted sequencing of cell-free nucleic acids from biopsies utilizing the designed and synthesized nucleic acid probes; (d) detecting staged variants in sequencing results representing circulating tumor nucleic acid sequences; (e) implementing clinical interventions based on the detection of circulating tumor nucleic acid sequences;

[0214] In some embodiments of any one of the methods disclosed herein, at least a portion of the identified one or more cell-free nucleic acid molecules containing multiple stepwise variants can be further analyzed to determine the subject's status. In such analysis, (i) the identified one or more cell-free nucleic acid molecules and (ii) other cell-free nucleic acid molecules of the multiple cell-free nucleic acid molecules that do not contain multiple stepwise variants can be analyzed as different variables. In some cases, the ratio of (i) the number of identified one or more cell-free nucleic acid molecules to (ii) the number of other cell-free nucleic acid molecules of the multiple cell-free nucleic acid molecules that do not contain multiple stepwise variants can be used as a factor for determining the subject's status. In some cases, (i) the position(s) of the identified one or more cell-free nucleic acid molecules relative to a reference genome sequence and (ii) the position(s) of the other cell-free nucleic acid molecules of the multiple cell-free nucleic acid molecules that do not contain multiple stepwise variants relative to a reference genome sequence can be used as a factor for determining the subject's status.

[0215] Alternatively, in some cases, analysis of one or more identified cell-free nucleic acid molecules comprising a plurality of graded variants to determine a subject's status may not, and need not, be based on other cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules that do not comprise a plurality of graded variants. As disclosed herein, non-limiting examples of information or characteristics of one or more cell-free nucleic acid molecules comprising a plurality of graded variants can include (i) the total number of such cell-free nucleic acid molecules, and (ii) the average number of graded variations per each nucleic acid molecule in the population of identified cell-free nucleic acid molecules.

[0216] Thus, in some embodiments of any one of the methods disclosed herein, the number of multiple stepwise variants from one or more cell-free nucleic acid molecules identified as having multiple stepwise variants can indicate the status of a subject. In some cases, the ratio of (i) the number of multiple stepwise variants from one or more cell-free nucleic acid molecules to (ii) the number of single-base variants from one or more cell-free nucleic acid molecules can indicate the status of a subject. For example, a particular condition (e.g., follicular lymphoma) may exhibit a signature ratio that differs from the signature ratio of another condition (e.g., breast cancer). In some examples, for cancer or solid tumors, the ratios disclosed herein can be from about 0.01 to about 0.20. In some examples, for cancer or solid tumors, the ratios disclosed herein can be about 0.01, about 0.02, about 0.03, about 0.04, about 0.05, about 0.06, about 0.07, about 0.08, about 0.09, about 0.10, about 0.11, about 0.12, about 0.13, about 0.14, about 0.15, about 0.16, about 0.17, about 0.18, about 0.19, or about 0.20. In some examples, for cancers or solid tumors, the ratios disclosed herein can be at least or at most about 0.01, at least or at most about 0.02, at least or at most about 0.03, at least or at most about 0.04, at least or at most about 0.05, at least or at most about 0.06, at least or at most about 0.07, at least or at most about 0.08, at least or at most about 0.09, at least or at most about 0.10, at least or at most about 0.11, at least or at most about 0.12, at least or at most about 0.13, at least or at most about 0.14, at least or at most about 0.15, at least or at most about 0.16, at least or at most about 0.17, at least or at most about 0.18, at least or at most about 0.19, or at least or at most about 0.20.

[0217] In some embodiments of any one of the methods disclosed herein, the frequency of multiple graded variants in one or more identified cell-free nucleic acid molecules can indicate the state of the subject. In some cases, based on the sequencing data disclosed herein, the average frequency of multiple graded variants per predetermined bin length (e.g., bins of about 50 base pairs) in each of the identified cell-free nucleic acid molecules can indicate the state of the subject. In some cases, based on the sequencing data disclosed herein, the average frequency of multiple graded variants per predetermined bin length (e.g., bins of about 50 base pairs) in each of the identified cell-free nucleic acid molecules related to a specific gene (e.g., BCL2, PIM1) can indicate the state of the subject. The bin size can be about 30, about 40, about 50, about 60, about 70, or about 80.

[0218] In some examples, a first condition (e.g., Hodgkin's lymphoma or HL) can exhibit a first average frequency, and a second condition (e.g., DLBCL) can exhibit a different average frequency, thereby enabling identification and / or determination of whether a subject has or is suspected of having a particular condition. In some examples, a first subtype of a disease can exhibit a first average frequency, and a second subtype of the same disease can exhibit a different average frequency, thereby enabling identification and / or determination of whether a subject has or is suspected of having a particular subtype of a disease. For example, a subject may have DLBCL, and as disclosed herein, one or more cell-free nucleic acid molecules derived from germinal center B cell (GCB) DLBCL or activated B cell (ABC) DLBCL may have different average frequencies of multiple graded variants per predetermined bin length.

[0219] In some examples, a subject's condition can have a predetermined number of graded variants (i.e., a predetermined frequency of graded variants) across a predetermined genomic locus. If the predetermined frequency of graded variants matches the frequency of a plurality of graded variants in one or more cell-free nucleic acid molecules identified from a plurality of cell-free nucleic acid molecules from the subject, it can indicate that the subject has such a condition.

[0220] In some embodiments of any one of the methods disclosed herein, one or more cell-free nucleic acid molecules identified as containing multiple staged variants can be analyzed to determine their genomic origin (e.g., which gene loci they originate from). Because different diseases can have multiple staged variants in different signature genes, the genomic origin of the identified one or more cell-free nucleic acid molecules can indicate the subject's condition. For example, a subject can have GCB DLBCL, and one or more cell-free nucleic acid molecules from the subject's GCB can have predominant staged variants in the BCL2 gene, while one or more cell-free nucleic acid molecules from the same subject's ABC can not contain as many staged variants in the BCL2 gene as those from the GCB. On the other hand, a subject can have ABC DLBCL, and one or more cell-free nucleic acid molecules from the subject's ABC can have predominant staged variants in the PIM1 gene, while one or more cell-free nucleic acid molecules from the same subject's GCB can not contain as many staged variants in the PIM1 gene as those from the ABC.

[0221] In some embodiments of any one of the methods disclosed herein, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, at least or up to about 50%, at least or up to about 55%, at least or up to about 60%, at least or up to about 65%, at least or up to about 70%, at least or up to about 75%, at least or up to about 80%, at least or up to about 85%, at least or up to about 90%, at least or up to about 95%, at least or up to about 99%, or about 100% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least two nucleotides away from an adjacent SNV.

[0222] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least 3 nucleotides away from an adjacent SNV.

[0223] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least 4 nucleotides away from an adjacent SNV.

[0224] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least 5 nucleotides away from an adjacent SNV.

[0225] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least 6 nucleotides away from an adjacent SNV.

[0226] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single nucleotide variant (SNV) that is at least 7 nucleotides away from an adjacent SNV.

[0227] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least 8 nucleotides away from an adjacent SNV.

[0228] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least 9 nucleotides away from an adjacent SNV.

[0229] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants can comprise a single base variant (SNV) that is at least 10 nucleotides away from an adjacent SNV. C. Reference genome sequence

[0230] In some embodiments of any one of the methods disclosed herein, the reference genome sequence may be at least a portion of a nucleic acid sequence database (i.e., a reference genome), which is constructed from genetic data and is intended to represent the genome of a reference cohort. In some cases, the reference cohort may be a collection of individuals from a particular or diverse genotype, haplotype, demographic, gender, nationality, age, ethnicity, relatives, physical condition (e.g., healthy or diagnosed with the same or different condition, such as a particular type of cancer), or other grouping. The reference genome sequence disclosed herein may be a mosaic (or consensus sequence) of the genomes of two or more individuals. The reference genome sequence may include at least a portion of a publicly available reference genome or an unofficial reference genome. Non-limiting examples of human reference genomes include hg19, hg18, hg17, hg16, and hg38.

[0231] In some examples, the reference genome sequence is at least or up to about 500 nucleic acid bases, at least or up to about 1 kilobase (kb), at least or up to about 2 kb, at least or up to about 3 kb, at least or up to about 4 kb, at least or up to about 5 kb, at least or up to about 6 kb, at least or up to about 7 kb, at least or up to about 8 kb, at least or up to about 9 kb, at least or up to about 10 kb, at least or up to about 20 kb, at least or up to about 30 kb, at least or up to about 40 kb, at least or up to about 50 kb, at least or up to about 60 kb, at least or up to about 70 kb, at least or up to about 80 kb, at least or up to about 90 kb, at least or up to about 100 kb, at least or up to about 200 kb, at least or up to about 300 kb, at least or up to about 400 kb, at least or up to about 500 kb, at least or up to about 600 kb, At least or up to about 700 kb, at least or up to about 800 kb, at least or up to about 900 kb, at least or up to about 1,000 kb, at least or up to about 2,000 kb, at least or up to about 3,000 kb, at least or up to about 4,000 kb, at least or up to about 5,000 kb, at least or up to about 6,000 kb, at least or up to about 7,000 kb, at least or up to about 8,000 kb, at least or up to about 9,000 kb 00 kb, at least or up to about 10,000 kb, at least or up to about 20,000 kb, at least or up to about 30,000 kb, at least or up to about 40,000 kb, at least or up to about 50,000 kb, at least or up to about 60,000 kb, at least or up to about 70,000 kb, at least or up to about 80,000 kb, at least or up to about 90,000 kb, or at least or up to about 100,000 kb.

[0232] In some cases, the reference genome sequence can be the entire reference genome or a portion of the genome (e.g., a portion related to a condition of interest). For example, the reference genome sequence can consist of at least 1, 2, 3, 4, 5, or more genes that undergo abnormal somatic hypermutation under a specific type of cancer. In some cases, the reference genome sequence can be the entire chromosome sequence or a fragment thereof. In some cases, the reference genome sequence can include two or more (e.g., at least 2, 3, 4, 5, or more) different portions of the reference genome that are not adjacent to each other (e.g., within the same chromosome or from different chromosomes).

[0233] In some embodiments of any one of the methods disclosed herein, the reference genome sequence can be at least a portion of the reference genome of a selected individual, e.g., a healthy individual or a subject of any of the methods disclosed herein.

[0234] In some cases, the reference genome sequence may be derived from an individual other than the subject (e.g., a healthy control individual). Alternatively, in some cases, the reference genome sequence may be derived from a sample from the subject. In some examples, the sample may be a healthy sample from the subject. The healthy sample from the subject may be any healthy subject's cells, such as healthy white blood cells. By comparing the sequencing data of a plurality of cell-free nucleic acid molecules (e.g., cfDNA molecules) from a subject with at least a portion of the genome sequence of healthy cells from the same subject, one or more cell-free nucleic acid molecules containing multiple step variants can be identified and analyzed as disclosed herein. In some examples, the sample may be a disease sample from the subject, such as disease cells (e.g., tumor cells) or a solid tumor. The reference genome sequence may be obtained by sequencing at least a portion of the subject's disease cells or by sequencing multiple cell-free nucleic acid molecules obtained from the subject's solid tumor. Once a subject is diagnosed with a particular condition (e.g., a disease), the subject's reference genome sequence containing multiple step variants can be used to determine whether the subject will still exhibit the same step variants at a future time point. In this regard, any new incremental variants identified between a subject's "affected" reference genomic sequence and new cell-free nucleic acid molecules obtained or derived from the subject may indicate a reduction (e.g., at least partial remission) in the degree of aberrant somatic hypermutation, particularly at genomic loci.

[0235] In various embodiments, the diagnostic scan is for the detection of acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), anal cancer, astrocytoma, basal cell carcinoma, bile duct cancer, bladder cancer, breast cancer, Burkitt's lymphoma, cervical cancer, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myeloproliferative neoplasms, colorectal cancer, diffuse large B-cell lymphoma, endometrial cancer, ependymoma, esophageal cancer, esthesioneuroblastoma, Ewing's sarcoma, fallopian tube cancer, follicular lymphoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, hairy cell leukemia, hepatocellular carcinoma, Hodgkin's lymphoma, hypopharyngeal cancer, capsulloblastoma, leukemia ... It can be performed for any neoplasm type, including, but not limited to, sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, Merkel cell carcinoma, mesothelioma, oral cancer, neuroblastoma, non-Hodgkin's lymphoma, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumors, pharyngeal cancer, pituitary tumors, prostate cancer, rectal cancer, renal cell carcinoma, retinoblastoma, skin cancer, small cell lung cancer, small intestine cancer, squamous cell cervical carcinoma, T-cell lymphoma, testicular cancer, thymoma, thyroid cancer, uterine cancer, vaginal cancer, and vascular tumors.

[0236] In some embodiments, diagnostic scans are utilized to provide early detection of cancer. In some embodiments, diagnostic scans detect cancer in individuals with stage I, II, or III cancer. In some embodiments, diagnostic scans are utilized to detect MRD or tumor burden. In some embodiments, diagnostic scans are utilized to determine treatment progress (e.g., progression or regression). Clinical procedures and / or treatments can be performed based on diagnostic scans. D. Nucleic Acid Probes

[0237] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes can be designed based on any of the subject reference genome sequences disclosed herein.In some cases, the set of nucleic acid probes can be designed based on a plurality of stage variants identified by comparing (i) sequencing data from the subject's solid tumor with (ii) sequencing data from the subject's or a healthy cohort's healthy cells, as disclosed herein.The set of nucleic acid probes can be designed based on a plurality of stage variants identified by comparing (i) sequencing data from the subject's solid tumor with (ii) sequencing data from the subject's healthy cells.The set of nucleic acid probes can be designed based on a plurality of stage variants identified by comparing (i) sequencing data from the subject's solid tumor with (ii) sequencing data from the subject's healthy cells, as disclosed herein.

[0238] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes is designed to hybridize with the sequence of a genomic locus associated with a condition.As disclosed herein, if a subject has a condition, the genomic locus associated with the condition can be determined to undergo or exhibit abnormal somatic hypermutation.Alternatively, the set of nucleic acid probes is designed to hybridize with the sequence of a typical region.

[0239] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes may be designed to hybridize to at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or about 100% of the genomic regions identified in Table 1.

[0240] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes can be designed to hybridize to at least a portion of cell-free nucleic acid (e.g., cfDNA) molecules derived from at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or about 100% of the genomic regions identified in Table 1.

[0241] In some embodiments of any one of the methods disclosed herein, each nucleic acid probe of the set of nucleic acid probes can have at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90% sequence identity, at least about 95% sequence identity, at least about 99%, or about 100% sequence identity to a probe sequence selected from Table 6.

[0242] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes may comprise at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or about 100% of the probe sequences in Table 6.

[0243] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes is at least or up to about 500 nucleic acid bases, at least or up to about 1 kilobase (kb), at least or up to about 2 kb, at least or up to about 3 kb, at least or up to about 4 kb, at least or up to about 5 kb, at least or up to about 6 kb, at least or up to about 7 kb, at least or up to about 8 kb, at least or up to about 9 kb, at least or up to about 10 kb, at least or up to about 20 kb. b, at least or up to about 30 kb, at least or up to about 40 kb, at least or up to about 50 kb, at least or up to about 60 kb, at least or up to about 70 kb, at least or up to about 80 kb, at least or up to about 90 kb, at least or up to about 100 kb, at least or up to about 200 kb, at least or up to about 300 kb, at least or up to about 400 kb, or at least or up to about 500 kb.

[0244] In some embodiments of any one of the methods disclosed herein, the target genomic region (e.g., target genomic locus) of the one or more target genomic regions is at most about 200 nucleic acid bases, at most about 300 nucleic acid bases, 400 nucleic acid bases, at most about 500 nucleic acid bases, at most about 600 nucleic acid bases, at most about 700 nucleic acid bases, at most about 800 nucleic acid bases, at most about 900 nucleic acid bases, at most about 1 kb, at most about 2 kb, at most about 3 kb, at most about 4 kb, at most about 5 kb, at most about 6 kb, at most about 7 kb, at most about 8 kb, at most about 9 kb, at most about 10 kb, at most about 11 kb, at most about 12 kb, at most about 13 kb, at most about 14 kb, at most about 15 kb, at most about 16 kb, at most about 17 kb, at most about 18 kb, at most about 19 kb, at most about 20 kb, at most about 21 kb, at most about 22 kb, at most about 23 kb, at most about 24 kb, at most about 25 kb, at most about 26 kb, at most about 27 kb, at most about 28 kb, at most about 29 kb, at most about 30 kb, at most about 31 kb, at most about 32 kb, at most about 33 kb, at most about 34 kb, at most about 35 kb, at most about 36 kb, at most about 37 kb, at most about 38 kb, at most about 39 kb, at most about 40 kb, at most about 41 kb, at most about 42 kb, at most about 43 k In some embodiments, the fragment length may be up to about 5 kb, up to about 6 kb, up to about 7 kb, up to about 8 kb, up to about 9 kb, up to about 10 kb, up to about 11 kb, up to about 12 kb, up to about 13 kb, up to about 14 kb, up to about 15 kb, up to about 16 kb, up to about 17 kb, up to about 18 kb, up to about 19 kb, up to about 20 kb, up to about 25 kb, up to about 30 kb, up to about 35 kb, up to about 40 kb, up to about 45 kb, up to about 50 kb, or up to about 100 kb.

[0245] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes can include at least or up to about 10, at least or up to about 20, at least or up to about 30, at least or up to about 40, at least or up to about 50, at least or up to about 60, at least or up to about 70, at least or up to about 80, at least or up to about 90, at least or up to about 100, at least or up to about 200, at least or up to about 300, at least or up to about 400, at least or up to about 500, at least or up to about 600, at least or up to about 700, at least or up to about 800, at least or up to about 900, at least or up to about 1,000, at least or up to about 2,000, at least or up to about 3,000, at least or up to about 4,000, or at least or up to about 5,000 different nucleic acid probes designed to hybridize to different target nucleic acid sequences.

[0246] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes can have a length of at least or up to about 50, at least or up to about 55, at least or up to about 60, at least or up to about 65, at least or up to about 70, at least or up to about 75, at least or up to about 80, at least or up to about 85, at least or up to about 90, at least or up to about 95, or at least or up to about 100 nucleotides.

[0247] In one aspect, the present disclosure provides a composition comprising a bait set, which comprises any one of the nucleic acid probe sets disclosed herein.The composition comprising such a bait set can be used in any of the methods disclosed herein.In some cases, the nucleic acid probe set can be designed to pull down (or capture) cfDNA molecules.In some cases, the nucleic acid probe set can be designed to pull down (or capture) cfRNA molecules.

[0248] In some embodiments, a bait set can include a set of nucleic acid probes designed to pull down cell-free nucleic acid (e.g., cfDNA) molecules from the genomic regions specified in Table 1. The set of nucleic acid probes can include at least or up to about 1%, at least or up to about 2%, at least or up to about 3%, at least or up to about 4%, at least or up to about 5%, at least or up to about 6%, at least or up to about 7%, at least or up to about 8%, at least or up to about 9%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about Can be designed to pull down the cell-free nucleic acid molecules from 35%, at least or up to about 40%, at least or up to about 45%, at least or up to about 50%, at least or up to about 55%, at least or up to about 60%, at least or up to about 65%, at least or up to about 70%, at least or up to about 75%, at least or up to about 80%, at least or up to about 85%, at least or up to about 90%, at least or up to about 95%, at least or up to about 99% or about 100%.In some cases, nucleic acid probe set can be designed to pull down cfDNA molecules.In some cases, nucleic acid probe set can be designed to pull down cfRNA molecules.

[0249] In some embodiments of any one of the compositions disclosed herein, each nucleic acid probe (or each nucleic acid probe) of the set of nucleic acid probes can include a pull-down tag. The pull-down tag can be used to enrich a sample (e.g., a sample containing multiple nucleic acid molecules obtained from or derived from a subject) for a specific subset (e.g., in the case of cell-free nucleic acid molecules containing multiple graded variants as disclosed herein).

[0250] In some cases, the pull-down tag can include a nucleic acid barcode (e.g., on one or both sides of the nucleic acid probe).By utilizing a bead or substrate containing a nucleic acid sequence complementary to the nucleic acid barcode, the nucleic acid barcode can be used to pull down and enrich any nucleic acid probe that hybridizes to the target cell-free nucleic acid molecule.Alternatively, or in addition, the nucleic acid barcode can be used to identify the target cell-free nucleic acid molecule from any sequencing data (e.g., amplification sequencing) obtained by using any of the sets of nucleic acid probes disclosed herein.

[0251] In some cases, the pull-down tag can include an affinity target moiety that can be specifically recognized and bound by an affinity binding moiety. The affinity binding moiety can specifically bind to the affinity target moiety to form an affinity pair. In some cases, by utilizing beads or substrates containing the affinity binding moiety, the affinity target moiety can be used to pull down and enrich any nucleic acid probe that hybridizes to the target cell-free nucleic acid molecule. Alternatively, the pull-down tag can include the affinity binding moiety, while the bead / substrate can include the affinity target moiety. Non-limiting examples of affinity pairs can include biotin / avidin, antibody / antigen, biotin / streptavidin, metal / chelator, ligand / receptor, nucleic acid and binding protein, and complementary nucleic acid. In one example, the pull-down tag can include biotin.

[0252] In some embodiments of any one of the compositions disclosed herein, the length of the target cell-free nucleic acid (e.g., cfDNA) molecule pulled down by any subject nucleic acid probe can be about 100 to about 200 nucleotides. The length of the target cell-free nucleic acid molecule can be at least about 100 nucleotides. The length of the target cell-free nucleic acid molecule can be up to about 200 nucleotides. The length of the target cell-free nucleic acid molecule may be from about 100 nucleotides to about 110 nucleotides, from about 100 nucleotides to about 120 nucleotides, from about 100 nucleotides to about 130 nucleotides, from about 100 nucleotides to about 140 nucleotides, from about 100 nucleotides to about 150 nucleotides, from about 100 nucleotides to about 160 nucleotides, from about 100 nucleotides to about 170 nucleotides, from about 100 nucleotides to about 180 nucleotides, from about 100 nucleotides to about 190 nucleotides, from about 100 nucleotides to about 200 nucleotides, from about 110 nucleotides to about 120 nucleotides, from about 110 nucleotides to about 130 nucleotides, from about 110 nucleotides to about 140 nucleotides, from about 110 nucleotides to about 150 nucleotides, from about 110 nucleotides to about 160 nucleotides, from about 110 nucleotides to about 170 nucleotides, from about 110 nucleotides to about 180 nucleotides, from about 110 nucleotides to about 190 nucleotides nucleotides, about 110 nucleotides to about 200 nucleotides, about 120 nucleotides to about 130 nucleotides, about 120 nucleotides to about 140 nucleotides, about 120 nucleotides to about 150 nucleotides, about 120 nucleotides to about 160 nucleotides, about 120 nucleotides to about 170 nucleotides, about 120 nucleotides to about 180 nucleotides, about 120 nucleotides to about 190 nucleotides, about 120 nucleotides to about 200 nucleotides, about 130 nucleotides to about 140 nucleotides, about 130 nucleotides to about 150 nucleotides, about 130 nucleotides to about 160 nucleotides, about 130 nucleotides to about 170 nucleotides, about 130 nucleotides to about 180 nucleotides, about 130 nucleotides to about 190 nucleotides, about 130 nucleotides to about 200 nucleotides, about 140 nucleotides to about 150 nucleotides, about 140 nucleotides to about 160 nucleotides,It can be about 140 nucleotides to about 170 nucleotides, about 140 nucleotides to about 180 nucleotides, about 140 nucleotides to about 190 nucleotides, about 140 nucleotides to about 200 nucleotides, about 150 nucleotides to about 160 nucleotides, about 150 nucleotides to about 170 nucleotides, about 150 nucleotides to about 180 nucleotides, about 150 nucleotides to about 190 nucleotides, about 150 nucleotides to about 200 nucleotides, about 160 nucleotides to about 170 nucleotides, about 160 nucleotides to about 180 nucleotides, about 160 nucleotides to about 190 nucleotides, about 160 nucleotides to about 200 nucleotides, about 170 nucleotides to about 180 nucleotides, about 170 nucleotides to about 190 nucleotides, about 170 nucleotides to about 200 nucleotides, about 180 nucleotides to about 190 nucleotides, about 180 nucleotides to about 200 nucleotides, or about 190 nucleotides to about 200 nucleotides. The length of the target cell-free nucleic acid molecule can be about 100 nucleotides, about 110 nucleotides, about 120 nucleotides, about 130 nucleotides, about 140 nucleotides, about 150 nucleotides, about 160 nucleotides, about 170 nucleotides, about 180 nucleotides, about 190 nucleotides, or about 200 nucleotides. In some examples, the length of the target cell-free nucleic acid molecule can range from about 100 nucleotides to about 180 nucleotides.

[0253] In some embodiments of any one of the compositions disclosed herein, the genomic region can be associated with a condition.If a subject has the condition, the genomic region can be determined to show abnormal somatic hypermutation.For example, the condition can include B-cell lymphoma or its subtypes, such as diffuse large B-cell lymphoma, follicular lymphoma, Burkitt's lymphoma and B-cell chronic lymphocytic leukemia.Further details of the condition are provided below.

[0254] In some embodiments of any one of the compositions disclosed herein, the composition further comprises a plurality of cell-free nucleic acid (e.g., cfDNA) molecules obtained or derived from the subject. E. Diagnostic or Therapeutic Uses

[0255] Some embodiments are directed to performing a diagnostic scan on an individual's cell-free nucleic acid and then, based on the results of the scan indicating cancer, performing further clinical procedures and / or treating the individual. According to various embodiments, numerous types of neoplasms can be detected.

[0256] In some embodiments of any one of the methods disclosed herein, the method may include determining that a subject has a condition, or determining the degree or state of the condition, based on one or more cell-free nucleic acid molecules containing multiple stepwise variants. In some cases, the method may further include determining that one or more cell-free nucleic acid molecules (each identified as containing multiple stepwise variants) are derived from a sample associated with a condition (e.g., cancer) based on statistical model analysis (i.e., molecular analysis). For example, the method may include using one or more algorithms (e.g., Monte Carlo simulation) to determine a first probability (e.g., 80%) that a cell-free nucleic acid identified as having multiple stepwise variants is associated with or derived from a first condition, and a second probability (e.g., 20%) that the same cell-free nucleic acid is associated with or derived from a second condition (or derived from a healthy cell). In some cases, the method may include determining the likelihood or probability that the subject has one or more conditions based on an analysis of one or more cell-free nucleic acid molecules each identified as containing multiple graded variants (i.e., a macro- or global analysis). For example, the method may include using one or more algorithms (including, e.g., one or more mathematical models disclosed herein, such as binomial sampling) to analyze multiple cell-free nucleic acid molecules each identified as containing multiple graded variants, thereby determining a first probability (e.g., 80%) that the subject has a first condition and a second probability (e.g., 20%) that the subject has a second condition (or is healthy).

[0257] The statistical model analysis disclosed herein may be an approximate solution using numerical approximations such as a binomial model, a ternary model, a Monte Carlo simulation, a finite difference method, etc. In one example, the statistical model analysis used herein may be a Monte Carlo statistical analysis. In another example, the statistical model analysis used herein may be a binomial model analysis or a ternary model analysis.

[0258] In some embodiments of any one of the methods disclosed herein, the method may include monitoring the progression of the subject's condition based on the identified one or more cell-free nucleic acid molecules, such that each of the identified cell-free nucleic acid molecules comprises multiple stage variants. In some cases, the progression of the condition may be a worsening of the condition (e.g., progressing from stage I cancer to stage III cancer), as described herein. In some cases, the progression of the condition may be at least a partial remission of the condition (e.g., downstaging from stage IV cancer to stage II cancer), as described herein. Alternatively, in some cases, the progression of the condition may remain substantially the same between two different time points, as described herein. In one example, the method may include determining the likelihood or probability of a different progression of the subject's condition. For example, the method may include using one or more algorithms (e.g., including one or more mathematical models disclosed herein, such as binomial sampling) to determine a first probability (e.g., 20%) that the subject's condition is worse than before, a second probability (e.g., 70%) of at least partial remission of the condition, and a third probability (e.g., 10%) that the subject's condition is the same as before.

[0259] In some embodiments of any one of the methods disclosed herein, the method can include performing a different procedure (e.g., a follow-up diagnostic procedure) to confirm the subject's condition, the condition being determined, and / or its progression being monitored, as provided herein. Non-limiting examples of different procedures can include physical examination, medical imaging, genetic testing, mammography, endoscopy, stool sampling, Pap test, alpha-fetoprotein blood test, CA-125 test, prostate-specific antigen (PSA) test, biopsy extraction, bone marrow aspiration, and tumor marker detection tests. Medical imaging includes, but is not limited to, X-ray, magnetic resonance imaging (MRI), computed tomography (CT), ultrasound, and positron emission tomography (PET). Endoscopy includes, but is not limited to, bronchoscopy, colonoscopy, colposcopy, cystoscopy, esophagoscopy, gastroscopy, laparoscopy, neuroendscopy, proctoscopy, and sigmoidoscopy.

[0260] In some embodiments of any one of the methods disclosed herein, the method can include determining a treatment for a subject's condition based on one or more identified cell-free nucleic acid molecules, each of which contains multiple stage variants. In some cases, the treatment can be determined based on (i) the determined subject's condition and / or (ii) the determined progression of the subject's condition. Furthermore, the treatment can be determined based on one or more additional factors such as the subject's gender, nationality, age, ethnicity, and other physical condition. In some examples, the treatment can be determined based on one or more characteristics of the multiple stage variants of the identified cell-free nucleic acid molecules as disclosed herein.

[0261] In some embodiments of any one of the methods disclosed herein, the subject may not have received any treatment for the condition, e.g., the subject may not have been diagnosed with the condition (e.g., lymphoma). In some embodiments of any one of the methods disclosed herein, the subject may have been subjected to treatment for the condition prior to any of the subject methods of the present disclosure. In some cases, the methods disclosed herein can be performed to monitor the progression of the condition for which the subject has been diagnosed, thereby (i) determining the effectiveness of previous treatment, and (ii) assessing whether to maintain treatment, change treatment, or discontinue treatment in favor of a new treatment.

[0262] In some embodiments of any one of the methods disclosed herein, non-limiting examples of treatments (e.g., pre-treatments, new treatments determined based on the methods of the present disclosure, etc.) can include chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, adoptive cell therapy (e.g., chimeric antigen receptor (CAR) T-cell therapy, CAR NK-cell therapy, modified T-cell receptor (TCR) T-cell therapy, etc.), hormone therapy, targeted drug therapy, surgery, transplant, blood transfusion, or medical surveillance.

[0263] In some embodiments of any one of the methods disclosed herein, the condition may include a disease. In some embodiments of any one of the methods disclosed herein, the condition may include a neoplasm, cancer, or tumor. In one example, the condition may include a solid tumor. In another example, the condition may include a lymphoma, such as a B-cell lymphoma (BCL). Non-limiting examples of BCL may include diffuse large B-cell lymphoma (DLBCL), follicular lymphoma (FL), Burkitt's lymphoma (BL), B-cell chronic lymphocytic leukemia (CLL), marginal zone B-cell lymphoma (MZL), and mantle cell lymphoma (MCL).

[0264] As disclosed herein, treating a condition in a subject can include administering one or more therapeutic agents to the subject. The one or more therapeutic agents can be administered to the subject by one or more of the following: orally, intraperitoneally, intravenously, intraarterially, transdermally, intramuscularly, liposomally, via local delivery by catheter or stent, subcutaneously, intraadiposely, and intrathecally.

[0265] Non-limiting examples of therapeutic drugs can include cytotoxic agents, chemotherapeutic agents, growth inhibitory agents, agents used in radiation therapy, anti-angiogenic agents, apoptotic agents, anti-tubulin agents, and other agents for treating cancer, such as anti-CD20 antibodies, anti-PD1 antibodies (e.g., pembrolizumab), platelet-derived growth factor inhibitors (e.g., GLEEVEC™ (imatinib mesylate)), COX-2 inhibitors (e.g., celecoxib), interferons, cytokines, antagonists (e.g., neutralizing antibodies) that bind to one or more of the following targets: PDGFR-β, BlyS, APRIL, BCMA receptor, TRAIL / Apo2, other bioactive agents, and organic chemical agents.

[0266] Non-limiting examples of cytotoxic agents may include radioactive isotopes (e.g., At211, I131, I125, Y90, Re186, Re188, Sm153, Bi212, P32, and radioactive isotopes of Lu), chemotherapeutic agents such as methotrexate, adriamycin, vinca alkaloids (vincristine, vinblastine, etoposide), doxorubicin, melphalan, mitomycin C, chlorambucil, daunorubicin or other intercalating agents, enzymes and fragments thereof, such as nucleases, antibiotics, and toxins, such as small molecule toxins or enzymatically active toxins of bacterial, fungal, plant or animal origin.

[0267] Non-limiting examples of chemotherapeutic agents include alkylating agents such as thiotepa and CYTOXAN® cyclophosphamide, alkyl sulfonates, e.g., busulfan, improsulfan, and piposulfan; aziridines, such as benzodopa, carboquone, meturedopa, and uredopa; altretamine, triethylenemelamine, triethylenephosphoramide, triethylenethiophosphoramide, and trimethylolmelamine. ethylenimines and methylameramines, including acetogenins (especially bullatacin and bullatacinone); delta-9-tetrahydrocannabinol (dronabinol, MARINOL®); beta-lapachone; lapachol; colchicine; betulinic acid; camptothecins (including synthetic analogs topotecan (HYCAMTIN®), CPT-11 (irinotecan, CAMPTOSAR®), acetylcamptothecin, scopolectin, 9-aminocamptothecin); bryostatins amines; kallistatins; CC-1065 (including its synthetic analogs adozelesin, carzelesin, and bizelesin); podophyllotoxins; podophyllic acid; teniposide; cryptophycins (especially cryptophycin 1 and cryptophycin 8); dolastatins; duocarmycins (including synthetic analogs, KW-2189 and CB1-TM1); eleutherobin; pancratistatin; sarcodictyin; spongistatins; chlorambucil, chlornaphazine, cyclophosphamide, estramustine Nitrogen mustards such as benzodiazepine, ifosfamide, mechlorethamine, mechlorethamine oxide hydrochloride, melphalan, nobembine, phenesterine, prednimustine, trofosfamide, and uracil mustard; nitrosoureas such as carmustine, chlorozotocin, fotemustine, lomustine, nimustine, and ranimnustine; antibiotics such as enediyne antibiotics; dynemycin (including dynemycin A); espiramicina;Also included are neocarzinostatin chromophores and related enediyne antibiotic chromophores), aclacinomycin, actinomycin, anthramycin, azaserine, bleomycin, cactinomycin, carabicin, carminomycin, carzinophilin, chromomycin, dactinomycin, daunorubicin, detrevicin, 6-diazo-5-oxo-L-norleucine, ADRIAMYCIN® doxorubicin (morpholino-doxorubicin, cyanomorpholino-doxorubicin, 2-pyrrolino ... and deoxydoxorubicin), mitomycins such as epirubicin, esorubicin, idarubicin, marcelomycin, and mitomycin C, mycophenolic acid, nogalamycin, olivomycin, peplomycin, potfilomycin, puromycin, chelamycin, rodorubicin, streptonigrin, streptozocin, tubercidin, ubenimex, zinostatin, and zorubicin; antimetabolites such as methotrexate and 5-fluorouracil (5-FU); denopterin, methotrexate, pteropterin, and trimetrexate folic acid analogues such as; purine analogues such as fludarabine, 6-mercaptopurine, thiamiprine, thioguanine; pyrimidine analogues such as ancitabine, azacitidine, 6-azauridine, carmofur, cytarabine, dideoxyuridine, doxifluridine, enocitabine, floxuridine; androgens such as calusterone, dromostanolone propionate, epithiostanol, mepitiostane, testolactone; antiadrenergic agents such as aminoglutethimide, mitotane, trilostane; folic acid supplements such as folinic acid; aceglatone; aldofos Famidoglycosides; aminolevulinic acid; eniluracil; amsacrine; bestravcil; bisantrene; edatraxate; defofamine; demecolcine; diaziconazole; eflornithine; elliptinium acetate; epothilone; etoglucide; gallium nitrate; hydroxyurea; lentinan; lonidynin; maytansinoids such as maytansine and ansamitocin; mitoguazone; mitoxantrone; mopidanol; nitraerine; pentostatin; phenamet; pirarubicin; losoxantrone; 2-ethylhydrazide;procarbazine; PSK® polysaccharide complex (JHS Natural Products, Eugene, Oreg.); razoxane; rhizoxin; schizofiran; spirogermanium; tenuazonic acid; triazicon; 2,2',2''-trichlorotriethylamine; trichothecenes (especially T-2 toxin, verrucarin A, roridin A, and anguidine); urethane; vindesine (ELDISINE®, FILDESIN®); dacarbazine; mannomustine; mitobronitol; mitolactol; pipobroman; gacytosine; arabinoside ("Ara-C"); thiotepa; taxoids, such as TAXOL® paclitaxel (Bristol-Myers Squibb Oncology, Princeton, NJ), ABRAXANE™ cremophor-free albumin-engineered nanoparticle formulation of paclitaxel (American Pharmaceutical Partners, Schaumberg, lll.), and TAXOTERE® docetaxel (Rhone-Poulenc taxanes, including those listed in Table 1 (Rorer, Antony, France); chlorambucil; gemcitabine (GEMZAR®); 6-thioguanine; mercaptopurine; methotrexate; platinum analogs such as cisplatin and carboplatin; vinblastine (VELBAN®); platinum; etoposide (VP-16); ifosfamide; mitoxantrone; vincristine (ONCOVIN®); oxaliplatin; leucovorin; vinorelbine (NAVELBINE®); novantrone; edatrexate; daunomycin; aminopterin; ibandronate; the topoisomerase inhibitor RFS2000; difluoromethylornithine (DMFO); retinoids, such as retinoic acid; capecitabine (XELODA®); pharmaceutically acceptable salts, acids, or derivatives of any of the above;and combinations of two or more of the above, such as CHOP, an abbreviation for the combination therapy of cyclophosphamide, doxorubicin, vincristine, and prednisolone, and FOLFOX, an abbreviation for the treatment regimen using oxaliplatin (ELOXATIN®) in combination with 5-FU and leucovorin;

[0268] Examples of chemotherapeutic agents may also include "antihormonal agents" or "endocrine therapeutic agents" that act to regulate, reduce, block, or inhibit the effects of hormones that can promote cancer growth, often in the form of systemic or systemic treatment. They may be hormones themselves. Examples include antiestrogens and selective estrogen receptor modulators (SERMs), such as tamoxifen (including NOLVADEX® tamoxifen), EVISTA® raloxifene, droloxifene, 4-hydroxytamoxifen, trioxifene, ketoxifene, LY117018, onapristone, and FARESTON® toremifene; antiprogesterones; estrogen receptor downregulators (ERDs); agents that function to suppress or shut down the ovaries, such as LUPRON® and ELIGARD® leuprolide acetate, goserelin acetate, and acetaminophen. Luteinizing hormone-releasing hormone (LHRH) agonists, such as buserelin acid and triptorelin; other antiandrogens, such as flutamide, nilutamide, and bicalutamide; and aromatase inhibitors, which inhibit the enzyme aromatase, which regulates estrogen production in the adrenal glands, such as, for example, 4(5)-imidazole, aminoglutethimide, MEGASE® megestrol acetate, AROMASIN® exemestane, formestanie, fadrozole, RIVISOR® vorozole, FEMARA® letrozole, and ARIMIDEX® anastrozole.Further included in such definition of chemotherapeutic agent are bisphosphonates, such as clodronate (e.g., BONEFOS® or OSTAC®), DIDROCAL® etidronate, NE-58095, ZOMETA® zoledronic acid / zoledronate, FOSAMAX® alendronate, AREDIA® pamidronate, SKELID® tiludronate, or ACTONEL® risedronate, as well as troxacitabine (a 1,3-dioxolane nucleoside cytosine analog); antisense oligonucleotides, particularly those inhibiting, for example, PKC-alpha, Raf, H-Ras, and epithelial growth factor receptors. those that inhibit the expression of genes in signaling pathways involved in aberrant cell growth, such as the EGFR and EGFR-associated protein receptor (EGFR); vaccines such as the THERATOPE® vaccine and gene therapy vaccines, for example, the ALLOVECTIN® vaccine, the LEUVECTIN® vaccine, and the VAXID® vaccine; the LURTOTECAN® topoisomerase 1 inhibitor; ABARELIX® rmRH; lapatinib ditosylate (also known as the ErbB-2 and EGFR dual tyrosine kinase small molecule inhibitor, GW572016); and pharmaceutically acceptable salts, acids, or derivatives of any of the above.

[0269] Examples of chemotherapeutic agents may also include antibodies such as alemtuzumab (Campath), bevacizumab (AVASTIN®, Genentech), cetuximab (ERBITUX®, Imclone); panitumumab (VECTIBIX®, Amgen), rituximab (RITUXAN®, Genentech / Biogen Idec), pertuzumab (OMNITARG®, 2C4, Genentech), trastuzumab (HERCEPTIN®, Genentech), tositumomab (Bexxar, Corixia), and the antibody-drug conjugate gemtuzumab ozogamicin (MYLOTARG®, Wyeth). Additional humanized monoclonal antibodies with therapeutic potential as agents in combination with the compounds of the invention include apolizumab, aselizumab, atlizumab, bapineuzumab, bivatuzumab mertansine, cantuzumab mertansine, cedelizumab, certolizumab pegol, cidfusituzumab, cidtuzumab, daclizumab, eculizumab, efalizumab, epratuzumab, erlizumab, femuzumab, fontolizumab, gemtuzumab ozogamicin, inotuzumab ozogamicin, ipilimumab, labetuzumab, lintuzumab, matuzumab, mepolizumab, motavizumab, nat ... tuzumab, nimotuzumab, norobizumab, numavizumab, ocrelizumab, ocrelizumab, palivizumab, pascolizumab, pecfusituzumab, pectuzumab, pexelizumab, ralivizumab, ranibizumab, reslivizumab, reslizumab, resivizumab, rovelizumab, ruplizumab, sibrotuzumab, siplizumab, sontuzumab, tacatuzumab tetraxetan, tadocizumab, talizumab, tefibazumab, tocilizumab, toralizumab, tucotuzumab celmoleukin, tucitutuzumab, umavizumab, urtoxazumab, ustekinumab, visilizumab, and interleukin-12 and anti-interleukin-12 (ABT-874 / J695, Wyeth Research and Abbott Laboratories), a recombinant exclusively human sequence full-length IgG 1λ antibody genetically engineered to recognize the p40 protein.

[0270] Examples of chemotherapeutic agents include "tyrosine kinase inhibitors" such as EGFR-targeted agents (e.g., small molecules, antibodies, etc.); small molecule HER2 tyrosine kinase inhibitors such as TAK165 available from Takeda; CP-724, 714 (Pfizer and OSI), which are oral selective inhibitors of ErbB2 receptor tyrosine kinase; dual HER inhibitors such as EKB-569 (available from Wyeth), which preferentially binds to EGFR but inhibits both HER2 and EGFR-overexpressing cells; lapatinib (GSK572016; available from Glaxo-SmithKline), an oral HER2 and EGFR tyrosine kinase inhibitor; PKI-166 (available from Novartis); pan-HER inhibitors such as canertinib (CI-1033; Pharmacia); Raf-1 inhibitors such as the antisense agent ISIS-5132 available from ISIS Pharmaceuticals, which inhibits Raf-1 signaling; imatinib mesylate (Glaxo non-HER-targeted TK inhibitors such as GLEEVEC® available from SmithKline; multi-targeted tyrosine kinase inhibitors such as sunitinib (SUTENT®, available from Pfizer); VEGF receptor tyrosine kinase inhibitors such as vatalanib (PTK787 / ZK222584 available from Novartis / Schering AG); MAPK extracellular regulated kinase I inhibitor CI-1040 (available from Pharmacia); quinazolines, e.g., PD 153035, 4-(3-chloroanilino)quinazoline; pyridopyrimidines; pyrimidopyrimidines; CGP 59326, CGP 60261, and CGP Pyrrolopyrimidines such as 62706; pyrazolopyrimidine, 4-(phenylamino)-7H-pyrrolo[2,3-d]pyrimidine; curcumin (diferuloylmethane, 4,5-bis(4-fluoroanilino)phthalimide); tyrphostins containing a nitrothiophene moiety; PD-0183805 (Warner-Lambert); antisense molecules (e.g., those that bind to HER-encoding nucleic acids); quinoxalines (U.S. Patent No. 5,804,396); tryphostin (U.S. Patent No. 5,804,396); ZD6474 (AstraZeneca); PTK-787 (Novartis / Schering AG);Pan-HER inhibitors such as CI-1033 (Pfizer); Affinitac (ISIS 3521; Isis / Lilly); imatinib mesylate (GLEEVEC®); PKI 166 (Novartis); GW2016 (GlaxoSmithKline); CI-1033 (Pfizer); EKB-569 (Wyeth); semaxinib (Pfizer); ZD6474 (AstraZeneca); PTK-787 (Novartis / Schering AG); INC-1C11 (Imclone); and rapamycin (sirolimus, RAPAMUNE®).

[0271] Examples of chemotherapeutic agents include dexamethasone, interferon, colchicine, metoprine, cyclosporine, amphotericin, metronidazole, alemtuzumab, alitretinoin, allopurinol, amifostine, arsenic trioxide, asparaginase, live BCG, bevacuzimab, bexarotene, cladribine, clofarabine, darbepoetin alfa, denileukin, dexrazoxane, epoetin alfa, erotinib, filgrastim, histrelin acetate, ibritumomab, interferon alfa-2a, interferon alfa- 2b, lenalidomide, levamisole, mesna, methoxsalen, nandrolone, nelarabine, nofetumomab, oprelvekin, palifermin, pamidronate, pegademase, pegaspargase, pegfilgrastim, pemetrexed disodium, plicamycin, porfimer sodium, quinacrine, rasburicase, sargramostim, temozolomide, VM-26, 6-TG, toremifene, tretinoin, ATRA, valrubicin, zoledronate and zoledronic acid, and pharmaceutically acceptable salts thereof.

[0272] Examples of chemotherapeutic agents include hydrocortisone, hydrocortisone acetate, cortisone acetate, tixocortol pivalate, triamcinolone acetonide, triamcinolone alcohol, mometasone, amcinonide, budesonide, desonide, fluocinonide, fluocinolone acetonide, betamethasone, betamethasone sodium phosphate, dexamethasone, dexamethasone sodium phosphate, fluocortolone, and hydrocortisone-17-butyrate. , hydrocortisone-17-valerate, aclomethasone dipropionate, betamethasone valerate, betamethasone dipropionate, prednicarb, clobetasone-17-butyrate, clobetasol-17-propionate, fluocortolone caproate, fluocortolone pivalate, and fluprednidene acetate: phenylalanine-glutamine-glycine (FEG) and its D-isomer (feG) (IMULAN) Immunoselective anti-inflammatory peptides (ImSAIDs) such as those from BioTherapeutics, LLC; antirheumatic drugs such as azathioprine, cyclosporine (cyclosporine A), D-penicillamine, gold salts, hydroxychloroquine, leflunomide minocycline, and sulfasalazine; etanercept (ENBREL®), infliximab (REMICADE®), adalimumab (HUMIRA®), and certolizumab pegol (CIMZI®); A®), tumor necrosis factor alpha (TNFα) blockers such as golimumab (SIMPONI®), interleukin 1 (IL-1) blockers such as anakinra (KINERET®), T-cell costimulation blockers such as abatacept (ORENCIA®), interleukin 6 (IL-6) blockers such as tocilizumab (ACTEMERA®); interleukin 13 (IL-13) blockers such as lebrikizumab; interferon alpha (IFN) blockers such as rontalizumab; beta7 integrin blockers such as rhuMAb Beta7; IgE pathway blockers such as anti-M1 prime; secreted homotrimeric LTa3 and membrane-bound heterotrimeric LTa / β2 blockers, e.g., anti-lymphotoxin alpha (LTa);Various investigational drugs, such as thioplatin, PS-341, phenylbutyrate, ET-18-OCH3, or famexyltransferase inhibitors (L-739749, L-744832); polyphenols such as quercetin, resveratrol, piceatannol, epigallocatechin gallate, theaflavin, flavanols, procyanidins, betulinic acid and its derivatives; autophagy inhibitors such as chloroquine; delta-9-tetrahydrocannabinol (dronabinol, MARINOL®); beta-lapachone; lapachol; colchicine; betulinic acid; acetylcamptothecin, scopolectin, 9-aminocamptothecin; podophyllotoxin; tegafur (UFTORAL®); bexarotene (TARGRETIN®); clodronate (e.g., BONEFOS® or OSTA C®), etidronate (DIDROCAL®), NE-58095, zoledronic acid / zoledronate (ZOMETA®), alendronate (FOSAMAX®), pamidronate (AREDIA®), tiludronate (SKELID®), or risedronate (ACTONEL®); and epidermal growth factor receptor (EGF-R); vaccines, such as the THERATOPE® vaccine; perifosine, COX-2 inhibitors (e.g., celecoxib or etoricoxib), proteosome inhibitors (e.g., PS341); CCI-779; tipifamib (R11577); orafenib, ABT510; Bcl-2 inhibitors such as oblimersen sodium (GENASENSE®); pixantrone; lonafamib (SCH 6636, SARASAR™); and pharmaceutically acceptable salts, acids, or derivatives of any of the above; and combinations of two or more of the above.

[0273] According to many embodiments, once a diagnosis of cancer is made, several treatments can be performed, including, but not limited to, surgery, resection, chemotherapy, radiation therapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, and blood transfusion. In some embodiments, anti-cancer and / or chemotherapeutic agents are administered, including, but not limited to, alkylating agents, platinum agents, taxanes, vinca agents, antiestrogens, aromatase inhibitors, ovarian suppressants, endocrine / hormonal agents, bisphosphonates, and targeted biological therapeutic agents. Agents include cyclophosphamide, fluorouracil (or 5-fluorouracil or 5-FU), methotrexate, thiotepa, carboplatin, cisplatin, taxanes, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifene, toremifene, fulvestrant, gemcitabine, irinotecan, ixabepilone, temozolomide, topotecan, bisphosphonates, and bisphosphonates. ancristine, vinblastine, eribulin, mutamycin, capecitabine, anastrozole, exemestane, letrozole, leuprolide, abarelix, buserelin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, zoledronate, Tykerb, daunorubicin, doxorubicin, epirubicin, idarubicin, valrubicinMitoxantrone, bevacizumab, cetuximab, ipilimumab, ado-trastuzumab emtansine, afatinib, aldesleukin, alectinib, alemtuzumab, atezolizumab, avelumab, axtinib, belimumab, belinostat, bevacizumab, blinatumomab, bortezomib, bosutinib, brentuximab vedotin, brigatinib, cabozantinib , canakinumab, carfilzomib, certinib, cetuximab, cobimetinib, crizotinib, dabrafenib, daratumumab, dasatinib, denosumab, dinutuximab, durvalumab, elotuzumab, enasidenib, erlotinib, everolimus, gefitinib, ibritumomab tiuxetan, ibrutinib, idelalisib, imatinib, ipilimumab, ixazo Mibu, lapatinib, lenvatinib, midostaurin, necitumumab, neratinib, nilotinib, niraparib, nivolumab, obinutuzumab, ofatumumab, olaparib, olaratumab, osimertinib, palbociclib, panitumumab, panobinostat, pembrolizumab, pertuzumab, ponatinib, ramucirumab, regorafenib, ribociclib, rituximab, These include, but are not limited to, midepsin, rucaparib, ruxolitinib, siltuximab, sipuleucel-T, sonidegib, sorafenib, temsirolimus, tocilizumab, tofacitinib, tositumomab, trametinib, trastuzumab, vandetanib, vemurafenib, venetoclax, vismodegib, vorinostat, and dib-aflibercept. According to various embodiments, individuals can be treated with a single drug or a combination of drugs as described herein. A common treatment combination is cyclophosphamide, methotrexate, and 5-fluorouracil (CMF).

[0274] In some embodiments of any one of the methods disclosed herein, any cell-free nucleic acid molecule (for example, cfDNA, cfRNA) can be derived from cells.For example, cell sample or tissue sample can be obtained from subject and processed to remove all cells from sample, thereby producing cell-free nucleic acid molecule derived from sample.

[0275] In some embodiments of any one of the methods disclosed herein, the reference genome sequence may be derived from cells of an individual. The individual may be a healthy control or a subject being subjected to a method disclosed herein to determine or monitor the progression of a condition.

[0276] The cell may be a healthy cell. Alternatively, the cell may be a diseased cell. The diseased cell may have altered metabolism, gene expression, and / or morphological characteristics. The diseased cell may be a cancer cell, a diabetic cell, or an apoptotic cell. The diseased cell may be a cell derived from a diseased subject. Exemplary diseases may include blood disorders, cancer, metabolic disorders, eye disorders, organ disorders, musculoskeletal disorders, heart diseases, etc.

[0277] The cell may be a mammalian cell or may be derived from a mammalian cell. The cell may be a rodent cell or may be derived from a rodent cell. The cell may be a human cell or may be derived from a human cell. The cell may be a prokaryotic cell or may be derived from a prokaryotic cell. The cell may be a bacterial cell or may be derived from a bacterial cell. The cell may be an archaeal cell or may be derived from an archaeal cell. The cell may be a eukaryotic cell or may be derived from a eukaryotic cell. The cell may be a pluripotent stem cell. The cell may be a plant cell or may be derived from a plant cell. The cell may be an animal cell or may be derived from an animal cell. The cell may be an invertebrate cell or may be derived from an invertebrate cell. The cell may be a vertebrate cell or may be derived from a vertebrate cell. The cell may be a microbial cell or may be derived from a microbial cell. The cell may be a fungal cell or may be derived from a fungal cell. The cells may be derived from a particular organ or tissue.

[0278] Non-limiting examples of cell(s) include lymphoid cells, such as B cells, T cells (cytotoxic T cells, natural killer T cells, regulatory T cells, T helper cells), natural killer cells, cytokine-induced killer (CIK) cells; myeloid cells, such as granulocytes (basophilic granulocytes, eosinophilic granulocytes, neutrophilic granulocytes / hypersegmented neutrophils), monocytes / macrophages, erythrocytes (reticulocytes), mast cells, platelets / megakaryocytes, dendritic cells; cells from the endocrine system, including thyroid (thyroid epithelial cells, parafollicular cells), parathyroid (parathyroid chief cells, eosinophilic cells), adrenal gland (chromaffin cells), and pineal gland (pineocyte) cells; glial cells Cells of the nervous system, including astrocytes, microglia, giant cell neurosecretory cells, stellate cells, Boettcher cells, and pituitary gland (gonadotropes, corticotropes, thyrotropes, somatotropes, and lactotropes); cells of the respiratory system, including alveolar cells (type I pneumocytes, type II pneumocytes), Clara cells, goblet cells, and dust cells; cells of the circulatory system, including cardiac myocytes and pericytes; cells of the digestive system, including stomach (chief cells, parietal cells), goblet cells, Paneth cells, G cells, D cells, ECL cells, I cells, K cells, and S cells; and enterochromaffin cells (enterochromaffin Enteroendocrine cells, including APUD cells, liver (hepatocytes, Kupffer cells), cartilage / bone / muscle; bone cells, including osteoblasts, osteocytes, osteoclasts, and teeth (cementoblasts, ameloblasts); chondrocytes, including chondrocytes and chondrocytes; skin cells, including trichocytes, keratinocytes, and melanocytes (nevus cells); muscle cells, including myocytes; urinary system cells, including podocytes, juxtaglomerular cells, intraglomerular / extraglomerular mesangial cells, renal proximal tubule brush border cells, and macula densa cells; and spermatids. , Sertoli cells, Leydig cells, germ line cells including oocytes; adipocytes, fibroblasts, tendon cells, epidermal keratinocytes (differentiated epidermal cells), epidermal basal cells (stem cells), fingernail and toenail keratinocytes, nail bed basal cells (stem cells), medullary hair stem cells, cortical hair stem cells, cuticle hair stem cells, cuticle root sheath cells, root sheath cells of Huxley's layer, root sheath cells of Henle's layer, outer root sheath cells, hair matrix cells (stem cells), wet stratified barrier epithelial cells, surface epithelial cells of the stratified squamous epithelium of the cornea, tongue, oral cavity, esophagus, anal canal, distal urethra and vagina, cornea,These may include basal cells (stem cells) of the epithelium of the tongue, oral cavity, esophagus, anal canal, distal urethra and vagina, urinary epithelial cells (lining the bladder and ureters), exocrine secretory epithelial cells, salivary gland mucous cells (secreting polysaccharides), salivary gland serous cells (secreting glycoprotein enzymes), von Ebner's gland cells of the tongue (cleansing the taste buds), mammary gland cells (secreting milk), lacrimal gland cells (secreting tears), ceruminous gland cells in the ear (secreting wax), eccrine sweat gland dark cells (secreting glycoproteins), eccrine sweat gland clear cells (secreting small molecules). Apocrine sweat gland cells (secreting odor, sensitive to sex hormones), Molar glands of the eyelids (specialized sweat glands), Sebaceous gland cells (secreting lipid-rich sebum), Bowman's gland cells of the nose (washing the olfactory epithelium), Brunner's gland cells of the duodenum (enzymes and alkaline mucus), seminal vesicle cells (secreting seminal fluid components including fructose for sperm swimming), prostate cells (secreting seminal fluid components), bulbourethral gland cells (secreting mucus), Bartholin's gland cells (secreting vaginal lubrication), Littré gland cells ( Mucus-secreting), endometrial cells (carbohydrate-secreting), isolated goblet cells of the respiratory and digestive tract (mucus-secreting), stomach lining mucosal cells (mucus-secreting), gastric zymogen cells (pepsinogen-secreting), gastric acid-secreting cells (hydrochloric acid-secreting), pancreatic acinar cells (bicarbonate and digestive enzyme secreting), Paneth cells of the small intestine (lysozyme-secreting), type II alveolar cells of the lung (surfactant-secreting), Clara cells of the lung, hormone-secreting cells, anterior pituitary cells, growth hormone-producing cells, lactogenic hormone-secreting cells monocytes, thyrotropes, gonadotropes, corticotropes, intermediate pituitary cells, magnocellular neurosecretory cells, intestinal and airway cells, thyroid cells, thyroid epithelial cells, parafollicular cells, parathyroid cells, parathyroid chief cells, eosinophilic cells, adrenal cells, chromaffin cells, Leydig cells of the testis, theca cells of the follicle, lutein cells of the ruptured follicle, granulosa lutein cells, theca cells of the follicle, juxtaglomerular cells (renin secretion), kidney Macula densa cells, metabolic and storage cells, barrier function cells (lung, intestine, exocrine glands and urogenital tract), kidney, type I alveolar cells (lining the air spaces of the lungs), pancreatic duct cells (central acinar cells), non-striated duct cells (such as those of sweat glands, salivary glands and mammary glands), ductal cells (such as those of seminal vesicles and prostate), epithelial cells lining closed body cavities, ciliated cells with propulsive functions, extracellular matrix secreting cells, contractile cells; skeletal muscle cells, stem cells, cardiac muscle cells, blood and immune system cells, erythrocytes (red blood cells),Megakaryocytes (platelet precursor cells), monocytes, connective tissue macrophages (various types), epidermal Langerhans cells, osteoclasts (in bone), dendritic cells (in lymphoid tissue), microglial cells (central nervous system), neutrophil granulocytes, eosinophil granulocytes, basophil granulocytes, mast cells, helper T cells, suppressor T cells, cytotoxic T cells, natural killer T cells, B cells, natural killer cells, reticulocytes, stem cells and committed progenitor cells (various types) for the blood and immune systems, pluripotent stem cells Examples of such cells include totipotent stem cells, induced pluripotent stem cells, adult stem cells, sensory transducer cells, autonomic neuron cells, sensory organ and peripheral neuron supporting cells, central nervous system neurons and glial cells, lens cells, pigment cells, melanocytes, retinal pigment epithelial cells, germ cells, oogonia / oocytes, spermatids, spermatocytes, spermatogonia (stem cells for sperm cells), sperm, nurse cells, ovarian follicle cells, Sertoli cells (testes), thymic epithelial cells, interstitial cells, and interstitial kidney cells.

[0279] In some embodiments of any one of the methods disclosed herein, the condition can be a cancer or tumor. Non-limiting examples of such conditions include acanthoma, acinar cell carcinoma, acoustic neuroma, acral lentiginous melanoma, acral hidradenoma, acute eosinophilic leukemia, acute lymphoblastic leukemia, acute megakaryoblastic leukemia, acute monocytic leukemia, acute myeloblastic leukemia with maturation, acute myeloid dendritic cell leukemia, acute myeloid leukemia, acute promyelocytic leukemia, adamantinoma, adenocarcinoma, adenoid cystic carcinoma, adenoma, adenomatous odontogenic tumor, adrenocortical carcinoma, adult T-cell leukemia, aggressive NK-cell leukemia, AIDS-related cancer, AIDS-related Lymphoma, alveolar soft part sarcoma, myeloblastic fibroma, anal cancer, anaplastic large cell lymphoma, anaplastic thyroid cancer, angioimmunoblastic T-cell lymphoma, angiomyolipoma, angiosarcoma, appendix cancer, astrocytoma, atypical teratoid rhabdoid tumor, basal cell carcinoma, basal-like carcinoma, B-cell leukemia, B-cell lymphoma, Bellini duct carcinoma, biliary tract cancer, bladder cancer, blastoma, bone cancer, bone tumor, brainstem glioma, brain tumor, breast cancer, Brenner tumor, bronchial tumor, bronchoalveolar carcinoma, brown tumor, Burkitt lymphoma, cancer of unknown primary site, carcinoma Id tumor, carcinoma, carcinoma in situ, penile cancer, carcinoma of unknown primary site, carcinosarcoma, Castleman's disease, central nervous system embryonal tumor, cerebellar astrocytoma, cerebral astrocytoma, cervical cancer, bile duct carcinoma, chondroma, chondrosarcoma, chordoma, choriocarcinoma, choroid plexus papilloma, chronic lymphocytic leukemia, chronic monocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative disorder, chronic neutrophilic leukemia, clear cell tumor, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, Degos disease, dermatofibrosarcoma protuberans, dermoid cyst, desmoplastic small round cell tumor, diffuse large cell carcinoma Alveolar B-cell lymphoma, dysembryoplastic neuroepithelial tumor, embryonal carcinoma, endodermal sinus tumor, endometrial cancer, endometrial uterine cancer, endometrioid tumor, enteropathy-associated T-cell lymphoma, ependymoblastoma, ependymoma, epithelioid sarcoma, erythroleukemia, esophageal cancer, esthesioneuroblastoma, Ewing family tumors, Ewing family sarcoma, Ewing sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, extramammary Paget's disease, fallopian tube cancer, inclusion fetus, fibroma, fibrosarcoma, follicular lymphoma, follicular thyroid cancer, gallbladder cancer, ganglioglioma, ganglioneuroma, gastric cancer, gastric lymphoma, gastrointestinal cancer,Gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, gastrointestinal stromal tumor, germ cell tumor, germ cell tumor, gestational choriocarcinoma, gestational trophoblastic tumor, giant cell tumor of bone, glioblastoma multiforme, glioma, gliomatosis cerebri, glomus tumor, glucagonoma, gonadoblastoma, granulosa cell tumor, hairy cell leukemia, hairy cell leukemia, head and neck cancer, cardiac cancer, hemangioblastoma, hemangiopericytoma, angiosarcoma, hematologic malignancies, hepatocellular carcinoma, hepatosplenic T-cell lymphoma, hereditary breast and ovarian cancer syndrome, Hodgkin's lymphoma, Hodgkin's lymphoma, hypopharyngeal cancer, hypothalamic glioma, inflammatory breast cancer, intraocular cataract Chromoma, Islet cell carcinoma, Islet cell tumor, Juvenile myelomonocytic leukemia, Kaposi's sarcoma, Kaposi's sarcoma, Kidney cancer, Kratskin tumor, Krukenberg's tumor, Laryngeal cancer, Laryngeal carcinoma, Lung cancer, Lung cancer, Lip and oral cavity cancer, Liposarcoma, Lung cancer, Luteoma, Lymphangioma, Lymphangiosarcoma, Lymphoepithelioma, Lymphocytic leukemia, Lymphoma, Macroglobulinemia, Malignant fibrous histiocytoma, Malignant fibrous histiocytoma of bone, Malignant glioma, Malignant mesothelioma, Malignant peripheral nerve sheath tumor, Malignant rhabdoid tumor, Malignant Triton tumor, MALT lymphoma, Mantle cell lymphoma lymphoma, mast cell leukemia, mediastinal germ cell tumor, mediastinal tumor, medullary thyroid cancer, medulloblastoma, medulloepithelioma, melanoma, melanoma, meningioma, Merkel cell carcinoma, mesothelioma, mesothelioma, metastatic squamous cell neck cancer of unknown primary, metastatic urothelial carcinoma, mixed Müllerian tumor, monocytic leukemia, oral cancer, mucinous tumor, multiple endocrine neoplasia syndrome, multiple myeloma, mycosis fungoides, mycosis fungoides, myelodysplasia, myelodysplastic syndrome, myeloid leukemia, myeloid sarcoma, myeloproliferative disorder, myxoma, nasal cavity cancer, nasopharyngeal carcinoma, neoplasm, schwannoma, neuroblastoma, neuroblastoma Cystoma, Neurofibroma, Neuroma, Nodular melanoma, Non-Hodgkin's lymphoma, Non-Hodgkin's lymphoma, Non-melanoma skin cancer, Non-small cell lung cancer, Ocular oncology, Oligoastrocytoma, Oligodendroglioma, Oncocytoma, Optic nerve sheath meningioma, Oral cavity cancer, Oral cancer, Oropharyngeal cancer, Osteosarcoma, Osteosarcoma, Ovarian cancer, Ovarian cancer, Ovarian epithelial cancer, Ovarian germ cell tumor, Ovarian low malignant potential tumor, Paget's disease of the breast, Pancoast tumor, Pancreatic cancer, Pancreatic cancer, Papillary thyroid cancer, Papillomatosis, Paraganglioma, Paranasal sinus cancer, Parathyroid cancer, Penile cancer, Perivascular epithelioid cell tumor, Pheochromocytoma,Moderately differentiated pineal parenchymal tumor, pineoblastoma, pituitary cell tumor, pituitary adenoma, pituitary tumor, plasma cell neoplasm, pleuropulmonary blastoma, polygerminoma, precursor T lymphoblastic lymphoma, primary central nervous system lymphoma, primary effusion lymphoma, primary hepatocellular carcinoma, primary liver cancer, primary peritoneal cancer, primitive neuroectodermal tumor, prostate cancer, pseudomyxoma peritonei, rectal cancer, renal cell carcinoma, respiratory cancer involving the NUT gene on chromosome 15, retinoblastoma, rhabdomyoma, rhabdomyosarcoma, Richter's transformation, sacrococcygeal teratoma, salivary gland cancer, sarcoma, schwannomatosis, sebaceous gland carcinoma, secondary neoplasm, seminoma, serous tumor, Sertoli-Leydig cell tumor, sex cord-stromal tumor, Sézary syndrome, signet ring cell carcinoma, skin cancer, small blue round cell tumor tumor), small cell carcinoma, small cell lung cancer, small cell lymphoma, small intestine cancer, soft tissue sarcoma, somatostatinoma, smoke warts, spinal cord tumor, spinal tumor tumor), splenic marginal zone lymphoma, squamous cell carcinoma, gastric cancer, superficial spreading melanoma, supratentorial primitive neuroectodermal tumor, surface epithelial-stromal tumor, synovial sarcoma, T-cell acute lymphoblastic leukemia, T-cell large granular lymphocytic leukemia, T-cell leukemia, T-cell lymphoma, T-cell prolymphocytic leukemia, teratoma, end-stage lymphoid cancer, testicular cancer, theca tumor, throat cancer, thymic carcinoma, thymoma, thyroid cancer, transitional cell carcinoma of the renal pelvis and ureter, transitional cell carcinoma, urachal cancer, urethral cancer, genitourinary neoplasms, uterine sarcoma, uveal melanoma, vaginal cancer, Verner-Morrison syndrome, verrucous carcinoma, visual pathway glioma, vulvar cancer, Waldenstrom's macroglobulinemia, Warthin's tumor, and Wilms' tumor.

[0280] According to various embodiments, the present invention relates to treatment of acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), anal cancer, astrocytoma, basal cell carcinoma, bile duct cancer, bladder cancer, breast cancer, Burkitt's lymphoma, cervical cancer, chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), chronic myeloproliferative neoplasms, colorectal cancer, diffuse large B-cell lymphoma, endometrial cancer, ependymoma, esophageal cancer, esthesioneuroblastoma, Ewing's sarcoma, fallopian tube cancer, follicular lymphoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, hairy cell leukemia, hepatocellular carcinoma, Hodgkin's lymphoma, hypopharyngeal cancer, Kaposi's sarcoma, and the like. The test can detect numerous types of neoplasms, including, but not limited to, thyroid cancer, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, Merkel cell carcinoma, mesothelioma, oral cancer, neuroblastoma, non-Hodgkin's lymphoma, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumors, pharyngeal cancer, pituitary tumors, prostate cancer, rectal cancer, renal cell carcinoma, retinoblastoma, skin cancer, small cell lung cancer, small intestine cancer, squamous cell cervical carcinoma, T-cell lymphoma, testicular cancer, thymoma, thyroid cancer, uterine cancer, vaginal cancer, and vascular tumors.

[0281] Many embodiments relate to diagnostic or companion diagnostic scans performed during an individual's cancer treatment. When diagnostic scans are performed during treatment, the ability of a drug to treat cancer growth can be monitored. Most anti-cancer therapeutics result in the death and necrosis of neoplastic cells, which should release more nucleic acids from these cells into the sample being tested. Therefore, the level of circulating tumor nucleic acids can be monitored over time, as it should rise during initial treatment and begin to decline as the number of cancerous cells decreases. In some embodiments, treatment is adjusted based on the therapeutic effect on cancer cells. For example, if the treatment is not cytotoxic to neoplastic cells, the dosage can be increased or a drug with greater cytotoxicity can be administered. Alternatively, if the cytotoxicity to cancer cells is good but undesirable side effects are high, the dosage can be reduced or a drug with fewer side effects can be administered.

[0282] Various embodiments also relate to diagnostic scans performed after treatment of an individual to detect residual disease and / or recurrence of cancer. If the diagnostic scan indicates residual and / or recurrence of cancer, further diagnostic tests and / or treatments can be performed as described herein. If the cancer and / or individual is prone to recurrence, diagnostic scans can be performed frequently to monitor any potential recurrence. F. Computer Systems

[0283] In one aspect, the present disclosure provides a computer program product, the computer program product including a non-transitory computer-readable medium having computer-executable code encoded therein, the computer-executable code adapted and executed to implement any one of the aforementioned methods.

[0284] The present disclosure provides a computer system programmed to implement the method of the present disclosure. In some cases, the system can include components such as a processor, an input module for inputting sequencing data or data derived therefrom, a computer-readable medium containing instructions that, when executed by the processor, execute an algorithm on the input related to one or more cell-free nucleic acid molecules, and an output module that provides one or more indicators related to a condition.

[0285] Figure 27 shows a computer system 2701 programmed or otherwise configured to implement some or all of the methods disclosed herein. Computer system 2701 can coordinate various aspects of the present disclosure, such as (i) identifying one or more cell-free nucleic acid molecules containing multiple phase variants from sequencing data derived from multiple cell-free nucleic acid molecules, (ii) analyzing any of the identified cell-free nucleic acid molecules, (iii) determining a subject's condition based at least in part on the identified cell-free nucleic acid molecules, (iv) monitoring the progression of the subject's condition based at least in part on the identified cell-free nucleic acid molecules, (v) identifying a subject based at least in part on the identified cell-free nucleic acid molecules, or (vi) determining an appropriate treatment for a subject's condition based at least in part on the identified cell-free nucleic acid molecules. Computer system 2701 can be a user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.

[0286] The computer system 2701 includes a central processing unit (CPU, herein "processor" and "computer processor") 2705, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 2701 also includes memory or memory locations 2710 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 2715 (e.g., a hard disk), a communication interface 2720 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 2725, such as cache, other memory, data storage, and / or electronic display adapters. The memory 2710, storage unit 2715, interface 2720, and peripheral devices 2725 communicate with the CPU 2705 via a communication bus (solid lines), such as a motherboard. The storage unit 2715 can be a data storage unit (or data repository) for storing data. The computer system 2701 can be operatively coupled to a computer network ("network") 2730 with the aid of the communication interface 2720. Network 2730 can be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. Network 2730 is, in some cases, a telecommunications and / or data network. Network 2730 can include one or more computer servers, which can enable distributed computing such as cloud computing. Network 2730 can, in some cases, implement a peer-to-peer network, which can enable devices coupled to computer system 2701 to operate as clients or servers, with the aid of computer system 2701.

[0287] The CPU 2705 may execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 2710. The instructions may be directed to the CPU 2705, which may then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 2705 may include fetch, decode, execute, and writeback.

[0288] The CPU 2705 may be part of a circuit, such as an integrated circuit. One or more other components of the system 2701 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0289] The storage unit 2715 can store files such as drivers, libraries, and saved programs. The storage unit 2715 can store user data, e.g., user preferences and user programs. The computer system 2701 can optionally include one or more additional data storage units external to the computer system 2701, such as located on a remote server that communicates with the computer system 2701 via an intranet or the Internet.

[0290] Computer system 2701 can communicate with one or more remote computer systems via network 2730. For example, computer system 2701 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 2701 via network 2730.

[0291] The methods described herein can be implemented by machine (e.g., a computer processor) executable code stored in an electronic storage location of the computer system 2701, such as memory 2710 or electronic storage unit 2715. The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 2705. In some cases, the code can be retrieved from the storage unit 2715 and stored in memory 2710 for easy access by the processor 2705. In some situations, the electronic storage unit 2715 can be omitted, and machine-executable instructions are stored in memory 2710.

[0292] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be supplied in a programming language that can be selected to allow the code to be executed in a pre-compiled or compiled manner.

[0293] Aspects of the systems and methods provided herein, such as computer system 2701, can be embodied in programming. Various aspects of the present technology can be considered "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage"-type media can include tangible memory of a computer, processor, etc., or any or all of its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or portions of the software may from time to time be communicated via the Internet or various other telecommunications networks. Such communication can, for example, enable loading of the software from one computer or processor to another, e.g., from an administrative server or host computer to an application server computer platform. Thus, another type of medium that may carry software elements includes optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and over various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0294] Thus, a machine-readable medium such as a computer-executable code may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices of any computer(s), such as may be used to implement the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and optical fiber, including the wires that comprise a bus within a computer system. Carrier wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punched card paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave carrying data or instructions, a cable or link carrying such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0295] Computer system 2701 can include or communicate with an electronic display 2735 that includes a user interface (UI) 2740 for providing, for example, (i) an analysis of any of the identified cell-free nucleic acid molecules, (ii) a determined condition of the subject based at least in part on the identified cell-free nucleic acid molecules, (iii) a determined progression of the subject's condition based at least in part on the identified cell-free nucleic acid molecules, (iv) an identified subject suspected of having a condition based at least in part on the identified cell-free nucleic acid molecules, or (v) a determined treatment of the subject's condition based at least in part on the identified cell-free nucleic acid molecules. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0296] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithm can be implemented by software when executed by the central processing unit 2705. The algorithm can, for example, (i) identify one or more cell-free nucleic acid molecules comprising a plurality of graded variants from sequencing data derived from a plurality of cell-free nucleic acid molecules, (ii) analyze any of the identified cell-free nucleic acid molecules, (iii) determine a subject's condition based at least in part on the identified cell-free nucleic acid molecules, (iv) monitor the progression of the subject's condition based at least in part on the identified cell-free nucleic acid molecules, (v) identify a subject based at least in part on the identified cell-free nucleic acid molecules, or (vi) determine an appropriate treatment for the subject's condition based at least in part on the identified cell-free nucleic acid molecules. [Example]

[0297] The following illustrative examples represent embodiments of the stimuli, systems, and methods described herein and are not meant to be limiting in any way. Example 1: Genomic distribution of phased variants

[0298] An alternative method to double-stranded sequencing to reduce background error rates is described, including the detection of "phased variants" (PVs), in which two or more mutations occur in cis (i.e., on the same strand of DNA, Figure 1A and Figure 1E). Similar to double-stranded sequencing, this method offers a lower error profile due to the concordant detection of two distinct, non-referenced events in individual molecules. However, unlike double-stranded sequencing, both events occur in the same sequencing read pair, thereby increasing the efficiency of genome recovery. Phased mutations are present in diverse cancer types but occur in a typological portion of the genome in B-cell malignancies, likely due to on-target and aberrant somatic hypermutation (aSHM) driven by activation-induced deaminase (AID). The most common regions of aSHM in B-cell non-Hodgkin's lymphoma (NHL) are identified. Described herein is phased variant enrichment and detection sequencing (PhasED-Seq), a novel method for detecting ctDNA via phased variants relative to tumor fractions on the order of approximately one in a million. Described herein is a demonstration that PhasED-Seq can significantly improve the detection of ctDNA in clinical samples both during treatment and before disease recurrence.

[0299] To identify malignant tumors in which PVs could potentially improve disease detection, we assessed the frequency of PVs across cancer types. Publicly available whole-genome sequencing data was analyzed to identify a set of variants occurring at distances less than 170 bp apart, representing the typical length of a single cfDNA fragment consisting of a single core nucleosome and associated linker. The frequencies of these "putative stage variants" (Example 10), which control the total number of SNVs, were identified and summarized (Figure 1B, Figure 5, and Table 1) from 2538 tumors across 24 cancer histologies, including solid tumors and hematological malignancies. PVs were most significantly enriched in two B-cell lymphomas (DLBCL and follicular lymphoma, FL; P<0.05 vs. all other histologies), a group of diseases with hypermutation caused by AID / AICDA. Example 2: Mutational mechanisms underlying PV

[0300] To investigate the origin of PVs, we compared single-base substitution (SBS) mutation signatures that contribute to SNVs occurring within 170 bp of another SNV with those that occur alone (e.g., without other SNVs within 170 bp) (Example 10). As expected, PVs were highly enriched in several mutation signatures associated with clustered mutations. Clustered mutation signatures associated with AID activity (SBS84 and SBS85) were significantly enriched in PVs from B-cell lymphoma and CLL, while signatures associated with APOBEC3B activity (SBS2 and SBS13), another mechanism of kataegis hypermutation, were significantly enriched in PVs from multiple solid cancer tissues, including ovarian, pancreatic, prostate, and breast adenocarcinoma (Figure 1C and Figure 6A-6WW). Clustered mutation signatures associated with AID activity (SBS84 and SBS85) were enriched in PVs found in lymphomas and CLL, whereas signatures associated with APOBEC3B activity (SBS2 and SBS13) were significantly enriched in breast cancer (Figure 1C and Figure 6A-6WW). PVs from multiple tumor types were also associated with SBS4, a tobacco use-related signature. Furthermore, among PVs across multiple tumor histologies, de novo enrichment of several other signatures without apparent mechanism of association was observed (e.g., SBS24, SBS37, SBS38, and SBS39). In contrast, age-associated mutation signatures such as SBS1 and SBS5 were significantly enriched in segregating SNVs. Example 3: PV occurs in genomic regions typical of lymphoid cancers

[0301] To assess the genomic distribution of putative PVs, we first binned these events into 1-kb regions to visualize their frequency across tumor types. We observed a strikingly typological distribution of PVs in individual lymphoid neoplasms (e.g., DLBCL, FL, Burkitt's lymphoma (BL), and chronic lymphocytic leukemia (CLL); Figures 1D and 7). In contrast, non-lymphoid cancers generally did not show substantial recurrence of clustered PVs in typological regions. This lack of typology in PV location held true even when considering melanoma and lung cancer, diseases with high frequencies of PVs.

[0302] Notably, the majority of hypermutated regions were shared across all three lymphoma subtypes, with the highest density found in known targets of aSHM, including BCL2, BCL6, and MYC, as well as the immunoglobulin (Ig) loci encoding the heavy and light chains IGH, IGK, and IGL (Table 2). Strikingly, specific regions within the Ig loci were densely mutated in nearly all lymphoma patients, as well as in CLL patients (Figure 1D). Among lymphoma subtypes, DLBCL tumors possessed the greatest number of 1-kb regions recurrently containing PVs (Figure 8A), consistent with the highest number of recurrently mutated genes observed in this tumor type. In total, 1,639 unique 1-kb regions recurrently containing PVs were identified in B-lymphoid malignancies. Of these lymphoma-associated 1-kb regions, nearly one-third were classified as genomic regions previously associated with physiological or aberrant SHM in B cells. Specifically, 19% (315 / 1639) were located in the Ig region, and 13% (218 / 1639) were among the 68 previously identified targets of aSHM (Table 2). Although most PVs fell into non-coding regions of the genome, further recurrently affected loci not previously described as targets of aSHM were also identified, including XBP1, LPP, and AICDA, among others.

[0303] The distribution of PVs within each lymphoid malignancy correlated with oncogenic features associated with distinct pathophysiology of the corresponding disease. For example, FL cases, in which >90% of tumors harbored oncogenic BCL2 fusions, were significantly more likely to contain staged variants in BCL2 than other lymphoid malignancies (Figure 1D and Figure 8B). Similarly, Burkitt lymphoma (BL) had significantly more PVs in MYC and ID3, two driver genes strongly associated with BL pathogenesis, than other lymphoid malignancies (Figure 1D and Figure 8C-8D). DLBCL molecular subtypes associated with different cells of origin also demonstrated distinct distributions of PVs (Table 2). Specifically, although germinal center B cell-like (GCB) and activated B cell-like (ABC) DLBCLs had similar frequencies of PVs overall (median 798 vs. 516, P = 0.37), a significant enrichment of PVs in the telomeric IGH class switch regions (Sγ1 and Sγ3) of ABC-DLBCLs was found, consistent with a previous report (Figure 8E). Conversely, GCB-DLBCLs had more phased haplotypes in the centromeric IGH class switch regions (Sα2 and Sε) and BCL2. Example 4: Design and validation of a PhaseED-Seq panel for lymphoma

[0304] To validate these PV-rich regions and evaluate their utility for disease detection from ctDNA, we designed a sequencing panel targeting putative PVs identified within WGS from three independent cohorts of DLBCL and CLL patients (Figure 2A and Example 10). This final phased variant enrichment and detection sequencing (PhasED-Seq) panel targeted approximately 115 kb of genomic space focused on PVs, along with an additional approximately 200 kb of target genes that are recurrently mutated in B-NHL (Table 3). The 115 kb space dedicated to PV capture targets only 0.0035% of the human genome, yet captures 26% of the phased variants observed in mature B-cell neoplasms characterized by WGS (Figure 9A), thus resulting in approximately 7,500-fold PV enrichment by PhaseED-Seq over WGS.

[0305] We compared predicted SNV and PV recovery with a previously reported CAPP-Seq selector designed to maximize SNVs per patient in B-cell lymphoma (Figure 9A-C). Considering diverse B-NHLs with available WGS data, PhaseED-Seq recovered 3.0-fold more SNVs (81 vs. 27) and 2.9-fold more PVs (50 vs. 17) in median cases than the previous CAPP-Seq panel. This observation highlights the importance of including non-coding portions of the genome for maximal mutation recovery. To experimentally validate these yield improvements, we profiled 16 pre-treatment tumor or plasma DNA samples (Table 4) from DLBCL patients. Both the CAPP-Seq and PhaseED-Seq panels were applied in parallel to each specimen, which were then sequenced to high unique molecular depth (Figure 2B). Compared to the expected enrichment established from WGS, we observed a similar improvement in SNV yield with PhaseED-Seq compared to CAPP-Seq (2.7-fold; median 304.5 vs. 114). However, when enumerating the PVs observed in individual sequenced DNA fragments, we observed an improvement in favor of PhaseED-Seq over the improvement expected from WGS (7.7-fold; median 5554 vs. 719.5 PVs / case). This improvement is potentially due to either 1) higher sequencing depth in targeted sequencing, which results in improved rare allele detection, or 2) enumeration of higher-order PVs in targeted sequencing by either PhaseED-Seq or CAPP-Seq, which was not considered in the WGS design (i.e., more than two SNVs per fragment; Figure 9D-F). Furthermore, across the 1 kb window of the panel, a strong correlation was observed between the frequency of estimated PVs in the WGS data and PVs from targeted sequencing by PhasED-Seq across 101 DLBCL samples ( Figure 2C ), further validating the frequency and distribution of PVs in B-cell malignancies. Example 5: Phase variant differences between lymphoma subtypes

[0306] After validating the PhasED-Seq panel, we investigated biological differences in PVs among various B-cell malignancies, including DLBCL (n = 101), primary mediastinal large B-cell lymphoma (PMBCL) (n = 16), and classical Hodgkin lymphoma (cHL) (n = 23). The number of SNVs identified per case did not differ significantly between lymphoma subtypes (Figures 9G-9K). However, when considering mutational haplotypes, cHL had a significantly lower PV burden than either DLBCL or PMBCL. In addition to this quantitative difference, we also observed differences in the genomic location of PVs among different B-cell lymphoma subtypes (Figures 2D-2E and 10-12). This included previously established biological associations among DLBCL subtypes, such as a higher number of BCL2 PVs in GCB than ABC DLBCL, with an inverse correlation observed for PIM1. We also observed more frequent PVs in CIITA in PMBCL compared with DLBCL, a gene whose breakpoints are common in PMBCL. Relative enrichment was also observed across the IGH locus, with more frequent PVs in the Sγ3 and Sγ1 regions in ABC-DLBCL (compared to GCB-DLBCL) and, interestingly, more frequent PVs in the Sε locus in cHL compared with DLBCL (Figure 2E and Figure S13). Overall, after correcting for multiple hypothesis testing, we found significant relative enrichment in 25 loci between ABC-DLBCL and GCB-DLBCL, 24 between DLBCL and PMBCL, and 40 between DLBCL and cHL (Figures S10-S12). Example 6: Recovery of phased variants by PhaseED-Seq

[0307] Efficient recovery of DNA molecules is desirable to facilitate the detection of ctDNA using PV. Hybrid capture sequencing is potentially sensitive to DNA mismatches, and increasing mutations decrease hybridization efficiency. Indeed, AID hotspots can contain local mutation rates of 5-10%, with even higher rates in specific regions of IGH. To empirically evaluate the impact of mutation rate on capture efficiency, we simulated in silico 150-mer DNA hybridization with varying mutation rates. As expected, the predicted binding energy decreased with increasing mutation number (Figure 14A). Notably, randomly distributed mutations had a greater impact on binding energy than clustered mutations. To evaluate the impact of this decreased binding affinity, we synthesized 150-mer DNA oligonucleotides with 0-10% differences from the reference sequence at MYC and BCL6, two loci targeted by aSHM. To assess the worst-case scenario for hybridization, non-reference bases were randomly distributed rather than clustered (Example 10). An equimolar mixture of these oligonucleotides was then captured with a PhasED-Seq panel. Consistent with in silico predictions, increasing mutation rates resulted in decreased capture efficiency (Figure 3A). Molecules with a 5% mutation rate were captured with 85% efficiency relative to their fully wild-type counterparts, whereas molecules with 10% mutations were captured with only 27% relative efficiency. To assess the prevalence of this level of mutation in human tumors, we calculated the proportion of mutated bases in overlapping 151-bp windows (Example 10) and examined the distribution of variants in the panel in 140 patients with B-cell lymphoma. Only 7% (10 / 140) of patients had any 151-bp window with a mutation rate exceeding 10% (Figures 14B-C). Indeed, experiments using synthetic oligonucleotides recovered 5% mutation rates with nearly the same efficiency as wild-type sequences. In more than half of all cases examined, no locus had a mutation rate greater than 5% in any window, but in all cases, more than 90% of the windows had mutations less than 5%.Overall, these observations indicate that despite hybridization bias, the majority of stepwise mutations can be recovered by efficient hybrid capture. Example 7: Error Profiles and Detection Limits for Phased Variant Sequencing

[0308] Previous methods for highly error-suppressed sequencing applied to cfDNA have utilized either a combination of molecular and in silico methods for error suppression (e.g., integrated digital error suppression, iDES) or double-stranded molecule recovery. However, each of these has limitations in either detecting events at ultra-low tumor rates or efficiently recovering original DNA molecules, which are important considerations for cfDNA analysis where input DNA is limited. We compared the error profile and recovery of input genomes from plasma cfDNA samples from 12 healthy adults by PhasED-Seq with both iDES-CAPP-Seq and double-stranded sequencing. While iDES-enhanced CAPP-Seq had a lower background error profile than barcode deduplication alone, double-stranded sequencing provided the lowest background error rate for non-reference single-base substitutions (Figure 3B, 3.3 × 10). -5 vs. 1.2 x 10 -5 , P<0.0001). However, the rate of stepwise errors (e.g., multiple non-reference bases occurring on the same sequencing fragment) was significantly lower than the rate of single errors in either the iDES-enhanced CAPP-Seq or double-stranded sequencing data. This was true for the occurrence of both two (2× or "doublet" PV) or three (3× or "triplet" PV) substitutions on the same DNA molecule (Figure 3B, 8.0 × 10, respectively). -7 and 3.4 × 10 -8, P<0.0001). Stepwise errors, including C-to-T or T-to-C transition substitutions, were more common than other types of PVs (Figure 14D). Notably, doublet PV error rates in cfDNA also correlated with the distance between positions, with the highest PV error rates consisting of adjacent SNVs (e.g., DNVs), and the error rate decreased as the distance between the constituent variants increased (Figure 14E). Considering unique molecular depth, double-stranded sequencing recovered only 19% of all unique cfDNA fragments (Figure 3C). In contrast, the unique depth of PVs within a genomic distance of less than 20 bp was nearly identical to the depth of individual positions (e.g., molecules covering individual SNVs). Similarly, PVs up to 80 bp in size had depths exceeding 50% of the median unique molecular depth of the samples. Importantly, nearly half (48%) of all PVs were within 80 bp of each other, demonstrating their utility for disease detection from input-limited cfDNA samples (Figure 3D).

[0309] To quantitatively compare the performance of PhasED-Seq with alternative methods for ctDNA detection, we generated limiting dilutions of ctDNA from three lymphoma patients into healthy control cfDNA, yielding expected tumor fractions ranging from 0.1% to 0.00005% (1 in 2,000,000) (Example 10). We compared the expected tumor fractions to the estimated tumor content in each of these dilutions using PhasED-Seq to track tumor-derived PVs, as well as error-suppression detection methods tailored to individual SNVs (e.g., iDES-enhanced CAPP-Seq or double-stranded sequencing; Figure 3E). All methods performed equally well down to tumor fractions of 0.01% (1 in 10,000). However, below this level (e.g., 0.001%, 0.0002%, 0.0001%, and 0.00005%), both PhaseED-Seq and double-stranded sequencing significantly outperformed iDES-enhanced CAPP-Seq (P<0.0001 for double-stranded, "2x" PhaseED-Seq, and "3x" PhaseED-Seq; Figure 3E). Furthermore, when compared with double-stranded sequencing, tracking two or three in-phase variants (e.g., 2x and 3x PhaseED-Seq) more accurately identified expected tumor content and had superior linearity by up to 2,000,000 times (P=0.005 for double-stranded vs. 2x PhaseED-Seq, P=0.002 for 3x PhaseED-Seq) (Example 10). We assessed the specificity of PVs by searching for evidence of tumor-derived SNVs or PVs in cfDNA samples from 12 unrelated healthy control subjects and healthy control subjects used for limiting dilution. Again, both 2x- and 3x-PhasED-Seq demonstrated significantly lower background signal levels than CAPP-Seq and double-stranded sequencing (Figure 3F). This low error rate and background from PVs improves the detection limit for ctDNA disease detection. In some instances, the sequencing-based cfDNA assay methods described herein (e.g., the methods shown in Figures 3E and 3F) do not require molecular barcodes to achieve sophisticated error suppression and low detection limits.Signals assessed by the barcode-free method used a limiting dilution series from 1:1,000 to 5:10 million and a "blank" control (Figures 23A-B).

[0310] This dilution series was used to assess the detection limit for a given number of PVs (Figures 3G-3I). When considering a set of PVs within a 150-base pair (bp) region, the probability of detection for a given sample can be accurately modeled by binomial sampling, taking into account both the sequencing depth and the number of 150-bp regions with PVs (Example 10). Example 8: Improvements in the detection of low-burden minimal residual disease

[0311] To test the utility of the lower LOD obtained by PhaseED-Seq for detecting ultralow-load MRD from cfDNA, serial cell-free DNA samples were sequenced from a patient undergoing frontline treatment for DLBCL (Figure 4A). Using CAPP-Seq, this patient had undetectable ctDNA after only one cycle of treatment and remained undetectable in multiple subsequent samples during and after treatment. This patient subsequently had the reappearance of detectable ctDNA more than 250 days after the start of treatment and had final clinical and radiological disease progression 5 months later, demonstrating serial false-negative measurements by CAPP-Seq. Remarkably, all four of the plasma samples that were undetectable by CAPP-Seq during and after treatment had detectable ctDNA levels by PhaseED-Seq, with a mean allele fraction as low as 6 per million. This increase in sensitivity improved the lead time for disease detection by ctDNA compared with radiological surveillance, from 5 months with CAPP-Seq to 10 months with PhaseED-Seq.

[0312] We next evaluated the performance of PhaseED-Seq ctDNA detection in a cohort of 107 patients with large B-cell lymphoma and blood samples available after one or two cycles of standard immunochemotherapy. Importantly, ctDNA levels measured by PhaseED-Seq were highly correlated with levels measured by CAPP-Seq. In total, 443 tumor, germline, and cell-free DNA samples were evaluated, including cfDNA before treatment (n = 107) and after one or two cycles of treatment (n = 82 and 89, respectively). Before treatment, patient-specific PVs were detectable by PhaseED-Seq in 98% of samples, with 95% specificity in cfDNA from healthy controls (Figures 15 and 16A). Importantly, when considering both pre- and post-treatment samples, ctDNA levels measured by PhaseED-Seq were highly correlated with those measured by CAPP-Seq (Spearman rho = 0.91, Figure 16B). Next, we compared quantitative ctDNA levels measured by PhaseED-Seq and CAPP-Seq from cfDNA samples after treatment initiation. In total, 72% (78 / 108) of samples with detectable ctDNA by PhaseED-Seq after one or two cycles were also detected by conventional CAPP-Seq (Figure 4B). Among the 108 samples detected by PhasED-Seq, disease burden was significantly lower for those with detectable (72%) versus undetectable (28%) ctDNA levels using conventional CAPP-Seq, with a difference in median ctDNA levels exceeding 10x (tumor fraction 2.2 × 10 -4 vs. 1.2 x 10 -5 , P<0.001, Figure 4B). In total, an additional 16% (13 / 82) of samples after one cycle of treatment and 19% (17 / 89) of samples after two cycles of treatment had detectable ctDNA when comparing PhasED-Seq with CAPP-Seq (Figure 4C).

[0313] ctDNA molecular response criteria have been previously described for DLBCL patients using CAPP-Seq, including a major molecular response (MMR), defined as a 2.5-log reduction in ctDNA after two cycles of treatment. While MMR at this time point is prognostic for outcome, many patients have undetectable ctDNA by CAPP-Seq at this landmark (Figure 4D-E). Importantly, even in patients with undetectable ctDNA by CAPP-Seq, detection of subclinical ultralow ctDNA levels by PhasED-Seq was prognostic for outcomes, including recurrence-free survival and overall survival (Figure 4D). Indeed, among 89 patients with available samples from this time point, 58% (52 / 89) had undetectable ctDNA by CAPP-Seq at their interim MMR assessment after completing two of the planned six cycles of treatment. Using PhasED-Seq, 33% (17 / 52) of the samples not detected by CAPP-Seq had evidence of ctDNA, as evidenced by PV, at levels as low as approximately 3:1,000,000 (Figures 17A-17D). These 17 cases, further detected by PhasED-Seq, represent potential false-negative results by CAPP-Seq. Similar results were seen at the early molecular response (EMR) time point (i.e., after one cycle of treatment, Figures 18A-18H).

[0314] Although detection of ctDNA in DLBCL after one or two cycles of therapy is a known adverse prognostic marker, outcomes for patients with undetectable ctDNA at these time points are heterogeneous (Figure 4E and Figure 18F). Importantly, even in patients with undetectable ctDNA by CAPP-Seq after one or two cycles of therapy, detection of ultralow ctDNA levels by PhasED-Seq was strongly prognostic for outcomes, including relapse-free survival (Figure 4F, Figures 17C-17D, Figures 18C-18D, and Figure 18G). When combining detection by PhasED-Seq with previously described MMR thresholds, patients could be stratified into three groups: patients who did not achieve MMR, patients who achieved MMR but had persistent ctDNA, and patients with undetectable ctDNA (Figure 4G). Interestingly, patients who did not achieve MMR were at particularly high risk for early events despite additional planned first-line therapy (e.g., within the first year of treatment), whereas patients with persistently low levels of ctDNA appeared to be at higher risk for later recurrence or progression events. In contrast, patients with undetectable ctDNA after two cycles of treatment with PhasED-Seq had overwhelmingly favorable outcomes, with 95% recurrence-free and 97% overall survival at 5 years. Similar results were seen at the time of EMR after one cycle of treatment (Figure 18H). Example 9: Exemplary embodiment of mutation detection using next generation sequen...

Claims

1. 1. A method of using a comparison between a first state of a condition in a subject and a second state of the condition in the subject as an indicator for progression of the condition in the subject, comprising: (a) identifying a first set of one or more cell-free nucleic acid molecules from a first plurality of cell-free nucleic acid molecules obtained or derived from the subject; (b) identifying a second set of one or more cell-free nucleic acid molecules from a second plurality of cell-free nucleic acid molecules obtained or derived from the subject, identifying the second plurality of cell-free nucleic acid molecules obtained from the subject after obtaining the first plurality of cell-free nucleic acid molecules from the subject; each of the one or more cell-free nucleic acid molecules comprises a plurality of graded variants relative to a reference genome sequence separated by at least one nucleotide, the plurality of graded variants comprising a first graded variant and a second graded variant, the first graded variant of the plurality of graded variants and the second graded variant of the plurality of graded variants being within 170 base pairs of each other; the presence of said plurality of graded variants is indicative of a first aspect and a second aspect of said condition in said subject; a comparison between the first status of the condition of the subject and the second status of the condition of the subject indicates the progression of the condition of the subject; the one or more cell-free nucleic acid molecules are captured from among the plurality of cell-free nucleic acid molecules with a set of nucleic acid probes, the set of nucleic acid probes configured to hybridize to at least a portion of the cell-free nucleic acid molecule that includes one or more genomic regions associated with the condition; The set of nucleic acid probes (i) comprises the nucleic acid probes shown in Table 1 below. Genomic regions identified in (ii) The second table below genomic regions identified in, or (iii) The third table below Genomic regions identified in is designed to hybridize to at least 5% of The method, wherein the condition is a B-cell lymphoma and the subject is a human subject.

2. The method of claim 1 , wherein the progression of the condition is a worsening of the condition.

3. 10. The method of claim 1, wherein said progression of said condition is at least partial remission of said condition.

4. 4. The method of any one of claims 1-3, wherein the second plurality of cell-free nucleic acid molecules is obtained from the subject at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 4 weeks, at least about 2 months, or at least about 3 months after obtaining the first plurality of cell-free nucleic acid molecules from the subject.

5. 5. The method of any one of claims 1 to 4, wherein the subject is subjected to treatment for the condition (i) before obtaining the second plurality of cell-free nucleic acid molecules from the subject, and (ii) after obtaining the first plurality of cell-free nucleic acid molecules from the subject.

6. 6. The method of any one of claims 1 to 5, wherein said progression of said condition indicates minimal residual disease of said condition in said subject.

7. The method of any one of claims 1 to 6, wherein the progression of the condition is indicative of tumor or cancer burden in the subject.

8. 8. The method of any one of claims 1-7, wherein at least 50% of the one or more cell-free nucleic acid molecules comprising a plurality of graded variants comprise a single base variant (SNV) that is separated by at least two nucleotides from an adjacent SNV.

9. The method of any one of claims 1 to 8, wherein the reference genome sequence is derived from a sample of the subject.

10. The method of claim 9, wherein the sample is a healthy sample.

11. 11. The method of any one of claims 1 to 10, wherein the set of nucleic acid probes is designed based on the plurality of graded variants identified by comparing (i) sequencing data from B-cell lymphoma and (ii) sequencing data from healthy cells of the subject or a healthy cohort.

12. The method of any one of claims 1 to 11, wherein the subject is undergoing treatment for the condition prior to (a).

13. 13. The method of claim 5 or 12, wherein the treatment comprises chemotherapy, radiation therapy, chemoradiotherapy, immunotherapy, adoptive cell therapy, hormone therapy, targeted drug therapy, surgery, transplantation, or blood transfusion.

14. The method of any one of claims 1 to 13, wherein the plurality of cell-free nucleic acid molecules is derived from a bodily sample of the subject.

15. 15. The method of claim 14, wherein the body sample comprises plasma, serum, or blood.

16. 16. The method of any one of claims 1-15, wherein the condition comprises a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt's lymphoma, and B-cell chronic lymphocytic leukemia.

17. 17. The method of any one of claims 1 to 16, wherein the set of nucleic acid probes is designed to hybridize to at least 10% of (i) the genomic regions identified in the first table, (ii) the genomic regions identified in the second table, or (iii) the genomic regions identified in the third table.

18. 18. The method of any one of claims 1 to 17, wherein the set of nucleic acid probes is designed to hybridize to at least 20% of (i) the genomic regions identified in the first table, (ii) the genomic regions identified in the second table, or (iii) the genomic regions identified in the third table.

19. 19. The method of any one of claims 1 to 18, wherein the set of nucleic acid probes is designed to hybridize to at least 40% of (i) the genomic regions identified in the first table, (ii) the genomic regions identified in the second table, or (iii) the genomic regions identified in the third table.

20. 20. The method of any one of claims 1 to 19, wherein the set of nucleic acid probes is designed to hybridize to at least 60% of (i) the genomic regions identified in the first table, (ii) the genomic regions identified in the second table, or (iii) the genomic regions identified in the third table.

21. 21. The method of any one of claims 1 to 20, wherein the set of nucleic acid probes is designed to hybridize to at least 80% of (i) the genomic regions identified in the first table, (ii) the genomic regions identified in the second table, or (iii) the genomic regions identified in the third table.

22. 22. The method of any one of claims 1 to 21, wherein the set of nucleic acid probes is designed to hybridize to about 100% of (i) the genomic regions identified in the first table, (ii) the genomic regions identified in the second table, or (iii) the genomic regions identified in the third table.

Citation Information

Patent Citations

  • Analysis of nucleic acid sequences

    JP2017522866A

  • Deep sequencing of peripheral blood plasma DNA is highly reliable in confirming the diagnosis of myelodysplastic syndrome

    JP2017533714A

  • diagnostic method

    JP2018522531A

  • Myeloma treatment or monitoring of progression

    JP2018537128A

  • Detection and diagnosis of cancer evolution

    JP2019512823A