Methods and systems for analyzing nucleic acid molecules
By analyzing cell-free nucleic acids for phased variants using sequencing data and nucleic acid probes, the method addresses the limitations of current MRD detection, enhancing sensitivity and reliability in cancer diagnosis and monitoring.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
- Filing Date
- 2026-02-16
- Publication Date
- 2026-05-26
AI Technical Summary
Current methods for detecting minimal residual disease (MRD) in cancer patients using cell-free nucleic acids are limited by low input DNA volume and high background error rates, leading to insufficient detection of residual disease, particularly in conditions like diffuse large B-cell lymphoma, colorectal, and breast cancer, resulting in false-negative test outcomes.
A method involving the analysis of cell-free nucleic acids using sequencing data to identify phased variants separated by at least one nucleotide, processed by a computer system to enhance detection sensitivity and specificity, with a detection limit of less than 1 out of 50,000 observations, utilizing nucleic acid probes that hybridize to stepwise variants for improved accuracy.
The method achieves enhanced sensitivity and reliability in detecting cancer-derived nucleic acids, allowing for improved monitoring and treatment decisions with a significantly reduced false-negative rate.
Smart Images

Figure 2026086762000186 
Figure 2026086762000187 
Figure 2026086762000188
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the interests of U.S. Provisional Patent Application No. 62 / 931,688, filed on November 6, 2019, which is incorporated herein by reference in its entirety. Sequence List
[0002] This application includes an electronically submitted sequence listing in ASCII format, the entirety of which is incorporated herein by reference. The name of the above ASCII copy, created on 3 November 2020, is 58626-702_601_SL.txt, and its size is 307,199 bytes. Government rights
[0003] This invention was made with government support under CA233975, CA241076, and CA188298, awarded by the National Institutes of Health. The government has certain rights to this invention. [Background technology]
[0004] background Non-invasive blood tests that can detect somatic changes (e.g., mutant nucleic acids) based on the analysis of cell-free nucleic acids (e.g., cell-free deoxyribonucleic acid (cfDNA) and cell-free ribonucleic acid (cfRNA)) are attractive candidates for cancer screening applications because biological samples (e.g., biological fluids) are relatively easy to obtain. Circulating tumor nucleic acids (e.g., ctDNA or ctRNA, i.e., nucleic acids derived from cancerous cells) can be susceptible and specific biomarkers in numerous cancer subtypes. However, current methods for minimal residual disease (MRD) detection from ctDNA may be limited by one or more factors, such as low input DNA volume and high background error rates.
[0005] Recent approaches have improved ctDNA MRD performance by tracking multiple somatic mutations using error-suppressed sequencing, resulting in a low detection limit of 4 parts per hundred thousand from a limited cfDNA input. Detecting residual disease during or after treatment is a powerful tool, and detectable MRD represents an adverse prognostic sign, even during radiological remission. However, current detection limits may be insufficient to universally detect residual disease in patients where disease relapse or progression is anticipated. This “loss of detection” is exemplified in diffuse large B-cell lymphoma (DLBCL), where ctDNA detection after two cycles of curative treatment is a strong prognostic marker. Nevertheless, nearly one-third of patients experiencing disease progression lack detectable ctDNA at this landmark, representing a “false-negative” test. Similar false-negative rates have been observed in colorectal and breast cancer. [Overview of the project] [Means for solving the problem]
[0006] overview This disclosure provides methods and systems for analyzing cell-free nucleic acids (e.g., cfDNA, cfRNA) from a subject. The methods and systems of this disclosure can utilize sequencing results derived from the subject to detect cancer-derived nucleic acids (e.g., ctDNA, ctRNA) for disease diagnosis, disease monitoring, or treatment decisions. The methods and systems of this disclosure may demonstrate improved sensitivity, specificity, and / or reliability in the detection of cancer-derived nucleic acids.
[0007] In one embodiment, the disclosure includes (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from a subject using a computer system, and (b) processing the sequencing data using a computer system to identify one or more cell-free nucleic acid molecules from the plurality of cell-free nucleic acid molecules, each of the one or more cell-free nucleic acid molecules comprising a plurality of phased variants with respect to a reference genome sequence, and at least about 10% of the one or more cell-free nucleic acid molecules comprising a plurality of phased variants The present invention provides a method comprising (c) identifying a cell-free nucleic acid molecule, comprising (a) a first stepwise variant of a stepwise variant and a second stepwise variant of a plurality of stepwise variants separated by at least one nucleotide, and (b) analyzing the identified cell-free nucleic acid molecule by a computer system to determine the state of the substance.
[0008] In some embodiments of any one of the methods disclosed herein, at least about 10% of the cell-free nucleic acid molecules comprises at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of 1 or more cell-free nucleic acid molecules.
[0009] In one embodiment, the Disclosure provides a method comprising: (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from a subject using a computer system; (b) processing the sequencing data using a computer system to identify one or more cell-free nucleic acid molecules from the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules includes a plurality of stepwise variants to a reference genome sequence separated by at least one nucleotide; and (c) analyzing the identified one or more cell-free nucleic acid molecules using a computer system to determine the state of the subject.
[0010] In one aspect, the present disclosure provides a method comprising: (a) obtaining sequencing data derived from a plurality of cell-free nucleic acid molecules obtained from or derived from a subject; (b) processing the sequencing data to identify one or more cell-free nucleic acid molecules out of the plurality of cell-free nucleic acid molecules at a detection limit of less than about 1 out of 50,000 observations from the sequencing data; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine the state of the subject.
[0011] In some embodiments of any one of the methods disclosed herein, the detection limit of the identification step is less than about 1 out of 100,000 observations, less than about 1 out of 500,000 observations, less than about 1 out of 1,000,000 observations, less than about 1 out of 1,500,000 observations, or less than about 1 out of 2,000,000 observations from the sequencing data.
[0012] In some embodiments of any one of the methods disclosed herein, each of the one or more cell-free nucleic acid molecules comprises a plurality of stepwise variants relative to a reference genomic sequence. In some embodiments of any one of the methods disclosed herein, the first stepwise variant of the plurality of stepwise variants and the second stepwise variant of the plurality of stepwise variants are separated by at least one nucleotide.
[0013] In some embodiments of any one of the methods disclosed herein, processes (a)-(c) are performed by a computer system.
[0014] In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on nucleic acid amplification. In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on polymerase chain reaction. In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on amplicon sequencing.
[0015] In some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on next-generation sequencing (NGS). Alternatively, in some embodiments of any one of the methods disclosed herein, the sequencing data is generated based on non-hybridization-based NGS.
[0016] In some embodiments of any one of the methods disclosed herein, the sequencing data is generated without using molecular barcoding of at least a portion of the plurality of cell-free nucleic acid molecules. In some embodiments of any one of the methods disclosed herein, the sequencing data is obtained without using sample barcoding of at least a portion of the plurality of cell-free nucleic acid molecules.
[0017] In some embodiments of any one of the methods disclosed herein, the sequencing data is obtained without in-silico removal or suppression of (i) background errors or (ii) sequencing errors.
[0018] In one aspect, the present disclosure presents a method of treating a condition of a subject, comprising: (a) identifying the subject for treatment of the condition, wherein the subject is determined to have the condition based on the identification of one or more cell-free nucleic acid molecules obtained from or derived from the subject, each of the identified one or more cell-free nucleic acid molecules comprising a plurality of stepwise variants relative to a reference genomic sequence separated by at least one nucleotide, the presence of the plurality of stepwise variants indicating the condition of the subject; and (b) subjecting the subject to a treatment based on the identification in (a).
[0019] In one embodiment, the Disclosure provides a method for monitoring the progression of a state of a subject, comprising: (a) determining a first state of the subject based on the identification of a first set of one or more cell-free nucleic acid molecules from a first set of cell-free nucleic acid molecules obtained from or derived from the subject; (b) determining a second state of the subject based on the identification of a second set of one or more cell-free nucleic acid molecules from a second set of cell-free nucleic acid molecules obtained from or derived from the subject, wherein the second set of cell-free nucleic acid molecules is obtained from the subject after the first set of cell-free nucleic acid molecules are obtained from the subject; and (c) determining the progression of the state based on the first and second states of the state, wherein each of the one or more cell-free nucleic acid molecules comprises a set of stepwise variants to a reference genome sequence separated by at least one nucleotide.
[0020] In some embodiments of any one of the methods disclosed herein, the progression of a state is the deterioration of a state.
[0021] In some embodiments of any one of the methods disclosed herein, the progression of the condition is at least partial remission of the condition.
[0022] In some embodiments of any one of the methods disclosed herein, the presence of multiple stepwise variants indicates a first or second state of the state in question.
[0023] In some embodiments of any one of the methods disclosed herein, a second plurality of cell-free nucleic acid molecules are obtained from a subject at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 4 weeks, at least about 2 months, or at least about 3 months after obtaining the first plurality of cell-free nucleic acid molecules.
[0024] In some embodiments of any one of the methods disclosed herein, the subject is subjected to treatment for a condition (i) before obtaining a second plurality of cell-free nucleic acid molecules from the subject, and (ii) after obtaining a first plurality of cell-free nucleic acid molecules from the subject.
[0025] In some embodiments of any one of the methods disclosed herein, the progression of the condition represents minimal residual disease of the condition in question. In some embodiments of any one of the methods disclosed herein, the progression of the condition represents tumor or cancerous burden of the condition in question.
[0026] In some embodiments of any one of the methods disclosed herein, one or more cell-free nucleic acid molecules are captured from among a plurality of cell-free nucleic acid molecules by a set of nucleic acid probes, the set of nucleic acid probes is configured to hybridize to at least a portion of the cell-free nucleic acid molecules, which include one or more genomic regions associated with the state.
[0027] In one embodiment, the Disclosure provides (a) a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained from or derived from a subject, wherein each nucleic acid probe of the set of nucleic acid probes is designed to hybridize to at least a portion of target cell-free nucleic acid molecules comprising a plurality of stepwise variants to a reference genome sequence separated by at least one nucleotide, wherein each nucleic acid probe comprises an activatable reporter, and the activation of the activatable reporter is selected from the group comprising (i) hybridization of the individual nucleic acid probe to the plurality of stepwise variants and (ii) dehybridization of at least a portion of the individual nucleic acid probes hybridized to the plurality of stepwise variants; (b) detecting the activated activatable reporter to identify one or more cell-free nucleic acid molecules of a plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of stepwise variants; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine the state of the subject.
[0028] In one embodiment, the disclosure provides a mixture comprising (a) (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained from or derived from a subject, wherein each nucleic acid probe of the set of nucleic acid probes is designed to hybridize to at least a portion of a target cell-free nucleic acid molecule comprising a plurality of stepwise variants with respect to a reference genome sequence, and each nucleic acid probe comprises an activatable reporter, the activation of which (i) hybridization of the individual nucleic acid probe to the plurality of stepwise variants and (ii) hybridization of the plurality of stepwise variants The present invention provides a method comprising: (b) detecting an activated activatable reporter agent to identify one or more cell-free nucleic acid molecules among a plurality of cell-free nucleic acid molecules, each of which comprises a plurality of stepwise variants, and the detection limit of the identification step is less than about one of 50,000 cell-free nucleic acid molecules among the plurality of cell-free nucleic acid molecules; and (c) analyzing the identified one or more cell-free nucleic acid molecules to determine the state of the subject.
[0029] In some embodiments of any one of the methods disclosed herein, the detection limit of the identification step is less than about 1 out of 100,000, less than 1 out of 500,000, less than 1 out of 1,000,000, less than 1 out of 1,500,000, or less than 1 out of 2,000,000 cell-free nucleic acid molecules.
[0030] In some embodiments of any one of the methods disclosed herein, the first stepwise variant of the plurality of stepwise variants and the second stepwise variant of the plurality of stepwise variants are separated by at least one nucleotide.
[0031] In some embodiments of any one of the methods disclosed herein, an activatable reporter agent is activated during the hybridization of individual nucleic acid probes into multiple stepwise variants.
[0032] In some embodiments of any one of the methods disclosed herein, an activatable reporter agent is activated during the dehybridization of at least some of the individual nucleic acid probes hybridized into multiple stepwise variants.
[0033] In some embodiments of any one of the methods disclosed herein, the method further comprises (1) mixing a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules.
[0034] In some embodiments of any one of the methods disclosed herein, the activatable reporter agent is a fluorophore.
[0035] In some embodiments of any one of the methods disclosed herein, the analysis of one or more identified cell-free nucleic acid molecules includes analyzing (i) one or more identified cell-free nucleic acid molecules and (ii) other cell-free nucleic acid molecules that do not include multiple stepwise variants, as different variables.
[0036] In some embodiments of any one of the methods disclosed herein, the analysis of one or more identified cell-free nucleic acid molecules is not based on other cell-free nucleic acid molecules of multiple cell-free nucleic acid molecules that do not include multiple stepwise variants.
[0037] In some embodiments of any one of the methods disclosed herein, the number of stepwise variants from one or more identified cell-free nucleic acid molecules indicates the state of interest. In some embodiments, the ratio of (i) the number of stepwise variants from one or more cell-free nucleic acid molecules to (ii) the number of single nucleotide variants (SNVs) from one or more cell-free nucleic acid molecules indicates the state of interest.
[0038] In some embodiments of any one of the methods disclosed herein, the frequency of multiple stepwise variants in one or more identified cell-free nucleic acid molecules indicates the condition of interest. In some embodiments, the frequency indicates disease cells associated with the condition. In some embodiments, the condition is diffuse large B-cell lymphoma, and the frequency indicates whether one or more cell-free nucleic acid molecules originate from germinal center B cells (GCBs) or activated B cells (ABCs).
[0039] In some embodiments of any one of the methods disclosed herein, the genomic origin of one or more identified cell-free nucleic acid molecules indicates the state of the subject.
[0040] In some embodiments of any one of the methods disclosed herein, the first and second stepwise variants are separated by at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 nucleotides. In some embodiments of any one of the methods disclosed herein, the first and second stepwise variants are separated by up to about 180 nucleotides, up to about 170 nucleotides, up to about 160 nucleotides, up to about 150 nucleotides, or up to about 140 nucleotides.
[0041] In some embodiments of any one of the methods disclosed herein, at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants contain a single nucleotide variant (SNV) located at least two nucleotides away from an adjacent SNV.
[0042] In some embodiments of any one of the methods disclosed herein, the stepwise variant comprises at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, or at least 25 stepwise variants within the same cell-free nucleic acid molecule.
[0043] In some embodiments of any one of the methods disclosed herein, the identified one or more cell-free nucleic acid molecules comprise at least two, at least three, at least four, at least five, at least ten, at least 50, at least 100, at least 500, or at least 1,000 cell-free nucleic acid molecules.
[0044] In some embodiments of any one of the methods disclosed herein, the reference genome sequence is derived from a reference cohort. In some embodiments, the reference genome sequence includes a consensus sequence from a reference cohort. In some embodiments, the reference genome sequence includes at least a portion of the HG19 human genome, HG18 genome, HG17 genome, HG16 genome, or HG38 genome.
[0045] In some embodiments of any one of the methods disclosed herein, the reference genome sequence is derived from the sample of interest.
[0046] In some embodiments of any one of the methods disclosed herein, the sample is a healthy sample. In some embodiments, the sample comprises healthy cells. In some embodiments, the healthy cells comprise healthy leukocytes.
[0047] In some embodiments of any one of the methods disclosed herein, the sample is a disease sample. In some embodiments, the disease sample comprises disease cells. In some embodiments, the disease cells comprises tumor cells. In some embodiments, the disease sample comprises a solid tumor.
[0048] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes is designed based on a plurality of stepwise variants identified by comparing (i) sequencing data from a solid tumor, lymphoma, or hematological malignancy of interest with (ii) sequencing data from healthy cells of interest or a healthy cohort. In some embodiments, the healthy cells are from interest. In some embodiments, the healthy cells are from a healthy cohort.
[0049] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes is designed to hybridize to at least a portion of the sequences of a state-associated genomic locus. In some embodiments, the state-associated genomic locus is known to exhibit abnormal somatic hypermutation when the subject has the state.
[0050] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes is designed to hybridize at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of a genomic region identified in (i) Table 1, (ii) Table 3, or (iii) Table 3 as having multiple stepwise variants.
[0051] In some embodiments of any one of the methods disclosed herein, each nucleic acid probe in the set of nucleic acid probes has at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% sequence identity with respect to a probe sequence selected from Table 6.
[0052] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes comprises at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the probe sequences in Table 6.
[0053] In some embodiments of any one of the methods disclosed herein, the method determines whether a subject has a condition or the degree or state of the subject's condition based on one or more identified cell-free nucleic acid molecules comprising a plurality of stepwise variants. The method further includes determining the status. In some embodiments, the method further includes determining that one or more cell-free nucleic acid molecules originate from the sample associated with the status, based on performing a statistical model analysis of the identified one or more cell-free nucleic acid molecules. In some embodiments, the statistical model analysis includes Monte Carlo statistical analysis.
[0054] In some embodiments of any one of the methods disclosed herein, the method further includes monitoring the progression of the condition in question based on one or more identified cell-free nucleic acid molecules.
[0055] In some embodiments of any one of the methods disclosed herein, the method further includes performing different procedures to confirm the condition of the subject. In some embodiments, the different procedures include blood tests, genetic tests, medical imaging, physical examinations, or tissue biopsies.
[0056] In some embodiments of any one of the methods disclosed herein, the method further includes determining a treatment for a condition in question based on one or more identified cell-free nucleic acid molecules.
[0057] In some embodiments of any one of the methods disclosed herein, the subject is subjected to treatment of the condition prior to (a).
[0058] In some embodiments of any one of the methods disclosed herein, the treatment includes chemotherapy, radiotherapy, chemoradiotherapy, immunotherapy, adoptive cell therapy, hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusion, or medical surveillance.
[0059] In some embodiments of any one of the methods disclosed herein, the plurality of cell-free nucleic acid molecules comprises a plurality of cell-free deoxyribonucleic acid (DNA) molecules.
[0060] In some embodiments of any one of the methods disclosed herein, the condition includes a disease.
[0061] In some embodiments of any one of the methods disclosed herein, a plurality of cell-free nucleic acid molecules are derived from a body sample of interest. In some embodiments, the body sample includes plasma, serum, blood, cerebrospinal fluid, lymph, saliva, urine, or feces.
[0062] In some embodiments of any one of the methods disclosed herein, the subject is a mammal. In some embodiments of any one of the methods disclosed herein, the subject is a human.
[0063] In some embodiments of any one of the methods disclosed herein, the condition includes a neoplasm, cancer, or tumor. In some embodiments, the condition includes a solid tumor. In some embodiments, the condition includes a lymphoma. In some embodiments, the condition includes a B-cell lymphoma. In some embodiments, the condition includes a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt lymphoma, and B-cell chronic lymphocytic leukemia.
[0064] In some embodiments of any one of the methods disclosed herein, multiple stepwise variants have been previously identified as tumors derived from sequencing of a previous tumor sample or a cell-free nucleic acid sample.
[0065] In one embodiment, the present disclosure provides a composition comprising a bait set comprising a set of nucleic acid probes designed to capture cell-free DNA molecules derived from at least about 5% of the genomic regions specified in (i) the genomic regions identified in Table 1, (ii) the genomic regions identified in Table 3, or (iii) the genomic regions identified in Table 3 as having multiple stepwise variants.
[0066] In some embodiments of any of the compositions disclosed herein, the set of nucleic acid probes is designed to pull down cell-free DNA molecules from at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of the genomic regions specified in (i) the genomic regions identified in Table 1, (ii) the genomic regions identified in Table 3, or (iii) the genomic regions identified in Table 3 as having multiple stepwise variants.
[0067] In some embodiments of any of the compositions disclosed herein, the set of nucleic acid probes is designed to capture one or more cell-free DNA molecules derived from up to about 10%, up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 80%, up to about 90%, or up to about 100% of the genomic regions specified in (i) the genomic regions specified in Table 1, (ii) the genomic regions specified in Table 3, or (iii) the genomic regions specified in Table 3 to have multiple stepwise variants.
[0068] In some embodiments of any of the compositions disclosed herein, the bait set comprises up to 5, up to 10, up to 50, up to 100, up to 500, up to 1000, or up to 2000 nucleic acid probes.
[0069] In some embodiments of any of the compositions disclosed herein, each nucleic acid probe in a set of nucleic acid probes includes a pull-down tag.
[0070] In some embodiments of any of the compositions disclosed herein, the pull-down tag includes a nucleic acid barcode.
[0071] In some embodiments of any of the compositions disclosed herein, the pull-down tag contains biotin.
[0072] In some embodiments of any of the compositions disclosed herein, each cell-free DNA molecule is about 100 to about 180 nucleotides long.
[0073] In some embodiments of any of the compositions disclosed herein, the genomic region is state-related.
[0074] In some embodiments of any of the compositions disclosed herein, if the subject has a condition, the genomic region exhibits abnormal somatic hypermutation.
[0075] In some embodiments of any of the compositions disclosed herein, the condition comprises B-cell lymphoma. In some embodiments, the condition comprises a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt lymphoma, and B-cell chronic lymphocytic leukemia.
[0076] In some embodiments of any of the compositions disclosed herein, the composition further comprises a plurality of cell-free DNA molecules obtained from or derived from a subject.
[0077] In one embodiment, the Disclosure provides a method for performing a clinical procedure on an individual, a method for performing a clinical procedure on an individual, comprising: (a) obtaining, or having obtained, a target sequencing result of a cell-free nucleic acid molecule collection, wherein the cell-free nucleic acid molecule collection is supplied from a liquid or waste biopsy of the individual, and the target sequencing is performed using a nucleic acid probe to pull down sequences of genomic loci known to experience abnormal somatic hypermutation in B-cell cancer; (b) identifying, or having identified, multiple in-phase variants within the cell-free nucleic acid sequencing result; (c) determining, or having determined, that the cell-free nucleic acid sequencing result contains neoplasm-derived nucleotides, using a statistical model and the identified stepwise variants; and (d) performing a clinical procedure on the individual to confirm the presence of B-cell cancer, based on the determination that the cell-free nucleic acid sequencing result contains nucleic acid sequences that may originate from B-cell cancer.
[0078] In some embodiments of any of the compositions disclosed herein, the biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces.
[0079] In some embodiments of any of the compositions disclosed herein, the genomic locus is selected from (i) the genomic regions identified in Table 1, (ii) the genomic regions identified in Table 3, or (iii) the genomic regions identified in Table 3 as having multiple stepwise variants.
[0080] In some embodiments of any of the compositions disclosed herein, the nucleic acid probe sequence is selected from Table 6.
[0081] In some embodiments of any of the compositions disclosed herein, the clinical procedure is a blood test, medical imaging, or physical examination.
[0082] In one embodiment, the Disclosure provides a method for treating an individual for B-cell cancer, comprising: (a) obtaining, or having obtained, a targeted sequencing result of a cell-free nucleic acid molecule collection, wherein the cell-free nucleic acid molecule collection is supplied from the individual's liquid or waste biopsy, and the targeted sequencing is performed using a nucleic acid probe to pull down sequences of genomic loci known to experience abnormal somatic hypermutation in B-cell cancer; (b) identifying, or having identified, multiple in-phase variants within the cell-free nucleic acid sequencing result; (c) determining, or having determined, that the cell-free nucleic acid sequencing result contains neoplasm-derived nucleotides, using a statistical model and the identified stepwise variants; and (d) treating the individual to reduce the B-cell cancer based on the determination that the cell-free nucleic acid sequencing result contains nucleic acid sequences derived from B-cell cancer.
[0083] In some embodiments of any of the compositions disclosed herein, the biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces.
[0084] In some embodiments of any of the compositions disclosed herein, the genomic locus is selected from (i) the genomic regions identified in Table 1, (ii) the genomic regions identified in Table 3, or (iii) the genomic regions identified in Table 3 as having multiple stepwise variants.
[0085] In some embodiments of any of the compositions disclosed herein, the nucleic acid probe sequence is selected from Table 6.
[0086] In some embodiments of any of the compositions disclosed herein, the treatment is chemotherapy, radiotherapy, immunotherapy, hormone therapy, targeted drug therapy, or medical surveillance.
[0087] In one embodiment, the Disclosure provides a method for detecting minimal residual cancer in an individual and treating the individual for cancer, comprising: (a) obtaining, or having obtained, a cell-free nucleic acid molecule collection, the cell-free nucleic acid molecule collection being supplied from a liquid or waste biopsy of the individual, the liquid or waste biopsy being supplied after a series of treatments to detect minimal residual disease, and the targeted sequencing being performed using a nucleic acid probe to pull down the sequence of a genomic locus determined to contain multiple variants in the same phase, as determined by previous sequencing results for a previous biopsy derived from cancer; (b) identifying, or having identified, at least one set of multiple variants in the same phase within the cell-free nucleic acid sequencing result; and (c) treating the individual to reduce the cancer based on the determination that the cell-free nucleic acid sequencing result contains nucleic acid sequences derived from cancer.
[0088] In some embodiments of any of the compositions disclosed herein, the liquid or waste biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces.
[0089] In some embodiments of any of the compositions disclosed herein, the treatment is chemotherapy, radiotherapy, immunotherapy, hormone therapy, targeted drug therapy, or medical surveillance.
[0090] In one embodiment, the Disclosure provides a computer program product comprising a non-temporary computer-readable medium having computer executable code encoded therein, wherein the computer executable code is adapted to be executed in any one of the methods disclosed herein.
[0091] In one embodiment, the Disclosure provides a system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, when executed by the one or more computer processors, implements any one of the methods disclosed herein.
[0092] Further aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, which shows and describes only exemplary embodiments of the present disclosure. As will be understood, other different embodiments of the present disclosure are possible, and some of their details can be modified in various obvious ways without departing from the present disclosure. Accordingly, the drawings and description should be considered illustrative and not limiting in nature. Embedding by reference
[0093] All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as each individual publication, patent, and patent application is specifically and individually incorporated by reference. To the extent that any publications and patents or patent applications incorporated by reference conflict with any disclosures contained herein, this specification is intended to supersede and / or take precedence over such conflicting material. Brief explanation of the drawing
[0094] Various features of the present invention are described in detail in the appended claims. A better understanding of the features and advantages of the present invention can be obtained by referring to the following detailed description, which describes exemplary embodiments in which the principles of the present invention are utilized, and to the appended drawings (also referred to herein as "Figure" and "FIG."). In certain embodiments, for example, the following items are provided: (Item 1) (a) Obtaining sequencing data from multiple cell-free nucleic acid molecules obtained from or derived from the subject using a computer system, (b) The computer system processes the sequencing data to identify one or more of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises a plurality of stepwise variants with respect to a reference genome sequence, and at least about 10% of the one or more cell-free nucleic acid molecules comprises a first stepwise variant of the plurality of stepwise variants and a second stepwise variant of the plurality of stepwise variants separated by at least one nucleotide. (c) A method comprising analyzing the identified one or more cell-free nucleic acid molecules using the computer system to determine the state of the subject. (Item 2) The method according to item 1, wherein at least about 10% of the cell-free nucleic acid molecules comprises at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of the 1 or more cell-free nucleic acid molecules. (Item 3) (a) Obtaining sequencing data from multiple cell-free nucleic acid molecules obtained from or derived from the subject using a computer system, (b) Processing the sequencing data using the computer system to identify one or more of the cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules includes a plurality of stepwise variants to a reference genome sequence separated by at least one nucleotide. (c) A method comprising analyzing the identified one or more cell-free nucleic acid molecules using the computer system to determine the state of the subject. (Item 4) (a) Obtain sequencing data from multiple cell-free nucleic acid molecules obtained from or derived from the subject, (b) Processing the sequencing data to identify one or more cell-free nucleic acid molecules among the plurality of cell-free nucleic acid molecules with a detection limit of less than approximately one of 50,000 observations from the sequencing data, (c) A method comprising analyzing one or more identified cell-free nucleic acid molecules to determine the state of the subject. (Item 5) The method according to item 4, wherein the detection limit of the identification step is less than approximately 1 in 100,000, less than approximately 1 in 500,000, less than approximately 1 in 1,000,000, less than approximately 1 in 1,500,000, or less than approximately 1 in 2,000,000, based on observations from the sequencing data. (Item 6) The method according to any one of items 4 to 5, wherein each of the one or more cell-free nucleic acid molecules comprises multiple stepwise variants with respect to a reference genome sequence. (Item 7) The method according to item 6, wherein the first step variant of the plurality of step variants and the second step variant of the plurality of step variants are separated by at least one nucleotide. (Item 8) (a) to (c) are performed by a computer system, as described in any one of items 4 to 7. (Item 9) The method according to any one of the preceding items, wherein the sequencing data is generated based on nucleic acid amplification. (Item 10) The method according to any one of the preceding items, wherein the sequencing data is generated based on a polymerase chain reaction. (Item 11) The method according to any one of the preceding items, wherein the sequencing data is generated based on amplicon sequencing. (Item 12) The method according to any one of the preceding items, wherein the sequencing data is generated based on next-generation sequencing (NGS). (Item 13) The method according to any one of the preceding items, wherein the sequencing data is generated based on non-hybridization-based NGS. (Item 14) The method according to any one of the preceding items, wherein the sequencing data is generated without using molecular barcoding of at least some of the plurality of cell-free nucleic acid molecules. (Item 15) The method according to any one of the preceding items, wherein the sequencing data is obtained without using sample barcoding of at least some of the plurality of cell-free nucleic acid molecules. (Item 16) The method according to any one of the preceding items, wherein the sequencing data is obtained without (i) background errors or (ii) in silico removal or suppression of sequencing errors. (Item 17) A method for treating the condition of the subject, (a) Identifying the subject for treatment of the condition, wherein the subject is determined to have the condition based on the identification of one or more cell-free nucleic acid molecules from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject, Each of the one or more identified cell-free nucleic acid molecules comprises multiple stepwise variants of a reference genome sequence separated by at least one nucleotide. The existence of the aforementioned multiple stepwise variants indicates the state of the subject, and identifies it. (b) A method comprising applying the treatment based on the identification in (a) to the subject. (Item 18) A method for monitoring the progression of the target state, (a) Determining a first state of the state of the subject based on the identification of a first set of one or more cell-free nucleic acid molecules from a first plurality of cell-free nucleic acid molecules obtained from or derived from the subject, (b) Determining a second state of the state of the subject based on the identification of a second set of one or more cell-free nucleic acid molecules from a second plurality of cell-free nucleic acid molecules obtained from or derived from the subject, The determination is that the second plurality of cell-free nucleic acid molecules are obtained from the subject after the first plurality of cell-free nucleic acid molecules are obtained from the subject, (c) Determining the progression of the state based on the first condition and the second condition of the state, A method comprising determining that each of the one or more cell-free nucleic acid molecules comprises a plurality of stepwise variants to a reference genome sequence separated by at least one nucleotide. (Item 19) The method according to item 18, wherein the progression of the aforementioned state is a deterioration of the aforementioned state. (Item 20) The method according to item 18, wherein the progression of the aforementioned condition is at least partial remission of the aforementioned condition. (Item 21) The method according to any one of items 18 to 20, wherein the presence of the plurality of stepwise variants indicates a first or second state of the state of the subject. (Item 22) The method according to any one of items 18 to 21, wherein the second plurality of cell-free nucleic acid molecules are obtained from the subject at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 4 weeks, at least about 2 months, or at least about 3 months after obtaining the first plurality of cell-free nucleic acid molecules from the subject. (Item 23) The method according to any one of items 18 to 22, wherein the subject is subjected to treatment for the condition (i) before obtaining the second plurality of cell-free nucleic acid molecules from the subject, and (ii) after obtaining the first plurality of cell-free nucleic acid molecules from the subject. (Item 24) The method according to any one of items 18 to 23, wherein the progression of the aforementioned state indicates the minimum residual disease of the aforementioned state in the subject. (Item 25) The method according to any one of items 18 to 24, wherein the progression of the aforementioned condition indicates a tumor burden or cancer burden on the subject. (Item 26) The method according to any one of the preceding items, wherein one or more cell-free nucleic acid molecules are captured from among the plurality of cell-free nucleic acid molecules by a set of nucleic acid probes, and the set of nucleic acid probes is configured to hybridize to at least a portion of the cell-free nucleic acid molecules comprising one or more genomic regions associated with the state. (Item 27) (a) To provide a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained from or derived from a subject, Each nucleic acid probe in the set of nucleic acid probes is designed to hybridize to at least a portion of a target cell-free nucleic acid molecule that includes multiple stepwise variants of a reference genome sequence separated by at least one nucleotide. To provide a solution in which each nucleic acid probe comprises an activatable reporter agent, and the activation of the activatable reporter agent is selected from the group consisting of (i) hybridization of each nucleic acid probe into the plurality of stepwise variants, and (ii) dehybridization of at least a portion of the individual nucleic acid probes hybridized into the plurality of stepwise variants. (b) detecting the activated activatable reporter agent to identify one or more of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules includes the plurality of stepwise variants. (c) A method comprising analyzing one or more identified cell-free nucleic acid molecules to determine the state of the subject. (Item 28) (a) To provide a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained from or derived from a subject, Each nucleic acid probe in the set of nucleic acid probes is designed to hybridize to at least a portion of a target cell-free nucleic acid molecule containing multiple stepwise variants relative to a reference genome sequence. To provide a solution in which each nucleic acid probe comprises an activatable reporter agent, and the activation of the activatable reporter agent is selected from the group consisting of (i) hybridization of the individual nucleic acid probe into the plurality of stepwise variants, and (ii) dehybridization of at least a portion of the individual nucleic acid probe hybridized into the plurality of stepwise variants. (b) Detecting the activated activatable reporter agent to identify one or more of the plurality of cell-free nucleic acid molecules, wherein each of the one or more cell-free nucleic acid molecules comprises the plurality of stepwise variants, and the detection limit of the identification step is less than approximately one of the 50,000 cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules. (c) A method comprising analyzing one or more identified cell-free nucleic acid molecules to determine the state of the subject. (Item 29) The method according to item 28, wherein the detection limit of the identification step is less than about 1 out of 100,000, less than 1 out of 500,000, less than 1 out of 1,000,000, less than 1 out of 1,500,000, or less than 1 out of 2,000,000 of the plurality of cell-free nucleic acid molecules. (Item 30) The method according to item 28 or 29, wherein the first stepwise variant of the plurality of stepwise variants and the second stepwise variant of the plurality of stepwise variants are separated by at least one nucleotide. (Item 31) The method according to any one of items 27 to 30, wherein the activatable reporter agent is activated during the hybridization of the individual nucleic acid probes into the plurality of stepwise variants. (Item 32) The method according to any one of items 27 to 30, wherein the activatable reporter agent is activated upon dehybridization of at least some of the individual nucleic acid probes hybridized to the plurality of stepwise variants. (Item 33) The method according to any one of items 27 to 32, further comprising (1) a set of nucleic acid probes and (2) mixing the plurality of cell-free nucleic acid molecules. (Item 34) The method according to any one of items 27 to 33, wherein the activatable reporter agent is a fluorophore. (Item 35) The method according to any one of the preceding items, wherein analyzing the identified one or more cell-free nucleic acid molecules comprises (i) the identified one or more cell-free nucleic acid molecules and (ii) other cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules that do not include the plurality of stepwise variants, as different variables. (Item 36) The method according to any one of the preceding items, wherein the analysis of the identified one or more cell-free nucleic acid molecules is not based on other cell-free nucleic acid molecules of the plurality of cell-free nucleic acid molecules that do not contain the plurality of stepwise variants. (Item 37) The method according to any one of the preceding items, wherein the number of the multiple stepwise variants from the identified one or more cell-free nucleic acid molecules indicates the state of the subject. (Item 38) The method according to item 37, wherein (i) the number of the plurality of stepwise variants from the one or more cell-free nucleic acid molecules and (ii) the number of single nucleotide variants (SNVs) from the one or more cell-free nucleic acid molecules indicate the state of the subject. (Item 39) The method according to any one of the preceding items, wherein the frequency of the multiple stepwise variants in the identified one or more cell-free nucleic acid molecules indicates the state of the subject. (Item 40) The method according to item 39, wherein the frequency indicates disease cells associated with the condition. (Item 41) The method according to item 40, wherein the condition is diffuse large B-cell lymphoma, and the frequency indicates whether the 1 or more cell-free nucleic acid molecules originate from germinal center B cells (GCB) or activated B cells (ABC). (Item 42) The method according to any one of the preceding items, wherein the genomic origin of the identified one or more cell-free nucleic acid molecules indicates the state of the subject. (Item 43) The method according to any one of the preceding items, wherein the first stepwise variant and the second stepwise variant are separated by at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 nucleotides. (Item 44) The method according to any one of the preceding items, wherein the first stepwise variant and the second stepwise variant are separated by a maximum of about 180 nucleotides, a maximum of about 170 nucleotides, a maximum of about 160 nucleotides, a maximum of about 150 nucleotides, or a maximum of about 140 nucleotides. (Item 45) The method according to any one of the preceding items, wherein at least about 10%, at least about 20%, at least about 30%, at least about 40%, or at least about 50% of the one or more cell-free nucleic acid molecules containing multiple stepwise variants contain a single nucleotide variant (SNV) located at least 2 nucleotides away from an adjacent SNV. (Item 46) The method according to any one of the preceding items, wherein the plurality of stepwise variants include at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, or at least 25 stepwise variants within the same cell-free nucleic acid molecule. (Item 47) The method according to any one of the preceding items, wherein the identified 1 or more cell-free nucleic acid molecules comprise at least 2, at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 500, or at least 1,000 cell-free nucleic acid molecules. (Item 48) The method according to any one of the preceding items, wherein the aforementioned reference genome sequence is derived from a reference cohort. (Item 49) The method according to item 48, wherein the reference genome sequence includes a consensus sequence from the reference cohort. (Item 50) The method according to item 48, wherein the reference genome sequence includes at least a portion of the HG19 human genome, the HG18 genome, the HG17 genome, the HG16 genome, or the HG38 genome. (Item 51) The method according to any one of the preceding items, wherein the reference genome sequence is derived from the sample of interest. (Item 52) The method described in item 51, wherein the sample is a health sample. (Item 53) The method described in item 52, wherein the sample contains healthy cells. (Item 54) The method according to item 53, wherein the healthy cells include healthy white blood cells. (Item 55) The method described in item 51, wherein the sample is a disease sample. (Item 56) The method according to item 55, wherein the disease sample contains disease cells. (Item 57) The method according to item 56, wherein the diseased cells include tumor cells. (Item 58) The method according to item 55, wherein the disease sample includes a solid tumor. (Item 59) The method according to any one of the preceding items, wherein the set of nucleic acid probes is designed based on a plurality of stepwise variants identified by comparing (i) sequencing data from the subject solid tumor, lymphoma, or hematological malignancy with (ii) sequencing data from healthy cells of the subject or a healthy cohort. (Item 60) The method according to item 59, wherein the healthy cells are from the subject. (Item 61) The method according to item 59, wherein the healthy cells are from the healthy cohort. (Item 62) The method according to any one of the preceding items, wherein the set of nucleic acid probes is designed to hybridize to at least a portion of the sequences of genomic loci associated with the state. (Item 63) The method according to item 62, wherein the genomic locus associated with the aforementioned condition is known to exhibit abnormal somatic hypermutation when the subject has the aforementioned condition. (Item 64) The method according to any one of the preceding items, wherein the set of nucleic acid probes is designed to hybridize at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of the genomic regions identified in (i) Table 1, (ii) Table 3, or (iii) Table 3 as having multiple stepwise variants. (Item 65) The method according to any one of the preceding items, wherein each nucleic acid probe in the set of nucleic acid probes has at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% sequence identity with respect to a probe sequence selected from Table 6. (Item 66) The method according to any one of the preceding items, wherein the set of nucleic acid probes comprises at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the probe sequences in Table 6. (Item 67) The method of any one of the preceding items, further comprising determining that the subject has the state, or determining the degree or state of the state of the subject, based on the identified one or more cell-free nucleic acid molecules comprising the plurality of stepwise variants. (Item 68) The method of item 67, further comprising determining, based on performing a statistical model analysis of the one or more identified cell-free nucleic acid molecules, that the one or more cell-free nucleic acid molecules originate from a sample associated with the condition. (Item 69) The method described in item 68, wherein the statistical model analysis includes Monte Carlo statistical analysis. (Item 70) The method according to any one of the preceding items, further comprising monitoring the progression of the state of the subject based on the identified one or more cell-free nucleic acid molecules. (Item 71) The method of any one of the preceding items, further comprising performing different steps to confirm the state of the subject. (Item 72) The method according to item 71, wherein the aforementioned different procedures include blood tests, genetic tests, medical imaging, physical examinations, or tissue biopsies. (Item 73) The method according to any one of the preceding items, further comprising determining a course of action for the condition of the subject based on the identified one or more cell-free nucleic acid molecules. (Item 74) The method of any one of the preceding items, wherein the subject is subjected to treatment for the condition before (a). (Item 75) The method according to any one of the preceding items, wherein the procedure includes chemotherapy, radiotherapy, chemoradiotherapy, immunotherapy, adoptive cell therapy, hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusion, or medical surveillance. (Item 76) The method according to any one of the preceding items, wherein the plurality of cell-free nucleic acid molecules include a plurality of cell-free deoxyribonucleic acid (DNA) molecules. (Item 77) The method described in any one of the preceding items, wherein the aforementioned condition includes a disease. (Item 78) The method according to any one of the preceding items, wherein the plurality of cell-free nucleic acid molecules are derived from the subject body sample. (Item 79) The method according to item 78, wherein the body sample includes plasma, serum, blood, cerebrospinal fluid, lymph, saliva, urine, or feces. (Item 80) The method according to any one of the preceding items, wherein the subject is a mammal. (Item 81) The method described in any one of the preceding items, wherein the subject is a human. (Item 82) The method described in any one of the preceding items, wherein the condition described above includes a neoplasm, cancer, or tumor. (Item 83) The method according to item 82, wherein the aforementioned condition includes a solid tumor. (Item 84) The method described in item 82, wherein the aforementioned condition includes lymphoma. (Item 85) The method described in item 84, wherein the condition includes B-cell lymphoma. (Item 86) The method according to item 85, wherein the condition includes a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt lymphoma, and B-cell chronic lymphocytic leukemia. (Item 87) The method according to any one of the preceding items, wherein the aforementioned stepwise variants have been previously identified as tumors derived from sequencing of a previous tumor sample or a cell-free nucleic acid sample. (Item 88) A composition comprising a bait set containing a set of nucleic acid probes designed to capture cell-free DNA molecules from at least about 5% of the genomic regions specified in (i) the genomic regions identified in Table 1, (ii) the genomic regions identified in Table 3, or (iii) the genomic regions identified in Table 3 as having multiple stepwise variants. (Item 89) The composition according to item 88, wherein the set of nucleic acid probes is designed to pull down cell-free DNA molecules derived from at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% of the genomic regions specified in (i) the genomic regions specified in Table 1, (ii) the genomic regions specified in Table 3, or (iii) the genomic regions specified in Table 3 to have multiple stepwise variants. (Item 90) The composition according to any one of items 88 to 89, wherein the set of nucleic acid probes is designed to capture one or more cell-free DNA molecules derived from up to about 10%, up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 80%, up to about 90%, or up to about 100% of the genomic regions specified in (i) the genomic regions specified in Table 1, (ii) the genomic regions specified in Table 3, or (iii) the genomic regions specified in Table 3 to have multiple stepwise variants. (Item 91) The composition according to any one of items 88 to 90, wherein the bait set comprises up to 5, up to 10, up to 50, up to 100, up to 500, up to 1000, or up to 2000 nucleic acid probes. (Item 92) The composition according to any one of items 88 to 91, wherein each nucleic acid probe in the set of nucleic acid probes includes a pull-down tag. (Item 93) The composition according to any one of items 88 to 92, wherein the pull-down tag includes a nucleic acid barcode. (Item 94) The aforementioned pull-down tag is a composition according to any one of items 88 to 93, comprising biotin. (Item 95) The composition according to any one of items 88 to 94, wherein each of the cell-free DNA molecules is about 100 nucleotides to about 180 nucleotides in length. (Item 96) The composition according to any one of items 88 to 95, wherein the genomic region is related to the state. (Item 97) The composition according to any one of items 88 to 96, wherein, when the subject has the aforementioned condition, the genomic region exhibits abnormal somatic hypermutation. (Item 98) The composition according to any one of items 88 to 97, wherein the condition described above includes B-cell lymphoma. (Item 99) The composition according to item 98, wherein the condition includes a subtype of B-cell lymphoma selected from the group consisting of diffuse large B-cell lymphoma, follicular lymphoma, Burkitt lymphoma, and B-cell chronic lymphocytic leukemia. (Item 100) A composition according to any one of items 88 to 99, further comprising a plurality of cell-free DNA molecules obtained from or derived from the subject. (Item 101) A method for performing clinical procedures on an individual, Obtaining or having obtained targeted sequencing results for a collection of cell-free nucleic acid molecules, The cell-free nucleic acid molecule collection is supplied from a solid liquid or waste biopsy, and The aforementioned targeted sequencing is performed using nucleic acid probes to pull down sequences of genomic loci known to experience abnormal somatic hypermutation in B-cell cancer, or to obtain, or to obtain. Within the cell-free nucleic acid sequencing results, identifying multiple variants in the same phase, or having identified them, Using a statistical model and the identified stepwise variants, it is determined that the cell-free nucleic acid sequencing result contains nucleotides derived from the neoplasm, or that it has been determined that, A method comprising: determining that the cell-free nucleic acid sequencing result contains a nucleic acid sequence that may originate from the B-cell cancer; and performing a clinical procedure on the individual to confirm the presence of the B-cell cancer. (Item 102) The method according to item 101, wherein the biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces. (Item 103) The method according to item 101, wherein the genomic locus is selected from (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple stepwise variants. (Item 104) The method according to item 101, wherein the sequence of the nucleic acid probe is selected from Table 6. (Item 105) The method according to item 101, wherein the clinical procedure is a blood test, medical imaging, or physical examination. (Item 106) A method for treating an individual with B-cell carcinoma, Obtaining or having obtained targeted sequencing results for a collection of cell-free nucleic acid molecules, The cell-free nucleic acid molecule collection is supplied from a solid liquid or waste biopsy, and The aforementioned targeted sequencing is performed using nucleic acid probes to pull down sequences of genomic loci known to experience abnormal somatic hypermutation in B-cell cancer, or to obtain, or to obtain. Within the cell-free nucleic acid sequencing results, identifying multiple variants in the same phase, or having identified them, Using a statistical model and the identified stepwise variants, it is determined that the cell-free nucleic acid sequencing result contains nucleotides derived from the neoplasm, or that it has been determined that, A method comprising treating the individual to reduce the B-cell cancer based on the determination that the cell-free nucleic acid sequencing result contains a nucleic acid sequence derived from the B-cell cancer. (Item 107) The method according to item 106, wherein the biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces. (Item 108) The method according to item 106, wherein the genomic locus is selected from (i) a genomic region identified in Table 1, (ii) a genomic region identified in Table 3, or (iii) a genomic region identified in Table 3 as having multiple stepwise variants. (Item 109) The method according to item 106, wherein the sequence of the nucleic acid probe is selected from Table 6. (Item 110) The method according to item 106, wherein the treatment is chemotherapy, radiotherapy, immunotherapy, hormone therapy, targeted drug therapy, or medical monitoring. (Item 111) A method for detecting minimal residual cancer in an individual and treating the individual for cancer, Obtaining or having obtained targeted sequencing results for a collection of cell-free nucleic acid molecules, The cell-free nucleic acid molecule collection is supplied from a solid liquid or waste biopsy. The aforementioned liquid or waste biopsy is supplied after a series of procedures to detect minimal residual disease, and The targeted sequencing is performed using nucleic acid probes to pull down the sequences of genomic loci determined to contain multiple variants in the same phase, as determined by previous sequencing results for a previous biopsy derived from the cancer, or to obtain, Identifying, or having identified, at least one set of the plurality of variants in the same phase within the cell-free nucleic acid sequencing results. A method comprising: treating the individual to reduce the tumor based on the determination that the cell-free nucleic acid sequencing result contains a nucleic acid sequence derived from the tumor. (Item 112) The method according to item 111, wherein the liquid or waste biopsy is one of blood, serum, cerebrospinal fluid, lymph, urine, or feces. (Item 113) The method according to item 111, wherein the treatment is chemotherapy, radiotherapy, immunotherapy, hormone therapy, targeted drug therapy, or medical surveillance. (Item 114) A computer program product comprising a non-temporary computer-readable medium having computer-executable code encoded therein, wherein the computer-executable code is adapted and executed to implement the method described in any one of the preceding items. (Item 115) A system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, when executed by the one or more computer processors, implements the method described in any one of the preceding items. [Brief explanation of the drawing]
[0095] [Figure 1-1]Figure 1 illustrates the discovery of stepwise variants and their mutational signatures through the analysis of whole-genome sequencing data. Figure 1A is a schematic diagram showing the difference in detection between single nucleotide variants (SNVs) on individual cell-free DNA molecules (top) and multiple "in-phase" variants (stepwise variants, PVs; bottom). Theoretically, the detection of PVs is a more specific event than the detection of a single SNV. Figure 1B is a scatter plot showing the distribution of PV counts from WGS data for 24 different histological features of cancer, normalized by the total number of SNVs. The bars represent the median and interquartile range. (FL-NHL, follicular lymphoma; DLBCL-NHL, diffuse large B-cell lymphoma; Burkitt-NHL, Burkitt lymphoma; Lung-SCC, squamous cell lung cancer; Lung-Adeno, lung adenocarcinoma; Kidney-RCC, renal cell carcinoma; Bone-Os teosarc, osteosarcoma; Liver-HCC, hepatocellular carcinoma; Breast-Adeno, pancreatic adenocarcinoma; Head-SCC, head and neck squamous cell carcinoma; Ovary-Adeno, ovarian adenocarcinoma; Eso-Adeno, esophageal adenocarcinoma; Uterus-Ad Figure 1C is a heatmap showing the enrichment of PV single nucleotide substitution (SBS) mutation signatures for single SNVs across multiple cancer types. Blue represents signatures where PV is enriched in a specific histological picture. Darker gray represents signatures where a single SNV is enriched without phases. Red represents SNVs occurring alone. Only signatures with a significant difference between PV and non-phased SNVs after adjusting for multiple hypotheses are shown. Other signatures are gray. Signatures related to smoking, AID / AICDA, and APOBEC are shown.Figure 1D demonstrates a bar plot showing the distribution of PVs occurring in genomically typological regions in B lymphoid malignancies and lung adenocarcinoma. In this plot, the genome was divided into 1000 bp bins, and the percentage of samples of a given histological image having PVs in each 1000 bp bin was calculated. Only bins with a recurrence frequency of at least 2% in any cancer subtype are shown. Important genomic loci are also labeled. Figure 1E is a comparison of double-strand sequencing and stepwise variant sequencing. A schematic diagram comparing error-suppressed sequencing by double-strand sequencing versus recovery of stepwise variants. Double-strand sequencing requires the recovery of a single SNV observed on both strands of the original DNA double helix (i.e., in trans). This requires the independent recovery of two molecules by sequencing, as the positive and negative strands of the original DNA molecule pass through library preparation and PCR independently. In contrast, PV recovery requires multiple SNVs observed on the same single strand of DNA (i.e., in cis). Therefore, recovery of only the positive or negative chain (but not both) is sufficient for identifying PV. [Figure 1-2] Same as above. [Figure 1-3] Same as above.
[0096] [Figure 2-1]Figure 2 illustrates the design, validation, and application of stepwise variant enrichment sequencing. Figure 2A is a schematic diagram of the design for PhasED-Seq. WGS data from DLBCL tumor samples were aggregated (left) to identify regions of estimated recurrent PVs (center). An assay was then designed to capture the most recurrently PV-containing genomic regions (right), resulting in approximately 7500-fold enrichment of PVs compared to WGS. The upper right panel shows the estimated in silico number of PVs per case per kilobase for increasing the panel size (x axis). The dashed line indicates selected regions in the PhasED-Seq panel. The lower right panel shows the total estimated number of PVs per case (y axis, assessed in silico from WGS data) for increasing the panel size (y axis). The dark areas indicate selected regions in the PhasED-Seq panel. Figure 2B illustrates two panels showing the yield of SNVs (left) and PVs (right) for sequencing tumor DNA and matched germline using previously established lymphoma CAPP-Seq or PhaseD-Seq panels. Values are assessed in silico by limiting WGS to the target space of interest. PVs reported in the right-hand panel include doublet, triplet, and quadruplet phased events. Figure 2C, similar to Figure 2B, shows the yield of SNVs (left) and PVs (right) from experimental sequencing of tumor and / or cell-free DNA from CAPP-Seq versus PhaseD-Seq. Figure 2D is a scatter plot showing the frequency of PVs by genomic location (in a 1000 bp bin) for patients with DLBCL identified by WGS or PhaseD-Seq. The PVs of IGH, BCL2, MYC, and BCL6 are highlighted. Figure 2E illustrates a scatter plot comparing the frequency of PVs by genomic location (in a 50 bp bin) for patients with different types of lymphoma. The colored circles show the relative frequency of PVs in a 50 bp bin from the specific gene of interest. The other (gray) circles show the relative frequency of PVs in a 50 bp bin from the rest of the PhasED-Seq sequencing panel.Figure 2F illustrates volcano plots summarizing the differences in relative PV frequencies at specific loci between lymphoma types, including ABC-DLBCL vs. GCB-DLBCL (dark gray, left), PMBCL vs. DLBCL (dark gray, center), and HL vs. DLBCL (dark gray, right). The x-axis shows the relative enrichment of PV at a specific locus, and the y-axis shows the statistical significance of this association. (Example 10). [Figure 2-2] Same as above. [Figure 2-3] Same as above. [Figure 2-4] Same as above.
[0097] [Figure 3-1]Figure 3 illustrates the technical performance of PhasED-Seq for disease detection. Figure 3A illustrates a bar plot showing the performance of hybrid capture sequencing for the recovery of synthetic 150 bp oligonucleotides from two loci (MYC and BCL6) with increased levels of mutated / non-reference bases. Error bars represent 95% confidence intervals (n = 3 replicates for each condition in separate samples). Figure 3B illustrates a plot revealing the background error rates (Example 10) for different types of error suppression from 12 healthy control cell-free DNA samples sequenced with a PhasED-Seq panel. "PhasED-Seq 2x" or "doublet" represents the detection of two homeophase mutations on the same DNA molecule. "PhasED-Seq 3x" or "triplet" represents the detection of three homeophase mutations on the same DNA molecule. Figure 3C illustrates a bar plot showing the depth of intrinsic molecule recovery (e.g., depth after barcode-mediated PCR replica removal) from sequencing data from 12 cell-free DNA samples for different types of error suppression, including barcode deduplication, double-strand sequencing, and PV recovery where the maximum distance between in-phase SNVs increases. Figure 3D illustrates a bar plot showing the cumulative percentage of PVs with a maximum distance between SNVs less than the number of base pairs shown on the x-axis. Figure 3E illustrates a plot showing the results of a limiting dilution series simulating cell-free DNA samples containing patient-specific tumor percentages of 1 × 10⁻³ to 0.5 × 10⁻⁶. cfDNA from three independent patient samples was used for each dilution. The same sequencing data was analyzed using various error suppression methods for the recovery of expected tumor percentages, including iDES, double-strand sequencing, and PhasED-Seq (both for the recovery of doublet and triplet molecules). Points and error bars represent the mean, minimum, and maximum across the three patient-specific tumor mutations considered. The difference between the observed tumor rate and the predicted tumor rate for a sample <1:10,000 was compared using a paired t-test. *, P<0.05; **, P<0.005; ***, P<0.0005.Figure 3F illustrates plots showing background signals for the detection of tumor-specific alleles in 12 unrelated healthy cell-free DNA samples and healthy cfDNA samples used in the limiting dilution series (total n=13 samples). For each sample, tumor-specific SNVs or PVs from three patient samples used in the limiting dilution experiment shown in Figure 3E were evaluated for a total of 39 assessments. The bars represent the arithmetic mean across all 39 assessments. Statistical comparisons were performed by the Wilcoxon rank-sum test. *, P<0.05, **, P<0.005, ***, P<0.0005. Figure 3G illustrates plots showing the theoretical detection rate of samples with a given number of PV-containing regions by simple binomial sampling. This plot is constructed by assuming an intrinsic sequencing depth of 5000x (line) along with a variety of independent 150 bp PV-containing regions ranging from 3 regions (blue) to 67 regions (purple). The confidence envelope considers a depth of 4000–6000x. A false positive rate of 5% is also assumed. Figure 3H illustrates a plot showing the observed detection rate (y-axis) for a given true tumor percentage (x-axis) with varying numbers of PV-containing regions. For each number of tumor reporter regions ranging from 3 to 67, a 150 bp window of this number was randomly sampled 25 times from each of three patient-specific PV reporter lists and used to evaluate tumor detection in each dilution. Filled points represent "wet" dilution series experiments, and open points represent in silico dilution experiments. Points and error bars represent the mean, minimum, and maximum of the three patient-specific PV reporter lists used in the original sampling. Figure 3I illustrates a scatter plot comparing the predicted and observed detection rates for samples from the dilution series shown in the panels of Figures 3G and 3H. Further details of this experiment are provided in Example 10. [Figure 3-2] Same as above. [Figure 3-3] Same as above.
[0098] [Figure 4-1]Figure 4 illustrates the clinical application of PhasED-Seq for ultra-sensitive disease detection and response monitoring in DLBCL. Figure 4A illustrates a plot showing ctDNA levels in DLBCL patients who responded to first-line immunochemotherapy and subsequently relapsed. Levels measured by CAPP-Seq are shown by darker gray circles, while levels measured by PhasED-Seq are shown by lighter gray circles. White circles represent levels undetectable by CAPP-Seq. Figure 4B illustrates a univariate scatter plot showing the mean tumor allele percentage measured by PhasED-Seq for clinical samples at the time of minimal disease (i.e., after 1 or 2 cycles of treatment). The plot is separated into samples detected and undetected by standard CAPP-Seq. P-values from the Wilcoxon rank-sum test. Figure 4C illustrates a bar plot showing the percentage of DLBCL patients with detectable ctDNA by CAPP-Seq after one or two cycles of treatment (dark gray bars), and similarly, the percentage of further patients with detectable disease when PhasED-Seq is added to standard CAPP-Seq (medium gray bars). The P-value represents Fisher's direct test for detection by CAPP-Seq alone compared to the combination of PhasED-Seq and CAPP-Seq in 171 samples after one or two cycles of treatment. Figure 4D illustrates a waterfall plot showing the change in ctDNA levels measured by CAPP-Seq after two cycles of first-line treatment in DLBCL patients. Patients with ctDNA undetectable by CAPP-Seq are shown in darker colors as "ND" ("not detected"). The bar color also indicates the final clinical outcome for these patients. Figure 4E illustrates a Kaplan-Meier plot showing relapse-free survival for 52 DLBCL patients with undetectable ctDNA measured by CAPP-Seq after two cycles.Figure 4F illustrates a Kaplan-Meier plot showing relapse-free survival for 52 patients (with ctDNA undetectable by CAPP-Seq) shown in Figure 4E, stratified by ctDNA detection by PhasED-Seq at the same point in time (Cycle 3, Day 1). Figure 4G illustrates a Kaplan-Meier plot showing relapse-free survival for 89 DLBCL patients stratified by ctDNA in Cycle 3, with Day 1 classified into three strata: patients who did not achieve a major molecular response (dark gray), patients with a major molecular response and ctDNA still detectable by PhasED-Seq and / or CAPP-Seq (light gray), and patients with stringent molecular remission (ctDNA undetectable by both PhasED-Seq and CAPP-Seq; medium gray). [Figure 4-2] Same as above.
[0099] [Figure 5]Figure 5 illustrates an enumeration of SNVs and PVs in various cancers from WGS. Figures 5A–5C illustrate univariate scatter plots showing the number of SNVs (Figure 5A), PVs (Figure 5B), and PVs controlling the total number of SNVs (Figure 5C) from WGS data for 24 different histological cancers. The bars indicate the median and interquartile range. (FL-NHL, follicular lymphoma; DLBCL-NHL, diffuse large B-cell lymphoma; Burkitt-NHL, Burkitt lymphoma; Lung-SCC, squamous cell lung cancer; Lung-Adeno, lung adenocarcinoma; Kidney-RCC, renal cell carcinoma; Bone-Osteosarc, osteosarcoma; Liver-HCC, hepatocellular carcinoma; Breast-Adeno, mammary adenocarcinoma; Panc-Adeno, pancreatic adenocarcinoma; Head-SCC, head and neck squamous cell carcinoma; Ovary-A deno, ovarian adenocarcinoma; Eso-Adeno, esophageal adenocarcinoma; Uterus-Adeno, Stomach-Adeno, gastric adenocarcinoma; CLL, chronic lymphocytic leukemia; ColoRect-Adeno, colorectal adenocarcinoma; Prost-Ade no, prostate adenocarcinoma; CNS-GBM, glioblastoma multiforme; Panc-Endocrine, pancreatic neuroendocrine tumor; Thy-Adeno, thyroid adenocarcinoma; CNS-PiloAstro, pilocytic astrocytoma; CNS-Medullo, medulloblastoma).
[0100] [Figure 6-1] Figure 6 illustrates the contribution of mutation signatures to stepped and non-stepped SNVs in WGS (Figures 6A–6WW). Scatter plots show the contribution of established single nucleotide substitution (SBS) mutation signatures to SNVs found in PV (shown in dark colors) and SNVs found outside possible stepping relationships from WGS (shown in light colors). This is presented for 49 SBS mutation signatures across 24 cancer subtypes. Mutation signatures showing a significant difference in contribution between stepped and non-stepped SNVs after multiple hypothesis testing adjustments are indicated by a*. These figures represent the raw data summarized in Figure 1C. [Figure 6-2] Same as above. [Figure 6-3] Same as above. [Figure 6-4] Same as above. [Figure 6-5] Same as above. [Figure 6-6] Same as above.
[0101] [Figure 7] Figure 7 illustrates the distribution of PVs in typological regions across the genome. The bar plot shows the distribution of PVs occurring in typological regions across the genomes of multiple cancer types. In this plot, the genome was divided into 1000 bp bins, and the percentage of samples of a given histological image having PVs in each 1000 bp bin was calculated. Only bins with a recurrence rate of at least 2% in any given cancer subtype are shown. The histological images shown are similar to those in Figure 1E. Activated B cell (ABC) and germinal center B cell (GCB) subtypes of DLBCL are also shown.
[0102] [Figure 8] Figure 8 illustrates the amount and genomic location of PVs from WGS in lymphoid malignancies. Figure 8A is a bar plot showing the number of independent 1000 bp regions across the genome that repeatedly contain PVs for DLBCL, FL, BL, and CLL (n=68, 74, 36, and 151, respectively). Figures 8B–8D are plots showing the frequency of PVs for multiple lymphoid malignancies with specific loci, including Figure 8B: BCL2, Figure 8C: MYC, and Figure 8D: ID 3. The location of the transcript of a given gene is shown below the gray plot. Exons are shown in darker gray. * Indicates regions with significantly more PVs in a given cancer histological image compared to all other histological images by Fisher's direct test (P<0.05). As with Figures 8E, 8B–8D, these plots show the frequency of PVs across lymphoma subtypes. Here, the IGH gene loci, consisting of the IGHV, IGHD, and IGHJ portions, are shown for ABC and GCB subtype DLBCL (n=25 and 25, respectively). The coding regions of the Ig portion, including the Ig constant region and the V gene, are shown. (DLBCL, diffuse large B-cell lymphoma; FL, follicular lymphoma; BL, Burkitt lymphoma, CLL, chronic lymphocytic leukemia).
[0103] [Figure 9-1] Figure 9 illustrates the performance of PhasED-Seq for PV recovery across lymphoma. Figure 9A shows a univariate scatter plot showing the proportion of total PVs across the genome (n=79) identified by WGS recovered by a previously reported lymphoma CAPP-Seq panel 8 (left) compared to PhasED-Seq (right). Figure 9B illustrates the expected yield per case of SNVs identified from WGS using a previously established lymphoma CAPP-Seq panel or PhasED-Seq panel. Figure 9C illustrates the expected yield per case of PVs identified from WGS using a previously established lymphoma CAPP-Seq panel or PhasED-Seq panel. Data from three independent, publicly available cohorts are shown in Figures 9A–9C. Figures 9D–9F illustrate plots showing the improvement in PV recovery by PhasED-Seq compared to CAPP-Seq in 16 patients sequenced by both assays. This includes improvements in d) two SNVs in the same phase (e.g., 2x or "doublet PV"), e) three SNVs in the same phase (3x or "triplet PV"), and f) four SNVs in the same phase (e.g., 4x or "quadruplet PV"). Figures 9G–9K illustrate panels showing the number of SNVs and PVs identified for patients with different types of lymphoma. These panels show the number of SNVs, h) doublet PVs, i) triplet PVs, j) quadrupleplet PVs, and k) all PVs. *, P<0.05; **, P<0.01, ***, P<0.001. (DLBCL, diffuse large B-cell lymphoma; GCB, germinal center B-cell-like DLBCL; ABC, activated B-cell-like DLBCL; PMBCL, primary mediastinal B-cell lymphoma; HL, Hodgkin lymphoma). [Figure 9-2] Same as above.
[0104] [Figure 10-1]Figure 10 illustrates the site-specific differences in PV between ABC-DLBCL and GCB-DLBC (Figures 10A-10Y). Similar to Figure 2D, these scatter plots compare the frequency of PV by genomic location (in 50 bp bins) for patients with different types of lymphoma. This figure shows the difference between ABC-DLBCL and GCB-DLBCL. Red circles indicate the relative frequency of PV in 50 bp bins from the specific gene of interest. Other (gray) circles indicate the relative frequency of PV in 50 bp bins from the rest of the PhasED-Seq sequencing panel. Only genes with statistically significant differences in PV between ABC-DLBCL and GCB-DLBCL are shown. The P-value represents the Wilcoxon rank-sum test of 50 bp bins from a given gene against all other 50 bp bins. See Example 10. [Figure 10-2] Same as above. [Figure 10-3] Same as above.
[0105] [Figure 11-1] Figure 11 illustrates the location-specific differences in PV between DLBCL and PMBCL (Figures 11A–11X). Similar to Figure 2D, these scatter plots compare the frequency of PV by genomic location (in 50 bp bins) for patients with different types of lymphoma. This figure shows the differences between DLBCL and PMBCL. Blue circles represent the relative frequency of PV in 50 bp bins from the specific gene of interest. Other (gray) circles represent the relative frequency of PV in 50 bp bins from the rest of the PhaseD-Seq sequencing panel. Only genes with statistically significant differences in PV between DLBCL and PMBCL are shown. The P-value represents the Wilcoxon rank-sum test of 50 bp bins from a given gene against all other 50 bp bins. See Example 10. [Figure 11-2] Same as above. [Figure 11-3] Same as above.
[0106] [Figure 12-1]Figure 12 illustrates the site-specific differences in PV between DLBCL and HL. Similar to Figure 2D, the scatter plots in Figures 12A–12NN compare the frequency of PV by genomic location (in 50 bp bins) in patients with different types of lymphoma. This figure shows the difference between DLBCL and HL. Green circles represent the relative frequency of PV in 50 bp bins from the specific gene of interest. Other (gray) circles represent the relative frequency of PV in 50 bp bins from the rest of the PhasED-Seq sequencing panel. Only genes with statistically significant differences in PV between DLBCL and HL are shown. The P-value represents the Wilcoxon rank-sum test of 50 bp bins from a given gene against all other 50 bp bins. See Example 10. [Figure 12-2] Same as above. [Figure 12-3] Same as above. [Figure 12-4] Same as above. [Figure 12-5] Same as above.
[0107] [Figure 13] Figure 13 illustrates the differences in PVs between lymphoma types in mutations at the IGH locus. This figure shows the frequency of PVs from PhasED-Seq across the @IGH locus for different types of B-cell lymphoma. The bottom track shows the structure of the @IGH locus and the gene region containing the Ig constant gene and the V gene. The next (outlined) track shows the frequency of PVs in this genomic region from WGS data (ICGC cohort). The remaining tracks show the frequency of PVs from PhasED-Seq targeted sequencing data, including 1) DLBCL, GCB-DLBCL, ABC-DLBCL, PMBCL, and HL. Regions targeted by the PhasED-Seq panel are shown at the top. Specific histological images label selected immunoglobulin regions with enriched PVs (i.e., IGHV4-34, Sε, Sγ3, and Sγ1).
[0108] [Figure 14-1]Figure 14 illustrates the technical aspects of PhasED-Seq by hybrid capture sequencing. Figure 14A shows a plot of theoretical binding energies for a typical 150-mer across the entire genome, with an increasing percentage of mutated bases from the reference genome. Mutations were spread across the entire 150-mer, either clustered at one end of the sequence, in the middle of the sequence, or randomly clustered across the entire sequence. Points and error bars represent the median and interquartile range from 10,000 in silico simulations. Figure 14B illustrates a plot showing two histograms of summary metrics for 151bp window mutation rates across the PhasED-Seq panel across all patients in this study. The light gray histogram shows the maximum percentage of mutated bases in any 151bp window for all patients in this study. The dark gray histogram shows the 95th percentile mutation rate across all mutated 151bp windows. Figure 14C is a plot showing the percentiles of mutation rates across all mutated 151bp windows across all patients in this study. Figure 14D illustrates a heatmap showing the relative error rates (as log10(error rate)) for single SNVs (left, "RED"), doublet PVs (center, "YELLOW"), and triplet PVs (right, "BLUE"). Figure 14D demonstrates that analysis based on multiple stepwise variants (e.g., double or triplet PVs) yields lower error rates than analysis based on single SNVs. Furthermore, Figure 14D demonstrates that analysis using a larger set of stepwise variants (e.g., triplet PVs labeled "BLUE") yields lower error rates than analysis based on a smaller set of stepwise variants (e.g., doublet PVs labeled "YELLOW"). Error rates for single SNVs from sequencing using multiple error suppression methods, including barcode deduplication, iDES, and double-stranded sequencing, are shown. Error rates are aggregated by mutation type. In the case of triplet PVs, the x and y axes of the heatmap represent the first and second types of base changes in the PV.The third change is averaged across all 12 possible base changes. Figure 14E illustrates a plot showing the error rate for doublet / 2× PV as a function of genomic distance between component SNVs. [Figure 14-2] Same as above.
[0109] [Figure 15] Figure 15 illustrates the comparison of ctDNA quantification by PhasED-Seq with CAPP-Seq and clinical applications. Figure 15 shows the detection rates of ctDNA from pre-treatment samples across 107 patients with large B-cell lymphoma using standard CAPP-Seq (green), as well as PhasED-Seq with doublets (light blue), triplets (medium blue), and quadrupletts (dark blue). The specificity of ctDNA detection is also shown. The two plots below show the false detection rates in 40 reserved healthy control cfDNA samples. The size of each bar in these two plots represents the detection rate of patient-specific cfDNA mutations in these 40 reserved controls across all 107 cases. [Figure 16]Figure 16 illustrates a comparison of ctDNA quantification by PhasED-Seq with CAPP-Seq and clinical applications. Figure 16A shows a table summarizing the sensitivity and specificity of ctDNA detection in pre-treatment samples by CAPP-Seq and PhasED-Seq using doublet, triplet, and quadruplet spreads, as shown in Panel A. Sensitivity is calculated across all 107 cases, while specificity is calculated across 40 reserved control samples, evaluated for a total of 4280 independent tests for each of the 107 independent patient-specific mutation lists. Figure 16B shows a scatter plot showing the amount of ctDNA measured by CAPP-Seq versus PhasED-Seq in individual samples (measured as log10 (haploid genomic equivalents / mL)). Samples collected before cycle 1 (i.e., pre-treatment), before cycle 2, and before cycle 3 of RCHOP treatment are shown in separate colors (blue, green, and red, respectively; 278 samples in total). Undetectable levels are on the axis. Spearman correlation and p-value are shown.
[0110] [Figure 17]Figure 17 illustrates the detection of ctDNA after two cycles of systemic therapy. Figure 17A illustrates a scatter plot showing the log-multiple change in ctDNA after two cycles of treatment (i.e., major molecular response or MMR) measured by CAPP-Seq or PhaseD-Seq for patients receiving RCHOP therapy. The dotted line represents the previously established threshold for a 2.5-log decrease in ctDNA relative to MMR. Undetectable samples are on the axis, and the correlation coefficient represents the Spearman RHO for 33 samples detected by both CAPP-Seq and PhaseD-Seq. Figure 17B illustrates a 2x2 table summarizing the detection rates of ctDNA samples after two cycles of treatment by PhaseD-Seq versus CAPP-Seq. Patients who ultimately progressed to disease are shown in the lower panel, and patients who did not ultimately progress to disease are shown in the upper panel. Figure 17C illustrates a bar plot showing the area under the receiver operator curve (AUC) for patient classifications of relapse-free survival at 24 months based on CAPP-Seq (light color) or PhaseD-Seq (dark color) after two cycles of treatment. Classifications for all patients (n=89, left) and patients who achieved MMR (n=69, right) are both shown. Figure 17D illustrates a Kaplan-Meier plot showing relapse-free survival for 69 patients who achieved MMR, stratified by ctDNA detection by CAPP-Seq (top) or PhaseD-Seq (bottom).
[0111] [Figure 18-1]Figure 18 illustrates the detection of ctDNA after one cycle of systemic therapy. Figure 18A illustrates a scatter plot showing the log-multiple change in ctDNA after one cycle of treatment (i.e., initial molecular response or EMR), measured by CAPP-Seq or PhaseD-Seq for patients receiving RCHOP therapy. The dotted line indicates the previously established threshold for the 2-log decrease of ctDNA against EMR. Undetectable samples are on the axis, and the correlation coefficient represents the Spearman RHO for 45 samples detected by both CAPP-Seq and PhaseD-Seq. Figure 18B illustrates a 2x2 table summarizing the detection rates of ctDNA samples after one cycle of treatment by PhaseD-Seq versus CAPP-Seq. Patients with final disease progression are shown in red, and patients without final disease progression are shown in blue. Figure 18C illustrates a bar plot showing the area under the receiver operating curve (AUC) for patient classifications for relapse-free survival at 24 months based on CAPP-Seq (light color) or PhaseD-Seq (dark color) after one cycle of treatment. Classifications for all patients (n=82, left) and patients who achieved EMR (n=63, right) are both shown. Figure 18D illustrates a Kaplan-Meier plot showing relapse-free survival for 63 patients who achieved EMR, stratified by ctDNA detection by CAPP-Seq (top) or ctDNA detection by PhaseD-Seq (bottom). Figure 18E illustrates a waterfall plot showing the change in ctDNA levels measured by CAPP-Seq after one cycle of first-line treatment in DLBCL patients. Patients with ctDNA undetectable by CAPP-Seq are shown in darker colors as "ND" ("not detected"). The bar colors also indicate the final clinical outcome for these patients. Figure 18F illustrates Kaplan-Meier plots showing relapse-free survival for 33 DLBCL patients with undetectable ctDNA measured by CAPP-Seq after one cycle of treatment.Figure 18G illustrates a Kaplan-Meier plot showing relapse-free survival for 33 patients (with ctDNA undetectable by CAPP-Seq) shown in Figure 18F, stratified by ctDNA detection by PhasED-Seq at the same point in time (Cycle 2, Day 1). Figure 18H illustrates a Kaplan-Meier plot showing relapse-free survival for 82 DLBCL patients stratified by ctDNA in Cycle 2, classified into three strata on Day 1: patients who did not achieve an initial molecular response, patients with an initial molecular response and still with ctDNA detectable by PhasED-Seq and / or CAPP-Seq, and patients with stringent molecular remission (ctDNA undetectable by both PhasED-Seq and CAPP-Seq). [Figure 18-2] Same as above.
[0112] [Figure 19] Figure 19 illustrates the proportion of patients for whom PhasED-Seq would achieve a lower LOD than double-strand sequencing tracking SNVs based on PCAWG data (whole-genome sequencing) that quantified the number of SNVs and stepwise variants (PVs) in different tumor types.
[0113] [Figure 20] Figure 20 illustrates the improved LOD achieved in lung cancer (adenocarcinoma, abbreviated as "A" and squamous cell carcinoma, abbreviated as "S") compared to double-strand sequencing of whole-genome sequencing data.
[0114] [Figure 21] Figure 21 illustrates empirical data from an experiment in which a custom panel was designed for 5 patients with solid tumors (5 lung cancers) to investigate and compare the LOD of custom CAPP-Seq versus PhasED-Seq, which showed approximately 10-fold lower LOD using PhasED-Seq in 5 / 5 patients. WGS was performed on tumor tissue.
[0115] [Figure 22]Figure 22A illustrates a proof-of-principle example patient photograph comparing the use of custom CAPP-Seq and PhaseD-Seq for disease surveillance in lung cancer, demonstrating earlier detection of recurrence using PhaseD-Seq. Figure 22B illustrates a proof-of-principle example patient photograph comparing the use of custom CAPP-Seq and PhaseD-Seq for early detection of disease in breast cancer, demonstrating earlier detection of disease with PhaseD-Seq.
[0116] [Figure 23] Figures 23A to 23B illustrate that the methods described herein (for example, the methods shown to generate Figures 3E and 3F) do not require barcode-mediated error suppression.
[0117] [Figure 24] Figure 24 illustrates a flowchart of a process in which clinical intervention and / or treatment is performed on an individual based on the detection of circulating tumor nucleic acid sequences in sequencing results, according to one embodiment.
[0118] [Figure 25A] Figures 25A to 25C show illustrative flowcharts of a method for determining the state of a target based on one or more cell-free nucleic acid molecules containing multiple variants.
[0119] [Figure 25B] Figures 25A to 25C show illustrative flowcharts of a method for determining the state of a target based on one or more cell-free nucleic acid molecules containing multiple variants. [Figure 25C] Figures 25A to 25C show illustrative flowcharts of a method for determining the state of a target based on one or more cell-free nucleic acid molecules containing multiple variants.
[0120] [Figure 25D]Figure 25D shows an illustrative flowchart of a method for treating a condition in question based on one or more cell-free nucleic acid molecules containing multiple variants.
[0121] [Figure 25E] Figure 25E shows an illustrative flowchart of a method for determining the progression (e.g., progression or regression) of a target state based on one or more cell-free nucleic acid molecules containing multiple variants.
[0122] [Figure 25F] Figures 25F and 25G show illustrative flowcharts of a method for determining the state of a target based on one or more cell-free nucleic acid molecules containing multiple variants. [Figure 25G] Figures 25F and 25G show illustrative flowcharts of a method for determining the state of a target based on one or more cell-free nucleic acid molecules containing multiple variants.
[0123] [Figure 26] Figures 26A and 26B schematically illustrate different fluorescent probes for identifying one or more cell-free nucleic acid molecules, including multiple stepwise variants.
[0124] [Figure 27] Figure 27 shows a computer system programmed or otherwise configured to implement the methods provided herein. [Modes for carrying out the invention]
[0125] Detailed explanation While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Those skilled in the art can make numerous modifications, changes, and substitutions without departing from the present invention. It should be understood that various alternative forms to the embodiments of the present invention described herein may be used.
[0126] The terms “about” or “approximately” generally mean within an acceptable margin of error for a given value, which may depend in part on how the value is measured or determined, for example, on the limitations of the measurement system. For example, “about” may mean within or above one standard deviation, according to convention in the art. Alternatively, “about” may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly with respect to biological systems or processes, the term may mean within one order of magnitude of a value, preferably up to five times, and more preferably up to two times. Where a particular value is described in this application and claims, unless otherwise specified, the term “about” can be assumed to mean within an acceptable margin of error for that particular value.
[0127] The terms “phased variants,” “variants in phase,” “PV,” or “somatic variant in phase,” as used interchangeably herein, generally refer to two or more mutations (e.g., SNVs or indels) occurring in the cis (i.e., on the same strand of the nucleic acid molecule) within a single cell-free nucleic acid molecule. In some cases, the cell-free nucleic acid molecule may be a cell-free deoxyribonucleic acid (cfDNA) molecule. In some cases, the cfDNA molecule may originate from diseased tissue, such as a tumor (e.g., a circulating tumor DNA (ctDNA) molecule).
[0128] The terms “biological sample” or “body sample,” as used interchangeably herein, generally refer to tissue or fluid samples derived from a subject. A biological sample may be obtained directly from a subject, or it may be derived from a subject (for example, by processing an initial biological sample obtained from the subject). A biological sample may be, or contain, one or more nucleic acid molecules, such as DNA or ribonucleic acid (RNA) molecules. A biological sample may be derived from any organ, tissue, or biological fluid. A biological sample may include, for example, fluid or solid tissue samples. An example of a solid tissue sample is, for example, a tumor sample from a solid tumor biopsy. Non-limiting examples of fluids include blood, serum, plasma, tumor cells, saliva, urine, cerebrospinal fluid, lymph, prostatic fluid, semen, milk, sputum, feces, tears, and their derivatives. In some cases, one or more cell-free nucleic acid molecules, as disclosed herein, may be derived from a biological sample.
[0129] As used herein, the term “subject” generally refers to any animal, mammal, or human. A subject may, potentially, or be suspected of having one or more conditions, such as a disease. In some cases, the subject’s condition may be cancer, cancer-related symptoms, or asymptomatic or undiagnosed with respect to cancer (e.g., not diagnosed with cancer). In some cases, a subject may have cancer, may exhibit cancer-related symptoms, may not have cancer-related symptoms, or may not be diagnosed with cancer. In some examples, the subject is human.
[0130] The terms “cell-free DNA” or “cfDNA,” as used interchangeably herein, generally refer to DNA fragments that circulate freely in the bloodstream of the subject. Cell-free DNA fragments may have dinucleosome protection (e.g., fragment size of at least 240 base pairs ("bp")). These cfDNA fragments with dinucleosome protection may not be cleaved between nucleosomes, resulting in longer fragment lengths (e.g., they have a typical size distribution centered around 334 bp). Cell-free DNA fragments may have mononucleosome protection (e.g., fragment size of less than 240 base pairs ("bp")). These cfDNA fragments with mononucleosome protection may be cleaved between nucleosomes, resulting in shorter fragment lengths (e.g., they have a typical size distribution centered around 167 bp).
[0131] As used herein, the term “sequencing data” generally refers to “raw sequence reads” and / or “consensus sequences” of nucleic acids, such as cell-free nucleic acids or their derivatives. Raw sequence reads are the output of a DNA sequencer and typically contain redundant sequences of the same parent molecule, for example, after amplification. A “consensus sequence” is a sequence derived from redundant sequences of the parent molecule, intended to represent the sequence of the original parent molecule. Consensus sequences can be constructed by voting (where each majority nucleotide in the sequence, e.g., the nucleotide most commonly observed at a given base position, is the consensus nucleotide) or by other approaches such as comparison with a reference genome. In some cases, consensus sequences can be constructed by tagging the original parent molecule with a unique or non-unique molecular tag, enabling tracking of the tag and / or tracking of the offspring sequence (e.g., after amplification) by using information within the sequence read.
[0132] As used herein, the term “reference genome sequence” generally refers to the nucleotide sequence on which the nucleotide sequence in question is compared.
[0133] As used herein, the term “genomic region” generally refers to any region of the genome (e.g., a range of base pair positions), such as the entire genome, a chromosome, a gene, or an exon. Genomic regions may be continuous or discontinuous. A “genetic locus” (or “locus”) may be part or all of a genomic region (e.g., a gene, a portion of a gene, or a single base of a gene).
[0134] As used herein, the term "likelihood" generally refers to probability, relative probability, existence or absence, or degree.
[0135] As used herein, the term “liquid biopsy” generally refers to a non-invasive or minimally invasive laboratory test or assay (e.g., of a biological sample or cell-free nucleic acid). A “liquid biopsy” assay can report the detection or measurement (e.g., minor allele frequency, gene expression, or protein expression) of one or more marker genes associated with the condition of interest (e.g., cancer or tumor-related marker genes). A. Introduction
[0136] Modifications of genomic DNA (e.g., mutations) may occur in the formation and / or progression of one or more conditions (e.g., diseases such as cancer or tumors) in a subject. This disclosure provides methods and systems for analyzing cell-free nucleic acid molecules, such as cfDNA, from a subject to determine the presence or absence of a condition in the subject, the prognosis of the diagnosed condition in the subject, the progression of the condition in the subject over time, therapeutic treatment for the diagnosed condition in the subject, or the predicted treatment outcome for the condition in the subject.
[0137] The analysis of cell-free nucleic acids, such as cfDNA, is being developed for a wide range of applications, for example, in prenatal testing, organ transplantation, infectious diseases, and oncology. In terms of detecting or monitoring target diseases such as cancer, circulating tumor DNA (ctDNA) can be a highly sensitive and specific biomarker in numerous cancer types. In some cases, ctDNA can be used to detect the presence of minimal residual disease (MRD) or tumor burden after procedures such as chemotherapy or surgical resection of solid tumors. However, the detection limit (LOD) of ctDNA analysis can be limited by several factors, including (i) the low input DNA volume from typical blood samples and (ii) the background error rate from sequencing.
[0138] In some cases, ctDNA-based cancer detection can be improved by tracking multiple somatic mutations using error-suppressed sequencing, for example, with an LOD of about 2 parts per 100,000 from a cfDNA input, while using commercially available panels or personalized assays. However, in some cases, the current LOD of the ctDNA of interest may be insufficient to broadly detect MRD in patients where disease recurrence or progression is anticipated. Such “loss of detection” can be exemplified, for example, in diffuse large B-cell lymphoma (DLBCL). In DLBCL, provisional ctDNA detection after treatment aimed at cure with only two cycles can represent the major molecular response (MMR) and can be a strong prognostic marker of the final clinical outcome. Nevertheless, nearly one-third of patients who ultimately experience disease progression do not have ctDNA detectable at this provisional landmark using available technologies (e.g., Cancer Personalized Profiling by Deep Sequencing (CAPP-Seq)), and therefore represent “false negative” measurements. Such high false-negative rates have also been observed in DLBCL patients using alternative methods such as ctDNA monitoring via immunoglobulin gene rearrangement. Therefore, improved methods for ctDNA-based cancer detection with higher sensitivity are needed.
[0139] Using somatic variants detected on both complementary strands of the parental DNA double helix can reduce the line of error (LOD) of ctDNA detection, thereby advantageously increasing the sensitivity of ctDNA detection. Such "double-strand sequencing" can reduce the background error profile due to the requirement of two matching events for the detection of single nucleotide variants (SNVs). However, since the recovery of both original strands may occur in only a small number of all recovered molecules, the double-strand sequencing approach alone can be limited by the inefficient recovery of the DNA double helix. Therefore, double-strand sequencing may be suboptimal and inefficient for real-world ctDNA detection with limited starting samples where the input DNA from actual blood volume (e.g., approximately 4,000 to 8,000 genomes per standard 10 ml (mL) blood collection tube) is limited and maximum genome recovery is essential.
[0140] Therefore, there remains a significant unmet need for the detection and analysis of ctDNA with low LOD (e.g., resulting in high sensitivity) to determine, for example, the presence or absence of a target disease, the prognosis of the disease, the treatment of the disease, and / or the expected outcome of the treatment. B. Methods and systems for determining or monitoring a state
[0141] This disclosure describes methods and systems for detecting and analyzing cell-free nucleic acids having multiple stepwise variants as a characteristic of the state of the subject. In some embodiments, the cell-free nucleic acid molecules may include cfDNA molecules such as ctDNA molecules. The methods and systems disclosed herein can utilize sequencing data derived from multiple cell-free nucleic acid molecules of the subject to identify subsets of multiple cell-free nucleic acid molecules having multiple stepwise variants, thereby determining the state of the subject. The methods and systems disclosed herein can directly detect such subsets of multiple cell-free nucleic acid molecules exhibiting multiple stepwise variants, and in some cases pull down (or capture), thereby determining the state of the subject with or without sequencing. The methods and systems disclosed herein can reduce the background error rate often involved in the detection and analysis of cell-free nucleic acid molecules such as cfDNA.
[0142] In some embodiments, methods and systems for cell-free nucleic acid sequencing and cancer detection are provided. In some embodiments, cell-free nucleic acids (e.g., cfDNA or cfRNA) can be extracted from a liquid biopsy of an individual and prepared for sequencing. The sequencing results of the cell-free nucleic acids can be analyzed to detect in-phase somatic variants (i.e., stepwise variants as disclosed herein) as indicators of circulating tumor nucleic acid (ctDNA or ctRNA) sequences (i.e., nucleic acids derived from cancer cells). ) or sequences originating from the nucleic acids of cancer cells. Therefore, in some cases, cancer can be detected in an individual by extracting a liquid biopsy from the individual and sequencing the cell-free nucleic acids derived from that liquid biopsy to detect circulating tumor nucleic acid sequences, and the presence of circulating tumor nucleic acid sequences can indicate that the individual has cancer (e.g., a specific type of cancer). In some cases, clinical interventions and / or treatments can be determined and / or implemented for the individual based on the detection of cancer.
[0143] As disclosed herein, the presence of in-phase somatic variants can be a strong indicator that nucleic acids containing such stepwise variants originate from a body sample having a condition such as cancer cells (or that nucleic acids originate from a body sample obtained from or derived from an object having a condition such as cancer). Since mutations are unlikely to occur within a small gene window, which is the approximate size of a typical cell-free nucleic acid molecule (e.g., about 170 bp or less), stepwise detection of somatic variants can improve the signal-to-noise ratio of cell-free nucleic acid detection methods (e.g., by reducing or eliminating false "noise" signals).
[0144] In some embodiments, particularly in various cancers, such as lymphoma, certain genomic regions can be used as hotspots for detecting stepwise variants. In some cases, enzymes (e.g., AID, Apobec3a) can typologically mutate DNA at specific genes and locations, leading to the development of specific cancers. Thus, cell-free nucleic acids derived from such hotspot genomic regions can be captured or targeted (e.g., with or without deep sequencing) for cancer detection and / or monitoring. Alternatively, capture or targeted sequencing can be performed to detect cancer in a particular individual against regions where stepwise variants have been previously detected from the individual's cancerous source (e.g., tumor).
[0145] In some embodiments, capture sequencing of cell-free nucleic acids can be performed as a screening diagnosis. In some cases, a screening diagnosis can be developed and used to detect circulating tumor nucleic acids for cancers having typological regions of stepwise variants. In some cases, capture sequencing of cell-free nucleic acids is performed as a diagnosis to detect MRD or tumor burden and determine whether a particular disease is present during or after treatment. In some cases, capture sequencing of cell-free nucleic acids can be performed as a diagnosis to determine the progression of treatment (e.g., progression or regression).
[0146] In some embodiments, cell-free nucleic acid sequencing results can be analyzed to detect whether stepwise somatic single nucleotide variants (SNVs) or other mutations or variants (e.g., indels) are present in the cell-free nucleic acid sample. In some cases, the presence of a particular somatic SNV or other variant may indicate a circulating tumor nucleic acid sequence and thus indicate a tumor present in the subject. In some cases, at least two variants can be detected in phase on the cell-free nucleic acid molecule. In some cases, at least three variants can be detected in phase on the cell-free nucleic acid molecule. In some cases, at least four variants can be detected in phase on the cell-free nucleic acid molecule. In some cases, at least five or more variants can be detected in phase on the cell-free nucleic acid molecule. In some cases, a greater number of stepwise variants detected on the cell-free nucleic acid molecule is more likely to originate from cancer, in contrast to detecting harmless sequences of somatic variants resulting from molecular preparation of the sequence library or random biological errors. Thus, the likelihood of false-positive detection can decrease with the detection of more in-phase variants within the molecule (e.g., thereby increasing the specificity of the detection).
[0147] In some embodiments, cell-free nucleic acid sequencing results can be analyzed to detect whether one or more nucleic acid base insertions or deletions (i.e., indels) are present in the cell-free nucleic acid sample relative to, for example, a reference genome sequence. While not wishing to be constrained by theory, in some cases, the presence of indels in a cell-free nucleic acid molecule (e.g., cfDNA) can indicate a condition in question, such as a disease, like cancer. In some cases, the resulting genetic variation of an indel can be treated as a variant or mutation, and therefore, two indels can be treated as two stepwise variants, as disclosed herein. In some examples, within a cell-free nucleic acid molecule, a first genetic variation from a first indel (first phase variant) and a second genetic variation from a second indel (second phase variant) can be separated from each other by at least one nucleotide.
[0148] As disclosed herein, within a single cell-free nucleic acid molecule (e.g., a single cfDNA molecule), the first stepwise variant may be an SNV, and the second stepwise variant may be part of a different small nucleotide polymorphism, e.g., another SNV or a multinucleotide variant (MNV). A multinucleotide variant may be a cluster of two or more (e.g., at least two, three, four, five, or more) adjacent variants present within the same stand of the nucleic acid molecule. In some cases, the first and second stepwise variants may be part of the same MNV within a single cell-free nucleic acid molecule. In some cases, the first and second stepwise variants may originate from two different MNVs within a single cell-free nucleic acid molecule.
[0149] In some embodiments, statistical methods can be used to calculate the likelihood that the detected stepwise variants originate from cancer and are neither random nor artificial (e.g., from sample preparation or sequencing errors). In some cases, Monte Carlo sampling methods can be used to determine the likelihood that the detected stepwise variants originate from cancer and are neither random nor artificial.
[0150] Aspects of this disclosure provide the identification or detection of cell-free nucleic acids (e.g., cfDNA molecules) having multiple stepwise variants, for example, from a liquid biopsy of a subject. In some cases, the first stepwise variant of the multiple stepwise variants and the second stepwise variant of the multiple stepwise variants can be directly adjacent to each other (e.g., adjacent SNVs). In some cases, the first stepwise variant of the multiple stepwise variants and the second stepwise variant of the multiple stepwise variants can be separated by at least one nucleotide. The interval between the first stepwise variant and the second stepwise variant can be limited by the length of the cell-free nucleic acid molecule.
[0151] As disclosed herein, within a single cell-free nucleic acid molecule (e.g., a single cfDNA molecule), the first stepwise variant and the second stepwise variant are at least or up to about 1 nucleotide, at least or up to about 2 nucleotides, at least or up to about 3 nucleotides, at least or up to about 4 nucleotides, at least or up to about 5 nucleotides, at least or up to about 6 nucleotides, at least or up to about 7 nucleotides, at least or up to about 8 nucleotides, at least or up to about 9 nucleotides, at least or up to about 10 nucleotides, at least or up to about 11 nucleotides, at least or up to about 12 nucleotides, at least or up to about 13 nucleotides, at least or up to about 14 nucleotides, at least or up to about 15 nucleotides, at least or up to about 20 nucleotides, at least or up to about 25 They may be separated from each other by nucleotides, at least or up to approximately 30 nucleotides, at least or up to approximately 35 nucleotides, at least or up to approximately 40 nucleotides, at least or up to approximately 45 nucleotides, at least or up to approximately 50 nucleotides, at least or up to approximately 60 nucleotides, at least or up to approximately 70 nucleotides, at least or up to approximately 80 nucleotides, at least or up to approximately 90 nucleotides, at least or up to approximately 100 nucleotides, at least or up to approximately 110 nucleotides, at least or up to approximately 120 nucleotides, at least or up to approximately 130 nucleotides, at least or up to approximately 140 nucleotides, at least or up to approximately 150 nucleotides, at least or up to approximately 160 nucleotides, at least or up to approximately 170 nucleotides, or at least or up to approximately 180 nucleotides. Alternatively, or in addition, within a single cell-free nucleic acid molecule, the first stepwise variant and the second stepwise variant may not be separated by one or more nucleotides, or do not need to be separated, and may therefore be directly adjacent to each other.
[0152] A single cell-free nucleic acid molecule (e.g., a single cfDNA molecule) disclosed herein can contain at least or up to about 2 stepwise variants, at least or up to about 3 stepwise variants, at least or up to about 4 stepwise variants, at least or up to about 5 stepwise variants, at least or up to about 6 stepwise variants, at least or up to about 7 stepwise variants, at least or up to about 8 stepwise variants, at least or up to about 9 stepwise variants, at least or up to about 10 stepwise variants, at least or up to about 12 stepwise variants, at least or up to about 12 stepwise variants, at least or up to about 13 stepwise variants, at least or up to about 14 stepwise variants, at least or up to about 15 stepwise variants, at least or up to about 20 stepwise variants, or at least or up to about 25 stepwise variants within the same molecule.
[0153] From the multiple cell-free nucleic acid molecules obtained (e.g., from the fluid biopsy of the subject), each cell-free nucleic acid molecule identified as containing multiple stepwise variants contained, on average, at least or up to approximately 2 stepwise variants, at least or up to approximately 3 stepwise variants, at least or up to approximately 4 stepwise variants, at least or up to approximately 5 stepwise variants, at least or up to approximately 6 stepwise variants, at least or up to approximately 7 stepwise variants, at least or up to approximately 8 stepwise variants, at least or up to approximately 9 stepwise variants, and at least Alternatively, two or more (e.g., 10 or more, 1,000 or more, 10,000 or more) cell-free nucleic acid molecules can be identified that have up to approximately 10 stepwise variants, at least or up to approximately 12 stepwise variants, at least or up to approximately 13 stepwise variants, at least or up to approximately 14 stepwise variants, at least or up to approximately 15 stepwise variants, at least or up to approximately 20, or at least or up to approximately 25 stepwise variants.
[0154] In some cases, multiple cell-free nucleic acid molecules (e.g., cfDNA molecules) can be obtained from a biological sample of the subject (e.g., a solid tumor or a fluid biopsy). Among the multiple cell-free nucleic acid molecules, there may be at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least fifteen, at least twenty, at least twenty-five, at least thirty, at least thirty-five, at least forty, at least forty-five, at least fifty, at least sixty, at least seventy, at least eighty-five, at least fifty-five, at least sixty, at least seventy, at least eighty-five, At least or up to 90, at least or up to 100, at least or up to 150, at least or up to 200, at least or up to 300, at least or up to 400, at least or up to 500, at least or up to 600, at least or up to 700, at least or up to 800, at least or up to 900, at least or up to 1,000, at least or up to 5,000, at least or up to 10,000, at least or up to 50,000, or at least or up to 100,000 cell-free nucleic acid molecules can be identified, each identified cell-free nucleic acid molecule comprising multiple stepwise variants as disclosed herein.
[0155] In some cases, multiple cell-free nucleic acid molecules (e.g., cfDNA molecules) can be obtained from a biological sample of a subject (e.g., a solid tumor or a liquid biopsy). Among the multiple cell-free nucleic acid molecules, at least or at most 1, at least or at most 2, at least or at most 3, at least or at most 4, at least or at most 5, at least or at most 6, at least or at most 7, at least or at most 8, at least or at most 9, at least or at most 10, at least or at most 15, at least or at most 20, at least or at most 25, at least or at most 30, at least or at most 35, at least or at most 40, at least or at most 45, at least or at most 50, at least or at most 60, at least or at most 70, at least or at most 80, at least or at most 90, at least or at most 100, at least or at most 150, at least or at most 200, at least or at most 300, at least or at most 400, at least or at most 500, at least or at most 600, at least or at most 700, at least or at most 800, at least or at most 900, or at least or at most 1,000 cell-free nucleic acid molecules can be identified from a target genomic region (e.g., a target genomic locus), and each identified cell-free nucleic acid molecule contains multiple stepwise variants as disclosed herein.
[0156] Figures 1A and 1E schematically illustrate examples of (i) a cfDNA molecule containing SNVs and (ii) another cfDNA molecule containing multiple stepwise variants. Each variant identified within the cfDNA can indicate the presence of another genetic mutation in the cell from which the cfNDA is derived. In alternative embodiments, one or more stepwise variants can be insertions or deletions (indels) instead of SNVs.
[0157] In one embodiment, the present disclosure provides a method for determining the state of a subject, as shown by flowchart 2510 in Figure 25A. The method may include (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject by a computer system (process 2512). The method may further include (b) processing the sequencing data by a computer system to identify one or more cell-free nucleic acid molecules from a plurality of cell-free nucleic acid molecules, each of the identified one or more cell-free nucleic acid molecules containing a plurality of stepwise variants with respect to a reference genome sequence (process 2514). In some cases, as disclosed herein, at least a portion of the one or more cell-free nucleic acid molecules may include a first stepwise variant of a plurality of stepwise variants separated by at least one nucleotide and a second stepwise variant of a plurality of stepwise variants. The method may optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules by a computer system to determine the state of the subject (process 2516).
[0158] In some cases, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, at least or up to about 50%, at least or up to about 60%, at least or up to about 70%, at least or up to about 80%, at least or up to about 90%, at least or up to about 95%, at least or up to about 99%, or about 100% of one or more cell-free nucleic acid molecules may include a first step variant of a plurality of step variants and a second step variant of a plurality of step variants, separated by at least one nucleotide, as disclosed herein. In some examples, multiple stepwise variants within a single cfDNA molecule may consist of (i) a first set of stepwise variants separated from each other by at least one nucleotide, and (ii) a second set of stepwise variants adjacent to each other (e.g., two stepwise variants within an MNV). In some examples, multiple stepwise variants within a single cfDNA molecule may consist of stepwise variants separated from each other by at least one nucleotide.
[0159] In one embodiment, the present disclosure provides a method for determining the state of a subject, as shown by flowchart 2520 in Figure 25B. The method may include (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject by a computer system (process 2522). The method may further include (b) processing the sequencing data by a computer system to identify one or more cell-free nucleic acid molecules from a plurality of cell-free nucleic acid molecules, where each of the one or more cell-free nucleic acid molecules contains a plurality of stepwise variants with respect to a reference genome sequence (process 2524). In some cases, as disclosed herein, a first stepwise variant of the plurality of stepwise variants and a second stepwise variant of the plurality of stepwise variants may be separated by at least one nucleotide. The method may optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules by a computer system to determine the state of the subject (process 2526).
[0160] In one embodiment, the present disclosure provides a method for determining the state of a subject, as shown by flowchart 2530 in Figure 25C. This method may include (a) obtaining sequencing data from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject (process 2532). This method may further include (b) processing the sequencing data to identify one or more cell-free nucleic acid molecules having LOD, which is less than about one (or cell-free nucleic acid molecule) out of 50,000 observations from the sequencing data (process 2534). In some cases, each of the one or more cell-free nucleic acid molecules contains multiple stepwise variants with respect to a reference genome sequence. This method may optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules to determine the state of the subject (process 2536).
[0161] In some cases, as disclosed herein, the LOD of an operation to identify one or more cell-free nucleic acid molecules may be observed to be less than approximately 1 in 60,000, less than 1 in 70,000, less than 10 in 80,000, less than 1 in 90,000, less than 1 in 100,000, less than 1 in 150,000, less than 1 in 200,000, less than 1 in 300,000, less than 1 in 400,000, less than 1 in 500,000, less than 1 in 600,000, less than 1 in 700,000, less than 1 in 800,000, less than 1 in 900,000, less than 1 in 1,0200,000, less than 1 in 1,300,000, less than 1 in 1,400,000, less than 1 in 1,500,000, or less than 1 in 2,000,000 from the sequencing data.
[0162] In some cases, as disclosed herein, at least one of the identified cell-free nucleic acid molecules may comprise a first stepwise variant of a plurality of stepwise variants and a second stepwise variant of the plurality of stepwise variants separated by at least one nucleotide.
[0163] In some cases, one or more operations (a) to (c) of the subject method can be performed by a computer system. For example, all of operations (a) to (c) of the subject method can be performed by a computer system.
[0164] The sequencing data disclosed herein can be obtained from one or more sequencing methods. The sequencing methods may be first-generation sequencing methods (e.g., Maxam-Gilbert sequencing, Sanger sequencing). The sequencing methods may also be high-throughput sequencing methods, such as next-generation sequencing (NGS) (e.g., synthetic sequencing). High-throughput sequencing methods can simultaneously (or substantially simultaneously) sequence at least about 10,000, at least about 100,000, at least about 1,000,000, at least about 10,000,000, at least about 100,000,000, at least about 1,000,000,000,000 polynucleotide molecules (e.g., cell-free nucleic acid molecules or their derivatives) or more. NGS may be sequencing techniques of any number of generations (e.g., second-generation sequencing techniques, third-generation sequencing techniques, fourth-generation sequencing techniques, etc.). Non-limiting examples of high-throughput sequencing methods include large-scale parallel signature sequencing, Polony sequencing, pyro-sequencing, synthetic sequencing, combinatorial probe anchor synthesis (cPAS), ligation sequencing (e.g., oligonucleotide ligation and detection (SOLiD) sequencing), semiconductor sequencing (e.g., ion torrent semiconductor sequencing), DNA nanoball sequencing, single-molecule sequencing, and hybridization sequencing.
[0165] In some embodiments of any one of the methods disclosed herein, sequencing data can be obtained based on any of the disclosed sequencing methods that utilize nucleic acid amplification (e.g., polymerase chain reaction (PCR)). Non-limiting examples of such sequencing methods may include 454 pyrosequencing, Polony sequencing, and SoLiD sequencing. In some cases, sequencing data can be generated by generating amplicons (e.g., derivatives of multiple cell-free nucleic acid molecules obtained from or derived from a subject, as disclosed herein) corresponding to a genomic region of interest (e.g., a genomic region associated with a disease) by PCR, pooling them as needed, and then sequencing them. In some examples, since the region of interest is amplified into amplicons by PCR before sequencing, the nucleic acid sample is already enriched with the region of interest, and therefore additional pooling (e.g., hybridization) before sequencing may not be required, and is not required (e.g., non-hybridization-based NGS). Alternatively, pooling by hybridization can be further performed for further enrichment before sequencing. Alternatively, sequencing data can be obtained without generating PCR copies, for example, by cPAS sequencing.
[0166] Some embodiments utilize capture hybridization techniques to perform targeted sequencing. When sequencing cell-free nucleic acids, library products can be captured by hybridization before sequencing to increase the resolution of specific genomic loci. Capture hybridization can be particularly useful when attempting to detect rare and / or somatic stepwise variants from a sample at specific genomic loci. In some situations, the detection of rare and / or somatic stepwise variants indicates a source of nucleic acids, including nucleic acids originating from a cancer source. Thus, capture hybridization is a tool that can enhance the detection of circulating tumor nucleic acids within cell-free nucleic acids.
[0167] Various types of cancer repeatedly experience abnormal somatic hypermutation, particularly at genomic loci. For example, enzyme activation-inducible deaminase induces abnormal somatic hypermutation in B cells, resulting in a variety of B-cell lymphomas, including (but not limited to) diffuse large B-cell lymphoma (DLBCL), follicular lymphoma (FL), Burkitt lymphoma (BL), and B-cell chronic lymphocytic leukemia (CLL). Therefore, in numerous embodiments, probes are designed to pull down (or capture) genomic loci known to experience abnormal somatic hypermutation in lymphomas. Figure 1D and Table 1 describe some regions that experience abnormal somatic hypermutation in DLBCL, FL, BL, and CLL. Provided in Table 6 is a list of nucleic acid probes that can be used to pull down (or capture) genomic loci to detect abnormal somatic hypermutation in B-cell cancers.
[0168] Capture sequencing can also be performed using personalized nucleic acid probes designed to detect the presence of cancer in an individual. An individual with cancer can have the cancer biopsied and sequenced to detect somatic stepwise variants accumulated in the tumor. Based on the sequencing results, according to several embodiments, nucleic acid probes are designed and synthesized that can pull down genomic loci containing the locations where stepwise variants reside. These personalized, designed, and synthesized nucleic acid probes can be used to detect circulating tumor nucleic acids from a fluid biopsy of that individual. Therefore, personalized nucleic acid probes may be useful for determining treatment response and / or detecting post-treatment MRD.
[0169] In some embodiments of any one of the methods disclosed herein, sequencing data can be obtained based on any sequencing method utilizing an adapter. A nucleic acid sample (e.g., multiple cell-free nucleic acid molecules derived from a subject, as disclosed herein) can be conjugated with one or more adapters (or adapter sequences) for recognition (e.g., by hybridization) of the sample or any derivative (e.g., an amplicon). In some examples, for example, a nucleic acid sample can be tagged with a molecular barcode so that each cell-free nucleic acid molecule of multiple cell-free nucleic acid molecules can have a unique barcode. Alternatively, or in addition to this, a nucleic acid sample can be tagged with a sample barcode so that multiple cell-free nucleic acid molecules derived from a subject (e.g., multiple cell-free nucleic acid molecules obtained from a specific body tissue of a subject) can have the same barcode.
[0170] In alternative embodiments, the method for identifying one or more cell-free nucleic acid molecules comprising multiple stepwise variants disclosed herein can be carried out without molecular barcoding, without sample barcoding, or without molecular and sample barcoding, at least in part due to the high specificity and low LOD achieved by relying on the identification of stepwise variants, in contrast to, for example, a single SNV.
[0171] In any one or several embodiments of the methods disclosed herein, sequencing data can be obtained and analyzed at least in part due to high specificity and low LOD achieved by relying on identifying stepwise variants, as opposed to single SNVs or indels, without (i) background errors and / or (ii) sequencing errors being removed or suppressed in silico.
[0172] In some embodiments of any one of the methods disclosed herein, using multiple variants as conditions for identifying target cell-free nucleic acid molecules with desired specific mutations without in silico error suppression can yield background error rates at least about 5 times, at least about 10 times, at least about 20 times, at least about 30 times, at least about 40 times, at least about 50 times, at least about 60 times, at least about 70 times, at least about 80 times, at least about 90 times, at least about 100 times, at least about 200 times, at least about 400 times, at least about 600 times, at least about 800 times, or at least about 1,000 times lower than the background error rates of (i) barcode deduplication, (ii) integrated digital error suppression, or (iii) double-strand sequencing. This approach can favorably increase the signal-to-noise ratio for identifying target cell-free nucleic acid molecules with desired specific mutations (and thereby increase sensitivity and / or specificity).
[0173] In some embodiments of any one of the methods disclosed herein, the background error rate can be reduced by at least about 5 times, at least about 10 times, at least about 20 times, at least about 30 times, at least about 40 times, at least about 50 times, at least about 60 times, at least about 70 times, at least about 80 times, at least about 90 times, or at least about 100 times by increasing the minimum number of stepwise variants per cell-free nucleic acid molecule required as a condition for identifying a target cell-free nucleic acid molecule having the specific mutation of the desired object (e.g., increasing from at least two stepwise variants to at least three stepwise variants). This approach can favorably increase the signal-to-noise ratio for identifying a target cell-free nucleic acid molecule having the specific mutation of the desired object (and thereby increase sensitivity and / or specificity).
[0174] In one embodiment, the present disclosure provides a method for treating a condition of a subject, as shown in flowchart 2540 of Figure 25D. The method may include (a) identifying a subject for treatment of a condition, the subject being determined to have a condition based on the identification of one or more cell-free nucleic acid molecules from a plurality of cell-free nucleic acid molecules obtained from or derived from the subject (process 2542). Each of the identified one or more cell-free nucleic acid molecules may contain a plurality of stepwise variants with respect to a reference genome sequence. As disclosed herein, at least a portion (e.g., partially or all) of the plurality of stepwise variants can be separated by at least one nucleotide such that the first stepwise variant of the plurality of stepwise variants and the second stepwise variant of the plurality of stepwise variants are separated by at least one nucleotide. In some cases, the presence of the plurality of stepwise variants indicates a condition of the subject (e.g., a disease such as cancer). The method may further include (b) subjecting the subject to treatment based on step (a) (process 2544). Examples of such treatments of a condition of a subject are disclosed elsewhere in the present disclosure.
[0175] In one embodiment, the present disclosure provides a method for monitoring the progression (e.g., progression or regression) of a state of a subject, as shown in flowchart 2550 of Figure 25E. The method may include (a) determining a first state of the subject based on the identification of a first set of one or more cell-free nucleic acid molecules from a first set of cell-free nucleic acid molecules obtained from or derived from the subject (process 2552). The method may further include (b) determining a second state of the subject based on the identification of a second set of one or more cell-free nucleic acid molecules from a second set of cell-free nucleic acid molecules obtained from or derived from the subject (process 2554). The second set of cell-free nucleic acid molecules can be obtained from the subject after the first set of cell-free nucleic acid molecules have been obtained from the subject. The method may optionally include (c) determining the progression (e.g., progression or regression) of the state based at least in part on the first state and the second state (process 2556). In some cases, each of the identified one or more cell-free nucleic acid molecules (e.g., each of the first set of identified one or more cell-free nucleic acid molecules, each of the second set of identified one or more cell-free nucleic acid molecules) may contain multiple stepwise variants with respect to a reference genome sequence. As disclosed herein, at least a portion (e.g., partially or all) of the identified one or more cell-free nucleic acid molecules can be separated by at least one nucleotide. In some cases, the presence of multiple stepwise variants may indicate the state of the subject.
[0176] In some cases, a first plurality of cell-free nucleic acid molecules can be obtained from a subject (e.g., by a blood biopsy), analyzed, and used to determine (e.g., diagnose) a first situation of the subject's condition (e.g., a disease such as cancer). The first plurality of cell-free nucleic acid molecules can be analyzed using any of the methods disclosed herein (e.g., with or without sequencing) to identify a first set of one or more cell-free nucleic acid molecules that include a plurality of stepwise variants, and the presence or characteristics of the first set of one or more cell-free nucleic acid molecules can be used to determine a first situation (e.g., an initial diagnosis) of the subject's condition. Based on the first situation of the determined condition, the subject can be subjected to one or more treatments (e.g., chemotherapy) disclosed herein. Following the one or more treatments, a second plurality of cell-free nucleic acid molecules can be obtained from the subject.
[0177] In some cases, the subject may be subjected to at least or up to approximately one treatment, at least or up to approximately two treatments, at least or up to approximately three treatments, at least or up to approximately four treatments, at least or up to approximately five treatments, at least or up to approximately six treatments, at least or up to approximately seven treatments, at least or up to approximately eight treatments, at least or up to approximately nine treatments, or at least or up to approximately ten treatments, based on the first condition of the determined state. In some cases, the subject may receive multiple treatments based on a first condition of the determined state, and the first treatment of multiple treatments and the second treatment of multiple treatments may be separated for at least or up to approximately 1 day, at least or up to approximately 7 days, at least or up to approximately 2 weeks, at least or up to approximately 3 weeks, at least or up to approximately 4 weeks, at least or up to approximately 2 months, at least or up to approximately 3 months, at least or up to approximately 4 months, at least or up to approximately 5 months, at least or up to approximately 6 months, at least or up to approximately 12 months, at least or up to approximately 2 years, at least or up to approximately 3 years, at least or up to approximately 4 years, at least or up to approximately 5 years, or at least or up to approximately 10 years. The multiple treatments for the subject may be the same. Alternatively, the multiple treatments may differ in terms of drug type (e.g., different chemotherapy drugs), drug dosage (e.g., increased dosage, decreased dosage), presence or absence of co-treatments (e.g., chemotherapy and immunotherapy), mode of administration (e.g., intravenous versus oral administration), frequency of administration (e.g., daily, weekly, monthly), etc.
[0178] In some cases, the subject may not be treated for a state between the determination of the first state of the condition and the determination of the second state of the condition, and it may not be necessary to treat it. For example, without intervention of treatment, a second set of cell-free nucleic acid molecules derived from the subject (e.g., by liquid biopsy) may be included to confirm whether the subject still exhibits signs of the first state of the condition.
[0179] In some cases, a second set of cell-free nucleic acid molecules from a subject can be obtained (e.g., by blood biopsy) at least or up to approximately 1 day, at least or up to approximately 7 days, at least or up to approximately 2 weeks, at least or up to approximately 3 weeks, at least or up to approximately 4 weeks, at least or up to approximately 2 months, at least or up to approximately 3 months, at least or up to approximately 4 months, at least or up to approximately 5 months, at least or up to approximately 6 months, at least or up to approximately 12 months, at least or up to approximately 2 years, at least or up to approximately 3 years, at least or up to approximately 4 years, at least or up to approximately 5 years, or at least or up to approximately 10 years after obtaining the first set of cell-free nucleic acid molecules from the subject.
[0180] In some cases, as disclosed herein, different samples containing at least or up to about 2, at least or up to about 3, at least or up to about 4, at least or up to about 5, at least or up to about 6, at least or up to about 7, at least or up to about 8, at least or up to about 9, or at least or up to about 10 nucleic acid molecules (e.g., at least a plurality of first cell-free nucleic acid molecules and a plurality of second cell-free nucleic acid molecules) can be obtained over time (e.g., once a month for 6 months, once every 2 months for 1 year, once every 3 months for 1 year, once every 6 months for 1 year or more) to monitor the progression of the condition in question.
[0181] In some cases, the step of determining the progression of a condition based on a first and second state of the condition may include comparing one or more features of the first and second states of the condition, such as (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants in each state (e.g., per equal weight or volume of the original biological sample, per equal number of initial cell-free nucleic acid molecules analyzed), (ii) the average number of multiple stepwise variants per cell-free nucleic acid molecule identified as containing multiple stepwise variants (i.e., two or more stepwise variants), or (iii) the number of cell-free nucleic acid molecules identified as containing multiple stepwise variants divided by the total number of cell-free nucleic acid molecules containing mutations overlapping with some of the multiple stepwise variants (i.e., stepwise variant allele frequencies). Based on such comparisons, the MRD of the condition in question (e.g., cancer or tumor) can be determined. For example, the tumor burden or cancer burden of the condition can be determined based on such comparisons.
[0182] In some cases, progression of a condition can be a progression or worsening of the condition. In one example, worsening of a condition may include the development of cancer from an early stage to a later stage, such as from stage I cancer to stage III cancer. In another example, worsening of a condition may include an increase in the size (e.g., volume) of a solid tumor. In yet another example, worsening of a condition may include cancer metastasis from one location to another within the subject's body.
[0183] In some cases, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from a second state of the target state was at least or up to approximately 0.1 times, at least or up to approximately 0.2 times, at least or up to approximately 0.3 times, at least or up to approximately 0.4 times, at least or up to approximately 0.5 times, at least or up to approximately 0.6 times, at least or up to approximately 0.7 times, at least or up to approximately 0.8 times, at least or up to approximately 0.9 times, at least or up to approximately 1 time, at least or up to approximately 2 times, at least or up to approximately 3 times, at least or up to approximately 4 times less than (ii) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from a first state of the target state. It can become at least approximately 5 times, at least approximately 6 times, at least approximately 7 times, at least approximately 8 times, at least approximately 9 times, at least approximately 10 times, at least approximately 15 times, at least approximately 20 times, at least approximately 30 times, at least approximately 40 times, at least approximately 50 times, at least approximately 60 times, at least approximately 70 times, at least approximately 80 times, at least approximately 90 times, at least approximately 100 times, at least approximately 200 times, at least approximately 300 times, at least approximately 400 times, or at least approximately 500 times.
[0184] In some cases, (i) the average number of stepwise variants per cell-free nucleic acid molecule identified as containing multiple stepwise variants from the second state of the target state is at least or up to approximately 0.1 times, at least or up to approximately 0.2 times, at least or up to approximately 0.3 times, at least or up to approximately 0.4 times, at least or up to approximately 0.5 times, at least or up to approximately 0.6 times, at least or up to approximately 0.7 times, at least or up to approximately 0.8 times, at least or up to approximately 0.9 times, at least or up to approximately 1 time, at least or up to approximately 2 times, at least or up to approximately 3 times, less It can be no larger or up to approximately 4 times larger, at least or up to approximately 5 times larger, at least or up to approximately 6 times larger, at least or up to approximately 7 times larger, at least or up to approximately 8 times larger, at least or up to approximately 9 times larger, at least or up to approximately 10 times larger, at least or up to approximately 15 times larger, at least or up to approximately 20 times larger, at least or up to approximately 30 times larger, at least or up to approximately 40 times larger, at least or up to approximately 50 times larger, at least or up to approximately 60 times larger, at least or up to approximately 70 times larger, at least or up to approximately 80 times larger, at least or up to approximately 90 times larger, at least or up to approximately 100 times larger, at least or up to approximately 200 times larger, at least or up to approximately 300 times larger, at least or up to approximately 400 times larger, or at least or up to approximately 500 times larger.
[0185] In some cases, the progression of the condition may be a regression or at least partial remission. In one example, at least partial remission of the condition may involve downstaging the cancer from a later stage to an earlier stage, such as from stage IV cancer to stage II cancer. Alternatively, at least partial remission of the condition may be complete remission from the cancer. In another example, at least partial remission of the condition may involve a decrease in the size (e.g., volume) of a solid tumor.
[0186] In some cases, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from a second state of the target state was at least or up to approximately 0.1 times, at least or up to approximately 0.2 times, at least or up to approximately 0.3 times, at least or up to approximately 0.4 times, at least or up to approximately 0.5 times, at least or up to approximately 0.6 times, at least or up to approximately 0.7 times, at least or up to approximately 0.8 times, at least or up to approximately 0.9 times, at least or up to approximately 1 time, at least or up to approximately 2 times, at least or up to approximately 3 times, at least or up to approximately 4 times less than (ii) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from a first state of the target state. It can become at least or up to approximately 5 times smaller, at least or up to approximately 6 times smaller, at least or up to approximately 7 times smaller, at least or up to approximately 8 times smaller, at least or up to approximately 9 times smaller, at least or up to approximately 10 times smaller, at least or up to approximately 15 times smaller, at least or up to approximately 20 times smaller, at least or up to approximately 30 times smaller, at least or up to approximately 40 times smaller, at least or up to approximately 50 times smaller, at least or up to approximately 60 times smaller, at least or up to approximately 70 times smaller, at least or up to approximately 80 times smaller, at least or up to approximately 90 times smaller, at least or up to approximately 100 times smaller, at least or up to approximately 200 times smaller, at least or up to approximately 300 times smaller, at least or up to approximately 400 times smaller, or at least or up to approximately 500 times smaller.
[0187] In some cases, (i) the average number of stepwise variants per cell-free nucleic acid molecule identified as containing multiple stepwise variants from the second state of the target state is at least or up to approximately 0.1 times, at least or up to approximately 0.2 times, at least or up to approximately 0.3 times, at least or up to approximately 0.4 times, at least or up to approximately 0.5 times, at least or up to approximately 0.6 times, at least or up to approximately 0.7 times, at least or up to approximately 0.8 times, at least or up to approximately 0.9 times, at least or up to approximately 1 time, at least or up to approximately 2 times, at least or up to approximately 3 times, less It can be no smaller or up to approximately 4 times smaller, at least or up to approximately 5 times smaller, at least or up to approximately 6 times smaller, at least or up to approximately 7 times smaller, at least or up to approximately 8 times smaller, at least or up to approximately 9 times smaller, at least or up to approximately 10 times smaller, at least or up to approximately 15 times smaller, at least or up to approximately 20 times smaller, at least or up to approximately 30 times smaller, at least or up to approximately 40 times smaller, at least or up to approximately 50 times smaller, at least or up to approximately 60 times smaller, at least or up to approximately 70 times smaller, at least or up to approximately 80 times smaller, at least or up to approximately 90 times smaller, at least or up to approximately 100 times smaller, at least or up to approximately 200 times smaller, at least or up to approximately 300 times smaller, at least or up to approximately 400 times smaller, or at least or up to approximately 500 times smaller.
[0188] In some cases, the progression of states may remain substantially the same between the two states of the state in question. In some examples, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from the second state of the state in question may be approximately the same as (ii) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants from the first state of the state in question. In some examples, (i) the average number of multiple stepwise variants per cell-free nucleic acid molecule identified as containing multiple stepwise variants from the second state of the state in question may be approximately the same as (ii) the average number of multiple stepwise variants per cell-free nucleic acid molecule identified as containing multiple stepwise variants from the first state of the state in question.
[0189] In some embodiments of any one of the methods disclosed herein, one or more cell-free nucleic acid molecules containing multiple stepwise variants can be identified from multiple cell-free nucleic acid molecules by one or more sequencing methods. Alternatively, or in addition to this, one or more cell-free nucleic acid molecules containing multiple stepwise variants can be identified by pulling down (or capturing from among) multiple cell-free nucleic acid molecules using a set of nucleic acid probes. A pull-down (or capture) method via a set of nucleic acid probes may be sufficient to identify one or more cell-free nucleic acid molecules of interest without sequencing. In some cases, a set of nucleic acid probes may be configured to hybridize to at least some of the cell-free nucleic acid (e.g., cfDNA) molecules from one or more genomic regions associated with the state of interest. Thus, the presence of one or more cell-free nucleic acid molecules pulled down by a set of nucleic acid probes may be an indicator that one or more cell-free nucleic acid molecules originate from a state (e.g., ctDNA or ctRNA). Further details of the set of nucleic acid probes are disclosed elsewhere in this disclosure.
[0190] In some embodiments of any one of the methods disclosed herein, based on sequencing data derived from multiple cell-free nucleic acid molecules (e.g., cfDNA) obtained from or derived from a subject, (i) one or more cell-free nucleic acid molecules identified as containing multiple stepwise variants can be separated in silico from (ii) one or more other cell-free nucleic acid molecules not identified as containing multiple stepwise variants (or one or more other cell-free nucleic acid molecules not containing multiple stepwise variants). In some cases, the method may further include generating additional data containing sequencing information only for (i) one or more cell-free nucleic acid molecules identified as containing multiple stepwise variants. In some cases, the method may further include generating different data containing sequencing information only for (ii) one or more other cell-free nucleic acid molecules not identified as containing multiple stepwise variants (or one or more other cell-free nucleic acid molecules not containing multiple stepwise variants).
[0191] In one embodiment, the present disclosure provides a method for determining the state of a subject, as shown by flowchart 2560 in Figure 25F. This method may include (a) providing a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained from or derived from a subject (process 2562). In some cases, individual nucleic acid probes in a set of nucleic acid probes can be designed to hybridize to a target cell-free nucleic acid molecule comprising a plurality of stepwise variants to a reference genome sequence separated by at least one nucleotide. Thus, as disclosed herein, a first stepwise variant of a plurality of stepwise variants and a second stepwise variant of a plurality of stepwise variants can be separated by at least one nucleotide. In some cases, individual nucleic acid probes can include an activatable reporter agent. The activatable reporter agent can be activated by either (i) hybridization of an individual nucleic acid probe to a plurality of stepwise variants, or (ii) dehybridization of at least a portion of the individual nucleic acid probes hybridized to the plurality of stepwise variants. This method may further include (b) detecting an activated reporter agent to identify one or more cell-free nucleic acid molecules among several cell-free nucleic acid molecules (process 2564). Each of the one or more cell-free nucleic acid molecules may include multiple stepwise variants. This method may optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules to determine the state of the subject (process 2566).
[0192] In one embodiment, the present disclosure provides a method for determining the state of a subject, as shown by flowchart 2570 in Figure 25G. The method may include (a) providing a mixture comprising (1) a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules obtained from or derived from a subject (process 2572). In some cases, individual nucleic acid probes in the set of nucleic acid probes can be designed to hybridize to a target cell-free nucleic acid molecule comprising a plurality of stepwise variants with respect to a reference genome sequence. In some cases, individual nucleic acid probes may include an activatable reporter agent. The activatable reporter agent may be activated by either (i) hybridization of the individual nucleic acid probe to the plurality of stepwise variants, or (ii) dehybridization of at least some of the individual nucleic acid probes hybridized to the plurality of stepwise variants. The method may further include (b) detecting the activated reporter agent to identify one or more cell-free nucleic acid molecules among the plurality of cell-free nucleic acid molecules (process 2574). Each of the one or more cell-free nucleic acid molecules may include multiple stepwise variants, and as disclosed herein, the LOD of the identification step may be less than about 1 of the 50,000 cell-free nucleic acid molecules of the multiple cell-free nucleic acid molecules. The method may optionally include (c) analyzing at least a portion of the identified one or more cell-free nucleic acid molecules to determine the state of the subject (process 2576).
[0193] In some cases, as disclosed herein, the first stepwise variant of a plurality of stepwise variants and the second stepwise variant of a plurality of stepwise variants are separated by at least one nucleotide.
[0194] In some cases, the LOD of the step of identifying one or more cell-free nucleic acid molecules as disclosed herein is less than approximately 1 in 60,000, less than 1 in 70,000, less than 10 in 80,000, less than 1 in 90,000, less than 1 in 100,000, less than 1 in 150,000, less than 1 in 200,000, less than 1 in 300,000, less than 1 in 400,000, less than 1 in 500,000, less than 1 in 600,000, and 7 The number of cell-free nucleic acid molecules may be less than 1 per 0,000, less than 1 per 800,000, less than 1 per 900,000, less than 1 per 1,0200,000, less than 1 per 1,300,000, less than 1 per 1,400,000, less than 1 per 1,500,000, less than 2,000,000, less than 1 per 2,500,000, less than 1 per 3,000,000, less than 1 per 4,000,000, or less than 1 per 5,000,000. Generally, detection methods with lower LODs have higher sensitivity for such detections.
[0195] In some embodiments of any one of the methods disclosed herein, the method may further include (1) mixing a set of nucleic acid probes and (2) a plurality of cell-free nucleic acid molecules.
[0196] In some embodiments of any one of the methods disclosed herein, an activatable reporter agent for a nucleic acid probe may be activated during hybridization of individual nucleic acid probes into multiple stepwise variants. Non-limiting examples of such nucleic acid probes include molecular beacons, eclipse probes, and amplified phosphor probes. Examples include scorpion PCR primers and light-up extension fluorogenic PCR primers (LUX primers).
[0197] For example, a nucleic acid probe may be a molecular beacon, as shown in Figure 26A. The molecular beacon may be a fluorescently labeled (e.g., dye-labeled) oligonucleotide probe containing complementarity to a target cell-free nucleic acid molecule 2603 within a region containing multiple stepwise variants. The molecular beacon may have a length of approximately 25 to 50 nucleotides. The molecular beacon may also be designed to be partially self-complementary, forming a hairpin structure with a stem 2601a and a loop 2601b. The 5' and 3' ends of the molecular beacon probe may have complementary sequences (e.g., approximately 5-6 nucleotides) that form the stem structure 2601a. The loop portion 2601b of the hairpin may be designed to specifically hybridize to a portion of the target sequence (e.g., approximately 15-30 nucleotides) containing two or more stepwise variants. The hairpin may be designed to hybridize to a portion containing at least two, three, four, five or more stepwise variants. A fluorescent reporter molecule can be bound to the 5' end of a molecular beacon probe, and a quencher that quenches the fluorescence of the fluorescent reporter can be bound to the 3' end of the molecular beacon probe. Thus, hairpin formation allows the fluorescent reporter and quencher to be bound together, resulting in no fluorescence emission. However, during the annealing operation of an amplification reaction of multiple cell-free nucleic acid molecules obtained from or derived from a target, the loop portion of the molecular beacon can bind to its target sequence and denature the stem. Thus, the reporter and quencher can be separated, the quench can be neutralized, and the fluorescent reporter can be activated and detected. Since the fluorescence of the fluorescent reporter is emitted from the molecular beacon probe only when the probe is bound to the target sequence, the amount or level of fluorescence detected may be proportional to the amount of target in the reaction (e.g., as disclosed herein, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants in each situation, or (ii) the average number of multiple stepwise variants per cell-free nucleic acid molecule identified as containing multiple stepwise variants).
[0198] In some embodiments of any one of the methods disclosed herein, an activatable reporter agent may be activated upon dehybridization of at least some of the individual nucleic acid probes hybridized into multiple stepwise variants. In other words, when individual nucleic acid probes hybridize into a portion of a target cell-free nucleic acid molecule containing multiple stepwise variants, at least some of the individual nucleic acid probes and target cell-free nucleic acids are activated. Some dehybridizations can activate activatable reporter agents. Non-exclusive examples of such nucleic acid probes may include hydrolysis probes (e.g., TaqMan prob), bihybridization probes, and QZyme PCR primers.
[0199] For example, a nucleic acid probe may be a hydrolysis probe, as shown in Figure 26B. The hydrolysis probe 2611 may be a fluorescently labeled oligonucleotide probe that can specifically hybridize to a portion (e.g., about 10 to about 25 nucleotides) of the target cell-free nucleic acid molecule 2613, the hybridized portion containing two or more stepwise variants. The hydrolysis probe 2611 may be labeled with a fluorescent reporter at its 5' end and a quencher at its 3' end. If the hydrolysis probe is intact (e.g., not cleaved), the fluorescence of the reporter is quenched due to its proximity to the quencher (Figure 26B). During the annealing operation of the amplification reaction of multiple cell-free nucleic acid molecules obtained from or derived from the subject, the 5'→3' exonuclease activity of a specific thermostable polymerase (e.g., Taq or Tth) is increased. The amplification reaction of multiple cell-free nucleic acid molecules obtained from or derived from a subject may involve a combined annealing / extension operation in which a hydrolysis probe hybridizes to the target cell-free nucleic acid molecule, and the dsDNA-specific 5'→3' exonuclease activity of a thermostable polymerase (e.g., Taq or Tth) cleaves the fluorescent reporter from the hydrolysis probe. As a result, the fluorescent reporter is separated from the quencher, producing a fluorescent signal proportional to the amount of target in the sample (e.g., as disclosed herein, (i) the total number of cell-free nucleic acid molecules identified as containing multiple stepwise variants in each situation, or (ii) the average number of multiple stepwise variants per cell-free nucleic acid molecule identified as containing multiple stepwise variants).
[0200] In some embodiments of any one of the methods disclosed herein, the reporter agent may include a fluorescent reporter. Non-limiting examples of fluorescent reporters include fluorescein amidite (FAM, 2-[3-(dimethylamino)-6-dimethyliminioxanthene-9-yl]benzoate TAMRA, (2E)-2-[(2E,4E)-5-(2-tert-butyl-9-ethyl-6,8,8-trimethylpyrano[3,2-g]quinoline-1-ium-4-yl)penta-2,4-dienylidene]-1-(6-hydroxy-6-oxohexyl)-3,3-dimethyl-indoline-5-sulfonate Dy 750, 6-carboxy-2',4,4',5',7,7'-hexachlorofluorescein, 4,5,6,7-tetrachlorofluorescein TET (trademark), sulforhodamine 101 acid chloride succinimidyl ester Texas Red-X, ALEXA Examples include Dyes, Bodipy Dyes, Cyanine Dyes, Rhodamine 123 (hydrochloride), Well RED Dyes, MAX, and TEX 613. In some cases, the reporter agent further comprises a quencher such as those disclosed herein. Non-limiting examples of quenchers include Black Hole Quencher, Iowa Black Quencher, and 4-dimethylaminoazobenzene-4'-sulfonyl chloride (DABCYL).
[0201] In some embodiments of any one of the methods disclosed herein, any PCR reaction utilizing a set of nucleic acid probes can be performed using real-time PCR (qPCR). Alternatively, a PCR reaction utilizing a set of nucleic acid probes can be performed using digital PCR (dPCR).
[0202] Figure 24 provides an example flowchart of the process for implementing clinical interventions and / or treatments based on the detection of circulating tumor nucleic acids in an individual's biological sample. In some embodiments, the detection of circulating tumor nucleic acids is determined by the detection of congenital somatic variants in a cell-free nucleic acid sample. In many embodiments, the detection of circulating tumor nucleic acids indicates the presence of cancer, and therefore appropriate clinical interventions and / or treatments can be taken.
[0203] Referring to Figure 24, process 2400 can begin with obtaining, preparing, and sequencing cell-free nucleic acids obtained from a non-invasive biopsy (e.g., liquid or waste biopsy) using a capture sequencing approach across a region shown to have multiple in-phase gene mutations or variants (2401). In some embodiments, cfDNA and / or cfRNA are extracted from plasma, blood, lymph, saliva, urine, feces, and / or other suitable bodily fluids. Cell-free nucleic acids can be isolated and purified by any suitable means. In some embodiments, column purification is utilized (e.g., QIAamp Circulating Nucleic Acid Kit from Qiagen, Hilden, Germany). In some embodiments, isolated RNA fragments can be converted to complementary DNA for further downstream analysis.
[0204] In some embodiments, the biopsy is extracted before any signs of cancer appear. In some embodiments, the biopsy is extracted to provide an initial screening for detecting cancer. In some embodiments, the biopsy is extracted to detect whether residual cancer is present after treatment. In some embodiments, the biopsy is extracted during treatment to determine whether the treatment is providing the desired response. Screening for any specific cancer can be performed. In some embodiments, the screening is performed to detect cancers that express somatic stepwise variants in typological regions of the genome, such as lymphoma (e.g.). In some embodiments, the screening is performed using a previously extracted cancer biopsy to detect cancers in which somatic stepwise variants have been found.
[0205] In some embodiments, biopsies are drawn from individuals determined to be at risk of developing cancer, such as individuals with a family history of the disorder or those with determined risk factors (e.g., exposure to carcinogens). In many embodiments, biopsies are drawn from any individual within the general population. In some embodiments, biopsies are drawn from individuals within a specific age group with a higher risk of cancer, such as older individuals over 50 years of age. In some embodiments, biopsies are drawn from individuals who have been diagnosed with cancer and treated for cancer.
[0206] In some embodiments, the extracted cell-free nucleic acids are prepared for sequencing. Thus, the cell-free nucleic acids are converted into a molecular library for sequencing. In some embodiments, adapters and / or primers are bound to the cell-free nucleic acids to facilitate sequencing. In some embodiments, targeted sequencing of specific genomic loci should be performed, and thus specific sequences corresponding to specific loci are captured by hybridization before sequencing (e.g., capture sequencing). In some embodiments, capture sequencing is performed using a set of probes that pull down (or capture) regions found to have a common stepwise variant for a particular cancer (e.g., lymphoma). In some embodiments, capture sequencing is performed using a set of probes that pull down (or capture) regions found to have a stepwise variant that has been previously determined by sequencing a cancer biopsy. A more detailed discussion of capture sequencing and probes is provided in the section titled “Capture Sequencing”.
[0207] In some embodiments, any suitable sequencing technique capable of detecting stepwise variants representing circulating tumor nucleic acids can be utilized. Examples of sequencing techniques include, but are not limited to, 454 sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent sequencing, single-read sequencing, and paired-end sequencing.
[0208] Process 2400 analyzes the results of cell-free nucleic acid sequencing (2403) to detect circulating tumor nucleic acid sequences determined by the detection of somatic variants occurring in the same phase. Because cancer is actively growing and expanding, neoplastic cells often release biomolecules (particularly nucleic acids) into the vascular, lymphatic, and / or waste systems. Furthermore, due to biophysical constraints in their local environment, neoplastic cells often rupture, releasing their internal cellular contents into the vascular, lymphatic, and / or waste systems. Therefore, it is possible to detect distal primary tumors and / or metastases from fluid biopsies or waste biopsies.
[0209] Detection of circulating tumor nucleic acid sequences indicates the presence of cancer in the individual being examined. Therefore, clinical interventions and / or treatments can be performed based on the detection of circulating tumor nucleic acids (2405). In some embodiments, clinical procedures such as (e.g.) blood tests, genetic testing, medical imaging, physical examinations, tumor biopsies, or any combination thereof are performed. In some embodiments, diagnoses are performed to determine a specific stage of cancer. In some embodiments, treatments such as (e.g.) chemotherapy, radiotherapy, chemoradiotherapy, immunotherapy, hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusions, medical surveillance, or any combination thereof are performed. In some embodiments, the individual is examined by a doctor, physician, physician's assistant, nurse practitioner, nurse, caregiver. It is evaluated and / or treated by a medical professional such as a nutritionist.
[0210] Various embodiments of this disclosure relate to the use of cancer detection to carry out clinical interventions. In some embodiments, an individual has a fluid or waste biopsy screened and processed by the method herein to indicate that the individual has cancer and therefore should be intervened. Clinical interventions include clinical procedures and treatments. Clinical procedures include, but are not limited to, blood tests, genetic tests, medical imaging, physical examinations, and tumor biopsies. Treatments include, but are not limited to, chemotherapy, radiotherapy, chemoradiotherapy, immunotherapy, hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusions, and medical monitoring. In some embodiments, a diagnosis is made to determine a specific stage of cancer. In some embodiments, an individual is evaluated and / or treated by a healthcare professional such as a physician, doctor, physician's assistant, nurse practitioner, nurse, caregiver, or dietitian.
[0211] In some embodiments described herein, cancer can be detected using sequencing results of cell-free nucleic acids derived from blood, serum, cerebrospinal fluid, lymph, urine, or feces. In many embodiments, cancer is detected if the sequencing results contain one or more somatic variants present in phase within a short gene window, such as the length of the cell-free molecule (e.g., about 170 bp). In numerous embodiments, statistical methods are used to determine whether the presence of stepwise variants originates from a cancerous source (as opposed to molecular artifacts or other biological sources). Various embodiments utilize Monte Carlo sampling as a statistical method to determine whether the sequencing results of cell-free nucleic acids contain sequences of circulating tumor nucleic acids, based on a score determined by the presence of stepwise variants. Thus, in some embodiments, cell-free nucleic acids are extracted, processed, sequenced, and the sequencing results are analyzed to detect cancer. This process is particularly useful in clinical practice to provide diagnostic scans.
[0212] An exemplary procedure for a diagnostic scan of an individual for B-cell carcinoma is as follows: (a) Extract liquid or waste biopsy from a solid, (b) Prepare and perform targeted sequencing of cell-free nucleic acids from biopsy using nucleic acid probes specific to B-cell cancer. (c) Detect stepwise variants in sequencing results showing circulating tumor nucleic acid sequences. (d) Clinical interventions are implemented based on the detection of circulating tumor nucleic acid sequences.
[0213] The following is an exemplary procedure for an individualized diagnostic scan of a previously sequenced cancer to detect stepwise variants at a specific genomic locus: Extract cancer biopsies from individual sequences to detect stepwise variants accumulated in the cancer. (a) Design and synthesize nucleic acid probes for genomic loci that include the location of the detected stepwise variant. (b) Extract liquid or waste biopsy from solid, (c) Prepare and perform targeted sequencing of cell-free nucleic acids from biopsy using designed and synthesized nucleic acid probes. (d) Detect stepwise variants in sequencing results showing circulating tumor nucleic acid sequences. (e) Clinical interventions are implemented based on the detection of circulating tumor nucleic acid sequences.
[0214] In some embodiments of any one of the methods disclosed herein, at least a portion of the identified one or more cell-free nucleic acid molecules containing multiple stepwise variants can be further analyzed to determine the state of the subject. In such analysis, (i) the identified one or more cell-free nucleic acid molecules and (ii) other cell-free nucleic acid molecules of the multiple cell-free nucleic acid molecules that do not contain multiple stepwise variants can be analyzed as different variables. In some cases, (i) the ratio of the number of identified one or more cell-free nucleic acid molecules to (ii) the number of other cell-free nucleic acid molecules of the multiple cell-free nucleic acid molecules that do not contain multiple stepwise variants can be used as a factor for determining the state of the subject. In some cases, a comparison of (i) the position(s) of the identified one or more cell-free nucleic acid molecules relative to a reference genome sequence and (ii) the position(s) of other cell-free nucleic acid molecules of the multiple cell-free nucleic acid molecules that do not contain multiple stepwise variants relative to a reference genome sequence can be used as a factor for determining the state of the subject.
[0215] Alternatively, in some cases, the analysis of one or more identified cell-free nucleic acid molecules containing multiple stepwise variants to determine the state of a subject may not be based on, and may not be based on, other cell-free nucleic acid molecules that do not contain multiple stepwise variants. Non-limiting examples of information or characteristics of one or more cell-free nucleic acid molecules containing multiple stepwise variants, as disclosed herein, may include (i) the total number of such cell-free nucleic acid molecules, and (ii) the average number of multiple stepwise variations per nucleic acid molecule in the population of identified cell-free nucleic acid molecules.
[0216] Accordingly, in some embodiments of any one of the methods disclosed herein, the number of multiple stepwise variants from one or more cell-free nucleic acid molecules identified as having multiple stepwise variants can indicate the state of interest. In some cases, the ratio of (i) the number of multiple stepwise variants from one or more cell-free nucleic acid molecules to (ii) the number of single-nucleotide variants from one or more cell-free nucleic acid molecules can indicate the state of interest. For example, a particular state (e.g., follicular lymphoma) may exhibit a different signature ratio than that of another state (e.g., breast cancer). In some examples, for cancer or solid tumors, the ratios disclosed herein may range from about 0.01 to about 0.20. In some examples, for cancer or solid tumors, the ratios disclosed herein may be about 0.01, about 0.02, about 0.03, about 0.04, about 0.05, about 0.06, about 0.07, about 0.08, about 0.09, about 0.10, about 0.11, about 0.12, about 0.13, about 0.14, about 0.15, about 0.16, about 0.17, about 0.18, about 0.19, or about 0.20. In some examples, for cancer or solid tumors, the ratios disclosed herein may be at least or up to about 0.01, at least or up to about 0.02, at least or up to about 0.03, at least or up to about 0.04, at least or up to about 0.05, at least or up to about 0.06, at least or up to about 0.07, at least or up to about 0.08, at least or up to about 0.09, at least or up to about 0.10, at least or up to about 0.11, at least or up to about 0.12, at least or up to about 0.13, at least or up to about 0.14, at least or up to about 0.15, at least or up to about 0.16, at least or up to about 0.17, at least or up to about 0.18, at least or up to about 0.19, or at least or up to about 0.20.
[0217] In some embodiments of any one of the methods disclosed herein, the frequency of multiple stepwise variants in one or more identified cell-free nucleic acid molecules may indicate the state of the subject. In some cases, based on the sequencing data disclosed herein, the average frequency of multiple stepwise variants per given bin length (e.g., a bin of about 50 base pairs) in each of the identified cell-free nucleic acid molecules may indicate the state of the subject. In some cases, based on the sequencing data disclosed herein, the average frequency of multiple stepwise variants per given bin length (e.g., a bin of about 50 base pairs) in each of the identified cell-free nucleic acid molecules associated with a particular gene (e.g., BCL2, PIM1) may indicate the state of the subject. The bin size may be about 30, about 40, about 50, about 60, about 70, or about 80.
[0218] In some cases, a first condition (e.g., Hodgkin lymphoma or HL) may exhibit a first mean frequency, and a second condition (e.g., DLBCL) may exhibit a different mean frequency, thereby enabling the identification and / or determination of whether a subject has or is suspected of having a particular condition. In some cases, a first subtype of a disease may exhibit a first mean frequency, and a second subtype of the same disease may exhibit a different mean frequency, thereby enabling the identification and / or determination of whether a subject has or is suspected of having a particular subtype of the disease. For example, a subject may have DLBCL, and as disclosed herein, one or more cell-free nucleic acid molecules derived from germinal center B cell (GCB) DLBCL or activated B cell (ABC) DLBCL may have different mean frequencies of multiple stepwise variants per given bin length.
[0219] In some cases, a subject's condition may have a predetermined number of stepwise variants (i.e., a predetermined frequency of stepwise variants) spanning a given genomic locus. If the predetermined frequency of stepwise variants matches the frequency of multiple stepwise variants in one or more cell-free nucleic acid molecules identified from multiple cell-free nucleic acid molecules from the subject, it may indicate that the subject has such a condition.
[0220] In some embodiments of any one of the methods disclosed herein, one or more cell-free nucleic acid molecules identified as containing multiple stepwise variants can be analyzed to determine their genomic origin (e.g., which locus they originate from). Since different diseases may have multiple stepwise variants in different signature genes, the genomic origin of the identified one or more cell-free nucleic acid molecules can indicate the condition of the subject. For example, a subject may have GCB DLBCL, and one or more cell-free nucleic acid molecules derived from the subject's GCB may have a dominant stepwise variant in the BCL2 gene, while one or more cell-free nucleic acid molecules derived from the same subject's ABC may not contain as many stepwise variants in the BCL2 gene as those derived from GCB. Conversely, a subject may have ABC DLBCL, and one or more cell-free nucleic acid molecules derived from the subject's ABC may have a dominant stepwise variant in the PIM1 gene, while one or more cell-free nucleic acid molecules derived from the same subject's GCB may not contain as many stepwise variants in the PIM1 gene as those derived from ABC.
[0221] In some embodiments of any one of the methods disclosed herein, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, at least or up to about 50%, at least or up to about 55%, at least or up to about 60%, at least or up to about 65%, at least or up to about 70%, at least or up to about 75%, at least or up to about 80%, at least or up to about 85%, at least or up to about 90%, at least or up to about 95%, at least or up to about 99%, or about 100% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain a single nucleotide variant (SNV) located at least 2 nucleotides away from an adjacent SNV.
[0222] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 3 nucleotides away from adjacent SNVs.
[0223] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 4 nucleotides away from adjacent SNVs.
[0224] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 5 nucleotides away from adjacent SNVs.
[0225] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 6 nucleotides away from adjacent SNVs.
[0226] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 7 nucleotides away from adjacent SNVs.
[0227] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 8 nucleotides away from adjacent SNVs.
[0228] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 9 nucleotides away from adjacent SNVs.
[0229] In some embodiments of any one of the methods disclosed herein, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 25%, at least or up to about 30%, at least or up to about 35%, at least or up to about 40%, at least or up to about 45%, or at least or up to about 50% of one or more cell-free nucleic acid molecules containing multiple stepwise variants may contain single nucleotide variants (SNVs) located at least 10 nucleotides away from adjacent SNVs. C. Reference genome sequence
[0230] In some embodiments of any one of the methods disclosed herein, the reference genome sequence may be at least part of a nucleic acid sequence database (i.e., a reference genome), which is constructed from genetic data and intended to represent the genome of a reference cohort. In some cases, the reference cohort may be a collection of individuals from a specific or diverse genotype, haplotype, demographic, sex, nationality, age, ethnicity, kinship, health condition (e.g., healthy, or diagnosed with the same or different condition, e.g., having a particular type of cancer), or other grouping. The reference genome sequences disclosed herein may be mosaics (or consensus sequences) of genomes from two or more individuals. The reference genome sequence may include at least part of a publicly available reference genome or an informal reference genome. Non-limiting examples of human reference genomes include hg19, hg18, hg17, hg16, and hg38.
[0231] In some cases, the reference genome sequence contains at least or up to approximately 500 nucleic acid bases, at least or up to approximately 1 kilobase (kb), at least or up to approximately 2 kb, at least or up to approximately 3 kb, at least or up to approximately 4 kb, at least or up to approximately 5 kb, at least or up to approximately 6 kb, at least or up to approximately 7 kb, at least or up to approximately 8 kb, at least or up to approximately 9 kb, at least or up to approximately 10 kb, at least or up to approximately 20 kb, at least or up to approximately 30 kb, at least or up to approximately 40 kb, at least or up to approximately 50 kb, at least or up to approximately 60 kb, at least or up to approximately 70 kb, at least or up to approximately 80 kb, at least or up to approximately 90 kb, at least or up to approximately 100 kb, at least or up to approximately 200 kb, at least or up to approximately 300 kb, at least or up to approximately 400 kb, at least or up to approximately 500 kb, at least or up to approximately 600 kb, and less None or up to approximately 700kb, at least or up to approximately 800kb, at least or up to approximately 900kb, at least or up to approximately 1,000kb, at least or up to approximately 2,000kb, at least or up to approximately 3,000kb, at least or up to approximately 4,000kb, at least or up to approximately 5,000kb, at least or up to approximately 6,000kb, at least or up to approximately 7,000kb, at least or up to approximately 8,000kb, at least or up to approximately 9,0 It may include 00kb, at least or up to approximately 10,000kb, at least or up to approximately 20,000kb, at least or up to approximately 30,000kb, at least or up to approximately 40,000kb, at least or up to approximately 50,000kb, at least or up to approximately 60,000kb, at least or up to approximately 70,000kb, at least or up to approximately 80,000kb, at least or up to approximately 90,000kb, or at least or up to approximately 100,000kb.
[0232] In some cases, a reference genome sequence can be the entire reference genome or a portion of the genome (e.g., a portion related to the desired condition). For example, a reference genome sequence may consist of at least one, two, three, four, five, or more genes that experience abnormal somatic hypermutation under a particular type of cancer. In some cases, a reference genome sequence can be an entire chromosome sequence or a fragment thereof. In some cases, a reference genome sequence may consist of two or more (e.g., at least two, three, four, five, or more) different portions of the reference genome that are not adjacent to each other (e.g., within the same chromosome or from different chromosomes).
[0233] In some embodiments of any one of the methods disclosed herein, the reference genome sequence may be at least a portion of the reference genome of a selected individual, such as a healthy individual, or of the subject of any of the methods disclosed herein.
[0234] In some cases, the reference genome sequence may originate from an individual that is not the subject (e.g., a healthy control individual). Alternatively, in some cases, the reference genome sequence may originate from a sample of the subject. In some examples, the sample may be a healthy sample of the subject. A healthy sample of the subject may be any healthy subject cell, e.g., a healthy leukocyte. By comparing sequencing data of multiple cell-free nucleic acid molecules (e.g., cfDNA molecules) of the subject with the genome sequences of at least a portion of healthy cells of the same subject, one or more cell-free nucleic acid molecules containing multiple stepwise variants can be identified and analyzed, as disclosed herein. In some examples, the sample may be diseased cells (e.g., tumor cells) or diseased samples of the subject, such as a solid tumor. The reference genome sequence can be obtained from sequencing at least a portion of diseased cells of the subject, or from sequencing multiple cell-free nucleic acid molecules obtained from a solid tumor of the subject. Once the subject is diagnosed with a particular condition (e.g., disease), the reference genome sequence of the subject containing multiple stepwise variants can be used to determine whether the subject will still exhibit the same stepwise variants at a future point in time. In this regard, any novel stepwise variants identified between the "affected" reference genome sequence of the subject and novel cell-free nucleic acid molecules obtained from or derived from the subject may exhibit abnormal somatic hypermutation, particularly a reduction in the degree of genomic loci (e.g., at least partial remission).
[0235] In various embodiments, the diagnostic scan can detect acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), anal cancer, astrocytoma, basal cell carcinoma, cholangiocarcinoma, bladder cancer, breast cancer, Burkitt lymphoma, cervical cancer, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myeloproliferative neoplasms, colorectal cancer, diffuse large B-cell lymphoma, endometrial cancer, ependymoma, esophageal cancer, sensory neuroblastoma, Ewing's sarcoma, fallopian tube cancer, follicular lymphoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumors, hairy cell leukemia, hepatocellular carcinoma, Hodgkin lymphoma, hypopharyngeal cancer, and kapo It can be performed for any type of neoplasm, including but not limited to disarcoma, renal cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, Merkel cell carcinoma, mesothelioma, oral cancer, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumor, pharyngeal cancer, pituitary tumor, prostate cancer, rectal cancer, renal cell carcinoma, retinoblastoma, skin cancer, small cell lung cancer, small intestine cancer, squamous cell carcinoma of the neck, T-cell lymphoma, testicular cancer, thymoma, thyroid cancer, uterine cancer, vaginal cancer, and hemangiomas.
[0236] In some embodiments, diagnostic scans are used to provide early detection of cancer. In some embodiments, diagnostic scans detect cancer in individuals with stage I, II, or III cancer. In some embodiments, diagnostic scans are used to detect MRD or tumor burden. In some embodiments, diagnostic scans are used to determine the progression of treatment (e.g., progression or regression). Clinical procedures and / or treatments can be performed based on the diagnostic scan. D. Nucleic acid probes
[0237] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes can be designed based on any of the reference genome sequences of interest disclosed herein. In some cases, a set of nucleic acid probes can be designed based on a plurality of stepwise variants identified by comparing (i) sequencing data from a solid tumor of interest with (ii) sequencing data from healthy cells of interest or a healthy cohort, as disclosed herein. A set of nucleic acid probes can be designed based on a plurality of stepwise variants identified by comparing (i) sequencing data from a solid tumor of interest with (ii) sequencing data from healthy cells of interest. A set of nucleic acid probes can be designed based on a plurality of stepwise variants identified by comparing (i) sequencing data from a solid tumor of interest with (ii) sequencing data from healthy cells of a healthy cohort.
[0238] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes is designed to hybridize to the sequence of a state-associated genomic locus. As disclosed herein, if the subject has a state, the state-associated genomic locus may be determined to experience or exhibit abnormal somatic hypermutation. Alternatively, a set of nucleic acid probes is designed to hybridize to the sequence of a typological region.
[0239] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes may be designed to hybridize to at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or about 100% of the genomic regions identified in Table 1.
[0240] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes may be designed to hybridize to at least a portion of cell-free nucleic acid (e.g., cfDNA) molecules derived from at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or about 100% of the genomic regions identified in Table 1.
[0241] In some embodiments of any one of the methods disclosed herein, each nucleic acid probe in a set of nucleic acid probes may have at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or about 100% sequence identity with a probe sequence selected from Table 6.
[0242] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes may comprise at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or about 100% of the probe sequences in Table 6.
[0243] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes comprises at least or up to about 500 nucleic acid bases, at least or up to about 1 kilobase (kb), at least or up to about 2 kb, at least or up to about 3 kb, at least or up to about 4 kb, at least or up to about 5 kb, at least or up to about 6 kb, at least or up to about 7 kb, at least or up to about 8 kb, at least or up to about 9 kb, at least or up to about 10 kb, at least or up to about 20 kb b. It may be designed to cover one or more target genomic regions including at least or up to approximately 30kb, at least or up to approximately 40kb, at least or up to approximately 50kb, at least or up to approximately 60kb, at least or up to approximately 70kb, at least or up to approximately 80kb, at least or up to approximately 90kb, at least or up to approximately 100kb, at least or up to approximately 200kb, at least or up to approximately 300kb, at least or up to approximately 400kb, or at least or up to approximately 500kb.
[0244] In some embodiments of any one of the methods disclosed herein, one or more target genomic regions (e.g., target genomic loci) may be up to about 200 nucleic acid bases, up to about 300 nucleic acid bases, 400 nucleic acid bases, up to about 500 nucleic acid bases, up to about 600 nucleic acid bases, up to about 700 nucleic acid bases, up to about 800 nucleic acid bases, up to about 900 nucleic acid bases, up to about 1 kb, up to about 2 kb, up to about 3 kb, up to about 4 kb, and up to It may include approximately 5kb, up to approximately 6kb, up to approximately 7kb, up to approximately 8kb, up to approximately 9kb, up to approximately 10kb, up to approximately 11kb, up to approximately 12kb, up to approximately 13kb, up to approximately 14kb, up to approximately 15kb, up to approximately 16kb, up to approximately 17kb, up to approximately 18kb, up to approximately 19kb, up to approximately 20kb, up to approximately 25kb, up to approximately 30kb, up to approximately 35kb, up to approximately 40kb, up to approximately 45kb, up to approximately 50kb, or up to approximately 100kb.
[0245] In some embodiments of any one of the methods disclosed herein, a set of nucleic acid probes may include at least or up to about 10, at least or up to about 20, at least or up to about 30, at least or up to about 40, at least or up to about 50, at least or up to about 60, at least or up to about 70, at least or up to about 80, at least or up to about 90, at least or up to about 100, at least or up to about 200, at least or up to about 300, at least or up to about 400, at least or up to about 500, at least or up to about 600, at least or up to about 700, at least or up to about 800, at least or up to about 900, at least or up to about 1,000, at least or up to about 2,000, at least or up to about 3,000, at least or up to about 4,000, or at least or up to about 5,000 different nucleic acid probes designed to hybridize to different target nucleic acid sequences.
[0246] In some embodiments of any one of the methods disclosed herein, the set of nucleic acid probes may have a length of at least or up to about 50, at least or up to about 55, at least or up to about 60, at least or up to about 65, at least or up to about 70, at least or up to about 75, at least or up to about 80, at least or up to about 85, at least or up to about 90, at least or up to about 95, or at least or up to about 100 nucleotides.
[0247] In one embodiment, the disclosure provides a composition comprising a bait set containing one of the sets of nucleic acid probes disclosed herein. Such a composition comprising a bait set can be used in any of the methods disclosed herein. In some cases, the set of nucleic acid probes may be designed to pull down (or capture) cfDNA molecules. In some cases, the set of nucleic acid probes may be designed to pull down (or capture) cfRNA molecules.
[0248] In some embodiments, the bait set may include a set of nucleic acid probes designed to pull down cell-free nucleic acid (e.g., cfDNA) molecules derived from the genomic regions specified in Table 1. The set of nucleic acid probes may include at least or up to approximately 1%, at least or up to approximately 2%, at least or up to approximately 3%, at least or up to approximately 4%, at least or up to approximately 5%, at least or up to approximately 6%, at least or up to approximately 7%, at least or up to approximately 8%, at least or up to approximately 9%, at least or up to approximately 10%, at least or up to approximately 15%, at least or up to approximately 20%, at least or up to approximately 25%, at least or up to approximately 30%, at least or up to approximately The nucleic acid probe set can be designed to pull down cell-free nucleic acid molecules derived from 35%, at least or up to approximately 40%, at least or up to approximately 45%, at least or up to approximately 50%, at least or up to approximately 55%, at least or up to approximately 60%, at least or up to approximately 65%, at least or up to approximately 70%, at least or up to approximately 75%, at least or up to approximately 80%, at least or up to approximately 85%, at least or up to approximately 90%, at least or up to approximately 95%, at least or up to approximately 99%, or approximately 100%. In some cases, the nucleic acid probe set can be designed to pull down cfDNA molecules. In some cases, the nucleic acid probe set can be designed to pull down cfRNA molecules.
[0249] In some embodiments of any one of the compositions disclosed herein, individual nucleic acid probes (or each nucleic acid probe) of a set of nucleic acid probes may include pull-down tags. Pull-down tags can be used to enrich a sample (e.g., a sample containing multiple nucleic acid molecules obtained from or derived from a subject) of a particular subset (e.g., a cell-free nucleic acid molecule containing multiple stepwise variants, as disclosed herein).
[0250] In some cases, the pull-down tag may include a nucleic acid barcode (e.g., on one or both sides of a nucleic acid probe). By utilizing beads or substrates containing nucleic acid sequences complementary to the nucleic acid barcode, any nucleic acid probe that hybridizes to a target cell-free nucleic acid molecule can be pulled down and enriched using the nucleic acid barcode. Alternatively, or in addition to the above, the target cell-free nucleic acid molecule can be identified from any sequencing data (e.g., sequencing by amplification) obtained by using any of the sets of nucleic acid probes disclosed herein.
[0251] In some cases, a pull-down tag may contain an affinity target moiety that can be specifically recognized and bound by an affinity binding moiety. The affinity binding moiety can specifically bind to the affinity target moiety to form an affinity pair. In some cases, by utilizing beads or substrates containing an affinity binding moiety, any nucleic acid probe that hybridizes to a target cell-free nucleic acid molecule can be pulled down and enriched using the affinity target moiety. Alternatively, the pull-down tag may contain an affinity binding moiety, while the beads / substrate may contain an affinity target moiety. Non-limiting examples of affinity pairs include biotin / avidin, antibody / antigen, biotin / streptavidin, metal / chelator, ligand / receptor, nucleic acids and binding proteins, and complementary nucleic acids. In one example, the pull-down tag may contain biotin.
[0252] In some embodiments of any of the compositions disclosed herein, the length of the target cell-free nucleic acid (e.g., cfDNA) molecule pulled down by any target nucleic acid probe may be about 100 to about 200 nucleotides. The length of the target cell-free nucleic acid molecule may be at least about 100 nucleotides. The length of the target cell-free nucleic acid molecule may be up to about 200 nucleotides. The lengths of the target cell-free nucleic acid molecules are approximately 100 to 110 nucleotides, 100 to 120 nucleotides, 100 to 130 nucleotides, 100 to 140 nucleotides, 100 to 150 nucleotides, 100 to 160 nucleotides, 100 to 170 nucleotides, 100 to 180 nucleotides, 100 to 190 nucleotides, 100 to 200 nucleotides, 110 to 120 nucleotides, 110 to 130 nucleotides, 110 to 140 nucleotides, 110 to 150 nucleotides, 110 to 160 nucleotides, 110 to 170 nucleotides, 110 to 180 nucleotides, and 110 to 190 nucleotides. Nucleotides, approximately 110 nucleotides to approximately 200 nucleotides, approximately 120 nucleotides to approximately 130 nucleotides, approximately 120 nucleotides to approximately 140 nucleotides, approximately 120 nucleotides to approximately 150 nucleotides, approximately 120 nucleotides to approximately 160 nucleotides, approximately 120 nucleotides to approximately 170 nucleotides, approximately 120 nucleotides to approximately 180 nucleotides, approximately 120 nucleotides to approximately 190 nucleotides, approximately 120 nucleotides to approximately 200 nucleotides, approximately 130 nucleotides to approximately 140 nucleotides, approximately 130 nucleotides to approximately 150 nucleotides, approximately 130 nucleotides to approximately 160 nucleotides, approximately 130 nucleotides to approximately 170 nucleotides, approximately 130 nucleotides to approximately 180 nucleotides, approximately 130 nucleotides to approximately 190 nucleotides, approximately 130 nucleotides to approximately 200 nucleotides, approximately 140 nucleotides to approximately 150 nucleotides, approximately 140 nucleotides to approximately 160 nucleotides,It may be approximately 140 to 170 nucleotides, approximately 140 to 180 nucleotides, approximately 140 to 190 nucleotides, approximately 140 to 200 nucleotides, approximately 150 to 160 nucleotides, approximately 150 to 170 nucleotides, approximately 150 to 180 nucleotides, approximately 150 to 190 nucleotides, approximately 150 to 200 nucleotides, approximately 160 to 170 nucleotides, approximately 160 to 180 nucleotides, approximately 160 to 190 nucleotides, approximately 160 to 200 nucleotides, approximately 170 to 180 nucleotides, approximately 170 to 190 nucleotides, approximately 170 to 200 nucleotides, approximately 180 to 190 nucleotides, approximately 180 to 200 nucleotides, or approximately 190 to 200 nucleotides. The length of the target cell-free nucleic acid molecule may be approximately 100 nucleotides, 110 nucleotides, 120 nucleotides, 130 nucleotides, 140 nucleotides, 150 nucleotides, 160 nucleotides, 170 nucleotides, 180 nucleotides, 190 nucleotides, or 200 nucleotides. In some examples, the length of the target cell-free nucleic acid molecule may range from approximately 100 nucleotides to approximately 180 nucleotides.
[0253] In some embodiments of any one of the compositions disclosed herein, the genomic region may be condition-related. The genomic region can be determined to exhibit abnormal somatic hypermutation if the subject has that condition. For example, the condition may include B-cell lymphoma or its subtypes, such as diffuse large B-cell lymphoma, follicular lymphoma, Burkitt lymphoma, and B-cell chronic lymphocytic leukemia. Further details of the conditions are provided below.
[0254] In some embodiments of any of the compositions disclosed herein, the composition further comprises a plurality of cell-free nucleic acid (e.g., cfDNA) molecules obtained from or derived from a subject. E. Diagnostic or therapeutic use
[0255] Some embodiments involve performing a diagnostic scan on an individual's cell-free nucleic acid, and then, based on the scan results indicating cancer, performing further clinical procedures and / or treating the individual. According to various embodiments, numerous types of neoplasms can be detected.
[0256] In some embodiments of any one of the methods disclosed herein, the method may include determining whether a subject has a state, or determining the degree or state of a state, based on one or more cell-free nucleic acid molecules containing multiple stepwise variants. In some cases, the method may further include determining, based on statistical model analysis (i.e., molecular analysis), that one or more cell-free nucleic acid molecules (each identified as containing multiple stepwise variants) originate from a sample associated with the state (e.g., cancer). For example, the method may include using one or more algorithms (e.g., Monte Carlo simulation) to determine a first probability (e.g., 80%) that the cell-free nucleic acid identified as having multiple stepwise variants is associated with or originates from a first state, and a second probability (e.g., 20%) that the same cell-free nucleic acid is associated with or originates from a second state (or from healthy cells). In some cases, the method may involve determining the likelihood or probability that a subject has one or more states based on an analysis of one or more identified cell-free nucleic acid molecules, each of which contains multiple stepwise variants (i.e., macro-analysis or global analysis). For example, the method may involve analyzing multiple cell-free nucleic acid molecules, each of which has been identified as containing multiple stepwise variants, using one or more algorithms (including one or more mathematical models disclosed herein, such as binomial sampling), thereby determining a first probability (e.g., 80%) that the subject has a first state and a second probability (e.g., 20%) that the subject has a second state (or is healthy).
[0257] The statistical model analyses disclosed herein may be approximate solutions obtained through numerical approximations such as binomial models, ternary models, Monte Carlo simulations, and finite difference methods. In one example, the statistical model analysis used herein may be a Monte Carlo statistical analysis. In another example, the statistical model analysis used herein may be a binomial model analysis or a ternary model analysis.
[0258] In some embodiments of any one of the methods disclosed herein, the method may include monitoring the progression of a condition in question based on one or more identified cell-free nucleic acid molecules, such that each identified cell-free nucleic acid molecule includes a plurality of stepwise variants. In some cases, the progression of the condition may be a deterioration of the condition (e.g., progression from stage I cancer to stage III cancer), as described in this disclosure. In some cases, the progression of the condition may be at least partial remission of the condition (e.g., downstaging from stage IV cancer to stage II cancer), as described in this disclosure. Or, in some cases, the progression of the condition may remain substantially the same between two different points in time, as described in this disclosure. In one example, the method may include determining the likelihood or probability of different progressions of the condition in question. For example, this method may involve using one or more algorithms (including one or more mathematical models disclosed herein, such as binomial sampling) to determine a first probability (e.g., 20%) that the condition of the subject is worse than before, a second probability (e.g., 70%) that the condition is at least partially improved, and a third probability (e.g., 10%) that the condition of the subject is the same as before.
[0259] In some embodiments of any one of the methods disclosed herein, the method may include performing different procedures (e.g., follow-up diagnostic procedures) to confirm the condition of the subject, as provided herein, that the condition has been determined, and / or its progression has been monitored. Non-limiting examples of different procedures may include physical examination, medical imaging, genetic testing, mammography, endoscopy, stool sampling, Pap tests, alpha-fetoprotein blood tests, CA-125 tests, prostate-specific antigen (PSA) tests, biopsy extractions, bone marrow aspirations, and tumor marker detection tests. Medical imaging includes, but is not limited to, X-ray, magnetic resonance imaging (MRI), computed tomography (CT), ultrasound, and positron emission tomography (PET). Endoscopy includes, but is not limited to, bronchoscopy, colonoscopy, vaginoscopy, cystoscopy, esophagoscopy, gastroscopy, laparoscopy, neuroendoscopy, proctoscopy, and sigmoidoscopy.
[0260] In some embodiments of any one of the methods disclosed herein, the method may include determining a treatment for a condition of a subject based on one or more identified cell-free nucleic acid molecules, each identified cell-free nucleic acid molecule comprising multiple stepwise variants. In some cases, the treatment may be determined based on (i) the determined condition of the subject and / or (ii) the progression of the determined condition of the subject. Furthermore, the treatment may be determined based on one or more additional factors of the subject's sex, nationality, age, ethnicity, and other physical condition. In some examples, the treatment may be determined based on one or more characteristics of multiple stepwise variants of the identified cell-free nucleic acid molecule, as disclosed herein.
[0261] In some embodiments of any one of the methods disclosed herein, the subject may not have received any treatment for the condition, for example, the subject may not have been diagnosed with the condition (e.g., lymphoma). In some embodiments of any one of the methods disclosed herein, the subject may be subjected to treatment for the condition prior to any subjective method of this disclosure. In some cases, the methods disclosed herein can be performed to monitor the progression of the diagnosed condition in the subject, thereby (i) determining the effectiveness of previous treatments, and (ii) evaluating whether to maintain, modify, or discontinue treatment in favor of a new treatment.
[0262] In some embodiments of any one of the methods disclosed herein, non-limiting examples of treatments (e.g., pretreatment, new treatment determined based on the method of this disclosure) may include chemotherapy, radiotherapy, chemoradiotherapy, immunotherapy, adoptive cell therapy (e.g., chimeric antigen receptor (CAR) T cell therapy, CAR NK cell therapy, modified T cell receptor (TCR) T cell therapy, etc.), hormone therapy, targeted drug therapy, surgery, transplantation, blood transfusion, or medical surveillance therapy.
[0263] In some embodiments of any one of the methods disclosed herein, the condition may include a disease. In some embodiments of any one of the methods disclosed herein, the condition may include a neoplasm, cancer, or tumor. In one example, the condition may include a solid tumor. In another example, the condition may include a lymphoma, such as B-cell lymphoma (BCL). Non-limiting examples of BCL may include diffuse large B-cell lymphoma (DLBCL), follicular lymphoma (FL), Burkitt lymphoma (BL), chronic B-cell lymphocytic leukemia (CLL), marginal zone B-cell lymphoma (MZL), and mantle cell lymphoma (MCL).
[0264] As disclosed herein, treatment of a condition of a subject may include administering one or more therapeutic agents to the subject. One or more therapeutic agents may be administered to the subject by one or more of the following: orally, intraperitoneally, intravenously, intraarterially, percutaneously, intramuscularly, via liposomes, subcutaneously, intrafatally, and intrathecally, via local delivery by catheter or stent.
[0265] Non-limiting examples of therapeutic agents include cytotoxic agents, chemotherapeutic agents, growth inhibitors, agents used in radiotherapy, anti-angiogenic agents, apoptotic agents, antitubulin agents, and other agents for treating cancer, such as anti-CD20 antibodies, anti-PD1 antibodies (e.g., pembrolizumab), platelet-derived growth factor inhibitors (e.g., GLEEVEC® (imatinib mesylate)), COX-2 inhibitors (e.g., celecoxib), interferons, cytokines, antagonists (e.g., neutralizing antibodies) that bind to one or more of the following targets: PDGFR-β, BlyS, APRIL, BCMA receptor, TRAIL / Apo2, other bioactive and organic chemical agents.
[0266] Non-limiting examples of cytotoxic agents include radioisotopes (e.g., At211, I131, I125, Y90, Re186, Re188, Sm153, Bi212, P32, and radioisotopes of Lu), chemotherapeutic agents (e.g., methotrexate, adriamycin, vinca alkaloids (vincristine, vinblastine, etoposide), doxorubicin, melphalan, mitomycin C, chlorambucil, daunorubicin or other inserts), enzymes and their fragments (e.g., nucleases), antibiotics, and toxins (e.g., small molecule toxins or enzymatically active toxins of bacterial, fungal, plant, or animal origin).
[0267] Non-exclusive examples of chemotherapeutic agents include alkylating agents such as thiotepa and CYTOXAN® cyclophosphamide, alkyl sulfonates such as busulfan, improsulfan and piposulfan; aziridines such as benzodopa, carbocon, metsuredopa, and uredopa; altretamine, triethylenemelamine, triethylenephosphoramide, triethylenethiophosphoramide, and trime Ethyleneimines and methylamelamine containing tyrolmelamine; acetogenins (especially bratacin and bratacinone); delta-9-tetrahydrocannabinol (dronabinol, MARINOL®); beta-lapacone; lapachol; colchicine; betulinic acid; camptothecin (including synthetic analogs topotecan (HYCAMTIN®), CPT-11 (irinotecan, CAMPTOSAR®), acetylcamptothecin, scopolectin, and 9-aminocamptothecin); bryostatin; calistatin; CC-1065 (including its synthetic analogs adzeresin, karzeresin, and bizeresin); podophyllotoxin; podophyllic acid; teniposide; cryptophycin (especially cryptophycin 1 and cryptophycin 8) ); Dorastatin; Duocalmycin (including synthetic analogs, KW-2189 and CB1-TM1); Eleuterobin; Pancratistatin; Sarcodictiin; Spongestatin; Chlorambucil, Chlornafadin, Cyclophosphamide, Estramustine, Ifosfamide, Mechloretamine, Mechloretamine Oxide Hydrochloride, Melphalan, Nobenbitin, Fenesterine, Prednimustine, Trophosphamide, Nitrogen Mustards such as Uracil Mustard; Nitrosoureas such as Carmustine, Chlorozotosine, Fotemustine, Lomustine, Nimustine, Ranimnustine; Antibiotics such as Endiyne antibiotics; Dynemycin (including Dynemycin A); Espiramicina;Similarly, neocardinostatin chromophore and related pigment protein enediin antibiotic chromophore), acrasinomycin, actinomycin, anthramycin, azaserin, bleomycin, kactinomycin, carabicin, carminomycin, cardinophilin, chromomycin, dactinomycin, daunorubicin, detrevicin, 6-diazo-5-oxo-L-norleucine, ADRIAMYCIN® doxorubicin (morpholino-doxorubicin, cyanomorpholino-doxorubicin, 2-pyrrolino-doxorubicin) Mitomycins such as malceromycin, mitomycin C, mycophenolic acid, nogaramycin, olibomycin, peplomycin, potophyllomycin, puromycin, keramycin, rhodorubicin, streptonigrin, streptozocin, tubercidine, ubenimex, dinostatin, zorubicin; antimetabolites such as methotrexate and 5-fluorouracil (5-FU); denopterin, methotrexate, pteropterin, trimethrexate Folic acid analogs such as; purine analogs such as fludarabine, 6-mercaptopurine, thiamiprine, and thioguanine; pyrimidine analogs such as ancitabine, azacitidine, 6-azauridine, carmofur, cytarabine, dideoxyuridine, doxifluridine, enocitabine, and phloxuridine; androgens, such as carsterone, dromostanolone propionate, epithiostanol, mepitiostane, and testolactone; antiadrenergic agents, such as aminoglutethimide, mitotane, and trilostane; folic acid supplements such as folic acid; acegraton; and aldofos Famide glycoside; aminolevulinic acid; enyluracil; amsacrin; bestrabusil; bisanthren; edatraxate; defofamin; demecolsin; diazicon; eflornithine; eriptinium acetate; epotilon; etogluside; gallium nitrate; hydroxyurea; lentinan; ronidynin; meitansinoids such as meitansin and anthamitosin; mitogluazone; mitoxantrone; mopidamol; nitraerine; pentostatin; fenamet; pirarubicin; losoxantrone; 2-ethylhydrazide;Procarbazine; PSK® polysaccharide complex (JHS Natural Products, Eugene, Oreg.); Lazoxane; Rhizoxin; Schizophyllan; Spirogermanium; Tenuazonic acid; Triadicone; 2,2',2''-Trichlorotriethylamine; Trichothecene (especially T-2 toxin, bergalin A, loridine A and anguidine); Urethane; Vindesine (ELD; ISINE®, FILDESIN®); dacarbazine; mannomustine; mitobronitol; mitractol; pipobromane; gasitosine; arabinoside ("Ara-C"); thiotepa; taxoids, e.g., TAXOL® paclitaxel (Bristol-Myers Squibb Oncology, Princeton, NJ), albumin-modified nanoparticle formulations of paclitaxel without ABRAXANE® cremofol (American Pharmaceutical Partners, Schaumberg, III.), and TAXOTERE® docetaxel (Rhone-Poulenc). Taxanes (including Rorer, Antony, France); chlorambucil; gemcitabine (GEMZAR®); 6-thioguanine; mercaptopurine; methotrexate; platinum analogs such as cisplatin and carboplatin; vinblastine (VELBAN®); platinum; etoposide (VP-16); ifosfamide; mitoxantrone; vincristine (ONCOVIN®); oxaliplatin; leucobobin; vinorelbine (NAVELBINE®); novantrone; edatrexate; daunomycin; aminopterin; ibandronate; Examples of combinations of the above include topoisomerase inhibitors RFS2000; difluoromethylornithine (DMFO); retinoids such as retinoic acid; capecitabine (XELODA®); any pharmaceutically acceptable salts, acids, or derivatives of the above; and two or more combinations of the above, such as CHOP, an abbreviation for combination therapy of cyclophosphamide, doxorubicin, vincristine, and prednisolone, and FOLFOX, an abbreviation for treatment regimens using oxaliplatin (ELOXATIN®) in combination with 5-FU and leucovorin.
[0268] Examples of chemotherapy agents may also include “anti-hormone agents” or “endocrine therapies” that act to modulate, reduce, block or inhibit the effects of hormones that can promote cancer growth and are often systemic or in the form of systemic treatment. They may be hormones themselves. Examples include anti-estrogen drugs and selective estrogen receptor modulators (SERMs), such as tamoxifen (including NOLVADEX® tamoxifen), EVISTA® raloxifene, droloxifene, 4-hydroxytamoxifen, trioxyfen, keoxyfen, LY117018, onapristone and FARESTON® toremifene; anti-progesterone agents; estrogen receptor down regulators (ERDs); and drugs that function to suppress or shut down the ovaries, such as LUPRON (registered trademark). Luteinizing hormone-releasing hormone (LHRH) agonists such as leuprolide acetate (T) and ELIGARD), goserelin acetate, buserelin acetate, and triptorelin; other antiandrogens such as flutamide, nilutamide, and bicalutamide; and aromatase inhibitors that inhibit aromatase, an enzyme that regulates estrogen production in the adrenal gland, such as 4(5)-imidazole, aminoglutethimide, MEGASE® megestrol acetate, AROMASIN® exemestane, formestanie, fadrozol, and RIVISO. Such definitions of chemotherapeutic agents include bisphosphonates, e.g., clodronate (e.g., BONEFOS® or OSTAC®), DIDROCAL® etidronate, NE-58095, ZOMETA® zoledronic acid / zoledronate, FOSAMAX® alendronate, AREDIA® pamidronate, SKELID® tildronate, or ACTONEL® risedronate, as well as troxacitabine (1,3-dioxolane nucleoside cytosine analog); antisense oligonucleotides, in particular, e.g., PKC-alpha, Raf, H-Ras, and epidermal growth This includes inhibitors of gene expression in signaling pathways involved in abherant cell proliferation, such as factor receptor (EGFR); vaccines such as THERATOPE® vaccine and gene therapy vaccines, e.g., ALLOVECTIN® vaccine, LEUVECTIN® vaccine and VAXID® vaccine; LURTOTECAN® topoisomerase 1 inhibitor; ABARELIX® rmRH; lapatinib ditosylate (ErbB-2 and EGFR bityrosine kinase small molecule inhibitor, also known as GW572016); and any pharmaceutically acceptable salts, acids, or derivatives of any of the above.
[0269] Examples of chemotherapy agents may include alemtuzumab (Campath), bevacizumab (AVASTIN®, Genentech), cetuximab (ERBITUX®, Imclone); panitumumab (VECTIBIX®, Amgen), rituximab (RITUXAN®, Genentech / Biogen Idec), pertuzumab (OMNITARG®, 2C4, Genentech), trastuzumab (HERCEPTIN®, Genentech), tositumomab (Bexxar, Corixia), and antibodies such as the antibody-drug conjugate gemtuzumab ozogamicin (MYLOTARG®, Wyeth). Further humanized monoclonal antibodies that have therapeutic potential as drugs in combination with the compounds of the present invention include apolizumab, aselizumab, atlizumab, bapinuzumab, vibatuzumab meltansine, cantuzumab meltansine, sedelizumab, certolizumab pegol, sidofcituzumab, sidotuzumab, daclizumab, eculizumab, efalizumab, epratuzumab, erulizumab, femzumab, fontrizumab, gemtuzumab ozogamicin, inotuzumab ozogamicin, ipilimumab, rabetuzumab, lintuzumab, matsuzumab, mepolizumab, motabizumab, motobizumab, and nata Lizumab, nimotuzumab, norovizumab, numabizumab, ocrelizumab, ocrelizumab, palivizumab, pascolizumab, pecufcituzumab, pectuzumab, pexerizumab, larivizumab, ranibizumab, reslivizumab, reslizumab, resivizumab, loberizumab, luprizumab, cibrotuzumab, ciprizumab, sontuzumab, tacutuzumab tetraxetan, tadocizumab, talizumab, tefibazumab, tocilizumab, tralizumab, tucotsuzumab cermoloykin, tucitutsuzumab, umavizumab, urtoxazumab, ustekinumab, bicilizumab, and interleukin-12 This includes anti-interleukin-12 (ABT-874 / J695, Wyeth Research and Abbott Laboratories), a recombinant exclusive human full-length IgG 1λ antibody genetically modified to recognize the p40 protein.
[0270] Examples of chemotherapy agents include "tyrosine kinase inhibitors" such as EGFR targeting agents (e.g., small molecules, antibodies, etc.); small molecule HER2 tyrosine kinase inhibitors such as TAK165 available from Takeda; oral selective inhibitors of ErbB2 receptor tyrosine kinase such as CP-724 and 714 (Pfizer and OSI); dual HER inhibitors such as EKB-569 (available from Wyeth) which preferentially binds to EGFR but inhibits both HER2 and EGFR overexpressing cells; lapatinib (GSK572016; available from Glaxo-SmithKline), an oral HER2 and EGFR tyrosine kinase inhibitor; PKI-166 (available from Novartis); pan-HER inhibitors such as canertinib (CI-1033; Pharmacia); Raf-1 inhibitors such as the antisense agent ISIS-5132 available from ISIS Pharmaceuticals, which inhibits Raf-1 signaling; and imatinib mesylate (Glaxo Non-HER-targeted TK inhibitors such as GLEEVEC® (available from SmithKline); multi-target tyrosine kinase inhibitors such as sunitinib (SUTENT®, available from Pfizer); VEGF receptor tyrosine kinase inhibitors such as batalanib (PTK787 / ZK222584, available from Novartis / Schering AG); MAPK extracellular regulatory kinase I inhibitor CI-1040 (available from Pharmacia); quinazolines, e.g., PD 153035, 4-(3-chloroanilino)quinazoline; pyridopyrimidines; pyrimidopyrimidines; CGP 59326, CGP 60261 and CGP Pyrrolopyrimidines such as 62706; pyrazolopyrimidines, 4-(phenylamino)-7H-pyrrolo[2,3-d]pyrimidines; curcumin (diferuloylmethane, 4,5-bis(4-fluoroanilino)phthalimide); tilphostin containing the nitrothiophene moiety; PD-0183805 (Warner-Lamber); antisense molecules (e.g., those that bind to HER coding nucleic acids); quinoxaline (US Patent No. 5,804,396); triphostin (US Patent No. 5,804,396); ZD6474 (AstraZeneca); PTK-787 (Novartis / Schering AG);This may also include pan-HER inhibitors such as CI-1033 (Pfizer); Affinitac (ISIS 3521; Isis / Lilly); Imatinib mesylate (GLEEVEC®); PKI 166 (Novartis); GW2016 (Glaxo SmithKline); CI-1033 (Pfizer); EKB-569 (Wyeth); Semaxinib (Pfizer); ZD6474 (AstraZeneca); PTK-787 (Novartis / Schering AG); INC-1C11 (Imclone); and rapamycin (sirolimus, RAPAMUNE®).
[0271] Examples of chemotherapy agents include dexamethasone, interferon, colchicine, methoprine, cyclosporine, amphotericin, metronidazole, alemtuzumab, alitretinoin, allopurinol, amifostine, arsenic trioxide, asparaginase, live BCG, becuzimab, besarotene, cladribine, clofarabine, darbepoetin alfa, denileukin, dexrazoxane, epoetin alfa, erotinib, filgrastim, histrelin acetate, ibritumomab, interferon alfa-2a, interferon alfa- 2b may include lenalidomide, levamisol, mesna, methoxsalen, nandrolone, nelarabine, nofetumomab, oprelbequin, parifermin, pamidronate, pegademase, pegaspargase, pegfilgrastim, pemetrexed disodium, plicamycin, porfimer sodium, quinacrine, rasburicase, salglamostim, temozolomide, VM-26, 6-TG, toremifene, tretinoin, ATRA, barrubicin, zoledronate and zoledronic acid, and pharmaceutically acceptable salts thereof.
[0272] Examples of chemotherapeutic agents include hydrocortisone, hydrocortisone acetate, cortisone acetate, thixocortol pivalate, triamcinolone acetonide, triamcinolone alcohol, mometasone, amcinonide, budesonide, desonide, fluocinonide, fluocinolone acetonide, betamethasone, betamethasone sodium phosphate, dexamethasone, dexamethasone sodium phosphate, fluocortone, and hydrocortisone-17-butyrate. Hydrocortisone-17-valerate, acromethasone dipropionate, betamethasone valerate, betamethasone dipropionate, prednicarbate, clobetazone-17-butyrate, clobetasol-17-propionate, fluocortron caproate, fluocortron pivalate and flupredniden acetate: phenylalanine-glutamine-glycine (FEG) and its D-isomer (feG) (IMULAN Immunoselective anti-inflammatory peptides (ImSAIDs) such as BioTherapeutics, LLC; antirheumatic drugs such as azathioprine, cyclosporine (cyclosporine A), D-penicillamine, gold salt, hydroxychloroquine, leflunomideminocycline, sulfasalazine, etanercept (ENBREL®), infliximab (REMICADE®), adalimumab (HUMIRA®), certolizumab pegol (CIMZI Tumor necrosis factor α (TNFα) blockers such as A(registered trademark), golimumab (SIMPONI(registered trademark)), interleukin-1 (IL-1) blockers such as anakinra (KINERET(registered trademark)), T-cell costimulatory blockers such as abatacept (ORENCIA(registered trademark)), interleukin-6 (IL-6) blockers such as tocilizumab (ACTEMERA(registered trademark)); interleukin-13 (IL-13) blockers such as lebrikizumab; interferon-alpha (IFN) blockers such as lontalizumab; beta-7 integrin blockers such as rhuMAb Beta7; IgE pathway blockers such as anti-M1 prime; secreted homotrimeric LTa3 and membrane-bound heterotrimeric LTa / β2 blockers, such as antilymphotoxin α(LTa);Various investigational drugs, such as thioplatin, PS-341, phenylbutyrate, ET-18-OCH3, or famesyltransferase inhibitors (L-739749, L-744832); polyphenols such as quercetin, resveratrol, piceatannol, epigallocatechin gallate, theaflavin, flavanol, procyanidin, betulinic acid and its derivatives; autophagy inhibitors such as chloroquine; delta-9-tetrahydrocannabinol (dronabinol, MARINOL®); beta-lapacon; lapachol; colchicine; betulinic acid; acetylcamptothecin, scopolectin, 9-aminocamptothecin); podophyllotoxin; tegafur (UFTORAL®); bexarotene (TARGRETIN®); clodronate (e.g., BONEFOS® or OSTA) Bisphosphonates such as C(registered trademark), etidronate (DIDROCAL(registered trademark)), NE-58095, zoledronic acid / zoledronate (ZOMETA(registered trademark)), alendronate (FOSAMAX(registered trademark)), pamidronate (AREDIA(registered trademark)), tildronate (SKELID(registered trademark)), or risedronate (ACTONEL(registered trademark)); and epidermal growth factor receptor (EGF-R); vaccines, e.g., THERATOPE(registered trademark) vaccine; perifosine, COX-2 inhibitors (e.g., celecoxib or etoricoxib), proteosome inhibitors (e.g., PS341); CCI-779; tipifamib (R11577); olafenib, ABT510; Bcl-2 inhibitors such as oblimersen sodium (GENASENSE(registered trademark)); pixantrone; ronafamib (SCH Examples include famesyltransferase inhibitors such as 6636, SARASAR (trademark); and pharmaceutically acceptable salts, acids, or derivatives of any of the above; as well as combinations of two or more of the above.
[0273] In many embodiments, upon diagnosis of cancer, several procedures can be performed, including (but not limited to) surgery, resection, chemotherapy, radiotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, and blood transfusion. In some embodiments, anticancer agents and / or chemotherapeutic agents are administered, including (but not limited to) alkylating agents, platinum agents, taxanes, vinca agents, anti-estrogens, aromatase inhibitors, ovarian suppressants, endocrine / hormone agents, bisphosphonate agents, and targeted biological agents. These agents include cyclophosphamide, fluorouracil (or 5-fluorouracil or 5-FU), methotrexate, thiotepa, carboplatin, cisplatin, taxanes, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifen, toremifene, fulvestrant, gemcitabine, irinotecan, ixabépirone, temozolomide, topotecan, and vitrexate. Cristine, vinblastine, eribulin, mutamycin, capecitabine, capecitabine, anastrozole, exemestane, letrozole, leuprolide, abalerix, buserelin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, zoledronate, tykerb, daunorubicin, doxorubicin, epirubicin, idarubicin, barurubicinMitoxantrone, bevacizumab, cetuximab, ipilimumab, ado-trastuzumab emtansine, afatinib, aldesleukin, alectinib, alemtuzumab, atezolizumab, avelumab, axutinib, belimumab, bellinostat, bevacizumab, blinatumomab, bortezomib, bosutinib, brentuximab vedotin, brigatinib, cabozantinib Canakinumab, carfilzomib, cerutinib, cetuximab, cobimetinib, crizotinib, dabrafenib, daratumumab, dasatinib, denosumab, dinutuximab, durvalumab, elotuzumab, enasidenib, erlotinib, everolimus, gefitinib, ibritumomab tiuxetan, ibrutinib, idelalisib, imatinib, ipilimumab, ixazo Mib, lapatinib, lenvatinib, midostaurin, necitumumab, neratinib, nilotinib, niraparib, nivolumab, obinutuzumab, ofatumumab, olaparib, olalatumab, osimertinib, palbociclib, panitumumab, panobinostat, pembrolizumab, pertuzumab, ponatinib, ramucirumab, regorafenib, ribociclib, rituximab, ro These include, but are not limited to, midepsin, rucaparib, ruxolitinib, siltuximab, ciproisel-T, sonidegib, sorafenib, temsirolimus, tocilizumab, tofacitinib, tocitumomab, trametinib, trastuzumab, vandetanib, vemurafenib, venetoclax, bismodegib, vorinostat, and zib-aflibercept. According to various embodiments, individuals may be treated with a single agent or combination of agents described herein. A common combination of agents is cyclophosphamide, methotrexate, and 5-fluorouracil (CMF).
[0274] In some embodiments of any one of the methods disclosed herein, any cell-free nucleic acid molecules (e.g., cfDNA, cfRNA) may be derived from cells. For example, a cell sample or tissue sample can be obtained from a subject and processed to remove all cells from the sample, thereby generating cell-free nucleic acid molecules derived from the sample.
[0275] In some embodiments of any one of the methods disclosed herein, the reference genome sequence may be derived from the cells of an individual. The individual may be a healthy control or the subject of the methods disclosed herein for determining or monitoring the progression of a condition.
[0276] A cell may be a healthy cell, or it may be a diseased cell. Diseased cells may have altered metabolism, gene expression, and / or morphological features. Diseased cells may be cancer cells, diabetic cells, and apoptotic cells. Diseased cells may be cells derived from a disease target. Exemplary diseases may include blood disorders, cancer, metabolic disorders, eye disorders, organ disorders, musculoskeletal disorders, and heart diseases.
[0277] Cells may be mammalian cells or may originate from mammalian cells. Cells may be rodent cells or may originate from rodent cells. Cells may be human cells or may originate from human cells. Cells may be prokaryotic cells or may originate from prokaryotic cells. Cells may be bacterial cells or may originate from bacterial cells. Cells may be archaeal cells or may originate from archaeal cells. Cells may be eukaryotic cells or may originate from eukaryotic cells. Cells may be pluripotent stem cells. Cells may be plant cells or may originate from plant cells. Cells may be animal cells or may originate from animal cells. Cells may be invertebrate cells or may originate from invertebrate cells. Cells may be vertebrate cells or may originate from vertebrate cells. Cells may be microbial cells or may originate from microbial cells. Cells may be fungal cells or may originate from fungal cells. Cells can originate from a specific organ or tissue.
[0278] Non-limiting examples of cells include lymphoid cells, e.g., B cells, T cells (cytotoxic T cells, natural killer T cells, regulatory T cells, T helper cells), natural killer cells, cytokine-induced killer (CIK) cells; myeloid cells, e.g., granulocytes (basophil granulocytes, eosinophil granulocytes, neutrophil granulocytes / hypersegmented neutrophils), monocytes / macrophages, erythrocytes (reticulocytes), mast cells, platelets / megakaryocytes, dendritic cells; cells from the endocrine system, including thyroid (thyroid epithelial cells, parafollicular cells), parathyroid (chief parathyroid cells, eosinophilic cells), adrenal (chromaffin cells), and pineal (pineal cells); and glial cells. Cells of the nervous system including (astroglial cells, microglia), giant cell neurosecretory cells, astrocytic cells, Boettcher cells and pituitary gland (gonadotropin-producing cells, corticotropin-producing cells, thyroid-stimulating hormone-producing cells, growth hormone-producing cells, mammary gland-stimulating hormone-secreting cells); cells of the respiratory system including alveolar cells (type I alveolar cells, type II alveolar cells), Clara cells, goblet cells, and dust cells; cells of the circulatory system including cardiomyocytes and pericytes; cells of the digestive system including stomach (gastrocnemiocytes, parietal cells), goblet cells, Paneth cells, G cells, D cells, ECL cells, I cells, K cells, and S cells; enteroendocrine cells including enterochromaffin cells, APUD cells, and liver (hepatocytes) , Kupffer cells), cartilage / bone / muscle; osteoblasts, osteocytes, osteoclasts, teeth (cementoblasts, ameloblasts); chondrocytes, including chondrocytes; skin cells, including trichocytes, keratinocytes, and melanocytes (nevus cells); muscle cells, including myocytes; urinary tract cells, including podocytes, juxtaglomerular cells, intraglomerular mesangial cells / extraglomerular mesangial cells, renal proximal tubular brush margin cells, and macula densa cells; germline cells, including spermatids, Sertoli cells, Leydig cells, and oocytes; adipocytes, fibroblasts, tendinocytes, epidermal keratinocytes (differentiated epidermal cells), epidermal basal cells (stem cells), keratinocytes of fingernails and toenails, nail bed basal cells (stem cells), medullary hair stem cells, cortical hair stem cells, cuticle hair stem cells, Examples include cycla hair follicle sheath cells, Huxley layer hair follicle sheath cells, Henle layer hair follicle sheath cells, outer hair follicle sheath cells, hair matrix cells (stem cells), Wet layered barrier epithelial cells, surface epithelial cells of layered squamous epithelium of the cornea, tongue, oral cavity, esophagus, anal canal, distal urethra and vagina, basal cells (stem cells) of the epithelium of the cornea, tongue, oral cavity, esophagus, anal canal, distal urethra and vagina, urinary epithelial cells (lining the bladder and ureters), exocrine secretory epithelial cells, salivary gland mucus cells (polysaccharide-rich secretion), salivary gland serous cells (glycoprotein enzyme-rich secretion), von Ebner's gland cells of the tongue (washing the taste buds), mammary gland cells (milk secretion), lacrimal gland cells (tear secretion), ceruminous gland cells in the ear (wax secretion), eccrine sweat gland dark cells (glycoprotein secretion), and eccrine sweat gland clear cells (small molecule secretion). Apocrine sweat gland cells (odor secretion, sex hormone sensitive), Mollian cells of the eyelids (specialized sweat glands), sebaceous gland cells (lipid-rich sebum secretion), Bowman's gland cells of the nose (washes the olfactory epithelium), Brunner's gland cells of the duodenum (enzymes and alkaline mucus), seminal vesicle cells (secretes seminal fluid components containing fructose to aid sperm movement), prostate cells (secretes seminal fluid components), bulbourethral gland cells (mucus secretion), Bartoli's gland Glandular cells (vaginal lubrication secretion), Littre cells (mucus secretion), endometrial cells (carbohydrate secretion), isolated goblet cells of the respiratory and digestive tract (mucus secretion), gastric mucosal cells (mucus secretion), gastric gland enzyme progenitor cells (pepsinogen secretion), gastric gland acid-secreting cells (hydrochloric acid secretion), pancreatic acinar cells (bicarbonate and digestive enzyme secretion), Paneth cells of the small intestine (lysozyme secretion), type II alveolar cells of the lungs (surfactant secretion),Clara cells of the lung, hormone-secreting cells, anterior pituitary cells, growth hormone-producing cells, mammary gland-stimulating hormone-producing cells, thyroid-stimulating hormone-producing cells, gonadotropin-producing cells, adrenocorticotropic hormone-producing cells, intermediate pituitary cells, giant cell neurosecretory cells, intestinal and airway cells, thyroid cells, thyroid epithelial cells, parafollicular cells, parathyroid cells, parathyroid chief cells, eosinophilic cells, adrenal cells, chromaffin cells, Leydig cells of the testis, endofollicular membrane cells of ovarian follicles, luteal cells of ruptured follicles, granulosa luteal cells, Follicular luteal cells, juxtaglomerular cells (renin secretion), macula compacta cells of the kidney, metabolic and storage cells, barrier function cells (lungs, intestines, exocrine glands and urogenital tract), kidney, type I alveolar cells (lining the air spaces of the lungs), pancreatic duct cells (acinate central cells), non-striatal duct cells (sweat glands, salivary glands, mammary glands, etc.), glandular tubule cells (seminal vesicles, prostate, etc.), epithelial cells lining closed body lumens, ciliated cells with propulsive function, extracellular matrix secretory cells, contractile cells; skeletal muscle cells, stem cells, cardiomyocytes, blood and immune system cells, red blood Chromocytes (red blood cells), megakaryocytes (platelet progenitor cells), monocytes, connective tissue macrophages (various types), epidermal Langerhans cells, osteoclasts (in bone), dendritic cells (in lymphoid tissue), microglia (central nervous system), neutrophil granulocytes, eosinophil granulocytes, basophil granulocytes, mast cells, helper T cells, suppressor T cells, cytotoxic T cells, natural killer T cells, B cells, natural killer cells, reticulocytes, stem cells and fate-determined progenitor cells (various types) for the blood and immune system, poly Examples include pluripotent stem cells, totipotent stem cells, induced pluripotent stem cells, adult stem cells, sensory transducer cells, autonomic neuron cells, sensory organ and peripheral neuron supporting cells, central nervous system neurons and glial cells, lens cells, pigment cells, melanocytes, retinal pigment epithelial cells, germ cells, oogonia / oocytes, spermatids, spermatocytes, spermatogonia (stem cells for spermatids), sperm, nurse cells, ovarian follicular cells, Sertoli cells (testis), thymic epithelial cells, stromal cells, and interstitial kidney cells.
[0279] In some embodiments of any one of the methods disclosed herein, the condition may be cancer or tumor. Non-limiting examples of such conditions include acanthoma, asin cell carcinoma, acoustic neuroma, acral lentigo melanoma, hidradenoma, acute eosinophilic leukemia, acute lymphoblastic leukemia, acute megakaryoblastic leukemia, acute monocytic leukemia, mature acute myeloblastic leukemia, acute myeloid dendritic cell leukemia, acute myeloid leukemia, acute promyelocytic leukemia, adamantinoma, adenocarcinoma, adenoid cystic carcinoma, adenoma, adenomatous odontogenic tumor, adrenocortical carcinoma, adult T-cell leukemia, aggressive NK-cell leukemia, AIDS-related cancer, AIDS-related Lymphoma, alveolar soft tissue sarcoma, myeloblastic fibroma, anal cancer, anaplastic large cell lymphoma, anaplastic thyroid cancer, angioimmunoblastic T-cell lymphoma, angiomyolipoma, angiosarcoma, appendiceal cancer, astrocytoma, atypical teratogenic rhabdoid tumor, basal cell carcinoma, basaloid carcinoma, B-cell leukemia, B-cell lymphoma, Bellini duct carcinoma, biliary tract cancer, bladder cancer, blastoma, bone cancer, bone tumor, brainstem glioma, brain tumor, breast cancer, Brenner tumor, bronchial tumor, bronchoalveolar carcinoma, pheochromocytoma, Burkitt lymphoma, cancer of unknown primary site, carcinoma Idoid tumor, cancer, carcinoma in situ, penile cancer, cancer of unknown primary site, carcinosarcoma, Castleman disease, central nervous system embryonic tumor, cerebellar astrocytoma, cerebral astrocytoma, cervical cancer, cholangiocarcinoma, chondroma, chondrosarcoma, chordoma, choriocarcinoma, choroid plexus papilloma, chronic lymphocytic leukemia, chronic monocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, chronic neutrophilic leukemia, clear cell tumor, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, Degos disease, dermatofibrosarcoma protuberans, dermoid cyst, fibrinogenic round cell tumor, diffuse large cell tumor B-cell lymphoma, germinal dysplastic neuroepithelial tumor, embryonic carcinoma, endoderm sinus tumor, endometrial cancer, endometrial uterine cancer, endometrioid tumor, enteropathy-associated T-cell lymphoma, ependymoblastoma, ependymomas, epithelioid sarcomas, erythroleukemia, esophageal cancer, sensory neuroblastoma, Ewing family tumors, Ewing family sarcomas, Ewing sarcomas, extracranial germ cell tumors, extragonadal germ cell tumors, extrahepatic cholangiocarcinoma, extramammary Paget's disease, fallopian tube cancer, inclusion fetus, fibroma, fibrosarcoma, follicular lymphoma, follicular thyroid cancer, gallbladder cancer, ganglioglioma, ganglioneuroma, gastric cancer, gastric lymphoma, gastrointestinal cancer,Gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, germ cell tumor, choriocarcinoma of pregnancy, trophoblastic tumor of pregnancy, giant cell tumor of bone, glioblastoma pleomorphoni, glioma, cerebral gliomatosis, glomus tumor, glucagonoma, gonadoblastoma, granulosa cell tumor, hairy cell leukemia, hairy cell leukemia, head and neck cancer, heart cancer, hemangioblastoma, perivascular cell tumor, angiosarcoma, hematological malignancies, hepatocellular carcinoma, hepatosplenic T-cell lymphoma, hereditary breast and ovarian cancer syndrome, Hodgkin lymphoma, hypopharyngeal cancer, hypothalamic glioma, inflammatory breast cancer, intraocular black Pancreatic sarcoma, islet cell carcinoma, islet cell tumor, juvenile myelomonocytic leukemia, Kaposi's sarcoma, Kaposi's sarcoma, kidney cancer, Krukenberg's tumor, laryngeal cancer, malignant lentigo melanoma, leukemia, lip and oral cancer, liposarcoma, lung cancer, luteoma, lymphangioma, lymphangiosarcoma, lymphoepithelioma, lymphocytic leukemia, lymphoma, macroglobulinemia, malignant fibrous histiocytoma, malignant fibrous histiocytoma of bone, malignant glioma, malignant mesothelioma, malignant peripheral nerve sheath tumor, malignant rhabdoid tumor, malignant Triton's tumor, MALT lymphoma, mantle cell lymphoma Thyroid cancer, mast cell leukemia, mediastinal germ cell tumor, mediastinal tumor, medullary thyroid cancer, medulloblastoma, medullary cell tumor, melanoma, melanoma, meningioma, Merkel cell carcinoma, mesothelioma, mesothelioma, metastatic squamous cell carcinoma of unknown primary origin, metastatic urothelial carcinoma, mixed Müller tumor, monocytic leukemia, oral cancer, myxoid neoplasm, multiple endocrine neoplasm syndrome, multiple myeloma, multiple myeloma, mycosis fungoides, myelodysplasia, myelodysplastic syndrome, myeloid leukemia, myeloid sarcoma, myeloproliferative disorder, myxoma, nasal cavity cancer, nasopharyngeal cancer, neoplasm, schwannoma, neuroblastoma, neuroblastoma Cystoma, neurofibroma, neuroma, nodular melanoma, non-Hodgkin lymphoma, non-melanoma skin cancer, non-small cell lung cancer, ocular oncology, oligoastrocytoma, oligodendroglioma, pallocyte tumor, optic nerve sheath meningioma, oral cancer, oral cancer, oropharyngeal cancer, osteosarcoma, osteosarcoma, ovarian cancer, ovarian epithelial carcinoma, ovarian germ cell tumor, low-grade ovarian tumor, Paget's disease of the breast, Pancoast tumor, pancreatic cancer, pancreatic cancer, papillary thyroid cancer, papillomatosis, paraganglioma, paranasal sinus cancer, parathyroid cancer, penile cancer, perivascular epithelioid cell tumor, pharyngeal cancer, pheochromocytoma,Moderately differentiated pineal parenchymal tumors, pineoblastomas, pituitary cell tumors, pituitary adenomas, pituitary tumors, plasma cell neoplasms, pleuropulmonary blastomas, polygermomas, precursor T lymphoblastic lymphomas, primary central nervous system lymphomas, primary exudative lymphomas, primary hepatocellular carcinomas, primary liver cancers, primary peritoneal cancers, primitive neuroectodermal tumors, prostate cancers, pseudomyxoma peritonei, rectal cancers, renal cell carcinomas, respiratory cancers involving the NUT gene on chromosome 15, retinoblastomas, rhabdomyomas, rhabdomyosarcomas, Richter's transformation, sacrococcygeal teratomas, salivary gland cancers, sarcomas, schwannomatoses, sebaceous gland carcinomas, secondary neoplasms, seminomas, serous tumors, Sertoli-Leydig cell tumors, sex cord-stromal tumors, Sézary syndrome, signet ring cell carcinomas, skin cancers, and small blue round cell tumors. tumor), small cell carcinoma, small cell lung cancer, small cell lymphoma, Tumors, small intestine cancer, soft tissue sarcoma, somatostatinoma, smoke warts, spinal cord tumors, spinal tumors, splenic marginal zone lymphoma, squamous cell carcinoma, gastric cancer, superficial spreading melanoma, tentorium This includes primitive neuroectodermal tumors, surface epithelial stromal tumors, synovial sarcomas, T-cell acute lymphoblastic leukemia, T-cell macrogranule lymphocyte leukemia, T-cell leukemia, T-cell lymphoma, T-cell prelymphocytic leukemia, teratomas, end-stage lymphoid cancer, testicular cancer, theca, pharyngeal cancer, thymic carcinoma, thymoma, thyroid cancer, transitional cell carcinoma of the renal pelvis and ureter, transitional cell carcinoma, urachal cancer, urethral cancer, urogenital neoplasms, uterine sarcoma, uveal melanoma, vaginal cancer, Verner-Morrison syndrome, verrucous carcinoma, visual pathway glioma, vulvar cancer, Waldenstrom macroglobulinemia, Warsin tumor, and Wilms tumor.
[0280] According to various embodiments, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), anal cancer, astrocytoma, basal cell carcinoma, cholangiocarcinoma, bladder cancer, breast cancer, Burkitt lymphoma, cervical cancer, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myeloproliferative neoplasm, colorectal cancer, diffuse large B-cell lymphoma, endometrial cancer, ependymoma, esophageal cancer, sensory neuroblastoma, Ewing's sarcoma, fallopian tube cancer, follicular lymphoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, hairy cell leukemia, hepatocellular carcinoma, Hodgkin lymphoma, hypopharyngeal cancer, Kaposi's sarcoma It can detect a wide range of neoplasms, including (but not limited to) tumors, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, Merkel cell carcinoma, mesothelioma, oral cancer, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumor, pharyngeal cancer, pituitary tumor, prostate cancer, rectal cancer, renal cell carcinoma, retinoblastoma, skin cancer, small cell lung cancer, small intestine cancer, squamous cell carcinoma of the neck, T-cell lymphoma, testicular cancer, thymoma, thyroid cancer, uterine cancer, vaginal cancer, and hemangiomas.
[0281] Many embodiments relate to diagnostic or companion diagnostic scans performed during cancer treatment of an individual. When diagnostic scans are performed during treatment, the ability of the drug to treat cancer growth can be monitored. Most anticancer drugs cause the death and necrosis of neoplastic cells, which should release more nucleic acids from these cells into the sample being tested. Therefore, levels of circulating tumor nucleic acids can be monitored over time, as they should rise during initial treatment and begin to decrease as the number of cancer cells decreases. In some embodiments, the treatment is adjusted based on the treatment effect on cancer cells. For example, if the treatment is not cytotoxic to neoplastic cells, the dose can be increased, or a drug with higher cytotoxicity can be administered. Alternatively, if the cytotoxicity of cancer cells is good, but undesirable side effects are high, the dose can be decreased, or a drug with fewer side effects can be administered.
[0282] Various embodiments also relate to diagnostic scans performed after treatment of an individual to detect residual disease and / or cancer recurrence. If the diagnostic scan indicates residual and / or cancer recurrence, further diagnostic tests and / or treatments may be performed as described herein. If the cancer and / or the individual is prone to recurrence, diagnostic scans may be performed more frequently to monitor any potential recurrences. F. Computer System
[0283] In one embodiment, the Disclosure provides a computer program product comprising a non-temporary computer-readable medium having computer executable code encoded therein, wherein the computer executable code is adapted to be executed in a manner that implements any one of the methods described above.
[0284] This disclosure provides a computer system programmed to implement the method of this disclosure. The system may in some cases include components such as a processor, an input module for inputting sequencing data or data derived therefrom, a computer-readable medium containing instructions that, when executed by the processor, execute an algorithm on inputs relating to one or more cell-free nucleic acid molecules, and an output module that provides one or more indices relating to a state.
[0285] Figure 27 shows a computer system 2701 configured by program or other means to implement some or all of the methods disclosed herein. The computer system 2701 can adapt various aspects of the disclosure, for example, (i) identifying one or more cell-free nucleic acid molecules containing multiple stepwise variants from sequencing data derived from multiple cell-free nucleic acid molecules; (ii) analyzing any of the identified cell-free nucleic acid molecules; (iii) determining the state of a subject based at least partially on the identified cell-free nucleic acid molecules; (iv) monitoring the progression of the state of a subject based at least partially on the identified cell-free nucleic acid molecules; (v) identifying a subject based at least partially on the identified cell-free nucleic acid molecules; or (vi) determining appropriate treatment for the state of a subject based at least partially on the identified cell-free nucleic acid molecules. The computer system 2701 may be a computer system located remotely on a user's electronic device or on an electronic device. The electronic device may be a mobile electronic device.
[0286] The computer system 2701 includes a central processing unit (CPU, "processor" and "computer processor" as herein) 2705, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 2701 also includes memory or memory locations 2710 (e.g., random-access memory, read-only memory, flash memory), an electronic storage unit 2715 (e.g., a hard disk), a communication interface 2720 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 2725 such as cache, other memory, data storage, and / or electronic display adapters. The memory 2710, storage unit 2715, interface 2720, and peripheral devices 2725 communicate with the CPU 2705 via a communication bus (solid line), such as a motherboard. The storage unit 2715 can be a data storage unit (or data repository) for storing data. The computer system 2701 can be operably coupled to a computer network ("network") 2730 with the help of the communication interface 2720. Network 2730 can be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. Network 2730 may, in some cases, be a telecommunications and / or data network. Network 2730 may include one or more computer servers that can enable distributed computing, such as cloud computing. Network 2730 may, in some cases, implement a peer-to-peer network that, with the help of computer system 2701, can enable devices coupled to computer system 2701 to act as clients or servers.
[0287] The CPU 2705 can execute a set of machine-readable instructions that can be embodied in a program or software. Instructions can be stored in memory locations such as memory 2710. Instructions can target the CPU 2705, which can then be programmed or otherwise configured to implement the methods of this disclosure. Examples of operations performed by the CPU 2705 may include fetching, decoding, executing, and writing back.
[0288] The CPU 2705 can be part of a circuit, such as an integrated circuit. One or more other components of System 2701 can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0289] The storage unit 2715 can store files such as drivers, libraries, and saved programs. The storage unit 2715 can also store user data, such as user preferences and user programs. The computer system 2701 may, in some cases, include one or more additional data storage units located outside the computer system 2701, such as on a remote server that communicates with the computer system 2701 via an intranet or the internet.
[0290] Computer system 2701 can communicate with one or more remote computer systems via network 2730. For example, computer system 2701 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone®, Android-enabled devices, Blackberry®), or personal digital assistants. Users can access computer system 2701 via network 2730.
[0291] The methods described herein can be implemented by machine-executable code (e.g., a computer processor) stored in an electronic storage location of a computer system 2701, such as memory 2710 or an electronic storage unit 2715. The machine-executable code or machine-readable code can be provided in the form of software. In use, the code can be executed by the processor 2705. In some cases, the code may be retrieved from the storage unit 2715 and stored in memory 2710 for easy access by the processor 2705. In some situations, the electronic storage unit 2715 may be omitted, and machine-executable instructions are stored in memory 2710.
[0292] The code can be pre-compiled and configured for use in a machine with a processor adapted to run the code, or it can be compiled at runtime. The code can be supplied in a programming language that can be chosen to allow the code to run either pre-compiled or as-compiled.
[0293] Embodiments of systems and methods provided herein, such as computer system 2701, can be embodied in programming. Various embodiments of the art can typically be considered “products” or “manufactured goods” in the form of machine (or processor) executable code and / or associated data carried on or embodied thereon on some kind of machine-readable medium. Machine executable code can be stored in memory (e.g., read-only memory, random-access memory, flash memory) or electronic storage units such as hard disks. The “storage” type medium can include any or all of tangible memory such as computers and processors, or associated modules such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-temporary storage at any time for software programming. All or part of the software may be communicated from time to time over the Internet or various other telecommunication networks. Such communication can enable, for example, the loading of software from one computer or processor to another computer or processor, or from a management server or host computer to an application server computer platform, for example. Therefore, other types of media that can carry software elements include optical waves, electrical waves, and electromagnetic waves, such as those used across physical interfaces between local devices, via wired and optical fixed telephone networks, and via various air links. Physical elements that carry such waves, such as wired or wireless links and optical links, can also be considered media that carry software. As used herein, unless limited to non-temporary and tangible “storage” media, terms such as “readable media” of a computer or machine refer to any medium involved in providing instructions to a processor for execution.
[0294] Therefore, machine-readable media such as computer executable code can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any storage device such as any computer(s), which may be used to implement databases, etc., as shown in drawings. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include copper wires and optical fibers, including coaxial cables; wires with buses within computer systems. Carrier media can take the form of electrical or electromagnetic signals, or sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card paper tapes, any other physical storage media having a pattern of holes, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carriers that carry data or instructions, cables or links that carry such carriers, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media can be involved in transporting one or more sequences of one or more instructions to a processor for execution.
[0295] The computer system 2701 includes, or can communicate with, an electronic display 2735 including a user interface (UI) 2740 for providing, for example, (i) analysis of any of the identified cell-free nucleic acid molecules, (ii) a determined state of an object based at least partially on the identified cell-free nucleic acid molecules, (iii) a determined progression of the state of an object based at least partially on the identified cell-free nucleic acid molecules, (iv) an identified object suspected to have a state based at least partially on the identified cell-free nucleic acid molecules, or (v) a determined treatment of an object based at least partially on the identified cell-free nucleic acid molecules. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0296] The methods and systems of this disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by the central processing unit 2705. For example, the algorithms can (i) identify one or more cell-free nucleic acid molecules containing multiple stepwise variants from sequencing data derived from multiple cell-free nucleic acid molecules; (ii) analyze any of the identified cell-free nucleic acid molecules; (iii) determine the state of a subject based at least partially on the identified cell-free nucleic acid molecules; (iv) monitor the progression of the state of a subject based at least partially on the identified cell-free nucleic acid molecules; (v) identify a subject based at least partially on the identified cell-free nucleic acid molecules; or (vi) determine appropriate treatment for the state of a subject based at least partially on the identified cell-free nucleic acid molecules. [Examples]
[0297] The following illustrative examples represent embodiments of the stimuli, systems, and methods described herein and are not intended to limit them in any way. Example 1: Genomic distribution of stepwise variants
[0298] An alternative to double-strand sequencing is described for reducing background error rates, including the detection of “stepped variants” (PVs) in which two or more mutations occur in cis (i.e., on the same strand of DNA in Figures 1A and 1E). Similar to double-strand sequencing, this method provides a lower error profile for the matched detection of two distinct non-reference events in individual molecules. However, unlike double-strand sequencing, both events occur in the same sequencing read pair, thereby increasing the efficiency of genome retrieval. Stepped variants are present in a variety of cancer types but occur in typological regions of the genome of B-cell malignancies, likely due to on-target and aberrant somatic hypermutations (aSHMs) driven by activation-induced deaminase (AID). The most common regions of aSHMs in B-cell non-Hodgkin lymphoma (NHL) are identified. Described herein is a novel method for detecting ctDNA via stepped variants to tumor proportions on the order of approximately one in a million: Stepped Variant Enrichment and Detection Sequencing (PhasED-Seq). This specification demonstrates that PhasED-Seq can significantly improve the detection of ctDNA in clinical samples both during treatment and before disease recurrence.
[0299] To identify malignancies in which PV may potentially improve disease detection, we evaluated the frequency of PV across cancer types. We analyzed publicly available whole-genome sequencing data to identify a set of variants occurring at distances less than 170 bp apart, representing the typical length of a single cfDNA fragment consisting of a single core nucleosome and associated linker. We identified and summarized the frequencies (Figure 1B, Figure 5, and Table 1) of these “presumptive stepwise variants” (Example 10) controlling the total number of SNVs from 2538 tumors across 24 cancer histological features, including solid tumors and hematological malignancies. PV was most significantly enriched in two B-cell lymphomas (DLBCL and follicular lymphoma, FL, P<0.05 for all other histological features), which are a group of hypermutated diseases caused by AID / AICDA. Example 2: The underlying mutation mechanism of PV
[0300] To investigate the origin of PVs, we compared single nucleotide substitution (SBS) mutation signatures contributing to SNVs occurring within 170 bp of another SNV with SNVs occurring alone (e.g., without another SNV within 170 bp) (Example 10). As expected, PVs were highly enriched with several mutation signatures associated with clustered mutations. Clustered mutation signatures associated with AID activity (SBS84 and SBS85) were significantly enriched in PVs from B-cell lymphoma and CLL, while signatures associated with APOBEC3B activity (SBS2 and SBS13), another mechanism of kataegis hypermutation, were significantly enriched in PVs from multiple solid tumor histologies, including ovarian, pancreatic, prostate, and mammary gland adenocarcinomas (Figure 1C and Figures 6A-6WW). Clustered mutation signatures associated with AID activity (SBS84 and SBS85) were enriched in PVs found in lymphoma and CLL, while signatures associated with APOBEC3B activity (SBS2 and SBS13) were significantly enriched in breast cancer (Figure 1C and Figures 6A-6WW). PVs from multiple tumor types were also associated with SBS4, a signature associated with tobacco use. Furthermore, novel enrichment was observed in several other signatures with no apparent associated mechanisms among PVs across multiple tumor histological features (e.g., SBS24, SBS37, SBS38, and SBS39). In contrast, age-related mutation signatures such as SBS1 and SBS5 were significantly enriched in isolated SNVs. Example 3: PV occurs in the typical genomic regions of lymphoid cancers.
[0301] To assess the genomic distribution of estimated PVs, these events were first binned into 1kb regions to visualize their frequencies across tumor types. Significantly typological PV distributions were observed in individual lymphoid neoplasms (e.g., DLBCL, FL, Burkitt lymphoma (BL), and chronic lymphocytic leukemia (CLL); Figures 1D and 7). In contrast, non-lymphoid cancers generally did not show substantial recurrence of clustered PVs in typological regions. This lack of typology in PV location was true even when considering melanoma and lung cancer, diseases with high PV frequencies.
[0302] Notably, the majority of hypermutated regions were shared across all three lymphoma subtypes, with the highest densities observed at known targets of aSHM, including BCL2, BCL6, and MYC, as well as at immunoglobulin (Ig) loci encoding heavy and light chain IGH, IGK, and IGL (Table 2). Surprisingly, specific regions within the Ig loci were densely mutated in almost all lymphoma and CLL patients (Figure 1D). Among the lymphoma subtypes, DLBCL tumors had the most 1kb regions containing recurrent PV (Figure 8A), coinciding with the most frequently recurrently mutated genes observed in this tumor type. In total, 1639 unique 1kb regions containing recurrent PV were identified in B lymphoid malignancies. Of these lymphoma-associated 1kb regions, nearly one-third were classified as genomic regions previously associated with physiological or abnormal SHM in B cells. Specifically, 19% (315 / 1639) were located in Ig regions, and 13% (218 / 1639) were in some of the 68 previously identified targets of aSHM (Table 2). While most PVs were classified as non-coding regions of the genome, more recurrently affected loci not previously described as targets of aSHM were also identified, including XBP1, LPP, and AICDA.
[0303] The distribution of PVs within each lymphoid malignancy correlated with oncogenic features associated with the different pathophysiologies of the corresponding diseases. For example, cases of FL in which over 90% of tumors had oncogenic BCL2 fusions were significantly more likely to contain stepwise variants in BCL2 than other lymphoid malignancies (Figure 1D and Figure 8B). Similarly, Burkitt lymphoma (BL) had significantly more PVs in MYC and ID3, two driver genes strongly associated with BL pathogenesis, than other lymphoid malignancies (Figure 1D and Figures 8C-8D). DLBCL molecular subtypes associated with different cell origins also demonstrated different distributions of PVs (Table 2). Specifically, germinal center B-cell-like (GCB) and activated B-cell-like (ABC) DLBCLs had similar overall PV frequencies (median 798 vs. 516, P=0.37), but significant enrichment of PV in the telomere IGH class switch regions (Sγ1 and Sγ3) of ABC-DLBCLs was found, consistent with previous reports41 (Figure 8E). Conversely, GCB-DLBCLs had more stepwise haplotypes in the centromere IGH class switch regions (Sα2 and Sε) and BCL2. Example 4: Design and validation of a PhasED-Seq panel for lymphoma
[0304] To validate these PV-rich regions and evaluate their usefulness for disease detection from ctDNA, we designed sequencing panels targeting putative PVs identified within WGS from three independent cohorts of DLBCL and CLL patients (Figure 2A and Example 10). This final stepwise variant enrichment and detection sequencing (PhasED-Seq) panel targeted approximately 115kb of genomic space focused on PVs, along with approximately 200kb of additional target genes that are recurrently mutated in B-NHL (Table 3). Although the 115kb space dedicated to PV capture targets only 0.0035% of the human genome, it captures 26% of the stepwise variants observed in mature B-cell neoplasms characterized by WGS (Figure 9A), and therefore, PhasED-Seq provides approximately 7500-fold greater PV enrichment than WGS.
[0305] We compared expected SNV and PV retrieval with a previously reported CAPP-Seq selector designed to maximize patient-to-patient SNVs in B-cell lymphoma (Figures 9A–9C). Considering the diverse B-NHL with available WGS data, PhaseD-Seq retrieved 3.0 times more SNVs (81 vs. 27) and 2.9 times more PVs (50 vs. 17) in median terms than the previous CAPP-Seq panel. This observation highlights the importance of including non-coding regions of the genome for maximum mutation retrieval. To experimentally validate these yield improvements, we profiled 16 pre-treated tumor or plasma DNA samples from DLBCL patients (Table 4). Both CAPP-Seq and PhaseD-Seq panels were applied in parallel to each sample, and they were then sequenced to high intrinsic molecular depth (Figure 2B). A similar improvement in SNV yield with PhasED-Seq compared to CAPP-Seq (2.7 times; median 304.5 vs. 114) was observed compared to the expected enrichment established from WGS. However, when enumerating the PVs observed in individual sequenced DNA fragments, a more favorable improvement for PhasED-Seq was observed (7.7 times; median 5554 vs. 719.5 PV / case) than the improvement expected from WGS was seen. This improvement is potentially attributable to either 1) a higher sequencing depth in targeted sequencing resulting in improved detection of rare alleles, or 2) a higher enumeration of higher-order PVs in targeted sequencing by PhasED-Seq or CAPP-Seq, which was not considered in the WGS design (i.e., more than two SNVs per fragment; Figures 9D-9F). Furthermore, a strong correlation was observed between the estimated frequency of PVs in WGS data and PVs from targeted sequencing by PhasED-Seq across 101 DLBCL samples (Figure 2C) across a 1kb window of the panel, further exploring the frequency and distribution of PVs in B-cell malignancies. Example 5: Differences in stepwise variants between lymphoma subtypes
[0306] After validating the PhasED-Seq panel, we investigated biological differences in PV among various B-cell malignancies, including DLBCL (n=101), primary mediastinal large B-cell lymphoma (PMBCL) (n=16), and classical Hodgkin lymphoma (cHL) (n=23). The number of SNVs identified per case did not differ significantly between lymphoma subtypes (Figure 9G-9K). However, considering the mutation haplotype, cHL had a significantly lower PV burden than either DLBCL or PMBCL. In addition to this quantitative difference, differences in the genomic location of PVs between different B-cell lymphoma subtypes were also observed (Figures 2D-2E and 10-12). This included already established biological associations within DLBCL subtypes, such as a higher proportion of BCL2 PV in GCB-type DLBCL than in ABC-type DLBCL, with an inverse correlation for PIM1. Compared to DLBCL, where the breakpoint is common in PMBCL, more frequent PVs were also observed in CIITA in PMBCL. Relative enrichment was also observed across the entire IGH locus, with more frequent PVs in the Sγ3 and Sγ1 regions of ABC-DLBCL (compared to GCB-DLBCL), and interestingly, more frequent PVs in the Sε locus of cHL compared to DLBCL (Figures 2E and 13). Overall, after adjusting for multiple hypotheses, significant relative enrichments were found at 25 loci between ABC-DLBCL and GCB-DLBCL, 24 between DLBCL and PMBCL, and 40 between DLBCL and cHL (Figures 10-12). Example 6: Stepwise Variant Recovery by PhasED-Seq
[0307] Efficient recovery of DNA molecules is desirable to facilitate ctDNA detection using PV. Hybrid capture sequencing is potentially sensitive to DNA mismatches, and hybridization efficiency decreases as mutations increase. In fact, AID hotspots can contain local mutation rates of 5–10%, and even higher rates in specific regions of IGH. To empirically assess the effect of mutation rate on capture efficiency, we simulated 150-mer DNA hybridization with varying mutation rates in silico. As expected, the predicted binding energy decreased with increasing mutation rate (Figure 14A). Notably, randomly distributed mutations had a greater impact on binding energy than clustered mutations. To evaluate the effect of this decrease in binding affinity, we synthesized 150-mer DNA oligonucleotides with 0–10% differences from the reference sequence at two loci targeted by MYC and BCL6, aSHM. To assess the worst-case scenario for hybridization, the non-reference bases were randomly distributed rather than clustered (Example 10). Next, equimolar mixtures of these oligonucleotides were captured using a PhasED-Seq panel. Consistent with in silico predictions, increased mutation rates resulted in decreased capture efficiency (Figure 3A). Molecules with a 5% mutation rate were captured with 85% efficiency compared to their fully wild-type counterparts, while molecules with a 10% mutation rate were captured with only 27% relative efficiency. To assess the prevalence of this degree of mutation in human tumors, the proportion of mutated bases in overlapping 151 bp windows was calculated (Example 10) to examine the variant distribution in the panel of 140 B-cell lymphoma patients. Only 7% (10 / 140) of patients had any 151 bp window with a mutation rate exceeding 10% (Figures 14B-14C). Indeed, experiments using synthetic oligonucleotides recovered 5% mutation rates with nearly the same efficiency as wild-type sequences. In more than half of all cases examined, no loci had a mutation rate exceeding 5% in any window; however, in all cases, more than 90% of the windows had mutation rates of less than 5%.Overall, these observations suggest that, despite hybridization bias, the majority of stepwise mutations can be recovered through efficient hybrid capture. Example 7: Error profile and detection limit for stepwise variant sequencing.
[0308] Previous methods for highly error-suppressed sequencing applied to cfDNA have utilized either a combination of molecular and in silico methods for error suppression (e.g., integrated digital error suppression, iDES) or double-stranded molecular recovery. However, each of these has limitations in either detecting events at very low tumor rates or efficiently recovering the original DNA molecule, which are important considerations for cfDNA analysis where input DNA is limited. Error profiling and recovery of input genomes from plasma cfDNA samples from 12 healthy adults using PhasED-Seq were compared with both iDES-CAPP-Seq and double-stranded sequencing. iDES-enhanced CAPP-Seq had a lower background error profile than barcode deduplication alone, but double-stranded sequencing provided the lowest background error rate for non-reference single nucleotide substitutions (Figure 3B, 3.3 × 10⁻⁶). -5 vs 1.2 × 10 -5 (P<0.0001). However, the rate of stepwise errors (e.g., multiple non-reference bases occurring on the same sequencing fragment) was significantly lower than the rate of single errors in either iDES-enhanced CAPP-Seq or double-stranded sequencing data. This was true for the incidence of both two (2× or "doublet" PVs) or three (3× or "triplet" PVs) substitutions on the same DNA molecule (Figure 3B, 8.0×10⁶ each). -7 and 3.4 × 10 -8(P<0.0001). Stepwise errors, including C-to-T or T-to-C transposition substitutions, were more common than other types of PVs (Figure 14D). Notably, the doublet PV error rate in cfDNA also correlated with the distance between locations, with the highest PV error rates consisting of adjacent SNVs (e.g., DNVs), and the error rate decreased as the distance between constituent variants increased (Figure 14E). Considering intrinsic molecular depth, double-strand sequencing recovered only 19% of all intrinsic cfDNA fragments (Figure 3C). In contrast, the intrinsic depth of PVs within genomic distances of less than 20 bp was nearly identical to the depth of individual locations (e.g., molecules covering individual SNVs). Similarly, PVs with a size of up to 80 bp had depths exceeding 50% of the median intrinsic molecular depth of the sample. Importantly, nearly half (48%) of all PVs were within 80 bp of each other, demonstrating their usefulness for disease detection from input-restricted cfDNA samples (Figure 3D).
[0309] To quantitatively compare the performance of PhasED-Seq with alternative methods for ctDNA detection, limiting dilutions of ctDNA from three lymphoma patients were generated from healthy control cfDNA, and expected tumor fractions of 0.1% to 0.00005% (1 in 2,000,000) were calculated. A match was obtained (Example 10). The expected tumor percentage was compared to the estimated tumor content in each of these dilutions using PhasED-Seq to track tumor-derived PVs, as well as error-suppressed detection methods corresponding to individual SNVs (e.g., iDES-enhanced CAPP-Seq or double-strand sequencing; Figure 3E). All methods were performed using a 0.01% (1 / 10,000) tumor percentage. The results were equally thorough up to this point. However, below this level (e.g., 0.001%, 0.0002%, 0.0001%, and 0.00005%), both PhasED-Seq and double-strand sequencing significantly outperformed iDES-enhanced CAPP-Seq (P<0.0001 for double-strand, "2x" PhasED-Seq, and "3x" PhasED-Seq; Figure 3E). Furthermore, when compared to dual sequencing, tracking two or three inophase variants (e.g., 2x and 3x PhasED-Seq) allowed for more accurate identification of expected tumor content, down to 1 in 2,000,000. Excellent linearity was observed for double-stranded pairs (P=0.005 for 2×PhasED-Seq and P=0.002 for 3×PhasED-Seq) (Example 10). PV specificity was assessed by searching for evidence of tumor-derived SNVs or PVs in cfDNA samples from 12 unrelated healthy controls and healthy controls used for limiting dilutions. Here again, both 2x- and 3x-PhasED-Seq showed significantly lower background signal levels than CAPP-Seq and double-stranded sequencing (Figure 3F). This low error rate and background from PV improves the detection limit for ctDNA disease detection. In some examples, the sequencing-based cfDNA assay methods described herein (e.g., methods shown in Figures 3E and 3F) do not require molecular barcoding to achieve sophisticated error suppression and low detection limits. Signals evaluated using a non-barcode method were obtained using a limiting dilution series of 1:10 million to 5:10 million and a "blank" control (Figures 23A to 23B).
[0310] This dilution series was used to evaluate the detection limit for a given number of PVs (Figure 3G-3I). When considering a set of PVs within a 150 base pair (bp) region, the probability of detection for a given sample can be accurately modeled by binomial sampling, taking into account both the sequencing depth and the number of 150 bp regions containing PVs (Example 10). Example 8: Improvement in the detection of low-load, minimal residual disease
[0311] To test the usefulness of lower LODs obtained by PhaseD-Seq for the detection of ultra-low loading MRD from cfDNA, serial cell-free DNA samples were sequenced from a patient receiving frontline treatment for DLBCL (Figure 4A). Using CAPP-Seq, this patient had undetectable ctDNA after only one cycle of treatment, and multiple subsequent samples during and after treatment also remained undetectable. This patient subsequently had detectable ctDNA reappearance beyond 250 days after treatment initiation, and final clinical and radiological disease progression at 5 months, showing serial false-negative measurements by CAPP-Seq. Surprisingly, all four plasma samples that were undetectable by CAPP-Seq during and after treatment had detectable ctDNA levels by PhaseD-Seq, with a mean allele percentage as low as 6 parts per million. This increased sensitivity improved the lead time for disease detection by ctDNA compared to radiological monitoring from 5 months with CAPP-Seq to 10 months with PhaseD-Seq.
[0312] Next, we evaluated the performance of PhaseED-Seq ctDNA detection in a cohort of 107 patients with large B-cell lymphoma and blood samples available after one or two cycles of standard immunochemotherapy. Importantly, ctDNA levels measured by PhaseED-Seq correlated highly with levels measured by CAPP-Seq. In total, 443 tumor, germline, and cell-free DNA samples, including cfDNA, were evaluated pre-treatment (n=107) and post-treatment (n=82 and 89) (n=82 and 89). Pre-treatment, patient-specific PV was detectable by PhaseED-Seq in 98% of samples, and cfDNA from healthy controls had 95% specificity (Figures 15 and 16A). Importantly, considering both pre-treated and post-treated samples, ctDNA levels measured by PhasED-Seq correlated highly with ctDNA levels measured by CAPP-Seq (Spearman rho = 0.91, Figure 16B). Next, we compared the quantitative levels of ctDNA measured by PhasED-Seq and CAPP-Seq from cfDNA samples after the initiation of treatment. In total, 72% (78 / 108) of samples with ctDNA detectable by PhasED-Seq after 1 or 2 cycles were also detectable by conventional CAPP-Seq (Figure 4B). Among the 108 samples detected by PhasED-Seq, disease burden was significantly lower for those with undetectable (28%) ctDNA levels compared to detectable (72%) using conventional CAPP-Seq, with a difference in median ctDNA levels exceeding 10x (tumor proportion 2.2 × 10⁻⁶). -4 vs 1.2 × 10 -5 (P<0.001, Figure 4B). When comparing PhasED-Seq and CAPP-Seq, a total of 16% (13 / 82) of samples after one cycle of treatment and 19% (17 / 89) of samples after two cycles of treatment had detectable ctDNA (Figure 4C).
[0313] The ctDNA molecular response criterion has been previously described for DLBCL patients using CAPP-Seq, and includes Major Molecular Response (MMR), defined as a 2.5 log decrease in ctDNA after two cycles of treatment. While MMR at this point is prognostic for outcomes, many patients have ctDNA undetectable by CAPP-Seq at this landmark (Figures 4D-4E). Importantly, even in patients with CAPP-Seq undetectable ctDNA, detection of subclinical, very low ctDNA levels by PhasED-Seq was prognostic for outcomes including relapse-free survival and overall survival (Figure 4D). Indeed, in 89 patients with available samples from this point onward, 58% (52 / 89) had ctDNA undetectable by CAPP-Seq at their provisional MMR assessment after completing two cycles of the planned six-cycle treatment. Using PhasED-Seq, 33% (17 / 52) of the samples not detected by CAPP-Seq had evidence of ctDNA as demonstrated by PV, at low levels of approximately 3:1,000,000 (Figures 17A-17D)—further detected by PhasED-Seq. These 17 cases represent potential false-negative tests by CAPP-Seq. Similar results were observed at the initial molecular response (EMR) time (i.e., after one cycle of treatment, Figures 18A-18H).
[0314] While detection of ctDNA in DLBCL after one or two cycles of treatment is a known adverse prognostic marker, outcomes for patients with undetectable ctDNA at these time points are heterogeneous (Figures 4E and 18F). Importantly, even in patients with undetectable ctDNA by CAPP-Seq after one or two cycles of treatment, detection of very low ctDNA levels by PhasED-Seq was a strong prognostic indicator for outcomes, including relapse-free survival (Figures 4F, 17C-17D, 18C-18D, and 18G). When detection by PhasED-Seq was combined with the previously described MMR threshold, patients could be stratified into three groups: patients who had not achieved MMR, patients who had achieved MMR but had persistent ctDNA, and patients with undetectable ctDNA (Figure 4G). Interestingly, patients who did not achieve MMR had a particularly high risk of early events despite additional planned first-line treatment (e.g., within the first year of treatment), while patients with persistent low levels of ctDNA appeared to be at higher risk of later relapse or progression events. In contrast, patients with undetectable ctDNA after two cycles of treatment with PhaseD-Seq had overwhelmingly favorable outcomes, with 95% being relapse-free and a 97% overall survival rate at 5 years. Similar results were observed at the time of EMR after one cycle of treatment (Figure 18H). Example 9: Exemplary embodiment of mutation detection using next-generation sequencing (NGS) when the mutation is a pair of mutations rather than a single nucleotide substitution.
[0315] In many cases, limitations in cfDNA tracking may be limited by the number of molecules available for detection. Furthermore, there are several potential limitations in tracking tumor molecules from cell-free DNA, including not only the sequencing error profile but also the number of molecules available for detection. The number of molecules available for detection—referred here to be the number of “evaluable fragments”—can be thought of as a function of both the number of unique genomes recovered (e.g., the unique depth of sequencing) and the number of somatic mutations being tracked. More specifically, the number of evaluable fragments is equal to EF = d * n.
[0316] Here, d = intrinsic molecular depth considered, and n = number of somatic changes tracked. In a typical cell-free DNA sample, fewer than 10,000 intrinsic genomes are often recovered (d), requiring any highly sensitive method to track multiple changes (n). Furthermore, as mentioned above, a major limitation of double-strand sequencing is the difficulty in recovering a sufficient intrinsic molecular depth (d). Thus, from a typical plasma sample with approximately 1,500x double-strand depth, only 150,000 evaluable fragments are available, even after 100 somatic changes. Therefore, in this scenario, sensitivity ...
Claims
[Claim 1] The invention described herein.