Methods for identifying tumor nucleic acids
The method of circularizing and concatemer-based sequencing addresses the limitations of current tumor nucleic acid detection by achieving low detection limits and efficient error correction, facilitating rapid and sensitive tumor-specific nucleic acid detection in blood samples across multiple cancer types.
Patent Information
- Application Number
- JP2025526410
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-28
- Filing Date
- 2023-11-07
- Publication Date
- 2025-11-26
AI Technical Summary
Current methods for detecting tumor nucleic acids in blood samples face challenges in achieving low detection limits and efficient error correction, particularly in tumor-informed approaches, which require deep sequencing and individualized designs, leading to high costs and logistical complexities.
A method involving circularization and concatemer-based sequencing is employed, where nucleic acids are circularized, amplified to generate concatemers, and sequenced at low depths to identify tumor-specific sequence variants, utilizing techniques like rolling circle amplification and negative/positive selection to enhance accuracy and reduce errors.
This approach enables rapid and sensitive detection of tumor-specific nucleic acids in blood samples, achieving low detection limits and reducing logistical challenges by using concatemer sequencing for genome-wide error suppression, suitable for various cancer types.
Smart Images

Figure 2025538165000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 382,944, filed November 9, 2023, and U.S. Provisional Application No. 63 / 492,690, filed March 28, 2023, each of which is incorporated by reference in its entirety. [Background technology]
[0002] During tumor development, tumor-derived nucleic acids are often released into the bloodstream by tumors. Apoptosis, necrosis, and active cell secretion are thought to contribute to high levels of circular nucleic acids in the blood of some cancer patients. Summary of the Invention
[0003] In one aspect, the present disclosure provides a method for detecting tumor nucleic acid in acellular biological samples from a subject.In some cases, the method comprises: circularizing the nucleic acid from acellular biological samples to produce circularized nucleic acid.In some cases, the method comprises: amplifying the circularized nucleic acid to generate a concatemer comprising at least two copies of the sequence of the circularized nucleic acid.In some cases, the method comprises: sequencing the concatemer or its derivative to obtain the sequence of the concatemer, and the sequencing depth is 18 reads or less.In some cases, the sequencing depth is 18 reads or less per original nucleic acid.In some cases, the sequencing depth is 1 read or less per concatemer.In some cases, the sequencing depth is 1 read or less per circularized nucleic acid.In some cases, the method comprises: processing the sequence of the concatemer to identify at least two occurrences of the tumor-specific sequence variant of the subject. In some cases, the method comprises identifying the nucleic acid as having at least one tumor-specific sequence variant when identifying at least two occurrences of tumor-specific sequence variants in the sequence of the concatemer.In some cases, the method further comprises obtaining tumor-specific sequence variants from the subject.In some cases, the step of obtaining tumor-specific sequence variants comprises sequencing nucleic acids derived from the subject's tumor.In some cases, the step of obtaining tumor-specific sequence variants comprises sequencing nucleic acids derived from the subject's healthy tissue, and comparing the sequence of the nucleic acid derived from the tumor with the sequence of the nucleic acid derived from the healthy tissue.In some cases, the step of obtaining tumor-specific sequence variants comprises sequencing nucleic acids derived from the subject's low-tumor burden tissue or non-tumor burden tissue, and comparing the sequence of the nucleic acid derived from the tumor with the sequence derived from the subject's low-tumor burden tissue or non-tumor burden tissue.In some cases, the sequencing of the nucleic acid derived from the subject's tumor is at a depth of more than 20 reads.In some cases, the sequencing of the nucleic acid derived from the subject's tumor is at a depth of more than 20 reads per original tumor nucleic acid molecule. In some cases, the sequencing of nucleic acids from the subject's tumor is to a depth of greater than 20 reads per nucleotide position.In some cases, the sequencing of the concatemers is at a depth of 10 reads or less. In some cases, the sequencing of the concatemers is at a depth of 5 reads or less. In some cases, the sequencing of the concatemers is at a depth of 2 reads or less. In some cases, the sequencing depth is measured by reads per concatemer. In some cases, the sequencing depth is measured by reads per original nucleic acid molecule. In some cases, the sequencing of the concatemers comprises at least 10 gigabases of sequence. In some cases, the sequencing of the concatemers comprises at least 10 gigabases of the total sequence of the sample. In some cases, the nucleic acid derived from the tumor is subjected to selection prior to sequencing. In some cases, the nucleic acid derived from the healthy tissue is subjected to selection prior to sequencing. In some cases, the selection comprises negative selection to remove non-target sequences from the nucleic acid. In some cases, the selection comprises positive selection to select target sequences from the nucleic acid. In some cases, the method further comprises subjecting the nucleic acid derived from the acellular biological sample to selection prior to circularization. In some cases, the selection comprises negative selection to remove non-target sequences from the nucleic acid. In some cases, the selection comprises positive selection to select target sequences from the nucleic acid. In some cases, circularizing comprises ligating the ends of the nucleic acid or a derivative thereof to each other. In some cases, circularizing comprises attaching an adapter to the 5' end, the 3' end, or the 5' and 3' ends of the nucleic acid or a derivative thereof. In some cases, amplifying the circularized nucleic acid is performed with a polymerase having strand displacement activity. In some cases, amplifying the circularized nucleic acid is performed with a polymerase having 5' to 3' exonuclease activity. In some cases, amplifying is performed with at least one primer from a plurality of random primers. In some cases, amplifying is performed with at least one primer from a plurality of primers designed for whole genome amplification. In some cases, the nucleic acid is single-stranded. In some cases, the nucleic acid is double-stranded.In some cases, the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some cases, sequencing includes (i) contacting a concatemer or a derivative thereof with a plurality of nucleotides in the presence of a polymerase to incorporate one or more nucleotides of the plurality of nucleotides into a growing strand complementary to the concatemer or derivative thereof, and (ii) detecting one or more signals indicative of the incorporation of the one or more nucleotides into the growing strand. In some cases, sequencing includes sequencing by ligation. In some cases, the tumor-specific sequence variants include single-base variants, fusions, insertions, deletions, or epigenetic modifications. In some cases, the acellular biological sample is a bodily fluid. In some cases, the bodily fluid includes urine, saliva, blood, serum, or plasma. In some cases, the tumor is colorectal cancer, pancreatic cancer, ovarian cancer, breast cancer, prostate cancer, bladder cancer, lung cancer, skin cancer, or blood cancer.
[0004] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0005] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, the computer memory comprising machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.
[0006]
[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0007] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]
[0008] The features and advantages of the present invention will be understood by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Figure 1] FIG. 1 shows an exemplary method for tumor-informed disease detection and monitoring using shallow sequencing. [Figure 2] FIG. 1 shows an exemplary method for tumor-informed disease detection and monitoring using shallow sequencing with concatemer-based error correction. [Figure 3] FIG. 1 illustrates a computer system that is programmed or otherwise configured to implement the methods provided herein. [Figure 4] FIG. 1 shows an exemplary method for genome complexity reduction. [Figure 5A] Figure 1 shows an exemplary method for identifying tumor-specific mutations.Tumor tissue and normal tissue (e.g., leukocyte) from the same individual are sequenced, and variants are identified by comparing with the human reference genome database.The variants identified in both normal tissue and tumor tissue are subtracted from the list of tumor tissue variants, and the variants found only in tumor tissue are used as tumor-specific variants for minimal residual disease detection in plasma samples from the same individual. [Figure 5B]Illustrates an exemplary method for identifying tumor-specific mutations. Tumor tissue and post-treatment plasma samples from the same individual are sequenced. Plasma samples are sequenced to a specific depth (i.e., more than 40x, more than 50x, or more than 60x). Variants are identified by comparison with the human reference genome database. Variants identified in plasma with multiple molecular (or lead) support that are also found in tumor tissue are subtracted from the list of tumor tissue variants, and variants found in tumor tissue but not in plasma samples with multiple molecular (lead) support are used as tumor-specific variants for minimal residual disease detection in plasma samples from the same individual. [Figure 6] Figure 1 shows a sequencing workflow without selection: variants are detected in the final step by single-read error correction per molecule. [Figure 7] Figure 1 shows the negative selection sequencing workflow. Blocking oligonucleotides with modified 5' and / or 3' ends that cannot be ligated or extended are added to the ligation mixture at high concentration. The blocking oligos bind to DNA regions with complementary sequences, forming double-stranded regions. These hybrid DNA molecules do not circularize and are optionally removed by DNA exonuclease. Only the circularized DNA is amplified by rolling circle amplification and will be sequenced in the next step. This process can be used to selectively exclude regions in a sequencing library. [Figure 8] 1 shows a positive selection sequencing workflow. Primers targeting target regions are added to a WGS reaction mixture containing random primers. These primers bind to target sequences in the RCA reaction, improving the amplification of these target regions. Compared to the standard workflow, target regions with spike-in primers are amplified more than those without spike-in primers, resulting in these regions receiving more sequencing reads in the final sequencing data. [Figure 9A] FIG. 1 shows the workflow for whole genome sequencing with concatemer error correction. [Figure 9B] FIG. 1 shows the WGS error rates for unfiltered reads, read 1 read 2 corrected reads and healthy human cfDNA samples (N=3) measured by AccuScan. [Figure 10A] FIG. 1 shows the limits of detection for various assay conditions. [Figure 10B] FIG. 10 is a diagram showing the false positive rate when VAF=0. [Figure 10C] FIG. 1 shows the analytical sensitivity of AccuScan using a healthy sample mixture. [Figure 10D] FIG. 1 shows titration of cancer samples. [Figure 11A] Figure 11 shows two workflows for AccuScan MRD detection: Figure 11A shows a comparison of tumor tissue and blood cells. [Figure 11B] Figure 11B shows two workflows for AccuScan MRD detection: Figure 11B shows a comparison between tumor tissue and post-treatment plasma; [Figure 11C] FIG. 1 shows the total number of tumor-specific markers identified with and without leukocytes. [Figure 11D] FIG. 1 shows the mutation profile of tumor-specific markers identified with and without leukocytes. [Figure 11E] FIG. 1 shows a comparison of VAF measured in plasma using a tumor-WBC workflow versus a tumor-plasma workflow. [Figure 11F] FIG. 1 shows MRD calling using tumor-WBC workflow versus tumor-plasma workflow. [Figure 12A] FIG. 1 shows the number of tumor-specific variants identified in CRC, ESCC, and melanoma. [Figure 12B] FIG. 1 shows the VAF of pre-treatment plasma samples. [Figure 12C]FIG. 1 shows the results of ESCC MRD detection in samples taken one week after surgery. [Figure 12D] FIG. 1 shows all pre-imaging CRC recurrences detected by AccuScan. [Figure 12E] Kaplan-Meier disease-free survival analysis of CRC and ESCC surgery patients. Patients with ctDNA+ in post-surgery plasma samples showed significantly shorter disease-free survival. [Figure 13A] FIG. 1 shows AccuScan for IO monitoring. [Figure 13B] FIG. 1 shows AccuScan for IO monitoring: dynamic changes of ctDNA over time. [Figure 14A] Figure 14 shows the analytical sensitivity and specificity of AccuScan. Figure 14A shows a simulation to predict the theoretical detection rate under different sequencing coverage as a function of cTAF, using 5000, 20000, 40000, and 80000 markers and two different error rates. The 4.2 × 10-7 error rate showed higher sensitivity than the 2.8 × 10-5 error rate under the same sequencing depth. The detection rate is calculated as the proportion of tests considered MRD-positive, with the nominal specificity set at 99%. [Figure 14B] Figure 14B shows the analytical sensitivity and specificity of AccuScan. Figure 14B shows a simulation predicting theoretical specificity using 5,000, 20,000, 40,000, and 80,000 markers and two different error rates, with a nominal specificity set at 99%. Specificity is calculated as the proportion of tests considered MRD-negative when the cTAF is 0. [Figure 15A] Figure 15A shows ddPCR of melanoma cancer cfDNA samples. Figure 15A shows ddPCR of the original melanoma cancer cfDNA samples. [Figure 15B] Figure 15A shows ddPCR of melanoma cancer cfDNA samples. Figure 15B shows ddPCR of healthy plasma samples. [Figure 15C]Figure 15C shows ddPCR of melanoma cancer cfDNA samples. Figure 15C shows ddPCR of diluted melanoma cancer cfDNA samples in a healthy plasma background with an expected cTAF of 0.1%. [Figure 16] FIG. 1 shows the VAF of all ctDNA-positive plasma samples. [Figure 17] Figure 1 shows AccuScan error rates. Overall error rate and error rate for each variant type from AccuScan WGS data for healthy human cfDNA samples (N=3) sequenced with paired-end 150 read length or single-end 300 read length using 300 cycles of sequencing reagents. DETAILED DESCRIPTION OF THE INVENTION
[0009] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be utilized.
[0010] As used herein, the terms "about" or "approximately" mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which may depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, depending on practice in the art, "about" can mean within one or more standard deviations. As another example, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. With respect to biological systems or processes, the term "about" can mean within an index, such as within 5-fold of a value or within 2-fold of a value. When specific values are described in this application and claims, unless otherwise specified, the term "about" means within an acceptable error range for the particular value.
[0011] As used herein, the terms "polynucleotide," "nucleotide," "nucleotide sequence," "nucleic acid," and "oligonucleotide" are used interchangeably and generally refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs thereof. Polynucleotides may have any three-dimensional structure and may perform any function. The following are non-limiting examples of polynucleotides: Cell-free nucleic acids, cell-free DNA (cfDNA), cell-free RNA (cfRNA), circular tumor DNA (ctDNA), circular tumor RNA (ctRNA), coding or non-coding regions of genes or gene fragments, loci (locuses) defined by linkage analysis (locuses), exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide elements. Polynucleotides may be further modified after polymerization, such as by conjugation with a labeling element.
[0012] The term "subject" as used herein generally refers to a vertebrate, such as a mammal (e.g., a human). Mammals include, but are not limited to, mice, monkeys, humans, farm animals, sport animals, and pets (e.g., dogs or cats). Also encompassed are tissues, cells, and products of biological entities obtained in vivo or cultured in vitro. A subject may be a patient. A subject may be symptomatic for a disease (e.g., cancer). Alternatively, a subject may be asymptomatic for a disease.
[0013] As used herein, the term "biological sample" generally refers to a sample derived from or obtained from a subject, such as a mammal (e.g., a human). Biological samples may include, but are not limited to, hair, fingernails, skin, sweat, tears, ocular fluid, nasal or nasopharyngeal swabs, sputum, throat swabs, saliva, mucus, blood, serum, plasma, placental fluid, amniotic fluid, umbilical cord blood, emphatic fluids, body cavity fluids, earwax, oil, glandular secretions, bile, lymph, pus, bacterial flora, meconium, breast milk, bone marrow, bone, CNS tissue, cerebrospinal fluid, adipose tissue, synovial fluid, stool, gastric juice, urine, semen, vaginal secretions, stomach, small intestine, large intestine, rectum, pancreas, liver, kidney, bladder, lung, and other tissues and fluids derived from or obtained from a subject. A biological sample may be a cell-free (or acellular) biological sample.
[0014] The term "acellular biological sample" as used herein generally refers to a sample that does not contain cells, derived from or obtained from a subject. Acellular biological samples may include, but are not limited to, blood, serum, plasma, nasal swab or nasopharyngeal wash, saliva, urine, gastric juice, tears, stool, mucus, sweat, earwax, oil, glandular secretions, bile, lymph, cerebrospinal fluid, tissue, semen, vaginal fluid, interstitial fluid (including interstitial fluid derived from tumor tissue), intraocular fluid, spinal fluid, throat swab, exhaled breath, hair, fingernail, skin, biopsy, placental fluid, amniotic fluid, umbilical cord blood, sputum, body cavity fluid, sputum, pus, bacterial flora, meconium, breast milk, and / or other excretions.
[0015] When the terms "at least," "greater than," or "greater than or equal to" precede the first number in a series of two or more numbers, the terms "at least," "greater than," or "greater than or equal to" always apply to each and every number in the series. For example, 1, 2, or 3 or more is equal to 1 or more, 2 or more, or 3 or more.
[0016] When the term "no more than," "less than," or "less than or equal to" appears before the first number in a series of two or more numbers, the term "no more than," "less than," or "less than or equal to" always applies to each and every number in the series. For example, 3, 2, or 1 or less is equal to 3 or less, 2 or less, or 1 or less.
[0017] Molecular residual disease (MRD) refers to cancer cells that persist after therapeutic treatment. Timely and sensitive measurement of MRD is important for recurrence risk assessment, treatment prognosis, and patient stratification. Circular tumor DNA (ctDNA), released by cancer cells and possessing a short half-life (less than 2 hours), has emerged as a promising real-time biomarker for MRD detection and monitoring. Studies have shown that the level of cancer-specific somatic mutations in ctDNA correlates with tumor stage, burden, and response to treatment across tumor types. Compared with other blood-based cancer biomarkers, such as circulating tumor cells and cancer antigens, ctDNA provides a more sensitive and specific measurement of MRD.
[0018] Currently, there are two main strategies for ctDNA-based MRD detection: 1) tumor-naive approaches, which test MRD samples for alterations known to be enriched in tumors, such as common somatic mutations and methylation changes, and 2) tumor-informed approaches, which require tumor samples to identify patient-specific variants and then test MRD samples for those variants.
[0019] Tumor-naive approaches are logistically simpler, using universal panels to test plasma samples for the presence of cancer signals without the need to obtain and sequence tumor samples. These tests offer operational convenience but tend to have moderate limits of detection (LODs). Methylation-based cancer detection tests claim a 50% sensitivity for a circular tumor allelic fraction (cTAF) of 3.1 x 10-4. Detection of 55.6% of CRC recurrences has also been demonstrated using plasma collected at a landmark timepoint (4 weeks after surgery) with a panel combining methylation and mutation signals.
[0020] Tumor-informed approaches incorporate patient-specific somatic mutation information from tumor tissue into MRD analysis, which can lead to ultralow detection limits. Factors that affect sensitivity include the accuracy of somatic mutation determination from tissue and plasma samples and the total number of cfDNA molecules interrogated, which is the product of the number of somatic variants tracked and the inherent molecular depth achieved by sequencing.
[0021] Tumor-informed approaches can use either custom MRD tests or off-the-shelf MRD tests. Custom MRD assays are designed after tumor results are available and follow a limited number of variants through ultra-deep sequencing. Sequencing of custom panels can be exhaustive, and therefore the inherent molecular depth is most often limited by the amount of input material available. For example, Signatera, a tumor-informed NGS-based multiplex PCR assay tracking 16 personalized markers, achieved a 10-fold increase in MRD accuracy when using up to 66 ng of DNA. -4 Achieved analytical sensitivities of 81.3% to 96.1% with limits of detection (LOD) of 10. Tumor-informed personalized MRD assays targeting multiple markers and employing error correction using UMI or duplex sequencing have been shown to be effective in 10 -4Phase-Seq uses multiple somatic mutations in individual DNA fragments to detect ctDNA, which reduces background noise by 10%. -6 The sensitivity has been reduced to below 1000 kJ / s, lowering the claimed detection limit to PPM levels given sufficient phase change variants. Tumor-informed, custom-tailored MRD approaches can achieve very high sensitivity, but the requirement for individualized design significantly increases turnaround time (TAT) and poses significant logistical challenges.
[0022] Tumor-informed, off-the-shelf methods use the same assay for both tumor and plasma in every patient. Without the need for patient-specific reagents, these methods share the low turnaround time of tumor-naive approaches and offer much simpler logistics than custom-built methods. The challenge is generating off-the-shelf assays with sufficient genome coverage and a sufficiently low error rate. Predesigned MRD panels targeting cancer-related genes typically use UMI with deep sequencing to achieve high accuracy in variant identification, but the number of markers tracked by these panels for each patient is sparse. For example, a 130-kb panel covering 139 important lung cancer-related genes captures a median of only two mutations per patient (range: 1–8 mutations).
[0023] In recent years, whole genome sequencing (WGS) assays have emerged as an innovative approach for cancer screening and MRD detection. Tumor-informed WGS MRD assays use genome breadth to complement sequencing depth for sensitivity and overcome limitations in input sample quantity. UMI-based error correction, which relies on having multiple reads per input molecule, would be too costly at the WGS scale. Some estimate the WGS somatic single-nucleotide variant (SNV) error rate at 4.96 × 10 -5 A read-centric SVM model was used to reduce the number of somatic mutations to 10 by utilizing the cumulative signal of thousands of somatic mutations observed in the tumor genome. -4reported an analytical sensitivity of 95% in a tumor fraction of 10. Other whole-genome techniques using duplex sequencing have yielded results in 10 -7 Although ultra-low error rates have been demonstrated at levels below 100 kJ / s, these methods suffer from low conversion rates, making it difficult to achieve low LODs. Therefore, efficient and cost-effective genome-wide error correction methods are needed to enable WGS for MRD detection at low LODs.
[0024] DNA concatemers generated by rolling circle amplification (RCA) physically link DNA copies, enabling error correction at the single-read level. The combination of RCA with replicate confirmation eliminates both PCR and sequencing errors. Compared to UMI methods, concatemer sequencing has shown higher error correction efficiency when applied to genomic DNA. Recently, concatemer sequencing has been adapted for liquid biopsies, demonstrating the feasibility of applying this technology to treatment selection and cancer screening. A WGS solution for ctDNA detection utilizing concatemer sequencing for genome-wide single-read error suppression is provided herein, enabling rapid and sensitive MRD detection and monitoring in plasma samples from cancer patients.
[0025] In some embodiments, a method for detecting tumor nucleic acid in a biological sample from a subject is provided herein. Optionally, the method includes detecting tumor nucleic acid in an acellular biological sample from the subject. Optionally, the method includes circularizing nucleic acid from a biological sample, such as an acellular biological sample, to produce a circularized nucleic acid. The method can then include amplifying the circularized nucleic acid to generate a concatemer comprising at least two copies of the sequence of the circularized nucleic acid. The concatemer or its derivative can then be sequenced to obtain the sequence of the concatemer. Optionally, the sequencing is performed at a depth of 18 reads or less. The sequence of the concatemer is then processed to identify at least two occurrences of the tumor-specific sequence variant of the subject. When identifying at least two occurrences of the tumor-specific sequence variant in the sequence of the concatemer, the method can include identifying the nucleic acid as having at least one tumor-specific sequence variant. The method can further include obtaining tumor-specific sequence variants from the subject, for example, by sequencing nucleic acid derived from the subject's tumor. In some cases, the method further comprises sequencing the nucleic acid from the healthy tissue of the subject, and comparing the sequence from the nucleic acid from the tumor with the sequence from the healthy tissue of the subject.In some cases, the sequencing of the nucleic acid from the tumor is carried out at an appropriate depth, measured by per molecule or per read, which are used interchangeably herein.In some cases, the sequencing of the nucleic acid from the tumor of the subject is at a depth of more than 20 reads.In some cases, the sequencing of the nucleic acid from the tumor of the subject is at a depth of more than 25 reads.In some cases, the sequencing of the nucleic acid from the tumor of the subject is at a depth of more than 30 reads.In some cases, the sequencing of the nucleic acid from the tumor of the subject is at a depth of more than 35 reads.In some cases, the sequencing of the nucleic acid from the tumor of the subject is at a depth of more than 40 reads.
[0026] In another embodiment of the method for detecting tumor nucleic acid in a biological sample herein, the concatemer is sequenced to an appropriate depth, measured in reads per molecule or reads, which are used interchangeably herein.In some cases, the concatemer is sequenced to a depth of 18 reads or less.In some cases, the concatemer is sequenced to a depth of 15 reads or less.In some cases, the concatemer is sequenced to a depth of 12 reads or less.In some cases, the concatemer is sequenced to a depth of 10 reads or less.In some cases, the concatemer is sequenced to a depth of 9 reads or less.In some cases, the concatemer is sequenced to a depth of 8 reads or less.In some cases, the concatemer is sequenced to a depth of 7 reads or less.In some cases, the concatemer is sequenced to a depth of 6 reads or less.In some cases, the concatemer is sequenced to a depth of 5 reads or less.In some cases, the concatemer is sequenced to a depth of 4 reads or less. In some cases, the sequencing of the concatemers is 3 reads deep or less. In some cases, the sequencing of the concatemers is 2 reads deep or less. In some cases, the sequencing of the concatemers is 1 read deep or less. In some cases, the sequencing of the concatemers is whole genome sequencing. In some cases, the sequencing of the concatemers includes at least 10 gigabases of sequence.
[0027] In another embodiment of detecting tumor nucleic acids in a biological sample herein, nucleic acids derived from a tumor are subjected to selection before sequencing. Optionally, nucleic acids derived from a healthy tissue are subjected to selection before sequencing. Optionally, nucleic acids derived from an acellular biological sample are subjected to selection before circularizing the nucleic acids. Optionally, the selection comprises negative selection to remove non-target sequences from the nucleic acids. Optionally, the negative selection comprises contacting the nucleic acids with a blocker that binds to the non-target sequences, and amplifying, ligating, or capturing nucleic acids that are not bound to the blocker. Optionally, the blocker comprises an oligonucleotide. Optionally, the negative selection comprises contacting the nucleic acids with a nuclease that specifically cleaves the non-target sequences. Optionally, the nuclease is a clustered regularly interspaced short palindromic repeats (CRISPR) nuclease. Optionally, the selection comprises positive selection to select target sequences from the nucleic acids. Optionally, the positive selection comprises hybrid capture. Optionally, the positive selection comprises amplification. Optionally, the amplification comprises polymerase chain reaction (PCR).
[0028] In another embodiment of the method for detecting tumor nucleic acid in a biological sample herein, optionally, circularizing the nucleic acid derived from the biological sample comprises ligating the ends of the nucleic acid or its derivative to each other.Optionally, circularizing the nucleic acid derived from the biological sample comprises attaching an adaptor to the 5' end, the 3' end, or the 5' end and the 3' end of the nucleic acid or its derivative.
[0029] In another embodiment of the method for detecting tumor nucleic acid in a biological sample herein, optionally, amplifying the circularized nucleic acid to generate concatemers a is performed using a polymerase with strand displacement activity. Optionally, amplifying the circularized nucleic acid to generate concatemers a is performed using a polymerase with 5'-3' exonuclease activity. Optionally, amplifying is performed using at least one primer from a plurality of random primers. Optionally, amplifying is performed using at least one primer from a plurality of primers designed for whole genome amplification.
[0030] In another embodiment of the method for detecting tumor nucleic acid in biological sample herein, optionally, the nucleic acid in biological sample is single-stranded.Optionally, the nucleic acid is double-stranded.Optionally, the nucleic acid in biological sample is a mixture of single-stranded nucleic acid and double-stranded nucleic acid.Optionally, the nucleic acid is made single-stranded before circularization.Optionally, the nucleic acid is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination of DNA and RNA.Optionally, the nucleic acid is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination of DNA and RNA.Optionally, the nucleic acid in biological sample is a mixture of single-stranded nucleic acid and double-stranded nucleic acid.Optionally, the nucleic acid is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination of DNA and RNA.Optionally, the nucleic acid is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or double-stranded nucleic acid ...
[0031] In another aspect of the methods of detecting tumor nucleic acids in a biological sample herein, optionally, sequencing the concatemers comprises contacting the concatemers, or derivatives thereof, with a plurality of nucleotides in the presence of a polymerase to incorporate one or more nucleotides of the plurality of nucleotides into a growing strand complementary to the concatemers, or derivatives thereof, and detecting one or more signals indicative of the incorporation of the one or more nucleotides into the growing strand. Sequencing the concatemers may comprise any suitable method provided herein. Alternatively, or in combination, sequencing the concatemers comprises sequencing by ligation. Sequencing the concatemers may comprise any suitable method provided herein.
[0032] In another aspect of the methods of detecting tumor nucleic acid in a biological sample provided herein, optionally the tumor-specific sequence variant comprises a single nucleotide variant, a fusion, an insertion, a deletion, an epigenetic modification, or any combination thereof.
[0033] In another aspect of the method for detecting tumor nucleic acid in a biological sample provided herein, the biological sample may be an acellular biological sample. In some cases, the acellular biological sample is a bodily fluid. In some cases, the bodily fluid includes urine, saliva, blood, serum, or plasma. In some cases, the biological sample, acellular biological sample, or bodily fluid may be any suitable sample provided herein.
[0034] In another embodiment of the methods for detecting tumor nucleic acid in a biological sample provided herein, optionally the tumor is colorectal cancer, pancreatic cancer, ovarian cancer, breast cancer, prostate cancer, bladder cancer, lung cancer, skin cancer, or blood cancer. Optionally, the tumor is any cancer suitable for detection provided herein.
[0035] Library preparation and amplification methods In certain cases, the methods herein include amplifying polynucleotides present in a sample from a subject. The amplification method used herein often includes rolling circle amplification. Alternatively, or in combination, the amplification method used herein includes PCR. In some cases, the amplification method used herein includes linear amplification. In many cases, the amplification does not target a single gene or set of genes, but rather the entire nucleic acid sample is amplified. In some cases, the method includes (a) circularizing a plurality of individual polynucleotides to form a plurality of circular polynucleotides, each of which has a linkage between its 5' and 3' end, and (b) amplifying the circular polynucleotides of (a) to produce amplified polynucleotides. In further cases, the amplification method includes (c) fragmenting the amplified polynucleotides to produce fragmented polynucleotides, each of which includes one or more fragmentation points at its 5' and / or 3' end. In some cases, the method does not include enrichment of the target sequence.
[0036] Generally, the ends of a polynucleotide are joined together to form a circular polynucleotide (either directly or using one or more intermediate adaptor oligonucleotides), producing a junction having a junction sequence. When the 5' and 3' ends of a polynucleotide are joined via an adaptor polynucleotide, the term "junction" refers to the junction between the polynucleotide and the adaptor (e.g., one of the 5'-end junction or the 3'-end junction), or the junction between the 5' and 3' ends of the polynucleotide formed by and including the adaptor oligonucleotide. When the 5' and 3' ends of a polynucleotide are joined without an intervening adaptor (e.g., the 5' and 3' ends of a single-stranded DNA), the term "junction" refers to the point where these two ends are joined. A junction may be identified by the sequence of nucleotides comprising the junction (also referred to as the "junction sequence").
[0037] As used herein, samples include polynucleotides with a mixture of ends formed by natural degradative processes (such as cell lysis, cell death, and other processes in which polynucleotides, such as DNA and RNA, are released from cells into their surrounding environment (e.g., cell-free polynucleotides, such as cell-free DNA and cell-free RNA), where they may be further degraded); fragmentation that is a by-product of sample processing (such as fixation, staining, and / or storage procedures); and fragmentation by methods that cleave DNA without being restricted to a specific target sequence (e.g., mechanical fragmentation, such as by sonication, or treatment with non-sequence-specific nucleases, such as DNase I or fragmentase). When a sample contains polynucleotides with a mixture of ends, it is unlikely that two polynucleotides will have identical 5' or 3' ends, and it is even less likely that two polynucleotides will independently have both identical 5' and 3' ends. Thus, in some embodiments, binding moieties may be used to distinguish between different polynucleotides, even if they contain portions with the same target sequence. When the ends of polynucleotides are joined without an intervening adaptor, the sequence of the junction can be identified by aligning with a reference sequence.For example, if the sequence order of two elements appears to be reversed relative to the reference sequence, the point where this reversal appears to occur may indicate the presence of a junction at that point.When the ends of polynucleotides are joined via one or more adaptor sequences, the junction can be identified by being close to a known adaptor sequence, or by aligning as described above when sequencing reads are long enough to obtain sequences from both the 5'-end and the 3'-end of the circularized polynucleotide.In some embodiments, the formation of a specific junction is a sufficiently rare event so that it is unique among the circularized polynucleotides of a sample.
[0038] In some embodiments, circularization of individual polynucleotides in (a) is performed by subjecting the plurality of polynucleotides to a ligation reaction. The ligation reaction may include a ligase enzyme. In some cases, the ligase enzyme is a single-stranded DNA ligase or an RNA ligase. In some cases, the ligase enzyme is a double-stranded DNA ligase. In some embodiments, the ligase enzyme is degraded prior to amplification in (b). Degrading the ligase prior to amplification in (b) can increase the recovery rate of amplifiable polynucleotides. In some embodiments, the plurality of circularized polynucleotides is not purified or isolated prior to (b). In some embodiments, non-circularized linear polynucleotides are degraded prior to amplification. In some cases, the plurality of polynucleotides is denatured prior to circularization to generate single-stranded polynucleotides; in some cases, the plurality of polynucleotides is not denatured prior to circularization.
[0039] In some cases, circularizing in (a) includes attaching an adaptor polynucleotide to the 5' end, the 3' end, or both the 5' and 3' ends of a polynucleotide of the plurality of polynucleotides. As discussed above, when the 5' and / or 3' end of a polynucleotide is attached via an adaptor polynucleotide, the term "junction" refers to the junction between the polynucleotide and the adaptor (e.g., one of the 5' end attachment or the 3' end attachment), or the junction between the 5' end and the 3' end of the polynucleotide formed by and including the adaptor oligonucleotide.
[0040] In some cases, polynucleotides are subjected to a selection step. In some cases, polynucleotides having a desired sequence are subjected to a positive selection step to enrich for polynucleotides having the desired sequence. Alternatively, polynucleotides having an unwanted sequence are subjected to a negative selection step to remove polynucleotides having the unwanted sequence. In some cases, the negative selection includes denaturing the polynucleotide to generate single-stranded polynucleotides, annealing one or more blocking oligonucleotides to the polynucleotide to generate double-stranded polynucleotides and single-stranded polynucleotides having the unwanted sequence, and circularizing the single-stranded polynucleotides. In some cases, the blocking oligonucleotides have modified 5' and / or 3' ends that do not allow ligation. In some cases, the blocking oligonucleotides have modified 5' and / or 3' ends that do not allow extension. In some cases, the linear double-stranded polynucleotides are removed using an exonuclease. The circularized polynucleotides can be used in subsequent rolling circle amplification and sequencing steps.
[0041] In one aspect, a method for identifying sequence variants in a plurality of polynucleotides is provided herein, comprising: denaturing a plurality of polynucleotides; annealing one or more blocking oligonucleotides to a polynucleotide having an unwanted sequence; and circularizing the resulting single-stranded polynucleotide. Optionally, the remaining linear polynucleotides annealed to the blocking oligonucleotides are degraded using a nuclease, such as a DNA exonuclease. The circularized polynucleotides can then be amplified by rolling circle amplification, resulting in concatemers containing multiple copies of the original polynucleotide. Optionally, rolling circle amplification is performed using random primers. Optionally, rolling circle amplification is performed using target-specific primers. The concatemers are then sequenced to obtain sequencing reads. These sequencing reads are used to identify variants. Optionally, variants are identified when present on multiple copies of a polynucleotide in a concatemer. Optionally, variants are identified when present on two different concatemers.
[0042] The circularized polynucleotide may optionally be amplified, for example, after degradation of the ligase enzyme, to obtain an amplified polynucleotide. The amplification of the circular polynucleotide in (b) may be performed by a polymerase. Optionally, the polymerase is a polymerase with strand displacement activity. Optionally, the polymerase is Phi29 DNA polymerase. Alternatively, the polymerase is a polymerase without strand displacement activity. Optionally, the polymerase is T4 DNA polymerase or T7 DNA polymerase. Alternatively, or in combination, the polymerase is Taq polymerase or a polymerase of the Taq polymerase family. Optionally, the amplification comprises rolling circle amplification (RCA). The amplified polynucleotide resulting from RCA may comprise a linear concatemer or a polynucleotide comprising multiple copies of a target sequence (e.g., a subunit sequence) from the template polynucleotide. In some embodiments, the amplification comprises exposing the circular polynucleotide to an amplification reaction mixture comprising random primers. In some cases, amplifying comprises exposing the circular polynucleotide to an amplification reaction mixture comprising one or more primers, each of which specifically hybridizes to a different target sequence by sequence complementarity. In some cases, amplifying comprises exposing the circular polynucleotide to an amplification reaction mixture comprising a reverse primer.
[0043] The amplified polynucleotides are fragmented, in some cases producing fragmented polynucleotides that are shorter in length than the unfragmented polynucleotides. Two or more fragmented polynucleotides that originate from the same linear concatemer may have the same junction sequence but may have different 5' and / or 3' ends (e.g., fragmented ends).
[0044] Cell-free polynucleotides from a sample may be any of a variety of polynucleotides, including, but not limited to, DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), short interfering RNA (siRNA), fragments of any of these, or combinations of any two or more of these. In some embodiments, the sample contains DNA. In some embodiments, the sample contains cell-free genomic DNA. In some embodiments, the sample contains DNA produced by amplification, such as by a primer extension reaction using any appropriate combination of primers and DNA polymerase, including, but not limited to, polymerase chain reaction (PCR), reverse transcription, and combinations thereof. When the template for the primer extension reaction is RNA, the product of reverse transcription is referred to as complementary DNA (cDNA). Primers useful in primer extension reactions can contain one or more target-specific sequences, random sequences, partially random sequences, and combinations thereof. In some cases, the primers contain a mixture of random sequences and one or more target-specific sequences. Generally, the sample polynucleotide includes any polynucleotide present in the sample, which may or may not include a target polynucleotide. The polynucleotide may be single-stranded, double-stranded, or a combination thereof. In some embodiments, the polynucleotide subjected to the method of the present disclosure is a single-stranded polynucleotide, which may or may not be present in the presence of a double-stranded polynucleotide. In some embodiments, the polynucleotide is single-stranded DNA. The single-stranded DNA (ssDNA) may be ssDNA isolated in single-stranded form, or DNA isolated in double-stranded form and then made single-stranded for the purpose of one or more steps in the method of the present disclosure.
[0045] In one aspect, the present disclosure provides a method for identifying sequence variants in a plurality of polynucleotides, comprising: denaturing polynucleotides; circularizing the resulting linear polynucleotides; and amplifying the resulting circular polynucleotides; wherein the amplifying step is used to enrich the sequence of interest, for example, by adding one or more primers that bind to the sequence of interest to an amplification reaction that contains random primers.Using random primers and the primers that bind to the sequence of interest, circular polynucleotides are amplified by rolling circle amplification to create concatemers.The concatemers are then sequenced to obtain sequencing reads.These sequencing reads are used to identify variants.In some cases, variants are identified when they exist on multiple copies of polynucleotides in concatemers.In some cases, variants are identified when they exist on two different concatemers.
[0046] In some embodiments, polynucleotides are subjected to subsequent steps (e.g., circularization and amplification) without extraction and / or purification steps. For example, a liquid sample may be treated to remove cells without extraction steps to produce a purified liquid sample and a cell sample, followed by isolation of DNA from the purified liquid sample. Various procedures are available for isolating polynucleotides, such as by precipitation or nonspecific binding to a substrate, followed by washing the substrate to release bound polynucleotides. When polynucleotides are isolated from a sample without a cell extraction step, the majority of the polynucleotides are extracellular or "cell-free" polynucleotides, such as cell-free DNA and cell-free RNA, and represent dead or damaged cells. The identity of such cells may be used to characterize derived cells or cell populations, such as tumor cells (e.g., in cancer detection), fetal cells (e.g., in prenatal diagnosis), cells from transplanted tissue (e.g., in early detection of transplant failure), or members of a microbial community.
[0047] When sample is processed to extract polynucleotides from cells in the sample, for example, various extraction methods are available.For example, nucleic acid can be purified by organic extraction using phenol, phenol / chloroform / isoamyl alcohol, or similar preparations including Trizol and Tri reagent.Other non-limiting examples of extraction techniques include: (1) organic extraction followed by ethanol precipitation, for example, using phenol / chloroform organic reagent (Ausubel et al., 1993, the entire contents of which are incorporated herein by reference), with or without the use of an automated nucleic acid extractor, for example, Model 341 DNA Extractor available from Applied Biosystems (Foster City, California); (2) solid phase adsorption method (U.S. Patent No. 5,234,809, Walsh et al., 1991, each of which is incorporated herein by reference in its entirety); and (3) salt-induced nucleic acid precipitation method, typically referred to as "salting out" method (Miller et al., (1988), the entire contents of which are incorporated herein by reference). Another example of nucleic acid isolation and / or purification includes the use of magnetic particles to which nucleic acids can be specifically or nonspecifically bound, followed by isolating the beads using a magnet, washing the beads, and eluting the nucleic acids from the beads (see, e.g., U.S. Pat. No. 5,705,628, incorporated herein by reference in its entirety). In some embodiments, the isolation method may be followed by an enzymatic digestion step, such as digestion with proteinase K or other similar proteases, to remove unwanted proteins from the sample. See, e.g., U.S. Pat. No. 7,001,724, incorporated herein by reference in its entirety. If necessary, an RNase inhibitor may be added to the lysis buffer. For certain cell or sample types, it may be desirable to add a protein denaturation / digestion step to the protocol. The purification method may involve isolating DNA, RNA, or both. If both DNA and RNA are isolated together during or after the extraction procedure, additional steps may be used to purify one or both separately from the other.Subfractions of extracted nucleic acids can also be produced by purification, for example, by size, sequence, or other physical or chemical properties. In addition to the initial nucleic acid isolation step, nucleic acid purification can be performed after any step of the disclosed method, for example, to remove excess or unnecessary reagents, reactants, or products. Various methods are available for determining the amount and / or purity of nucleic acids in a sample, such as by absorbance (e.g., absorbance at 260 nm, 280 nm, and ratios thereof) and detection of labels (e.g., fluorescent dyes and intercalators such as Cyber Green, Cyber Blue, DAPI, propidium iodide, Hoechst stain, Cyber Gold, and ethidium bromide).
[0048] In some cases, the methods herein include preparing a DNA library from polynucleotides. For example, the methods herein include preparing a single-stranded DNA library. Any suitable method for preparing a single-stranded DNA library may be used in the methods herein. For example, a method for preparing a single-stranded DNA library includes denaturing a DNA sample to generate multiple ssDNAs, ligating an adaptor to the 3' end of the ssDNA molecules or extending the 3' end of the ssDNA molecules by non-template synthesis, synthesizing a second strand using a primer complementary to the adaptor or the 3' extension sequence, ligating a double-stranded adaptor to the extension product, amplifying the second strand using primers targeting the first adaptor and the second adaptor (for example, using PCR), and sequencing the library on a sequencer. For example, a further method of single-stranded library preparation includes denaturing a DNA sample to generate multiple ssDNA molecules, ligating adapters to the 3' ends of the ssDNA molecules, synthesizing the second strand by using a primer complementary to the adapter, ligating a double-stranded adapter to the extension product, amplifying the second strand (e.g., using PCR) using primers targeting the first adapter and the second adapter, optionally enriching for regions of interest using hybridization with a capture probe, amplifying the capture products (e.g., using PCR), and sequencing the library on a sequencer.
[0049] Further examples of single-stranded library preparation include treating DNA with a heat-sensitive phosphatase to remove residual phosphate groups from the 5' and 3' ends of DNA strands, removing deoxyuracil from the DNA strands resulting from cytosine deamination, ligating a 5' phosphorylated adapter oligonucleotide with a 3' biotin-labeled spacer arm of about 10 nucleotides and longer to the 3' end of the DNA strands, immobilizing the adapter-ligated molecules on streptavidin beads, replicating the template strand using Bst polymerase and a 5' tail primer complementary to the adapter, washing away excess primer, removing the 3' overhang using T4 DNA polymerase, joining a second adapter to the newly synthesized strand using blunt-end ligation, washing away excess adapter, releasing the library molecules by heat denaturation, and adding the full-length adapter sequence including the barcode by amplification using the tail primer. and sequencing the library as described in J. Immunol. 1999, 14, 1143-1146, 1999.
[0050] In a further embodiment, the method herein includes preparing a double-stranded DNA library. Any suitable method for preparing a double-stranded DNA library may be used in the method herein. For example, a method for preparing a double-stranded DNA library includes ligating sequencing adaptors to the 5' and 3' ends of multiple DNA fragments and sequencing the library on a sequencer. A further method for preparing a double-stranded DNA library includes ligating adaptors to the 5' and 3' ends of DNA fragments, linking the full-length adaptor sequences to the ligated fragments by PCR using primers complementary to the ligated adaptors, and sequencing the library on a sequencer. A further method includes ligating adaptors to the 5' and 3' ends of DNA fragments, amplifying the ligated products by PCR complementary to the ligated adaptors, optionally enriching the target region by hybridization with a capture probe, PCR-amplifying the capture products, and sequencing the library on a sequencer. Further methods of double-stranded library preparation include ligating adapters to the 5' and 3' ends of the DNA fragments, amplifying the ligated products by PCR using primers complementary to the ligated adapters, circularizing the double-stranded PCR products or denaturing and circularizing the single-stranded PCR products, optionally enriching for regions of interest by PCR using primers targeting specific genes, and sequencing the library on a sequencer.
[0051] Further examples of double-stranded library preparation include the safe sequencing system described by Kinde et al. (Kinde et al. 2011. Proc. Natl. Acad. Sci., USA, 108(23)9530-9535, which is incorporated herein by reference in its entirety), which involves assigning a unique identifier (UID) to each template molecule, amplifying each uniquely tagged template molecule, and redundantly sequencing the amplification products to generate a UID family. Further examples include the circular single molecule amplification and differential sequencing technology (cSMART) described by Lv et al. (Lv et al. 2015. Clin. Chem., 61(1)172-181, which is incorporated herein by reference in its entirety), which tags a single molecule with a unique barcode, circularizes it, targets alleles for replication by inverse PCR, and then sequences the prepared library to count the alleles present.
[0052] In another library preparation method, the cfDNA fragments with certain characteristics are selected using antibody.In some cases, methylated or hypermethylated cfDNA fragments are selected using antibody.Then, the selected cfDNA fragments are used in any library preparation method described herein, including circularization, single-stranded DNA library preparation and double-stranded DNA library preparation.By sequencing this isolated cfDNA fragment, information about the characteristics present in cfDNA is obtained, including modification such as methylation or hypermethylation.
[0053] According to some embodiments, polynucleotides in a plurality of polynucleotides from a sample are circularized. Circularization may involve ligating the 5' end of a polynucleotide to the 3' end of the same polynucleotide, to the 3' end of another polynucleotide in the sample, or to the 3' end of a polynucleotide from a different source (e.g., an artificial polynucleotide such as an oligonucleotide adaptor). In some embodiments, the 5' end of a polynucleotide is ligated to the 3' end of the same polynucleotide (also referred to as "self-ligation"). In some embodiments, the conditions of the circularization reaction are selected to promote self-ligation of polynucleotides within a specific length range to produce a population of circularized polynucleotides of a specific average length. For example, the circularization reaction conditions may be selected to promote self-ligation of polynucleotides shorter than about 5000, about 2500, about 1000, about 750, about 500, about 400, about 300, about 200, about 150, about 100, or about 50 nucleotides in length. In some embodiments, fragments having lengths of 50 to 5,000 nucleotides, 100 to 2,500 nucleotides, or 150 to 500 nucleotides are preferred, with the average length of the circularized polynucleotide falling within each range. In some embodiments, 80% or more of the circularized fragments are 50 to 500 nucleotides in length, e.g., 50 to 200 nucleotides in length. Reaction conditions that may be optimized include the length of time allotted for the ligation reaction, the concentrations of various reagents, and the concentration of the polynucleotides to be ligated. In some embodiments, the circularization reaction preserves the distribution of fragment lengths present in the sample before circularization. For example, one or more of the mean, median, mode, and standard deviation of the fragment lengths in the sample before circularization and the circularized polynucleotides are within 75%, 80%, 85%, 90%, 95%, or more of each other.
[0054] In some cases, rather than preferentially forming a self-ligated circularized product, one or more adapter oligonucleotides are used, so that the 5' and 3' ends in the sample are joined by one or more intervening adapter oligonucleotides to form a circular polynucleotide. For example, the 5' end of a polynucleotide can be joined to the 3' end of an adapter, and the 5' end of the same adapter can be joined to the 3' end of the same polynucleotide. Adapter oligonucleotides include any oligonucleotides whose sequence is at least a portion of which is known and can be joined to a sample polynucleotide. Adapter oligonucleotides can include DNA, RNA, nucleotide analogs, non-standard nucleotides, labeled nucleotides, modified nucleotides, or combinations thereof. Adapter oligonucleotides can be single-stranded, double-stranded, or partially duplexed. Generally, partially duplexed adapters contain one or more single-stranded regions and one or more double-stranded regions. A double-stranded adapter can comprise two separate oligonucleotides (also referred to as "oligonucleotide duplexes") hybridized to each other; the hybridization may leave one or more blunt ends, one or more 3' overhangs, one or more 5' overhangs, one or more bulges resulting from mismatched and / or unpaired nucleotides, or any combination thereof. When two hybridized regions of an adapter are separated from each other by an unhybridized region, a "bubble" structure results. Different types of adapters can be used in combination, such as adapters of different sequences. Different adapters can be attached to sample polynucleotides in sequential reactions or simultaneously. In some embodiments, the same adapter is added to both ends of a target polynucleotide. For example, a first adapter and a second adapter can be added in the same reaction. The adapters can be manipulated before combining with the same sample polynucleotide. For example, terminal phosphates can be added or removed.
[0055] When adapter oligonucleotides are used, the adapter oligonucleotides may contain one or more of a variety of sequence elements, including, but not limited to, one or more amplification primer annealing sequences or their complementary sequences, one or more sequencing primer annealing sequences or their complementary sequences, one or more barcode sequences, one or more common sequences shared among a plurality of different adapters or a subset of different adapters, one or more restriction enzyme recognition sites, one or more overhangs complementary to one or more target polynucleotide overhangs, one or more probe binding sites (e.g., for attachment to a sequencing platform such as a flow cell for massively parallel sequencing, such as a flow cell developed by Illumina, Inc.), one or more random or near-random sequences (e.g., one or more nucleotides randomly selected from a series of two or more different nucleotides at one or more positions, wherein each of the different nucleotides selected at the one or more positions is represented by a pool of adapters containing the random sequence), and combinations thereof. In some cases, the adapters may be used to purify the circles containing the adapters, for example, by using beads (especially magnetic beads for ease of handling) coated with oligonucleotides containing sequences complementary to the adapters and capable of "capturing" closed circles with the correct adapters through hybridization, washing away those circles that do not contain adapters and any unligated elements, and then releasing the captured circles from the beads. Furthermore, in some cases, the hybridized capture probe complex and target circle can be used directly to produce concatemers, such as by direct rolling circle amplification (RCA). In some embodiments, the adapters within the circle can also be used as sequencing primers. Two or more sequence elements may be non-adjacent (e.g., separated by one or more nucleotides), adjacent to each other, partially overlapping, or completely overlapping.For example, the annealing sequence of an amplification primer can also function as the annealing sequence of a sequencing primer. The sequence element can be located at or near the 3' end, at or near the 5' end, or internal to the adapter oligonucleotide. The sequence element can be any suitable length, such as about 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or less nucleotides in length. The adapter oligonucleotides can have any suitable length, at least sufficient to accommodate the one or more sequence elements they contain. In some embodiments, the adapters are about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, 200, or less nucleotides in length. In some embodiments, the adapter oligonucleotides are in the range of about 12 to 40 nucleotides in length, such as about 15 to 35 nucleotides in length.
[0056] In some embodiments, adapter oligonucleotides attached to fragmented polynucleotides from a single sample contain one or more sequences common to all adapter oligonucleotides and a barcode unique to the adapter attached to the polynucleotides of a particular sample, and the barcode sequence can be used to distinguish polynucleotides originating from one sample or adapter ligation reaction from polynucleotides originating from another sample or adapter ligation reaction. In some embodiments, the adapter oligonucleotides contain 5' overhangs, 3' overhangs, or both that are complementary to one or more target polynucleotide overhangs. The complementary overhangs can be one or more nucleotides in length, including, but not limited to, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more nucleotides in length. The complementary overhangs may comprise a fixed sequence. The complementary overhang of the adaptor oligonucleotide may comprise a random sequence of one or more nucleotides, where one or more nucleotides are randomly selected from a set of two or more different nucleotides at one or more positions, such that each of the different nucleotides selected at one or more positions is represented in the pool of adaptors having a complementary overhang comprising the random sequence. In some embodiments, the adaptor overhang is complementary to a target polynucleotide overhang produced by restriction endonuclease digestion. In some embodiments, the adaptor overhang consists of adenine or thymine.
[0057] Various methods are available for circularizing polynucleotides. In some embodiments, circularization involves an enzymatic reaction, such as the use of a ligase (e.g., RNA ligase or DNA ligase). Various ligases are available, including, but not limited to, Circligase™ (Epicentre; Madison, WI), RNA ligase, and T4 RNA ligase 1 (an ssRNA ligase that acts on both DNA and RNA). Furthermore, in the absence of a dsDNA template, T4 DNA ligase can also ligate to ssDNA, although this is generally a slower reaction. Other non-limiting examples of ligases include NAD-dependent ligases, including Taq DNA ligase, Thermus filiformis DNA ligase, E. coli DNA ligase, Tth DNA ligase, Thermus scotoductus DNA ligase (I and II), thermostable ligase, Ampligase thermostable DNA ligase, VanC-type ligase, 9°N DNA ligase, Tsp DNA ligase, and novel ligases discovered by biospectrochemical engineering; ATP-dependent ligases, including T4 RNA ligase, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Pfu DNA ligase, DNA ligase 1, DNA ligase III, DNA ligase IV, and novel ligases discovered by biospectrochemical engineering, as well as wild-type, mutant isoforms, and engineered variants thereof. If self-ligation is desired, the concentrations of polynucleotides and enzymes can be adjusted to promote intramolecular circle formation rather than intermolecular structures. The reaction temperature and reaction time can also be adjusted. In some embodiments, 60°C is used to promote intramolecular circularization. In some embodiments, the reaction time is 12 to 16 hours. Reaction conditions may be those specified by the manufacturer of the selected enzyme. In some embodiments, an exonuclease step can be included to digest any unlinked nucleic acids after the circularization reaction. That is, closed circles do not contain free 5' or 3' ends, so the introduction of a 5' or 3' exonuclease will not digest the closed circle but will digest the unlinked elements.This may find particular use in multiplex systems.
[0058] Generally, the ends of a polynucleotide are joined together to form a circular polynucleotide (either directly or using one or more intermediate adaptor oligonucleotides), producing a junction having a junction sequence. When the 5' and 3' ends of a polynucleotide are joined via an adaptor polynucleotide, the term "junction" refers to the junction between the polynucleotide and the adaptor (e.g., one of the 5'-end junction or the 3'-end junction), or the junction between the 5' and 3' ends of the polynucleotide formed by and including the adaptor oligonucleotide. When the 5' and 3' ends of a polynucleotide are joined without an intervening adaptor (e.g., the 5' and 3' ends of a single-stranded DNA), the term "junction" refers to the point where these two ends are joined. A junction may be identified by the sequence of nucleotides comprising the junction (also referred to as the "junction sequence"). In some embodiments, a sample contains polynucleotides with a mixture of ends formed by natural degradative processes (such as cell lysis, cell death, and other processes in which DNA is released from cells into its surrounding environment (e.g., into cell-free polynucleotides such as cell-free DNA and cell-free RNA) where it can be further degraded), fragmentation that is a by-product of sample processing (such as fixation, staining, and / or storage procedures), and fragmentation by methods that cleave DNA without being restricted to a specific target sequence (e.g., mechanical fragmentation such as by sonication, treatment with non-sequence-specific nucleases such as DNase I or fragmentase). When a sample contains polynucleotides with a mixture of ends, it is unlikely that two polynucleotides will have identical 5' or 3' ends, and even less likely that two polynucleotides will independently have both identical 5' and 3' ends. Thus, in some embodiments, junctions may be used to distinguish between different polynucleotides even when the two polynucleotides contain portions with identical target sequences. When the ends of the polynucleotides are joined without intervening adaptors, the sequence of the junction can be identified by alignment with a reference sequence.For example, if the sequence order of two elements appears to be reversed relative to the reference sequence, the point where this reversal appears to occur may indicate the presence of a junction at that point. When polynucleotide ends are joined via one or more adapter sequences, the junction may be identified by proximity to a known adapter sequence, or by alignment as described above when the sequencing read is long enough to obtain sequence from both the 5' and 3' ends of the circularized polynucleotide. In some embodiments, the formation of a particular junction is a sufficiently rare event that it is unique among the circularized polynucleotides of a sample.
[0059] Sequencing methods According to some embodiments of the method for detecting tumor nucleic acid described herein, linear and / or circularized polynucleotides (or their amplification products, which may optionally be enriched) are subjected to a sequencing reaction to generate sequencing reads.The sequencing depth is selected based on the requirements of the sample to be sequenced.In some cases, sequencing is low-depth or fewer reads or reads per molecule, which are used interchangeably herein.In some cases, sequencing is high-depth or more reads or reads per molecule, which are used interchangeably herein.The sequencing reads generated in this way can be used according to other methods disclosed herein.A variety of sequencing methodologies are available, particularly high-throughput sequencing methodologies. Examples include, but are not limited to, sequencing systems manufactured by Illumina (sequencing systems such as HiSeq® and MiSeq®), Life Technologies (Ion Torrent®, SOLiD®, etc.), Roche's 454 Life Sciences systems, Pacific Biosciences systems, Oxford Nanopore Technologies, nanoball sequencing, sequencing by hybridization, polymerized colony (POLONY) sequencing, nanogrid rolling circle sequencing (ROLONY), and the like. In some embodiments, sequencing involves the use of HiSeq® and MiSeq® systems to generate reads that are about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 250, about 300 or more nucleotides in length, or more than about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 250, about 300 or more nucleotides in length. In some embodiments, sequencing involves a sequencing by synthesis process, where individual nucleotides are repeatedly identified as they are added to a growing primer extension product.Pyrosequencing is an example of a sequence by synthesis process in which the incorporation of a nucleotide is identified by assaying the resulting synthesis mixture for the presence of a by-product of the sequencing reaction, i.e., pyrophosphate. Specifically, the primer / template / polymerase complex is contacted with a single nucleotide. If this nucleotide is incorporated, the polymerization reaction cleaves the nucleoside triphosphate between the α- and β-phosphates of the triphosphate chain, releasing pyrophosphate. The presence of released pyrophosphate is then identified using a chemiluminescent enzyme reporter system that uses AMP to convert pyrophosphate to ATP, followed by measurement of ATP using a luciferase enzyme, which produces a measurable light label. If light is detected, the base has been incorporated; if light is not detected, the base has not been incorporated. After appropriate washing steps, various bases are periodically contacted with the complex, resulting in the identification of the base following the template sequence. See, for example, U.S. Patent No. 6,210,891.
[0060] In a related sequencing process, the primer / template / polymerase complex is immobilized on a substrate and contacted with a labeled nucleotide. The immobilization of the complex may be due to the primer sequence, the template sequence, and / or the polymerase enzyme, and may be due to a covalent or non-covalent bond. For example, the immobilization of the complex may be due to the bond between the polymerase or the primer and the substrate surface. In an alternative configuration, the nucleotide is provided with and without a removable terminator group. Upon introduction, the label is detectable because it is bound to the complex. In the case of nucleotides with terminators, all four different nucleotides with individually identifiable labels are contacted with the complex. The introduction of the labeled nucleotide terminates elongation due to the presence of the terminator, adding a label to the complex, thereby enabling the identification of the introduced nucleotide. The label and terminator are then removed from the introduced nucleotide, and the process is repeated after appropriate washing steps. In the case of non-transcriptionally terminated nucleotides, a single type of labeled nucleotide is added to the complex to determine whether it is introduced, similar to pyrosequencing. After removing the labeling group on the nucleotide and appropriate washing steps, various different nucleotides are circulated through the reaction mixture in the same process. For example, see U.S. Patent No. 6,833,246, the entire contents of which are incorporated herein by reference for all purposes. For example, the Illumina Genome Analyzer System is based on the technology described in International Publication No. 98 / 44151, in which DNA molecules are bound to a sequencing platform (flow cell) via anchor probe binding sites (alternatively referred to as flow cell binding sites) and amplified in situ on a glass slide. The solid surface on which DNA molecules are amplified typically contains a plurality of first and second binding oligonucleotides, where the first binding oligonucleotide is complementary to the sequence near or at one end of the target polynucleotide, and the second binding oligonucleotide is complementary to the sequence at or at the other end of the target polynucleotide.This configuration allows for bridge amplification, such as that described in U.S. Patent Application Publication No. 20140121116. The DNA molecules are then annealed to a sequencing primer and sequenced in parallel, base by base, using a reversible terminator approach. Hybridization of a sequencing primer can occur after cleaving one strand of the double-stranded bridge polynucleotide at a cleavage site in one of the binding oligonucleotides that anchor the bridge, leaving one single strand unbound to the solid substrate, which can be removed by denaturation, and the other strand bound and available for hybridization to the sequencing primer. Typically, the Illumina Genome Analyzer System utilizes an eight-channel flow cell, generates sequencing reads of 18-36 bases long, and generates over 1.3 Gbp of high-quality data per run (see www.illumina.com).
[0061] In another further arrangement of the synthesis process, the incorporation of differentially labeled nucleotides is observed in real time as template-dependent synthesis is carried out. Individual immobilized primer / template / polymerase complexes may be observed as fluorescently labeled nucleotides are introduced, allowing for real-time identification of each added base as it is added. In this process, a labeling group may be attached to a portion of the nucleotide that is cleaved during incorporation. For example, by attaching a labeling group to a portion of the phosphate chain that is removed during incorporation, i.e., the β, γ, or other terminal phosphate group on the nucleoside polyphosphate, the label is not incorporated into the nascent strand, and instead, native DNA is produced. Observation of individual molecules may involve optical trapping of the complex within a very small amount of illumination. Optical trapping of the complex can create a monitoring region where nucleotides may be present randomly diffusing for a very short period of time, but the introduced nucleotide can be retained within the observation volume for a longer period of time as it is being introduced. This can produce a distinctive signal associated with the incorporation event, characterized by a signal profile that is characteristic of the base being added. The polymerase or other portion of the complex and the introduced nucleotide may be provided with interactive labeling elements, such as a fluorescent resonant energy transfer (FRET) dye pair, such that the introduction event brings the labeling elements into interactive proximity, and again the characteristic signal is also characteristic of the base being introduced (see, e.g., U.S. Pat. Nos. 6,917,726, 7,033,764, 7,052,847, 7,056,676, 7,170,050, 7,361,466, and 7,416,844, and U.S. Patent Application Publication No. 20070134128, each of which is incorporated by reference in its entirety).
[0062] In some embodiments, the nucleic acid in sample can be sequenced by ligation.This method typically uses DNA ligase enzyme to identify target sequence, as for example used in Polony method and SOLiD technology (Applied Biosystems, now Invitrogen).Generally, provide a pool of all possible fixed length oligonucleotides, and they are labeled according to the position of sequencing.Oligonucleotides are annealed and ligated, and by the preferred ligation of DNA ligase to match the sequence, the signal corresponding to the complementary sequence at that position is generated.
[0063] The sequencing method of the present disclosure can provide useful information for various applications, such as identifying a subject's disease (e.g., cancer) or determining whether a subject is at risk of having (or developing) a disease. Sequencing can provide the sequence of a polymorphic region. Sequencing can also provide the length of a polynucleotide such as DNA (e.g., cfDNA). Furthermore, sequencing can provide the sequence of the breakpoint or end of DNA such as cfDNA. Sequencing can also provide the sequence of the boundary of a protein binding site or the boundary of a DNase-sensitive site.
[0064] sample In some embodiments of the various methods described herein, the sample is from a subject. The subject may be any animal, including, but not limited to, cows, pigs, mice, rats, chickens, cats, dogs, etc., but is typically a mammal such as a human. Sample polynucleotides are often isolated from a cell-free sample from a subject, such as a tissue sample, body fluid sample, or organ sample, including, for example, a blood sample or liquid sample (e.g., saliva) containing nucleic acids. In some cases, the sample is processed to remove cells or isolated without a cell extraction step (e.g., to isolate cell-free polynucleotides, such as cell-free DNA). Other examples of sample sources include blood, urine, feces, nasal cavity, lung, intestine, other body fluids or excretions, materials derived therefrom, or combinations thereof. In some embodiments, the sample is a blood sample or a portion thereof (e.g., plasma or serum). Serum and plasma may be of particular interest due to the relative enrichment of tumor DNA, which is associated with high rates of malignant cell death in such tissues. In some embodiments, a sample from a single individual is divided into multiple separate samples (for example, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more separate samples) that are independently subjected to the method of the present disclosure, such as repeating twice, repeating three times, repeating four or more times.When a sample is from a subject, the consensus sequence from the sample under analysis, or the sequence reference sequence of the polynucleotide from another sample or tissue of the same subject, can also be from the subject.For example, a blood sample can be analyzed for cfDNA mutations, while the cellular DNA from another sample (for example, buccal sample or skin sample) is analyzed to determine the reference sequence.
[0065] Polynucleotides may be extracted from a sample according to any suitable method. Various kits are available for polynucleotide extraction, the choice of which may depend on the type of sample and the type of nucleic acid to be isolated. Examples of extraction methods are provided herein, such as those described with respect to any of the various embodiments disclosed herein. In one example, the sample may be a blood sample, such as a sample collected in an EDTA tube (e.g., a BD Vacutainer). Plasma may be separated from peripheral blood cells by centrifugation (e.g., at 1900 g for 10 minutes at 4°C). Plasma separation performed in this manner on a 6 mL blood sample typically yields 2.5 to 3 mL of plasma. Circular cell-free DNA can be extracted from the plasma sample, for example, by using the QIAmp Circulating Nucleic Acid Kit (Qiagene) according to the manufacturer's protocol. DNA may then be quantified (e.g., on an Agilent 2100 Bioanalyzer equipped with a High Sensitivity DNA kit (Agilent)). As an example, the yield of blood DNA from such plasma samples from healthy humans ranges from 1 ng to 10 ng per mL of plasma, and is significantly higher in samples from patients with disease (e.g., cancer).
[0066] In some embodiments, the plurality of polynucleotides comprises cell-free polynucleotides, such as cell-free DNA (cfDNA), cell-free RNA (cfRNA), circular tumor DNA (ctDNA), or circular tumor RNA (ctRNA). Cell-free DNA circulates in both healthy and diseased individuals. Cell-free RNA circulates in both healthy and diseased individuals. cfDNA (ctDNA) from tumors is not limited to any particular cancer type, but it appears to be a common finding across a variety of malignancies. Some measurements suggest that plasma concentrations of free circular DNA range from approximately 14-18 ng / mL in control subjects to approximately 180-318 ng / mL in tumor-bearing patients. Apoptotic and necrotic cell death contribute to cell-free circular DNA in bodily fluids. For example, significantly elevated blood DNA levels have been observed in the plasma of patients with prostate cancer and other prostate diseases, such as benign prostatic hyperplasia and prostatitis. Furthermore, circular tumor DNA is present in fluids originating from the organ where the primary tumor arises. Furthermore, circulating tumor DNA is present in fluids originating from the organ where the primary tumor arises. Thus, breast cancer detection is achieved using breast ductal washings, colorectal cancer detection using stool, lung cancer detection using sputum, and prostate cancer detection using urine or semen. Cell-free DNA can be obtained from a variety of sources. One common source is a subject's blood sample. However, cfDNA or other fragmented DNA can be derived from a variety of other sources. For example, urine and stool samples can be sources of cfDNA, including ctDNA. Cell-free RNA can be obtained from a variety of sources.
[0067] In some embodiments, polynucleotides are subjected to subsequent steps (e.g., circularization and amplification) without an extraction step and / or a purification step. For example, a liquid sample may be treated to remove cells without an extraction step to produce a purified liquid sample and a cell sample, followed by isolation of DNA from the purified liquid sample. Various procedures are available for isolating polynucleotides, such as by precipitation or nonspecific binding to a substrate, followed by washing the substrate to release bound polynucleotides. When polynucleotides are isolated from a sample without a cell extraction step, the polynucleotides are predominantly extracellular or "cell-free" polynucleotides. For example, cell-free polynucleotides may include cell-free DNA (also referred to as "circular" DNA). In some embodiments, the circular DNA is circular tumor DNA (ctDNA) from tumor cells, such as from bodily fluids or excretions (e.g., blood samples). The cell-free polynucleotides may include cell-free RNA (also referred to as "circular" RNA). In some embodiments, the circular RNA is circular tumor RNA (ctRNA) from tumor cells. Tumors may undergo apoptosis or necrosis, such that tumor nucleic acids are released into the body, including the subject's bloodstream, in different forms and at different levels through various mechanisms. Typically, ctDNA can range in size from highly concentrated smaller fragments, generally 70-200 nucleotides in length, to less concentrated larger fragments up to several thousand kilobases in length.
[0068] cancer In some cases, the methods for detecting tumor nucleic acids provided herein include staging the cancer. The staging of cancer varies depending on the cancer type, with each cancer type having its own classification system. Examples of cancer staging or classification systems are described in more detail below.
[0069] [Table 1]
[0070] [Table 2]
[0071] Table 3-1
[0072] Table 3-2
[0073] Table 4
[0074] Table 5
[0075] Table 6
[0076] Table 7
[0077] Table 8
[0078] Table 9
[0079] Table 10
[0080] Table 11
[0081] Table 12
[0082] Table 13
[0083] Table 14
[0084] Table 15
[0085] Table 16
[0086] Table 17
[0087] Table 18-1
[0088] Table 18-2
[0089] Table 19
[0090] Table 20
[0091] Table 21-1
[0092] Table 21-2
[0093] Table 22
[0094] Table 23-1
[0095] Table 23-2
[0096] Table 24
[0097] Table 25-1
[0098] Table 25-2
[0099] Table 25-3
[0100] Table 26
[0101] Table 27
[0102] Methods for detecting tumor nucleic acids are provided herein. Examples of cancers that can be detected in connection with the methods disclosed herein include, but are not limited to, acanthoma, acinic cell carcinoma, acoustic neuroma, acral lentiginous melanoma, acral hidradenoma, acute eosinophilic leukemia, acute lymphocytic leukemia, acute megakaryoblastic leukemia, acute monocytic leukemia, acute myeloblastic leukemia with maturation, acute myeloid dendritic cell leukemia, acute myeloid leukemia, acute promyelocytic leukemia, adamantinoma, adenocarcinoma, adenoid cystic carcinoma, adenoma, adenoid odontogenic tumor, adrenocortical carcinoma, adult T-cell leukemia, aggressive NK-cell leukemia, AIDS-related cancer ... S-related leukemia, alveolar soft part sarcoma, ameloblastic fibroma, anal cancer, anaplastic large cell lymphoma, anaplastic thyroid cancer, angioimmunoblastic T-cell lymphoma, angiomyolipoma, angiosarcoma, appendix cancer, astrocytoma, atypical teratoid rhabdoid tumor, basal cell carcinoma, basal cell-like carcinoma, B-cell leukemia, B-cell lymphoma, Bellini duct carcinoma, biliary tract cancer, bladder cancer, blastoma, bone cancer, bone tumor, brain stem glioma, brain tumor, breast cancer, Brenner tumor, bronchial tumor, bronchioloalveolar carcinoma, Brown tumor, Burkitt lymphoma, cancer of unknown primary site, carcinoid tumor, cytoplasmic Cyst, carcinoma in situ, penile cancer, tumor of unknown primary tumor, carcinosarcoma, Castleman's disease, central nervous system embryonal tumor, cerebellar astrocytoma, cerebral astrocytoma, cervical cancer, cholangiocarcinoma, chondroma, chondrosarcoma, chordoma, choriocarcinoma, choroid plexus papilloma, chronic lymphocytic leukemia, chronic monocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, chronic neutrophilic leukemia, clear cell tumor, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, Degos disease, dermatofibrosarcoma protuberans, dermoid cyst, desmoplastic small round cell tumor, diffuse large B-cell lymphoma, dysembryoplastic neuroepithelial cell tumor, embryonal carcinoma, yolk sac tumor, endometrial cancer, endometrial uterine cancer, endometrioid tumor, enteropathy-associated T-cell lymphoma, ependymoblastoma, ependymoma, epithelioid sarcoma, erythroleukemia, esophageal cancer, esthesioneuroblastoma, Ewing family tumors, Ewing family sarcoma, Ewing sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, extramammary Paget's disease, fallopian tube cancer, inclusion fetus, fibroma, fibrosarcoma, follicular lymphoma, follicular thyroid cancer, gallbladder cancer, gallbladder cancer, ganglioglioma, ganglioneuroma, gastric cancer, gastric lymphoma, gastrointestinal cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor,Gastrointestinal stromal tumor, germ cell tumor, germ cell tumor, gestational choriocarcinoma, gestational trophoblastic tumor, giant cell tumor of bone, glioblastoma multiforme, glioma, gliomatosis cerebri, glomus tumor, glucagon-producing tumor, gonadoblastoma, granulosa cell tumor, hairy cell leukemia, hairy cell leukemia, head and neck cancer, cardiac cancer, hemangioblastoma, hemangiopericytoma, angiosarcoma, blood tumor, hepatocellular carcinoma, hepatosplenic T-cell lymphoma, hereditary breast and ovarian cancer Cancer syndrome, Hodgkin's lymphoma, Hodgkin's lymphoma, hypopharyngeal cancer, hypothalamic glioma, inflammatory breast cancer, intraocular melanoma, pancreatic islet cell carcinoma, pancreatic islet tumor, juvenile myelomonocytic leukemia, Kaposi's sarcoma, Kaposi's sarcoma, kidney cancer, Krukenberg's tumor, pharyngeal cancer, pharyngeal carcinoma, lentigo maligna melanoma, leukemia, leukemia, lip and oral cancer, liposarcoma, lung cancer, luteinoma, lymphangioma, lymphangiocarcinoma tumor, lymphoepithelioma, lymphocytic leukemia, lymphoma, macroglobulinemia, malignant fibrous histiocytoma, primary malignant fibrous histiocytoma of bone, malignant glioma, malignant mesothelioma, malignant peripheral nerve sheath tumor, malignant rhabdoid tumor, malignant Triton tumor, MALT lymphoma, mantle cell lymphoma, mast cell leukemia, mediastinal germ cell tumor, mediastinal tumor, medullary thyroid carcinoma, medulloblastoma, medulloepithelioma, melanoma, meningioma, Merkel cell carcinoma Cellular carcinoma, mesothelioma, metastatic squamous cell carcinoma of the neck of unknown primary site, metastatic urothelial carcinoma, mixed Müllerian tumor, monocytic leukemia, oral cavity cancer, mucinous tumor, multiple endocrine neoplasia syndrome, multiple myeloma, mycosis fungoides, myelodysplastic disease, myelodysplastic syndrome, myeloid leukemia, myeloid sarcoma, myeloproliferative disease, myxoma, nasal cavity cancer, nasopharyngeal carcinoma, nasopharyngeal carcinoma carcinoma), neoplasms, schwannoma, neuroblastoma, neuroblastoma, neurofibroma, neuroma, nodular melanoma, non-Hodgkin's lymphoma, non-melanoma skin cancer, non-small cell lung cancer, ophthalmic oncology, oligoastrocytoma, oligodendroglioma, oncocytoma, optic nerve sheath meningioma, oral cavity cancer, oral cancer, oropharyngeal cancer, osteosarcoma, osteosarcoma, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, ovarian low malignant potential tumor, Paget's disease of the breast, Pancoast tumor, pancreatic cancer, pancreatic cancer, papillary thyroid cancer, papilloma, paraganglioma, sinus malignancy, parathyroid cancer, penile cancer, perivascular epithelioid cell tumor, pharyngeal cancer, pheochromocytoma, intermediate pineal parenchymal tumor,Pineoblastoma, pituitary cell tumor, pituitary adenoma, pituitary tumor, plasma cell tumor, pleuropulmonary blastoma, polyblastoma, T-lymphoblastic lymphoma, primary central nervous system lymphoma, primary effusion lymphoma, primary hepatocellular carcinoma, primary liver cancer, primary peritoneal cancer, primary neuroectodermal tumor, prostate cancer, pseudomyxoma peritonei, rectal cancer, renal cell carcinoma, respiratory tract cancer related to the NUT gene on chromosome 15, retinoblastoma, rhabdomyoma, rhabdomyosarcoma, Richter transformation, sacrococcygeal teratoma, salivary gland cancer, sarcoma, schwannomatosis, sebaceous gland carcinoma, secondary tumors, seminoma, serous tumor, Sertoli-Leydig cell tumor, sex cord-stromal tumor, Sézary syndrome, signet ring cell carcinoma, skin cancer, small round cell tumor, small cell carcinoma, small cell lung cancer, small cell lymphoma, small intestine cancer, These include soft tissue sarcoma, somatostatinoma, sooty warts, spinal cord tumors, spinal tumors, splenic marginal zone lymphoma, squamous cell carcinoma, gastric cancer, superficial spreading melanoma, primary neuroectodermal tumors, superficial epithelial stromal tumors, synovial sarcoma, T-cell acute lymphoblastic leukemia, large granular lymphocyte leukemia, T-cell leukemia, T-cell lymphoma, T-cell prolymphocytic leukemia, teratoma, terminal lymphoma, testicular cancer, thecoma, pharyngeal cancer, thymic carcinoma, thymoma, thyroid cancer, renal pelvis and ureteral transitional cell carcinoma, transitional cell carcinoma, urachal cancer, urethral cancer, genitourinary tumors, uterine sarcoma, uveal melanoma, vaginal cancer, Werner-Morrison syndrome, verrucous carcinoma, optic nerve glioma, vulvar cancer, Waldenstrom hypergammaglobulinemia, Warthin's tumor, Wilms' tumor, and combinations thereof.
[0103] Computer Systems The present disclosure provides a computer system programmed to implement tumor nucleic acid detection. Figure 3 shows a computer system (301) programmed or otherwise configured to detect tumor nucleic acid. The computer system (301) can coordinate various aspects of the tumor nucleic acid detection method of the present disclosure, such as detecting tumor nucleic acid in cell-free nucleic acid using a low-depth sequencing method. The computer system (301) can be a user's electronic device or a computer system located remotely from the electronic device. The electronic device can be a mobile electronic device.
[0104] The computer system (301) includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") (305), which may be a single-core or multi-core processor for parallel processing, or multiple processors. The computer system (301) also includes memory or memory locations (310) (e.g., random access memory, read-only memory, flash memory), electronic storage (315) (e.g., hard disk), a communication interface (320) (e.g., network adapter) for communicating with one or more other systems, and peripheral devices (325), such as cache, other memory, data storage, and / or electronic display adapters. The memory (310), storage (315), interface (320), and peripheral devices (325) communicate with the CPU (305) via a communication bus (solid lines), such as a motherboard. The storage (315) may be a data storage device (or data repository) for storing data. The computer system 301 may be operatively coupled to a computer network ("network") 330 with the aid of the communication interface 320. The network 330 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. In some cases, the network 330 is a telecommunications and / or data network. The network 330 may include one or more computer servers, which may enable distributed computing, such as cloud computing. In some cases, the network 330 may implement a peer-to-peer network with the aid of the computer system 301, allowing devices coupled to the computer system 301 to function as clients or servers.
[0105] The CPU (305) can execute a series of machine-readable instructions, which may be implemented in a program or software. The instructions may be stored in a memory location, such as the memory (310). The instructions may be directed to the CPU (305), which may be later programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU (305) may include fetch, decode, execute, and writeback.
[0106] The CPU 305 may be part of a circuit, such as an integrated circuit. One or more other components of the system 301 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0107] The storage device (315) can store files such as drivers, libraries, and saved programs. The storage device (315) can store user data such as user preferences and user programs. In some cases, the computer system (301) can include one or more additional data storage devices external to the computer system (301), such as located on a remote server that communicates with the computer system (301) over an intranet or the Internet.
[0108] The computer system (301) can communicate with one or more remote computer systems via a network (330). For example, the computer system (301) can communicate with a remote computer system of a user (e.g., a person desiring to detect tumor nucleic acids). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate PC or tablet PC remote computer system (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. The user can access the computer system (301) via the network (330).
[0109] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on electronic storage locations of the computer system (301), such as on memory (310) or electronic storage (315). The machine-executable or machine-readable code can be provided in software form. During use, the code can be executed by the processor (305). In some cases, the code can be retrieved from storage (315) and stored on memory (310) for ready access by the processor (305). In some situations, the electronic storage (315) can be eliminated, and the machine-executable instructions are stored on memory (310).
[0110] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled at run time. The code can be provided in a programming language that can be selected to execute the code in a pre-compiled or compiled manner.
[0111] Aspects of the systems and methods provided herein, such as the computer system (301), may be implemented with programming. Various aspects of the technology can be considered "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data executed or embodied on some type of machine-readable medium. The machine-executable code can be stored in electronic storage devices, such as memory (e.g., read-only memory, random-access memory, flash memory) or hard disks. "Storage"-type media can include any or all of various semiconductor memories, tape drives, disk drives, and the like, as well as tangible memory of a computer, processor, or their associated modules, which may provide non-transitory storage for software programming at any time. All or portions of the software may subsequently be communicated via the Internet or various other telecommunications networks. Such communication may allow, for example, loading of the software from one computer or processor to another, such as from a management server or host computer to an application server computer platform. Thus, other types of media that software elements may carry include light waves, radio waves, and electromagnetic waves used through physical interfaces between local devices by wired and optical cable networks and across various air links. The physical elements that carry such waves, such as wired or wireless connections, optical links, etc., are considered to be media bearing the software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0112] Thus, a machine-readable medium such as a computer-executable code may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices of any computer, or those used to implement the databases shown in the figures, etc. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, e.g., copper wire and fiber optics, which comprise the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer readable media therefore include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper cards, any other physical storage media with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves transmitting data or instructions, cables or links which transmit such transmissions, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0113] The computer system (301) may include or be in communication with an electronic display (335) that includes a user interface (UI) (340) for displaying, for example, tumor nucleic acid detection. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0114] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by a central processing unit (305). The algorithms may, for example, detect tumor nucleic acids.
[0115] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided herein. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it is to be understood that all aspects of the invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be utilized in practicing the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. Therefore, it is contemplated that the present invention also encompasses any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and equivalents thereof be embraced therein. [Example]
[0116] The following examples are given for the purpose of illustrating various embodiments of the present invention and are not meant to limit the present invention in any manner. The examples, together with the methods described herein, represent presently preferred embodiments and are exemplary and are not intended to limit the scope of the present invention. Those skilled in the art will envision modifications thereof and other uses that are encompassed within the spirit of the present invention, as defined by the scope of the claims.
[0117] Example 1: Detection of tumor nucleic acids in acellular samples A tumor-informed disease detection and monitoring method is used to detect tumor nucleic acid in a sample. This method begins by performing whole-genome sequencing (e.g., more than 10 gigabases) of DNA obtained from tumor tissue to identify a list of tumor-specific somatic variants in the subject. Tumor-specific variants can be identified by comparing the sequence of tumor DNA with sequences from healthy tissue DNA or plasma DNA after surgery. Simultaneously or subsequently, cell-free DNA is obtained from a sample from the subject. The cell-free DNA is circularized and amplified using rolling circle amplification to obtain concatemeric copies of cell-free DNA containing two or more copies of the sequence of the cell-free DNA molecule. Whole-genome sequencing is performed on the concatemers at a low sequencing depth, e.g., less than two reads per molecule, to determine whether the sample is positive or negative for tumor nucleic acid based on the presence or absence of tumor-specific variants. Determining whether the sample is positive or negative for tumor-specific variants uses error correction, where a variant is determined only when multiple occurrences of the variant are observed in the concatemer. This method is illustrated in Figures 1, 2 and 5A-5B.
[0118] Example 2: Tumor nucleic acid detection by genome complexity reduction A tumor-informed disease detection and monitoring method is used to detect tumor nucleic acids in a sample. This method begins by reducing genomic complexity from approximately 3 billion positions to approximately 30 million positions by using positive selection against target sequences in tumor and normal nucleic acid samples (Figure 4). Whole-genome sequencing is then performed to identify a list of tumor-specific somatic variants in the subject. Cell-free DNA is also obtained from the sample from the subject. The cell-free DNA is circularized and amplified using rolling circle amplification to obtain concatemeric copies of cell-free DNA containing two or more copies of the sequence of the cell-free DNA molecule. Whole-genome sequencing is performed on the amplified cell-free DNA at a low sequencing depth to determine whether the sample is positive or negative for tumor nucleic acids based on the presence or absence of tumor-specific variants. The determination of whether a sample is positive or negative for tumor-specific variants uses error correction, where a variant is identified only if multiple occurrences of the variant are observed in the concatemer.
[0119] Example 3: Negative Selection Workflow A nucleic acid sample, such as a cell-free DNA sample, containing a sequence of interest and unwanted sequences is subjected to negative selection. The nucleic acids in the sample are subjected to a denaturation step, and then blocking oligonucleotides with modified 5' and / or 3' ends are annealed to the unwanted sequences in the nucleic acid sample. The blocking oligonucleotides cannot be ligated or extended with a polymerase. Therefore, only single-stranded nucleic acids in the sample are subjected to circularization. The remaining linear nucleic acids bound to the blocking oligonucleotides are then optionally degraded with a DNA exonuclease. The circularized nucleic acids are amplified by rolling circle amplification to produce concatemers. The concatemers are sequenced to obtain the nucleic acid sequences, and variants are detected based on their presence in one or more copies of the sequence in the concatemers. This method is shown in Figure 7. A workflow without selection is shown in Figure 6.
[0120] Example 4: Positive Selection Workflow A nucleic acid sample, such as a cell-free DNA sample, containing a sequence of interest and an unwanted sequence is subjected to positive selection. The nucleic acid in the sample is subjected to a denaturation step that causes circularization. The circularized nucleic acid is amplified by rolling circle amplification to generate concatemers using a combination of random primers and target-specific primers, resulting in improved amplification of the sequence of interest. The concatemers are sequenced to obtain the sequence of the nucleic acid, and variants are detected based on their presence in one or more copies of the sequence in the concatemer. This method is shown in Figure 8. A workflow without selection is shown in Figure 6.
[0121] Example 5: Ultrasensitive circular tumor DNA detection by whole genome sequencing with single-read error correction Whole-genome sequencing (WGS) of cell-free DNA (cfDNA) is useful for detecting circular tumor DNA (ctDNA) and assessing tumor burden. While the breadth of WGS can compensate for the lack of cfDNA, its sensitivity is limited by the error rate. Herein, we report a genome-wide error suppression at the single-read level of 4.2 × 10 -7 AccuScan, an efficient cfDNA WGS technique, is described that can achieve an error rate of 10, which is more than an order of magnitude lower than read-centric denoising methods. When applied to molecular residual disease (MRD) detection, this method achieves a 10-fold error rate with 99% sample-level specificity. -6 We have demonstrated analytical sensitivity down to the circular tumor allele fraction (cTAF). In colorectal cancer, AccuScan demonstrated a landmark sensitivity of 90% for predicting recurrence. AccuScan also demonstrated robust MRD performance in esophageal cancer and high prognostic value for immunotherapy monitoring in melanoma patients using samples collected as early as 1 week after surgery. Overall, AccuScan provides a highly accurate WGS solution that enables ctDNA detection in the ppm range without deep sequencing or personalized reagents.
[0122] Genome-wide error suppression enables detection of ultra-low ctDNA levels.
[0123] The AccuScan assay workflow (Figure 9A) is optimized for low-input cfDNA, efficiently capturing double-stranded, single-stranded, and nicked DNA in samples. After denaturing and circularizing cfDNA by ligation, whole-genome amplification is performed using RCA to generate concatemeric molecules containing multiple tandem copies of the original template. These concatemeric products are sequenced using a PE150 read length and aligned to the human reference genome. The sequences of each copy within a read pair are compared. Variations from the reference that are consistent across all copies are putative variants, while discrepant variations are likely PCR or sequencing errors, which are removed. To evaluate the efficiency of error correction by AccuScan, identical cfDNA samples from healthy donors (N=3) were sequenced using both regular WGS and AccuScan WGS. Figure 9B shows a comparison of the measured error rates. The observed overall error rate was approximately 4.2 x 10 for AccuScan. -7 and 3.3 × 10 using regular WGS with qscore filter. -4 This suggests that the error correction of the concatemers reduced noise by approximately 1000 times.
[0124] Simulations were performed to predict the effect of error rates on ctDNA detection under different cTAFs, sequencing depths, and marker numbers (Figures 10A-10D, Figures 14A-14B). A statistical model that calculates the probability of observing an expected base call at a particular marker locus is used to predict the presence of ctDNA. Sensitivity is calculated as the number of predicted positives over the total number of simulations for each combination, with a nominal specificity setting of 99%.
[0125] Figure 10A shows the 1 x 10 -4 or 5×10 7The sensitivity of the simulations using 10,000 markers at either an error rate of 5 x 10 or 5 x 10. Reducing the error rate or increasing the sequencing depth both improves the detection rate. -7 and at a sequencing depth of 10x error rate, 4.5 × 10 -5 Although the detection rate (LOD95) of cTAF is 95%, the error rate is 1×10 -4 For example, 100x more sequencing is required to achieve a similar LOD95 for the same cTAF. At 60x sequencing depth, the LOD95 is 1x10 -5 Figure 10B shows that when cTAF is set to 0, the false positive rate remains below 1.2% under all conditions, which is consistent with the nominal specificity setting.
[0126] The analytical sensitivity of AccuScan was measured using a mixture of healthy samples (Figure 10C). cfDNA from three different healthy "test" donors was collected at 1 x 10 -4 ~1×10 -6 The 21 cfDNA mixtures were individually titrated into cfDNA from different healthy "background" donors at seven different concentrations ranging from 0.1 to 0.2. These 21 cfDNA mixtures were then sequenced 60x using AccuScan with 10 ng of input DNA per reaction to assess the ability to detect SNPs in the "test" donors from the background. Of the more than 100,000 SNVs in each test and background sample pair, randomly selected subsets of 5,000, 10,000, or 2,000 SNVs (SNVs selected to have a variant profile similar to that of CRC tumors) were tested. This was repeated 1,000 times per condition, and MRD testing was performed with a nominal specificity of 99% (Table 28). The observed specificity was less than 99% for the 5,000, 10,000, or 20,000 marker conditions. -5 The observed sensitivity at cTAF levels of 1 × 10 and above is greater than 99% for all conditions tested. -5At 10 parts per million, corresponding to a cTAF of 1, the average detection rate for 5,000 markers was 77%, testing of 10,000 markers showed an average sensitivity of 96%, and testing of 20,000 markers maintained 100% sensitivity across all replicates.
[0127] [Table 28]
[0128] The analytical sensitivity of AccuScan was further confirmed by mixing cfDNA from melanoma patients with cfDNA from healthy donors. The original cancer cfDNA sample had 1.1% cTAFs, as measured by ddPCR of the BRAF V600E mutation found in the primary tumor. 1 x 10 -3 ~2×10 -6 Dilute the 5 different expected frequencies of 1 x 10 and perform ddPCR. -3 The diluted BRAFV600E VAF was confirmed. Diluted cancer samples were sequenced by AccuScan with an input of 10 ng per reaction. The observed detection rate was 1 x 10 -3 , 1×10 -4 , and 1 × 10 -5 100%, 5 × 10 for samples with cTAF -6 So 67% (2 / 3), 2 x 10 -6 In the case of CFDNA from a healthy donor, the percentage was 33% (1 / 3) (Figure 10D). AccuScan sequencing of the negative control (cfDNA from a healthy donor) was negative in both replicates.
[0129] To measure AccuScan's sample-level specificity, we used tumor-specific variants from 60 different cancer patients, including CRC, ESCC, and melanoma, and randomly sampled 5K, 10K, and 20K equivalent variants to test MRD calls in mismatched patient plasma samples. For each combination of variant count level and plasma sample, 2000 random samples of mismatched variants were performed. Mean sample-level specificity was calculated as the fraction of MRD tests characterized by a negative MRD call. The observed values were similar to the nominal specificity: 99.3%, 99.1%, and 98.9% for the 5K, 10K, and 20K variant count levels, respectively. These results suggest that the AccuScan assay and analysis perform as intended with patient plasma using tumor variants.
[0130] Identifying tumor-specific variants using a white blood cell (WBC)-free workflow
[0131] Tumor-informative MRD testing uses tumor-specific variants as markers to track disease. Sequencing of tumor tissue uncovers not only cancer mutations but also germline SNPs and other types of variants, such as clonal hematopoietic with undetermined potential (CHIP) variants, that interfere with MRD analysis. One common strategy for filtering non-cancer mutations is to remove variants found in matched WBCs from the same patient. However, this method requires additional sample processing and sequencing (Figure 11A). To simplify the MRD workflow, we investigated the effect of skipping WBC sequencing and using information from post-treatment plasma samples to remove germline and CHIP variants (Figure 11B).
[0132] When low-tumor-burden plasma samples are sequenced at 40x, germline and CHIP mutations can be found at the two- or more-molecule level, whereas tumor-specific variants will only be found at the single-molecule level. Therefore, variants found at two or more molecules in post-treatment plasma can be removed from tumor tissue sequencing results, yielding a list of tumor-specific mutations. To test the feasibility of this approach, the performance of tumor-WBC and tumor-plasma pairs was compared using matched tumor tissue, WBC, and plasma samples from 20 cancer patient samples. The number of tumor-specific variants and variant type profiles found by the two different workflows are shown in Figures 11C-11D. Overall, the number of mutations identified by both methods and the mutation profiles of the identified variants are similar. AccuScan MRD analysis of plasma samples (n = 48) from these 20 patients returned identical MRD calls under either workflow (Figure 11F), and cTAF values are strongly correlated (R 2 =0.99, Figure 11E). These results suggest that it is feasible to use post-treatment plasma instead of WBCs for the identification of tumor-specific variants.
[0133] MRD detection and prognostic value in surgical patients
[0134] We then evaluated the performance of AccuScan for detecting MRD in postoperative gastrointestinal cancer, including 32 patients with colorectal cancer (CRC) and 17 patients with esophageal squamous cell carcinoma (ESCC).
[0135] The ESCC cohort included patients with stage I-III disease (18% stage I, 53% stage II, 29% stage III) who underwent curative-intent surgery. Formalin-fixed, paraffin-embedded (FFPE), WBC, pre-operative plasma, and early post-operative (1 week) plasma samples were collected from all patients. Using tumor and WBC samples, we identified a median of 6768 tumor-specific variants per patient (Figure 12A).
[0136] ctDNA was detected in all 17 preoperative samples, with a median cTAF of 0.27% [interquartile range (IQR): 0.13%–0.55%]. There was no significant trend toward higher cTAF in later-stage patients (Figure 12B). In postoperative plasma samples, ctDNA was detected in 35.29% (6 / 17) of patients, with a median cTAF of 1.3 × 10 -4 (IQR: 1.9 × 10 -5 ~1.1×10 -2 ). The follow-up period for this ESCC cohort ranged from 4.03 months to 24 months. All patients (6 / 6, 100%) whose postoperative specimens were ctDNA-positive (ctDNA+) experienced disease recurrence within 2 years of surgery, and 5 of 6 patients (83%) experienced disease recurrence within 1 year (Figure 12C). In contrast, of 11 patients whose postoperative specimens were ctDNA-negative, only 3 experienced disease recurrence, and all 8 disease-free patients were followed for 24 months. ctDNA detection 1 week after surgery had a sensitivity of 66.67% (95% CI: 29.93%-92.51%), a specificity of 100% (95% CI: 63.06%-100%), and an accuracy of 82.35% (95% CI: 56.57%-96.27%) in predicting ESCC recurrence.
[0137] CRC patients were at different clinical stages (stage I 22%, stage II 38%, stage III 34%, stage IV 6%) and underwent curative surgery. Formalin-fixed, paraffin-embedded (FFPE) samples were available from all patients. For 15 patients with available WBCs, tumor-specific variants were identified using the tumor-WBC workflow; for all other patients, plasma samples from the first surgery were used for the WBC-free workflow (Figure 11B). A median of 5,820 tumor-specific variants per patient (2,148–265,800, Figure 12A) were found, corresponding to approximately 2 mutations / Mb.
[0138] Of the 32 patients, 26 had plasma samples collected preoperatively and 28 had plasma samples collected at landmark locations (within 1 month of surgery). All preoperative cfDNA samples had a median cTAF count of 5.2 × 10 -4 (IQR: 6.6 × 10 -5 ~2.4×10 -3 ctDNA was detected in 100 patients, with no significant trend toward higher cTAF in later-stage patients (Figure 12B). The median follow-up period in this CRC cohort was 24.13 months (IQR: 18.5–36). 34.4% (11 / 32) of patients had detectable ctDNA in postoperative specimens. All patients who were ctDNA-positive in postoperative specimens relapsed within 3 years after surgery (Figure 12D). The median disease-free survival (DFS) in the ctDNA+ patient group was 10.8 months (IQR: 5.8–12.7), with 63.64% (7 / 11) of ctDNA+ patients relapsed within 1 year and 90.91% (10 / 11) of ctDNA+ patients relapsed within 2 years. One ctDNA+ patient, patient #11, was ctDNA- at the first landmark time point, converted to ctDNA+ 6 months after surgery, and then relapsed at 32 months. Patients who were ctDNA-negative (ctDNA-) at all postoperative time points remained progression-free throughout follow-up (up to 36 months) (Figure 12D). Collectively, these results suggest a 90% (95% CI: 55.5%-99.8%) sensitivity for landmark detection, a 100% sensitivity for long-term monitoring, a 100% (95% CI: 80.5%-100%) specificity, and a 96.3% (95% CI: 81%-99.9%) accuracy for predicting CRC recurrence.
[0139] Looking at early post-surgery samples, MRD+ patients had shorter DFS than MRD- patients in ESCC (hazard ratio, HR, 8.68, 95% CI: 1.63–46.32, log-rank p=0.0001) and CRC (HR, 45.54, 95% CI: 9.78–212, log-rank p<0.0001) (Figure 12E).
[0140] ctDNA monitoring during immunotherapy
[0141] Advances in immune checkpoint blockade (ICB) have significantly improved survival rates for patients with advanced melanoma. However, only a small fraction of patients (<20%) respond to ICB. There is an urgent need for prognostic and monitoring tools for patients receiving immunotherapy. Therefore, we investigated the use of AccuScan to monitor patient response to ICB in a pilot study for advanced melanoma (N=8). A total of 22 plasma samples were collected, including 6 pre-treatment and 16 during treatment. WGS of paired tumor and WBC DNA samples identified a median of 34,323 SNVs, with an average of 90,006 tumor-specific SNVs per patient (Figure 12A). All six pre-treatment samples were ctDNA-positive, with cTAF levels of 3.06 × 10 -6 (Figure 12B). Of 16 samples obtained during treatment, 10 were ctDNA-positive.
[0142] For patients 1-5 and 8, radiographic changes matched the changes in cTAF measured by AccuScan (Figures 13A-13B). Patients 1, 2, and 3 were ctDNA positive before treatment, converted to ctDNA negative after treatment, and maintained a complete response or were free of disease recurrence throughout the monitoring period. Patient 8 had persistently low cTAF (approximately 1 x 10) in samples obtained during treatment. -5 ctDNA was positive (pretreatment samples were not available), and CT scans showed stable lung nodules with no evidence of disease recurrence. Patients 4 and 5 had very high cTAF levels (0.4-16%) in all plasma samples, with cTAF at the second time point approximately double that at the first time point. Both patients experienced tumor progression.
[0143] The correlation between ctDNA dynamics and radiological information is complicated for patients 6 and 7. In patient 6, AccuScan revealed a high cTAF (2 × 10) before treatment. -4ctDNA was detected by CT scan and showed clearance of ctDNA 3 months after surgery and after two cycles of ICI treatment (before cycle 3 of ICI treatment) (Figure 13B), but CT scan detected lymphadenopathy 4 months later (after cycle 6 of ICI treatment). The patient was switched to a TKI and showed a complete response (Figure 13A). Despite the discrepancy between ctDNA and imaging data during treatment, the observed clearance of ctDNA is consistent with the clinical outcome. Either MRD levels did not reflect tumor burden or imaging-based clinical response was delayed, and the patient was responding even before switching to a TKI.
[0144] Patient 7 did not have a pretreatment sample, but the first sample during treatment was ctDNA-negative, followed by four subsequent ctDNA-positive samples with stably low levels (Figure 13B). Three additional CTs obtained during treatment showed continued tumor progression, but the fourth obtained after the final ctDNA test showed an excellent partial response, with the patient reaching near CR one year later (Figure 13A). This is another example of discordance between imaging and ctDNA testing, where ctDNA levels remained stably low while imaging indicated tumor progression. Dynamic changes in ctDNA combined with imaging data may better predict patient outcome than either imaging or isolated ctDNA results alone.
[0145] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments described herein may be utilized. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. 1. A method for detecting tumor nucleic acid in an acellular biological sample from a subject, comprising: (a) circularizing nucleic acid derived from the cell-free biological sample to produce a circularized nucleic acid; (b) amplifying the circularized nucleic acid to produce concatemers comprising at least two copies of the sequence of the circularized nucleic acid; (c) sequencing the concatemers or derivatives thereof to obtain the sequence of the concatemers, wherein the sequencing is to a depth of 18 reads or less; (d) processing the sequences of the concatemers to identify at least two occurrences of a tumor-specific sequence variant of the subject; (e) identifying the nucleic acid as having at least one tumor-specific sequence variant upon identifying the at least two occurrences of the tumor-specific sequence variant in the sequence of the concatemer; A method comprising:
2. 10. The method of claim 1, further comprising obtaining the tumor-specific sequence variant from the subject.
3. 3. The method of claim 2, wherein obtaining the tumor-specific sequence variants comprises sequencing nucleic acid derived from the subject's tumor.
4. 4. The method of claim 3, wherein obtaining the tumor-specific sequence variants further comprises sequencing nucleic acid from healthy tissue of the subject.
5. 5. The method of claim 4, wherein the healthy tissue has low or no tumor content.
6. 6. The method of claim 4 or 5, wherein the healthy tissue comprises post-treatment plasma.
7. The method of any one of claims 1 to 6, wherein the sequencing in (c) is at a depth of 18 reads or less per concatemer.
8. 7. The method of any one of claims 1 to 6, wherein the sequencing in (c) is to a depth of 18 reads or less per circularized nucleic acid.
9. The method of any one of claims 1 to 6, wherein the sequencing in (c) is at a depth of no more than one read per concatemer.
10. 7. The method of any one of claims 1 to 6, wherein the sequencing in (c) is to a depth of no more than one read per circularized nucleic acid.
11. 11. The method of any one of claims 3 to 10, wherein the sequencing of nucleic acids from the tumor or healthy tissue of the subject is to a depth of greater than 20 reads.
12. 12. The method of any one of claims 1 to 11, wherein the sequencing in (c) is to a depth of 10 reads or less.
13. 13. The method of any one of claims 1 to 12, wherein the sequencing in (c) is to a depth of 5 reads or less.
14. 14. The method of any one of claims 1 to 13, wherein the sequencing in (c) is to a depth of 2 reads or less.
15. 12. The method of any one of claims 1 to 11, wherein the sequencing in (c) comprises a sequence of at least 10 gigabases.
16. 16. The method of any one of claims 3 to 15, wherein the nucleic acid derived from the tumor is subjected to selection prior to sequencing.
17. The method of any one of claims 4 to 16, wherein the nucleic acids derived from the healthy tissue are subjected to selection prior to sequencing.
18. 18. The method of claim 16 or 17, wherein the selection comprises a negative selection to remove non-target sequences from the nucleic acid.
19. 19. The method of claim 18, wherein negative selection comprises annealing one or more blocking oligonucleotides to unwanted sequences in the nucleic acid derived from the tumor or the nucleic acid derived from the healthy tissue, and circularizing remaining single-stranded nucleic acid.
20. 20. The method of claim 19, wherein the blocking oligonucleotide has a modified 5' end, a modified 3' end, or modified 5' and 3' ends.
21. 21. The method of claim 19 or 20, further comprising the step of degrading the resulting linear double-stranded nucleic acid using an exonuclease.
22. 18. The method of claim 16 or 17, wherein selecting comprises positive selection to select a target sequence from the nucleic acid.
23. 23. The method of claim 22, wherein positive selection comprises amplifying the nucleic acid from the tumor or the nucleic acid from the healthy tissue using a plurality of random primers and a plurality of target-specific primers.
24. The method of any one of claims 1 to 23, further comprising, prior to (a), subjecting the nucleic acid from the cell-free biological sample to selection.
25. The method of any one of claims 1 to 23, further comprising, prior to (c), subjecting the nucleic acid from the cell-free biological sample to selection.
26. 26. The method of claim 24 or 25, wherein the selection comprises a negative selection to remove non-target sequences from the nucleic acid.
27. 26. The method of claim 24 or 25, wherein selecting comprises positive selection to select a target sequence from the nucleic acid.
28. 28. The method of any one of claims 1 to 27, wherein (a) comprises ligating the ends of the nucleic acid or derivative thereof to each other.
29. The method of any one of claims 1 to 27, wherein (a) comprises attaching an adaptor to the 5' end, the 3' end, or the 5' end and the 3' end of the nucleic acid or derivative thereof.
30. The method of any one of claims 1 to 29, wherein (b) is carried out by a polymerase having strand displacement activity.
31. The method of any one of claims 1 to 30, wherein (b) is carried out by a polymerase having 5' to 3' exonuclease activity.
32. 32. The method of any one of claims 1 to 31, wherein the amplifying is performed by at least one primer from a plurality of random primers.
33. 32. The method of any one of claims 1 to 31, wherein the amplifying is performed with at least one primer from a plurality of primers designed for whole genome amplification.
34. 34. The method of any one of claims 1 to 33, wherein the nucleic acid is single-stranded.
35. 34. The method of any one of claims 1 to 33, wherein the nucleic acid is double-stranded.
36. 36. The method of any one of claims 1 to 35, wherein the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
37. 37. The method of any one of claims 1-36, wherein (c) comprises (i) contacting the concatemer or a derivative thereof with a plurality of nucleotides in the presence of a polymerase to incorporate one or more nucleotides of the plurality of nucleotides into a growing strand complementary to the concatemer or derivative thereof, and (ii) detecting one or more signals indicative of incorporation of the one or more nucleotides into the growing strand.
38. 38. The method of any one of claims 1 to 37, wherein (c) comprises sequencing by ligation.
39. 39. The method of any one of claims 1 to 38, wherein the tumor-specific sequence variant comprises a single nucleotide variant, a fusion, an insertion, a deletion, or an epigenetic modification.
40. 40. The method of any one of claims 1 to 39, wherein the acellular biological sample is a body fluid.
41. 41. The method of claim 40, wherein the bodily fluid comprises urine, saliva, blood, serum, or plasma.
42. 42. The method of any one of claims 1 to 41, wherein the tumor is colorectal cancer, pancreatic cancer, ovarian cancer, breast cancer, prostate cancer, bladder cancer, lung cancer, skin cancer, or blood cancer.
43. 43. The method of any one of claims 1 to 42, further comprising determining the subject as minimal residual disease (MRD) positive if the nucleic acid has at least one of the tumor-specific sequence variants.