Tumor nucleic acid recognition method
Through multi-converged sequencing technology, nucleic acids are cyclized and amplified in cell-free biological samples, combined with error correction and selective recognition, the problems of ctDNA detection height limit and long turnover time in the prior art are solved, and efficient and rapid MRD detection and monitoring are achieved.
Patent Information
- Application Number
- CN202380090901.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-28
- Filing Date
- 2023-11-07
- Publication Date
- 2025-08-15
AI Technical Summary
In the detection of circulating tumor nucleic acids (ctDNA), the prior art has high detection limits, long turnover time and logical challenges, making it difficult to achieve high sensitivity and efficient minimum residual disease (MRD) detection and monitoring.
Multi-conjunction sequencing technology is used to cyclize and amplify nucleic acids in cell-free biological samples, generate multi-conjunctions, and error correction is performed at a single read level. Combined with negative and positive selection, tumor-specific sequence variants are identified to achieve whole-genome sequencing.
It realizes rapid and sensitive detection and monitoring of ctDNA under low detection limits, improves the efficiency and accuracy of MRD detection, and reduces cost and turnover time.
Smart Images

Figure CN120500545A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 382,944, filed on November 9, 2023, and U.S. Provisional Application No. 63 / 492,690, filed on March 28, 2023, each of which is incorporated herein by reference in its entirety. Background Art
[0003] During tumor progression, nucleic acids from tumors are often released into the bloodstream by the tumor. Apoptosis, necrosis, and active cell secretion are believed to contribute to the high levels of circulating nucleic acids in the blood of some subjects with cancer. Summary of the Invention
[0004] In one aspect, provided herein is a method for detecting tumor nucleic acids in a cell-free biological sample from a subject. In some cases, the method includes circularizing a nucleic acid derived from the cell-free biological sample to produce a circularized nucleic acid. In some cases, the method includes amplifying the circularized nucleic acid to generate a concatemer, the concatemer including at least two copies of the sequence of the circularized nucleic acid. In some cases, the method includes sequencing the concatemer or its derivative to obtain the sequence of the concatemer, wherein the depth of the sequencing is no more than 18 reads. In some cases, the depth of the sequencing is no more than 18 reads per original nucleic acid. In some cases, the depth of the sequencing is no more than 1 read per concatemer. In some cases, the depth of the sequencing is no more than 1 read per circularized nucleic acid. In some cases, the method includes processing the sequence of the concatemer to identify at least two occurrences of a tumor-specific sequence variant of the subject. In some cases, the method includes identifying the nucleic acid as having the at least one tumor-specific sequence variant when identifying the at least two occurrences of the tumor-specific sequence variant in the sequence of the concatemer. In some cases, the method further comprises obtaining the tumor-specific sequence variants from the subject. In some cases, obtaining the tumor-specific sequence variants comprises sequencing a nucleic acid derived from the subject's tumor. In some cases, obtaining the tumor-specific sequence variants comprises sequencing a nucleic acid derived from healthy tissue of the subject and comparing the sequence of the nucleic acid derived from the tumor with the sequence of the nucleic acid derived from the healthy tissue. In some cases, obtaining the tumor-specific sequence variants comprises sequencing a nucleic acid derived from low tumor burden tissue or no tumor burden tissue of the subject and comparing the sequence of the nucleic acid derived from the tumor with the sequence of the nucleic acid derived from the low tumor burden tissue or no tumor burden tissue of the subject. In some cases, the depth of sequencing the nucleic acid derived from the subject's tumor is greater than 20 reads. In some cases, the depth of sequencing the nucleic acid derived from the subject's tumor is greater than 20 reads per original tumor nucleic acid molecule. In some cases, the depth of sequencing the nucleic acid derived from the subject's tumor is greater than 20 reads per nucleotide position. In some cases, the depth of sequencing the concatemer is no greater than ten reads. In some cases, the depth of the sequencing of the concatemer is no more than five reads. In some cases, the depth of the sequencing of the concatemer is no more than two reads. In some cases, sequencing depth is measured by the reading of each concatemer. In some cases, sequencing depth is measured by the reading of each original nucleic acid molecule. In some cases, the sequencing of the concatemer includes at least 10 gigabases of sequence. In some cases, the sequencing of the concatemer includes at least 10 gigabases in the total sequence of the sample.In some cases, the nucleic acid derived from the tumor is subjected to selection before sequencing. In some cases, the nucleic acid derived from the healthy tissue is subjected to selection before sequencing. In some cases, the selection includes negative selection to remove non-target sequences from the nucleic acid. In some cases, the selection includes positive selection to select the target sequence from the nucleic acid. In some cases, the method further includes, before circularization, subjecting the nucleic acid derived from the cell-free biological sample to selection. In some cases, the selection includes negative selection to remove non-target sequences from the nucleic acid. In some cases, the selection includes positive selection to select the target sequence from the nucleic acid. In some cases, circularization includes ligating the ends of the nucleic acid or its derivative to each other. In some cases, circularization includes coupling an adapter to the 5' end, the 3' end, or both the 5' end and the 3' end of the nucleic acid or its derivative. In some cases, amplifying the circularized nucleic acid is achieved by a polymerase having strand displacement activity. In some cases, amplifying the circularized nucleic acid is achieved by a polymerase having 5' to 3' exonuclease activity. In some cases, the amplification is achieved by at least one primer from a plurality of random primers. In some cases, the amplification is achieved by designing at least one primer in a plurality of primers for whole genome amplification. In some cases, the nucleic acid is single-stranded. In some cases, the nucleic acid is double-stranded. In some cases, the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some cases, sequencing includes (i) contacting the concatemer or its derivative with a plurality of nucleotides in the presence of a polymerase to incorporate one or more nucleotides in the plurality of nucleotides into a growing chain complementary to the concatemer or its derivative, and (ii) detecting one or more signals indicating that the one or more nucleotides are incorporated into the growing chain. In some cases, sequencing includes ligation sequencing. In some cases, the tumor-specific sequence variants include single nucleotide variants, fusions, insertions, deletions, or epigenetic modifications. In some cases, the acellular biological sample is a body fluid. In some cases, the body fluid includes urine, saliva, blood, serum, or plasma. In some cases, the tumor is colorectal cancer, pancreatic cancer, ovarian cancer, breast cancer, prostate cancer, bladder cancer, lung cancer, skin cancer, or blood cancer.
[0005] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0006] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, wherein the computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any method described above or elsewhere herein.
[0007] Other aspects and advantages of the present disclosure will become apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0008] Incorporation by reference
[0009] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] An understanding of the features and advantages of the present invention will be gained by referring to the following detailed description which sets forth illustrative embodiments utilizing the principles of the invention and the accompanying drawings in which:
[0011] Figure 1 Example methods for tumor-informed disease detection and monitoring using shallow sequencing are shown.
[0012] Figure 2 An example method for tumor-informed disease detection and monitoring using concatemer-based error-corrected shallow sequencing is shown.
[0013] Figure 3 A computer system is shown that is programmed or otherwise configured to implement the methods provided herein.
[0014] Figure 4 An example method for genome complexity reduction is shown.
[0015] Figure 5A An example method for identifying tumor-specific mutations is shown. Tumor tissue and normal tissue (e.g., leukocytes) from the same individual are sequenced; variants are identified by comparison with a human reference genome database. Variants identified in both normal and tumor tissues are subtracted from the tumor tissue variant list; variants found only in tumor tissue are used as tumor-specific variants for minimal residual disease detection in plasma samples from the same individual.
[0016] Figure 5BAn example method for identifying tumor-specific mutations is shown. Tumor tissue and plasma samples after treatment from the same individual are sequenced. Plasma samples are sequenced to a certain depth (i.e., >40x, or >50x, or >60x). Variants are identified by comparison with a human reference genome database. Variants identified in plasma with more than one molecule (or read) support and also found in tumor tissue are deducted from the tumor tissue variant list; variants found in tumor tissue but not found in plasma samples with more than one molecule (read) support are used as tumor-specific variants for minimal residual disease detection in plasma samples from the same individual.
[0017] Figure 6 The sequencing workflow without selection is shown. Variants are detected in the final step by single-read error correction per molecule.
[0018] Figure 7 The sequencing workflow with negative selection is shown. Blocking oligonucleotides with modified 5' ends and / or 3' ends that cannot be connected or extended are added to the ligation mixture at high concentrations. The blocking oligonucleotides bind to the DNA region through complementary sequences and form double-stranded regions. These hybrid DNA molecules will not cyclize and are optionally removed by DNA exonucleases. Only the cyclized DNA will be amplified by rolling circle amplification and sequenced in the following steps. This process can be used to selectively exclude regions in the sequencing library.
[0019] Figure 8 The sequencing workflow with positive selection is shown. Primers targeting the region of interest are spiked into the WGS reaction mixture comprising random primers. These primers are combined with the target sequence in the RCA reaction and enhance the amplification of these regions of interest. Compared with the standard workflow, the regions of interest with primer spike-in are amplified more than the regions without primer spike-in, and therefore, these regions obtain more sequencing reads in the final sequencing data.
[0020] Figure 9A The workflow for whole-genome sequencing with concatemer error correction is shown.
[0021] Figure 9B Shown are the error rates for WGS on healthy human cfDNA samples (N=3) measured by unfiltered reads, read1 read2 corrected reads, and AccuScan.
[0022] Figure 10A The detection limits for various assay conditions are shown.
[0023] Figure 10B The false positive rate when VAF=0 is shown.
[0024] Figure 10C The analytical sensitivity of AccuScan using a mixture of healthy samples is shown.
[0025] Figure 10D Titration of cancer samples is shown.
[0026] Figures 11A-11B Two workflows for AccuScan MRD detection are shown. Figure 11A A comparison between tumor tissue and blood cells is shown. Figure 11B Comparison between tumor tissue and post-treatment plasma is shown.
[0027] Figure 11C The total number of tumor-specific markers identified in the presence and absence of leukocytes is shown.
[0028] Figure 11D Shown are the variant profiles of tumor-specific markers identified in the presence and absence of leukocytes.
[0029] Figure 11E Shown is a comparison of VAF measured in plasma using the tumor-WBC workflow and the tumor-plasma workflow.
[0030] Figure 11F MRD calls using the tumor-WBC workflow versus the tumor-plasma workflow are shown.
[0031] Figure 12A Shown are the numbers of tumor-specific variants identified in CRC, ESCC, and melanoma.
[0032] Figure 12B VAFs of pre-treatment plasma samples are shown.
[0033] Figure 12C Detection of ESCC MRD in samples one week after surgery is shown.
[0034] Figure 12D Shown are all CRC recurrences detected by AccuScan prior to imaging.
[0035] Figure 12E Figure 3 shows Kaplan-Meier disease-free survival analysis of patients undergoing CRC and ESCC surgery. Patients who were ctDNA+ in their plasma samples after surgery showed significantly shorter disease-free survival.
[0036] Figure 13A AccuScan is shown for IO monitoring.
[0037] Figure 13B AccuScan for IO monitoring: dynamic changes of ctDNA over time are shown.
[0038] Figures 14A-14B The analytical sensitivity and specificity of AccuScan are shown. Figure 14A The figure shows simulations using 5000, 20000, 40000, and 80000 markers and two different error rates to predict the relationship between theoretical detection rate and cTAF at different sequencing coverages. -7 The error rate is 2.8x10 -5 The error rate of 99% indicates higher sensitivity. The detection rate was calculated as the fraction of tests that were considered MRD positive, with a nominal specificity set at 99%. Figure 14B Shown are simulations using 5,000, 20,000, 40,000, and 80,000 markers and two different error rates to predict theoretical specificity with a nominal specificity setting of 99%. Specificity is calculated as the fraction of tests that are considered MRD negative when cTAF is 0.
[0039] Figures 15A-15C Shown is ddPCR of melanoma cancer cfDNA samples. Figure 15A Shown is ddPCR of primary melanoma cancer cfDNA samples. Figure 15B Shown is ddPCR of healthy plasma samples. Figure 15C Shown is ddPCR of a diluted melanoma cancer cfDNA sample in a healthy plasma background with an expected cTAF of 0.1%.
[0040] Figure 16 The VAFs of all ctDNA-positive plasma samples are shown.
[0041] Figure 17 AccuScan error rates are shown. Overall error rates and error rates per variant type are from AccuScan WGS data of healthy human cfDNA samples (N=3) sequenced with paired-end 150 reads or single-end 300 reads using 300 cycle sequencing reagents. DETAILED DESCRIPTION
[0042] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and substitutions may occur to those skilled in the art without departing from the present invention. It will be understood that various alternatives to the embodiments of the present invention described herein may be employed.
[0043] As used herein, the term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which may depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to the practice in the art, "about" can mean within 1 or more than 1 standard deviation. As another example, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. With respect to biological systems or processes, the term "about" can mean within an order of magnitude, such as within 5 times or 2 times a value. Where specific values are described in the application and claims, unless otherwise stated, the term "about" means within an acceptable error range for the specific value.
[0044] As used herein, the terms "polynucleotide," "nucleotide," "nucleotide sequence," "nucleic acid," and "oligonucleotide" are used interchangeably and generally refer to a polymeric form of nucleotides of any length, whether deoxyribonucleotides (DNA) or ribonucleotides (RNA) or analogs thereof. A polynucleotide can have any three-dimensional structure and can perform any function. The following are non-limiting examples of polynucleotides: cell-free nucleic acid, cell-free DNA (cfDNA), cell-free RNA (cfRNA), circulating tumor DNA (ctDNA), circulating tumor RNA (ctRNA), coding or non-coding regions of a gene or gene fragment, a locus (locus) defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide can contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after polymer assembly. The nucleotide sequence can be interrupted by non-nucleotide components. The polynucleotide can be further modified after polymerization, such as by conjugation with a labeling component.
[0045] As used herein, the term "subject" generally refers to a vertebrate, such as a mammal (e.g., a human). Mammals include, but are not limited to, rodents, monkeys, humans, farm animals, competitive animals, and pets (e.g., dogs or cats). Also included are tissues, cells, and progeny of biological entities obtained in vivo or cultured in vitro. The subject can be a patient. The subject may have symptoms associated with a disease (e.g., cancer). Alternatively, the subject may not have symptoms associated with the disease.
[0046] As used herein, the term "biological sample" generally refers to a sample derived from or obtained from a subject such as a mammal (e.g., a human). A biological sample can include, but is not limited to, hair, nails, skin, sweat, tears, eye fluid, nasal swab or nasopharyngeal wash, sputum, throat swab, saliva, mucus, blood, serum, plasma, placental fluid, amniotic fluid, umbilical cord blood, emphasizing fluid (emphatic fluid), cavity fluid, earwax, oil, glandular secretions, bile, lymph, pus, microbial flora, meconium, breast milk, bone marrow, bone, central nervous system (CNS) tissue, cerebrospinal fluid, adipose tissue, synovial fluid, feces, gastric juice, urine, semen, vaginal secretions, stomach, small intestine, large intestine, rectum, pancreas, liver, kidney, bladder, lung and other tissues and liquids derived from or obtained from a subject. A biological sample can be a cell-free (or cell-free) biological sample.
[0047] As used herein, the term "cell-free biological sample" generally refers to a sample that is derived from or obtained from a subject and that does not contain cells. Cell-free biological samples can include, but are not limited to, blood, serum, plasma, nasal swabs or nasopharyngeal washes, saliva, urine, gastric juice, tears, feces, mucus, sweat, cerumen, oils, glandular secretions, bile, lymph, cerebrospinal fluid, tissue, semen, vaginal fluid, interstitial fluid (including interstitial fluid from tumor tissue), eye fluid, spinal fluid, throat swabs, breath, hair, nails, skin, biopsy tissue sections, placental fluid, amniotic fluid, umbilical cord blood, focal fluids, cavity fluids, sputum, pus, microbiota, meconium, breast milk, and / or other excretions.
[0048] Whenever the term "at least," "greater than," or "greater than or equal to" precedes the first value in a series of two or more values, the term "at least," "greater than," or "greater than or equal to" applies to every value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0049] Whenever the term "not greater than," "less than," or "less than or equal to" precedes the first value in a series of two or more values, the term "not greater than," "less than," or "less than or equal to" applies to every value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0050] Molecular residual disease (MRD) refers to the persistence of cancer cells after curative treatment. Timely and sensitive measurement of MRD is crucial for relapse risk assessment, treatment prognosis and patient stratification. Circulating tumor DNA (ctDNA), which is released by cancer cells and has a short half-life (<2 hours), has emerged as a promising real-time biomarker for MRD detection and monitoring. Studies have shown that the level of cancer-specific somatic mutations in ctDNA is associated with tumor stage, burden and response to treatment across tumor types. Compared with other blood-based cancer biomarkers (such as circulating tumor cells and cancer antigens), ctDNA provides a more sensitive and specific measurement of MRD.
[0051] There are currently two main strategies for ctDNA-based MRD detection: 1) tumor naivety Methods that test MRD samples for alterations known to be enriched in tumors, such as common somatic mutations and methylation changes; and 2) tumor-informed methods that require tumor samples to identify patient-specific variants and then test these variants in MRD samples.
[0052] The tumor naive approach is logistically simple, eliminating the need to collect and sequence tumor samples and use a universal panel to test plasma samples for the presence of cancer signals. While these tests offer operational convenience, they tend to have moderate limits of detection (LOD). A methylation-based cancer detection test claimed to have a 50% sensitivity of 3.1x10-4 circulating tumor allele fraction (cTAF). It was also shown that using plasma collected at a landmark time point (4 weeks after surgery), 55.6% of CRC recurrences were detected by combining a panel of methylation and mutation signals.
[0053] Tumor-informed approaches incorporate patient-specific somatic mutation information from tumor tissue into MRD analysis, which can result in ultra-low limits of detection. Factors influencing their sensitivity include the accuracy of somatic mutation calls from tissue and plasma samples, and the total number of cfDNA molecules interrogated, which is the product of the number of somatic variants tracked and the depth of unique molecules obtained by sequencing.
[0054] Tumor-informed approaches can use custom or off-the-shelf MRD tests. Custom MRD assays are designed after tumor results become available and track a limited number of variants through ultra-deep sequencing. Sequencing of custom panels can be exhaustive; therefore, the depth of unique molecules is primarily limited by the amount of available input material. For example, Signatera, a tumor-informed NGS-based multiplex PCR assay, detects MRD at a limit of detection (LOD) of 10 when using up to 66 ng of DNA. -4When tracking 16 personalized markers, the analytical sensitivity reached 81.3%-96.1%. Tumor-informed personalized MRD assays targeting a large number of markers and using UMI or duplex sequencing for error correction showed LODs below 10 -4 Phase-Seq uses multiple somatic mutations in a single DNA fragment to detect ctDNA, which reduces background noise to less than 10 -6 , and claimed that the detection limit dropped to the PPM level given sufficient phased variants. Although tumor-informed customized MRD methods can achieve very high sensitivity, the requirement of personalized design significantly increases the turnaround time (TAT) and poses considerable logistical challenges.
[0055] A tumor-informed, off-the-shelf approach uses the same assay for both tumor and plasma from all patients. Requiring no patient-specific reagents, it has the low TAT of a tumor-naive approach and offers much simpler logic than custom approaches. The challenge is to generate an off-the-shelf assay that covers enough of the genome with a low enough error rate. Predesigned MRD panels targeting cancer-associated genes typically use UMIs with deep sequencing to achieve high accuracy in variant calling, but these panels track a small number of markers per patient. For example, a 130 kb panel covering 139 key lung cancer-associated genes captured only a median of 2 mutations per patient (range: 1-8 mutations).
[0056] Recently, whole genome sequencing (WGS) assays have emerged as an innovative approach for cancer screening and MRD detection. Tumor-informed WGS MRD assays use genome breadth to complement sequencing depth for sensitivity, overcoming limitations in input sample size. UMI-based error correction, which relies on having multiple reads per input molecule, is prohibitively expensive at the scale of WGS. Some have used read-centric SVM models to reduce the WGS somatic single nucleotide variant (SNV) error rate to 4.96x10 -5 By exploiting the cumulative signals of thousands of somatic mutations observed in tumor genomes, they -4 Analytical sensitivity of 95% has been reported for tumor fractions of <10. -7 However, these methods suffer from low conversion rates, making low LODs difficult to achieve. An efficient and cost-effective genome-wide error correction method is needed to enable WGS for MRD detection with low LODs.
[0057] DNA concatemers, which physically connect DNA copies, produced via rolling circle amplification (RCA), allow error correction at the single read level. The combination of RCA and repeat confirmation eliminates both PCR and sequencing errors. Compared to UMI methods, concatemer sequencing shows higher efficiency in error correction when applied to genomic DNA. Recently, concatemer sequencing has been applied to liquid biopsies to demonstrate the feasibility of applying this technology to treatment selection and cancer screening. This article provides a WGS solution for ctDNA detection that utilizes concatemer sequencing for whole-genome single read error suppression, enabling rapid and sensitive MRD detection and monitoring in cancer patient plasma samples.
[0058] In one aspect, the present invention provides a method for detecting tumor nucleic acids in a biological sample from a subject. In some cases, the method includes detecting tumor nucleic acids in a cell-free biological sample from a subject. In some cases, the method includes cyclizing a nucleic acid derived from a biological sample (such as a cell-free biological sample) to produce a circularized nucleic acid. Next, the method may include amplifying the circularized nucleic acid to generate a concatemer, the concatemer including the sequence of at least two copies of the circularized nucleic acid. Then, the concatemer or its derivatives may be sequenced to obtain the sequence of the concatemer. In some cases, the depth of sequencing is no more than 18 reads. Next, the sequence of the concatemer is processed to identify at least two occurrences of a tumor-specific sequence variant of the subject. When at least two occurrences of a tumor-specific sequence variant are identified in the sequence of the concatemer, the method may include identifying the nucleic acid as having at least one tumor-specific sequence variant. The method may further include obtaining a tumor-specific sequence variant from the subject, for example, by sequencing a nucleic acid derived from a tumor of the subject. In some cases, the method further includes sequencing a nucleic acid derived from a healthy tissue of the subject, and comparing the sequence from the nucleic acid derived from the tumor with the sequence from the nucleic acid derived from the healthy tissue of the subject. In some cases, sequencing of nucleic acids derived from a tumor is performed at a suitable depth, measured as reads per molecule or reads, which are used interchangeably herein. In some cases, the depth of sequencing of nucleic acids derived from a tumor in a subject is greater than 20 reads. In some cases, the depth of sequencing of nucleic acids derived from a tumor in a subject is greater than 25 reads. In some cases, the depth of sequencing of nucleic acids derived from a tumor in a subject is greater than 30 reads. In some cases, the depth of sequencing of nucleic acids derived from a tumor in a subject is greater than 35 reads. In some cases, the depth of sequencing of nucleic acids derived from a tumor in a subject is greater than 40 reads.
[0059] In another aspect of the method for detecting tumor nucleic acids in a biological sample herein, sequencing of the concatemers is performed at a suitable depth, measured as reads per molecule or reads, as used interchangeably herein. In some cases, the depth of sequencing of the concatemers is no greater than 18 reads. In some cases, the depth of sequencing of the concatemers is no greater than 15 reads. In some cases, the depth of sequencing of the concatemers is no greater than 12 reads. In some cases, the depth of sequencing of the concatemers is no greater than 10 reads. In some cases, the depth of sequencing of the concatemers is no greater than nine reads. In some cases, the depth of sequencing of the concatemers is no greater than eight reads. In some cases, the depth of sequencing of the concatemers is no greater than seven reads. In some cases, the depth of sequencing of the concatemers is no greater than six reads. In some cases, the depth of sequencing of the concatemers is no greater than five reads. In some cases, the depth of sequencing of the concatemers is no greater than four reads. In some cases, the depth of sequencing of the concatemers is no greater than three reads. In some cases, the depth of sequencing of the concatemers is no greater than two reads. In some cases, the depth of sequencing of the concatemers is no greater than one read. In some cases, the sequencing of the concatemers is whole genome sequencing. In some cases, the sequencing of the concatemers includes at least 10 gigabases of sequence.
[0060] In another aspect of detecting tumor nucleic acids in biological samples herein, nucleic acids derived from tumors are selected before sequencing. In some cases, nucleic acids derived from healthy tissues are selected before sequencing. In some cases, nucleic acids derived from cell-free biological samples are selected before cyclizing nucleic acids. In some cases, selection includes negative selection to remove non-target sequences from the nucleic acid. In some cases, negative selection includes contacting the nucleic acid with a blocker that binds to the non-target sequence and amplifying, connecting or capturing nucleic acids that are not bound to the blocker. In some cases, the blocker includes an oligonucleotide. In some cases, negative selection includes contacting the nucleic acid with a nuclease that specifically cuts the non-target sequence. In some cases, the nuclease is a clustered regularly interspaced short palindromic repeat (CRISPR) nuclease. In some cases, selection includes positive selection to select a target sequence from the nucleic acid. In some cases, positive selection includes hybridization capture. In some cases, positive selection includes amplification. In some cases, amplification includes polymerase chain reaction (PCR).
[0061] In further aspects of the methods of detecting tumor nucleic acids in a biological sample herein, in some cases, circularizing the nucleic acid derived from the biological sample comprises ligating the ends of the nucleic acid or its derivative to each other. In some cases, circularizing the nucleic acid derived from the biological sample comprises coupling an adaptor to the 5' end, the 3' end, or both the 5' end and the 3' end of the nucleic acid or its derivative.
[0062] In another aspect of the method for detecting tumor nucleic acid in a biological sample herein, in some cases, amplifying the circularized nucleic acid to generate concatemers is achieved by a polymerase having strand displacement activity. In some cases, amplifying the circularized nucleic acid to generate concatemers is achieved by a polymerase having 5' to 3' exonuclease activity. In some cases, amplification is achieved by at least one primer from a plurality of random primers. In some cases, amplification is achieved by at least one primer from a plurality of primers designed for whole genome amplification.
[0063] In another aspect of the method for detecting tumor nucleic acid in a biological sample herein, in some cases, the nucleic acid in the biological sample is single-stranded. In some cases, the nucleic acid is double-stranded. In some cases, the nucleic acid in the biological sample is a mixture of single-stranded nucleic acid and double-stranded nucleic acid. In some cases, the nucleic acid is made into a single strand before cyclization. In some cases, the nucleic acid is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination of DNA and RNA.
[0064] In another aspect of the methods herein for detecting tumor nucleic acids in a biological sample, in some cases, sequencing the concatemers comprises contacting the concatemers or derivatives thereof with a plurality of nucleotides in the presence of a polymerase to incorporate one or more nucleotides of the plurality of nucleotides into a growing chain that is complementary to the concatemers or derivatives thereof, and detecting one or more signals indicative of incorporation of the one or more nucleotides into the growing chain. Alternatively or in combination, sequencing the concatemers comprises ligation sequencing. Sequencing of the concatemers may comprise any suitable method provided herein.
[0065] In another aspect of the methods provided herein for detecting tumor nucleic acids in a biological sample, in some cases, the tumor-specific sequence variants comprise single nucleotide variants, fusions, insertions, deletions, epigenetic modifications, or any combination thereof.
[0066] In another aspect of the methods provided herein for detecting tumor nucleic acids in a biological sample, in some cases, the biological sample is a cell-free biological sample. In some cases, the cell-free biological sample is a body fluid. In some cases, the body fluid comprises urine, saliva, blood, serum, or plasma. In some cases, the biological sample, cell-free biological sample, or body fluid is any suitable sample provided herein.
[0067] In another aspect of the methods for detecting tumor nucleic acids in a biological sample provided herein, in some cases, the tumor is colorectal cancer, pancreatic cancer, ovarian cancer, breast cancer, prostate cancer, bladder cancer, lung cancer, skin cancer, or blood cancer. In some cases, the tumor is any cancer suitable for detection as provided herein.
[0068] Library preparation and amplification methods
[0069] In some cases, the method for detecting tumor nucleic acid provided herein includes amplifying polynucleotides present in a sample from a subject. The amplification method used herein generally includes rolling circle amplification. Alternatively or in combination, the amplification method used herein includes PCR. In some cases, the amplification method herein includes linear amplification. Amplification is generally not directed to a gene or a group of genes, but rather to amplifying the entire nucleic acid sample. In some cases, the method includes (a) cyclizing a single polynucleotide in a plurality of polynucleotides to form a plurality of circular polynucleotides, each of which has a junction between the 5' end and the 3' end; and (b) amplifying the circular polynucleotide of (a) to produce amplified polynucleotides. In other cases, the amplification method includes (c) shearing the amplified polynucleotides to produce sheared polynucleotides, each sheared polynucleotide comprising one or more shearing points at the 5' end and / or 3' end. In some cases, the method does not include enrichment of the target sequence.
[0070] Typically, the ends of a polynucleotide are joined to each other to form a circular polynucleotide (directly joined, or joined by one or more intermediate adapter oligonucleotides) to produce a junction with a junction sequence. When the 5' end and 3' end of a polynucleotide are joined by an adapter polynucleotide, the term "junction" can refer to the junction between the polynucleotide and the adapter (e.g., one of the 5' end junction or the 3' end junction), or refers to the junction formed between the 5' end and the 3' end of a polynucleotide and comprising the adapter polynucleotide. When the 5' end and 3' end of a polynucleotide are joined without an intervening adapter (e.g., the 5' end and 3' end of a single-stranded DNA), the term "junction" refers to the point where the two ends join. A junction can be identified by the nucleotide sequence (also referred to as "junction sequence") constituting the junction.
[0071] The samples herein comprise polynucleotides having a mixture of ends formed by natural degradation processes (such as cell lysis, cell death, and other processes in which polynucleotides such as DNA and RNA are released from cells into their surroundings, where they can further degrade, such as cell-free polynucleotides, such as cell-free DNA and cell-free RNA), fragmentation as a byproduct of sample processing (such as fixation, staining, and / or storage processes), and fragmentation by methods that cut DNA that are not limited to a specific target sequence (e.g., mechanical fragmentation, such as by sonication; treatment with non-sequence-specific nucleases, such as DNase I, fragmentase). When a sample comprises polynucleotides having a mixture of ends, the probability that two polynucleotides have the same 5' end or 3' end is low, and the probability that two polynucleotides independently have the same 5' end and 3' end is even lower. Therefore, in some embodiments, even if two polynucleotides comprise portions having the same target sequence, junctions can be used to distinguish between different polynucleotides. When polynucleotide ends are joined without an intervening adapter, the joined sequence can be identified by alignment with a reference sequence. For example, where the order of two component sequences appears to be reversed relative to a reference sequence, the point at which the reversal appears to have occurred can represent a junction at that point. When the polynucleotide ends are joined by one or more adapter sequences, the junction can be identified by proximity to known adapter sequences, or, if the sequencing read is long enough to obtain sequence from both the 5' and 3' ends of the circularized polynucleotide, by alignment as described above. In some embodiments, the formation of a particular junction is a very rare event such that it is unique among the circularized polynucleotides of a sample.
[0072] In some embodiments, the single polynucleotide in cyclization (a) is achieved by subjecting multiple polynucleotides to ligation. Ligation can include a ligase. In some cases, the ligase is a single-stranded DNA or RNA ligase. In some cases, the ligase is a double-stranded DNA ligase. In some embodiments, the ligase is degraded before the amplification in (b). Degrading the ligase before the amplification in (b) can increase the recovery rate of amplifiable polynucleotides. In some embodiments, multiple cyclized polynucleotides are not purified or separated before (b). In some embodiments, uncyclized linear polynucleotides are degraded before amplification. In some cases, multiple polynucleotides are denatured to produce single-stranded polynucleotides before cyclization; in some cases, multiple polynucleotides are not denatured before cyclization.
[0073] In some cases, the circularization in (a) includes the step of joining an adapter polynucleotide to the 5' end, the 3' end, or both the 5' end and the 3' end of a polynucleotide in the plurality of polynucleotides. As previously described, when the 5' end and / or the 3' end of a polynucleotide is joined by an adapter polynucleotide, the term "junction" can refer to the junction between the polynucleotide and the adapter (e.g., one of the 5' end junction or the 3' end junction), or to the junction formed between the 5' end and the 3' end of the polynucleotide and comprising the adapter polynucleotide.
[0074] In some cases, polynucleotide is subjected to selection step.In some cases, polynucleotide with sequence of interest is subjected to positive selection step, has polynucleotide of sequence of interest with enrichment.Alternatively, polynucleotide with unwanted sequence is subjected to negative selection step to remove polynucleotide with unwanted sequence.In some cases, negative selection comprises and makes polynucleotide denaturation to produce single-stranded polynucleotide, one or more blocking oligonucleotides are annealed to polynucleotide, to produce double-stranded polynucleotide and single-stranded polynucleotide with unwanted sequence, and cyclization single-stranded polynucleotide.In some cases, blocking oligonucleotide has the 5 ' end of modification that does not allow connection and / or the 3 ' end of modification.In some cases, blocking oligonucleotide has the 5 ' end of modification that does not allow extension and / or the 3 ' end of modification.In some cases, exonuclease is used to remove linear double-stranded polynucleotide.Cyclation polynucleotide can be used in the subsequent step of rolling circle amplification and order-checking.
[0075] In one aspect, this paper provides a method for identifying sequence variants in multiple polynucleotides, including making multiple polynucleotides denaturation, annealing the blocking oligonucleotide to the polynucleotide with unwanted sequence, and cyclizing the resulting single-stranded polynucleotides. In some cases, for example, using a nuclease (such as a DNA exonuclease), the remaining linear polynucleotides annealed to the blocking oligonucleotides are degraded. Next, the cyclized polynucleotides can be amplified by rolling circle amplification to produce a concatemer comprising the original polynucleotides of more than one copy. In some cases, rolling circle amplification is achieved by random primers. In some cases, rolling circle amplification is achieved by target-specific primers. Next, concatemers are subjected to sequencing to obtain sequencing reads. These sequencing reads are used to identify variants. In some cases, when variants are present in concatemers on the polynucleotides of more than one copy, they are identified. In some cases, when variants are present on two different concatemers, they are identified.
[0076] In some cases, such as after ligase degradation, amplification cyclization polynucleotides are used to produce amplified polynucleotides. (b) Amplification of circular polynucleotides can be achieved by polymerase. In some cases, polymerase is a polymerase with strand displacement activity. In some cases, polymerase is Phi29 DNA polymerase. Or, polymerase is a polymerase without strand displacement activity. In some cases, polymerase is T4 DNA polymerase or T7 DNA polymerase. Alternatively or in combination, polymerase is Taq polymerase, or a polymerase in the Taq polymerase family. In some cases, amplification includes rolling circle amplification (RCA). The polynucleotides of the amplification produced by RCA can include linear concatemers, or include polynucleotides of a target sequence (for example, a subunit sequence) more than one copy from the template polynucleotide. In some embodiments, amplification includes subjecting the circular polynucleotides to an amplification reaction mixture comprising random primers. In some cases, amplification includes subjecting the circular polynucleotides to an amplification reaction mixture comprising one or more primers, each of which specifically hybridizes to different target sequences via sequence complementarity. In some cases, amplifying comprises subjecting the circular polynucleotide to an amplification reaction mixture comprising a reverse primer.
[0077] In some cases, the amplified polynucleotide is sheared to produce a sheared polynucleotide of shorter length relative to the unsheared polynucleotide. Two or more sheared polynucleotides derived from the same linear concatemer may have the same junction sequence but may have different 5' and / or 3' ends (e.g., sheared ends).
[0078] The cell-free polynucleotides from the sample can be any of a variety of polynucleotides, including but not limited to DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), small interfering RNA (siRNA), fragments of any of these, or any combination of two or more of these. In some embodiments, the sample comprises DNA. In some embodiments, the sample comprises cell-free genomic DNA. In some embodiments, the sample comprises DNA produced by amplification, such as by a primer extension reaction using any suitable combination of primers and DNA polymerase, including but not limited to polymerase chain reaction (PCR), reverse transcription, and combinations thereof. When the template for the primer extension reaction is RNA, the reverse transcription product is referred to as complementary DNA (cDNA). The primers used for the primer extension reaction can include sequences specific to one or more targets, random sequences, partially random sequences, and combinations thereof. In some cases, the primers include a mixture of random sequences and sequences specific to one or more targets. Typically, the sample polynucleotides include any polynucleotides present in the sample, which may or may not include the target polynucleotides. The polynucleotides can be single-stranded, double-stranded, or a combination thereof. In some embodiments, the polynucleotides that can be subjected to the disclosed methods are single-stranded polynucleotides, which may or may not be double-stranded polynucleotides. In some embodiments, the polynucleotides are single-stranded DNA. Single-stranded DNA (ssDNA) can be ssDNA isolated in single-stranded form, or DNA isolated in double-stranded form and subsequently made single-stranded for the purpose of one or more steps in the disclosed methods.
[0079] In one aspect, the present invention provides a method for identifying sequence variants in multiple polynucleotides, comprising denaturing the polynucleotides, cyclizing the resulting linear polynucleotides, and amplifying the resulting circular polynucleotides, wherein the amplification step is used to enrich the sequence of interest, such as by adding one or more primers that bind the sequence of interest to an amplification reaction comprising random primers. Random primers and primers that bind the sequence of interest are used to amplify the circular polynucleotides to generate concatemers by rolling circle amplification. Next, the concatemers are subjected to sequencing to obtain sequencing reads. These sequencing reads are used to identify variants. In some cases, when a variant is present in a polynucleotide more than one copy in a concatemer, it is identified. In some cases, when a variant is present on two different concatemers, it is identified.
[0080] In some embodiments, polynucleotide is made to undergo subsequent steps (such as cyclization and amplification) in the absence of an extraction step and / or in the absence of a purification step. For example, a fluid sample can be processed to remove cells in the absence of an extraction step, thereby producing a liquid sample and a cell sample of purification, and then DNA is separated from the fluid sample of purification. There are multiple methods for separating polynucleotides available, such as by precipitation, or with substrate non-specific binding and subsequently washing the substrate to release the polynucleotides of combination. When separating polynucleotides from a sample in the absence of a cell extraction step, polynucleotides will mainly be extracellular or "cell-free" polynucleotides, such as cell-free DNA and cell-free RNA, which can correspond to dead or damaged cells. The characteristic of these cells can be used to characterize the cell or cell mass from which it is derived, such as tumor cells (when for example cancer detection), fetal cells (when for example prenatal diagnosis), cells from transplanted tissue (when for example early detection of transplantation failure) or members of microbial communities.
[0081] If the sample is processed to extract polynucleotides, such as from cells in the sample, a variety of extraction methods can be used. For example, nucleic acids can be purified by organic extraction with phenol, phenol / chloroform / isoamyl alcohol, or similar formulations (including TRIzol and TriReagent). Other non-limiting examples of extraction techniques include: (1) organic extraction followed by ethanol precipitation, for example, using phenol / chloroform organic reagent (Ausubel et al., 1993, which is incorporated herein by reference in its entirety), with or without an automated nucleic acid extractor, such as the 341 DNA extractor available from Applied Biosystems (Foster city, Calif); (2) stationary phase adsorption (U.S. Pat. No. 5,234,809; Walsh et al., 1991, each of which is incorporated herein by reference in its entirety); and (3) salt-induced nucleic acid precipitation (Miller et al., 1988, which is incorporated herein by reference in its entirety), which is generally referred to as a "salting-out" method. Another example of nucleic acid separation and / or purification includes using magnetic particles that can specifically or non-specifically bind nucleic acid, then using magnet separation beads, and washing and eluting nucleic acid from beads (see, for example, U.S. Patent number 5,705,628, which is incorporated herein by reference in its entirety). In some embodiments, an enzymatic digestion step can be carried out before the above-mentioned separation method to help remove unwanted protein from the sample, for example, digested with proteinase K or other similar proteases. See, for example, U.S. Patent number 7,001,724, which is incorporated herein by reference in its entirety. If necessary, an Rnase inhibitor can be added to the lysis buffer. For specific cells or sample types, it may be necessary to increase protein denaturation / digestion steps in the scheme. Purification methods can be directed to separating DNA, RNA or both. When DNA and RNA are separated together during or after the extraction procedure, further steps can be used to purify one or both separately from another. The subfraction of the nucleic acid extracted can also be generated, for example, purified according to size, sequence or other physical or chemical properties. In addition to the initial nucleic acid isolation step, purification of nucleic acids can also be performed after any step of the disclosed methods, such as to remove excess or unwanted reagents, reactants, or products. A variety of methods for determining the amount of nucleic acid in a sample and / or the purity of the nucleic acid are available, such as by absorbance (e.g., absorbance at 260 nm, 280 nm, and ratios thereof) and detection of markers (e.g., fluorescent dyes and intercalators, such as SYBR green, SYBR blue, DAPI, propidium iodide, Hoechst stain, SYBR gold, ethidium bromide).
[0082] In some cases, the methods herein include preparing a DNA library from a polynucleotide. For example, the methods herein include preparing a single-stranded DNA library. Any suitable method for preparing a single-stranded DNA library can be used in the methods herein. For example, the method for preparing a single-stranded DNA library includes denaturing a DNA sample to produce a plurality of ssDNAs; connecting an adapter to the 3' end of an ssDNA molecule or extending the 3' end of an ssDNA molecule by non-templated synthesis; synthesizing a second strand using a primer complementary to the adapter or 3' extension sequence; connecting a double-stranded adapter to the extension product; amplifying the second strand using primers targeting the first adapter and the second adapter (e.g., using PCR); and sequencing the library on a sequencer. Additional methods for preparing single-stranded libraries include denaturing a DNA sample to generate a plurality of ssDNAs; ligating adapters to the 3' ends of the ssDNA molecules; synthesizing a second strand using primers complementary to the adapters; ligating a double-stranded adapter to the extension product; amplifying the second strand using primers targeting the first adapter and the second adapter (e.g., by PCR); optionally enriching for regions of interest using hybridization with capture probes; amplifying the captured products (e.g., by PCR); and sequencing the library on a sequencer.
[0083] A further example of single-stranded library preparation includes a method comprising the steps of treating DNA with a thermolabile phosphatase to remove residual phosphate groups from the 5' and 3' ends of the DNA strand; removing deoxyuracil resulting from cytosine deamination from the DNA strand; ligating a 5'-phosphorylated adaptor oligonucleotide having approximately 10 nucleotides and a long 3' biotinylated spacer arm to the 3' end of the DNA strand; immobilizing the adaptor-ligated molecules on streptavidin beads; replicating the template strand using a 5' tailing primer complementary to the adaptor using Bst polymerase; washing away excess primer; removing the 3' overhang using T4 DNA polymerase; joining a second adaptor to the newly synthesized strand using blunt-end ligation; washing away excess adaptor; releasing the library molecules by heat denaturation; adding the full-length adaptor sequence including the barcode by amplification using the tailing primer; and sequencing the library as described in Gansauge et al. 2013. Nature Protocols. 8(4)737-748, which is incorporated herein by reference in its entirety.
[0084] In another embodiment, the method herein includes preparing a double-stranded DNA library. Any suitable method for preparing a double-stranded DNA library can be used in the method herein. For example, a method for preparing a double-stranded DNA library includes connecting sequencing adapters to the 5' and 3' ends of multiple DNA fragments and sequencing the library on a sequencer. Additional methods for preparing a double-stranded DNA library include connecting adapters to the 5' and 3' ends of multiple DNA fragments; using primers complementary to the connected adapters, attaching the complete adapter sequence to the connected fragments by PCR; and sequencing the library on a sequencer. A further method includes connecting adapters to the 5' and 3' ends of multiple DNA fragments; amplifying the connected products by PCR complementary to the connected adapters; optionally enriching the region of interest by hybridization with a capture probe; PCR amplifying the captured products; and sequencing the library on a sequencer. Additional methods for preparing double-stranded libraries include ligating adapters to the 5' and 3' ends of multiple DNA fragments; amplifying the ligated products by PCR using primers complementary to the ligated adapters; circularizing the double-stranded PCR products or denaturing and circularizing the single-stranded PCR products; optionally, enriching regions of interest by PCR using primers targeted to specific genes; and sequencing the library on a sequencer.
[0085] Further examples of double-stranded library preparation include the Safe-Sequencing System described in Kinde et al. (Kinde et al., 2011. Proc. Natl. Acad. Sci., USA, 108(23)9530-9535, which is incorporated herein by reference in its entirety), which includes assigning a unique identifier (UID) to each template molecule; amplifying each uniquely labeled template molecule to create a UID family; and redundant sequencing of the amplified products. Additional examples include the circularized single molecule amplification and resequencing technology (cSMART) described in Lv et al. (Lv et al., 2015. Clin. Chem., 61(1)172-181, which is incorporated herein by reference in its entirety), which labels individual molecules with unique barcodes, circularizes, targets alleles for replication by inverse PCR, and then sequences the prepared library and counts the alleles present.
[0086] In additional library preparation methods, antibodies are used to select cfDNA fragments with certain characteristics. In some cases, antibodies are used to select methylated or hypermethylated cfDNA fragments. The selected cfDNA fragments are then used in any of the library preparation methods described herein, including circularization, single-stranded DNA library preparation, and double-stranded DNA library preparation. Sequencing such isolated cfDNA fragments can provide information about the characteristics present in the cfDNA, including modifications such as methylation or hypermethylation.
[0087] In some embodiments, the polynucleotides from the multiple polynucleotides of the sample are cyclized.Cyclization can include joining the 5' end of the polynucleotide to the 3' end of the same polynucleotide, joining to the 3' end of another polynucleotide in the sample, or joining to the 3' end of the polynucleotides from different sources (for example, artificial polynucleotides, such as oligonucleotide adapters). In some embodiments, the 5' end of the polynucleotide is joined to the 3' end of the same polynucleotide (also referred to as "self-engagement"). In some embodiments, the conditions of the cyclization reaction are selected to facilitate the self-engagement of the polynucleotides within the range of a specific length, so as to generate a cyclization polynucleotide colony with a specific average length. For example, cyclization reaction conditions can be selected to facilitate the self-engagement of the polynucleotides of a length shorter than about 5000, 2500, 1000, 750, 500, 400, 300, 200, 150, 100, 50 or less nucleotides. In some embodiments, the fragment that is biased towards a length of 50-5000 nucleotide, 100-2500 nucleotide or 150-500 nucleotide is so that the average length of the cyclized polynucleotide falls within the scope of each other. In some embodiments, the length of the cyclized fragments of 80% or more is 50-500 nucleotide, such as a length of 50-200 nucleotide. The reaction conditions that can be optimized include the concentration of the time span assigned to the joining reaction, the concentration of various reagents and the polynucleotide to be joined. In some embodiments, the cyclization reaction remains on the distribution of the fragment length present in the sample before cyclization. For example, one or more of the fragment length in the sample before cyclization and the mean value, median, mode and standard deviation of the cyclized polynucleotide are within 75%, 80%, 90%, 95% or higher percentages.
[0088] In some cases, one or more adapter oligonucleotides are used so that the 5' end and 3' end of the polynucleotide in the sample are joined to form a circular polynucleotide by one or more intervening adapter oligonucleotides, rather than preferentially forming a self-engagement cyclization product. For example, the 5' end of the polynucleotide can be joined to the 3' end of the adapter, and the 5' end of the same adapter can be joined to the 3' end of the same polynucleotide. The adapter oligonucleotide includes any oligonucleotide with a certain sequence, at least a portion of which is known and can be connected to the sample polynucleotide. The adapter oligonucleotide can comprise DNA, RNA, nucleotide analogs, atypical nucleotides, labeled nucleotides, modified nucleotides or a combination thereof. The adapter oligonucleotide can be single-stranded, double-stranded or partially duplexed. Typically, the partially duplexed adapter comprises one or more single-stranded regions and one or more double-stranded regions. Double-stranded adapters can comprise two separated oligonucleotides (also referred to as "oligonucleotide duplexes") that hybridize to each other, and the hybridization can leave one or more blunt ends, one or more 3' overhangs, one or more 5' overhangs, one or more protrusions caused by mismatched and / or unpaired nucleotides, or any combination thereof. When the two hybridization regions of the adapter are separated from each other by a non-hybridized region, a "bubble" structure is produced. Different types of adapters, such as adapters with different sequences, can be used in combination. Different adapters can be joined to the sample polynucleotide in sequential reactions or simultaneously. In some embodiments, the same adapter is added to both ends of the target polynucleotide. For example, the first and second adapters can be added to the same reaction. The adapter can be operated before being combined with the sample polynucleotide. For example, terminal phosphates can be added or removed.
[0089] Where adaptor oligonucleotides are used, the adaptor oligonucleotides may comprise one or more of a variety of sequence elements, including, but not limited to, one or more amplification primer annealing sequences or their complements, one or more sequencing primer annealing sequences or their complements, one or more barcode sequences, one or more common sequences shared between multiple different adaptors or a subset of different adaptors, one or more restriction enzyme recognition sites, one or more overhangs complementary to one or more target polynucleotide overhangs, one or more probe binding sites (e.g., for attachment to a sequencing platform, such as a flow cell for massively parallel sequencing, such as a flow cell developed by Illumina, Inc.), one or more random or near-random sequences (e.g., one or more nucleotides randomly selected at one or more positions from a set of two or more different nucleotides, wherein each of the different nucleotides selected at the one or more positions is represented in a set of adaptors comprising a random sequence), or a combination thereof. In some cases, adapters can be used to purify these adapter-containing rings, for example by using beads coated with oligonucleotides containing adapter complements (for ease of handling, particularly magnetic beads), which can "capture" closed rings with the correct adapters by hybridizing therewith, washing away those rings that do not contain adapters and any unconnected components, and then releasing the captured rings from the beads. In addition, in some cases, the complex of the hybridized capture probe and the target ring can be used directly to generate concatemers, such as by direct rolling circle amplification (RCA). In some embodiments, the adapters in the ring can also be used as sequencing primers. Two or more sequence elements can be non-adjacent to each other (e.g., separated by one or more nucleotides), adjacent to each other, partially overlapping, or completely overlapping. For example, an amplification primer annealing sequence can also be used as a sequencing primer annealing sequence. The sequence element can be located at or near the 3' end, at or near the 5' end, or inside the adapter oligonucleotide. The sequence element can be any suitable length, such as about or less than about 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides in length. The adapter oligonucleotide can have any suitable length, at least enough to accommodate the one or more sequence elements it comprises. In some embodiments, the length of the adapter is about or less than about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, 200 or more nucleotides. In some embodiments, the length of the adapter oligonucleotide is in the range of about 12 to 40 nucleotides, such as a length of about 15 to 35 nucleotides.
[0090] In some embodiments, the adapter oligonucleotides joined to the fragmented polynucleotides from a sample comprise a sequence common to one or more adapter oligonucleotides and a barcode unique to the adapters joined to the polynucleotides of the particular sample, such that the barcode sequence can be used to distinguish polynucleotides derived from one sample or adapter ligation reaction from polynucleotides derived from another sample or adapter ligation reaction. In some embodiments, the adapter oligonucleotides comprise a 5' overhang, a 3' overhang, or both that are complementary to one or more target polynucleotide overhangs. The complementary overhang can be a length of one or more nucleotides, including but not limited to a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more nucleotides. The complementary overhang can comprise a fixed sequence. The complementary overhang of the adapter oligonucleotide can comprise a random sequence of one or more nucleotides, such that one or more nucleotides are randomly selected from a group of two or more different nucleotides at one or more positions, wherein each of the different nucleotides selected at one or more positions is represented in an adapter having a complementary overhang that comprises a random sequence. In some embodiments, the adapter overhang is complementary to a target polynucleotide overhang generated by restriction endonuclease digestion. In some embodiments, the adapter overhang consists of adenine or thymine.
[0091] A variety of methods for cyclizing polynucleotides are available. In some embodiments, cyclization comprises an enzymatic reaction, such as using a ligase (e.g., RNA or DNA ligase). A variety of ligases are available, including but not limited to, Circligase TM(Epicentre; Madison, WI), RNA ligase, T4 RNA ligase 1 (ssRNA ligase, which acts on both DNA and RNA). In addition, T4 DNA ligase can also ligate ssDNA if a dsDNA template is not present, although this is generally a slow reaction. Other non-limiting examples of ligases include: NAD-dependent ligases, including Taq DNA ligase, Thermus filiformis DNA ligase, E. coli DNA ligase, Tth DNA ligase, Thermus scotoductus DNA ligase (I and II), thermostable ligases, Ampligase thermostable DNA ligase, VanC-type ligases, 9°N DNA ligase, Tsp DNA ligase, and novel ligases discovered through bioprospecting; ATP-dependent ligases, including T4 RNA ligase, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Pfu DNA ligase, DNA ligase I, DNA ligase III, DNA ligase IV, and novel ligases discovered through bioprospecting; and wild-type, mutant isoforms, and genetically engineered variants thereof. When self-ligation is desired, the concentrations of the polynucleotide and enzyme can be adjusted to promote the formation of intramolecular loops rather than intermolecular structures. The reaction temperature and time can also be adjusted. In some embodiments, 60°C is used to promote intramolecular ring formation. In some embodiments, the reaction time is 12-16 hours. The reaction conditions can be those specified by the manufacturer of the selected enzyme. In some embodiments, an exonuclease step can be included to digest any unligated nucleic acids after the circularization reaction. That is, the closed ring does not contain free 5' or 3' ends, so introducing a 5' or 3' exonuclease will not digest the closed ring but will digest the unligated components. This is particularly useful in multiplex systems.
[0092] Typically, the ends of a polynucleotide are joined to form a circular polynucleotide (directly or using one or more intervening adapter oligonucleotides) to produce a junction with a junction sequence. When the 5' end and 3' end of a polynucleotide are joined by an adapter polynucleotide, the term "junction" refers to the junction between the polynucleotide and the adapter (e.g., one of the 5' end junction or the 3' end junction), or refers to the junction formed between the 5' end and the 3' end of a polynucleotide and comprising the adapter polynucleotide. When the 5' end and 3' end of a polynucleotide are joined without the use of an intervening adapter (e.g., the 5' end and 3' end of a single-stranded DNA), the term "junction" can refer to the point at which the two ends are joined. A junction can be identified by the sequence of nucleotides comprising the junction (also referred to as a "junction sequence"). In some embodiments, the sample comprises polynucleotides having a mixture of ends formed by natural degradation processes (such as cell lysis, cell death, and other processes in which DNA is released from cells into their surroundings, where it can be further degraded, such as in cell-free polynucleotides, such as cell-free DNA and cell-free RNA), fragmentation as a byproduct of sample processing (such as fixation, staining, and / or storage processes), and fragmentation by methods that cut DNA that are not limited to a specific target sequence (e.g., mechanical fragmentation, such as by sonication; treatment with non-sequence-specific nucleases, such as DNase I, fragmentase). When a sample comprises polynucleotides having a mixture of ends, the probability that two polynucleotides have the same 5' end or 3' end is low, and the probability that two polynucleotides independently have both the same 5' end and 3' end is extremely low. Therefore, in some embodiments, even when two polynucleotides comprise portions having the same target sequence, junctions can be used to distinguish between different polynucleotides. When polynucleotide ends are joined without the use of intervening adapters, the joined sequence can be identified by comparison with a reference sequence. For example, when the order of the two component sequences appears to be reversed relative to a reference sequence, the point at which the reversal occurs can indicate a junction at that point. When the polynucleotide ends are joined by one or more adapter sequences, the junction can be identified by proximity to known adapter sequences, or, if the sequencing read is long enough to obtain sequence from both the 5' and 3' ends of the circularized polynucleotide, by alignment as described above. In some embodiments, the formation of a particular junction is a sufficiently rare event that it is unique among the circularized polynucleotides of a sample.
[0093] Sequencing methods
[0094] According to some embodiments of the methods for detecting tumor nucleic acids provided herein, linear and / or circularized polynucleotides (or amplification products thereof that may have been optionally enriched) are subjected to sequencing reactions to generate sequencing reads. The sequencing depth is selected based on the needs of the sample being sequenced. In some cases, the depth of sequencing is low or reads are fewer or reads per molecule are used interchangeably herein. In some cases, the depth of sequencing is high or reads are more or reads per molecule are used interchangeably herein. The sequencing reads generated by such methods can be used according to other methods disclosed herein. A variety of sequencing methods are available, particularly high-throughput sequencing methods. Examples include, but are not limited to, sequencing systems manufactured by Illumina (such as and Sequencing systems manufactured by Life Technologies (Ion etc.), Roche's 454 Life Science system, Pacific Biosciences system, Oxford Nanopore Technologies, nanoball sequencing, hybridization sequencing, polymerase cloning (POLONY) sequencing, nanogrid rolling circle sequencing (ROLONY), etc. In some embodiments, sequencing includes using and The system generates a length of about or more than about 50, 75, 100, 125, 150, 175, 200, 250, 300 or more nucleotides. In some embodiments, sequencing includes a sequencing-by-synthesis process, wherein as a single nucleotide is added to the primer extension product of growth, the nucleotide is iteratively identified. Pyrophosphate sequencing is an example of sequencing-by-synthesis, which identifies the incorporation of nucleotides by analyzing the presence of sequencing reaction byproducts, i.e., pyrophosphate, in the synthesis mixture produced. In particular, the primer / template / polymerase complex contacts a type of nucleotide. If the nucleotide is incorporated, the polymerization reaction cuts the nucleoside triphosphate between the α and β phosphates of the triphosphate chain, thereby releasing pyrophosphate. The presence of the released pyrophosphate is then identified using a chemical luciferase reporter system, which converts the pyrophosphate containing AMP into ATP, and then measures ATP with luciferase to generate a measurable light signal. When light is detected, the base has been incorporated, and when light is not detected, the base is not incorporated. After appropriate washing steps, various bases are periodically contacted with the complex to successively identify subsequent bases in the template sequence. See, e.g., U.S. Patent No. 6,210,891.
[0095] In the relevant sequencing process, primer / template / polymerase complex is fixed on substrate, and this complex is contacted with the nucleotide of labeling.The fixing of complex can be carried out by primer sequence, template sequence and / or polymerase, and can be covalent or non-covalent.For example, the fixing of complex can be realized by the connection between polymerase or primer and substrate surface.In alternative configuration, nucleotide has and does not have removable terminator group.After incorporation, label is coupled with complex, and is therefore detectable.In the case of nucleotide carrying terminator, all four kinds of different nucleotides that carry label that can be identified separately are contacted with complex.The incorporation of the nucleotide of labeling has stopped extension due to the existence of terminator, and label is added to complex, thereby allows identification of the nucleotide of incorporation.Then label and terminator are removed from the nucleotide of incorporation, and the process is repeated after appropriate washing step.In the case of unterminated nucleotide, as pyrophosphate sequencing, the nucleotide of a type of label is added to complex to determine whether it will incorporation. After removing the labeling group on nucleotide and appropriate washing steps, various nucleotides are circulated by reaction mixture in the same process.See, for example, U.S. Patent number 6,833,246, which is incorporated herein by reference in its entirety for all purposes. For example, Illumina genome analysis system (Illumina Genome Analyzer System) is based on the technology described in WO 98 / 44151, wherein DNA molecules are attached to sequencing platform (flow cell) by anchor probe binding site (also referred to as flow cell binding site), and amplified in situ on a glass slide. The solid surface of amplified DNA molecules thereon generally includes a plurality of first and second binding oligonucleotides, the first being complementary to a sequence near or positioned at one end of the target polynucleotide, and the second being complementary to a sequence near or positioned at the other end of the target polynucleotide. This arrangement allows bridge amplification, such as described in US20140121116. DNA molecules are then annealed with sequencing primers, and are sequenced in parallel one by one using a reversible terminator method. Prior to hybridization of the sequencing primer, one strand of the double-stranded bridge polynucleotide can be cleaved at a cleavage site in one of the binding oligonucleotides anchoring the double-stranded bridge, leaving one single strand unbound to the solid substrate, which can be removed by denaturation, while the other strand is bound and available for hybridization with the sequencing primer. Typically, the Illumina Genome Analyzer system utilizes a flow cell with 8 channels, generating sequencing reads with lengths of 18 to 36 bases, producing >1.3 Gbp of high-quality data per run (see www.illumina.com).
[0096] In the process of another synthetic sequence, along with the carrying out of template-dependent synthesis, the incorporation of differentially labeled nucleotides is observed in real time. Along with the incorporation of fluorescently labeled nucleotides, a single fixed primer / template / polymerase complex can be observed, thereby allowing each interpolated base to be identified in real time when interpolated. In this process, a labeling group can be attached to a part of the nucleotide that is cut during the incorporation process. For example, by attaching a labeling group to a part of the phosphate chain removed during the incorporation process, that is, α, β, γ or other terminal phosphate groups on the nucleoside polyphosphate, the labeling will not be incorporated into the nascent chain, but will produce natural DNA. The observation of a single molecule may involve optical confinement of the complex in a very small illumination volume. By optionally carrying out optical confinement of the complex, a monitored region can be created, wherein the randomly diffused nucleotide can exist for a very short time, and the nucleotide incorporated can retain a longer time in the observation volume when incorporated. This may result in a characteristic signal associated with an incorporation event, which is also characterized by a signal spectrum unique to the added base. Interacting labeling components, such as fluorescence resonance energy transfer (FRET) dye pairs, can be provided with the polymerase or other parts of the complex and the incorporated nucleotide such that the incorporation event brings the labeling components into proximity with each other and produces a characteristic signal outcome that is also unique to the incorporated base (see, e.g., U.S. Patent Nos. 6,917,726, 7,033,764, 7,052,847, 7,056,676, 7,170,050, 7,361,466, and 7,416,844; and US20070134128, each of which is herein incorporated by reference in its entirety).
[0097] In some embodiments, the nucleotides in the sample can be sequenced by ligation. This method generally uses DNA ligase to identify the target sequence, for example, as used in polymerase cloning (polony) methods and in SOLiD technology (Applied Biosystems, currently Invirogen). Typically, a set of all possible oligonucleotides of a fixed length is provided, labeled according to the sequencing position. The oligonucleotides are annealed and ligated; the preferential ligation of matching sequences by DNA ligase generates a signal corresponding to the complementary sequence at that site.
[0098] The sequencing methods disclosed herein can provide information useful for various applications, such as, for example, identifying a disease (e.g., cancer) in a subject or determining that a subject is at risk of (or developing) a disease. Sequencing can provide the sequence of a polymorphic region. Sequencing can provide the length of a polynucleotide such as DNA (e.g., cfDNA). In addition, sequencing can provide the sequence of the breakpoints or ends of DNA such as breakpoint cfDNA. Sequencing can provide the sequence of protein binding site boundaries or DNase hypersensitive site boundaries.
[0099] sample
[0100] In some embodiments of the various methods described herein, sample is from object.Object can be any animal, including but not limited to cattle, pigs, mice, rats, chickens, cats, dogs etc., and is typically mammals, such as people.Sample polynucleotides are usually separated from the cell-free sample from the object, such as tissue samples, body fluid samples or organ samples, including for example blood samples or fluid samples (such as saliva) containing nucleic acid.In some cases, sample is processed to remove cells, or in the case of no cell extraction step, polynucleotides are separated (for example, separating cell-free polynucleotides, such as cell-free DNA).Other examples of sample sources include from blood, urine, feces, nostrils, lungs, intestines, other body fluids or excreta, material derived therefrom or its combination.In some embodiments, sample is a blood sample or a part thereof (such as blood plasma or serum).Serum and blood plasma may be particularly interesting because the tumor DNA associated with the higher mortality rate of malignant cells in such tissues is relatively enriched. In some embodiments, a sample from a single individual is divided into multiple separate samples (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more separate samples) that are independently subjected to the methods of the present disclosure, such as analysis in duplicate, triplicate, quadruplicate, or more. In the case where the sample is from an object, the reference sequence can also be derived from the object, such as a consensus sequence from the analyzed sample or a polynucleotide sequence from another sample or tissue of the same object. For example, a blood sample can be analyzed for cfDNA mutations while simultaneously analyzing cellular DNA from another sample (e.g., oral or skin sample) to determine the reference sequence.
[0101] Polynucleotides can be extracted from a sample according to any suitable method. There are a variety of kits available for extracting polynucleotides, the selection of which may depend on the type of sample or the type of nucleic acid to be separated. Examples of extraction methods are provided herein, such as those described herein in relation to any of the aspects disclosed herein. In one example, the sample can be a blood sample, such as a sample collected in an EDTA tube (e.g., BD Vacutainer). Plasma can be separated from peripheral blood cells by centrifugation (e.g., at 1900xg, 10 minutes at 4°C). Plasma separation performed on a 6mL blood sample in this manner will typically produce 2.5 to 3mL of plasma. According to the manufacturer's protocol, circulating cell-free DNA can be extracted from a plasma sample, such as by using a QIAmp circulating nucleic acid kit (Qiagene). DNA can then be quantified (e.g., on an Agilent 2100 bioanalyzer with a high-sensitivity DNA kit (Agilent)). For example, the yield of circulating DNA from such a plasma sample of a healthy individual can be 1ng to 10ng per milliliter of plasma, which is significantly higher in disease (e.g., cancer) patient samples.
[0102] In some embodiments, the plurality of polynucleotides comprises cell-free polynucleotides, such as cell-free DNA (cfDNA), cell-free RNA (cfRNA), circulating tumor DNA (ctDNA), or circulating tumor RNA (ctRNA). Cell-free DNA circulates in healthy and diseased individuals. Cell-free RNA circulates in healthy and diseased individuals. cfDNA (ctDNA) from tumors is not limited to any particular cancer type, but appears to be a common finding in different malignancies. According to some measurements, the concentration of free circulating DNA in plasma is about 14-18 ng / ml in control subjects and about 180-318 ng / ml in tumor patients. Cell death by apoptosis and necrosis contributes to cell-free circulating DNA in body fluids. For example, significantly elevated levels of circulating DNA are observed in the plasma of prostate cancer patients and other prostate diseases such as benign prostatic hyperplasia and prostatitis. In addition, circulating tumor DNA is present in the body fluids of the organ where the primary tumor is located. Thus, breast cancer detection can be achieved in ductal lavage fluid; colorectal cancer detection is achieved in stool; lung cancer detection is achieved in sputum, and prostate cancer detection is detected in urine or semen. Cell-free DNA can be obtained from a variety of sources. A common source is a blood sample from a subject. However, cfDNA or other fragmented DNA can come from a variety of other sources. For example, urine and fecal samples can be sources of cfDNA, including ctDNA. Cell-free RNA can be obtained from a variety of sources.
[0103] In some embodiments, the polynucleotides are subjected to subsequent steps (e.g., cyclization and amplification) in the absence of an extraction step and / or a purification step. For example, a fluid sample can be processed to remove cells in the absence of an extraction step to produce a purified liquid sample and a cell sample, and then DNA is isolated from the purified fluid sample. There are various methods for isolating polynucleotides available, such as by precipitation, or non-specific binding to a substrate and subsequently washing the substrate to release the bound polynucleotides. When isolating polynucleotides from a sample in the absence of a cell extraction step, the polynucleotides will be primarily extracellular or "cell-free" polynucleotides. For example, cell-free polynucleotides may include cell-free DNA (also referred to as "circulating" DNA). In some embodiments, circulating DNA is circulating tumor DNA (ctDNA) from tumor cells, such as from body fluids or excreta (e.g., blood samples). Cell-free polynucleotides may include cell-free RNA (also referred to as "circulating" RNA). In some embodiments, circulating RNA is circulating tumor RNA (ctRNA) from tumor cells. Tumors can show apoptosis or necrosis, causing tumor nucleic acids to be released into the body in different forms and at different levels through various mechanisms, including in the bloodstream of the subject. Typically, ctDNA can range in size from smaller fragments (typically 70 to 200 nucleotides in length) at higher concentrations to larger fragments up to several kilobases in length at lower concentrations.
[0104] cancer
[0105] The methods for detecting tumor nucleic acids provided herein include, in some cases, staging of cancer. The staging of cancer depends on the type of cancer, wherein each type of cancer has its own classification system.
[0106] Examples of cancer staging or classification systems are described in more detail below.
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] In aspects of the methods for detecting tumor nucleic acids provided herein, examples of tumor nucleic acids that can be detected according to the methods disclosed herein include, but are not limited to, acanthoma, acinar cell carcinoma, acoustic neuroma, acral lentiginous melanoma, apical spirochete tumor, acute eosinophilic leukemia, acute lymphoblastic leukemia, acute megakaryoblastic leukemia, acute monocytic leukemia, mature acute myeloblastic leukemia, acute myeloid dendritic cell leukemia, acute myeloid leukemia, acute promyelocytic leukemia, adamantoma, adenocarcinoma, adenoid cystic carcinoma, adenoma, odontogenic adenomatous tumor, adrenocortical carcinoma, adult T-cell leukemia, aggressive NK-cell leukemia, AIDS-related cancers, AIDS-related lymphoma, alveolar sarcoma of soft tissue , ameloblastic fibroma, anal cancer, anaplastic large cell lymphoma, anaplastic thyroid cancer, angioimmunoblastic T-cell lymphoma, angiomyolipoma, angiosarcoma, appendix cancer, astrocytoma, atypical teratoid rhabdoid tumor, basal cell carcinoma, basaloid carcinoma, B-cell leukemia, B-cell lymphoma, Bellini duct carcinoma, biliary tract cancer, bladder cancer, blastoma, bone cancer, osteoma, brainstem glioma, brain tumor, breast cancer, Brenner tumor, bronchiolar carcinoma, bronchioalveolar carcinoma, brown tumor, Burkitt lymphoma, cancer of unknown primary site, carcinoid tumor, carcinoma, carcinoma in situ, penile cancer, cancer of unknown primary site, carcinosarcoma, Castleman disease, embryonal tumor of the central nervous system, cerebellar astrocytoma, cerebral astrocytoma Cervical cancer, cholangiocarcinoma, chondroma, chondrosarcoma, chordoma, choriocarcinoma, choroid plexus papilloma, chronic lymphocytic leukemia, chronic monocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disease, chronic neutrophilic leukemia, clear cell tumor, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, Degos disease, dermatofibrosarcoma protuberans, dermoid cyst, desmoplastic small round cell tumor, diffuse large B-cell lymphoma, dysembryonic neuroepithelial tumor, embryonal carcinoma, endodermal sinus tumor, endometrial cancer, endometrial carcinoma, endometrioid tumor, enteropathy-associated T-cell lymphoma, ependymoblastoma, ependymoma, epithelioid sarcoma, erythroleukemia, esophageal cancer , nasal glioma, Ewing family tumors, Ewing family sarcoma, Ewing sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, non-mammary Paget's disease, fallopian tube cancer, fetus in fetus, fibroma, fibrosarcoma, follicular lymphoma, follicular thyroid cancer, gallbladder cancer, gallbladder cancer, ganglioglioma, ganglioneuroma, gastric cancer, gastric lymphoma, gastrointestinal cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, gastrointestinal stromal tumor, germ cell tumor, germ cell tumor, gestational choriocarcinoma, gestational trophoblastic tumor, giant cell tumor of bone, glioblastoma multiforme, glioma, gliomatosis cerebri, glomus tumor, glucagonoma, gonadoblastoma, granulosa cell tumor, hairy cell leukemia, hairy cell leukemia, head and neck cancer, head and neck cancer, heart cancer,Hemangioblastoma, hemangiopericytoma, angiosarcoma, hematological malignancies, hepatocellular carcinoma, hepatosplenic T-cell lymphoma, hereditary breast-ovarian cancer syndrome, Hodgkin lymphoma, Hodgkin lymphoma, hypopharyngeal cancer, hypothalamic glioma, inflammatory breast cancer, intraocular melanoma, islet cell carcinoma, islet cell tumor, juvenile myelomonocytic leukemia, Kaposi's sarcoma, Kaposi's sarcoma, kidney cancer, Klatskin tumor, Krukenberg tumor, laryngeal cancer, laryngeal cancer, lentigo maligna, melanoma, leukemia, leukemia, lip and oral cancer, liposarcoma, lung cancer, luteoma, lymphangioma, lymphangiosarcoma, lymphoepithelioma, lymphoid leukemia, lymphoma, macroglobulinemia, malignant fibrous histiocytoma, malignant fibrous histiocytoma cell carcinoma, malignant fibrous histiocytoma of bone, malignant glioma, malignant mesothelioma, malignant peripheral nerve sheath tumor, malignant rhabdomyosarcoma, malignant Triton tumor, MALT lymphoma, mantle cell lymphoma, mast cell lymphoma, mediastinal germ cell tumor, mediastinal tumor, medullary thyroid carcinoma, medulloblastoma, medulloblastoma, medullary epithelioma, melanoma, melanoma, meningioma, Merkel cell carcinoma, mesothelioma, mesothelioma, squamous neck carcinoma with occult metastasis, metastatic urothelial carcinoma, mixed Mullerian tumor, monocytic leukemia, oral cancer, myxoma, multiple endocrine neoplasia syndrome, multiple myeloma, multiple myeloma, mycosis fungoides, mycosis fungoides, myelodysplastic disease, myelodysplastic syndrome, myeloid Leukemia, myeloid sarcoma, myeloproliferative disorders, myxoma, nasal cavity cancer, nasopharyngeal cancer, nasopharyngeal carcinoma, neoplasm, schwannoma, neuroblastoma, neuroblastoma, neurofibroma, neuroma, nodular melanoma, non-Hodgkin lymphoma, non-Hodgkin lymphoma, non-melanoma skin cancer, non-small cell lung cancer, eye tumor, oligoastrocytoma, oligodendroglioma, oncocytic adenoma, optic nerve sheath meningioma, oral cancer, oral cancer, oropharyngeal cancer, osteosarcoma, osteosarcoma, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, ovarian low malignant potential tumor, Paget's disease of the breast, Pancoast tumor, pancreatic cancer, pancreatic cancer, papillary thyroid cancer, papillomatosis, paraganglioma, paranasal sinus cancer, parathyroid cancer , penile cancer, perivascular epithelioid cell tumor, pharyngeal cancer, pheochromocytoma, moderately differentiated pineal parenchymal tumor, pinealoblastoma, pituitary cell tumor, pituitary adenoma, pituitary tumor, plasmacytoma, pleuropulmonary blastoma, polyembryoma, precursor T-cell lymphoblastic lymphoma, primary central nervous system lymphoma, primary effusion lymphoma, primary hepatocellular carcinoma, primary liver cancer, primary peritoneal cancer, primitive neuroectodermal tumor, prostate cancer, pseudomyxoma peritonei, rectal cancer, renal cell carcinoma, respiratory cancer involving the NUT gene on chromosome 15, retinoblastoma, rhabdomyosarcoma, rhabdomyosarcoma, Richter's transformation, sacrococcygeal teratoma, salivary gland cancer, sarcoma, Schwannomatosis,Sebaceous gland carcinoma, secondary tumors, seminoma, serous tumor, Sertoli-Leydig cell tumor, sex cord stromal tumor, Sezary syndrome, Signet cell carcinoma, skin cancer, small blue round cell tumor, small cell carcinoma, small cell lung cancer, small cell lymphoma, small intestinal cancer, soft tissue sarcoma, somatostatinoma, sooty wart, spinal cord tumor, spinal cord tumor, splenic marginal zone lymphoma, squamous cell carcinoma, gastric cancer, superficial spreading melanoma, supratentorial primitive neuroectodermal tumor, surface epithelial-stromal tumor, synovial sarcoma, T-cell acute lymphoblastic leukemia, T-cell large granular lymphocytic leukemia, T-cell leukemia, T-cell lymphoma, T-cell prolymphocytic leukemia, teratoma, advanced lymphoma, testicular cancer, thecoma, laryngeal cancer, thymic cancer, thymoma, thyroid cancer, transitional cell carcinoma of the renal pelvis and ureter, transitional cell carcinoma, urachal cancer, urethral cancer, genitourinary tumors, uterine sarcoma, uveal melanoma, vaginal cancer, Verner Morrison syndrome, verrucous carcinoma, optic pathway glioma, vulvar cancer, Waldenstrom's macroglobulinemia, Warthin's tumor, Wilms' tumor, and combinations thereof.
[0133] Computer system
[0134] The present disclosure provides a computer system programmed to implement a method for detecting tumor nucleic acids. Figure 3 A computer system 301 is shown that is programmed or otherwise configured to detect tumor nucleic acids. Computer system 301 can facilitate various aspects of the disclosed methods for detecting tumor nucleic acids, such as, for example, detecting tumor nucleic acids in cell-free nucleic acids using low-depth sequencing methods. Computer system 301 can be a user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.
[0135] Computer system 301 includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") 305, which can be a single-core processor or a multi-core processor, or multiple processors for parallel processing. Computer system 301 also includes memory or memory locations 310 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 315 (e.g., a hard disk), a communication interface 320 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 325 (such as cache, other memory, data storage, and / or an electronic display adapter). Memory 310, storage unit 315, interface 320, and peripheral devices 325 communicate with CPU 305 via a communication bus (solid line), such as a motherboard. Storage unit 315 can be a data storage unit (or data repository) for storing data. Computer system 301 can be operatively coupled to a computer network ("network") 330 via communication interface 320. Network 330 can be the Internet, the internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, network 330 is a telecommunications and / or data network. Network 330 can include one or more computer servers that can support distributed computing such as cloud computing. In some cases, network 330 can implement a peer-to-peer network with the aid of computer system 301, which can enable devices coupled to computer system 301 to act as clients or servers.
[0136] The CPU 305 can execute a series of machine-readable instructions embodied in a program or software. The instructions can be stored in a memory location such as the memory 310. The instructions can be directed to the CPU 305, which can then program or otherwise configure the CPU 305 to implement the methods of the present disclosure. Examples of operations performed by the CPU 305 can include fetching, decoding, executing, and writing back.
[0137] CPU 305 may be part of a circuit, such as an integrated circuit. One or more other components of system 301 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0138] Storage unit 315 can store files such as drivers, libraries, and saved programs. Storage unit 315 can store user data, such as user preferences and user programs. In some cases, computer system 301 can include one or more additional data storage units external to computer system 301, such as located on a remote server that communicates with computer system 301 via an intranet or the Internet.
[0139] Computer system 301 can communicate with one or more remote computer systems via network 330. For example, computer system 301 can communicate with a remote computer system of a user (e.g., a person who wishes to detect tumor nucleic acids). Examples of remote computer systems include personal computers (e.g., portable PCs), tablet or tablet PCs (e.g., iPad, Galaxy Tab), phones, smartphones (e.g. iPhone, Android-enabled devices, ) or personal digital assistant. Users can access computer system 301 via network 330.
[0140] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored in an electronic storage location of computer system 301, such as, for example, memory 310 or electronic storage unit 315. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by processor 305. In some cases, the code can be retrieved from storage unit 315 and stored in memory 310 for ready access by processor 305. In some cases, electronic storage unit 315 may not be included, and the machine executable instructions can be stored in memory 310.
[0141] The code may be precompiled and configured for use with a machine having a processor suitable for executing the code, or may be compiled at runtime. The code may be provided in a programming language that may be selected so that the code can be executed in a precompiled or compiled manner.
[0142] Aspects of the systems and methods (such as computer system 301) provided herein can be embodied in programming. Various aspects of the technology can be considered to be "products" or "articles" in the form of machine (or processor) executable code and / or associated data, typically carried or embodied on a type of machine-readable medium. Machine executable code can be stored on an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" type media can include any or all tangible memories of a computer, processor, etc., or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication can support loading software from one computer or processor to another, for example, from a management server or host computer to a computer platform of an application server. Therefore, another type of medium that can carry software elements includes optical waves, radio waves, and electromagnetic waves such as those used through physical interfaces between local devices, through wired and optical landline networks, and through various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered to be the medium that carries the software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0143] Thus, machine-readable media such as computer executable code may take a variety of forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks (such as any storage device in any computer, etc.), such as those that may be used to implement the databases shown in the accompanying drawings, etc. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that make up a bus within a computer system. Carrier transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punched card stock tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can participate in carrying one or more sequences of one or more instructions to a processor for execution.
[0144] The computer system 301 may include or communicate with an electronic display 335, which includes a user interface (UI) 340 for providing, for example, displaying sequencing results of tumor nucleic acid detection. Examples of UI include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0145] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by the central processing unit 305. For example, the algorithms can detect tumor nucleic acids.
[0146] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided as examples only. The present invention is not intended to be limited by the specific examples provided in the specification. Although the present invention has been described with reference to the foregoing description, the description and illustration of the embodiments herein are not intended to be interpreted as limiting. Without departing from the present invention, those skilled in the art will now appreciate many variations, changes, and replacements. In addition, it should be understood that all aspects of the present invention are not limited to the specific description, configuration, or relative proportions set forth herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein can be adopted when implementing the present invention. It is therefore contemplated that the present invention should also encompass any such substitution, modification, variation, or equivalent. The accompanying claims are intended to define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are thus covered.
[0147] Example
[0148] The following examples are given to illustrate various embodiments of the present invention and are not meant to limit the invention in any way. The examples and the methods described herein are presently representative of preferred embodiments and are exemplary and are not intended to limit the scope of the invention. Those skilled in the art will appreciate variations and other uses encompassed within the spirit of the invention as defined by the scope of the claims.
[0149] Example 1: Detection of tumor nucleic acids in cell-free samples
[0150] Methods for tumor-informed disease detection and monitoring are used to detect tumor nucleic acids in a sample. The method begins with whole genome sequencing (e.g., greater than 10 gigabases) of DNA obtained from tumor tissue to identify a list of tumor-specific somatic variants in the subject. Tumor-specific variants can be identified by comparing the sequence of the tumor DNA with sequences from healthy tissue DNA or from post-operative plasma DNA. Simultaneously or later, cell-free DNA is obtained from a sample from the subject. The cell-free DNA is circularized and amplified using rolling circle amplification to obtain concatemer copies of cell-free DNA, which contain the sequence of two or more copies of the cell-free DNA molecule. The concatemers are whole genome sequenced at a low sequencing depth (e.g., less than two reads per molecule) to determine whether the sample is positive or negative for tumor nucleic acids based on the presence or absence of tumor-specific variants. Error correction is used to determine whether the sample is positive or negative for tumor-specific variants, where a variant is called only if more than one occurrence of the variant is observed in the concatemer. The method is shown in Figure 1 、 Figure 2 and Figures 5A-5B middle.
[0151] Example 2: Detection of tumor nucleic acids through genome complexity reduction
[0152] The method for tumor-informed disease detection and monitoring is used to detect tumor nucleic acids in samples. The method begins by reducing the genomic complexity from approximately 3 billion positions to approximately 30 million positions ( Figure 4 ). Whole genome sequencing is then performed to identify a list of tumor-specific somatic variants in the subject. Cell-free DNA is also obtained from a sample from the subject. The cell-free DNA is circularized and amplified using rolling circle amplification to obtain concatemer copies of cell-free DNA, which contain the sequence of two or more copies of the cell-free DNA molecule. Whole genome sequencing is performed on the amplified cell-free DNA at a low sequencing depth to determine whether the sample is positive or negative for tumor nucleic acid based on the presence or absence of tumor-specific variants. Error correction is used to determine whether the sample is positive or negative for tumor-specific variants, wherein a variant is called only if more than one occurrence of the variant is observed in the concatemer.
[0153] Example 3: Negative selection workflow
[0154] A nucleic acid sample (such as a cell-free DNA sample with a sequence of interest and an unwanted sequence) is subjected to negative selection. The nucleic acid in the sample is subjected to a denaturation step, and then a blocking oligonucleotide with a modified 5' end and / or 3' end is annealed to the unwanted sequence in the nucleic acid sample. The blocking oligonucleotide cannot be connected and cannot be extended with a polymerase. Therefore, only the single-stranded nucleic acid in the sample will undergo cyclization. Next, the remaining linear nucleic acid bound to the blocking oligonucleotide is optionally degraded with a DNA exonuclease. The cyclized nucleic acid is amplified by rolling circle amplification to generate concatemers. The concatemers are subjected to sequencing to obtain the sequence of the nucleic acid, and variants are detected based on their presence in the sequence in the concatemers of one or more copies. The method is shown in Figure 7 middle. Figure 6 The workflow without selection is shown in .
[0155] Example 4: Positive selection workflow
[0156] A nucleic acid sample (such as a cell-free DNA sample having a sequence of interest and an unwanted sequence) is subjected to positive selection. The nucleic acids in the sample are subjected to a denaturation step and subjected to circularization. The circularized nucleic acids are amplified by rolling circle amplification using a combination of random primers and target-specific primers to generate concatemers, thereby enhancing the amplification of the sequence of interest. The concatemers are subjected to sequencing to obtain the sequence of the nucleic acid, and variants are detected based on their presence in the sequence in one or more copies of the concatemer. This method is shown in Figure 8 middle. Figure 6 The workflow without selection is shown in .
[0157] Example 5: Ultrasensitive circulating tumor DNA detection by whole-genome sequencing with single-read error correction Test
[0158] Whole-genome sequencing (WGS) of cell-free DNA (cfDNA) can be used to detect circulating tumor DNA (ctDNA) and assess tumor burden. Although the breadth of WGS can compensate for the scarcity of cfDNA, its sensitivity is limited by its error rate. This paper describes AccuScan, a highly efficient cfDNA WGS technology that achieves genome-wide error suppression of 4.2x10 at the single-read level. -7 When applied to molecular residual disease (MRD) detection, the method showed an error rate as low as 10 at a sample-level specificity of 99%. -6 AccuScan demonstrated a 90% landmark sensitivity for predicting recurrence in colorectal cancer using circulating tumor allele fraction (cTAF). It also demonstrated robust MRD performance in esophageal cancer using samples collected as early as 1 week after surgery, as well as strong prognostic value for immunotherapy monitoring in melanoma patients. Overall, AccuScan provides a highly accurate WGS solution capable of ctDNA detection in the ppm range without the need for deep sequencing or personalized reagents.
[0159] Genome-wide error suppression enables detection of ultra-low ctDNA levels
[0160] AccuScan assay workflow ( Figure 9A ) is optimized for low input cfDNA, effectively capturing double-stranded DNA, single-stranded DNA, and nicked DNA in the sample. The cfDNA is denatured and circularized by ligation, and then whole-genome amplification is performed using RCA to generate concatemer molecules containing multiple tandem copies of the original template. These concatemer products are sequenced using PE150 read lengths and aligned to the human reference genome. The sequences of each copy in the read pair are compared. Changes relative to the reference that are consistent in all copies are putative variants; while inconsistent changes may be PCR or sequencing errors and are removed. To evaluate the efficiency of AccuScan's error correction, the same cfDNA samples from healthy donors (N=3) were sequenced using both conventional WGS and AccuScan WGS. Figure 9B A comparison of the measured error rates is shown. The observed overall error rate for AccuScan was ~4.2 x 10 -7 , while the overall error rate using conventional WGS with qscore filter is 3.3x10 -4, which shows that the noise can be reduced to ∼1 / 1000 by multi-concatemer error correction.
[0161] Simulations were performed to predict the impact of error rates on ctDNA detection under different cTAFs, sequencing depths, and marker numbers ( Figures 10A-10D 、 Figures 14A-14B The presence of ctDNA was predicted using a statistical model that calculates the probability of observing an expected base call at a specific marker locus. Sensitivity was calculated as the number of positives predicted out of the total number of simulations for each combination at a nominal specificity setting of 99%.
[0162] Figure 10A The simulated sensitivity is shown for 10,000 markers with an error rate of 1x10 -4 or 5x10 7 Reducing the error rate or increasing the sequencing depth will improve the detection rate. -7 The error rate and 10x sequencing depth were 4.5x10 -5 There is a 95% detection rate (LOD95) under cTAF, but when the error rate is 1x10 -4 At 60x sequencing depth, the expected LOD95 is 1x10 -5 . Figure 10B It is shown that under all conditions consistent with the nominal specificity setting, the false positive rate with cTAF set to 0 remains below 1.2%.
[0163] The analytical sensitivity of AccuScan was measured using a healthy sample mixture ( Figure 10C cfDNA from three different healthy “test” donors was collected at a range of 1 × 10 -4 to 1x10 -6 Seven different concentrations of were independently titrated into cfDNA from different healthy "background" donors. These 21 cfDNA mixture samples were then sequenced to 60x using AccuScan with 10ng of input DNA per reaction, and the ability to detect "test" donor SNPs from the background was assessed. Of the more than 100,000 SNVs that differed in each test sample and background sample pair, a subset of 5000, 10000, or 2000 SNVs (the selected SNVs had a variant type spectrum similar to that of CRC tumors) was randomly selected for testing. Each condition was repeated 1000 times, and the MRD test was run with 99% nominal specificity (Table 28). Under the conditions of 5000, 10000, or 20000 markers, the observed specificity was >99%. Under all tested conditions, at 2.5x10 -5 The sensitivity observed at cTAF levels and above was greater than 99%.-5 At a cTAF fraction of 10 parts per million, the 5,000-marker test had an average detection rate of 77%, the 10,000-marker test showed an average sensitivity of 96%, and the test with 20,000 markers maintained 100% sensitivity across all replicates.
[0164]
[0165] AccuScan's analytical sensitivity was further demonstrated by mixing cfDNA from melanoma patients with cfDNA from healthy donors. The original cancer cfDNA sample had 1.1% cTAF as measured by ddPCR for the BRAF V600E mutation found in primary tumors. -3 to 2×10 -6 5 different expected frequencies were performed and ddPCR was performed to confirm 1x10 -3 Diluted BRAFV600E VAF. Diluted cancer samples were sequenced by AccuScan with 10 ng input per reaction. For cTAF, 1x10 -3 , 1x10 -4 and 1x10 -5 The observed detection rate was 100% for the 5x10 -6 The detection rate is 67% (2 / 3), and for 2x10 -6 The detection rate was 33% (1 / 3) ( Figure 10D AccuScan sequencing of the negative control (cfDNA from a healthy donor) was negative in both replicates.
[0166] To measure the sample level specificity of AccuScan, tumor-specific variants from 60 different cancer patients (including CRC, ESCC and melanoma) were used, and 5K, 10K and 20K equivalent variants were randomly sampled for testing MRD determination in unmatched patient plasma samples. Random sampling of 2000 unmatched variants was performed for each combination of variant count level and plasma sample. The average sample level specificity was calculated as the score characterized by negative MRD determination in the MRD test. The observed values for 5K, 10K and 20K variant count levels were similar to the nominal specificity, at 99.3%, 99.1% and 98.9%, respectively. These results show that AccuScan determination and analysis have the expected performance for patient plasma using tumor variants.
[0167] Identifying tumor-specific variants using a white blood cell (WBC)-free workflow
[0168] Tumor-informed MRD testing uses tumor-specific variants as markers to track disease. Sequencing of tumor tissue not only finds cancer mutations, but also germline SNPs and other types of variants (such as clonal hematopoiesis of uncertain potential (CHIP) variants) that will interfere with MRD analysis. A common strategy to filter non-cancer mutations is to remove variants found in matching WBCs from the same patient. However, this approach requires additional sample processing and sequencing ( Figure 11A To streamline the MRD workflow, the effect of skipping WBC sequencing and using information from post-treatment plasma samples to remove germline and CHIP variants was investigated ( Figure 11B ).
[0169] 40x sequencing of low tumor burden plasma samples can find germline and CHIP mutations at the ≥2 molecule level, while tumor-specific variants will only be found at the single molecule level. Therefore, variants with 2 or more molecules found in post-treatment plasma can be removed from tumor tissue sequencing results to obtain a list of tumor-specific mutations. To test the feasibility of this approach, the performance of tumor-WBC pairs and tumor-plasma pairs was compared using matched tumor tissue, WBC, and plasma samples collected from 20 cancer patient samples. Figure 11C-11D Figure 2 shows the number of tumor-specific variants and the spectrum of variant types found by these two different workflows. Overall, the number of mutations identified by the two methods was very similar, as was the mutation spectrum of the identified variants. AccuScan MRD analysis of plasma samples from these 20 patients (n=48) returned the same MRD call under either workflow ( Figure 11F ), and the cTAF values were strongly correlated (R2 = 0.99, Figure 11E These results suggest that it is feasible to use post-treatment plasma instead of WBC for tumor-specific variant identification.
[0170] MRD detection and prognostic value in surgical patients
[0171] Next, we evaluated the performance of AccuScan for MRD detection in postoperative gastrointestinal cancers, including 32 patients with colorectal cancer (CRC) and 17 patients with esophageal squamous cell carcinoma (ESCC).
[0172] The ESCC cohort included patients from stage I-III (18% stage I, 53% stage II, 29% stage III) who underwent curative surgery. Formalin-fixed paraffin-embedded (FFPE), WBC, preoperative plasma, and early postoperative (1 week) plasma samples were collected from all patients. Using tumor and WBC samples, we determined a median of 6768 tumor-specific variants per patient ( Figure 12A ).
[0173] In all 17 preoperative samples, ctDNA was detected, with a median cTAF of 0.27% [interquartile range (IQR): 0.13%-0.55%], with a nonsignificant trend toward higher cTAF in patients with more advanced disease ( Figure 12B In postoperative plasma samples, ctDNA was detected in 35.29% (6 / 17) of patients, with a median cTAF of 1.3x10 -4 (IQR: 1.9*10 -5 -1.1*10 -2 The follow-up time for this ESCC cohort ranged from 4.03 months to 24 months. All patients (6 / 6, 100%) whose postoperative samples were ctDNA-positive (ctDNA+) experienced disease recurrence within 2 years of surgery; 5 of 6 patients (83%) relapsed within one year. ( Figure 12C In contrast, of the 11 patients whose post-operative ctDNA samples were negative, only 3 experienced disease recurrence; all 8 disease-free patients were followed for 24 months. ctDNA testing performed 1 week after surgery had a sensitivity of 66.67% (95% CI: 29.93%-92.51%), a specificity of 100% (95% CI: 63.06%-100%), and an accuracy of 82.35% (95% CI: 56.57%-96.27%) in predicting ESCC recurrence.
[0174] CRC patients were at different clinical stages (22% stage I, 38% stage II, 34% stage III, and 6% stage IV) and underwent curative surgery. Formalin-fixed paraffin-embedded (FFPE) samples were available from all patients. For 15 patients with available WBCs, a tumor-WBC workflow was used to identify tumor-specific variants; for all other patients, a first post-operative plasma sample was used for a WBC-free workflow ( Figure 11B A median of 5820 tumor-specific variants per patient (range, 2148-265800, Figure 12A ), which corresponds to ~2 mutations / Mb.
[0175] Of these 32 patients, 26 had plasma samples collected before surgery, and 28 had plasma samples collected at the landmark (within one month after surgery). ctDNA was detected in all pre-operative cfDNA samples, with a median cTAF of 5.2 x 10 -4 (IQR: 6.6x10 -5 -2.4x10 -3 ), with a non-significant trend towards higher cTAF in more advanced patients ( Figure 12BThe median follow-up time for this CRC cohort was 24.13 months (IQR: 18.5-36). 34.4% (11 / 32) of patients had ctDNA detected in their postoperative samples. All patients who were ctDNA positive in their postoperative samples had a recurrence within 3 years after surgery ( Figure 12D ). The median disease-free survival (DFS) of the ctDNA+ patient group was 10.8 months (IQR: 5.8-12.7), of which 63.64% (7 / 11) of ctDNA+ patients relapsed within one year, and 90.91% (10 / 11) of ctDNA+ patients relapsed within two years. One ctDNA+ patient (patient 11), who was ctDNA- at the first landmark time point, converted to ctDNA+ 6 months after surgery and then relapsed at 32 months. Patients who were ctDNA-negative (ctDNA-) at all post-operative time points did not progress during the follow-up period (up to 36 months) ( Figure 12D Taken together, these results showed that for predicting CRC recurrence, the landmark had a sensitivity of 90% (95% CI: 55.5%-99.8%), a sensitivity of 100% for longitudinal monitoring, a specificity of 100% (95% CI: 80.5%-100%), and an accuracy of 96.3% (95% CI: 81%-99.9%).
[0176] When looking at early post-operative samples, MRD+ patients had shorter DFS times than MRD- patients in both ESCC (hazard ratio, HR, 8.68, 95% CI: 1.63-46.32, log-rank p = 0.0001) and CRC (HR, 45.54, 95% CI: 9.78-212, log-rank p < 0.0001). Figure 12E ).
[0177] ctDNA monitoring during immunotherapy
[0178] Advances in immune checkpoint blockade (ICB) have significantly improved the survival of patients with advanced melanoma. However, only a small proportion of patients (<20%) respond to ICB. There is an urgent need for methods to prognose and monitor patients undergoing immunotherapy. Therefore, in a pilot study of advanced melanoma (N=8), the use of AccuScan to monitor patient response to ICB was explored. A total of 22 plasma samples were collected, including 6 pre-treatment samples and 16 during treatment samples. WGS of paired tumor and WBC DNA samples identified a median of 34,323 SNVs, with an average of 90,006 tumor-specific SNVs per patient ( Figure 12A All six pre-treatment samples were ctDNA positive, with cTAF levels as low as 3.06x10 -6 ( Figure 12BOf the 16 samples collected during treatment, 10 were positive for ctDNA.
[0179] For patients 1 to 5 and 8, radiographic changes matched the changes in cTAF measured by AccuScan ( Figures 13A-13B Patients 1, 2, and 3 were ctDNA positive at the pre-treatment time point, converted to ctDNA negative after treatment, and had sustained complete responses or no disease recurrence during monitoring. Patient 8 was ctDNA positive in samples collected during treatment (no pre-treatment samples were available) and had persistently low cTAF (approximately 1×10 -5 ), and CT scans showed stable lung nodules with no evidence of disease recurrence. Patients 4 and 5 had very high cTAF levels in all plasma samples (0.4-16%), and cTAF at the second time point was approximately 2-fold or higher than that at the first time point. Both patients experienced tumor progression.
[0180] The correlation between ctDNA dynamics and radiographic information was complex for patients 6 and 7. In patient 6, AccuScan detected high cTAF (2x10 -4 ) of ctDNA and showed clearance of ctDNA 3 months after surgery and 2 cycles of ICI treatment (before the 3rd cycle of ICI treatment) ( Figure 13B ), but lymphadenopathy was detected on CT scan 4 months later (after 6 cycles of ICI treatment). The patient was switched to TKI and showed a complete response ( Figure 13A Despite the discordance between ctDNA and imaging data during treatment, the observed clearance of ctDNA was consistent with clinical outcomes. Either the MRD level failed to reflect the tumor burden, or the clinical response via imaging was delayed, and the patient was responding even before switching to TKIs.
[0181] Patient 7 had no pre-treatment sample, but the first sample during treatment was ctDNA negative, followed by four stable low-level ctDNA-positive samples ( Figure 13B However, 3 CT scans taken during treatment showed continued tumor progression, although the fourth CT scan taken after the final ctDNA test showed an excellent partial response, and the patient achieved a near CR after 1 year ( Figure 13A This is another example of discordance between imaging and ctDNA testing, where ctDNA levels remained stably low while imaging showed tumor progression. Dynamic changes in ctDNA combined with imaging data may be a better predictor of patient outcomes than imaging alone or isolated ctDNA results.
[0182] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided as examples only. Without departing from the present invention, those skilled in the art will now contemplate many variations, changes, and replacements. It should be understood that various alternatives to the embodiments described herein may be employed. The accompanying claims are intended to define the scope of the present invention, and methods and structures within the scope of these claims and their equivalents are thus covered.
Claims
1. A method for detecting tumor nucleic acids in a cell-free biological sample from a subject, the method comprising: (a) circularizing a nucleic acid derived from the cell-free biological sample to produce a circularized nucleic acid; (b) amplifying the circularized nucleic acid to generate a concatemer, the concatemer comprising at least two copies of the sequence of the circularized nucleic acid; (c) sequencing the concatemer or a derivative thereof to obtain a sequence of the concatemer, wherein the sequencing depth is no greater than 18 reads; (d) processing the sequences of the concatemer to identify at least two occurrences of a tumor-specific sequence variant of the subject; and (e) identifying the nucleic acid as having the at least one tumor-specific sequence variant upon identifying the at least two occurrences of the tumor-specific sequence variant in the sequence of the concatemer.
2. The method of claim 1, further comprising obtaining the tumor-specific sequence variant from the subject.
3. The method of claim 2, wherein obtaining the tumor-specific sequence variants comprises sequencing nucleic acid derived from a tumor of the subject.
4. The method of claim 3, wherein obtaining the tumor-specific sequence variants further comprises sequencing nucleic acid derived from healthy tissue of the subject.
5. The method of claim 4, wherein the healthy tissue has a low tumor content or no tumor content.
6. The method of claim 4 or claim 5, wherein the healthy tissue comprises post-treatment plasma.
7. The method of any one of claims 1 to 6, wherein the sequencing in (c) is to a depth of no greater than 18 reads per concatemer.
8. The method of any one of claims 1 to 6, wherein the depth of the sequencing in (c) is no greater than 18 reads per circularized nucleic acid.
9. The method of any one of claims 1 to 6, wherein the sequencing in (c) is to a depth of no greater than 1 read per concatemer.
10. The method of any one of claims 1 to 6, wherein the sequencing in (c) is to a depth of no greater than 1 read per circularized nucleic acid.
11. The method of any one of claims 3 to 10, wherein the sequencing of nucleic acid derived from the tumor or the healthy tissue of the subject is performed to a depth of greater than 20 reads.
12. The method of any one of claims 1 to 11, wherein the sequencing in (c) is to a depth of no greater than ten reads.
13. The method of any one of claims 1 to 12, wherein the sequencing in (c) is to a depth of no greater than five reads.
14. The method of any one of claims 1 to 13, wherein the sequencing in (c) is to a depth of no more than two reads.
15. The method of any one of claims 1 to 11, wherein the sequencing in (c) comprises at least 10 gigabases of sequence.
16. The method according to any one of claims 3 to 15, wherein the nucleic acid derived from the tumor is subjected to selection before sequencing.
17. The method according to any one of claims 4 to 16, wherein the nucleic acid derived from the healthy tissue is subjected to selection before sequencing.
18. The method of claim 16 or claim 17, wherein selecting comprises negative selection to remove non-target sequences from the nucleic acid.
19. The method of claim 18, wherein negative selection comprises annealing one or more blocking oligonucleotides to undesirable sequences in the nucleic acid derived from the tumor or the nucleic acid derived from the healthy tissue and circularizing the remaining single-stranded nucleic acid.
20. The method of claim 19, wherein the blocking oligonucleotide has a modified 5' end, a modified 3' end, or modified 5' and 3' ends.
21. The method of claim 19 or claim 20, further comprising degrading the resulting linear double-stranded nucleic acid using an exonuclease.
22. The method of claim 16 or claim 17, wherein selecting comprises positive selection to select target sequences from the nucleic acid.
23. The method of claim 22, wherein positive selection comprises amplifying the nucleic acid derived from the tumor or the nucleic acid derived from the healthy tissue using a plurality of random primers and a plurality of target-specific primers.
24. The method of any one of claims 1 to 23, further comprising, prior to (a), subjecting the nucleic acid derived from the cell-free biological sample to selection.
25. The method of any one of claims 1 to 23, further comprising, before (c), subjecting the nucleic acid derived from the cell-free biological sample to selection.
26. The method of claim 24 or claim 25, wherein selecting comprises negative selection to remove non-target sequences from the nucleic acid.
27. The method of claim 24 or claim 25, wherein selecting comprises positive selection to select target sequences from the nucleic acid.
28. The method according to any one of claims 1 to 27, wherein (a) comprises ligating the ends of the nucleic acid or derivative thereof to each other.
29. The method of any one of claims 1 to 27, wherein (a) comprises coupling an adaptor to the 5' end, the 3' end, or both the 5' end and the 3' end of the nucleic acid or derivative thereof.
30. The method according to any one of claims 1 to 29, wherein (b) is achieved by a polymerase having strand displacement activity.
31. The method of any one of claims 1 to 30, wherein (b) is achieved by a polymerase having 5' to 3' exonuclease activity.
32. The method of any one of claims 1 to 31, wherein the amplification is achieved by at least one primer of a plurality of random primers.
33. The method of any one of claims 1 to 31, wherein the amplification is achieved by at least one primer of a plurality of primers designed for whole genome amplification.
34. The method of any one of claims 1 to 33, wherein the nucleic acid is single-stranded.
35. The method of any one of claims 1 to 33, wherein the nucleic acid is double-stranded.
36. The method of any one of claims 1 to 35, wherein the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
37. The method of any one of claims 1 to 36, wherein (c) comprises (i) contacting the concatemer or derivative thereof with a plurality of nucleotides in the presence of a polymerase to incorporate one or more nucleotides of the plurality of nucleotides into a growing chain that is complementary to the concatemer or derivative thereof, and (ii) detecting one or more signals indicative of incorporation of the one or more nucleotides into the growing chain.
38. The method of any one of claims 1 to 37, wherein (c) comprises ligation sequencing.
39. The method of any one of claims 1 to 38, wherein the tumor-specific sequence variants comprise single nucleotide variants, fusions, insertions, deletions, or epigenetic modifications.
40. The method of any one of claims 1 to 39, wherein the cell-free biological sample is a body fluid.
41. The method of claim 40, wherein the body fluid comprises urine, saliva, blood, serum, or plasma.
42. The method of any one of claims 1 to 41, wherein the tumor is colorectal cancer, pancreatic cancer, ovarian cancer, breast cancer, prostate cancer, bladder cancer, lung cancer, skin cancer, or blood cancer.
43. The method of any one of claims 1 to 42, further comprising calling the subject as minimal residual disease (MRD) positive when the nucleic acid has the at least one tumor-specific sequence variant.
Citation Information
Patent Citations
Uniform surfaces for hybrid material substrate and methods for making and using same
US20070134128A1
System and Methods for Detecting Genetic Variation
US20140121116A1
Process for isolating nucleic acid
US5234809A
DNA purification and isolation using magnetic particles
US5705628A
Method of sequencing DNA
US6210891B1