Compositions and methods for detecting rare sequence variants
The method of rolling circle amplification and junction analysis in cyclic polynucleotides addresses the challenge of detecting rare sequence variants, improving sensitivity and accuracy in nucleic acid sequencing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-03-17
AI Technical Summary
Existing nucleic acid sequencing methods struggle to accurately detect rare sequence variants due to high error frequencies, leading to false negatives, especially in samples with low variant frequencies or complex genomic backgrounds.
A method involving rolling circle amplification (RCA) of cyclic polynucleotides, followed by sequencing and analysis of junctions between 5' and 3' ends to identify sequence variants, including the use of ligase enzymes, adapter polynucleotides, and polymerases for amplification, and optionally enriching target polynucleotides.
Enhances the detection of rare sequence variants by reducing false negatives, enabling sensitive identification of low-frequency mutations and contaminants in various samples, including clinical and environmental samples.
Smart Images

Figure 2026048677000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 375,396, filed Aug. 15, 2016, which is incorporated herein by reference in its entirety for all purposes.
Background Art
[0002] The identification of sequence variations within a population of complexes is a field that has grown exponentially, particularly with the advent of massively parallel nucleic acid sequencing. However, massively parallel sequencing is significantly limited in that the inherent error frequency of commonly used techniques is greater than the frequency of many of the actual sequence variations in the population. For example, error rates of 0.1 - 1% have been reported for standard high-throughput sequencing. When the frequency of a variant is low, such as at or below the error rate, the detection of rare sequence variants has a high false negative rate.
[0003] There is often an urgent need to detect rare sequence variants. For example, the detection of rare characteristic sequences can be used to identify and distinguish the presence of contaminants in harmful environments, such as bacterial taxa. A common method for characterizing bacterial taxa is to identify differences in highly conserved sequences, such as rRNA sequences. However, typical sequencing-based approaches to this face challenges regarding the degree of homology between the many different genomes and members in a given sample, presenting complex problems that still make the procedure cumbersome. Improvements to the procedure could enhance contamination detection in a variety of settings. For example, sterile rooms used to assemble components for satellites and other spacecraft can be investigated using this system and method to understand what microbial communities are present and to develop better decontamination and cleaning techniques to prevent the introduction of terrestrial microorganisms or their samples to other planets, or to develop methods for distinguishing data generated by terrestrial microbial contamination from data generated by putative extraterrestrial microorganisms. Food surveillance applications include periodic testing of production lines in food processing plants, inspections of slaughterhouses, examinations of kitchens and food storage areas in restaurants, hospitals, schools, and prisons, and other foodborne pathogen agencies. Water storage levels and processing plants may also be monitored.
[0004] The detection of rare variants can be crucial for the early detection of pathogenic mutations. For example, detecting cancer-related point mutations in clinical samples can improve the identification of minimal residual disease during chemotherapy and reveal the appearance of tumor cells in recurrent patients. The detection of rare point mutations is also important for assessing exposure to environmental mutagens, monitoring endogenous DNA repair, and studying the accumulation of somatic mutations in aging individuals. Furthermore, highly sensitive methods for detecting rare variants can enhance prenatal diagnosis, enabling the characterization of fetal cells in maternal blood. [Overview of the project]
[0005] In light of the foregoing, there is a need for improved methods for detecting rare sequence variants. The compositions and methods of the present disclosure address this need and also provide additional advantages. In particular, various aspects of the present disclosure result in highly sensitive detection of rare or low-frequency nucleic acid sequence variants (often referred to as mutations). This includes the identification and elucidation of low-frequency nucleic acid variants (including substitutions, insertions, and deletions) in samples that may contain low amounts of variant sequences against a background of normal sequences, as well as the identification of low-frequency mutations against a background of sequencing errors.
[0006] In one embodiment, the Disclosure provides a method for carrying out rolling circle amplification, the method comprising the steps of (a) cyclizing individual polynucleotides in a plurality of polynucleotides to form a plurality of cyclic polynucleotides using a ligase enzyme, wherein each polynucleotide in the plurality of polynucleotides has a 5' end and a 3' end before ligation; (b) degrading the ligase enzyme; and (c) amplifying the cyclic polynucleotides after degrading the ligase enzyme to produce amplified polynucleotides, wherein the polynucleotides are not purified or isolated between steps (a) and (c). In some embodiments, the method further comprises the step of degrading linear polynucleotides between steps (a) and (c). In some embodiments, the plurality of polynucleotides include single-stranded polynucleotides. In some embodiments, individual cyclic polynucleotides have characteristic junctions within the cyclic polynucleotide. In some embodiments, the cyclization step comprises the step of attaching adapter polynucleotides to the 5' end, 3' end, or both the 5' and 3' ends of the polynucleotides in the plurality of polynucleotides. In some embodiments, the amplification step includes exposing a cyclic polynucleotide to an amplification reaction mixture containing random primers. In some embodiments, the amplification step includes exposing a cyclic polynucleotide to an amplification reaction mixture containing one or more primers, each of which specifically hybridizes to a different target sequence by sequence complementarity. In some embodiments, the sample is a sample from a subject. In some embodiments, the sample is urine, feces, blood, saliva, tissue, or body fluid. In some embodiments, the sample contains tumor cells. In some embodiments, the sample is a formalin-fixed paraffin-embedded sample. In some embodiments, the method further includes diagnosing the subject based on a calling step and optionally treating it. In some embodiments, the sequence variant is a causative gene variant. In some embodiments, the sequence variant is associated with a type or stage of cancer.In some embodiments, the polynucleotides include cell-free polynucleotides. In some embodiments, the cell-free polynucleotides include circulating tumor DNA. In some embodiments, the cell-free polynucleotides include circulating tumor RNA. In some embodiments, the method further includes the step of sequencing the amplified polynucleotides to generate a plurality of sequencing reads. In some embodiments, the method further includes the step of identifying sequence differences between the sequencing reads and a reference sequence. In some embodiments, the method further includes the step of calling the sequence differences sequence variants in the plurality of polynucleotides only when (i) the sequence differences are identified on both strands of a double-stranded input molecule, (ii) the sequence differences occur in the consensus sequence of a concatemer formed by rolling circle amplification, and / or (iii) the sequence differences occur in two different molecules. In some embodiments, the sequence differences are identified as occurring in two different molecules when the sequence differences occur in at least two cyclic polynucleotides having different junctions formed between the 5' and 3' ends. In some embodiments, when reads corresponding to two different molecules have different 5' ends and different 3' ends, the sequence difference is identified as occurring in the two different molecules.
[0007] In other embodiments, the Disclosure provides a method for identifying sequence variants in a nucleic acid sample comprising a plurality of polynucleotides, each having a 5' end and a 3' end, the method comprising: (a) circularizing individual polynucleotides of the plurality of polynucleotides to form a plurality of circular polynucleotides, each of which polynucleotides has a junction between its 5' end and its 3' end; (b) amplifying the circular polynucleotides of (a) to produce amplified polynucleotides; (c) cleaving the amplified polynucleotides to produce cleaved polynucleotides, each cleaved polynucleotide having one or more cleavage points at its 5' end and / or 3' end; (d) sequencing the cleaved polynucleotides to produce a plurality of sequencing reads; (e) identifying sequence differences between the sequencing reads and a reference sequence; and (f) calling the sequence differences as sequence variants when the sequence differences occur in at least two different cleaved polynucleotides. In some embodiments, the step of calling a sequence difference as a sequence variant further includes (i) the sequence difference occurs in at least two cyclic polynucleotides having different junctions, (ii) the sequence difference is identified on both strands of a double-stranded input molecule, and / or (iii) the sequence difference occurs in the consensus sequence of a concatemer formed by amplification including rolling circle amplification. In some embodiments, the plurality of polynucleotides include single-stranded polynucleotides. In some embodiments, the cyclization step is achieved by subjecting the plurality of polynucleotides to a ligation reaction. In some embodiments, the sequence variant is a single nucleotide polymorphism. In some embodiments, the reference sequence is a consensus sequence formed by aligning sequencing reads with each other. In some embodiments, the reference sequence is a sequencing read. In some embodiments, the cyclization step includes conjugating adapter polynucleotides to the 5' end, 3' end, or both the 5' and 3' ends of the polynucleotides in the plurality of polynucleotides. In some embodiments, the amplification step is achieved by using a polymerase having strand displacement activity.In some embodiments, the amplification step includes exposing a cyclic polynucleotide to an amplification reaction mixture containing random primers. In some embodiments, the amplification step includes exposing a cyclic polynucleotide to an amplification reaction mixture containing one or more primers, each of which specifically hybridizes to a different target sequence by sequence complementarity. In some embodiments, the amplified polynucleotide is subjected to a sequencing step without enrichment. In some embodiments, the method further includes enriching one or more target polynucleotides in the amplified polynucleotide by performing an enrichment step before sequencing. In some embodiments, microbial contaminants are identified based on a calling step. In some embodiments, the sample is a sample from a subject. In some embodiments, the sample is urine, feces, blood, saliva, tissue, or body fluid. In some embodiments, the sample contains tumor cells. In some embodiments, the sample is a formalin-fixed paraffin-embedded sample. In some embodiments, the method further includes diagnosing the subject based on a calling step and optionally treating it. In some embodiments, the sequence variant is a causative gene variant. In some embodiments, sequence variants are associated with cancer type or stage. In some embodiments, the polynucleotides include cell-free polynucleotides. In some embodiments, the cell-free polynucleotides include circulating tumor DNA. Figure 14 provides a schematic diagram of an exemplary workflow.
[0008] In other embodiments, the Disclosure provides a reaction mixture for carrying out any of the methods of the Spec, the reaction mixture comprising: (a) a plurality of concatemers, each concatemer comprising a different junction formed by cyclizing individual polynucleotides having a 5' end and a 3' end; (b) a first primer comprising sequence A', the first primer specifically hybridizing to sequence A of the target sequence by sequence complementarity between sequence A and sequence A'; (c) a second primer comprising sequence B, the second primer specifically hybridizing to sequence B' present in a complementary polynucleotide comprising the complement of the target sequence by sequence complementarity between sequence B and sequence B'; and (d) a polymerase extending the first primer and the second primer to produce an amplified polynucleotide, wherein the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nt or less. In some embodiments, the first primer contains sequence C5' for sequence A', the second primer contains sequence D5' for sequence B, and neither sequence C nor sequence D hybridizes to two or more concatemers during the first amplification step of the amplification reaction.
[0009] In other embodiments, the Disclosure provides a system for detecting sequence variants, the system comprising: (a) a computer configured to receive a user request to carry out a detection reaction on a sample; and (b) an amplification system that, in response to a user request, carries out a nucleic acid amplification reaction on the sample or a portion thereof, wherein the amplification reaction comprises: (i) cyclizing individual polynucleotides in a plurality of polynucleotides to form a plurality of cyclic polynucleotides using a ligase enzyme, wherein each polynucleotide of the plurality of polynucleotides has a 5' end and a 3' end prior to ligation; and (ii) degrading the ligase enzyme; and (iii) The present invention comprises an amplification system comprising the step of amplifying a cyclic polynucleotide after degrading a ligase enzyme to produce amplified polynucleotides, wherein the polynucleotides are not purified or isolated between steps (i) and (iii); (c) a sequencing system that generates sequencing reads of the polynucleotides amplified by the amplification system, identifies sequence differences between the sequencing reads and a reference sequence, and calls sequence differences resulting in at least two cyclic polynucleotides having various junctions as sequence variants; and (d) a reporting device that transmits a report to a recipient, wherein the report includes the results of sequence variant detection. In some embodiments, the recipient is a user.
[0010] In other embodiments, the Disclosure provides a computer-readable medium containing code for performing a method for detecting sequence variants when performed by one or more processors, the method being performed: (a) receiving a customer request to perform a detection reaction on a sample; (b) performing a nucleic acid amplification reaction on the sample or a portion thereof in response to the customer request, wherein the amplification reaction (i) cyclizes individual polynucleotides in a plurality of polynucleotides to form a plurality of cyclic polynucleotides using a ligase enzyme, wherein each polynucleotide of the plurality of polynucleotides has a 5' end and a 3' end prior to ligation; and (ii) decomposing the ligase enzyme (i) a step of generating a cyclic polynucleotide after degrading a ligase enzyme in order to generate an amplified polynucleotide, wherein the polynucleotide is not purified or isolated between steps (i) and (iii), a step of performing a nucleic acid amplification reaction, (c) a step of generating a sequencing read of the polynucleotide amplified in the amplification reaction, (ii) a step of identifying sequence differences between the sequencing read and a reference sequence, and (iii) a step of calling sequence differences resulting from at least two cyclic polynucleotides having different junctions as sequence variants, and (d) a step of preparing a report including the results of the detection of sequence variants.
[0011] In other embodiments, the Disclosure provides a method for identifying sequence variants in a nucleic acid sample comprising a plurality of polynucleotides, each having a 5' end and a 3' end, the method comprising: (a) cyclizing individual polynucleotides of the plurality of polynucleotides to form a plurality of cyclic polynucleotides, each of which has a junction between its 5' end and its 3' end; (b) degrading a ligase enzyme; (c) amplifying the cyclic polynucleotides of (a) using random primers to produce amplified polynucleotides; (d) cleaving the amplified polynucleotides; (e) sequencing the cleaved polynucleotides to produce a plurality of sequencing reads; and (f) The process includes (g) identifying sequence differences between a column determination read and a reference sequence, and (i) calling a sequence difference as a sequence variant when (i) a sequence difference is identified on both strands of a double-stranded input molecule, (ii) a sequence difference occurs in a consensus sequence of a concatemer formed by rolling circle amplification, (iii) a sequence difference occurs in at least two different cleaved polynucleotides, and / or (iv) a sequence difference occurs in two different molecules, wherein the sequence difference is identified as occurring in two different molecules when the sequence difference occurs in at least two cyclic polynucleotides having different junctions formed between the 5' and 3' ends.
[0012] In other embodiments, the Disclosure provides a method for identifying sequence variants in a nucleic acid sample comprising a plurality of polynucleotides, each having a 5' end and a 3' end, the method comprising: (a) cyclizing individual polynucleotides of the plurality of polynucleotides to form a plurality of cyclic polynucleotides, each of which has a junction between its 5' end and its 3' end; (b) degrading a ligase enzyme; (c) amplifying the cyclic polynucleotides of (a) using one or more primers, each of which specifically hybridizes to a different target sequence by sequence complementarity to produce an amplified polynucleotide; and (d) cleaving the amplified polynucleotides. The process includes: (e) sequencing polynucleotides cleaved to generate multiple sequencing reads; (f) identifying sequence differences between the sequencing reads and a reference sequence; (g) calling the sequence differences as sequence variants when (i) the sequence differences are identified on both strands of a double-stranded input molecule, (ii) the sequence differences occur in the consensus sequence of a concatemer formed by rolling circle amplification, and / or (iii) the sequence differences occur in two different molecules; and identifying the sequence differences as occurring in two different molecules when the sequence differences occur in at least two cyclic polynucleotides having different junctions formed between the 5' and 3' ends.
[0013] In one embodiment, the Disclosure provides a method for identifying sequence variants in a nucleic acid sample comprising a plurality of polynucleotides, each having a 5' end and a 3' end, the method comprising: (a) circulating individual polynucleotides of the plurality of polynucleotides to form a plurality of circular polynucleotides, wherein a given circular polynucleotide of the plurality of polynucleotides has a conjugate sequence resulting from the circulation; (b) amplifying the circular polynucleotides of (a) to produce a plurality of amplified polynucleotides, wherein a first amplified polynucleotide of the plurality of polynucleotides and a second amplified polynucleotide of the plurality of polynucleotides include a conjugate sequence but have different sequences at their respective 5' and / or 3' ends; (c) sequencing the plurality of amplified polynucleotides or their amplification products to produce a plurality of sequencing reads corresponding to the first amplified polynucleotide and the second amplified polynucleotide; and (d) calling the sequence differences detected in the sequencing reads as sequence variants when the sequence differences occur in the sequencing reads corresponding to both the first amplified polynucleotide and the second amplified polynucleotide.
[0014] In some embodiments, the step of cyclizing individual polynucleotides in (a) is achieved by a ligase enzyme. In some embodiments, the ligase enzyme is degraded before (b). In some embodiments, the non-cyclized polynucleotides are degraded before (b). In some embodiments, the multiple cyclic polynucleotides are not purified or isolated before (b).
[0015] In some embodiments, the cyclicization step in (a) includes the step of attaching adapter polynucleotides to the 5' end, 3' end, or both the 5' and 3' ends of polynucleotides in a plurality of polynucleotides.
[0016] In some embodiments, the step of amplifying the cyclic polynucleotide in (b) is achieved by a polymerase having chain displacement activity. In some embodiments, the step of amplifying the cyclic polynucleotide in (b) includes rolling circle amplification (RCA). In some embodiments, the amplification step in (b) includes exposing the cyclic polynucleotide to an amplification reaction mixture containing random primers. In some embodiments, the individual random primers have sequences at their respective 5' and / or 3' ends that are different from each other. In some embodiments, the amplification step in (b) includes exposing the cyclic polynucleotide to an amplification reaction mixture containing target-specific primers. In some embodiments, the amplification step includes multiple cycles of denaturation, primer binding, and primer extension.
[0017] In some embodiments, the amplified polynucleotides are subjected to sequencing of (c) without enrichment. In some embodiments, the method further includes a step of enriching one or more target polynucleotides in the amplified polynucleotides or their amplification product by performing an enrichment step prior to sequencing of (c).
[0018] In some embodiments, the multiple polynucleotides include single-stranded polynucleotides. In some embodiments, the sequence variants are single nucleotide polymorphisms. In some embodiments, the sample is a sample from a subject. In some embodiments, the sample includes urine, feces, blood, saliva, tissue, or body fluids. In some embodiments, the sample includes tumor cells. In some embodiments, the sample includes a formalin-fixed, paraffin-embedded sample. In some embodiments, the multiple polynucleotides include cell-free polynucleotides. In some embodiments, the cell-free polynucleotides include cell-free DNA. In some embodiments, the cell-free polynucleotides include cell-free RNA. In some embodiments, the cell-free polynucleotides include circulating tumor DNA. In some embodiments, the cell-free polynucleotides include circulating tumor RNA.
[0019] In some embodiments, (d) the method includes calling the sequence differences detected in the sequencing reads as sequence variants when the sequence differences detected in the sequencing reads occur in at least 50% of the sequencing reads from the first amplified polynucleotide and at least 50% of the sequencing reads from the second amplified polynucleotide.
[0020] In one embodiment, the Disclosure provides a method for identifying sequence variants in a nucleic acid sample comprising a plurality of polynucleotides, each having a 5' end and a 3' end, the method comprising: (a) circulating individual polynucleotides of the plurality of polynucleotides to form a plurality of circular polynucleotides, wherein a given circular polynucleotide of the plurality of polynucleotides has a conjugate sequence resulting from the circulation; (b) amplifying the circular polynucleotides of (a) to produce a plurality of amplified polynucleotides; (c) cleaving the amplified polynucleotides to produce cleaved polynucleotides, wherein each cleaved polynucleotide has one or more cleavage points at its 5' end and / or 3' end; (d) sequencing the amplified products of the cleaved polynucleotides to produce a plurality of sequencing reads; and (e) calling the sequence differences detected in the sequencing reads as sequence variants when the sequence differences detected in the sequencing reads occur in the sequencing read corresponding to a first cleaved polynucleotide and the sequencing read corresponding to a second cleaved polynucleotide.
[0021] In some embodiments, the step of cyclizing individual polynucleotides in (a) is achieved by a ligase enzyme. In some embodiments, the ligase enzyme is degraded before (b). In some embodiments, the non-cyclized polynucleotides are degraded before (b). In some embodiments, the multiple cyclic polynucleotides are not purified or isolated before (b).
[0022] In some embodiments, the cyclicization step in (a) includes the step of attaching adapter polynucleotides to the 5' end, 3' end, or both the 5' and 3' ends of polynucleotides in a plurality of polynucleotides.
[0023] In some embodiments, the step of amplifying the cyclic polynucleotide in (b) is achieved by a polymerase having chain substitution activity. In some embodiments, the step of amplifying the cyclic polynucleotide in (b) includes rolling circle amplification (RCA). In some embodiments, the amplification step in (b) includes exposing the cyclic polynucleotide to an amplification reaction mixture containing random primers. In some embodiments, the individual random primers have sequences at their respective 5' and / or 3' ends that are different from each other. In some embodiments, the amplification step in (b) includes exposing the cyclic polynucleotide to an amplification reaction mixture containing target-specific primers.
[0024] In some embodiments, the amplified product of the cleaved polynucleotide is subjected to unenriched sequencing. In some embodiments, the method further includes enriching one or more target polynucleotides in the amplified product of the cleaved polynucleotide by performing an enrichment step prior to sequencing in (d). In some embodiments, the cleaving step in (c) includes subjecting the amplified polynucleotide to sonication. In some embodiments, the cleaving step in (c) includes subjecting the amplified polynucleotide to enzymatic cleavage.
[0025] In some embodiments, the plurality of polynucleotides comprises single-stranded polynucleotides. In some embodiments, the sequence variant is a single nucleotide polymorphism. In some embodiments, the sample is a sample from a subject. In some embodiments, the sample comprises urine, feces, blood, saliva, tissue, or body fluid. In some embodiments, the sample comprises tumor cells. In some embodiments, the sample comprises a formalin-fixed paraffin-embedded sample. In some embodiments, the plurality of polynucleotides comprises cell-free polynucleotides. In some embodiments, the cell-free polynucleotides comprise cell-free DNA. In some embodiments, the cell-free polynucleotides comprise cell-free RNA. In some embodiments, the cell-free polynucleotides comprise circulating tumor DNA. In some embodiments, the cell-free polynucleotides comprise circulating tumor RNA.
[0026] In some embodiments, in (e), the method comprises calling a difference in the sequence detected in the sequencing reads as a sequence variant when the difference in the sequence detected in the sequencing reads occurs in at least 50% of the sequencing reads from the first cleaved polynucleotide and at least 50% of the sequencing reads from the second cleaved polynucleotide. Incorporation by reference
[0027] All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety as if each individual publication, patent, or patent application was specifically and individually incorporated by reference herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The novel features of the invention are particularly pointed out in the appended claims. A better understanding of the features and advantages of the invention will be obtained by reference to the following detailed description that illustrates embodiments in which the principles of the invention are utilized, and to the following appended drawings. [Figure 1]Draw a schematic diagram of one embodiment of the method according to the present disclosure. The DNA strand is circularized, and target-specific primers corresponding to the gene under investigation are added together with polymerase, dNTPs, buffer, etc. so that rolling circle amplification (RCA) results in the formation of concatemers (e.g., "multimers") of template DNA (e.g., "monomers"). The concatemers are processed to synthesize the corresponding complementary strands, and then adapters are added to create a sequencing library. The resulting library is then sequenced using standard techniques and typically contains three species: nDNA ("normal" DNA) that does not contain rare sequence variants (e.g., mutations); nDNA that contains enzymatic sequencing errors, and DNA that contains multimers of "real" or actual sequence variants that were present in the sample polynucleotide prior to amplification. Effectively, the presence of multiple copies of rare mutations enables the detection and identification of sequence variants. [Figure 2] Draws a strategy similar to Figure 1 but with an adapter added to facilitate polynucleotide circularization. Figure 2 also shows the use of target-specific primers. [Figure 3] It is similar to Figure 2 except that an adapter primer is used in the amplification. [Figure 4] A - C depict three embodiments related to the formation of circularized single-stranded (ss) DNA. In the upper figure, single-stranded DNA (ssDNA) is circularized without an adapter, the central schematic depicts the use of an adapter, and the lower figure utilizes two adapter oligos (resulting in different sequences at each end) and may further include a splint oligo that hybridizes to both adapters to bring the two ends closer together. [Figure 5] Depicts an embodiment for circularizing a specific target by using a "molecular clamp" to spatially approximate the two ends of single-stranded DNA for ligation. [Figure 6] A and B depict two additional figures of the addition of adapters using the blocked ends of nucleic acids. [Figure 7]AC illustrates three different ways of stimulating a rolling circle amplification (RCA) reaction. A shows the use of target-specific primers (e.g., a specific target gene or a desired target sequence). This typically results in only the target sequence being amplified. B depicts the use of random primers to perform whole genome amplification (WGA), which typically amplifies all sample sequences, and the sequences are bioinformatically sorted during processing. C depicts the use of adapter primers when adapters are used, which also results in common non-target-specific amplification. [Figure 8] An example of double-stranded DNA circularization and amplification, in which both strands are amplified, is depicted according to one embodiment. [Figure 9] AD illustrates various schemes for achieving complementary chain synthesis for subsequent sequencing. A illustrates the use of random priming of the target chain and subsequent ligation. B illustrates the use of adapter priming of the target chain and subsequent ligation. C illustrates the use of a “loop” adapter, which has two complementary sequences that hybridize with each other to form a loop (e.g., a stem-loop structure). During ligation to the concatemer's terminals, the free ends of the loop serve as primers for the complementary chain. D shows the use of hyperbranched random primers to achieve second chain synthesis. [Figure 10] This document describes a PCR method embodying the sequencing of a cyclic polynucleotide or strand containing at least two copies of a target nucleic acid sequence using a pair of primers that are oriented away from each other when aligned within the monomer of the target sequence (also known as "back to back," e.g., oriented in two directions but not on the ends of the domain being amplified). In some embodiments, these primer sets are used after concatemers have been formed to make the amplicons higher-order multimers (e.g., dimers, trimers, etc.) of the target sequence. Optionally, the method may further include size sorting to remove amplicons smaller than dimers. [Figure 11]AD depicts an embodiment in which back-to-back (B2B) primers are used in conjunction with a "touch-up" PCR step, as amplification of short products (such as monomers) is less desirable. In this case, the primer has two domains: a first domain (gray or black arrow) that hybridizes to the target sequence, and a second domain that is a "universal primer" binding domain (bent rectangle; often also called an adapter) that does not hybridize to the original target sequence. In some embodiments, the first round of PCR is performed in a low-temperature annealing step (A) such that the gene-specific sequence is joined. Running at low temperatures yields PCR products of varying lengths, including short products (B). After several rounds, the annealing temperature is increased so that hybridization of the entire primer is favorable in both domains (C). As depicted, these are seen at the ends of the template, but the internal binding is not very stable. Therefore, short products are less favorable at high temperatures than at low temperatures in both domains, or unfavorable in single domains only (D). [Figure 12] Figures A and B illustrate two different methods for constructing sequencing libraries. Figure A illustrates an example of the Illumina® Nextera sample preparation system, which allows DNA to be simultaneously fragmented and tagged using sequencing adapters in a single step. In Figure B, concatemers are fragmented by sonication, then adapters are added to both ends (e.g., using a kit from KAPA Biosystems), and PCR amplification is performed. Other methods are also available. [Figure 13]AC provides a diagram illustrating the advantages of back-to-back (B2B) primer design compared to conventional PCR primer design. Conventional PCR primer design (left) places primers (arrows, A and B) in regions adjacent to the target sequence, which may be mutation hotspots (black stars). These are typically separate, at least 60 base pairs (bp), resulting in a typical footprint of approximately 100 bp. In this diagram, the B2B primer design (right) places the primers on one side of the target sequence. The two B2B primers face in opposite directions and may overlap to some extent (e.g., approximately 12 bp or less, approximately 10 bp or less, approximately 5 bp or less, or less). Depending on the length of the B2B primers, the total footprint in this diagram can be between 28 and 50 bp. Due to the large footprint, fragmentation events are more likely to disrupt primer binding in conventional designs, leading to loss of sequence information, whether linear fragments (A), circularized DNA (B), or amplified products (C). Furthermore, as shown in C, the B2B primer design captures a conjugate sequence (also known as a "natural barcode") that can be used to identify different polynucleotides. [Figure 14] This document illustrates a method for generating a template for detecting sequence variants, according to one embodiment (for example, an implementation that exemplifies a process using circularized polynucleotides). The DNA input is denatured into circularized ssDNA by ligation, and the non-circularized DNA is degraded by exonuclease digestion. Ligation efficiency is quantified by quantitative PCR (qPCR), comparing the amount of input DNA with the amount of circularized DNA, typically resulting in a ligation efficiency of at least approximately 80%. The circularized DNA is purified for buffer exchange and then subjected to whole-genome amplification (WGA) using randomized primers and Phi29 polymerase. The WGA product is purified and fragmented into short fragments of approximately 400 bp or less (e.g., by sonication). The on-target rate of the amplified DNA is quantified by qPCR comparing the amplified DNA with the same amount of reference genomic DNA, typically showing an average on-target rate of approximately 95% or higher. [Figure 15] AC illustrates further amplification using tailed B2B primers and performing a second "touch-up" step of PCR at high temperatures. The B2B primers contain a sequence-specific region (bold black line) and an adapter sequence (open box). At the first-step annealing temperature, the target-specific sequence is annealed to the template, producing initial monomers, and the PCR product contains tandem repeats (A). In the second amplification step at high temperatures, hybridization of both the target-specific and adapter sequences is preferable to hybridization of the target-specific sequence alone, reducing the degree to which shorter products are preferentially produced (B). Internal annealing using the target-specific sequence without supporting the entire primer rapidly increases the monomer fraction (C, left). [Figure 16] This illustrates a comparison between background noise (variant frequency) detected by a targeted sequencing method using a Q30 filter, with and without the requirement (bottom line) of sequence differences occurring on two different polynucleotides that can be counted as variants (e.g., identified by different junctions). Human genomic DNA (12878, Coriell Institute) was fragmented into 100-200 bp segments and contained 2% spike-ins compared to genomic DNA (19240, Coriell Institute) containing a known SNP (CYP2C19). True variant signals (marked peaks) did not significantly exceed the background (top, light gray plot). Background noise was reduced to approximately 0.1 by applying a validation filter (bottom, black plot). [Figure 17] The detection of sequence variants spiked at various low frequencies (2%, 0.2%, and 0.02%) in polynucleotide populations has been illustrated, and despite this, these are significantly above background when the methods of this disclosure are applied. [Figure 18] A and B illustrate the results of the analysis of ligation efficiency and on-target rate in embodiments of the present disclosure. [Figure 19] The methods according to the embodiments of this disclosure illustrate the preservation of allele frequencies and the substantial absence of bias. [Figure 20] According to one embodiment, the results for detecting sequence variants in a small input sample are illustrated. [Figure 21] This illustrates an example of high background in the detection of sequence variants obtained according to standard sequencing methods, without requiring the sequence difference to occur on two different polynucleotides. [Figure 22] A graph illustrating a comparison between the GC content distribution of a genome and the GC content distribution of sequencing results generated according to a method according to one embodiment of this disclosure (the method disclosed herein; left), wherein the sequencing results are obtained using alternative sequencing library construction kits (Rubicon, Rubicon Genomics; middle) and cell-free DNA (cfDNA) as commonly reported in the literature for 32ng (right). [Figure 23] This document provides a graph illustrating the size distribution of input DNA obtained from sequencing reads using a method according to one embodiment. [Figure 24] A graph illustrating uniform amplification across multiple targets by a random priming method according to one embodiment is provided. [Figure 25]A and B illustrate embodiments relating to the formation of polynucleotide multimers having identifiable junctions in a non-cyclic state. Polynucleotides (such as polynucleotide fragments or cell-free DNA) conjugate to form multimers having unnatural junctions that are useful for identifying independent polynucleotides according to embodiments of the present disclosure (also referred to herein as “auto-tags”). In A, polynucleotides are directly conjugated to each other by blunt-end ligation. In B, polynucleotides are conjugated by one or more intervening adapter oligonucleotides, which may also include barcode sequences. The multimers are then subjected to amplification in one of several ways, such as by random primers (whole-genome amplification), adapter primers, or one or more target-specific primers or primer pairs. [Figure 26] Figure 25 illustrates an example of variation in the process. Polynucleotides (e.g., cfDNA or other polynucleotide fragments) are end-repaired, A-tailed, and ligated with an adapter (e.g., using standard kits such as those from KAPABiosystems). The internally uracil (U)-labeled carrier DNA can be supplemented to raise the total DNA input to a desired level (e.g., approximately 20 ng or more). Detected sequence variants are indicated by a "star shape". Once ligation is complete, the carrier DNA can be degraded by adding the Uracil-Specific Excision Reagent (USER) enzyme, which is a mixture of uracil DNA glycosylase (UDG) and DNA glycosylase lyase terminal nuclease VIII. The product is purified to remove fragments of carrier DNA. The purified product is amplified (e.g., by PCR using primers targeting the adapter sequence). Any remaining carrier DNA is unlikely to be amplified by degradation and separation from the adapter at at least one end. The amplified product can be purified to remove short DNA fragments. [Figure 27]AE illustrates the variation in the process shown in Figure 25. Target-specific amplification primers contain a common 5' "tail" that acts as an adapter (gray arrow). Initial amplification (e.g., by PCR) continues for a few cycles (e.g., at least about 5, 10, or more). PCR products are annealed to other PCR products (e.g., when the annealing temperature drops in the second stage) to produce concatemers with identifiable junctions, which can then serve as primers. The second stage may involve a number of cycles (e.g., 5, 10, 15, or 20 or more) and may involve selection or changes in conditions favorable to concatemer formation and amplification. The method relating to this outline is also called “Relay Amp Seq” and may be seen in specific uses in partitioned settings (e.g., in droplets). [Figure 28] AE illustrates non-limiting examples of methods for circularizing polynucleotides. In A, a double-stranded polynucleotide (e.g., dsDNA) is denatured into a single strand and then undergoes direct circularization (e.g., autojoint ligation by CircLigase). In B, a polynucleotide (e.g., a DNA fragment) is repaired at the ends and an A tail is added (a single base extension of adenosine to the 3' end) to improve ligation efficiency, followed by denaturation into a single strand and circularization. In C, a polynucleotide is repaired at the ends, an A tail is added (if double-stranded), joined to an adapter with a thymidine (T) extension, denatured into a single strand, and then circularized. In D, the polynucleotide undergoes end repair and A-tail addition (if double-stranded), and both ends are ligated into an adapter with three elements (T extension for ligation, complementarity between adapters, and a 3' tail), the strand is denatured, and the single-stranded polynucleotide is cyclized (facilitated by complementarity between adapter sequences). In E, the double-stranded polynucleotide is denatured into a single-stranded form and cyclized in the presence of molecular clamps that bring the ends of the polynucleotide closer together to facilitate joining. [Figure 29]In particular, with respect to cyclic polynucleotides, we illustrate a workflow design that serves as an example of an amplification system for identifying sequence variants according to the method of this disclosure. [Figure 30] In particular, for linear polynucleotide inputs without a cyclization step, we illustrate a workflow design that serves as an example of an amplification system for identifying sequence variants according to the method of this disclosure. [Figure 31] A simplified diagram of an example workflow for identifying sequence variants according to the methods of this disclosure is provided. Along the “Linear Polynucleotide Analysis” (top) branch, the analysis may include digital PCR (e.g., digital droplet PCR, ddPCR), real-time PCR, enrichment by probe capture (capture seq) with analysis of conjugation sequences (self-tagged), sequencing based on inserted adapter sequences (barcoded insertions), or relay amplifier sequencing. Along the “Circular Polynucleotide Analysis” (bottom) branch, the analysis may include digital PCR (e.g., digital droplet PCR, ddPCR), real-time PCR, enrichment by probe capture (capture seq) with analysis of conjugation sequences (natural barcodes), enrichment by probe capture (capture seq) or targeted amplification (e.g., B2B amplification), and sequence analysis with a validation step to identify sequence variants as differences occurring in two different polynucleotides (e.g., polynucleotides with different junctions). [Figure 32] This is a diagram of a system according to one embodiment. [Figure 33] This example illustrates the efficiency of capture and coverage along a target region. >90% of the target bases are covered by 20x or higher, and >50% of the targeted bases have a coverage of >50x. [Figure 34] A and B illustrate an exemplary single-reaction assay workflow. [Figure 35] This illustrates sequence variant calling that uses junction information. [Figure 36A] The steps of various embodiments of the present invention are illustrated. [Figure 36B]The steps of various embodiments of the present invention are illustrated. [Figure 36C] The steps of various embodiments of the present invention are illustrated. [Figure 36D] The steps of various embodiments of the present invention are illustrated. [Figure 36E] The steps of various embodiments of the present invention are illustrated. [Figure 36F] The steps of various embodiments of the present invention are illustrated. [Figure 36G] The steps of various embodiments of the present invention are illustrated. [Figure 36H] The steps of various embodiments of the present invention are illustrated. [Figure 37A] Embodiments in which the target polynucleotide includes a single-stranded polynucleotide are illustrated. [Figure 37B] Embodiments in which the target polynucleotide includes a single-stranded polynucleotide are illustrated. [Figure 38] This indicates the complexity of the libraries of various workflows described in this specification. [Figure 39] An exemplary schematic diagram illustrates how concatemer breaks and junction sequences are used to uniquely index reads amplified from the same origin. [Figure 40] An illustrative schematic diagram illustrates how concatemer '5' and '3' sequences and junction sequences derived from random priming during amplification are used to uniquely index reads amplified from the same origin. [Figure 41A] An exemplary schematic diagram illustrates how the conjugate sequence and the 5' / 3' ends of the amplified polynucleotide are used to generate a read family. In Figure 41A, the group is counted as a "mutant" because all reads in the group show a mutant ("x") by concatemer confirmation. [Figure 41B]An exemplary schematic diagram illustrates how the conjugate sequence and the 5' / 3' ends of the amplified polynucleotide are used to generate the read family. In Figure 41B, the mutant is rejected, and the read family consensus is classified as wild-type because the majority of reads within the family do not exhibit the mutant ("x"). [Figure 41C] An illustrative schematic diagram illustrates how the conjugate sequence and the 5' / 3' ends of the amplified polynucleotide are used to generate a read family. In Figure 41C, a variant is defined as a sequence difference ("x") detected in at least two different read families. Sequence differences (circles) not detected in at least two different read families are not defined as variants. [Modes for carrying out the invention]
[0029] The implementation of some embodiments disclosed herein utilizes, unless otherwise specified, prior art in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA that is within the realm of those skilled in the art. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FM Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)).
[0030] The terms “about” or “approximately” mean within an acceptable margin of error of a particular value as determined by those skilled in the art, which in part depends on how that value is measured or determined, i.e., the limitations of the measuring system. For example, “about” may mean one or more standard deviations for the practice in the art. Alternatively, “about” may mean within a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biosystems or biological processes, the term may mean within one order of magnitude, preferably five times, more preferably twice, a given value. Where a particular value is described in this application and claims, unless otherwise specified, the term “about” should be assumed to mean within an acceptable margin of error of that particular value.
[0031] The terms “polynucleotide,” “nucleotide,” “nucleotide sequence,” “nucleic acid,” and “oligonucleotide” are used interchangeably. These refer to the macromolecular form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any unknown or known function. The following are not limited examples of polynucleotides: coding or non-coding regions of genes or gene fragments, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may also contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure, when present, may be given before or after the assembly of the polymer. The sequence of nucleotides can be blocked by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with labeling components.
[0032] Generally, the term “target polynucleotide” refers to a nucleic acid molecule or polynucleotide in a starting population of nucleic acid molecules that have a target sequence in which the presence, quantity, and / or nucleotide sequence, or one or more of these changes, are desired to be determined. Generally, the term “target sequence” refers to a nucleic acid sequence on a single strand of nucleic acid. The target sequence may be a gene, a regulatory sequence, genomic DNA, cDNA, mRNA, miRNA, rRNA, or part of RNA, including other things. The target sequence may be a target sequence from a sample, or a second target, such as a product of an amplification reaction.
[0033] Hybridization refers to a reaction in which one or more polynucleotides react to form a complex stabilized by hydrogen bonds between the bases of nucleotide residues. These hydrogen bonds can occur in a Watson-Crick base pairing, Hoogstein bond, or in a sequence-specific manner according to base complementarity. The complex can consist of two strands forming a double structure, three or more strands forming a multi-strand complex, a single self-hybridized strand, or any combination thereof. Hybridization reactions can constitute a step in larger processes such as the initiation of PCR or the enzymatic cleavage of polynucleotides by endonucleases. A second sequence complementary to a first sequence is called the "complement" of the first sequence. The term "hybridized," as applied to polynucleotides, refers to the ability of a polynucleotide to form a complex stabilized by hydrogen bonds between the bases of its nucleotide residues in a hybridization reaction.
[0034] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either through the conventional Watson-Crick method or other non-traditional types. Percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, and 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary, respectively). "Perfectly complementary" means that all adjacent residues in one nucleic acid sequence can hydrogen bond with the same number of adjacent residues in the second nucleic acid sequence. As used herein, “substantially complementary” means a degree of complementarity where there is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% of a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or two nucleic acids that hybridize under stringent conditions. Sequence identity, for purposes such as evaluating percent complementarity, may be measured by any suitable alignment algorithm, including, but not limited to, the Needleman-Wunsch algorithm (e.g., the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally with default settings), the BLAST algorithm (e.g., the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally with default settings), or the Smith-Waterman algorithm (e.g., the EMBOSS Water aligner available at www.ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally with default settings). Optimal alignment may also be evaluated using any suitable parameters of the selected algorithm, including default parameters.
[0035] Generally, "stringent conditions" for hybridization refer to conditions under which nucleic acids complementary to the target sequence predominantly hybridize to the target sequence and substantially do not hybridize to non-target sequences. Stringent conditions are usually sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which the sequence specifically hybridizes to its target sequence. Non-limiting examples of stringent conditions are described in detail in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology-Hybridization With Nucleic Acid Probes Part I, Second Chapter “Overview of principles of hybridization and the strategy of nucleic acid probe assay”, Elsevier, NY.
[0036] In one embodiment, the Disclosure provides a method for identifying sequence variants in a nucleic acid sample comprising a plurality of polynucleotides, each having a 5' end and a 3' end. In some cases, the method includes: (a) circulating individual polynucleotides of the plurality of polynucleotides to form a plurality of circular polynucleotides, wherein a given circular polynucleotide of the plurality of polynucleotides has a conjugate sequence resulting from the circulation; (b) amplifying the circulated polynucleotides of (a) to produce a plurality of amplified polynucleotides; (c) cleaving the amplified polynucleotides to produce cleaved polynucleotides, wherein each cleaved polynucleotide has one or more cleavage points at its 5' end and / or 3' end; (d) sequencing the cleaved polynucleotides and / or amplified products to produce a plurality of sequencing reads; and (e) calling the sequence differences detected in the sequencing reads as sequence variants when the sequence differences detected in the sequencing reads occur in the sequencing reads corresponding to a first cleaved polynucleotide and a second cleaved polynucleotide.
[0037] In some cases, the above method includes (a) circularizing individual polynucleotides of a plurality of polynucleotides to form a plurality of cyclic polynucleotides, wherein each polynucleotide has a junction between its 5' end and its 3' end; (b) amplifying the cyclic polynucleotides of (a) to produce amplified polynucleotides; (c) cleaving the amplified polynucleotides to produce cleaved polynucleotides, wherein each cleaved polynucleotide has one or more cleavage points at its 5' end and / or 3' end; (d) sequencing the cleaved polynucleotides to produce a plurality of sequencing reads; (e) identifying sequencing differences between the sequencing reads and a reference sequence; and (f) calling the sequence differences as sequence variants when the sequence differences occur in at least two different cleaved polynucleotides.
[0038] Generally, a junction containing a junction sequence is generated by joining the ends of polynucleotides together (either directly or using one or more intermediate adapter oligonucleotides) to form a cyclic polynucleotide. When the 5' and 3' ends of a polynucleotide are joined by an adapter polynucleotide, the term “junction” may refer to the junction between the polynucleotide and the adapter (e.g., one of the 5' or 3' junctions), or the junction between the 5' and 3' ends of a polynucleotide that is formed by and contains an adapter polynucleotide. When the 5' and 3' ends of a polynucleotide are joined without an intervening adapter (e.g., the 5' and 3' ends of single-stranded DNA), the term “junction” refers to the point where these two ends are joined. A junction may be identified by the sequence of nucleotides containing the junction (also called the “junction sequence”).
[0039] In some embodiments, the sample contains polynucleotides having a mixture of ends formed by natural degradation processes (cell lysis, cell death, and other processes in which polynucleotides such as DNA and RNA, such as cell-free DNA and cell-free RNA, are released from cells into their surrounding environment, where they may be further degraded), fragmentation as a byproduct of sample processing (fixation, staining, and / or storage procedures), and fragmentation by methods of cutting DNA into specific target sequences without restriction (e.g., mechanical fragmentation by sonication; non-sequence-specific nuclease treatment such as DNase I, fragmentase). When the sample contains polynucleotides having a mixture of ends, the probability of two polynucleotides having the same 5' or 3' end is low, and the probability of two polynucleotides independently having both the same 5' and 3' ends is even lower. Accordingly, in some embodiments, junctions may be used to distinguish different polynucleotides, in which case the two polynucleotides may even contain portions having the same target sequence. When polynucleotide ends are joined without an intervening adapter, the junction may be identified by alignment to a reference sequence. For example, if the order of two component sequences appears to be reversed with respect to the reference sequence, the point where the reversal appears to occur may indicate a junction at that point. When polynucleotide ends are joined by one or more adapter sequences, the junction may be identified by proximity to a known adapter sequence, or by the alignment described above, provided that the sequencing read is long enough to obtain both the 5' and 3' sequences of the circularized polynucleotide. In some embodiments, the formation of a particular junction is such a rare event that it becomes unique among the circularized polynucleotides of the sample.
[0040] In some embodiments, the step of cyclizing individual polynucleotides in (a) is achieved by exposing multiple polynucleotides to a ligation reaction. The ligation reaction may include a ligase enzyme. In some embodiments, the ligase enzyme is degraded before amplification in (b). Degradation of the ligase before amplification in (b) can increase the recovery rate of the amplified polynucleotides. In some embodiments, the multiple cyclized polynucleotides are not purified or isolated before (b). In some embodiments, the non-cyclized linear polynucleotides are degraded before amplification.
[0041] In some cases, the cyclicization step in (a) includes the step of attaching adapter polynucleotides to the 5' end, 3' end, or both the 5' and 3' ends of polynucleotides in a plurality of polynucleotides. When the 5' and 3' ends of a polynucleotide are attached by adapter polynucleotides, the term “junction” may refer to the junction between the polynucleotide and the adapter (e.g., one of the 5'-end junctions or the 3'-end junction), or the junction between the 5' and 3' ends of a polynucleotide that is formed by and contains an adapter polynucleotide.
[0042] Cyclized polynucleotides can be amplified, for example, after degradation by a ligase enzyme, to yield amplified polynucleotides. The step of amplifying the cyclic polynucleotide in (b) can be achieved by a polymerase having strand displacement activity. In some cases, the polymerase is Phi29 DNA polymerase. In some cases, the amplification involves rolling circle amplification (RCA). The amplified polynucleotide derived from RCA may include a linear concatemer or a polynucleotide containing two or more copies of a target sequence (e.g., a subunit sequence) from a template polynucleotide. In some embodiments, the amplification step involves exposing the cyclic polynucleotide to an amplification reaction mixture containing random primers. In some cases, the amplification step involves exposing the cyclic polynucleotide to an amplification reaction mixture containing one or more primers, each of which hybridizes specifically to a different target sequence by sequence complementarity.
[0043] Amplified polynucleotides can, in some cases, be cleaved to produce shorter cleaved polynucleotides relative to the uncleaved polynucleotide. Two or more cleaved polynucleotides starting from the same linear concatemer may have the same junction sequence but can have different 5' and / or 3' ends (e.g., by cleaving the ends).
[0044] Amplified polynucleotides can be cleaved using a variety of methods, including, but not limited to, physical, enzymatic, and chemical fragmentation. Non-limiting examples of physical fragmentation methods that can be used to fragment amplified polynucleotides include acoustic shearing, sonication, and hydrodynamic shearing. Acoustic shearing and sonication may be preferred in some cases. Non-limiting examples of enzymatic fragmentation methods that can be used to fragment amplified polynucleotides include the use of enzymes such as DNase I and other restriction endonucleases, including nonspecific nucleases and transposases. Non-limiting examples of chemical fragmentation methods that can be used to fragment amplified polynucleotides include the use of heat and divalent metal cations.
[0045] Shorter fragmented polynucleotides (also called sequenced polynucleotides) compared to unfragmented polynucleotides may be desirable to match the capabilities of the sequencing instrument used to generate sequencing reads, also known as sequence reads. For example, amplified polynucleotides may be fragmented, e.g., cleaved, to an optimal length determined by the downstream sequencing platform. Various sequencing instruments further described herein can accommodate nucleic acids of different lengths. In some cases, amplified polynucleotides are cleaved in the process of attaching adapters that are useful in the downstream sequencing platform, e.g., in flow cell attachment or sequencing primer binding. In some cases, cleaved polynucleotides are subjected to amplification to produce an amplified product of the cleaved polynucleotide before sequencing. Additional amplification may be desirable to produce a sufficient amount of polynucleotide for downstream analysis, e.g., sequencing analysis. The resulting amplified product may contain multiple copies of individual cleaved polynucleotides.
[0046] During sequencing, cleaved polynucleotides or their amplification products starting from the same amplified polynucleotide are sequenceable. Sequenced reads resulting from sequencing can be grouped into read families. A read family can contain any appropriate number of sequence reads. In some cases, a read family contains at least 5, 10, 15, 20, 25, 50, 75, or 100 sequence reads. In some cases, a group of sequence reads may not be identified as a read family if a minimum number of sequence reads are not present. For example, a read family can contain at least 2, 3, 4, 5, 7, 8, 9, or 10 sequence reads. In some cases, a read family contains at least 25 read sequences. In some cases, sequence reads may be classified into read families based on shared junctional sequences and shared sequences at the 5' and 3' ends. In some embodiments, sequence reads in a read family have the same junctional sequence. In some embodiments, sequence reads of a read family have the same sequence at their 5' and 3' ends, for example, the sequences may be identical by at least 5, 6, 7, 8, 9, or 10 bases at each of the 5' and 3' ends. In some cases, the sequences at the 5' and 3' ends are not identical across all sequence reads of a read family due to errors resulting from amplification and / or sequencing errors. For example, when compared by alignment, the sequenced reads of a read family may show overlap. In some cases, when optimally aligned, the sequenced reads of a read family show at least 75% identity. The term “percent (%) identity” refers to the percentage of identical residues shared between two sequences, e.g., a candidate sequence and a reference sequence, after aligning the sequences and introducing gaps to achieve maximum percentage identity, where necessary (i.e., gaps can be introduced into both the candidate and reference sequences for optimal alignment, and in some cases, non-homologous sequences can be ignored for comparison purposes).Alignment for the purpose of determining percent identity can be achieved in various ways using publicly available computer software, such as BLAST, ALIGN, or Megalign (DNASTAR) software. Percent identity of two sequences can be calculated by aligning a test sequence to a comparison sequence using BLAST, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to the same amino acids or nucleotides in the comparison sequence at the same position, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence. When optimally aligned, two sequencing reads in a family can exhibit at least 75% identity at any appropriate length of bases (e.g., at least 80%, 85%, 90%, or 95% identity). A first pair of sequencing reads in a read family may exhibit different percent identity than a second pair of sequencing reads in the same family. In some cases, % identity is determined for an alignment over a length of at least 50 bases (e.g., at least 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 bases). In some cases, the alignment extends over a length of approximately 25–250 bases, approximately 50–200 bases, approximately 75–175 bases, or approximately 100–150 bases. In some cases, the alignment extends over the entire length of the test sequence or comparison sequence. In some embodiments, when two sequencing reads of a read family are optimally aligned, they exhibit at least 75% identity (e.g., at least 80%, 85%, 90%, or 95% identity) over a length of at least 50 bases (e.g., at least 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 bases).
[0047] Amplified polynucleotides containing a linear concatemer of a cyclic polynucleotide template may contain multiple repeats or copies of the cyclic polynucleotide template sequence. The cleaved polynucleotides produced from the amplified polynucleotides can have various copies of the cyclic polynucleotide template sequence. A cleaved polynucleotide may have less than one copy of the repeat sequence, at least one copy, at least two copies, or at least three copies of the repeat sequence. The number of repeats in a cleaved polynucleotide may depend on the length of the repeat sequence. For example, for cleaved fragments of approximately the same size, a concatemer with relatively short repeats may produce cleaved fragments with more copies of the repeat sequence compared to a concatemer with longer repeats.
[0048] Sequence reads of cleaved polynucleotides or their amplification products may, in some cases, contain at least one copy of the repeat sequence. In some cases, the sequence reads may contain at least two copies of the repeat sequence (e.g., at least three, four, or five copies). The average number of repeat sequence copies from sequence reads in a read family may depend on the length of the polynucleotide in the nucleic acid sample.
[0049] Sequencing reads are classified into read families by first identifying the length and / or sequence of the repeating segments in the concatemer that correspond to the sequence of the cyclic polynucleotide template. In some cases, the identification of the length and / or sequence of the repeating segments involves alignment of the read to other reads or to a reference sequence. The conjugate sequence can then be identified, for example, by alignment to a reference sequence. The sequences of the 5' and 3' ends of the polynucleotide and their relative distances from the conjugate (e.g., in the bases) can be determined. Reads that have sequences shared with the same conjugate sequence at the 5' and 3' ends can be classified into read families representing sequencing reads of amplification products starting from the same cleaved polynucleotide.
[0050] Sequence differences observed in a read family can, in some cases, be called true sequence differences, contrary to the result of amplification and / or sequencing errors, by confirming that the sequence differences occur in a second read family that has the same conjugated sequence at each of its 5' and 3' ends but different sequences (e.g., at least two cleaved polynucleotides). Two read families with the same conjugated sequence but different 5' and / or 3' ends can correspond to two cleaved polynucleotides of the same linear concatemer. Observing sequence differences in two read families corresponding to two cleaved polynucleotides of the same amplified polynucleotide is one way to confirm that the sequence differences truly exist on the other cyclic polynucleotide and are not the result of amplification and / or sequencing errors in one of the cleaved polynucleotides.
[0051] In some cases, a sequence difference observed in a read family is considered a sequence difference if that difference occurs in the majority of the sequencing reads in the read family. In some cases, a sequence difference observed in a read family is considered a sequence difference if that difference occurs in at least 50% of the sequencing reads in the read family (e.g., at least 60%, 70%, 80%, 90%, or 95% of the sequencing reads). In some cases, a sequence difference observed in a read family is considered a sequence difference if that difference occurs in 100% of the sequencing reads in the read family. In some cases, a sequence difference detected in a sequencing read is called a sequence variant if that difference occurs in the majority of the sequencing reads from the first cleaved polynucleotide and the majority of the sequencing reads from the second cleaved polynucleotide. In some cases, a sequence difference detected in sequencing reads is called a sequence variant if the sequence difference occurs in at least 50% of the sequencing reads from the first cleaved polynucleotide (e.g., at least 60%, 70%, 80%, 90%, or 95% of the sequencing reads) and at least 50% of the sequencing reads from the second cleaved polynucleotide (e.g., at least 60%, 70%, 80%, 90%, or 95%). In some cases, a sequence difference detected in sequencing reads is called a sequence variant if the sequence difference occurs in 100% of the sequencing reads from the first cleaved polynucleotide and 100% of the sequencing reads from the second cleaved polynucleotide.
[0052] To confirm the presence of sequence differences identified from sequencing reads in a sample, the detection of sequence variants can be improved by using two different cleaved polynucleotides, which are two cleaved polynucleotides with the same conjugated sequence but different cleavage ends. True sequence variants are expected to be found in at least two cleaved polynucleotides starting from the same amplified polynucleotide, while errors are expected to be found in fewer than two cleaved polynucleotides. In some cases, the error rate of variant detection decreases. In some embodiments, the error rate of variant detection decreases by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%. In some cases, the sensitivity and / or specificity of variant detection increases. In some embodiments, the sensitivity of variant detection increases by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%. In some embodiments, the specificity of mutant detection increases by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%. In some cases, the rate of false positives is reduced.
[0053] In some cases, the sequence difference is further called a sequence variant when (i) the sequence difference occurs in at least two cyclic polynucleotides having different junctions, (ii) the sequence difference is identified on both strands of a double-stranded input molecule, and / or (iii) the sequence difference occurs in the consensus sequence of a concatemer formed by amplification, including rolling circle amplification (RCA). In some cases, the reference sequence is a sequencing read. In some cases, the reference sequence is a consensus sequence formed by aligning sequencing reads with each other.
[0054] In some cases, cleaved polynucleotides are subjected to sequencing without enrichment. However, if necessary, enriching one or more target polynucleotides in the amplified and / or cleaved polynucleotides can be carried out in an enrichment step before sequencing. An exemplary enrichment step may include using nucleic acids along with sequences complementary to the target sequence.
[0055] Sequence variants can be any variation from a reference sequence, as further described herein. Non-exclusive examples of sequence variants detectable using the methods described herein include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number polymorphisms (CNVs), short tandem repeats (STRs), simple repeat sequences (SSRs), variable number tandem repeats (VNTRs), amplified fragment length polymorphisms (AFLPs), retrotransposon-based insertion polymorphisms, sequence-specific amplified polymorphisms, and differences in epigenetic marks (e.g., methylation differences) detectable as sequence variants. In some cases, sequence variants are polymorphisms such as single nucleotide polymorphisms. In some cases, sequence variants are causative gene variants. In some cases, sequence variants are associated with cancer type or stage.
[0056] Nucleic acid samples may be samples from a subject. In some cases, the sample is from a human subject. In some cases, the sample contains urine, feces, blood, saliva, tissue, or bodily fluids from a subject such as a human subject. In some cases, the sample contains tumor cells. In some cases, the sample contains formalin-fixed paraffin-embedded samples. In some cases, multiple polynucleotides in the sample contain cell-free polynucleotides. Cell-free polynucleotides may contain cell-free DNA, and in some cases, circulating tumor DNA and / or circulating tumor RNA. Cell-free polynucleotides may contain cell-free RNA. In some embodiments, the method further includes the step of diagnosing and optionally treating the subject based on sequence mutation calls. In some cases, microbial contaminants in the sample are identified based on sequence mutation calls. In such cases, the sample may be from a subject, but may also be from a non-subject sample such as a soil sample or a food sample.
[0057] Multiple polynucleotides can be single-stranded. In some cases, polynucleotides are double-stranded and, for example, can be treated by denaturation to become single-stranded before cyclic formation. In some cases, double-stranded polynucleotides can be cyclically formed to obtain a double-stranded ring, and this double-stranded ring can be treated by denaturation to obtain a single-stranded ring.
[0058] In another aspect, the Disclosure provides a method for identifying sequence variants in a nucleic acid sample comprising a plurality of polynucleotides, each having a 5' end and a 3' end. In some embodiments, the method includes: (a) circulating individual polynucleotides of the plurality of polynucleotides to form a plurality of circular polynucleotides, wherein a given circular polynucleotide has a conjugate sequence resulting from the circulation; (b) amplifying the circular polynucleotides of (a) to produce a plurality of amplified polynucleotides, wherein a first amplified polynucleotide of the plurality and a second amplified polynucleotide of the plurality of polynucleotides include a conjugate sequence but have different sequences at their respective 5' and / or 3' ends; (c) sequencing the plurality of amplified polynucleotides and / or their amplification products to produce a plurality of sequencing reads corresponding to the first amplified polynucleotide and the second amplified polynucleotide; and (d) calling the sequence differences detected in the sequencing reads as sequence variants when the sequence differences occur in the sequencing reads corresponding to both the first amplified polynucleotide and the second amplified polynucleotide. In some embodiments, the step of cyclizing the individual polynucleotides in (a) is achieved by a ligase enzyme. In some embodiments, the ligase enzyme is degraded before amplification in (b). Degradation of the ligase before amplification in (b) can increase the recovery rate of the amplified polynucleotides. In some embodiments, the multiple cyclized polynucleotides are not purified or isolated before (b).
[0059] In some cases, the cyclicization step in (a) includes the step of attaching adapter polynucleotides to the 5' end, 3' end, or both the 5' and 3' ends of polynucleotides in a plurality of polynucleotides. When the 5' and 3' ends of a polynucleotide are attached by adapter polynucleotides, the term “junction” may refer to the junction between the polynucleotide and the adapter (e.g., one of the 5'-end junctions or the 3'-end junction), or the junction between the 5' and 3' ends of a polynucleotide that is formed by and contains an adapter polynucleotide.
[0060] After cyclicization, the cyclic polynucleotide is amplified. The step of amplifying the cyclic polynucleotide in (b) can be achieved by a polymerase having strand substitution activity. In some cases, the polymerase is Phi29 DNA polymerase. In some cases, the step of amplifying the cyclic polynucleotide in (b) involves rolling circle amplification (RCA). Rolling circle amplification can result in an amplified polynucleotide containing a linear concatemer of the template cyclic polynucleotide sequence. In some cases, the amplification step in (b) involves exposing the cyclic polynucleotide to the amplification reaction mixture using random primers. Random primers can hybridize nonspecifically (e.g., randomly) to the cyclic polynucleotide during amplification in (b). Random primers that can hybridize nonspecifically to the cyclic polynucleotide can hybridize to a common cyclic polynucleotide, multiple cyclic polynucleotides, or both. In some cases, two or more random primers may hybridize to the same circular polynucleotide (e.g., different regions of the same circular polynucleotide), resulting in an amplified polynucleotide having repeats of the same target sequence (or subunit sequence). Amplified polynucleotides of the same template (e.g., circular polynucleotides) may have the same junction sequence. In some embodiments, individual random primers may contain sequences at their respective different 5' and / or 3' ends, and the resulting amplified polynucleotides may contain sequences at their respective different 5' and / or 3' ends. Amplified polynucleotides of the same template may, in some cases, have different 5' and / or 3' ends, depending on where the primers were first bound and where nucleotide incorporation was completed. In some cases, the amplification step in (b) includes exposing the circular polynucleotide to an amplification reaction mixture containing target-specific primers. Target-specific primers may refer to primers that target a particular gene sequence, or in some cases, primers that target an adapter polynucleotide sequence.Amplified polynucleotides resulting from the use of target-specific primers may share a common first end (e.g., the primer) and, depending on where nucleotide incorporation is completed, may not share a second end. The amplification process may involve multiple cycles of denaturation, primer binding, and primer extension. In some cases, the amplified polynucleotide may be subjected to further amplification to produce an amplification product of the amplified polynucleotide. Additional amplification may be desirable to produce a sufficient amount of polynucleotide for downstream analysis, such as sequencing analysis. The resulting amplification product may contain multiple copies of the individual amplified polynucleotide.
[0061] The amplified polynucleotides and / or amplification products are then sequenced to obtain sequenced reads. In some cases, the amplified polynucleotides and / or amplification products are subjected to unenriched sequencing. However, if necessary, the step of enriching one or more target polynucleotides in the amplified polynucleotides and / or amplification products can be carried out in an enrichment step prior to sequencing.
[0062] Sequencing reads can be classified into read families. A read family can contain any appropriate number of sequence reads. In some cases, a read family contains at least 5, 10, 15, 20, 25, 50, 75, or 100 sequence reads. In some cases, a group of sequence reads may not be identified as a read family if a minimum number of sequence reads are not present. For example, a read family contains at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequence reads. In some cases, a read family contains at least 25 read sequences. In some embodiments, the sequence reads in a read family have the same conjugate sequence. In some embodiments, the sequence reads in a read family have the same sequence at the 5' and 3' ends, for example, the sequences may be identical at the 5' and 3' ends by at least 5, 6, 7, 8, 9, or 10 bases, respectively. In some cases, the sequences at the 5' and 3' ends are not identical in all sequence reads of a read family due to errors resulting from amplification and / or sequencing. For example, when compared by alignment, sequencing reads in a read family may show overlap. In some cases, when optimally aligned, sequencing reads in a read family will show at least 75% identity. When optimally aligned, two sequencing reads in a family can show at least 75% identity (e.g., at least 80%, 85%, 90%, or 95% identity) over any suitable length of bases. A first pair of sequencing reads in a read family may show a different % identity than a second pair of sequencing reads in the read family. In some cases, % identity is determined by alignment over a length of at least 50 bases (e.g., at least 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 bases). In some cases, the alignment can span lengths of approximately 25–250 bases, 50–200 bases, 75–175 bases, or 100–150 bases.In some cases, the alignment extends over the entire length of the test sequence or comparison sequence. In some embodiments, when two sequencing reads of a read family are optimally aligned, they exhibit at least 75% identity (e.g., at least 80%, 85%, 90%, or 95% identity) over a length of at least 50 bases (e.g., at least 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 bases).
[0063] Amplified polynucleotides containing linear concatemers of a shared cyclic polynucleotide template can produce multiple linear concatemers on multiple individual molecules, even though they are the same cyclic polynucleotide sequence. Sequence reads of the amplified polynucleotide or its amplification product may, in some cases, contain at least one copy of the repeat sequence. In some cases, the sequence reads may contain at least two copies of the repeat sequence (e.g., at least three, four, or five copies). The average number of repeat sequence copies from sequence reads in a read family may depend on the length of the polynucleotide in the nucleic acid sample. For example, a sample containing relatively long polynucleotides may yield concatemers with fewer repeats compared to a sample containing relatively short polynucleotides, given that the concatemers are of similar length.
[0064] Sequencing reads are classified into read families by first identifying the length and / or sequence of the repeating segments in the concatemer that correspond to the sequence of the cyclic polynucleotide template. In some cases, the identification of the length and / or sequence of the repeating segments involves alignment of the read to other reads or to a reference sequence. The conjugate sequence can then be identified, for example, by alignment to a reference sequence. The sequences of the 5' and 3' ends of the polynucleotide and their relative distances from the conjugate (e.g., in the bases) can be determined. Reads that have sequences shared with the same conjugate sequence at the 5' and 3' ends can be classified into read families representing sequencing reads of amplification products starting from the same amplified polynucleotide or the same molecular copy of the cyclic polynucleotide.
[0065] Sequence differences observed in a read family can, in some cases, be called true sequence differences, contrary to the result of amplification and / or sequencing errors, by confirming that the sequence differences occur in a second read family having the same but different conjugated sequences at their respective 5' and 3' ends. Two read families having the same conjugated sequence but different 5' and / or 3' ends can correspond to two amplified polynucleotides of the same circular polynucleotide. Observing sequence differences in two read families corresponding to the same circular polynucleotide can be one way to confirm that the sequence differences truly exist on the circular polynucleotide and are not the result of amplification and / or sequencing errors in one of the amplified polynucleotides.
[0066] In some cases, a sequence difference observed in a read family of sequence reads is considered a sequence difference if that sequence difference occurs in the majority of the sequencing reads in the read family. In some cases, a sequence difference observed in a read family of sequence reads is considered a sequence difference if that sequence difference occurs in at least 50% of the sequencing reads in the read family (e.g., at least 60%, 70%, 80%, 90%, or 95% of the sequencing reads). In some cases, a sequence difference observed in a read family of sequence reads is considered a sequence difference if that sequence difference occurs in 100% of the sequencing reads in the read family. In some cases, a sequence difference detected in a sequencing read is called a sequence variant if that sequence difference occurs in the majority of the sequencing reads from the first amplified polynucleotide and the majority of the sequencing reads from the second amplified polynucleotide. In some cases, a sequence difference detected in a sequencing read is called a sequence variant if the difference occurs in at least 50% of the sequencing reads from the first amplified polynucleotide (e.g., at least 60%, 70%, 80%, 90%, or 95% of the sequencing reads) and at least 50% of the sequencing reads from the second amplified polynucleotide (e.g., at least 60%, 70%, 80%, 90%, or 95% of the sequencing reads). In some cases, a sequence difference detected in a sequencing read is called a sequence variant if the difference occurs in 100% of the sequencing reads from the first amplified polynucleotide and 100% of the sequencing reads from the second amplified polynucleotide.
[0067] When performing the methods described herein, the detection of variants in samples containing multiple polynucleotides can be improved. In some cases, the error rate of variant detection decreases. In some embodiments, the error rate of variant detection decreases by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%. In some cases, the sensitivity and / or specificity of variant detection increases. In some embodiments, the sensitivity of variant detection increases by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%. In some embodiments, the specificity of variant detection increases by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50%. In some cases, the rate of false positives is reduced.
[0068] Sequence variants can be any variation from a reference sequence, as further described herein. Non-exclusive examples of sequence variants detectable using the methods described herein include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number polymorphisms (CNVs), short tandem repeats (STRs), simple repeat sequences (SSRs), variable number tandem repeats (VNTRs), amplified fragment length polymorphisms (AFLPs), retrotransposon-based insertion polymorphisms, sequence-specific amplified polymorphisms, and differences in epigenetic marks (e.g., methylation differences) that can be detected as sequence variants. In some cases, sequence variants are polymorphisms such as single nucleotide polymorphisms.
[0069] Nucleic acid samples may be samples from subjects. In some cases, samples are from human subjects. In some cases, samples may contain urine, feces, blood, saliva, tissue, or bodily fluids from subjects such as human subjects. In some cases, samples may contain tumor cells. In some cases, samples may contain formalin-fixed, paraffin-embedded samples. In some cases, multiple polynucleotides in a sample may contain cell-free polynucleotides. Cell-free polynucleotides may contain cell-free DNA, and in some cases, circulating tumor DNA. Cell-free polynucleotides may contain cell-free RNA, and in some cases, circulating tumor RNA.
[0070] As previously mentioned, multiple polynucleotides can be single-stranded. In some cases, polynucleotides are double-stranded and, for example, can be treated by denaturation to obtain a single-stranded form before cyclization. In some cases, double-stranded polynucleotides are cyclized to obtain a double-stranded ring, and this double-stranded ring can be treated by denaturation to obtain a single-stranded ring.
[0071] In other embodiments, the Disclosure provides a method for performing rolling circle amplification in a nucleic acid sample containing multiple polynucleotides. In some embodiments, each polynucleotide of the multiple polynucleotides has a 5' end and a 3' end, and the method includes: (a) circularizing the individual polynucleotides of the multiple polynucleotides to form multiple circular polynucleotides using a ligase enzyme, wherein each polynucleotide has a junction between its 5' end and its 3' end; (b) degrading the ligase enzyme; and (c) amplifying the circular polynucleotides of (a) after degradation of the ligase enzyme, wherein the polynucleotides are not purified or isolated between steps (a) and (c). In some embodiments, the method includes the additional steps of (d) sequencing the amplified polynucleotides to generate multiple sequencing reads; (e) identifying sequence differences between the sequencing reads and a reference sequence; and (f) calling sequence differences resulting from at least two circular polynucleotides having different junctions as sequence variants. In some embodiments, the method includes: identifying sequence differences between a sequencing read and a reference sequence; and calling sequence differences resulting from at least two cyclic polynucleotides having different junctions as sequence variants, wherein (a) the sequencing read corresponds to amplification products of at least two cyclic polynucleotides, and (b) each of the at least two cyclic polynucleotides includes different junctions formed by ligating the 5' and 3' ends of each polynucleotide.
[0072] In other embodiments, the Disclosure provides a method for performing rolling circle amplification in a nucleic acid sample containing a plurality of polynucleotides. In some embodiments, each of the plurality of polynucleotides has a 5' end and a 3' end, and the method includes: (a) cyclizing the individual polynucleotides of the plurality of polynucleotides using a ligase enzyme to form a plurality of cyclic polynucleotides, wherein each polynucleotide has a junction between its 5' end and its 3' end; (b) degrading the ligase enzyme; (c) amplifying the cyclic polynucleotides of (a) after degrading the ligase enzyme to produce amplified polynucleotides, wherein the polynucleotides are not purified or isolated between steps (a) and (c); and (d) cleaving the amplified polynucleotides to produce cleaved polynucleotides, wherein each cleaved polynucleotide has one or more cleavage points at its 5' end and / or 3' end. In some embodiments, the method includes (e) sequencing cleaved polynucleotides to generate a plurality of sequencing reads; (f) identifying sequence differences between the sequencing reads and a reference sequence; and (g) calling the sequence differences as sequence variants when the sequence differences occur in at least two different cleaved polynucleotides. (c) Ligase degradation before amplification can increase the recovery rate of amplified polynucleotides.
[0073] In some embodiments, the method comprises the steps of: identifying a sequence difference between a sequencing read and a reference sequence; and calling a sequence difference arising in at least two cyclic polynucleotides having different junctions as a sequence variant, wherein (a) the sequencing read corresponds to an amplification product of at least two cyclic polynucleotides; and (b) each of the at least two cyclic polynucleotides has different junctions formed by ligating the 5' and 3' ends of each polynucleotide. In some embodiments, the method further comprises: (i) calling a sequence difference as a sequence variant if the sequence difference arises in at least two cyclic polynucleotides having different junctions; and (ii) when the sequence difference is identified on both strands of a double-stranded input molecule, and / or (iii) when the sequence difference arises in a consensus sequence of a concatemer formed by amplification including rolling circle amplification.
[0074] Generally, the term “sequence variant” refers to any variation of a sequence relative to one or more reference sequences. Typically, sequence variants occur less frequently than the reference sequences of a given population of known individuals whose reference sequences are known. For example, a particular bacterial genus may have a consensus reference sequence for the 16S rRNA gene, but individual species within that genus may have one or more sequence variants in the gene (or part thereof) that helps identify that species within a bacterial population. As a further example, sequences from multiple individuals of the same species (or multiple sequencing reads from the same individual), when optimally aligned, may produce a consensus sequence, and sequence variants against that consensus may be used to identify mutants in a population that indicate dangerous contamination. Generally, “consensus sequence” refers to a nucleotide sequence that reflects the most common selection of bases at each position in the sequence when a set of related nucleic acids are subjected to thorough mathematical analysis and / or sequence analysis, such as optimal sequence alignment according to one of various sequence alignment algorithms. Various alignment algorithms are available, some of which are described herein. In some embodiments, the reference sequence is a single known reference sequence, such as the genome sequence of one individual. In some embodiments, the reference sequence is a consensus sequence formed by aligning multiple known sequences, such as the genome sequences of multiple individuals that serve as a reference population, or multiple sequencing reads of polynucleotides from the same individual. In some embodiments, the reference sequence is a consensus sequence formed by optimally aligning sequences from a sample under analysis so that sequence variants represent variations with respect to corresponding sequences in the same sample. In some embodiments, sequence variants occur at low frequencies in a population (also called “rare” sequence variants). For example, sequence variants may occur at frequencies of about 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or less. In some embodiments, sequence variants occur at a frequency of approximately 0.1% or less.
[0075] Sequence variants can be any variation relative to a reference sequence. Sequence variants may consist of changes, insertions, or deletions of a single nucleotide or multiple nucleotides (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides). When a sequence variant contains two or more nucleotide differences, the different nucleotides may be adjacent or discontinuous. Non-limiting examples of the types of sequence variants include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number polymorphisms (CNVs), short tandem repeats (STRs), simple repeat sequences (SSRs), variable number tandem repeats (VNTRs), amplified fragment length polymorphisms (AFLPs), retrotransposon-based insertion polymorphisms, sequence-specific amplified polymorphisms, and differences in epigenetic marks (e.g., methylation differences) that can be detected as sequence variants.
[0076] Nucleic acid samples that may be exposed to the methods described herein may originate from any suitable source. In some embodiments, the samples used are environmental samples. Environmental samples may be from any environmental source, e.g., naturally occurring or artificial atmospheres, water pipes, soil, or any other desired sample. In some embodiments, environmental samples may be obtained, for example, from airborne pathogen recovery systems, subsurface sediments, groundwater, ancient water deep underground, grassland plant root soil interfaces, coastal water, and sewage treatment plants.
[0077] The polynucleotides from a sample may be any of the following, but are not limited to: DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), fragments of any of these, or any combination of two or more of these. In some embodiments, the sample contains DNA. In some embodiments, the sample contains genomic DNA. In some embodiments, the sample contains mitochondrial DNA, chloroplast DNA, plasmid DNA, bacterial artificial chromosomes, yeast artificial chromosomes, oligonucleotide tags, or combinations thereof. In some embodiments, the sample contains DNA produced by amplification, such as by primer extension reactions using appropriate combinations of primers and DNA polymerase, including but not limited to polymerase chain reaction (PCR), reverse transcription, and combinations thereof. When the template for a primer extension reaction is RNA, the product of reverse transcription is called complementary DNA (cDNA). Primers useful for primer extension reactions may include sequences specific to one or more targets, random sequences, partially random sequences, and combinations thereof. Generally, a sample contains polynucleotides, which may or may not contain the target polynucleotide. Polynucleotides can be single-stranded, double-stranded, or a combination thereof. In some embodiments, the polynucleotides exposed to the methods of the present disclosure are single-stranded polynucleotides, which may or may not be in the presence of double-stranded polynucleotides. In some embodiments, the polynucleotides are single-stranded DNA. Single-stranded DNA (ssDNA) may be ssDNA isolated in single-stranded form, or DNA isolated in double-stranded form and then made single-stranded for the purposes of one or more steps of the methods of the present disclosure.
[0078] In some embodiments, polynucleotides are subjected to subsequent processes (e.g., cyclization and amplification) that do not involve extraction and / or purification steps. For example, a fluid sample may be treated to remove cells without an extraction step to produce a purified fluid sample and a cell sample, and DNA may then be isolated from the purified fluid sample. Various procedures are available for the isolation of polynucleotides, such as precipitation or nonspecific binding, or washing of the substrate to release the subsequent bound polynucleotides. When polynucleotides are isolated from a sample without a cell extraction step, the polynucleotides are mostly extracellular polynucleotides or "cell-free" polynucleotides such as cell-free DNA and cell-free RNA, which may correspond to dead or damaged cells. Such cell identity may be used to characterize the cells or populations of cells from which they originate, such as tumor cells (e.g., during cancer detection), fetal cells (e.g., during prenatal diagnosis), cells from transplanted tissue (e.g., during early detection of transplant failure), or members of a microbial community.
[0079] When a sample is treated to extract polynucleotides, such as cells, from the sample, various extraction methods are available. For example, nucleic acids can be purified by organic extraction with phenol, phenol / chloroform / isoamyl alcohol, or similar formulations including TRIzol and TriReagent. Other, not limited, examples of extraction techniques include: (1) organic extraction with or without the use of an automated nucleic acid extractor, e.g., the Model 341 DNA Extractor available from Applied Biosystems (Foster City, Calif.), using an organic reagent such as phenol / chloroform (Ausubel et al., 1993), followed by ethanol precipitation; (2) stationary phase adsorption (U.S. Patent No. 5,234,809; Walsh et al., 1991); and (3) salt-induced nucleic acid precipitation, such as precipitation methods typically referred to as "salting-out" methods (Miller et al., (1988)). Another example of nucleic acid isolation and / or purification involves the use of magnetic particles, to which nucleic acids can bind specifically or nonspecifically, and then a magnet can be used to isolate and wash the beads, and then the nucleic acids can be eluted from the beads (see, for example, U.S. Patent No. 5,705,628). In some embodiments, the above isolation method may begin with an enzymatic digestion step, e.g., digestion with proteinase K or other proteases, which helps to remove unwanted proteins from the sample (see, for example, U.S. Patent No. 7,001,724). If desired, an RNase inhibitor can be added to the lysis buffer. For specific cell or sample types, it may be desirable to add a protein denaturation / digestion step to the protocol. The purification method may aim to isolate DNA, RNA, or both. If both DNA and RNA are isolated together during or after the extraction procedure, further steps can be used to purify one or both separately from the other. For example, purification by size, sequence, or other physical or chemical properties can also generate fractions of the extracted nucleic acids.In addition to the initial nucleic acid isolation step, nucleic acid purification can be performed after any step in the disclosed method to remove excess or unwanted reagents, reactants, or products. Various methods for determining the amount and / or purity of nucleic acid in a sample are available, such as by the absorbance of labels (e.g., absorbance of light at 260 nm and 280 nm and their ratios) and detection of fluorescent dyes and inserts such as SYBR green, SYBR blue, DAPI, propidium iodide, Hoechst stain, SYBR gold, and ethidium bromide.
[0080] If desired, polynucleotides from a sample may be fragmented before further processing. Fragmentation may be achieved by any of a variety of methods, including chemical, enzymatic, and mechanical fragmentation. In some embodiments, fragments have an average or median length of about 10 to about 1,000 nucleotides, such as 10-800, 10-500, 50-500, 90-200, or 50-150 nucleotides. In some embodiments, fragments have an average or median length of about 100, 200, 300, 500, 600, 800, 1,000, or 1,500 or less nucleotides. In some embodiments, fragments extend to about 90-200 nucleotides and / or have an average length of about 150 nucleotides. In some embodiments, fragmentation is carried out mechanically, including exposure of the sample polynucleotides to acoustic sonication. In some embodiments, fragmentation includes the step of treating the sample polynucleotides with one or more enzymes under conditions suitable for one or more enzymes to produce double-stranded nucleic acid breaks. Enzymes useful for generating polynucleotide fragments include sequence-specific and sequence-independent nucleases. Non-restrictive examples of nucleases include DNase I, fragmentases, restriction endonucleases, their variants, and combinations thereof. For example, digestion with DNase I can induce random double-strand breaks in DNA in the absence of Mg++ and in the presence of Mn++. In some embodiments, fragmentation involves the step of treating a sample polynucleotide with one or more restriction endonucleases. Fragmentation can produce fragments having 5' overhangs, 3' overhangs, blunt ends, or combinations thereof. In some embodiments, such as when fragmentation involves the use of one or more restriction endonucleases, the cleavage of the sample polynucleotide leaves overhangs with predictable sequences. Fragmented polynucleotides may also be subjected to a step of selecting fragment size via standard methods such as column purification or isolation from agarose gels.
[0081] According to several embodiments, polynucleotides in multiple polynucleotides from a sample are cyclized. Cyclization may involve joining the 5' end of a polynucleotide to the 3' end of the same polynucleotide, to the 3' end of another polynucleotide, or to the 3' end of a polynucleotide from a different source (e.g., an artificial polynucleotide such as an oligonucleotide adapter). In some embodiments, the 5' end of a polynucleotide is joined to the 3' end of the same polynucleotide (also called "self-joining"). In some embodiments, the conditions of the cyclization reaction are selected to support the self-joining of polynucleotides within a specific length range to produce a population of cyclized polynucleotides of a particular average length. For example, the cyclization reaction conditions may be selected to support the self-joining of polynucleotides shorter than approximately 5000, 2500, 1000, 750, 500, 400, 300, 200, 150, 100, 50 nucleotides or less in length. In some embodiments, fragments having lengths of 50–5000 nucleotides, 100–2500 nucleotides, or 150–500 nucleotides are preferred so that the average length of the cyclized polynucleotide falls within each respective range. In some embodiments, more than 80% of the cyclized fragments are nucleotides with lengths of 50–500, such as nucleotides with lengths of 50–200. Reaction conditions that can be optimized include the length of time allocated to the conjugation reaction, the concentrations of various reagents, and the concentrations of the polynucleotides being conjugated. In some embodiments, the cyclization reaction maintains the distribution of fragment lengths present in the sample before cyclization. For example, the fragment lengths in the sample before cyclization, and one or more of the mean, median, mode, and standard deviation of the cyclized polynucleotides, are within the range of 75%, 80%, 85%, 90%, or 95% or more of each other.
[0082] In some cases, rather than preferentially forming a self-jointing cyclic product, one or more adapter oligonucleotides are used so that the 5' and 3' ends of a polynucleotide in the sample are joined by one or more intervening adapter oligonucleotides to form a cyclic polynucleotide. For example, the 5' end of a polynucleotide can join to the 3' end of an adapter, and the 5' end of the same adapter can join to the 3' end of the same polynucleotide. The adapter oligonucleotide includes any oligonucleotide having a sequence (at least a portion of which is known) that can join to the sample polynucleotide. The adapter oligonucleotide may include DNA, RNA, nucleotide analogs, non-standard nucleotides, labeled nucleotides, modified nucleotides, or combinations thereof. The adapter oligonucleotide may be single-stranded, double-stranded, or partially double-stranded. Generally, a partially double-stranded adapter includes one or more single-stranded regions and one or more double-stranded regions. A double-stranded adapter comprises two distinct oligonucleotides (also referred to as an "oligonucleotide duplex") hybridized with each other, and the hybridization may include one or more blunt ends, one or more 3' overhangs, one or more 5' overhangs, one or more bulges derived from mismatched and / or unpaired nucleotides, or any combination thereof. When the two hybridized regions of the adapter are separated from each other by the unhybridized region, a "bubble" structure is produced as a result. Various types of adapters, such as adapters with different sequences, can be used in combination. Different adapters can be conjugated to a sample polynucleotide in a sequential reaction or simultaneously. In some embodiments, identical adapters are added to both ends of a target polynucleotide. For example, a first and a second adapter can be added to the same reaction. The adapter is manipulable before conjugation to the sample polynucleotide. For example, terminal phosphates may be added or removed.
[0083] When an adapter oligonucleotide is used, the adapter oligonucleotide may include, but is not limited to, one or more of various sequence factors, including: one or more amplification primers that anneal to a sequence or its complement; one or more sequencing primers that anneal to a sequence or its complement; one or more barcode sequences; one or more common sequences shared among a number of different adapters or subsets of different adapters; one or more restriction enzyme recognition sites; one or more overhangs complementary to one or more target polynucleotide overhangs; one or more probe binding sites (for example, for joining to a sequencing platform such as a flow cell for large-scale parallel sequencing, such as the flow cell developed by Illumina, Inc.); one or more random or near-random sequences (for example, one or more nucleotides randomly selected from a set of two or more different nucleotides at one or more positions, each of which is selected at one or more positions represented in a pool of adapters containing random sequences); and combinations thereof. In some cases, the adapter may be used to purify such rings containing the adapter by using beads (especially magnetic beads, for ease of handling) coated with oligonucleotides containing a sequence complementary to the adapter, the adapter can "capture" the ring closed with the appropriate adapter by hybridization, wash away the ring that does not contain the adapter or unligated components, and then release the captured ring from the beads. In addition, in some cases, the complex of the hybridized capture probe and the target ring can be used directly to generate concatemers by direct rolling circle amplification (RCA), etc. In some embodiments, the ring adapter can also be used as a sequencing primer. Two or more sequence factors may not be adjacent to each other (e.g., separated by one or more nucleotides), may be adjacent to each other, partially overlap, or completely overlap. For example, an amplification primer that anneals a sequence can also serve as a sequencing primer that anneals a sequence.The sequence element can be positioned at or near the 3' end, at or near the 5' end, or inside the adapter oligonucleotide. The sequence element may be of any suitable length, such as less than or equal to about 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides. The adapter oligonucleotide may have any suitable length, at least sufficient to accommodate one or more sequence elements that comprise them. In some embodiments, the adapter is of about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, 200 or more nucleotides. In some embodiments, the adapter oligonucleotide is of a length in the range of about 12 to 40 nucleotides, such as about 15 to 35 nucleotides.
[0084] In some embodiments, an adapter oligonucleotide conjugated to a fragmented polynucleotide from one sample includes one or more sequences common to all adapter oligonucleotides and a barcode specific to the adapter conjugated to the polynucleotide of that particular sample, the barcode sequence being usable to distinguish a polynucleotide starting from one sample or adapter conjugation reaction from a polynucleotide starting from another sample or adapter conjugation reaction. In some embodiments, the adapter oligonucleotide includes one or more target polynucleotide overhangs and complementary 5' overhangs, 3' overhangs, or both. The complementary overhangs are nucleotides of length 1 or more, but are not limited to nucleotides of length 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more. The complementary overhangs may include fixed sequences. The complementary overhang of the adapter oligonucleotide may contain a random sequence of one or more nucleotides, such that one or more nucleotides are randomly selected from a set of two or more different nucleotides at one or more positions, and each of the different nucleotides is selected at one or more positions represented in the pool of adapters with complementary overhangs containing random sequences. In some embodiments, the adapter overhang is complementary to the overhang of the target polynucleotide produced by restriction endonuclease digestion. In some embodiments, the adapter overhang consists of adenine or thymine.
[0085] Various methods are available for cyclizing polynucleotides. Figure 28 AE illustrates a non-limiting example of a method for cyclizing polynucleotides. In some embodiments, cyclization involves enzymatic reactions, such as the use of ligases (e.g., RNA or DNA ligases). Various ligases are available, but are not limited to Circligase® (Epicentre; Madison, WI), RNA ligases, and T4 RNA ligase 1 (ssRNA ligase, which acts on both DNA and RNA). In addition, if a dsDNA template is not present, T4 DNA ligase can further ligate ssDNA, although this is usually a slow reaction. Other non-exclusive examples of ligases include NAD-dependent ligases, including Taq DNA ligase, Thermus filiformis DNA ligase, Escherichia coli DNA ligase, Tth DNA ligase, Thermus scotoductus DNA ligase (I and II), heat-stable ligases, Ampligase heat-stable DNA ligase, VanC-type ligase, 9°N DNA ligase, Tsp DNA ligase, and novel ligases discovered by bioprospecting; ATP-dependent ligases, including T4 RNA ligase, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Pfu DNA ligase, DNA ligase 1, DNA ligase III, DNA ligase IV, and novel ligases discovered by bioprospecting; and their wild types, mutant isoforms, and genetically engineered variants. If self-joining is desired, the concentrations of the polynucleotide and enzyme can be adjusted to promote intramolecular ring formation rather than intermolecular structure. The reaction temperature and time can be adjusted accordingly. In some embodiments, 60°C is used to promote intramolecular ring formation. In some embodiments, the reaction time is 12–16 hours. The reaction conditions may be those specified by the manufacturer of the selected enzyme.In some embodiments, an exonuclease step may be included to digest any unligated nucleic acids after the cyclization reaction. That is, the closed ring does not contain a free 5' or 3' end, and therefore, the introduction of a 5' or 3' exonuclease does not digest the closed ring but digests the unligated component. This may find specific applications in multiple systems.
[0086] Generally, a junction containing a junction sequence is generated by joining the ends of polynucleotides together (either directly or using one or more intermediate adapter oligonucleotides) to form a cyclic polynucleotide. When the 5' and 3' ends of a polynucleotide are joined by an adapter polynucleotide, the term “junction” may refer to the junction between the polynucleotide and the adapter (e.g., one of the 5' or 3' junctions), or the junction between the 5' and 3' ends of a polynucleotide that is formed by and contains an adapter polynucleotide. When the 5' and 3' ends of a polynucleotide are joined without an intervening adapter (e.g., the 5' and 3' ends of single-stranded DNA), the term “junction” refers to the point where these two ends are joined. A junction may be identified by the sequence of nucleotides containing the junction (also called the “junction sequence”). In some embodiments, the sample contains polynucleotides having a mixture of ends formed by natural degradation processes (cell lysis, cell death, and other processes in which DNA is released from the cell into its surrounding environment (where it may be further degraded in cell-free polynucleotides, cell-free DNA, and cell-free RNA, etc.)), fragmentation as a byproduct of sample processing (fixation, staining, and / or storage procedures), and fragmentation by methods of cutting DNA into specific target sequences without restriction (e.g., mechanical fragmentation such as sonication; non-sequence-specific nuclease treatment such as DNase I, fragmentase). When the sample contains polynucleotides having a mixture of ends, it is unlikely that two polynucleotides will have the same 5' or 3' end, and significantly unlikely that two polynucleotides will independently have both the same 5' and 3' ends. Accordingly, in some embodiments, junctions may be used to distinguish different polynucleotides, in which case the two polynucleotides may even contain portions having the same target sequence. When polynucleotide ends are joined without an intervening adapter, the joining sequence can sometimes be identified by alignment to a reference sequence.For example, if the order of two component sequences appears to be reversed compared to the reference sequence, the point where the reversal appears may indicate a junction. When polynucleotide ends are joined by one or more adapter sequences, the junction may be identified by proximity to a known adapter sequence, or by alignment as described above, provided that the sequencing read is long enough to obtain both the 5' and 3' sequences of the circularized polynucleotide. In some embodiments, the formation of a particular junction is such a rare event that it becomes unique among the circularized polynucleotides of the sample.
[0087] Figure 4A illustrates three non-limiting examples of methods for cyclizing polynucleotides. In the top (Figure 4A), the polynucleotide is cyclized without an adapter; the middle scheme (Figure 4B) depicts the use of an adapter; and the bottom scheme (Figure 4C) utilizes two adapters. When two adapters are used, one is attached to the 5' end of the polynucleotide, and the other adapter can be attached to the 3' end of the same polynucleotide. In some embodiments, adapter ligation may involve the use of two different adapters, along with a "splint" nucleic acid complementary to the two adapters that facilitate ligation. Branched or "Y" adapters may also be used. When two adapters are used, polynucleotides with the same adapter at both ends may be removed in a later step by self-annealing. Figures 1-3 illustrate embodiments of the method of this disclosure in which polynucleotides are cyclized without an adapter (Figure 1) and with adapters (Figures 2 and 3). Circularized polynucleotides with adapters (Figures 2 and 3) can be amplified by rolling circle amplification (RCA) using target-specific primers (Figure 2) or primers that hybridize to the adapter sequence (Figure 3).
[0088] Figures 6A and 6 illustrate a method that provides a further non-limiting example of cyclic ligation of polynucleotides, such as single-stranded DNA. The adapter can be asymmetrically added to either the 5' or 3' end of the polynucleotide. As shown in Figure 6A, the single-stranded DNA (ssDNA) may have a free hydroxyl group at its 3' end, and the adapter may have a closed 3' end such that, in the presence of a ligase, the preferred reaction junctions the 3' end of the ssDNA to the 5' end of the adapter. In this embodiment, it may be useful to use an agent such as polyethylene glycol (PEG) to drive intermolecular ligation of a single ssDNA fragment and a single adapter prior to intramolecular ligation to form the ring. The process can also be carried out in the reverse order of the ends (closed 3', free 5', etc.). Once linear ligation is performed, the ligated portion can be enzymatically treated to remove the occlusion, such as by the use of a kinase or other suitable enzyme or by chemical action. Once the occlusion is removed, the intramolecular reaction can form a circularized polynucleotide with the addition of a cyclizing enzyme such as CircLigase. As shown in Figure 6B, a double-stranded structure can be formed by using a double-stranded adapter with one strand occluded at the 5' or 3' end, which generates a double-stranded fragment with a nick during ligation. The two strands can then be separated, the occlusion can be removed, and the single-stranded fragment can be circularized to form a circularized polynucleotide. In some cases, as shown in Figure 8, the circularization of double-stranded DNA (dsDNA) results in a circularized double-stranded ring. The double-stranded ring can be denatured to allow primer binding and amplification of both strands.
[0089] In some embodiments, molecular clamps are used to merge two ends of a polynucleotide (e.g., single-stranded DNA) to enhance the rate of intramolecular cyclization. An example diagram of one such process is provided in Figure 5. This can be done with or without an adapter. The use of molecular clamps can be particularly useful when the average polynucleotide fragment is greater than approximately 100 nucleotides in length. In some embodiments, the molecular clamp probe comprises three domains: a first domain, an intervening domain, and a second domain. The first and second domains hybridize to the corresponding sequences in the target polynucleotide by sequence complementarity. The intervening domain of the molecular clamp probe may not hybridize significantly with the target sequence. Hybridization of the clamp using the target polynucleotide can bring the two ends of the target sequence closer together, thereby promoting intramolecular cyclization of the target sequence in the presence of a cyclizing enzyme. This is even more useful when, in some embodiments, the molecular clamp can similarly serve as an amplification primer.
[0090] After cyclization, the ligation enzyme is removed from the reaction product using a proteolytic step. In some embodiments, proteolytic includes a procedure to remove or degrade the ligase used in the cyclization reaction. In some embodiments, the procedure to degrade the ligase includes a procedure using a protease such as proteinase K. The procedure with proteinase K may follow the manufacturer's protocol or a standard protocol (e.g., provided in Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012)). In some embodiments, proteolytic includes a procedure using a low pH or acidic solution or buffer. In some embodiments, proteolytic includes heating the reaction, for example, heating the reaction to above 55°C, above 60°C, above 65°C, above 70°C, or higher. In some embodiments, the linear polynucleotide is degraded after cyclization. In some embodiments, the linear polynucleotide is degraded using an exonuclease. In some embodiments, the exonuclease includes lambda exonuclease. In some embodiments, the exonuclease comprises a RecJf nuclease. In some embodiments, the exonuclease is selected from at least one of ExoI, ExoIII, ExoV, ExoVII, and ExoT.
[0091] Circulation may directly involve sequencing of the circularized polynucleotide. Alternatively, sequencing may precede one or more amplification reactions. Generally, "amplification" refers to the process by which one or more copies are made of a target polynucleotide or a portion thereof. Various methods are available for amplifying polynucleotides (e.g., DNA and / or RNA). Amplification may be linear, exponential, or involve both linear and exponential phases in a multiphase amplification process. Amplification methods may involve temperature fluctuations, such as thermal denaturation steps, or they may be isothermal processes that do not require thermal denaturation. Polymerase chain reaction (PCR) uses multiple cycles of denaturation, annealing of primer pairs to the opposite strand, and primer extension to exponentially increase the copy number of the target sequence. Denaturation of annealed nucleic acid strands may be achieved by heating, increasing local metal ion concentration (e.g., U.S. Patent No. 6,277,605), ultrasonic radiation (e.g., WO / 2000 / 049176), application of voltage (e.g., U.S. Patents No. 5,527,670, 6,033,850, 5,939,291, and 6,333,157), and application of an electromagnetic field in combination with primers bound to a magnetically reactive material (e.g., U.S. Patent No. 5,545,540). In a variation called RT-PCR, reverse transcriptase (RT) is used to create complementary DNA (cDNA) from RNA, which is then amplified by PCR to produce multiple copies of DNA (e.g., U.S. Patents No. 5,322,770 and 5,310,652).One example of an isothermal amplification method is chain substitution amplification, commonly known as SDA, which uses a cycle of annealing the primer sequence to the opposite chain of the target sequence, primer extension in the presence of dNTPs to produce a double hemiphosphorothioated primer extension product, endonuclease-mediated nicking of the hemimodified restriction endonuclease recognition site, and polymerase-mediated primer extension from the 3' end of the nicking to replace the existing chain and generate the next round of chain substitution (e.g., U.S. Patents 5,270,184 and 5,455,166). Thermophilic SDA (tSDA) uses thermophilic endonucleases and polymerases at higher temperatures in essentially the same manner (European Patent No. 0,684 315). Other amplification methods include rolling circle amplification (RCA) (e.g., Lizardi, “Rolling Circle Replication Reporter Systems,” U.S. Patent No. 5,854,033); helicase-dependent amplification (HDA) (e.g., Kong et al., “Helicase Dependent Amplification Nucleic Acids,” U.S. Patent Application Publication 2004-0058378 A1); and loop-mediated isothermal amplification (LAMP) (e.g., Notomi et al., “Process for Synthesizing Nucleic Acid,” U.S. Patent No. 6,410,278). In some cases, isothermal amplification utilizes RNA polymerase transcription from a promoter sequence, which may be incorporated into oligonucleotide primers.Transcription-based amplification methods include nucleic acid sequence-based amplification called NASBA (e.g., U.S. Patent No. 5,130,238); methods that rely on the use of RNA replication enzymes, commonly called Qβ replicases, to amplify the probe molecule itself (e.g., Lizardi, P. et al. (1988) BioTechnol. 6, 1197-1202); auto-persistent sequence replication methods (e.g., Guatelli, J. et al. (1990) Proc. Natl. Acad. Sci. USA 87, 1874-1878; Landgren (1993) Trends in Genetics 9, 199-202; and HELEN H. LEE et al., NUCLEIC ACID AMPLIFICATION TECHNOLOGIES (1997)); and methods for generating additional transcription templates (e.g., U.S. Patents No. 5,480,784 and 5,399,491). Further methods for isothermal nucleic acid amplification include the use of primers containing non-standard nucleotides (e.g., uracil or RNA nucleotides) in combination with an enzyme that cleaves the nucleic acid with non-standard nucleotides (e.g., DNA glycosylase or RNaseH) to expose the binding sites of additional primers (e.g., U.S. Patents 6,251,639, 6,946,251, and 7,824,890). The isothermal amplification process may be linear or exponential.
[0092] In some embodiments, amplification involves rolling circle amplification (RCA). A typical RCA reaction mixture comprises one or more primers, a polymerase, and a dNTP to produce a concatemer. Typically, the polymerase in an RCA reaction is a polymerase with chain displacement activity. A variety of such polymerases are available, and non-limiting examples include the exonuclease minus DNA polymerase I large (Klenow) fragment, Phi29 DNA polymerase, and Taq DNA polymerase. Generally, a concatemer is a polynucleotide amplification product containing two or more copies of the target sequence from a template polynucleotide (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, or more copies of the target sequence; in some embodiments, about two or more copies). The amplification primer may be of any suitable length, such as approximately or at least approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, or more nucleotides, and some or all of them may be complementary to the corresponding target sequence that the primer hybridizes (e.g., approximately or at least approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides). Figure 7AC depicts three non-limiting examples of suitable primers. Figure 7A shows use without an adapter and target-specific primer, which can be used to detect the presence or absence of sequence variants within a particular target sequence. In some embodiments, multiple target-specific primers for multiple targets are used in the same reaction. For example, target-specific primers for approximately or at least approximately 10, 50, 100, 150, 200, 250, 300, 400, 500, 1000, 2500, 5000, 10000, 15000, or more different target sequences may be used in a single amplification reaction to amplify the corresponding number of target sequences (if any) in parallel. Multiple target sequences may correspond to different parts of the same gene, different genes, or non-gene sequences.When multiple primers target multiple target sequences within a single gene, the primers may be positioned at regular intervals along the gene sequence to cover all or all or specified portions of the target gene (e.g., approximately or at least approximately 50 nucleotides, every 50–150 nucleotides, or every 50–100 nucleotides). Figure 7C illustrates the use of a primer that hybridizes to an adapter sequence (which may, in some cases, be the adapter oligonucleotide itself).
[0093] Figure 7B illustrates an example of amplification using random primers. Generally, random primers contain one or more random or nearly random sequences (e.g., one or more nucleotides randomly selected from two or more different sets of nucleotides at one or more positions, each of which is represented by a pool of adapters containing random sequences). Thus, polynucleotides (e.g., all or substantially all cyclic polynucleotides) can be amplified in a sequence-nonspecific manner. Such procedures are sometimes called “whole-genome amplification” (WGA); however, typical WGA protocols (including the cyclicization step) do not efficiently amplify short polynucleotides such as the polynucleotide fragments intended by this disclosure. For a further illustrative discussion of WGA procedures, see, for example, Li et al (2006) J Mol. Diagn. 8(1):22-30.
[0094] When cyclic polynucleotides are amplified before sequencing, the amplified product may be subjected to sequencing directly without enrichment, or it may be subjected to one or more enrichment steps thereafter. Enrichment may include purifying one or more reaction components, such as retaining the amplified product or removing one or more reagents. For example, the amplified product may be purified by hybridization to multiple probes bound to a substrate, followed by the release of the captured polynucleotide by a washing step or the like. Alternatively, the amplified product may be labeled with a member of a binding pair, followed by binding to another member of the binding pair bound to a substrate, and washing to release the amplified product. Possible substrates include, but are not limited to, improved or functionalized glass, plastics (including acrylic resins, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon®, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glass, plastics, optical fiber bundles, and various other polymers. In some embodiments, the substrate is in the form of beads or other small discrete particles, which may be magnetic or paramagnetic beads to facilitate isolation via the application of a magnetic field. Generally, a “binding pair” refers to one of the first and second parts, the first and second parts having a specific binding affinity to each other. Suitable binding pairs include, but are not limited to, antigens / antibodies (e.g., digoxigenin / anti-digoxigenin, dinitrophenyl (DNP) / anti-DNP, dansyl-X-anti-dansyl, fluorescein / anti-fluorescein, Lucifer Yellow / anti-Lucifer Yellow, and rhodamine / anti-rhodamine); biotin / avidin (or biotin / streptavidin); calmodulin-binding protein (CBP) / calmodulin; hormones / hormone receptors; lectins / carbohydrates; peptides / cell membrane receptors; protein A / antibodies; haptens / anti-haptens; enzymes / cofactors; and enzymes / substrates.
[0095] In some embodiments, enrichment after amplification of a cyclic polynucleotide includes one or more additional amplification reactions. In some embodiments, enrichment includes amplifying a target sequence comprising sequence A and sequence B (oriented in the 5' to 3' direction) in an amplification reaction mixture, the amplification reaction mixture comprising (a) an amplified polynucleotide; (b) a first primer comprising sequence A', wherein the first primer specifically hybridizes to sequence A of the target sequence by sequence complementarity between sequence A and sequence A'; (c) a second primer comprising sequence B, wherein the second primer specifically hybridizes to sequence B' present in a complementary polynucleotide comprising the complement of the target sequence by sequence complementarity between sequence B and sequence B'; and (d) a polymerase extending the first primer and the second primer to produce an amplified polynucleotide, wherein the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nt or less. Figure 10 illustrates an example of a first and second primer configuration relative to a target sequence in the context of a single repeat (typically not amplified unless cyclic) and concatemers containing multiple copies of the target sequence. Considering the orientation of the primers relative to the monomer of the target sequence, this configuration may be called “back-to-back” (B2B) primers or “reverse” primers. Amplification using B2B primers promotes enrichment of cyclic and / or concatemeric amplification products. Furthermore, this orientation, combined with a relatively small footprint (the total distance covered by the pair of primers), allows for the amplification of a wide variety of fragmentation events around the target sequence because junctions are less likely to occur between primers than in the primer configuration seen in typical amplification reactions (facing each other and extending to the target sequence). Further embodiments and advantages of back-to-back primers are illustrated in AC of Figure 13.
[0096] In some embodiments, the distance between the 5' end of sequence A and the 3' end of sequence B is less than or equal to about 200, 150, 100, 75, 50, 40, 30, 25, 20, 15, or fewer nucleotides. In some embodiments, sequence A is the complement of sequence B. In some embodiments, multiple pairs of B2B primers directed to multiple different target sequences are used in the same reaction to amplify multiple different target sequences in parallel (e.g., about or at least about 10, 50, 100, 150, 200, 250, 300, 400, 500, 1000, 2500, 5000, 10000, 15000, or more different target sequences). The primers may be of any suitable length, as separately described herein. Amplification may also include any suitable amplification reaction under suitable conditions, such as the amplification reactions described herein. In some embodiments, amplification is a polymerase chain reaction.
[0097] In some embodiments, the B2B primer includes at least two sequence elements: a first element that hybridizes to the target sequence by sequence complementarity, and a 5' tail that does not hybridize to the target sequence during the first amplification step at a first hybridization temperature, where the first element hybridizes (for example, immediately 3' from the site where the first element junctions, due to a lack of sequence complementarity between the tail and a portion of the target sequence). For example, the first primer includes sequence C5' for sequence A', and the second primer includes sequence D5' for sequence B, where neither sequence C nor sequence D hybridizes to multiple concatemers during the first amplification step at the first hybridization temperature. In some embodiments, such tail-added primers are used, and amplification may involve a first and second step, the first step comprising a hybridization step at a first temperature during which the first and second primers hybridize to a concatemer (or cyclic polynucleotide) and a primer extension, and the second step comprising a hybridization step at a second temperature higher than the first temperature during which the first and second primers hybridize to an amplification product containing the extended first and second primers or their complement and the primer extension. Higher temperatures prefer hybridization between the first element and the tail element of the primer in the primer extension product over short fragments formed by hybridization between only the first element in the primer and the internal target sequence in the concatemer. Accordingly, two-step amplification may be used to reduce to some extent the preference for shorter amplification products, thereby maintaining a relatively higher proportion of amplification products having two or more copies of the target sequence. For example, after five cycles of hybridization of the second temperature and primer extension (e.g., at least 5, 6, 7, 8, 9, 10, 15, 20, or more cycles), at least 5% of the amplified polynucleotides in the reaction mixture (e.g., at least 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, or more) contain two or more copies of the target sequence.An example of an embodiment relating to this two-stage tail-added B2B primer amplification process is illustrated in Figure 11, part AD.
[0098] In some embodiments, enrichment involves amplification under distorted conditions to increase the length of the amplicon from the concatemer. For example, the primer concentration can be reduced so that not all priming sites hybridize the primer, thereby making the PCR product longer. Similarly, reducing the primer hybridization time during the cycle allows fewer primers to hybridize, thereby increasing the average PCR amplicon size. Furthermore, increasing the cycle temperature and / or extension time may also increase the average length of the PCR amplicon. Any combination of these techniques can also be used.
[0099] In some embodiments, particularly when amplification is performed using B2B primers, the amplification product is treated to filter the resulting amplicons based on size in order to reduce and / or remove monomers, which are a mixture containing concatemers. This can be done by various available techniques, but are not limited to, fragment excision from gel and gel filtration (e.g., for enriching fragments larger than approximately 300, 400, 500 or more nucleotides in length), as well as by using SPRI beads (Agencourt AMPure XP) for size sorting by fine-tuning the binding buffer concentration. For example, the use of 0.6x binding buffer during mixing with DNA fragments may be used to preferentially bind DNA fragments larger than approximately 500 base pairs (bp).
[0100] In some embodiments, if amplification yields a single-stranded concatemer, the single-stranded is converted to a double-stranded construct before or as part of the formation of a sequencing library generated for the sequencing reaction. A variety of suitable methods are available for generating a double-stranded construct from a single-stranded nucleic acid. Similarly, many other methods can be used, but many possible methods are illustrated in Figure 9A. As shown in Figure 9A, for example, the use of random primers, polymerase, dNTPs, and ligase yields a double-stranded structure. Figure 9B illustrates the synthesis of a second strand when the concatemer contains an adapter sequence that can be used as a primer in the reaction. Figure 9C illustrates the use of a “loop” when one end of a loop adapter is added to the end of the concatemer, and the loop adapter has a small portion of nucleic acid that self-hybridizes. In this case, the ligation of the loop adapter yields a loop that self-hybridizes and serves as a polymerase primer template. Figure 9D shows the use of hyperbranched primers, which are generally the most commonly used when the target sequence is known and multiple chains are formed, particularly when polymerases with strong chain displacement capabilities are used.
[0101] According to several embodiments, a cyclic polynucleotide (or its amplification product, which may optionally be enriched) is subjected to a sequencing reaction to generate sequencing reads. Sequence reads generated by such methods may be used in conjunction with other methods disclosed herein. A variety of sequencing methods are available, in particular high-throughput sequencing methods. Examples include, but are not limited to, sequencing systems manufactured by Illumina (such as HiSeq® and MiSeq® sequencing systems), Life Technologies (such as Ion Torrent® and SOLiD®), Roche's 454 Life Sciences systems, Pacific Biosciences systems, and others. In some embodiments, sequencing involves the use of HiSeq® and MiSeq® systems to generate reads of approximately 50, 75, 100, 125, 150, 175, 200, 250, 300, or more nucleotides in length. In some embodiments, sequencing involves sequencing by a synthetic process, where individual nucleotides are iteratively identified as they are added to a growing primer extension product. Pyrosequencing is an example of sequencing by a synthetic process that identifies nucleotide incorporation by assaying the resulting synthetic mixture for the presence of a byproduct of the sequencing reaction (i.e., pyrophosphate). In particular, a primer / template / polymerase complex is brought into contact with a single type of nucleotide. If the nucleotide is not incorporated, the polymerization reaction cleaves the nucleoside triphosphate between the α and β phosphates of the triphosphate chain, releasing a pyrophosphate. The presence of the released pyrophosphate is identified using a chemiluminescent enzyme reporter system, which converts the pyrophosphate to ATP with AMP, and then measures the ATP using a luciferase enzyme to generate a measurable light signal. If light is detected, the base is incorporated; if light is not detected, the base is not incorporated. After a suitable washing step, consecutive bases in the template sequence are sequentially identified by periodically contacting the complex with various bases.For example, see U.S. Patent No. 6,210,891.
[0102] In the relevant sequencing process, the primer / template / polymerase complex is immobilized on the substrate, and the complex comes into contact with the labeled nucleotide. Immobilization of the complex may be achieved by the primer sequence, template sequence, and / or polymerase enzyme, and may be covalent or non-covalent. For example, complex immobilization may occur via binding between the polymerase or primer and the substrate surface. In alternative configurations, the nucleotide is provided with or without a group of removable terminators. Upon incorporation, the label is fused to the complex and therefore detectable. In the case of terminators with nucleotides, all four distinct nucleotides, each with individually identifiable labels, come into contact with the complex. The incorporation of the labeled nucleotide is inhibited by the presence of the terminator, which adds the label to the complex, allowing for the identification of the incorporated nucleotide. The label and terminator are then removed from the incorporated nucleotide, and after appropriate washing steps, the process is repeated. In the case of non-terminal nucleotides, a single type of labeled nucleotide is added to the complex to determine whether it will be incorporated, as in pyrosequencing. After removal of the nucleotide-labeled group and appropriate washing steps, various different nucleotides are circulated by the reaction mixture in the same process. See, for example, U.S. Patent No. 6,833,246, which is incorporated herein by reference in its entirety for all purposes. For example, the Illumina Genome Analyzer System is based on the technique described in WO98 / 44151, in which a DNA molecule is bound to a sequencing platform (flow cell) by an anchor probe binding site (otherwise called a flow cell binding site) and amplified in Insights on a glass slide. The solid surface on which the DNA molecule is amplified typically contains several first and second bound oligonucleotides, the first bound oligonucleotide being complementary to a sequence near or at one end of the target polynucleotide, and the second bound oligonucleotide being complementary to a sequence near or at the other end of the target polynucleotide. This arrangement allows for the amplification of crosslinks, as described in U.S.20140121116.Subsequently, the DNA molecule is annealed to a sequencing primer and sequenced base by base in parallel using a reversible terminator approach. Prior to hybridization of the sequencing primer, one strand of the double-stranded crosslinked polynucleotide at the cleavage site may be preceded in one of the bound oligonucleotides that fix the crosslinking, so that one strand is not bound to the solid substrate which may be removed by denaturation, while the other strand is bound and available for hybridization to the sequencing primer. Typically, the Illumina Genome Analyzer System utilizes an eight-channel flow cell to generate sequencing reads of 18–36 bases in length, producing high-quality data of >1.3 Gbp per run (see www.illumina.com).
[0103] In further sequencing processes, the incorporation of differently labeled nucleotides is observed in real time when template-dependent synthesis occurs. In particular, when fluorescently labeled nucleotides are incorporated, individual fixed primer / template / polymerase complexes are observed, allowing for real-time identification of each added base at the time of incorporation. In this process, the label group is bound to a portion of the nucleotide that is cleaved during incorporation. For example, by attaching the label group to a portion of the phosphate chain removed during incorporation—i.e., to the β, γ, or other terminal phosphate groups on a nucleoside polyphosphate—the label is not incorporated into the nascent strand, and instead, native DNA is generated. Observation of individual molecules typically involves photoconfinement of the complex within very low illumination levels. By optically confining the complex, those skilled in the art can create a monitoring region where randomly diffusing nucleotides are present for only a very short time, while incorporated nucleotides are retained within the observation level for a longer period as they are incorporated. This results in a characteristic signal associated with the incorporation event, which also features a signal profile, characteristic of the added base. In relevant embodiments, interacting labeling components, such as fluorescence resonance energy transfer (FRET) dye pairs, are provided to the polymerase or other parts of the complex and to the uptake nucleotides, so that the uptake events interact, bringing the labeling components closer together, producing a characteristic signal that is also characteristic of the base being re-uptaken (see, for example, U.S. Patent Nos. 6,917,726, 7,033,764, 7,052,847, 7,056,676, 7,170,050, 7,361,466, and 7,416,844; and U.S. Patent No. 20070134128).
[0104] In some embodiments, nucleic acids in a sample can be sequenced by ligation. For example, as used in the Polony method and SOLiD technology (Applied Biosystems, now Invitrogen), this method typically uses DNA ligase enzymes to identify target sequences. Generally, a pool of any possible fixed-length oligonucleotides is labeled according to the sequenced position. The oligonucleotides are annealed and ligated; preferential ligation by DNA ligase to match the sequences yields a signal corresponding to a complementary sequence at that position.
[0105] In some embodiments, the sequencing library is constructed from DNA concatemers amplified prior to sequencing analysis. The amplified DNA concatemers are simultaneously fragmented and taggable with sequencing adapters, as illustrated in Figure 12A. In some cases, the amplified DNA concatemers are fragmented, for example, by sonication, and adapters are added to both ends of the fragments, as illustrated in Figure 12B.
[0106] According to some embodiments, a sequence difference between a sequencing read and a reference sequence is called a pure sequence variant (e.g., present in the sample before amplification or sequencing and not a result of either of these processes) if it occurs in at least two different polynucleotides (e.g., two different cyclic polynucleotides that are distinguishable because they have different junctions). Since sequence variants resulting from amplification or sequencing errors may not replicate accurately (e.g., in terms of position and type) on two different polynucleotides containing the same target sequence, adding this validation parameter significantly reduces the background of false sequence variants and simultaneously increases the sensitivity and accuracy of detecting actual sequence variants in the sample. In some embodiments, the values are approximately 5%, 4%, and 3%. Sequence variants with frequencies of 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or less, are sufficiently high above the background to enable accurate calling. In some embodiments, sequence variants occur at a frequency of about 0.1% or less. In some embodiments, the frequency of sequence variants is such that the background error rate is approximately 0.05%, 0.01%, 0.0 A value statistically significant above a p-value of 0.0001 or less is considered sufficiently high compared to the background. In some embodiments, the frequency of sequence variants is sufficiently high compared to the background (e.g., at least 5 times higher) if it is about or at least about 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 25 times, 50 times, 100 times, or more than the background error rate. In some embodiments, the background error rate for accurately determining a sequence at a given position is about 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001%, or less. In some embodiments, the error rate is lower than 0.001%.
[0107] In some embodiments, identifying pure sequence variants (also called “calling” or “making a call”) involves optimally aligning one or more sequencing reads to a reference sequence to identify differences between two and to identify junctions. Generally, alignment involves placing one sequence along another, repeatedly introducing gaps along each sequence, scoring how well the two sequences match, and preferably repeating at various positions along the reference sequence. The best-scoring match is considered the alignment and represents an inference about the degree of relationship between the sequences. In some embodiments, the reference sequence on which the sequencing reads are compared is a reference genome, such as the genome of a member of the same species as the subject. The reference genome may be complete or incomplete. In some embodiments, the reference genome consists only of regions containing the target polynucleotide, such as from the reference genome or from a consensus generated from the sequencing reads being analyzed. In some embodiments, the reference sequence consists of or comprises sequences of polynucleotides from one or more organisms, such as sequences from one or more bacteria, archaea, viruses, protists, fungi, or other organisms. In some embodiments, the reference sequence consists only of a portion of the reference genome, such as a region (e.g., one or more genes or parts thereof) corresponding to one or more target sequences under analysis. For example, for pathogen detection (e.g., in the case of contamination detection), the reference genome is the whole genome or a portion thereof of a pathogen (e.g., HIV, HPV, or a harmful strain, e.g., E. coli) that is useful in identifying a specific strain or serotype. In some embodiments, sequencing reads are aligned to multiple different reference sequences, such as for screening multiple different organisms or strains.
[0108] In a typical alignment, a base in the sequencing read on the side of a mismatched base in the reference sequence indicates that a substitution mutation occurred at that point. Similarly, if one sequence contains a gap next to a base in another sequence, it is inferred that an insertion or deletion mutation ("indel") has occurred. When it is desirable to explicitly indicate that one sequence is aligned to another, the alignment is often called a paired alignment. Many sequence alignments typically refer to the alignment of two or more sequences, including, for example, those by a series of paired alignments. In some embodiments, the scoring of an alignment includes values set for the likelihood of substitutions and indels. When individual bases are aligned, matches or mismatches contribute to the alignment score by substitution probability, which could be, for example, 1 for a match and 0.33 for a mismatch. Indels are deducted from the alignment score by a gap penalty, which could be, for example, -1. The gap penalty and substitution probability may be based on empirical knowledge or deductive assumptions about how the sequences change. Their values affect the resulting alignment. Examples of algorithms for performing alignment include, but are not limited to, the Smith-Waterman (SW) algorithm, the Needleman-Wunsch (NW) algorithm, algorithms based on the Burrows-Wheeler Transform (BWT), and Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (maq.sourceforge.Examples of hash function aligners include those available on .net. One exemplary alignment program that implements the BWT approach is the Burrows-Wheeler Aligner (BWA), available from the SourceForge website maintained by Geeknet (Fairfax, Va.). BWTs typically occupy 2 bits of memory per nucleotide, allowing indexing of nucleotide sequences up to 4G base pairs in length using a typical desktop or laptop computer. Preprocessing includes building the BWT (i.e., indexing the reference) and preliminary support data structures. BWA includes two different algorithms, both based on BWTs. Alignment with BWA is processed using the algorithm bwa-short and can be designed for short queries up to approximately 200 with a low error rate (<3%) (Li H. and Durbin R. Bioinformatics, 25:1754-60 (2009)). The second algorithm, BWA-SW, is designed for longer reads with more error (Li H. and Durbin R. (2010). Fast and accurate long-read alignment with Burrows-Wheeler Transform. Bioinformatics, Epub.). The bwa-sw aligner is often called "bwa-long," "bwa long algorithm," or similar. MUMmer is an alignment program that runs a version of the Smith-Waterman algorithm and is available from the SourceForge website maintained by Geeknet (Fairfax, Va.). MUMmer is a system for rapidly aligning whole genomes, whether in complete or draft form (Kurtz, S., et al., Genome Biology, 5:R12 (2004); Delcher, AL, et al., Nucl. Acids Res., 27:11 (1999)). For example, MUMmer 3.0 is based on version 2.On a 4GHz Linux® desktop computer, using 78MB of memory, it can find all 20 or more accurate base pairs or more between a pair of 5-megabase genomes in 13.7 seconds. MUMmer can further align incomplete genomes; it can handle hundreds or thousands of contigs from shotgun sequencing projects and align them to another set of contigs or genomes using the NUCmer program included in the system. Other non-exclusive examples of alignment programs include: BLAT from Kent Informatics (Santa Cruz, Calif.) (Kent, WJ, Genome Research 4: 656-664 (2002)); SOAP2 from Beijing Genomics Institute (Beijing, Conn.) or BGI Americas Corporation (Cambridge, Mass.); Bowtie (Langmead, et al., Genome Biology, 10:R25 (2009)); Efficient Large-Scale Alignment of Nucleotide Databases (ELAND) or the ELANDv2 component of the Consensus Assessment of Sequence and Variation (CASAVA) software (Illumina, San Diego, Calif.); RTG Investigator from Real Time Genomics, Inc. (San Francisco, Calif.); Novoalign from Novocraft (Selangor, Malaysia); and Exonerate, European Bioinformatics Institute (Hinxton, UK). (Slater, G., and Birney, E.ClustalW or ClustalX from University College Dublin (Dublin, Ireland) (Larkin MA, et al., Bioinformatics, 23, 2947-2948 (2007)); and FASTA, European Bioinformatics Institute (Hinxton, UK) (Pearson WR, et al., PNAS 85(8):2444-8 (1988); Lipman, DJ, Science 227(4693):1435-41 (1985)).
[0109] Examples of the process according to several embodiments are provided in Figures 36A–H, in particular for embodiments using a 3' tailing reaction. Figure 36A shows cell-free double-stranded polynucleotides 1, 2, 3…K(101) of a sample, each of which contains a locus (100) consisting of a single nucleotide that may be occupied by "G" or the rare variant "A". The sample containing such polynucleotides may be a patient tissue sample, such as a blood or plasma sample. Typically, a reference sequence (e.g., in a human genome database) is available for comparison with the polynucleotide sequence. Each polynucleotide has four sequence regions at each end, corresponding to the sequences of two complementary strands. Thus, for example, target polynucleotide 1 in Figure 36A has sequence regions n1(110) and n2(112) at each end of the strand and complementary sequence regions n1'(116) and n2'(108) at the ends of the complementary strand (120). While sequence regions of various polynucleotide chains have been exemplified as small parts of the chain, a sequence region can encompass the entire segment from the end of the chain to the gene locus (100).
[0110] In some embodiments, 3' tail-addition activity is added along nucleic acid monomers and / or other reaction components to a target polynucleotide of a sample in order to perform a tail-addition reaction (125) that extends the 3' end with one or more A's. In this embodiment, the predetermined nucleotide extension is indicated as "A…A", suggesting that one or more nucleotides are added, or that the exact number added to each chain is not determined (unless an exo-polymerase is used, as noted below). The expression of added nucleotides by "A…A" is not intended to limit the type of added nucleotide to A only. The added nucleotides are predetermined in the sense that the type of nucleotide precursor used in the tail-addition reaction is known and selected as an assay design choice. For example, a factor in the selection of the predetermined nucleotide type in a particular embodiment may be the efficiency of the cyclization step in terms of the type of nucleotide selected. In some embodiments, the nucleotide precursor may be, separately, any nucleoside triphosphate of any of the four nucleotides, resulting in the production of a homopolymer tail, or a mixture thereof, resulting in the production of a bi-nucleotide tail or a tri-nucleotide tail. In some cases, uracil and / or nucleotide analogs may be used in addition to or instead of the four natural DNA bases. In some embodiments in which the CircLigase® enzyme is used, the predetermined nucleotides may be A' and / or T'. In some embodiments, exopolymerase is used in a tail addition reaction, and only a single deoxyadenylate is added to the 3' end.
[0111] After tail addition and optional separation of the reaction products from the reaction mixture, the individual chains are cyclized using a cyclization reaction to generate a ring (132), as shown in Figure 36B, each containing a sequence element in the form of "nj-A…A-nj+1" (133). After cyclization and optional separation of the ring (132) from the reaction mixture, primers (134) are annealed to one or more primer-binding sites on the ring (132), and then they are extended to produce concatemers, each containing a copy of the respective nj-A…A-nj+1 sequence element, as illustrated in Figure 36E. After sequencing, complementary chains such as (136) and (138) may be identified by matching the sequence element components nj and nj+1 to their respective complements nj' and nj+1'. The selection of primer binding sites on ring (132) is a matter of design choice, or alternatively, random sequence primers may be used. In some embodiments, a single primer binding site is selected adjacent to locus (100); in other embodiments, multiple primer binding sites are each selected for separate primers so as to ensure amplification even if a boundary occurs incidentally at one of the primer binding sites. In some embodiments, two primers having separate primer binding sites are used to generate concatemers.
[0112] After identifying a pair of concatemers containing complementary strands, the concatemer sequences may be aligned, and base calls at corresponding positions on the two strands may be compared. At some positions in a concatemer pair, as illustrated by (140) in Figure 36F, a base called at a given position on one member of a pair may not be complementary to a base called on another member of the pair, indicating that an inaccurate call was made, for example, due to amplification errors, sequencing errors, etc. In this case, the uncertainty at a given position may be resolved by examining the base calls at the corresponding positions of other copies within the concatemer pair. For example, the base call at a given position may be obtained to match, or nearly match, the base call made for the individual copies in a pair of concatemers. Other methods for making such determinations are available to those skilled in the art and may be used instead of, or in addition to, these methods to supplement efforts to resolve base calls when the sequence information between complementary strands is not complementary. In some cases, if bases at specific positions in complementary strands starting from the same double-stranded molecule are not complementary (for example, as identified by the sequences at the 3' and 5' ends), the base call resolves in favor of the reference sequence being compared to the sample sequence, so that the difference is not identified as a true sequence variant with respect to such a reference sequence.
[0113] In other situations, the same error may appear in each copy of the target polynucleotide within the concatemer, as illustrated by (145) in Figure 36G. Such data would suggest that the target polynucleotide was corrupted before amplification or sequencing.
[0114] In yet another scenario, only a single concatemer may be identified; that is, a concatemer in which no match is found based on boundary information such as the length of a predetermined nucleotide segment and the sequences of adjacent 3' and 5' ends. Such a scenario is illustrated in Figures 37A and B, where the target polynucleotide (201) comprises a single-stranded polynucleotide 1 and a double-stranded polynucleotide 2, each containing a locus (200). A predetermined nucleotide (e.g., adenylate) may be bonded to both polynucleotides 1 and 2 during a tail-addition reaction (225) to form a 3' tail-added polynucleotide (220). As described above, polynucleotide (220) may then be circularized, amplified by RCA, and sequenced to obtain the concatemer sequence (230) shown in Figure 37B. If the observed variants are common in the DNA damage (e.g., C vs T or G vs T), such information from the unpaired concatemer is still useful in determining whether it is a true mutation versus mere DNA damage.
[0115] In some embodiments, as illustrated in Figures 36C and D, primers, each containing a molecular tag, e.g., MT1(150), MT2, etc., may be annealed to their respective single-stranded rings at predetermined primer-binding sites to produce concatemers, each with a unique tag. The presence of a unique molecular tag identifies single-stranded ring products that coincidentally have the same boundary or nj-A…A-nj+1 sequence elements. Such tags may also be used to count molecules to determine copy number variation at a locus, for example, according to the method described in Brenner et al., U.S. Patent No. 7,537,897, etc., incorporated herein by reference. In some embodiments, primers with molecular tags having binding sites on only one strand of the target polynucleotide may be selected, such that the concatemer with the molecular tag represents only one of the two strands of the target polynucleotide. In other embodiments, each of the complementary strands of the target polynucleotide may be amplified using a molecular-tagged primer (as shown in Figure 36C).
[0116] In some embodiments, the above steps for identifying the complementary strand of a target polynucleotide may be incorporated into a method for detecting rare variants at a locus. In some embodiments, the above method includes the steps of: (a) extending the 3' end of a polynucleotide by one or more predetermined nucleotides; (b) circularizing the individual strands of the polynucleotide to form a single-stranded polynucleotide ring, wherein one or more predetermined nucleotides define a boundary between the 3' and 5' sequences of each single-stranded polynucleotide ring; (c) amplifying the single-stranded polynucleotide ring by rolling circle replication (RCR) to form a concatemer; (d) sequencing the concatemer; (e) identifying pairs of concatemers containing the complementary strand of the polynucleotide by identifying 3' and 5' sequences adjacent to one or more predetermined nucleotides; and (f) sequencing a locus from the sequences of pairs of concatemers containing the complementary strand of the same polynucleotide. In another embodiment, the step of amplifying a single-stranded ring by RCR includes the steps of annealing a primer having a 5'-noncomplementary tail to the single-stranded ring, wherein such a primer contains a specific molecular tag in the 5'-noncomplementary tail, and extending the primer according to the RCR protocol. The resulting product is a concatemer containing the specific molecular tag, which may be counted together with other molecular tags bound to the ring from the same locus to yield a copy number measurement of the locus.
[0117] In some embodiments, the extension step may be carried out by tail-adding one or more predetermined nucleotides to the 3' end of the polynucleotide during the tail-addition reaction. In some embodiments, such tail-addition may be carried out by untemplated 3' nucleotide addition activity, such as TdT activity or exopolymerase activity.
[0118] Using the process described above, concatemer sequences can be identified from polynucleotide sequences. In large-scale parallel sequencing (also known as "next-generation sequencing" or NGS), reads containing concatemers can be identified and used to perform error correction and find sequence variants. The junction of the original input molecule (beginning and end of the DNA / RNA sequence) can be reconstructed from the concatemer by aligning the concatemer to the reference sequence. This junction can then be used to identify the original input molecule and remove sequencing duplication for more accurate measurements. Strand identity of each read that may contain a concatemer can be calculated by aligning the read to the reference sequence and checking the sequence element components nj and nj+1, as shown in Figure 36A. Variants found in both concatemers labeled as complementary strands have high statistical confidence and can be used to perform further error correction. Variant identification using chain identity may be performed by (but is not limited to) the following steps: a) Variants found in reads with complementary chain identity are considered more reliable; b) Reads carrying variants can be classified by their junction identification, and variants are more reliable if they are found in a group of reads with the same junction identification and complementary chain identity; c) Reads carrying variants can be classified by their molecular barcode, or a combination of their molecular barcode and junction identification. Variants are more reliable if they are found in a group of reads with the same molecular barcode and / or junction identification and complementary chain identity.
[0119] Error correction using molecular barcodes and junction identification can be used independently or in combination with error correction with concatemer sequencing as described in previous steps. a) Reads with various molecular barcodes (or junction identifications) can be classified into various read families representing reads derived from various input molecules; b) Consensus sequences can be constructed from families of reads; c) Consensus can be used for variant calling; d) Molecular barcodes and junction identification can be combined to form a synthetic ID for reads, which helps identify the original input molecule. In some embodiments, base calls found in various read families (e.g., sequence differences to a reference sequence) are given higher confidence. In some cases, sequence differences, when passed through one or more filters that increase the confidence of base calls as described above, are identified as true sequence variants representing the original source polynucleotide (against sample processing or analysis errors). In some embodiments, sequence differences are identified only as true sequence variants if (a) the sequence difference is identified on both strands of the double-stranded input molecule, (b) the sequence difference occurs in the consensus sequence for the concatemer from which it originates (e.g., 50%, 80%, or 90% or more of the repeats within the concatemer contain the sequence difference), and / or (c) the sequence difference occurs in two different molecules (e.g., identified by different 3' and 5' endpoints and / or by an exogenous tag sequence).
[0120] Determination of chain identity: 1) The junction of the original input molecule can be reconstructed from a read that may contain a concatemer sequence by aligning the sequence to a reference sequence; 2) The junction can be positioned on the read using alignment; 3) Sequence element components nj and nj+1 representing chain identity can be extracted from the sequence based on the junction position in the read, as shown in Figure 36A; and, in the case of a concatemer, the sequence may be found between junctions in the concatemer sequence; 4) Combined with the chain identity sequence in the read identified in step 3, the original chain from which the sequence variant originates can be identified using the chain (positive or negative) of the reference sequence to which the read is aligned, and the original chain from which the sequence variant originates. For example, suppose a strand identity sequence "AA" is added to the end of the original input DNA fragment strand; after sequencing, the read of the DNA fragment is aligned to the "+" strand of the reference sequence, the strand identity sequence in the read is "AA", and we know that the original input strand is "+"; if the strand identity sequence is "TT", the read is inversely complementary to the original input strand, and the original input strand is the "-" strand. Strand identity determination makes it possible to distinguish sequence variants from their inversely complementary counterparts, for example, distinguishing a C>T substitution from a G>A substitution. Accurate identification of allele changes can be used to perform allele-specific error reduction in variant calling. For example, DNA damage often occurs as a certain allele changes, and allele-specific error reduction can be performed to mitigate such damage; such error reduction can be performed, for example, by 1) calculating the distribution of various allele changes in the sequencing data (baseline) and then, 2) by a z-test or other statistical test to determine whether the observed allele changes differ from the baseline distribution.
[0121] In some embodiments, the Disclosure provides a method for identifying gene variants on a particular strand at a locus by comparing the frequency of a measured sequence, or one or more nucleotides, to the baseline frequency of the same sequence, or one or more nucleotides, resulting in nucleotide damage. In some embodiments, the method may include the steps of: (a) extending the 3' end of a polynucleotide by one or more predetermined nucleotides; (b) amplifying the individual strands of the extended polynucleotide; (c) sequencing the amplified individual strands of the extended polynucleotide; (d) identifying the complementary strand of the polynucleotide by the identity of the 3' and / or 5' sequences adjacent to one or more predetermined nucleotides, and identifying the nucleotides of each strand at the locus; and (e) determining the frequency of each of one or more nucleotides at the locus from the identified concatemer in order to identify gene variants. In some embodiments, this method may be used to distinguish gene variants from nucleotide damage by the following steps: whenever the frequency of a chain showing at least one nucleotide exceeds the baseline frequency of a chain having nucleotide damage resulting in the same nucleotide, call at least one of the one or more nucleotides at the locus on the chain identified as a gene variant by one or more predetermined nucleotides.
[0122] As mentioned above, in some embodiments, the amplification step is carried out by (i) circularizing the individual strands of polynucleotides to form a single-stranded polynucleotide ring, wherein one or more predetermined nucleotides define the boundary between the 3' and 5' sequences of the polynucleotides in each single-stranded polynucleotide ring, and (ii) amplifying the single-stranded polynucleotide ring by rolling circle replication to form a concatemer of the single-stranded polynucleotide ring.
[0123] The baseline frequency of strands with nucleotide damage may be based on previous measurements of samples from the same individual being tested by the method described above, or it may be based on previous measurements of a population of individuals other than the one being tested. The baseline frequency may also depend on and / or be specific to the type of process or protocol used when preparing the sample for analysis by the method of this disclosure. By comparing the measured frequency with the baseline frequency, a statistical measure may be obtained in terms of the likelihood (or confidence) that the measured or determined sequence is a pure gene variant and not due to damage or processing error.
[0124] Typically, sequencing data is obtained from large-scale parallel sequencing reactions. While other formats may be used, many next-generation high-throughput sequencing systems export data as FASTQ files. In some embodiments, sequences are typically analyzed by sequence alignment to identify repeat unit lengths (e.g., monomer lengths), junctions formed by circularization, and any true variations relative to the reference sequence. Identifying repeat unit lengths involves calculating the region of the repeat unit, finding the reference locus (e.g., when one or more sequences are the target of amplification, enrichment, and / or sequencing in particular), the boundaries of individual repeated regions, and / or the number of repeats within each sequencing run. Sequence analysis involves analyzing sequence data for both strands of a double helix. As noted above, in some embodiments, identical variants that appear to be sequences of different polynucleotide reads from a sample (e.g., circularized polynucleotides with different junctions) are considered confirmed variants. In some embodiments, sequence variants may also be considered confirmed or true variants if they occur in more than one repeat unit of the same polynucleotide. This is because the same sequence mutation is unlikely to occur similarly at the same location in a repeated target sequence within the same concatemer. Sequence quality scores may also be considered when identifying variants or confirmed variants; for example, sequences and bases with quality scores below a threshold may be excluded. Other bioinformatics methods can be used to increase the sensitivity and specificity of variant calling.
[0125] In some embodiments, statistical analysis may be applied to the determination of variants (mutations), and the proportion of variants in a complete DNA sample may also be quantified. The complete count of specific bases can be computed using sequencing data. For example, from alignment results calculated in a previous step, the number of "valid reads," i.e., the number of confirmed reads for each locus, can be calculated. The allele frequency of a variant can be normalized by the number of valid reads for a locus. The overall noise level, which is the average proportion of variants observed at all loci, can be calculated. The variant frequency and overall noise level, combined with other factors, can be used to determine the confidence interval for the variant call. Statistical models such as the Poisson distribution can be used to evaluate the confidence interval for the variant call. The allele frequency of a variant can also be used as an indicator of the relative amount of variants in a complete sample.
[0126] In some embodiments, microbial contaminants are identified based on a calling process. For example, certain sequence variants may indicate contamination by potentially infectious microorganisms. Sequence variants may be identified within highly conserved polynucleotides for the purpose of identifying microorganisms. Exemplary highly conserved polynucleotides useful for phylogenetic characterization and identification of microorganisms include nucleotide sequences found in the 16S rRNA gene, 23S rRNA gene, 5S rRNA gene, 5.8S rRNA gene, 12S rRNA gene, 18S rRNA gene, 28S rRNA gene, gyrB gene, rpoB gene, fusA gene, recA gene, coxl gene, and nifD gene. In eukaryotes, rRNA genes can be in the nucleus, mitochondria, or both. In some embodiments, sequence variants in the 16S-23S rRNA gene internal transcription spacer (ITS) can be used, with or without other rRNA genes, for the differentiation and identification of closely related taxa. Due to structural constraints of 16S rRNA, non-structural segments may exhibit high variability, but certain ranges throughout the gene contain highly conserved polynucleotide sequences. Using sequence variant identification, operational taxonomic units (OTUs) representing subgenera, genera, subfamilies, families, suborders, orders, subclasses, classes, subphylums, phylums, subkingdoms, or kingdoms can be identified, and their frequencies in populations can be determined as desired. Detection of specific sequence variants can be used to detect the presence and, optionally, the quantity (relative or absolute) of microbial contamination. Examples of applications include water quality testing for fecal or other contamination, testing for animal or human pathogens, accurately identifying sources of water pollution, testing of reclaimed or recovered water, testing of sewage discharges including ocean discharge upwellings, monitoring of aquaculture facilities for pathogens, monitoring of coastlines, swimming areas, or other water-related recreational facilities, and prediction of algal blooms.Applications of food surveillance include periodic testing of production lines in food processing plants, inspections of slaughterhouses, examinations of kitchens and food storage areas in restaurants, hospitals, schools, and prisons, and inspections of other foodborne pathogens such as E. coli strains O157:H7 or O111:B4, Listeria monocytogenes, and Salmonella enterica subsp. enterica serovar Enteritidis. Crustaceans and water-producing crustaceans may be investigated for algae involved in parasitic shellfish poisoning, neurotoxic shellfish poisoning, diarrhetic shellfish poisoning, and amnesic shellfish poisoning. Furthermore, imported food products can be screened at pre-release tariffs to ensure food safety. Plant pathogen monitoring applications include monitoring in horticulture and nurseries (e.g., monitoring for Phytophthora ramorum, microorganisms involved in Sudden Oak Death, crop pathogen monitoring and disease control, and forest pathogen monitoring and disease control). Where microbial contamination is a major safety concern, production environments for pharmaceuticals, medical devices, and other consumables or critical components can be investigated for the presence of specific pathogens such as Pseudomonas aeruginosa or Staphylococcus aureus, more common microorganisms related to humans, and microorganisms associated with the presence of water or other substances representing bioburdens previously identified in that particular or similar environment. Similarly, structural and assembly areas for high-precision equipment, including spacecraft, can be monitored for pre-identified microorganisms known to inhabit or most commonly introduced into such environments.
[0127] In some embodiments, the method includes the step of identifying sequence variants in a nucleic acid sample containing less than 50 ng of polynucleotides, each polynucleotide containing a 5' end and a 3' end. In some embodiments, the method includes: (a) cyclizing individual polynucleotides in the sample with a ligase to form a circular polynucleotide; (b) amplifying the circular polynucleotide to form a concatemer upon separation of the ligase from the circular polynucleotide; (c) sequencing the concatemer to generate a plurality of sequencing reads; (d) identifying sequence differences between the plurality of sequencing reads and a reference sequence; and (e) calling sequence differences occurring at a frequency of 0.05% or more in the plurality of reads from the nucleic acid sample containing less than 50 ng of polynucleotides as sequence variants.
[0128] The starting amount of polynucleotides in the sample may be small. In some embodiments, the amount of starting polynucleotides is less than 100 ng. In some embodiments, the amount of starting material is less than 75 ng. In some embodiments, the amount of starting material is less than 50 ng, such as 45 ng, 40 ng, 35 ng, 30 ng, 25 ng, 20 ng, 15 ng, 10 ng, 5 ng, 4 ng, 3 ng, 2 ng, 1 ng, 0.5 ng, and less than 0.1 ng. In some embodiments, the amount of starting polynucleotides is in the range of 0.1-100 ng, such as 1-75 ng, 5-50 ng, or 10-20 ng. Generally, lower starting material increases the importance of increased recovery from various processing steps. Processes that reduce the amount of polynucleotides in the sample for involvement in subsequent reactions reduce the sensitivity that can detect rare mutations. For example, the method described by Lou et al. (PNAS, 2013, 110 (49)) is expected to recover only 10-20% of the starting material. For large quantities of starting material (e.g., purified from bacteria cultured in a laboratory), this may not be a substantial obstacle. However, for samples with significantly low starting material, this low recovery range can be a substantial obstacle to the detection of sufficiently rare variants. Accordingly, in some embodiments, sample recovery from one step to another in the methods of this disclosure (e.g., mass fraction of input to a cyclization step available for input to a subsequent amplification or sequencing step) is approximately 50%, 60%, 75%, 80%, 85%, 90%, 95%, or higher. Recovery from a particular step may even be close to 100%. Recovery may also be for specific forms, such as the recovery of cyclic polynucleotides from non-circular polynucleotide inputs.
[0129] The polynucleotides may be derived from any suitable sample, such as the samples described herein for various aspects of this disclosure. The polynucleotides from a sample may be any of a variety of polynucleotides, including, but are not limited to, DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), fragments of any of these, or any combination of two or more of these. In some embodiments, the sample includes DNA. In some embodiments, the polynucleotides are single-stranded, as acquired or by treatment (e.g., denaturation). Furthermore, examples of suitable polynucleotides are described herein, such as for any of the various aspects of this disclosure. In some embodiments, the polynucleotides are subjected to subsequent processes (e.g., cyclization and amplification) that do not involve an extraction step and / or a purification step. For example, a fluid sample may be treated to remove cells without an extraction step to produce a purified fluid sample and a cell sample, and then DNA may be isolated from the purified fluid sample. Various procedures are available for the isolation of polynucleotides, such as precipitation or nonspecific binding, and subsequent washing of the substrate to release the bound polynucleotides. When polynucleotides are isolated from a sample without a cell extraction step, the polynucleotides are mostly extracellular polynucleotides or "cell-free" polynucleotides, such as cell-free DNA and cell-free RNA, which may correspond to dead or damaged cells. Such cell identity may be used to characterize the cells or cell populations from which they originate, such as microbial communities. When a sample is treated to extract polynucleotides from cells in the sample, various extraction methods are available, and examples thereof are provided herein (for example, with respect to any of the various aspects of this disclosure).
[0130] Sequence variants in nucleic acid samples can be any of several different sequence variants. Several non-limiting examples of sequence variants are described herein, for example, to any of the various aspects of this disclosure. In some embodiments, sequence variants are single nucleotide polymorphisms (SNPs). In some embodiments, sequence variants occur at a low frequency in a population (also called “rare” sequence variants). For example, sequence variants may occur at frequencies of about 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or less. In some embodiments, sequence variants occur at a frequency of about 0.1% or less.
[0131] According to several embodiments, the polynucleotides of a sample are cyclized, for example, by the use of a ligase. Cyclization may involve joining the 5' end of a polynucleotide to the 3' end of the same polynucleotide, to the 3' end of another polynucleotide, or to the 3' end of a polynucleotide from a different source (e.g., an artificial polynucleotide such as an oligonucleotide adapter). In some embodiments, the 5' end of a polynucleotide is joined to the 3' end of the same polynucleotide (also called "self-joining"). Non-limiting examples of cyclization processes (e.g., with and without adapter oligonucleotides), reagents (e.g., type of adapter and use of ligase), reaction conditions (e.g., supporting self-joining), and optional additional processing (e.g., post-reaction purification) are provided herein, including with respect to any of the various embodiments of this disclosure.
[0132] As previously described, junctions generally have a junction sequence when the ends of polynucleotides are joined together (either directly or using one or more intermediate adapter oligonucleotides) to form a cyclic polynucleotide. When the 5' and 3' ends of a polynucleotide are joined by an adapter polynucleotide, the term “junction” may refer to the junction between the polynucleotide and the adapter (e.g., one of the 5' or 3' junctions), or the junction between the 5' and 3' ends of a polynucleotide that is formed by and contains an adapter polynucleotide. When the 5' and 3' ends of a polynucleotide are joined without an intervening adapter (e.g., the 5' and 3' ends of single-stranded DNA), the term “junction” refers to the point where these two ends are joined. Junctions may be identified by the sequence of nucleotides containing the junction (also called the “junction sequence”). In some embodiments, the sample contains polynucleotides having a mixture of ends formed by natural degradation processes (cell lysis, cell death, and other processes in which DNA is released from the cell into its surrounding environment (where it may be further degraded in cell-free polynucleotides, cell-free DNA, and cell-free cells, etc.)), fragmentation as a byproduct of sample processing (fixation, staining, and / or storage procedures), and fragmentation by methods of cutting DNA into specific target sequences without restriction (e.g., mechanical fragmentation by sonication, etc.; non-sequence-specific nuclease treatment such as DNase I, fragmentase, etc.). When the sample contains polynucleotides having a mixture of ends, it is unlikely that two polynucleotides will have the same 5' or 3' end, and significantly unlikely that two polynucleotides will independently have both the same 5' and 3' ends. Accordingly, in some embodiments, junctions may be used to distinguish different polynucleotides, in which case the two polynucleotides may even contain portions having the same target sequence. When polynucleotide ends are joined without an intervening adapter, the joining sequence can sometimes be identified by alignment to a reference sequence.For example, if the order of two component sequences appears to be reversed compared to the reference sequence, the point where the reversal appears may indicate a junction. When polynucleotide ends are joined by one or more adapter sequences, the junction may be identified by proximity to a known adapter sequence, or by alignment as described above, provided that the sequencing read is long enough to obtain both the 5' and 3' sequences of the circularized polynucleotide. In some embodiments, the formation of a particular junction is such a rare event that it becomes unique among the circularized polynucleotides of the sample.
[0133] After cyclization, the reaction product may be purified before amplification or sequencing to increase the relative concentration or purity of the cyclized polynucleotide available for subsequent steps (e.g., isolation of the cyclized polynucleotide or removal of one or more other molecules in the reaction). For example, single-stranded (non-cyclized) polynucleotides may be removed by processing the cyclization reaction or its components, such as by treatment with an exonuclease. As a further example, the cyclization reaction or a part thereof may be subjected to stereoexclusion chromatography, which either retains and discards small amounts of reagent (e.g., unreacted adapters) or retains and releases the cyclized product in separate amounts. Various kits are available for purifying ligation reactions, such as the kits offered by Zymo Research's Zymo Oligo Purification Kit. In some embodiments, purification includes treatment to remove or degrade the ligase used in the cyclization reaction and / or to purify the cyclized polynucleotide from such ligase. In some embodiments, the treatment to degrade the ligase includes treatment with a protease. Suitable proteases are available from prokaryotes, viruses, and eukaryotes. Examples of proteases include proteinase K (from Tritirachium album), pronase E (from Streptomyces griceus), Bacillus polymixaprotease, thermolysin (from thermophilic bacteria), trypsin, subtilisin, and furin. In some embodiments, the protease is proteinase K. Treatment with proteases may follow the manufacturer's protocol or be subjected to standard conditions (e.g., provided in Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012)). Protease treatment may also involve extraction and precipitation.In one example, the cyclic polynucleotide is purified by proteinase K (Qiagen) treatment in the presence of 0.1% SDS and 20 mM EDTA, extracted with 1:1 phenol / chloroform and chloroform, and precipitated with ethanol or isopropanol. In some embodiments, the precipitation is in ethanol.
[0134] As described in other aspects of this disclosure, cyclization may directly involve subsequent sequencing of the cyclized polynucleotide. Alternatively, sequencing may precede one or more amplification reactions. Various methods are available for amplifying polynucleotides (e.g., DNA and / or RNA). Amplification may be linear, exponential, or involve both linear and exponential phases in a multiphase amplification process. Amplification methods may involve temperature fluctuations, such as thermal denaturation steps, or may be isothermal processes that do not require thermal denaturation. Non-limiting examples of suitable amplification processes are described herein in relation to any of the various aspects of this disclosure. In some embodiments, amplification involves rolling circle amplification (RCA). As described elsewhere herein, a typical RCA reaction mixture comprises one or more primers, polymerases, and dNTPs to produce concatemers. Typically, the polymerase in an RCA reaction is a polymerase with chain displacement activity. A variety of such polymerases are available, and non-limiting examples include the exonuclease minus DNA polymerase I large (Klenow) fragment, Phi29 DNA polymerase, and Taq DNA polymerase. Generally, concatemers are polynucleotide amplification products containing two or more copies of the target sequence from a template polynucleotide (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, or more copies of the target sequence; in some embodiments, about 2 or more copies). The amplification primers may be of any suitable length, such as approximately or at least approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, or more nucleotides, and some or all of them may be complementary to the corresponding target sequence that the primer hybridizes (e.g., approximately or at least approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides). Various examples of RCA processes, including the use of random primers, target-specific primers, and primers targeting adapters, are described herein, some of which are illustrated in AC of Figure 7.
[0135] When a cyclic polynucleotide is amplified before sequencing (for example, to generate concatemers), the amplified product may be subjected to sequencing directly without enrichment, or it may be subsequently subjected to one or more enrichment steps. Non-limiting examples of suitable enrichment processes are described herein in relation to any of the various aspects of this disclosure (e.g., the use of B2B primers in a second amplification step). According to some embodiments, a cyclic polynucleotide (or its amplified product, which may optionally be enriched) is subjected to a sequencing reaction to generate sequencing reads. Sequencing reads generated by such methods may be used in conjunction with other methods disclosed herein. A variety of sequencing methods are available, in particular high-throughput sequencing methods. Examples include, but are not limited to, sequencing systems manufactured by Illumina (such as HiSeq® and MiSeq® sequencing systems), Life Technologies (such as Ion Torrent® and SOLiD®), Roche's 454 Life Sciences systems, Pacific Biosciences systems, and others. In some embodiments, sequencing involves the use of HiSeq® and MiSeq® systems to generate reads of approximately 50, 75, 100, 125, 150, 175, 200, 250, 300, or more nucleotides in length. Additional non-limiting examples of amplification platforms and methods are described herein with respect to any of the various aspects of this disclosure.
[0136] According to some embodiments, a sequence difference between a sequencing read and a reference sequence is called a pure sequence variant (e.g., present in the sample before amplification or sequencing and not a result of either of these processes) if it occurs in at least two different polynucleotides (e.g., two different cyclic polynucleotides distinguishable by having different junctions, or two different polypeptides having different 5' ends and / or different 3' ends). Since sequence variants resulting from amplification or sequencing errors may not replicate accurately (e.g., in terms of position and type) on two different polynucleotides containing the same target sequence, adding this validation parameter significantly reduces the background of false sequence variants and simultaneously increases the sensitivity and accuracy of detecting actual sequence variants in the sample. In some embodiments, sequence variants having frequencies of approximately 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or less are sufficiently high above the background to enable accurate calling. In some embodiments, sequence variants occur at frequencies of approximately 0.1% or less. In some embodiments, the method includes calling sequence differences having frequencies in the range of approximately 0.0005%–approximately 3%, such as 0.001%–2% or 0.01%–1%, with pure sequence variants. In some embodiments, the frequency of sequence variants is sufficiently high above the background if it is statistically significant above the background error rate (e.g., p-values of approximately 0.05, 0.01, 0.001, 0.0001 or less). In some embodiments, the frequency of sequence variants is sufficiently high above the background (e.g., at least 5 times higher) when it is about or at least about 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 25 times, 50 times, 100 times, or more than the background error rate. In some embodiments, the background error rate for accurately determining the sequence at a given position is about 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001%, or 0.0005% or less.In some embodiments, the error rate is less than 0.001%. Methods for determining the frequency and error rate are described herein in relation to any of the various aspects of this disclosure.
[0137] In some embodiments, identifying pure sequence variants (also called “calling” or “making a call”) involves optimally aligning one or more sequencing reads to a reference sequence to identify differences between two and to identify junctions. Generally, alignment involves placing one sequence along another, repeatedly introducing gaps along each sequence, scoring how well the two sequences match, and preferably repeating this across various positions along the reference sequence. The best-scoring match is considered to be the alignment, representing an inference about the degree of relationship between the sequences. Various alignment algorithms and aligners are available to perform them, and non-limiting examples thereof are described herein in relation to any of the various aspects of this disclosure. In some embodiments, the reference sequence on which the sequencing reads are compared is a known reference sequence, such as a reference genome (e.g., the genome of a member of the same species as the subject). The reference genome may be complete or incomplete. In some embodiments, the reference genome consists only of regions containing the target polynucleotide, such as from the reference genome or from a consensus generated from the sequencing reads being analyzed. In some embodiments, the reference sequence comprises or consists of a polynucleotide sequence from one or more organisms, such as a sequence from one or more bacteria, archaea, viruses, protists, fungi, or other organisms. In some embodiments, the reference sequence consists only of a portion of a reference genome, such as a region (e.g., one or more genes or portions thereof) corresponding to one or more target sequences under analysis. For example, for the detection of a pathogen (e.g., in the case of contamination detection), the reference genome is the whole genome or a portion thereof of a pathogen (e.g., HIV, HPV, or a harmful strain, e.g., Escherichia coli) that is useful in identifying a particular strain or serotype. In some embodiments, sequencing reads are aligned to multiple different reference sequences, such as for screening multiple different organisms or strains. Additional non-limiting examples of reference sequences (and called sequence variants) in which sequence differences can be identified are described herein in relation to any of the various aspects of this disclosure.
[0138] In one embodiment, the present disclosure provides a method for amplifying a plurality of different concatemers containing two or more copies of a target sequence in a reaction mixture, wherein the target sequence includes sequences A and B oriented in the 5'-3' direction. In some embodiments, the method comprises the step of subjecting a reaction mixture to a nucleic acid amplification reaction, the reaction mixture comprising: (a) a plurality of concatemers, each concatemer comprising a different junction formed by cyclizing individual polynucleotides having a 5' end and a 3' end; (b) a first primer comprising sequence A', the first primer specifically hybridizing to sequence A of the target sequence by sequence complementarity between sequence A and sequence A'; (c) a second primer comprising sequence B, the second primer specifically hybridizing to sequence B' present in a complementary polynucleotide comprising the complement of the target sequence by sequence complementarity between sequence B and sequence B'; and (d) a polymerase extending the first primer and the second primer to produce an amplified polynucleotide, wherein the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nt or less.
[0139] In a related embodiment, the Disclosure provides a method for amplifying a plurality of different cyclic polynucleotides, including a target sequence, in a reaction mixture, wherein the target sequence includes sequences A and B oriented in the 5' to 3' direction. In some embodiments, the method comprises the step of subjecting a reaction mixture to a nucleic acid amplification reaction, the reaction mixture comprising: (a) a plurality of cyclic polynucleotides, where each cyclic polynucleotide comprises a different junction formed by the cyclization of individual polynucleotides having 5' and 3' ends; (b) a first primer comprising sequence A', which specifically hybridizes to sequence A of the target sequence via sequence complementarity between sequence A and sequence A'; (c) a second primer comprising sequence B, which specifically hybridizes to sequence B' present in a complementary polynucleotide containing the complement of the target sequence via sequence complementarity between sequence B and B'; and (d) a polymerase extending the first and second primers to produce amplified polynucleotides, where sequences A and B are endogenous sequences, and the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nt or less.
[0140] Amplification of either cyclic polynucleotides or concatemers may result in such polynucleotides originating from any suitable sample source (directly or indirectly, such as through amplification). A variety of suitable sample sources, optional extraction processes, types of polynucleotides, and types of sequence variants are described herein for any of the various aspects of this disclosure. Cyclic polynucleotides may originate from the cyclization of acyclic polynucleotides. Non-limiting examples of cyclization processes (e.g., with and without adapter oligonucleotides), reagents (e.g., types of adapters and use of ligases), reaction conditions (e.g., preference for self-conjugation), optional additional processing (e.g., post-reaction purification), and the junctions formed thereby are provided herein for any of the various aspects of this disclosure. Concatemers may originate from the amplification of cyclic polynucleotides. A variety of methods for amplifying polynucleotides (e.g., DNA and / or RNA) are available, non-limiting examples of which are also described herein. In some embodiments, concatemers are produced by rolling-circle amplification of cyclic polynucleotides.
[0141] Figure 10 illustrates an example of the placement of first and second primers to a target sequence in the context of a single repeat (typically circular, which will not be amplified) and concatemers containing multiple copies of the target sequence. As noted in relation to other embodiments described herein, this primer placement may be referred to as “back-to-back (B2B)” or “reverse” primers. Amplification with B2B primers facilitates enrichment of circular and / or concatemer templates. Furthermore, since junctions are likely not as far between the primers as in typical amplification reactions (facing each other and extending to the target sequence), this orientation, combined with a relatively small footprint (total distance extending to the pair of primers), allows for a wider range of fragmentation events around the target sequence. In some embodiments, the distance between the 5' end of sequence A and the 3' end of sequence B is approximately 200, 150, 100, 75, 50, 40, 30, 25, 20, 15, or fewer nucleotides. In some embodiments, sequence A is the complement of sequence B. In some embodiments, multiple pairs of B2B primers directed to multiple different target sequences are used in the same reaction to amplify multiple different target sequences in parallel (e.g., about, or at least about 10, 50, 100, 150, 200, 250, 300, 400, 500, 1000, 2500, 5000, 10000, 15000, or more different target sequences). The primers may be of an appropriate length, such as those separately described herein. Amplification may include an appropriate amplification reaction under appropriate conditions, such as the amplification reactions described herein. In some embodiments, amplification is a polymerase chain reaction.
[0142] In some embodiments, the B2B primer comprises at least two sequence elements: a first element that hybridizes to the target sequence via sequence complementarity, and a 5' "tail" that does not hybridize to the target sequence during a first amplification step at a first hybridization temperature, in which the first element hybridizes (for example, due to a lack of sequence complementarity between the tail and the portion of the target sequence adjacent to the 3' site of binding of the first element). For example, the first primer comprises sequence C5' for sequence A', the second primer comprises sequence D5' for sequence B, and both sequences C and D hybridize to a plurality of concatemers (or cyclic polynucleotides) during a first amplification step at a first hybridization temperature. In some embodiments in which such tailed primers are used, amplification may comprise a first and a second stage; the first stage comprises a hybridization step at a first temperature, during which the first and second primers hybridize to a concatemer (or cyclic polynucleotide) and a primer extension; the second stage comprises a hybridization step at a second temperature higher than the first temperature, during which the first and second primers hybridize to an amplification product comprising an extended first or second primer, or its complement, and a primer extension. The number of amplification cycles at each of the two temperatures can be adjusted based on the desired product. Typically, the first temperature is used for a relatively small number of cycles, such as about 15, 10, 9, 8, 7, 6, 5, or fewer. The number of cycles at higher temperatures can be chosen independently of the number of cycles at the first temperature, but is typically about the same or more, such as approximately 5, 6, 7, 8, 9, 10, 15, 20, 25, or more cycles. At higher temperatures, hybridization between the first element and tail element of the primer in the primer extension product is preferred, over shorter fragments formed by hybridization between only the first element in the primer and the internal target sequence within the concatemer.Therefore, two-step amplification can be used to reduce the degree to which shorter amplification products are preferred in other cases, thereby maintaining a relatively high proportion of amplification products containing two or more copies of the target sequence. For example, after five cycles of hybridization at a second temperature and primer extension (e.g., at least 5, 6, 7, 8, 9, 10, 15, 20, or more cycles), at least 5% of the amplified polynucleotide in the reaction mixture (e.g., at least 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, or more) contains two or more copies of the target sequence. As an example of an embodiment following this two-step process, a tail-addition B2B primer amplification process is illustrated in Figure 11 AD. Further examples of implementations are provided in Figure 15 AC.
[0143] In some embodiments, amplification is carried out under conditions that are distorted to increase the length of the amplicon from the concatemer. For example, the primer concentration can be reduced so that no stimulation site hybridizes the primer, and therefore the PCR product is longer. Similarly, reducing the primer hybridization time during the cycle allows fewer primers to hybridize, thus increasing the average size of the PCR amplicon. Furthermore, increasing the temperature and / or extension time during the cycle may also increase the average length of the PCR amplicon. Any combination of these techniques can also be used.
[0144] In some embodiments, particularly when amplification is performed with B2B primers, the amplification product is processed to filter the resulting amplicons based on size in order to reduce and / or remove the number of monomers in the concatemer-containing mixture. This can be done using various available techniques, including but not limited to fragment excision from gel and gel filtration (for example, to enrich fragments larger than approximately 300, 400, 500 or more nucleotides in length); similarly, it can be done using SPRI beads (Agencourt AMPure XP) for size selection by fine-tuning the binding buffer concentration. For example, the use of 0.6x binding buffer during mixing with DNA fragments may be preferred for binding DNA fragments larger than approximately 500 base pairs (bp).
[0145] In some embodiments, the first primer comprises sequence C5' for sequence A', and the second primer comprises sequence 5' for sequence B, wherein neither sequence C nor sequence D hybridizes to multiple cyclic polynucleotides during the first amplification step at a first hybridization temperature. Amplification may comprise a first and a second step; the first step comprises a hybridization step at a first temperature, during which the first and second primers hybridize to their cyclic polynucleotides or amplification product before primer extension; and the second step comprises a hybridization step at a second temperature higher than the first temperature, during which the first and second primers hybridize to an amplification product comprising the extended first or second primer or its complement. For example, the first temperature may be selected as the approximate Tm of sequence A', sequence B, or their average, or as a temperature higher than 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, or one of these Tm values. In this example, the second temperature may be selected as the approximate Tm of the combined sequence (A'+C), combined sequence (B+D), or their average, or as a temperature higher than 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, or one of these Tm values. The term "Tm" is also called the "melting temperature" and generally represents the temperature at which 50% of the oligonucleotide consisting of a reference sequence (which may actually be a subsequence within a larger polynucleotide) and its complementary sequence are hybridized (or separated). Generally, Tm increases with increasing length, and therefore, it is presumed that the Tm of sequence A' is lower than the Tm of the combination sequence (A'+C).
[0146] In one embodiment, the present disclosure provides a system for detecting sequence variants. In some embodiments, the system comprises: (a) a computer configured to receive user requests to perform a detection reaction on a sample; (b) an amplification system that performs a nucleic acid amplification reaction on the sample or a portion thereof in response to a user request, the amplification reaction comprising: (i) cyclizing a plurality of polynucleotides into individual polynucleotides in order to form a plurality of circular polynucleotides using a ligase enzyme, wherein each of the plurality of polynucleotides has a junction between its 3' end and 5' end prior to ligation; (ii) degrading the ligase enzyme; and (iii) amplifying the circular polynucleotides after degrading the ligase enzyme to produce amplified polynucleotides; the polynucleotides are not purified or isolated between steps (i) and (iii); (c) a sequencing system that generates sequencing reads for the polynucleotides amplified by the amplification system, identifies sequence differences between the sequencing reads and a reference sequence, and calls sequence differences occurring in at least two circular polynucleotides having different junctions as sequence variants; and (d) a reporting device that sends a report to the recipient including results regarding the detection of sequence variants. In some embodiments, the recipient is the user. Figure 32 illustrates a non-limiting example of a system useful in the method of this disclosure. Figures 29 and 30 provide an exemplary overview of an exemplary workflow design.
[0147] The computer used in this system may include one or more processors. A processor may be associated with one or more control units, computing units, and / or other units of the computer system, or may be embedded in firmware as desired. If implemented in software, routines may be stored in any computer-readable memory, such as RAM, ROM, flash memory, magnetic disks, laser disks (registered trademark), or other suitable storage media. Similarly, this software may be delivered to computing via any known delivery method, such as over communication channels like telephone lines, the internet, or wireless connections, or via mobile media such as computer-readable disks or flash drives. Various processes may be implemented as various blocks, operations, tools, modules, and techniques that can be implemented sequentially in hardware, firmware, software, or any combination of hardware, firmware, and / or software. When implemented in hardware, some or all of the blocks, operations, and techniques can be implemented, for example, in custom integrated circuits (ICs), application-specific storage circuits (ASICs), custom field-programmable logic arrays (FPGAs), programmable logic arrays (PLAs), etc. A client-server relational database structure can be used in embodiments of this system. A client-server architecture is a network architecture in which computers or processes on a network are either clients or servers. A server computer may typically be a high-performance computer dedicated to managing disk drives (file servers), printers (print servers), or network traffic (network servers). A client computer may include a PC (personal computer) or a workstation on which a user runs applications, as well as output devices such as those disclosed herein. A client computer depends on a server computer for resources such as files, devices, and processing power.The server computer handles all aspects of database functionality. Client computers can have software that handles all front-end data management and can receive data input from users.
[0148] This system may be configured to receive user requests to perform detection reactions on a sample. User requests may be direct or indirect. Examples of direct requests include those sent via input devices such as keyboards, mice, or touchscreens. Examples of indirect requests include those sent via communication media such as the internet (wired or wireless).
[0149] The system may further include an amplification system that performs nucleic acid amplification reactions on a sample or a portion thereof upon user request. Various methods for amplifying polynucleotides (e.g., DNA and / or RNA) are available. Amplification may be linear, exponential, or involve both linear and exponential processes in a multiphase amplification process. Amplification methods may involve temperature changes, such as thermal denaturation steps, or may be isothermal processes that do not require thermal denaturation. Non-limiting examples of suitable amplification processes are described herein, including those relating to any of the various aspects of this disclosure. In some embodiments, amplification includes rolling circle amplification (RCA). Various systems for amplifying polynucleotides are available and may vary depending on the type of amplification reaction performed. For example, with respect to amplification methods involving cycles of temperature changes, the amplification system may include a thermocycler. Amplification systems may include real-time amplification and detection devices, such as systems manufactured by Applied Biosystems, Roche, and Strategene. In some embodiments, the amplification reaction comprises (i) a step of cyclizing individual polynucleotides to form a plurality of cyclic polynucleotides, each of which has a junction between its 5' and 3' ends; and (ii) a step of amplifying the cyclic polynucleotides. The sample, polynucleotide, primer, polymerase, and other reagents may be any of those described herein, including those relating to any of the various embodiments of this disclosure. Non-limiting examples of cyclization processes (e.g., with and without adapter oligonucleotides), reagents (e.g., type of adapter and use of ligase), reaction conditions (e.g., preference for self-junction), optional additional processing (e.g., post-reaction purification), and junctions formed thereby are provided herein, including those relating to any of the various embodiments of this disclosure. Systems for carrying out such methods may be selected and designed.
[0150] The system may further include a sequencing system that generates sequencing reads for polynucleotides amplified by the amplification system, identifies sequence differences between the sequencing reads and a reference sequence, and calls sequence differences occurring in at least two cyclic polynucleotides with different junctions as sequence variants. The sequencing system and the amplification system may include the same or overlapping equipment. For example, both the amplification system and the sequencing system may utilize the same thermocycle. Various sequencing platforms for use with this system are available and may be selected based on the chosen sequencing method. Examples of sequencing methods are described herein. Amplification and sequencing may involve the use of a liquid handler. Various commercially available liquid handler systems can be used to automate these processes (see, for example, liquid handlers from Perkin-Elmer, Beckman Coulter, Caliper Life Sciences, Tecan, Eppendorf, Apricot Design, and Velocity 11). Various automated sequencing machines are commercially available, including sequencers manufactured by Life Technologies (SOLiD platform and pH-based detection), Roche (454 platform), and Illumina (e.g., flow cell-based systems, Genome Analyzer devices). Transfer between two, three, four, five, or more automated devices (e.g., between a liquid handler and one or more sequencing devices) can be manual or automated.
[0151] Methods for identifying sequence differences and calling sequence variants relative to a reference sequence are described herein in relation to any of the various aspects of this disclosure. Sequence determination systems typically include software for performing these steps in response to input of sequence determination data and desired parameters (e.g., selection of a reference genome). Examples of alignment algorithms and aligners for performing these algorithms are described herein, but are not limited to, the Needleman-Wunsch algorithm (e.g., see the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html with optional default settings), the BLAST algorithm (e.g., see the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi with optional default settings), or the Smith-Waterman algorithm (e.g., see the EMBOSS Water aligner available at www.ebi.ac.uk / Tools / psa emboss_water / nucleotide.html with optional default settings). The optimal alignment can be evaluated using any appropriate parameters of a selected algorithm, including default parameters. Such an alignment algorithm can form part of a sequencing system.
[0152] The system further includes a reporting device that sends a report to the recipient containing results regarding the detection of sequence variants. The report may be generated in real time, such as during sequencing reads or while sequencing data is being analyzed, with periodic updates as the process progresses. In addition, or alternatively, the report may be generated at the end of the analysis. The report may be generated automatically, in which case the sequencing system completes the step of calling all sequence variants. In some embodiments, the report is generated in response to instructions from the user. In addition to the results of sequence variant detection, the report may also include an analysis based on one or more sequence variants. For example, if one or more sequence variants are associated with a particular impurity or phenotype, the report may include information about this association, such as the likelihood and level of the presence of the impurity or phenotype, and optionally suggestions based on this information (e.g., retesting, monitoring, or corrective action). The report may take any of the various forms. Data relating to this disclosure may be transmitted over such networks or communications (connections) (or other appropriate means for transmitting information, including, but not limited to, mailing physical reports such as printouts) for receipt and / or confirmation by the recipient. The recipient may be, but is not limited to, an individual or an electronic device (e.g., one or more computers and / or one or more servers).
[0153] In one embodiment, the disclosure provides a computer-readable medium including code that performs a method for detecting sequence variants when executed by one or more processors. In some embodiments, the method implemented includes: (a) receiving a customer request to perform a detection reaction on a sample; (b) performing a nucleic acid amplification reaction on the sample or a portion thereof in response to the customer request, the amplification reaction comprising: (i) cyclizing a plurality of polynucleotides into individual polynucleotides in order to form a plurality of circular polynucleotides using a ligase enzyme, wherein each of the plurality of polynucleotides has a 5' end and a 3' end before ligation; (ii) degrading the ligase enzyme; and (iii) amplifying the circular polynucleotides after degrading the ligase enzyme to produce amplified polynucleotides; the polynucleotides are not purified or isolated between steps (i) and (iii); the amplification system; (c) performing sequencing analysis, comprising: (i) generating sequencing reads of the polynucleotides amplified in the amplification reaction; (ii) identifying the differences between the sequencing reads and a reference sequence; and (iii) calling sequence differences occurring in at least two circular polynucleotides having different junctions as sequence variants; and (d) generating a report including results relating to the detection of sequence variants.
[0154] Machine-readable media containing computer executable code may take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any storage device in a computer, which can be used to implement databases, etc. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include copper wires and optical fibers, including coaxial cables and wires including buses in computer systems. Carrier transmission media may take the form of electrical signals or electromagnetic signals, or sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tapes, other magnetic media, CD-ROMs, DVDs or DVD-ROMs, other optical media, punch cards, paper tapes, other physical storage media having a pattern of holes, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, other memory chips or cartridges, carriers for transporting data or instructions, cables or links for transmitting such carriers, or other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be related to transporting one or more sequences of one or more instructions to a processor for execution.
[0155] The computer executable code of the subject can be executed on any suitable device including a processor, such as a server, PC, or mobile device such as a smartphone or tablet. Any control device or computer may optionally include a monitor, which may be a cathode ray tube ("CRT") display, a flat panel display (e.g., an active-matrix liquid crystal display, a liquid crystal display, etc.), or other. Computer circuits are often housed in a box containing numerous integrated circuit chips, such as a microprocessor, memory, interface circuits, and others. The box may also optionally include high-capacity removable drives such as hard disk drives, floppy disk drives, writable CD-ROMs, and other common peripheral elements. Input devices such as a keyboard, mouse, or touch-sensitive screen may optionally provide user input. A computer may include software suitable for receiving user commands, in the form of user input to a set of parameter fields, e.g., a GUI, or in the form of pre-programmed instructions, e.g., pre-programmed for various different specific operations.
[0156] In some embodiments of any of the various aspects disclosed herein, the methods, compositions, and systems have therapeutic applications, such as characterizing patient samples and optionally diagnosing diseases in subjects. Therapeutic applications may also include steps to inform the patient of the most responsive treatment (also referred to as “theranostic”) and the selection of the actual treatment required for the subject, based on the results of the methods described herein. In particular, the methods and compositions disclosed herein may be used to diagnose the presence, progression, and / or metastasis of tumors, especially when the analyzed polynucleotides include, or consist of, cfDNA, ctDNA, cfRNA, or fragmented tumor DNA. In some embodiments, the subject is monitored for the treatment effect. For example, by monitoring ctDNA over time, a decrease in ctDNA can be used as an indicator of effective treatment, while an increase can facilitate the selection of different treatments or different doses. Other uses include the assessment of organ rejection in transplant recipients (in which case an increase in the amount of circulating DNA corresponding to the transplant donor genome is used as an early sign of transplant rejection), and genotyping / isotyping of pathogenic infections such as viral or bacterial infections. Detection of sequence variants in circulating fetal DNA may be used to diagnose fetal diseases.
[0157] As used herein, “treatment,” “to treat,” “to alleviate,” or “to induce remission” are interchangeable. These terms refer to methods for obtaining beneficial or desired outcomes, including, but not limited to, therapeutic and / or preventive effects. Therapeutic effect means an improvement in, or an effect on, one or more diseases, illnesses, or symptoms under treatment. With respect to preventive effects, a composition may be administered to subjects at risk of progression of a disease, illness, or symptom, even if the disease, illness, or symptom has not yet become apparent, or to subjects reporting one or more physiological symptoms of a disease. Typically, preventive effects include reducing the occurrence and / or exacerbation of one or more diseases, illnesses, or symptoms under treatment (e.g., between a treated group and an untreated group, or between a treated state and an untreated state of a subject). Improvement of therapeutic outcomes may include diagnosing the subject’s condition to identify a subject as one or more who will benefit from, or will not benefit from, treatment by one or more therapeutic agents or other therapeutic interventions (such as surgery). In such diagnostic applications, the overall success rate of treatment with one or more therapeutic agents may be improved compared to the effectiveness among patients grouped without diagnosis according to the method of this disclosure (e.g., as an improvement of at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more in measures of therapeutic effect).
[0158] The terms “subject,” “individual,” and “patient” are used interchangeably herein to refer to vertebrates, preferably mammals, more preferably humans. Examples include, but are not limited to, mice, monkeys, humans, livestock, sport animals, and pets. Tissues, cells, and their biological offspring obtained in vivo or cultured in vitro are also included.
[0159] The terms “therapeutic agent,” “therapeutic drug,” or “treatment agent” are used interchangeably and refer to molecules or compounds that produce a certain degree of beneficial effect upon administration to a subject. Beneficial effects include the possibility of diagnostic determination; improvement of disease, symptoms, disorders, or pathological conditions; reduction or prevention of the onset of disease, symptoms, disorders, or conditions; and generally, neutralization of disease, symptoms, disorders, or pathological conditions.
[0160] In some embodiments of the various methods described herein, the sample is derived from a subject. The subject is any living organism, and non-limiting examples include plants, animals, fungi, protists, Monera, viruses, mitochondria, and chloroplasts. The sample polynucleotide is isolated from the subject, such as a cell sample, tissue sample, bodily fluid sample, or organ sample (or a cell culture derived from any of these), and examples include cultured cell lines, biopsies, blood samples, oral mucosal specimens, or fluid samples containing cells (e.g., saliva). In some cases, the sample does not contain intact cells and is treated to remove cells, or the polynucleotide is isolated without using a cell extraction step (e.g., isolating cell-free polynucleotides such as cell-free DNA). Other examples of sample sources include blood, urine, feces, nasal passages, lungs, intestines, other bodily fluids or excretions, those obtained therefrom, or those derived from combinations thereof. The subjects are animals, but are not limited to, cattle, pigs, mice, rats, birds, cats, dogs, etc., and are usually mammals such as humans. In some embodiments, the sample includes tumor cells in a sample of tumor tissue from the subject. In some embodiments, the sample is a blood sample or a portion thereof (e.g., plasma or serum). Serum and plasma may be of particular interest due to the relative enrichment of tumor DNA associated with a high rate of malignant cell death between such tissues. The sample may be a fresh sample or a sample that has been subjected to one or more storage processes (e.g., paraffin-embedded samples, in particular formalin-fixed paraffin-embedded (FFPE) samples). In some embodiments, a sample from one individual is divided into multiple distinct samples (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more distinct samples) that are independently subjected to the methods of this disclosure, such as analysis repeated two, three, four, or more times. If the sample is derived from the subject, the reference sequence may also be derived from the subject, such as the consensus sequence of the sample being analyzed, or a polynucleotide sequence from another sample or tissue of the same subject.For example, a blood sample may be analyzed for ctDNA mutations, while intracellular DNA from another sample (e.g., a buccal or skin sample) is analyzed to determine a reference sequence.
[0161] Polynucleotides can be extracted from a sample with or without extraction from cells in the sample, according to an appropriate method. Various kits are available for polynucleotide extraction, which may depend on the type of sample, or for the type of nucleic acid to be isolated. Examples of extraction methods are provided herein, such as those described for any of the various embodiments disclosed herein. In one example, the sample may be a blood sample, such as a sample collected in an EDTA tube (e.g., a BD Vacutainer). Plasma can be separated from peripheral blood cells by centrifugation (e.g., 1900 × g for 10 minutes at 4°C). Plasma separation performed in this manner on a 6 mL blood sample typically yields 2.5 to 3 mL of plasma. Circulating cell-free DNA may be extracted from the plasma sample, for example, by using the QIAmp Circulating Nucleic Acid Kit (Qiagene), according to the manufacturer's protocol. The DNA can then be quantified (e.g., on an Agilent 2100 Bioanalyzer with High Sensitivity DNA kit (Agilent)). For example, the yield of circulating DNA in such plasma samples from healthy individuals ranges from 1 ng to 10 ng per mL of plasma, and is significantly higher in samples from cancer patients.
[0162] Polynucleotides can also originate from stored, frozen, or archived samples. One common method for storing samples is formalin fixation and paraffin embedding. However, this process is also associated with nucleic acid degradation. Polynucleotides processed and analyzed from FFPE samples may include short polynucleotides, such as fragments in the range of 50–200 base pairs or shorter. Numerous techniques exist for the purification of nucleic acids from fixed paraffin-embedded samples, including those described in WO2007133703, and methods described in Foss, et al Diagnostic Molecular Pathology, (1994) 3:148–155 and Paska, C., et al Diagnostic Molecular Pathology, (2004) 13:234–240. Commercially available kits may also be used to purify polynucleotides from FFPE samples, such as Ambion's Recoverall Total Nucleic acid Isolation kit. A typical method begins with the step of removing paraffin from the tissue via extraction with xylene or other organic solvents, followed by heat and treatment with protease-like proteinase K, which helps cleave the tissue and proteins and release genomic material from the tissue. The released nucleic acids can then be captured on a membrane or precipitated from the solution and washed to remove impurities, and in the case of mRNA isolation, a DNase treatment step is sometimes added to degrade unwanted DNA. Other methods for extracting FFPE DNA are available and can be used in the method of this disclosure.
[0163] In some embodiments, the polynucleotides include cell-free polynucleotides such as cell-free DNA (cfDNA), cell-free RNA (cfRNA), circulating tumor DNA (ctDNA), or circulating tumor RNA (ctRNA). Cell-free DNA circulates in both healthy and diseased individuals. Cell-free RNA circulates in both healthy and diseased individuals. cfDNA from tumors (ctDNA) is considered a common finding across different malignant lesions, although it is not limited to any specific cancer type. According to several measurements, the concentration of free circulating DNA in plasma is approximately 14–18 ng / ml in control subjects and approximately 180–318 ng / ml in patients with tumorigenesis. Apoptosis and necrotic cell death are attributed to cell-free circulating DNA in body fluids. For example, significantly increased circulating DNA levels have been observed in the plasma of prostate cancer patients and patients with other prostate diseases such as benign benign prostatic hyperplasia and prostatitis (Prostatits). In addition, circulating tumor DNA is present in fluids originating from organs where primary tumors arise. Therefore, breast cancer detection can be achieved through ductal lavage; colorectal cancer detection in stool; lung cancer detection in saliva; and prostate cancer detection in urine or semen. Cell-free DNA can be obtained from various sources. One common source is a subject's blood sample. However, cfDNA or other fragmented DNA can also originate from various other sources. For example, urine and fecal samples can be sources of cfDNA, including ctDNA. Cell-free RNA can also be obtained from various sources.
[0164] In some embodiments, polynucleotides are subjected to subsequent processes (e.g., cyclization and amplification) that do not involve extraction and / or purification. For example, a fluid sample may be treated to remove cells without an extraction step to produce a purified liquid sample and a cell sample, and then to remove DNA from the purified fluid sample. Various procedures for the isolation of polynucleotides are available, including precipitation or nonspecific binding to a substrate, and subsequent washing of the substrate to release the bound polynucleotides. When polynucleotides are isolated from a sample without a cell extraction step, the majority of the polynucleotides are extracellular or “cell-free” polynucleotides. For example, cell-free polynucleotides may include cell-free DNA (also referred to as “circulating” DNA). In some embodiments, circulating DNA is circulating tumor DNA (ctDNA) derived from tumor cells from body fluids or excretions (e.g., blood samples). Cell-free polynucleotides may include cell-free RNA (also referred to as “circulating” RNA). In some embodiments, circulating RNA is circulating tumor RNA (ctRNA) derived from tumor cells. Tumors frequently undergo apoptosis or necrosis, thereby releasing tumor nucleic acids into the body, including the subject's bloodstream, through various mechanisms in different forms and at different levels. Typically, ctDNA sizes can range from higher concentrations of smaller fragments, usually 70 to 200 nucleotides in length, to lower concentrations of larger fragments up to several thousand kilobases.
[0165] In some embodiments of any of the various aspects described herein, detecting sequence variants involves detecting mutations (e.g., rare somatic mutations) relative to a reference sequence or in a background without mutations, where the sequence variant is associated with a disease. Generally, sequence variants for which there is statistical, biological, and / or functional evidence of association with a disease or trait are referred to as “causative gene variants.” A single causative gene variant may be associated with one or more diseases or traits. In some embodiments, causative gene variants may be associated with Mendelian traits, non-Mendelian traits, or both. Causative gene variants may manifest as polynucleotide variations, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more sequence differences (e.g., between a polynucleotide containing a causative gene variant and a polynucleotide lacking the causative gene variant at the same relative genomic location). Non-limiting examples of causative gene variant types include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number polymorphisms (CNVs), short tandem repeats (STRs), restriction fragment length polymorphisms (RFLPs), simple repeat sequences (SSRs), tandem repeats of varying lengths (VNTRs), random amplification polymorphic DNA (RAPDs), amplification fragment length polymorphisms (AFLPs), retrotransposon-inter-amplification polymorphisms (IRAPs), long dispersed intranuclear repeat sequences and dispersed intranuclear repeat sequences (LINE / SINEs), long tandem repeats (LTRs), mobile elements, retrotransposon-microsatellite amplification polymorphisms, retrotransposon-based insertion polymorphisms, sequence-specific amplification polymorphisms, and heritable epigenetic modifications (e.g., DNA methylation). A causative gene variant can also be a set of closely related causative gene variants. Some causative gene variants may also affect as sequence variations in RNA polynucleotides. At this level, some causative gene variants are also indicated by the presence or absence of a particular type of RNA polynucleotide. Furthermore, several causative gene variants result in sequence modifications in protein polypeptides. Numerous causative gene variants have been reported. One example of a causative gene variant that is a SNP is the Hb S variant of hemoglobin, which causes sickle cell anemia.An example of a causative gene variant that is DIP is the delta 508 mutation in the CFTR gene, which causes cystic fibrosis. An example of a causative gene variant that is CNV is trisomy 21, which causes Down syndrome. An example of a causative gene variant that is STR is the tandem repeat, which causes Huntington's disease. A non-limiting list of causative gene variants and the diseases they are associated with is provided in Table 1. Additional non-limiting lists of causative gene variants are described in WO2014015084. Further examples of genes in which mutations are associated with a disease and sequence variants can be detected according to the methods of this disclosure are provided in Table 2.
[0166] [Table 1-1]
[0167] [Table 1-2]
[0168] [Table 1-3]
[0169] [Table 1-4]
[0170] [Table 1-5]
[0171] [Table 1-6]
[0172] [Table 1-7]
[0173] Table 1-8
[0174] Table 1-9
[0175] Table 1-10
[0176] Table 1-11
[0177] Table 1-12
[0178] Table 1-13
[0179] Table 1-14
[0180] Table 1-15
[0181] Table 1-16
[0182] Table 1-17
[0183] Table 1-18
[0184] Table 1-19
[0185] Table 1-20
[0186] Table 1-21
[0187] Table 1-22
[0188] Table 1-23
[0189] Table 1-24
[0190] Table 1-25
[0191] Table 2
[0192] In some embodiments, the method further includes a step of diagnosing a subject based on a calling step, such as diagnosing a subject with a disease associated with a detected causative gene variant, or reporting the possibility that a patient has or is likely to develop such a disease. Examples of diseases, associated genes, and associated sequence variants are provided herein. In some embodiments, the results are reported via a reporting device such as those described herein.
[0193] In some embodiments, one or more causative gene variants are sequence variants associated with a specific type or stage of cancer, or cancer with specific characteristics (e.g., metastatic potential, drug resistance, drug responsiveness). In some embodiments, the disclosure provides methods for prognosis, such as those in which specific mutations are known to be associated with patient outcomes. For example, ctDNA has been shown to be a superior biomarker for breast cancer than conventional cancer antigen 53 (CA-53) and a list of circulating tumor cells (see, e.g., Dawson, et al., N Engl J Med 368:1199 (20 13)). In addition, the methods of the disclosure can be used in the progression of cancer treatment and clinical trials, as well as in therapeutic decisions, guidance, and monitoring. For example, treatment effectiveness can be monitored by comparing patient ctDNA samples before, during, and after treatment with specific therapies, such as molecular targeted therapy (monoclonal drugs), chemotherapy drugs, radiotherapy protocols, or combinations thereof. For example, ctDNA can be monitored to determine whether a particular mutation increases or decreases, or whether new mutations appear, after a procedure, which would allow a physician to modify the procedure (e.g., continue, stop, or change the procedure) in a much shorter timeframe than would be possible with monitoring methods that track the patient's symptoms. In some embodiments, the method further includes a step of diagnosing a subject based on a calling step, such as diagnosing a subject with a particular stage or type of cancer associated with the detected sequence variant, or reporting that the patient has or is likely to develop such cancer.
[0194] For example, for therapies specifically targeted to a patient based on molecular markers (e.g., Herceptin and HER2 / neu status), patients may be tested to find out if specific mutations are present in their tumors. Such mutations can be used to predict response to or resistance to the therapy and to guide the decision of whether or not to use the therapy. Therefore, detecting and monitoring ctDNA during treatment can be very useful in guiding the selection of treatment. Several primary (pre-treatment) or secondary (post-treatment) cancer mutations have been shown to be responsible for cancer resistance to certain therapies (Misale et al., Nature 486(7404):532 (2012)).
[0195] Various sequence variants are known that are associated with one or more types of cancer, which may be useful in determining diagnosis, prognosis, or treatment. Appropriate target sequences of oncological significance for use in the methods of this disclosure include, but are not limited to, modifications in the TP53 gene, ALK gene, KRAS gene, PIK3CA gene, BRAF gene, EGFR gene, and KIT gene. Target sequences that can be specifically amplified and / or specifically analyzed for sequence variants may be all or some of the cancer-related genes. In some embodiments, one or more sequence variants are identified in the TP53 gene. TP53 is one of the most frequently mutated genes in human cancers; for example, TP53 mutations are found in 45% of ovarian cancers, 43% of colorectal cancers, and 42% of upper respiratory tract and gastrointestinal cancers (see, e.g., M. Olivier, et, al. TP53 Mutations in Human Cancers: Origins, Consequences, and Clinical Use. Cold Spring Harb Perspect Biol. 2010 January; 2(1)). Characterizing TP53 mutational states can aid in clinical diagnosis, provide prognostic values, and influence the treatment of cancer patients. For example, TP53 mutations may be used as predictors of poor prognosis in patients with glial cell-derived CNS tumors and as predictors of rapid disease progression in patients with chronic lymphocytic leukemia (see, e.g., McLendon RE, et al. Cancer. 2005 Oct 15; 1 04(8): 1693-9; Dicker F, et al. Leukemia. 2009 Jan; 23(1): 117-24). Sequence variations can occur anywhere within a gene. Therefore, all or part of the TP53 gene can be evaluated herein. That is, as described throughout this specification, when target-specific components (e.g., target-specific primers) are used, multiple TP53-specific sequences can be used to amplify and detect fragments extending into the gene, rather than simply one or more selected subsequences (such as mutation "hotspots"), which may be used for a selected target, for example.Alternatively, target-specific primers may be designed to hybridize upstream or downstream of one or more selected subsequences (such as nucleotides or nucleotide regions associated with an increased rate of mutation in a subject classification, as included by the term “hotspot”). Standard primers extending to such subsequences may be designed, and / or B2B primers may be designed to hybridize upstream or downstream of such subsequences.
[0196] In some embodiments, one or more sequence variants are identified in all or part of the ALK gene. ALK fusions have been reported in as many as 7% of lung tumors, some of which are associated with EGFR tyrosine kinase inhibitor (TKI) resistance (see, e.g., Shaw et al., J Clin Oncol. Sep 10, 2009; 27(26): 4247-4253). Up to 2013 different point mutations across the entire ALK tyrosine kinase domain have been found in patients with secondary resistance to ALK tyrosine kinase inhibitors (TKIs) (Katayama R 2012 Sci Transl Med. 2012 Feb 8;4(120)). Therefore, mutation detection in the ALK gene can be used to assist in cancer treatment decisions.
[0197] In some embodiments, one or more sequence variants are identified in all or part of the KRAS gene. Approximately 15–25% of lung adenocarcinoma patients and 40% of colorectal cancer patients have been reported to have tumor-associated KRAS mutations (see, e.g., Neuman 2009, Pathol Res Pract. 2009;205(12):858–62). The majority of mutations are located at codons 12, 13, and 61 of the KRAS gene. These mutations activate the KRAS signaling pathway, which causes tumor cell growth and proliferation. Some studies have shown that patients with tumors carrying KRAS mutations may benefit from anti-EGFR antibody therapy alone or in combination with chemotherapy (see, e.g., Amado et al. 2008 J Clin On col. 2008 Apr 1;26(1 0):1626-34, Bokemeyer et al. 2009 J Clin Oncol. 2009 Feb 10;27(5):663-71). One specific “hotspot” for sequence variations that can be targeted to identify sequence variants is located at gene position 35. Identifying KRAS sequence variants can be used in treatment selection, such as treatment selection for colorectal cancer subjects.
[0198] In some embodiments, one or more sequence variants are identified in all or part of the PIK3CA gene. Somatic mutations in PIK3CA are frequently found in various types of cancer, e.g., 10–30% of colorectal cancers (see, e.g., Samuels et al. 2004 Science. 2004 Apr 23;304(5670):554). These mutations are most commonly located in two “hotspot” regions within exon 9 (helical domain) and exon 20 (kinase domain), which can be specifically targeted for amplification and / or analysis of the detected sequence variants. Position 3140 can also be specifically targeted.
[0199] In some embodiments, one or more sequence variants are identified in all or part of the BRAF gene. Approximately 50% of all malignant melanomas have been reported to carry somatic mutations in BRAF (see, e.g., Maldonado et al., J Natl Cancer Inst. 2003 Dec 17;95(24):1878-90). BRAF mutations have been found in all melanoma subtypes, but were most frequently found in melanomas of non-injured skin origin, induced by chronic sun exposure. The most common BRAF mutation in melanoma is the missense mutation V600E, which substitutes valine at position 600 with glutamine. BRAF V600E mutations are associated with the clinical efficacy of BRAF inhibitor therapy. Detection of BRAF mutations can be used in the selection of treatments for melanoma and in the study of resistance to targeted therapies.
[0200] In some embodiments, one or more sequence variants are identified in all or part of the EGFR gene. EGFR mutations are frequently associated with non-small cell lung cancer (approximately 10% in the US and 35% in East Asia; see, e.g., Pao et al., Proc Natl Acad Sci US A. 2004 Sep 7;101(36):13306-11). These mutations typically occur within EGFR exons 18-21 and are usually heterozygous. Approximately 90% of these mutations are exon 19 deletions or exon 21 L858R point mutations.
[0201] In some embodiments, one or more sequence variants are identified in all or part of the KIT gene. Approximately 85% of gastrointestinal stromal tumors (GISTs) have been reported to have KIT mutations (see, e.g., Heinrich et al. 2003 J Clin Oncol. 2003 Dec I;21 (23):4342-9). The majority of KIT mutations are found in the near-membrane domain (exon 11, 70%), the extracellular dimerization motif (exon 9, 10-15%), the tyrosine kinase I (TKI) domain (exon 13, 1-3%), and the tyrosine kinase 2 (TK2) domain and activation loop (exon 17, 1-3%). Secondary KIT mutations are generally identified after the targeted therapeutic agent imatinib and after patients have developed resistance to the treatment.
[0202] Further non-limiting examples of all or some genes associated with cancer, whose sequence variants can be analyzed according to the methods described herein, include, but are not limited to, PTEN;ATM;ATR;EGFR;ERBB2;ERBB3;ERBB4;Notch1;Notch2;Notch3;Notch4;AKT;AKT2;AKT3;HIF;HIF1a;HIF3a;Met;HRG;Bcl2;PPARα;PPARγ;WT1 (Wilms tumor);FGF receptor family members (5 members: 1, 2, 3, 4, 5);CDKN2a;APC;RB (retinoblastoma);MEN1;VHL;BRCA1;BRCA2;AR;(androgen receptor);TSG101;IGF;IGF receptor;Igf1 (4 variants);Igf2 (3 variants);Igf 1 receptor;Igf This includes two receptors; Bax; Bcl2; the caspase family (nine members: 1, 2, 3, 4, 6, 7, 8, 9, 12); Kras; and Apc. Further examples are provided throughout this specification. Examples of cancers that can be diagnosed based on calling one or more sequence variants according to the methods disclosed herein include, but are not limited to, acanthoma, acinar cell tumor, acoustic neuroma, acral lentiginous melanoma, acromyelitis, acute eosinophilic leukemia, acute lymphoblastic leukemia, acute megakaryoblastic leukemia, acute monocytic leukemia, mature acute myeloblastic leukemia, acute myeloid dendritic cell leukemia, acute myeloid leukemia, acute promyelocytic leukemia, adamantinoma, adenocarcinoma, adenoid cystic carcinoma, adenoma, adenoid odontogenic tumor, adrenocortical carcinoma, adult T-cell leukemia, invasive NK-cell leukemia, AIDS-related cancer, AIDS-related lymphoma, alveolar soft part sarcoma, ameloblastic fibroma, anal cancer, Anaplastic large cell lymphoma, histoplastic thyroid cancer, angioimmunoblastic T-cell lymphoma, angiomyolipoma, angiosarcoma, appendiceal cancer, astrocytoma, atypical teratomatoid rhabdoid tumor, basal cell carcinoma, basaloid carcinoma, B-cell leukemia, B-cell lymphoma, Bellini duct carcinoma, biliary tract cancer, bladder cancer, blastoma, bone cancer, bone tumor, brainstem glioma, brain tumor, breast cancer, Brenner tumor, bronchial tumor, bronchioloalveolar carcinoma, pheochromocytoma, Burkitt lymphoma, cancer of unknown primary site, carcinoid tumor, carcinoma, carcinoma in situ, penile carcinoma, carcinoma of unknown primary site, carcinosarcoma, Castlman's disease, central nervous system germ cell tumor, cerebellar astrocytoma, cerebral astrocytoma, cervical cancer, intrahepatic cholangiocarcinoma, chondroma, chondrosarcoma, chordoma,Choriocarcinoma, choroid plexus papilloma, chronic lymphocytic leukemia, chronic monocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, chronic neutrophilic leukemia, clear cell tumor, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, Degos disease, dermatofibrosarcoma protuberans, dermoid cyst, fibrinogenic round cell tumor, diffuse large B-cell lymphoma, germinal dysplastic neuroepithelial tumor, embryonal carcinoma, endodermal sinus tumor, endometrial cancer, endometrial uterine cancer, endometrial tumor, intestinal disease-related T-cell lymphoma, ependymoblastoma, ependymoma, epithelioid sarcoma, erythroleukemia, esophageal cancer, sensory neuroblastoma, Ewing family of tumors, Ewing family sarcoma, Ewing's sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic cholangiocarcinoma, extramammary Paget's disease, Fallopian duct carcinoma, inclusion malformed fetus, fibroma, fibrosarcoma, follicular lymphoma, follicular thyroid cancer, gallbladder cancer, gallbladder cancer, neuronal glioma, ganglionoma, gastric cancer, gastric lymphoma, digestive system cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, germ cell tumor, germ cell tumor, gestational trophoblastic carcinoma, giant cell tumor of bone, glioblastoma multiforme, glioma, gliomatosis, glomus tumor, glucagonoma, gonadoblastoma, granulosa cell tumor, hairy cell leukemia Hairy cell leukemia, head and neck cancer, heart cancer, hemangioblastoma, hemangiopericytoma, angiosarcoma, hematological malignancies, hepatocellular carcinoma, hepatosplenic T-cell lymphoma, hereditary breast and ovarian cancer syndrome, Hodgkin lymphoma, hypopharyngeal cancer, hypothalamic glioma, inflammatory breast cancer, intraocular melanoma, islet cell carcinoma, islet cell tumor, juvenile myelomonocytic leukemia, Kaposi's sarcoma, kidney cancer, Kratzkin tumor, Kruckenberg tumor, laryngeal cancer, lentigo malignant melanoma, leukemia, lip cancer and oral cancer, liposarcoma, lung cancer, luteal malformation, lymphangioma, lymphangiosarcoma, lymphoepithelioma, lymphoblastoma Pascal leukemia, lymphoma, macroglobulinemia, malignant fibrous histiocytoma, malignant fibrous histiocytoma of bone, malignant glioma, malignant mesothelioma, malignant peripheral nerve sheath tumor, malignant rhabdoid tumor, malignant Triton tumor, MALT lymphoma, mantle cell lymphoma, mast cell leukemia, mediastinal germ cell tumor, mediastinal tumor, bone marrow thyroid cancer, medulloblastoma, medullary epithelioma, melanoma, melanoma, meningioma, Merkel cell tumor, mesothelioma, mesothelioma, metastatic cervical squamous cell carcinoma of unknown primary origin, metastatic urothelial carcinoma, Müllerian mixed tumor, monocytic leukemia, oral cancer, mucinous tumor, multiple endocrine neoplasia syndrome,Multiple myeloma, multiple myeloma, mycosis fungoides, myelodysplastic disorder, myelodysplastic syndrome, myeloid leukemia, myeloid sarcoma, myeloproliferative disorder, myxoma, nasal cavity cancer, nasopharyngeal cancer, nasopharyngeal cancer, neoplasm, neurinoma, neuroblastoma, neurofibroma, neuroma, nodular melanoma, non-Hodgkin lymphoma, non-melanoma skin cancer, non-small cell lung cancer, ocular tumor, oligodendronoma, oligodendronoma, pallocyte tumor, optic nerve sheath meningioma, oral cancer, oral cancer, oral pharyngeal cancer Head cancer, osteosarcoma, osteosarcoma, ovarian cancer, ovarian epithelial carcinoma, ovarian germ cell tumor, low-grade ovarian cancer, Paget's disease of the breast, Pancoast tumor, pancreatic cancer, pancreatic cancer, papillary thyroid carcinoma, papillomatosis, paraganglioma, paranasal sinus cancer, parathyroid cancer, penile cancer, perivascular epithelioid cell tumor, pharyngeal cancer, chromaffin cell tumor, intermediate pineal parenchymal tumor, pineoblastoma, pituitary cell tumor, pituitary adenoma, pituitary tumor, plasma cell tumor, pleuropulmonary blastoma, polygermoma, precursor T-lymphoblastic lymphoma, primary central nervous system Lymphoma, primary exudative lymphoma, primary hepatocellular carcinoma, primary liver cancer, primary peritoneal cancer, primitive neuroepithelial tumor, prostate cancer, pseudomyxoma peritonei, rectal cancer, renal cell carcinoma, airway carcinoma involving the NUT gene on chromosome 15, retinoblastoma, rhabdomyoma, rhabdomyosarcoma, Richter transformation, sacrococcygeal teratoma, salivary gland carcinoma, sarcoma, schwannomatosis, sebaceous gland carcinoma, secondary neoplasm, seminomas, serous tumors, Sertoli-Leydig cell tumor, genital cord-stromal tumor, Sézary syndrome, ring cell carcinoma of Signet Skin cancer, small round blue cell tumor, small cell carcinoma, small cell lung cancer, small cell lymphoma, small intestine cancer, soft tissue sarcoma, somatostatin-producing tumor, sooty verruca, spinal cord tumor, vertebral tumor, splenic marginal zone lymphoma, squamous cell carcinoma, gastric cancer, superficial spreading melanoma, supratentorial primitive neuroepithelial tumor, surface epithelial stromal tumor, synovial sarcoma, T-cell acute lymphoblastic leukemia, T-cell macrogranular lymphocyte leukemia, T-cell leukemia, T-cell lymphoma, T-cell pre-lymphocytic leukemia, teratoma, peripheral lymphadenoma (Terminal Lymphatic cancer, testicular cancer, theca, laryngeal cancer, thymic cancer, thymoma, thyroid cancer, transitional cell carcinoma of the renal pelvis and ureter, transitional cell carcinoma, urachal cancer, ureteral cancer, urogenital tumors, uterine sarcoma, uveal melanoma, vaginal cancer, Verner-Morrison syndrome, verrucous carcinoma, optic nerve glioma, vulvar cancer, Valdenström macroglobulinemia, Warthin's tumor, Wilms' tumor,and combinations thereof. Non-exclusive examples of specific sequence variants associated with cancer are provided in Table 3.
[0203] [Table 3-1]
[0204] [Table 3-2]
[0205] [Table 3-3]
[0206] In addition, the methods and compositions disclosed herein may be useful in discovering novel and rare mutations associated with one or more cancer types, stages, or cancer features. For example, a population of individuals sharing a feature under analysis (e.g., a specific disease, cancer type, or cancer stage) may be subjected to a sequence variant detection method according to this disclosure to identify sequence variants or types of sequence variants (e.g., mutations in a specific gene or portion of a gene). Sequence variants identified occurring in a population of individuals sharing the feature at a statistically significantly greater frequency than in individuals without the feature may be assigned some degree of association with that feature. The sequence variants or types of sequence variants thus identified may then be used in diagnosing or treating individuals found to be carrying them.
[0207] Other therapeutic uses include use in non-invasive fetal diagnostic methods. Fetal DNA can be found in the blood of a pregnant woman. The methods and compositions described herein can be used to identify sequence variants in circulating fetal DNA and, therefore, may be used to diagnose one or more genetic diseases in the fetus, such as those associated with one or more causative gene variants. Non-exclusive examples of causative gene variants are described herein and include trisomal, cystic fibrosis, sickle cell anemia, and Tay-Sachs disease. In this embodiment, the mother may provide a control sample and a blood sample to be used for comparison. The control sample is a suitable tissue, typically a process for extracting intracellular DNA, which can later be sequenced to provide a reference sequence. The sequence of cfDNA corresponding to the fetal genomic DNA can then be identified as a sequence variant to the mother's reference. The father may also provide a reference sample to assist in the identification of the fetal sequence and sequence variants.
[0208] Further therapeutic applications include the detection of exogenous polynucleotides from pathogens (e.g., bacteria, viruses, fungi, and microorganisms) where the information can inform the selection of diagnostic and treatment methods. For example, several HIV subtypes correlate with drug resistance (see, e.g., hivdb.stanford.edu / pages / genotype-rx). Similarly, HCV classification, subclassification, and isotype mutations can also be performed using the methods and compositions of this disclosure. Furthermore, if an HPV subtype correlates with the risk of cervical cancer, such a diagnosis can further inform the assessment of cancer risk. Furthermore, non-exclusive examples of viruses that may be detected include: Hepadnaviridae hepatitis B virus (HBV), woodchuck hepatitis virus, squirrel (Hepadnaviridae) hepatitis virus, duck hepatitis B virus, heron hepatitis B virus, herpes simplex virus (HSV) types 1 and 2, varicella-zoster virus, cytomegalovirus (CMV), human cytomegalovirus (HCMV), mouse cytomegalovirus (MCMV), guinea pig cytomegalovirus (GPCMV), Epstein-Barr virus (EBV), human herpesvirus 6 (HHV variants A and B), human herpesvirus 7 (HHV-7), human herpesvirus 8 (HHV-8), Kaposi's sarcoma-associated herpesvirus (KSHV), B virus, poxvirus, vaccinia virus, smallpox virus, and staghorn fistula. Smallpox virus, cowpox virus, camelpox virus, ectromelia virus, mousepox virus, rabbitpox virus, raccoonpox virus, molluscum contagiosum virus, Orff virus, Milker nodule virus, bovine papular stomatitis virus, sheeppox virus, goatpox virus, Lumpy Skin disease virus, fowlpox virus, canarypox virus, pigeonpox virus, sparrowpox virus, myxoma virus, tulareviformis virus Rabbit fibroma virus, squirrel fibroma virus, swine pox virus, peach pox virus, yana pox virus, flavivirus, dengue virus, hepatitis C virus (HCV), hepatitis GB virus (GBV-A, GBV-B, and GBV-C), West Nile virus, yellow fever virus, St. Louis encephalitis virus, Japanese encephalitis virus, Poissan virus, tick-borne encephalitis virus, Kyasanur forest disease virus, togavirus,Venezuelan encephalitis (VEE) virus, Chikungunya virus, Ross River virus, Mayarovirus, Sindbis virus, rubella virus, retrovirus, human immunodeficiency virus (HIV) types 1 and 2, human T-cell leukocyte virus (HTLV) types 1, 2, and 5, mouse mammary tumor virus (MMTV), Roussarcoma virus (RSV), lentivirus, coronavirus, severe acute respiratory syndrome (SARS) virus, filovirus, Ebola virus, Marburg virus, metapneumovirus (MPV) including human metapneumovirus (HMPV), rhabdovirus, rabies virus, vesicular stomatitis virus, bunyavirus, Crimean-Congo hemorrhagic fever virus This includes, among others, Rift Valley fever virus, Lacrosse virus, Hantavirus, Orthomyxovirus, Influenza virus (types A, B, and C), Paramyxovirus, Parainfluenza virus (types PIV 1, 2, and 3), Respiratory rash virus (types A and B), Measles virus, Mumps virus, Arenavirus genus, Lymphochoroidal meningitis virus, Junin virus, Macpo virus, Guanalitovirus, Lassa fever virus, Ampari virus, Flexal virus, Yippee virus, Mobala virus, Mopeia virus, Latinovirus, Paranavirus, Pikindevirus, Puntatrovirus (PTV), Takaribe virus, and Tamiamivirus.
[0209] Examples of pathogens that can be detected by the methods disclosed herein, and specific examples of pathogens, include, but are not limited to, Acinetobacter baumannii, Actinabacillus sp., Actinomycetes, Actinomyces sp. (such as Actinomycetes japonica and Actinomycetes neslund's), Aeromonas sp. (such as Aeromonas hydrophylla, Aeromonas veroni subsubspecies sobria (Aeromonas sobria), and Aeromonas caviar), Anaplasma phagocytephyllum, Achromobacter xylosoxidance, Acinetobacter baumannii, and Actinobasila Bacillus actinomycetemucomitans, Bacillus sp. (Bacillus anthrasis, Bacillus cereus, Bacillus subtilis, Bacillus slingiensis, and Bacillus stearotermophilus, etc.), Bacteroides sp. (Bacteroides fragilis, etc.), Bartonella sp. (Bartonella basiliformis and Bartonella henselae, etc.), Bifidobacterium sp., Bordetella sp. (Bordetella patassis, Bordetella parapatassis, and Bordetella bronchiseptica, etc.), Borelli Aspergillus sp. (including Borrelia relapsing and Borrelia burgdorferi), Brucella sp. (including Brucella avoltus, Brucella canis, Brucella melitensis, and Brucella suisse), Burkholderia sp. (including Burkholderia pseudomalei and Burkholderia cepacia), Campylobacter sp. (including Campylobacter jejuni, Campylobacter coli, Campylobacter lari, and Campylobacter phytus), Capnositophaga sp., Cardiobacterium hominis, Lamidia trachomatis, Chlamydophila pneumoniae, Chlamydophila sitassi, Citrobacter sp., Coxiella brunetii, Corynebacterium sp. (Corynebacterium diphtheriae, Corynebacterium jakeum, and Corynebacterium, etc.), Clostridium sp. (Clostridium perfringens, Clostridium difficile, Clostridium botulinum, and Clostridium tetani, etc.), Eichenella collodens, Enterobacter sp.(Enterobacter aerogenes, Enterobacter agglomerans, Enterobacter cloacae, and opportunistic Escherichia coli, enterotoxigenic Escherichia coli, enteroinvasive Escherichia coli, pathogenic Escherichia coli, enterohemorrhagic Escherichia coli, enteroaggregative Escherichia coli, and Escherichia coli that causes urinary tract diseases, etc.), Enterococcus sp. (Enterococcus faecalis and Enterococcus faecium, etc.), Ehrlichia sp. (Ehrlichia shafensis and Ehrlichia canis, etc.), Erysipelas swine, Eubacterium sp., Francisella tularensis, Fusobacter Thelium nucleatum, Gardnerella vaginalis, Gemella morbilorum, Haemophilus sp. (Haemophilus influenzae, Haemophilus ducleyi, Haemophilus aegyptius, Haemophilus parainfluenza, Haemophilus haemolyticus, and Haemophilus parahemolyticus), Helicobacter sp. (Helicobacter pylori, Helicobacter cinedii, and Helicobacter phenellae, etc.), Kingella kingii, Klebsiella sp. (Klebsiella pneumoniae, Klebsiella granulosa) (Matisse and Klebsiella oxytoka, etc.), Lactobacillus sp., Listeria monocytogenes, Leptospira interlogans, Regenera pneumophila, Leptospira interlogans, Peptostreptococcus sp., Moraxella catarrhalis, Morganella sp., Mobiluncus sp., Micrococcus sp., Mycobacterium sp. (Mycobacterium leprae, Mycobacterium tuberculosis, Mycobacterium intracellulare, Mycobacterium avium, Mycobacterium bovis, and Mycobacterium Cobacterium marinum, etc.), Mycoplasma sp. (Mycoplasma pneumoniae, Mycoplasma hominis, and Mycoplasma genitalium, etc.), Nocardia sp. (Nocardia asteroides, Nocardia siriasigeorgica, and Nocardia brasiliensis, etc.), Neisseria sp. (Neisseria gonorea and Neisseria meningitidis, etc.), Pasteurella maltosida, Presiomonas sigeoides, Prevotella sp., Porphyromonas sp., Prevotella melaninogenica, Proteus sp.(Proteus vulgaris and Proteus mirabilis, etc.), Providencia sp. (Providencia alkaline fasciens, Providencia rettgeri, and Providencia stuartii, etc.), Pseudomonas erginosa, Propionibacterium acnes, Rhodococcus equii, Rickettsia sp. (Rickettsia rickettii, Rickettsia akari and Rickettsia prowatzekii, Orientia tsutsugamushi (formerly: Rickettsia tsutsugamushi) and Rickettsia cifi), Rhodococcus sp., Serratia marcescens, Stenotrophomonas maltophilia, Sal Monera sp. (Salmonella enterica, Salmonella cifida, Salmonella paracifi, Salmonella enteritidis, Salmonella choleresuis, and Salmonella tiphimurum, etc.), Serratia sp. (Seratia marcescens and Serratia lycfaciens, etc.), Shigella sp. (Shigella shigaensis, Shigella flexneri, Shigella boyi, and Shigella sonei, etc.), Staphylococcus sp. (Staphylococcus aureus, Staphylococcus epidermides, Staphylococcus haemolyticus, Staphylococcus saprophyticus, etc.), Streptococcus sp.(Streptococcus pneumoniae (e.g., chloramphenicol-resistant serotype 4 streptococcus pneumoniae, spectinomycin-resistant serotype 6B streptococcus pneumoniae, streptomycin-resistant serotype 9V streptococcus pneumoniae, erythromycin-resistant serotype 14 streptococcus pneumoniae, optohyne-resistant serotype 14 streptococcus pneumoniae, rifampicin-resistant serotype 18C streptococcus pneumoniae, tetracycline-resistant serotype 19F streptococcus pneumoniae, penicillin-resistant serotype 19F streptococcus pneumoniae, and trimethoprim-resistant serotype 23F streptococcus pneumoniae Nicole-resistant serotype 4 Streptococcus pneumoniae, spectinomycin-resistant serotype 6B Streptococcus pneumoniae, streptomycin-resistant serotype 9V Streptococcus pneumoniae, optohypin-resistant serotype 14 Streptococcus pneumoniae, rifampicin-resistant serotype 18C Streptococcus pneumoniae, penicillin-resistant serotype 19F Streptococcus pneumoniae, or trimethoprim-resistant serotype 23F Streptococcus pneumoniae), Streptococcus agalactie, Streptococcus mutans, Streptococcus pyogenes, Group A Streptococcus, Streptococcus pyogenes, Group B Streptococcus, Stre Ptococcus agalactie, Group C Streptococcus, Streptococcus anginosus, Streptococcus excimilis, Group D Streptococcus, Streptococcus bovis, Group F Streptococcus, and Streptococcus anginosus, Group G Streptococcus, etc.), Rat-bite Spirillum, Streptobacillus moniliform, Treponema sp. (Pintaflambenzia, Treponema pertenu, Treponema pallidum, and Treponema endemicum, etc.), Troferima howipli, Ureaplasma urearichium, Bayonella sp., Vibrio sp. (B This includes one or more (or a combination thereof) of the following: *Brio cholerae*, *Vibrio parahaemollicus*, *Vibrio vulnificus*, *Vibrio parahaemollicus*, *Vibrio vulnificus*, *Vibrio arginolicus*, *Vibrio mimicus*, *Vibrio horise*, *Vibrio fulvialis*, *Vibrio methiconicophye*, *Vibrio damsella*, and *Vibrio furnisii*, *Yersinia sp.* (such as *Yersinia enterocolitica*, *Yersinia pestis*, and *Yersinia pseudotuberculosis*), and *Xanthomonas maltophilia*.
[0210] In some embodiments, the methods and compositions of this disclosure are used in monitoring organ transplant recipients. Typically, polynucleotides from donor cells are found circulating in the background of polynucleotides from recipient cells. Levels of donor circulating DNA are usually stable once the organ has been well accepted, and a rapid increase in donor DNA (e.g., as frequency in a given sample) can be used as an early sign of transplant rejection. Treatment can be given at this stage to prevent transplant failure. Rejection of a donor organ has been shown to result in an increase in donor DNA in the blood; see Snyder et al., PNAS 108(15):6629 (2011). This disclosure offers a significant improvement in sensitivity over prior art in this field. In this embodiment, recipient control samples (e.g., oral mucosal specimens) and donor control samples can be used for comparison. The recipient sample can be used to provide its reference sequence, while the sequence corresponding to the donor genome can be identified as a sequence variant to that reference. Monitoring may include the step of obtaining samples (e.g., blood samples) from the recipient over a period of time. Early samples (e.g., within the first few weeks) can be used to establish a baseline for donor cfDNA fractionation. Later samples can be compared to the baseline. In some embodiments, an increase in donor cfDNA fractionation of approximately, or at least approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 250%, 500%, 1000%, or more may serve as an indicator that the recipient is in the process of reflecting the donor tissue.
[0211] In some embodiments, methods for detection are provided. [Examples]
[0212] The following examples are given for the purpose of illustrating various embodiments of the invention and are not intended to limit the invention in any way. Together with the methods described herein, these examples are representative and typical of preferred embodiments and are not intended to limit the scope of the invention. Variations and other uses encompassed within the spirit of the invention as defined by the claims will be anticipated by those skilled in the art.
[0213] Example 1: Preparation of a sequencing library for tandem repeats for mutation detection Starting with >10 ng of ~150 bp DNA fragments in 12 μL of water or 10 mM Tris-HCl pH 8.0, 2 μL of 10X CircLigase buffer mixture was added, and the mixture was heated to 95°C over 2 minutes and cooled on ice for 5 minutes. To this, 4 μL of 5 M betaine, 1 μL of 50 mM MnCl2, and 1 μL of CircLigase II were added. The reaction mixture was incubated at 60°C for at least 12 hours. Next, 2 μL of RCA primer mixture (to final concentrations of 50 nM and 5 nM, respectively) was added and mixed. The mixture was heated to 95°C over 2 minutes and cooled to 42°C over 2 hours. The CirLigation product was purified using the Zymo oligonucleotide purification kit. Following the manufacturer's instructions, 28 μL of water was added to 22 μL of CircLigation product for a total volume of 50 μL. This was mixed with 100 μL of oligo-binding buffer and 400 μL of ethanol. The mixture was rotated at >10,000 xg for 30 seconds, and the pass-through fraction was discarded. 750 μL of DNA wash buffer was added, and the column was then rotated at >10,000 xg for 30 seconds, the pass-through fraction was discarded, and the column was rotated at full speed for another minute. The column was transferred to a new Eppendorf tube and eluted with 17 μL of water (the final eluted volume was approximately 15 μL).
[0214] Rolling circle amplification was performed on approximately 50 μL of sample. To 15 μL of elution sample, 5 μL of 10X RepliPHI buffer (Epicentre), 1 μL of 25 mM dNTPs, 2 μL of 100 mM DTT, 1 μL of 100 U / μL RepliPHI Phi29, and 26 μL of water were added. The reaction mixture was incubated at 30°C for 1 hour. The RCA product was purified by adding 80 μL of Ampure beads according to the manufacturer's instructions for the remaining washing step. For elution, 22.5 μL of elution buffer was added and the beads were incubated at 65°C for 5 minutes. After a brief rotation, the tube was returned to the magnet.
[0215] Approximately 20 μL of eluted product from the RCA reaction was mixed with 25 μL of 2X Phusion Master mix, 2.5 μL of DMSO, and 0.5 μL of 10 μM each B2B primer mixture. Amplification was performed using the following PCR program: 5 expansion cycles at 95°C for 1 minute (15 seconds at 95°C, 15 seconds at 55°C, 1 minute at 72°C), 13–18 replication cycles (15 seconds at 95°C, 15 seconds at 68°C, 1 minute at 72°C), and a final expansion at 72°C for 7 minutes. The size of the PCR product was confirmed by running an E-gel. If the range was 100–500 bp, the 0.6X Ampure beads were purified and enriched to 300–500 bp to obtain 1–2 ng for another PCR with small RNA library adapter primers. If the product size range was >1000 bp, the product was generated with 1.6X Ampure beads, yielding 2-3 ng for Nextera XT amplicon library preparation, and then enriched in size to the 400-1000 bp range by 0.6X Ampure bead purification.
[0216] To perform bioinformatics on sequencing data, FASTQ files were obtained from MiSeq runs. The sequences were aligned in the FASTQ files to reference genome sequences containing targeted sequences (e.g., KRAS and EGFR) using BWA. The region and length of repeat units and their reference locations were identified for each sequence (both reads) using the alignment results. Variants at all loci were identified using the alignment results and information for the repeat units of each sequence. The results of two reads were combined. The standardized frequency and noise level of the variants were calculated. Several further criteria were applied to variant calls from confirmed variants, including a q score > 30 and a p value < 0.0001. Confirmed variants that passed these criteria were reported as true variants (mutations). The process can be automated using a computer language (e.g., Python).
[0217] Example 2: Creation of a tandem repeat sequencing library for detecting sequence variants 10 ng of DNA fragments with an average length of 150 bp in a 12 μL volume were used to construct a tandem repeat sequencing library. The DNA was pre-treated with T4 polynucleotide kinase (New England Biolabs) to add a phosphate group at the 5' end and leave a hydroxyl group at the 3' end. For DNA fragments generated by DNase I or enzymatic fragmentation, or extracted from serum or plasma, the endpoint treatment step was skipped. The DNA was mixed with 2 μL of 10X CircLigase buffer (Epicentre CL9021K). The mixture was heated to 95°C over 2 minutes, cooled on ice for 5 minutes, and then 4 μL of betaine, 1 μL of 50 mM MnCl2, and 1 μL of CircLigase II (Epicentre CL9021K) were added. The ligation reaction was carried out at 60°C for at least 12 hours. Each RCA primer mixture (up to a final concentration of 10 nM) was added to the ligation product in 1 μL at 200 nM and mixed. The mixture was then heated to 96°C over 1 minute, cooled to 42°C, and incubated at 42°C for 2 hours.
[0218] CircLigation products containing hybridized RCA primers were purified using the Zymo Oligonucleotide Purification Kit (Zymo Research, D4060). To do this, 21 μL of the product was diluted to 50 μL with 28 μL of water and 1 μL of carrier RNA (Sigma-Aldrich, R5636, diluted to 200 ng / μL with 1X TE buffer). The diluted sample was mixed with 100 μL of oligobinding buffer and 400 μL of 100% ethanol. The mixture was placed on a column and centrifuged at >10,000 xg for 30 seconds. The pass-through fraction was discarded. The column was washed with 750 μL of DNA washing buffer by centrifugation at >10,000 xg for 30 seconds, discarding the pass-through fraction, and then centrifugation at full speed for a further 1 minute. The column was transferred to a new 1.5 mL Eppendorf tube, and the DNA was eluted with 17 μL of elution buffer (10 mM Tris=Cl pH 8.0, final elution volume approximately 15 μL).
[0219] 5 μL of 10X RepliPHI buffer, 2 μL of 25 mM dNTPs, 2 μL of 100 mM DTT, 1 μL of 100 U / μL RepliPHI Phi29, and 25 μL of water (Epicentre, RH040210) were added to 15 μL of eluted sample from the column for a total reaction volume of 50 μL. The reaction mixture was incubated at 30°C for 2 hours. The RCA product was purified by adding 80 μL of Ampure XP beads (Beckman Coulter, A63881). The washing process followed the manufacturer's instructions. The RCA product was eluted after incubation at 65°C for 5 minutes in 22.5 μL of elution buffer. The tube was briefly centrifuged before returning to the magnet.
[0220] Approximately 20 μL of eluted product from the RCA reaction was mixed with 25 μL of 2X Phusion Master mix (New England Biolabs M0531S), 2.5 μL of water, 2.5 μL of DMSO, and 0.5 μL of B2B primer (10 μM each). Amplification was performed with the following thermal cycling program: 2 minutes at 95°C, 5 cycles of expansion (30 seconds at 95°C, 15 seconds at 55°C, 1 minute at 72°C), 18 cycles of replication (15 seconds at 95°C, 15 seconds at 68°C, 1 minute at 72°C), and a final expansion for 7 minutes at 72°C. The size of the PCR product was confirmed by electrophoresis. Once the long PCR product was confirmed by electrophoresis, the PCR product was mixed with 30 μL of Ampure beads (0.6X volume) for purification to enrich for PCR products >500 bp. The purified product was quantified on a Qubit 2.0 Quantification Platform (Invitrogen). Approximately 1 ng of purified DNA was used for the Nextera XT Amplicon library preparation (Illumina FC-131-1024). Library elements with an insert size >500 bp were enriched by purification with 0.6X Ampure beads.
[0221] The concentration and size distribution of the amplified library were analyzed using the Agilent DNA High Sensitivity Kit for the 2100 Bioanalyzer (Agilent Technologies Inc., Santa Clara, CA). Sequencing was performed using an Illumina MiSeq with a 2 - 250 bp MiSeq sequencing kit. According to the MiSeq manual, a 12 pM denatured library was loaded onto the sequencing run.
[0222] In a variation of this procedure, Illumina adapters were used in the library preparation instead of the Nextera preparation. To do this, approximately 1 ng of similarly purified DNA was used for PCR amplification with a pair of primers that included the universal portion of the B2B primer and the Illumina adapter sequences (P5 and P7; 5’CAAGCAGAAGACGGCATACGA3’ and 5’ACACTCTTTCCCTACACGACGCTCTTCCGATCT3’). Using Phusion Master Mix, a 12-cycle replication process (30 seconds at 95°C, 15 seconds at 55°C, 60 seconds at 72°C) was performed. The purpose of this amplification step was to add the Illumina adapter for amplicon sequencing. Amplicons >500 bp in length were enriched with 0.6X Ampure beads. The concentration and size distribution of the amplicon library were analyzed using the Agilent DNA High Sensitivity Kit for the 2100 Bioanalyzer (Agilent Technologies Inc., Santa Clara, CA). Sequencing was performed using an Illumina MiSeq with the 2×250 bp MiSeq sequencing kit. The universal portion of the B2B primer also functions as the sequencing primer sequence, and a custom sequencing primer is added if the primer is not included in the Illumina kit. A 12 pM denatured library was loaded onto the sequencing run.
[0223] The coverage of the target region in one example of the analysis is illustrated in Figure 33. Table 6 below describes the results regarding the analysis of the target region.
[0224] Table 4 provides examples of RCA primers useful in the methods of the present disclosure. Table 5 provides examples of B2B primers useful in the methods of the present disclosure.
[0225]
Table 4-1
[0226] [Table 4-2]
[0227] [Table 5-1]
[0228] [Table 5-2]
[0229] [Table 5-3]
[0230] [Table 6]
[0231] Example 3: Fragmentation of genomic DNA for sequencing library construction 1 μL of genomic DNA was processed using the NEBNext dsDNA Fragmentase kit (New England Biolabs) according to the manufacturer's protocol. The incubation time was extended to 45 minutes at 37°C. The fragmentation reaction was stopped by adding 5 μL of 0.5 M EDTA pH 8.0, and the DNA was purified by adding twice the volume of Ampure XP beads (Beckman Coulter, A63881) according to the manufacturer's protocol. The fragmented DNA was analyzed on a Bioanalyzer with a High Sensitivity DNA kit (Agilent). The size range of the fragmented DNA was typically approximately 100 bp to 200 bp, with a peak of approximately 150 bp.
[0232] Example 4: Library Preparation Procedure In this embodiment, the KAPA Library Preparation Kit (KK8230) was used for illustrative purposes.
[0233] For the bead purification process, AMPure XP beads (cat#A63881) were equilibrated to room temperature and thoroughly re-mixed before mixing with the sample. After thoroughly mixing with the sample on a vortex mixer, this was incubated at room temperature for 15 minutes to allow DNA to bind to the beads. The beads were then placed on a magnetic stand until the liquid became clear. The beads were then washed twice with 200 ul of 80% ethanol and dried at room temperature for 15 minutes.
[0234] To perform the end-repair reaction, up to 50 μL (2-10 ng) of cell-free DNA was mixed with 20 μL of end-repair master mix (8 μL of water, 7 μL of 10X KAPA end-repair buffer, and 5 μL of KAPA end-repair enzyme mixture) and incubated at 20°C for 30 minutes. Subsequently, 120 μL of AMPure XP beads were added to 70 μL of the end-repair reaction mixture. The sample was then purified as described above.
[0235] To perform the A-tailing reaction, dried beads containing end-repaired DNA fragments were mixed with the A-tailing master mix (42 μL of water, 5 μL of 10X KAPA A tailing buffer, and KAPA A tailing enzyme). The reaction mixture was incubated at 30°C for 30 minutes. After adding 90 μL of PEG solution (20% PEG 8000, 2.5 M NaCl), the mixture was washed according to the bead purification protocol described above. This A-tailing step was skipped in the blunt-end ligation reaction.
[0236] For linker ligation, two oligonucleotides (5' to 3') with the following sequence were used to form an adapter polynucleotide replica: / 5Phos / CCATTTCATTACCTCTTTCTCCGCACCCGACATAGAT * T and / 5Phos / ATCTATGTCGGGTGCGGAGAAAGAGGTAATGAAATGG *Dried beads containing T. end repair (for blunt ligation) or A tail addition (for linker-based ligation) were mixed with 45 μL of ligation master mix (30 μL of water, 10 μL of 5x KAPA ligation buffer, and 5 μL of KAPA T4 DNA ligase), and 5 μL of water (for blunt-end ligation), or 5 μL of a molar equivalent mixture of linker oligonucleotides (for linker-based ligation). The beads were resuspended thoroughly and incubated at 20 °C for 15 minutes. After adding 50 μL of PEG solution (see above), the mixture was washed according to the bead purification protocol described above.
[0237] Multiple displacement amplification (MDA) was performed using the Illustra Genomiphi V2 DNA amplification kit. Dried beads containing the ligated single-stranded fragments were resuspended in 9 μL of buffer containing random hexamers, heated to 95 °C over 3 minutes, and then rapidly cooled on ice. After adding 1 μL of the enzyme mixture, the cooled sample was incubated at 30 °C for 90 minutes. The reaction was then stopped by heating at 65 °C for 10 minutes. After adding 30 μL of PEG solution (see above), the mixture was washed according to the purification protocol described above and resuspended in 200 μL of TE (by incubation at 65 °C for 5 minutes). If desired, the purified product can be quantified by quantitative PCR, digital droplet PCR (ddPCR), or proposed for next-generation sequencing (NGS).
[0238] After MDA, long ligated fragment strands (e.g., >2kb) were sonicated to ~300bp using a Covaris S220 in a total volume of 130 μL. The manufacturer's protocol indicated a peak power of 140 W, a duty factor of 10%, 200 cycles per burst, and a treatment time of 80 seconds. A fragment length of ~300 bp was selected to increase the likelihood of maintaining intact initial cell-free DNA fragments. Using the standard library preparation protocol, adapters could be placed on the sonicated DNA fragments for sequencing if desired. Various read compositions were returned from paired-end sequencing runs on an Illumina sequencer (either HiSeq or MiSeq). Barcoding of the desired sequences was performed using reads where the junction (self-junction, or adapter junction if the adapter was included in the ligation process) was internal to the read (adjacent to the non-adapter sequence at 5' and 3').
[0239] Example 5: Circularization and amplification This provides an example of a cyclization and amplification procedure according to the method described herein. This procedure used the following supplies: PCR Machine (e.g., MJ research PTC-200 Peltier thermal cycler); Circligase II, ssDNA ligase Epicentre cat# CL9025K; Exonuclease (e.g., Exol, NEB Biolabs cat # M0293S; Exol II, NEB Biolabs cat# M0206S); T4 polynucleotide kinase (NEB Biolab cat # M0201S); Whole genome amplification kit (e.g., GE Healthcare, Illustra, Ready-To-Go, Genomiphi, V3 DNA amplification kit); GlycoBlue (e.g., Ambion cat# AM9515); Microcentrifuge (e.g., Eppendrof 5415D); DNA purification beads (e.g., Agencourt, AMpure XP, Beckman Coulter cat# A63881); Magnetic stand (e.g., The MagnaRack® Invitrogen cat# CS15000); QubitR 2.0 Fluorometer (Invitrogen, cat#Q32866); molecular probe ds DNA HS assay kit (Life Technology cat #032854); and bioanalyzer (Agilent 2100), as well as high-sensitivity DNA reagent (cat#5067-4626).
[0240] For amplification of DNA fragments lacking 5'-terminus phosphate (e.g., cell-free DNA), the first step was single-strand end repair and formation. DNA was denatured at 96°C for 30 seconds (e.g., on a PCR machine). A polynucleotide kinase (PNK) reaction product was prepared by combining 40 μL of DNA with 5 μL of 10X PNK reaction buffer, followed by incubation at 37°C for 30 minutes. 1 mM ATP and PNK enzyme were added to the reaction product and incubated at 37°C for 45 minutes. Buffer exchange was performed by precipitation and resuspension of DNA. 50 μL of DNA from the PNK reaction product was combined with 5 μL of 0.5 M sodium acetate pH 5.2, 1 μL of GlycoBlue, 1 μL of oligonucleotides (100 ng / μL), and 150 μL of 100% ethanol. The mixture was incubated at -80°C for 30 minutes and centrifuged at 16 k rpm for 5 minutes to pelletize the DNA. The DNA pellet was washed with 500 μL of 70% ethanol, air-dried at room temperature for 5 minutes, and then suspended in 12 μL of 10 mM Tris=Cl pH 8.0.
[0241] Subsequently, the resuspended DNA was circularized by ligation. The DNA was denatured at 96°C for 30 seconds, the samples were cooled on ice for 2 minutes, and a ligase mixture (2 μL of 10X CircLigase buffer, 4 μL of 5M betaine, 1 μL of 50 mM MnCl2, and 1 μL of CircLigase II) was added. The ligation reaction was incubated on a PCR machine at 60°C for 16 hours. Unligated polynucleotides were digested by exonuclease digestion. For this purpose, the DNA was denatured at 80°C for 45 seconds, and 1 μL of the exonuclease mixture (ExoI 20U / μL:ExoIII 100U / μL = 1:2) was added to each tube. This was mixed by five pipette aspirates and spits, and then briefly rotated. The digested mixture was incubated at 37°C for 45 minutes. The volume was increased to 50 μL by adding 30 μL of water, and further buffer exchange was performed by precipitation and resuspension as described above.
[0242] To perform whole-genome amplification (WGA), purified DNA was first denatured at 65°C for 5 minutes. 10 μL of denaturation buffer from the GE WGA kit was added to 10 μL of purified DNA. The DNA was cooled on a cold block or ice for 2 minutes. 20 μL of DNA was added to the Ready-To-Go GenomiPhi V3 cake (WGA). The WGA reaction mixture was incubated at 30°C for 1.5 hours, followed by thermal inactivation at 65°C for 10 minutes.
[0243] The sample was purified using AmpureXP magnetic beads (1.6X). The beads were agitated, and 80 μL was divided equally into 1.5 mL tubes. Then, 30 μL of water, 20 μL of amplified DNA, and 80 μL of beads were combined and incubated at room temperature for 3 minutes. The tubes were placed on a magnetic stand for 2 minutes, and a clear solution was pipetted out. The beads were washed twice with 80% ethanol. The DNA was eluted by adding 200 μL of 10 mM Tris-Cl pH 8.0. The DNA-bead mixture was incubated at 65°C for 5 minutes. The tubes were returned to the magnetic stand for 2 minutes. 195 μL of DNA was transferred to a new tube. 1 μLw was used for quantification using Qbit. Finally, 130 μL of the WGA product was sonicated using a Covaris S220 to a size of approximately 400 bp.
[0244] Example 6: Analysis of ligation efficiency and on-target rate cfDNA that had been circularized and subjected to whole-genome amplification as in the above examples was analyzed by quantitative PCR (qPCR). The results of the qPCR amplification curves for the sample target (using KRAS primers) are shown in Figures 18A and 18B. As shown in Figure 18A, qPCR amplification of 1 / 10 of the input cfDNA yielded an average Ct (cycle threshold) of 31.75, and 1 / 10 of the sample ligation product yielded an average Ct of 31.927, indicating a high ligation efficiency of approximately 88%. Ligation efficiencies can also reach approximately 70%, 80%, 90%, 95%, or even 100%. Linear DNA that was not circularized was removed in some examples so that almost all of the DNA could be amplified from the circular DNA. Each sample was replicated and run three times. As shown in Figure 18B, the amplification curves for 10 ng of WGA product and reference genomic DNA (gDNA) (12878, 10 ng) substantially overlap each other. The mean Ct of the WGA samples was 26.655, while the mean of the gDNA samples was 26.605, indicating a high on-target rate of over 96%. The number of KRAS in a given amount of amplified DNA is comparable to that of unamplified gDNA, indicating an unbiased amplification process. Each sample was replicated and tested three times. For comparison, the circularization protocol provided by Lou et al. (PNAS, 2013, 110 (49)) was also tested. Using Lou's method, which lacked the precipitation and purification steps of the described examples, only 10-30% of linear input DNA was converted to circular DNA. Such low recovery presents a hindrance to downstream sequencing and variant detection.
[0245] Example 7: Analysis of circularized DNA amplified by ddPCR We evaluated the conservation and bias of allele frequencies in whole-genome amplification products generated from cyclic polynucleotides using droplet digital PCR (ddPCR). Generally, ddPCR refers to a digital PCR assay that measures absolute quantities by counting nucleic acid molecules encapsulated in separate, volumetrically defined water-in-oil droplet compartments that support PCR amplification (Hinson et al, 2011, Anal. Chem. 83:8604-8610; Pinheiro et al, 2012, Anal. Chem. 84: 1003-1011). A single ddPCR reaction may consist of at least 20,000 compartmentalized droplets per well. Droplet digital PCR may be performed using a platform that performs a digital PCR assay that measures absolute quantities by counting nucleic acid molecules encapsulated in separate, volumetrically defined water-in-oil droplet compartments that support PCR amplification. A typical strategy for droplet digital PCR may be summarized as follows: the sample is diluted and divided into thousands to millions of separate reaction vessels (water-in-oil droplets), so that each does not contain one or more copies of the ...
Claims
1. A method for performing rolling circle amplification, The method is, (a) A step of cyclizing individual polynucleotides in a plurality of polynucleotides in order to form a plurality of cyclic polynucleotides using a ligase enzyme, wherein each polynucleotide of the plurality of polynucleotides has a 5' end and a 3' end before ligation, (b) A step of breaking down the ligase enzyme, (c) The step of amplifying a cyclic polynucleotide after degrading a ligase enzyme in order to produce amplified polynucleotides, A method in which polynucleotides are not purified or isolated between steps (a) and (c).
2. The method according to claim 1, further comprising the step of degrading a linear polynucleotide between steps (a) and (c).
3. The method according to claim 1, wherein multiple polynucleotides are in a single strand.
4. The method according to claim 1, wherein each cyclic polynucleotide has a characteristic junction within the cyclic polynucleotide.
5. The method according to claim 1, wherein the cyclicization step includes a step of attaching an adapter polynucleotide to the 5' end, 3' end, or both the 5' end and 3' end of a polynucleotide in a plurality of polynucleotides.
6. The method according to claim 1, wherein the amplification step includes exposing a cyclic polynucleotide to an amplification reaction mixture containing random primers.
7. The method according to claim 1, wherein the amplification step comprises exposing a cyclic polynucleotide to an amplification reaction mixture containing one or more primers, each of which specifically hybridizes to a different target sequence by sequence complementarity.
8. The method according to claim 1, wherein the sample is a sample from a subject.
9. The method according to claim 8, wherein the sample is urine, feces, blood, saliva, tissue, or bodily fluid.
10. The method according to claim 8, wherein the sample comprises tumor cells.
11. The method according to claim 8, wherein the sample is a formalin-fixed paraffin-embedded (FFPE) sample.
12. The method according to claim 8, further comprising the steps of diagnosing the subject based on the call process and performing a treatment as appropriate.
13. The method according to claim 1, wherein the sequence variant is a causative gene variant.
14. The method according to claim 1, wherein the sequence variant is associated with the type or stage of cancer.
15. The method according to claim 1, wherein the plurality of polynucleotides include cell-free polynucleotides.
16. The method according to claim 15, wherein the cell-free polynucleotides contain circulating tumor DNA.
17. The method according to claim 1, further comprising the step of sequencing amplified polynucleotides to generate multiple sequencing reads.
18. The method according to claim 17, further comprising the step of identifying sequence differences between a sequencing read and a reference sequence.
19. The method according to claim 18, further comprising the steps of (i) identifying sequence differences in both strands of a double-stranded input molecule, (ii) calling the sequence differences as sequence variants in a plurality of polynucleotides only when the sequence differences occur in two different molecules, the sequence differences are identified in both strands of a double-stranded input molecule, (ii) resulting in a consensus sequence of a concatemer formed by rolling circle amplification, and / or (iii) resulting in a sequence difference of a plurality of polynucleotides.
20. The method according to claim 19, wherein when the sequence difference occurs in at least two cyclic polynucleotides having different junctions formed between their 5' and 3' ends, the sequence difference is identified as occurring in two different molecules.
21. The method according to claim 19 or 20, wherein when reads corresponding to two different molecules have different 5' ends and different 3' ends, the sequence difference is identified as occurring in the two different molecules.
22. The method according to claim 1, further comprising the step of cleaving the amplified nucleotides.
23. A method for identifying sequence variants in a nucleic acid sample containing multiple polynucleotides, each having a 5' end and a 3' end, The method is: (a) A step of cyclicizing individual polynucleotides of a plurality of polynucleotides in order to form a plurality of cyclic polynucleotides, wherein each polynucleotide has a junction between its 5' end and its 3' end, (b) A step of amplifying the cyclic polynucleotide of (a) to produce amplified polynucleotides, (c) A step of cleaving an amplified polynucleotide to produce cleaved polynucleotides, wherein each cleaved polynucleotide has one or more cleavage points at its 5' end and / or 3' end, (d) A step of sequencing the cleaved polynucleotides in order to generate multiple sequencing reads, (e) A step of identifying the sequence differences between the sequencing read and the reference sequence, (f) A method comprising the step of calling a sequence difference as a sequence variant when the sequence difference occurs in at least two different cleaved polynucleotides.
24. The method according to claim 23, wherein the sequence difference is called a sequence variant when (i) the sequence difference occurs in at least two cyclic polynucleotides having different junctions, (ii) the sequence difference is identified in both strands of a double-stranded input molecule, and / or (iii) the sequence difference occurs in the consensus sequence of a concatemer formed by amplification including rolling circle amplification.
25. The method according to claim 23, wherein multiple polynucleotides are in a single strand.
26. The method according to claim 23, wherein the cyclization step is achieved by subjecting a plurality of polynucleotides to a ligation reaction.
27. The method according to claim 23, wherein the sequence variant is a single nucleotide polymorphism.
28. The method according to claim 23, wherein the reference sequence is a consensus sequence formed by aligning sequencing reads with each other.
29. The method according to claim 23, wherein the reference sequence is a sequencing read.
30. The method according to claim 23, wherein the cyclicization step includes a step of attaching an adapter polynucleotide to the 5' end, the 3' end, or both the 5' end and the 3' end of a polynucleotide in a plurality of polynucleotides.
31. The method according to claim 23, wherein the amplification step is achieved by using a polymerase having chain substitution activity.
32. The method according to claim 23, wherein the amplification step includes exposing a cyclic polynucleotide to an amplification reaction mixture containing random primers.
33. The method according to claim 23, wherein the amplification step comprises exposing a cyclic polynucleotide to an amplification reaction mixture comprising one or more primers, each of which specifically hybridizes to a different target sequence by sequence complementarity.
34. The method according to claim 23, wherein the amplified polynucleotides are subjected to a sequencing process that does not involve enrichment.
35. The method according to claim 23, further comprising the step of enriching one or more target polynucleotides in the amplified polynucleotide by performing an enrichment step before sequencing.
36. The method according to claim 23, wherein microbial contaminants are identified based on a coal process.
37. The method according to claim 23, wherein the sample is a sample from a subject.
38. The method according to claim 37, wherein the sample is urine, feces, blood, saliva, tissue, or bodily fluid.
39. The method according to claim 37, wherein the sample includes tumor cells.
40. The method according to claim 37, wherein the sample is a formalin-fixed paraffin-embedded (FFPE) sample.
41. The method according to claim 37, further comprising the steps of diagnosing the subject based on the call process and performing treatment as appropriate.
42. The method according to claim 23, wherein the sequence variant is a causative gene variant.
43. The method according to claim 23, wherein the sequence variant is associated with the type or stage of cancer.
44. The method according to claim 23, wherein the multiple polynucleotides include cell-free polynucleotides.
45. The method according to claim 44, wherein the cell-free polynucleotides contain circulating tumor DNA.
46. A reaction mixture for carrying out any one of the methods of claims 1 to 43, The reacting mixture is (a) Multiple concatemers, wherein each concatemer of the multiple concatemers includes a different junction formed by cyclizing individual polynucleotides having a 5' end and a 3' end, (b) A first primer comprising sequence A', wherein the first primer specifically hybridizes to sequence A of the target sequence by sequence complementarity between sequence A and sequence A', (c) A second primer comprising sequence B, wherein the second primer specifically hybridizes to sequence B' present in a complementary polynucleotide containing the complement of the target sequence due to sequence complementation between sequence B and sequence B', (d) comprising a polymerase that extends the first and second primers to produce amplified polynucleotides, A reaction mixture in which the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nt or less.
47. The reaction mixture according to claim 46, wherein the first primer comprises sequence C5' for sequence A', the second primer comprises sequence D5' for sequence B, and neither sequence C nor sequence D hybridizes to two or more concatemers during the first amplification step of the amplification reaction.
48. A system for detecting sequence variants, wherein the system is (a) A computer configured to receive user requests to perform a detection reaction on a sample, (b) An amplification system that performs a nucleic acid amplification reaction on a sample or a portion thereof in response to a user's request, Amplification reaction (i) A step of cyclizing individual polynucleotides in a plurality of polynucleotides in order to form a plurality of cyclic polynucleotides using a ligase enzyme, wherein each polynucleotide of the plurality of polynucleotides has a 5' end and a 3' end before ligation, (ii) The process of breaking down the ligase enzyme, (iii) an amplification system comprising the step of amplifying a cyclic polynucleotide after degrading a ligase enzyme to produce an amplified polynucleotide, wherein the polynucleotide is not purified or isolated between steps (i) and (iii), (c) A sequencing system that generates sequencing reads of polynucleotides amplified by an amplification system, identifies sequence differences between the sequencing reads and a reference sequence, and calls sequence differences arising in at least two cyclic polynucleotides having various junctions as sequence variants. (d) A system comprising a report generation device for sending a report to a recipient, wherein the report includes the results of sequence variant detection.
49. The system according to claim 48, wherein the recipient is the user.
50. A computer-readable medium containing code that performs a method for detecting sequence variants when executed by one or more processors, The method to be implemented is: (a) A step of receiving a customer request to perform a detection reaction on a sample, (b) A step of performing a nucleic acid amplification reaction on a sample or a portion thereof in response to a customer request, Amplification reaction (i) A step of cyclizing individual polynucleotides in a plurality of polynucleotides in order to form a plurality of cyclic polynucleotides using a ligase enzyme, wherein each polynucleotide of the plurality of polynucleotides has a 5' end and a 3' end before ligation, (ii) The process of breaking down the ligase enzyme, (iii) A step of performing a nucleic acid amplification reaction, which includes a step of amplifying a cyclic polynucleotide after degrading a ligase enzyme in order to produce amplified polynucleotides, wherein the polynucleotides are not purified or isolated between step (i) and (iii), (c) A step of performing sequencing analysis, which includes (i) generating sequencing reads of polynucleotides amplified by an amplification reaction, (ii) identifying sequence differences between the sequencing reads and a reference sequence, and (iii) calling sequence differences arising from at least two cyclic polynucleotides having different junctions as sequence variants. (d) A computer-readable medium comprising the step of preparing a report including the results of the detection of sequence variants.
51. A method for identifying sequence variants in a nucleic acid sample containing multiple polynucleotides, each having a 5' end and a 3' end, The method is: (a) A step of cyclicizing individual polynucleotides of a plurality of polynucleotides in order to form a plurality of cyclic polynucleotides, wherein a predetermined cyclic polynucleotide of the plurality of polynucleotides has a conjugate sequence resulting from the cyclicization, (b) A step of amplifying the cyclic polynucleotide of (a) to produce a plurality of amplified polynucleotides, wherein the first amplified polynucleotide of the plurality and the second amplified polynucleotide of the plurality include a conjugate sequence, but each includes different sequences at its 5' and / or 3' ends. (c) A step of sequencing a plurality of amplified polynucleotides or their amplification products in order to generate a plurality of sequencing reads corresponding to a first amplified polynucleotide and a second amplified polynucleotide, (d) A method comprising calling a sequence difference detected in a sequencing read as a sequence variant when the sequence difference detected in the sequencing read occurs in the sequencing read corresponding to both the first amplified polynucleotide and the second amplified polynucleotide.
52. The method according to claim 51, wherein the step of cyclizing the individual polynucleotides in (a) is achieved by a ligase enzyme.
53. The method according to claim 52, wherein the ligase enzyme is broken down before (b).
54. The method according to claim 51, wherein the acyclic polynucleotide is degraded before (b).
55. The method according to any one of claims 51-54, wherein the multiple cyclic polynucleotides are not purified or isolated prior to (b).
56. The method according to claim 51, wherein the cyclic formation step in (a) includes the step of attaching an adapter polynucleotide to the 5' end, the 3' end, or both the 5' end and the 3' end of a polynucleotide in a plurality of polynucleotides.
57. The method according to claim 51, wherein the step of amplifying the cyclic polynucleotide in (b) is achieved by a polymerase having chain substitution activity.
58. The method according to claim 51, wherein the step of amplifying the cyclic polynucleotide in (b) includes rolling circle amplification (RCA).
59. The method according to claim 51, wherein the amplification step in (b) includes exposing a cyclic polynucleotide to an amplification reaction mixture containing random primers.
60. The method according to claim 59, wherein each random primer has a sequence at each of its different 5' and / or 3' ends.
61. The method according to claim 51, wherein the amplification step in (b) includes exposing a cyclic polynucleotide to an amplification reaction mixture containing a target-specific primer.
62. The method according to claim 51, wherein the amplification step comprises multiple cycles of denaturation, primer binding, and primer extension.
63. The method according to claim 51, wherein the amplified polynucleotides are subjected to sequencing of (c) without enrichment.
64. The method according to claim 51, further comprising the step of enriching one or more target polynucleotides in the amplified polynucleotide or its amplification product by carrying out an enrichment step prior to sequencing in (c).
65. The method according to claim 51, wherein the plurality of polynucleotides include single-stranded polynucleotides.
66. The method according to claim 51, wherein the sequence variant is a single nucleotide polymorphism.
67. The method according to claim 51, wherein the sample is a sample from a subject.
68. The method according to claim 67, wherein the sample includes urine, feces, blood, saliva, tissue, or bodily fluids.
69. The method according to claim 67, wherein the sample comprises tumor cells.
70. The method according to claim 67, wherein the sample includes a formalin-fixed paraffin-embedded (FFPE) sample.
71. The method according to claim 51, wherein the multiple polynucleotides include cell-free polynucleotides.
72. The method according to claim 71, wherein the cell-free polynucleotide includes cell-free DNA.
73. The method according to claim 71, wherein the cell-free polynucleotide includes cell-free RNA.
74. The method according to claim 71, wherein the cell-free polynucleotides contain circulating tumor DNA.
75. The method according to claim 71, wherein the cell-free polynucleotides contain circulating tumor RNA.
76. The method according to claim 51, wherein, in (d), when the sequence differences detected in the sequencing reads occur in at least 50% of the sequencing reads from the first amplified polynucleotide and at least 50% of the sequencing reads from the second amplified polynucleotide, the sequence differences detected in the sequencing reads are called sequence variants.
77. A method for identifying sequence variants in a nucleic acid sample containing multiple polynucleotides, Each of the multiple polynucleotides has a 5' end and a 3' end. The aforementioned method: (a) A step of cyclizing individual polynucleotides of a plurality of polynucleotides in order to form a plurality of cyclic polynucleotides, wherein a predetermined cyclic polynucleotide of the plurality of polynucleotides has a conjugate sequence resulting from the cyclization, (b) A step of amplifying the cyclic polynucleotide of (a) to produce multiple amplified polynucleotides, (c) A step of cleaving an amplified polynucleotide to produce cleaved polynucleotides, wherein each cleaved polynucleotide has one or more cleavage points at its 5' end and / or 3' end, (d) A step of sequencing the amplification products of polynucleotides that have been cleaved to generate multiple sequencing reads, (e) A method comprising calling a sequence difference detected in a sequencing read as a sequence variant when the sequence difference detected in the sequencing read occurs in the sequencing read corresponding to a first cleaved polynucleotide and a second cleaved polynucleotide.
78. The method of claim 77, wherein the step of cyclizing the individual polynucleotides in (a) is achieved by a ligase enzyme.
79. The method according to claim 78, wherein the ligase enzyme is broken down before (b).
80. The method according to claim 77, wherein the acyclic polynucleotide is degraded before (b).
81. The method according to any one of claims 77-80, wherein the multiple cyclic polynucleotides are not purified or isolated prior to (b).
82. The method according to claim 77, wherein the cyclic formation step in (a) includes the step of attaching an adapter polynucleotide to the 5' end, the 3' end, or both the 5' end and the 3' end of a polynucleotide in a plurality of polynucleotides.
83. The method according to claim 77, wherein the step of amplifying the cyclic polynucleotide in (b) is achieved by a polymerase having chain substitution activity.
84. The method according to claim 77, wherein the step of amplifying the cyclic polynucleotide in (b) includes rolling circle amplification (RCA).
85. The method according to claim 77, wherein the amplification step in (b) includes exposing a cyclic polynucleotide to an amplification reaction mixture containing random primers.
86. The method according to claim 85, wherein each random primer has a sequence at each of its different 5' and / or 3' ends.
87. The method according to claim 77, wherein the amplification step in (b) includes exposing a cyclic polynucleotide to an amplification reaction mixture containing a target-specific primer.
88. The method according to claim 77, wherein the amplified product of the cleaved polynucleotide is subjected to unenriched sequencing.
89. The method according to claim 77, further comprising the step of enriching one or more target polynucleotides in the amplification product of the cleaved polynucleotides by performing an enrichment step prior to sequencing of (d).
90. The method according to claim 77, wherein the plurality of polynucleotides include single-stranded polynucleotides.
91. The method according to claim 77, wherein the sequence variant is a single nucleotide polymorphism.
92. The method according to claim 77, wherein the sample is a sample from a subject.
93. The method according to claim 92, wherein the sample includes urine, feces, blood, saliva, tissue, or bodily fluids.
94. The method according to claim 92, wherein the sample includes tumor cells.
95. The method according to claim 92, wherein the sample includes a formalin-fixed paraffin-embedded (FFPE) sample.
96. The method according to claim 77, wherein the multiple polynucleotides include cell-free polynucleotides.
97. The method according to claim 96, wherein the cell-free polynucleotide includes cell-free DNA.
98. The method according to claim 96, wherein the cell-free polynucleotide includes cell-free RNA.
99. The method according to claim 96, wherein the cell-free polynucleotides contain circulating tumor DNA.
100. The method according to claim 96, wherein the cell-free polynucleotides contain circulating tumor RNA.
101. The method according to claim 77, wherein, in (e), when the sequence differences detected in the sequencing reads occur in at least 50% of the sequencing reads from the first cleaved polynucleotide and at least 50% of the sequencing reads from the second cleaved polynucleotide, the sequence differences detected in the sequencing reads are called sequence variants.
102. The method according to any one of claims 77-101, wherein the cleavage step in (c) includes a step of exposing the amplified polynucleotide to sonication.
103. The method according to any one of claims 77-101, wherein the cleavage step in (c) includes a step of exposing the amplified polynucleotide to enzymatic cleavage.