Methods and reagents for nucleic acid sequencing and related uses
The use of adapter molecules with hairpin shapes in duplex sequencing improves conversion efficiency and reduces costs, enabling high-precision sequencing for clinical and diagnostic applications by amplifying and error-correcting both strands of nucleic acid molecules.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-04
AI Technical Summary
Existing duplex sequencing methods face limitations in conversion efficiency, leading to insufficient sequence information production and increased costs due to the need for extensive reagents and steps, which hampers their application in preclinical and clinical testing.
A method involving the use of adapter molecules with hairpin shapes to enhance duplex sequencing by amplifying and sequencing both strands of double-stranded nucleic acid molecules, followed by error correction through strand comparison, thereby increasing conversion efficiency and reducing reagent and step requirements.
This approach provides high-precision sequencing at a lower cost and faster rate, enhancing the accuracy and efficiency of duplex sequencing for applications in diagnostics and clinical testing.
Smart Images

Figure 2026035740000001 
Figure 2026035740000002 
Figure 2026035740000003
Abstract
Description
[Technical Field]
[0001] The present technology generally relates to methods and related reagents for providing highly accurate (e.g., error-corrected) nucleic acid sequences. In particular, some embodiments are directed to adapter molecules comprising hairpin shapes and methods of using such adapters in duplex sequencing and other sequencing applications.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 881,936, filed August 1, 2019, the disclosure of which is incorporated herein by reference in its entirety. [Background technology]
[0003] Duplex sequencing is an error-correction method that achieves exceptional sequence accuracy by comparing the sequence information from both strands of individual double-stranded nucleic acid molecules. Regarding the efficiency of the duplex sequencing process or other high-precision sequencing methods, conversion efficiency can be defined as the fraction of unique nucleic acid molecules input into a sequencing library preparation reaction that produces at least one double-stranded consensus sequence read (or other high-precision sequence read). In some cases, shortcomings in conversion efficiency can limit the usefulness of high-precision sequencing for some applications that would otherwise be very suitable. For example, low conversion efficiency can lead to a situation where the copy number of the target double-stranded nucleic acid is limited, which can result in less than desired amount of sequence information being produced. There is a need for a cost-effective manufacturing method for synthesizing raw sequence reads of nucleic acid molecules for use in various applications, including duplex sequencing applications. Summary of the Invention
[0004] The technology of the present invention generally relates to methods and related reagents for nucleic acid sequencing. In particular, some aspects of the technology are directed to methods for achieving high-precision sequencing reads, which are provided at a faster rate (e.g., with fewer steps) and / or at a lower cost (e.g., using fewer reagents), and provide more desired data. Other aspects of the technology are directed to methods and reagents for increasing the conversion efficiency for duplex sequencing. Various aspects of the technology have many applications in preclinical and clinical testing and diagnostics, as well as other applications.
[0005] In some aspects, the disclosure provides a method for sequencing a double-stranded target nucleic acid molecule, comprising: (a) amplifying physically linked nucleic acid complexes on a surface to generate physically linked nucleic acid complex amplicons bound to the surface in both forward and reverse orientations, the physically linked nucleic acid complexes comprising: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain on a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion on a second end of the double-stranded target nucleic acid molecule; (b) removing either (i) the physically linked nucleic acid complex amplicons bound to the surface in the reverse orientation or (ii) the physically linked nucleic acid complex amplicons bound to the surface in the forward orientation; and (c) removing any remaining bound physically linked nucleic acid complex amplicons. (d) sequencing the subset of single-stranded amplicons to provide sequencing reads derived from the original strand of the double-stranded target nucleic acid molecule; (e) amplifying the subset of physically linked nucleic acid complex amplicons on a surface; (f) removing the physically linked nucleic acid complex amplicons in the other orientation; (g) cleaving the remaining bound physically linked nucleic acid complex amplicons to provide single-stranded amplicons containing information from the other strand; and (h) sequencing the single-stranded amplicons to provide sequencing reads derived from the other original strand of the double-stranded target nucleic acid molecule.
[0006] In some aspects, the disclosure provides a method for sequencing a double-stranded target nucleic acid molecule, the method comprising: (a) amplifying physically linked nucleic acid complexes on a surface to generate clusters of surface-bound physically linked nucleic acid complex amplicons, the physically linked nucleic acid complexes comprising: (i) a double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain on one end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end of the double-stranded target nucleic acid molecule; and (b) amplifying the physically linked nucleic acid complexes on a surface to generate clusters of surface-bound physically linked nucleic acid complex amplicons, the physically linked nucleic acid complexes comprising: (i) a double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain on one end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end of the double-stranded target nucleic acid molecule. (i) removing physically linked nucleic acid complex amplicons bound to a surface at either (i) the 5' end of the bound nucleic acid complex amplicon, or (ii) the 3' end of the bound nucleic acid complex amplicon; (c) cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons at the cleavage site to provide single-stranded amplicons containing sequence information derived from one original strand of the double-stranded target nucleic acid molecule; and (d) sequencing the single-stranded amplicons to provide sequencing reads derived from one original strand of the double-stranded target nucleic acid molecule. In some aspects, the method further comprises cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons and comprises preserving at least one surface-bound physically linked nucleic acid complex amplicon. In some aspects, the method further comprises: (e) amplifying at least one physically linked nucleic acid complex amplicon on the surface and rearranging the cluster of surface-bound physically linked nucleic acid complex amplicons; (f) removing physically linked nucleic acid complex amplicons in the other orientation not removed in (b); (g) cleaving the remaining bound physically linked nucleic acid complex amplicons to provide single-stranded amplicons containing information derived from the other original strand of the double-stranded target nucleic acid molecule; and (h) sequencing the single-stranded amplicons to provide sequencing reads derived from the other original strand of the double-stranded target nucleic acid molecule.
[0007] In some aspects, the method further comprises comparing sequence reads from one original strand with sequence reads from the other original strand to generate a consensus sequence of the double-stranded target nucleic acid molecule. In some aspects, the method further comprises identifying sequence variations in the sequence reads from one original strand and the sequence reads from the other original strand, wherein the sequence variations from one original strand and the other original strand are consistent sequence variations, or excluding or ignoring sequence variations that occur in one original strand but not in the other original strand. In some aspects, the method further comprises comparing the sequence reads from one original strand with the sequence reads from the other original strand, identifying nucleotide positions that do not match between the sequence reads from one original strand and the sequence reads from the other original strand, and generating an error-corrected sequence of the double-stranded target nucleic acid molecule by ignoring, excluding, or correcting the nucleotide positions identified as not matching.
[0008] In some aspects, the disclosure provides a method for sequencing a collection of double-stranded target nucleic acid molecules, each of which comprises a first strand and a second strand, comprising the steps of: (a) amplifying a plurality of physically linked nucleic acid complexes on a surface to generate a plurality of clonal clusters, each of which comprises a plurality of physically linked nucleic acid complex amplicons, each of which comprises a first strand amplicon and a second strand amplicon, each physically linked nucleic acid complex comprising: (i) a double-stranded target nucleic acid molecule from the collection; (ii) a first adaptor comprising a linker domain attached to a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion attached to a second end of the double-stranded target nucleic acid molecule; (b) removing any physically linked nucleic acid complex amplicons from each surface-bound clonal cluster in either (i) the reverse direction or (ii) the forward direction; (c) cleaving a portion of the remaining surface-bound physically linked nucleic acid complex amplicons remaining after (b), thereby physically separating first-strand amplicons and second-strand amplicons; (d) removing any unbound physically separated first- or second-strand amplicons; and (e) sequencing the remaining surface-bound physically separated first- or second-strand amplicons to generate first-strand or second-strand nucleic acid sequence reads for each clonal cluster on the surface. In some embodiments, cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons comprises preserving at least one physically linked nucleic acid complex amplicon in at least some of the surface-bound clonal clusters.In some embodiments, the method further comprises: (f) amplifying at least one physically linked nucleic acid complex amplicon on the surface in at least some of the clonal clusters to rearrange clonal clusters of surface-bound physically linked nucleic acid complex amplicons; (g) removing physically linked nucleic acid complex amplicons in the other direction from step (b); (h) removing unbound physically separated first or second strand amplicons; (i) cleaving the remaining bound physically linked nucleic acid complex amplicons remaining after (h), thereby physically separating the first strand amplicons and second strand amplicons; and (j) sequencing the remaining surface-bound physically separated first or second strand amplicons to generate first strand or second strand nucleic acid sequence reads for each clonal cluster on the surface.
[0009] In some aspects, the disclosure provides a method for sequencing a collection of double-stranded target nucleic acid molecules, each of which comprises a first strand and a second strand, comprising: (a) amplifying a plurality of physically linked nucleic acid complexes bound on a surface to generate a plurality of clusters, each cluster comprising a plurality of physically linked nucleic acid complex amplicons representing an original double-stranded target nucleic acid molecule, each physically linked nucleic acid complex amplicon comprising a first strand amplicon and a second strand amplicon, each physically linked nucleic acid complex comprising a double-stranded target nucleic acid molecule from the collection connected to (i) a first adaptor comprising a linker domain at one end between the first and second strands, and (ii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end; and (b) amplifying the surface-bound physically linked nucleic acid complex amplicons to generate a plurality of clusters, each cluster comprising a plurality of physically linked nucleic acid complex amplicons representing an original double-stranded target nucleic acid molecule, each physically linked nucleic acid complex amplicon comprising a first strand amplicon and a second strand amplicon, each physically linked nucleic acid complex comprising a double-stranded target nucleic acid molecule from the collection connected to (i) a first adaptor comprising a linker domain at one end between the first and second strands, and (ii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end. (c) cleaving the surface-bound physically separated first-strand amplicons, thereby physically separating the first-strand amplicons and the second-strand amplicons; (c) removing unbound physically separated first-strand amplicons and / or unbound physically separated second-strand amplicons, wherein the remaining amplicons bound to the surface comprise (i) physically separated first-strand amplicons and (ii) physically separated second-strand amplicons; (d) sequencing the surface-bound physically separated first-strand amplicons to generate first-strand nucleic acid sequence reads for each cluster on the surface; and (e) sequencing the surface-bound physically separated second-strand amplicons to generate second-strand nucleic acid sequence reads for each cluster on the surface.
[0010] In some embodiments, for at least some of the clusters on the surface, the method further comprises comparing the first-strand nucleic acid sequence reads with the second-strand nucleic acid sequence reads to generate error-corrected sequence reads of the original double-stranded target nucleic acid molecules. In some embodiments, the method further comprises using a unique molecular identifier (UMI) to associate the first-strand nucleic acid sequence reads of the original double-stranded target nucleic acid molecules from the collection with the second-strand nucleic acid sequence reads of the same original double-stranded target nucleic acid molecules. In some embodiments, the UMI comprises a physical location on the surface. In other embodiments, the UMI comprises a tag sequence, a molecular-specific feature, a cluster location on the surface, or a combination thereof. In some embodiments, the molecular-specific feature comprises nucleic acid mapping information relative to a reference sequence, sequence information at or near the ends of the double-stranded target nucleic acid molecules, the length of the double-stranded target nucleic acid molecules, or a combination thereof.
[0011] In some embodiments, the method further comprises using a strand-defining element (SDE) to distinguish nucleic acid sequence reads of a first strand of an original double-stranded target nucleic acid molecule from nucleic acid sequence reads of a second strand from the same original double-stranded target nucleic acid molecule. In some embodiments, the SDE is an association of steps (e) and (j), or steps (d) and (e), with sequence read information. In some embodiments, the SDE comprises a portion of an adapter sequence.
[0012] In some embodiments, sequencing the physically separated first strand amplicons or second strand amplicons comprises sequencing by synthesis.
[0013] In some embodiments, the method further comprises preparing physically linked nucleic acid complexes by ligating a first adaptor and a second adaptor to each of a plurality of double-stranded target nucleic acid molecules in the collection, and presenting the physically linked nucleic acid complexes on a surface, the surface having a plurality of bound oligonucleotides at least partially complementary to the single-stranded portions of the second adaptor, such that the plurality of physically linked nucleic acid complexes are captured on the surface via hybridization to the plurality of bound oligonucleotides. In some embodiments, the method further comprises amplifying the physically linked nucleic acid complexes prior to the presenting step. In some embodiments, amplifying the physically linked nucleic acid complexes prior to the presenting step comprises PCR amplification or circle amplification. In other embodiments, the physically linked nucleic acid complexes are captured on the surface in both the forward and reverse directions.
[0014] In some aspects, the amplifying step comprises bridge amplification.
[0015] In some embodiments, the method further includes, for at least some of the double-stranded target nucleic acid molecules in the collection, (i) comparing sequence reads from the first strand with sequence reads from the second strand; (ii) identifying nucleotide positions that do not match between the sequence reads from the first strand and the sequence reads from the second strand; and (iii) generating error-corrected sequence reads for the double-stranded target nucleic acid molecules by disregarding, excluding, or correcting the identified nucleotide positions that do not match.
[0016] In some embodiments, the first adapter comprises a cleavable site or motif. In some embodiments, the first adapter and the second adapter each comprise a sequencing primer binding site and, optionally, a single molecule identifier (SMI) sequence. In some embodiments, the second adapter comprises a sequencing primer binding site, an amplification primer binding site, an index sequence, or any combination thereof. In some embodiments, the linker domain comprises a cleavage site. In some embodiments, the first adapter comprises a cleavable domain. In some embodiments, the first adapter comprises a hairpin loop structure comprising a self-complementary stem portion and a single-stranded nucleotide loop portion. In some embodiments, the single-stranded nucleotide loop portion comprises a cleavable domain. In some embodiments, the stem portion comprises a cleavable domain. In some embodiments, the cleavable domain comprises an enzyme recognition site. In some embodiments, the enzyme recognition site is an endonuclease recognition site. In some embodiments, the endonuclease is a restriction enzyme or a targeting endonuclease.
[0017] In some embodiments, the second adaptor is a "Y"-shaped adaptor. In some embodiments, one or both arms of the Y-shaped adaptor are capable of hybridizing to a surface-bound oligonucleotide.
[0018] In some embodiments, the single-stranded portion of the second adapter comprises a first arm having a first primer binding site and a second arm having a second primer binding site. In some embodiments, upon denaturation, the physically linked double-stranded nucleic acid complex comprises, 5' to 3' or 3' to 5', the first primer binding site, the first strand, the first adapter comprising a linker domain, the second strand, and the second primer binding site.
[0019] In some embodiments, the surface is a sequencing surface. In some embodiments, the surface is a flow cell. In other embodiments, the surface is the surface of a bead.
[0020] In some embodiments, the amplification is selected from the group consisting of PCR amplification, isothermal amplification, polony amplification, cluster amplification, and bridge amplification. In some embodiments, the amplification is bridge amplification on a surface.
[0021] In some embodiments, one or more of the plurality of first-strand amplicons and / or the plurality of second-strand amplicons are bound to the surface in a forward orientation, hi some embodiments, one or more of the plurality of first-strand amplicons and / or the plurality of second-strand amplicons are bound to the surface in a reverse orientation.
[0022] In some aspects, the method further comprises flowing a plurality of physically linked double-stranded nucleic acid complexes over a surface prior to amplification.
[0023] In some embodiments, the surface comprises a plurality of one or more bound oligonucleotides that are at least partially complementary to one or more regions of the second adaptor, hi some embodiments, the plurality of one or more bound oligonucleotides are at least partially complementary to a single-stranded portion of the second adaptor.
[0024] In some embodiments, the first and second strands of the physically linked nucleic acid complexes are amplified by multiple amplification reactions to generate clusters of physically linked nucleic acid complex amplicons on the surface, hi some embodiments, the first and second strands of each of the multiple physically linked nucleic acid complexes are amplified to simultaneously generate multiple clusters on the surface.
[0025] In some embodiments, cleaving a portion of the bound physically linked nucleic acid complex amplicons comprises inefficiently cleaving the cleavable site in the first adaptor, resulting in both cleaved and uncleaved nucleic acid complexes within each cluster on the surface. In some embodiments, the ratio of uncleaved nucleic acid complexes to all nucleic acid complexes within each cluster on the flow cell is 1%, 5%, 10%, 20%, 30%, 40%, 45%, or 50%. In some embodiments, the cleaved nucleic acid complexes are cleaved at the cleavable site in the linker domain of the first adaptor by a cleavage promoter. In some embodiments, the cleavage is a site-directed enzymatic reaction. In some embodiments, the cleavage promoter is an endonuclease. In some embodiments, the endonuclease is a restriction site endonuclease or a targeting endonuclease. In some embodiments, the cleavage promoting factor is selected from the group consisting of a ribonucleoprotein, a Cas enzyme, a Cas9-like enzyme, a meganuclease, a transcription activator-like effector-based nuclease (TALEN), a zinc finger nuclease, an Argonaute nuclease, or a combination thereof. In some embodiments, the cleavage promoting factor comprises a CRISPR-associated enzyme. In some embodiments, the cleavage promoting factor comprises Cas9, or CPF1, or a derivative thereof. In other embodiments, the cleavage promoting factor comprises a nickase or a nickase variant. In some embodiments, the cleavage promoting factor comprises a chemical process.
[0026] In some embodiments, the amount of uncleaved nucleic acid complex remaining on the surface can be expanded by controlling the amount or concentration of the cleavage promoter introduced for site-directed cleavage, or by controlling the amount of time the cleavage promoter is introduced for site-directed cleavage. In some embodiments, the uncleaved nucleic acid complex is protected by adding an anti-cleavage promoter before or during the cleavage step. In some embodiments, the anti-cleavage promoter comprises an anti-cleavage motif in the linker domain of the first adaptor. In some embodiments, the cleavable site is already present in the linker domain of the first adaptor, and the anti-cleavage motif is created by hybridization of an oligonucleotide comprising a sequence at least partially complementary to the linker domain of the first adaptor.
[0027] In some embodiments, cleaving the portion of the bound physically linked nucleic acid complex amplicon further comprises (i) introducing an anti-cleavage promoting factor, and (ii) introducing a cleavage promoting factor either after (i) or simultaneously with (i), wherein interaction with the anti-cleavage promoting factor protects the physically linked nucleic acid complex amplicon from cleavage. In some embodiments, the cleavable site is created by hybridization of an oligonucleotide comprising a sequence at least partially complementary to the linker domain of the first adaptor, and the physically linked nucleic acid complex amplicon that is not hybridized to the oligonucleotide is not cleaved. In some embodiments, the cleavable site is created by hybridization of a first oligonucleotide comprising a sequence at least partially complementary to the linker domain of the adapter, and the anti-cleavage motif is created by hybridization of a second oligonucleotide comprising a sequence at least partially complementary to the linker domain of the adapter. Cleavage of the portion of the bound physically linked nucleic acid complex amplicon further comprises (i) introducing a mixture of the first and second oligonucleotides and (ii) introducing a cleavage-promoting factor. In some embodiments, either the first oligonucleotide or the second oligonucleotide is methylated. In some embodiments, hybridization can be extended by controlling the amount or concentration of the oligonucleotides introduced for hybridization or by controlling the amount of time the oligonucleotides are introduced for hybridization. In some embodiments, the anti-cleavage motif comprises an oligonucleotide sequence with bulky appendages or side chains that prevent access to the cleavage site. In some embodiments, the anti-cleavage motif comprises an oligonucleotide sequence with one or more mismatches that prevent the cleavage-promoting factor from recognizing the cleavage site. In some embodiments, the cleavage-resistant motif comprises one or more of a nucleoside analog, an abasic site, a nucleotide analog, and an oligonucleotide sequence having a peptide-nucleic acid linkage.
[0028] In some embodiments, the cleaved nucleic acid complex is cleaved at the cleavable site in the first adaptor by a catalytically active enzyme, and the uncleaved nucleic acid complex is protected from cleavage in the first adaptor by a catalytically inactive enzyme. In some embodiments, the cleavage site is in the self-complementary portion of the first adaptor or in the single-stranded portion of the first adaptor. In some embodiments, the cleavage site is available when the physically linked nucleic acid complex amplicon is in a self-hybridized configuration on a surface. In some embodiments, the cleavage site is available when the physically linked nucleic acid complex amplicon is in a double-stranded bridge amplified configuration.
[0029] In some aspects, the method further comprises, prior to step (a), selectively enriching physically linked nucleic acid complexes comprising one or more targeted genomic regions to provide a plurality of enriched physically linked nucleic acid complexes.
[0030] Many aspects of the present disclosure can be better understood by reference to the following figures, which together comprise the drawings. These figures are for purposes of illustration only, and not of limitation. The components in the figures are not necessarily to scale. Emphasis instead is placed upon clearly illustrating the principles of the present disclosure. [Brief explanation of the drawings]
[0031] [Figure 1A-1B] 1 is a conceptual diagram of various duplex sequencing method steps in accordance with an embodiment of the present technology. [Figure 2A-2B] 1 illustrates a nucleic acid adapter molecule for use with an embodiment of the present technology, and the formation of a double-stranded adapter-nucleic acid complex as a result of such an adapter being attached to target a double-stranded nucleic acid fragment, in accordance with another embodiment of the present technology. [Figures 3A-3D] 1 illustrates steps in a method for sequencing a double-stranded adaptor-nucleic acid complex, in accordance with one embodiment of the present technology. [Figures 4A-4E]1 illustrates steps in a method for sequencing a double-stranded adaptor-nucleic acid complex in accordance with another embodiment of the present technology. [Figures 5A-5E] 1 illustrates steps in a method for sequencing a double-stranded adaptor-nucleic acid complex in accordance with a further embodiment of the present technology. [Figure 6] 1 illustrates various adapters and their uses in accordance with embodiments of the present technology. [Figure 7] 1 illustrates various adapters and their uses in accordance with embodiments of the present technology. [Figure 8A-8B] 1 illustrates various adapters and their uses in accordance with embodiments of the present technology. [Figure 9A-9B] 1 illustrates various adapters and their uses in accordance with embodiments of the present technology. [Figures 10A-10B] 1 illustrates various adapters and their uses in accordance with embodiments of the present technology. [Figures 11A-11B] 1 illustrates various adapters and their uses in accordance with embodiments of the present technology. [Figures 12A-12C] 1 shows a method for cleaving a double-stranded adaptor-nucleic acid complex in accordance with yet another embodiment of the present technology.
[0032] definition In order that this disclosure may be more readily understood, certain terms are first defined below. Additional definitions for these and other terms are set forth throughout the specification.
[0033] In this application, unless otherwise clear from the context, the term "a" may be understood to mean "at least one." As used in this application, the term "or" may be understood to mean "and / or." As used in this application, the terms "comprising" and "including" may be understood to encompass the itemized element or step, whether stated by itself or with one or more additional elements or steps. When ranges are presented herein, their ends are inclusive. As used in this application, the term "comprise" and variations of this term, such as "comprising" and "comprises," are not intended to exclude other appendants, elements, integers, or steps.
[0034] About: The term "about," when used herein with reference to a value, refers to a similar value in the context of the referenced value. Generally, a person of ordinary skill in the art familiar with the context will understand the appropriate degree of variation encompassed by "about" in that context. For example, in some embodiments, the term "about" can encompass values within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referenced value.
[0035] Analog: As used herein, the term "analog" refers to a substance that shares one or more specific structural features, elements, components, or moieties with a reference substance. Typically, an "analog" exhibits significant structural similarity to the reference substance (e.g., shares a core or consensus structure), but also differs in certain distinct ways. In some embodiments, an analog is a substance that can be produced from a reference substance, e.g., by chemical manipulation of the reference substance. In some embodiments, an analog is a substance that can be produced through the performance of a synthetic process that is substantially similar (e.g., shares multiple steps) to the synthetic process that produces the reference substance. In some embodiments, an analog is produced, or can be produced, through the performance of a synthetic process that is different from that used to produce the reference substance.
[0036] Biological sample: As used herein, the term "biological sample" or "sample" typically refers to a sample obtained from or derived from a biological source of interest (e.g., a tissue or organism or cell culture), as described herein. In some embodiments, the source of interest includes an organism, such as an animal or a human. In other embodiments, the source of interest includes a microorganism, such as a bacterium, a virus, a protozoan, or a fungus. In further embodiments, the source of interest may be a synthetic tissue, organism, cell culture, nucleic acid, or other substance. In still further embodiments, the source of interest may be a plant-based organism. In yet another embodiment, the sample may be an environmental sample, such as, for example, a water sample, a soil sample, an archaeological sample, or other sample collected from a non-living source. In other embodiments, the sample may be a multi-organism sample (e.g., a mixed biological sample). In some embodiments, the biological sample is or includes a biological tissue or a biological fluid. In some embodiments, the biological sample may be or include bone marrow, blood, blood cells, ascites, tissue or fine needle biopsy sample, cell-containing bodily fluid, free-floating nucleic acids, sputum, saliva, urine, cerebrospinal fluid, peritoneal fluid, pleural effusion, feces, lymphatic fluid, gynecological fluid, skin swab, vaginal swab, Pap smear, oral swab, nasal swab, stain or washing such as ductal washing or bronchopulmonary lavage, vaginal fluid, aspirate, scraping, bone marrow specimen, tissue biopsy specimen, fetal tissue or fluid, surgical specimen, feces, other bodily fluid, secretion, and / or excretion, and / or cells therefrom, etc. In some embodiments, the biological sample is or includes cells obtained from an individual. In some embodiments, the obtained cells are or include cells from the individual from whom the sample is obtained. In certain embodiments, the biological sample is a liquid biopsy obtained from a subject. In some embodiments, a sample is a "primary sample" obtained directly from a source of interest by any suitable means. For example, in some embodiments, a primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of bodily fluids (e.g., blood, lymph, stool, etc.), and the like.In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing a primary sample (e.g., by removing one or more components thereof and / or adding one or more agents thereto), e.g., by filtering using a semi-permeable membrane. Such a "processed sample" may contain, for example, nucleic acids or proteins extracted from the sample or obtained by subjecting the primary sample to techniques such as mRNA amplification or reverse transcription, isolation and / or purification of specific components, etc. Cleavage site: Also referred to as a "cleavage motif" and "nick site," this is a bond or bond pair between nucleotides in a nucleic acid molecule. In the case of double-stranded nucleic acid molecules (e.g., double-stranded DNA), the cleavage site may involve bonds (typically phosphodiester bonds) immediately adjacent to each other in the double-stranded molecule so that a "blunt" end is formed after cleavage. The cleavage site may also involve two nucleotide bonds on each single strand of the pair that are not immediately opposite each other, so that a "sticky end" is left when cleaved, thereby leaving a region of single-stranded nucleotides at the end of the molecule. Cleavage sites can be defined by specific nucleotide sequences that can be recognized by enzymes (e.g., restriction enzymes) or other endonucleases with sequence recognition capabilities, such as CRISPR / Cas9. Cleavage sites can be within the recognition sequence of such enzymes (i.e., type 1 restriction enzymes) or adjacent to these enzymes by a defined nucleotide interval (i.e., type 2 restriction enzymes). Cleavage sites can also be defined by the location of modified nucleotides that can be recognized by specific nucleases. For example, abasic sites can be recognized and cleaved by endonuclease VII and the enzyme FPG. Uracil systems can be recognized and rendered into abasic sites by the enzyme UDG. Ribose-containing nucleotides in other DNA sequences can be recognized and cleaved by RNAse H2 when annealed to complementary DNA sequences.
[0037] Determining: Many methodologies described herein include a "determining" step. Those skilled in the art will understand, upon reading this specification, that such "determining" can be utilized or accomplished through the use of any of a variety of techniques available to those skilled in the art, including, for example, the specific techniques explicitly mentioned herein. In some embodiments, determining involves physical manipulation of the sample. In some embodiments, determining involves reviewing and / or manipulating data or information, for example, using a computer or other processing unit adapted to perform the relevant analysis. In some embodiments, determining involves receiving relevant information and / or materials from a source. In some embodiments, determining involves comparing one or more characteristics of the sample or entity to a comparable reference.
[0038] Duplex Sequencing (DS): As used herein, "Duplex Sequencing (DS)" refers, in its broadest sense, to an error-correction method that achieves exceptional accuracy by comparing sequences from both strands of individual DNA molecules.
[0039] Error corrected: As used herein, the term "error corrected" or "error correction" refers to the product or process that results in identifying and then disregarding, excluding, or otherwise correcting one or more nucleotide errors in a region of a nucleic acid molecule where the two strands of a double-stranded portion of the nucleic acid molecule are not fully complementary to one another (e.g., due to a nucleotide mismatch). In some embodiments, the mismatch may be the result of a point mutation, deletion, insertion, or chemical modification. In some embodiments, the mismatch includes base pairs on opposite strands of the sequence (e.g., but not limited to, AA, CC, TT, GG, AC, AG, TC, TG), or the inverse of these pairs (which are equivalent, i.e., AG is equivalent to GA), deletions, insertions, or other modifications to one or more of the bases. The mismatch may be biologically derived, DNA synthetically derived, or a damaged or modified nucleotide base caused by the mismatch. In some embodiments, the damaged or modified nucleotide base is present in one or both strands and has been converted to a mismatch by an enzymatic process (e.g., a DNA polymerase, a DNA glycosylase, or another nucleic acid-modifying enzyme or chemical process). In some embodiments, this mismatch can be used to infer the presence of nucleic acid damage or nucleotide modification prior to the enzymatic process or chemical treatment.
[0040] Expression: As used herein, "expression" of a nucleic acid sequence refers to one or more of the following events: (1) production of an RNA template from a DNA sequence (e.g., by transcription), (2) processing of the RNA transcript (e.g., by splicing, editing, 5' capping and / or 3' end formation), (3) translation of the RNA into a polypeptide or protein, and / or (4) post-translational modification of the polypeptide or protein.
[0041] Functionalized Surface: As used herein, the term "functionalized surface" refers to a solid surface, bead, or other fixed structure capable of binding or immobilizing a nucleic acid molecule or other capture moiety. In some embodiments, the functionalized surface comprises a binding moiety capable of capturing a target nucleic acid. In some embodiments, the binding moiety is directly linked to a surface. In some embodiments, an oligonucleotide at least partially complementary to the target nucleic acid serves as the binding moiety. In some embodiments, the oligonucleotide is covalently attached to the surface. In some embodiments, the functionalized surface may include controlled pore glass (CPG), magnetic porous glass (MPG), among other glass or non-glass surfaces. In one embodiment, the functionalized surface may be a sequencing surface, such as the surface of a flow cell. Chemical functionalization may involve ketone modification, aldehyde modification, thiol modification, azide modification, and alkyne modification, among others. In some embodiments, the functionalized surfaces and oligonucleotides used for hybridization capture are linked using one or more of the following immobilization chemistries: amide, alkylamine, thiourea, diazo, and hydrazine bond-forming, among other surface chemistries. In some embodiments, the functionalized surfaces and oligonucleotides used for hybridization capture are linked using one or more of the following reagents: EDAC, NHS, sodium periodate, glutaraldehyde, pyridyl disulfide, nitric acid, and biotin, among other linking reagents.
[0042] gRNA: As used herein, "gRNA" or "guide RNA" refers to a short RNA molecule that includes a scaffold sequence suitable for binding a targeting endonuclease (such as a Cas enzyme (e.g., Cas9 or Cpfl) or another ribonucleoprotein with similar properties) to a substantially target-specific sequence that facilitates cleavage of a specific region of DNA or RNA.
[0043] Mutation: As used herein, the term "mutation" refers to a change in nucleic acid sequence or structure relative to a reference sequence. Mutations to a polynucleotide sequence may include point mutations (e.g., single-base mutations), multiple-nucleotide mutations, nucleotide deletions, sequence rearrangements, nucleotide insertions, and DNA sequence duplications in a sample, among other complex multi-nucleotide changes. Mutations may occur on both strands of a double-stranded DNA molecule as complementary base changes (i.e., true mutations), or as mutations present in one strand but not the other, which can either be repaired, disrupted, or misrepaired / converted into a true double-strand mutation (i.e., heteroduplexes). The reference sequence may be present in a database (i.e., the HG38 human reference genome) or in the sequence of another sample to which the sequence is compared. Mutations are also known as genetic variants.
[0044] Nucleic acid: As used herein, in its broadest sense, refers to any compound and / or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, nucleic acids are compounds and / or substances that are or can be incorporated into an oligonucleotide chain via a phosphodiester bond. As will be clear from the context, in some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides), and in some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, "nucleic acid" is or comprises RNA, and in some embodiments, "nucleic acid" is or comprises DNA. In some embodiments, a nucleic acid is, comprises, or consists of one or more naturally occurring nucleic acid residues. In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleic acid analogs. In some embodiments, a nucleic acid analog differs from a nucleic acid in that it does not utilize a phosphodiester backbone. For example, in some embodiments, the nucleic acid is, comprises, or consists of one or more "peptide nucleic acids," which have peptide bonds instead of phosphodiester bonds in the backbone, and are known in the art and considered within the scope of the present technology. Alternatively or additionally, in some embodiments, the nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoramidite linkages rather than phosphodiester linkages. In some embodiments, the nucleic acid is, comprises, or consists of one or more naturally occurring nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine).In some embodiments, the nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynylcytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, the nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, hexose, or locked nucleic acids) compared to sugars commonly present in natural nucleic acids. In some embodiments, the nucleic acid has a nucleotide sequence that encodes a functional gene product such as RNA or a protein. In some embodiments, the nucleic acid comprises one or more introns. In some embodiments, the nucleic acid may be a non-protein-coding RNA product, e.g., a microRNA, a ribosomal RNA, or a CRISPR / Cas9 guide RNA. In some embodiments, the nucleic acid serves a regulatory purpose in a genome. In some embodiments, the nucleic acid does not originate from a genome. In some embodiments, the nucleic acid comprises an intergenic sequence. In some embodiments, the nucleic acid is derived from an extrachromosomal element or a non-nuclear genome (such as a mitochondria or chloroplast). In some embodiments, the nucleic acid is prepared by one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), reproduction in a recombinant cell or system, and chemical synthesis.In some embodiments, the nucleic acid is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues in length. In some embodiments, the nucleic acid is partially or completely single-stranded, and in some embodiments, the nucleic acid is partially or completely double-stranded. In some embodiments, the nucleic acid has a nucleotide sequence that encodes a polypeptide or includes at least one element that is the complement of a sequence that encodes a polypeptide. In some embodiments, the nucleic acid has enzymatic activity. In some embodiments, the nucleic acid performs a mechanical function, for example, in a ribonucleic acid protein complex or transfer RNA. In some embodiments, the nucleic acid functions as an aptamer. In some embodiments, the nucleic acid may be used for data storage. In some embodiments, the nucleic acid may be chemically synthesized in vitro.
[0045] Reference: As used herein, describes a standard or control relative to which a comparison is made. For example, in some embodiments, an agent, animal, individual, collection, sample, sequence, or value of interest is compared to a reference or control agent, animal, individual, collection, sample, sequence, or value. In some embodiments, the reference or control is tested and / or determined substantially simultaneously with the test or determination of interest. In some embodiments, the reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as will be understood by those skilled in the art, a reference or control is determined or characterized under conditions or circumstances comparable to those being evaluated. Those skilled in the art will understand when sufficient similarity exists to justify reliance on and / or comparison to a particular possible reference or control.
[0046] Sequence read: As used herein, the term "sequence read" or "sequencing read" refers to nucleic acid sequence data corresponding to a reference or target nucleic acid molecule. In some embodiments, the data is a predicted sequence of base pairs (or base pair probabilities) corresponding to all or a portion (e.g., a fragment or portion thereof) of the reference or target nucleic acid molecule processed by a sequencing platform. Sequence read lengths can range from a few base pairs (bp) to hundreds of kilobases (kb). Sequence read length can be affected by the size or length of the reference or target nucleic acid molecule and the sequencing platform used. In some embodiments, the sequence reads are generated using a sequencing technology, such as, but not limited to, a next-generation sequencing platform, such as Illumina® HiSeq®, Illumina® NovaSeq®, Illumina® NextSeq®, Illumina® MiSeq®, Illumina® iSeq®, Oxford Nanopore Sequencing System, ThermoFisher® Ion Torrent® Sequencing System, Roche 454 GS System®, Illumina Genome Analyzer®, Applied Biosystems SOLiD System®, Helicos Heliscope®, Complete Genomics®, and Pacific Biosciences SMRT®.
[0047] Single Molecular Identifier (SMI): As used herein, the term "single molecular identifier" or "SMI" (which may be referred to by other names, such as "tag," "barcode," "molecular barcode," "unique molecular identifier," or "UMI") refers to any substance (e.g., nucleotide sequence, nucleic acid molecular feature) capable of distinguishing an individual molecule within a larger heterogeneous collection of molecules. In some embodiments, the SMI may be or include an exogenously applied SMI. In some embodiments, the exogenously applied SMI may be or include a degenerate or semi-degenerate sequence. In some embodiments, a substantially degenerate SMI may be known as a random unique molecular identifier (R-UMI). In some embodiments, the SMI may include a code (e.g., a nucleic acid sequence) from within a pool of known codes. In some embodiments, a predetermined SMI code is known as a defined unique molecular identifier (DUMI). In some embodiments, the SMI may be or include an endogenous SMI. In some embodiments, endogenous SMI may be or include information related to specific shear points of a target sequence or features associated with the ends of individual molecules comprising the target sequence. In some embodiments, SMI may be related to sequence variations in nucleic acid molecules due to random or semi-random damage, chemical, enzymatic, or other modifications to the nucleic acid molecule. In some embodiments, the modification may be deamination of methylcytosine. In some embodiments, the modification may involve the site of a nucleic acid nick. In some embodiments, SMI may include both exogenous and endogenous elements. In some embodiments, SMI may include physically adjacent SMI elements. In some embodiments, SMI elements may be spatially distinct within the molecule. In some embodiments, SMI may be non-nucleic acid. In some embodiments, SMI may include two or more different types of SMI information.Various embodiments of the SMI are further disclosed in International Patent Publication No. WO 2017 / 100441, which is incorporated herein by reference in its entirety.
[0048] Strand Defining Element (SDE): As used herein, the term "strand defining element" or "SDE" refers to any substance that allows for the identification of a particular strand of double-stranded nucleic acid material and therefore its differentiation from other / complementary strands (e.g., any substance that renders the amplification products of two single-stranded nucleic acids obtained from a target double-stranded nucleic acid substantially distinguishable from one another after sequencing or other nucleic acid interrogation). In some embodiments, an SDE may be or include one or more segments of substantially non-complementary sequence within an adapter sequence. In certain embodiments, the segment of substantially non-complementary sequence within an adapter sequence may be provided by an adapter molecule comprising a Y-shape or "loop" shape. In other embodiments, the segment of substantially non-complementary sequence within an adapter sequence may form an unpaired "bubble" in the middle of adjacent complementary sequences within the adapter sequence. In other embodiments, an SDE may include a nucleic acid modification. In some embodiments, an SDE may include physical separation of paired strands into physically separate reaction compartments. In some embodiments, an SDE may include a chemical modification. In some embodiments, the SDE may comprise a modified nucleic acid. In some embodiments, the SDE may be associated with sequence variation in a nucleic acid molecule due to random or semi-random damage, chemical modification, enzymatic modification, or other modification to the nucleic acid molecule. In some embodiments, the modification may be deamination of methylcytosine. In some embodiments, the modification may involve the site of a nucleic acid nick. Various embodiments of SDEs are further disclosed in International Patent Publication No. 2017 / 100441, the entire contents of which are incorporated herein by reference.
[0049] Subject: As used herein, the term "subject" refers to an organism, typically a mammal (e.g., a human, including, in some embodiments, prenatal human forms). In some embodiments, the subject is afflicted with a relevant disease, disorder, or condition. In some embodiments, the subject is predisposed to a disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of a disease, disorder, or condition. In some embodiments, the subject does not exhibit any symptoms or characteristics of a disease, disorder, or condition. In some embodiments, the subject is someone who has one or more characteristics characteristic of a susceptibility to, or risk for, a disease, disorder, or condition. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual for whom and / or after diagnosis and / or therapy is to be performed.
[0050] Substantially: As used herein, the term "substantially" refers to a qualitative state indicating the total or nearly total extent or degree of a desired characteristic or property. Those skilled in the art of biology will understand that biological and chemical phenomena rarely, if ever, go completely to completion and / or move toward complete completion, or achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.
[0051] Variant: As used herein, the term "variant" refers to an entity that exhibits significant structural identity with a reference entity but differs structurally from the reference entity in the presence or level of one or more chemical moieties relative to the reference entity. In the context of nucleic acids, a variant nucleic acid may have a distinctive sequence element consisting of multiple nucleotide residues that have a designated position relative to another nucleic acid in linear or three-dimensional space. Homologous sequences differ by one or more variants. For example, a variant polynucleotide (e.g., DNA) may differ from a reference polynucleotide as a result of one or more differences in the nucleic acid sequence. In some embodiments, a variant polynucleotide sequence contains an insertion, deletion, substitution, or mutation relative to another sequence (e.g., a reference sequence or another polynucleotide (e.g., DNA) sequence in a sample). Examples of variants include SNPs, SNVs, CNVs, CNPs, MNVs, MNPs, mutations, cancer mutations, driver mutations, passenger mutations, and inherited polymorphisms. DETAILED DESCRIPTION OF THE INVENTION
[0052] The present technology generally relates to methods for providing error-corrected sequence reads for nucleic acid material using duplex sequencing, and related reagents for use in such methods. Some embodiments of the present technology are directed to methods for achieving highly accurate sequencing reads that are provided at a faster rate (e.g., in fewer steps) and / or at a lower cost (e.g., utilizing fewer reagents), yielding more desired data. Other aspects of the present technology are directed to methods and reagents for increasing the conversion efficiency (i.e., the percentage of nucleic acid molecules for which sequences are generated) for duplex sequencing. Various aspects of the present technology have many applications in preclinical and clinical testing and diagnostics, as well as other applications.
[0053] Specific details of several embodiments of the present technology are described below with reference to Figures 1A-12C. While many of the embodiments are described herein with reference to duplex sequencing, other sequencing modes capable of generating error-corrected sequencing reads and providing sequence information, in addition to those described herein, are within the scope of the present technology. Furthermore, other embodiments of the present technology may have different configurations, components, or procedures than those described herein. Thus, those skilled in the art will understand that the present technology may include other embodiments having additional elements, and may also include other embodiments that do not include some of the features shown and described below with reference to Figures 1A-12C.
[0054] Regarding the efficiency of duplex sequencing process or other high-precision sequencing methods, conversion efficiency can be defined as the fraction of unique nucleic acid molecules that are input into sequencing library preparation reaction and produce at least one double-stranded consensus sequence read (or other high-precision sequence read).In some cases, the drawback of conversion efficiency may limit the usefulness of high-precision duplex sequencing for some applications that would otherwise be very well suited.For example, low conversion efficiency may result in a situation where the copy number of target double-stranded nucleic acid is limited, which may result in less than desired amount of sequence information being produced.A non-limiting example of this concept includes DNA from circulating tumor cells or cell-free DNA from tumors, or prenatal infants that flow into body fluids such as plasma and mix with excess DNA from other tissues. Other non-limiting examples include forensic materials, such as those left at crime scenes in limited quantities; ancient DNA, such as that found in archaeological sites; very small biopsies, such as those obtained by needle biopsy, aspirates, or endoscopies; small amounts of formalin-fixed clinical material; microdissected samples; small biological regions or samples from human or non-human organisms; samples or hair, blood spots, or other biological materials (including single cells or a small number of cells) produced by or derived from multicellular or unicellular organisms in limited quantities. Duplex sequencing typically has an accuracy capable of resolving one mutant molecule among more than 100,000 unmutated molecules. However, for example, if there are only 10,000 available molecules in a sample (e.g., 10,000 genome equivalents in the case of a single-copy gene or locus), even if the ideal efficiency of converting these into double-stranded consensus sequence reads is 100%, the minimum measurable mutation frequency will be 1 / (10,000*100%) = 1 / 10,000. For clinical diagnostics, it may be important to have maximum sensitivity for detecting low-level signals of cancer or treatment- or diagnostic-related mutations, so a relatively low conversion efficiency would be undesirable in this regard. Similarly, forensic applications often have very little DNA available for testing.When only nanogram or picogram quantities can be recovered from a crime scene or natural disaster scene, and / or when DNA from multiple individuals is mixed, it can be important to have maximum conversion efficiency in being able to detect the presence of DNA from all individuals in the mixture.
[0055] Methods incorporating duplex sequencing and other sequencing modalities may include connecting (e.g., ligating) one or more sequencing adapters to a target double-stranded nucleic acid molecule to generate a double-stranded target nucleic acid complex. Such adapter molecules may include one or more of a variety of features suitable for massively parallel sequencing platforms, such as a sequencing primer recognition site, an amplification primer recognition site, a barcode (e.g., a single molecule identifier (SMI)) sequence (also known as a unique molecular identifier (UMI)), an index sequence, a single-stranded portion, a double-stranded portion, a strand-distinguishing element or feature, etc. As noted above, obtaining duplex sequencing information requires successful recovery of sequence information from both strands of an original double-stranded molecule. Aspects of the present disclosure provide methods and reagents for generating and linking sequencing information from both strands of an original double-stranded molecule via physically linking the strands prior to amplification and sequencing.
[0056] I. Selected Embodiments of Duplex Sequencing Methods and Associated Adapters and Reagents Duplex sequencing is a method for generating error-corrected DNA sequences from double-stranded nucleic acid molecules, and was originally described in International Patent Publication No. 2013 / 142389 and US Patent No. 9,752,188, both of which are incorporated herein by reference in their entirety.In a specific embodiment of the present technology, duplex sequencing can be used to sequence both strands of individual DNA molecules in such a way that derivative sequence reads can not only be recognized as coming from the same double-stranded nucleic acid parent molecule during massively parallel sequencing (MPS) (also commonly known as next-generation sequencing (NGS)), but can also be distinguished from each other as distinct entities after sequencing.The sequence reads obtained from each strand are then compared to obtain the error-corrected sequence of original double-stranded nucleic acid molecules.
[0057] 1 is a conceptual diagram of various steps of a duplex sequencing method according to an embodiment of the present technology. In certain embodiments, a method incorporating duplex sequencing may include ligating one or more sequencing adaptors to a plurality of target double-stranded nucleic acid molecules, each comprising a first-strand target nucleic acid sequence and a second-strand target nucleic acid sequence, to generate a plurality of double-stranded target nucleic acid complexes (FIG. 1A). After the double-stranded nucleic acid library preparation is generated, the complexes may be subjected to DNA amplification (e.g., using PCR) or any other biochemical method of DNA amplification (e.g., rolling circle amplification, multiple displacement amplification, isothermal amplification, bridge amplification, polony amplification, isothermal amplification, or surface-bound amplification), resulting in the generation of one or more copies of the first-strand target nucleic acid sequence and one or more copies of the second-strand target nucleic acid sequence (e.g., FIG. 1A). One or more amplified copies of the first strand target nucleic acid molecule and one or more amplified copies of the second target nucleic acid molecule can then be subjected to DNA sequencing, preferably using a "next generation" massively parallel DNA sequencing platform (e.g., Figure 1A).
[0058] After sequencing, sequence reads generated from a first strand of a target nucleic acid molecule are compared with sequence reads generated from a second strand of the same target nucleic acid molecule. In some embodiments, more than one sequence read may be generated from the first and second strands. Upon comparison, an error-corrected target nucleic acid molecule sequence can be generated (e.g., FIG. 1B). For example, nucleotide positions where bases from both the first and second strand target nucleic acid sequences match are considered to be true sequences, while nucleotide positions where there is no match between the two strands are recognized as potential sites of technical error, which may be disregarded, excluded, corrected, or otherwise identified. In some embodiments, if a nucleotide position does not match, the site may be identified as unknown (e.g., indicated as "N" in FIG. 1B). In this manner, an error-corrected sequence of the original double-stranded target nucleic acid molecule can be generated (as shown in FIG. 1B). Optionally, in some embodiments, after separately grouping each of the sequencing reads generated from the first strand target nucleic acid molecule and the second strand target nucleic acid molecule, a single-stranded consensus sequence can be generated for each of the first strand and the second strand. The single-stranded consensus sequences from the first strand target nucleic acid molecule and the second strand target nucleic acid molecule can then be compared to generate an error-corrected target nucleic acid molecule sequence (e.g., FIG. 1B).
[0059] Alternatively, in some embodiments, a sequence mismatch between the two strands can be recognized as a potential biologically derived mismatch in the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, a sequence mismatch between the two strands can be recognized as a potential DNA synthesis-derived mismatch in the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, a sequence mismatch between the two strands can be recognized as a potential site where a damaged or modified nucleotide base is present in one or both strands and converted into a mismatch by an enzymatic process (e.g., DNA polymerase, DNA glycosylase, or another nucleic acid-modifying enzyme or chemical process). In some embodiments, the modified nucleotide base is 5-methyl-cytosine, 8-oxo-guanine, ribose base, abasic nucleotide, or uracil nucleotide. In some embodiments, this latter finding can be used to infer the presence of nucleic acid damage or nucleotide modification prior to enzymatic or chemical processing.
[0060] In certain embodiments, as described in U.S. Pat. No. 9,752,188 and International Patent Publication No. 2017 / 100441, first-strand sequencing reads and second-strand sequencing reads from each original double-stranded nucleic acid molecule can be associated (e.g., grouped) using (a) single molecule identifier (SMI) sequences associated with adapters during library preparation, (b) fragment features associated with the original double-stranded molecule, e.g., sequences at, near, or to the fragment ends, and (c) combinations thereof.
[0061] In one embodiment, the generation of raw sequence reads for use in duplex sequencing embodies the use of a target double-stranded nucleic acid molecule with a hairpin adapter attached to one end of the molecule and a "Y"-shaped adapter attached to the other end of the molecule. This ligated or double-stranded complex, including both the first and second strands of the original double-stranded nucleic acid molecule, can be further amplified using any type of amplification (e.g., PCR or bridging) and then subjected to massively parallel sequencing (e.g., sequencing-by-synthesis, next-generation sequencing (NGS)), etc.) to generate sequence reads for use in duplex sequencing. Adapter-double-stranded nucleic acid complexes with hairpin adapters (i.e., "looped" or "U"-shaped) allow for the generation of sequence reads from both the original first and second strands of the target double-stranded nucleic acid molecule in a non-limiting example, in a manner that allows the sequence reads to be grouped by the nature of the position of the sequencing reaction on the flow cell surface (in the case of sequencing by synthesis) or by the nature of the position of the sequencing reaction / process in other ways.
[0062] Aspects of the present technology are directed to methods and reagents for associating and / or grouping first and second strand sequencing reads by physically linking the first and second strands in such a way that the sequencing information from both strands is associated with each other by the nature of their physical linkage (e.g., for error correction). In certain embodiments, a method for preparing a sequencing library for use in duplex sequencing may include ligating a hairpin adapter to one end of a target double-stranded nucleic acid molecule and ligating a "Y"-shaped adapter to the opposite end of the same target double-stranded nucleic acid molecule. In one embodiment, the hairpin adapter molecule includes a cleavable hairpin adapter element for targeted separation of the first and second strands of the target double-stranded nucleic acid molecule.
[0063] In some embodiments, the association of the first-strand sequence read and the second-strand sequencing read can be achieved during or after the sequencing reaction on a sequencer. For example, in certain embodiments, the first and second strands of a double-stranded nucleic acid molecule are linked by an intervening linker domain, such as a hairpin adapter sequence. In one embodiment, sequence information from both strands of the original nucleic acid molecule is generated within the same clonal cluster on an MPS sequencer (e.g., on a flow cell). Sequencing the linked first and second strands on a sequencer presents challenges because self-complementary hairpin sequences may preferentially hybridize on the sequencing surface or in solution, impairing polymerase extension. Certain aspects of the present technology disclose methods for overcoming these challenges associated with self-complementary hybridization of the linked first and second strands while still allowing sequencing reads to be obtained from both the first and second strands within the same clonal cluster on a sequencer.
[0064] Adapters and adapter sequences Adapter molecules containing primer sites, flow cell sequences, and / or other features such as SMIs (e.g., molecular barcodes) or SDEs in various configurations are contemplated for use with many of the embodiments disclosed herein. In some embodiments, the provided adapters may be or include one or more sequences that are complementary or at least partially complementary to PCR primers (e.g., primer sites) that have at least one of the following properties: (1) high target specificity, (2) amenable to multiplexing, and (3) exhibit robust and minimally biased amplification.
[0065] In some embodiments, adapter molecules may be "Y"-shaped, "U"-shaped, "hairpin"-shaped, have a bubble (e.g., a portion of non-complementary sequence), or other features. In other embodiments, adapter molecules may include a "Y"-shaped, "U"-shaped, "hairpin"-shaped, or bubble. For purposes of this disclosure, "U"-shaped or "hairpin"-shaped adapters may both be used to collectively refer to adapters having a linker domain that ligates or connects the first strand of a target double-stranded nucleic acid molecule to the second strand of the same molecule. Certain adapters may contain modified or non-standard nucleotides, restriction sites, or other features for engineering structure or function in vitro. Adapter molecules may be ligated to a variety of nucleic acid entities with termini. For example, the adapter molecule may have a T-overhang, an A-overhang, a CG-overhang, a multi-nucleotide overhang (also referred to herein as a "sticky end" or "sticky overhang"), or a single-stranded overhang region of a known nucleotide length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides), a dehydroxylated base, or be suitable for ligation to a blunt end of a nucleic acid entity, where the end of the molecule is dephosphorylated 5' of the target or otherwise blocked from conventional ligation. In other embodiments, the adapter molecule may contain a dephosphorylated or other ligation-preventing modification on the 5' strand of the ligation site. In the latter two embodiments, such a strategy may be useful to prevent dimerization of the library fragments or the adapter molecule.
[0066] Figure 2A illustrates a nucleic acid adapter molecule for use with some embodiments of the present technology, as well as a double-stranded adapter-nucleic acid complex resulting from ligating the adapter molecule to a double-stranded nucleic acid fragment in accordance with an embodiment of the present technology. As shown in Figure 2A, the first adapter molecule (Adapter 1) may be a Y-shaped adapter molecule having a first primer site and a second primer site (labeled Primer Site 1 and Primer Site 2) suitable for ligation to a double-stranded nucleic acid fragment via a T-overhang. The second adapter molecule (Adapter 2) suitable for ligation to a target nucleic acid fragment via a T-overhang is shown as a hairpin adapter containing a single-stranded linking domain. Sequencing library generation of a collection of double-stranded nucleic acid fragments can include ligating a pool of adapters containing both Adapter 1 and Adapter 2 to the collection of double-stranded nucleic acid fragments. Figure 2A illustrates one product resulting from the described ligation reaction. Other products include an adapter-nucleic acid complex containing Adapter 1 at both ends and an adapter-nucleic acid complex containing Adapter 2 at both ends. In various embodiments described herein, it is desirable to generate an adaptor-nucleic acid complex as shown in Figure 2A for use with duplex sequencing methods.
[0067] FIG. 2B illustrates another embodiment in which a target double-stranded nucleic acid fragment comprises sticky end 1 at one end of the fragment and sticky end 2 at the opposite end of the fragment. By design, the sequence of sticky end 1 (the overhang at the 5' end of the target fragment) is known. Similarly, the sequence of sticky end 2 (the overhang at the 3' end of the target fragment) is known. In one embodiment, the sequence of sticky end 1 is different from the sequence of sticky end 2. In another embodiment, the sequence of sticky end 1 is a different length than the sequence of sticky end 2. In a further embodiment, sticky end 1 is a 5' overhang and sticky end 2 is a 3' overhang. Specific adapters comprising substantially complementary sequences can be synthesized so that fragments can be joined to the adapters at both ends. In one embodiment, the adapters can be different (e.g., adapter 1 can comprise a Y-shape and adapter 2 can comprise a U-shape). In other embodiments (not shown), the adapters can be the same type of adapter (e.g., adapters comprising a Y-shape, a U-shape, a barcoded adapter, etc.). As shown in Figure 2B, this design allows each target double-stranded nucleic acid molecule to have a Y-shaped adapter at one end and a hairpin (e.g., an adapter with a linking domain) at the other end. Thus, upon denaturation, the adapter-nucleic acid complex contains a single-stranded molecule comprising a first primer site, a first strand, a linking domain, a second strand, and a second primer site. There may be advantages in other applications to designing specific adapters to be located at either the 5' or 3' end of a fragment. The specificity of the substantially unique sticky ends on the target fragments facilitates these types of applications. Furthermore, positive selection of successfully cleaved, adapter-ligated target fragments can ensure amplification and sequencing of only enriched target nucleic acid regions.
[0068] Thus, in some embodiments, a set of adaptor molecules can comprise sticky overhangs that are different, unique, or semi-unique with respect to other sets of adaptor molecules. The number of different types of sticky ends can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. The number of different types of sticky ends can be about 11, 12, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 150, 200, 300, 400, 500, 750, 1000, or more. In a particular example, the hairpin adapter molecule can include a first sticky overhang suitable for ligating to the sticky end of a first complementary fragment, and the Y-shaped adapter can include a second sticky overhang suitable for ligating to the sticky end of a second complementary fragment. Thus, preparing a sequencing library of a collection of nucleic acid molecules can include generating nucleic acid fragments having first and second sticky ends and ligating the nucleic acid fragments to the hairpin and Y-shaped adapters. The resulting sequencing library can include a plurality of double-stranded adapter-nucleic acid fragment complexes, each having a hairpin adapter at its first end and a Y-shaped adapter at its second end.
[0069] amplification In one embodiment, the method can include amplifying an adaptor-nucleic acid complex containing both the first and second strands on a sequencer surface, such as the surface of a flow cell. In some embodiments, amplification on a surface, e.g., bridge amplification on the surface of a flow cell, involves generating clusters or multiple copies of the bound nucleic acid template. In certain embodiments, the linked first and second strand nucleic acid templates can be bridge-amplified on the surface of a flow cell to generate, for example, multiple clonal clusters, each containing nucleic acid template copies derived from both the original first and second strands of the original double-stranded nucleic acid molecule. Some of the clonal copies in a cluster are forward-sense and the remaining copies are reverse-sense. Those skilled in the art will appreciate various embodiments, such as polony amplification, cluster amplification, and bridge amplification, that use amplification involving flowing the adaptor-nucleic acid complex over a surface and providing bound oligonucleotides that are at least partially complementary to a region of the Y-shaped adaptor. The surface can provide one or more oligonucleotides complementary to portions of the adaptor. In practice, both arms of the Y-shaped adaptor can hybridize to the surface of the flow cell.
[0070] Bridge amplification (not shown) can be used to generate multiple copies of the complex to form colonies or clusters (also referred to herein as clonal clusters), each of which contains multiple copies of the original molecule (e.g., adapter-nucleic acid complex) in both the forward and reverse directions.
[0071] In one embodiment, the sequencing reaction can proceed when either the forward or reverse copy is cleaved and removed. Figure 3A illustrates a step in the process after bridge amplification of an adaptor-nucleic acid complex (e.g., a double-stranded nucleic acid complex) and after the copy containing the forward copy (e.g., nucleic acid sequence "2" is attached to the surface of the flow cell) is cleaved and removed. As shown in Figure 3A, the remaining complex is in the reverse orientation (e.g., nucleic acid sequence "1" is attached to the surface of the flow cell, e.g., with the 3' end of the molecule attached to the surface). In one embodiment, the nucleic acid sequence of the first strand readily hybridizes to the complementary nucleic acid sequence of the second strand, making synthesis of longer complexes difficult for sequencing. The bound copy of the illustrated complex contains a linker domain provided by a hairpin adaptor (e.g., adaptor 2, Figures 2A and 2B). In some embodiments, the linker domain contains a cleavable site or motif ("C"). Cleavable site C may comprise a nucleotide sequence, a single nucleotide base, a modified base, or other enzymatic or non-enzymatic cleavable feature.
[0072] As shown in FIG. 3B, the process may include a step comprising cleaving cleavable site C to separate the first strand sequence from the second strand sequence. In one embodiment, the cleavage event at site C may be facilitated by a cleavage-facilitating agent (e.g., an enzyme, a chemical, etc.). In one embodiment, the cleavage step may be inefficient, such that only a portion of the complex is cleaved at site C. As such, a portion of the complex (e.g., about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 45%, about 50%, or more or less, between about 1% and about 10%, between about 10% and about 25%, between about 25% and about 45%, more than 50%, less than 10%, etc.) may remain uncleaved, and the first and second strand sequences remain linked. In some embodiments, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the complex is cleaved, e.g., at site C.
[0073] Cleavage at site C separates the first strand from the second strand, washing away any unbound strand (e.g., adjacent nucleic acid sequence 2). For example, as shown in Figure 3C, a portion of the complex cleaved at site C contains only the nucleotide sequence of the first strand and a portion of the hairpin adapter. Because the complex no longer self-hybridizes, a sequencing reaction can be performed to generate sequencing reads for the remaining first strand in the clonal cluster using a primer specific to the adapter (e.g., at or around nucleotide sequence 1 at the 3' end of the bound molecule) (Figure 3D). Index reads can also be generated (not shown). Note that the first-strand sequencing reads are single-end sequence reads. Complexes that remain uncleaved in the clonal cluster likely will not be successfully sequenced during the sequencing reaction due to the difficulty of displacing the longer second strand with the sequencing primer (Figure 3D).
[0074] After obtaining sequencing information from the first strand present in the clonal cluster, the next step in the process involves a second round of amplification (e.g., bridge amplification) to provide more copies of the uncleaved complex. Bridge amplification requires the presence of both nucleic acid sequence 1 and nucleic acid sequence 2 present on the full-length complex. Only the remaining uncleaved complex still has both adapter sequences present. Therefore, the clonal cluster can be relocated by bridge amplification using the remaining oligonucleotides bound to the surface of the flow cell (Figure 4A).
[0075] After amplification, when any of the reverse copies are cleaved and removed, a second sequencing reaction can proceed. Figure 4B shows a step in the process after bridge amplification of an adapter-nucleic acid complex (e.g., a double-stranded nucleic acid complex) and after the copy containing the reverse copy (e.g., nucleic acid sequence "1" is attached to the surface of the flow cell) is cleaved and removed. As shown in Figure 4B, the remaining complex is in the forward direction (e.g., nucleic acid sequence "2" is attached to the surface of the flow cell, e.g., the 5' end of the molecule is attached to the surface). As mentioned above, the nucleic acid sequences of the first and second strands hybridize easily, making synthesis of longer complexes difficult for sequencing.
[0076] As shown in FIG. 4C, the process may include a step comprising cleaving cleavable site C to separate the second strand sequence from the first strand sequence. In one embodiment, the cleavage event at site C may be facilitated by a cleavage-facilitating agent (e.g., an enzyme, a chemical, etc.). As discussed above, the cleavage step may be inefficient, such that only a portion of the complex is cleaved at site C. As such, a portion of the complex (e.g., about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 45%, about 50%, or more or less, between about 1% and about 10%, between about 10% and about 25%, between about 25% and about 45%, more than 50%, less than 10%, etc.) may remain uncleaved, and the first and second strand sequences remain linked. Alternatively, the cleavage step may be efficient and cleave the entire complex (eg, as shown in Figure 4C).
[0077] When the second strand is separated from the first strand by cleavage at site C, the unbound strand (e.g., adjacent nucleic acid sequence 1) is washed away. For example, as shown in Figure 4D, the portion of the complex cleaved at site C contains only the nucleotide sequence of the second strand and a portion of the hairpin adapter. Because the complex no longer self-hybridizes, a sequencing reaction can be performed using a primer specific to the remaining portion of the hairpin adapter to generate sequencing reads for the second strand remaining in the clonal cluster (Figure 4E). Index reads can also be generated (not shown). Note that the sequencing reads for the second strand are single-ended sequence reads. Once sequence reads derived from both the first and second strands (e.g., within the same clonal cluster) are generated, they can be compared for error correction.
[0078] Figures 5A-5E illustrate another embodiment of double-stranded complex sequencing to provide duplex sequencing information on a sequencing surface (e.g., a flow cell). In the embodiment shown in Figures 5A-5E, sequence reads from both the first and second strands of the original adaptor-nucleic acid complex can be generated without a second bridge amplification step. As described above, each double-stranded complex can be independently bridge-amplified on a surface to generate clonal clusters containing multiple copies of the double-stranded complex, each having both a first strand and a complementary second strand with an intervening hairpin linker domain containing a cleavable site (Figure 5A). The copies may be in both the forward and reverse directions, as described above.
[0079] As shown in Figure 5B, in one embodiment, the double-stranded complex may be cleaved at cleavage site C (e.g., via a cleavage-promoting agent, discussed further herein). After cleavage at site C, the unbound strand is removed. Referring to Figure 5C, the remaining molecules bound to the surface of the flow cell comprise (a) the sequence of the first strand in the reverse orientation (e.g., adjacent to primer site "1") and (b) the sequence of the second strand in the forward orientation (e.g., adjacent to primer site "2").
[0080] In the next step, a first sequencing reaction using a reverse-specific primer is used to obtain sequencing information for the first strand (Figure 5D). The primer used in the first sequencing reaction can be washed away. In the next step, a second sequencing reaction using a forward-specific primer is used to obtain sequencing information for the second strand (Figure 5E). The embodiment shown in Figures 5D and 5E illustrates sequential sequencing of the first and second strands. It will be understood that in another embodiment, the first and second strands can be sequenced simultaneously (e.g., in the same sequencing reaction), for example, by multicolor chemistry (e.g., four-color chemistry) followed by deconvolution of the sequencing / color frequency signals to determine the origin of a particular sequencer base call or signal.
[0081] Once the sequencing reads from the first and second strands are generated, the first strand sequencing reads can be compared with the second strand sequencing reads to provide duplex error correction. The embodiments described herein overcome some of the challenges associated with conversion efficiency described above in that the sequencing information from each clonal cluster provides both a first strand sequencing read and a second strand sequencing read.
[0082] II. Method and Reagent Embodiments for Cleavage of Hairpin Adapters Traditionally, sequencing reactions of hairpin-linked adapter-nucleic acid complexes can be challenging because the polymerase must displace the self-complementary hybridized region. For example, polymerase-based sequencing of such structures remains a barrier to providing duplex sequencing data of physically linked strands due to the high melting temperatures (Tm) of the complementary portions of the first and second strands resulting from the close proximity of the self-complementary portions of the adapter-nucleic acid complex.
[0083] As mentioned above, embodiments of the present technology incorporate the use of hairpin adapters that have a cleavable site or motif so that the first and second strand nucleic acid sequences can be separated from each other during a sequencing reaction.
[0084] In certain embodiments, as shown in FIG. 6, a hairpin adapter (e.g., in the single-stranded or double-stranded portion) can include a cleavage motif that allows for subsequent cleavage of the hairpin DNA molecule by an enzyme (e.g., an endonuclease) or other cleavage-promoting agent (chemical or non-enzymatic process). Referring to FIG. 7, in one embodiment, a single strand of the hairpin adapter (e.g., the linker region) can be cleaved using an endonuclease (e.g., a restriction site endonuclease, a targeting endonuclease, etc.). For example, FIG. 7 shows a single-stranded cleavage site (e.g., a nucleic acid sequence) digestible by an endonuclease (e.g., a restriction enzyme). Referring to FIGS. 3A-5E and 7, after bridge amplification of the double-stranded complex, an enzyme can be introduced (e.g., flowed through a flow cell) to cleave at the cleavage site. In some embodiments, inefficient cleavage is desired (e.g., it is desirable for some remaining uncleaved double-stranded complexes to serve as the starting point for a second round of bridge amplification). In some embodiments, the enzymatic reaction can be controlled for time or concentration so that a portion of the double-stranded complex is cleaved and a portion remains uncleaved. For example, a limited amount of restriction enzyme can be flowed across a functionalized surface to cleave most, but not all, of the hairpin DNA molecules. In another embodiment, the restriction enzyme can be flowed across the surface for a limited time to cleave most, but not all, of the hairpin DNA molecules. In another embodiment, a mixture of enzymes, most of which are catalytically active and a small amount of which are catalytically inactive, can be flowed across a functionalized surface to cleave most, but not all, of the hairpin DNA molecules.
[0085] Figures 8A and 8B show another embodiment for providing a cleavable site in the linker domain of a hairpin adaptor in a manner that allows for inefficient cleavage of double-stranded complexes in clonal clusters. In this example, prior to introduction of an endonuclease, the method can provide for the introduction of an oligonucleotide at least partially complementary to the linker domain of the hairpin adaptor. As shown in Figure 8B, hybridization of the introduced oligonucleotide can prevent cleavage by the endonuclease (e.g., provide an anti-cleavage motif "AC"). Double-stranded complexes without a hybridized oligo (Figure 8A) remain susceptible to cleavage by the endonuclease. The concentration of oligonucleotides provided to the sequencing flow cell prior to enzymatic cleavage (or simultaneously with endonuclease introduction) can be scalable to leave a desired number of uncleaved complexes within each clonal cluster on the flow cell. For example, a small amount of an oligonucleotide sequence containing an anti-cleavage motif can be flowed across a functionalized surface, allowing the oligonucleotide sequence to hybridize to a (e.g., limited) subset of hairpin DNA molecules in each clonal cluster (Figure 8B). The majority of hairpin DNA molecules (containing a cleavage motif within the hairpin) are not hybridized to the oligonucleotide sequence containing the anti-cleavage motif. Therefore, the majority of hairpin DNA molecules (not hybridized to the oligonucleotide sequence containing the anti-cleavage motif) can be cleaved at the single-strand cleavage motif within the hairpin adapter. Hairpin DNA molecules hybridized to the oligonucleotide sequence containing the anti-cleavage motif remain uncleaved by the enzyme.
[0086] In one embodiment, the cleavage motif in the hairpin adapter can be methylated, and the anti-cleavage motif in the oligonucleotide sequence can be unmethylated. An enzyme that cleaves only methylated DNA can then be flowed across the functionalized surface. In another embodiment, the cleavage motif in the hairpin adapter can be unmethylated, and the anti-cleavage motif in the oligonucleotide sequence can be methylated. An enzyme that cleaves only unmethylated DNA can then be flowed across the functionalized surface. In another embodiment, the anti-cleavage motif in the oligonucleotide sequence can be a side chain that prevents cleavage of the hairpin DNA molecule. In another embodiment, the anti-cleavage motif in the oligonucleotide sequence can be a bulky adduct that prevents cleavage of the hairpin DNA molecule. In another embodiment, the anti-cleavage motif in the oligonucleotide sequence can be one or more mismatches that prevent the enzyme from cleaving the hairpin DNA molecule. In another embodiment, the anti-cleavage motif can be an abasic site that prevents cleavage. In another embodiment, the anti-cleavage motif can be a nucleotide analog that prevents cleavage. In another embodiment, the anti-cleavage motif can be a peptide-nucleic acid bond that prevents cleavage.
[0087] In another embodiment, shown in Figures 9A-9B, an oligonucleotide containing a sequence at least partially complementary to the linker domain of a hairpin adapter can be provided and hybridized to the linker domain to form a cleavage site / motif. For example, an endonuclease that recognizes a double-stranded cleavage site can be used to cleave the linker region containing the double-stranded region provided by the hybridized oligonucleotide (Figure 9A). For example, the oligonucleotide can be flowed across a functionalized surface to allow hybridization of the oligonucleotide sequence to the linker region of the hairpin adapter, thereby providing a double-stranded cleavage motif in a portion of the hairpin DNA molecule (Figure 9A). In one embodiment, a limited amount of oligonucleotide can be flowed across the functionalized surface such that hybridization between the oligonucleotide sequence and the hairpin DNA molecule occurs in some, but not all, of the hairpin DNA molecule. In another embodiment, the oligonucleotide can be flowed across the functionalized surface for a limited time such that hybridization between the oligonucleotide sequence and the hairpin DNA molecule occurs in some, but not all, of the hairpin DNA molecule. Hairpin DNA molecules that hybridize to an oligonucleotide sequence, thereby providing a cleavage motif, are cleaved after flowing an endonuclease over the functionalized surface. Hairpin DNA molecules that do not hybridize to an oligonucleotide sequence containing the cleavage motif remain uncleaved.
[0088] In yet another embodiment, as shown in Figures 10A-10B, a pool of oligonucleotides comprising sequences at least partially complementary to the linker domain of a hairpin adapter can be provided and hybridized to the linker domain. The pool of oligonucleotides can include a subset of oligonucleotides that, upon hybridization, provide a cleavage site / motif (e.g., for a suitable endonuclease) (Figure 10A). The pool of oligonucleotides can also include a subset of oligonucleotides that, upon hybridization, provide an anti-cleavage motif (and / or prevent cleavage, e.g., by disrupting site recognition by the endonuclease) (Figure 10B). In one example, the pool of oligonucleotides can be flowed across a functionalized surface. Hairpin DNA molecules hybridized to oligonucleotide sequences containing the cleavage motif are cleaved, while hairpin DNA molecules hybridized to oligonucleotide sequences containing the anti-cleavage motif remain uncleaved. In one embodiment, one subset of oligonucleotides can be methylated, and a second subset of oligonucleotides can be unmethylated. In one embodiment, an enzyme that cleaves only methylated DNA can then be flowed across the functionalized surface. In another embodiment, an enzyme that cleaves only unmethylated DNA can be flowed across the functionalized surface. In another embodiment, the oligonucleotide providing the anti-cleavage motif can include a side chain that prevents cleavage of the hairpin DNA molecule. In another embodiment, the anti-cleavage motif within the oligonucleotide sequence can be a bulky adduct that prevents cleavage of the hairpin DNA molecule. In another embodiment, the anti-cleavage motif within the oligonucleotide sequence can be one or more mismatches that prevent the enzyme from cleaving the hairpin DNA molecule. In another embodiment, the anti-cleavage motif can be an abasic site that prevents cleavage. In another embodiment, the anti-cleavage motif can be a nucleotide analog that prevents cleavage. In another embodiment, the anti-cleavage motif can be a peptide-nucleic acid bond that prevents cleavage.Those skilled in the art will recognize other biochemical means for providing a subset of oligonucleotides that prevent or promote cleavage by selected endonucleases or other enzymes.
[0089] In yet a further embodiment, as shown in Figures 11A and 11B, inefficient cleavage of some of the clonal copies of a double-stranded nucleic acid complex can be achieved by using a mixed pool of endonucleases having catalytically active enzyme portions (striped, Figure 11A) and catalytically inactive enzyme portions (dotted black, Figure 11B).
[0090] In some embodiments, the endonuclease is or comprises a targeting endonuclease. In some embodiments, the targeting endonuclease is or comprises at least one of the following restriction endonucleases (i.e., restriction enzymes) that cleave DNA at or near a recognition site (e.g., EcoRI, BamHI, XbaI, HindIII, AluI, AvaiI, BsaJI, BstNI, DsaV, Fnu4HI, HaeIII, MaeIII, NlaIV, NSil, MspJI, FspEI, NaeI, Bsu36I, NotI, HinFI, Sau3AI, PvuII, SmaI, HgaI, AluI, EcoRV, etc.). A list of some restriction endonucleases is available in both printed and computer-readable form and is provided by many commercial suppliers (e.g., New England Biolabs, Ipswich, MA). Those skilled in the art will understand that any restriction endonuclease can be used in accordance with various embodiments of the present technology. In other embodiments, the targeting endonuclease is or includes at least one of a ribonucleoprotein complex, such as a CRISPR-associated (Cas) enzyme / guide RNA complex (e.g., Cas9 or Cpf1) or a Cas9-like enzyme. In other embodiments, the targeting endonuclease is or includes a homing endonuclease, a zinc finger nuclease, a TALEN and / or a meganuclease (e.g., a megaTAL nuclease), an Argonaute nuclease, or a combination thereof. In some embodiments, the targeting endonuclease includes Cas9 or CPF1, or a derivative thereof. In another embodiment, the nuclease can cut at a forked nucleic acid region (e.g., FEN1). In some embodiments, more than one (eg, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) targeting endonucleases may be used.
[0091] In some embodiments, the cleavage site is or comprises a user-directed recognition sequence for a targeting endonuclease (e.g., a CRISPR or CRISPR-like endonuclease) or other tunable endonuclease. In some embodiments, cleaving the nucleic acid material may include at least one of enzymatic digestion, enzymatic cleavage, enzymatic cleavage of one strand, enzymatic cleavage of both strands, incorporation of a modified nucleic acid followed by an enzymatic treatment causing cleavage of one or both strands, incorporation of a replication-blocking nucleotide, incorporation of a chain terminator, incorporation of a photocleavable linker, incorporation of uracil, incorporation of a ribose base, incorporation of an 8-oxo-guanine adduct, use of a restriction endonuclease, use of a ribonucleoprotein endonuclease (e.g., a Cas enzyme, e.g., Cas9 or CPF1), or other programmable endonucleases (e.g., homing endonucleases, zinc finger nucleases, TALENs, meganucleases (e.g., megaTAL nucleases), Argonaute nucleases, etc.), and any combination thereof.
[0092] Targeting endonucleases (e.g., CRISPR-associated ribonucleoprotein complexes, e.g., Cas9 or Cpfl, homing nucleases, zinc finger nucleases, TALENs, megaTAL nucleases, Argonaute nucleases, and / or derivatives thereof) can be used to selectively cleave targeted portions of nucleic acid materials. In some embodiments, the targeting endonuclease may be modified (e.g., have amino acid substitutions) to, for example, improve thermostability, salt tolerance, and / or pH tolerance, or to provide improved specificity or alternative PAM site recognition, or higher binding affinity. In other embodiments, the targeting endonuclease may be biotinylated, fused to streptavidin, and / or incorporate other affinity-based (e.g., bait / prey) technologies. In certain embodiments, the targeting endonuclease may have altered recognition site specificity (e.g., SpCas9 variants with altered PAM site specificity). In other embodiments, the targeting endonuclease may be catalytically inactive so that cleavage does not occur upon binding to the targeting portion of the nucleic acid material. In some embodiments, the targeting endonuclease is modified (e.g., a nickase variant) to cleave a single strand of the targeting portion of the nucleic acid material, thereby generating a nick in the nucleic acid material. CRISPR-based targeting endonucleases are further discussed herein, providing further detailed, non-limiting examples of the use of targeting endonucleases. Note that the nomenclature surrounding such targeting endonucleases is still in flux. For purposes herein, the term "CRISPR-based" is used to generally refer to endonucleases that contain a nucleic acid sequence that can be modified to redefine the nucleic acid sequence to be cleaved. Cas9 and CPF1 are examples of such targeting endonucleases currently in use, but many more are likely to exist in different locations in nature, and the availability of a variety of different such targeted, easily tunable nucleases is expected to grow rapidly in the coming years. For example, Casl2a, Casl3, CasX, etc. are contemplated for use in various embodiments.Similarly, multiple engineered variants of these enzymes are becoming available to improve or modify their properties. The use of substantially functionally similar targeting endonucleases not expressly described herein or yet to be discovered to achieve similar purposes as those disclosed herein is expressly contemplated herein.
[0093] It is specifically contemplated that any of a variety of restriction endonucleases (i.e., enzymes) may be used. In general, restriction enzymes are typically produced by particular bacteria / other prokaryotes and cut at, near, or between specific sequences in a given segment of DNA.
[0094] It will be apparent to one of skill in the art that a restriction enzyme may be selected to cleave at a specific site, or alternatively, at a site engineered to create a restriction site for cleavage. In some embodiments, the restriction enzyme is a synthetic enzyme. In some embodiments, the restriction enzyme is not a synthetic enzyme. In some embodiments, the restriction enzymes used herein are modified to introduce one or more changes into their own genome. In some embodiments, the restriction enzyme generates a double-stranded break between defined sequences within a given portion of DNA.
[0095] While any restriction enzyme (e.g., Type I, Type II, Type III and / or Type IV) may be used according to some embodiments, the following represents a non-limiting list of restriction enzymes that may be used: AluI, ApoI, AspHI, BamHI, BfaI, BsaI, CfrI, DdeI, DpnI, DraI, EcoRI, EcoRII, EcoRV, HaeII, HaeIII, HgaI, HindII, HindIII, HinFI, HPYCH4III, KpnI, MamI, MNLI, MseI, MstI, M stII, NcoI, NdeI, NotI, Pad, PstI, PvuI, PvuII, RcaI, RsaI, SacI, SacII, SalI, Sau3AI, SeaI, SmaI, SpeI, SphI, StuI, TaqI, XbaI, XhoI, XhoII, XmaI, XmaII, and any combination thereof. Extensive, but non-exhaustive, lists of suitable restriction enzymes can be found in publicly available catalogs or on the Internet (e.g., available from New England Biolabs, Ipswich, MA, USA). Those skilled in the art will appreciate that various enzymes, ribozymes, or other nucleic acid-modifying enzymes that can be used alone or in combination to target phosphodiester backbone cleavage of nucleic acid molecules and achieve the same purpose may not be included in the above list or may have yet to be discovered. Various nucleic acid-modifying enzymes can recognize base modifications (e.g., CpG methylation) that can be used to target further modifications (e.g., to generate abasic sites) in adjacent nucleic acid sequences that can be cleaved (e.g., by an enzyme with lyase activity). Thus, substantial sequence specificity of cleavage can be achieved based on the recognition of DNA or RNA modifications, which can be used alone or in combination with a targeting endonuclease to achieve fragmentation of the targeted nucleic acid. Other embodiments of cleavage promoters can include non-enzyme promoters. For example, pH changes or hydrolysis can be used to cleave at the cleavage site. Photocleavage is also an approach to disrupting this backbone.For example, incorporation of modified nucleotides in a hairpin adapter sequence or hybridization of a complementary or partially complementary oligonucleotide with a light-sensitive moiety can create a recognition site for other chemical or enzymatic processes that cleave the opposite strand (e.g., upon exposure to light).
[0096] In some embodiments (e.g., those described above), cleavage site C is provided when the physically linked adaptor-molecule complex is in a self-hybridized configuration on a surface (e.g., Figures 6, 7, 8A, 9A, 10A, and 11A). In yet further embodiments, as shown in Figures 12A-12C, cleavage site C is available for cleavage by a cleavage promoter when the physically linked nucleic acid complex is in a double-stranded bridge-amplified configuration. For example, cleavage site C is a double-stranded motif provided by the double-stranded configuration after duplex formation across the "bridge" on the surface but before denaturation (Figure 12A). Upon cleavage, the first-strand sequence amplicon is separated from the second-strand amplicon while still bound to the surface (Figure 12B). After denaturation and removal of unbound amplicons (Figure 12C), both the first-strand and second-strand single-stranded amplicons remain bound and available for sequencing. In one embodiment, sequencing of the first and second strand amplicons can proceed using sequencing reactions such as those described with respect to Figures 5D and 5E.
[0097] adapter As discussed above, adapter molecules may be or include a "Y"-shaped, "U"-shaped, or "hairpin"-shaped, and may have a bubble (e.g., a portion of sequence that is non-complementary), or other feature. A "U"-shaped or "hairpin"-shaped adapter may refer to an adapter having a linker domain that links or connects a first strand of a target double-stranded nucleic acid molecule to a second strand of the same molecule. Certain hairpin adapters may be, for example, cleavable hairpin adapters and / or may include modified or non-standard nucleotides, restriction sites, or other features for engineering structure or function in vitro.
[0098] Adapter molecules may be ligated to various nucleic acid materials with ends. For example, the adapter molecule may have a T-overhang, an A-overhang, a CG-overhang, a multi-nucleotide overhang (also referred to herein as a "sticky end" or "sticky overhang"), or a single-stranded overhang region of a known nucleotide length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides), a dehydroxylated base, or a blunt end suitable for ligation to a nucleic acid material, where the end of the molecule is dephosphorylated at the 5' end of the target or otherwise blocked from conventional ligation. In other embodiments, the adapter molecule may contain a modification in the 5' strand of the ligation site that prevents dephosphorylation or other ligation. In the latter two embodiments, such a strategy may be useful for preventing dimerization of the library fragments or adapter molecules.
[0099] The ligation domain of the adapter can be cleaved with an endonuclease (e.g., a restriction endonuclease, a targeting endonuclease, etc.) enzyme, leaving a 3' "T" overhang compatible with ligation to a 3' "A" overhang in the prepared library fragment. In certain embodiments, the resulting ligation domain is a single base-pair thymine (T) overhang on the 3' end of the extended extension strand, but in other embodiments, it can be a blunt end, or a different type or "sticky" end of the 3' or 5' overhang. In this particular example, "cut" implies cleavage using a sequence-specific endonuclease, e.g., a restriction enzyme, in a manner that inherently creates ligatable ends. In other embodiments, ligatable ends can be created after cleavage by further enzymatic or chemical treatment, such as terminal transferase.
[0100] Referring back to FIG. 2A, the ligatable ends are shown as T-overhangs, however, it will be apparent to one of skill in the art that the ligatable ends may be in any of a variety of forms, such as "sticky" ends comprising, among others, a blunt end, an A-3' overhang, a 1 nucleotide 3' overhang, a 2 nucleotide 3' overhang, a 3 nucleotide 3' overhang, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotide 3' overhang, a 1 nucleotide 5' overhang, a 2 nucleotide 5' overhang, a 3 nucleotide 5' overhang, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotide 5' overhang (FIG. 2B). The 5' base of the ligation site may be phosphorylated and the 3' base may have a hydroxyl group, or either may be dephosphorylated or dehydrated, alone or in combination, or may be further chemically modified to facilitate ligation or improvement of one strand, optionally preventing ligation of the other strand until a later point in time.
[0101] In some embodiments, the adapter molecule may comprise a capture moiety suitable for isolating a desired target nucleic acid molecule ligated to the capture moiety.
[0102] An adapter sequence can refer to a single-stranded sequence, a double-stranded sequence, a complementary sequence, a non-complementary sequence, a partially complementary sequence, an asymmetric sequence, a sequence that binds to a primer, a flow cell sequence, a ligation sequence, or other sequence provided by an adapter molecule. In certain embodiments, an adapter sequence can refer to a sequence used for amplification by a complement to an oligonucleotide.
[0103] In some embodiments, the provided methods and compositions include at least one adapter sequence (e.g., two adapter sequences, one at each of the 5' and 3' ends of the nucleic acid material). In some embodiments, the provided methods and compositions may include two or more adapter sequences (e.g., 3, 4, 5, 6, 7, 8, 9, 10 or more). In some embodiments, at least two of the adapter sequences differ from each other (e.g., by sequence). In some embodiments, each adapter sequence differs from each other (e.g., by sequence). In some embodiments, at least one adapter sequence is at least partially non-complementary (e.g., non-complementary by at least one nucleotide) to at least a portion of at least one other adapter sequence.
[0104] In some embodiments, the adapter sequence comprises at least one non-standard nucleotide, in some embodiments, the non-standard nucleotide is an abasic site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5'nitroindole, 5-hydroxymethyl-2'-deoxycytidine, iso-cytosine, 5'-methyl-isocytosine, or isoguanosine, a methylated nucleotide, an RNA nucleotide, a ribose nucleotide, an 8-oxo-guanine, a photocleavable linker, a biotinylated nucleotide, a desthiobiotin nucleotide, a thiol-modified nucleotide, or a nucleotide sequence. The spacer is selected from a nucleotide selected from the group consisting of nucleotides selected from ...
[0105] In some embodiments, the adapter sequence includes a portion having magnetic properties (i.e., a magnetic portion). In some embodiments, the magnetic property is paramagnetic. In some embodiments, where the adapter sequence includes a magnetic portion (e.g., a nucleic acid material ligated to an adapter sequence including a magnetic portion), when a magnetic field is applied, the adapter sequence including the magnetic portion is substantially separated from the adapter sequence that does not include the magnetic portion (e.g., a nucleic acid material ligated to an adapter sequence that does not include a magnetic portion).
[0106] In some embodiments, at least one adapter sequence is located 5' to the SMI. In some embodiments, at least one adapter sequence is located 3' to the SMI.
[0107] In some embodiments, the adapter sequence may include one or more linker domains. In some embodiments, the linker domain may be composed of nucleotides. In some embodiments, the linker domain may include at least one modified nucleotide or non-nucleotide molecule (e.g., as described elsewhere in this disclosure). In some embodiments, the linker domain may be or include a loop.
[0108] In some embodiments, the adapter sequences on one or both ends of each strand of a double-stranded nucleic acid material may further comprise one or more elements that provide an SDE, which in some embodiments may be or may include an asymmetric primer site contained within the adapter sequence.
[0109] In some embodiments, the adapter sequence may be or may include at least one SDE and at least one ligation domain (i.e., at least one domain amendable to ligase activity, e.g., a domain suitable for ligating to a nucleic acid material through ligase activity). In some embodiments, from 5' to 3', the adapter sequence may be or may include a primer binding site, an SDE, and a ligation domain.
[0110] Various methods for synthesizing duplex sequencing adapters have been previously described, for example, in U.S. Pat. No. 9,752,188, International Patent Publication No. 2017 / 100441, and International Patent Publication No. PCT / US18 / 59908 (filed November 8, 2018), all of which are incorporated herein by reference in their entireties.
[0111] Various methods for synthesizing duplex sequencing adapters have been described (e.g., U.S. Pat. No. 9,752,188 and U.S. Pat. No. PCT / US19 / 17908, which are incorporated herein by reference). For example, in one embodiment, one oligonucleotide can be hybridized to another oligonucleotide containing a degenerate or semi-degenerate nucleotide sequence at a non-complementary region. The hybridized oligonucleotides can then be chemically linked, or can be two portions of a contiguous oligonucleotide that form a "loop" or "U" shape (hairpin adapter) when hybridized. An enzyme capable of polymerizing nucleotides can then be used to copy the single-stranded degenerate or semi-degenerate region to synthesize its complement. This generates a complementary double-stranded degenerate or semi-degenerate sequence, which can function as at least one SMI element during duplex sequencing. The ligation site on the adapter molecule may be modified from this extension product by enzymatic or chemical manipulation (e.g., by restriction digestion, terminal transferase activity of a polymerase, or other enzymes or any other method known in the art).
[0112] Primer In some embodiments, one or more PCR primers having at least one of the following characteristics are contemplated for use in various embodiments according to the present technology: (1) high target specificity, (2) multiplexing capability, and (3) robust amplification with minimal bias. Many conventional tests and commercial products have designed primer mixtures that meet these specific criteria for conventional PCR-CE. However, it should be noted that these primer mixtures are not always suitable for use with MPS. In fact, developing a highly multiplexed primer mixture can be a challenging and time-consuming process. Fortunately, both Illumina and Promega have recently developed multiplex-compatible primer mixtures for the Illumina platform that demonstrate robust and efficient amplification of various standard and non-standard STR and SNP loci. Because these kits use PCR to amplify their target regions before sequencing, the 5' end of each read in paired-end sequencing data corresponds to the 5' end of the PCR primer used to amplify the DNA. In some embodiments, the provided methods and compositions include primers designed to ensure uniform amplification, which may involve varying reaction concentrations, melting temperatures, and minimizing secondary structure and intra- or inter-primer interactions. Many techniques have been described for optimizing highly multiplexed primers for MPS applications. In particular, these techniques are often known as ampliseq methods, as are well described in the art.
[0113] amplification The provided methods and compositions, in various embodiments, utilize or employ at least one amplification step to amplify nucleic acid material (or a portion thereof, e.g., a specific target region or locus) to form an amplified nucleic acid material (e.g., some member of an amplicon product).
[0114] In some embodiments, amplifying the nucleic acid material includes amplifying the nucleic acid material derived from each of the first and second nucleic acid strands from the original double-stranded nucleic acid material using at least one single-stranded oligonucleotide that is at least partially complementary to a sequence present in the first adapter sequence. The amplifying step further includes amplifying each strand of interest using a second single-stranded oligonucleotide, which may (a) be at least partially complementary to the target sequence of interest or (b) be at least partially complementary to a sequence present in the second adapter sequence such that the at least one single-stranded oligonucleotide and the second single-stranded oligonucleotide are oriented in a manner that effectively amplifies the nucleic acid material.
[0115] In some embodiments, amplifying nucleic acid material in a sample may include amplifying nucleic acid material in a "tube" (e.g., a PCR tube), in an emulsion droplet, a microchamber, and other examples described above or other known containers. In some embodiments, amplifying nucleic acid material may include amplifying nucleic acid material in two or more (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 or more samples) physically separated samples (e.g., tubes, droplets, chambers, containers, etc.).
[0116] While any amplification reaction suitable for any application is contemplated to be compatible with some embodiments, by way of example, in some embodiments the amplification step may be or may include polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), isothermal amplification, polony amplification in an emulsion, bridge amplification on a surface, on a bead, or in a hydrogel, and any combination thereof.
[0117] In some embodiments, amplification on a surface, such as bridge amplification on the surface of a flow cell, involves generating clusters or multiple copies of linked nucleic acid templates. In certain embodiments, linked first and second strand nucleic acid templates can be bridge-amplified on the surface of a flow cell to generate, for example, multiple clonal clusters, each of which contains nucleic acid template copies derived from both the original first and second strands of the original double-stranded nucleic acid molecule. Some of the clonal copies in a cluster are forward-directed, and the rest are reverse-directed. First, either the forward copy or the reverse copy is cleaved and removed, and then the sequencing reaction can proceed.
[0118] In some embodiments, amplifying the nucleic acid material involves the use of single-stranded oligonucleotides that are at least partially complementary to regions of an adapter sequence on the 5' and 3' ends of each strand of the nucleic acid material. In some embodiments, amplifying the nucleic acid material involves the use of at least one single-stranded oligonucleotide that is at least partially complementary to a target region or sequence of interest (e.g., a genomic sequence, a mitochondrial sequence, a plasmid sequence, a synthetically produced target nucleic acid, etc.) and a single-stranded oligonucleotide that is at least partially complementary to a region of the adapter sequence (e.g., a primer site).
[0119] In general, stable amplification (e.g., PCR amplification) can be highly dependent on reaction conditions. Multiplex PCR can be sensitive to, for example, buffer composition, monovalent or divalent cation concentration, detergent concentration, crowding agent (e.g., PEG, glycerol, etc.) concentration, primer concentration, primer Tm, primer design, primer GC content, primer modified nucleotide characteristics, and cycling conditions (e.g., temperature and extension time, and temperature change rate). Optimizing buffer conditions can be a difficult and time-consuming process. In some embodiments, amplification reactions may use at least one of buffers, primer pool concentrations, and PCR conditions according to already known amplification protocols. In some embodiments, new amplification protocols may be created and / or amplification reaction optimization may be used. As a specific example, in some embodiments, PCR optimization kits, such as Promega® PCR optimization kits, may be used, which include several pre-formulated buffers partially optimized for various PCR applications (e.g., multiplex, real-time, GC-rich, and inhibitor-resistant amplification). These preformulated buffers can be quickly replenished with various Mg2+ and primer concentrations, as well as primer pool ratios. Additionally, in some embodiments, various cycling conditions (e.g., thermal cycling) may be evaluated and / or used. When assessing whether a particular embodiment is suitable for a particular desired application, one or more of the following may be evaluated: specificity, allele coverage ratio for heterozygous loci, interlocus balance, and depth, among other aspects. Measurement of amplification success may include DNA sequencing of the products, evaluation of products by gel or capillary electrophoresis or HPLC or other size separation methods followed by visualization of fragments, melting curve analysis using double-stranded nucleic acid binding dyes or fluorescent probes, mass spectrometry, or other methods known in the art.
[0120] In some embodiments, at least one amplifying step comprises at least one primer that is or includes at least one non-standard nucleotide, hi some embodiments, the non-standard nucleotide is selected from uracil, methylated nucleotides, RNA nucleotides, ribose nucleotides, 8-oxo-guanine, biotinylated nucleotides, locked nucleic acids, peptide nucleic acids, high Tm nucleic acid variants, allele-distinguishing nucleic acid variants, any other nucleotide or linker variant described elsewhere herein, or any combination thereof.
[0121] Nucleic acid material kinds Any of a variety of nucleic acid materials may be used according to various embodiments. In some embodiments, the nucleic acid material may include at least one modification to a polynucleotide within the canonical sugar-phosphate backbone. In some embodiments, the nucleic acid material may include at least one modification within any base in the nucleic acid material. For example, and not by way of limitation, in some embodiments, the nucleic acid material is or includes at least one of double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, peptide nucleic acid (PNA), and locked nucleic acid (LNA).
[0122] source It is contemplated that the nucleic acid material may be derived from any of a variety of sources. For example, in some embodiments, the nucleic acid material is provided from a sample from at least one subject (e.g., a human or animal subject) or other biological source. In some embodiments, the nucleic acid material is provided from a preserved / stored sample. In some embodiments, the sample is blood, serum, sweat, saliva, cerebrospinal fluid, mucus, uterine washing, vaginal swab, nasal swab, oral swab, tissue scraping, hair, fingerprint, urine, stool, vitreous fluid, peritoneal washing, sputum, bronchial lavage, oral wash, pleural lavage, gastric lavage, gastric juice, bile, pancreatic duct lavage, bile duct lavage, common bile duct lavage, gallbladder fluid, synovial fluid, infected wound, non-infected wound, archaeological sample, forensic sample, water sample, tissue sample, food sample, bioreactor sample, etc. In other embodiments, the sample is or comprises at least one of: tissue, plant sample, nail polish, semen, prostatic fluid, fallopian tube washing, cell-free nucleic acid, intracellular nucleic acid, metagenomics sample, implanted foreign body wash, nasal wash, intestinal fluid, epithelial brushing, epithelial wash, tissue biopsy, autopsy sample, necrosis sample, organ sample, human identification ampoule, artificially created nucleic acid sample, synthetic genetic sample, nucleic acid data repository sample, tumor tissue, and any combination thereof. In other embodiments, the sample is or comprises at least one of: microorganism, plant-based organism, or any collected environmental sample (e.g., water, soil, archaeological sample, etc.).
[0123] qualification According to various embodiments, the nucleic acid material may undergo one or more modifications prior to, substantially simultaneously with, or after any particular step, depending on the application for which the particular provided method or composition is to be used.
[0124] In some embodiments, the modification may be or may include repair of at least a portion of the nucleic acid material. Although it is contemplated that methods suitable for any application of nucleic acid repair are compatible with some embodiments, certain exemplary methods and compositions are described below and in the Examples.
[0125] As a non-limiting example, in some embodiments, DNA repair enzymes such as uracil-DNA glycosylase (UDG), formamidopyrimidine DNA glycosylase (FPG), and 8-oxoguanine DNA glycosylase (OGGI) can be used to repair DNA damage (e.g., in vitro DNA damage). In some embodiments, these DNA repair enzymes are glycosylases that remove damaged bases from DNA. For example, UDG removes uracil resulting from cytosine deamination (resulting from spontaneous hydrolysis of cytosine), and FPG removes 8-oxo-guanine (e.g., a common DNA lesion resulting from reactive oxygen species). FPG also has lyase activity and can generate one-base gaps at abasic sites. Such abasic sites cannot be subsequently amplified by PCR, for example, because the polymerase cannot copy the template. Therefore, the use of such DNA damage repair enzymes can effectively remove damaged DNA that does not harbor true mutations but may not otherwise be detected as errors after sequencing and duplex sequence analysis.
[0126] In further embodiments, sequencing reads generated from the processing steps described herein can be further filtered to eliminate false mutations by trimming the ends of reads that are most prone to artifacts. For example, DNA fragmentation can generate single-stranded portions at the ends of double-stranded molecules. These single-stranded portions may be blunted during end repair (e.g., by Klenow). In some cases, polymerases may make copying errors in these end-repaired regions, resulting in the creation of "false double-stranded molecules." These artifacts may appear to be true mutations when sequenced. Errors resulting from these end-repair mechanisms can be excluded from post-sequencing analysis by trimming the ends of sequencing reads to exclude possible mutations, thereby reducing the number of false mutations. In some embodiments, such trimming of sequencing reads can be achieved automatically (e.g., as a routine processing step). In some embodiments, mutation frequencies can be assessed for fragment end regions, and if a threshold level of mutations is observed in the fragment end regions, the sequencing reads can be trimmed before generating double-stranded consensus sequence reads for the DNA fragments.
[0127] Some embodiments of duplex sequencing methods provide a PCR-based targeted enrichment strategy that is compatible with the use of cleavable hairpin adapters for error correction. For example, sequencing enrichment strategies utilizing steps of the Separated PCRs of Linked Templates for Sequencing ("SPLiT-DS") method can also benefit from pre-enriched nucleic acid material using one or more of the embodiments described herein. SPLiT-DS was originally described in International Patent Publication No. 2018 / 175997, the entire contents of which are incorporated herein by reference. The SPLiT-DS approach may begin with labeling (e.g., tagging) fragmented double-stranded nucleic acid material (e.g., from a DNA sample) with molecular barcodes in a manner similar to that described above for standard duplex sequencing library construction protocols. In some embodiments, the double-stranded nucleic acid material may be fragmented (e.g., cell-free DNA, damaged DNA, etc.). However, in other embodiments, various steps may include fragmenting the nucleic acid material using mechanical shearing, such as sonication, or other DNA cleavage methods (e.g., those described further herein). Labeling of the fragmented double-stranded nucleic acid material may include end-repair and 3'-dA tailing, followed by ligation of the double-stranded nucleic acid fragments to duplex sequencing adapters (e.g., cleavable hairpin adapters, Y-shaped adapters, etc.) if required for a particular application. In other embodiments, endogenous SMI sequences, or a combination of exogenous and endogenous SMI sequences, for uniquely associated information from both strands of the original nucleic acid molecule may also be used in conjunction with physical linkage of the first and second strands. After ligation of the adapter molecule to the double-stranded nucleic acid material, the method may be followed by amplification (e.g., PCR amplification, rolling circle amplification, multiple displacement amplification, isothermal amplification, bridge amplification, surface-bound amplification, etc.).
[0128] Reagent-based kits Aspects of the present technology further encompass kits for carrying out various aspects of duplex sequencing methods (also referred to herein as "DS kits"). In some embodiments, the kits may include various reagents along with instructions for carrying out one or more of the methods or method steps disclosed herein for nucleic acid extraction, nucleic acid library preparation, amplification (e.g., PCR, bridge amplification), cleavage of ligated nucleic acid complexes, and sequencing. In one embodiment, the kits further include a computer program product (e.g., coded algorithms for execution on a computer, access code to a cloud-based server for executing one or more algorithms) for analyzing sequencing data (e.g., raw sequencing data, sequencing reads, etc.) to determine, for example, variant alleles, mutations, etc. associated with a sample according to aspects of the present technology. The kits may include other forms of DNA standards, as well as positive and negative controls.
[0129] In some embodiments, the DS kit may include reagents or combinations of reagents (e.g., enzymes, dNTPs, wash buffers, etc.) suitable for performing various aspects of sample preparation (e.g., tissue manipulation, DNA extraction, DNA fragmentation), nucleic acid library preparation, amplification, cleavage, and on-sequencer surface treatment steps, and sequencing. For example, the DS kit may optionally include one or more DNA extraction reagents (e.g., buffers, columns, etc.) and / or tissue extraction reagents. Optionally, the DS kit may further include one or more reagents or tools for fragmenting double-stranded DNA, for example, by physical means (e.g., tubing to facilitate acoustic shearing or sonication, nebulizer units, etc.) or enzymatic means (e.g., enzymes for random or semi-random genomic shearing and appropriate reaction enzymes). For example, the kit may include DNA fragmentation reagents for enzymatically fragmenting double-stranded DNA, including one or more of enzymes for targeted digestion (e.g., restriction endonucleases, CRISPR / Cas endonucleases and RNA guides, and / or other endonucleases), double-stranded fragmentase cocktails, single-stranded DNase enzymes (e.g., mung bean nuclease, SI nuclease) to fragment DNA predominantly double-stranded and / or to break single-stranded DNA, and appropriate buffers and solutions to facilitate such enzymatic reactions.
[0130] In one embodiment, the DS kit includes primers and adapters for preparing a nucleic acid sequence library from a sample suitable for performing a duplex sequencing process step to generate error-corrected (e.g., highly accurate) sequences of double-stranded nucleic acid molecules in the sample. For example, the kit may include at least one pool of adapter molecules containing a linker domain (e.g., hairpin adapters), at least one pool of adapter molecules containing a double-stranded portion and a single-stranded portion (e.g., "Y"-shaped adapters), or tools for users to create such (e.g., single-stranded oligonucleotides). In some embodiments, the pool of adapter molecules includes a single molecule identifier (SMI) sequence or a suitable number of substantially unique SMI sequences, either alone or in combination with a unique feature of the fragment to be ligated, so that multiple nucleic acid molecules in the sample can be substantially uniquely labeled after the adapter molecules are attached. Those skilled in the art will understand that what constitutes a "suitable" number of SMI sequences in molecular tagging can vary by several orders of magnitude depending on various specific factors (e.g., input DNA, type of DNA fragmentation, average fragment size, complexity vs. repetitiveness of the sequences to be sequenced in the genome, etc.). Optionally, the adapter molecule further comprises one or more PCR primer binding sites, one or more sequencing primer binding sites, or both. In another embodiment, the DS kit does not include an adapter molecule comprising an SMI sequence or barcode, but instead comprises a conventional adapter molecule (e.g., a Y-shaped sequencing adapter, etc.), and various method steps may utilize endogenous SMI and / or physical location on the sequencing surface to associate molecular sequence reads. In some embodiments, the adapter molecule is an index adapter and / or comprises an index sequence. In other embodiments, the index is added to a particular sample through PCR "tailing in" using primers provided in the kit.
[0131] In one embodiment, the DS kit includes a set of adapter molecules, each having a non-complementary region and / or some other strand-defining element (SDE), or tools for the user to generate them (e.g., single-stranded oligonucleotides). In another embodiment, the kit includes at least one set of adapter molecules, wherein at least a subset of the adapter molecules each include at least one SMI and at least one SDE, or tools for generating them. In some embodiments, the subset of adapter molecules may be configured to have ligatable ends (e.g., blunt ends, overhangs, substantially or partially inherent sticky ends, etc.). Further features of primers and adapters for preparing nucleic acid sequencing libraries from samples suitable for performing duplex sequencing process steps are described above and disclosed in U.S. Pat. No. 9,752,188, International Patent Publication No. 2017 / 100441, and International Patent Application No. PCT / US18 / 59908 (filed November 8, 2018), all of which are incorporated herein by reference in their entireties.
[0132] In one embodiment, the DS kit includes reagents for processing steps that occur on the sequencing surface, such as cleavage promoters (e.g., enzymes, non-enzyme solutions, light, hybridization oligonucleotides, etc.) and anti-cleavage promoters (e.g., enzymes, including catalytically inactive enzymes, hybridization oligonucleotides, etc.), and other wash solutions for carrying out the various steps of the method.
[0133] Additionally, the kit may further include DNA quantification materials, such as, for example, SYBR™ green or SYBR™ gold (available from Thermo Fisher Scientific, Waltham, MA) for use with a Qubit™ fluorometer (available from Thermo Fisher Scientific, Waltham, MA), or a DNA-binding dye such as PicoGreen™ dye (available from Thermo Fisher Scientific, Waltham, MA) for use with a suitable fluorescence spectrometer or real-time or digital droplet PCR machine. Other reagents suitable for DNA quantification on other platforms are also contemplated. Further embodiments include kits that include one or more of nucleic acid size selection reagents (e.g., Solid Phase Reversible Immobilization (SPRI) magnetic beads, gels, columns), columns for target DNA capture using bait / prey hybridization, qPCR reagents (e.g., for copy number determination), and / or digital droplet PCR reagents. In some embodiments, the kit may optionally include one or more of library preparation enzymes (ligase, polymerase, endonuclease, e.g., reverse transcriptase for RNA interrogation), dNTPs, buffers, capture reagents (e.g., beads, surfaces, coated tubes, columns, etc.), index primers, amplification primers (PCR primers), and sequencing primers. In some embodiments, the kit may include reagents for assessing the type of DNA damage, such as error-prone DNA polymerases and / or high-fidelity DNA polymerases. Additional additives and reagents are contemplated for PCR or ligation reactions in specific conditions (e.g., GC-rich genomes / targets).
[0134] In one embodiment, the kit further comprises reagents, such as DNA error correcting enzymes, that repair errors in DNA sequences that interfere with the polymerase chain reaction (PCR) process (as opposed to repairing disease-causing mutations). By way of non-limiting example, enzymes include the following, among other glycosylases, lyases, endonucleases and exonucleases: monofunctional uracil-DNA glycosylase (hSMUG1), uracil-DNA glycosylase (UDG), N-glycosylase / AP-lyase NEIL1 protein (hNEIL1), formamidopyrimidine DNA glycosylase (FPG), 8-oxoguanine DNA glycosylase (OGG1), human apurinic / apyrimidinic endonuclease (APE 1), endonuclease III (Endo III), endonuclease IV (Endo IV), endonuclease V (Endo V), endonuclease VIII (Endo VIII), T7 endonuclease I (T7 Endo I), T4 pyrimidine dimer glycosylase (T4 These DNA repair enzymes include one or more of the following: human PDG (human single-strand-selective human alkyldenine DNA glycosylase (hAAG) and the like, which can be used to correct DNA damage (e.g., in vitro or in vivo DNA damage). Some of these DNA repair enzymes are glycosylases that remove damaged bases from DNA. For example, UDG removes uracil resulting from cytosine deamination (resulting from spontaneous hydrolysis of cytosine), and FPG removes 8-oxo-guanine (e.g., a common DNA lesion resulting from reactive oxygen species). FPG also has lyase activity and can generate one-base gaps at abasic sites. Such abasic sites cannot be subsequently amplified by PCR, for example, because polymerases cannot copy the template. Therefore, such DNA damage repair enzymes and / or others listed herein and known in the art can be used to effectively remove damaged DNA that does not contain true mutations but may otherwise go undetected as errors.
[0135] The kit may further include appropriate controls, such as DNA amplification controls, nucleic acid (template) quantification controls, sequencing controls, and nucleic acid molecules derived from similar biological sources (e.g., healthy subjects). In some embodiments, the kit may include a control collection of cells. Thus, the kit can include suitable reagents (e.g., test compounds, nucleic acids, control sequencing libraries, etc.) to provide controls and provide expected duplex sequencing results for samples containing rare genetic variants (e.g., nucleic acid molecules containing disease-associated variants / mutations that can be added or included in sample preparation steps) to determine the reliability of the protocol. In some embodiments, the kit may include reference sequence information. In some embodiments, the kit may include sequence information useful for identifying one or more DNA variants in a collection of cells or a cell-free DNA sample. In one embodiment, the kit includes a container for transporting the sample, storage materials for stabilizing the sample, materials for freezing the sample, a cell sample for analysis to detect DNA variants in a subject sample, etc. In another embodiment, the kit may include a nucleic acid contamination control standard (eg, a hybridization capture probe with affinity for a genomic region in an organism different from the test or target organism).
[0136] The kit may further include one or more other containers containing materials desirable from a commercial and user standpoint, including PCR and sequencing buffers, diluents, subject sample extraction tools (e.g., syringes, swabs, etc.), and a package insert containing instructions for use. In addition, a label may be provided on the container with instructions for use as described above, and / or the instructions and / or other information may be included on a package insert included with the kit and / or via a website address provided therewith. The kit may also include labware, such as, for example, sample tubes, plate sealers, microcentrifuge tube openers, labels, magnetic particle separators, foam inserts, ice packs, dry ice packs, insulation, etc.
[0137] The kit may further include a pre-packaged or application-specific functionalized surface for use in amplifying a sequencing library. In one embodiment, the functionalized surface may include a surface suitable for performing a sequencing reaction therein. The functionalized surface may be pre-configured with bound oligonucleotides suitable for bridge amplification of a sequencing library (e.g., the surface includes a dispersed lawn of bound oligonucleotides complementary to sequence domains in one or more adapter sets). In one embodiment, the functionalized surface is a flow cell configured for use with a sequencing system described below.
[0138] The kit may further include a computer program product installable on an electronic computing device (e.g., a laptop / desktop computer, tablet, etc.) or accessible via a network (e.g., a remote server, cloud computing), where the computing device or remote server comprises one or more processors configured to execute instructions for performing operations including the duplex sequencing analysis step. For example, the processor may be configured to execute instructions for processing raw sequencing reads or unanalyzed sequencing reads to generate duplex sequencing data. In additional embodiments, the computer program product may include a database containing subject or sample records (e.g., information about a particular subject or sample or group of samples) and empirically derived information about targeted regions of DNA. The computer program product is embodied in a non-transitory computer-readable medium that, when executed on a computer, performs the steps of the methods disclosed herein.
[0139] The kit may further include instructions and / or access codes / passwords, etc. for accessing a remote server (including a cloud-based server) for uploading and downloading data (e.g., sequencing data, reports, other data), or software to be installed on a local device. All computational work may be performed on the remote server and accessed by the user / kit user via an internet connection, etc.
[0140] The kit may be suitable for use with a sequencing system optimized for use with the methods and reagents described herein. For example, the sequencing system and associated sequencing reagents may be configured to perform a stepwise sequencing reaction that provides intervening processing steps. In one embodiment, the sequencing system may provide delivery systems for cleavage promoter delivery, anti-cleavage promoter delivery, enzyme solution delivery, oligonucleotide delivery, wash buffers, etc. Similarly, the sequencing system may include appropriate controls (e.g., manual, automatic, semi-automatic, etc.) and internal programming for processing step time, temperature, pH, concentration, etc. [Example]
[0141] In addition to the various aspects, embodiments, examples, etc. described herein, the present disclosure includes the following exemplary aspects (“E”), numbered E1 through E87. This list of aspects is presented as an exemplary list, and the present application is not limited to these aspects. E1. A method for sequencing a double-stranded target nucleic acid molecule, the method comprising: (a) amplifying a physically linked nucleic acid complex on a surface to produce a physically linked nucleic acid complex amplicon bound to the surface in both forward and reverse orientations, the physically linked nucleic acid complex comprising: (i) a double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain on a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion on a second end of the double-stranded target nucleic acid molecule; (b) removing either (i) the physically linked nucleic acid complex amplicons bound to the surface in the reverse orientation, or (ii) the physically linked nucleic acid complex amplicons bound to the surface in the forward orientation; (c) cleaving a portion of the remaining bound physically linked nucleic acid complex amplicons to provide a subset of single-stranded amplicons containing information from one strand and a subset of physically linked nucleic acid complex amplicons; (d) sequencing a subset of the single-stranded amplicons to provide sequencing reads derived from the original strand of the double-stranded target nucleic acid molecule; (e) amplifying a subset of the physically linked nucleic acid complex amplicons on the surface; (f) removing the physically linked nucleic acid complex amplicons in the other direction; (g) cleaving the remaining bound physically linked nucleic acid complex amplicons to provide single-stranded amplicons containing information from the other strand; (h) sequencing the single-stranded amplicon to provide sequencing reads derived from the other original strand of the double-stranded target nucleic acid molecule. E2. A method for sequencing a double-stranded target nucleic acid molecule, the method comprising: (a) amplifying physically linked nucleic acid complexes on a surface to produce clusters of surface-bound physically linked nucleic acid complex amplicons, the physically linked nucleic acid complexes comprising: (i) a double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain on one end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion on the other end of the double-stranded target nucleic acid molecule; (b) removing either (i) the physically linked nucleic acid complex amplicon at the 5' end of the physically linked nucleic acid complex amplicon, or (ii) the physically linked nucleic acid complex amplicon bound to the surface at the 3' end of the physically linked nucleic acid complex amplicon; (c) cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons at the cleavage site to provide single-stranded amplicons that contain sequence information derived from one original strand of the double-stranded target nucleic acid molecule; (d) sequencing the single-stranded amplicon to provide sequencing reads derived from one original strand of the double-stranded target nucleic acid molecule. E3. The method of E2, wherein cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons comprises preserving at least one physically linked nucleic acid complex amplicon bound to the surface. E4. (e) amplifying at least one physically linked nucleic acid complex amplicon on the surface and rearranging the clusters of physically linked nucleic acid complex amplicons bound to the surface; (f) removing physically linked nucleic acid complex amplicons in the other direction that are not removed in (b); and (g) cleaving the remaining bound physically linked nucleic acid complex amplicons to provide single-stranded amplicons containing information derived from the other original strand of the double-stranded target nucleic acid molecule; (h) sequencing the single-stranded amplicon to provide sequencing reads derived from the other original strand of the double-stranded target nucleic acid molecule. E5. The method of any one of the preceding Examples, further comprising comparing sequence reads from one original strand with sequence reads from the other original strand to generate a consensus sequence for the double-stranded target nucleic acid molecule. E6. Identifying sequence variations in sequence reads from one original strand and sequence reads from the other original strand, wherein the sequence variations from one original strand and the other original strand are consistent sequence variations; or Excluding or not taking into account sequence variations that occur in one original strand but not in the other original strand; Any one of the methods E1-E4, further comprising: E7. comparing sequence reads from one original strand with sequence reads from the other original strand; identifying nucleotide positions that do not match between sequence reads from one original strand and sequence reads from the other original strand; generating an error-corrected sequence of the double-stranded target nucleic acid molecule by disregarding, excluding, or correcting the nucleotide positions identified as mismatched; Any one of the methods E1-E4, further comprising: E8. A method for sequencing a collection of double-stranded target nucleic acid molecules, each comprising a first strand and a second strand, the method comprising: (a) amplifying a plurality of physically linked nucleic acid complexes on a surface to generate a plurality of clonal clusters, each clonal cluster comprising a plurality of physically linked nucleic acid complex amplicons, each amplicon comprising a first strand amplicon and a second strand amplicon, each physically linked nucleic acid complex comprising: (i) a double-stranded target nucleic acid molecule from a collection; (ii) a first adaptor comprising a linker domain attached to a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion attached to a second end of the double-stranded target nucleic acid molecule; (b) removing either (i) the physically linked nucleic acid complex amplicons from each clonal cluster bound to the surface in the reverse direction or (ii) the forward direction; (c) cleaving the portion of the remaining surface-bound physically linked nucleic acid complex amplicons remaining after (b), thereby physically separating the first strand amplicons and the second strand amplicons; (d) removing unbound, physically separated first or second strand amplicons; (e) sequencing the remaining physically separated first or second strand amplicons bound to the surface to generate first or second strand nucleic acid sequence reads for each clonal cluster on the surface; A method comprising: E9. The method of E8, wherein cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons comprises preserving at least one physically linked nucleic acid complex amplicon in at least some of the surface-bound clonal clusters. E10. (f) amplifying at least one physically linked nucleic acid complex amplicon on the surface in at least some of the clonal clusters and rearranging the clonal clusters of physically linked nucleic acid complex amplicons bound to the surface; (g) removing physically linked nucleic acid complex amplicons in the opposite direction to step (b); (h) removing unbound, physically separated first or second strand amplicons; (i) cleaving the remaining bound physically linked nucleic acid complex amplicons remaining after (h), thereby physically separating the first strand amplicons and the second strand amplicons; (j) sequencing the remaining physically separated first or second strand amplicons bound to the surface to generate first or second strand nucleic acid sequence reads for each clonal cluster on the surface; The method of E9, further comprising: E11. A method for sequencing a collection of double-stranded target nucleic acid molecules, each comprising a first strand and a second strand, the method comprising: (a) amplifying a plurality of physically linked nucleic acid complexes bound on a surface to generate a plurality of clusters, each cluster comprising a plurality of physically linked nucleic acid complex amplicons representing an original double-stranded target nucleic acid molecule, each physically linked nucleic acid complex amplicon comprising a first strand amplicon and a second strand amplicon, each physically linked nucleic acid complex comprising a double-stranded target nucleic acid molecule from the collection attached to (i) a first adaptor comprising a linker domain at one end between the first strand and the second strand, and (ii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end; (b) cleaving the surface-bound physically linked nucleic acid complex amplicons, thereby physically separating the first strand amplicons and the second strand amplicons; (c) removing unbound physically separated first strand amplicons and / or unbound physically separated second strand amplicons, wherein the remaining amplicons bound to the surface comprise (i) physically separated first strand amplicons and (ii) physically separated second strand amplicons; (d) sequencing the surface-bound physically separated first-strand amplicons to generate first-strand nucleic acid sequence reads for each cluster on the surface; (e) sequencing the surface-bound, physically separated second-strand amplicons to generate second-strand nucleic acid sequence reads for each cluster on the surface; A method comprising: E12. The method of E10 or E11, further comprising comparing the first strand nucleic acid sequence reads with the second strand nucleic acid sequence reads for at least some of the clusters on the surface to generate error-corrected sequence reads of the original double-stranded target nucleic acid molecule. E13. The method of any one of E10-E12, further comprising associating nucleic acid sequence reads of a first strand of an original double-stranded target nucleic acid molecule from the collection with nucleic acid sequence reads of a second strand of the same original double-stranded target nucleic acid molecule using unique molecular identifiers (UMIs). E14. The method of E13, wherein the UMI comprises a physical location on a surface. E15. The method of E14, wherein the UMI comprises a tag sequence, a molecular specific feature, a cluster location on a surface, or a combination thereof. E16. The method of E15, wherein the molecular-specific characteristic comprises nucleic acid mapping information relative to a reference sequence, sequence information at or near the ends of the double-stranded target nucleic acid molecule, the length of the double-stranded target nucleic acid molecule, or a combination thereof. E17. The method of any one of E10-E16, further comprising using a strand defining element (SDE) to distinguish nucleic acid sequence reads of a first strand of an original double-stranded target nucleic acid molecule from nucleic acid sequence reads of a second strand from the same original double-stranded target nucleic acid molecule. E18. The method of E17, wherein the SDE is an association of sequence read information with step (e) and step (j) of E10, or step (d) and (e) of E11. E19. The method of E17, wherein the SDE comprises a portion of an adapter sequence. E20. The method of any one of E8-E19, wherein sequencing the physically separated first strand amplicon or second strand amplicon comprises sequencing by synthesis. E21. Preparing a physically linked nucleic acid complex by ligating a first adaptor and a second adaptor to each of a plurality of double-stranded target nucleic acid molecules in a collection; presenting the physically linked nucleic acid complexes on a surface, the surface having a plurality of bound oligonucleotides that are at least partially complementary to the single-stranded portion of the second adaptor, such that a plurality of the physically linked nucleic acid complexes are captured on the surface via hybridization to the plurality of bound oligonucleotides; The method of any one of E8 to E20, further comprising: E22. The method of E21, further comprising amplifying the physically linked nucleic acid complex prior to the displaying step. E23. The method of E22, wherein amplifying the physically linked nucleic acid complex prior to the presenting step comprises PCR amplification or circle amplification. E24. The method of any one of E21-E23, wherein the physically linked nucleic acid complex is captured on the surface in both the forward and reverse orientations. E25. The method of any one of E8 to E24, wherein the amplifying step of (a) comprises bridge amplification. E26. For at least some of the double-stranded target nucleic acid molecules in the collection, (i) comparing sequence reads from a first strand with sequence reads from a second strand; (ii) identifying nucleotide positions that do not match between the sequence reads from the first strand and the sequence reads from the second strand; (iii) generating error-corrected sequence reads of the double-stranded target nucleic acid molecule by disregarding, excluding, or correcting the mismatched identified nucleotide positions; and The method of any one of E8 to E25, further comprising: E27. The method of any one of E1-E26, wherein the first adaptor comprises a cleavable site or motif. E28. The method of any one of E1-E27, wherein the first adaptor and the second adaptor each comprise a sequencing primer binding site and, optionally, a single molecule identifier (SMI) sequence. E29. The method of any one of E1-E27, wherein the second adapter comprises a sequencing primer binding site, an amplification primer binding site, an index sequence, or any combination thereof. E30. The method of any one of E1-E29, wherein the linker domain comprises a cleavage site. E31. The method of any one of E1-E29, wherein the first adaptor comprises a cleavable domain. E32. The method of any one of E1-E31, wherein the first adaptor comprises a hairpin loop structure comprising a self-complementary stem portion and a single-stranded nucleotide loop portion. E33. The method of E32, wherein the single-stranded nucleotide loop portion comprises a cleavable domain. E34. The method of E32, wherein the stem portion comprises a cleavable domain. E35. The method of E33 or E34, wherein the cleavable domain comprises an enzyme recognition site. E36. The method of E35, wherein the enzyme recognition site is an endonuclease recognition site. E37. The method of E36, wherein the endonuclease is a restriction enzyme or a targeting endonuclease. E38. The method of any one of E1 to E37, wherein the second adapter is a "Y" shaped adapter. E39. The method of E38, wherein one or both arms of the Y-shaped adaptor are capable of hybridizing to a surface-bound oligonucleotide. E40. The method of any one of E1-E39, wherein the single-stranded portion of the second adapter comprises a first arm having a first primer binding site and a second arm having a second primer binding site. E41. The method of E40, wherein upon denaturation, the physically linked double-stranded nucleic acid complex comprises, 5' to 3', or 3' to 5', a first primer binding site, a first strand, a first adaptor comprising a linker domain, a second strand, and a second primer binding site. E42. The method of any one of E1-E41, wherein the surface is a sequencing surface. E43. The method of any one of E1-E42, wherein the surface is a flow cell. E44. The method of any one of E1-E43, wherein the surface is a surface of a bead. E45. The method of any one of E1-E44, wherein the amplification is selected from the group consisting of PCR amplification, isothermal amplification, polony amplification, cluster amplification, and bridge amplification. E46. The method of any one of E1-E45, wherein the amplification is bridge amplification on a surface. E47. The method of any one of E8-E46, wherein one or more of the plurality of first strand amplicons and / or the plurality of second strand amplicons are bound to a surface in a forward orientation. E48. The method of any one of E8-E46, wherein one or more of the plurality of first strand amplicons and / or the plurality of second strand amplicons are bound to the surface in reverse orientation. E49. The method of any one of E8-E48, further comprising, prior to amplifying in (a), flowing the plurality of physically linked double-stranded nucleic acid complexes over a surface. E50. The method of any one of E1-E49, wherein the surface comprises a plurality of one or more attached oligonucleotides that are at least partially complementary to one or more regions of the second adaptor. E51. The method of E50, wherein a plurality of the one or more attached oligonucleotides is at least partially complementary to a single-stranded portion of the second adaptor. E52. The method of any one of E1-E51, wherein the first strand and the second strand of the physically linked nucleic acid complex are amplified by multiple amplification reactions in step (a) to generate clusters of physically linked nucleic acid complex amplicons on the surface. E53. The method of any one of E8-E52, wherein the first strand and the second strand of each of the plurality of physically linked nucleic acid complexes are amplified in step (a) to simultaneously generate a plurality of clusters on the surface. E54. The method of any one of E1-E8 and E12-E53, wherein cleaving a portion of the bound physically linked nucleic acid complex amplicons comprises inefficiently cleaving a cleavable site in the first adaptor, resulting in both cleaved and uncleaved nucleic acid complexes within each cluster on the surface. E55. The method of E54, wherein the proportion of uncleaved nucleic acid complexes among all nucleic acid complexes in each cluster on the flow cell is 1%, 5%, 10%, 20%, 30%, 40%, 45%, or 50%. E56. The method of E54 or E55, wherein the cleaved nucleic acid complex is cleaved at a cleavable site in the linker domain of the first adaptor by a cleavage promoting agent. E57. The method of E56, wherein the cleavage is a site-directed enzymatic reaction. E58. The method of E56 or E57, wherein the cleavage enhancing agent is an endonuclease. E59. The method of E58, wherein the endonuclease is a restriction site endonuclease or a targeting endonuclease. E60. The method of E56 or E57, wherein the cleavage-facilitating factor is selected from the group consisting of a ribonucleoprotein, a Cas enzyme, a Cas9-like enzyme, a meganuclease, a transcription activator-like effector-based nuclease (TALEN), a zinc finger nuclease, an Argonaute nuclease, or a combination thereof. E61. The method of E56 or E57, wherein the cleavage-promoting factor comprises a CRISPR-associated enzyme. E62. The method of E56 or E57, wherein the cleavage promoting factor comprises Cas9, or CPF1, or a derivative thereof. E63. The method of E56 or E57, wherein the cleavage-facilitating factor comprises a nickase or nickase variant. E64. The method of E56, wherein the cleavage promoting factor comprises a chemical process. E65. Any one of the methods E54-E64, wherein the amount of uncleaved nucleic acid complex remaining on the surface can be enhanced by controlling the amount or concentration of the cleavage enhancing factor introduced for site-directed cleavage, or by controlling the amount of time the cleavage enhancing factor is introduced for site-directed cleavage. E66. The method of any one of E54-E63, wherein the uncleaved nucleic acid complex is protected by the addition of an anti-cleavage promoting agent prior to or during the cleavage step. E67. The method of E66, wherein the anti-cleavage promoting factor comprises an anti-cleavage motif in the linker domain of the first adaptor. E68. The method of E67, wherein the cleavable site is already present in the linker domain of the first adaptor, and the anti-cleavage motif is created by hybridization of an oligonucleotide comprising a sequence at least partially complementary to the linker domain of the first adaptor. E69. Cleavage of a portion of a bound physically linked nucleic acid complex amplicon (i) introducing an anti-cleavage promoting factor; (ii) introducing a cleavage-promoting factor either after or simultaneously with (i); further comprising The method of any one of E66-E68, wherein interaction with the anti-cleavage promoting factor protects the physically linked nucleic acid complex amplicon from cleavage. E70. The method of any one of E54-E63, wherein the cleavable site is created by hybridization of an oligonucleotide comprising a sequence at least partially complementary to the linker domain of the first adaptor, and wherein a physically linked nucleic acid complex amplicon that is not hybridized to the oligonucleotide is not cleaved. E71. A method for cleaving a portion of the bound physically linked nucleic acid complex amplicon, wherein the cleavable site is created by hybridization of a first oligonucleotide comprising a sequence at least partially complementary to the linker domain of the adapter, and the anti-cleavage motif is created by hybridization of a second oligonucleotide comprising a sequence at least partially complementary to the linker domain of the adapter; (i) introducing a mixture of a first and a second oligonucleotide; (ii) introducing a cleavage-promoting factor; The method of any one of E54 to E63, further comprising: E72. The method of E71, wherein either the first oligonucleotide or the second oligonucleotide is methylated. E73. The method of E70 or E71, wherein hybridization can be enhanced by controlling the amount or concentration of oligonucleotide introduced for hybridization, or by controlling the amount of time the oligonucleotide is introduced for hybridization. E74. The method of any one of E67, E68 or E71-E73, wherein the anti-cleavage motif comprises an oligonucleotide sequence having bulky appendages or side chains that prevent access to the cleavage site. E75. The method of any one of E67, E68, or E71-E73, wherein the anti-cleavage motif comprises an oligonucleotide sequence having one or more mismatches that prevent the cleavage-promoting factor from recognizing the cleavage site. E76. The method of any one of E67, E68, or E71-E73, wherein the anti-cleavage motif comprises one or more of a nucleoside analog, an abasic site, a nucleotide analog, and an oligonucleotide sequence having a peptide-nucleic acid linkage. E77. The method of any one of E54-E63, wherein the cleaved nucleic acid complex is cleaved at the cleavable site in the first adaptor by a catalytically active enzyme, and the uncleaved nucleic acid complex is protected from cleavage in the first adaptor by a catalytically inactive enzyme. E78. The method of any one of E54-E63, wherein the cleavage site is in a self-complementary portion of the first adaptor or in a single-stranded portion of the first adaptor. E79. The method of E78, wherein the cleavage site is available when the physically linked nucleic acid complex amplicon is in a self-hybridized configuration on a surface. E80. The method of any one of E54-E63, wherein the cleavage site is available when the physically linked nucleic acid complex amplicon is in a double-stranded bridge amplified configuration. E81. The method of any one of E8-E80, further comprising, prior to step (a), selectively enriching physically linked nucleic acid complexes having one or more targeted genomic regions to provide a plurality of enriched physically linked nucleic acid complexes. E82. A kit that can be used for error-corrected duplex sequencing of double-stranded nucleic acid molecules, the kit comprising: at least one set of sequencing primers; a first set of adaptor molecules comprising a linker domain; a second set of adapter molecules comprising a double-stranded portion and a single-stranded portion configured to be immobilized on a surface for amplification; Including, The primer and adapter molecules can be used in error-corrected duplex sequencing experiments, A kit comprising instructions on how to use the kit in performing error-corrected duplex sequencing of nucleic acids extracted from a biological sample. E83. The kit of E82, further comprising a cleavage-promoting factor. E84. A kit of E82 or E83, wherein the linker domain has a cleavable motif. E85. The kit of any one of E82 to E84, further comprising an anti-cleavage promoting factor. E86. The kit of any one of E82-E85, further comprising a computer program product embodied in a non-transitory computer-readable medium that, when executed on a computer or a remote computing server, performs steps for determining error-corrected duplex sequencing reads for one or more double-stranded nucleic acid molecules in a sample. E87. A sequencing system comprising: a sequencing surface comprising covalently attached oligonucleotides; a delivery system for delivering sequencing reagents to the sequencing surface; a delivery system for delivering the cleavage enhancing agent to the sequencing surface; a computing network for transmitting information related to sequencing data, the information including one or more of raw sequencing data, duplex sequencing data, and sample information; A sequencing system comprising:
[0142] conclusion The above detailed description of embodiments of the present technology is not intended to be exhaustive or to limit the present technology to the precise form described above. Specific embodiments and examples of the present technology have been described above for illustrative purposes, but those skilled in the relevant art will recognize that various equivalent modifications are possible within the scope of the present technology. For example, while steps are presented in a given order, alternative embodiments may perform the steps in a different order. Various embodiments described herein may also be combined to provide further embodiments. All references cited herein are incorporated by reference as if fully set forth herein.
[0143] From the above description, it should be understood that, although specific embodiments of the present technology are described herein for illustrative purposes, well-known structures and functions have not been described or shown in detail to avoid unnecessarily obscuring the description of the embodiments of the present technology. Where the context permits, singular or plural terms may also include plural or singular terms, respectively.
[0144] Furthermore, unless the word "or" is expressly limited in relation to a list of two or more items to mean only a single item exclusive of the other items, the use of "or" in such a list shall be interpreted as including (a) any single item in that list, (b) all items in that list, or (c) any combination of items in that list. In addition, the term "comprising" is used throughout to mean the inclusion of at least the recited features, and does not exclude any greater number of the same features and / or additional types of other features. While particular embodiments are described herein for illustrative purposes, it should be understood that various changes can be made without departing from the present technology. Furthermore, while advantages associated with particular embodiments of the present technology are described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments necessarily exhibit such advantages to fall within the scope of the present technology. Accordingly, the present disclosure and related technology may encompass other embodiments not expressly shown or described herein.
Claims
1. 1. A method for sequencing a double-stranded target nucleic acid molecule, said method comprising: (a) amplifying a physically linked nucleic acid complex on a surface to produce a physically linked nucleic acid complex amplicon bound to the surface in both forward and reverse orientations, the physically linked nucleic acid complex comprising: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain on a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion on a second end of the double-stranded target nucleic acid molecule; (b) removing either (i) the physically linked nucleic acid complex amplicons bound to the surface in the reverse orientation, or (ii) the physically linked nucleic acid complex amplicons bound to the surface in the forward orientation; (c) cleaving a portion of the remaining bound physically linked nucleic acid complex amplicons to provide a subset of single-stranded amplicons containing information from one strand and a subset of physically linked nucleic acid complex amplicons; (d) sequencing a subset of the single-stranded amplicons to provide sequencing reads derived from the original strand of the double-stranded target nucleic acid molecule; (e) amplifying a subset of the physically linked nucleic acid complex amplicons on the surface; (f) removing said physically linked nucleic acid complex amplicons in the other orientation; (g) cleaving the remaining bound physically linked nucleic acid complex amplicons to provide single-stranded amplicons containing information from the other strand; (h) sequencing the single-stranded amplicon to provide sequencing reads derived from the other original strand of the double-stranded target nucleic acid molecule; A method comprising:
2. 1. A method for sequencing a double-stranded target nucleic acid molecule, said method comprising: (a) amplifying physically linked nucleic acid complexes on a surface to produce clusters of surface-bound physically linked nucleic acid complex amplicons, the physically linked nucleic acid complexes comprising: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain on one end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion on the other end of the double-stranded target nucleic acid molecule; (b) removing either (i) the physically linked nucleic acid complex amplicon bound to the surface at the 5' end of the physically linked nucleic acid complex amplicon, or (ii) the 3' end of the physically linked nucleic acid complex amplicon; (c) cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons at the cleavage site to provide single-stranded amplicons that contain sequence information derived from one original strand of said double-stranded target nucleic acid molecule; (d) sequencing the single-stranded amplicon to provide sequencing reads derived from said one original strand of the double-stranded target nucleic acid molecule; A method comprising:
3. 3. The method of claim 2, wherein cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons comprises preserving at least one physically linked nucleic acid complex amplicon bound to the surface.
4. (e) amplifying the at least one physically linked nucleic acid complex amplicon on the surface to rearrange the clusters of physically linked nucleic acid complex amplicons bound to the surface; (f) removing the physically linked nucleic acid complex amplicons in the other orientation that are not removed in (b); (g) cleaving the remaining bound physically linked nucleic acid complex amplicons to provide single-stranded amplicons containing information derived from the other original strand of the double-stranded target nucleic acid molecule; (h) sequencing the single-stranded amplicon to provide sequencing reads derived from the other original strand of the double-stranded target nucleic acid molecule; The method of claim 3 further comprising:
5. 10. The method of any one of the preceding claims, further comprising comparing the sequence reads from the one original strand with the sequence reads from the other original strand to generate a consensus sequence for the double-stranded target nucleic acid molecule.
6. identifying sequence variations in the sequence reads from the one original strand and the sequence reads from the other original strand, wherein the sequence variations from the one original strand and the other original strand are consistent sequence variations; or excluding or disregarding sequence variations that occur in said one original strand and not in said other original strand; The method of any one of claims 1 to 4, further comprising:
7. comparing the sequence reads from the one original strand with the sequence reads from the other original strand; identifying nucleotide positions that do not match between the sequence reads from the one original strand and the sequence reads from the other original strand; generating an error-corrected sequence of said double-stranded target nucleic acid molecule by disregarding, excluding or correcting said nucleotide positions identified as mismatched; The method of any one of claims 1 to 4, further comprising:
8. 1. A method for sequencing a collection of double-stranded target nucleic acid molecules, each of which comprises a first strand and a second strand, the method comprising: (a) amplifying a plurality of physically linked nucleic acid complexes on a surface to generate a plurality of clonal clusters, each clonal cluster comprising a plurality of physically linked nucleic acid complex amplicons, each amplicon comprising a first strand amplicon and a second strand amplicon, each physically linked nucleic acid complex comprising: (i) a double-stranded target nucleic acid molecule from the collection; (ii) a first adaptor comprising a linker domain attached to a first strand of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion attached to a second end of the double-stranded target nucleic acid molecule; (b) removing either the physically linked nucleic acid complex amplicons from each clonal cluster bound to the surface in either (i) the reverse direction or (ii) the forward direction; (c) cleaving the portion of the remaining surface-bound physically linked nucleic acid complex amplicons remaining after (b), thereby physically separating the first strand amplicons and the second strand amplicons; (d) removing unbound, physically separated first strand or second strand amplicons; (e) sequencing the remaining physically separated first or second strand amplicons bound to the surface to generate the first strand or the second strand nucleic acid sequence reads for each clonal cluster on the surface; A method comprising:
9. 9. The method of claim 8, wherein cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicons comprises preserving at least one physically linked nucleic acid complex amplicon in at least some of the clonal clusters bound to the surface.
10. (f) amplifying the at least one physically linked nucleic acid complex amplicon on the surface for at least some of the clonal clusters to rearrange the clonal clusters of the physically linked nucleic acid complex amplicons bound to the surface; (g) removing said physically linked nucleic acid complex amplicons in the other orientation from step (b); (h) removing unbound, physically separated first or second strand amplicons; (i) cleaving the remaining bound physically linked nucleic acid complex amplicons remaining after (h), thereby physically separating said first strand amplicons and said second strand amplicons; (j) sequencing the remaining physically separated first or second strand amplicons bound to the surface to generate the first strand or the second strand nucleic acid sequence reads for each clonal cluster on the surface; 10. The method of claim 9, further comprising:
11. 1. A method for sequencing a collection of double-stranded target nucleic acid molecules, each of which comprises a first strand and a second strand, the method comprising: (a) amplifying a plurality of physically linked nucleic acid complexes bound on a surface to generate a plurality of clusters, each cluster comprising a plurality of physically linked nucleic acid complex amplicons representing an original double-stranded target nucleic acid molecule, each physically linked nucleic acid complex amplicon comprising a first strand amplicon and a second strand amplicon, each physically linked nucleic acid complex comprising a double-stranded target nucleic acid molecule from the collection connected to (i) a first adaptor comprising a linker domain at one end between the first strand and the second strand, and (ii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end; (b) cleaving the surface-bound physically linked nucleic acid complex amplicons, thereby physically separating said first strand amplicons and said second strand amplicons; (c) removing unbound physically separated first strand amplicons and / or unbound physically separated second strand amplicons, wherein the remaining amplicons bound to the surface include (i) the physically separated first strand amplicons and (ii) the physically separated second strand amplicons; (d) sequencing the physically separated first-strand amplicons bound to the surface to generate the first-strand nucleic acid sequence reads for each cluster on the surface; (e) sequencing the physically separated second-strand amplicons bound to the surface to generate second-strand nucleic acid sequence reads for each cluster on the surface; A method comprising:
12. 12. The method of Claim 10 or Claim 11, further comprising comparing the nucleic acid sequence reads of the first strand with the nucleic acid sequence reads of the second strand for at least some of the clusters on the surface to generate error-corrected sequence reads of the original double-stranded target nucleic acid molecule.
13. 13. The method of any one of claims 10-12, further comprising associating nucleic acid sequence reads of first strands of original double-stranded target nucleic acid molecules from the collection with nucleic acid sequence reads of second strands of the same original double-stranded target nucleic acid molecules using unique molecular identifiers (UMIs).
14. The method of claim 13 , wherein the UMI comprises a physical location on the surface.
15. 15. The method of claim 14, wherein the UMI comprises a tag sequence, a molecular specific feature, a cluster location on the surface, or a combination thereof.
16. 16. The method of claim 15, wherein the molecule-specific features comprise nucleic acid mapping information relative to a reference sequence, sequence information at or near the ends of the double-stranded target nucleic acid molecule, the length of the double-stranded target nucleic acid molecule, or a combination thereof.
17. 17. The method of any one of claims 10-16, further comprising using a strand defining element (SDE) to distinguish the nucleic acid sequence reads of the first strand of an original double-stranded target nucleic acid molecule from the nucleic acid sequence reads of the second strand from the same original double-stranded target nucleic acid molecule.
18. 18. The method of claim 17, wherein the SDE is the association of sequence read information with steps (e) and (j) of claim 10, or steps (d) and (e) of claim 11.
19. The method of claim 17 , wherein the SDE comprises a portion of an adapter sequence.
20. 20. The method of any one of claims 8-19, wherein sequencing the physically separated first-strand amplicons or second-strand amplicons comprises sequencing by synthesis.
21. preparing the physically linked nucleic acid complex by ligating the first adaptor and the second adaptor to each of a plurality of double-stranded target nucleic acid molecules in the collection; presenting the physically linked nucleic acid complexes on a surface, the surface having a plurality of bound oligonucleotides at least partially complementary to the single-stranded portion of the second adaptor, such that a plurality of physically linked nucleic acid complexes are captured on the surface via hybridization to a plurality of bound oligonucleotides; The method of any one of claims 8 to 20, further comprising:
22. The method of any one of claims 8 to 21, wherein the amplifying step of (a) comprises bridge amplification.
23. for at least some of the double-stranded target nucleic acid molecules in the collection, (i) comparing the sequence reads from the first strand with the sequence reads from the second strand; (ii) identifying nucleotide positions that do not match between the sequence reads from the first strand and the sequence reads from the second strand; (iii) generating error-corrected sequence reads of the double-stranded target nucleic acid molecule by disregarding, excluding, or correcting the identified nucleotide positions that do not match; and The method of any one of claims 8 to 22, further comprising:
24. 24. The method of any one of claims 1 to 23, wherein the first adaptor comprises a cleavable site or motif.
25. 25. The method of any one of claims 1 to 24, wherein the first adaptor comprises a cleavable domain.
26. 26. The method of any one of claims 1 to 25, wherein the first adaptor comprises a hairpin loop structure comprising a self-complementary stem portion and a single-stranded nucleotide loop portion.
27. 27. The method of claim 26, wherein the cleavable domain is in the single-stranded nucleotide loop portion or the stem portion.
28. 34. The method of claim 33, wherein the cleavable domain comprises an enzyme recognition site.
29. 29. The method of claim 28, wherein the enzyme recognition site is targeted by a restriction enzyme or a targeting endonuclease.
30. 30. The method of any one of claims 1 to 29, wherein the single-stranded portion of the second adapter comprises a first arm having a first primer binding site and a second arm having a second primer binding site.
31. 31. The method of claim 30, wherein upon denaturation, the physically linked double-stranded nucleic acid complex comprises, from 5' to 3' or 3' to 5', the first primer binding site, the first strand, the first adaptor comprising the linker domain, the second strand, and the second primer binding site.
32. 10. The method of any one of the preceding claims, wherein the surface is a sequencing surface.
33. 33. The method of any one of claims 8-32, further comprising, prior to said amplifying in (a), flowing said plurality of physically linked double-stranded nucleic acid complexes over said surface.
34. 10. The method of any one of the preceding claims, wherein the surface comprises a plurality of one or more bound oligonucleotides that are at least partially complementary to one or more regions of the second adaptor.
35. 35. The method of claim 34, wherein one or more of the plurality of attached oligonucleotides is at least partially complementary to the single-stranded portion of the second adaptor.
36. 36. The method of any one of claims 1 to 35, wherein the first and second strands of the physically linked nucleic acid complexes are amplified in step (a) by multiple amplification reactions to generate clusters of the physically linked nucleic acid complex amplicons on the surface.
37. 37. The method of any one of claims 8 to 36, wherein the first strand and the second strand of each of the plurality of physically linked nucleic acid complexes are amplified in step (a) to simultaneously generate the plurality of clusters on the surface.
38. 38. The method of any one of claims 1-8 and 12-37, wherein cleaving a portion of the bound physically linked nucleic acid complex amplicons comprises inefficiently cleaving a cleavable site in the first adaptor, resulting in both cleaved and uncleaved nucleic acid complexes within each cluster on the surface.
39. 39. The method of claim 38, wherein the proportion of uncleaved nucleic acid complexes of all nucleic acid complexes in each cluster on the flow cell is 1%, 5%, 10%, 20%, 30%, 40%, 45%, or 50%.
40. 40. The method of Claim 38 or 39, wherein the cleaved nucleic acid complex is cleaved at a cleavable site in the linker domain of the first adaptor by a cleavage-promoting agent.
41. 41. The method of claim 40, wherein the cleavage is a site-directed enzymatic reaction.
42. 42. The method of claim 40 or claim 41, wherein the cleavage-facilitating agent is an endonuclease.
43. 42. The method of claim 40 or claim 41, wherein the cleavage-facilitating factor comprises a CRISPR-associated enzyme.
44. 42. The method of claim 40 or claim 41, wherein the cleavage-facilitating factor comprises a nickase or a nickase variant.
45. 41. The method of claim 40, wherein the cleavage-facilitating agent comprises a chemical process.
46. 46. The method of any one of claims 38-45, wherein the amount of uncleaved nucleic acid complex remaining on the surface can be expanded by controlling the amount or concentration of the cleavage facilitating factor introduced for site-directed cleavage, or by controlling the amount of time the cleavage facilitating factor is introduced for site-directed cleavage.
47. 46. The method of any one of claims 38 to 45, wherein the uncleaved nucleic acid complex is protected by the addition of an anti-cleavage promoting agent before or during the cleavage step.
48. cleaving a portion of the bound physically linked nucleic acid complex amplicon; (i) introducing an anti-cleavage promoting factor; (ii) introducing the cleavage-facilitating factor either after (i) or simultaneously with (i); further comprising 48. The method of claim 47, wherein interaction with the anti-cleavage promoting factor protects the physically linked nucleic acid complex amplicon from cleavage.
49. 45. The method of any one of claims 38 to 44, wherein the cleavable site is created by hybridization of an oligonucleotide comprising a sequence at least partially complementary to the linker domain of the first adapter, and physically linked nucleic acid complex amplicons that are not hybridized to the oligonucleotide are not cleaved.
50. wherein the cleavable site is created by hybridization of a first oligonucleotide comprising a sequence at least partially complementary to the linker domain of the adapter, and a cleavage-resistant motif is created by hybridization of a second oligonucleotide comprising a sequence at least partially complementary to the linker domain of the adapter, and cleaving a portion of the bound physically linked nucleic acid complex amplicon; (i) introducing a mixture of the first and second oligonucleotides; (ii) introducing the cleavage-facilitating factor; The method of any one of claims 38 to 44, further comprising:
51. 45. The method of any one of claims 38 to 44, wherein the cleaved nucleic acid complex is cleaved at a cleavable site in the first adaptor by a catalytically active enzyme, and the uncleaved nucleic acid complex is protected from cleavage in the first adaptor by a catalytically inactive enzyme.
52. 45. The method of any one of claims 38 to 44, wherein the cleavage site is in a self-complementary portion of the first adaptor or in a single-stranded portion of the first adaptor.
53. 53. The method of claim 52, wherein the cleavage site is available when the physically linked nucleic acid complex amplicon is in a self-hybridized configuration on the surface.
54. 45. The method of any one of claims 38 to 44, wherein the cleavage site is available when the physically linked nucleic acid complex amplicon is in a double-stranded bridge amplified configuration.