Methods and reagents for nucleic acid sequencing and related applications
By amplifying and cleaving physically connected nucleic acid complexes on the surface, and combining them with adaptors and strand-limiting elements, the problem of insufficient double-stranded sequencing conversion efficiency is solved, enabling faster, lower-cost, high-accuracy sequencing and enhancement of nucleic acid sequence information.
Patent Information
- Application Number
- CN202080055766.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-01
- Filing Date
- 2020-08-01
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2040-08-01
AI Technical Summary
The conversion efficiency of existing double-stranded sequencing technologies is insufficient, resulting in a limited copy number of the target double-stranded nucleic acid, which affects the practicality of high-accuracy sequencing. Cost-effective and efficient methods are needed to synthesize the original sequence reads of nucleic acid molecules.
By amplifying physically linked nucleic acid complexes on the surface, cutting and sequencing single-stranded amplicones, combining adaptors and strand-limiting elements, the efficiency of double-stranded sequencing conversion is improved. The original strand sequence is distinguished using unique molecular identifiers and strand-limiting elements, and error correction is performed.
It achieves faster, lower-cost, and more accurate sequencing, improves double-strand sequencing conversion efficiency, and enhances the ability to generate nucleic acid sequence information.
Smart Images

Figure BDA0003495593660000121 
Figure HDA0003495593670000011 
Figure HDA0003495593670000021
Abstract
Description
Technical Field
[0001] This invention generally relates to methods and related reagents for providing highly accurate (e.g., error-corrected) nucleic acid sequences. Specifically, several embodiments relate to hairpin-shaped adaptor molecules and methods for using such adaptors in double-stranded sequencing and other sequencing applications.
[0002] Cross-reference to related applications
[0003] This application claims priority and benefit to U.S. Provisional Patent Application No. 62 / 881,936, filed August 1, 2019, the disclosure of which is hereby incorporated in its entirety. Background Technology
[0004] Double-stranded sequencing is an error-correction method that achieves superior sequence accuracy by comparing the sequence information from the two strands of various double-stranded nucleic acid molecules. Regarding the efficiency of double-stranded sequencing or other high-accuracy sequencing methods, conversion efficiency can be defined as the fraction of unique nucleic acid molecules input into the sequencing library preparation reaction that produces at least one double-stranded common sequence read (or other high-accuracy sequence read) from the sequencing library preparation reaction. In some cases, insufficient conversion efficiency can limit the practicality of high-accuracy sequencing in applications that it would otherwise be well-suited for. For example, low conversion efficiency will result in a limited copy number of the target double-stranded nucleic acid, which may lead to less sequence information produced than desired. Cost-effective and efficient methods are needed to synthesize raw sequence reads of nucleic acid molecules for various applications, including double-stranded sequencing. Summary of the Invention
[0005] This invention generally relates to methods and related reagents for nucleic acid sequencing. Specifically, some aspects of the technology relate to methods for achieving high-accuracy sequencing reads and generating increased desired data at faster rates (e.g., with fewer steps) and / or at lower costs (e.g., using fewer reagents). Other aspects of the technology relate to methods and reagents for improving double-strand sequencing conversion efficiency. These various aspects of the invention have numerous applications in preclinical and clinical testing and diagnostics, as well as other applications.
[0006] In some aspects, this disclosure provides a method for sequencing a double-stranded target nucleic acid molecule, the method comprising the steps of: (a) amplifying a physically linked nucleic acid complex on a surface to generate a physically linked nucleic acid complex amplicon that binds to the surface in both a forward orientation and a reverse orientation, wherein the physically linked nucleic acid complex comprises: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor including a linker domain on a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion on a second end of the double-stranded target nucleic acid molecule; (b) removing (i) the physically linked nucleic acid complex amplicon bound to the surface in the reverse orientation or (ii) the physically linked nucleus bound to the surface in the forward orientation. (c) Cleaving a portion of the remaining physically bound nucleic acid complex amplicon to provide a subset of single-stranded amplicon including information from one strand and a subset of physically bound nucleic acid complex amplicon; (d) Sequencing the subset of single-stranded amplicon to provide a sequencing read from the original strand of the double-stranded target nucleic acid molecule; (e) Amplifying the subset of physically bound nucleic acid complex amplicon on the surface; (f) Removing the physically bound nucleic acid complex amplicon in other orientations; (g) Cleaving the remaining physically bound nucleic acid complex amplicon to provide a single-stranded amplicon including information from the other strand; and (h) Sequencing the single-stranded amplicon to provide a sequencing read from the other original strand of the double-stranded target nucleic acid molecule.
[0007] In some aspects, this disclosure provides a method for sequencing a double-stranded target nucleic acid molecule, the method comprising the steps of: (a) amplifying physically linked nucleic acid complexes on a surface to generate clusters of physically linked nucleic acid complex amplicones bound to the surface, wherein the physically linked nucleic acid complexes comprise: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor including an adapter domain at one end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end of the double-stranded target nucleic acid molecule; (b) removing the physically linked nucleic acid complex amplicon bound to the surface at (i) the 5' end of the physically linked nucleic acid complex amplicon or (ii) the 3' end of the physically linked nucleic acid complex amplicon; (c) cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicones at a cleavage site to provide single-stranded amplicon including sequence information of one original strand of the double-stranded target nucleic acid molecule; and (d) sequencing the single-stranded amplicon to provide a sequencing read of the original strand of the double-stranded target nucleic acid molecule. In some aspects, the method further includes cleaving at least a portion of the remaining physically linked nucleic acid complex amplicon, including retaining at least one physically linked nucleic acid complex amplicon bound to the surface. In some aspects, the method further includes the steps of: (e) amplifying the at least one physically linked nucleic acid complex amplicon on the surface to re-proliferate clusters of the physically linked nucleic acid complex amplicon bound to the surface; (f) removing nucleic acid complex amplicon with other orientations that were not removed in (b); (g) cleaving the remaining physically linked nucleic acid complex amplicon to provide a single-stranded amplicon including information from another original strand of the double-stranded target nucleic acid molecule; and (h) sequencing the single-stranded amplicon to provide a sequencing read from the other original strand of the double-stranded target nucleic acid molecule.
[0008] In some aspects, the method further includes the steps of: comparing a sequence read from one original strand with a sequence read from another original strand to generate a common sequence of the double-stranded target nucleic acid molecule. In some aspects, the method further includes the steps of: identifying sequence variations in the sequence reads from one original strand and the sequence reads from the other original strand, wherein the sequence variations from the one original strand and the other original strand are consistent sequence variations; or eliminating or ignoring sequence variations occurring in one original strand but not in the other original strand. In some aspects, the method further includes the steps of: comparing a sequence read from one original strand with a sequence read from another original strand; identifying inconsistent nucleotide positions between the sequence reads from one original strand and the sequence reads from the other original strand; and generating a miscorrected sequence of the double-stranded target nucleic acid molecule by ignoring, eliminating, or correcting the identified inconsistent nucleotide positions.
[0009] In some aspects, this disclosure provides a method for sequencing a population of double-stranded target nucleic acid molecules, each double-stranded target nucleic acid step comprising a first strand and a second strand, the method comprising the steps of: (a) amplifying multiple physically linked nucleic acid complexes on a surface to generate multiple clonal clusters, each clonal cluster comprising multiple physically linked nucleic acid complex amplicones, each nucleic acid complex amplicon comprising a first-strand amplicon and a second-strand amplicon, wherein each physically linked nucleic acid complex comprises: (i) a double-stranded target nucleic acid molecule from the population; (ii) a first adaptor including an adapter domain attached to a first end of the double-stranded target nucleic acid molecule; and (iii) a double-stranded target nucleic acid molecule having a double-stranded target nucleic acid complex attached to a second end of the double-stranded target nucleic acid molecule. (a) a second adapter for the single-stranded portion and the single-stranded portion; (b) removing the physically linked nucleic acid complex amplicon from each clonal cluster bound to the surface in the reverse orientation described in (i) or the forward orientation described in (ii); (c) cleaving a portion of the remaining surface-bound physically linked nucleic acid complex amplicon after (b), thereby physically separating the first-stranded amplicon and the second-stranded amplicon; (d) removing unbound physically separated first-stranded or second-stranded amplicon; and (e) sequencing the remaining physically separated first-stranded or second-stranded amplicon bound to the surface to generate a nucleic acid sequence read of the first strand or the second strand for each clonal cluster on the surface. In some aspects, cleaving at least a portion of the remaining bound physically linked nucleic acid complex amplicon includes retaining at least one physically linked nucleic acid complex amplicon from at least some of the clonal clusters bound to the surface. In some aspects, the method further includes the following steps: (f) amplifying the at least one physically linked nucleic acid complex amplicon on the surface in at least some of the clone clusters to re-proliferate the clone clusters of physically linked nucleic acid complex amplicon bound to the surface; (g) removing the physically linked nucleic acid complex amplicon in other orientations from step (b); (h) removing unbound physically separated first-strand amplicon or second-strand amplicon; (i) cleaving the remaining physically linked nucleic acid complex amplicon after (h), thereby physically separating the first-strand amplicon and the second-strand amplicon; and (j) sequencing the remaining physically separated first-strand amplicon or second-strand amplicon bound to the surface to generate a nucleic acid sequence read of the first strand or the second strand for each clone cluster on the surface.
[0010] In some aspects, this disclosure provides a method for sequencing a population of double-stranded target nucleic acid molecules, each double-stranded target nucleic acid step comprising a first strand and a second strand, the method comprising the steps of: (a) amplifying multiple physically linked nucleic acid complexes bound on a surface to generate a plurality of clusters, each cluster comprising a plurality of physically linked nucleic acid complex amplicones representing the original double-stranded target nucleic acid molecule, wherein each physically linked nucleic acid complex amplicon comprises a first-strand amplicon and a second-strand amplicon, and wherein each physically linked nucleic acid complex comprises a double-stranded target nucleic acid molecule from the population, the double-stranded target nucleic acid molecule: (i) is linked at one end to a first adapter comprising a adapter structural domain between the first strand and the second strand; and (ii) is linked at the other end to a second adapter having a double-stranded portion and a single-stranded portion. (a) Connecting the surface to the adapter; (b) cleaving the physically connected nucleic acid complex amplicon bound to the surface, thereby physically separating the first-strand amplicon and the second-strand amplicon; (c) removing unbound physically separated first-strand amplicon and / or unbound physically separated second-strand amplicon, wherein the remaining amplicon bound to the surface comprises: (i) the physically separated first-strand amplicon; and (ii) the physically separated second-strand amplicon; (d) sequencing the physically separated first-strand amplicon bound to the surface to generate a nucleic acid sequence read of the first strand for each cluster on the surface; and (e) sequencing the physically separated second-strand amplicon bound to the surface to generate a nucleic acid sequence read of the second strand for each cluster on the surface.
[0011] In some aspects, for at least some of the clusters on the surface, the method further includes the step of comparing the nucleic acid sequence reads of the first strand with the nucleic acid sequence reads of the second strand to generate error-corrected sequence reads of the original double-stranded target nucleic acid molecule. In some aspects, the method further includes the step of using a unique molecular identifier (UMI) to associate the nucleic acid sequence reads of the first strand of the original double-stranded target nucleic acid molecule from the population with the nucleic acid sequence reads of the second strand of the same original double-stranded target nucleic acid molecule. In some aspects, the UMI includes a physical location on the surface. In another aspect, the UMI includes a tag sequence, a molecular-specific feature, a cluster location on the surface, or a combination thereof. In some aspects, the molecular-specific feature includes nucleic acid mapping information against a reference sequence, sequence information at or near the ends of the double-stranded target nucleic acid molecule, the length of the double-stranded target nucleic acid molecule, or a combination thereof.
[0012] In some aspects, the method further includes the step of: using a strand-defining element (SDE) to distinguish the nucleic acid sequence read of the first strand of the original double-stranded target nucleic acid molecule from the nucleic acid sequence read of the second strand of the same original double-stranded target nucleic acid molecule. In some aspects, the SDE is sequence read information associated with steps (e) and (j) or steps (d) and (e). In some aspects, the SDE includes a portion of an adaptor sequence.
[0013] In some respects, sequencing the physically separated first-strand amplicon or second-strand amplicon includes synthetic sequencing.
[0014] In some aspects, the method further includes the steps of: preparing the physically linked nucleic acid complex by linking the first and second adaptors to each of a plurality of double-stranded target nucleic acid molecules in the population; and presenting the physically linked nucleic acid complex to the surface having a plurality of binding oligonucleotides at least partially complementary to the single-stranded portion of the second adaptor, such that the plurality of physically linked nucleic acid complexes are captured on the surface by hybridization with the plurality of binding oligonucleotides. In some aspects, the method further includes the step of: amplifying the physically linked nucleic acid complex prior to the presentation step. In some aspects, amplifying the physically linked nucleic acid complex prior to the presentation step includes PCR amplification or circular amplification. In other aspects, the physically linked nucleic acid complex is captured on the surface in both forward and reverse orientations.
[0015] In some respects, the amplification step includes bridging amplification.
[0016] In some aspects, the method for at least some double-stranded target nucleic acid molecules in the population further includes the steps of: (i) comparing a sequence read from the first strand with a sequence read from the second strand; (ii) identifying inconsistent nucleotide positions between the sequence read from the first strand and the sequence read from the second strand; and (iii) generating an erroneously corrected sequence read of the double-stranded target nucleic acid molecule by ignoring, eliminating, or correcting the identified inconsistent nucleotide positions.
[0017] In some aspects, the first adaptor includes a cleavable site or motif. In some aspects, the first and second adaptors each include a sequencing primer binding site and optionally a single-molecule identifier (SMI) sequence. In some aspects, the second adaptor includes a sequencing primer binding site, an amplification primer binding site, an index sequence, or any combination thereof. In some aspects, the adapter domain includes a cleavable site. In some aspects, the first adaptor includes a cleavable domain. In some aspects, the first adaptor includes a hairpin loop structure comprising a self-complementary stem portion and a single-stranded nucleotide loop portion. In some aspects, the single-stranded nucleotide loop portion includes a cleavable domain. In some aspects, the stem portion includes a cleavable domain. In some aspects, the cleavable domain includes an enzyme recognition site. In some aspects, the enzyme recognition site is an endonuclease recognition site. In some aspects, the endonuclease is a restriction enzyme or a targeted endonuclease.
[0018] In some respects, the second adaptor is a "Y"-shaped adaptor. In some respects, one or both arms of the Y-shaped adaptor can hybridize with an oligonucleotide bound to the surface.
[0019] In some aspects, the single-stranded portion of the second adaptor includes a first arm having a first primer binding site and a second arm having a second primer binding site. In some aspects, when denatured, the physically linked double-stranded nucleic acid complex, from 5' to 3' or from 3' to 5', comprises: the first primer binding site, the first strand, the first adaptor including the adapter domain, the second strand, and the second primer binding site.
[0020] In some respects, the surface is a sequencing surface. In some respects, the surface is a flow cell. In other respects, the surface is the surface of a bead.
[0021] In some aspects, the amplification is selected from the group consisting of PCR amplification, isothermal amplification, clonal amplification, cluster amplification, and bridge amplification. In some aspects, the amplification is bridge amplification on the surface.
[0022] In some aspects, one or more of the plurality of first-strand amplicons and / or the plurality of second-strand amplicons bind to the surface in a forward orientation. In some aspects, one or more of the plurality of first-strand amplicons and / or the plurality of second-strand amplicons bind to the surface in a reverse orientation.
[0023] In some aspects, the method further includes the step of: allowing the multiple physically linked double-stranded nucleic acid complexes to flow through the surface prior to the amplification.
[0024] In some aspects, the surface includes a plurality of one or more bound oligonucleotides that are at least partially complementary to one or more regions of the second adaptor. In some aspects, the plurality of one or more bound oligonucleotides are at least partially complementary to the single-stranded portion of the second adaptor.
[0025] In some aspects, the first and second strands of the physically linked nucleic acid complexes are amplified through multiple amplification reactions to generate clusters of the physically linked nucleic acid complex amplicones on the surface. In some aspects, the first and second strands of each of the plurality of physically linked nucleic acid complexes are amplified to simultaneously generate the plurality of clusters on the surface.
[0026] In some aspects, cleaving a portion of the physically linked nucleic acid complex amplicon comprises inefficient cleavage at a cleavable site in the first adapter, thereby producing both cleaved and uncleaved nucleic acid complexes within each cluster on the surface. In some aspects, the ratio of uncleaved nucleic acid complexes in all nucleic acid complexes within each cluster on the flow cell is 1%, 5%, 10%, 20%, 30%, 40%, 45%, or 50%. In some aspects, the cleaved nucleic acid complex is cleaved at a cleavable site in the linker domain of the first adapter by a cleavage promoter. In some aspects, the cleavage is a site-directed enzymatic reaction. In some aspects, the cleavage promoter is a nuclease. In some aspects, the nuclease is a restriction site nuclease or a targeted nuclease. In some aspects, the cleavage promoter is selected from the group consisting of: ribonucleoproteins, Cas enzymes, Cas9-like enzymes, broad-spectrum nucleases, transcription activator-like effector-based nucleases (TALENs), zinc finger nucleases, argonaute nucleases, or combinations thereof. In some aspects, the cleavage accelerator comprises a CRISPR-associated enzyme. In some aspects, the cleavage accelerator comprises Cas9 or CPF1 or derivatives thereof. In other aspects, the cleavage accelerator comprises a cleavage enzyme or a variant of a cleavage enzyme. In some aspects, the cleavage accelerator comprises a chemical process.
[0027] In some aspects, the amount of uncut nucleic acid complex remaining on the surface can be scaled by controlling the amount or concentration of the cleavage accelerator introduced for site-specific cleavage or by controlling the amount of time the cleavage accelerator is introduced for site-specific cleavage. In some aspects, the uncut nucleic acid complex is protected by adding an anti-cleavage accelerator before or during the cleavage step. In some aspects, the anti-cleavage accelerator includes an anti-cleavage motif in the linker domain of the first adapter. In some aspects, the cleavable site is already present in the linker domain of the first adapter, and the anti-cleavage motif is generated by hybridization with an oligonucleotide comprising a sequence at least partially complementary to the linker domain of the first adapter.
[0028] In some aspects, cleaving a portion of the physically linked nucleic acid complex amplicon further includes the steps of: (i) introducing the anti-cleavage promoter; and (ii) introducing the cleavage promoter after or simultaneously with (i), wherein the interaction with the anti-cleavage promoter protects the physically linked nucleic acid complex amplicon from cleavage. In some aspects, the cleavable site is generated by hybridization with an oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the first adaptor, and wherein physically linked nucleic acid complex amplicon that does not hybridize with the oligonucleotide is not cleaved. In some aspects, the cleavable site is generated by hybridization with a first oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the adaptor, and the anti-cleavage motif is generated by hybridization with a second oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the adaptor, and wherein cleaving a portion of the physically linked nucleic acid complex amplicon further includes: (i) introducing a mixture of the first oligonucleotide and the second oligonucleotide; and (ii) introducing the cleavage promoter. In some aspects, the first oligonucleotide or the second oligonucleotide is methylated. In some aspects, the hybridization can be scaled by controlling the amount or concentration of the oligonucleotide introduced for hybridization or by controlling the amount of time the oligonucleotide is introduced for hybridization. In some aspects, the anti-cleavage motif comprises an oligonucleotide sequence having a bulk adduct or side chain that prevents entry into the cleavage site. In some aspects, the anti-cleavage motif comprises one or more mismatched oligonucleotide sequences that prevent cleavage promoters from recognizing the cleavage site. In some aspects, the anti-cleavage motif comprises one or more of the following: an oligonucleotide sequence having a nucleoside analog, a base-free site, a nucleotide analog, and a peptide-nucleic acid bond.
[0029] In some aspects, the cleaved nucleic acid complex is cleaved by a catalytically active enzyme at a cleavable site in the first adaptor, and the uncleaved nucleic acid complex is protected from cleavage by a catalytically inactivating enzyme in the first adaptor. In some aspects, the cleavage site is located in the self-complementary portion or the single-stranded portion of the first adaptor. In some aspects, the cleavage site is available when the physically connected nucleic acid complex amplicon is in a self-hybridization configuration on the surface. In some aspects, the cleavage site is available when the physically connected nucleic acid complex amplicon is in a double-stranded bridge amplification configuration.
[0030] In some aspects, the method further includes the step of selectively enriching nucleic acid complexes with physical links to one or more target genomic regions prior to step (a) to provide a plurality of enriched physically linked nucleic acid complexes. Attached Figure Description
[0031] Many aspects of this disclosure can be better understood with reference to the accompanying drawings, which are shown in the diagrams below. These drawings are for illustrative purposes only and are not intended to be limiting. The components in the drawings are not necessarily to scale. Rather, the focus is on clearly demonstrating the principles of this disclosure.
[0032] Figure 1A and 1B This is a conceptual illustration of various double-stranded sequencing method steps according to embodiments of the present invention.
[0033] Figure 2A and 2B Nucleic acid adaptor molecules for use with embodiments of the present invention are shown, as well as double-stranded adaptor-nucleic acid complexes formed by the linkage of such adaptors with target double-stranded nucleic acid fragments according to another embodiment of the present invention.
[0034] Figures 3A-3D The steps in a method for sequencing a double-stranded adapter-nucleic acid complex according to embodiments of the present invention are illustrated.
[0035] Figures 4A-4E The steps in a method for sequencing a double-stranded adapter-nucleic acid complex according to another embodiment of the present invention are shown.
[0036] Figures 5A-5E Steps in a method for sequencing a double-stranded adapter-nucleic acid complex according to another embodiment of the present invention are shown.
[0037] Figure 6-11B Various connectors according to embodiments of the present invention and their uses are illustrated.
[0038] Figures 12A-12C A method for cleaving a double-stranded linker-nucleic acid complex is shown according to yet another embodiment of the technology of the present invention.
[0039] definition
[0040] To facilitate understanding of this disclosure, certain terms are defined below. Further definitions for these and other terms are set forth throughout the specification.
[0041] In this application, unless otherwise specified in the context, the term "a" is to be understood as meaning "at least one". As used herein, the term "or" is to be understood as meaning "and / or". In this application, the terms "comprising" and "including" are to be understood as including the listed components or steps, whether presented individually or together with one or more other components or steps. Where the scope provided herein applies, the term includes the endpoint. As used herein, the term "comprise" and variations thereof, such as "comprising" and "comprises", are not intended to exclude other additives, components, wholes, or steps.
[0042] Approximately: When used herein as a reference value, the term "approximately" refers to a value similar to the reference value in the context. Generally, those skilled in the art will understand the degree of variation implied by "approximately" in the context. For example, in some embodiments, the term "approximately" may cover values within a range of 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the reference value.
[0043] Analog: As used herein, the term "analog" refers to a substance that shares one or more specific structural features, elements, components, or portions with a reference substance. Typically, an "analog" exhibits significant structural similarity to a reference substance, such as sharing a core or consensus structure, but also differs in certain discrete ways. In some embodiments, an analog is a substance that can be generated from a reference substance, for example, by chemical manipulation of the reference substance. In some embodiments, an analog is a substance that can be generated by performing a synthetic process substantially similar to (e.g., sharing multiple steps) the synthetic process used to generate the reference substance. In some embodiments, an analog is generated by, or may be generated by, performing a synthetic process different from the synthetic process used to generate the reference substance.
[0044] Biological Samples: As used herein, the term "biological sample" or "sample" generally refers to a sample obtained or derived from a biological source of interest (e.g., tissue or organism or cell culture) as described herein. In some embodiments, the source of interest includes organisms such as animals or humans. In other embodiments, the source of interest includes microorganisms such as bacteria, viruses, protozoa, or fungi. In yet another embodiment, the source of interest may be synthetic tissues, organisms, cell cultures, nucleic acids, or other materials. In yet another embodiment, the source of interest may be plant-based organisms. In yet another embodiment, the sample may be an environmental sample, such as a water sample, soil sample, archaeological sample, or other sample collected from a non-biological source. In other embodiments, the sample may be a multi-organism sample (e.g., a mixed-organism sample). In some embodiments, a biological sample is or includes biological tissue or fluid. In some embodiments, a biological sample may be or include bone marrow; blood; blood cells; ascites; tissue or fine-needle biopsy samples; cell-containing body fluids; free-floating nucleic acids; sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; feces; lymph; gynecological fluids; skin swabs; vaginal swabs; Pap smears, oral swabs; nasal swabs; irrigation or lavage fluids, such as catheter lavage fluid or bronchoalveolar lavage fluid; vaginal fluids, aspirates; waste; bone marrow samples; tissue biopsy samples; fetal tissue or fluids; surgical samples; feces, other body fluids, secretions and / or excretions; and / or cells therefrom. In some embodiments, a biological sample is or includes cells obtained from an individual. In some embodiments, the obtained cells are or comprise cells from the individual from whom the sample was obtained. In a particular embodiment, the biological sample is a liquid biopsy sample obtained from a subject. In some embodiments, the sample is a “primary sample” obtained directly from the source of interest by any suitable means. For example, in some embodiments, primary biological samples are obtained by methods selected from the group consisting of: biopsy (e.g., fine-needle aspiration or tissue biopsy), surgery, and collection of bodily fluids (e.g., blood, lymph, feces, etc.). In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing a primary sample (e.g., by removing one or more components of the primary sample and / or by adding one or more agents to the primary sample). For example, using a semi-permeable membrane filtration. Such "processed samples" can include, for example, nucleic acids or proteins extracted from a sample or obtained by techniques such as subjecting the primary sample to amplification or reverse transcription of mRNA, separation and / or purification of certain components, etc. Cleavage site: also known as a "cleavage motif" and "cutting site," is a bond or bond pair between nucleotides in a nucleic acid molecule. In the case of double-stranded nucleic acid molecules (such as double-stranded DNA), the cleavage site may require bonds (typically phosphodiester bonds) that are adjacent to each other in the double-stranded molecule, such that a "blunt" end is formed after cleavage.Cleavage sites can also be defined by two nucleotide bonds on each strand of a single strand that are not directly opposite each other, leaving "sticky ends" upon cleavage, thus preserving the region of the single-stranded nucleotide at the end of the molecule. Cleavage sites can be defined by specific nucleotide sequences that can be recognized by enzymes such as restriction enzymes or other sequence-recognizing endonucleases such as CRISPR / Cas9. Cleavage sites can be within the recognition sequences of such enzymes (i.e., type 1 restriction enzymes) or adjacent to them via defined nucleotide intervals (i.e., type 2 restriction enzymes). Cleavage sites can also be defined by the position of modified nucleotides that can be recognized by certain nucleases. For example, abase-free sites can be recognized and cleaved by endonuclease VII and the enzyme FPG. Uracil bases can be recognized by the enzyme UDG and converted into abase-free sites. Ribonucleotides in additional DNA sequences can be recognized and cleaved by RNaseH2 when annealed to complementary DNA sequences.
[0045] Determination: Many methods described herein include a step of “determination.” Those skilled in the art who read this specification will understand that such “determination” can be achieved using or by any of a variety of techniques available to those skilled in the art, including, for example, specific techniques explicitly mentioned herein. In some embodiments, determination involves operations involving a physical sample. In some embodiments, determination involves consideration and / or manipulation of data or information, for example, using a computer or other processing unit suitable for performing relevant analyses. In some embodiments, determination involves receiving relevant information and / or materials from a source. In some embodiments, determination involves comparing one or more features of a sample or entity with a comparable reference.
[0046] Double-stranded sequencing (DS): As used in this article, “double-stranded sequencing (DS)” in its broadest sense refers to a method of error correction that achieves superior accuracy by comparing sequences from the two strands of each DNA molecule.
[0047] Error-correcting: As used herein, the term "error-correcting" or "error correction" refers to the resulting product or process (e.g., due to nucleotide mismatch) that identifies and subsequently ignores, eliminates, or otherwise corrects one or more nucleotide errors in a region of a nucleic acid molecule where the two strands of the double-stranded portion of the molecule are not perfectly complementary to each other. In some aspects, mismatches can be the result of point mutations, deletions, insertions, or chemical modifications. In some aspects, mismatches involve base pairs of opposite strands having a sequence, such as, but not limited to, AA, CC, TT, GG, AC, AG, TC, TG, or the inverse of these pairs (which are equivalent, i.e., AG is equivalent to GA), resulting in the deletion, insertion, or other modification of one or more bases. Mismatches can be of biological origin, of DNA synthesis origin, or caused by damaged or modified nucleotide bases. In some aspects, damaged or modified nucleotide bases are present on one or both strands and are converted into mismatches by an enzymatic process (e.g., DNA polymerase, DNA glycosylase, or another nucleic acid modifying enzyme or chemical process). In some respects, this mismatch can be used to infer the presence of nucleic acid damage or nucleotide modification prior to enzymatic processes or chemical treatments.
[0048] Expression: As used herein, “expression” of a nucleic acid sequence means one or more of the following events: (1) the generation of an RNA template from a DNA sequence (e.g., by transcription); (2) the processing of RNA transcripts (e.g., by splicing, editing, 5' cap formation and / or 3' end formation); (3) the translation of RNA into a polypeptide or protein; and / or (4) post-translational modifications of a polypeptide or protein.
[0049] Functionalized Surfaces: As used herein, the term "functionalized surface" refers to a solid surface, bead, or other immobilization structure capable of binding or immobilizing nucleic acid molecules or other capture portions. In some embodiments, a functionalized surface includes a binding portion capable of capturing target nucleic acids. In some embodiments, the binding portion is directly connected to the surface. In some embodiments, an oligonucleotide at least partially complementary to the target nucleic acid acts as the binding portion. In some embodiments, the oligonucleotide is covalently bound to the surface. In some embodiments, a functionalized surface may include controlled-pore glass (CPG), magnetic porous glass (MPG), and other glass or non-glass surfaces. In one embodiment, a functionalized surface may be a sequencing surface, such as the surface of a flow cell. Chemical functionalization may require ketone modification, aldehyde modification, thiol modification, azide modification, and alkyne modification, etc. In some embodiments, the functionalized surface and the oligonucleotide for hybridization capture are linked using one or more sets of immobilization chemicals that form amide bonds, alkylamine bonds, thiourea bonds, diazo bonds, hydrazine bonds, and other surface chemicals. In some embodiments, one or more of a group of reagents are used to link functionalized surfaces and oligonucleotides for hybridization capture, said reagents including EDAC, NHS, sodium periodate, glutaraldehyde, pyridyl disulfide, nitrite, biotin, and other linking reagents.
[0050] gRNA: As used herein, “gRNA” or “guide RNA” refers to a short RNA molecule containing a scaffold sequence suitable for targeting a nuclease (e.g., a Cas enzyme such as Cas9 or Cpf1 or another ribonucleoprotein with similar properties), the targeted nuclease binding to a substantially target-specific sequence that facilitates the cleavage of a specific region of DNA or RNA.
[0051] Mutation: As used herein, the term “mutation” refers to an alteration of a nucleic acid sequence or structure relative to a reference sequence. Mutations in polynucleotide sequences can include point mutations (e.g., single-base mutations), polynucleotide mutations, nucleotide deletions, sequence rearrangements, nucleotide insertions, and duplications of DNA sequences in a sample, as well as complex polynucleotide variations. Mutations can occur on both strands of a double-stranded DNA molecule as complementary base changes (i.e., true mutations) or as mutations on one strand but not the other (i.e., heteroduplex mutations), which can potentially be repaired, disrupted, or incorrectly repaired / converted into true double-stranded mutations. The reference sequence can exist in a database (i.e., the HG38 human reference genome) or as a sequence from another sample to which the sequence is compared. Mutations are also referred to as genetic variants.
[0052] Nucleic acid: As used herein, in its broadest sense, means any compound and / or substance incorporated into or potentially incorporated into an oligonucleotide chain. In some embodiments, nucleic acid is a compound and / or substance incorporated into or potentially incorporated into an oligonucleotide chain via phosphodiester bonds. As will be clear from the context, in some embodiments, “nucleic acid” refers to a single nucleic acid residue (e.g., a nucleotide and / or nucleoside); in some embodiments, “nucleic acid” refers to an oligonucleotide chain comprising a single nucleic acid residue. In some embodiments, “nucleic acid” is or includes RNA; in some embodiments, “nucleic acid” is or includes DNA. In some embodiments, nucleic acid is one or more natural nucleic acid residues, including or consisting of them. In some embodiments, nucleic acid is one or more nucleic acid analogs, including or consisting of them. In some embodiments, nucleic acid analogs differ from nucleic acids because they do not utilize a phosphodiester backbone. For example, in some embodiments, nucleic acid is one or more “peptide nucleic acids,” including or consisting of them, which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, and are considered to be within the scope of the present invention. Alternatively or additionally, in some embodiments, the nucleic acid has one or more thiophosphate and / or 5'-N-phosphamide bonds, rather than phosphodiester bonds. In some embodiments, the nucleic acid is, comprises, or is composed of one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine). In some embodiments, the nucleic acid is, includes, or consists of one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiopyrimidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, C-5-propynyl-cytidine, C-5-propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazoadenosine, 7-deazoguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, the nucleic acid includes one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, hexose, or locked nucleic acids) compared to the nucleic acids in naturally occurring nucleic acids. In some embodiments, the nucleic acid has a nucleotide sequence encoding a functional gene product, such as RNA or a protein. In some embodiments, the nucleic acid contains one or more introns. In some embodiments, the nucleic acid may be a non-protein-coding RNA product, such as microRNA, ribosomal RNA, or CRISPR / Cas9 guide RNA. In some embodiments, the nucleic acid plays a regulatory role in the genome. In some embodiments, the nucleic acid is not derived from the genome. In some embodiments, the nucleic acid comprises intergenetic sequences.In some embodiments, nucleic acids are derived from extrachromosomal elements or non-nuclear genomes (mitochondria, chloroplasts, etc.). In some embodiments, nucleic acids are prepared by one or more of the following methods: isolation from natural sources, enzymatic synthesis (in vivo or in vitro) via complementary template-based polymerization, replication in recombinant cells or systems, and chemical synthesis. In some embodiments, the nucleic acid is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues in length. In some embodiments, the nucleic acid is partially or entirely single-stranded; in some embodiments, the nucleic acid is partially or entirely double-stranded. In some embodiments, the nucleic acid has an elemental nucleotide sequence comprising at least one encoding polypeptide, or a complement of a sequence encoding a polypeptide. In some embodiments, the nucleic acid has enzymatic activity. In some embodiments, the nucleic acid performs a mechanical function, for example, in a ribonucleoprotein complex or transfer RNA. In some embodiments, the nucleic acid acts as an adaptor. In some embodiments, the nucleic acid can be used for data storage. In some embodiments, the nucleic acid can be chemically synthesized in vitro.
[0053] References: As used herein, standards or controls are described relative to which comparisons are made. For example, in some embodiments, the agent, animal, individual, population, sample, sequence, or value of interest is compared with a reference or control agent, animal, individual, population, sample, sequence, or value. In some embodiments, a reference or control is tested and / or determined substantially simultaneously with the test or determination of interest. In some embodiments, the reference or control is a historical reference or control, optionally contained in a tangible medium. Generally, as those skilled in the art will understand, the reference or control is determined or characterized under conditions or conditions comparable to those of the condition or environment being evaluated. Those skilled in the art will understand when sufficient similarity exists to justify dependence on and / or comparison with a particular possible reference or control.
[0054] Sequence read: As used herein, the term "sequence read" or "sequencing read" refers to nucleic acid sequence data corresponding to a reference nucleic acid molecule or a target nucleic acid molecule. In some aspects, the data is an inferred sequence of base pairs (or base pair probabilities) corresponding to all or part (e.g., fragments or portions) of a reference nucleic acid molecule or target nucleic acid molecule processed by a sequencing platform. Sequence read lengths can range from several base pairs (bp) to hundreds of kilobases (kb). Sequence read lengths can be affected by the size or length of the reference or target nucleic acid molecule and the sequencing platform used. In some aspects, sequence reads are generated using sequencing technologies, such as, but not limited to, next-generation sequencing platforms, for example, Oxford Nanopore sequencing systems Ion Sequencing system, Roche 454GS IlluminaGenome Applied Biosystems SOLiD Helicos Complete and Pacific Biosciences
[0055] Single-molecule identifier (SMI): As used herein, the term "single-molecule identifier" or "SMI" (which may be referred to as a "tag," "barcode," "molecular barcode," "unique molecular identifier," or "UMI," and other names) refers to any material (e.g., a nucleotide sequence, nucleic acid molecule feature) capable of distinguishing a single molecule within a large population of heterogeneous molecules. In some embodiments, an SMI may be or include an exogenously applied SMI. In some embodiments, an exogenously applied SMI may be or include a degenerate or semi-degenerate sequence. In some embodiments, a substantially degenerate SMI may be referred to as a random unique molecular identifier (R-UMI). In some embodiments, an SMI may include a code from a known code pool (e.g., a nucleic acid sequence). In some embodiments, a predefined SMI code is referred to as a defined unique molecular identifier (DUMI). In some embodiments, an SMI may be or include an endogenous SMI. In some embodiments, an endogenous SMI may be or include information associated with a specific cleavage site of a target sequence or with a feature associated with the ends of a single molecule including the target sequence. In some embodiments, an SMI may relate to sequence variations in a nucleic acid molecule caused by random or semi-random damage, chemical modification, enzymatic modification, or other modifications to the nucleic acid molecule. In some embodiments, the modification may be deamination of methylcytosine. In some embodiments, the modification may require a site for nucleic acid cleavage. In some embodiments, the SMI may include both exogenous and endogenous elements. In some embodiments, the SMI may include physically adjacent SMI elements. In some embodiments, the SMI elements may be spatially distinct within the molecule. In some embodiments, the SMI may be non-nucleic acid. In some embodiments, the SMI may include two or more different types of SMI information. Various embodiments of the SMI are further disclosed in International Patent Publication No. WO2017 / 100441, which is incorporated herein by reference in its entirety.
[0056] Chain-Defining Element (SDE): As used herein, the term "chain-defining element" or "SDE" refers to any material that allows the identification of a specific strand of a double-stranded nucleic acid material and thus distinguishes it from another / complementary strand (e.g., any material that, after sequencing or other nucleic acid interrogation, makes the amplified products of each of the two single-stranded nucleic acids produced from the target double-stranded nucleic acid substantially distinguishable from each other). In some embodiments, the SDE may be or include one or more fragments of substantially non-complementary sequences in the adaptor sequence. In certain embodiments, fragments of substantially non-complementary sequences in the adaptor sequence may be provided by adaptor molecules comprising Y-shaped or "loop" shapes. In other embodiments, fragments of substantially non-complementary sequences in the adaptor sequence may form unpaired "bubbles" in the middle of adjacent complementary sequences in the adaptor sequence. In other embodiments, the SDE may cover nucleic acid modifications. In some embodiments, the SDE may include physically separated paired strands into physically separated reaction chambers. In some embodiments, the SDE may include chemical modifications. In some embodiments, the SDE may include modified nucleic acids. In some embodiments, the SDE may involve sequence variations in nucleic acid molecules caused by random or semi-random damage to the nucleic acid molecule, chemical modifications, enzymatic modifications, or other modifications. In some embodiments, the modification may be the deamination of methylcytosine. In some embodiments, the modification may require a cleavage site. Various embodiments of SDE are further disclosed in International Patent Publication No. WO2017 / 100441, which is incorporated herein by reference in its entirety.
[0057] Subject: As used herein, the term "subject" refers to an organism, typically a mammal (e.g., a human, including prenatal human forms in some embodiments). In some embodiments, the subject suffers from a relevant disease, condition, or symptom. In some embodiments, the subject is susceptible to a disease, condition, or symptom. In some embodiments, the subject exhibits one or more symptoms or characteristics of a disease, condition, or symptom. In some embodiments, the subject does not exhibit any symptoms or characteristics of a disease, condition, or symptom. In some embodiments, the subject is a person having one or more characteristics of susceptibility or risk to a disease, condition, or symptom. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual to whom and / or to whom a diagnosis and / or therapy has been administered.
[0058] Essentially: As used herein, the term “essentially” refers to a qualitative condition that exhibits all or nearly all of the range or degree of the feature or property of interest. Those skilled in the art of biology will understand that biological and chemical phenomena rarely (if any) reach completion and / or proceed to completion or achieve or avoid an absolute result. Therefore, the term “essentially” is used herein to capture the inherent completeness potentially lacking in many biological and chemical phenomena.
[0059] Variants: As used herein, the term "variant" refers to an entity that exhibits significant structural identity with a reference entity but is structurally different from the reference entity in the presence or level of one or more chemical motifs compared to the reference entity. In the context of nucleic acids, variant nucleic acids may have characteristic sequence elements comprising multiple nucleotide residues that occupy designated positions relative to another nucleic acid in linear or three-dimensional space. Homologous sequences differ due to one or more variants. For example, a variant polynucleotide (e.g., DNA) may differ from a reference polynucleotide due to one or more differences in the nucleic acid sequence. In some embodiments, the variant polynucleotide sequence contains insertions, deletions, substitutions, or mutations relative to another sequence (e.g., a reference sequence in a sample or other polynucleotide (e.g., DNA) sequence). Examples of variants include SNPs, SNVs, CNVs, CNPs, MNVs, MNPs, mutations, cancer mutations, driver mutations, passenger mutations, and genetic polymorphisms. Detailed Implementation
[0060] This invention generally relates to methods for providing error-corrected sequence reads of nucleic acid materials using double-stranded sequencing and related reagents for such methods. Some embodiments of the technology relate to methods for achieving high-accuracy sequencing reads and generating increased desired data at a faster rate (e.g., with fewer steps) and / or at less cost (e.g., using fewer reagents). Other aspects of the technology relate to methods and reagents for improving the conversion efficiency of double-stranded sequencing (i.e., the proportion of nucleic acid molecules that produce the sequence). Various aspects of this invention have numerous applications in preclinical and clinical testing and diagnostics, as well as other applications.
[0061] The following text and references Figure 1A-12C Specific details of several embodiments of the present invention are described herein. Although many embodiments of double-stranded sequencing are described herein, other sequencing methods capable of generating error-corrected sequencing reads and other sequencing methods for providing sequence information, in addition to those described herein, are also within the scope of the present invention. Furthermore, the configurations, components, and procedures of other embodiments of the present invention may differ from those described herein. Therefore, those skilled in the art will understand that the present invention may include other embodiments with additional elements, and the present invention may include embodiments not referred to below. Figure 1A-12C Other embodiments of several features shown and described.
[0062] Regarding the efficiency of double-stranded sequencing processes or other high-accuracy sequencing methods, conversion efficiency can be defined as the fraction of unique nucleic acid molecules input into the sequencing library preparation reaction that yields at least one double-stranded common sequence read (or other high-accuracy sequence read) from the sequencing library preparation reaction. In some cases, insufficient conversion efficiency may limit the practicality of high-accuracy double-stranded sequencing in applications that it would otherwise be well-suited for. For example, low conversion efficiency will result in a limited copy number of the target double-stranded nucleic acid, which may lead to less sequence information produced than expected. Non-limiting examples of this concept include DNA from circulating tumor cells or cell-free DNA derived from tumors, or DNA shed into bodily fluids such as plasma and mixed with excess DNA from other tissues in prenatal infants. Other non-limiting examples include limited quantities of forensic material left at a crime scene, ancient DNA that may be found at archaeological sites, very small biopsies obtained by needle biopsy, aspiration, or endoscopy, small amounts of formalin-fixed clinical material, microdissected samples, samples from small biological regions or human or non-human objects, samples or hair, bloodstains, or a limited number of other biological materials produced or derived from multicellular or single-celled organisms, containing single cells or a small number of cells. Although the typical accuracy of double-stranded sequencing is capable of resolving one mutated molecule among more than 100,000 unmutated molecules, the lowest measurable mutation frequency is 1 / (10,000) if, for example, only 10,000 molecules are available in the sample (e.g., 10,000 genomic equivalents in the case of a single-copy gene or locus), and even if these are converted to double-stranded shared sequence reads with an ideal efficiency of 100%, the lowest measurable mutation frequency is 1 / (10,000). * (100%) = 1 / 10,000. In clinical diagnosis, maximum sensitivity for detecting low-level cancer signals or treatment- or diagnostically relevant mutations is often important, and therefore, relatively low conversion efficiency would be undesirable in this context. Similarly, in forensic applications, the amount of DNA available for testing is typically limited. Maximum conversion efficiency may be crucial for detecting the presence of DNA from all individuals within a mixture when only nanograms or picograms can be recovered from a crime scene or natural disaster site, and / or when DNA from multiple individuals is mixed together.
[0063] Methods incorporating double-stranded sequencing and other sequencing approaches may involve ligating (e.g., linking) one or more sequencing adaptors to a target double-stranded nucleic acid molecule to generate a double-stranded target nucleic acid complex. Such adaptor molecules may contain one or more features suitable for massively parallel sequencing platforms, such as sequencing primer recognition sites, amplification primer recognition sites, barcode (e.g., single-molecule identifier (SMI)) sequences (also known as unique molecular identifiers (UMI)), index sequences, single-stranded portions, double-stranded portions, strand-distinguishing elements, or features, etc. As discussed above, to obtain double-stranded sequencing information, sequence information needs to be successfully recovered from the two strands of the original double-stranded molecule. Aspects of this disclosure provide methods and reagents for generating and associating sequencing information from the two strands of an original double-stranded molecule by physically linking the strands prior to amplification and sequencing.
[0064] I. Selected Examples of Double-Strand Sequencing Methods and Related Interchangeants and Reagents
[0065] Double-stranded sequencing is a method for generating error-corrected DNA sequences from double-stranded nucleic acid molecules, and was originally described in International Patent Publication No. WO 2013 / 142389 and U.S. Patent No. 9,752,188, both of which are incorporated herein by reference in their entirety. In some aspects of the technique, double-stranded sequencing can be used to sequence both strands of individual DNA molecules in a manner that, during massively parallel sequencing (MPS) (also commonly referred to as next-generation sequencing (NGS)), derived sequence reads can be identified as originating from the same double-stranded nucleic acid parent molecule, but also as distinguishable entities after sequencing. The resulting sequence reads from each strand are then compared to obtain the error-corrected sequence of the original double-stranded nucleic acid molecule.
[0066] Figure 1 is a conceptual illustration of various double-stranded sequencing method steps according to embodiments of the present invention. In some embodiments, the method of incorporating double-stranded sequencing may include ligating one or more sequencing adaptors to a plurality of target double-stranded nucleic acid molecules, each target double-stranded nucleic acid molecule including a first-stranded target nucleic acid sequence and a second-stranded target nucleic acid sequence to generate a plurality of double-stranded target nucleic acid complexes. Figure 1A Once a formulation of a double-stranded nucleic acid library is formed, the complex can be amplified using DNA amplification methods such as PCR or any other DNA amplification biochemical method (e.g., rolling circle amplification, multiple displacement amplification, isothermal amplification, bridge amplification, polyclonal amplification, isothermal amplification, or surface binding amplification), resulting in one or more copies of the first-strand target nucleic acid sequence and one or more copies of the second-strand target nucleic acid sequence (e.g., Figure 1A Then, DNA sequencing can be performed on one or more amplified copies of the first-strand target nucleic acid molecule and one or more amplified copies of the second target nucleic acid molecule, preferably using a "next-generation" massively parallel DNA sequencing platform (e.g., Figure 1A ).
[0067] After sequencing, the sequence reads generated from the first strand of the target nucleic acid molecule are compared with the sequence reads generated from the second strand of the same target nucleic acid molecule. In some embodiments, more than one sequence read can be generated from the first and second strands. Once compared, the error-corrected target nucleic acid molecule sequence can be generated (e.g., Figure 1B For example, nucleotide positions where the bases from the first-strand target nucleic acid sequence and the second-strand target nucleic acid sequence are identical are considered true sequences, while nucleotide positions where the bases are inconsistent between the two strands are considered potential sites of technical error, which can be ignored, eliminated, corrected, or otherwise identified. In some embodiments, when nucleotide positions are inconsistent, the site can be identified as unknown (e.g., in...). Figure 1B (shown as "N" in the diagram). Therefore, error-corrected sequences of the original double-stranded target nucleic acid molecule can be generated (in... Figure 1B (As shown in the figure). Optionally, and in some embodiments, and after each sequencing read generated from the first-strand target nucleic acid molecule and the second-strand target nucleic acid molecule is grouped separately, a single-stranded common sequence can be generated for each of the first and second strands. The single-stranded common sequences from the first-strand target nucleic acid molecule and the second-strand target nucleic acid molecule can then be compared to generate an error-corrected target nucleic acid molecule sequence (e.g., Figure 1B ).
[0068] Alternatively, in some embodiments, the sequence inconsistency site between the two strands can be identified as a potential site of a bio-derived mismatch in the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, the sequence inconsistency site between the two strands can be identified as a potential site of a mismatch originating from DNA synthesis in the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, the sequence inconsistency site between the two strands can be identified as a potential site in which a damaged or modified nucleotide base is present on one or both strands and is converted into a mismatch by an enzymatic process (e.g., DNA polymerase, DNA glycosylase, or another nucleic acid modifying enzyme or chemical process). In some embodiments, the modified nucleotide base is 5-methyl-cytosine, 8-oxo-guanine, a ribose base, a baseless nucleotide, or a uracil nucleotide. In some embodiments, this later finding can be used to infer the presence of nucleic acid damage or nucleotide modification prior to the enzymatic process or chemical treatment.
[0069] In some embodiments, and as described in U.S. Patent No. 9,752,188 and International Patent Publication No. WO2017 / 100441, the following can be used to associate (e.g., group) first-strand sequencing reads and second-strand sequencing reads from a single original double-stranded nucleic acid molecule: (a) a single-molecule identifier (SMI) sequence associated with the adaptor during library preparation; (b) fragment features associated with the original double-stranded molecule, such as sequences located at or near or relative to the fragment ends; and (c) combinations thereof.
[0070] In one embodiment, the generation of raw sequence reads for double-stranded sequencing reflects the purpose of the target double-stranded nucleic acid molecule, wherein a hairpin adaptor is attached to one end of the molecule and a “Y”-shaped adaptor is attached to the other end. This ligation or double-stranded complex, comprising the first and second strands of the raw double-stranded nucleic acid molecule, can be further amplified using any type of amplification (e.g., PCR or bridging) and can then be subjected to massively parallel sequencing (e.g., sequencing by synthesis, next-generation sequencing (NGS), etc.) to generate sequence reads for double-stranded sequencing. In a non-limiting example, an adaptor double-stranded nucleic acid complex with a hairpin adaptor (i.e., a “loop” or “U” shape) allows for the generation of sequence reads from the raw first and second strands of the target double-stranded nucleic acid molecule in such a way that the sequence reads are grouped according to the nature of the sequencing reaction position on the flow cell surface (if by sequencing by synthesis) or otherwise at the position of the sequencing reaction / process.
[0071] Various aspects of this invention relate to methods and reagents for associating and / or grouping first-strand and second-strand sequencing reads by physically linking the first and second strands in a manner that makes sequencing information derived from both strands correlated with each other (e.g., for error correction). In some embodiments, methods for preparing sequencing libraries for use in double-strand sequencing may include attaching a hairpin adaptor to one end of a target double-stranded nucleic acid molecule and attaching a Y-shaped adapter to the opposite end of the same target double-stranded nucleic acid molecule. In one embodiment, the hairpin adaptor molecule includes cleavable hairpin adaptor elements for targeted separation of the first and second strands of the target double-stranded nucleic acid molecule.
[0072] In some embodiments, the association of first-strand and second-strand sequencing reads can be performed during or after the sequencing reaction on a sequencer. For example, in some embodiments, the first and second strands of a double-stranded nucleic acid molecule are joined by an intermediate linker domain, such as a hairpin adaptor sequence. In one embodiment, sequence information from the two strands derived from the original nucleic acid molecule is generated within the same clonal cluster on an MPS sequencer (e.g., a flow cell). Sequencing the joined first and second strands on a sequencer presents challenges because self-complementary hairpin sequences can preferentially hybridize on the sequencing surface or in solution, thereby weakening polymerase extension. Certain aspects of the present invention disclose methods for overcoming these challenges associated with self-complementary hybridization of the joined first and second strands, while being able to obtain sequencing reads from the first and second strands within the same clonal cluster on a sequencer.
[0073] Connector and Connector Sequence
[0074] In various permutations, adaptor molecules comprising primer sites, flow cell sequences, and / or other features such as SMI (e.g., molecular barcodes) or SDE are envisioned for use in many embodiments disclosed herein. In some embodiments, the provided adaptor may be or comprise one or more sequences complementary to or at least partially complementary to PCR primers (e.g., primer sites) having at least one of the following properties: 1) high target specificity; 2) multiplexing capability; and 3) exhibiting robust and minimally biased amplification.
[0075] In some embodiments, the adaptor molecule may be Y-shaped, U-shaped, hairpin-shaped, have bubbles (e.g., part of a non-complementary sequence), or have other features. In other embodiments, the adaptor molecule may include Y-shaped, U-shaped, hairpin-shaped, or bubble-shaped elements. For the purposes of this disclosure, both U-shaped and hairpin-shaped adaptors can be used collectively as adaptors having a linker domain that links (or connects) the first strand of a target double-stranded nucleic acid molecule to the second strand of the same molecule. Some adaptors may include modified or non-standard nucleotides, restriction sites, or other features for in vitro manipulation of structure or function. The adaptor molecule can be linked to a variety of terminal nucleic acid materials. For example, the adaptor molecule may be adapted to attach to a T-terminal, A-terminal, CG-terminal, polynucleotide-terminal (also referred to herein as a “sticky end” or “sticky protrusion”) or a single-stranded protrusion with a known nucleotide length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides), a dehydroxylated base, a blunt end of nucleic acid material, or the end of a molecule, wherein the 5' end of the target is dephosphorylated or otherwise blocked from conventional linking. In other embodiments, the adaptor molecule may contain dephosphorylated or otherwise link-preventing modifications on the 5' strand at the linking site. In the latter two embodiments, such strategies can be used to prevent dimerization of the library fragment or the adaptor molecule.
[0076] Figure 2A Nucleic acid adaptor molecules for use with embodiments of the present invention are shown, as well as double-stranded adaptor nucleic acid complexes generated by the linkage of the adaptor molecules with double-stranded nucleic acid fragments according to embodiments of the present invention. Figure 2A As shown, the first adaptor molecule (adapter 1) can be a Y-shaped adaptor molecule having a first primer site and a second primer site (labeled primer site 1 and primer site 2), and is adapted to ligate to a double-stranded nucleic acid fragment via a T-protrusion. The second adaptor molecule (adapter 2) adapted to ligate to the target nucleic acid fragment via a T-protrusion is shown as a hairpin adaptor including a single-stranded linker domain. Sequencing library generation from a population of double-stranded nucleic acid fragments can include ligating a pool of adaptors, including both adaptor 1 and adaptor 2, to a population of double-stranded nucleic acid fragments. Figure 2A This illustration shows one product obtained from the ligation reaction described herein. Other products will comprise an adaptor nucleic acid complex comprising adaptor 1 at both ends and an adaptor nucleic acid complex comprising adaptor 2 at both ends. In the various embodiments described herein, it is desirable to produce products such as... Figure 2A The adapter nucleic acid complex shown is intended for use with double-stranded sequencing methods.
[0077] Figure 2B Another embodiment is shown, wherein the target double-stranded nucleic acid fragment includes a sticky end 1 at one end of the fragment and a sticky end 2 at the opposite end of the fragment. By design, the sequence of sticky end 1 (the 5' protrusion of the target fragment) is known. Similarly, the sequence of sticky end 2 (the 3' protrusion of the target fragment) is known. In one embodiment, the sequence of sticky end 1 differs from the sequence of sticky end 2. In another embodiment, the sequence lengths of sticky end 1 and sticky end 2 are different. In yet another embodiment, sticky end 1 is a 5' protrusion, and sticky end 2 is a 3' protrusion. Specific adaptors comprising substantially complementary sequences can be synthesized, allowing the fragment to be attached to the adaptors at both ends. In one embodiment, the adaptors can be different (e.g., adaptor 1 may include a Y-shape and adaptor 2 may include a U-shape). In other embodiments (not shown), the adaptors can be of the same type (e.g., including Y-shaped, U-shaped, barcode-shaped, etc.). Figure 2B As demonstrated, this design allows each target double-stranded nucleic acid molecule to have a Y-shaped adaptor at one end and a hairpin (e.g., an adaptor with a linker domain) at the other end. Thus, upon denaturation, the adaptor nucleic acid complex comprises a single-stranded molecule including a first primer site, a first strand, a linker domain, a second strand, and a second primer site. In other applications, designing specific adaptors to locate at the 5' or 3' end of the fragment may be advantageous. The specificity of the essentially unique sticky ends on the target fragment facilitates these types of applications. Furthermore, positive selection of the target fragment by successful cleavage and adaptor ligation ensures that only the target-enriched nucleic acid region is amplified and sequenced.
[0078] Therefore, in some embodiments, the connector molecule set may include a different, unique, or semi-unique adhesive protrusion relative to other connector molecule sets. The number of adhesive protrusions of different types may be 2 or 3, 4, 5, 6, 7, 8, 9, or 10 or more. It may be about 11 or 12 or 15 or 20 or 25 or 30 or 35 or 40 or 45 or 50 or 60 or 70 or 80 or 90 or 100 or 120 or 140 or 150 or 200 or 300 or 400 or 500 or 750 or 1000 or more. In a particular example, a hairpin connector molecule may include a first adhesive protrusion adapted to attach to a first complementary segment adhesive end, and a Y-shaped connector may include a second adhesive protrusion adapted to attach to a second complementary segment adhesive end. Thus, the preparation of sequencing libraries for a population of nucleic acid molecules may include generating nucleic acid fragments with first and second sticky ends, and ligating the nucleic acid fragments to hairpins and Y-shaped adaptors. The resulting sequencing library may include multiple double-stranded adaptor nucleic acid fragment complexes, each complex having a hairpin adaptor at the first end and a Y-shaped adaptor at the second end.
[0079] Amplification
[0080] In one embodiment, the method may include amplifying an adaptor nucleic acid complex comprising both a first strand and a second strand on a sequencer surface, such as a flow cell surface. In some embodiments, amplification on a surface, such as bridging amplification on a flow cell surface, comprises generating clusters or multiple copies of a binding nucleic acid template. In a particular embodiment, the coupled first and second strand nucleic acid templates may be bridge-amplified on a flow cell surface, for example, to generate multiple clonal clusters, wherein each clonal cluster comprises a copy of the original first and second strand nucleic acid template derived from the original double-stranded nucleic acid molecule. Some of the clonal copies in the cluster will be forward-oriented, while the remainder will be reverse-oriented. Those skilled in the art will understand that various embodiments using amplification, such as polyclonal amplification, cluster amplification, bridging amplification, etc., include the step of flowing the adaptor nucleic acid complex through a surface that provides a binding oligonucleotide at least partially complementary to the region of the Y-shaped adaptor. The surface may provide one or more oligonucleotides partially complementary to the adaptor. In practice, both arms of the Y-shaped adaptor may hybridize with the surface of the flow cell.
[0081] Bridge amplification (not shown) can be used to generate multiple copies of the complex to form colonies or clusters (also referred to herein as clonal clusters). Each clonal cluster comprises the multiple copies derived from the original molecule (e.g., the adaptor nucleic acid complex) in both forward and reverse orientations.
[0082] In one embodiment, a sequencing reaction can be performed when a copy in the forward orientation or a copy in the reverse orientation is cut and removed. Figure 3A The steps in the process are illustrated after bridging amplification of the adaptor nucleic acid complex (e.g., a double-stranded nucleic acid complex) and after the cleavage and removal of the forward-oriented copy (e.g., where nucleic acid sequence "2" binds to the surface of the flow cell). Figure 3A As shown, the remaining complex is in reverse orientation (e.g., where nucleic acid sequence "1" binds to the surface of the flow cell; e.g., the 3' end of the molecule binds to the surface). In one embodiment, the nucleic acid sequence of the first strand readily hybridizes to the complementary nucleic acid sequence of the second strand, making sequencing by synthesizing longer complexes difficult. The binding copy of the complex shown includes a hairpin adaptor (e.g., adaptor 2, ...). Figure 2A and 2B The adapter domain provided by the adapter. In some embodiments, the adapter domain includes a cleavable site or motif (“C”). The cleavable site C may include a nucleotide sequence, a single nucleotide base, a modified base, or other enzymatically or non-enzymatically cleavable features.
[0083] like Figure 3B As shown, the process may include a step of cleavage at a cleavable site C to separate the first strand sequence from the second strand sequence. In one embodiment, the cleavage event at site C may be facilitated by a cleavage promoter (e.g., an enzyme, a chemical, etc.). In one embodiment, the cleavage step may be inefficient, such that only a portion of the complex is cleaved at site C. Thus, a portion of the complex (e.g., about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 45%, about 50% or more or less; about 1% to about 10%; about 10% to about 25%; about 25% to about 45%; greater than about 50%, less than about 10%) may remain uncleavage, and the first strand sequence and the second strand sequence remain connected. In some respects, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the complex is cleaved, for example, at site C.
[0084] After separating the first and second strands by cleavage at site C, unbound strands (e.g., adjacent nucleic acid sequences 2) are washed away. For example, as... Figure 3CAs shown, the portion of the complex cleaved at site C comprises only the nucleotide sequence of the first strand and a portion of the hairpin adaptor. Since the complex will no longer self-hybridize, sequencing reactions using primers specific to the adaptor (e.g., binding to the 3' end of the molecule at or near nucleotide sequence 1) can be performed to generate sequencing reads of the remaining first strand in the clonal cluster. Figure 3D Index reads (not shown) can also be generated. Note that the sequencing reads of the first strand are single-end sequence reads. Uncut complexes in clonal clusters remain self-hybridized and are likely to fail to sequence successfully during the sequencing reaction due to the difficulty in replacing the longer second strand with sequencing primers. Figure 3D ).
[0085] After obtaining sequencing information from the first strand present in the clonal cluster, the next step in the process involves a second round of amplification (e.g., bridging amplification) to provide additional copies of the uncut complex. Bridging amplification requires nucleic acid sequence 1 and nucleic acid sequence 2 present in the full-length complex. Only the remaining uncut complex has the two still-present adaptor sequences. Thus, the clonal cluster can be reproduced via bridging amplification using the remaining oligonucleotides bound to the flow cell surface. Figure 4A ).
[0086] After amplification, a second sequencing reaction can be performed when the reverse-oriented copy is cut and removed. Figure 4B The steps in the process are illustrated after bridging amplification of the adaptor nucleic acid complex (e.g., a double-stranded nucleic acid complex) and after the reverse-oriented copy (e.g., where the nucleic acid sequence "1" binds to the surface of the flow cell) is cleaved and removed. Figure 4B As shown, the remaining complex is in the forward orientation (e.g., where nucleic acid sequence "2" is bound to the surface of the flow cell; e.g., where the 5' end of the molecule is bound to the surface). As described above, the nucleic acid sequences of the first and second strands readily hybridize, making sequencing by synthesizing longer complexes difficult.
[0087] like Figure 4CAs shown, the process may include a step of cleavage at a cleavable site C to separate the second strand sequence from the first strand sequence. In one embodiment, the cleavage event at site C may be promoted by a cleavage accelerator (e.g., an enzyme, a chemical, etc.). As discussed above, the cleavage step may be inefficient, such that only a portion of the complex is cleaved and located at site C. Thus, a portion of the complex (e.g., about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 45%, about 50% or more or less; about 1% to about 10%; about 10% to about 25%; about 25% to about 45%; greater than about 50%, less than about 10%) may remain uncleaved, and the first and second strand sequences remain connected. Alternatively, the cleavage step may be efficient and may cleave all of the complex (e.g., as shown in the diagram). Figure 4C (As shown).
[0088] After separating the second strand from the first strand by cleaving at site C, unbound strands (e.g., adjacent nucleic acid sequence 1) are washed away. For example, as... Figure 4D As shown, the portion of the complex cleaved at site C comprises only the second-strand nucleotide sequence and a portion of the hairpin adaptor. Since the complex will no longer self-hybridize, sequencing reactions using primers specific to the remaining portion of the hairpin adaptor can be performed to generate sequencing reads of the remaining second strand in the clonal cluster. Figure 4E Index reads (not shown) can also be generated. Note that the sequencing reads from the second strand are single-end sequence reads. Once sequence reads originating from both the first and second strands (e.g., within the same clonal cluster) are generated, they can be compared for error correction.
[0089] Figures 5A-5E Another embodiment for sequencing double-stranded complexes that provide double-stranded sequencing information on a sequencing surface (e.g., a flow cell) is illustrated. Figures 5A-5E In the illustrated embodiments, sequence reads from the first and second strands of the original adaptor nucleic acid complex can be generated without a second bridging amplification step. As discussed above, each double-stranded complex can be independently bridged amplified on its surface to generate a clonal cluster comprising multiple copies of the double-stranded complex having both a first strand and a complementary second strand, wherein the intermediate hairpin adapter domain has a cleavable site ( Figure 5A As discussed above, copying can be forward-oriented or reverse-oriented.
[0090] like Figure 5BAs shown, and in one embodiment, the double-stranded complex can be cleaved at cleavage site C (e.g., by a cleavage accelerator discussed further herein). Following cleavage at site C, the non-binding strand is removed. Reference Figure 5C The remaining molecules that bind to the surface of the flow cell contain (a) a first strand sequence in reverse orientation (e.g., adjacent to primer site "1") and (b) a second strand sequence in forward orientation (e.g., adjacent to primer site "2").
[0091] In the next step, a first sequencing reaction using primers specific for reverse orientation is used to obtain sequencing information for the first strand. Figure 5D The primers used in the first sequencing reaction can be washed away. In the next step, a second sequencing reaction using primers specific for forward orientation is used to obtain sequencing information for the second strand. Figure 5E ). Figure 5D and 5E The illustrated embodiments demonstrate sequential sequencing of the first and second strands. It will be understood that in another embodiment, the first and second strands may be sequenced simultaneously (e.g., in the same sequencing reaction) using, for example, multicolor chemistry (e.g., 4-color chemistry) followed by deconvolution of the sequencing / color frequency signals to determine the source of a particular sequencer base response or signal.
[0092] Once sequencing reads from the first and second strands are generated, the first-strand sequencing reads can be compared with the second-strand sequencing reads to provide double-strand error correction. The embodiments described herein overcome some of the challenges associated with the aforementioned transformation efficiency because the sequencing information from each clonal cluster provides both first-strand and second-strand sequencing reads.
[0093] II. Examples of methods and reagents for cutting hairpin connectors.
[0094] Typically, sequencing of hairpin-linked adaptor nucleic acid complexes can be challenging because the polymerase must replace the self-complementary hybridization region. For example, polymerase-based sequencing of such structures remains an obstacle to providing physically linked double-stranded sequencing data due to the close proximity of the self-complementary regions of the adaptor nucleic acid complex and the high melting temperatures (Tm) of the complementary regions of the first and second strands.
[0095] As discussed above, various aspects of the present invention incorporate the use of hairpin adaptors with cleavable sites or motifs, enabling the first-strand nucleic acid sequence and the second-strand nucleic acid sequence to be separated from each other during the sequencing reaction.
[0096] In some embodiments, and as Figure 6As shown, the hairpin adaptor may (e.g., in the single-stranded or double-stranded portion) include a cleavage motif that allows subsequent cleavage of the hairpin DNA molecule by an enzyme (e.g., an endonuclease) or other cleavage accelerator (chemical or non-enzymatic process). Reference Figure 7 Furthermore, in one embodiment, the single strand of the hairpin adapter (e.g., the connector region) can be cleaved using a nuclease (e.g., restriction site nuclease, target nuclease, etc.). For example, Figure 7 This demonstrates single-strand cleavage sites (e.g., nucleic acid sequences) that can be digested by endonucleases (e.g., restriction enzymes). Reference Figures 3A-5E Following double-stranded complex bridging amplification, an enzyme can be introduced (e.g., flow through a flow cell) to cleave at the cleavage site. In some embodiments, inefficient cleavage is desired (e.g., some uncleaved double-stranded complex is desired for seeding a second round of bridging amplification). In some embodiments, the enzymatic reaction can be time- or concentration-controlled, such that a portion of the double-stranded complex is cleaved and a portion remains uncleaved. For example, a limited amount of restriction enzyme can be flowed through a functionalized surface to cleave most, but not all, of the hairpin DNA molecules. In another embodiment, the restriction enzyme can flow through the surface for a limited amount of time to cleave most, but not all, of the hairpin DNA molecules. In yet another embodiment, a mixture of enzymes, mostly catalytically active and a small amount non-catalytically active, can flow through a functionalized surface to cleave most, but not all, of the hairpin DNA molecules.
[0097] Figure 8A and 8B Another embodiment is shown that provides a cleavage site in the adapter domain of a hairpin adaptor in a manner that allows for inefficient cleavage of the double-stranded complex in the clonal cluster. In this example, and prior to the introduction of the endonuclease, the method can provide an oligonucleotide that is at least partially complementary to the adapter domain of the hairpin adaptor. Figure 8B As shown, the hybridization of the introduced oligonucleotides can prevent cleavage by endonucleases (e.g., by providing the cleavage-resistant motif "AC"). The double-stranded complex of oligonucleotides that do not exhibit hybridization ( Figure 8A Oligonucleotides remain readily cleaved by endonucleases. The concentration of oligonucleotides supplied to the sequencing flow cell prior to enzymatic cleavage (or concurrently with the introduction of endonucleases) can be scaled up to retain a desired amount of uncut complex within each clonal cluster on the flow cell. For example, a small amount of oligonucleotide sequence containing a cleavage-resistant motif can flow over a functionalized surface, resulting in hybridization of the oligonucleotide sequence with a subset (e.g., a limited amount) of hairpin DNA molecules in each clonal cluster. Figure 8BMost hairpin DNA molecules (containing a cleavage motif within the hairpin) will not hybridize with oligonucleotide sequences containing cleavage-resistant motifs. Thus, most hairpin DNA molecules (not hybridized with oligonucleotide sequences containing cleavage-resistant motifs) can be cleaved at the single-stranded cleavage motif within the hairpin adaptor. Hairpin DNA molecules that hybridize with oligonucleotide sequences containing cleavage-resistant motifs remain uncut by enzymes.
[0098] In one embodiment, the cleavage motif within the hairpin adaptor may be methylated, and the anti-cleavage motif within the oligonucleotide sequence may be unmethylated. An enzyme that cleaves only methylated DNA can then flow through the functionalized surface. In another embodiment, the cleavage motif within the hairpin adaptor may be unmethylated, and the anti-cleavage motif within the oligonucleotide sequence may be methylated. An enzyme that cleaves only unmethylated DNA can then flow through the functionalized surface. In another embodiment, the anti-cleavage motif within the oligonucleotide sequence may be a side chain that prevents the hairpin DNA molecule from being cleaved. In another embodiment, the anti-cleavage motif within the oligonucleotide sequence may be a bulk adduct that prevents the hairpin DNA molecule from being cleaved. In another embodiment, the anti-cleavage motif within the oligonucleotide sequence may be one or more mismatches that prevent the enzyme from cleaving the hairpin DNA molecule. In another embodiment, the anti-cleavage motif may be a baseless site that prevents cleavage. In another embodiment, the anti-cleavage motif may be a nucleotide analog that prevents cleavage. In another embodiment, the anti-cleavage motif may be a peptide-nucleic acid bond that prevents cleavage.
[0099] exist Figures 9A-9B In another embodiment shown, an oligonucleotide comprising a sequence at least partially complementary to the adapter domain of a hairpin adaptor can be provided to hybridize with the adapter domain and form a cleavage site / motif. For example, an endonuclease that recognizes a double-stranded cleavage site can be used to cleave the adapter region comprising the double-stranded region provided by the hybrid oligonucleotide. Figure 9A For example, oligonucleotides can flow across a functionalized surface, causing the oligonucleotide sequence to hybridize with the adapter region of a hairpin adaptor, thereby providing a double-stranded cleavage motif within a portion of the hairpin DNA molecule. Figure 9A In one embodiment, a limited amount of oligonucleotides may be allowed to flow through the functionalized surface to allow hybridization between the oligonucleotide sequence and hairpin DNA molecules, some but not all of them. In another embodiment, the oligonucleotides may flow through the functionalized surface for a limited amount of time to allow hybridization between the oligonucleotide sequence and hairpin DNA molecules, some but not all of them. Hairpin DNA molecules that hybridize with the oligonucleotide sequence, thereby providing a cleavage motif, are cleaved after an endonuclease flows through the functionalized surface. Hairpin DNA molecules that do not hybridize with the oligonucleotide sequence containing the cleavage motif remain uncleaved.
[0100] exist Figures 10A-10B In yet another embodiment shown, an oligonucleotide pool may be provided comprising a sequence at least partially complementary to the adapter domain of a hairpin adaptor for hybridization with the adapter domain. The oligonucleotide pool may contain a subset of oligonucleotides that, upon hybridization, provide cleavage sites / motifs (e.g., for suitable endonucleases). Figure 10A The oligonucleotide pool may also contain a subset of oligonucleotides that, once hybridized, provide cleavage-resistant motifs (and / or prevent cleavage by, for example, disrupting the site recognition of endonucleases). Figure 10B In one example, an oligonucleotide pool can flow through a functionalized surface. Hairpin DNA molecules hybridizing with an oligonucleotide sequence containing a cleavage motif are cleaved, while hairpin DNA molecules hybridizing with an oligonucleotide sequence containing an anti-cleavage motif remain uncleaved. In one embodiment, a subset of the oligonucleotides may be methylated, and a second subset of the oligonucleotides may be unmethylated. In one embodiment, an enzyme that cleaves only methylated DNA can then flow through the functionalized surface. In another embodiment, an enzyme that cleaves only unmethylated DNA can flow through the functionalized surface. In another embodiment, the oligonucleotide providing the anti-cleavage motif may include a side chain that prevents the hairpin DNA molecule from being cleaved. In another embodiment, the anti-cleavage motif within the oligonucleotide sequence may be a bulk adduct that prevents the hairpin DNA molecule from being cleaved. In another embodiment, the anti-cleavage motif within the oligonucleotide sequence may be one or more mismatches that prevent the enzyme from cleaving the hairpin DNA molecule. In another embodiment, the anti-cleavage motif may be a baseless site that prevents cleavage. In another embodiment, the anti-cleavage motif may be a nucleotide analog that prevents cleavage. In another embodiment, the anti-cleavage motif may be a peptide-nucleic acid bond that prevents cleavage. Those skilled in the art will recognize other biochemical means for providing oligonucleotide subsets that will prevent or facilitate the cleavage of selected endonucleases or other enzymes.
[0101] In yet another embodiment, and as Figure 11A and 11B As shown, it can be achieved by using enzymes with partial catalytic activity (stripes); Figure 11A ) and some catalytically inactivated enzymes (black with dots); Figure 11B A mixture of endonucleases is used to achieve inefficient cleavage of clone copies of double-stranded nucleic acid complexes.
[0102] In some embodiments, the restriction endonuclease is or includes a targeted endonuclease. In some embodiments, the targeted endonuclease is or includes at least one of restriction endonucleases (i.e., restriction enzymes) that cleave DNA at or near a recognition site (e.g., EcoRI, BamHI, XbaI, HindIII, AluI, AvaII, BsaJI, BstNI, DsaV, Fnu4HI, HaeIII, MaeIII, N1aIV, NSiI, MspJI, FspEI, NaeI, Bsu36I, NotI, HinF1, Sau3AI, PvuII, SmaI, HgaI, AluI, EcoRV, etc.). A list of several restriction endonucleases is provided in print and computer-readable form and is available from numerous commercial vendors (e.g., New England Biolabs, Ipswich, MA). Those skilled in the art will understand that any restriction endonuclease can be used according to various embodiments of the invention. In other embodiments, the targeted endonuclease is or includes at least one of a ribonucleoprotein complex, such as a CRISPR-associated (Cas) enzyme / guide RNA complex (e.g., Cas9 or Cpf1) or a Cas9-like enzyme. In other embodiments, the targeted endonuclease is or includes a homing endonuclease, a zinc finger nuclease, TALEN and / or a broad range of nucleases (e.g., megaTAL nucleases), an argonaute nuclease, or a combination thereof. In some embodiments, the targeted endonuclease includes Cas9 or CPF1 or derivatives thereof. In another embodiment, the nuclease can cleave at a bifurcated nucleic acid region (e.g., FEN1). In some embodiments, more than one targeted endonuclease may be used (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more).
[0103] In some embodiments, the cleavage site is or includes a user-directed recognition sequence for targeting a nuclease (e.g., CRISPR or a CRISPR-like nuclease) or other tunable nucleases. In some embodiments, cleaving nucleic acid material may include at least one of the following: enzymatic digestion, enzymatic cleavage, one-strand enzymatic cleavage, two-strand enzymatic cleavage, incorporation of modified nucleic acid followed by enzymatic treatment (which results in cleavage of one or two strands), incorporation of replication-blocking nucleotides, incorporation of chain terminators, incorporation of photocleavable adapters, incorporation of uracil, incorporation of ribose bases, incorporation of 8-oxo-guanine adducts, use of restriction endonucleases, use of ribonucleoprotein endonucleases (e.g., Cas enzymes, such as Cas9 or CPF1) or other programmable endonucleases (e.g., homing endonucleases, zinc finger nucleases, TALENs, broad-spectrum nucleases (e.g., megaTAL nucleases), arginine nucleases, etc.), and any combination thereof.
[0104] Targeted endonucleases (e.g., CRISPR-associated ribonucleoprotein complexes such as Cas9 or Cpf1, homing nucleases, zinc finger nucleases, TALENs, megaTAL nucleases, argonaute nucleases, and / or derivatives thereof) can be used to selectively cleave target moieties of nucleic acid materials. In some embodiments, the targeted endonuclease may be modified, such as having amino acid substitutions, to provide, for example, enhanced thermostability, salt tolerance, and / or pH tolerance, or enhanced specificity or alternative PAM site recognition, or higher binding affinity. In other embodiments, the targeted endonuclease may be biotinylated, fused with streptavidin, and / or incorporated into other affinity-based (e.g., decoy / prey) technologies. In some embodiments, the targeted endonuclease may have altered recognition site specificity (e.g., a SpCas9 variant with altered PAM site specificity). In other embodiments, the targeted endonuclease may be catalytically inactive, so that cleavage does not occur once it binds to the target moieties of the nucleic acid material. In some embodiments, a targeted endonuclease is modified to cleave a single strand of a targeted portion of nucleic acid material (e.g., a nicking enzyme variant), thereby creating a nick in the nucleic acid material. CRISPR-based targeted endonucleases are further discussed herein to provide additional detailed, non-limiting examples of the use of targeted endonucleases. It is noted that the nomenclature surrounding such targeted endonucleases is still evolving. For the purposes of this document, the term “CRISPR-based” generally refers to an endonuclease comprising a nucleic acid sequence that can be modified to redefine the nucleic acid sequence to be cleaved. Cas9 and CPF1 are examples of such targeted endonucleases currently in use, but appear to be more prevalent in different parts of nature, and the availability of various varieties of such targeted and easily modifiable endonucleases is expected to grow rapidly in the coming years. For example, Cas12a, Cas13, CasX, etc., are contemplated for use in various embodiments. Similarly, various engineered variants of these enzymes for enhancing or modifying their properties are becoming available. In this document, functionally substantially similar targeted endonucleases not explicitly described or discovered herein are explicitly envisioned for the purpose of achieving similar objectives to those described in the disclosure.
[0105] It is specifically envisioned that any of a variety of restriction endonucleases (i.e., enzymes) can be used. Typically, restriction enzymes are produced by certain bacteria / other prokaryotes and cut at, near, or between specific sequences within a given DNA segment.
[0106] It will be apparent to those skilled in the art that restriction enzymes are selected to cut at a specific site or at a generated site to produce a restriction site for cleavage. In some embodiments, the restriction enzyme is a synthetic enzyme. In some embodiments, the restriction enzyme is not a synthetic enzyme. In some embodiments, the restriction enzyme as used herein has been modified to introduce one or more variations within the genome of the enzyme itself. In some embodiments, the restriction enzyme produces a double-strand cut between defined sequences within a given DNA motif.
[0107] Although any restriction enzyme (e.g., type I, type II, type III, and / or type IV) may be used according to some embodiments, the following represents a non-restrictive list of restriction enzymes that may be used: AluI, ApoI, AspHI, BamHI, BfaI, BsaI, CfrI, DdeI, DpnI, DraI, EcoRI, EcoRII, EcoRV, HaeII, HaeIII, HgaI, HindII, HindIII, HinFI, HPYCH4III, KpnI, MamI, MNL1, MseI, MstI, MstII, NcoI, NdeI, NotI, PacI, PstI, PvuI, PvuII, RcaI, RsaI, SacI, SacII, SalI, Sau3AI, ScaI, SmaI, SpeI, SphI, StuI, TaqI, XbaI, XhoI, XhoII, XmaI, XmaII, and any combination thereof. A broad but not exhaustive list of suitable restriction enzymes can be found in publicly available catalogs and on the Internet (e.g., at the New England Biological Laboratory in Ipswich, Massachusetts). Those skilled in the art will understand that many enzymes, ribozymes, or other nucleic acid-modifying enzymes that can be used alone or in combination to target phosphodiester backbone cleavage of nucleic acid molecules that can achieve the same purpose may not be included in the above list or have not yet been found therein. Many nucleic acid-modifying enzymes can recognize base modifications (e.g., CpG methylation) that can be used to target additional modifications (e.g., for the generation of base-free sites) of adjacent nucleic acid sequences that can be cleaved (e.g., by enzymes with lysin activity). Thus, substantial sequence-specific cleavage can be achieved based on the recognition of DNA or RNA modifications, and this can be used alone or in combination with targeted endonucleases to achieve targeted nucleic acid fragmentation. Other embodiments of cleavage accelerators may include non-enzymatic accelerators. For example, pH changes or hydrolysis can be used to cleave at the cleavage site. Photocleavage is also a method of breaking this backbone. For example, hybridization of modified nucleotides or complementary or partially complementary oligonucleotides with photosensitive portions into hairpin adaptor sequences can generate recognition sites for other chemical or enzymatic processes that will cleave (e.g., upon exposure to light).
[0108] In some embodiments, such as those described above, a cleavage site C is provided when the physically linked connective molecule complex is in a self-hybridization configuration on the surface (e.g., Figure 6 , 7 (e.g., 8A, 9A, 10A, and 11A). In yet another embodiment, and as... Figure 12A As indicated by -C, cleavage site C is available for cleavage via cleavage promoters when the nucleic acid complex is physically linked or in a double-stranded bridge amplification configuration. For example, cleavage site C is a double-stranded motif provided by the double-stranded configuration after the formation of the double strand across a "bridge" on the surface but before denaturation. Figure 12A Once cleaved, the first-strand amplicon will separate from the second-strand amplicon while still binding to the surface. Figure 12B After denaturation and removal of unbound amplicon ( Figure 12C The single-stranded amplicon of both the first and second strands remains bound and can be used for sequencing. In one embodiment, sequencing of the first and second strand amplicon can be performed via a sequencing reaction, as described above. Figure 5D and 5E Those described.
[0109] connector
[0110] As described herein, adaptor molecules can be or include “Y”-shaped, “U”-shaped, “hairpin”-shaped, bubble-shaped (e.g., part of a non-complementary sequence), or other features. “U”-shaped or “hairpin”-shaped adaptors can refer to adaptors having a linker domain that links (or connects) the first strand of a target double-stranded nucleic acid molecule to the second strand of the same molecule. Some hairpin adaptors, for example, can be cleavable hairpin adaptors and / or may include modified or non-standard nucleotides, restriction sites, or other features for in vitro manipulation of structure or function.
[0111] Adaptor molecules can be linked to a variety of terminal nucleic acid materials. For example, an adaptor molecule can be adapted to link to T-terminals, A-terminals, CG-terminals, polynucleotide protrusions (also referred to herein as “sticky ends” or “sticky protrusions”), or single-stranded protrusions of known nucleotide lengths (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides), dehydroxylated bases, blunt ends of nucleic acid materials, and the ends of molecules, wherein the 5' end of the target is dephosphorylated or otherwise blocked from conventional linking. In other embodiments, the adaptor molecule may contain dephosphorylated or otherwise linkage-preventing modifications on the 5' strand at the linking site. In the latter two embodiments, such strategies can be used to prevent dimerization of the library fragment or the adaptor molecule.
[0112] The linker domain of the adaptor can be cleaved with a nuclease (e.g., restriction endonuclease, targeted endonuclease, etc.) to leave a 3' "T" overhang compatible with the 3' "A" overhang of the prepared library fragment. In some embodiments, the resulting linker domain is a single base pair thymine (T) overhang at the 3' end of an extended strand, but in other embodiments, it can be a blunt end, or a "sticky" end of a different type, either 3' or 5' overhang. In this particular example, "CUT" means cleavage performed using a sequence-specific nuclease, such as a restriction enzyme, in a manner that inherently produces linkable ends. In other embodiments, additional enzymatic or chemical treatment, such as using a terminal transferase, following cleavage can produce linkable ends.
[0113] Return to reference Figure 2A The connectable end is shown as a T-protrusion; however, it will be apparent to those skilled in the art that the connectable end can be any of a variety of forms, such as a blunt end, an A-3' protrusion, or a "sticky" end including: a 3' protrusion of one nucleotide, a 3' protrusion of two nucleotides, a 3' protrusion of three nucleotides, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides, a 5' protrusion of one nucleotide, a 5' protrusion of two nucleotides, a 5' protrusion of three nucleotides, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides, etc. (e.g., Figure 2BThe 5' base at the linkage site may be phosphorylated, and the 3' base may have a hydroxyl group, or may be dephosphorylated or dehydrated alone or in combination, or further chemically modified to promote the linkage of one chain to prevent the linkage of another chain, optionally until a later point in time.
[0114] In some embodiments, the adaptor molecule may include a capture portion adapted to separate the desired target nucleic acid molecule to which it is attached.
[0115] The adaptor sequence can refer to a single-stranded sequence, double-stranded sequence, complementary sequence, non-complementary sequence, partially complementary sequence, asymmetric sequence, primer-binding sequence, flow cell sequence, linker sequence, or other sequence provided by the adaptor molecule. In a particular embodiment, the adaptor sequence can refer to a sequence used for amplification via complementary oligonucleotides.
[0116] In some embodiments, the provided methods and compositions comprise at least one adaptor sequence (e.g., two adaptor sequences, one at each of the 5' and 3' ends of nucleic acid material). In some embodiments, the provided methods and compositions may include two or more adaptor sequences (e.g., 3, 4, 5, 6, 7, 8, 9, 10 or more). In some embodiments, at least two of the adaptor sequences are distinct from each other (e.g., by sequence). In some embodiments, each adaptor sequence is distinct from each other (e.g., by sequence). In some embodiments, at least one adaptor sequence is at least partially non-complementary to at least a portion of at least one other adaptor sequence (e.g., non-complementary to at least one nucleotide).
[0117] In some embodiments, the adaptor sequence includes at least one non-standard nucleotide. In some embodiments, the non-standard nucleotide is selected from a base-free site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5'-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, isocytosine, 5'-methyl-isocytosine or isoguanosine, methylated nucleotides, RNA nucleotides, ribonucleotides, 8-oxoguanine, photolyzable linkers, biotinylated nucleotides, dethiobiotinylated nucleotides, and thiol-modified nucleotides. Modified nucleotides, acrylate-modified nucleotides, iso-dC, iso-dG, 2'-O-methyl nucleotides, inosine nucleotides, peptide nucleotides, 5-methyldC, 5-bromodeoxyuridine, 2,6-diaminopurine, 2-aminopurine nucleotides, abase-free nucleotides, 5-nitroindole nucleotides, adenosylnucleotides, azide nucleotides, digitalisin nucleotides, I-linkers, 5'-hexynyl-modified nucleotides, 5-octadiynyl-dU, photolytically cleavable spacers, non-photolytically cleavable spacers, click-chemically compatible modified nucleotides, and any combination thereof.
[0118] In some embodiments, the adapter sequence includes a portion having magnetic properties (i.e., a magnetic portion). In some embodiments, this magnetic property is paramagnetic. In some embodiments, the adapter sequence includes a magnetic portion (e.g., attached to nucleic acid material containing an adapter sequence that includes a magnetic portion), and when a magnetic field is applied, the adapter sequence including the magnetic portion is substantially separated from the adapter sequence that does not include a magnetic portion (e.g., attached to nucleic acid material containing an adapter sequence that does not contain a magnetic portion).
[0119] In some embodiments, at least one adaptor sequence is located at the 5' of the SMI. In some embodiments, at least one adaptor sequence is located at the 3' of the SMI.
[0120] In some embodiments, the adaptor sequence may include one or more adapter domains. In some embodiments, the adapter domain may contain a nucleotide. In some embodiments, the adapter domain may contain at least one modified nucleotide or non-nucleotide molecule (e.g., as described elsewhere in this disclosure). In some embodiments, the adapter domain may be or include a loop.
[0121] In some embodiments, the adaptor sequence at either end or both ends of each strand of the double-stranded nucleic acid material may further include one or more elements providing an SDE. In some embodiments, the SDE may be or include an asymmetric primer site contained in the adaptor sequence.
[0122] In some embodiments, the adaptor sequence may be or include at least one SDE and at least one linker domain (i.e., a domain modified according to the activity of at least one ligase, for example, a domain adapted to be linked to nucleic acid material by the activity of a ligase). In some embodiments, from 5' to 3', the adaptor sequence may be or include a primer binding site, an SDE, and a linker domain.
[0123] Various methods for synthesizing double-stranded sequencing adapters have been previously described, for example, in U.S. Patent No. 9,752,188, International Patent Publication No. WO2017 / 100441, and International Patent Application No. PCT / US18 / 59908 (filed November 8, 2018), all of which are incorporated herein by reference in their entirety.
[0124] Various methods for synthesizing double-stranded sequencing adaptors have been previously described (e.g., U.S. Patent No. 9,752,188 and U.S. Patent No. PCT / US19 / 17908, all of which are incorporated herein by reference). For example, and in one embodiment, an oligonucleotide may hybridize with another oligonucleotide containing a degenerate or semi-degenerate nucleotide sequence at a non-complementary region. The hybridized oligonucleotides may then be chemically linked, or may be two parts of a continuous oligonucleotide that, when hybridized, form a “loop” or “U” shape (hairpin adaptor). The single-stranded degenerate or semi-degenerate region may then be copied using an enzyme capable of polymerizing nucleotides, thereby synthesizing complement. This results in a now complementary double-stranded degenerate or semi-degenerate sequence that can be used as at least one SMI element during double-stranded sequencing. The linking site on the adaptor molecule may be modified by enzymatic or chemical manipulation (e.g., by restriction digestion, terminal transferase activity of polymerases or other enzymes, or any other method known in the art) to extend the product.
[0125] Primers
[0126] In some embodiments, one or more PCR primers having at least one of the following properties are intended for use in various embodiments of the various aspects of the invention: 1) high target specificity; 2) multiplexing capability; and 3) exhibiting robust and minimally biased amplification. Many previous studies and commercial products have been designed primer mixtures to meet some of these standards for conventional PCR-CE. However, it has been noted that these primer mixtures are not always the optimal choice for use with MPS. In fact, developing highly multiplexed primer mixtures can be a challenging and time-consuming process. Conveniently, both Illumina and Promega have recently developed multiplex-compatible primer mixtures for the Illumina platform, demonstrating robust and efficient amplification of a variety of standard and non-standard STR and SNP loci. Because these kits use PCR to amplify their target regions prior to sequencing, the 5' end of each read in the paired end sequencing data corresponds to the 5' end of the PCR primer used to amplify the DNA. In some embodiments, the provided methods and compositions comprise primers designed to ensure uniform amplification, which may require variations in reaction concentration, melting temperature, and minimization of secondary structures and intra- and inter-primer interactions. Several techniques have been described for highly complex primer optimization for MPS applications. In particular, these techniques are often referred to as ampliseq methods, as described in the art.
[0127] Amplification
[0128] In various embodiments, the provided methods and compositions utilize or are used for at least one amplification step, wherein nucleic acid material (or a portion thereof, such as a specific target region or locus) is amplified to form amplified nucleic acid material (e.g., some amplicon products).
[0129] In some embodiments, amplifying nucleic acid material includes the step of amplifying nucleic acid material derived from each of the first and second nucleic acid strands from original double-stranded nucleic acid material using at least one single-stranded oligonucleotide, said at least one single-stranded oligonucleotide being at least partially complementary to a sequence present in a first adaptor sequence. The amplification step further includes using a second single-stranded oligonucleotide to amplify each strand of interest, and such second single-stranded oligonucleotide may (a) be at least partially complementary to a target sequence of interest, or (b) be at least partially complementary to a sequence present in a second adaptor sequence, such that said at least one single-stranded oligonucleotide and the second single-stranded oligonucleotide are oriented in a manner that efficiently amplifies the nucleic acid material.
[0130] In some embodiments, the nucleic acid material in the amplified sample may comprise amplification "tubes" (e.g., PCR tubes), emulsion droplets, microchambers, and other examples or other known containers described above. In some embodiments, the amplified nucleic acid material may comprise amplified nucleic acid material in two or more (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 or more samples) physically separated samples (e.g., tubes, droplets, chambers, containers, etc.).
[0131] While any suitable amplification reaction is considered compatible with some embodiments, as specific examples, in some embodiments the amplification step may be or include polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), isothermal amplification, polymerase clonal amplification in emulsion, bridging amplification on a surface, on the surface of beads or in a hydrogel, and any combination thereof.
[0132] In some embodiments, surface amplification, such as bridging amplification on the surface of a flow cell, comprises generating clusters or multiple copies of a nucleic acid template bound to it. In a particular embodiment, ligated first-stranded and second-stranded nucleic acid templates can be bridge-amplified on the surface of a flow cell, for example, to generate multiple clonal clusters, each clonal cluster comprising copies of both the original first and second strands of the original double-stranded nucleic acid molecule. Some of the clonal copies in the cluster will be in the forward orientation, while the remainder will be in the reverse orientation. When the forward-oriented or reverse-oriented copies are first cleaved and removed, a sequencing reaction can be performed.
[0133] In some embodiments, the amplified nucleic acid material comprises using single-stranded oligonucleotides that are at least partially complementary to regions of adaptor sequences at the 5' and 3' ends of each strand of the nucleic acid material. In some embodiments, the amplified nucleic acid material comprises using at least one single-stranded oligonucleotide that is at least partially complementary to a target region or a target sequence of interest (e.g., a genomic sequence, mitochondrial sequence, plasmid sequence, synthetically produced target nucleic acid, etc.) and at least partially complementary to regions of adaptor sequences (e.g., primer sites).
[0134] Typically, robust amplification, such as PCR amplification, can be highly dependent on reaction conditions. For example, multiplex PCR can be sensitive to buffer composition, monovalent or divalent cation concentration, detergent concentration, congestant (i.e., PEG, glycerol, etc.) concentration, primer concentration, primer Tms, primer design, primer GC content, nucleotide properties of primer modifications, and cycling conditions (i.e., temperature and extension time, and the rate of temperature change). Optimizing buffer conditions can be a difficult and time-consuming process. In some embodiments, the amplification reaction may use at least one of the buffer, primer pool concentration, and PCR conditions according to a previously known amplification protocol. In some embodiments, new amplification protocols may be created, and / or amplification reaction optimization may be used. As a specific example, in some embodiments, PCR optimization kits, such as those from [unclear - likely a specific source], may be used. This PCR optimization kit contains a number of pre-formulated buffers that are partially optimized for various PCR applications, such as multiplex, real-time, GC-rich, and inhibitor-resistant amplification. These pre-formulated buffers can be rapidly replenished with different Mg2+ and primer concentrations, as well as primer pool ratios. Additionally, in some embodiments, various cycling conditions (e.g., thermal cycling) can be evaluated and / or used. When evaluating whether a particular embodiment is suitable for a specific desired application, one or more of the following aspects can be evaluated: specificity, allele coverage of heterozygous loci, locus balance and depth, and others. Measurements of amplification success may include DNA sequencing of the product, evaluation of the product by gel or capillary electrophoresis or HPLC or other size separation methods, followed by fragment visualization, melt flow analysis using double-stranded nucleic acid binding dyes or fluorescent probes, mass spectrometry, or other methods known in the art.
[0135] In some embodiments, at least one amplification step comprises at least one primer, which is or includes at least one non-standard nucleotide. In some embodiments, the non-standard nucleotide is selected from uracil, methylated nucleotides, RNA nucleotides, ribonucleotides, 8-oxoguanine, biotinylated nucleotides, locked nucleic acids, peptide nucleic acids, high Tm nucleic acid variants, allele-recognition nucleic acid variants, any other nucleotide or adapter variants described elsewhere herein, and any combination thereof.
[0136] Nucleic acid materials
[0137] type
[0138] According to various embodiments, any of a variety of nucleic acid materials can be used. In some embodiments, the nucleic acid material may include at least one modification of a polynucleotide within a typical sugar-phosphate backbone. In some embodiments, the nucleic acid material may include at least one modification within any base of the nucleic acid material. For example, as a non-limiting example, in some embodiments, the nucleic acid material is or includes at least one of double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, peptide nucleic acid (PNA), and locked nucleic acid (LNA).
[0139] source
[0140] It is envisioned that nucleic acid materials can be derived from any of a variety of sources. For example, in some embodiments, nucleic acid materials are provided from samples from at least one subject (e.g., a human or animal subject) or other biological sources. In some embodiments, nucleic acid materials are provided from stocked / stored samples. In some embodiments, the sample is or includes at least one of the following: blood, serum, sweat, saliva, cerebrospinal fluid, mucus, uterine lavage fluid, vaginal swab, nasal swab, oral swab, tissue scraping, hair, fingerprint, urine, feces, vitreous fluid, peritoneal lavage fluid, sputum, bronchial lavage fluid, oral lavage fluid, pleural lavage fluid, gastric lavage fluid, gastric juice, bile, pancreatic duct lavage fluid, bile duct lavage fluid, common bile duct lavage fluid, gallbladder fluid, synovial fluid, infected wound, uninfected wound, archaeological samples, forensic samples. Samples include water samples, tissue samples, food samples, bioreactor samples, plant samples, nail scrapings, semen, prostatic fluid, fallopian tube lavage fluid, cell-free nucleic acids, intracellular nucleic acids, metagenomic samples, lavage fluid for implanted foreign bodies, nasal lavage fluid, intestinal fluid, epithelial brush scrapings, epithelial lavage fluid, tissue biopsy samples, autopsy samples, necropsy samples, organ samples, human identification samples, artificially generated nucleic acid samples, synthetic gene samples, nucleic acid data storage samples, tumor tissue, and any combination thereof. In other embodiments, the sample is or includes at least one of microorganisms, plant-based organisms, or any collected environmental sample (e.g., water, soil, archaeological, etc.).
[0141] Modification
[0142] According to various embodiments, nucleic acid materials may be modified one or more times before, substantially simultaneously with, or after any particular step, depending on the application of the particular provided method or composition.
[0143] In some embodiments, the modification may be or include the repair of at least a portion of the nucleic acid material. Although any suitable nucleic acid repair method is considered compatible with some embodiments, certain exemplary methods and compositions are therefore described below and in examples.
[0144] As a non-limiting example, in some embodiments, DNA repair enzymes, such as uracil-DNA glycosylase (UDG), formamide-pyrimidine DNA glycosylase (FPG), and 8-oxoguanine DNA glycosylase (OGG1), can be used to correct DNA damage (e.g., in vitro DNA damage). In some embodiments, these DNA repair enzymes are, for example, glycosylases that remove damaged bases from DNA. For example, UDG removes uracil, which is produced by the deamination of cytosine (caused by spontaneous hydrolysis of cytosine), and FPG removes 8-oxoguanine (e.g., the most common DNA lesion caused by reactive oxygen species). FPG also has lysin activity, which can create a one-base vacancy at a base-free site. Such a base-free site will subsequently be unable to be amplified by PCR, for example, because the polymerase cannot replicate the template. Therefore, the use of such DNA damage repair enzymes can effectively remove damaged DNA without actual mutations, but may not be detected as an error in other ways after sequencing and double-stranded sequence analysis.
[0145] In another embodiment, the sequencing reads generated from the processing steps discussed herein can be further filtered to eliminate false mutations by trimming the ends of reads most prone to artifact generation. For example, DNA fragmentation can generate single-stranded portions at the ends of double-stranded molecules. These single-stranded portions can be filled during end repair (e.g., via Klenow). In some cases, polymerases cause replication errors in these end-repaired regions, resulting in the generation of “false double-stranded molecules.” Once sequenced, these artifacts may appear to be genuine mutations. As a result of the end repair mechanism, these errors can be eliminated from post-sequencing analysis by trimming the ends of the sequencing reads to exclude any possible mutations, thereby reducing the number of erroneous mutations. In some embodiments, such trimming of sequencing reads can be performed automatically (e.g., as part of normal process steps). In some embodiments, the mutation frequency in the fragment end regions can be assessed, and if a threshold level of mutation is observed in the fragment end regions, sequencing read trimming can be performed before generating double-stranded common sequence reads of the DNA fragment.
[0146] Some embodiments of double-stranded sequencing methods provide PCR-based targeted enrichment strategies compatible with error correction using cleavable hairpin adaptors. For example, sequencing enrichment strategies using isolated PCR for sequencing with ligated templates (“SPLiT-DS”) can also benefit from pre-enriched nucleic acid material using one or more embodiments described herein. SPLiT-DS was originally described in International Patent Publication No. WO / 2018 / 175997, which is incorporated herein by reference in its entirety. The SPLiT-DS method can begin with fragmented double-stranded nucleic acid material (e.g., from a DNA sample) labeled with molecular barcoding (e.g., tagging) in a manner similar to that described above and with reference to standard double-stranded sequencing library construction protocols. In some embodiments, the double-stranded nucleic acid material may be fragmented (e.g., using cell-free DNA, damaged DNA, etc.); however, in other embodiments, the steps may include fragmenting the nucleic acid material using mechanical shearing such as acoustic treatment or other DNA cutting methods (as further described herein). The labeled fragmented double-stranded nucleic acid material may include end repair and 3'-dA-tailing (if required in a particular application), followed by ligation of the double-stranded nucleic acid fragments using a double-stranded sequencing adaptor (e.g., a cleavable hairpin adaptor, a Y-shaped adaptor, etc.). In other embodiments, combinations of endogenous or exogenous and endogenous SMI sequences used to uniquely correlate information from the two strands of the original nucleic acid molecule may also be used in conjunction with the physical ligation of the first and second strands. After ligating the adaptor molecule to the double-stranded nucleic acid material, the method can proceed with amplification (e.g., PCR amplification, rolling circle amplification, multiplex displacement amplification, isothermal amplification, bridging amplification, surface binding amplification, etc.).
[0147] Kits containing reagents
[0148] The present invention further encompasses kits (also referred to herein as “DS kits”) for various aspects of double-stranded sequencing methods. In some embodiments, the kit may include various reagents and instructions for performing one or more of the methods and method steps disclosed herein for nucleic acid extraction, nucleic acid library preparation, amplification (e.g., PCR, bridging amplification), cutting of ligated nucleic acid complexes, and sequencing. In one embodiment, the kit may further include computer program products (e.g., coding algorithms running on a computer, access codes for running one or more algorithms on a cloud-based server, etc.) for analyzing sequencing data (e.g., raw sequencing data, sequencing reads, etc.) to determine, for example, sample-associated variant alleles, mutations, etc., according to various aspects of the present invention. The kit may include DNA standards and other forms of positive and negative controls.
[0149] In some embodiments, the DS kit may include reagents or combinations of reagents (e.g., enzymes, dNTPs, wash buffers, etc.) suitable for performing various aspects of sample preparation (e.g., tissue manipulation, DNA extraction, DNA fragmentation), nucleic acid library preparation, amplification, cutting, and sequencer surface treatment steps, and sequencing. For example, the DS kit may optionally include one or more DNA extraction reagents (e.g., buffers, columns, etc.) and / or tissue extraction reagents. Optionally, the DS kit may further include one or more reagents or tools for fragmenting double-stranded DNA, such as by physical means (e.g., tubes, nebulizer units, etc. for promoting acoustic shearing or sonication) or enzymatic means (e.g., enzymes and appropriate reactive enzymes for random or semi-random genomic shearing). For example, a kit may contain a DNA fragmentation reagent for enzymatically fragmenting double-stranded DNA, comprising one or more enzymes for targeted digestion (e.g., restriction endonucleases, CRISPR / Cas endonucleases, and RNA guides and / or other endonucleases), a mixture of double-stranded fragment enzymes, single-stranded DNases (e.g., mung bean nuclease, S1 nuclease) for making the DNA fragments predominantly double-stranded and / or for disrupting single-stranded DNA, and appropriate buffers and solutions to facilitate such enzymatic reactions.
[0150] In one embodiment, the DS kit includes primers and adaptors for preparing a nucleic acid sequence library from a sample, the nucleic acid sequence library being adapted to perform a double-stranded sequencing process step to generate error-correcting (e.g., high-precision) sequences of double-stranded nucleic acid molecules in the sample. For example, the kit may include at least one pool of adaptor molecules including an adapter domain (e.g., a hairpin adaptor), at least one pool of adaptor molecules including a double-stranded portion and a single-stranded portion (e.g., a “Y”-shaped adaptor), or a user-created tool (e.g., a single-stranded oligonucleotide). In some embodiments, the adaptor molecule pool will include single-molecule identifier (SMI) sequences or an appropriate number of substantially unique SMI sequences such that, after ligation of the adaptor molecules, multiple nucleic acid molecules in the sample can be substantially uniquely labeled individually or with a unique combination of characteristics of the fragment to which they are ligated. Those skilled in the art of molecular markers will recognize that the “appropriate” number of SMI sequences required will vary by orders of magnitude depending on various specific factors (input DNA, type of DNA fragmentation, average size of the fragment, complexity and reproducibility of the sequenced sequence within the genome, etc.). Optionally, the adaptor molecule further includes one or more PCR primer binding sites, one or more sequencing primer binding sites, or both. In another embodiment, the DS kit does not contain an adaptor molecule including an SMI sequence or barcode, but instead contains a conventional adaptor molecule (e.g., a Y-sequencing adaptor, etc.), and various method steps can utilize endogenous SMIs on the sequencing surface and / or their physical locations to correlate molecular sequence reads. In some embodiments, the adaptor molecule is an index adaptor and / or includes an index sequence. In other embodiments, the index is added to a specific sample via PCR "tailingin" using primers supplied in the kit.
[0151] In one embodiment, the DS kit includes a set of adaptor molecules, each adaptor molecule having a non-complementary region and / or some other chain-defining element (SDE), or a tool for the user to create them (e.g., single-stranded oligonucleotides). In another embodiment, the kit includes at least one set of adaptor molecules, wherein at least one subset of the adaptor molecules each includes at least one SMI and at least one SDE, or a tool for generating them. In some embodiments, the subset of adaptor molecules may be configured to have connectable ends (e.g., blunt ends, protruding ends, substantially or partially unique sticky ends, etc.). Further features of primers and adaptors used for preparing nucleic acid sequencing libraries from samples suitable for performing double-stranded sequencing process steps have been described above and disclosed in U.S. Patent No. 9,752,188, International Patent Publication No. WO2017 / 100441, and International Patent Application No. PCT / US18 / 59908 (filed November 8, 2018), all of which are incorporated herein by reference in their entirety.
[0152] In one embodiment, the DS kit includes reagents for processing steps occurring on the sequencing surface, such as cleavage promoters (e.g., enzymes, non-enzymatic solutions, light, hybrid oligonucleotides, etc.) and anti-cleavage promoters (e.g., enzymes containing catalytically inactivating enzymes, hybrid oligonucleotides, etc.), as well as other washing solutions for performing the various steps of the method.
[0153] Additionally, the kit may further include DNA quantification materials, such as DNA binding dyes, like those used with Qubit. TM SYBR used with fluorometer TM Green or SYBR TM Gold (available from Thermo Fisher Scientific, Waltham, MA), or PicoGreen for use on a suitable fluorescence spectrometer, real-time PCR machine, or digital droplet PCR machine. TM Dyes (available from Thermo Fisher Scientific, Waltham, Massachusetts). Other reagents suitable for DNA quantification on other platforms are also envisioned. Additional embodiments include kits comprising one or more of the following: nucleic acid size selection reagents (e.g., solid-phase reversible immobilization (SPRI) magnetic beads, gels, columns), columns for capturing target DNA using bait / prey hybridization, qPCR reagents (e.g., for copy number determination), and / or digital droplet PCR reagents. In some embodiments, the kit may optionally comprise one or more of the following: library preparation enzymes (ligases, polymerases, endonucleases, reverse transcriptases for, for example, RNA interrogation), dNTPs, buffers, capture reagents (e.g., beads, surfaces, coating tubes, columns, etc.), index primers, amplification primers (PCR primers), and sequencing primers. In some embodiments, the kit may comprise reagents for assessing the type of DNA damage, such as error-prone DNA polymerases and / or high-fidelity DNA polymerases. Additional additives and reagents (e.g., highly GC-enriched genomes / targets) for use under specific conditions for PCR or ligation reactions are envisioned.
[0154] In one embodiment, the kit further includes reagents such as DNA error-correcting enzymes that repair DNA sequence errors that interfere with the polymerase chain reaction (PCR) process (as opposed to repairing mutations that cause disease). As a non-limiting example, the enzyme includes one or more of the following: monofunctional uracil-DNA glycosylase (hSMUG1), uracil-DNA glycosylase (UDG), N-glycosylase / AP-lyase NEIL 1 protein (hNEIL1), formamidopyrimidine DNA glycosylase (FPG), 8-oxoguanine DNA glycosylase (OGG1), human apurinic / depyrimidine endonuclease (APE1), endonuclease III (Endo III), endonuclease IV (Endo IV), endonuclease V (Endo V), endonuclease VIII (EndoVIII), T7 endonuclease I (T7 Endo I), T4 pyrimidine dimer glycosylase (T4 PDG), human single-stranded selective human alkyladenine DNA glycosylase (hAAG), and other glycosylases, lyases, endonucleases, and exonucleases; and can be used to correct DNA damage (e.g., in vitro or in vivo DNA damage). For example, some of these DNA repair enzymes are glycosylases that remove damaged bases from DNA. For instance, UDG removes uracil, produced by the deamination of cytosine (caused by spontaneous hydrolysis of cytosine), and FPG removes 8-oxoguanine (e.g., the most common DNA lesion caused by reactive oxygen species). FPG also possesses lysin activity, which can create a one-base vacancy at a base-free site. Such base-free sites will subsequently be unable to be amplified by PCR, for example, because the polymerase cannot replicate the template. Therefore, the use of such DNA damage repair enzymes and / or other enzymes listed herein and known in the art can effectively remove damaged DNA that does not have a true mutation but may not have been detected as an error.
[0155] The kit may further include appropriate controls, such as DNA amplification controls, nucleic acid (template) quantification controls, sequencing controls, and nucleic acid molecules derived from similar biological sources (e.g., healthy subjects). In some embodiments, the kit may contain a control cell population. Therefore, the kit may contain suitable reagents (test compounds, nucleic acids, control sequencing libraries, etc.) to provide controls that produce the expected double-stranded sequencing results, which will determine the protocol authenticity of samples including rare genetic variants (e.g., nucleic acid molecules including disease-related variants / mutations that may be incorporated into or included in the sample preparation steps). In some embodiments, the kit may contain reference sequence information. In some embodiments, the kit may contain sequence information for identifying one or more DNA variants in a cell population or cell-free DNA sample. In one embodiment, the kit includes a container for transporting samples; storage materials for stabilizing samples; materials for freezing samples such as cell samples; and materials for analysis to detect DNA variants in subject samples. In another embodiment, the kit may contain nucleic acid contamination control criteria (e.g., hybridization capture probes with affinity for genomic regions in organisms different from the test or subject organism).
[0156] The kit may further include one or more other containers comprising materials desired from a commercial and user perspective, containing PCR and sequencing buffers, diluents, subject sample extraction tools (e.g., syringes, swabs, etc.), and a packaging insert with instructions for use. Additionally, labels with instructions for use, such as those described above, may be provided on the containers; and / or instructions and / or other information may also be included on the insert contained with the kit; and / or via a website address provided therein. The kit may also include laboratory tools, such as sample tubes, plate sealers, microcentrifuge tube openers, labels, magnetic particle separators, foam inserts, ice packs, dry ice packs, insulating materials, etc.
[0157] The kit may further include pre-packaged or application-specific functionalized surfaces for amplifying sequencing libraries. In one embodiment, the functionalized surface may include a surface adapted to perform sequencing reactions therein. The functionalized surface may be pre-configured to have binding oligonucleotides suitable for bridging amplification of sequencing libraries (e.g., the surface includes a distribution of oligonucleotides that bind to sequence domains complementary to one or more adaptor groups). In one embodiment, the functionalized surface is a flow cell configured for use with a sequencing system as described below.
[0158] The kit may further include a computer program product that can be mounted on an electronic computing device (e.g., a laptop / desktop computer, tablet, etc.) or accessed via a network (e.g., a remote server, cloud computing), wherein the computing device or remote server includes one or more processors configured to execute instructions to perform operations including double-stranded sequencing analysis steps. For example, the processor may be configured to execute instructions for processing raw or unanalyzed sequencing reads to generate double-stranded sequencing data. In another embodiment, the computer program product may include a database comprising subject or sample records (e.g., information about a particular subject or sample or sample set) and empirically derived information about DNA target regions. The computer program product is embodied in a non-transitory computer-readable medium that, when executed on a computer, performs the steps of the methods disclosed herein.
[0159] The kit may further include instructions and / or access codes / passwords for accessing a remote server (including cloud-based servers) to upload and download data (e.g., sequencing data, reports, other data) or software to be installed on a local device. All computations can reside on the remote server and be accessed by the user / kit user via an internet connection, etc.
[0160] The kit may be adapted to be used with sequencing systems that are compatible with the methods and reagents described herein. For example, the sequencing system and associated sequencing reagents may be configured to perform stepwise sequencing reactions that provide interventional treatment steps. In one embodiment, the sequencing system may provide a delivery system for delivery of cleavage promoters, anti-cleavage promoters, enzyme solutions, oligonucleotides, wash buffers, etc. Similarly, the sequencing system may include appropriate controls (e.g., manual, automatic, semi-automatic, etc.) and internal programs for treatment step times, temperatures, pH, concentrations, etc.
[0161] Example
[0162] In addition to the various aspects, embodiments, examples, etc., described herein, this disclosure also includes exemplary aspects numbered E1 through E87 (“E”). This list of aspects is presented as an exemplary list and this application is not limited to these aspects.
[0163] E1. A method for sequencing a double-stranded target nucleic acid molecule, the method comprising:
[0164] (a) Amplifying physically linked nucleic acid complexes on a surface to generate a physically linked nucleic acid complex amplicon that binds to the surface in both a forward and reverse orientation, wherein the physically linked nucleic acid complex comprises: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor including a linker domain at a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion at a second end of the double-stranded target nucleic acid molecule;
[0165] (b) Removing (i) the physically linked nucleic acid complex amplicon that is bound to the surface in the reverse orientation or (ii) the physically linked nucleic acid complex amplicon that is bound to the surface in the forward orientation;
[0166] (c) Cut a portion of the remaining physically linked nucleic acid complex amplicon to provide a subset of single-stranded amplicon that includes information from one strand and a subset of physically linked nucleic acid complex amplicon;
[0167] (d) Sequencing a subset of the single-stranded amplicon to provide sequencing reads of the original strand of the double-stranded target nucleic acid molecule;
[0168] (e) Amplifying a subset of physically linked nucleic acid complex amplicones on the surface;
[0169] (f) Remove the nucleic acid complex amplicon that is in another orientation from the physically linked nucleic acid complex;
[0170] (g) Cleavage the remaining physically bound nucleic acid complex amplicon to provide a single-stranded amplicon that includes information from the other strand; and
[0171] (h) Sequencing the single-stranded amplicon to provide a sequencing read from the other original strand of the double-stranded target nucleic acid molecule.
[0172] E2. A method for sequencing a double-stranded target nucleic acid molecule, the method comprising:
[0173] (a) Amplifying physically linked nucleic acid complexes on a surface to generate a cluster of physically linked nucleic acid complex amplicones that bind to the surface, wherein the physically linked nucleic acid complexes comprise: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor including a linker domain at one end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion at the other end of the double-stranded target nucleic acid molecule;
[0174] (b) Remove the physically linked nucleic acid complex amplicon that is bound to the surface at the 5' end of the physically linked nucleic acid complex amplicon described in (i) or the 3' end of the physically linked nucleic acid complex amplicon described in (ii);
[0175] (c) Cleaving at least a portion of the remaining physically linked nucleic acid complex amplicon at the cleavage site to provide a single-stranded amplicon including sequence information from one original strand of the double-stranded target nucleic acid molecule; and
[0176] (d) Sequencing the single-stranded amplicon to provide a sequencing read of one of the original strands of the double-stranded target nucleic acid molecule.
[0177] E3. The method according to E2, wherein cleaving at least a portion of the remaining physically linked nucleic acid complex amplicon includes retaining at least one physically linked nucleic acid complex amplicon bound to the surface.
[0178] E4. The method according to E3 further includes:
[0179] (e) Amplify the at least one physically linked nucleic acid complex amplicon on the surface to further proliferate the cluster of the physically linked nucleic acid complex amplicon bound to the surface;
[0180] (f) Remove the physically linked nucleic acid complex amplicon that was not removed in (b) and is in a different orientation;
[0181] (g) Cleavage the remaining physically bound nucleic acid complex amplicon to provide a single-stranded amplicon including information from the other original strand of the double-stranded target nucleic acid molecule; and
[0182] (h) Sequencing the single-stranded amplicon to provide a sequencing read from the other original strand of the double-stranded target nucleic acid molecule.
[0183] E5. The method according to any one of the foregoing examples further includes: comparing a sequence read from one original strand with a sequence read from another original strand to generate a common sequence of the double-stranded target nucleic acid molecule.
[0184] E6. The method according to any one of E1 to E4, further comprising:
[0185] Identify sequence variations in sequence reads from one original chain and sequence reads from another original chain, wherein the sequence variations from one original chain and the other original chain are identical sequence variations; or
[0186] Eliminate or ignore sequence variations that occur in one original chain but not in another.
[0187] E7. The method according to any one of E1 to E4, further comprising:
[0188] The sequence reads from one of the original chains are compared with the sequence reads from another original chain;
[0189] Identify nucleotide positions where sequence reads from one original strand do not match those from another original strand; and
[0190] The miscorrected sequence of the double-stranded target nucleic acid molecule is generated by ignoring, eliminating, or correcting the identified inconsistent nucleotide positions.
[0191] E8. A method for sequencing a population of double-stranded target nucleic acid molecules, each double-stranded target nucleic acid molecule comprising a first strand and a second strand, the method comprising:
[0192] (a) Amplifying multiple physically linked nucleic acid complexes on a surface to generate multiple clonal clusters, each clonal cluster comprising multiple physically linked nucleic acid complex amplicones, each nucleic acid complex amplicon comprising a first-strand amplicon and a second-strand amplicon, wherein each physically linked nucleic acid complex comprises: (i) a double-stranded target nucleic acid molecule from the population; (ii) a first adaptor comprising an adapter domain connected to a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion connected to a second end of the double-stranded target nucleic acid molecule;
[0193] (b) Remove the physically linked nucleic acid complex amplicon from each clonal cluster that is bound to the surface in the reverse orientation described in (i) or the forward orientation described in (ii);
[0194] (c) Cutting a portion of the remaining surface-bound, physically linked nucleic acid complex amplicon after (b), thereby physically separating the first strand amplicon and the second strand amplicon;
[0195] (d) Removal of unbound, physically separated first-strand or second-strand amplicon; and
[0196] (e) Sequencing the remaining physically separated first-strand or second-strand amplicon bound to the surface to generate nucleic acid sequence reads of the first-strand or second-strand for each cloning cluster on the surface.
[0197] E9. The method according to E8, wherein cleaving at least a portion of the remaining physically linked nucleic acid complex amplicon comprises retaining at least one physically linked nucleic acid complex amplicon from at least some of the clonal clusters bound to the surface.
[0198] E10. The method according to E9 further includes:
[0199] (f) In at least some of the clone clusters, amplify the at least one physically linked nucleic acid complex amplicon on the surface to further proliferate the clone cluster of physically linked nucleic acid complex amplicon bound to the surface;
[0200] (g) Remove the nucleic acid complex amplicon in other orientations from step (b);
[0201] (h) Remove unbound, physically separated first-strand or second-strand amplicon;
[0202] (i) the remaining physically linked nucleic acid complex amplicon after cleavage (h), and thereby physically separating the first-strand amplicon and the second-strand amplicon; and
[0203] (j) Sequencing the remaining physically separated first-strand or second-strand amplicon that is bound to the surface to generate a nucleic acid sequence read of the first-strand or second-strand for each cloning cluster on the surface.
[0204] E11. A method for sequencing a population of double-stranded target nucleic acid molecules, each double-stranded target nucleic acid molecule comprising a first strand and a second strand, the method comprising:
[0205] (a) Amplifying multiple physically linked nucleic acid complexes bound on a surface to generate multiple clusters, each cluster comprising multiple physically linked nucleic acid complex amplicones representing an original double-stranded target nucleic acid molecule, wherein each physically linked nucleic acid complex amplicon comprises a first-strand amplicon and a second-strand amplicon, and wherein each physically linked nucleic acid complex comprises a double-stranded target nucleic acid molecule from the cluster, the double-stranded target nucleic acid molecule being: (i) linked at one end to a first adapter comprising a linker domain between the first strand and the second strand; and (ii) linked at the other end to a second adapter having a double-stranded portion and a single-stranded portion;
[0206] (b) Cutting the surface-bound, physically connected nucleic acid complex amplicon, thereby physically separating the first-strand amplicon and the second-strand amplicon;
[0207] (c) Removal of unbound, physically separated first-strand amplicon and / or unbound, physically separated second-strand amplicon, wherein the remaining amplicon bound to the surface comprises: (i) the physically separated first-strand amplicon; and (ii) the physically separated second-strand amplicon;
[0208] (d) Sequencing the physically separated first-strand amplicon bound to the surface to generate nucleic acid sequence reads of the first strand for each cluster on the surface; and
[0209] (e) Sequencing the physically separated second-strand amplicon bound to the surface to generate a second-strand nucleic acid sequence read for each cluster on the surface.
[0210] E12. The method according to E10 or E11, further comprising: comparing, for at least some of the clusters on the surface, the nucleic acid sequence reads of the first strand with the nucleic acid sequence reads of the second strand to generate error-corrected sequence reads of the original double-stranded target nucleic acid molecule.
[0211] E13. The method according to any one of E10 to E12, further comprising: using a unique molecular identifier (UMI) to associate the nucleic acid sequence read of the first strand of an original double-stranded target nucleic acid molecule from the population with the nucleic acid sequence read of the second strand of the same original double-stranded target nucleic acid molecule.
[0212] E14. The method according to E13, wherein the UMI includes a physical location on the surface.
[0213] E15. The method according to E14, wherein the UMI includes a tag sequence, a molecular-specific feature, a cluster location on the surface, or a combination thereof.
[0214] E16. The method according to E15, wherein the molecular specific features include nucleic acid mapping information for a reference sequence, sequence information at or near the ends of the double-stranded target nucleic acid molecule, the length of the double-stranded target nucleic acid molecule, or a combination thereof.
[0215] E17. The method according to any one of E10 to E16, further comprising: using a chain defining element (SDE) to distinguish the nucleic acid sequence read of the first strand of the original double-stranded target nucleic acid molecule from the nucleic acid sequence read of the second strand of the same original double-stranded target nucleic acid molecule.
[0216] E18. The method according to E17, wherein the SDE is the sequence read information associated with steps (e) and (j) of E10 or with steps (d) and (e) of E11.
[0217] E19. The method according to E17, wherein the SDE includes a portion of a connector sequence.
[0218] E20. The method according to any one of E8 to E19, wherein sequencing the physically separated first-strand amplicon or the second-strand amplicon comprises synthetic sequencing.
[0219] E21. The method according to any one of E8 to E20, further comprising:
[0220] The physically linked nucleic acid complex is prepared by linking the first and second adaptors to each of the plurality of double-stranded target nucleic acid molecules in the population; and
[0221] The physically linked nucleic acid complex is presented onto the surface having a plurality of bound oligonucleotides that are at least partially complementary to the single-stranded portion of the second adaptor, such that the plurality of physically linked nucleic acid complexes are captured on the surface by hybridization with the plurality of bound oligonucleotides.
[0222] E22. The method according to E21, further comprising a step of amplifying the physically linked nucleic acid complex prior to the presentation step.
[0223] E23. The method according to E22, wherein amplifying the physically linked nucleic acid complex prior to the presentation step comprises PCR amplification or circular amplification.
[0224] E24. The method according to any one of E21 to E23, wherein the physically linked nucleic acid complex is captured on the surface in both forward and reverse orientations.
[0225] E25. The method according to any one of E8 to E24, wherein the amplification step in (a) comprises bridging amplification.
[0226] E26. The method according to any one of E8 to E25, further comprising:
[0227] For at least some of the double-stranded target nucleic acid molecules in the said population:
[0228] (i) Compare the sequence reads from the first chain with the sequence reads from the second chain;
[0229] (ii) Identifying nucleotide positions where there is a discrepancy between the sequence read from the first strand and the sequence read from the second strand; and
[0230] (iii) To generate an error-corrected sequence read of the double-stranded target nucleic acid molecule by ignoring, eliminating or correcting identified inconsistent nucleotide positions.
[0231] E27. The method according to any one of E1 to E26, wherein the first connector comprises a cleavable site or motif.
[0232] E28. The method according to any one of E1 to E27, wherein the first adapter and the second adapter each include a sequencing primer binding site and optionally a single-molecule identifier (SMI) sequence.
[0233] E29. The method according to any one of E1 to E27, wherein the second adaptor comprises a sequencing primer binding site, an amplification primer binding site, an index sequence, or any combination thereof.
[0234] E30. The method according to any one of E1 to E29, wherein the joint structure domain includes a cutting point.
[0235] E31. The method according to any one of E1 to E29, wherein the first connector comprises a cuttable structural domain.
[0236] E32. The method according to any one of E1 to E31, wherein the first adaptor comprises a hairpin loop structure, the hairpin loop structure comprising a self-complementary stem portion and a single-stranded nucleotide loop portion.
[0237] E33. The method according to E32, wherein the single-stranded nucleotide loop portion includes a cleavable domain.
[0238] E34. The method according to E32, wherein the stem portion includes a cuttable structural domain.
[0239] E35. According to the method of E33 or E34, the cleavable domain includes an enzyme recognition site.
[0240] E36. The method according to E35, wherein the enzyme recognition site is a nuclease recognition site.
[0241] E37. The method according to E36, wherein the endonuclease is a restriction enzyme or a targeted endonuclease.
[0242] E38. The method according to any one of E1 to E37, wherein the second connector is a "Y"-shaped connector.
[0243] E39. The method according to E38, wherein one or both arms of the Y-shaped adaptor can hybridize with an oligonucleotide bound to the surface.
[0244] E40. The method according to any one of E1 to E39, wherein the single-stranded portion of the second adapter comprises a first arm having a first primer binding site and a second arm having a second primer binding site.
[0245] E41. The method according to E40, wherein when denatured, the physically linked double-stranded nucleic acid complex from 5' to 3' or from 3' to 5' comprises: a first primer binding site, a first strand, a first adapter including the adapter domain, a second strand, and a second primer binding site.
[0246] E42. The method according to any one of E1 to E41, wherein the surface is a sequencing surface.
[0247] E43. The method according to any one of E1 to E42, wherein the surface is a flow cell.
[0248] E44. The method according to any one of E1 to E43, wherein the surface is the surface of a bead.
[0249] E45. The method according to any one of E1 to E44, wherein the amplification is selected from the group consisting of: PCR amplification, isothermal amplification, clonal amplification, cluster amplification and bridge amplification.
[0250] E46. The method according to any one of E1 to E45, wherein the amplification is a bridging amplification on the surface.
[0251] E47. The method according to any one of E8 to E46, wherein one or more of the plurality of first-strand amplicon and / or the plurality of second-strand amplicon bind to the surface in a forward orientation.
[0252] E48. The method according to any one of E8 to E46, wherein one or more of the plurality of first-strand amplicon and / or the plurality of second-strand amplicon bind to the surface in a reverse orientation.
[0253] E49. The method according to any one of E8 to E48, further comprising: passing the plurality of physically linked double-stranded nucleic acid complexes through the surface prior to the amplification in (a).
[0254] E50. The method according to any one of E1 to E49, wherein the surface comprises a plurality of one or more bound oligonucleotides that are at least partially complementary to one or more regions of the second adaptor.
[0255] E51. The method according to E50, wherein the plurality of one or more bound oligonucleotides are at least partially complementary to the single-stranded portion of the second adaptor.
[0256] E52. The method according to any one of E1 to E51, wherein the first and second strands of the physically linked nucleic acid complex are amplified in step (a) by multiple amplification reactions to generate a cluster of the physically linked nucleic acid complex amplicon on the surface.
[0257] E53. The method according to any one of E8 to E52, wherein the first and second strands of each of the plurality of physically linked nucleic acid complexes are amplified in step (a) to simultaneously generate the plurality of clusters on the surface.
[0258] E54. The method according to any one of E1 to E8 and E12 to E53, wherein cleaving a portion of the physically linked nucleic acid complex amplicon comprises inefficient cleavage at a cleavable site in the first adapter, thereby producing both cleaved and uncleaved nucleic acid complexes within each cluster on the surface.
[0259] E55. The method according to E54, wherein the ratio of uncut nucleic acid complexes in all nucleic acid complexes within each cluster on the flow cell is 1%, 5%, 10%, 20%, 30%, 40%, 45%, or 50%.
[0260] E56. The method according to E54 or E55, wherein the cleaved nucleic acid complex is cleaved by a cleavage accelerator at a cleavable site in the adapter domain of the first adapter.
[0261] E57. The method according to E56, wherein the cleavage is a site-directed enzymatic reaction.
[0262] E58. The method according to E56 or E57, wherein the cleavage accelerator is a nuclease.
[0263] E59. The method according to E58, wherein the endonuclease is a restriction site endonuclease or a targeted endonuclease.
[0264] E60. The method according to E56 or E57, wherein the cleavage promoter is selected from the group consisting of: ribonucleoproteins, Cas enzymes, Cas9-like enzymes, broad-spectrum nucleases, transcription activator-like effector-based nucleases (TALENs), zinc finger nucleases, argonaute nucleases, or combinations thereof.
[0265] E61. The method according to E56 or E57, wherein the cleavage promoter comprises a CRISPR-associated enzyme.
[0266] E62. The method according to E56 or E57, wherein the cutting accelerator comprises Cas9 or CPF1 or a derivative thereof.
[0267] E63. The method according to E56 or E57, wherein the cutting promoter comprises a cutting enzyme or a cutting enzyme variant.
[0268] E64. The method according to E56, wherein the cutting accelerator comprises a chemical process.
[0269] E65. The method according to any one of E54 to E64, wherein the amount of uncut nucleic acid complex remaining on the surface can be scaled by controlling the amount or concentration of the cleavage accelerator introduced for site-specific cleavage or by controlling the amount of time the cleavage accelerator is introduced for site-specific cleavage.
[0270] E66. The method according to any one of E54 to E63, wherein the uncut nucleic acid complex is protected by adding an anti-cleavage promoter before or during the cleavage step.
[0271] E67. The method according to E66, wherein the anti-cutting accelerator includes an anti-cutting motif in the joint structure domain of the first connector.
[0272] E68. The method according to E67, wherein the cleavable site is already present in the linker domain of the first adapter, and the cleavage-resistant motif is generated by hybridization with an oligonucleotide comprising a sequence at least partially complementary to the linker domain of the first adapter.
[0273] E69. The method according to E66 to E68, wherein cleaving a portion of the physically linked nucleic acid complex amplicon further comprises:
[0274] (i) Introducing the aforementioned anti-cutting accelerator; and
[0275] (ii) The cutting accelerator is introduced after or simultaneously with (i).
[0276] The interaction with the anti-cleavage promoter protects the physically linked nucleic acid complex amplicon from cleavage.
[0277] E70. The method according to E54 to E63, wherein the cleavable site is generated by hybridization with an oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the first adaptor, and wherein a physically linked nucleic acid complex amplicon that has not hybridized with the oligonucleotide is not cleaved.
[0278] E71. The method according to E54 to E63, wherein the cleavable site is generated by hybridization with a first oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the adaptor, and the cleavage-resistant motif is generated by hybridization with a second oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the adaptor, and wherein cleaving a portion of the physically linked nucleic acid complex amplicon further comprises:
[0279] (i) Introducing a mixture of the first oligonucleotide and the second oligonucleotide; and
[0280] (ii) Introduce the cutting accelerator.
[0281] E72. The method according to E71, wherein the first oligonucleotide or the second oligonucleotide is methylated.
[0282] E73. The method according to E70 or E71, wherein the hybridization can be scaled by controlling the amount or concentration of the oligonucleotide introduced for hybridization or by controlling the amount of time the oligonucleotide is introduced for hybridization.
[0283] E74. The method according to any one of E67, E68 or E71 to E73, wherein the anti-cleavage motif comprises an oligonucleotide sequence having a large adduct or side chain that prevents entry into the cleavage site.
[0284] E75. The method according to any one of E67, E68 or E71 to E73, wherein the anti-cleavage motif comprises one or more mismatched oligonucleotide sequences that prevent cleavage promoters from recognizing the cleavage site.
[0285] E76. The method according to any one of E67, E68 or E71 to E73, wherein the cleavage-resistant motif comprises one or more of the following: an oligonucleotide sequence having a nucleoside analog, a base-free site, a nucleotide analog, and a peptide nucleic acid bond.
[0286] E77. The method according to E54 to E63, wherein the cleaved nucleic acid complex is cleaved by a catalytically active enzyme at a cleavable site in the first adapter, and the uncleaved nucleic acid complex is protected from cleavage by a catalytically inactivating enzyme in the first adapter.
[0287] E78. The method according to any one of E54 to E63, wherein the cleavage site is located in the self-complementary portion of the first connector or the single-stranded portion of the first connector.
[0288] E79. The method according to E78, wherein the cleavage site is available when the physically linked nucleic acid complex amplicon is in a self-hybridization configuration on the surface.
[0289] E80. The method according to E54 to E63, wherein the cleavage site is available when the physically linked nucleic acid complex amplicon is in a double-stranded bridge amplification configuration.
[0290] E81. The method according to any one of E8 to E80, further comprising selectively enriching nucleic acid complexes having physical links to one or more target genomic regions prior to step (a) to provide a plurality of enriched physically linked nucleic acid complexes.
[0291] E82. A kit for error-correcting double-stranded sequencing of double-stranded nucleic acid molecules, the kit comprising:
[0292] At least one set of sequencing primers;
[0293] A set of first linker molecules, wherein the first linker molecules include a linker structural domain;
[0294] A set of second linker molecules, the second linker molecules comprising double-stranded portions and single-stranded portions configured to be fixed on the surface for amplification;
[0295] The primers and adaptor molecules described herein can be used for error-correcting double-stranded sequencing experiments; and
[0296] Instructions for use of the kit to perform error-corrected double-strand sequencing of nucleic acids extracted from biological samples.
[0297] E83. The kit according to E82 further includes a cutting accelerator.
[0298] E84. The kit according to E82 or E83, wherein the connector domain has a cleavable motif.
[0299] E85. The kit according to any one of E82 to E84, further comprising an anti-cutting accelerator.
[0300] E86. The kit according to any one of E82 to E85, further comprising a computer program product embodied in a non-transitory computer-readable medium, which, when executed on a computer or remote computing server, performs the step of determining error-corrected double-stranded sequencing reads for one or more double-stranded nucleic acid molecules in a sample.
[0301] E87. A sequencing system comprising:
[0302] A sequencing surface comprising covalently bound oligonucleotides;
[0303] A delivery system for delivering sequencing reagents to the sequencing surface;
[0304] Delivery system for delivering a cutting promoter to the sequencing surface; and
[0305] A computing network for transmitting information related to sequencing data, wherein the information includes one or more of raw sequencing data, double-stranded sequencing data, and sample information.
[0306] in conclusion
[0307] The above detailed description of embodiments of the present invention is not intended to be exhaustive or to limit the invention to the precise forms disclosed above. While specific embodiments and examples of the invention have been described above for illustrative purposes, various equivalent modifications are possible within the scope of the invention, as will be recognized by those skilled in the art. For example, although the steps are presented in a given order, alternative embodiments may perform the steps in a different order. The various embodiments described herein may also be combined to provide other embodiments. All of the documents described herein are fully set forth and incorporated by reference.
[0308] Based on the foregoing, it will be understood that specific embodiments of the invention have been described herein for illustrative purposes, but well-known structures and functions have not been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments of the invention. Where the context permits, singular or plural terms may also include plural or singular terms respectively.
[0309] Furthermore, unless the word “or” is explicitly limited to meaning only a single item other than those in a list having two or more items, its use in such a list should be interpreted as including (a) any single item in the list, (b) all items in the list, or (c) any combination of items in the list. Additionally, the term “comprising” is used throughout to mean at least one or more of the stated features, without excluding any larger number of the same features and / or other features of a different type. It should also be understood that specific embodiments have been described herein for illustrative purposes, but various modifications may be made without departing from the art. Furthermore, while advantages associated with certain embodiments of the new technology have been described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments need to exhibit such advantages to fall within the scope of the art. Therefore, this disclosure and associated techniques may cover other embodiments not explicitly shown or described herein.
Claims
1. A method for sequencing double-stranded target nucleic acid molecules, the method comprising: (a) Amplifying physically linked nucleic acid complexes on a surface to generate a physically linked nucleic acid complex amplicon that binds to the surface in both forward and reverse orientations, wherein the physically linked nucleic acid complexes comprise: (i) the double-stranded target nucleic acid molecule; (ii) a first adaptor comprising a linker domain and connected to a first end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion and connected to a second end of the double-stranded target nucleic acid molecule, wherein the first adaptor and / or the second adaptor comprises a primer binding site; (b) Remove (i) the physically linked nucleic acid complex amplicon that is bound to the surface in the reverse orientation or (ii) the physically linked nucleic acid complex amplicon that is bound to the surface in the forward orientation; (c) Cleaving a portion of the remaining physically linked nucleic acid complex amplicon bound to the surface at a cleavage site associated with the first and / or the second adapter to provide (i) a cleaved single-stranded amplicon bound to the surface and including information from one strand of the double-stranded target nucleic acid molecule and (ii) an uncleaved physically linked nucleic acid complex amplicon bound to the surface; (d) Sequencing the single-stranded amplicon that binds to the surface and includes information from one strand of the double-stranded target nucleic acid molecule to provide a sequencing read of the original strand of the double-stranded target nucleic acid molecule. (e) Amplify uncut, physically linked nucleic acid complex amplicones that bind to the surface to further proliferate clusters of physically linked nucleic acid complex amplicones that bind to the surface; (f) Remove the physically linked nucleic acid complex amplicon that is in a different orientation relative to the amplicon that was not removed in (b); (g) Cleaving the remaining physically linked nucleic acid complex amplicon bound to the surface to provide a cleaved single-stranded amplicon bound to the surface and including information from the other original strand of the double-stranded target nucleic acid molecule; and (h) Sequencing the single-stranded amplicon that binds to the surface and includes information from the other strand of the double-stranded target nucleic acid molecule to provide a sequencing read from the other original strand of the double-stranded target nucleic acid molecule.
2. A method for sequencing a population of double-stranded target nucleic acid molecules, the method comprising: (a) Amplifying multiple physically linked nucleic acid complexes on a surface to generate multiple clonal clusters, each clonal cluster comprising multiple physically linked nucleic acid complex amplicones bound to the surface, wherein each physically linked nucleic acid complex comprises: (i) a double-stranded target nucleic acid molecule from the population; (ii) a first adaptor comprising an adapter domain and connected to one end of the double-stranded target nucleic acid molecule; and (iii) a second adaptor having a double-stranded portion and a single-stranded portion and connected to the other end of the double-stranded target nucleic acid molecule, wherein the first adaptor and / or the second adaptor comprises a primer binding site; (b) Remove the physically linked nucleic acid complex amplicon from each clonal cluster that is bound to the surface at the 5' end of the physically linked nucleic acid complex amplicon in (i) or the 3' end of the physically linked nucleic acid complex amplicon in (ii); (c) Cleaving at least a portion of the remaining physically linked nucleic acid complex amplicon from each cloning cluster that is bound to the surface at a cleavage site associated with the first and / or the second adaptor to provide a cleaved single-stranded amplicon that is bound to the surface and includes sequence information of one original strand derived from the double-stranded target nucleic acid molecule, wherein after the cleavage, at least one uncleaved physically linked nucleic acid complex amplicon from each cloning cluster remains bound to the surface; and (d) Sequencing the cut single-stranded amplicon that binds to the surface and includes sequence information of one original strand of the double-stranded target nucleic acid molecule to provide each cloning cluster on the surface with a sequencing read of the original strand of the double-stranded target nucleic acid molecule.
3. The method according to claim 2, further comprising: (e) Amplify on the surface the at least one physically linked nucleic acid complex amplicon from each clone cluster that was not cleaved in (d) to reproliferate each clone cluster; (f) Remove the physically linked nucleic acid complex amplicon that is in a different orientation relative to the amplicon removed in (b); (g) Cutting the remaining physically linked nucleic acid complex amplicon from each cloning cluster bound to the surface to provide each cloning cluster on the surface with a cut single-stranded amplicon that is bound to the surface and includes information from the other original strand of the double-stranded target nucleic acid molecule; and (h) Sequencing of the cut single-stranded amplicon that binds to the surface and includes information from the other original strand of the double-stranded target nucleic acid molecule to provide each cloning cluster on the surface with a sequencing read from the other original strand of the double-stranded target nucleic acid molecule.
4. The method of claim 3, further comprising comparing a sequence read from one original strand with a sequence read from another original strand to generate a common sequence of the double-stranded target nucleic acid molecule.
5. The method of claim 1, further comprising: The sequence reads from one of the original chains are compared with the sequence reads from another original chain; Identify nucleotide positions where a sequence read from one original strand is inconsistent with a sequence read from another original strand; as well as The miscorrected sequence of the double-stranded target nucleic acid molecule is generated by ignoring, eliminating, or correcting the identified inconsistent nucleotide positions.
6. A method for sequencing a population of double-stranded target nucleic acid molecules, each double-stranded target nucleic acid molecule comprising a first strand and a second strand, the method comprising: (a) Amplifying multiple physically linked nucleic acid complexes bound on a surface to generate multiple clusters, each cluster comprising multiple physically linked nucleic acid complex amplicones representing an original double-stranded target nucleic acid molecule, wherein each physically linked nucleic acid complex amplicon comprises a first-strand amplicon and a second-strand amplicon, and wherein each physically linked nucleic acid complex comprises a double-stranded target nucleic acid molecule from the population, the double-stranded target nucleic acid molecule being: (i) linked to a first adaptor comprising a linker domain between the first strand and the second strand and connected to one end of the double-stranded target nucleic acid molecule; and (ii) linked to a second adaptor having a double-stranded portion and a single-stranded portion and connected to the other end of the double-stranded target nucleic acid molecule, wherein the first adaptor and / or the second adaptor comprises a primer binding site; (b) Cleaving the surface-bound, physically connected nucleic acid complex amplicon at a cleavage site associated with the first and / or the second adapter, and physically separating the first strand amplicon from the second strand amplicon; (c) Remove unbound physically separated first-strand amplicon from surface-bound second-strand amplicon and unbound physically separated second-strand amplicon from surface-bound first-strand amplicon, such that the remaining amplicon bound to the surface comprises: (i) surface-bound physically separated first-strand amplicon; and (ii) surface-bound physically separated second-strand amplicon; (d) Sequencing the physically separated first-strand amplicon bound to the surface to generate nucleic acid sequence reads of the first strand for each cluster on the surface; and (e) Sequencing the physically separated second-strand amplicon bound to the surface to generate a second-strand nucleic acid sequence read for each cluster on the surface.
7. The method of claim 6, further comprising using a unique molecular identifier (UMI) to associate the nucleic acid sequence read of the first strand of an original double-stranded target nucleic acid molecule from the population with the nucleic acid sequence read of the second strand of the same original double-stranded target nucleic acid molecule.
8. The method of claim 7, wherein the UMI includes a physical location on the surface.
9. The method of claim 8, wherein the UMI comprises a tag sequence, a molecular-specific feature, a cluster location on the surface, or a combination thereof.
10. The method of claim 6, further comprising using a strand defining element (SDE) to distinguish the nucleic acid sequence read of the first strand of the original double-stranded target nucleic acid molecule from the nucleic acid sequence read of the second strand of the same original double-stranded target nucleic acid molecule.
11. The method of claim 6, further comprising: For at least some of the double-stranded target nucleic acid molecules in the said population: (i) Compare the sequence reads from the first chain with the sequence reads from the second chain; (ii) Identify nucleotide positions where there is a discrepancy between the sequence read from the first strand and the sequence read from the second strand; as well as (iii) To generate an error-corrected sequence read of the double-stranded target nucleic acid molecule by ignoring, eliminating or correcting identified inconsistent nucleotide positions.
12. The method of claim 1, wherein the first connector includes the cutting site.
13. The method of claim 1, wherein the first adaptor comprises a hairpin loop structure, the hairpin loop structure comprising a self-complementary stem portion and a single-stranded nucleotide loop portion.
14. The method of claim 13, wherein the cleavage site is located in the single-stranded nucleotide loop portion or the self-complementary stem portion.
15. The method of claim 6, wherein the surface is a sequencing surface.
16. The method of claim 6, further comprising passing the plurality of physically linked double-stranded nucleic acid complexes through the surface prior to the amplification in (a).
17. The method of any one of claims 1 to 4, wherein the first adapter includes a cleavage site, and wherein cleaving a portion of the physically linked nucleic acid complex amplicon comprises inefficient cleavage at the cleavage site in the first adapter, thereby generating both cleaved and uncleaved nucleic acid complexes within each cluster on the surface.
18. The method of claim 17, wherein the ratio of uncut nucleic acid complexes in all nucleic acid complexes within each cluster on the flow cell is 1%, 5%, 10%, 20%, 30%, 40%, 45%, or 50%.
19. The method of claim 17, wherein the cleaved nucleic acid complex is cleaved by a cleavage accelerator.
20. The method of claim 19, wherein the cleaved nucleic acid complex is cleaved by a site-directed enzymatic reaction.
21. The method of claim 19, wherein the cleavage accelerator is a nuclease.
22. The method of claim 20, wherein the cleavage accelerator is a nuclease.
23. The method of claim 17, wherein the amount of uncut nucleic acid complex remaining on the surface can be scaled by controlling the amount or concentration of the cleavage accelerator introduced for site-specific cleavage or by controlling the amount of time the cleavage accelerator is introduced for site-specific cleavage.
24. The method of claim 17, wherein the uncut nucleic acid complex is protected by adding an anti-cleavage promoter before or during the cleavage step.
25. The method of claim 24, wherein cleaving a portion of the physically linked nucleic acid complex amplicon further comprises: (i) Introducing the aforementioned anti-cutting accelerator; as well as (ii) The cutting accelerator is introduced after or simultaneously with (i). The interaction with the anti-cleavage promoter protects the physically linked nucleic acid complex amplicon from cleavage.
26. The method of claim 17, wherein the cleavage site is generated by hybridization with an oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the first adapter, and wherein a physically linked nucleic acid complex amplicon that has not hybridized with the oligonucleotide is not cleaved.
27. The method of claim 17, wherein the cleavage site is generated by hybridization with a first oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the adaptor, and the anti-cleavage motif is generated by hybridization with a second oligonucleotide comprising a sequence at least partially complementary to the adapter domain of the adaptor, and wherein cleaving a portion of the physically linked nucleic acid complex amplicon further comprises: (i) Introducing a mixture of the first oligonucleotide and the second oligonucleotide; as well as (ii) Introduce the cutting accelerator.
28. The method of claim 17, wherein the cleavage site is available when the physically linked nucleic acid complex amplicon is in a self-hybridization configuration on the surface.
Citation Information
Patent Citations
Eye state detection apparatus and method of detecting open and closed states of eye
US20130142389A1
Methods of lowering the error rate of massively parallel DNA sequencing using duplex consensus sequencing
US9752188B2
Improved adapters, methods, and compositions for duplex sequencing
WO2017100441A1
Improved adapters, methods, and compositions for duplex sequencing
CN109072294A
Methods for targeted nucleic acid sequence enrichment with applications to error corrected nucleic acid sequencing
WO2018175997A1