Improved adapters, methods, and compositions for duplex sequencing
Patent Information
- Application Number
- CN202610964088.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2016-01-22
- Filing Date
- 2016-12-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]因此,存在并不涉及使用不对称引物结合位点的双重测序途径的未被满足的需要
[0042]本发明的其它特征、优点和修改将从图式、实施方式和权利要求书显而易见。前述描述意图说明且不限制本发明的范围。
Smart Images

Figure CN122833151A_ABST
Abstract
Description
[0001] This application is a divisional application of application number 201680080120.4, filed on December 8, 2016, entitled "Improved adaptor, method and composition for dual sequencing".
[0002] Cross-references to related applications
[0003] This application claims priority and interest in U.S. Provisional Application No. 62 / 264,822, filed December 8, 2015, and U.S. Provisional Application No. 62 / 281,917, filed January 22, 2016. Each of the foregoing applications is incorporated herein by reference in its entirety.
[0004] sequence list
[0005] This application contains a sequence list, which has been submitted by EFS-Web in ASCII format and incorporated herein by full reference. The ASCII copy created on December 8, 2016, is named TWIN-001_ST25.txt and is 11,778 bytes in size. Background Technology
[0006] Duplex sequencing significantly improves the accuracy of high-throughput DNA sequencing by separately amplifying and sequencing both strands of the double-helix DNA; therefore, amplification and sequencing errors can be eliminated when they are typically present on only one of the two strands. Duplex sequencing is first described using asymmetric (i.e., non-complementary) PCR primer binding sites introduced into Y-shaped or "circular" adaptors attached to the ends of DNA fragments. The asymmetric primer binding sites inherent in the adaptor itself produce separate products from both DNA strands, enabling error correction for each. In some cases, using asymmetric primer binding sites may not be optimal; for example, the free ends of the Y-adaptor may tend to be degraded by exonucleases, and these free ends may also anneal to other molecules, producing "daisy-chain" molecules. Furthermore, dual sequencing using Y-shaped or "circular" adaptors is most readily applicable to paired-end sequencing pathways; alternative pathways suitable for single-end sequencing simplify the broader application of dual sequencing across various sequencing platforms.
[0007] Therefore, there is an unmet need for dual sequencing approaches that do not involve the use of asymmetric primer binding sites. Summary of the Invention
[0008] This article describes alternative and superior approaches to dual sequencing that do not require the use of asymmetric primer binding sites. In practice, asymmetry between the two strands can be introduced by generating an adaptor in the DNA molecule to be sequenced or by differentiating the two strands by at least one nucleotide in the DNA sequence of the two strands, or by labeling the two strands differently by other methods, such as attaching the molecule to at least one of the strands (which allows the two strands to be physically separated).
[0009] In a first aspect, the present invention relates to adaptor nucleic acid sequence pairs for sequencing double-stranded target nucleic acid molecules, comprising a first adaptor nucleic acid sequence and a second adaptor nucleic acid sequence, wherein each adaptor nucleic acid sequence comprises a primer-binding domain, a strand-defining element (SDE), a single molecular identifier (SMI) domain, and a linker domain. The SDE of the first adaptor nucleic acid sequence may be at least partially non-complementary to the SDE of the second adaptor nucleic acid sequence.
[0010] In an embodiment of the first aspect, the two adaptor sequences may comprise two separate DNA molecules that are at least partially annealed together. The first and second adaptor nucleic acid sequences may be linked via a linker domain. The linker domain may be composed of nucleotides. The linker domain may comprise one or more modified nucleotides or non-nucleotide molecules. The one or more modified nucleotides or non-nucleotide molecules may be baseless, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5′-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, isocytosine, 5′-methyl-isocytosine, or isoguanosine. The linker domain may form a loop. The SDE of the first adaptor nucleic acid sequence may be non-complementary to the SDE of the second adaptor nucleic acid sequence. The primer-binding domain of the first adaptor nucleic acid sequence may be at least partially complementary to the primer-binding domain of the second adaptor nucleic acid sequence. In an embodiment, the primer-binding domain of the first adaptor nucleic acid sequence may be complementary to the primer-binding domain of the second adaptor nucleic acid sequence. The primer-binding domain of the first adaptor nucleic acid sequence may be at least partially non-complementary to the primer-binding domain of the second adaptor nucleic acid sequence. In an embodiment, at least one SMI domain may be an endogenous SMI, for example, involving a cleavage site (e.g., using the cleavage site itself, using the actual mapped location of the cleavage site (e.g., chromosome 3, positions 1, 234, 567), using a specified number of nucleotides in the DNA immediately adjacent to the cleavage site (e.g., eight nucleotides ten nucleotides away from the cleavage site, seven nucleotides away from the cleavage site, and six nucleotides after the cleavage site following the first appearance of "C")). In an embodiment, the SMI domain includes at least one degenerate or semi-degenerate nucleic acid. In an embodiment, the SMI domain may be non-degenerate. In embodiments, the sequence of the SMI domain may be contemplated to bind to sequences corresponding to the random or semi-random splice ends of the linked DNA to obtain SMI sequences capable of distinguishing individual DNA molecules from each other. The SMI domain of the first adaptor nucleic acid sequence may be at least partially complementary to the SMI domain of the second adaptor nucleic acid sequence. The SMI domain of the first adaptor nucleic acid sequence may be complementary to the SMI domain of the second adaptor nucleic acid sequence. The SMI domain of the first adaptor nucleic acid sequence may be at least partially non-complementary to the SMI domain of the second adaptor nucleic acid sequence. In embodiments, each SMI domain includes a primer binding site. In embodiments, each SMI domain may be located distal to its linker domain. The SMI domain of the first adaptor nucleic acid sequence may be non-complementary to the SMI domain of the second adaptor nucleic acid sequence. In embodiments, each SMI domain includes between about 1 and about 30 degenerate or semi-degenerate nucleic acids. The linker domain of the first adaptor nucleic acid sequence may be at least partially complementary to the linker domain of the second adaptor nucleic acid sequence.In embodiments, each linker domain may be capable of linking to one strand of a double-stranded target nucleic acid sequence. In embodiments, one of the linker domains includes a T-protrusion, an A-protrusion, a CG-protrusion, a blunt end, or another linkable nucleic acid sequence. In embodiments, both linker domains contain blunt ends. In embodiments, at least one of the linker domains includes a modified nucleic acid. The modified nucleotide may be a base-free site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5′-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, isocytosine, 5′-methyl-isocytosine, or isoguanosine. In embodiments, at least one of the linker domains includes a dephosphorylated base. In embodiments, at least one of the linker domains includes a dehydroxylated base. In embodiments, at least one of the linker domains has been chemically modified to make it non-linkable. The SDE of the first adaptor nucleic acid sequence differs from and / or may be non-complementary to the SDE of the second adaptor nucleic acid sequence. In embodiments, at least one nucleotide may be omitted from the SDE of the first adaptor nucleic acid sequence or the SDE of the second adaptor nucleic acid by an enzymatic reaction. The enzymatic reaction includes a polymerase, endonuclease, glycosylation enzyme, or lyase. At least one nucleotide may be a modified nucleotide or include a labeled nucleotide. The modified nucleotide or includes a labeled nucleotide may be a baseless site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5′-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, isocytosine, 5′-methyl-isocytosine, or isoguanosine. The SDE of the first adaptor nucleic acid sequence includes a self-complementary domain capable of forming a hairpin loop. The first adaptor nucleic acid sequence, distal from its linker domain, may link to the second adaptor nucleic acid sequence, possibly distal from its linker domain, thereby forming a loop. The loop includes a restriction enzyme recognition site. In an embodiment, at least the first adaptor nucleic acid sequence further includes a second SDE. The second SDE may be located at the end of the first adaptor nucleic acid sequence. The second adaptor nucleic acid sequence further includes a second SDE. The second SDE may be located at the end of the second adaptor nucleic acid sequence. The second SDE of the first adaptor nucleic acid sequence may be at least partially non-complementary to the second SDE of the second adaptor nucleic acid sequence. The second SDE of the first adaptor nucleic acid sequence differs from and / or may be non-complementary to at least one nucleotide from the second SDE of the second adaptor nucleic acid sequence. In an embodiment, at least one nucleotide may be omitted from the second SDE of the first adaptor nucleic acid sequence or the second SDE of the second adaptor nucleic acid sequence by an enzymatic reaction. The enzymatic reaction includes a polymerase, endonuclease, glycosylation enzyme, or lyase.The second SDE of the first adaptor nucleic acid sequence may not be complementary to the second SDE of the second adaptor nucleic acid sequence. The SDE of the first adaptor nucleic acid sequence may be directly linked to the second SDE of the second adaptor nucleic acid sequence. The primer-binding domain of the first adaptor nucleic acid sequence may be located at the 5′ position of the first SDE. The first SDE of the first adaptor nucleic acid sequence may be located at the 5′ position of the SMI domain. The first SDE of the first adaptor nucleic acid sequence may be located at the 3′ position of the SMI domain. The first SDE of the first adaptor nucleic acid sequence may be located at the 5′ position of the SMI domain and may also be located at the 3′ position of the primer-binding domain. The first SDE of the first adaptor nucleic acid sequence may be located at the 3′ position of the SMI domain (which may also be located at the 3′ position of the primer-binding domain). The SMI domain of the first adaptor nucleic acid sequence may be located at the 5′ position of the linker domain. The 3′ end of the first adaptor nucleic acid sequence includes a linker domain. The first adaptor nucleic acid sequence from 5′ to 3′ includes a primer-binding domain, a first SDE, an SMI domain, and a linker domain. The first adaptor nucleic acid sequence, from 5′ to 3′, includes a primer-binding domain, an SMI domain, a first SDE, and a linker domain. In the embodiments, the first or second adaptor nucleic acid sequence includes a modified nucleotide or non-nucleotide molecule. The modified nucleotide or non-nucleotide molecule may be colicin E2, Im2, glutathione, glutathione S-transferase (GST), nickel, polyhistidine, FLAG-tag, myc-tag, or biotin. Biotin can be in the form of biotin-16-aminoallyl-2'-deoxyuridine-5'-triphosphate, biotin-16-aminoallyl-2'-deoxycytidine-5'-triphosphate, biotin-16-aminoallylcytidine-5'-triphosphate, N4-biotin-OBEA-2'-deoxycytidine-5'-triphosphate, biotin-16-aminoallyluridine-5'-triphosphate, biotin-16-7-denitro-7-aminoallyl-2'-deoxyguanosine-5'-triphosphate, desulfobiotin-6-aminoallyl-2'-deoxycytidine-5'-triphosphate, 5'-biotin-G-monophosphate, 5'-biotin-A-monophosphate, 5'-biotin-dG-monophosphate, or 5'-biotin-dA-monophosphate. Biotin can bind to streptavidin attached to a matrix. In an embodiment, when biotin binds to the antibiotic streptavidin attached to the matrix, the first adaptor nucleic acid sequence can be separated from the second adaptor nucleic acid sequence. In an embodiment, the first or second adaptor nucleic acid sequence includes an affinity label selected from small molecules, nucleic acids, peptides, and uniquely binding portions (which may be capable of binding to an affinity coupler). In an embodiment, when the affinity coupler is attached to a solid matrix and binds to the affinity label, the adaptor nucleic acid sequence including the affinity label can be separated from the adaptor nucleic acid sequence not including the affinity label. The solid matrix can be a solid surface, beads, or another fixed structure.Nucleic acid may be DNA, RNA, or a combination thereof, and optionally includes peptide nucleic acid or locked nucleic acid. The affinity marker may be located at the end of the adaptor or within a domain of a first adaptor nucleic acid sequence that may not be perfectly complementary to a domain in the second adaptor nucleic acid sequence. In embodiments, the first or second adaptor nucleic acid sequence includes a physical group having magnetic, charge-dependent, or insoluble properties. In embodiments, when the physical group has magnetic properties and a magnetic field is applied, the adaptor nucleic acid sequence including the physical group separates from the adaptor nucleic acid sequence not containing the physical group. In embodiments, when the physical group has charge-dependent properties and an electric field is applied, the adaptor nucleic acid sequence including the physical group separates from the adaptor nucleic acid sequence not containing the physical group. In embodiments, when the physical group has insoluble properties and the adaptor nucleic acid sequence pair is contained in a solution insoluble to the physical group, the adaptor nucleic acid sequence including the physical group precipitates from the adaptor nucleic acid sequence not containing the physical group and remains in the solution. The physical group may be located at the end of the adaptor or within a domain of a first adaptor nucleic acid sequence that may not be perfectly complementary to a domain in the second adaptor nucleic acid sequence. The second adaptor nucleic acid sequence includes at least one phosphate-thioester bond. The double-stranded target nucleic acid sequence may be DNA or RNA. In embodiments, each adaptor nucleic acid sequence includes a linker domain at each of its ends. The first or second adaptor nucleic acid sequence may be at least partially single-stranded. The first or second adaptor nucleic acid sequence may be single-stranded. The first and second adaptor nucleic acid sequences may be single-stranded.
[0011] In a second aspect, the present invention relates to a composition comprising at least one pair of adaptor nucleic acid sequences and a second pair of adaptor nucleic acid sequences of the first aspect, wherein each strand of the second pair of adaptor nucleic acid sequences comprises at least one primer binding site and a linker domain.
[0012] The second aspect further relates to a composition comprising at least two pairs of adaptor nucleic acid sequences of the first aspect, wherein the SDE of a first adaptor nucleic acid sequence from the first pair of adaptor nucleic acid sequences is different from the SDE of a first adaptor nucleic acid sequence from at least a second pair of adaptor nucleic acid sequences.
[0013] The second aspect also relates to a composition comprising at least two pairs of adaptor nucleic acid molecules of the first aspect, wherein the SMI domain of the first adaptor nucleic acid molecule from the first pair of adaptor nucleic acid molecules is different from the SMI domain of the first adaptor nucleic acid molecule from at least the second pair of adaptor nucleic acid molecules.
[0014] In an embodiment of the second aspect, the composition further comprises an SMI domain in each strand of the second pair of adaptor nucleic acid sequences. The composition may further comprise a primer binding site in each strand of the second pair of adaptor nucleic acid sequences. The SMI domain of the first adaptor nucleic acid molecule from the first pair of single-stranded adaptor nucleic acid molecules may be of the same length as the SMI domain of the first single-stranded adaptor nucleic acid molecule from at least the second pair of single-stranded adaptor nucleic acid molecules. The SMI domain of the first adaptor nucleic acid molecule from the first pair of single-stranded adaptor nucleic acid molecules may have a different length than the SMI domain of the first single-stranded adaptor nucleic acid molecule from at least the second pair of single-stranded adaptor nucleic acid molecules. In an embodiment, each SMI domain comprises one or more fixed bases within or at a site where an SMI is attached. In an embodiment, at least one first double-stranded complex nucleic acid comprising the first pair of adaptor nucleic acid molecules of the first aspect is attached to a first end of a double-stranded target nucleic acid molecule, and a second pair of adaptor nucleic acid molecules of the first aspect is attached to a second end of a double-stranded target nucleic acid molecule. The first pair of adaptor nucleic acid molecules may be different from the second pair of adaptor nucleic acid molecules. The first strand of the first pair of adaptor nucleic acid molecules includes a first SMI domain, and the first strand of the second pair of adaptor nucleic acid molecules includes a second SMI domain. In an embodiment, the composition includes at least a second double-stranded complex nucleic acid.
[0015] In a third aspect, the present invention relates to a pair of adaptor nucleic acid sequences for sequencing a double-stranded target nucleic acid molecule comprising a first adaptor nucleic acid sequence and a second adaptor nucleic acid sequence. In the third aspect, each adaptor nucleic acid sequence comprises a primer-binding domain and a single molecular identifier (SMI) domain.
[0016] In an embodiment of the third aspect, at least one of the first or second adaptor nucleic acid sequences further includes a domain comprising at least one modified nucleotide. The first and second adaptor nucleic acid sequences further comprise a domain comprising at least one modified nucleotide. In an embodiment, at least one of the first or second adaptor nucleic acid sequences further includes a linker domain. The first and second adaptor nucleic acid sequences may include a linker domain. The at least one modified nucleotide may be a base-free site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5′-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, isocytosine, 5′-methyl-isocytosine, or isoguanosine. The two adaptor sequences may comprise two separate DNA molecules that are at least partially annealed together. The first and second adaptor nucleic acid sequences may be linked via a linker domain. The linker domain may be composed of nucleotides. The linker domain may include one or more modified nucleotides or non-nucleotide molecules. In embodiments, at least one modified nucleotide or non-nucleotide molecule may be a base-free site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5′-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, isocytosine, 5′-methyl-isocytosine, or isoguanosine. The linker domain may form a loop. The primer-binding domain of the first adaptor nucleic acid sequence may be at least partially complementary to the primer-binding domain of the second adaptor nucleic acid sequence. The primer-binding domain of the first adaptor nucleic acid sequence may be complementary to the primer-binding domain of the second adaptor nucleic acid sequence. The primer-binding domain of the first adaptor nucleic acid sequence may not be complementary to the primer-binding domain of the second adaptor nucleic acid sequence. In embodiments, at least one SMI domain is an endogenous SMI, for example, relating to a cleavage site (e.g., using the cleavage site itself, using the actual mapped location of the cleavage site (e.g., chromosome 3, positions 1, 234, 567), using a specified number of nucleotides in the DNA immediately adjacent to the cleavage site (e.g., eight nucleotides ten nucleotides from the cleavage site, seven nucleotides away from the cleavage site, and six nucleotides after the cleavage site following the first appearance of "C")). The SMI domain includes at least one degenerate or semi-degenerate nucleic acid. The SMI domain may be non-degenerate. The sequence of the SMI domain may be considered to bind to sequences corresponding to random or semi-random splice ends of the linked DNA to obtain SMI sequences capable of distinguishing individual DNA molecules from each other. The SMI domain of the first adaptor nucleic acid sequence may be at least partially complementary to the SMI domain of the second adaptor nucleic acid sequence.The SMI domain of the first adaptor nucleic acid sequence may be complementary to the SMI domain of the second adaptor nucleic acid sequence. The SMI domain of the first adaptor nucleic acid sequence may be at least partially non-complementary to the SMI domain of the second adaptor nucleic acid sequence. The SMI domain of the first adaptor nucleic acid sequence may be non-complementary to the SMI domain of the second adaptor nucleic acid sequence. In embodiments, each SMI domain comprises about 1 to about 30 degenerate or semi-degenerate nucleic acids. The linker domain of the first adaptor nucleic acid sequence may be at least partially complementary to the linker domain of the second adaptor nucleic acid sequence. In embodiments, each linker domain may be capable of linking to one strand of the double-stranded target nucleic acid sequence. In embodiments, one of the linker domains includes a T-protrusion, an A-protrusion, a CG-protrusion, a blunt end, or another linkable nucleic acid sequence. In embodiments, both linker domains contain blunt ends. In embodiments, each SMI domain includes a primer binding site. In embodiments, at least the first adaptor nucleic acid sequence further includes an SDE. The SDE may be located at the end of the first adaptor nucleic acid sequence. The second adaptor nucleic acid sequence further includes an SDE. The SDE may be located at the end of the second adaptor nucleic acid sequence. The SDE of the first adaptor nucleic acid sequence may be at least partially non-complementary to the SDE of the second adaptor nucleic acid sequence. The SDE of the first adaptor nucleic acid sequence may be non-complementary to the SDE of the second adaptor nucleic acid sequence. The SDE of the first adaptor nucleic acid sequence may be directly linked to the SDE of the second adaptor nucleic acid sequence. The SDE of the first adaptor nucleic acid sequence differs from and / or may be non-complementary to at least one nucleotide from the SDE of the second adaptor nucleic acid sequence. At least one nucleotide may be omitted from the SDE of the first adaptor nucleic acid sequence or the SDE of the second adaptor nucleic acid by an enzymatic reaction. The enzymatic reaction may include a polymerase or an endonuclease. At least one nucleotide may be a modified nucleotide or include a labeled nucleotide. The modified nucleotide or the labeled nucleotide may be a base-free site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5′-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, isocytosine, 5′-methyl-isocytosine, or isoguanosine. The SDE of the first adaptor nucleic acid sequence may contain a self-complementary domain capable of forming a hairpin loop. The end of the first adaptor nucleic acid sequence remote from its linker domain may be linked to the end of the second adaptor nucleic acid sequence remote from its linker domain, thereby forming a loop. The loop may include a restriction enzyme recognition site. The primer-binding domain of the first adaptor nucleic acid sequence may be located at the 5′ of the SMI domain. The domain of at least one modified nucleotide including the first adaptor nucleic acid sequence may be located at the 5′ of the SMI domain. The domain of at least one modified nucleotide, including the first adaptor nucleic acid sequence, may be located at the 3′ position of the SMI domain.The domain of at least one modified nucleotide comprising the first adaptor nucleic acid sequence may be located at the 5′ of the SMI domain and at the 3′ of the primer-binding domain. The domain of at least one modified nucleotide comprising the first adaptor nucleic acid sequence may be located at the 3′ of the SMI domain (which may be located at the 3′ of the primer-binding domain). The SMI domain of the first adaptor nucleic acid sequence may be located at the 5′ of the linker domain. The 3′ end of the first adaptor nucleic acid sequence may include a linker domain. In an embodiment, the first adaptor nucleic acid sequence from 5′ to 3′ includes a primer-binding domain, a domain comprising at least one modified nucleotide, an SMI domain, and a linker domain. In an embodiment, the first adaptor nucleic acid sequence from 5′ to 3′ includes a primer-binding domain, an SMI domain, a domain comprising at least one modified nucleotide, and a linker domain. In an embodiment, the first adaptor nucleic acid sequence or the second adaptor nucleic acid sequence comprises a modified nucleotide or a non-nucleotide molecule. The modified nucleotide or non-nucleotide molecule may be coliform E2, Im2, glutathione, glutathione S-transferase (GST), nickel, polyhistidine, FLAG-tag, myc-tag, or biotin. Biotin can be in the form of biotin-16-aminoallyl-2'-deoxyuridine-5'-triphosphate, biotin-16-aminoallyl-2'-deoxycytidine-5'-triphosphate, biotin-16-aminoallylcytidine-5'-triphosphate, N4-biotin-OBEA-2'-deoxycytidine-5'-triphosphate, biotin-16-aminoallyluridine-5'-triphosphate, biotin-16-7-denitro-7-aminoallyl-2'-deoxyguanosine-5'-triphosphate, desulfobiotin-6-aminoallyl-2'-deoxycytidine-5'-triphosphate, 5'-biotin-G-monophosphate, 5'-biotin-A-monophosphate, 5'-biotin-dG-monophosphate, or 5'-biotin-dA-monophosphate. Biotin can bind to streptavidin attached to a matrix. In an embodiment, when biotin binds to the antibiotic streptavidin attached to the matrix, the first adaptor nucleic acid sequence can be separated from the second adaptor nucleic acid sequence. The second adaptor nucleic acid sequence may include at least one phosphate-thioester bond. The double-stranded target nucleic acid sequence may be DNA or RNA. In an embodiment, the first or second adaptor nucleic acid sequence includes an affinity label selected from small molecules, nucleic acids, peptides, and uniquely binding portions (capable of binding to affinity couplers). In an embodiment, when the affinity coupler is attached to a solid matrix and binds to the affinity label, the adaptor nucleic acid sequence including the affinity label can be separated from the adaptor nucleic acid sequence not including the affinity label. The solid matrix may be a solid surface, beads, or another immobilized structure. The nucleic acid may be DNA, RNA, or a combination thereof, and optionally include peptide nucleic acids or locked nucleic acids.The affinity marker may be located at the end of the adaptor or within a domain of the first adaptor nucleic acid sequence that may not be completely complementary to the relative domain of the second adaptor nucleic acid sequence. In embodiments, the first or second adaptor nucleic acid sequence includes a physical group having magnetic, charge-dependent, or insoluble properties. In embodiments, when the physical group has magnetic properties and a magnetic field is applied, the adaptor nucleic acid sequence including the physical group separates from the adaptor nucleic acid sequence not containing the physical group. In embodiments, when the physical group has charge-dependent properties and an electric field is applied, the adaptor nucleic acid sequence including the physical group separates from the adaptor nucleic acid sequence not containing the physical group. In embodiments, when the physical group has insoluble properties and the adaptor nucleic acid sequence pair is contained in a solution insoluble to the physical group, the adaptor nucleic acid sequence including the physical group precipitates from the adaptor nucleic acid sequence not containing the physical group and remains in the solution. The physical group may be located at the end of the adaptor or within a domain of the first adaptor nucleic acid sequence that may not be completely complementary to the relative domain of the second adaptor nucleic acid sequence. The first or second adaptor nucleic acid sequence may be at least partially single-stranded. The first or second adaptor nucleic acid sequence may be single-stranded. In an embodiment, at least one of the linker domains includes a dehydroxylated base. In an embodiment, at least one of the linker domains has been chemically modified to make it non-linkable.
[0017] In a fourth aspect, the present invention relates to a composition comprising at least two pairs of adaptor nucleic acid molecules of the third aspect, wherein the SMI domain of the first adaptor nucleic acid molecule from the first pair of adaptor nucleic acid molecules is different from the SMI domain of the first adaptor nucleic acid molecule from at least the second pair of adaptor nucleic acid molecules.
[0018] In an embodiment of the fourth aspect, the SMI domain of the first adaptor nucleic acid molecule from the first pair of single-stranded adaptor nucleic acid molecules may have the same length as the SMI domain of the first single-stranded adaptor nucleic acid molecule from at least the second pair of single-stranded adaptor nucleic acid molecules. The SMI domain of the first adaptor nucleic acid molecule from the first pair of single-stranded adaptor nucleic acid molecules may have a different length than the SMI domain of the first single-stranded adaptor nucleic acid molecule from at least the second pair of single-stranded adaptor nucleic acid molecules. In an embodiment, each SMI domain includes one or more fixed bases within the SMI or at a site where the SMI is adjacent.
[0019] In a fifth aspect, the present invention relates to a composition comprising at least one first double-stranded complex nucleic acid, comprising a first pair of adaptor nucleic acid molecules attached to a third aspect of a first end of a double-stranded target nucleic acid molecule and a second pair of adaptor nucleic acid molecules attached to a third aspect of a second end of a double-stranded target nucleic acid molecule.
[0020] In an embodiment of the fifth aspect, the first pair of adaptor nucleic acid molecules may differ from the second pair of adaptor nucleic acid molecules. The first strand of the adaptor target nucleic acid molecule in the first pair of adaptor nucleic acids may include a first SMI domain, and the first strand of the adaptor target nucleic acid molecule in the second pair of adaptor nucleic acids may include a second SMI domain. In an embodiment, the composition comprises at least a second double-stranded complex nucleic acid.
[0021] In a sixth aspect, the present invention relates to a composition comprising at least one pair of adaptor nucleic acid molecules of the first aspect and at least one pair of adaptor nucleic acid molecules of the third aspect.
[0022] In a seventh aspect, the present invention relates to a composition comprising at least one first double-stranded complex nucleic acid, comprising a first pair of adaptor nucleic acid molecules attached to a first aspect of a first end of a double-stranded target nucleic acid molecule and a second pair of adaptor nucleic acid molecules attached to a third aspect of a second end of a double-stranded target nucleic acid molecule.
[0023] In an eighth aspect, the present invention relates to a method for sequencing a double-stranded target nucleic acid, comprising the following steps: (1) linking a pair of adaptor nucleic acid sequences of the first aspect to at least one end of a double-stranded target nucleic acid molecule, thereby forming a double-stranded nucleic acid molecule comprising a first-strand adaptor target nucleic acid sequence and a second-strand adaptor target nucleic acid sequence; (2) amplifying the first-strand adaptor target nucleic acid sequence, thereby producing a first set of amplified products comprising multiple first-strand adaptor target nucleic acid sequences and multiple complementary molecules thereof; (3) amplifying the second-strand adaptor target nucleic acid sequence, thereby producing a second set of amplified products comprising multiple second-strand adaptor target nucleic acid sequences and multiple complementary molecules thereof, wherein the second set of amplified products is distinguishable from the first set of amplified products; (4) sequencing the first set of amplified products; and (5) sequencing the second set of amplified products.
[0024] In an embodiment of the eighth aspect, at least one end may be two ends. Amplification may be performed by PCR, by multiple displacement amplification, or by isothermal amplification. The adaptor nucleic acid sequence pair attached to the first end of the double-stranded target nucleic acid sequence has the same structure as the adaptor nucleic acid sequence pair attached to the second end of the double-stranded target nucleic acid sequence. In an embodiment of the eighth aspect, the first-strand adaptor target nucleic acid sequence, arranged in 5′ to 3′ order, comprises: (a) the first adaptor nucleic acid sequence, (b) the first strand of the double-stranded target nucleic acid, and (c) the second adaptor nucleic acid sequence. In an embodiment of the eighth aspect, the second-strand adaptor target nucleic acid sequence may be arranged in 3′ to 5′ order, comprising: (a) the first adaptor nucleic acid sequence, (b) the second strand of the double-stranded target nucleic acid, and (c) the second adaptor nucleic acid sequence. The adaptor nucleic acid sequence pair attached to the first end of the double-stranded target nucleic acid sequence may be different from the adaptor nucleic acid sequence pair attached to the second end of the double-stranded target nucleic acid sequence. The adaptor nucleic acid sequence pair attached to the first end of the double-stranded target nucleic acid sequence has a first SMI domain, and the adaptor nucleic acid sequence pair attached to the second end of the double-stranded target nucleic acid sequence has a second SMI domain, wherein the first SMI domain may be different from the second SMI domain. In an embodiment of the eighth aspect, the first strand adaptor target nucleic acid sequence may be ordered from 5′ to 3′ to include: (a) a first adaptor nucleic acid sequence including a first SDE, (b) a first SMI domain, (c) a first strand of the double-stranded target nucleic acid, and (d) a second adaptor nucleic acid sequence. In an embodiment of the eighth aspect, the second strand adaptor target nucleic acid sequence may be ordered from 5′ to 3′ to include: (a) a first adaptor nucleic acid sequence including a first SDE, (b) a second SMI domain, (c) a second strand of the double-stranded target nucleic acid, and (d) a second adaptor nucleic acid sequence. In an embodiment, a common sequence for the first set of amplification products may be compared with a common sequence for the second set of amplification products, and the difference between the two common sequences may be considered an artifact.
[0025] In a ninth aspect, the present invention relates to a method for sequencing double-stranded target nucleic acids, comprising the following steps: (1) ligating a pair of adaptor nucleic acid sequences of the third aspect to at least one end of a double-stranded target nucleic acid molecule, thereby forming a double-stranded nucleic acid molecule comprising a first-strand adaptor target nucleic acid sequence and a second-strand adaptor target nucleic acid sequence; (2) amplifying the first-strand adaptor target nucleic acid molecule, thereby generating a first set of amplified products comprising a plurality of first-strand adaptor target nucleic acid molecules and a plurality of their complementary molecules; (3) amplifying the second-strand adaptor target nucleic acid molecule, thereby generating a second set of amplified products comprising a plurality of second-strand adaptor target nucleic acid molecules and a plurality of their complementary molecules; (4) sequencing the first set of amplified products, thereby obtaining a common sequence for the first set of amplified products; and (5) sequencing the second set of amplified products, thereby obtaining a common sequence for the second set of amplified products.
[0026] In an embodiment of the ninth aspect, the second set of amplification products may be distinguished from the first set of amplification products. Amplification may be performed by PCR, by multiple displacement amplification, or by isothermal amplification. In an embodiment of the ninth aspect, the method further includes, after step (1), contacting the double-stranded nucleic acid molecule with at least one enzyme (e.g., a glycosylation enzyme) that alters at least one modified nucleotide to another chemical structure. The adaptor nucleic acid sequence pair attached to the first end of the double-stranded target nucleic acid molecule may be identical to the adaptor nucleic acid sequence pair attached to the second end of the double-stranded target nucleic acid molecule. The adaptor nucleic acid sequence pair attached to the first end of the double-stranded target nucleic acid molecule may be different from the adaptor nucleic acid sequence pair attached to the second end of the double-stranded target nucleic acid molecule. In an embodiment, a pair of adaptor nucleic acid sequences may be attached to the first end of the double-stranded target nucleic acid molecule and may be used with primers corresponding to a portion of the DNA sequence of the target DNA molecule to amplify the DNA molecule. In an embodiment of the ninth aspect, the first strand adaptor target nucleic acid sequence, arranged in 5′ to 3′ order, comprises: (a) a first adaptor nucleic acid sequence including at least one modified nucleotide or at least one base-free site, (b) the first strand of the double-stranded target nucleic acid, and (c) a second adaptor nucleic acid sequence. In an embodiment of the ninth aspect, the second-stranded adaptor target nucleic acid sequence, arranged in 3′ to 5′ order, comprises: (a) a first adaptor nucleic acid sequence, (b) the second strand of the double-stranded target nucleic acid, and (c) a second adaptor nucleic acid sequence. The adaptor nucleic acid sequence pair attached to the first end of the double-stranded target nucleic acid molecule may differ from the adaptor nucleic acid sequence pair attached to the second end of the double-stranded target nucleic acid molecule. The adaptor nucleic acid sequence pair attached to the first end of the double-stranded target nucleic acid molecule has a first SMI domain, and the adaptor nucleic acid sequence pair attached to the second end of the double-stranded target nucleic acid sequence has a second SMI domain, wherein the first SMI domain may differ from the second SMI domain. In an embodiment of the ninth aspect, the first-stranded adaptor target nucleic acid sequence, arranged in 5′ to 3′ order, comprises: (a) a first adaptor nucleic acid sequence comprising at least one modified nucleotide or at least one base-free site and a first SMI domain, (b) the first strand of the double-stranded target nucleic acid, and (c) a second adaptor nucleic acid sequence comprising a second SMI domain. In an embodiment, when at least one modified nucleotide may be 8-oxo-G, the second adaptor nucleic acid sequence includes cytosine at the position corresponding to 8-oxo-G. In an embodiment of the ninth aspect, the second-strand adaptor target nucleic acid sequence, ordered from 3′ to 5′, comprises: (a) a first adaptor nucleic acid sequence including a first SMI domain, (b) a second strand of the double-stranded target nucleic acid, and (c) a second adaptor nucleic acid sequence including a second SMI domain. In an embodiment, at least one modified nucleotide may be 8-oxo-G, and the second adaptor nucleic acid sequence includes cytosine at the position corresponding to 8-oxo-G.In an embodiment, during the amplification in step (2) or step (3), at least one base-free site may be converted into thymidine in the corresponding amplification product after amplification, thereby introducing SDE. In an embodiment of the ninth aspect, during the amplification in step (2) or step (3), at least one modified nucleotide site encodes adenosine in the corresponding amplification product.
[0027] In a tenth aspect, the present invention relates to a method in which distinguishable amplification products are available from each of the two strands of an individual DNA molecule, and a common sequence for a first set of amplification products is comparable to a common sequence for a second set of amplification products, wherein the difference between the two common sequences can be considered as a pseudo-error.
[0028] In an embodiment of the tenth aspect, by means of sharing the same SMI sequence, it can be determined that the amplification products originated from the same initial DNA molecule. In an embodiment, by means of carrying dissimilar SMI sequences that may be known to correspond to each other, based on a database generated and bound to the SMI adaptor library during synthesis, it can be determined that the amplification products originated from dissimilar strands of the same initial double-stranded DNA sequence. In an embodiment, by means of at least one nucleotide sequence difference introduced through SDE, it can be determined that the amplification products originated from dissimilar strands of the same initial double-stranded DNA sequence.
[0029] In an eleventh aspect, the present invention relates to a method in which distinguishable amplification products are available from each of the two strands of an individual DNA molecule, and a sequence of the amplification product corresponding to one of the two initial DNA strands of a single DNA molecule is compared with an amplification product corresponding to the second of the two initial DNA strands, and the difference between the two sequences is considered a pseudo-difference.
[0030] In a twelfth aspect, the present invention relates to a method wherein, when a sequence of an amplified product corresponding to one of two initial DNA strands of a single DNA molecule is compared with an amplified product corresponding to the second of the two initial DNA strands, indistinguishable amplified products can be obtained from the two strands of an individual DNA molecule and no difference can be identified between the two sequences.
[0031] In an embodiment of the twelfth aspect, by means of a shared SMI sequence, based on a database generated and bound to the SMI adaptor library during synthesis, the amplification product can be determined to originate from the same initial double-stranded DNA molecule. In an embodiment, the amplification product can be determined to originate from dissimilar strands of the same initial double-stranded DNA sequence via at least one nucleotide sequence difference introduced by an SDE. In an embodiment, the method further includes a step of monomolecular dilution after the DNA double helix has been thermally or chemically melted into its component single strands. The single strands can be diluted into a plurality of physically separated reaction chambers to minimize the probability of two initially paired strands sharing the same container. The physically separated reaction chambers can be selected from containers, sleeves, wells, and at least one pair of non-connected droplets. In an embodiment, PCR amplification can be performed for each physically separated reaction chamber, preferably using primers for each chamber carrying different tag sequences. In an embodiment, each tag sequence acts as an SDE. In an embodiment, a series of paired sequences corresponding to the two strands of the same initial DNA can be compared with each other, and at least one sequence of a series of products can be selected as the correct sequence most likely representing the initial DNA molecule. This can be at least partly attributed to having the minimum number of mismatches between the products obtained from the two DNA strands, selecting the product that most likely represents the correct sequence of the initial DNA molecule. This can also be at least partly attributed to having the minimum number of mismatches relative to the reference sequence, selecting the product that most likely represents the correct sequence of the initial DNA molecule.
[0032] In a thirteenth aspect, the present invention relates to a composition comprising at least two pairs of adaptor nucleic acid sequences, wherein the first pair of adaptor nucleic acid sequences comprises a primer-binding domain, a strand-limiting element (SDE), and a linker domain, and wherein the second pair of adaptor nucleic acid sequences comprises a primer-binding domain, a single molecular identifier (SMI) domain, and a linker domain.
[0033] In a fourteenth aspect, the present invention relates to a double-stranded complex nucleic acid comprising: (1) a first pair of adaptor nucleic acid sequences including a primer-binding domain and an SDE; and (2) a double-stranded target nucleic acid; and (3) a second pair of adaptor nucleic acid sequences including a primer-binding domain and a single molecular identifier (SMI) domain, wherein the first pair of adaptor nucleic acid molecules is connectable to a first end of the double-stranded target nucleic acid molecule, and the second pair of adaptor nucleic acid molecules is connectable to a second end of the double-stranded target nucleic acid molecule. In embodiments of the fourteenth aspect, the first pair of adaptor nucleic acid sequences and / or the second pair of adaptor nucleic acid sequences may further include a linker domain.
[0034] In a fifteenth aspect, the present invention relates to an adaptor nucleic acid sequence pair for sequencing a double-stranded target nucleic acid molecule comprising a first adaptor nucleic acid sequence and a second adaptor nucleic acid sequence, wherein each adaptor nucleic acid sequence comprises: a primer-binding domain, an SDE, and a linker domain, wherein the SDE of the first adaptor nucleic acid sequence may be at least partially non-complementary to the SDE of the second adaptor nucleic acid sequence.
[0035] In a sixteenth aspect, the present invention relates to a double-stranded circular nucleic acid comprising a pair of adaptor nucleic acid molecules having a first aspect attached to a first end of a double-stranded target nucleic acid molecule and a second aspect attached to a second end of a double-stranded target nucleic acid molecule.
[0036] In a seventeenth aspect, the present invention relates to a double-stranded circular nucleic acid comprising a pair of adaptor nucleic acid molecules of a third aspect, which are attached to a first end of a double-stranded target nucleic acid molecule and to a second end of a double-stranded target nucleic acid molecule.
[0037] In an eighteenth aspect, the present invention relates to a double-stranded circular nucleic acid comprising a pair of adaptor nucleic acid molecules of a first aspect attached to a first end of a double-stranded target nucleic acid molecule and an annealed primer-binding domain pair attached to a second end of the double-stranded target nucleic acid molecule, wherein the annealed primer-binding domain pair is connectable to the pair of adaptor nucleic acid molecules.
[0038] In a nineteenth aspect, the present invention relates to a double-stranded circular nucleic acid comprising a pair of adaptor nucleic acid molecules on a third aspect attached to a first end of a double-stranded target nucleic acid molecule and an annealed primer-binding domain pair attached to a second end of the double-stranded target nucleic acid molecule, wherein the annealed primer-binding domain pair is connectable to the adaptor nucleic acid molecule pair.
[0039] In a twentieth aspect, the present invention relates to a double-stranded complex nucleic acid comprising: (1) a pair of adaptor nucleic acid sequences, each comprising a primer-binding domain, a chain-limiting element (SDE), and a single molecular identifier (SMI) domain; (2) a double-stranded target nucleic acid; and (3) an annealed primer-binding domain pair, wherein the adaptor nucleic acid pair is ligable to a first end of the double-stranded target nucleic acid molecule and the annealed primer-binding domain pair is ligable to a second end of the double-stranded target nucleic acid molecule. In an embodiment of the twentieth aspect, the adaptor nucleic acid sequence pair and / or the annealed primer-binding domain pair further comprises a linker domain.
[0040] Dual sequencing is also described in WO2013142389A1 and Schmitt et al., PNAS 2012, each of which is incorporated herein by reference in full.
[0041] Any of the above aspects and embodiments can be combined with any other aspects or embodiments disclosed in the present invention, drawings and / or implementation (including the following specific non-limiting examples / embodiments of the invention).
[0042] Other features, advantages, and modifications of the invention will be apparent from the drawings, embodiments, and claims. The foregoing description is intended to illustrate but not limit the scope of the invention. Attached Figure Description
[0043] The above and other features will become clearer from the following detailed description taken in conjunction with the accompanying drawings.
[0044] Figures 1A to 1I This illustrates dual sequencing using the initial description of the Y-adaptor. An exemplary Y-adaptor is shown. Figure 1A ), double-stranded DNA molecules connected to such adaptors ( Figure 1B ), and PCR products derived from it ( Figure 1C and Figure 1D ) and the resulting sequencing reads ( Figures 1E to 1I ).
[0045] Figures 2A to 2K This illustrates the dual sequencing of the present invention using a non-complementary "bubble" adaptor. An exemplary "bubble" adaptor is shown. Figure 2A and Figures 2H to 2K ), connected to Figure 2A The double-stranded DNA molecule of the adaptor ( Figure 2B ), and PCR products derived from it ( Figure 2C and Figure 2D ) and the resulting sequencing reads ( Figures 2E to 2G ).
[0046] Figures 3A to 3G This invention illustrates dual sequencing using an adaptor with a non-complementary "bubble" shaped single-molecule identifier (SMI) (which together serve as a molecular identifier) and an asymmetric introduction of a chain-limiting element (SDE). An exemplary "bubble" adaptor is shown. Figure 3A ), connected to Figure 3A The double-stranded DNA molecule of the adaptor ( Figure 3B ), and PCR products derived from it ( Figure 3C and Figure 3D ) and the resulting sequencing reads ( Figure 3E and Figure 3F ). Figure 3G This shows the grouping of specific SMI sequences and their corresponding non-complementary mating bodies. Figure 3E and Figure 3F The sequencing reads.
[0047] Figures 4A to 4HThis invention illustrates dual sequencing using an adaptor having a nucleotide or nucleotide analogue (which initially forms paired strands of DNA but subsequently confers DNA mismatch after a subsequent biochemical reaction). An exemplary adaptor containing 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G) is shown. Figure 4A ), connected to Figure 4A The double-stranded DNA molecule of the adaptor ( Figure 4B ), Figure 4C This is shown after treatment with an enzyme that produces a base-free site that replaces the 8-oxo-G base and the resulting mismatch in the adaptor. Figure 4B Double-stranded DNA molecules; PCR products derived from them ( Figure 4D and Figure 4E ); and the resulting sequencing reads ( Figures 4F to 4H ).
[0048] Figures 5A to 5H This invention illustrates the use of combinations of dual sequencing adaptors designed to introduce different primer sites at opposite ends of a DNA molecule in dual sequencing. An exemplary dual sequencing adaptor is shown. Figure 5A ) and "standard" connector ( Figure 5B );when Figure 5A and Figure 5B The three types of double-stranded DNA molecules produced when an adaptor is attached to a DNA molecule ( Figures 5C to 5E ); PCR products derived from it ( Figure 5F and Figure 5G ); and the resulting sequencing reads ( Figure 5H ).
[0049] Figures 6A to 6I This invention illustrates the dual sequencing method using combinations of dual sequencing adaptors that allow for the design of two reads on non-end-of-end platforms. The "standard" adaptor is shown. Figure 6A ) and exemplary dual sequencing adaptors ( Figure 6B );when Figure 6A and Figure 6B The preferred double-stranded DNA molecule produced when the adaptor is attached to a DNA molecule ( Figure 6C ); PCR products derived from it ( Figure 6D and Figure 6E ); used for sequencing derived from the "top" strand ( Figure 6F ) and "bottom" chain ( Figure 6G The arrangement of the template strand; and the resulting sequencing reads ( Figure 6H and Figure 6I ).
[0050] Figures 7A to 7IThis invention illustrates dual sequencing using combinations of dual sequencing adaptors that allow for the design of two reads on non-end-of-end platforms. It also shows adaptors that further include degenerate or semi-degenerate SMI sequences. Figure 7A ) and exemplary dual sequencing adaptors ( Figure 7B );when Figure 7A and Figure 7B The preferred double-stranded DNA molecule produced when the adaptor is attached to a DNA molecule ( Figure 7C ); PCR products derived from it ( Figure 7D and Figure 7E ); used for sequencing derived from the "top" strand ( Figure 7F ) and "bottom" chain ( Figure 7G The arrangement of the template strand; and the resulting sequencing reads ( Figure 7H and Figure 7I ).
[0051] Figures 8A to 8J This invention illustrates dual sequencing using a Y-shaped dual sequencing adaptor with asymmetric SMI. An exemplary dual sequencing adaptor is shown. Figure 8A ),when Figure 8A The double-stranded DNA molecule produced when the adaptor is attached to a DNA molecule ( Figure 8B ), and PCR products derived from it ( Figure 8C and Figure 8D ) and the resulting sequencing reads ( Figure 8E and Figure 8F ). Figure 8G This shows the grouping of specific SMI sequences and their corresponding non-complementary mating bodies. Figure 8E and Figure 8F The sequencing reads. Figures 8H to 8J This illustrates an alternative connector design suitable for this embodiment.
[0052] Figures 9A to 9G This invention illustrates dual sequencing using Y-shaped or circular dual sequencing adaptors with asymmetric SMI located in a region without a single-stranded tail. An exemplary dual sequencing adaptor is shown. Figure 9A ),when Figure 9A The preferred double-stranded DNA molecule produced when the adaptor is attached to a DNA molecule ( Figure 9B ), and PCR products derived from it ( Figure 9C and Figure 9D The orientation of sequencing primer sites and index primer sites is displayed in Figure 9E and Figure 9F middle. Figure 9G show Figure 9E and Figure 9F The grouped sequencing reads obtained by the method shown.
[0053] Figures 10A to 10EThis invention describes dual sequencing, wherein all elements necessary for dual sequencing are included in a single molecule rather than in two paired adaptors. Figure 10A This configuration is shown before the ligation of double-stranded DNA molecules and Figure 10B Shown after the ligation of double-stranded DNA molecules Figure 10A Configuration. Figures 10C to 10E Some alternatives to this embodiment are shown.
[0054] Figures 11A to 11D This illustrates dual sequencing via asymmetric chemical tagging and strand separation. An exemplary dual sequencing adaptor with a chemical tag (biotin in this case) is shown. Figure 11A ) and second connector ( Figure 11B );when Figure 11A connector and Figure 11B The preferred double-stranded DNA molecule produced when the adaptor is attached to a DNA molecule ( Figure 11C ), and other steps in the method, including separating the chemically tagged strands from other strands and amplifying and sequencing them independently. Figure 11D ).
[0055] Figures 12A to 12M This invention illustrates dual sequencing, in which SDE is introduced via nick translation. Figures 12A to 12D The connector design is shown, where the SDE is lost after cutout translation. The Ion Torrent™-compatible connector suitable for this embodiment is shown. Figure 12E and Figure 12F );when Figure 12E and Figure 12F The preferred double-stranded DNA molecule produced when the adaptor is attached to a DNA molecule ( Figure 12G Misincorporation of terminal nucleotides (); Figure 12H ); and its derived derivatives, which exhibit mismatch ( Figure 12I ); derived from Figure 12I PCR products of the molecule ( Figure 12J and Figure 12K ); and the resulting sequencing reads ( Figure 12L and Figure 12M ).
[0056] Figures 13A to 13G This invention illustrates dual sequencing, wherein the SDE is introduced after nick translation. A dual sequencing adaptor containing a dephosphorylated 5′ end is shown. Figure 13A );when Figure 13A The double-stranded DNA molecule produced when the adaptor is attached to a DNA molecule ( Figure 13B ); the structure after chain substitution synthesis ( Figure 13C ); Figure 13C The extended products of the structure ( Figure 13DThe results showed no mismatches; the structure including the nick was observed after treatment with uracil DNA glycosylase and appropriate AP endonuclease. Figure 13E After the gap has been filled with mismatched nucleotides and the connection is closed; Figure 13E Structure ( Figure 13F ); and the resulting sequencing reads ( Figure 13G ).
[0057] Figures 14A to 14I This invention illustrates dual sequencing, wherein mismatches are introduced by extending a polymerase into the DNA molecule to be sequenced. The image shows the double-stranded DNA molecule to be sequenced (…). Figure 14A Endonuclease treatment that has already left the 5′ overhang Figure 14A Double-stranded DNA molecules ( Figure 14B ); Figure 14B Part of the double-stranded DNA molecule was treated to introduce two mismatches ( Figure 14C ), Figure 14C The extended products of the structure ( Figure 14D ), which now includes "bubbles" at every mismatch; shown in Figure 14E A pair of connectors in; when Figure 14E The connector is connected to Figure 14D DNA molecules are produced Figure 14F The structure; derived from Figure 14F PCR products of the molecule ( Figure 14G and Figure 14H ); and the resulting sequencing reads ( Figure 14I ). Detailed Implementation
[0058] First, duplex sequencing is described using asymmetric primer binding sites for separate amplification of the two DNA strands. Alternative and preferred approaches to duplex sequencing that do not necessarily require the use of asymmetric primer binding sites are described herein. In practice, asymmetry between the two strands can be introduced by generating a difference in at least one nucleotide in the DNA sequence between the two strands (e.g., mismatches, extra nucleotides, and omitted nucleotides) within the adaptor or elsewhere in the DNA molecule to be sequenced, by substituting at least one nucleotide with a modified nucleotide (e.g., a nucleotide lacking a base or having an atypical base), and / or by including at least one labeled nucleotide (e.g., a biotinylated nucleotide) that can physically separate the two strands. Table 1 illustrates exemplary options for assembling adaptors for duplex sequencing as disclosed herein.
[0059] Table 1:
[0060]
[0061] Note:
[0062] (i) All of these adapter designs may have additional optional elements added (e.g., two adapter strands linked together and utilizing PCR primer sites in various configurations).
[0063] (ii) Whenever an SMI is used, it can be random / degenerate, semi-random / semi-degenerate, or predefined. Furthermore, if the SMI contains two chains, the two chains can be complementary, non-complementary, or partially complementary.
[0064] (iii) A fully adapted molecular complex containing at least one SDE and at least one SMI may be present in the adaptor and / or the DNA ligated before attachment may be generated after ligation or may be combined with it.
[0065] The adaptor design and approach for dual sequencing described in this article do not depend on using a Y adaptor with a complementary SMI sequence.
[0066] Some designs are directly applicable to single-end sequencing. The approaches disclosed herein share two common features: (1) labeling each single strand of an individual half-double-helix DNA molecule in such a way that the sequence ultimately derived from each of the two strands is identifiable as being associated with the same DNA double helix; and (2) labeling each single strand of an individual double-helix DNA molecule in such a way that the sequence ultimately derived from each of the two strands is identifiable as being different from those derived from the opposite strand. The molecular features that provide these corresponding functions are named in this document as single-molecule identifiers (SMIs) and strand-defining elements (SDEs).
[0067] This is the first publicly disclosed method for introducing strand-defined asymmetry via internal non-complementary "bubble" sequences of different forms. One such embodiment involves introducing a non-complementary "bubble" sequence not located within the amplification primer site; the sequence, distinct from the two strands of the "bubble," will subsequently produce separate labels for the two strands.
[0068] This document discloses how chain-limited asymmetry can be similarly introduced into adapted DNA molecules by using modified DNA bases as the SDE. In an example, the asymmetry is introduced by including one or more nucleotide analogs that initially produce a complementary sequence but can subsequently be converted to a non-complementary sequence.
[0069] The method for designing a non-Y-shaped asymmetric adaptors that can be applied to sequencing platforms that require different primer sequences at opposite ends of each DNA molecule is also disclosed.
[0070] This paper discloses an alternative approach in which different types of SMI tags and SDEs can be distributed in two different primer sites containing adaptors to maximize read length and SMI tag diversity.
[0071] This paper also discloses additional designs for dual-sequencing adaptors containing Y or circular tails that are readily adaptable to paired-end sequencing, but in which the SMI tag is a non-complementary sequence, thus allowing for significant design flexibility.
[0072] This paper demonstrates how the introduction of such asymmetry can enable the differentiation of products from the two DNA strands via duplex sequencing for error correction purposes. Furthermore, this paper demonstrates how some embodiments facilitate the description of duplex sequencing on single-end read platforms.
[0073] A method for introducing primer sites and SMI and SDE sites for dual sequencing of a single adaptor to form a circular adaptor-DNA molecular complex is further disclosed.
[0074] In addition, a completely different approach to introducing SDE is disclosed, which relies on asymmetric chemical labeling, allowing paired chains to be physically / mechanically separated into different reaction compartments for independent analysis, rather than molecular labeling based on the differential sequence of the two chains.
[0075] This article discloses examples of adaptor design, particularly for the Ion Torrent™ (Life Technologies®) sequencing platform.
[0076] This paper discloses a variant of a linker that can be attached to two single strands at each end of a double helix molecule and a design that allows single-strand linkage followed by “cutaway translation”, which retains the necessary SMI and SDE elements in the final molecule.
[0077] This article discloses how SDEs can be incorporated into DNA molecules independently of adaptor ligation.
[0078] Finally, this paper discloses a streamlined alternative algorithmic approach for dual sequencing that can be used with any double-helix adaptor design that eliminates the need for the aforementioned single-stranded common sequence (SSCS) generation.
[0079] In some embodiments, a portion of the nucleotide sequence may be "degenerate." In a degenerate sequence, each position may be any nucleotide, that is, each position represented by "X," "N," or "M" may be adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U), or any other natural or non-natural DNA or RNA nucleotide or nucleotide analogue or analogue having base-pairing properties (e.g., xanthoside, inosine, hypoxanthine, xanthine, 7-methylguanine, 7-methylguanosine, 5,6-dihydrouracil, 5-methylcytosine, dihydrouridine, isocytosine, isoguanine, deoxynucleosides, nucleosides, peptide nucleic acids, locked nucleic acids, diol nucleic acids, and threonic acid). Alternatively, a portion of the nucleotide sequence may be incompletely degenerate such that the sequence includes at least one predefined nucleotide or at least one predefined polynucleotide and position, which may be any nucleotide or one or more positions including only a subset of possible nucleotide combinations. Possible subsets of nucleotides may include: any three of the following: A, C, G, and T; any two of the following: A, C, G, T, and U; or U plus any three of the following: A, C, G, and T. Such subsets may additionally include any other natural or non-natural DNA or RNA nucleotides or nucleotide analogs or substitutes thereof that have base-pairing properties. The stoichiometric ratio between any of these nucleotides in the molecular population may be about 1:1 or any other ratio; such sequences are referred to herein as “semi-degenerate.” In some embodiments, a “semi-degenerate” sequence refers to a group of two or more sequences in which two or more sequences differ at at least one nucleotide position. In embodiments, a semi-degenerate sequence is a sequence in which not every nucleotide is random relative to its adjacent nucleotides (immediately adjacent or within two or more nucleotides). In embodiments, as used herein, the terms degenerate and semi-degenerate may have the same meaning as commonly understood and used by one of ordinary skill in the art to which this application pertains; such art is incorporated herein by reference in its entirety.
[0080] In this embodiment, the sequence need not contain all possible bases at every position. Degenerate or semi-degenerate n-mer sequences can be generated by polymerase-mediated methods or by preparing and annealing individual oligonucleotide libraries of known sequences. Alternatively, any degenerate or semi-degenerate n-mer sequence can be a randomly or non-randomly fragmented double-stranded DNA molecule from any alternative source different from the target DNA source. In some embodiments, the alternative source is a genome or plasmid derived from bacteria, an organism other than the target DNA, or a combination of such alternative organisms or sources. Randomly or non-randomly fragmented DNA can be introduced into an SMI adaptor to act as a variable tag. This can be achieved via enzyme ligation or any other method known in the art.
[0081] Unless the context clearly specifies otherwise, as used in this specification and the appended claims, the singular forms “a (an)” and “the” include a plurality of indicators.
[0082] Unless otherwise stated or the context clearly indicates, as used herein, the term “or” is understood to be inclusive and encompasses both “or” and “and”.
[0083] The terms "one or more", "at least one", "more than one" and similar expressions are understood to mean, but are not limited to, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 13 0, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149 or 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more and any value in between.
[0084] Conversely, the term "not exceeding" includes every value less than the stated value. For example, "not exceeding 100 nucleotides" includes 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 54, 53, 52, 51, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, and 0 nucleotides.
[0085] The terms "multiple," "at least two," "two or more," "at least second," and similar expressions are understood to include, but are not limited to, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 9 0, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 13 0, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149 or 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more and any value in between.
[0086] Throughout this specification, the word “comprising” or variations thereof, such as “comprises” or “comprising”, shall be understood to imply inclusion of the stated element, integer or step, or group of elements, integers or steps, but not to exclude any other element, integer or step, or group of elements, integers or steps.
[0087] Unless specifically stated or obvious from the context, the term "about" as used herein should be understood as being within the general tolerances in the field, such as within 2 standard deviations of the mean. "About" can be understood as being within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, or 0.001% of the stated value. All numerical values provided herein are modified by the term "about" unless otherwise clearly apparent from the context.
[0088] Although similar or equivalent methods and materials to those described herein may be used in the practice or testing of this invention, suitable methods and materials are described below. All disclosures, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety. No prior art to the claimed invention is acknowledged in connection with any references cited herein. In the event of any conflict, this specification (including definitions) shall prevail. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be restrictive.
[0089] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood and used by one of ordinary skill in the art to which this application pertains; such art is incorporated herein by reference in its entirety.
[0090] Any of the above aspects and embodiments may be combined with any other aspects or embodiments disclosed in the Summary of the Invention, Drawings and / or Description of the Invention section, including the following examples / embodiments.
[0091] Specific non-limiting examples / embodiments of the present invention
[0092] Disadvantages of using Y-adaptors for dual sequencing
[0093] Duplex sequencing with Y-shaped adaptors is most readily performed using paired-end sequencing reads, as previously described (WO2013142389A1 and Schmitt et al., PNAS 2012, each of which is incorporated herein by reference in full). However, not all sequencing platforms are compatible with paired-end sequencing reads. When using the previously described Y or circular adaptors, where the asymmetric primer sites are located in single-stranded regions opposite the connectable ends of the adaptor, dual sequencing with single-end sequencing reads requires the sequencing reads to extend completely across the DNA molecule. This necessitates capturing SMI tag sequences at both ends of the molecule, which need to be able to distinguish the sequencing reads from the two derivative strands. This requirement is explained below.
[0094] The previously described Y-shaped dual sequencing adaptor was shown in Figure 1A In. Figure 1A In this context, features A and B represent different primer binding sites; α and α′ represent degenerate or semi-degenerate sequences and their reverse complementary sequences; β represents different degenerate or semi-degenerate sequences; and α and β represent two arbitrary sequences from the pool of degenerate or semi-degenerate sequences. Together, these serve as single-molecule identifiers (SMIs).
[0095] As previously described (e.g., WO2013142389A1), SMIs are used to distinguish individual molecules within a larger pool. It is necessary that the population of these encoded in the adaptor library be sufficiently large such that it is statistically unlikely that any two DNA molecules would be labeled with the same SMI sequence. Furthermore, as previously described, breakpoints introduced during library generation can be used as endogenous SMIs, in some cases independently or in combination with exogenous SMIs encoded in the adaptor sequence. In this invention, only exogenous SMI domains are shown in examples of different adaptor designs; however, it should be understood (and included in this invention) that exogenous SMI domains can be substituted by DNA cleavage sites that act as endogenous SMIs or by their amplification.
[0096] After the adaptor is attached to each end of the double-stranded DNA fragment from the library, the structure will be presented as shown in Figure 1B In order to make it clear to track derivatives in subsequent diagrams, the “left” and “right” ends, as well as the “top” and “bottom” strands of a particular DNA insertion sequence are indicated.
[0097] Following PCR, the double-stranded product derived from the "top" strand was observed... Figure 1C In the middle. (L) and (R) indicate the corresponding "left" and "right" ends of the starting DNA molecule:
[0098] Double-stranded PCR products derived from the "bottom" strand are shown in Figure 1D middle.
[0099] It should be noted that α and β are configured differently relative to A and B in the "top" and "bottom" strand products. In the case of paired-end sequencing reads (i.e., reads from the two primer sites A and B of each PCR product), it is possible to distinguish products derived from each strand because the α tag is present in the A read of one strand and β in the B read, and they are reciprocal in other strands. See also Figure 1E .
[0100] Double helix sequence correction is possible using paired end reads as described above. However, if the sequencing reads are long enough to capture the SMI sequences at both ends, then double helix sequences can only be obtained using single-end sequencing reads (i.e., reads from primer site A or primer site B but not from two specific molecule). If sequencing primer A is used, then full-length sequencing reads derived from different strands (i.e., long enough to include both SMI sequences) will produce... Figure 1F The two sequences shown are similar. Similarly, using sequencing primer B with a full-length sequencing read will produce... Figure 1G The following two sequences are shown. In both cases above, the products derived from the "top" and bottom strands can be distinguished from each other by having SMIs in opposite orientations (α-β in one and β-α in the other). However, double sequencing with single-end sequencing is not readily feasible without sequencing reads long enough to capture both SMI sequences. This is because the two sequencing reads do not each contain both α and β tags. Another way to look at this problem is that for the terminal portion of the DNA molecule, the complementary sequence can be omitted from sequencing so that information about the second strand is absent, thus preventing comparison.
[0101] To illustrate this, the two types of sequences generated when using a non-full-length single-end sequencing read from primer A are shown in... Figure 1H Similarly, the corresponding sequences generated when using a non-full-length single-end sequencing read from primer B are shown in [the original text]. Figure 1I In the middle. It should be noted that, for Figure 1H and Figure 1I The two sequencing reads shown, with the "left" and "right" ends of each DNA fragment sequenced only after a given primer, make dual sequencing impossible. This is because there is no opposite strand sequence to compare against. Therefore, even if the amplified population of each sequencing molecule were obtained using two different primers, there would be no information about the second strand, revealing a specific set of reads A and B sequences derived from the same derivative molecule.
[0102] When using single-end sequencing, the need for “reading through” the entire DNA molecule can present some technical challenges on sequencing platforms where read lengths are limited.
[0103] For dual sequencing and single-end sequencing compatible with full-length sequencing reads, alternative adaptor design is essential. In the absence of paired end sequencing reads on Y-shaped adaptors and asymmetric primer sites, other forms of asymmetry must be introduced into the strand-differentiated DNA molecule. Examples of such designs are disclosed below.
[0104] Introducing chain-bound asymmetry using non-complementary "bubbles"
[0105] Figure 2A This invention discloses an exemplary design of a non-Y-shaped adaptor (the present invention) that allows for dual sequencing using unpaired end sequencing (i.e., “bubble adaptors”). Unlike previously described Y-shaped adaptors with two primer sites, this invention has only a single primer site (P) with an inverse complementary sequence (P′). α and its complementary sequence α′ represent degenerate or semi-degenerate single-molecule identifier (SMI) sequences; X and Y represent the two halves of a strand-defining element (SDE), which are segments of non-component sequences forming unpaired “bubbles” among adjacent complementary sequences within the adaptor. Ultimately, the adaptor has a ligatable sequence. The asymmetry introduced by the SDE in this adaptor design distinguishes it from... Figures 2B to 2G The sequencing reads for each strand are shown.
[0106] In similar Figure 2A After the adaptors shown are attached to each end of the DNA fragment, they produce... Figure 2B The structure is shown in the diagram. The second adaptor is shown with SMI sequences β and β′ to illustrate that the SMI sequence of the second adaptor is generally different from that of the first adaptor. Alternatively, the same adaptor can be attached to both ends of a DNA molecule.
[0107] Following PCR amplification, the double-stranded product derived from the "top" strand is displayed on... Figure 2C The double-stranded product derived from the "bottom" chain is shown in Figure 2D middle.
[0108] Because the primer site sequences are identical at both ends of the molecule in this example, depending on the sequencing half of each strand, two different types of sequence reads will be obtained from the single-end sequencing reads of the PCR product from each strand. Reads derived from the "top" strand PCR product are shown in... Figure 2E The reads derived from the "bottom" strand of the PCR product are shown in Figure 2F middle.
[0109] For analysis, such as Figure 2GAs shown, sequencing reads are grouped together by those containing a specific SMI, in this case α or β. Sequences appearing from a given single DNA molecule can be grouped together by means of sequences having the same SMI. It is evident that within each SMI group, two types of sequences are visible: one labeled with SDE X and one labeled with SDE Y. These definitions are derived from sequencing reads of opposite strands (i.e., "top" and "bottom"). For example, when sequences with SMI tag α are grouped together, the resulting sequence is X-α-DNA (…). Figure 2E ) and Y′-α-DNA ( Figure 2F A shared sequence consisting of the sequences generated from the “top” strand of the initial DNA molecule can be obtained by grouping the X-α-DNA sequences together. Similarly, the shared sequence of the “bottom” strand can be obtained by grouping the Y′-α-DNA sequences together. Ultimately, the shared sequence of the two strands can be obtained by comparing the sequences generated from both strands together (i.e., those labeled with sequence X will be compared with those labeled with sequence Y′). Together, these comparisons allow for inclusion as part of dual sequencing analysis.
[0110] Similar results can be achieved by switching the order of the SMI and SDE sequences. An example of such an adaptor is shown in... Figure 2H middle.
[0111] As described above and in WO2013142389A, in some embodiments, the SMI contained within the adaptor sequence can be omitted in place of the endogenous SMI sequence containing the DNA molecule's own splice site sequence. The structure of such an adaptor design would require the following... Figure 2A However, α and α′ are excluded.
[0112] In some applications, Figure 2H The orientation shown is preferred. For example, in some sequencing platforms, such as those currently manufactured by Illumina®, a certain number of bases are available at the start of the sequencing operation for cluster identification and "invariant bases," i.e., reads that are identical to all or substantially many of the bases being sequenced may affect the efficiency of the method. In this case, a degenerate or semi-degenerate SMI sequence immediately following the start of the sequencing operation may therefore be more desirable.
[0113] In other applications Figure 2AThe orientation shown is preferred. As described in the initial description of dual sequencing (i.e., WO2013142389A1), complementary double-stranded SMI sequences can be most suitably generated by extending single-stranded degenerate or semi-degenerate sequences with polymerase primers, or by separately synthesizing and annealing oligonucleotides containing different SMI sequences, and subsequently merging these together to produce different adaptor libraries. If the polymerase extension method is chosen, the presence of an SMI sequence at the end of the linker domain of the adaptor may be advantageous for the extension reaction. On some sequencing platforms, such as those manufactured by Ion Torrent™, the presence of a modified 3′ overhang at the non-linkable end of the adaptor may not be readily compatible with polymerase synthesis; thus, synthesis of adaptors via the polymerase extension pathway is most readily possible using sequences located at the end of the linker domain of the adaptor. Figure 2A The SMI sequence of the connectable ends of the connector shown is performed.
[0114] As a specific example of how this approach can be practiced, consider the Ion Torrent™ sequencing platform, which can handle the following adapter pairs:
[0115] Connector P1
[0116] (SEQ ID NO: 1)
[0117] (SEQ IDNO: 2)
[0118] Connector A
[0119] (SEQ ID NO: 3)
[0120] (SEQ ID NO: 4)
[0121] The asterisk "*" represents a thiophosphate bond.
[0122] The sequencing primers anneal to adaptor A, and the sequence information is read from the DNA fragment originating at the 3' end of adaptor A. Adaptor A can be converted into a sequence applicable to [other applications] using the following sequence. Figures 2A to 2K The formation of the pathway described in the text:
[0123] (SEQ ID NO: 5)
[0124] (SEQ IDNO: 6)
[0125] NNNN refers to four types of nucleotide sequences: degenerate or semi-degenerate; MMMM refers to their complementary sequences; and GC base pairs are included downstream of the degenerate sequence to facilitate ligation, but other forms of linker domains may also be used.
[0126] In this illustration, both adaptor P1 and adaptor A are ligated to the target DNA molecule to be sequenced. For simplicity, the same adaptors ligated to both ends of the DNA molecule can be ignored. However, Ion Torrent™ adaptors utilize different adaptors at each end of the molecule. After initial ligation, individual DNA molecules can be ligated to adaptors in various configurations, such as A-DNA-P1, A-DNA-A, or P1-DNA-P1. The correct configuration of A-DNA-P1 can be used for sequencing reactions by amplification in emulsion PCR using primers targeting sites A and P1. Alternatively, other selection methods known in the art for molecules ligated to only two different adaptors can be used.
[0127] After amplification and sequencing, the following products will be obtained:
[0128] [DNA sequence]
[0129] [DNA sequence]
[0130] It should be noted that these correspond to the products X-α-DNA and Y′-α-DNA, such as Figure 2G As shown in the image.
[0131] Products from both strands obtained via duplex sequencing as previously described can then be co-matched for data processing (see, for example, WO2013142389A1). Specifically, a co-sequence can be prepared from reads beginning with the sequence GCGC NNNN to obtain a co-sequence for the “top” strand. A separate co-sequence can be prepared from reads beginning with the sequence TATA NNNN to obtain a co-sequence for the “bottom” strand. The two single-stranded co-sequences can then be compared to obtain a double-helix co-sequence of the starting DNA molecule. Alternative data processing pathways are disclosed below; see “Alternative Data Processing Flowcharts for Duplex Sequencing”.
[0132] The above approach enables dual sequencing on platforms that utilize short reads that cannot be paired end-to-end, such as in this embodiment, where DNA sequence information only needs to come from one of the two ends of the DNA fragment.
[0133] An alternative embodiment of this approach would be to introduce asymmetry into the SMI sequence itself by using a double-stranded non-complementary or partially non-complementary SMI. Although the SMI sequence itself will not be complementary, by means of a predetermined pairing, the product generated by the non-complementary SMI sequence can be determined to originate from the same starting double-stranded DNA molecule.
[0134] As a specific example of this embodiment, consider a series of Ion Torrent™ “adaptor A” molecules having the following sequence:
[0135] Connector 1:
[0136] (SEQ ID NO: 7)
[0137] (SEQ IDNO: 8)
[0138] Connector 2:
[0139] (SEQ ID NO: 9)
[0140] (SEQ IDNO: 10)
[0141] Connector 3:
[0142] (SEQ ID NO: 11)
[0143] (SEQ IDNO: 12)
[0144] Connector 4:
[0145] (SEQ ID NO: 13)
[0146] (SEQ IDNO: 14)
[0147] For simplicity, only four types of adaptors are listed above, but in practice, a larger pool of such adaptors may be required. It should be noted that in this example, a complementary sequence is included downstream of the non-complementary sequence to form a double-stranded region that will facilitate attachment to the DNA molecule.
[0148] Individual DNA fragments are attached to individual adaptors, resulting in asymmetric labeling of the two DNA strands. Specifically, after sequencing, the sequence of the "top strand" of the starting DNA molecule is labeled with the sequence of the "top strand" of the adaptor. The sequence of the "bottom strand" of the starting DNA molecule is labeled with the reverse complementary sequence of the sequence of the "bottom strand" of the adaptor.
[0149] As a specific example, the two DNA strands linked to adaptor 1 will be labeled AAAT (top strand) and CCCG (bottom strand). Again, it should be noted that the bottom strand, after sequencing, produces an inverse complementary sequence to the sequence initially present in the bottom strand of the adaptor. Similarly, for sequences linked to other adaptors, molecular identifiers can be mutually paired by means of their matching tags. A computer program can then use a table of known tag sequences from the adaptors to compile them into reads generated from the complementary strands of individual DNA molecules. Table 2 shows how the resulting sequence reads will be labeled based on the specific non-complementary identifier sequences shown in the examples above.
[0150] Table 2.
[0151]
[0152] These are merely specific examples of specific embodiments. Those skilled in the art will appreciate that SMI tags can be of any arbitrary length, and SMIs can be completely random or composed of completely predefined sequences. When the SMI sequence is in both strands of a double-stranded molecule, the two SMI sequences can be completely complementary (as described in the first example mentioned above), partially non-complementary, or completely non-complementary. In some embodiments, exogenous molecular identifier tags are not required at all. In some cases, the ends of randomly spliced DNA molecules can be used as unique identifiers, provided there is some classification of asymmetry (including SDEs) that allows differentiation of products generated from the two independent strands of a given single molecule of double-stranded DNA.
[0153] In any aspect disclosed herein or in embodiments of the invention (and not limited to those described herein), in both single-stranded and double-stranded SMIs, the SMI tag set may be designed with different edit distances between tags such that errors in the synthesis, amplification, or sequencing of the SMI sequences will not cause one SMI sequence to be transformed into another (see, for example, Shiroguchi et al., Proceedings of the National Academy of Sciences of the United States of America, 109(4):1347-1352). Incorporating edit distances between SMI sequences allows for the removal of SMI errors, for example, by using an alternative method of error correction known in the field, such as identifying Hamming distance, Hamming code, or other methods. All SMIs from the set may be of the same length; or a mixture of two or more SMIs of different lengths may be used within a set of SMIs. Using a mixture of SMI lengths can facilitate the design of adaptors that utilize SMI sequences and additionally have one or more fixed bases at sites within or alongside the SMI, because using more than one SMI length within the ensemble will result in invariant bases not always present at the same read position during sequencing (see, for example, Hummelen R et al., PLoSOne, 5(8):e12078 (2010)). This approach avoids potential problems on sequencer platforms where suboptimal performance (e.g., difficulty in cluster identification) may occur when invariant bases are present at specific read positions.
[0154] Those skilled in the art will also appreciate that sequences capable of introducing asymmetry can be introduced anywhere within the sequencing adaptor, before or after the SMI sequence, or within the single-stranded “tail” sequence in an adaptor design having such a sequence, including, for example, as internal “bubble” sequences as shown above. These sequences, along with any associated SMI sequences, can be read directly as sequencing read portions or can be determined by a non-independent sequencing reaction (e.g., in index reads). These sequences can also be used in conjunction with Y-shaped adaptors, “loop” adaptors, or any other adaptor designs known in the art.
[0155] In fact, adaptors with SMI sequences, SDE sequences and primer binding sites having different relative orientations are envisioned and included in this invention.
[0156] Figure 2A and Figure 2H The connector design shown indicates that the non-connecting end is flat. However, this end can be protruding, recessed, or have modified bases or chemical groups to prevent degradation or unwanted connections.
[0157] Additionally, the two chains of the linker can connect to form a closed "loop," which in some applications can be desirable to prevent degradation or unwanted connections. See, for example... Figure 2I . Figure 2I The closed “loop” connection (marked at position “S”) can be achieved via a conventional phosphodiester bond or via any other natural or non-natural chemical linker group. This bond can be cleaved chemically or enzymatically to obtain an “open” end before, during, or after ligation; the loop may need to be cleaved before PCR amplification to prevent rolling circle amplicons. Non-standard bases such as uracil can be used here, and a combination of enzymatic steps can be used before, during, or after adapter ligation to cleave the phosphodiester backbone. For example, in the case of uracil, a combination of uracil DNA glycosylase (to form a base-free site) and endonuclease VIII (to cleave the backbone) can be used. Alternatively, a large chemical group or other non-transferable modified base at this linker site can be used to prevent polymerase translocation beyond the loop end and for the same purpose.
[0158] In any aspect or embodiment of the invention disclosed herein (and not limited to those described herein) for the design of adaptors using double-stranded SMI sequences (whether complementary, partially non-complementary, or completely non-complementary), a particular advantage of synthesizing the adaptor as a linear molecule annealed to a “loop” form is that the “top” and “bottom” SMI sequences will be present in a 1:1 ratio within the molecule itself. This approach can be advantageous relative to annealing individual “top” and “bottom” oligonucleotide pairs to form a double-stranded SMI, as in such an approach, if the concentrations of the oligonucleotides used for the “top” and “bottom” chains are not in a perfect 1:1 ratio, then there may be an excess of one adaptor chain or the other, which could be problematic in downstream steps (e.g., the extra single-stranded oligonucleotides may cause improper initiation during PCR amplification, or may anneal with other single-stranded oligonucleotides that may be present, potentially producing an adaptor molecule in which the two SMI chains are improperly paired).
[0159] In some cases, it may be necessary to prevent the entire circular sequence from replicating itself, where modified sequence locations may optionally include replication-blocking sites. This could be bases that can be removed enzymatically (e.g., uracil, which can be removed by uracil DNA glycosylase) or regions that partially or completely inhibit DNA replication (e.g., base-free sites).
[0160] Alternatively, restriction endonuclease sites may be introduced (in...) Figure 2I The position “T” in the middle is marked, which can be used to obtain an “opening” structure, ultimately releasing the smaller hairpin segment.
[0161] It should be obvious that different configurations of base asymmetry between the two linker chains equally function as chain-limiting elements. When relative to Figure 2J When the other complementary strand shown in the linker is inserted with one or more nucleotides, bubbles can be formed in the linker strand. Figure 2K The interleaver is displayed, wherein more than one nucleotide insertion includes a self-complementary portion; thereafter the interleaver provides a functional group similar to a simple difference between two strands involving one or more nucleotide positions.
[0162] Chain-bound asymmetry is introduced using non-complementary SMI sequences.
[0163] Figure 2A to 2K The adaptor design shown contains two key features for enabling tag-based duplex sequencing. One is a unique molecular identifier (i.e., SMI) and the other is a component that introduces asymmetry between the two DNA strands (i.e., SDE). In the initial description of duplex sequencing, a Y-shaped adaptor and paired-end sequencing reads are utilized. The introduction of asymmetry between the two DNA strands is achieved by means of the asymmetric tail itself. Figure 3A The different and excellent dual sequencing adaptor designs shown include non-complementary “bubble”-shaped SMIs, which together serve as molecular identifiers and introduce SDEs for asymmetry.
[0164] In this design, P and P′ represent the primer site and its complementary sequence, respectively, and αi and αii represent two degenerate or semi-degenerate sequences that are non-complementary for all or part of their length. The synthesis of this form of adaptor is most readily achieved by synthesizing and hybridizing oligonucleotide pairs with different degenerate or semi-degenerate sequences individually before combining two or more of these to form a diversity pool. Because the oligonucleotides are synthesized and annealed individually, the relationship between given αi and αii sequences will be known and documented in a database, from which the corresponding coupler SMI sequences can be retrieved during post-sequencing analysis.
[0165] After the adaptor is ligated to the double-stranded DNA fragment, it produces Figure 3B The structure shown is illustrated. In this structure, the pair of non-complementary SMI sequences βi and βii are generally different from αi and αii, but the same connective substructure can be connected to both ends.
[0166] Following PCR amplification, the double-stranded product derived from the "top" strand is displayed on... Figure 3C The double-stranded product derived from the "bottom" chain is shown in Figure 3D middle.
[0167] Because the primer site sequences are identical at both ends of the molecule (in this example), two different types of sequence reads will be obtained from single-end sequencing reads of the PCR product from each strand, depending on the single strand being sequenced. Single-end sequencing reads from the "top" strand PCR product are shown in... Figure 3E The single-end sequencing reads from the "bottom" strand are shown in Figure 3F middle.
[0168] During the analysis, reads can subsequently be grouped based on specific SMI sequences and their corresponding non-complementary pairs, according to database relationships known from those generated during synthesis from and associated with the SMI joiner sub-library. For example... Figure 3G As shown, the paired “top” and “bottom” chain sequences of the initial molecule are labeled with αi and αii (for the reads that originate at one end of the molecule) and βi and βii (for those that originate at the opposite ends).
[0169] Asymmetry is defined by introducing modified or non-standard nucleotides into the chain.
[0170] Another way to introduce strand asymmetry into dual-sequencing adaptors is by first forming paired strands of DNA, but then resulting in mismatched nucleotides or nucleotide analogs after another biochemical step. One example is DNA polymerase misincorporation. Misincorporation can occur inherently during amplification or after conversion to the mismatched region via chemical or enzymatic steps.
[0171] For some applications, this form of SDE may preferably be the “bubble-shaped” sequence disclosed above, as it avoids problems that may arise from non-single-stranded regions, such as misannealing to other DNA oligonucleotides and degradation by exonucleases / endonucleases.
[0172] Many non-standard nucleotides known in the field can be used for this purpose. Non-limiting examples of such modified nucleotides include tetrahydrofuran; 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A); 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G); deoxyinosine; 5′-nitroindole; 5-hydroxymethyl-2'-deoxycytidine; isocytosine; 5′-methyl-isocytosine; and isoguanosine and other nucleotides known in the field.
[0173] Duplex sequencing adaptors containing 8-oxo-G were shown in Figure 4A In the middle, the 8-oxo bases are paired with their complementary cytosine bases without bubble formation. As in the examples above and below, the relative order of the SMI sequence (α in this case) and the SDE site (8-oxo-G site in this case) can be switched as needed. P and P′ represent the primer site and its complementary sequence.
[0174] After the adaptor is ligated to the double-stranded DNA fragment, it produces Figure 4B The structure shown.
[0175] This can then be followed by glycosylation with an enzyme such as oxoguanine glycosylation enzyme (OGG1) (which potentially binds to DNA ligases to repair any resulting gaps that may occur with glycosylation enzymes that have cleavage activity). Figure 4B Treatment of double-stranded DNA. This treatment, when introducing base-free sites, will produce a complete phosphodiester DNA backbone, such as... Figure 4C As shown in the diagram. Each of the two strands can subsequently be copied, for example, with a polymerase. Under appropriate reaction conditions, certain thermostable polymerases preferentially insert into the A-opposite abase-free site (Belousova EA et al., *Biochim Biophys Acta*, 2006), producing a G->T mutation. In contrast, the reciprocal strands retain the C nucleotide present in the adaptor at the time of ligation. This treatment results in strand asymmetry, thus allowing for the differentiation of the two-strand products.
[0176] During PCR or other forms of DNA amplification, under certain conditions using a specific polymerase, adenine will preferentially insert relative to the abase-free strand during copying. With subsequent rounds of copying, this adenine will pair with thymine, eventually replacing the initial 8-oxo-G site with T. Furthermore, glycosylation enzyme treatment is not mandatory. Under appropriate reaction conditions, the polymerase can insert relative to the 8-oxo-G without the shown abase-free intermediate (Sikorsky JA et al., *Biochem Biophys Res Commun*, 2007). In either case, after PCR amplification, the double-stranded product derived from the "top" strand will appear as... Figure 4D As shown, and the double-stranded product derived from the "bottom" chain will be as follows: Figure 4E As shown in the image.
[0177] Because the primer site sequences are identical at both ends of the molecule in this (non-restrictive) example, depending on the single strand being sequenced, two different types of sequence reads will be obtained from the single-end sequencing reads of the PCR product from each strand. Those PCR products derived from the "top" strand PCR product will be as follows: Figure 4F As shown, and those PCR products derived from the "bottom" strand will be as follows: Figure 4G As shown in the image.
[0178] During analysis, sequencing reads can be grouped by those containing a specific SMI, in this case α or β. See also Figure 4H The products of T and G markers within each SMI group define the origin chain and allow for comparison of double helix sequences.
[0179] Those skilled in the art will also find it obvious that the modified nucleotide or other analogue, as described above, can be placed anywhere within the sequencing adapter, provided that the sequence obtained from the modified nucleotide or other analogue during DNA sequencing is recoverable.
[0180] Those skilled in the art will appreciate that a variety of other nucleotide analogs can be used to achieve the same purpose. Other examples include tetrahydrofuran and 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A). Any nucleotide modification that can inherently produce mis-incorporation of different nucleotides by DNA polymerase or that can be converted into miscoded lesions or mismatched bases by enzymatic or chemical steps or spontaneously over time can be used in the adaptor of this embodiment.
[0181] Furthermore, non-nucleotide molecules can be incorporated to asymmetrically label the two chains. For example, biotin can be incorporated into one of the two linker chains, which will facilitate the separate analysis of the two chains by physically separating the biotin-containing chain from the non-biotin-containing chain using streptavidin. This embodiment is disclosed in detail below.
[0182] Combinations of dual-sequencing adaptors are designed to introduce different primer sites at opposite ends of the DNA molecule.
[0183] The aforementioned examples of non-Y-shaped adaptors demonstrate that the same type of adaptor is symmetrically linked to both ends of a DNA molecule. Currently, most sequencing platforms require adapted DNA molecules with different primer sites at either end to allow, for example, cluster amplification on surfaces or beads. For sequencing platforms that do not routinely use Y-shaped adaptors to generate these different primer sites (e.g., Ion Torrent™ (Thermo®), SOLiD (Applied Biosystems®), and 454 (Roche®)), a mixture of two different adaptors is ligated, and molecules containing one of each primer site are subsequently selected; this is most commonly via bead-based emulsion PCR methods.
[0184] The following describes a simple approach to generating asymmetric primer sites using non-Y-shaped dual sequencing adaptors.
[0185] This produces a mixture of a double-helix adaptor and a standard adaptor, each containing a different PCR primer site. The double-helix adaptor can be any design described above or below herein or as known in the art.
[0186] An exemplary double-helix connector is shown in Figure 5A In this structure, there is a primer site P and a complementary sequence P′, followed by an SDE composed of mismatched sequences X and Y, each containing one or more nucleotides, and then a degenerate or semi-degenerate SMI sequence α. Figure 5B The other linkers shown are "standard" linkers containing different primer sites O and complementary sequences O′.
[0187] After this adapter mixture is ligated into a DNA library, three different types of products are produced, such as Figures 5C to 5E As shown in the diagram. On average, half of the successfully adapted molecules will carry a different adaptor sequence at each end (as shown in the diagram). Figure 5C One-quarter will have two double-helix connectors ( Figure 5D ), and one-quarter will have two standard connectors ( Figure 5E Under appropriate selection conditions, molecules with only one primer site P and one primer site O will amplify in clusters. Therefore, the latter two (unsuitable) types of products can be ignored thereafter and are not shown in the subsequent description.
[0188] After PCR amplification, the double-stranded product derived from the "top" strand will... Figure 5F As shown, and the double-stranded product derived from the "bottom" chain will be as follows: Figure 5G As shown in the image.
[0189] Sequencing at primer site P will produce the following sequences derived from the "top" and "bottom" strands. These can be distinguished by carrying SDE X or Y markers. See also Figure 5H .
[0190] Obviously, any other form of non-Y-shaped double-helix adaptor described herein or known in the art can be used for the same purpose as that used in this embodiment. For example, instead of one double-helix adaptor and one standard adaptor, it is possible to use two double-helix adaptors carrying different primer sites. After ligation and PCR, the amplification products can be separated, and one portion can be sequenced with primer P and the other with primer O. This allows for duplex sequencing of both ends of each adapted molecule. Because reads from different primer sites are not actually paired, they cannot be easily linked together for any particular molecule. However, for applications with a very limited amount of DNA to be sequenced, the additional sequence information obtained from duplex sequencing of both ends of the molecule can be advantageous.
[0191] Using two reads on an unpaired platform can maximize read length during duplex sequencing.
[0192] Pairwise sequencing, such as that performed on Illumina® instruments, typically requires a sequencing platform capable of sequencing one strand from a primer site at one end of the adaptor DNA molecule, subsequently generating a reverse complementary sequence strand, and then sequencing the other end of the molecule from a different primer site. The technical challenges include the complementary strand production process, which is why not all platforms are readily compatible with this pairwise sequencing.
[0193] However, it is possible to sequence two distinct portions of a adapted DNA molecule to a limited extent without generating complementary strands. This can be achieved by using a second primer site contained within a second adaptor attached to the opposite end of the DNA molecule relative to the first adaptor, allowing the sequencing read to advance away from both the DNA molecule and the first adaptor, thereby generating a sequencing read of the second adaptor itself. In some cases, this capability may be desirable. For example, because the SMI and SDE sequences required for dual sequencing consume a portion of the inherently limited achievable read length, it may be helpful to be able to move these elements to the opposite adaptor read during the second, shorter read when a maximum read length is required. A similar benefit can be achieved through index barcode sequence repositioning, which is typically used for sample multiplexing.
[0194] To achieve this process, two different connectors can be used. For example... Figure 6A The first one shown contains a simple primer site P opposite its complementary sequence P′.
[0195] like Figure 6B The other adaptor sequence shown contains the features necessary for duplex sequencing but without the Y-tail: SMI and SDE. This double-helix adaptor can be any of the designs described herein, where SMI and SDE are separate sequence elements, merged into the same sequence element as unpaired SMI, or where SDE consists of modified bases.
[0196] exist Figure 6B In the example shown, SDE requires mismatched sequences X and Y adjacent to the degenerate or semi-degenerate SMI sequence α. PCR primer site O and complementary sequence O′ are located at the unjoined end of the adaptor. This adaptor design features a second primer site P2 adjacent to a joinable end but oriented such that the annealing primer extends into the adaptor molecule itself rather than toward the complementary sequence P2′ of the DNA fragment.
[0197] After this mixture of adaptors was ligated into a DNA library, three different products were produced. Those with two of the same adaptor types at opposite ends are negligible because the product containing only one of each adaptor type (containing two primer sites, P and O, such as...) Figure 6C (As shown in the diagram) will successfully amplify and sequence the cluster.
[0198] After PCR amplification, the double-stranded product derived from the "top" strand will... Figure 6D As shown, and the double-stranded product derived from the "bottom" chain will be as follows: Figure 6E As shown in the image.
[0199] The following shows the orientation of annealed sequencing primers P1 and P2 and the regions that can be sequenced one by one. These reads will be sequenced optimally with one primer, followed by the other. This will be achieved by introducing a sequencing primer and undergoing the first sequencing read; then introducing the second primer after the first sequencing read is complete. If a "read 2" is performed first (e.g., Figure 6F and Figure 6G As shown in the diagram, sequencing can proceed until the molecular ends are reached and will self-terminate. If "read 1" is performed first, the sequencing reaction will need to be stopped before adding primer P2 to begin "read 2". This can be achieved by introducing modified dNTPs that do not extend further after incorporation or by melting the strand synthesized during the initial sequencing reaction away from the template strand with heat or chemicals and washing it off before adding the next sequencing primer.
[0200] The arrangement of the sequencing template strand derived from the "top" strand is as follows: Figure 6F As shown in the diagram, the arrangement of the sequencing template strands derived from the "bottom" strand is as follows: Figure 6G As shown in the image.
[0201] Sequencing reads from templates derived from the "top" strand will be as follows Figure 6H As shown in the diagram, sequencing reads from templates derived from the "bottom" strand will be as follows: Figure 6I As shown in the image.
[0202] It is evident that sequencing read pairs from different initial chain molecules are distinguishable by using SDE X or SDE Y markers.
[0203] Using two reads on a non-paired platform maximizes tag diversity in dual sequencing.
[0204] The potential advantages of using the dual reads in the form disclosed above go beyond simply preserving read length. In the initial description of tag-based dual sequencing using Y-adaptors, an SMI sequence is attached to each end of a adapted DNA molecule. This design has practical advantages in certain situations to efficiently generate a sufficiently large population of adaptors containing diverse SMIs to ensure that each DNA molecule can be uniquely labeled.
[0205] As an example, if a fully degenerate four-nucleotide SMI sequence is introduced into the initial Y-shaped adaptor design and ligated into a DNA fragment library (such as... Figure 1B (As shown in the diagram) and sequencing is performed using paired-end reads, then the total number of possible methods for labelable molecules is 4. 4 * 4 4 = 65,536. If a fully degenerate 8-base-pair SMI sequence is incorporated into a double-helix adaptor and ligated to a DNA library for single-end reading (e.g.) Figure 5CAs shown in the diagram, the same 65,536 tag combinations can be achieved. Both methods are equally feasible when generating complementary SMI tags using polymerase extension; however, the situation is different when generating an adaptor library using separately synthesized oligonucleotides. In the first scenario, a total of 4... 4 ×2 = 512 oligonucleotides. In the latter scenario, it will be necessary to generate and anneal 4 separately. 8 ×2=131,072; this will greatly increase the required financial costs and workload.
[0206] For some embodiments of dual sequencing, the oligonucleotide synthesis method for generating SMI adaptors is preferred, and it may be nearly impossible to achieve a population of adaptors containing a sufficient variety of SMIs with only a single SMI at one end of the molecule as disclosed above.
[0207] The dual-read method described above on unpaired end-compatible platforms can overcome this limitation by including the SMI sequence in both adaptors used for sequencing in two identical reaction steps. This is explained below.
[0208] This requires two types of adaptors, each with different amplification primer sites. At least one must contain an SDE, and in this example, both will contain degenerate or semi-degenerate SMI sequences. Figure 7A As shown, the first connector is similar to Figure 6A The linker differs in that it additionally includes an SMI sequence (identified here as "β"). For example... Figure 7B As shown, the second connector is similar to Figure 6B The linker shown contains an SMI sequence (identified here as "α").
[0209] Those skilled in the art will readily recognize that the relative configurations of the SMI and SDE features in the two connectors are interchangeable to achieve the same result. Instead, the SDE shown in the latter connector above can be placed in the former. Any form of SDE or SMI previously described can be replaced by its equivalent used in this example.
[0210] After this adapter mixture was ligated to the DNA library, the product successfully interacted with the DNA library. Figure 7C One combination of each connector subtype shown.
[0211] After PCR amplification, the double-stranded product derived from the "top" strand will... Figure 7D As shown, and the double-stranded product derived from the "bottom" chain will be as follows: Figure 7E As shown in the image.
[0212] As described in the foregoing embodiments, the orientation of sequencing primer sites P1 and P2 and the regions sequenced by their respective sequences are as follows: Figure 7F As shown (for the "top" chain) and as Figure 7G As shown in the diagram (for the bottom chain).
[0213] The reading segment of the template derived from the "top" chain will be as follows Figure 7H The read segment shown and derived from the template of the "bottom" chain will be as follows Figure 7I As shown in the image.
[0214] Furthermore, the products of the two strands can be easily distinguished by their different X and Y SDE markers. For dual sequencing analysis, the sequences of SMI α and SMI β can be combined into a single identifying tag sequence.
[0215] Asymmetric SMI in Y-shaped dual sequencing adaptors
[0216] Several currently available sequencing platforms require different primer sites at opposite ends of the DNA molecule to allow for cluster amplification and sequencing. This can be achieved using Y-shaped or bubble-shaped adaptors with asymmetric primer binding sites, or via the two adaptor ligation methods described in the preceding three examples. Y-shaped adaptors have been most commonly used in paired-end sequencing compatible platforms, such as those manufactured by Illumina®; however, they can be used on other platforms.
[0217] A general advantage of library-prepared Y- or "bubble-shaped" adaptors is that, theoretically, each doubly adapted DNA molecule will be capable of being sequenced. However, when using methods with two different adaptors, only half of the generated molecules will be capable of being sequenced, having one of each adaptor type, while the other half will have two copies of the same adaptor. In certain situations, such as when the input DNA is limited, the higher conversion rate of Y-shaped adaptors may be desirable.
[0218] However, as shown in the first embodiment described above (the previously described dual sequencing method), the previously described Y-shaped double helix adaptor does not readily allow dual sequencing with single-end reads in the absence of paired end reads or full read capability.
[0219] However, using sequencing primer sites in the complementary "stem" sequence of the Y-shaped adaptor allows for single-end reads for dual sequencing, but only if asymmetry is introduced elsewhere via at least one SDE in the adaptor sequence. A simplified illustration is shown below.
[0220] exist Figure 8AThe diagram shows a Y-shaped adaptor containing an unpaired SMI containing sequences αi and αii. This sequence in this design will also serve as the SDE. Three primer sites are present: A and B, which are PCR primers on the free tail; and C (and C′), which include the sequencing primer site (and its complementary sequence).
[0221] After the adaptor is linked to the DNA fragment, it produces Figure 8B The structure shown has two connectors with two different non-complementary SMIs attached to either end.
[0222] After PCR amplification using primers complementary to sites A and B, the double-stranded product derived from the "top" strand will appear as follows: Figure 8C As shown in the diagram, the double-stranded PCR product derived from the "bottom" strand will be as follows: Figure 8D As shown in the image.
[0223] Following sequencing from primer site C, two different types of sequencing reads will be obtained from the single-end reads of each strand of the PCR product, depending on the half of the single strand being sequenced. Sequencing reads from the "top" strand PCR product are as follows: Figure 8E As shown in the image, and the sequencing reads derived from the "bottom" strand PCR product are as follows: Figure 8F As shown in the image.
[0224] During analysis, sequencing reads can be grouped based on specific SMI sequences and their corresponding non-complementary pairs, according to database relationships known from those generated during synthesis from the SMI adaptor library and bound to it. For example, Figure 8G As shown, the paired “top” and “bottom” chain sequences of the initial molecule are labeled with αi and αii (for the reads that originate at one end of the molecule) and βi and βii (for those that originate at the opposite ends).
[0225] This allows for double helix sequence analysis, similar to the analysis described in the example above titled "Introducing Chain-Limited Asymmetry Using Non-Complementary SMI Sequences".
[0226] This type of Y-shaped connector is as follows: Figure 8H Alternative designs illustrated include closed loops, which help prevent exonuclease digestion or potentially nonspecific ligation to the free arm of Y and the "daisy chain" of the free arm. The closed "loop" bond (marked with arrows) can be achieved via a conventional phosphodiester bond or via any other natural or non-natural chemical linker. This ligation can be chemically or enzymatically cleaved to obtain an "open" end after ligation has taken place, as is often required before PCR amplification to prevent rolling circle amplicons. Alternatively, a large chemical group or modified nucleotide at this ligation site can be used to prevent polymerase lateral migration beyond the loop end and for the same purpose. Alternatively, as... Figure 8IAs illustrated, a restriction endonuclease recognition site is introduced at the hairpin complementary region within the loop (marked by an arrow); this can be used to obtain an "opening" structure, ultimately releasing a smaller hairpin fragment.
[0227] In some cases, it is preferable that no additional enzymatic step is required after adaptor ligation and before PCR. Where the need for an additional step does not exist, such as... Figure 8J The adaptor design illustrated in the example (where the adaptor tails are complementary but not covalently linked) can still overcome the problems caused by free unpaired DNA tails.
[0228] Asymmetric SMI in Y-shaped dual sequencing adaptors
[0229] Another variation of the concept of unpaired SMIs in Y-shaped or circular adaptors includes those located in the free single-stranded tail region between the PCR primer site and the complementary stem. One advantage of this design is that it allows the SMI to be fully sequenced as part of a “double-indexed” read, as is available on select Illumina® sequencing systems (Kircher et al. (2012), Nucleic Acid Res., Vol. 40, No. 1, e3). For applications requiring particularly long reads, excluding the SMI from the main sequencing read will maximize the read length of the DNA insert sequence. An example follows.
[0230] Figure 9A Displays a Y-shaped double-helix sequencing adaptor containing unpaired PCR primer sites A and B. αi and αii represent a pair of at least partially non-complementary degenerate or semi-degenerate SMIs. P and P′ are the sequencing primer sites and their complementary sequences.
[0231] After the adaptor is linked to the DNA fragment, it produces Figure 9B The structure shown here has two connectors of two at least partially non-complementary SMIs attached to either end.
[0232] After PCR amplification using primers complementary to sites A and B, the double-stranded product derived from the "top" strand will appear as follows: Figure 9C As shown, and the double-stranded product derived from the "bottom" chain will be as follows: Figure 9D As shown in the image.
[0233] On the Illumina® platform, as an example, when using paired-end sequencing with dual indexes, after completing one sequencing read and one index read, a complementary strand can be generated and the corresponding sequencing and index reads of the other strand can be performed.
[0234] However, it should be noted that paired-end sequencing or dual-indexing techniques themselves allow for dual sequencing. Although the two single strands of a given PCR product are effectively sequenced together, each PCR product is derived from only one of the two strands of the initial DNA double helix, and therefore sequencing both strands of the PCR product is not equivalent to sequencing both strands of the initial DNA double helix.
[0235] Sequencing primers and index primers, along with their sequencing regions, are shown in Figure 9E (for reads of PCR products derived from the "top" strand in both directions) and shown in Figure 9F (Reads for PCR products derived from the "bottom" strand in both directions).
[0236] It will also satisfy the requirement of sequencing both the SMI and the sequence itself in a single sequencing read, rather than two separate reads. It is evident that a variety of primer configurations and numbers can be used to sequence the SMI and read sequences. In some embodiments, such as nanopore sequencing, sequencing of the SMI and / or DNA sequence may not require specific primer sites at all. Furthermore, this example is described using PCR, but this and other embodiments can be amplified using any other method known in the art, including rolling circle amplification and other pathways. See Kircher et al. (2012).
[0237] When in chains derived from "top" and "bottom" (such as...) Figure 9G When comparing the sequences of different patterns in all four reads (as shown), it is evident that they are distinguishable from each other, since one carries SMI tags αi′ and βi and the other carries tags αii and βii′. Although the two chains do not share any common tags in this non-limiting instance, they can still be correlated with each other, since the relationships between αi and αii and between βi and βii are known from the database, which is an analytical component, during the preparation of the connector and can thus be looked up from the database.
[0238] Primer sites, SMI, and SDE for dual sequencing were introduced using a single circular vector.
[0239] Figures 10A to 10E The paper describes alternative structures for introducing all the elements necessary for dual sequencing into a single molecule rather than two paired adaptors.
[0240] In this embodiment, the ring structure is formed by attaching the two ends of a linear double-stranded molecule (containing elements necessary for dual sequencing) to the two ends of a DNA fragment having compatible ligation sites.
[0241] exist Figure 10AIn this context, A / A′ and B / B′ represent two different primer sites and their reverse complementary sequences; α and α′ require degenerate or semi-degenerate SMI sequences; and X and Y are the corresponding non-complementary halves of the SDE.
[0242] Double-stranded DNA fragments are ligated Figure 10A After entering the double-chain molecule, a closed ring is formed, such as... Figure 10B As shown in the image.
[0243] In production Figure 10B After ligation, the product is amplified by PCR from the primer sites. Alternatively, rolling circle amplification can be performed first. Selective disruption of unligated libraries and adaptors can be advantageously achieved using 5′–3′ or 3′–5′ exonucleases. The circular design uniquely provides these opportunities, which are unlikely to be easily achieved in many other designs.
[0244] It is obvious that any of the forms of SMI and SDE described above and below can replace those shown or their rearrangement.
[0245] As an example of another embodiment, such as Figure 10C As shown, a single element that is close to serving as a connection site for SMI and SDE can be used, as discussed in the embodiment entitled "Introducing Chain-Defined Asymmetry Using Non-Complementary SMI Sequences".
[0246] Alternatively, such as Figure 10D and Figure 10E As shown, SDE and SMI can be designed into sequences close to each of the adaptor connection sites to facilitate paired end sequencing.
[0247] In this design, it should be noted that the SMI sequences on opposite chains do not necessarily have to be complementary (e.g., Figure 10E As shown in the figure, it is sufficient to know the relationship between the corresponding sequences (i.e., αi and αii) and to be searchable in the database during the analysis.
[0248] Dual sequencing via asymmetric chemical labeling and strand separation
[0249] As discussed above, duplex sequencing fundamentally relies on sequencing the two strands of the DNA double helix in a distinguishable manner. In the previously described duplex sequencing embodiments (in WO2013142389A1), the two strands can be linked together with a hairpin sequence to jointly sequence the paired strand. WO2013142389A1 and several embodiments disclosed above describe methods in which the two strands of the unique DNA double helix can be distinguished using DNA markers. This subsequent approach involves labeling each DNA molecule with a unique DNA sequence (an endogenous SMI containing the coordinates of one or both ends of a DNA fragment, or an exogenous SMI containing a degenerate or semi-degenerate sequence) and introducing asymmetry (e.g., asymmetric primer sites with paired end reads, “bubble” sequences, non-complementary SMI sequences, and non-standard nucleotides that are naturally or chemically converted into mismatches) via at least one form of SDE.
[0250] The following describes another approach for performing dual sequencing, which involves asymmetric chemical labeling of the two strands in the double helix so that they can be physically separated for sequencing in a reaction-independent manner. An example is given below.
[0251] like Figure 11A As shown, two different adaptors are used. The first adaptor contains a primer site P with its complementary sequence P′ and an SMI sequence α with its complementary sequence α′. One strand of the first adaptor additionally carries a chemical tag capable of binding to or interacting with known substances (e.g., solid surfaces, beads, immobilized structures, and binding complexes) in a way that the other DNA strand would not. Figure 11A As shown, the chemical label is biotin, which has affinity for binding complexes and for streptavidin.
[0252] Other binding pairs known in the art can be used, preferably in the form of small molecules, peptides, or any other uniquely binding moieties. This tag can also be in the form of nucleic acid sequences (e.g., DNA, RNA, or combinations thereof; and modified nucleic acids, such as peptide nucleic acids or locked nucleic acids), preferably in single-stranded form, wherein substantially complementary "bait" sequences attached to a solid matrix (e.g., solid surfaces, beads, or other similar immobilized structures) can be used to bind and selectively trap and isolate one strand of the molecule linked by the adaptor from the other.
[0253] The second linker does not carry the chemical tag found in this non-limiting example. For example... Figure 11B As shown, the second linker carries different primer sites O and complementary sequences O′.
[0254] exist Figure 11A and Figure 11B After the adaptor is attached to the DNA fragment, it produces... Figure 11C The (preferred) structure shown.
[0255] Additionally, two other types of structures will be generated: one with an adaptor containing two primer sites P and another connected to an adaptor containing two primer sites O. As discussed in the example titled "Combinations of Adaptor Designs Using Dual Sequencing to Introduce Different Primer Sites at Opposite Ends of DNA Molecules," the preferred structure concentration can be routinely achieved before sequencing using specific amplification conditions relative to the other two types of structures, making the other two types of structures negligible.
[0256] like Figure 11D As shown, after ligation, the DNA strands can be melted apart by heat or chemical means, and subsequently, the strand with the chemical tag (which has selective affinity for a specific binding complex (in this case, streptavidin, such as binding to paramagnetic beads)) can be separated from the other strand. The two now separated strands can be sequenced independently, optionally using the aforementioned steps, wherein the two separated strands are amplified independently (sequencing can occur in physically different reactions or in the same reaction after each strand has been indexed differently, for example, with labeled PCR primers and recombination).
[0257] Alternatively, the two strands can be tagged with different chemical tags that have affinity for two different types of bait. The tag visible in one sequencing reaction or index group can then be compared with a corresponding tag in the other population and subjected to duplex sequencing analysis. In this example, SDE is used again, but it requires chemical tags that can be used to physically separate the asymmetric attachment of the strands. Their physically distinct compartmentalization allows the two strands to be sequenced separately or to undergo a subsequent differential labeling step (e.g., using PCR with primers carrying different index sequences at their tails) before merging, and the merged sequencing can later be deconvolved using a wavebeam.
[0258] Another embodiment of this concept involves using markers with other properties (i.e., physical groups) that allow for chain separation by means other than chemical affinity. For example, nucleic acid chains containing molecules with a strong positive charge (e.g., physical groups with charge properties) can be preferentially separated from their paired unlabeled strands by applying an electric field (e.g., by electrophoresis), or nucleic acid chains containing molecules with a strong magnetic permeability (e.g., physical groups with magnetic properties) can be preferentially separated from their paired unlabeled strands by applying a magnetic field. In solution, under certain application conditions, nucleic acid chains containing precipitation-sensitive chemical groups (e.g., physical groups with insoluble properties) can be preferentially separated from their paired unlabeled strands so that the DNA itself is soluble, but the DNA containing the physical groups is insoluble.
[0259] Another variation of the concept of physical separation of paired strands after the application of SMI (exogenous tags within the sequencer as ligands or endogenous SMIs containing unique cut sites of DNA fragments) is the use of dilution after the DNA double helix has been thermally or chemically melted into its component single strands. Single strands are diluted into multiple (i.e., two or more) physically separated reaction chambers to reduce the probability of two originally paired strands sharing the same container, rather than applying a purifiable chemical label to one strand to separate it from another. For example, if a mixture is randomly separated in one hundred containers, only about 1% of the paired strands will be placed in the same container. Containers may require a set of physical containers, such as wells in containers, test tubes, or microplates, or physically separated non-connected droplets, such as aqueous / hydrophobic emulsions. Any other method can be used in which two or more spatially dissimilar volumes of fluid or solid contents containing nucleic acid molecules are prevented from substantially intermixing with the nucleic acid molecules. In each container, PCR amplification is preferably performed using primers carrying different tag sequences. This unique tag sequence, added using different primers in each container, will be optimally positioned at a recordable location during sequencing index reads (see, for example, [link to documentation]). Figure 9E These markers will act as SDEs. In this example, approximately 99% of the mating strands carrying the same SMI marker will be designated with a different SDE marker than their mating strands. Only about 1% will be designated with the same marker. Duplex sequencing analysis and concordant sequence preparation can be performed using SMI and these SDEs as usual. In rare cases (where the mating strands accidentally acquire the same SDE), these molecules will be inherently ignored during double-helix analysis and will not provide false mutations.
[0260] Introducing SDE during incision translation
[0261] In some setups, such as in commercially available kits for adaptor ligation for the Ion Torrent™ platform, a double-stranded adaptor is ligated to the double-stranded target DNA molecule to be sequenced. However, here, only one of the two strands of the target DNA molecule is ligated to the adaptor. This is a common practice when the 5′ strand of the ligation domain is non-phosphorylated. In a process generally known as “nick translation,” a polymerase with strand displacement activity is then used to copy the sequence from the ligated strand to the unligated strand. If the adaptor design disclosed herein is used in this manner without modification, in many cases SDE will be lost during the nick translation step; thus, double sequencing is prevented. This is illustrated below.
[0262] Figure 12A The image shows a type of dual-sequencing adaptor. N′ represents a degenerate or semi-degenerate SMI sequence; TT, opposite to GG, is a non-complementary SDE region; and the asterisk system indicates non-linked dephosphorylated 5′ bases.
[0263] exist Figure 12A After the adaptor is ligated to the double-stranded DNA molecule, an unligated nick remains, such as... Figure 12B As shown in the image.
[0264] Using the standard "cut translation" pathway, strand-moving polymerase extends the 3' end of the library DNA molecule and replaces the unligated strand of the adaptor. This is shown in... Figure 12C In the middle. After the extension, non-complementary SDEs such as Figure 12D The loss is shown in the figure. When SDE is lost, duplex sequencing cannot be performed because the strands are indistinguishable.
[0265] One approach that allows the use of cut-out translation methods with connectors while preserving the SDE is as follows.
[0266] Figure 12E The image shows an example of an Ion Torrent™ adaptor “A” that has been modified to include a degenerate or semi-degenerate SMI sequence. Note that an SDE is not present. “A” is the primer site. The asterisk system indicates a non-phosphorylated 5′ base. Figure 12F The image shows an example of the Ion Torrent™ P1 primer. P1 indicates the primer site. The asterisk system indicates dephosphorylation of the 5′ base.
[0267] exist Figure 12E and Figure 12F Each adaptor connects to double-stranded DNA, forming Figure 12G The structure is shown. Products with two P1 or two A primer sites are not shown because they will not amplify in clusters. For clarity, unconnected linker strands are also not shown.
[0268] Subsequently, the chain-moving polymerase is added according to a typical nick translation protocol (e.g., Bst polymerase, as used in some industrial kits, due to its strong chain displacement activity). However, as... Figure 12H As shown, in this example of dGTP, only one of the four dNTPs is added first, and T-dGTP mis-incorporation will occur (it is worth noting that this mis-incorporation event can occur with a variety of DNA polymerases under appropriate reaction conditions; see, for example, McCulloch and Kunkel, Cell Research 18:148-161 (2008) and the references cited therein).
[0269] Although mismatch incorporation can be extremely efficient under certain conditions, the extension and generation of a second mismatch are extremely inefficient (McCulloch and Kunkel, 2008). Therefore, under appropriate conditions, nucleotide incorporation will cease after a mismatch occurs. At this point, the remaining three dNTPs can be added so that the polymerase can utilize all four dNTPs. The remainder of the adaptor sequence is copied to form... Figure 12I The structure shown has non-complementary positions so that the amplification product of the "top" chain can be distinguished from the amplification product of the "bottom" chain.
[0270] After PCR, the product generated from the initial "top" strand will be as follows: Figure 12J As shown in the diagram, the PCR product generated from the initial "bottom" strand will be as follows: Figure 12K As shown in the image.
[0271] Sequencing of the "top" chain product will produce Figure 12L The structure shown, and sequencing of the "bottom" chain product will produce Figure 12M The structure shown.
[0272] It should be noted that sequencing products can be distinguished from each other based on the introduced mismatches.
[0273] The following section shows a specific instance of weakening this concept to practice with the Ion Torrent™ connector.
[0274] The Ion Torrent™ connector can use the following sequences:
[0275] Connector P1
[0276] (SEQ ID NO:15)
[0277] (SEQ IDNO: 16)
[0278] Connector A
[0279] (SEQ ID NO: 17)
[0280] (SEQ ID NO: 18)
[0281] The asterisk system “*” represents a thiophosphate bond.
[0282] The sequence of adaptor A can be modified as follows. NNNN indicates a degenerate or semi-degenerate SMI sequence (showing four nucleotides, but this sequence length is arbitrary), and MMMM indicates the complementary sequence of NNNN. As previously described, double sequencing can be performed without using the SMI sequence, but here we show a specific instance of the concept of applying a double-stranded molecular marker to SMI.
[0283] Modified connector A
[0284] (SEQ ID NO: 19)
[0285] (SEQ ID NO:20)
[0286] The adaptors A and P1 are attached to opposite ends of the DNA molecule to be sequenced. For simplicity, only the end of the molecule with adaptor A is shown, and for the same simplicity, the two strands are shown as X′ and Y′, respectively. Any DNA sequence of any length can be used, as long as the length of the sequencing fragment is compatible with the sequencing method used.
[0287] If the "top" chain is connected, but the "bottom" chain is not connected, leave a cut (displayed as |).
[0288] (SEQ IDNO: 21)
[0289] (SEQ ID NO: 22)
[0290] Chain displacement polymerase is added along with dGTP. G is incorporated at the first position encountered in the 5′–3′ direction (correct incorporation of G opposite C) and the second position encountered (incorrect incorporation of G opposite A). Because incorrect base extension is inefficient after mismatch, under appropriate polymerase concentration, reaction time, and buffer conditions, and with the polymerase ceasing operation and no further incorporation occurring, it is important to note that the first two nucleotides of the “bottom” adaptor strand are substituted during this reaction, and the adaptor-DNA construct is shown in the following schematic diagram. Newly incorporated bases are indicated in bold.
[0291]
[0292]
[0293] \
[0294] TG
[0295] (Top: SEQ ID NO:23 and Bottom: SEQ ID NO:24)
[0296] Now, dCTP, dATP, and dTTP are added to the reaction so that all four nucleotides are available to the polymerase. For illustrative purposes, chain substitution synthesis can be performed using the intermediates shown below:
[0297]
[0298]
[0299] \
[0300]
[0301] (Top: SEQ ID NO: 25, Middle: SEQ ID NO: 26, and Bottom: SEQ ID NO: 27)
[0302] After reaching the end of the template, the initial “bottom” chain of the linker is completely replaced (not shown) and the fully synthesized “bottom” chain exists with a single non-complementary base pair (A:G base pair, underlined).
[0303] (SEQ IDNO: 28)
[0304] (SEQ IDNO: 29)
[0305] This construct can then be used for PCR amplification and sequencing according to a typical Ion Torrent™ protocol. Notably, PCR amplification produces products from the “top” and “bottom” strands, and these products are distinguishable from each other by non-complementary base pairs introduced during nick translation.
[0306] The product generated from the "top" chain will have the following form (the positions of base mismatches are underlined):
[0307] (SEQ ID NO: 30)
[0308] In contrast, the product generated by the "bottom" chain will have the following form (the positions of the base mismatches are underlined):
[0309] (SEQ ID NO: 31)
[0310] It should be noted that the "bottom" strand product is the reverse complementary sequence of the sequence that first exists in the "bottom" strand of the DNA connected by the adaptor (and therefore, the G nucleotide is read out during sequencing of C nucleotides, which is a base misinsertion introduced during nick translation).
[0311] Now, for error correction, the amplified replicas produced by each of the two strands can be compared with each other. The "top strand" product produced from a given molecule of double-stranded DNA will have the tag sequence NNNNAAC. In contrast, the "bottom strand" product will have the tag sequence NNNNACC. Thus, for error correction purposes, the replicas of both strands can be decomposed, as previously described (Schmitt et al., PNAS 2012).
[0312] Mismatch introduced after incision translation
[0313] The aforementioned alternative approach would involve a complete cleavage translation using all four available nucleotides, followed by changing the bases in the template strand to different bases.
[0314] The adaptor containing the primer sequence and its complementary sequence (P / P′), the UA base pair (U = uracil), and the single-stranded SMI sequence and its complementary sequence (α / α′) is shown in Figure 13A In the middle; the asterisk system indicates the dephosphorylated 5′ end.
[0315] exist Figure 13A After the adaptor is ligated to the double-stranded DNA molecule to be sequenced, the single-strand nick remains at the dephosphorylation site, such as... Figure 13B As shown in the diagram. Here, the "top" strand is linked by a 5' phosphate ester in the target DNA molecule, but the "bottom" strand, due to the absence of a 5' phosphate ester in the adaptor, does not link to the target DNA, leaving a nick.
[0316] Chain substitution synthesis can be performed using polymerases (such as Bst polymerase) and all four dNTPs, producing... Figure 13C The structure shown.
[0317] The resulting extended product now reappears just as it did in the initial connective. For example... Figure 13D As shown in the figure, but there are no asymmetric sites.
[0318] A purification step can be performed to remove polymerase and dNTPs. Uracil can then be removed from the "top" strand by adding uracil DNA glycosylase and an appropriate AP endonuclease. Figure 13D The structure shown is removed, resulting in a product as shown in the image. Figure 13E The single nucleotide gap shown in the figure.
[0319] Subsequently, a non-strand substitution polymerase (e.g., *Sulphurella* DNA polymerase IV, which is highly error-prone and facilitates base mis-incorporation) and a single nucleotide, such as dGTP, are added, but no other nucleotides. In this example, this will produce a G-relative to A mis-incorporation. The resulting nick can be sealed with DNA ligase, producing a product with a mismatch in the adaptor, such as... Figure 13F As shown in the image.
[0320] like Figure 13G As shown, after amplification and sequencing, based on sequencing reads carrying the same SMI sequence, the products produced by the "top" strand can be distinguished from those produced by the "bottom" strand by means of having G or T.
[0321] This example is illustrated using GA mismatches, but it will be obvious that any other mismatch of one or more bases at any position in the molecule will have the same effect.
[0322] The following section shows a specific example of applying this concept on the Ion Torrent™ platform.
[0323] Consider the following “modified adaptor A”, where the standard sequence (U = uracil) is added in bold:
[0324] (SEQ ID NO: 32)
[0325] (SEQ ID NO:33)
[0326] The adaptor is linked to the target DNA molecule as described above, where the cleavage site is shown as "|":
[0327] (SEQ IDNO: 34)
[0328] (SEQ ID NO: 35)
[0329] Now, in the presence of all four dNTPs, chain substitution polymerase is used to allow full chain substitution of the "bottom chain" of the adaptor (newly incorporated bases are in bold; the initial bottom adaptor chain is substituted and not shown):
[0330] (SEQ IDNO: 36)
[0331] (SEQ IDNO: 37)
[0332] The product was purified to remove dNTPs, followed by the addition of uracil DNA glycosylase and AP endonuclease to remove uracil from the "top" strand, leaving a single nucleotide gap:
[0333] (SEQ IDNO: 38)
[0334] (SEQ IDNO: 39)
[0335] Subsequently, a non-strand substitution-prone polymerase (e.g., *Sulphurella* DNA polymerase IV) is added along with dGTP, which causes G incorporation relative to A at the single nucleotide nick; a ligase can then be added to produce a complete adaptor-DNA product on the "top" strand. This results in non-complementary base pairs (positions underlined).
[0336] (SEQ IDNO: 40)
[0337] (SEQ IDNO: 41)
[0338] This product can be used for error correction, utilizing methods similar to those described in the preceding embodiments.
[0339] Mismatch introduced after incision translation
[0340] The example titled "Introduction of SDEs During Cut Translation" demonstrates how asymmetric SDEs can be introduced into the adaptor sequence during cut translation. The same principle can be applied to DNA libraries themselves to incorporate asymmetric sites (SDEs) into the library molecule, possibly even before adaptor addition. This can be achieved in several ways. The following is merely one example.
[0341] Double-stranded DNA molecules with "top" and "bottom" strands are shown in Figure 14A DNA molecules can be fragmented using various methods for library preparation. For example, some DNA sources, such as cell-free DNA from plasma, are already in small fragments and do not require a separate fragmentation step. Acoustic shearing is a commonly used method. Semi-random enzymatic shearing methods can be used. Non-random endonucleases that cut at specified recognition sites are another method. In this example, endonucleases leaving the 5' overhang are used to generate libraries with similar 5' overhang fragments, such as... Figure 14B As shown in the image.
[0342] This asymmetric state can be converted into sequence asymmetry by using a polymerase in the presence of a single nucleotide that is not complementary to the first nucleotide copied by the polymerase. In this example, dGTP is used, which will result in T-dGTP mis-incorporation (such mis-incorporation can be carried out with a variety of DNA polymerases under appropriate reaction conditions; see McCulloch and Kunkel, Cell Research 18:148-161 (2008), and the references cited therein). Partial double-stranded DNA molecules including two mismatches are shown in Figure 14C middle.
[0343] Subsequently, all four nucleotides were added to the reaction and copied continuously to extend the DNA molecule to a double helix. Mismatch bubbles were generated at the ends of each fragment, forming two SDEs, such as... Figure 14D As shown in the image.
[0344] The dual-sequencing adaptor can then be ligated into the DNA molecule. Figure 14E The exemplary linker shown has primer site P and complementary sequence P′, different primer sites O and complementary sequence O′, and degenerate or semi-degenerate SMI α and complementary sequence α′.
[0345] exist Figure 14D Double-stranded DNA molecules and Figure 14E Connect the connectors to generate Figure 14F The structure. As discussed in the foregoing embodiments, products linked to two identical adapter sequences can be ignored because they will not amplify under appropriate conditions.
[0346] Following PCR, products derived from the "top" strand, such as Figure 14G As shown, and products derived from the "bottom" chain, such as Figure 14H As shown in the image.
[0347] Sequencing using primer P will be performed by Figure 14I The corresponding chains shown produce the following sequences.
[0348] It should be noted that the presence of C and T after the SMI sequence allows the “top” chain reads to be distinguished from those derived from the “bottom” chain.
[0349] Similar SDE markers can be achieved by filling the 3′ indentation gap with mutagenic nucleotide analogs or other methods.
[0350] Other cutting methods can be used, and the 3′ indentation can be generated by an exonuclease in a manner that produces SDEs before filling.
[0351] In a broader sense, this example illustrates that SDEs can be introduced independently of the adaptor itself. For duplex sequencing, only one form of SMI and SDE in each ultimately adapted molecule allows the sequences derived from each strand of the double helix to be correlated with each other, yet also definitively distinguished from one another. These elements appear in a variety of forms as considered above and can be introduced before, during, or after adaptor ligation.
[0352] Changes in assembling molecules suitable for dual sequencing
[0353] The embodiments disclosed above illustrate an improved method for duplex sequencing, wherein the assembled final molecule comprises at least one strand-defining element (SDE) and at least one single-molecule identifier (SMI) sequence; both the SDE and SMI are attached to a double-stranded or partially double-stranded molecule of the DNA to be sequenced. However, the SMI and SDE do not need to be included in a single adaptor; they simply need to be present in the final molecule, ideally before or during any amplification and / or sequencing steps.
[0354] For example, SDE can be generated in the adaptor after enzymatic ligation, such as... Figure 4D As shown in the diagram. Similarly, as previously described (in WO2013142389A1), in some embodiments, a specific sequence at the cut site of an individual DNA library fragment can serve as an endogenous SMI sequence, without the need to add an exogenous SMI included within the adaptor. A “cut site” can be considered as the mapped coordinates of either end of a DNA fragment when the fragment is aligned with a reference genome. The coordinates of either end or both ends can be used as an “endogenous SMI” to distinguish different DNA molecules from each other, either alone or in combination with one or more exogenous SMI sequences.
[0355] The following list includes non-restrictive variants of this type of connective:
[0356] --SDE exists in both strands, but SMI and primer binding sites exist in only one adaptor strand. These elements are then copied to the other strand using polymerase.
[0357] --SDE is absent; SMI and primer binding site are in only one strand. The polymerase is used with only one incorrect dNTP present to produce SDE, and then the remaining dNTP is added to allow the polymerase to produce a double strand of SMI and primer binding site.
[0358] --The linker domain exists only in one linker strand (so that the second linker strand is not attached). The new second linker strand is then copied from the first linker strand using polymerase. This produces the SMI and primer-binding domain. As described above, only one incorrect dNTP is added first to produce the SDE; the remaining dNTPs are then added. This pathway is shown in the examples disclosed above.
[0359] --The linker domain exists only in one linker strand (so that the second linker strand is not attached); this linker strand contains uracil. The new second linker strand is then copied from the first linker strand using a polymerase with all four nucleotides present. Subsequently, the uracil bases in the initial linker strand are enzymatically removed using uracil DNA glycosylase and an appropriate AP endonuclease. DNA polymerase is then used with the single mismatched nucleotide present to insert the mismatch into the gap in the DNA, and the gap is subsequently ligated using DNA ligase. This pathway is shown in more detail above regarding... Figures 4A to 4H In the disclosed embodiments.
[0360] --The first attaching linker has an SMI domain in both strands. The second linker is then attached to it, which has a primer-binding domain and an SDE, also in both strands.
[0361] --The first attaching linker has SMI and SDE domains in both strands. The second attaching linker has primer-binding domains in both strands.
[0362] --The first attached linker has an SMI domain in both strands. A second "Y linker" is then attached, which has two non-complementary or partially non-complementary primer-binding domains.
[0363] --The first-attached adaptor has an SMI in both strands and the single-stranded region, similarly with the linker domain. The oligonucleotide anneals and is linked to the single-stranded region; mismatches are included within the oligonucleotide that generates the SDE domain.
[0364] --In other embodiments, the position of the bubble can be changed, the length of the n-mer can be changed, and the n-mer can be eliminated along with a copy of the cleavage site at each strand of the identified DNA molecule, rather than at the end of the DNA molecule. Variant nucleotides or nucleotide-like molecules can be used within DNA (e.g., locked nucleic acids (LNAs) and peptide nucleic acids (PNAs), as well as RNA).
[0365] Each of the variants disclosed herein is included in this invention.
[0366] In each of these variants, the same general concept applies: the final molecule of duplex sequencing contains core elements of an SDE and an SMI attached to the strand of the DNA to be sequenced. It should also be noted that the same general concept applies to the initial description of duplex sequencing (in WO2013142389A1), where duplex sequencing is performed with an adaptor containing two asymmetric primer binding sites (e.g., in the “Y” configuration) (which in this case acts as the SDE) and an SMI sequence attached to the double-stranded DNA molecule. These components can be assembled onto the target DNA molecule in various ways, provided the necessary components are present in the final molecule, ideally before or during any amplification or sequencing step.
[0367] Alternative data processing workflows for dual sequencing
[0368] Two single-stranded common sequences can be obtained by duplex sequencing of the "common sequences" of amplified copies produced from each of the two individual DNA strands, followed by comparison of the resulting single-stranded common sequences to obtain the double-helix common sequence. In some settings (e.g., in cases of recurrent amplification errors that may occur at a given location in severely damaged DNA), this approach of "averaging" the sequences of amplified copies of individual molecules position-by-position may not be necessary, and therefore reliable results can be obtained in various settings across different data processing pipelines.
[0369] Alternative approaches include the following:
[0370] --In molecules with given tag sequences corresponding to the "top" and "bottom" chains, arbitrarily select one "top" chain and one "bottom" chain, and compare the sequences of the two chains. Maintain positions where both chains are identical; mark positions where they are inconsistent as unrestricted. The resulting sequence read is a double helix read.
[0371] -- Repeat this method for any chosen "top" and "bottom" chains that share the same tag sequence to obtain a series of "double helix reads".
[0372] --From the resulting "double helix reads" with a given tag sequence, select the double helix read with, for example, the fewest sequence variations relative to a reference sequence and / or the fewest unrestricted positions within the read. This read can then be considered the read most likely to represent the true sequence of the originating DNA double helix.
[0373] In one embodiment, such an approach may be implemented, in particular, using the algorithm described below. It should be understood that this is merely a single example for illustrative purposes, and many other algorithms can be used to form double-helix shared reads. Furthermore, while an example of a specific embodiment of duplex sequencing is shown, many other similar examples suitable for duplex sequencing can be prepared.
[0374] The following steps can be used in the embodiments disclosed herein, which use “bubble” sequences to generate a “top” chain labeled GCGC and a “bottom” chain labeled TATA, wherein the two chains share the same single-molecule identifier (SMI) sequence.
[0375] 1. Create a file containing all sequencing reads from the experiment;
[0376] 2. Divide the file into two files: one file is called the read segment containing "GCGC" marked with GCGC, and the second file is called the read segment containing "TATA" marked with TATA;
[0377] 3. Select an arbitrary segment in the “GCGC” file, read its SMI tag, and search for a matching SMI tag in the “TATA” file;
[0378] 4. If a match is found, a new sequence is generated from the two sequences. In the new sequence, all sequence positions within the consistent read segment are maintained, and all inconsistent positions between the two read segments are marked as unqualified. This new sequence is written to a file called the "double helix," and both sequences are removed from the "GCGC" and "TATA" files.
[0379] If no match is found: then the sequence is removed from the "GCGC" file and written to a file called "No Match";
[0380] 5. Select another arbitrary segment from the "GCGC" file, and repeat steps 3 and 4; and
[0381] 6. Continue until there are no more read segments remaining in the "GCGC" file.
[0382] Within the resulting "double helix" file, assume all reads have matching SMI tag sequences. In some cases, multiple "double helix" reads with the same SMI tag may exist (these can be attributed, for example, to multiple PCR replicas of a single starting DNA molecule). These can be converted into a single double helix read through any of the following pathways:
[0383] --Of these reads, select the read with the fewest mismatches relative to the reference genome sequence and discard the remaining reads.
[0384] --Alternatively, select the read with the fewest non-restricted positions relative to the reference genome sequence and discard the remaining reads.
[0385] --Alternatively, a common sequence is generated in the reads with a common SMI tag sequence to generate a double helix common sequence read.
[0386] Those skilled in the art will readily recognize that the combination of the above options can be used to generate double-helix shared sequence reads, or several other methods not described above can be used.
[0387] Other embodiments
[0388] Although the invention has been described in conjunction with its embodiments, the foregoing description is intended to illustrate and not limit the scope of the invention as defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the appended claims.
Claims
1. An adaptor nucleic acid sequence pair for sequencing double-stranded target nucleic acid molecules, comprising a first adaptor nucleic acid sequence and a second adaptor nucleic acid sequence, wherein each adaptor nucleic acid sequence comprises: Primer binding domain, Chain-limited element (SDE). Single Molecular Identifier (SMI) domain, and Connect structural domains; The SDE of the first adaptor nucleic acid sequence is at least partially non-complementary to the SDE of the second adaptor nucleic acid sequence.
2. The adaptor nucleic acid sequence pair according to claim 1, wherein the two adaptor sequences consist of two separate DNA molecules that are at least partially annealed together.
3. The adapter nucleic acid sequence pair according to claim 1, wherein the first adapter nucleic acid sequence and the second adapter nucleic acid sequence are connected via a linker domain.
4. The adaptor nucleic acid sequence pair according to claim 3, wherein the linker domain is composed of nucleotides.
5. The adaptor nucleic acid sequence pair according to claim 3, wherein the adaptor domain contains one or more modified nucleotides or non-nucleotide molecules.
6. The adaptor nucleic acid sequence pair according to claim 4, wherein the one or more modified nucleotides or non-nucleotide molecules are selected from the following: nucleotides without a base site; uracil; tetrahydrofuran; 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A); 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G); deoxyinosine; 5′-nitroindole; 5-hydroxymethyl-2'-deoxycytidine; isocytosine; 5′-methyl-isocytosine; or isoguanosine.
7. The adaptor nucleic acid sequence pair according to any one of claims 3 to 5, wherein the connector domain forms a loop.
8. The adaptor nucleic acid sequence pair according to any one of claims 1 to 7, wherein the SDE of the first adaptor nucleic acid sequence is not complementary to the SDE of the second adaptor nucleic acid sequence.
9. The adapter nucleic acid sequence pair according to any one of claims 1 to 7, wherein the primer-binding domain of the first adapter nucleic acid sequence is at least partially complementary to the primer-binding domain of the second adapter nucleic acid sequence.
10. The adapter nucleic acid sequence pair according to any one of claims 1 to 9, wherein the primer-binding domain of the first adapter nucleic acid sequence is complementary to the primer-binding domain of the second adapter nucleic acid sequence.
Citation Information
Patent Citations
Methods of lowering the error rate of massively parallel DNA sequencing using duplex consensus sequencing
WO2013142389A1