Hairpin adapter constructs and primers to generate next-generation sequencing library molecules
Hairpin duplex molecule constructs and primers improve DNA sequencing accuracy by enabling two-pass sequencing, addressing the challenges of brief signal detection and enhancing data quality in sequencing libraries.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- F HOFFMANN LA ROCHE & CO AG
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing DNA sequencing methods suffer from reduced accuracy due to brief signal presence in detectors, necessitating multiple reads to achieve consensus sequences, which can be time-consuming and costly.
The development of hairpin duplex molecule constructs and primers for two-pass sequencing, utilizing oligonucleotide adapters with modified stem regions and hairpin structures to facilitate consensus sequencing while retaining methylation information.
Enhances sequencing accuracy by allowing for consensus sequence determination with retained methylation information, improving data quality and efficiency in generating sequencing libraries.
Smart Images

Figure EP2025081169_07052026_PF_FP_ABST
Abstract
Description
HAIRPIN ADAPTER CONSTRUCTS AND PRIMERS TO GENERATE NEXTGENERATION SEQUENCING LIBRARY MOLECULESSEQUENCE LISTING INCORPORATION BY REFERENCEBACKGROUND OF THE DISCLOSURE
[0001] DNA sequencing is a fundamental tool in biological and medical research. The importance of DNA sequencing has increased dramatically from its inception four decades ago. It is recognized as a crucial technology for most areas of biology and medicine and as the underpinning for the new paradigm of personalized and precision medicine. Information on individuals' genomes and epigenomes can help reveal their propensity for disease, clinical prognosis, and response to therapeutics, but routine application of genome sequencing in medicine requires comprehensive data delivered in a timely and cost-effective manner.
[0002] Compiling the sequential reads during sequencing allows one to construct the sequence of the sample. For single molecule sequencing methods, the sequencing accuracy is diminished because of the brief time that a signal representing a base is present in the detector. To improve accuracy, the sample DNA is often read multiple times at the same sensor, and consensus reads are compiled for each sample fragment. Indeed, one way to reduce sequencing error is to determine a consensus sequence by sequencing the target many times, thereby achieving a desired consensus sequence accuracy.BRIEF SUMMARY OF THE DISCLOSURE
[0003] Applicant has developed methods for preparing libraries including nucleic acid molecules suitable for two-pass sequencing. In particular, the disclosed methods facilitate the formation of a library including one or more hairpin duplex molecules which are suitable for two- pass sequencing. The disclosed methods use current state-of-the art library preparation steps while permitting for standard PCR amplification and target enrichment methods prior to conversion to into hairpin duplex molecules for two-pass sequencing.
[0004] The methods of the present disclosure facilitate the generation of consensus sequencing reads using the sequencing information from the original template and the copied product. As discussed herein, this facilitates the retention of methylation information from the original template strand while assessing the copied strand for additional sequence information.
[0005] A first aspect of the present disclosure is an oligonucleotide adapter comprising a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity;wherein a 5' portion of the first strand and a 3' portion of the second strand are each single stranded and non-complementary; and wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence.
[0006] In some embodiments, the substantially double-stranded stem region includes one or more modifications to prevent digestion from an exonuclease. In some embodiments, the substantially double-stranded stem region includes one or more phosphorothioate bonds. In some embodiments, the substantially double-stranded stem region includes one or more phosphorothioate bonds, wherein the one or more phosphorothioate bonds are positioned at end of the first strand and / or the second strand of the double-stranded region. In some embodiments, the substantially double-stranded stem region includes two or more phosphorothioate bonds.
[0007] In some embodiments, the 3' portion of the second strand comprises a region including a primer binding sequence. In some embodiments, the region including the primer binding sequence comprises a recognition site for a nicking enzyme. In some embodiments, the optional first loop sequence includes a primer binding sequence, which may optionally include a recognition site for a nicking enzyme.
[0008] In some embodiments, the first hairpin region forms a first loop-like structure at temperatures below about 50°C. In some embodiments, the first hairpin region forms a first looplike structure at temperatures above about 50°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures below about 45°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures below about 40°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures below about 35°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures of about 58°C.
[0010] In some embodiments, the 3' portion of the second strand comprises a second hairpin region including two second duplex sequences, wherein each of the two second duplex sequences are separated by an optional second loop sequence. In some embodiments, the two first duplex sequences have different nucleotide sequences than the two second duplex sequences. In some embodiments, each of the first and second hairpin regions form independent loop-like structures at a temperature below about 50°C, such as below about 45°C, such as below about 40°C, such as below about 35°C, etc. In some embodiments, each of the first and second hairpin regions form independent loop-like structures at a temperature above about 50°C, such as above about 55°C, such as about 58°C, etc.
[0011] In some embodiments, the oligonucleotide adapter further includes one or more barcodes. In some embodiments, the one or more barcodes are located within the substantiallydouble-stranded stem region. In some embodiments, the one or more barcodes are located within either the first or second hairpin regions. In some embodiments, one or more barcodes are included within the substantially double-stranded stem region; and one or more barcodes are located within either the first and / or second hairpin regions.
[0012] In some embodiments, the oligonucleotide adapters are utilized in the preparation of one or more nucleic acid molecules for sequencing, such as to facilitate two-pass sequencing.
[0013] A second aspect of the present disclosure is a double stranded nucleic acid molecule ligated to an oligonucleotide adapter, wherein the oligonucleotide adapter comprises a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are each single stranded and non- complementary; and wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence.
[0014] In some embodiments, the substantially double-stranded stem region includes one or more modifications to prevent digestion from an exonuclease. In some embodiments, the substantially double-stranded stem region includes one or more phosphorothioate bonds.
[0015] In some embodiments, the 3' portion of the second strand comprises a region including a primer binding sequence. In some embodiments, the region including the primer binding sequence comprises a recognition site for a nicking enzyme. In some embodiments, the optional first loop sequence includes a primer binding sequence.
[0016] In some embodiments, the oligonucleotide adapters are utilized in the preparation of one or more nucleic acid molecules for sequencing, such as to facilitate two-pass sequencing.
[0017] A third aspect of the present disclosure is an oligonucleotide adapter comprising a first strand and a second strand, wherein a 5' portion of the first strand and a 3' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 3' portion of the first strand and a 5' portion of the second strand are each single stranded and non-compl ementary; and wherein the 3' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence.
[0018] In some embodiments, the substantially double-stranded stem region includes one or more modifications to prevent digestion from an exonuclease. In some embodiments, the substantially double-stranded stem region includes one or more phosphorothioate bonds.
[0019] In some embodiments, the 5' portion of the second strand comprises a region including a primer binding sequence. In some embodiments, the region including the primerbinding sequence comprises a recognition site for a nicking enzyme. In some embodiments, the optional first loop sequence includes a primer binding sequence.
[0020] In some embodiments, the first hairpin region forms a first loop-like structure at temperatures below about 50°C. In some embodiments, the first hairpin region forms a first looplike structure at temperatures below about 45°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures below about 40°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures above about 50°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures above about 55°C. In some embodiments, the first hairpin region forms a first loop-like structure at temperatures of about 58°C.
[0021] In some embodiments, the 5' portion of the second strand comprises a second hairpin region including two second duplex sequences separated by an optional second loop sequence. In some embodiments, the two first duplex sequences have different nucleotide sequences than the two second duplex sequences. In some embodiments, each of the first and second hairpin regions form independent loop-like structures at a temperature below about 50°C, such as below about 45°C, such as below about 40°C, such as below about 35°C, etc. In some embodiments, each of the first and second hairpin regions form independent loop-like structures at a temperature above about 50°C, such as above about 55°C, such as at about 58°C, etc.
[0022] In some embodiments, the adapter further includes one or more barcodes. In some embodiments, the one or more barcodes are located within the substantially double-stranded stem region. In some embodiments, the one or more barcodes are located within either the first or second hairpin regions.
[0023] A fourth aspect of the present disclosure is a double stranded nucleic acid molecule ligated to an oligonucleotide adapter, wherein the oligonucleotide adapter comprises a first strand and a second strand, wherein a 5' portion of the first strand and a 3' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 3' portion of the first strand and a 5' portion of the second strand are each single stranded and non- complementary; and wherein the 3' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence. In some embodiments, the substantially double-stranded stem region includes one or more modifications to prevent digestion from an exonuclease. In some embodiments, the substantially double-stranded stem region includes one or more phosphorothioate bonds.
[0024] In some embodiments, the 3' portion of the second strand comprises a region including a primer binding sequence. In some embodiments, the region including the primerbinding sequence comprises a recognition site for a nicking enzyme. In some embodiments, the optional first loop sequence includes a primer binding sequence.
[0025] A fifth aspect of the present disclosure is an oligonucleotide adapter comprising a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are each single stranded and non-complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by a first optional loop sequence; and wherein the 3' portion of the second strand comprises a second hairpin region including two second duplex sequences that are substantially complementary to each other, wherein the two second duplex sequences are separated by a second optional loop sequence. In some embodiments, the oligonucleotide adapters of the fifth aspect of the present disclosure may be ligated to one or more nucleic acid molecules. In some embodiments, the oligonucleotide adapters are utilized in the preparation of one or more nucleic acid molecules for sequencing, such as to facilitate two-pass sequencing.
[0026] A sixth aspect of the present disclosure is a double stranded nucleic acid molecule ligated to an oligonucleotide adapter, wherein the oligonucleotide adapter comprises a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are each single stranded and non- complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by a first optional loop sequence; and wherein the 3' portion of the second strand comprises a second hairpin region including two second duplex sequences that are substantially complementary to each other, wherein the two second duplex sequences are separated by a second optional loop sequence. In some embodiments, each of the first and second hairpin regions form independent loop-like structures at a temperature below about 50°C, such as below about 45°C, such as below about 40°C, such as below about 35°C, etc. In some embodiments, each of the first and second hairpin regions form independent loop-like structures at a temperature above about 50°C, such as above about 55°C, such as at about 58°C, etc.
[0027] A seventh aspect of the present disclosure is a composition comprising one or more polynucleotides having the formula:
[0028] [first adapter] - [insert] - [second adapter];
[0029] wherein the insert is double stranded, a region of the first adapter in proximity to the insert is double stranded, a region of the second adapter in proximity to the insert is double stranded, a region of the first adapter distal to the insert includes two single strands each having an end, and a region of the second adapter distal to the insert comprises two single strands each having an end, wherein at least one strand of the two single strands of the first and second adapters includes a hairpin-forming sequence. In some embodiments, another strand of the two single strands of the first and second adapters includes a primer binding site. In some embodiments, the insert is derived from a DNA sample, such as a DNA sample which has been fragmented.
[0030] In some embodiments, both strands of the two single strands of the first and second adapters include hairpin-forming sequences. In some embodiments, the hairpin-forming sequences include two at least partially complementary duplex sequences separated by an optional loop sequence. In some embodiments, each of the hairpin-forming sequences form a loop-like structure at each end of each strand of the polynucleotide at a temperature below about 50°C, such as at a temperature below about 45°C, at a temperature below about 40°C, at a temperature below about 35°C, etc. In some embodiments, the hairpin-forming sequences include two at least partially complementary duplex sequences separated by an optional loop sequence. In some embodiments, each of the hairpin-forming sequences form a loop-like structure at each end of each strand of the polynucleotide at a temperature above about 50°C, such as at a temperature above about 55°C, at a temperature at about 58°C, etc. In some embodiments, each end of the two single strands of the first and second adapters are modified to prevent digestion by an exonuclease. In some embodiments, the each of the two single strands of the first and second adapters comprises a phosphorothioate bond. In some embodiments, each of the first and second adapters include one or more barcodes.
[0031] An eighth aspect of the present disclosure is a method for preparing a library of nucleic acid molecules comprising: (a) obtaining a sample including one or more nucleic acid molecules; (b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acid molecules, where each of the first and second adapters include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non-complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; and wherein the 3' portion of the second strand includes a primer binding site; (c) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleicacid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and (d) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
[0032] In some embodiments, the method further comprises amplifying the one or more adapter ligated nucleic acid molecules prior to the synthesizing of the hairpin duplex molecule. In some embodiments, the amplification comprises isothermal amplification.
[0033] In some embodiments, the method further comprises performing target enrichment.
[0034] In some embodiments, each duplex sequence has a Tm ranging from between about23°C to about 50°C.
[0035] In some embodiments, the synthesizing of the hairpin duplex molecule comprises (i) subjecting the sample including the one or more adapter ligated nucleic acid molecules to conditions in which a self-priming event occurs between the substantially complementary sequences of the two first duplex sequences to form a hairpin between the two first duplex sequences; and (ii) extending the formed hairpin. In some embodiments, the conditions are a temperature below about 50°C. In some embodiments, the conditions are a temperature below about 45°C. In some embodiments, the conditions are a temperature below about 40°C. In some embodiments, the conditions are a temperature above about 50°C. In some embodiments, the conditions are a temperature above about 55°C. In some embodiments, the conditions are a temperature at about 58°C.
[0036] In some embodiments, the formed hairpin is extended with a polymerase. In some embodiments, the polymerase is a low temperature polymerase. In some embodiments, the polymerase is selected form the group consisting of 029 DNA polymerase, E.coli DNA polymerase, BSU, and BST. In some embodiments, the extension is carried out in the presence of 5-Methylcytidine-5'-Triphosphate.
[0037] In some embodiments, the method further comprises converting unmethylated cytosine to uracil. In some embodiments, the converting of the unmethylated cytosine to uracil comprises performing a bisulfite treatment. In some embodiments, methylated cytosines are converted to one of 5hmC, 5fC, or 5caC.
[0038] In some embodiments, generating of the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing comprises introducing a sequencing primer to the sample including the one or more hairpin duplex molecules. In some embodiments, the sequencing primer is substantially complementary to the primer binding site.
[0039] In some embodiments, the substantially double stranded stem region of the first and second adapters includes one or more exonuclease blocking elements; and wherein the generatingof the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing further comprises introducing an exonuclease to the sample including the one or more hairpin duplex molecules. In some embodiments, the exonuclease has 5' to 3' activity. In some embodiments, the exonuclease having 5' to 3' activity is a T7 exonuclease.
[0040] In some embodiments, the substantially double stranded stem region of the first and second adapters includes a recognition site for a nicking enzyme; and wherein the generating of the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing further comprises introducing a nicking enzyme specific to the recognition site. In some embodiments, the nicking enzyme is selected from the group consisting of N.Bst9I, N.BstSEI, Nb.BbvCI(NEB), Nb. BpulOI(Fermantas), Nb.BsmI(NEB), Nb.BsrDI(NEB), Nb.BtsI(NEB), Nt.AlwI(NEB), Nt.BbvCI(NEB), Nt.BpulOI(Fermentas), Nt.BsmAI, Nt.BspD6I, Nt.BspQI(NEB), Nt.BstNBI(NEB), and Nt.CviPII(NEB).
[0041] In some embodiments, the optional loop sequence is present and wherein the loop sequence includes a primer binding site; and wherein the generating of the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing further comprises introducing a polymerase to open the hairpin at the 5' end of the one or more hairpin duplex molecules. In some embodiments, the polymerase is selected from the group consisting of BST, BSU, and Phi29.
[0042] In some embodiments, the method further comprises sequencing the library including the one or more synthesized hairpin duplex molecules which primed for sequencing. In some embodiments, the sequencing comprises next-generation sequencing. In some embodiments, the sequencing comprises nanopore sequencing. In some embodiments, the sequencing comprises Sequencing by Expansion. In some embodiments, the sequencing comprises bisulfite sequencing. In some embodiments, the sequencing comprises methylation sequencing.
[0043] A ninth aspect of the present disclosure is a method for preparing a library of nucleic acid molecules comprising: (a) obtaining a sample including one or more nucleic acid molecules; (b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acid molecules, where each of the first and second adapters include a first strand and a second strand, wherein a 5' portion of the first strand and a 3' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 3' portion of the first strand and a 5' portion of the second strand are single stranded and non-complementary; wherein the 3' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; and wherein the 5' portion of the second strand includes a primer binding site;(c) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and (d) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
[0044] A tenth aspect of the present disclosure is a method for preparing a library of nucleic acid molecules comprising: (a) obtaining a sample including one or more nucleic acid molecules;(b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acid molecules, where each of the first and second adapters include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non-complementary; wherein a 5' end of the first strand comprises a hairpin precursor region; and wherein a 3' portion of the second strand comprises a region including a primer binding site; (c) amplifying the one or more adapter ligated nucleic acid molecules to provide a double stranded molecule having a 5' end including a hairpin region and a 3' end including a region having a primer bind site; (d) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and (e) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
[0045] An eleventh aspect of the present disclosure is a method for preparing a library of nucleic acid molecules comprising: (a) obtaining a sample including one or more nucleic acid molecules; (b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acid molecules, where each of the first and second adapters include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non-complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; and wherein the 3' portion of the second strand includes a primer binding site;(c) amplifying the one or more adapter ligated nucleic acid molecules to provide a double stranded molecule having a 5' end including a hairpin region and a 3' end including a region having a primerbind site; (d) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and (e) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
[0046] A twelfth aspect of the present disclosure is a method of preparing a hairpin duplex molecule comprising (a) obtaining a double stranded nucleic acid molecule ligated to an oligonucleotide adapter, wherein the oligonucleotide adapter comprises a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are each single stranded and non- complementary; and wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; (b) synthesizing the hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site, wherein the synthesizing of the hairpin duplex molecule comprises initiating a self-priming event to form the hairpin at the 5 ’ end, and extending the formed hairpin. In some embodiments, the extension is performed with a polymerase. In some embodiments, the method further comprises sequencing the synthesized hairpin duplex molecule. In some embodiments, the sequencing comprises two- pass sequencing.
[0047] A thirteenth aspect of the present disclosure is an oligonucleotide primer including a 5’ portion and a 3’ portion, in which the 5’ portion includes a hairpin region including two duplex sequences that are substantially complementary to each other, wherein the two duplex sequences are separated by a loop sequence.
[0048] In some embodiments, the hairpin region forms a loop-like structure at temperatures below about 50°C. In some embodiments, the hairpin region forms a loop-like structure at temperatures above about 50°C. In some embodiments, the primer further includes one or more barcodes. In certain embodiments, the one or more barcodes are located within the 5’ portion. In other embodiments, the one or more barcodes are located within the 3’ portion. In some embodiments, the oligonucleotide primer further includes a 5’ biotin moiety. In other embodiments, the oligonucleotide primer further includes a 5’ phosphate moiety. In some embodiments, the 3’ portion includes a sequence that is complementary to a sequence in a universal adapter. In certain embodiments, the universal adapter is an Illumina adapter.
[0049] A fourteenth aspect of the present disclosure is A method for preparing a library of nucleic acid molecules including (a) obtaining a sample including one or more nucleic acid molecules, in which the one or more nucleic acid molecules include a first end including a first known sequence and a second end including a second known sequence; (b) amplifying the one or more nucleic acid molecules with a first oligonucleotide primer and a second oligonucleotide primer, in which the first oligonucleotide primer includes a 5’ portion and a 3’ portion, in which the 5’ portion includes a hairpin region including two duplex sequences that are substantially complementary to each other, in which the two duplex sequences are separated by a loop sequence, in which the sequence of the 3’ portion is complementary to the first known sequence, and in which the second primer includes a sequence complementary to the second known sequence to provide a sample of amplified one or more nucleic acid molecules; (c) synthesizing a hairpin duplex molecule from each of the amplified one or more nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, in which each of the one or more hairpin duplex molecules is double stranded and includes a first end including a hairpin and a second end lacking a hairpin; and (d) generating a library of one or more synthesized hairpin duplex molecules which are capable of hybridizing to an extension oligonucleotide.BRIEF DESCRIPTION OF THE FIGURES
[0050] For a general understanding of the features of the disclosure, reference is made to the drawings. In the drawings, like reference numerals have been used throughout to identify identical elements.
[0051] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0052] FIG. 1 A illustrates a "full-length" adapter, including a substantially double stranded region, a single-stranded hairpin region, and a single-stranded region including a primer binding site.
[0053] FIG. IB further illustrates the "full-length" adapter of FIG. 1A, showing that the single-stranded hairpin region includes a first and second duplex sequence separated by a loop sequence. In some embodiments, the loop sequence is optional, and the single-stranded hairpin region includes only first and second duplex sequences.
[0054] FIG. 1C illustrates a "full-length" adapter, including a substantially double stranded region, a single-stranded hairpin region, and a single-stranded region including a primer binding site.
[0055] FIG. ID further illustrates the "full-length" adapter of FIG. 1C, showing that the single-stranded hairpin region includes a first and second duplex sequence separated by a loop sequence. In some embodiments, the loop sequence is optional, and the single-stranded hairpin region includes only first and second duplex sequences.
[0056] FIG. 2A illustrates a "full-length" adapter, including a substantially double stranded region and two single-stranded hairpin regions.
[0057] FIG. 2B further illustrates the "full-length" adapter of FIG. 2A, showing that the two single-stranded hairpin regions each include first and second duplexes sequences separated by a loop sequence. In some embodiments, the loop sequences are optional and the single-stranded hairpin regions includes only duplex sequences.
[0058] FIG. 2C illustrates a "full-length" adapter, including a substantially double stranded region and two single-stranded hairpin regions.
[0059] FIG. 2D further illustrates the "full-length" adapter of FIG. 2D, showing that the two single-stranded hairpin regions each include first and second duplexes sequences separated by a loop sequence. In some embodiments, the loop sequences are optional and the single-stranded hairpin regions includes only duplex sequences.
[0060] FIG. 3 A illustrates a "truncated" adapter, including a substantially double stranded region, a single-stranded hairpin precursor region, and a region including a primer binding site.
[0061] FIG. 3B illustrates that the single-stranded hairpin region includes a single duplex sequence (as compared with the hairpin regions of FIGS. IB and 2B, which each include two duplex sequences separated by a loop sequence).
[0062] FIG. 3C illustrates a "truncated" adapter, including a substantially double stranded region, a single-stranded hairpin precursor region, and a region including a primer binding site.
[0063] FIG. 3D illustrates that the single-stranded hairpin region includes a single duplex sequence (as compared with the hairpin regions of FIGS. ID and 2D, which each include two duplex sequences separated by a loop sequence).
[0064] FIGS. 4A and 4B illustrate adapter ligated nucleic acid molecules including "full- length" adapters ligated to an insert molecule. Each strand of the adapter ligated nucleic acid molecule includes a 5' hairpin region and a 3' region including a primer binding site. In some embodiments, the loop sequence is optional, and the single-stranded hairpin region includes only first and second duplex sequences.
[0065] FIGS. 4C and 4D illustrate the structure of an adapter ligated nucleic acid molecule including a "full-length" adapter including one hairpin region. Each strand of the adapter ligated nucleic acid molecule includes a 5' hairpin region and a 3' region including a primer binding site. FIG. 4C further illustrates that the first and second duplex sequences of the hairpin regions (seeFIG. 4B) are each capable of forming a loop-like structures (such as during a self-priming event) at terminal ends of one of the strands of the nucleic acid molecule. Notably, the loop-like structures do not bridge the two strands of the adapter ligated nucleic acid molecule in the adapter ligated nucleic acid molecules.
[0066] FIG. 4E is similar to FIG. 4D but illustrates that the hairpin regions are provided at the 3' ends of the adapter ligated nucleic acid molecules. As such, each strand of the adapter ligated nucleic acid molecule includes a 3' hairpin region and a 5' region including a primer binding site.
[0067] FIGS. 5A and 5B illustrate adapter ligated nucleic acid molecules including "full- length" adapters ligated to an insert molecule. Each strand of the adapter ligated nucleic acid molecule includes a 5' hairpin region and a 3' hairpin region. The 5' hairpin region and the 3' hairpin region may be the same or different.
[0068] FIGS. 5C and 5D illustrates the structure of an adapter ligated nucleic acid molecule including a "full-length" adapter including two hairpin regions.
[0069] FIG. 6 A illustrates the extension of a single strand of an adapter ligated nucleic acid molecule includes a "full-length" adapter. In some embodiments, the loop sequence is optional. In some embodiments, a first primer is at least partially complementary to portions of the optional loop sequence and / or the duplex sequence. In some embodiments, a second primer is at least partially complementary to the region including a primer binding site.
[0070] FIG. 6B illustrates the extension of a single strand of an adapter ligated nucleic acid molecule including a "truncated" adapter, where after extension the "truncated" adapter includes all the elements of a "full-length" adapter. In some embodiments, the loop sequence is optional. In some embodiments, a first primer is at least partially complementary to portions of the duplex sequence. In some embodiments, a second primer is at least partially complementary to the region including a primer binding site.
[0071] FIG. 7 depicts the conversion of a single strand of an adapter ligated molecule into a hairpin duplex molecule in accordance with one embodiment of the present disclosure. In some embodiments, the hairpin duplex molecule is a double stranded molecule that includes a hairpin on end, where the hairpin bridges the first and second strands, forming a "closed" end. The opposite end of the double stranded nucleic acid molecule is "open." The "open" end is capable of being ligated to an adapter molecule (see, e.g., FIG. 12).
[0072] FIG. 8 provides a flowchart illustrating the methods of preparing a library for two- pass sequencing.
[0073] FIG. 9 depicts one method of introducing a sequencing primer to a hairpin duplex molecule, the method including introducing a 5' to 3' double stranded DNA exonuclease to the sample, which removes at least a portion of the region including a primer binding site from the 5'end of a second strand of the hairpin duplex molecule. The double stranded stem region is not degraded by the exonuclease given the inclusion of one or more blocking elements in the stem region (e.g., at a terminal end of the nucleotide sequence of the stem region).
[0074] FIG. 10 depicts another method of introducing a sequencing primer to a hairpin duplex molecule, the method including introducing a nicking enzyme to the sample, which removes at least a portion of a region including a primer binding site from the 5' end of a second strand of the hairpin duplex molecule. In these embodiments, the region including the primer binding site includes a recognition site for a nicking enzyme.
[0075] FIG. 11 depicts another method of introducing a sequencing primer to a hairpin duplex molecule, the method including introducing a primer having a least partial complementarity to a sequence with a Loop Sequence of a hairpin region of the hairpin duplex molecule. FIG. 10 further depicts that a polymerase may be used to open the hairpin duplex molecule, creating a free 3' end. Subsequently, a sequencing primer may be introduced.
[0076] FIG. 12 depicts another method of introducing a priming site to a hairpin duplex molecule, the method including ligating an adapter (e.g., a sequencing adapter, such as a sequencing adapter including a sequencing primer binding site) to the "open" end of the hairpin duplex molecule, namely the end opposite the hairpin. Subsequently, a sequencing primer may be annealed to a portion of the ligated adapter.
[0077] FIG. 13 provides a flowchart illustrating the methods of preparing a library for two- pass sequencing.
[0078] FIG. 14 depicts one embodiment of an amplicon-based method of producing a library of duplex template constructs for Sequencing by Expansion using the loop-tail primers of the present invention.
[0079] FIG. 15 A and 15B depict one embodiment of a method of designing and using primers for conversion of an Illumina library to a duplex sequencing library.
[0080] FIG. 16 depicts one embodiment of a genomic DNA-based method of producing a library of duplex template constructs for Sequencing by Expansion using the loop-tail primers of the present invention.DETAILED DESCRIPTION
[0081] It should also be understood that, unless clearly indicated to the contrary, in any methods claimed herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.
[0082] As used herein, the singular terms "a," "an," and "the" include plural referents unless context clearly indicates otherwise. Similarly, the word "or" is intended to include "and"unless the context clearly indicates otherwise. The term "includes" is defined inclusively, such that "includes A or B" means including A, B, or A and B.
[0083] As used herein in the specification and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as "only one of or "exactly one of," or, when used in the claims, "consisting of," will refer to the inclusion of exactly one element of a number or list of elements. In general, the term "or" as used herein shall only be interpreted as indicating exclusive alternatives (i.e., "one or the other but not both") when preceded by terms of exclusivity, such as "either," "one of," "only one of or "exactly one of." "Consisting essentially of," when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0084] The terms "comprising," "including," "having," and the like are used interchangeably and have the same meaning. Similarly, "comprises," "includes," "has," and the like are used interchangeably and have the same meaning. Specifically, each of the terms is defined consistent with the common United States patent law definition of "comprising" and is therefore interpreted to be an open term meaning "at least the following," and is also interpreted not to exclude additional features, limitations, aspects, etc. Thus, for example, "a device having components a, b, and c" means that the device includes at least components a, b, and c. Similarly, the phrase: "a method involving steps a, b, and c" means that the method includes at least steps a, b, and c. Moreover, while the steps and processes may be outlined herein in a particular order, the skilled artisan will recognize that the ordering steps and processes may vary.
[0085] As used herein in the specification and in the claims, the phrase "at least one," in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase "at least one" refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B," or, equivalently "at least one of A and / or B") can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet anotherembodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0086] As used herein, the term "about" refers to a range of values including the specified value, which a person of ordinary skill in the art would consider reasonably similar to the specified value. In some embodiments, the term "about" means within a standard deviation using measurements generally acceptable in the art. In some embodiments, about means a range extending to + / — 10% of the specified value.
[0087] As used herein, the term "adapter" refers to a nucleotide sequence that may be added to another sequence to import additional properties to that sequence. An adapter can be single- or double-stranded or may have both a single-stranded portion and a double-stranded portion. The ligation of an adapter to a target polynucleotide or a target polynucleotide strand of interest enables the generation of amplification-ready products of the target polynucleotide or the target polynucleotide strand of interest. The target polynucleotide molecules may be fragmented or not prior to the addition of adaptors.
[0088] As used herein "amplification" refers to a process in which a copy number increases. Amplification may be a process in which replication occurs repeatedly over time to form multiple copies of a template. Amplification can produce an exponential or linear increase in the number of copies as amplification proceeds. Exemplary amplification strategies include polymerase chain reaction (PCR), loop-mediated isothermal amplification (LAMP), rolling circle replication (RCA), cascade-RCA, nucleic acid-based amplification (NASB A), and the like. Also, amplification can utilize a linear or circular template. Amplification can be performed under any suitable temperature conditions, such as with thermal cycling or isothermally. Furthermore, amplification can be performed in an amplification mixture (or reagent mixture), which is any composition capable of amplifying a nucleic acid target, if any, in the mixture. PCR amplification relies on repeated cycles of heating and cooling (i.e., thermal cycling) to achieve successive rounds of replication. PCR can be performed by thermal cycling between two or more temperature setpoints, such as a higher denaturation temperature and a lower annealing / extension temperature, or among three or more temperature setpoints, such as a higher denaturation temperature, a lower annealing temperature, and an intermediate extension temperature, among others. PCR can be performed with a thermostable polymerase, such as Taq DNA polymerase. PCR produces an exponential increase in the amount of a product amplicon over successive cycles. PCR is described, for example, in U.S. Pat. No. 4,683,202; U.S. Pat. No. 4,683,195; U.S. Pat. No. 4,000,159; U.S. Pat. No. 4,965,188; U.S. Pat. No. 5,176,995), the disclosures of each are hereby incorporated by reference herein in their entirety.
[0089] As used herein, the term "biological sample," "tissue sample," "specimen" or the like refers to any sample including a biomolecule (such as a protein, a peptide, a nucleic acid, a lipid, a carbohydrate, or a combination thereof) that is obtained from any organism including viruses. Other examples of organisms include mammals (such as humans; veterinary animals like cats, dogs, horses, cattle, and swine; and laboratory animals like mice, rats, and primates), insects, annelids, arachnids, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (such as tissue sections and needle biopsies of tissue), cell samples (such as cytological smears such as Pap smears or blood smears or samples of cells obtained by microdissection), or cell fractions, fragments, or organelles (such as obtained by lysing cells and separating their components by centrifugation or otherwise). Other examples of biological samples include blood, serum, urine, semen, fecal matter, cerebrospinal fluid, interstitial fluid, mucous, tears, sweat, pus, biopsied tissue (for example, obtained by a surgical biopsy or a needle biopsy), nipple aspirates, cerumen, milk, vaginal fluid, saliva, swabs (such as buccal swabs), or any material containing biomolecules that is derived from a first biological sample. In certain embodiments, the term "biological sample" as used herein refers to a sample (such as a homogenized or liquefied sample) prepared from a tumor or a portion thereof obtained from a subject.
[0090] As used herein, the term "end" or "ends" refer to the regions of sequence at (or proximal to) either end of a nucleic acid sequence. As used herein, the term "3' region" refers to a region of a nucleotide strand that includes the 3' end of the strand. As used herein, the term "3' end" designates the end of a nucleotide strand that has the hydroxyl group of the third carbon in the sugar-ring of the deoxyribose at its terminus. As used herein, the term "5' region" refers to a region of a nucleotide strand that includes the 5' end of the strand. As used herein, the term "5' end" designates the end of a nucleotide strand that has the fifth carbon in the sugar-ring of the deoxyribose at its terminus.
[0091] As used herein, the term "exonuclease activity" is used in accordance with its ordinary meaning in the art and refers to the removal of a nucleotide from a nucleic acid by a DNA polymerase.
[0092] As used herein, the term "extension" refers to synthesis by a polymerase of a new polynucleotide strand complementary to a template strand by adding free nucleotides (e.g., dNTPs) from a reaction mixture that are complementary to the template in the 5'-to-3 ' direction. Extension includes condensing the 5 '-phosphate group of the dNTPs with the 3 '-hydroxy group at the end of the nascent (elongating) DNA strand.
[0093] As used herein, the term "hybridize" refers to the annealing of a nucleic acid sequence to another nucleic acid sequence (e.g., one single- stranded nucleic acid (such as a primer) to another nucleic acid) based on the well-understood principle of sequence complementarity. In an embodiment the other nucleic acid is a single-stranded nucleic acid. In some embodiments, one portion of a nucleic acid hybridizes to itself, such as in the formation of a hairpin structure. The propensity for hybridization between nucleic acids depends on the temperature and ionic strength of their milieu, the length of the nucleic acids and the degree of complementarity. The effect of these parameters on hybridization is described in, for example, Sambrook J., Fritsch E. F., Maniatis T., Molecular cloning: a laboratory manual, Cold Spring Harbor Laboratory Press, New York (1989). As used herein, hybridization of a primer, or of a DNA extension product, respectively, is extendable by creation of a phosphodiester bond with an available nucleotide or nucleotide analogue capable of forming a phosphodiester bond, therewith. For example, hybridization can be performed at a temperature ranging from 15° C. to 95° C. In some embodiments, the hybridization is performed at a temperature of about 20° C., about 25° C., about 30° C., about 35° C., about 40° C., about 45° C., about 50° C., about 55° C., about 60° C., about 65° C., about 70° C., about 75° C., about 80° C., about 85° C., about 90° C., or about 95° C. In other embodiments, the stringency of the hybridization can be further altered by the addition or removal of components of the buffered solution.
[0094] As used herein, the term "ligation" refers to a condensation reaction joining two nucleic acid strands wherein a 5'-phosphate group of one molecule reacts with the 3'-hydroxyl group of another molecule. Ligation is typically an enzymatic reaction catalyzed by a ligase or a topoisomerase. Ligation may join two single strands to create one single-stranded molecule. Ligation may also join two strands each belonging to a double-stranded molecule thus joining two double-stranded molecules. Ligation may also join both strands of a double-stranded molecule to both strands of another double-stranded molecule thus joining two double-stranded molecules. Ligation may also join two ends of a strand within a double-stranded molecule thus repairing a nick in the double-stranded molecule.
[0095] As used herein, the term "nanopore" refers to a pore, channel, or passage formed or otherwise provided in a membrane or other barrier material that has a characteristic width or diameter of about 0.1 nm to about 1000 nm. A nanopore can be made of a naturally occurring pore-forming protein, such as a-hemolysin from S. aureus, or a mutant or variant of a wild-type pore-forming protein, either non-naturally occurring (i.e., engineered) such as a-HL-C46, or naturally occurring. A membrane may be an organic membrane, such as a lipid bilayer, or a synthetic membrane made of a non-naturally occurring polymeric material. The nanopore may bedisposed adjacent or in proximity to a sensor, a sensing circuit, or an electrode coupled to a sensing circuit, such as, for example, a complementary metal-oxide semiconductor (CMOS) or field effect transistor (FET) circuit.
[0096] As used herein, the terms "nucleic acid" or "nucleic acid molecule" as used herein, refer to a high-molecular-weight biochemical macromolecule composed of nucleotide chains that convey genetic information. The most common nucleic acids are deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). The monomers from which nucleic acids are constructed are called nucleotides. Each nucleotide consists of three components: a nitrogenous heterocyclic base, either a purine or a pyrimidine (also known as a nucleobase); and a pentose sugar. Different nucleic acid types differ in the structure of the sugar in their nucleotides; DNA contains 2-deoxyribose while RNA contains ribose.
[0097] As used herein, the term "next generation sequencing" refers to sequencing technologies having high-throughput sequencing as compared to traditional Sanger- and capillary electrophoresis-based approaches, wherein the sequencing process is performed in parallel, for example producing thousands or millions of relatively small sequence reads at a time. Some examples of next generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. These technologies produce shorter reads (anywhere from about 25 - about 500 bp) but many hundreds of thousands or millions of reads in a relatively short time. Examples of such sequencing devices available from Illumina (San Diego, CA) include, but are not limited to iSEQ, MiniSEQ, MiSEQ, NextSEQ, NoveSEQ.
[0098] It is believed that the Illumina next-generation sequencing technology uses clonal amplification and sequencing by synthesis (SBS) chemistry to enable rapid sequencing. The process simultaneously identifies DNA bases while incorporating them into a nucleic acid chain. Each base emits a unique fluorescent signal as it is added to the growing strand, which is used to determine the order of the DNA sequence. A non-limiting example of a sequencing device available from ThermoFisher Scientific (Waltham, MA) includes the Ion Personal Genome Machine™ (PGM™) System.
[0099] It is believed that Ion Torrent sequencing measures the direct release of H+ (protons) from the incorporation of individual bases by DNA polymerase. A non-limiting example of a sequencing device available from Pacific Biosciences (Menlo Park, CA) includes the PacBio Sequel Systems. A non-limiting example of a sequencing device available from Roche (Pleasanton, CA) is the Roche 454. Next-generation sequencing methods may also include nanopore sequencing methods. In general, three nanopore sequencing approaches have been pursued: strand sequencing in which the bases of DNA are identified as they pass sequentially through a nanopore, exonuclease-based nanopore sequencing in which nucleotides areenzymatically cleaved one-by-one from a DNA molecule and monitored as they are captured by and pass through the nanopore, and a nanopore sequencing by synthesis (SBS) approach in which identifiable polymer tags are attached to nucleotides and registered in nanopores during enzyme- catalyzed DNA synthesis. Common to all these methods is the need for precise control of the reaction rates so that each base is determined in order.
[0100] Strand sequencing requires a method for slowing down the passage of the DNA through the nanopore and decoding a plurality of bases within the channel, ratcheting approaches, taking advantage of molecular motors, have been developed for this purpose. Exonuclease-based sequencing requires the release of each nucleotide close enough to the pore to guarantee its capture and its transit through the pore at a rate slow enough to obtain a valid ionic current signal. In addition, both methods rely on distinctions among the four natural bases, two relatively similar purines and two similar pyrimidines.
[0101] The nanopore SBS approach utilizes synthetic polymer tags attached to the nucleotides that are designed specifically to produce unique and readily distinguishable ionic current blockade signatures for sequence determination. In some embodiments, sequencing of nucleic acid molecules includes via nanopore sequencing includes preparing nanopore sequencing complexes and determining polynucleotide sequences. Methods of preparing nanopores and nanopore sequencing are described in U.S. Patent Application Publication No. 2017 / 0268052, and PCT Publication Nos. WO2014 / 074727, W02006 / 028508, WO2012 / 083249, and WO / 2014 / 074727, the disclosures of which are hereby incorporated by reference herein in their entireties. In some embodiments, tagged nucleotides may be used in the determination of the polynucleotide sequences (see, e.g., PCT Publication No. WO / 2020 / 131759, WO / 2013 / 191793, and WO / 2015 / 148402, the disclosures of which are hereby incorporated by reference herein in their entireties).
[0102] Analysis of the data generated by sequencing is performed using software and / or statistical algorithms that perform various data conversions, e.g., conversion of signal emissions into base calls, conversion of base calls into consensus sequences for a nucleic acid template, etc. Such software, statistical algorithms, and the use of such are described in detail, in U.S. Patent Application Publication Nos. 2009 / 0024331 2017 / 0044606 and in PCT Publication No. WO / 2018 / 034745, the disclosures of which are hereby incorporated by reference herein in their entireties.
[0103] As used herein, the term "nucleotide" refers to a nucleoside-5'-oligophosphate compound, or structural analog of a nucleoside-5'-oligophosphate, which can act as a substrate or inhibitor of a nucleic acid polymerase. Exemplary nucleotides include, but are not limited to, nucleoside-5'-triphosphates (e.g., dATP, dCTP, dGTP, dTTP, and dUTP); nucleosides (e.g., dA,dC, dG, dT, and dU) with 5'-oligophosphate chains of 4 or more phosphates in length (e.g., 5'- tetraphosphosphate, 5'-pentaphosphosphate, 5'-hexaphosphosphate, 5'-heptaphosphosphate, 5'- octaphosphosphate); and structural analogs of nucleoside-5'-triphosphates that can have a modified base moiety (e.g., a substituted purine or pyrimidine base), a modified sugar moiety (e.g., an O-alkylated sugar), and / or a modified oligophosphate moiety (e.g., an oligophosphate comprising a thio-phosphate, a methylene, and / or other bridges between phosphates).
[0104] As used herein, the "polymerase" as used herein, refers to an enzyme that catalyzes the process of replication of nucleic acids. More specifically, DNA polymerase catalyzes the polymerization of deoxyribonucleotides alongside a DNA strand, which the DNA polymerase "reads" and uses as a template. The newly polymerized molecule is complementary to the template strand and identical to the template's partner strand.
[0105] As used herein, the term "primer" refers to an oligonucleotide, either natural or synthetic, that is capable, upon forming a duplex with a polynucleotide template, of acting as a point of initiation of nucleic acid synthesis and being extended from its 3' end along the template so that an extended duplex is formed. Extension of a primer is usually carried out with a nucleic acid polymerase, such as a DNA or RNA polymerase. In some embodiments, the sequence of nucleotides added in the extension process is determined by the sequence of the template polynucleotide. Usually, primers are extended by a DNA polymerase. In some embodiments, primers have a length in the range of from 14 to 40 nucleotides, or in the range of from 18 to 36 nucleotides. Primers are employed in a variety of nucleic amplification reactions, for example, linear amplification reactions using a single primer, or polymerase chain reactions, employing two or more primers. Guidance for selecting the lengths and sequences of primers for particular applications is well known to those of ordinary skill in the art, as evidenced by the following reference that is incorporated by reference herein in its entirety: Dieffenbach, editor, PCR Primer: A Laboratory Manual, 2ndEdition (Cold Spring Harbor Press, New York, 2003).
[0106] As used herein, the phrase "primer binding site" refers to a region or site including a sequence that can be used for amplifying and / or sequencing a nucleic acid molecule (or a fragment thereof) ligated to an adapter.
[0107] As used herein, the term "sequence," when used in reference to a nucleic acid molecule, refers to the order of nucleotides (or bases) in the nucleic acid molecules. In embodiments where different species of nucleotides are present in the nucleic acid molecule, the sequence may include an identification of the species of nucleotide (or base) at respective positions in the nucleic acid molecule. A sequence is a property of all or part of a nucleic acid molecule.The term can be used similarly to describe the order and positional identity of monomeric units in other polymers such as amino acid monomeric units of protein polymers.
[0108] As used herein, the term "sequence complementarity" refers to a property shared between two nucleic acid sequences, such that when they are aligned antiparallel to each other, the nucleotide bases at each position will be complementary.
[0109] As used herein, the term "sequencing" refers to the determination of the order and position of bases in a nucleic acid molecule. More particularly, the term "sequencing" refers to biochemical methods for determining the order of the nucleotide bases, adenine, guanine, cytosine, and thymine, in a DNA oligonucleotide. Sequencing, as the term is used herein, can include without limitation parallel sequencing or any other sequencing method known of those skilled in the art, for example, chain-termination methods, rapid DNA sequencing methods, wandering-spot analysis, Maxam-Gilbert sequencing, dye- terminator sequencing, or using any other modern automated DNA sequencing instruments.
[0110] As used herein, the term "universal" refers to a nucleic acid molecule (e.g., primer or other oligonucleotide) that can be added to any nucleic acid molecule and perform its function irrespectively of the sequence of the nucleic acid molecule. The universal molecule may perform its function by hybridizing to the complement, e.g., a universal primer to a universal primer binding site in a universal primer.OVERVIEW
[0111] The present disclosure is directed to adapters and primers suitable for use in generating a library of nucleic acid molecules suitable for sequencing (e.g., with a next-generation sequencing technique), wherein the nucleic acid molecules suitable for sequencing include a hairpin, thereby facilitating two-pass sequencing. The present disclosure is also directed to methods of generating a library including one or more hairpin duplex molecules. As described herein, the methods of generating the library including the one or more hairpin duplex molecules comprises ligating the disclosed "full-length" or "truncated" adapters to one or more nucleic acid molecules derived from a sample. The methods of generating the library including the one or more hairpin duplex molecules also comprise amplifying one or more nucleic acid molecules derived from a sample with a loop tail primer. The present disclosure is also directed to methods of converting a library of nucleic acid molecules lacking hairpin duplex molecule into a library of nucleic acid molecules including one or more hairpin duplex molecules with a loop tail primer. The generated library including the one or more hairpin duplex molecules may then be sequenced.COMPOSITIONS
[0112] One aspect of the present disclosure is directed to adapters and primers including at least one hairpin region (referred to herein as "'full-length' adapters"). Another aspect of the present disclosure is directed to adapters including a pre-cursor to a hairpin region (referred to herein as "truncated adapters"). Yet another aspect of the present disclosure is directed to an adapter ligated nucleic acid molecule including one or more adapters (such as any of the "full- length" or "truncated" adapters described herein), where the one or more adapters include at least one hairpin region. Yet another aspect of the present disclosure is direct to amplified nucleic acid products including one or more primers described herein. These and other compositions are described further herein."Full-Length" Adapters
[0113] The present disclosure is directed to adapters capable of being ligated to double stranded nucleic acid molecules (such as insert fragments isolated from a sample, disclosed herein). In some embodiments, the "full-length" adapters include a region including a substantially double stranded nucleic acid (referred to herein as a "a substantially double stranded stem region"); and a region of single-stranded non-complementary nucleic acid strands (see, e.g., FIGS. 1 A and 2A). Said another way, the "full-length" adapters of the present disclosure include a substantially double stranded stem region, and a region including two non-complementary nucleic acid strands.
[0114] In some embodiments, the region of single-stranded non-complementary nucleic acid strands includes at least one hairpin region. In some embodiments, one of the single-stranded non-complementary nucleic acid strands includes a hairpin region (see, e.g., FIGS. 1 A - ID). In other embodiments, both of the single-stranded non-complementary nucleic acid strands include hairpin regions (see, e.g., FIGS. 2A - 2D).
[0115] The skilled artisan will appreciate that each strand of the substantially double stranded stem region is at least partially complementary to each other. Complementarity, however, need not be perfect. Stable duplexes, for example, may include mismatched base pairs or unmatched bases. As such, each strand of the of the substantially double stranded stem region may have 100% complementarity with each other; or have less than 100% complementarity, such as less than about 99% complementarity; such as less than about 98% complementarity; such as less than about 96% complementarity; such as less than about 94% complementarity; such as less than about 92% complementarity; such as less than about 90% complementarity; such as less than about 85% complementarity; such as less than about 80% complementarity; etc.
[0116] In some embodiments, the substantially double stranded stem region includes a first end and a second end, whereby the first end is adjacent to a hairpin region or a region including a primer binding site (described herein); and whereby the second end is capable of being ligated to a nucleic acid molecule or insert (see, e.g., FIG. 4A).
[0117] In some embodiments, each strand of the substantially double stranded stem region includes between about 8 and about 30 nucleotides, such as between about 10 and about 25 nucleotides, such as between about 10 and about 20 nucleotides, or such as between about 10 and about 15 nucleotides. In some embodiments, each strand of the substantially double stranded stem region includes about 10 nucleotides, such as about 11 nucleotides, such as about 12 nucleotides, such as about 13 nucleotides, such as about 14 nucleotides, such as about 15 nucleotides, such as about 16 nucleotides, etc.
[0118] In some embodiments, the Tmof each strand of the substantially double stranded stem region ranges from between about 30°C to about 60°C, such as between about 33°C to about 60°C, such as between about 35°C to about 60°C.
[0119] In some embodiments, the substantially double stranded stem region includes one or more exonuclease blocking elements, such as one or more elements which limit exonuclease degradation. In some embodiments, the one or more elements which limit exonuclease degradation are included at the first end of the substantially double stranded stem region.
[0120] In some embodiments, the one or more exonuclease blocking elements are one or more phosphorothioate bonds. In some embodiments, the substantially double stranded stem region includes between 1 and about 10 phosphorothioate bonds, such as between 1 and about 8 phosphorothioate bonds, such as between 1 and about 6 phosphorothioate bonds, such as between 1 and about 4 phosphorothioate bonds. In some embodiments, the phosphorothioate bonds are located at an end of the substantially double stranded stem region, such as an end adjacent to an insert (e.g., a fragment of a nucleic acid derived from a sample). For instance, the last 1, 2, 3, 4, 5, 6, etc. nucleotides within the stem region may include a phosphorothioate bond.
[0121] In other embodiments, the one or more exonuclease blocking elements are one or more 2'-modified nucleosides. In other embodiments, the two or more exonuclease blocking elements are one or more 2'-modified nucleosides.
[0122] In yet other embodiments, the one or more exonuclease blocking elements are one or more 2'-O-Methyl groups. In yet other embodiments, the two or more exonuclease blocking elements are one or more 2'-O-Methyl groups.
[0123] In even further embodiments, the one or more exonuclease blocking elements are one or more dideoxynucleotides. In even further embodiments, the two or more exonuclease blocking elements are one or more dideoxynucleotides.
[0124] In some embodiments, the "full-length" adapters include at least one "hairpin region." In some embodiments, the hairpin region is capable of forming a loop or loop-like structure when complementary portions of the sequence of the hairpin region anneal to each other (such as during a self-priming event). In some embodiments, a hairpin region includes two substantially complementary duplex sequences which are separated by a loop sequence. For instance, the hairpin region may have the structure:Duplex Sequence 1] - [Loop Sequence]n - [Duplex Sequence 2],
[0125] Where
[0126] n is O or l;
[0127] each of the Duplex Sequences are substantially complementary to each other (such as 100% complementary, such as less than 100% complementary, such as less than about 99% complementarity; such as less than about 98% complementarity; such as less than about 96% complementarity; such as less than about 94% complementarity; such as less than about 92% complementarity; such as less than about 90% complementarity; such as less than about 85% complementarity; such as less than about 80% complementarity; etc.).
[0128] In some embodiments, each of the Duplex Sequences may have the same size or different sizes. In some embodiments, each duplex sequences comprises between 6 (for a polymerase to recognize complementarity) and about 18, such as between about 6 and about 16 nucleotides, such as between about 6 and about 12 nucleotides. In some embodiments, the Tm(predicted, calculated, mean, average, or absolute) of each duplex sequence ranges from between about 23°C to about 50°C, such as from between about 25°C to about 50°C, such as from between about 25°C to about 45°C, such as from between about 25°C to about 40°C, etc. In some embodiments, at low temperatures (below temperatures commonly used for PCR, such as temperatures below 50°C), the duplex sequences at least partially hybridize to each other to form a hairpin. On the other hand, at higher temperatures (such as those used during PCR), each duplex sequence is unannealed, i.e., does not form a hairpin.
[0129] In some embodiments, the Loop Sequence, if present, may optionally include (i) a sample identifier, (ii) a unique molecular identifier; and / or (iii) a primer binding site. For instance, the loop sequence may include a primer binding site for a tailored PCR reaction where an adjacentterminal duplex sequence serves as a tail. In some embodiments, the Loop Sequence includes between about 5 and about 200 nucleotides, such as between about 5 and 100 nucleotides, such as between about 5 and 50 nucleotides, or such as between about 5 and 25 nucleotides.
[0130] In some embodiments, the adapters of the present disclosure include a single hairpin region (FIG. 1 A), wherein the single hairpin region forms or is capable of forming a hairpin structure with itself. In these embodiments, the hairpin region includes a loop sequence, a first duplex sequence, and a second duplex sequence, where the first and second duplex sequences are substantially complementary to each other and permit the formation of a hairpin or loop-like structure in the hairpin region (see FIG. IB). In some embodiments, the single hairpin region is capable of self-annealing and forming a loop-like structure at temperatures below about 50°C, such as at a temperature of below about 45°C, such as at a temperature of below about 40°C, below about 35°C, below about 30°C, etc. In some embodiments, the single hairpin region is capable of self-annealing and forming a loop-like structure at temperatures above about 50°C, such as at a temperature of above about 55°C, such as at a temperature of at about 58C, etc.
[0131] In those embodiments where only one of the single-stranded non-complementary nucleic acid strands includes a hairpin region, another of the single-stranded non-complementary nucleic acid strands includes a primer binding site or a region including a primer binding site (see FIGS. 1 A - ID). In some embodiments, the region including the primer binding site is designed such that it does not hybridize to any portion of the hairpin region of the adapter. In some embodiments, the region including the primer binding site comprises between 5 and about 30 nucleotides, such as between about 5 and about 25 nucleotides, such as between about 5 and about 20 nucleotides. In some embodiments, the primer binding site is a universal priming site. In some embodiments, the primer binding site is capable of hybridizing to a sequencing primer.
[0132] In some embodiments, the region including the primer binding site further includes a recognition site for a nicking enzyme, such as a nicking endonuclease (also referred to as a "nicking site"). As used herein, the term "nicking endonuclease" refers to an endonuclease having nicking activity that can recognize a specific nucleotide sequence and cleave only one strand of a double-stranded nucleic acid having the abovementioned nucleotide sequence. In some embodiments, the nicking endonuclease can cleave the phosphodiester bond of one strand of a double-stranded DNA molecule.
[0133] In some embodiments, the number of bases in the nicking endonuclease recognition site is at least 3, such as at least 4, such as at least 5, such as at least 6, or such as at least 7. Suitable endonucleases include, but are not limited to, Nb.BbvCI(the number of bases in the recognition site is 7: 5'-GC / TGAGG-3'), Nb.BsmI (the number of bases in the recognition site is 6: 5'-NG / CATTC-3'), Nb.BtsI (the number of bases in the recognition site is 6: 5'-NN / CACTGC-3'), Nb.BsrDI (the number of bases in the recognition site is 6: 5'- NN / CATTGC-3'), Nt.BspQI (the number of bases in the recognition site is 7: 5'-GCTCTTCN / -3'), Nt.BbvCI (the number of bases in the recognition site is 7: 5'-CC / TCAGC-3'), Nt.AIwI (the number of bases in the recognition site is 5: 5'-GGATCNNNN / N-3'), Nt.BsmAI (the number of bases in the recognition site is 5: 5'-GTCTCN / N-3'), Nt.BstNBI (the number of bases in the recognition site is 5: 5'-GAGTCNNNN / N-3'), Nt.CviPII (the number of bases in the recognition site is 3: 5'- / CCD-3'), Nb.Mval269I (the number of bases in the recognition site is 6: 5'-G / CATTC-3'), Nt.BpulOI (the number of bases in the recognition site is 7: 5'-CC / TNAGC- 3') and Nb.BssSI (the number of bases in the recognition site is 6: 5'-C / TCGTG-3') (the number of bases in the recognition site and its sequence is shown in each parenthesis above). Here,depicts a cleavage site; "N" is an A, T, G or C nucleotide; and "D" is an A, T or G nucleotide.
[0134] Additional nicking endonuclease recognition sites and methods of effectuating nicking with an endonuclease are disclosed by Joneja et. al., "Linear nicking endonuclease- mediated strand-displacement DNA amplification." Anal. Biochem. 2011 Jul 1 ;414( 1 ): 58-69, the disclosure of which is hereby incorporated by reference herein in its entirety. Yet additional nicking enzyme recognition sequences are identified in U.S. PatentNo. 10,570,441, the disclosure of which is hereby incorporated by reference herein in its entirety.
[0135] While FIGS. 1 A - ID depict a hairpin region at a 5' end of one strand of an adapter; and a region including a primer binding site at a 3' end of another strand of the adapter, the hairpin region may be at a 3' end while the region including the primer binding site may be at the 5' end (see also FIG. 4E).
[0136] In other embodiments, the adapters of the present disclosure include two hairpin regions (FIG. 2A), i.e., both of the single-stranded non-complementary nucleic acid strands include hairpin regions. In those embodiments which include two hairpin regions, the hairpin regions are designed such that they do not hybridize to each other, i.e., a hairpin region in a top strand of the adapter is designed such that it does not hybridize with a hairpin region of a bottom strand of the adapter. Instead, each hairpin region is designed such that it forms or is capable of a forming separate hairpin structure with itself. In some embodiments, each hairpin region is capable of self-annealing and forming a loop-like structure at temperatures below 50°C. In some embodiments, each hairpin region is capable of self-annealing and forming a loop-like structure at temperatures below 45°C. In some embodiments, each hairpin region is capable of self-annealing and forming a loop-like structure at temperatures below 40°C. In some embodiments, each hairpinregion is capable of self-annealing and forming a loop-like structure at temperatures below 35°C. In some embodiments, each hairpin region is capable of self-annealing and forming a loop-like structure at temperatures above about 50°C.
[0137] By way of example, a hairpin region of a top strand of the adapter of FIG. 2A may form a first hairpin or loop-like structure, while a hairpin region of a bottom strand of the adapter of FIG. 2A may form a second hairpin or loop-like structure. In the embodiment illustrated in FIG. 2 A, the sequences of each of hairpin region 1 and hairpin region 2 are different, i.e., comprising different duplex sequences and / or loop sequences. In these embodiments, the hairpin region of the top strand includes a first loop sequence, a first duplex sequence, and a second duplex sequence, where the first and second duplex sequences are substantially complementary to each other and form or are capable of forming a first hairpin or first loop-like structure (see FIG. 2B). Likewise, the hairpin region of the bottom strand includes a second loop sequence, a third duplex sequence, and a fourth duplex sequence, where the third and fourth duplex sequences are substantially complementary to each other and form or are capable of forming a second hairpin or second loop-like structure (see FIG. 2B).
[0138] In some embodiments, the full-length adapters include one or more barcodes, which may be included in either the hairpin region, the stem region, or in both the hairpin region and the stem region. In some embodiments, each of the top and bottom strands in the stem region include one or more barcodes. In other embodiments, the loop sequence of the hairpin region includes one or more barcodes. In yet other embodiments, the top and bottom strands of the stem region and the loop sequence each include one or more barcodes.
[0139] As used herein, the term "barcode" refers to a nucleic acid sequence that can be detected and identified. In some embodiments, the barcodes include between about 5 and about 20 nucleotides, such that in a sample, the nucleic acids incorporating the barcodes can be distinguished or grouped according to the barcodes. In some embodiments, the barcodes include between about 5 and about 15 nucleotides. In some embodiments, the barcodes include between about 5 and about 10 nucleotides. In some embodiments, the barcodes include between about 10 and about 15 nucleotides. In some embodiments, the barcodes include about 8, about 9, about 10, about 11, about 12, about 13, about 14, or about 15 nucleotides. Non-limiting examples of barcodes and / or unique molecular identifiers (UMIs) are described in U.S. Publication No. 2020 / 0032244, and in U.S. Patent Nos. 7,393,665, 8,168,385, 8,481,292, 8,685,678, and 8,722,368, and in PCT Publication No. WO / 2018 / 138237, the disclosures of which are hereby incorporated by reference herein in their entireties.
[0140] In some embodiments, UMIs may be incorporated as part of an overall DNA amplification and sequencing workflow to perform error correction. In some embodiments, errors are introduced (1) by the polymerase during amplification, and (2) during sequencing (i.e., reading) of the amplified molecules. In some embodiments, UMIs ligated to nucleic acid molecules reduce the impact of one or both sources of error. For instance, UMIs incorporate a unique barcode onto each molecule within a given sample library. By incorporating individual barcodes on each original DNA fragment, variant alleles present in the original sample (true variants) can be distinguished from errors introduced during library preparation, target enrichment, or sequencing.
[0141] In some embodiments, the UMIs have the one of the general Formulas set forth below:(W)(N)(N)(N)(N)(N)(W)(N)(N)(N)(N)(N)(W), or (N)(W)(N)(N)(N)(W)(N)(N)(N),
[0142] where N includes (in the aggregate) about 25% adenosine; about 25% guanine; about 25% cytosine; and about 25% thymine; and W includes (in the aggregate) about 50% adenosine and about 50% thymine.Truncated Adapters
[0143] In some embodiments, the present disclosure provides for truncated adapters which are pre-cursors to the full-length adapters of the present disclosure. Like the "full-length" adapters described above, the "truncated" adapters include a substantially double stranded stem region, and a region including two non-complementary nucleic acid strands. In some embodiments, one of the single-stranded non-complementary nucleic acid strands includes a hairpin precursor region; while another of the single-stranded non-complementary nucleic acid strands includes a region including a primer binding site. In some embodiments, the hairpin precursor region includes a single duplex sequence. Examples of such truncated adapters are depicted in FIGS. 3A and 3C.
[0144] As illustrated in FIGS. 3B and 3D, the hairpin precursor region includes a first duplex sequence. This single duplex sequence is designed such it may not self-anneal. As such, the hairpin precursor region of a "truncated" adapter is designed such that in its initial state it cannot form a loop-like structure. The skilled artisan will appreciate, however, that in some embodiments, a reverse compliment of at least the single duplex sequence of the truncated adapter (hairpin precursor region) may be used to generate a hairpin region (such as a hairpin region present in a "full-length" adapter) following polymerase extension with a suitable primer and polymerase (see FIG. 6B).
[0145] In some embodiments, the truncated adapter includes one or more barcodes. By way of example, the truncated adapter may include one or more barcodes in a stem region.Loop-Tail Primers
[0146] In one aspect, the present disclosure is directed to primers capable of hybridizing to single stranded nucleic acid molecules that include a known sequence (e.g., nucleic acid strands derived from a nucleic acid library construct or a PCR amplicon product). In some embodiments, the primers are referred to as “loop-tail” primers and include a 3’ portion of substantially single stranded nucleic acid (referred to herein as the primer region) and a 5’ a hairpin portion (i.e., the loop tail region). As disclosed herein, the hairpin portion is capable of forming a loop or loop-like structure when complementary portions of the sequence of the hairpin portion anneal to each other to form a duplex region. Non-limiting examples of hairpin structures are discussed further herein, e.g., with reference to full-length adapters.
[0147] As mentioned, in certain embodiments, the sequence of the 3’ primer portion of the loop tail primer is designed to hybridize to a known sequence in a nucleic acid library construct or a PCR amplicon product. As used herein, the term “construct” refers to an engineered nucleic acid molecule created by joining a nucleic acid fragment of interest (e.g., a genomic DNA insert, or library fragment) to one or more heterologous adapters. The final construct may be used for sequencing or other applications. Commonly, a nucleic acid construct is amplified to increase the copy number, using primers that hybridize to known sequences in the adapter regions of the construct. The products of an amplification reaction are fully double stranded PCR amplicons that include the known sequences of the amplification primers at each end.
[0148] In certain embodiments, the sequence of the 3 ’ primer portion of the loop tail primer may be complementary to a sequence in a region of a commercially available adapter. For example, many library prep kits used for next generation, or other sequencing platforms are commercially available and well known in the art, for example, those offered by Illumina, Agilent Technologies, Thermo Fisher Scientific, Roche Sequencing Solutions, and the like. In one example, an Illumina library prep kit provides Y adapters that provide a known “P5” sequence in one single stranded arm region and a known “P7” sequence in the other single stranded arm region. As such, in one non-limiting embodiment, the sequence of the 3’ primer portion of the loop tail adapter may be complementary to one or more of the P5 or P7 sequences present in an Illumina sequencing adapter.
[0149] In some embodiments, the primer portion of the loop tail primer may be from around 10 to around 25 nucleotides in length. In certain embodiments, the primer portion may befrom around 12 to 20 nucleotides in length. In one embodiment, the primer portion may be around 15 nucleotides in length. In some embodiments, the first and second duplex regions of the loop tail portion of the primer may be from around 5 nucleotides to around 15 nucleotides in length. In certain embodiments, the first and second duplex regions may be from around 6 nucleotides to around 12 nucleotides in length. In one embodiment, the first and second duplex regions may be around 9 nucleotides in length. In some embodiments, the single stranded loop region may be from around 4 to around 15 nucleotides in length. In certain embodiments, the single stranded loop may be from around 6 to around 10 nucleotides in length. In one embodiment, the single stranded loop may be around 7 nucleotides in length.
[0150] In certain embodiments, the loop sequence of the loop-tail primer may include a SID and / or UMI sequence. The loop sequence may also include a suitable recognition sequence to “mark” the loop. In certain embodiments, the loop sequence may be used as a primer binding site for a tailed PCR reaction where the duplex stem provides the tail.
[0151] In some embodiments, the loop-tail primers of the present invention may include additional modifications or features to facilitate downstream steps of a library prep workflow. For example, in certain embodiments, a primer may include a 5’ modification, such as a biotin moiety to facilitate subsequent purification steps based on streptavidin-biotin interaction chemistry, e.g., removal excess primer or other undesired side-products from a ligation reaction. In other embodiments, a primer may optionally include a 5’ phosphate moiety to enhance removal single stranded nucleic acid molecules using exonuclease treatment.Adapter Ligated Nucleic Acid Molecules
[0152] The present disclosure also provides for adapter ligated nucleic acid molecules, where one or more of the adapters of the present disclosure are ligated to a double stranded nucleic acid molecule (an "insert"), such as a double stranded nucleic acid molecule derived from a sample (also referred to herein as an "insert"). For instance, FIGS. 4A and 4B illustrate adapters ligated to a double stranded insert, where each adapter includes a single hairpin region and a substantially double stranded stem region. In these embodiments, each strand of the adapter ligated nucleic acid molecule includes a 5' hairpin region and a 3' region including a primer binding site. As illustrated in FIG. 4B, each hairpin region of each adapter includes a loop sequence (which is optional) and two duplex sequences, whereby the two duplex sequences are substantially complementary to each other such that they may form a loop-like structure at an end of one of the strands of the double stranded nucleic acid molecule, i.e., the adapters of the present disclosure do not form a hairpin between the two different strands of the double stranded molecules to which they are ligated (see,e.g., FIGS. 4C - 4E). Said another way, the adapters of the present disclosure, when ligated to a nucleic acid molecule or a fragment thereof, do not form a closed-ended double stranded nucleic acid molecule.
[0153] By way of another example, FIGS. 5 A and 5B illustrate adapters ligated to a double stranded insert (e.g., a fragmented nucleic acid molecule), where each adapter includes two hairpin regions and a substantially double stranded stem region. As illustrated in FIG. 5B, each hairpin region of each adapter includes a loop sequence (which is optional) and two duplex sequences, whereby the two duplex sequences in each hairpin region are substantially complementary to each other such that they may each form a loop-like structure at each end of each strand of the double stranded nucleic acid molecule. Notably, the adapters including the two hairpin regions do not form a hairpin between the two different strands of the double stranded molecule, i.e., a close- ended, double stranded structure is not formed.METHODS OF PREPARING TWO PASS SEQUENCING TEMPLATES
[0154] Aspects of the present disclosure include method of preparing a sequencing library including one or more two pass sequencing templates, where the method includes obtaining a sample including one or more double stranded nucleic acid molecules (FIG. 8, step 101), ligating the "full-length" or "truncated" adapters disclosed herein to the one or more nucleic acid molecules within the obtained sample to provide one or more adapter ligated nucleic acid molecules (step 102), optionally amplifying the one or more adapter ligated nucleic acid molecules (step 103), optionally enriching the sample for one or more adapter ligated target nucleic acid molecules (step 104), converting a single strand of the adapter ligated target nucleic acid molecule in the sample into a hairpin duplex molecule (step 105), annealing a sequencing primer to the hairpin duplex molecule to provide a primed hairpin duplex molecule (also referred to as a “two pass sequencing template”) (step 106), and sequencing the primed hairpin duplex molecule (step 107).Sample Preparation
[0155] In some embodiments, a sample comprising one or more nucleic acid molecules is obtained (step 101) and prepared for downstream processing. In some embodiments, the sample comprises one or more double stranded nucleic acid molecules (also referred to herein as "inserts").
[0156] In some embodiments, samples may be obtained from any source including a target nucleic acid molecule having one or more modified nucleotides, e.g., tissue (including tumor tissue or formalin-fixed paraffin-embedded (FFPE) tissue), blood, skin, swab (e.g., buccal, vaginal), urine, saliva, etc. In some embodiments, the sample is derived from a subject or a patient, such as a subject or a patient diagnosed with a disease or suspected of having a disease. In someembodiments, the sample may include a fragment of a solid tissue, or a tumor sample derived from the subject or the patient, e.g., by biopsy. As used herein, the term "tumor sample" encompasses samples prepared from a tumor or from a sample potentially including or suspected of comprising cancer cells, or to be tested for the potential presence of cancer cells, such as a lymph node. As used herein, the term "tumor" refers to a mass or a neoplasm, which itself is defined as an abnormal new growth of cells that usually grow more rapidly than normal cells and will continue to grow if not treated sometimes resulting in damage to adjacent structures. Tumor sizes can vary widely. A tumor may be solid, or fluid filled. A tumor can refer to benign (not malignant, generally harmless), or malignant (capable of metastasis) growths. Some tumors can include neoplastic cells that are benign (such as carcinoma in situ) and, simultaneously, contain malignant cancer cells (such as adenocarcinoma). This should be understood to include neoplasms found in multiple locations throughout the body. Therefore, for purposes of the present disclosure, tumors include primary tumors, lymph nodes, lymphatic tissue, and metastatic tumors.
[0157] Methods for isolating nucleic acid molecules from obtained samples and / or purifying the obtained samples are known (see, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 2d ed., Cold Spring Harbor Laboratory Press, 1989; Sambrook et al., Molecular Cloning: A Laboratory Manual, 3d ed., Cold Spring Harbor Press, 2001) and several kits are commercially available (e.g., High Pure RNA Isolation Kit, High Pure Viral Nucleic Acid Kit, and MagNA Pure LC Total Nucleic Acid Isolation Kit, DNA Isolation Kit for Cells and Tissues, DNA Isolation Kit for Mammalian Blood, High Pure FFPET DNA Isolation Kit, available from Roche). In the context of the presently disclosed methods, nucleic acid molecules, including genomic DNA, can be collected, purified, and / or isolated.
[0158] It will be appreciated that nucleic acid molecules may be isolated from obtained samples using any of a variety of procedures known in the art, for example, MagMAX™ DNA Multi-Sample Ultra Kit (Applied Biosystems, Thermo Fisher Scientific), the MagMAX™ Express-96 Magnetic Particle Processor and the KingFisher™ Flex Magnetic Particle Processor (Thermo Fisher Scientific), a RecoverAll™ Total Nucleic Acid Isolation Kit for FFPE and PureLink™ FFPE RNA Isolation Kit (Ambion™, Thermo Fisher Scientific), the ABI Prism™ 6100 Nucleic Acid PrepStation and the ABI Prism™ 6700 Automated Nucleic Acid Workstation (Applied Biosystems, Thermo Fisher Scientific), and the like.
[0159] In some embodiments, the nucleic acid molecules within the obtained sample are selected from DNA molecules, genomic DNA molecules, cfDNA molecules, cDNA molecules, RNA molecules, mRNA molecules, rRNA molecules, mtDNA, siRNA molecules, or anycombination thereof. In some embodiments, the plurality of nucleic acid molecules comprises single stranded polynucleotides.
[0160] In some embodiments, the nucleic acid molecules within the obtained sample may be prepared for downstream processing by fragmenting, cutting, or shearing the nucleic acids. In some embodiments, the fragmenting, cutting, and / or shearing may be accomplished using such procedures as mechanical force, sonication, restriction endonuclease cleavage, or any method known in the art. In other embodiments, no fragmentation is necessary (some genomic samples may already consist of appropriately sized fragments and will not require additional fragmentation).
[0161] In some embodiments, following fragmentation the nucleic acid molecules within any obtained sample have a size ranging from between about 10 mer to about 1000 mer. In some embodiments, following fragmentation the nucleic acid molecules within any obtained sample have a size ranging from between about 10 mer to about 550 mer. In other embodiments, following fragmentation the nucleic acid molecules within any obtained sample have a size ranging from between about 15 mer to about 500 mer. In yet other embodiments, following fragmentation the nucleic acid molecules within any obtained sample have a size ranging from between about 15 mer to about 450 mer. In further embodiments, following fragmentation the nucleic acid molecules within any obtained sample have a size ranging from between about 15 mer to about 400 mer. In even further embodiments, following fragmentation the nucleic acid molecules within any obtained sample have a size ranging from between about 15 mer to about 350 mer. In yet even further embodiments, following fragmentation the nucleic acid molecules within any obtained sample have a size ranging from between about 15 mer to about 300 mer. In some embodiments, the nucleic acid molecules within the sample are fragmented to a platform-specific size range.
[0162] Following the fragmentation of the sample, in some embodiments, the fragmented nucleic acid molecules are end repaired and then a "tailing" reaction is performed. Tailing is an enzymatic method for adding a non-templated nucleotide to the 3' end of a blunt, double-stranded DNA molecule. In some embodiments, a Taq polymerase is utilized for A-tailing.Ligation Of Adapters
[0163] After end-repair tailing, one or more of the "full-length" or "truncated" adapters of the present disclosure are ligated to one or more double stranded nucleic acid molecules within the obtained sample to provide adapter ligated nucleic acid molecules (step 102). Examples of adapter ligated nucleic acid molecules generated after ligation are illustrated in FIGS. 4A - 4C, 5A - 5C, and 6.
[0164] In some embodiments, the adapters may be ligated to the double stranded nucleic acid molecules using any method known in the art. For instance, suitable methods of ligating adapters to nucleic acid molecules are described in U.S. Patent Publication Nos. 2017 / 0037459, 2018 / 0334709, 2018 / 0016630 and in PCT Publication No. WO2017021449, the disclosures of which are hereby incorporated by reference herein in their entireties. In some embodiments, the adapters are ligated to the double stranded nucleic acid molecules with any suitable ligase, such as T4 DNA ligase. In some embodiments, the ligase is a modified T4 ligase, a mutant of a T4 ligase, or a variant of a T4 ligase (collectively referred to herein as a "modified T4 ligase"), such as those disclosed in International Publication No. WO / 2024 / 123733, in U.S. Patent No. 10,837,009, or in U.S. Publication No. 2018 / 0320162, the disclosures of which are hereby incorporated by reference herein in their entireties.
[0165] In some embodiments, the adapters ligated to the one or more nucleic acid molecules (also referred to as "insert" or "fragment") include any of the "full-length" adapters described herein. In some embodiments, the adapters ligated to the one or more nucleic acid molecules each include one or more hairpin regions. In other embodiments, the adapters ligated to the one or more nucleic acid molecules in the sample include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non-complementary; and wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence.
[0166] In yet other embodiments, adapters ligated to the fragmented nucleic acid molecules in the sample include a first strand and a second strand, wherein a 5' portion of the first strand and a 3' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 3' portion of the first strand and a 5' portion of the second strand are single stranded and non-complementary; and wherein the 3' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence.
[0167] In further embodiments, the adapters ligated to the fragmented nucleic acid molecules in the sample include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the secondstrand are single stranded and non-complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; and wherein the 3' portion of the second strand comprises a second hairpin region including two second duplex sequences that are substantially complementary to each other, wherein the two second duplex sequences are separated by an optional second loop sequence.
[0168] In some embodiments, the adapters ligated to the one or more nucleic acid molecules include any of the "truncated" adapters described herein. In some embodiments, the adapters ligated to the nucleic acid molecules in the sample include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non-complementary; wherein a 5' end of the first strand comprises a hairpin precursor region; and wherein a 3' portion of the second strand comprises a region including a primer binding site.Optional Amplification
[0169] Following the ligation of the adapters ("full-length" or "truncated" adapters) to the double stranded nucleic acid molecule fragments in the sample to provide the one or more adapter ligated nucleic acid molecules, the sample is optionally amplified (step 103). In some embodiments, the one or more adapter ligated nucleic acid molecules are optionally amplified to generate one or more amplified double stranded nucleic acid molecules which may be utilized in one or more downstream processes, e.g., in a target enrichment process and / or converted to a hairpin duplex molecule. In some embodiments, the amplification comprises performing a template directed oligonucleotide primer extension using one or more polymerases. In other embodiments, the amplification comprises an isothermal amplification technique.
[0170] In some embodiments, the optional amplification of the one or more adapter ligated nucleic acid molecules provides one or more amplified double stranded nucleic acid molecules having a first strand including a 5' end including a first hairpin region and a second strand having a 3' end including a primer binding site. This result is regardless of whether the "full-length" adapters or "truncated" adapters described herein are ligated to the double stranded nucleic acid molecule fragments in the sample (compare FIGS. 6A and 6B).
[0171] The skilled artisan will be able to select an appropriate amplification method and appropriate primers for use during amplification based on the structure of the adapter ligated nucleic acid molecule in the sample. For instance, amplification of a "full-length" adapter ligatednucleic acid molecule may utilize: (i) a first primer corresponding to at least a portion of a duplex sequence and / or at least a portion of an optional loop sequence of a hairpin region of a "full-length" adapter; and (ii) a second primer corresponding to at least a portion of a region including a primer binding site of a "full-length" adapter (see, e.g., FIG. 6A).
[0172] Likewise, amplification of a "truncated" adapter ligated nucleic acid molecule may utilize a first primer which is at least partially complementary to one or more portions of the hairpin precursor region of the truncated adapter (e.g., to a duplex sequence and / or a stem region) and a second primer which is at least partially complementary to a region of the truncated adapter including the primer binding site and / or a stem region. In this manner, the primers selected for amplification of the "truncated" adapter ligated nucleic acid molecules will introduce into the one or more amplified double stranded nucleic acid molecules those sequence elements "missing" from the "truncated" adapter ligated nucleic acid molecules but present in the corresponding "full- length" adapter ligated nucleic acid molecules. As such, the "truncated" adapter ligated nucleic acid molecules in a sample may be "converted" to the equivalent of the "full-length" adapter ligated nucleic acid molecules (see, e.g., FIG. 6B).Template Directed Oligonucleotide Primer Extension
[0173] In some embodiments, the optional amplification of the one or more adapter ligated nucleic acid molecules is based on template directed oligonucleotide primer extension using one or more polymerases. For instance, the sample including the one or more adapter ligated nucleic acid molecules is contacted with a polymerase and / or other amplification reagents to provide one or more amplified nucleic acid molecules.
[0174] Non-limiting examples of polymerases include prokaryotic DNA polymerases (e.g., Pol I, Pol II, Pol III, Pol IV, and Pol V), eukaryotic DNA polymerase, archaeal DNA polymerase, etc. In some embodiments, suitable polymerases may be derived from: archaea (e.g., Thermococcus litoralis (Vent, GenBank: AAA72101), Pyrococcus furiosus (Pfu, GenBank: D12983, BAA02362), Pyrococcus woesii, Pyrococcus GB-D (Deep Vent, GenBank: AAA67131), Thermococcus kodakaraensis KODI (KOD, GenBank: BD175553, BAA06142; Thermococcus sp. strain KOD (Pfx, GenBank: AAE68738)), Thermococcus gorgonarius (Tgo, Pdb: 4699806), Sulfolobus solataricus (GenBank: NC002754, P26811), Aeropyrum pemix (GenBank: BAA81109), Archaeglobus fulgidus (GenBank: 029753), Pyrobaculum aerophilum (GenBank: AAL63952), Pyrodictium occultum (GenBank: BAA07579, BAA07580), Thermococcus 9 degree Nm (GenBank: AAA88769, Q56366), Thermococcus fumicolans (GenBank: CAA93738, P74918), Thermococcus hydrothermalis (GenBank: CAC18555), Thermococcus sp. GE8(GenBank: CAC12850), Thermococcus sp. JDF-3 (GenBank: AX135456; WO0132887), Thermococcus sp. TY (GenBank: CAA73475), Pyrococcus abyssi (GenBank: P77916), Pyrococcus glycovorans (GenBank: CAC12849), Pyrococcus horikoshii (GenBank: NP 143776), Pyrococcus sp. GE23 (GenBank: CAA90887), Pyrococcus sp. ST700 (GenBank: CAC 12847), Thermococcus pacificus (GenBank: AX411312.1), Thermococcus zilligii (GenBank: DQ3366890), Thermococcus aggregans, Thermococcus barossii, Thermococcus celer (GenBank: DD259850.1), Thermococcus profundus (GenBank: E14137), Thermococcus siculi (GenBank: DD259857.1), Thermococcus thioreducens, Thermococcus onnurineus NA1, Sulfolobus acidocaldarium, Sulfolobus tokodaii, Pyrobaculum calidifontis, Pyrobaculum islandicum (GenBank: AAF27815), Methanococcusjannaschii (GenBank: Q58295), Desulforococcus species TOK, Desulforococcus, Pyrolobus, Pyrodictium, Staphylothermus, Vulcanisaetta, Methanococcus (GenBank: P52025) and other archaeal B polymerases, such as GenBank AAC62712, P956901, BAAA07579)), thermophilic bacteria Thermus species (e.g., flavus, ruber, thermophilus, lacteus, rubens, aquaticus), Bacillus stearothermophilus, Thermotoga maritima, Methanothermus fervidus, KOD polymerase, TNA1 polymerase, Thermococcus sp. 9 degrees N-7, T4, T7, phi29, Pyrococcus furiosus, P. abyssi, T. gorgonarius, T. litoralis, T. zilligii, T. sp. GT, P. sp. GB-D, KOD, Pfu, T. gorgonarius, T. zilligii, T. litoralis and Thermococcus sp. 9N-7 polymerases.
[0175] To effectuate amplification, the one or more adapter ligated nucleic acid molecules including the one or more methylated nucleotides are heat denatured. Melting temperatures (Tm) for heat denaturation are dependent upon several variables including the GC content of the nucleic acid molecule and / or the size of the nucleic acid molecule, but in general may be about 95°C or higher, such as for about 15 seconds to about 2 minutes.
[0176] Following heat denaturation, oligonucleotide primers are annealed to the template sequence of the one or more adapter ligated nucleic acid molecules (typically between about 40°C and about 60°C, such as for about 30 to about 60 seconds). In some embodiments, one oligonucleotide primer anneals to portion of the hairpin region of the adapter, e.g., to a portion of a duplex sequence and / or a portion of a loop sequence. In some embodiments, another primer anneals to at least a portion of a primer binding site incorporated within the adapter.
[0177] In some embodiments, the annealing temperature, like the heat denaturation temperature, is dependent upon the GC content and / or length of the primers. In some embodiments, the oligonucleotides form stable associations ('anneal') with the single stranded DNA (hereinafter referred to as the template strand) and thus serve as primers for nucleic acid synthesis by a polymerase. Subsequently, a corresponding nucleic acid strand to the template is synthesized from the primer oligonucleotide through use of the polymerase and deoxynucleotidetriphosphates (dNTPs) (also referred to as "primer extension"). In some embodiments, the temperature is raised for the polymerase, which in the case of commonly used thermostable polymerases is about 74° C, primer extension then lasts approximately 1 to 2 minutes. Reactions take place in a PCR master mixture which includes the nucleic acid molecule, a polymerase, oligonucleotide primers, deoxynucleotide triphosphates (dNTPs), reaction buffer, magnesium and / or optional additives.Isothermal Amplification Techniques
[0178] In some embodiments, the one or more adapter ligated double stranded nucleic acid molecules are isothermally amplified. Instead of melting DNA strands apart at high temperatures, isothermal amplification takes advantage of DNA polymerases with high strand displacement activity, (e.g., Bst or phi29 DNA polymerases). As used herein, the term "strand displacement" refers to the ability of the enzyme to separate the DNA strands in a double-stranded DNA molecule during primer-initiated synthesis. The enzyme can be a complete enzyme or a biologically active fragment thereof. The enzyme can be isolated and purified or recombinant. In some embodiments, the enzyme is thermostable. Such an enzyme is stable at elevated temperatures (e.g., greater than 40°C) and heat resistant to the extent that it effectively polymerizes DNA at the temperature employed. In some embodiments, the strand displacing polymerase used during isothermal amplification is a +29-DNA polymerase derived from bacteriophage (Blanco et al., U.S. Pat. Nos. 5,198,543 and 5,001,050). Other examples strand displacing DNA polymerase include, but are not limited to, DNA polymerase of the Bst large fragment (Exo (-) Bst (Aliotta et al., Genet. Anal. (Holland) 12: 185-195 (1996) and Exo(-)BcaDNA polymerase (Walker and Linn, Clinical Chemistry 42: 1604-1608 (1996)), phage M2 DNA polymerase(Matsumoto et al., Gene 84: 247 (1989)), phage cpPRDl DNA polymerase (Jung et al., Proc. Natl. Acad. Sci. USA 84: 8287 (1987)), VENT® DNA polymerase (Kong et al., J. Biol. Chem. 268: 1965-1975 (1993)), Klenow fragment of DNA polymerase I (Jacobsen et al., Eur. J. Biochem. 45: 623-627 (1974)), T5 DNA polymerase (Chatterjee et al., Gene 97: 13-19 (1991)), SEQUENASE® (manufactured by US Biochemicals Corp.), PRD1 DNA polymerase (Zhu and Ito, Biochem. Biophys. Acta. 1219: 267-276 (1994)), T4 DNA polymerase holoenzyme (Kaboord and Benkovic, Curr. Biol. 5: 149- 157 (1995)), etc. In some embodiments, the isothermal amplification includes Loop-Mediated Isothermal Amplification (LAMP). In other embodiments, the isothermal amplification includes Recombinase Polymerase Amplification (RPA). In yet other embodiments, the isothermal amplification comprises Rolling Circle Amplification (RCA).
[0179] In some embodiments, an adapter ligated nucleic acid molecule including a 5' hairpin region may be isothermally amplified using a linear nicking endonuclease-mediated stranddisplacement DNA amplification methodology. In these embodiments, the hairpin region includes a loop sequence having a recognition site for a nicking endonuclease. Such isothermal amplification is described by Joneja A, Huang X. Linear nicking endonuclease-mediated stranddisplacement DNA amplification. Anal Biochem. 2011 Jul l;414(l):58-69. doi: 10.1016 / j.ab.2011.02.025. Epub 2011 Feb 20. PMID: 21342654; PMCID: PMC3108800, the disclosure of which is incorporated by reference herein in its entirety. In some embodiments, linear nicking endonuclease-mediated strand displacement DNA amplification utilizes nicking endonuclease-mediated strand displacement by a DNA polymerase. The nicking of one strand of a DNA target by the endonuclease produces a primer for the polymerase to initiate synthesis. As the polymerization proceeds, the downstream strand is displaced into a single-stranded form while the nicking site is also regenerated. The combined continuous repetitive action of nicking by the endonuclease and strand displacement synthesis by the polymerase results in linear amplification of one strand of the DNA molecule.
[0180] In some embodiments, a "full-length" adapter ligated nucleic acid molecule including two hairpin regions (see, e.g., FIGS. 2A - 2D) may be amplified using LAMP. LAMP a single-step amplification reaction utilizing a DNA polymerase with strand displacement activity (e.g., Notomi et al., NucL Acids. Res. 28: E63, 2000; Nagamine et al., Mol. Cell. Probes 16:223- 229, 2002; Mori et al., J. Biochem. Biophys. Methods 59: 145-157, 2004; see also Quyen TL, Vinayaka AC, Golabi M, Nguyen T, Ngoc HV, Bang DD, Wolff A. Multiplexed Detection of Pathogens Using Solid-Phase Loop-Mediated Isothermal Amplification on a Supercritical Angle Fluorescence Array for Point-of-Care Applications. ACS Sens. 2022 Nov 25;7(11):3343-3351. doi: 10.1021 / acssensors.2c01337. Epub 2022 Oct 25. PMID: 36284082, the disclosures of which are hereby incorporated by reference herein in their entireties). LAMP utilizes 4 to 6 primers recognizing 6 to 8 distinct regions of target DNA for a highly specific amplification reaction. In some embodiments, the primers include a forward outer primer (F3), a backward outer primer (B3), a forward inner primer (HP), and a backward inner primer (BIP). A forward loop primer (Loop F), and a backward loop primer (Loop B) can also be included in some embodiments. A strand-displacing DNA polymerase initiates synthesis and two specially designed primers form "loop" structures (with inverted repeats of the nucleic acid sequence) to facilitate subsequent rounds of amplification through extension on the loops and additional annealing of primers.
[0181] In some embodiments, a LAMP reaction mixture includes the one or more adapter ligated double stranded nucleic acid molecules, LAMP primers, a DNA polymerase with strand displacement activity, dNTPs, and a reaction buffer. Additionally, specific modifications can be made to the reaction mix to facilitate methylation-specific detection, such as the addition ofmethylation-sensitive restriction enzymes or methylation-specific DNA-binding proteins. The LAMP reaction is incubated at a constant temperature, typically about 60°C to about 65°C. The DNA polymerase initiates strand displacement DNA synthesis, resulting in the accumulation of large amounts of amplification products. The amplification process involves multiple steps, including DNA strand displacement and DNA annealing / elongation, leading to a rapid and exponential increase in DNA amplification.
[0182] In some embodiments, a "full-length" adapter ligated nucleic acid molecule including a single 5' hairpin region (see, e.g., FIGS. 1 A - ID) may be amplified using a Thermal asymmetric interlaced PCR (TAIL-PCR) method. In some embodiments, the adapters including two hairpin regions, such as those depicted in FIGS. 5A - 5D, may likewise be amplified using TAIL-PCR. In some embodiments, nested insertion-specific primers are used together with arbitrary degenerate primers (AD primers), which are designed to differ in their annealing temperatures. Alternating cycles of high and low annealing temperature yield specific products bordered by an insertion-specific primer on one side and an AD primer on the other. Further specificity is obtained through subsequent rounds of TAIL-PCR, using nested insertion-specific primers. (See Singer T, Burke E. High-throughput TAIL-PCR as a tool to identify DNA flanking insertions. Methods Mol Biol. 2003;236:241-72, the disclosure of which is hereby incorporated by reference herein in its entirety).Optional Target Enrichment
[0183] Following the optional amplification of the one or more adapter ligated nucleic acid molecules, the sample is optionally enriched for one or more target nucleic acid molecules (step 104). During the step of enrichment, non-target nucleic acid molecules are removed from the amplified sample to provide for an enriched sample, namely a sample enriched for the presence of target nucleic acid molecules.
[0184] Any method may be utilized to enrich the prepared sample for the presence of one or more target nucleic acid molecules. In some embodiments, a hybridization-based target enrichment workflow may be utilized to enrich the prepared sample. In hybridization-based target enrichment workflows, a target area of a target nucleic acid molecule is captured by one or more hybridization probes that can selectively bind to a capture surface. This capture allows the removal of non-target nucleic acids and subsequent release and collection of captured target molecules. Hybridization of target regions may occur either on a solid surface (microarray) or in solution. Hybridization-based target enrichment workflows are described in United States Patent No. 8,383,338, the disclosure of which is hereby incorporated by reference herein in its entirety.Commercial hybridization-based target enrichment workflows are available from Roche Sequencing Solutions, Inc. (e.g., KAPA HyperCap Workflow). Other commercial hybridizationbased target enrichment workflows include SECAP EZ Target Enrichment System (ROCHE) and SURESELECT Target Enrichment System (AGILENT).
[0185] By way of example, hybridization-based target enrichment may be performed by capturing the target nucleic acid molecules in a sample with one or more introduced target-specific probes. In some embodiments, the one or more target nucleic molecules in an obtained sample may be denatured and contacted with single-stranded target-specific probes. In some embodiments, the single-stranded target-specific probes may comprise a ligand for an affinity capture moiety such that following the formation of hybridization complexes, the hybridization complexes are captured by contacting the sample with the affinity capture moiety. In some embodiments, the affinity capture moiety is avidin or streptavidin and the ligand is biotin. In some embodiments, the moiety is bound to solid support. In some embodiments, the solid support may comprise superparamagnetic spherical polymer particles such as DYNABEADS™ magnetic beads or magnetic glass particles.
[0186] In other embodiments, a primer extension target enrichment (PETE) workflow may be utilized to enrich the prepared input sample (see, e.g., FIGS. 10 and 11). PETE workflows are described in United Patent Application Publication Nos. 2021 / 0207211 and 2020 / 0392483; in United States Patent Nos. 10,907,204 and 11,499,180; and in International Publication Nos. WO / 2018 / 013710 and WO / 2022 / 008578, the disclosures of which are each incorporated by reference herein in their entireties. Commercial PETE workflows are available from Roche (e.g., HAPA HyperPETE Workflow). By way of example only, a PETE workflow may be utilized to enrich a sample with one or more target nucleic acid molecules by: a) providing a reaction mixture comprising the sample and a first target-specific primer, wherein the sample comprises singlestranded target nucleic acid molecule having a 3' and a 5' end and non-target nucleic acid molecules; b) hybridizing a first target-specific primer to the single-stranded target nucleic acid molecules in the reaction mixture, wherein the first target-specific primer hybridizes at least 6 nucleotides from the 3' end of the single-stranded target nucleic acid molecule and comprises an affinity ligand; c) extending the hybridized first target-specific primer with a DNA polymerase to form a first double-stranded product comprising the target nucleic acid molecule hybridized to the extended first target-specific primer, wherein the hybridized target nucleic acid molecule comprises a single-stranded overhang region of at least 6 consecutive nucleotides at the 3' end; d) removing single-stranded target and non-target nucleic acid molecules from the reaction mixture by capturing the affinity ligand of the first double-stranded product; e) hybridizing a second target-specific primer to the single-stranded overhang region at the 3' end of the hybridized target polynucleotide of the captured first double stranded product, wherein the second target-specific primer comprises a 3' hybridizing region and a barcode region; and f) extending the hybridized second target-specific primer with a DNA polymerase, wherein the DNA polymerase comprises strand displacement activity, 5'-3' double stranded DNA exonuclease activity, or a combination thereof, thereby displacing or degrading the extended first target-specific primer and forming a second double-stranded product comprising a barcode, wherein the second double-stranded product comprises the target nucleic acid molecule hybridized to an extended second targetspecific primer, wherein the extended second target-specific primer comprises: i) a complement of at least a portion of the target nucleic acid molecule; and, ii) a single-stranded 5' overhang region comprising the barcode.Conversion of a Single Strand of an Adapter Ligated Nucleic Acid Molecule into a Hairpin Duplex Molecule
[0187] The methods of the present disclosure further include the step of converting a single strand of an adapter ligated nucleic acid molecule into a hairpin duplex molecule (step 105). An example of a structure of a hairpin duplex molecule is illustrated in FIG. 7. Essentially, the hairpin duplex molecule is a double stranded nucleic acid molecule which includes a hairpin at one end (e.g., at a 5' end), thereby providing a double stranded nucleic acid molecule including an "open" end and a "closed" end. The skilled artisan will appreciate that the inclusion of the hairpin "closes" one end of the double stranded nucleic acid molecule by bridging a first strand and a second strand of the double stranded nucleic acid molecule, permitting the first strand to be continuous with the second strand. In some embodiments, the "open" end (e.g., the 3' end) includes a primer binding sequence or primer binding site. In some embodiments, an adapter may be coupled to the "open" end, such as via ligation.
[0188] A method of generating a hairpin duplex molecule is depicted in FIG. 7. In some embodiments, a single strand of an adapter ligated nucleic acid molecule, namely one including a hairpin region as described herein, is subjected to conditions which initiate a 3' self-priming event, such as a self-priming event between complementary portions of first and second duplex sequences. This results in the formation of a loop-like structure within the hairpin region, where the loop-like structure includes at least two duplex sequences (e.g., "Duplex 1" and "Duplex 2" as noted herein, and optionally the "Loop Sequence"). In some embodiments, the sample including the adapter ligated nucleic acid molecule may be subjected conditions including a temperature below about 50°C, such as at a temperature of below about 45°C, such as at a temperature of below about 40°C, below about 35°C, below about 30°C, etc. to encourage formation of a hairpin at anend of a single strand of the adapter ligated nucleic acid molecules. In other embodiments, the sample including the adapter ligated nucleic acid molecule may be subjected to conditions including a temperature above about 50°C, such as at a temperature of above about 55°C around about 58 °C. In some embodiments, the structure of the substantially double stranded stem region (e.g., the melting temperature of the stem region, the length of the stem region, and / or the nucleotide sequence of the stem region, etc.) is tuned to encourage hairpin formation (e.g., by reaction temperature optimization, additive testing including monovalent or divalent salts, DMSO, Betaine, etc.).
[0189] Following the self-priming event (and the formation of the hairpin, such as the formation of a hairpin between two duplex sequences, such as at a 5' end of the adapter ligated nucleic acid molecule), a polymerase extension is performed to generate the hairpin duplex molecule (such as by extending the formed hairpin). In some embodiments, the polymerase extension is performed using a low temperature polymerase, such as 029 DNA polymerase, E.coli DNA polymerase, BSU, BST, etc.
[0190] In some embodiments, and to prepare adapter ligated double stranded nucleic acid molecules for MethylSeq, the polymerase extension is carried out in the presence of 5- Methylcytidine-5'-Triphosphate (5-Methyl-CTP or m5CTP) (see Gao, et. al., (2021). RNA 5- methylcytosine modification and its emerging role as an epitranscriptomic mark. RNA Biology, 7S(supl), 117-127; see also Yan et. al., Methyl- SNP-seq reveals dual readouts of methylome and variome at molecule resolution while enabling target enrichment. Genome Res. 2022 Nov-Dec;32(l l-12):2079-2091, the disclosures of which are each incorporated by reference herein in their entireties).
[0191] In some embodiments, and regardless of whether the polymerase extension is carried out in the presence of m5CTP, a bisulfite treatment may be optionally utilized to convert unmethylated cytosine to uracil, whereby methylated cytosines remain unchanged, thus facilitating mapping of m5C during sequencing. Bisulfite conversion involves the deamination of unmodified cytosines to uracil, leaving the modified bases 5-mC and 5-hmC. More specifically, treatment of denatured DNA (i.e., single-stranded DNA) with sodium bisulfite leads to deamination of unmethylated cytosine residues to uracil, leaving 5-mC or 5-hmC intact. The uracil nucleotides may be amplified in a subsequent a PCR reaction as thymines, whereas 5-mC or 5-hmC residues get amplified as cytosines.
[0192] In some embodiments, methylated cytosines are enzymatically protected via enzymatic oxidation of 5mC to 5hmC, 5fC, or 5caC, by a ten-eleven translocation (TET)methylcytosine dioxygenase 2 mutant and subsequent enzymatic glycosylation of all 5hmCs, by beta-glycosyltransferase (see Fiillgrabe et. al., Simultaneous sequencing of genetic and epigenetic bases in DNA. Nat Biotechnol. 2023 Oct;41(10): 1457-1464. doi: 10.1038 / s41587-022-01652-0. Epub 2023 Feb 6. PMID: 36747096; PMCID: PMC10567558). The unprotected Cs are then deaminated to uracil prior to sequencing as noted above.Methods of Preparing Two Pass Sequencing Templates using Loop Tail Primers
[0193] Aspects of the present disclosure include methods of preparing a sequencing library including one or more two pass sequencing templates, in which the sequencing library is prepared from a nucleic acid molecule that includes a known sequence at one or both ends. In certain embodiments, the method includes the steps of obtaining a sample including the one or more nucleic acid molecules with known sequences at one or both ends (FIG. 13, step 1301), amplifying the one or more nucleic acid molecules with a loop tail primer (FIG. 13, step 1302), converting a single strand of the amplified nucleic acid molecule in the sample into a hairpin duplex molecule (FIG. 13, step 1303), annealing sequencing primer to the hairpin duplex molecule to provide a primed hairpin duplex molecule (FIG. 13, step 1304), and sequencing the formed library (FIG. 13, step 1305).
[0194] In one embodiment, an exemplary method of preparing a two-pass sequencing template using a loop tail primer of the present invention is illustrated in simplified form in FIG. 14. As illustrated in step 1, a sample 1400 including one or more nucleic acid molecules is obtained and prepared for downstream processing. In certain embodiments, sample 1400 may include one or more nucleic acid amplicons. In this embodiment, the method is referred to as “amplicon-based library preparation”. As used herein, the term “amplicon” refers to a nucleic acid fragment that is the product of an amplification or replication process, either naturally occurring or artificially generated. In some embodiments, amplicons are the result of polymerase chain reaction (PCR) or other similar techniques used to make many copies of a specific nucleic acid sequence. In some embodiments, amplicons may be produced by ligating adapters to one or more double stranded nucleic acid molecule in which the adapters include known hybridization sites for forward and reverse amplification primers, e.g., “universal” primers. Amplification of the adapter-ligated double stranded nucleic acid molecules produces a library of double stranded amplicons with known sequences at both ends, which can be used as primer hybridization sites during subsequent replication or amplification reactions.
[0195] In step 2, amplicons 1400 are further replicated, or amplified with loop tail primer 1410 that hybridizes near the 3’ end of single strand 1405 derived from an amplicon product. As discussed herein, loop tail primer 1410 includes a 3’ primer region that hybridizes to a known sequence near the 3’ end of strand 1405. The free 3’ end of the loop tail primer provides an initiation site for nucleic acid synthesis by a nucleic acid polymerase during the replication or amplification reaction.
[0196] Various exemplary methods of nucleic replication and amplification, including reaction conditions and reagents suitable for the practice of the present invention, are disclosed further herein.
[0197] In this embodiment, in step 2, the replication or amplification reaction includes the loop tail “forward” primer but does not include a “reverse” primer. In this case, only one strand of the amplicon product will be replicated or amplified. In other embodiments, the reaction of step 2 includes the loop tail forward primer and a second reverse primer. The sequence of the reverse primer may be designed to include a known sequence near the 5’ end of single strand 1405 such that it is capable of hybridizing to a complementary copy of strand 1405 during a subsequent round of nucleic acid replication or amplification. In certain embodiments, the sequence of the loop tail primer may include a SID sequence or other feature to facilitate downstream steps of the workflow, as disclosed herein.
[0198] In some embodiments, the amplification reaction may include suitable additives to facilitate various downstream steps of the workflow. For example, in certain embodiments, the amplification reaction may include suitable dNTP analogs, e.g., 7-deaza-dGTP, and the like.
[0199] In some embodiments, the amplification reaction is a PCR reaction that includes from around four to around 30 amplification cycles.
[0200] The amplification (or replication) product of step 2 is double stranded nucleic acid product 1420 that includes the sequence of loop-tail primer 1410 at one end. One of the strands of double stranded nucleic acid product 1420 will include the sequence of the loop tail primer at the 3’ end (i.e., strand 1420a), while the other strand the double stranded nucleic acid product will include the sequence of the loop tail primer at the 5’ end (i.e., strand 1420b). In certain embodiments, both ends of double stranded nucleic acid product 1420 are blunt ends.
[0201] The methods of the present disclosure further include the step of converting a single strand (e.g., strand 1420a) of double stranded nucleic acid product 1420 of step 2 into hairpin duplex molecule 1430 (as depicted in step 3). As disclosed herein, the hairpin duplex molecule is a double stranded nucleic acid molecule that includes a hairpin at one end, thereby providing adouble stranded nucleic acid molecule including an "open" end and a "closed" end. The skilled artisan will appreciate that the inclusion of the hairpin "closes" one end of the double stranded nucleic acid molecule by bridging a first strand and a second strand of the double stranded nucleic acid molecule, permitting the first strand to be continuous with the second strand. In some embodiments, the "open" end includes a primer binding sequence or primer binding site. In some embodiments, an adapter may be coupled to the "open" end, such as via ligation.
[0202] As disclosed herein, in some embodiments, the single strands of double stranded nucleic acid molecule 1420, namely strands including a hairpin region as described herein, are subjected to conditions which initiate terminal self-priming events, such as self-priming events between complementary portions of first and second duplex sequences of the loop tail primer sequence. This results in the formation of loop-like structure 1425 within the hairpin region, as discussed with reference to the full-length adapter. In some embodiments, the sample including the double stranded nucleic acid molecule may be subjected conditions including a temperature below about 50°C, such as at a temperature of below about 45°C, such as at a temperature of below about 40°C, below about 35°C, below about 30°C, etc. to encourage formation of a hairpin at an end of a single strand of double stranded nucleic acid molecules. In other embodiments, the sample including the adapter ligated nucleic acid molecule may be subjected to conditions including a temperature above about 50°C, such as at a temperature of above about 55°C around about 58 °C to encourage formation of a hairpin at an end of a single strand of double stranded nucleic acid molecules. In some embodiments, the structure of the substantially double stranded stem region (e.g., the melting temperature of the stem region, the length of the stem region, and / or the nucleotide sequence of the stem region, etc.) is tuned to encourage hairpin formation (e.g., by reaction temperature optimization, additive testing including monovalent or divalent salts, DMSO, Betaine, etc.).
[0203] Following the self-priming event of strand 1420a, and the formation of the terminal hairpin, the 3’ end of the strand provides an initiation site for a nucleic acid synthesis reaction, or polymerase extension reaction. As such, a polymerase extension is performed to generate hairpin duplex molecule 1430 (such as by extending the formed hairpin). In some embodiments, the polymerase extension is performed using a low temperature polymerase, such as 029 DNA polymerase, E.coli DNA polymerase, BSU, BST, etc.
[0204] However, strand 1420b of double stranded nucleic product 1420 is incapable of being extended following the self-priming event as the free end of the hairpin structure is the 5’ end of the strand. In certain embodiments, strand 1420b may be removed by at least one of themethods described herein, e.g., SPRI bead size selection methods, removal of 5’ biotinylated strands with SA-bead interactions, and / or exonuclease treatment.
[0205] In certain embodiments, in step 4, an adapter, e.g., Y adapter 1435 is ligated to the free end of hairpin duplex molecule 1430. In some embodiments, prior to adapter ligation, the free end the of the hairpin duplex molecule may be A-tailed to facilitate alignment with and ligation to the adapter, as disclosed herein. In some embodiments, the adapters may be ligated to the double stranded nucleic acid molecules using any method known in the art, as disclosed herein.
[0206] In certain embodiments, Y adapter 1435 may include one or more features to mediate downstream steps of the workflow. For example, a single stranded region of the Y adapter (e.g., a portion of a single stranded arm) may include a hybridization site for a primer, e.g., an extension oligonucleotide for initiation of Xpandomer synthesis (in this depiction represented by the lettef’E”). In some embodiments, one or more single stranded regions of the Y adapter may include a nicking endonuclease recognition site for an isothermal amplification step, as disclosed further herein (in this depiction, represented by the triangle symbol). In some embodiments, the Y adapter may include suitable modifications at the 5’ and / or 3’ end of the single stranded arm regions; for example, the adapter may include thio modifications (i.e., phosphorothioate bond modifications) at the 5’ or 3’ ends of the single stranded arm regions to render the ends resistant to exonuclease-mediated digestion.
[0207] In the embodiment depicted in FIG. 14, adapter-ligated hairpin duplex molecule 1440 includes a Y adapter that provides an extension oligonucleotide hybridization site in one single stranded arm region and a nicking endonuclease recognition site in another single stranded arm region of the adapter.
[0208] In certain embodiments, as shown in step 5, adapter-ligated hairpin duplex molecule 1440 is amplified to provide a pool of amplified hairpin duplex molecules 1450. In some embodiments, the adapter ligated hairpin duplex molecules are isothermally amplified using a linear nicking endonuclease-mediated strand displacement DNA amplification methodology. Such isothermal amplification is described by Joneja A, Huang X. Linear nicking endonuclease- mediated strand-displacement DNA amplification. Anal Biochem. 2011 Jul l;414(l):58-69, the disclosure of which is incorporated by reference herein in its entirety. In these embodiments, the adapter region includes a recognition site for a nicking endonuclease. In this embodiment, the adapter ligated duplex hairpin molecule is first replicated to produce an entirely double stranded product, e.g., in which both single stranded arms of the Y adapters are replicated in the double stranded product. In some embodiments, the nicking of one strand of a DNA target by theendonuclease produces a primer for the polymerase to initiate synthesis. As the polymerization proceeds, the downstream strand is displaced into a single-stranded form while the nicking site is also regenerated. The combined continuous repetitive action of nicking by the endonuclease and strand displacement synthesis by the polymerase results in linear amplification of one strand of the DNA molecule.Sequencing Library Conversion
[0209] In one application of the present invention, the amplicon-based methods of preparing two pass sequencing templates may, advantageously, be used to convert a simplex (i.e., single pass) nucleic acid library into a duplex sequencing library. For example, a nucleic acid library prepared for conventional sequencing by synthesis (SBS) platforms may be converted for use with a sequencing platform that enables two-pass sequencing (i.e., “duplex” sequencing, in which sequencing reads include two passes of a DNA template). In one embodiment, an Illumina sequencing library may be converted into a library suitable for duplex Sequencing by Expansion (SBX-D).
[0210] One embodiment of a method of converting an Illumina library into a duplex sequencing library is depicted in simplified form in FIG. 15A and FIG. 15B. FIG. 15A shows one embodiment of how primers may be designed to amplify an Illumina library into PCR products that may be used to form the hairpin duplex molecules of the present invention. Here, exemplary Illumina library construct 1500 includes library insert 1515 flanked by terminal P5 adapter sequence 1505 and terminal P7 adapter sequence 1510. The structures and sequences of the Illumina P5 and P7 adapters are well known in the art and readily available. The adapter sequences may, in certain embodiments, include index sequences, e.g., the Illumina i5 and i7 sequences and other features.
[0211] As depicted in FIG. 15 A, loop tail primer 1520 is provided, which includes loop tail region 1521 and primer region 1523. As disclosed further herein, loop tail region 1521 includes a fist duplex sequence and a second duplex sequence, separated by a loop sequence. The first and second duplex regions are capable of self-hybridizing to form the double stranded region of the loop tail primer that, in certain embodiments, provides a free 3’ end. Suitable sequence features of the loop tail region are disclosed further herein. Importantly, the sequence of loop tail region 1521 is designed to not include any regions of substantial complementarity to the sequence of the Illumina library construct. In other words, the loop tail region should not hybridize to any sequence in the library construct. In contrast, the sequence of primer region 1523 is designed to include regions of substantial complementarity to the library construct. In certain embodiments,the primer region is designed to hybridize to a region in the P5 sequence of the P5 adapter. Suitable sequence features of the primer region are disclosed further herein. In some embodiments, the hybridization site for the primer regions may fall within any suitable region of the P5 sequence and / or other P5 adapter sequences.
[0212] Also depicted here is Y adapter (YAD) primer 1525. The (YAD) primer includes primer region 1527 and YAD arm region 1529. In certain embodiments, the primer region may include features that enable downstream steps of the library conversion workflow, e.g., one or more exonuclease blocking elements (represented in this depiction by the asterisk symbols). In this embodiment, the sequence of YAD arm region 1529 is designed to not include any regions of substantial complementarity to the sequence of the Illumina library construct. In other words, the YAD arm region should not hybridize to any sequence in the library construct. In some embodiments, the YAD arm region may be from around 50 to around 25 nucleotides in length. In certain embodiments, the sequence of the YAD arm region includes a hybridization site for an extension oligonucleotide to, e.g., enable initiation of Xpandomer synthesis. In contrast, the sequence of primer region 1525 is designed to include regions of substantial complementarity to the library construct. In certain embodiments, the primer region is designed to hybridize to a region in the P7 sequence of the P7 adapter. In some embodiments, primer region 1527 may be from abound 30 to around 10 nucleotides in length. In certain embodiments, primer region 1527 may be from around 23 to 15 nucleotides in length and may fall within any suitable region of the P7 sequence.
[0213] As further depicted in FIG. 15 A, loop tail primer 1520 and YAD primer 1525 are used to amplify library construct Illumina 1500 via a PCR reaction. Amplification of the library construct produces fully double stranded PCR product 1550 that includes complementary strands 1550a and 1550b. The sequence of the 5’ end of strand 1550a includes the sense sequence of loop tail primer 1520, while the sequence of the 3’ end of the strand includes the antisense sequence of the YAD primer. The sequence of the 5’ end of strand 1550b includes the sense sequence of the YAD primer, while the sequence of the 3’ end of the strand includes the antisense sequence of the loop tail primer. In this manner, the loop tail sequence of strand 1550b will provide a free 3’ end upon formation of the hairpin structure during self-hybridization of the duplex regions.
[0214] FIG. 15B summarizes one exemplary library conversion workflow of the present invention. Here, in step 1555 an Illumina library is provided (that includes library constructs 1500 of FIG. 15 A). Any suitable library may be converted; non-limiting examples of exemplary libraries include whole genome libraries or targeted libraries, e.g., exome libraries, target enrichment libraries, any form of RNA library, and methylSEQ libraries.
[0215] In step 1560, the library is subjected to conversion PCR using the method outlined in FIG. 15 A. Upon formation, strand 1550b may be converted to a hairpin duplex molecule according to the methods disclosed further herein. In some embodiments, the hairpin duplex molecule is generated late in the amplification process, as the supply of PCR primers is extinguished, with conditions that favor the self-priming event that forms the hairpin structure. The polymerase provided by the PCR reaction conditions is then capable of extending the free 3’ end of the hairpin to form that mature hairpin duplex molecule.
[0216] In step 1565, one or more optional purification steps can be included, e.g., a SPRI bead clean-up step to remove undesired reaction side-products.
[0217] In step 1570, the sample may be subjected to exonuclease treatment. Exonuclease- mediated digestion of single strands of DNA molecules enables 1) exposure of the extension oligonucleotide hybridization site in the hairpin duplex molecule and 2) removal of undesired single stranded nucleic acid side-products. In some embodiments, an exonuclease enzyme may digest a single strand of the hairpin duplex molecule from the 5’ end of the YAD arm region. Digestion proceeds until the exonuclease blocking elements are encountered (e.g., the one or more phosphorothioate bonds designed into the YAD primer) which exposes the extension oligonucleotide hybridization site in the undigested opposite strand. Any suitable exonuclease enzyme may be used in step 1570; certain non-limiting examples include 5’ to 3’ single stranded exonucleases, such as RecJf (a recombinant fusion protein of Red exonuclease and maltose binding protein) and 5’ to 3’ double stranded exonucleases, such as T7 and lambda exonucleases, In certain embodiments, step 1570 includes treatment with more than one exonuclease enzyme, e.g., RecJf and T7 or lambda exonucleases.
[0218] In step 175, the sample may optionally be subjected to one or more purification steps, e.g., a SPRI bead cleanup step.
[0219] In some embodiments, the library conversion workflow outlined in FIG. 15B may include one or more additional steps to enable removal of single stranded nucleic acid by products from the sample of hairpin duplex molecules. In one embodiment, one or both of the primers used in conversion PCR step 1560 may include a terminal biotin moiety. For example, loop tail primer 1520 may include a terminal biotin moiety. In this manner, strand 1550a of double stranded PCR product 1550 will include the terminal biotin moiety and can be removed at the end of the PCR reaction using streptavidin coated beads-based purification protocols. In certain embodiments, the biotinylated primer clean-up step may be carried out before, during, or after exonuclease digestion step 1570.
[0220] In related embodiments, an exemplary method of preparing a two-pass sequencing template using a loop tail primer of the present invention includes a genomic DNA library preparation step and is illustrated in simplified form in FIG. 16. As illustrated in step 1, genomic DNA sample 1600 is obtained and is processed to provide amplified genomic library constructs 1610. In certain embodiments, a genomic DNA library is prepared using conventional DNA fragmentation and adapter ligation techniques, as discussed further herein. The adapters may include art-recognized universal adapters, as discussed with reference to FIG. 15 A. In some embodiments, amplified genomic library constructs 1610 are subjected to a target enrichment step prior to step 2, as disclosed further herein.
[0221] In step 2, amplified genomic library constructs 1610 are further replicated, or amplified with loop tail primer 1615, which hybridizes near the 3’ end of single strand 1620, derived from a genomic library construct. As discussed herein, loop tail primer 1615 includes a 3’ primer sequence that hybridizes to a known sequence near the 3’ end of strand 1620. The free 3’ end of the loop tail primer provides an initiation site for nucleic acid synthesis by a nucleic acid polymerase during the replication or amplification reaction. In some embodiments, the loop sequence of the hairpin loop may include a SID sequence. In some embodiment, the loop tail primer may include a 5’ biotin moiety.
[0222] Various exemplary methods of nucleic replication and amplification, including reaction conditions and reagents suitable for the practice of the present invention, are disclosed further herein.
[0223] In some embodiments, in step 2 the replication or amplification reaction includes the loop tail “forward” primer but does not include a “reverse” primer. In this case, only one strand of the amplicon product will be replicated or amplified. In other embodiments, the reaction of step 2 includes the loop tail forward primer and a second reverse primer. The sequence of the reverse primer may be designed to include a known sequence near the 5’ end of single strand 1620 such that it is capable of hybridizing to a complementary copy of strand 1620 during a subsequent round of nucleic acid replication or amplification. In certain embodiments, the reverse primer may be referred to as an E-oligo primer and include features such as an exonuclease blocking element and / or a nickase endonuclease cleavage site (indicated in the depiction by the asterisk symbol), as disclosed herein.
[0224] In some embodiments, the amplification reaction may include suitable additives to facilitate various downstream steps of the workflow. For example, in certain embodiments, the amplification reaction may include suitable dNTP analogs, e.g., 7-deaza-dGTP, and the like.
[0225] In some embodiments, the amplification reaction is a PCR reaction that includes from around four to around 30 amplification cycles.
[0226] Amplification (or replication) product of step 2 is double stranded nucleic acid product 1625 that includes the sequence of loop-tail primer 1615 at one end. One of the strands of double stranded nucleic acid product 1625 will include the sequence of the loop tail primer at the 3’ end (i.e., strand 1625a), while the other strand the double stranded nucleic acid product will include the sequence of the loop tail primer at the 5’ end (i.e., strand 1625b). In certain embodiments, both ends of double stranded nucleic acid product 1625b are blunt ends.
[0227] The methods of the present disclosure further include the step of converting a single strand (e.g., strand 1625a) of double stranded nucleic acid product 1625 of step 2 into hairpin duplex molecule 1630 (as depicted in step 3). As disclosed herein, the hairpin duplex molecule is a double stranded nucleic acid molecule that includes a hairpin at one end, thereby providing a double stranded nucleic acid molecule including an "open" end and a "closed" end. The skilled artisan will appreciate that the inclusion of the hairpin "closes" one end of the double stranded nucleic acid molecule by bridging a first strand and a second strand of the double stranded nucleic acid molecule, permitting the first strand to be continuous with the second strand. In some embodiments, the "open" end includes a primer binding sequence or primer binding site. In some embodiments, an adapter may be coupled to the "open" end, such as via ligation.
[0228] As disclosed herein, in some embodiments, the single strands of double stranded nucleic acid molecule 1625, namely strands including a hairpin region as described herein, are subjected to conditions which initiate terminal self-priming events, such as self-priming events between complementary portions of first and second duplex sequences of the loop tail primer sequence. This results in the formation of loop-like structure 1627 within the hairpin region, as discussed with reference to the full-length adapter. In some embodiments, the sample including the double stranded nucleic acid molecule may be subjected conditions including a temperature below about 50°C, such as at a temperature of below about 45°C, such as at a temperature of below about 40°C, below about 35°C, below about 30°C, etc. to encourage formation of a hairpin at an end of a single strand of double stranded nucleic acid molecules. In other embodiments, the sample including the adapter ligated nucleic acid molecule may be subjected to conditions including a temperature above about 50°C, such as at a temperature of above about 55°C around about 58 °C . to encourage formation of a hairpin at an end of a single strand of double stranded nucleic acid molecules. In some embodiments, the structure of the substantially double stranded stem region (e.g., the melting temperature of the stem region, the length of the stem region, and / or the nucleotide sequence of the stem region, etc.) is tuned to encourage hairpin formation (e.g., byreaction temperature optimization, additive testing including monovalent or divalent salts, DMSO, Betaine, etc.).
[0229] Following the self-priming event of strand 1625a, and the formation of the terminal hairpin, the 3’ end of the strand provides an initiation site for a nucleic acid synthesis reaction, or polymerase extension reaction. As such, a polymerase extension is performed to generate hairpin duplex molecule 1630 (such as by extending the formed hairpin). In some embodiments, the polymerase extension is performed using a low temperature polymerase, such as 029 DNA polymerase, E.coli DNA polymerase, BSU, BST, etc.
[0230] However, strand 1625b of double stranded nucleic product 1625 is incapable of being extended following the self-priming event as the free end of the hairpin structure is the 5’ end of the strand. In certain embodiments, strand 1625b may be removed by at least one of the methods described herein, e.g., SPRI bead size selection methods, removal of 5’ biotinylated strands with SA-bead interactions, and / or exonuclease treatment.
[0231] In certain embodiments, in step 4, hairpin duplex molecule 1630 is treated with an exonuclease enzyme (or in other embodiments, treated with a nicking endonuclease followed by heat denaturation) to provide mature hairpin duplex molecule 1640 with exposed extension oligonucleotide site 1645, according to methods disclosed further herein.
[0232] In certain embodiments, as shown in step 5, mature hairpin duplex molecule 1640 is amplified to provide a pool of amplified hairpin duplex molecules 1650. In some embodiments, the mature hairpin duplex molecules are isothermally amplified using a linear nicking endonuclease-mediated strand displacement DNA amplification methodology, as disclosed further herein.SEQUENCING
[0233] Following the preparation of a sample including one or more hairpin duplex nucleic acid molecules, the one or more hairpin duplex molecules are prepared for sequencing and then sequenced, such as with a next-generation sequencing platform. In some embodiments, the sequencing comprises sequencing both strands of the hairpin duplex molecules prepared for sequencing.Introduction of a Sequencing Primer to the Hairpin Duplex Molecules
[0234] Following the preparation of the hairpin duplex molecule, a sequencing primer is annealed to a primer binding site or primer binding sequence of the hairpin duplex molecule (see, e.g., FIG. 7) to prepare the nucleic acid molecule for sequencing (step 106). In some embodiments,the method comprises generating of a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
[0235] In those embodiments where the substantially double stranded stem region includes one or more exonuclease blocking elements (such as one or more phosphorothioate bonds), an exonuclease may be introduced into the sample to digest the region including the primer binding site, but not the substantially double stranded stem region (see, e.g., FIG. 9, Panels A and B). As illustrated in Panel B of FIG. 9, the introduction of the exonuclease modifies the hairpin duplex molecule such that the "open" end includes a 3' overhang. Subsequently, a sequencing primer is introduced to the hairpin duplex molecule (see, e.g., FIG. 9, Panels B and C).
[0236] In some embodiments, the exonuclease is one that requires double stranded nucleic acid molecules as substrate. In some embodiments, the exonuclease may comprise activity for double-stranded DNA without nicking. In some embodiments, the exonuclease has 5' to 3' exonuclease activity. In some embodiments, the exonuclease is a T7 exonuclease.
[0237] In those embodiments where the substantially double stranded stem region includes a recognition site for a nicking enzyme, e.g., nicking endonuclease, a nicking enzyme specific for the recognition site included within the region including the primer binding site is introduced to the sample (see, e.g., FIG. 10, Panel A). The introduction of the nicking enzymes modifies the hairpin duplex molecule such that the "open" end includes a 3' overhang. Following nicking, a sequencing primer is introduced to the hairpin duplex molecule (see, e.g., FIG. 10, Panel B).
[0238] The nicking enzyme utilized is dependent upon the recognition sequence included within the region including the primer binding site. Exemplary nicking enzymes include, but are not limited to, N.Bst9I, N.BstSEI, Nb.BbvCI(NEB), Nb. Bpul0I(F ermantas), Nb.BsmI(NEB), Nb.BsrDI(NEB), Nb.BtsI(NEB), Nt.AlwI(NEB), Nt.BbvCI(NEB), Nt.BpulOI(Fermentas), Nt.BsmAI, Nt.BspD6I, Nt.BspQI(NEB), Nt.BstNBI(NEB), and Nt.CviPII(NEB). Examples of nicking enzyme recognition sequences acted upon by nicking endonucleases are identified in U.S. Patent No. 10,570,441, the disclosure of which is hereby incorporated by reference herein in its entirety. Methods of introducing such a nick into a nucleic acid molecule are disclosed by Joneja et. al., "Linear nicking endonuclease-mediated strand-displacement DNA amplification." Anal Biochem. 2011 Jul l;414(l):58-69, the disclosure of which is hereby incorporated by reference herein in its entirety.
[0239] In yet other embodiments, the primer binding site is prepared by opening the Loop Sequence included within a hairpin. For instance, if a hairpin region includes a Loop Sequence including a primer binding site, an oligonucleotide at least partially complementary to the primerbinding site of the Loop Sequence may be introduced along with a polymerase (see FIG. 11, Panel A). Polymerase extension permits the hair duplex to open, thereby creating a free 3' end (see FIG. 11, Panel B). A sequencing primer complementary to the 3' portion of the hairpin duplex may then be hybridized to the hairpin duplex molecule (see FIG. 11, Panel C). In some embodiments, a strand displacing polymerase, e.g., BST, BSU, Phi29, may be used to open the duplex. Once the duplex is open, a primer can be annealed to the single stranded end of the opened molecule.
[0240] FIG. 12 illustrates yet another method of preparing a hairpin duplex molecule for sequencing. In this particular embodiment, an adapter is ligated to the "open" end of the hairpin duplex molecule. A sequencing primer may then be introduced which is complementary to at least a portion of the ligated adapter.Sequencing
[0241] Following the preparation of the library including the one or more hairpin duplex molecules prepared for sequencing, the library may then be sequenced (step 107). In some embodiments, the library including the one or more hairpin duplex molecules prepared for sequencing may be sequenced by any suitable method or with nay suitable instrument including SMRT (single-molecule real-time sequencing), ion semiconductor, pyrosequencing, sequencing by synthesis, combinatorial probe anchor synthesis, and SOLiD sequencing (sequencing by ligation). Non-limiting sequencing platforms include those provided by Illumina® (e.g., the MiniSeq™, MiSeq™, NextSeq™, and / or NovaSeq™ sequencing systems); Ion Torrent™ (e.g., the Ion PGM™, Ion S5™, and / or Ion Proton™ sequencing systems); Pacific Biosciences (e.g., the PACBIO RS II and / or Sequel II System sequencing system); ThermoFisher (e.g., a SOLID® sequencing system); or BGI Genomics (e.g., DNBSeq™ sequencing systems). See, for example U.S. Pat. Nos. 7,211,390; 7,244,559; 7,264,929; 6,255,475; 6,013,445; 8,882,980; 6,664,079; and 9,416,409; the disclosures of which are hereby incorporated by reference herein in their entireties.
[0242] In other embodiments, the library including the one or more hairpin duplex molecules prepared for sequencing may be sequenced by sequencing-by-synthesis (SBS), pyrosequencing, sequencing by ligation (SBL), or sequencing by hybridization (SBH). Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into a nascent nucleic acid strand (Ronaghi, et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 11(1), 3-11 (2001); Ronaghi et al. Science 281(5375), 363 (1998); U.S. Pat. Nos. 6,210,891; 6,258,568; and 6,274,320, each of which is incorporated herein by reference in its entirety). In pyrosequencing, released PPi can be detected by being converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated can bedetected via light produced by luciferase. In this manner, the sequencing reaction can be monitored via a luminescence detection system. In both SBL and SBH methods, target nucleic acids, and amplicons thereof, that are present at features of an array are subjected to repeated cycles of oligonucleotide delivery and detection. SBL methods, include those described in Shendure et al. Science 309: 1728-1732 (2005); U.S. Pat. Nos. 5,599,675; and 5,750,341, each of which is incorporated herein by reference in its entirety; and the SBH methodologies are as described in Bains et al., Journal of Theoretical Biology 135(3), 303-7 (1988); Drmanac et al., Nature Biotechnology 16, 54-58 (1998); Fodor et al., Science 251(4995), 767-773 (1995); and WO 1989 / 10977, each of which is incorporated herein by reference in its entirety.
[0243] In SBS, extension of a nucleic acid primer along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process can be catalyzed by a polymerase, wherein fluorescently labeled nucleotides are added to a primer (thereby extending the primer) in a template dependent fashion such that detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template.
[0244] Nanopore sequencing refers to the approach in which tags that are attached to nucleotides can be distinguished by their effect on ionic currents passing through nanopores as these tagged nucleotide analogs are added to a growing (nascent) DNA strand. Measurements can be made while tagged nucleotides are still part of the ternary complex, or after their tags are released by the polymerase reaction.
[0245] Nanopore sequencing of a nucleic acid molecule may be achieved by strand sequencing and / or exosequencing of the polynucleotide sequence. In some embodiments, nanopores may be used to sequence nucleic acid molecules where a polymerized nucleic acid molecule does not pass through the nanopore during sequencing. In these embodiments, the nucleic acid molecule may be at least partially located in the vestibule of the nanopore, but not in the pore (i.e., narrowest portion) of the nanopore. The nucleic acid molecule may pass within any suitable distance from and / or proximity to the nanopore, and optionally within a distance such that byproducts released from nucleotide incorporation events, e.g., tags cleaved from tagged nucleotide analogs, are detected in the nanopore.
[0246] Nanopore sequencing utilizes different tagged nucleotide analogs each having a covalently attached tag moiety that provides an identifiable, and distinguishable signature when detected within or near a nanopore. In some embodiments, nanopore sequencing requires a set of at least the four-standard deoxy-nucleotides dA, dC, dG, and dT, wherein each nucleotide has anattached tag capable of being detected by a nanopore upon the nucleotide being incorporated by a strand extending enzyme. Examples of tagged nucleotide analogs are described in United States Patent No. 10,975,426, the disclosure of which is hereby incorporated by reference herein in its entirety.
[0247] In some embodiments, a strand extending enzyme (e.g., a DNA polymerase), such as one located in proximity to a nanopore, specifically binds a tagged nucleotide analog that is complimentary to a nucleotide of a nucleic acid molecule which is hybridized to a growing (nascent) nucleic acid strand at its active site. The strand extending enzyme (e.g., a DNA polymerase) then catalytically incorporates the complementary nucleotide moiety of the tagged nucleotide analog ("nucleotide incorporation event") to the end of the nascent nucleic acid strand. Nucleotide incorporation events are catalyzed by the enzyme, such as DNA polymerase or any mutant or variant thereof and use base pair interactions with a template molecule to choose amongst the available nucleotides for incorporation at each location. Completion of the catalytic incorporation event results in the release of the tag moiety and the oligophosphate moiety (minus the one phosphate incorporated into the growing strand) which then passes through the adjacent nanopore.
[0248] "Nucleotide incorporation events," as that term is used herein, means the incorporation of a tagged nucleotide analog into a growing polynucleotide chain. In some embodiments, byproducts of nucleotide incorporation events may be detected by the nanopore. In some embodiments, a byproduct may be correlated with the incorporation of a given type of modified or unmodified nucleotide. In some embodiments, the byproduct passes through the nanopore and / or generates a signal detectable in the nanopore. Released tag molecules are examples of byproducts of nucleotide incorporation events. Additional details pertaining to such nanopore-based sequencing systems and methods are described in United States Patent Nos. 9,605,309 and 9,557,294, the disclosures of which are hereby incorporated by reference herein in their entireties.
[0249] It is believed that sequencing using the SMRT platform allows the observation of single DNA polymerases reading individual molecules of DNA in real time. It is also believed that the kinetic characteristics of DNA polymerization are observable on a single-molecule basis.
[0250] In some embodiments the incorporation of differently labeled nucleotides is observed in real time as template dependent synthesis is carried out. In particular, an individual immobilized primer / template / polymerase complex is observed as fluorescently labeled nucleotides are incorporated, permitting real time identification of each added base as it is added.In this process, label groups are attached to a portion of the nucleotide that is cleaved during incorporation. For example, by attaching the label group to a portion of the phosphate chain removed during incorporation, i.e., a P, y, or other terminal phosphate group on a nucleoside polyphosphate, the label is not incorporated into the nascent strand, and instead, natural DNA is produced. Observation of individual molecules typically involves the optical confinement of the complex within a very small illumination volume. By optically confining the complex, a monitored region is created in which randomly diffusing nucleotides are present for a very short period of time, while incorporated nucleotides are retained within the observation volume for longer as they are being incorporated. This results in a characteristic signal associated with the incorporation event, which is also characterized by a signal profile that is characteristic of the base being added. In some embodiments, interacting label components, such as fluorescent resonant energy transfer (FRET) dye pairs, are provided upon the polymerase or other portion of the complex and the incorporating nucleotide, such that the incorporation event puts the labeling components in interactive proximity, and a characteristic signal results, that is again, also characteristic of the base being incorporated (see, e.g., U.S. Pat. Nos. 6,056,661, 6,917,726, 7,033,764, 7,052,847, 7,056,676, 7,170,050, 7,361,466, 7,416,844 and Published U.S. Patent Application No. 2007-0134128, the disclosures of which are each hereby incorporated by reference herein in their entireties).
[0251] The SMRT platform uses a polymerase enzyme, a template sequence, and a primer sequence complementary to a portion of the template sequence. These components are immobilized within a confined illumination volume. The reaction mixture surrounding the complex has the four different nucleotides (A, G, T and C) each labeled with a spectrally distinguishable fluorescent label attached through its terminal phosphate group. Because the illumination volume is small, nucleotides and their associated fluorescent labels diffuse in and out of the illumination volume quickly and thus provide only very short fluorescent signals. When a particular nucleotide is incorporated by the polymerase in a primer extension reaction, the fluorescent label associated with the nucleotide is retained within the illumination volume for a longer time. Once incorporated, the fluorescent label is cleaved from the base through the action of the polymerase, and the label diffuses away.
[0252] In some embodiments, the sequencing comprises sequencing by expansion (SBX). This chemistry translates the sequence of DNA into a simple to measure surrogate molecule called an Xpandomer. Much like with polymerase chain reaction, Xpandomer synthesis is based on the natural function of DNA replication where expandable nucleoside triphosphates (X-NTPs) act as substrates for template-dependent, polymerase-based replication. Four easily differentiated X-NTPs (also called High Signal-to-Noise Reporters) are used during Xpandomer synthesis, one for each DNA base, and engineered polymerases incorporate the X-NTPs into an Xpandomer, which serves as a surrogate for the complement of the nucleic acid template. As the Xpandomer molecule transits through a nanopore, the distinct electrical signal of each base reporter is easily identifiable to enable highly accurate and high throughput nanopore-based nucleic acid sequencing. SBX is described in U.S. Patent No. 7,939,259, 9,771,614, 10,774,105, and 11,530,392, the disclosures of which are hereby incorporated by reference herein in their entireties. The Xpandomer synthesis process is described in published PCT Application No. WO2025 / 172745 and PCT Application No. PCTUS25 / 43277, the disclosures of which are hereby incorporated by reference in their entireties.Kits
[0253] The present disclosure also provides for kits including any of the "full-length" or "truncated" adapters or “loop tail” primers of the present disclosure, and one or more additional components. In some embodiments, the kits include (i) one of the "full-length" or "truncated" adapters of the present disclosure, and (ii) a ligase. Non-limiting examples of ligases include, but are not limited to, bacteriophage T4 ligase, a modified T4 ligase, and E. coli ligase. Thermostable ligases include, but are not limited to, Afu ligase, Taq ligase, Tfl ligase, Tth ligase, Tth HB8 ligase, Thermus species AK16D ligase and Pfu ligase (see, for example, PCT Publication WOOO / 26381, the disclosure of which is hereby incorporated by reference herein in its entirety).
[0254] In some embodiments, the kit includes an exonuclease, such as a 5' to 3' exonuclease. In other embodiments, the kit includes a nicking endonuclease. In some embodiments, the nicking endonuclease is one or more of N.Bst9I, N.BstSEI, Nb.BbvCI(NEB), Nb. Bpul0I(F ermantas), Nb.BsmI(NEB), Nb.BsrDI(NEB), Nb.BtsI(NEB), Nt.AlwI(NEB), Nt.BbvCI(NEB), Nt.BpulOI(Fermentas), Nt.BsmAI, Nt.BspD6I, Nt.BspQI(NEB), Nt.BstNBI(NEB), and Nt.CviPII(NEB).
[0255] In some embodiments, the kit may include reagents for amplification (a master mix), e.g., polymerase, dNTPs, 5 -Methyl -C TPs, buffers, and / or other elements (e.g., cofactors or aptamers) appropriate for amplification. Typically, the reagent mixture(s) is concentrated, so that an aliquot is added to the final reaction volume, along with sample (e.g., RNA or DNA), enzymes, and / or water. In some embodiments, the kit further includes at least one polymerase. In some embodiments, the kit further includes at least two different polymerases. In some embodiments, the kit further includes a plurality of nucleotides.
[0256] In some embodiments, the kit further includes one or more buffer solutions and / or wash solutions. In some embodiments, the kit further includes beads having a functionalized surface. In some embodiments, the kit further includes one or more release primers. In some embodiments, the kit further comprises first and second amplification primers.
[0257] Although the present disclosure has been described with reference to several illustrative embodiments, it should be understood that numerous other modifications and embodiments can be devised by those skilled in the art that will fall within the spirit and scope of the principles of this disclosure. More particularly, reasonable variations and modifications are possible in the component parts and / or arrangements of the subject combination arrangement within the scope of the foregoing disclosure, the drawings, and the appended claims without departing from the spirit of the disclosure. In addition to variations and modifications in the component parts and / or arrangements, alternative uses will also be apparent to those skilled in the art.EXAMPLESExample 1Amplification-Free Duplex Library Prep with Full-Length Adapters
[0258] In this example, a library of hairpin duplex constructs was prepared using an exemplary full-length adapter structure and a PCR-free workflow. The library of hairpin adapter constructs was sequenced using Roche Sequencing Solutions’ proprietary Sequencing by Expansion platform. Steps of the library prep workflow implemented are outlined below; steps 1 - 3 utilized the KAPA EvoPlus v2 DNA library prep kit, commercially available from Roche Sequencing Solutions, Inc. The full-length adapter included the following strands: 49-mer strand with the sequence:5’ ’TTTTCAGACGTGTGCTCTTCCGATCTCAACAATGCACCTGACGTACGCT 3’ (SEQ ID NO:1) (the bonds between nucleotides 33-38 are phosphorothioate bonds to block exonuclease- mediated digestion of the strand; nucleotides 39-48 of the strand form the double stranded stem region of the adapter; the 5’ end of the oligonucleotide included a phosphate group); and 43mer strand with the sequence:5’GCGTACGTCAAGGAGTAACATCCCATACATGGATGTTACTCCG3’ (SEQ ID N0:2)(the first 10 nucleotides of the strand form the double stranded stem region of the adapter and the remainder of the nucleotides form the hairpin region of the adapter; the 5’ end of the oligonucleotide included a phosphate group).
[0259] Step 1 : Preparation of genomic DNA insert fragments - enzymatic fragmentationstarting sample of genomic DNA was obtained from the NA12878 cell line For enzymatic fragmentation and A-tailing, a 60pL reaction was prepared that included the following reagents: 500ng of genomic DNA and 25pL of KAPA EvoPlus v2 FragTail RM in nuclease-free water. The reaction was incubated at 37°C for 18 minutes, followed by 55°C for 30 minutes, then held at 4°C.
[0260] Step 2: Ligation. A 75 pL reaction was prepared that included the following reagents: the 60pL fragmentation and A-tailing reaction of stepl, 15pM full-length adapter, and lOpL KAPA EvoPlus v2 ligation ready mix. The ligation reaction was incubated at 20°C for 15 minutes. Following the ligation, the sample was cleaned-up using KAPA HyperPure Beads following the manufacturers’ IFU and the ligation products were eluted in 25 pL of nuclease-free water.
[0261] Step 3: Extension. A 50pL extension reaction was prepared that included the following reagents: the 25pL sample of adapter-ligated DNA insert fragments of step 2 and 25pL of KAPA EvoPlus v2 HiFi HS RM (polymerase). The extension reaction was incubated as follows: 98°C for 45 seconds (initial denaturation), 98°C for 15 seconds (denaturation), 60°C for 30 seconds (primer annealing), 72°C for 30 seconds (extension), and 72°C for one minute (final extension). The sample was then cleaned-up with KAPA HyperPure beads and the DNA products were eluted in 25 pL of nuclease-free water.
[0262] Step 4: Exonuclease digestion. A 50pL digestion reaction was prepared that included the following reagents: 20pL of the post-extension DNA product, 5pL of lOx lambda exonuclease buffer, and IpL of NEB lambda exonuclease. The sample was incubated at 37°C for 30 minutes and the reaction was then terminated by adding EDTA to lOmM. Following the digestion reaction, the sample was cleaned-up with KAPA HyperPure beads (twice) and the final DNA products were eluted in 20pL of nuclease-free water.
[0263] Twenty-four libraries were prepared according to this workflow. The final libraries were pooled and concentrated using a Speedvac. An aliquot of the concentrated library was quality checked using the Tapestation High Sensitivity D1000 kit and Molecular Beacon Kit. The library was then sequenced on the Roche SBX Sequencing Platform according to methods disclosed in Applicant’s published PCT Application No. WO2025 / 172745, the disclosure of which is herein incorporated by reference in its entirety.
[0264] Bioinformatic analysis of the resulting sequencing reads demonstrated 48.79% full- length duplex reads and 51.21% “oneplus” reads (e.g., reads longer than a single copy of the insert,but not the entire duplex read). These results provide proof-of-concept data that validate the “amp- free full-length adapter” workflow as a practical library prep option for duplex sequencing platforms.Example 2Conversion of Illumina Sequencing Libraries to SBX-D Sequencing Libraries
[0265] In this example, various Illumina libraries were converted into a library suitable for duplex Sequencing by Expansion (SBX-D). In other words, a library of simplex DNA constructs (i.e., including one copy of a target fragment) was converted into a library of duplex DNA constructs (i.e., including two copies of a target fragment). The conversion workflow is based on a first loop-tail primer, which includes a hairpin-forming sequence and a SID sequence (HPSID primer) and a second YAD primer, which includes a sequence for SBX extension oligonucleotide hybridization. The following loop tail oligonucleotide primers were evaluated (the “AN” number refers to the length of the stem region of the hairpin structure; each primer has a sequence of 15 nucleotides that hybridizes to the Illumina template construct; all oligonucleotides included a 5’ phosphate group):HPSID AN6-IL15:5’ACATCCATGTATGGGATGTTACTCCTCCTACACGACGCTCT3’ (SEQ ID N0:3) HPSID AN9-IL15:5’ GTAACATCCATGTATGGGATGTTACTCCTCCTACACGACGCTCT3’ (SEQ ID NO:4) HPSID AN12-IL15:5’GGAGTAACATCCATGTATGGGATGTTACTCCTCCTACACGACGCTCT3’(SEQ ID NO: 5)The following YAD oligonucleotide primers were evaluated (all the oligonucleotides included a 5’ phosphate group):YAD IL15:5’TTTTCAGACGTGTGCTCTTCCGATCTCAACAATGCACCTGACGTACGCTAGTTCAG ACGTGTGC3’ (SEQ ID NO: 6)(the bonds between the underlined nucleotides are phophorothioate bonds)Y ADtrunc IL 15 : 5’TTTTCAGACGTGTGCTCTTCCGATCTCAACAATAGTTCAGACGTGTGC3’ (SEQ ID NO: 7) (the bonds between the underlined nucleotides are phophorothioate bonds).
[0266] Preliminary results indicate that the primer pair of HPSID_AN9-IL15 and YADtrunc_IL15 are suitable for the following workflow.
[0267] Step 1 : Conversion PCR. A 50pL PCR reaction was prepared that included the following reagents: 50ng Illumina library, 20pM each primer, and 25pL KAPA HiFi HotStart ReadyMix. The PCR reaction was cycled as follows: 98°C for 45 seconds, then a variable number of cycles of: 98°C for 15 seconds, 60°C for 30 seconds, and 72° C for 30 seconds, then 72° C for one minute. The sample was then cleaned up with KAPA HyperPure beads.
[0268] Step 2: Exonuclease digestion. A digestion reaction was prepared that included the following reagents: 35pL of converted PCR reaction product, 5pL of 10X NEBuffer 4, I pL T7 exonuclease, I pL RecJf, and 8pL nuclease-free water. The reaction was cycled as follows: 25°C for 5 minutes, 37°C for 5 minutes, and 95°C for 5 minutes. The sample was then cleaned up with KAPA HyperPure beads.
[0269] The amount of PCR product present at various time points was measured during the course of the PCR reaction. Results showed an increase in product yield as the number of cycles increased, as expected. Similarly, more PCR product was observed to persist following exonuclease treatment, as the number of cycles increased, suggesting the accumulation of exonuclease-resistant HP duplex constructs (data not shown).
[0270] Several different Illumina libraries were subjected to the library conversion workflow described in this Example and sequenced on the SBX platform. Sequence data confirmed successful conversion of the Illumina libraries to SBX duplex libraries, thereby validating this conversion workflow as a convenient means to produce a duplex sequencing library from a conventional simplex library.
Claims
- 65 -PATENT CLAIMS1. An oligonucleotide adapter comprising a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are each single stranded and non-complementary; and wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence.
2. The oligonucleotide adapter of claim 1, wherein the substantially double-stranded stem region includes one or more modifications to prevent digestion from an exonuclease.
3. The oligonucleotide adapter of claim 1, wherein the substantially double-stranded stem region includes one or more phosphorothioate bonds.
4. The oligonucleotide adapter of claim 1, wherein the 3' portion of the second strand comprises a region including a primer binding sequence.
5. The oligonucleotide adapter of claim 4, wherein the region including the primer binding sequence comprises a recognition site for a nicking enzyme.
6. The oligonucleotide adapter of claim 4, wherein the first hairpin region forms a first loop-like structure at temperatures below or above about 50°C.
7. The oligonucleotide adapter of claim 1, wherein the 3' portion of the second strand comprises a second hairpin region including two second duplex sequences, wherein each of the two second duplex sequences are separated by an optional second loop sequence.
8. The oligonucleotide adapter of claim 7, wherein the two first duplex sequences have different nucleotide sequences than the two second duplex sequences.
9. The oligonucleotide adapter of claim 7, wherein each of the first and second hairpin regions form independent loop-like structures at a temperature below or above about 50°C.
10. The oligonucleotide adapter of any one of the preceding claims, wherein the adapter further includes one or more barcodes.
11. The oligonucleotide adapter of claim 10, wherein the one or more barcodes are located within the substantially double-stranded stem region.
12. The oligonucleotide adapter of any one of claims 10 and 11, wherein the one or more barcodes are located within either the first or second hairpin regions.
13. A double stranded nucleic acid molecule ligated to the oligonucleotide adapter of any one of the preceding claims.
14. An oligonucleotide adapter comprising a first strand and a second strand, wherein a 5' portion of the first strand and a 3' portion of the second strand form a substantially double-stranded- 66 - stem region by sequence complementarity; wherein a 3' portion of the first strand and a 5' portion of the second strand are each single stranded and non-complementary; and wherein the 3' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence.
15. The oligonucleotide adapter of claim 14, wherein the substantially double-stranded stem region includes one or more modifications to prevent digestion from an exonuclease.
16. The oligonucleotide adapter of claim 14, wherein the substantially double-stranded stem region includes one or more phosphorothioate bonds.
17. The oligonucleotide adapter of claim 14, wherein the 5' portion of the second strand comprises a region including a primer binding sequence.
18. The oligonucleotide adapter of claim 17, wherein the region including the primer binding sequence comprises a recognition site for a nicking enzyme.
19. The oligonucleotide adapter of claim 17, wherein the first hairpin region forms a first looplike structure at temperatures below or above about 50°C.
20. The oligonucleotide adapter of claim 14, wherein the 5' portion of the second strand comprises a second hairpin region including two second duplex sequences separated by an optional second loop sequence.
21. The oligonucleotide adapter of claim 20, wherein the two first duplex sequences have different nucleotide sequences than the two second duplex sequences.
22. The oligonucleotide adapter of claim 20, wherein each of the first and second hairpin regions form independent loop-like structures at a temperature below or above about 50 °C.
23. The oligonucleotide adapter of any one of claims 14 - 22, wherein the adapter further includes one or more barcodes.
24. The oligonucleotide adapter of claim 23, wherein the one or more barcodes are located within the substantially double-stranded stem region.
25. The oligonucleotide adapter of any one of claims 23 and 24, wherein the one or more barcodes are located within either the first or second hairpin regions.
26. A double stranded nucleic acid molecule ligated to the oligonucleotide adapter of any one of claims 14 - 25.
27. An oligonucleotide adapter comprising a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are each single stranded and non-complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences- 67 - that are substantially complementary to each other, wherein the two first duplex sequences are separated by a first optional loop sequence; and wherein the 3' portion of the second strand comprises a second hairpin region including two second duplex sequences that are substantially complementary to each other, wherein the two second duplex sequences are separated by a second optional loop sequence.
28. A double stranded nucleic acid molecule ligated to the oligonucleotide adapter of claim 27.
29. A composition comprising one or more polynucleotides having the formula:[first adapter] - [insert] - [second adapter]; wherein the insert is double stranded, a region of the first adapter in proximity to the insert is double stranded, a region of the second adapter in proximity to the insert is double stranded, a region of the first adapter distal to the insert includes two single strands each having an end, and a region of the second adapter distal to the insert comprises two single strands each having an end, wherein at least one strand of the two single strands of the first and second adapters includes a hairpin-forming sequence.
30. The composition of claim 29, wherein both strands of the two single strands of the first and second adapters include hairpin-forming sequences.
31. The composition of any one of claims 29 and 30, wherein the hairpin-forming sequences include two at least partially complementary duplex sequences separated by an optional loop sequence.
32. The composition of any one of claims 29 - 31, wherein each of the hairpin-forming sequences form a loop-like structure at each end of each strand of the polynucleotide at a temperature below or above about 50°C.
33. The composition of any one of claims 29 - 32, wherein each end of the two single strands of the first and second adapters are modified to prevent digestion by an exonuclease.
34. The composition of claim 33, wherein the each of the two single strands of the first and second adapters comprises a phosphorothioate bond.
35. The composition of any one of claims 29 - 34, wherein each of the first and second adapters include one or more barcodes.
36. A method for preparing a library of nucleic acid molecules comprising(a) obtaining a sample including one or more nucleic acid molecules;(b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acid molecules, where each of the first and second adapters include a first strand and a second- 68 - strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non- complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; and wherein the 3' portion of the second strand includes a primer binding site;(c) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and(d) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
37. The method of claim 36, further comprising amplifying the one or more adapter ligated nucleic acid molecules prior to the synthesizing of the hairpin duplex molecule.
38. The method of claim 36, wherein the amplification comprises isothermal amplification.
39. The method of claim 37 or 38, further comprising performing target enrichment.
40. The method of any one of claims 36 - 39, wherein each duplex sequence has a Tmranging from between about 23°C to about 50°C.
41. The method of any one of claims 36 - 40, wherein the synthesizing of the hairpin duplex molecule comprises (i) subjecting the sample including the one or more adapter ligated nucleic acid molecules to conditions in which a self-priming event occurs between the substantially complementary sequences of the two first duplex sequences to form a hairpin between the two duplex sequences; and (ii) extending the formed hairpin.
42. The method of claim 41, wherein the conditions include a temperature below or above about 50°C.
43. The method of claim 41, wherein the conditions include a temperature below about 40°C.
44. The method of any one of claims 41 - 43, wherein the formed hairpin is extended with a polymerase.- 69 -45. The method of claim 44, wherein the polymerase is a low temperature polymerase.
46. The method of claim 44, wherein the polymerase is selected form the group consisting of 29 DNA polymerase, E.coli DNA polymerase, BSU, and BST.
47. The method of any one of claims 44 - 46, wherein the extension is carried out in the presence of 5-Methylcytidine-5'-Triphosphate.
48. The method of any one of claims 44 - 47, further comprising converting unmethylated cytosine to uracil.
49. The method of claim 48, wherein the converting of the unmethylated cytosine to uracil comprises performing a bisulfite treatment.
50. The method of claim 47, wherein methylated cytosines are converted to one of 5hmC, 5fC, or 5caC.
51. The method of any one of claims 36 - 50, wherein generating of the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing comprises introducing a sequencing primer to the sample including the one or more hairpin duplex molecules.
52. The method of claim 51, wherein the sequencing primer is substantially complementary to the primer binding site.
53. The method of claim 51, wherein the substantially double stranded stem region of the first and second adapters includes one or more exonuclease blocking elements; and wherein the generating of the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing further comprises introducing an exonuclease to the sample including the one or more hairpin duplex molecules.
54. The method of claim 53, wherein the exonuclease has 5' to 3' activity.
55. The method of claim 54, wherein the exonuclease having 5' to 3' activity is a T7 exonuclease.
56. The method of claim 51, wherein the substantially double stranded stem region of the first and second adapters includes a recognition site for a nicking enzyme; and wherein the generating of the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing further comprises introducing a nicking enzyme specific to the recognition site.- 70 -57. The method of claim 56, wherein the nicking enzyme is selected from the group consisting of N.Bst9I, N.BstSEI, Nb.BbvCI(NEB), Nb. BpulOI(Fermantas), Nb.BsmI(NEB), Nb.BsrDI(NEB), Nb.BtsI(NEB), Nt.AlwI(NEB), Nt.BbvCI(NEB), Nt.BpulOI(Fermentas), Nt.BsmAI, Nt.BspD6I, Nt.BspQI(NEB), Nt.BstNBI(NEB), and Nt.CviPII(NEB).
58. The method of claim 51, wherein the optional loop sequence is present and wherein the loop sequence includes a primer binding site; and wherein the generating of the library of the one or more synthesized hairpin duplex molecules which are primed for sequencing further comprises introducing a polymerase to open the hairpin at the 5' end of the one or more hairpin duplex molecules.
59. The method of claim 58, wherein the polymerase is selected from the group consisting of BST, BSU, and Phi29.
60. The method of any one of claims 36 - 58, further comprising sequencing the library including the one or more synthesized hairpin duplex molecules which primed for sequencing. The method of claim 59, wherein the sequencing comprises next-generation sequencing.
61. The method of claim 59, wherein the sequencing comprises nanopore sequencing.
62. The method of claim 59, wherein the sequencing comprises Sequencing by Expansion.
63. The method of claim 59, wherein the sequencing comprises bisulfite sequencing.
64. The method of claim 59, wherein the sequencing comprises methylation sequencing.
65. A method for preparing a library of nucleic acid molecules comprising(a) obtaining a sample including one or more nucleic acid molecules;(b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acid molecules, where each of the first and second adapters include a first strand and a second strand, wherein a 5' portion of the first strand and a 3' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 3' portion of the first strand and a 5' portion of the second strand are single stranded and non- complementary; wherein the 3' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; and wherein the 5' portion of the second strand includes a primer binding site;(c) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and(d) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
66. A method for preparing a library of nucleic acid molecules comprising(a) obtaining a sample including one or more nucleic acid molecules;(b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acid molecules, where each of the first and second adapters include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non- complementary; wherein a 5' end of the first strand comprises a hairpin precursor region; and wherein a 3' portion of the second strand comprises a region including a primer binding site;(c) amplifying the one or more adapter ligated nucleic acid molecules to provide a double stranded molecule having a 5' end including a hairpin region and a 3' end including a region having a primer bind site;(d) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and(e) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
67. A method for preparing a library of nucleic acid molecules comprising(a) obtaining a sample including one or more nucleic acid molecules;(b) ligating first and second adapters to each of the one or more nucleic acid molecules in the sample to provide a sample including one or more adapter ligated nucleic acidmolecules, where each of the first and second adapters include a first strand and a second strand, wherein a 3' portion of the first strand and a 5' portion of the second strand form a substantially double-stranded stem region by sequence complementarity; wherein a 5' portion of the first strand and a 3' portion of the second strand are single stranded and non- complementary; wherein the 5' portion of the first strand comprises a first hairpin region including two first duplex sequences that are substantially complementary to each other, wherein the two first duplex sequences are separated by an optional first loop sequence; and wherein the 3' portion of the second strand includes a primer binding site;(c) amplifying the one or more adapter ligated nucleic acid molecules to provide a double stranded molecule having a 5' end including a hairpin region and a 3' end including a region having a primer bind site;(d) synthesizing a hairpin duplex molecule from each of the one or more adapter ligated nucleic acid molecules to provide a sample including one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and includes a 5' end including hairpin and a 3' end including the primer binding site; and(e) generating a library of one or more synthesized hairpin duplex molecules which are primed for sequencing.
68. An oligonucleotide primer comprising a 5’ portion and a 3’ portion, wherein the 5’ portion comprises a hairpin region comprising two duplex sequences that are substantially complementary to each other, wherein the two duplex sequences are separated by a loop sequence.
69. The oligonucleotide primer of claim 68, wherein the hairpin region forms a loop-like structure at temperatures above about 50°C.
70. The oligonucleotide primer of claim 68 or 69, wherein the primer further comprises one or more barcodes.
71. The oligonucleotide primer of claim 70, wherein the one or more barcodes are located within the 5’ portion.
72. The oligonucleotide primer of claim 70, wherein the one or more barcodes are located within the 3’ portion.- 73 -73. The oligonucleotide primer of any one of claim 68-72 further comprising a 5’ biotin moiety.
74. The oligonucleotide primer of any one of claim 68-72 further comprising a 5’ phosphate moiety.
75. The oligonucleotide primer of any one of claim 68-74, wherein the 3’ portion comprises a sequence that is complementary to a sequence in a universal adapter.
76. The oligonucleotide primer of claim 75, wherein the universal adapter is an Illumina adapter.
77. A method for preparing a library of nucleic acid molecules comprising(a) obtaining a sample comprising one or more nucleic acid molecules, wherein the one or more nucleic acid molecules comprise a first end comprising a first known sequence and a second end comprising a second known sequence;(b) amplifying the one or more nucleic acid molecules with a first oligonucleotide primer and a second oligonucleotide primer, wherein the first oligonucleotide primer comprises a 5’ portion and a 3’ portion, wherein the 5’ portion comprises a hairpin region comprising two duplex sequences that are substantially complementary to each other, wherein the two duplex sequences are separated by a loop sequence, wherein the sequence of the 3’ portion is complementary to the first known sequence, and wherein the second primer comprises a sequence complementary to the second known sequence to provide a sample of amplified one or more nucleic acid molecules;(c) synthesizing a hairpin duplex molecule from each of the amplified one or more nucleic acid molecules to provide a sample comprising one or more hairpin duplex molecules, wherein each of the one or more hairpin duplex molecules is double stranded and comprises a first end comprising a hairpin and a second end lacking a hairpin; and(d) generating a library of one or more synthesized hairpin duplex molecules which are capable of hybridizing to an extension oligonucleotide.
78. The method of claim 77, further comprising amplifying the library of one or more synthesized hairpin duplex molecules which are capable of hybridizing to an extension oligonucleotide, wherein the amplification comprises isothermal amplification.- 74 -79. The method of claim 77 or 78, wherein the method of generating a library of one or more synthesized hairpin duplex molecules which are capable of hybridizing to an extension oligonucleotide comprises ligating a Y adapter to the second end lacking a hairpin, wherein the Y adapter comprises a first single stranded arm region comprises a hybridization site for the extension oligonucleotide and a second single stranded arm optionally comprising a recognition site for a nicking enzyme.
80. The method of claim 77, wherein the second primer comprises one or more modifications to prevent exonuclease-mediated digestion.
81. The method of claim 80, wherein the one or more modifications to prevent exonuclease- mediated digestion comprises one or more phosphorothioate bonds.
82. The method of claim 80 or 81, wherein the step of generating a library of one or more synthesized hairpin duplex molecules which are capable of hybridizing to an extension oligonucleotide comprises treating the one or more hairpin duplex molecules with one or more 5’ to 3’ exonuclease enzymes.
83. The method of claim 82, wherein the one or more 5 ’ to 3 ’ exonuclease enzymes are selected from the group consisting of lambda exonuclease, T7 exonuclease, and RecJf.
84. The method of any one of claims 77-83, wherein the first primer comprises a 5’ biotin moiety.
85. The method of claim 84, further comprising the step of treating the sample of step (c) with streptavidin coated beads, wherein the streptavidin coated beads are capable of removing reaction by-products.
86. The method of any one of claims 77-85, further comprising performing target enrichment.
87. The method of any one of claims 77-86, wherein each duplex sequence has a Tmranging from between about 23°C to about 50°C.
88. The method of any one of claims 77-86, wherein the synthesizing of the hairpin duplex molecule comprises (i) subjecting the sample including the amplified one or more nucleic acid molecules to conditions in which a self-priming event occurs between the substantially complementary sequences of the two duplex sequences to form a hairpin between the two duplex sequences; and (ii) extending the formed duplex.
89. The method of claim 88, wherein the formed hairpin is extended with a polymerase.- 75 -90. The method of any one of claims 77-89, further comprising sequencing the library of one or more synthesized hairpin duplex molecules which are capable of hybridizing to an extension oligonucleotide.
91. The method of claim 90, wherein the sequencing comprises Sequencing by Expansion.
Citation Information
Patent Citations
Compositions and methods for detecting nicking enzyme and polymerase activity using a substrate molecule
US10570441B2
Phosphoroamidate esters, and use and synthesis thereof
US10774105B2
DNA ligase variants
US10837009B1
Primer extension target enrichment
US10907204B2
Polypeptide tagged nucleotides and use thereof in nucleic acid sequencing by nanopore detection
US10975426B2